Graphics processor, load balancing method and electronic equipment
By dividing the parallel processing unit of the graphics processor into multiple levels and setting up a multi-level equalization manager, the problem of unbalanced work task distribution in the prior art is solved, and more efficient task allocation and performance utilization are achieved.
Patent Information
- Application Number
- CN202580000643.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-15
AI Technical Summary
The performance of the task balance distributor in existing GPUs is insufficient under high load conditions, resulting in untimely and unbalanced task distribution, and the performance of the graphics processor cannot be fully utilized.
The parallel processing unit of the graphics processor is divided into multiple levels, and a multi-level balancing manager is set up at each level to record and allocate workloads respectively, including the first balancing manager being responsible for the load of the graphics processing cluster, and the second balancing manager being responsible for the load of the minimum level parallel processing unit, and the balanced allocation of work tasks is achieved through the collaborative work of the multi-level manager.
It improves the balance and timeliness of work task distribution, makes full use of the parallel processing unit performance of the graphics processor, and improves data processing efficiency.
Smart Images

Figure CN120500701A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of processors, and specifically provides a graphics processor, a load balancing method, and an electronic device. Background Art
[0002] A GPU (Graphics Processing Unit) consists of multiple parallel processing units (PPUs). Evenly distributing different workloads to these PPUs helps maximize the GPU's performance. Existing GPUs are equipped with a workload balancer, which distributes workloads to each PPU for execution.
[0003] The workload balancer needs to simultaneously record the workload of each parallel processing unit and distribute tasks based on the workload, as well as receive feedback data from parallel processing units after completing tasks and update the workload. However, due to the limited performance of the workload balancer, when the GPU has a large number of tasks, the workload balancer's performance may not meet the needs of workload distribution and feedback processing, resulting in untimely and uneven workload distribution, low GPU processing efficiency, and failure to fully utilize the GPU's performance. Summary of the Invention
[0004] In view of this, the present application aims to provide a graphics processor, a load balancing method, and an electronic device to improve the balance of work task distribution and improve the performance of parallel processing unit data of the graphics processor.
[0005] In a first aspect, an embodiment of the present application provides a graphics processor comprising: a plurality of graphics processing clusters, each graphics processing cluster comprising a plurality of minimum-level parallel processing units; the minimum-level parallel processing units comprising a plurality of target processing cores, the target processing cores comprising at least one of a shading processing core and a computing processing core, each of the target processing cores being used to execute an assigned work task; a first balancing manager, for recording a first workload of each of the graphics processing clusters, and upon receiving a current work task, allocating the current work task to each of the graphics processing clusters based on the first workload of each of the graphics processing clusters; each of the graphics processing clusters comprising a second balancing manager, the second balancing manager being used to record a second workload of each of the plurality of minimum-level parallel processing units within the graphics processing cluster, and receiving the current work task and redistributing the current work task to each of the minimum-level parallel processing units based on the second workload of each of the minimum-level parallel processing units, so that the target processing core of each of the minimum-level parallel processing units executes the current work task.
[0006] The target processing core of the graphics processor is controlled by a parallel processing unit. One parallel processing unit can control multiple target processing cores. In the embodiment of the present application, the parallel processing unit is further divided, and the multiple parallel processing units in the graphics processor are divided into at least two levels. One is the GPU Cluste (graphics processor processing cluster), and the other level is the minimum level parallel processing unit included in each graphics processor processing cluster. The minimum level parallel processing unit controls the target processing core. At the same time, two-level balancing managers are set up for the two levels of parallel processing units respectively. The first balancing manager is responsible for recording the first workload and task distribution of different graphics processing clusters, and the second balancing manager is responsible for recording the second workload and task distribution of each minimum level parallel processing unit. When there is a work task, the work task is first distributed to each graphics processing cluster by the first balancing manager, and then distributed to the minimum parallel processing unit for execution by the second balancing manager in the graphics processing cluster. Compared with the existing technology in which the work task balance distributor is responsible for recording the workload and task distribution of all parallel processing units, the multi-level, multiple balance managers (including the first balance manager and the second balance manager) approach reduces the workload required for each balance manager to record the workload and distribute work tasks, so that each balance manager can distribute work tasks in a timely manner to improve the balance of work task distribution, make full use of different parallel processing units, and improve the performance of parallel data processing of the graphics processor.
[0007] In one embodiment, the graphics processing cluster includes a third parallel processing unit; the third parallel processing unit includes at least two of the minimum-level parallel processing units; each of the graphics processing clusters includes multiple second balancing managers, each of which is respectively arranged in each of the third parallel processing units, and the second balancing manager is used to record the second workloads of the multiple minimum-level parallel processing units in the third parallel processing unit assigned to the work tasks; each of the graphics processing clusters also includes a third balancing manager; the third balancing manager is used to record the third workloads of the multiple third parallel processing units in the graphics processing cluster, and receive the current work tasks, and redistribute the current work tasks to each of the third parallel processing units based on the third workloads of each of the third parallel processing units.
[0008] The minimum-level parallel processing unit includes multiple target processing cores, and the workload includes the working conditions of the target processing cores. When the number of minimum-level parallel processing units in each graphics processing cluster is large, each second balancing manager needs to be responsible for recording more working conditions of the target processing cores. In an embodiment of the present application, the parallel processing units are further graded, and the minimum-level parallel processing units are set in the third parallel processing unit. Each parallel processing unit is configured with a corresponding second balancing manager, and a third balancing manager is also set in each graphics processing cluster to record the workload and allocate work tasks to all third parallel processing units, so that the number of parallel processing units that each balancing manager needs to manage is reduced, the workload of each balancing manager is reduced, and the timeliness of task allocation is improved, so that work tasks can be allocated more evenly.
[0009] In one embodiment, the graphics processing cluster includes multiple fourth parallel processing units; each of the fourth parallel processing units includes at least two of the third parallel processing units; each of the graphics processing cluster includes multiple third balancing managers, and each of the third balancing managers is respectively set in each of the fourth parallel processing units; each of the graphics processing cluster also includes a fourth balancing manager; the fourth balancing manager is used to record the fourth workload of each of the multiple fourth parallel processing units in the graphics processing cluster, and receive the current work task, and redistribute the work task to each of the third parallel processing units based on the fourth workload of each of the fourth parallel processing units.
[0010] In an embodiment of the present application, the parallel processing units in the graphics processor are divided into four levels to achieve balanced distribution of work tasks among the four-level parallel processing units. Different balancing managers are responsible for the multi-level parallel processing units, which record the workload of each level and distribute tasks respectively. Each balancing manager is hierarchical and works simultaneously, effectively improving the balance of work task distribution and improving the performance of parallel data processing of the graphics processor.
[0011] In one embodiment, the target processing core includes multiple registers, and the target processing core executes the assigned work tasks based on the registers; the first workload, the second workload, the third workload and the fourth workload respectively include the number of registers assigned to the work tasks; for any one of the minimum-level parallel processing units: after the registers of the target processing core of the minimum-level parallel processing unit complete the assigned work tasks, the minimum-level parallel processing unit feeds back the first number of registers that have completed the work tasks to the second balancing manager; for each of the third parallel processing units: the second balancing manager of the parallel processing unit receives the first number of registers fed back by all the minimum-level parallel processing units in the parallel processing unit and counts them as the second number of registers, and updates the second workload based on the second number of registers, and feeds back the second number of registers to the third parallel processing unit where the third parallel processing unit is located. A third balancing manager of four parallel processing units; for any one of the fourth parallel processing units: the third balancing manager of the fourth parallel processing unit receives the second register quantity fed back by all the third parallel processing units in the fourth parallel processing unit and counts it as the third register quantity, updates the third workload based on the third register quantity, and feeds back the third register quantity to the fourth balancing manager of the graphics processing cluster where the fourth parallel processing unit is located; for any one of the graphics processing clusters: the fourth balancing manager in the graphics processing cluster receives the third register quantity fed back by all the fourth parallel processing units in the graphics processing cluster and counts it as the fourth register quantity, updates the fourth workload based on the fourth register quantity, and feeds back the fourth register quantity to the first balancing manager; the first balancing manager is used to update the first workload based on the fourth register quantity fed back by all the graphics processing clusters.
[0012] When a work task is assigned to a minimum-level parallel processing unit and executed by a target processing core, the corresponding registers will be configured. After the work task is completed, the registers allocated for the work task will be released so that the registers can be allocated to a new work task. In the embodiment of the present application, each level of the balance manager feeds back to the upper level step by step. The lower-level balance manager first records the registers that have completed the task. The higher-level balance manager does not need to count the number of registers in the minimum-level parallel processing unit one by one. On the one hand, the amount of data fed back upward is reduced. On the other hand, the difficulty of data processing by the higher-level balance manager is reduced. Thus, the efficiency of data statistics of the number of registers after the task is completed can be improved, and work tasks can be assigned to parallel processing units with low loads in a timely manner, thereby improving the balance of work task allocation. In addition, the number of lower-level balance managers is the largest, and the efficiency of multiple lower-level balance managers in recording the registers that have completed the task is relatively high. Thus, the efficiency of counting the number of registers that have completed the task can be improved, and the timeliness and balance of task allocation can be improved.
[0013] In one embodiment, the first balancing manager is used to: determine the graphics processing cluster with the smallest load based on the first workload of each of the graphics processing clusters, and allocate the current work task to the graphics processing cluster with the smallest load; the second balancing manager is used to: determine the minimum-level parallel processing unit with the smallest load based on the second workload of each of the minimum-level parallel processing units in the graphics processing cluster, and allocate the current work task to the minimum-level parallel processing unit with the smallest load.
[0014] In the embodiment of the present application, the unit with the least load (including the graphics processing cluster and the minimum-level parallel processing unit) is determined, and the work tasks are assigned to the unit with the least load for execution, which helps to make the load more balanced and fully utilize the performance of the graphics processor.
[0015] In one embodiment, the first balancing manager is further used to: predict the derivative work tasks that will be generated based on the current work tasks and the preset work task execution logic; determine the graphics processing cluster with the smallest load based on the first workload; pre-allocate the graphics processing cluster with the smallest load to execute the derivative work tasks; and update the first workload based on the pre-allocation results of the derivative work tasks.
[0016] Work tasks are related to each other to a certain extent. Previous work tasks may generate new work tasks. The allocation of work tasks has a certain delay. If a derived work task is allocated after it is generated, the derived work task may not be allocated in a timely manner. Therefore, in the embodiments of the present application, the derived work tasks that may be generated by the current work task can be predicted, and graphics processing clusters can be pre-allocated to the derived work tasks in advance to reduce the delay in allocating the derived work tasks. Furthermore, after pre-allocating the derived work tasks, the first workload is updated so that when allocating other work tasks, the future workload of the graphics processing cluster can be known in advance, thereby reducing the situation where the execution of the derived work tasks is overloaded and improving the balance of work task allocation.
[0017] In one embodiment, the first balancing manager is further used to: predict the derivative work tasks that will be generated based on the current work tasks and the preset work task execution logic; calculate the target resource amount required to execute the derivative work tasks; determine the target graphics processing cluster with the largest remaining resources based on the first workload of each graphics processing cluster; pre-allocate the target graphics processing cluster to execute the derivative work tasks and update the target graphics processing cluster and update the first workload.
[0018] In the embodiment of the present application, pre-allocating derived work tasks based on resource amounts can reduce the situation where work task allocation is unbalanced due to performance differences between different parallel processing units when allocating work tasks, and improve the balance of work task allocation.
[0019] In one embodiment, the current work task includes multiple subtasks, and the target resource amount includes the number of registers required to execute the derived work task, or the number of subtasks of the derived work task.
[0020] In a second aspect, an embodiment of the present application provides a load balancing method for a graphics processor, which is applied to a graphics processor, wherein the graphics processor includes a first balancing manager and multiple graphics processing clusters, each graphics processing cluster includes multiple minimum-level parallel processing units, the minimum-level parallel processing units include multiple target processing cores, and the target processing cores include at least one of a shading processing core and a computing processing core; the load balancing method includes: receiving a current work task; obtaining a first workload of each of the graphics processing clusters through the first balancing manager; allocating the current work task to a target graphics processing cluster based on the first workload through the first balancing manager; upon receiving the current work task, the target graphics processing cluster obtains the second workload of each of the multiple minimum-level parallel processing units in the target graphics processing cluster through the target second balancing manager of the target graphics processing cluster; and redistributing the current work task to the target minimum-level parallel processing unit based on the second workload through the target second balancing manager, so that the target processing core of the target minimum-level parallel processing unit executes the current work task.
[0021] In one embodiment, the graphics processing cluster includes a third parallel processing unit; the third parallel processing unit includes at least two minimum-level parallel processing units; each of the graphics processing clusters includes multiple second balancing managers, each of the second balancing managers is respectively set in each of the third parallel processing units, and the second balancing manager is used to record the second workload of each of the multiple minimum-level parallel processing units in the third parallel processing unit; each of the graphics processing clusters also includes a third balancing manager; the third balancing manager is used to record the third workload of each of the multiple third parallel processing units in the graphics processing cluster; obtaining the second workload of each of the multiple minimum-level parallel processing units in the target graphics processing cluster through the target second balancing manager of the target graphics processing cluster includes: obtaining the third workload of each of the third parallel processing units in the target graphics processing cluster through the target third balancing manager in the target graphics processing cluster; allocating the current work task to the target third parallel processing unit based on the third workload through the target third balancing manager; obtaining the second workload of each of the multiple minimum-level parallel processing units in the target graphics processing cluster based on the target second balancing manager in the target third parallel processing unit.
[0022] In one embodiment, the graphics processing cluster includes a plurality of fourth parallel processing units; each of the fourth parallel processing units includes at least two of the third parallel processing units; each of the graphics processing cluster includes a plurality of third balancing managers, each of which is respectively disposed in each of the fourth parallel processing units;
[0023] Each of the graphics processing clusters also includes a fourth balancing manager; the fourth balancing manager is used to record the fourth workload of each of the multiple fourth parallel processing units in the graphics processing cluster; obtaining the third workload of each of the third parallel processing units in the target graphics processing cluster through the target third balancing manager in the target graphics processing cluster includes: obtaining the fourth workload of each of the fourth parallel processing units in the target graphics processing cluster through the target fourth balancing manager in the target graphics processing cluster; allocating the current work task to the target fourth parallel processing unit based on the fourth workload through the target fourth balancing manager; and obtaining the third workload based on the target third balancing manager in the target fourth parallel processing unit.
[0024] In one embodiment, the target processing core includes a plurality of registers, and the target processing core executes the assigned work tasks based on the registers; the first workload, the second workload, the third workload and the fourth workload respectively include the number of registers assigned to the work tasks; the method further includes: after the registers of the target processing core of any of the minimum-level parallel processing units complete the execution of the assigned work tasks, the minimum-level parallel processing unit feeds back the first number of registers that have completed the work tasks to the target upper-level second balancing manager; the target upper-level second balancing manager receives the first number of registers fed back by all minimum-level parallel processing units and counts them as the second number of registers; the target upper-level second balancing manager updates the second workload based on the second number of registers; the target upper-level second balancing manager counts the second The number of registers is fed back to the target superior third balancing manager; the target superior third balancing manager receives the second number of registers of all target superior second balancing managers and counts them as the third number of registers; the target superior third balancing manager updates the third workload based on the third number of registers; the target superior third balancing manager feeds back the third number of registers to the target superior fourth balancing manager; the target superior fourth balancing manager receives the third number of registers of all target superior third balancing managers and counts them as the fourth number of registers; the target superior fourth balancing manager updates the fourth workload based on the fourth number of registers; the fourth balancing manager feeds back the fourth number of registers to the first balancing manager; the first balancing manager updates the first workload based on the fourth number of registers fed back by all the fourth balancing managers.
[0025] In one embodiment, allocating the current work task to the target graphics processing cluster based on the first workload by the first balancing manager includes: the target first balancing manager determines the graphics processing cluster with the smallest load based on the first workload of each of the graphics processing clusters; allocating the current work task to the graphics processing cluster with the smallest load; and, reallocating the target work task and the current work task to the target minimum-level parallel processing unit based on the second workload by the target second balancing manager includes: the target second balancing manager determines the minimum-level parallel processing unit with the smallest load based on the second workload of each minimum-level parallel processing unit in the graphics processing cluster, and allocating the current work task to the minimum-level parallel processing unit with the smallest load.
[0026] In one embodiment, after receiving the current work task, the method further includes: the first balancing manager predicts the derivative work tasks that will be generated based on the current work task based on the current work task and the preset work task execution logic; determines the graphics processing cluster with the smallest load based on the first workload; pre-allocates the graphics processing cluster with the smallest load to execute the derivative work tasks; and updates the first workload based on the pre-allocation result of the derivative work tasks.
[0027] In one embodiment, after receiving the current work task, the method further includes: the first balancing manager predicts the derivative work tasks that will be generated based on the current work task based on the current work task and the preset work task execution logic; calculates the target resource amount required to execute the derivative work tasks; determines the target graphics processing cluster with the largest remaining resource amount based on the first workload of each graphics processing cluster; pre-allocates the target graphics processing cluster to execute the derivative work tasks and updates the target graphics processing cluster and updates the first workload.
[0028] In one embodiment, the current work task includes multiple subtasks, and the target resource amount includes the number of registers required to execute the derived work task, or the number of subtasks of the derived work task.
[0029] In a third aspect, an embodiment of the present application provides an electronic device comprising a graphics processor as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0031] Figure 1 A first structural diagram of a graphics processor provided in one embodiment of the present application;
[0032] Figure 2 A hierarchical diagram of a parallel processing unit provided in one embodiment of the present application;
[0033] Figure 3 A second structural diagram of a graphics processor provided in one embodiment of the present application;
[0034] Figure 4 A third structural diagram of a graphics processor provided in one embodiment of the present application;
[0035] Figure 5 This is a flowchart of a load balancing method for a graphics processor provided in one embodiment of the present application.
[0036] Icons: graphics processor 100; graphics processing cluster 110; first balancing manager 121; second balancing manager 122; third balancing manager 123; fourth balancing manager 124; minimum level parallel processing unit 130; third parallel processing unit 140; fourth parallel processing unit 150. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.
[0038] See also Figure 1 , Figure 1 The first structural diagram of a graphics processor 100 provided in one embodiment of the present application is shown in FIG. The graphics processor 100 includes a graphics processing cluster 110 and a first equalization manager 121 .
[0039] The graphics processor 100 includes multiple parallel processing units, such as hundreds or thousands of parallel processing units. Figure 2 , Figure 2A hierarchical schematic diagram of a parallel processing unit provided for one embodiment of the present application. In the parallel processing unit, each parallel processing unit includes multiple target processing cores. In the present application, the target processing cores include shader cores and computational processing cores. Some or all of the target processing cores in the parallel processing unit are divided into a batch, and the target processing cores of a batch are used together to complete a task. Each target processing core is configured with multiple temp registers, and the target processing core executes the assigned work tasks through multiple registers.
[0040] In the embodiment of the present application, the parallel processing units are divided into different levels. For example, the parallel processing units are first divided into multiple graphics processing clusters 110, each of which includes multiple parallel processing units. For example, a graphics processor 100 includes 64 parallel processing units, which can be divided into four graphics processing clusters 110, each of which includes 16 parallel processing units.
[0041] In the embodiment of the present application, there can be multiple graphics processing clusters 110, which can include all parallel processing units of the graphics processor 100. The parallel processing units within each graphics processing cluster 110 can be further subdivided. Continuing with the above example, each graphics processing cluster 110 includes 16 parallel processing units, and the 16 parallel processing units can be further divided into 4 groups of secondary parallel processing units, each group of secondary parallel processing units including 4 parallel processing units. Similarly, the parallel processing units that are not further divided are the minimum-level parallel processing units 130 described in the embodiment of the present application. Accordingly, the graphics processing cluster 110 can also be understood as the maximum-level graphics processor 100.
[0042] like Figure 1 As shown, in one embodiment of the present application, the parallel processing units can be divided into at least two levels, one level is the graphics processing cluster 110, and the other level is the minimum level parallel processing unit 130. Each graphics processing cluster 110 includes multiple minimum level parallel processing units 130.
[0043] The distribution and recording of work tasks are managed by the corresponding manager. In an embodiment of the present application, a first balancing manager 121 can be set for all graphics processing clusters 110. The first balancing manager 121 is used to record the first workload of each graphics processing cluster 110 and, upon receiving a current work task, distribute the current work task to each graphics processing cluster 110 based on the first workload of each graphics processing cluster 110. The first workload can be recorded as a workload table.
[0044] The load balancing manager can utilize an existing LBM (load balance manager). In this embodiment, the first load balancing manager 121 is responsible for recording the workloads of all graphics processing clusters 110 and obtaining a first workload corresponding to each graphics processing cluster 110. The first workload refers to the workload of each graphics processing cluster 110, and the term "first" is not intended to be limiting. Accordingly, the first load balancer records the workloads of the graphics processing clusters 110 and distributes tasks. The first load balancer can also be referred to as a Cluster LBM (CLBM).
[0045] The first workload includes, but is not limited to, the allocation of work tasks to all minimum-level parallel processing units 130 in the graphics processing cluster 110, the usage and idleness of resources such as the number of idle minimum-level parallel processing units 130. Furthermore, the first workload may also record the usage and idleness of target processing cores or registers within the graphics processing cluster 110, which is not limited here.
[0046] In an embodiment of the present application, a second balancing manager 122 is further provided in each graphics processing cluster 110. The second balancing manager 122 is used to record the second workload of each of the multiple minimum-level parallel processing units 130 in the graphics processing cluster 110, and is also used to allocate work tasks.
[0047] In the embodiment of the present application, the structure of the second balancing manager 122 can be the same as that of the first balancing manager 121, except that they are located in different locations and are responsible for different parallel processing units. The first balancing manager 121, as the largest-level balancing manager, is responsible for cluster-level workload recording and task distribution, while the second balancing manager 122 is located within the graphics processing cluster 110 and is responsible for task distribution and workload recording for each smallest-level parallel processing unit 130 within the graphics processing cluster 110.
[0048] In an embodiment of the present application, when the graphics processor 100 receives a current work task and needs to distribute it, the first balancing manager 121 can distribute the current work task to each graphics processing cluster 110 based on the first workload of each graphics processing cluster 110, such as distributing it to the graphics processing cluster 110 with the smallest workload. When a graphics processing cluster 110 receives the current work task, the second balancing manager 122 within the graphics processing cluster 110 can redistribute the current work task to each minimum-level parallel processing unit 130 based on the second workload of each minimum-level parallel processing unit 130, such as distributing it to the minimum-level parallel processing unit 130 with the smallest workload, so that the target processing core of the minimum-level parallel processing unit 130 executes the current work task.
[0049] Accordingly, when a minimum-level parallel processing unit 130 completes a task, the second balancing manager 122 releases the minimum-level parallel processing unit 130 assigned to the completed task, counts the number of first registers in the minimum-level parallel processing unit 130 that completed the task, and updates the second workload based on the first register count. The second balancing manager 122 then reports the number of first registers that completed the task to the first balancing manager 121, allowing the first balancing manager 121 to also update the first workload. The first balancing manager 121 only needs to know the total number of remaining resources in each graphics processing cluster 110. Therefore, the first register count can be the total number of registers required by all minimum-level parallel processing units 130 to complete the task. This approach reduces the data processing workload required of higher-level balancing managers. The large number of lower-level balancing managers and their ability to operate simultaneously also helps improve data processing efficiency. Updating the first and second workloads can include increasing the number of registers for completed tasks in the corresponding workloads.
[0050] Through the first balancing manager 121 and the second balancing manager 122, two-level workload recording can be achieved. The second balancing manager 122 only needs to record the second workload of each minimum-level parallel processing unit 130 within the graphics processing cluster 110, and then summarize and feed it back to the first balancing manager 121. The first balancing manager 121 does not need to specifically record the load status of each minimum-level parallel processing unit 130, but only needs to obtain the overall load status of each graphics processing cluster 110 calculated by the second balancing manager 122. The second balancing managers 122 of different graphics processing clusters 110 operate simultaneously, which can effectively improve the efficiency of load recording. Furthermore, when distributing work tasks, the first workload and the second workload can be obtained in a timely manner to distribute the work tasks, so that the work tasks can be evenly distributed to each graphics processing cluster 110 and minimum-level parallel processing unit 130, making the workload of the minimum-level parallel processing unit 130 more balanced.
[0051] In addition to the two-level structure provided in the above embodiment, in the embodiment of the present application, the parallel processing unit of the graphics processor 100 can also be divided into a structure with more levels. Figure 3 , Figure 3 This is a second structural diagram of the graphics processor 100 provided in one embodiment of the present application.
[0052] In an embodiment of the present application, the graphics processing cluster 110 may further include a third parallel processing unit 140 , and the third parallel processing unit 140 includes at least two minimum-level parallel processing units 130 .
[0053] Each graphics processing cluster 110 includes multiple second balancing managers 122. Each second balancing manager 122 is respectively set in each third parallel processing unit 140. The second balancing manager 122 is used to record the second workload of each of the multiple minimum-level parallel processing units 130 in the third parallel processing unit 140.
[0054] like Figure 3 As shown, in this embodiment, the level of graphics processing cluster 110 is greater than the level of third parallel processing unit 140, and the level of third parallel processing unit 140 is greater than the level of minimum-level parallel processing unit 130. For example, in some embodiments of the present application, minimum-level parallel processing unit 130 may be an HPPU (Half Parallel Processor Unit), and third parallel processing unit 140 may be a conventional PPU (Parallel Processor Unit).
[0055] Accordingly, since the second balancing manager 122 is divided into the third parallel processing units 140, each third parallel processing unit 140 also needs to record the workload and distribute work tasks. Therefore, in the embodiment of the present application, each graphics processing cluster 110 may further include a third balancing manager 123. The third balancing manager 123 is used to record the third workload of each of the plurality of third parallel processing units 140 in the graphics processing cluster 110, receive current work tasks, and redistribute the current work tasks to each third parallel processing unit 140 based on the third workload of each third parallel processing unit 140.
[0056] Due to the further hierarchical structure, in this embodiment, the second balancing manager 122 can reduce the number of minimum-level parallel processing units 130 that it needs to be responsible for. For example:
[0057] The graphics processor 100 has 64 parallel processing units. If divided into two levels, the 64 parallel processing units are divided into four graphics processing clusters 110. Each graphics processing cluster 110 includes 16 minimum-level parallel processing units 130. The number of first balancing managers 121 is 1, and the number of second balancing managers 122 is 4. Each second balancing manager 122 is responsible for 16 minimum-level parallel processing units 130.
[0058] If divided into three levels, the 64 parallel processing units are divided into four graphics processing clusters 110, each graphics processing cluster 110 includes four third parallel processing units 140, the number of first balancing managers 121 is 1, the number of third balancing managers 123 is 4, and the number of second balancing managers 122 is 4. Each third parallel processing unit 140 includes four minimum-level parallel processing units 130.
[0059] By further stratifying the hierarchy, the number of lower-level balancing managers is increased, reducing the number of parallel processing units each lower-level balancing manager is responsible for. This allows more balancing managers to work in parallel, and reduces the amount of data each balancing manager must process, effectively improving the efficiency of the balancing managers. This increases the timeliness of workload recording and task allocation for the parallel processing units at their respective levels, thereby helping to improve the balance of task distribution.
[0060] Similar to the aforementioned two-level structure, in the three-level structure, when the graphics processor 100 receives a current work task and needs to distribute it, the first balancing manager 121 can distribute the current work task to each graphics processing cluster 110 based on the first workload of each graphics processing cluster 110. When a graphics processing cluster 110 receives the current work task, the third balancing manager 123 within the graphics processing cluster 110 redistributes the current work task to each third parallel processing unit 140 based on the third workload of each third parallel processing unit 140. After receiving the current work task, the second balancing manager 122 within the third parallel processing unit 140 redistributes the current work task to each minimum-level parallel processing unit 130 based on the second workload of each minimum-level parallel processing unit 130 within the third parallel processing unit 140, so that the target processing core of the minimum-level parallel processing unit 130 executes the current work task.
[0061] Accordingly, when the registers of the target processing core of a minimum-level parallel processing unit 130 complete the assigned work task, the second balancing manager 122 releases the minimum-level parallel processing unit 130 that has completed the work task, receives the first register count of the completed work task from the minimum-level parallel processing unit 130, calculates the first register counts of all minimum-level parallel processing units 130 to obtain a second register count, and updates the second workload based on the second register count. The second balancing manager 122 then feedbacks the second register count of the completed work task to the third balancing manager 123. The third balancing manager 123 receives the second register counts reported by the second balancing managers 122 of all third parallel processing units 140 within the graphics processor cluster and calculates them as the third register count. Each third balancing manager 123 updates the third workload based on the third register count. Each third balancing manager 123 then feedbacks its respective third register count to the first balancing manager 121. After receiving all third register counts, the first balancing manager 121 updates the first workload.
[0062] Furthermore, the present invention provides a graphics processor 100 with a four-level parallel processing unit structure. Figure 4 , Figure 4 This is a third structural diagram of the graphics processor 100 provided in one embodiment of the present application. Figure 4 The provided graphics processor has a four-level parallel processing unit structure. In this four-level parallel processing unit structure, the graphics processing cluster 110 includes multiple fourth parallel processing units 150. Each fourth parallel processing unit 150 includes at least two third parallel processing units 140. In other words, the third parallel processing units 140 are further divided and configured in the fourth parallel processing units 150.
[0063] In this embodiment, the PPUs are, from largest to smallest, graphics processing cluster 110, fourth PPU 150, third PPU 140, and smallest PPU 130. In this embodiment, fourth PPU 150 may be a DPPU (Double Parallel Processor Unit), third PPU 140 may be a PPU, and smallest PPU 130 may be an HPPU.
[0064] Accordingly, each graphics processing cluster 110 includes multiple third balancing managers 123, each of which is respectively set in each fourth parallel processing unit 150. At the same time, each graphics processing cluster 110 also includes a fourth balancing manager 124, which records the fourth workload of each of the multiple fourth parallel processing units 150 in the graphics processing cluster 110 through the fourth balancing manager 124, receives current work tasks, and redistributes work tasks to each fourth parallel processing unit 150 based on the fourth workload of each fourth parallel processing unit 150.
[0065] In a four-level parallel processing unit architecture, when the graphics processor 100 receives a current work task and needs to distribute it, the first balancing manager 121 can distribute the current work task to each graphics processing cluster 110 based on the first workload of each graphics processing cluster 110. When a graphics processing cluster 110 receives the current work task, the fourth balancing manager 124 within that graphics processing cluster 110 distributes the current work task to each fourth parallel processing unit 150 based on the fourth workload. Then, the third balancing manager 123 within the fourth parallel processing unit 150, which receives the current work task, distributes the current work task to each third parallel processing unit 140 based on the third workload. After receiving the current work task, the second balancing manager 122 within the third parallel processing unit 140 redistributes the current work task to each minimum-level parallel processing unit 130 within that third parallel processing unit 140 based on the second workload of each minimum-level parallel processing unit 130 within that third parallel processing unit 140, so that the target processing core of each minimum-level parallel processing unit 130 executes the current work task.
[0066] When the registers of the target processing core complete the work task, the work performed in the graphics processor 100 is as follows:
[0067] For any of the minimum-level parallel processing units 130 , after the registers of the target processing core of the minimum-level parallel processing unit 130 complete the assigned work tasks, the minimum-level parallel processing unit 130 feeds back the first number of registers that have completed the work tasks to the second balancing manager 122 .
[0068] For each third parallel processing unit 140, the second balancing manager 122 of the parallel processing unit receives the first register quantity fed back by all the minimum-level parallel processing units 130 in the parallel processing unit and counts it as the second register quantity, updates the second workload based on the second register quantity, and feeds back the second register quantity to the third balancing manager 123 of the fourth parallel processing unit 150 where the third parallel processing unit 140 is located.
[0069] For any of the fourth parallel processing units 150, the third balancing manager 123 of the fourth parallel processing unit 150 receives the second register quantities fed back by all the third parallel processing units 140 in the fourth parallel processing unit 150 and counts them as the third register quantities, updates the third workload based on the third register quantities, and feeds back the third register quantities to the fourth balancing manager 124 of the graphics processing cluster 110 where the fourth parallel processing unit 150 is located.
[0070] For any one of the graphics processing clusters 110: the fourth balancing manager 124 in the graphics processing cluster 110 receives the third register quantity fed back by all the fourth parallel processing units 150 in the graphics processing cluster 110 and counts it as the fourth register quantity, updates the fourth workload based on the fourth register quantity, and feeds back the fourth register quantity to the first balancing manager 121.
[0071] Finally, the first balancing manager 121 is configured to update the first workload based on the fourth register quantities fed back by all graphics processing clusters 110 .
[0072] In an embodiment of the present application, each level of the balance manager provides feedback to the upper level step by step. The lower-level balance manager first records the registers that have completed the task. The higher-level balance manager does not need to count the number of registers in the minimum-level parallel processing unit 130 one by one. On the one hand, the amount of data fed back upward is reduced, and on the other hand, the difficulty of data processing by the higher-level balance manager is reduced. As a result, the efficiency of data statistics on the number of registers after completing the task can be improved, and work tasks can be promptly assigned to parallel processing units with low loads, thereby improving the balance of work task allocation. In addition, the number of lower-level balance managers is the largest, and multiple lower-level balance managers are more efficient in recording the registers that have completed the task. As a result, the efficiency of counting the number of registers that have completed the task can be improved, and the timeliness and balance of task allocation can be improved.
[0073] In the graphics processor 100 provided in this application, the parallel processing units are divided into four levels to achieve balanced distribution of work tasks among the four levels of parallel processing units. Different balance managers are responsible for each level of parallel processing units, which respectively record the workload of each level and perform task distribution. Each balance manager works in a hierarchical and simultaneous manner, effectively improving the balance of work task distribution and improving the parallel data processing performance of the graphics processor 100. It is understood that when the number of parallel processing units in the graphics processor 100 is large, the levels can be further divided, for example, into a five-level structure, a six-level structure, etc. When divided into two or more levels, each balance manager in the graphics processor 100 can achieve the purpose of improving the timeliness of workload updates and work task distribution, thereby improving the balance of work task distribution and improving the parallel data processing performance of the graphics processor 100. Therefore, the above is only an example and should not be construed as limiting the present application.
[0074] In the above embodiment, when allocating work tasks, the balancing manager at each level may allocate the work tasks to the parallel processing unit at the next level with the smallest load, so as to balance the load among the parallel processing units.
[0075] For example, when the graphics processor 100 has a four-level structure:
[0076] The first balancing manager 121 is configured to determine the graphics processing cluster 110 with the least load based on the first workloads of the graphics processing clusters 110 , and allocate the current work task to the graphics processing cluster 110 with the least load.
[0077] The fourth balancing manager 124 is configured to determine the fourth PPU 150 with the least load based on the fourth workloads of the fourth PPUs 150 in the GPU 110 , and allocate the current task to the fourth PPU 150 with the least load.
[0078] The third balancing manager 123 is configured to determine the third PPU 140 with the smallest load based on the third workloads of the third PPUs 140 in the fourth PPU 150 , and allocate the current work task to the third PPU 140 with the smallest load.
[0079] The second balancing manager 122 is configured to determine the minimum-level parallel processing unit 130 with the smallest load based on the second workload of each minimum-level parallel processing unit 130 in the third parallel processing unit 140 , and allocate the current work task to the minimum-level parallel processing unit 130 with the smallest load.
[0080] For example, when the graphics processor 100 is a two-level structure: the first balancing manager 121 is used to determine the graphics processing cluster 110 with the least load based on the first workload of each graphics processing cluster 110, and allocate the current work task to the graphics processing cluster 110 with the least load; the second balancing manager 122 is used to determine the minimum-level parallel processing unit 130 with the least load based on the second workload of each minimum-level parallel processing unit 130 in the graphics processing cluster 110, and allocate the current work task to the minimum-level parallel processing unit 130 with the least load.
[0081] The tertiary structure or other level structures can be deduced by analogy and will not be expanded here.
[0082] There are delays in the allocation of work tasks, affecting the timeliness of workload updates and work task allocation. Work tasks are somewhat correlated, and new work tasks may be generated when the original work tasks are completed. For example, after the vertex shader task is completed, a tessellation control shader task will be generated. Therefore, in the embodiments of the present application, it is possible to predict the derivative work tasks that will be generated and pre-assign corresponding parallel processing units to the derivative work tasks that will be generated.
[0083] For low-level PPUs, they may not have enough registers to execute derivative work tasks. Therefore, in an embodiment of the present application, pre-allocated derivative work tasks will be executed in high-level PPUs, for example, they can be pre-allocated in the graphics processing cluster 110 and the fourth PPU 150.
[0084] Taking the pre-allocation of the graphics processing cluster 110 as an example, in some embodiments of the present application, the first balancing manager 121 is also used to: predict the derivative work tasks that will be generated based on the current work tasks and the preset work task execution logic; determine the graphics processing cluster 110 with the smallest load based on the first workload; pre-allocate the graphics processing cluster 110 with the smallest load to execute the derivative work tasks; and update the first workload based on the pre-allocation results of the derivative work tasks.
[0085] In this embodiment, the task execution logic can be determined based on the relationship between the tasks. For example, there is a derivative relationship between the vertex shader task and the tessellation control shader task. The task execution logic may vary for different task types and will not be further elaborated here.
[0086] After the graphics processor 100 receives the current work task, it can predict the derivative work tasks that may be generated by the current work task based on the preset work task execution logic, and allocate the corresponding graphics processing cluster 110 to the derivative work tasks in advance, so that after the derivative work tasks are generated, the derivative work tasks can be directly sent to the pre-allocated graphics processing cluster 110 for subsequent allocation and execution.
[0087] In an embodiment of the present application, after pre-assigning the derived work tasks, the first workload of the corresponding graphics processing cluster 110 can also be updated so that the first workload records the workload to be generated in advance. Therefore, when allocating work tasks, since the future workload can be known in advance, work tasks will no longer be allocated to the graphics processing cluster 110 whose future load will already be very large.
[0088] In an embodiment of the present application, after the derivative work task is generated, since the first workload of the graphics processing cluster 110 has been updated in the pre-allocation stage, the first workload will no longer be updated after the derivative work task is generated.
[0089] Similarly, if pre-allocation is performed in the fourth parallel processing unit 150, in this embodiment, the fourth balancing manager 124 is also used to: predict the derivative work tasks that will be generated based on the current work tasks and the preset work task execution logic; determine the fourth parallel processing unit 150 with the smallest load based on the fourth workload; pre-allocate the fourth parallel processing unit 150 with the smallest load to execute the derivative work tasks; and update the first workload based on the pre-allocation results of the derivative work tasks.
[0090] Through the above-mentioned pre-allocation, the delay in allocating derivative work tasks can be reduced, and the first workload can be updated after pre-allocating derivative work tasks, so that when allocating other work tasks, the future workload of the graphics processing cluster 110 can be known in advance, thereby reducing the situation where the parallel processing unit is allocated to derivative work tasks and other work tasks at the same time, which makes the parallel processing unit overloaded, and improving the balance of work task allocation.
[0091] In an embodiment of the present application, during pre-allocation, different parallel processing units may be configured with different numbers of target processing cores, resulting in different performances of different parallel processing units. Even if the work tasks are assigned to the parallel processing unit with the smallest load, it may still experience a large and uneven load.
[0092] In one embodiment of the present application, the first balancing manager 121 is further used to: predict the derivative work tasks that will be generated based on the current work tasks and the preset work task execution logic; calculate the target resource amount required to execute the derivative work tasks; determine the target graphics processing cluster 110 with the largest remaining resource amount based on the first workload of each graphics processing cluster 110; pre-allocate the target graphics processing cluster 110 to execute the derivative work tasks and update the target graphics processing cluster 110 and update the first workload.
[0093] In an embodiment of the present application, derivative work tasks are allocated according to the target resource amount required for the derivative work tasks and the remaining resource amount of the graphics processing cluster 110. The performance differences between different parallel processing units can be taken into account to balance the distribution of derivative work tasks based on the parallel processing units, thereby improving the balanced distribution effect.
[0094] In the above embodiment, the resource can be the number of registers, the target resource is the required number of registers, and the remaining resource is the number of free registers. The resource can also be the number of tasks, the target resource is the number of subtasks within the current task, and the remaining resource is the number of tasks that the parallel processing unit can process.
[0095] Similarly, when the above-mentioned pre-allocation is performed according to the amount of resources, it can be performed on the third balancing manager and the fourth balancing manager 124 .
[0096] Based on the same inventive concept, an embodiment of the present application also provides a load balancing method for a graphics processor, which can be applied to the graphics processor provided by any of the aforementioned embodiments. The graphics processor includes a first balancing manager and multiple graphics processing clusters, each graphics processing cluster includes multiple minimum-level parallel processing units, the minimum-level parallel processing units include multiple target processing cores, and the target processing cores may include at least one of a shading processing core and a computing processing core.
[0097] See also Figure 5 , Figure 5 This is a flow chart of a load balancing method for a graphics processor provided in one embodiment of the present application. The load balancing method for a graphics processor includes:
[0098] S510, receiving the current work task.
[0099] S520: Obtain a first workload of each graphics processing cluster through a first balancing manager.
[0100] S530: Allocate the current work task to the target graphics processing cluster based on the first workload through the first balancing manager.
[0101] S540: When the target graphics processing cluster receives the current work task, the target graphics processing cluster obtains the second workload of each of the plurality of minimum-level parallel processing units in the target graphics processing cluster through the target second balancing manager of the target graphics processing cluster.
[0102] S550: reallocate the current work task to the target minimum-level parallel processing unit based on the second workload through the target second balancing manager, so that the target processing core of the target minimum-level parallel processing unit executes the current work task.
[0103] In one embodiment, if a graphics processing cluster includes a third parallel processing unit (TPU), and the third parallel processing unit includes at least two minimum-level TPUs, each graphics processing cluster includes multiple second balancing managers, each second balancing manager being respectively disposed within a third parallel processing unit, and the second balancing manager being configured to record the respective second workloads of the multiple minimum-level TPUs within the third parallel processing unit. Each graphics processing cluster also includes a third balancing manager; the third balancing manager is configured to record the respective third workloads of the multiple third parallel processing units within the graphics processing cluster. Obtaining the respective second workloads of the multiple minimum-level TPUs within the target graphics processing cluster through the target second balancing manager of the target graphics processing cluster includes: obtaining the third workload of each third parallel processing unit within the target graphics processing cluster through the target third balancing manager within the target graphics processing cluster; allocating the current work task to the target third parallel processing unit based on the third workload through the target third balancing manager; and obtaining the respective second workloads of the multiple minimum-level TPUs within the target graphics processing cluster based on the target second balancing manager within the target third parallel processing unit.
[0104] In one embodiment, the graphics processing cluster includes multiple fourth parallel processing units (PPUs); each of the fourth PPUs includes at least two third PPUs; each PPU includes multiple third balancing managers, each of which is respectively disposed within each of the fourth PPUs; each PPU further includes a fourth balancing manager; and the fourth balancing manager is configured to record the fourth workloads of the respective fourth PPUs within the PPU. Obtaining the third workload of each third PPU within the target graphics processing cluster through the target third balancing manager within the target graphics processing cluster includes: obtaining the fourth workload of each fourth PPU within the target graphics processing cluster through the target fourth balancing manager within the target graphics processing cluster; allocating the current work task to the target fourth PPU based on the fourth workload through the target fourth balancing manager; and obtaining the third workload based on the target third balancing manager within the target fourth PPU.
[0105] In one embodiment, the target processing core includes a plurality of registers, and the target processing core executes the assigned work tasks based on the registers; the first workload, the second workload, the third workload, and the fourth workload each include a number of registers assigned to the work tasks; and the method further includes:
[0106] After the registers of the target processing core of the minimum-level parallel processing unit complete the assigned work tasks, the minimum-level parallel processing unit feeds back the first number of registers that have completed the work tasks to the target upper-level second balancing manager; the target upper-level second balancing manager receives the first number of registers fed back by all minimum-level parallel processing units and counts them as the second number of registers; the target upper-level second balancing manager updates the second workload based on the second number of registers; the target upper-level second balancing manager feeds back the second number of registers to the target upper-level third balancing manager; the target upper-level third balancing manager receives the second number of registers of all target upper-level second balancing managers and counts them as the third number of registers; the target upper-level third balancing manager updates the third workload based on the third number of registers; the target upper-level third balancing manager feeds back the third number of registers to the target upper-level fourth balancing manager; the target upper-level fourth balancing manager receives the third number of registers of all target upper-level third balancing managers and counts them as the fourth number of registers; the target upper-level fourth balancing manager updates the fourth workload based on the fourth number of registers; the fourth balancing manager feeds back the fourth number of registers to the first balancing manager; the first balancing manager updates the first workload based on the fourth number of registers fed back by all fourth balancing managers.
[0107] In one embodiment, allocating the current work task to the target graphics processing cluster based on the first workload by the first balancing manager includes: determining the graphics processing cluster with the smallest load based on the first workload of each graphics processing cluster by the target first balancing manager, and allocating the current work task to the graphics processing cluster with the smallest load. Furthermore, allocating the current work task to the target minimum-level parallel processing unit based on the second workload by the target second balancing manager includes: determining the minimum-level parallel processing unit with the smallest load based on the second workload of each minimum-level parallel processing unit within the graphics processing cluster by the target second balancing manager, and allocating the current work task to the minimum-level parallel processing unit allocated to the minimum-level parallel processing unit with the smallest load.
[0108] In one embodiment, after receiving the current work task, the method further includes: predicting the derivative work task that will be generated based on the current work task based on the current work task and the preset work task execution logic; determining the graphics processing cluster with the smallest load based on the first workload; pre-allocating the graphics processing cluster with the smallest load to execute the derivative work task; and updating the first workload based on the pre-allocation result of the derivative work task.
[0109] In one embodiment, after receiving the current work task, the method further includes: predicting the derivative work tasks that will be generated based on the current work task based on the current work task and the preset work task execution logic; calculating the target amount of resources required to execute the derivative work tasks; determining the target graphics processing cluster with the largest amount of remaining resources based on the first workload of each graphics processing cluster; pre-allocating the target graphics processing cluster to execute the derivative work tasks and updating the target graphics processing cluster and updating the first workload.
[0110] In one embodiment, the current work task includes multiple subtasks, and the target resource amount includes the number of registers required to execute the derived work task, or the number of subtasks of the derived work task.
[0111] The load balancing method of the graphics processor can refer to the structure and power consumption of the aforementioned graphics processor, which will not be further elaborated here.
[0112] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, which includes the graphics processor provided by any of the aforementioned embodiments.
[0113] In the embodiments of the present application, the electronic device may be a computer, a server, a mobile phone, etc., which is not limited here.
[0114] The electronic device may also include common modules required by the electronic device, such as a memory and a communication module, which will not be introduced one by one in the embodiments of the present application.
[0115] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0116] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A graphics processor, characterized in that: include: a plurality of graphics processing clusters, each graphics processing cluster including a plurality of minimum-level parallel processing units; the minimum-level parallel processing units including a plurality of target processing cores, the target processing cores including at least one of a shading processing core and a computational processing core, each of the target processing cores being configured to execute an assigned work task; a first balancing manager configured to record a first workload of each of the graphics processing clusters, and, upon receiving a current workload, distribute the current workload to each of the graphics processing clusters based on the first workload of each of the graphics processing clusters; Each of the graphics processing clusters includes a second balancing manager, which is used to record the second workload of each of the multiple minimum-level parallel processing units in the graphics processing cluster, and to receive the current work task and redistribute the current work task to each of the minimum-level parallel processing units based on the second workload of each of the minimum-level parallel processing units, so that the target processing core of each of the minimum-level parallel processing units executes the current work task.
2. The graphics processor according to claim 1, wherein: The graphics processing cluster includes a third parallel processing unit; the third parallel processing unit includes at least two of the minimum-level parallel processing units; Each of the graphics processing clusters includes a plurality of second balancing managers, each of which is respectively disposed in a respective third parallel processing unit, and the second balancing manager is configured to record respective second workloads of the plurality of minimum-level parallel processing units in the third parallel processing unit; Each of the graphics processing clusters also includes a third balancing manager; the third balancing manager is used to record the third workload of each of the multiple third parallel processing units in the graphics processing cluster, receive the current work task, and redistribute the current work task to each of the third parallel processing units based on the third workload of each of the third parallel processing units.
3. The graphics processor according to claim 2, wherein: The graphics processing cluster includes a plurality of fourth parallel processing units; each of the fourth parallel processing units includes at least two of the third parallel processing units; Each of the graphics processing clusters includes a plurality of the third balancing managers, and each of the third balancing managers is respectively disposed in each of the fourth parallel processing units; Each of the graphics processing clusters also includes a fourth balancing manager; the fourth balancing manager is used to record the fourth workload of each of the multiple fourth parallel processing units in the graphics processing cluster, receive the current work tasks, and redistribute the work tasks to each of the third parallel processing units based on the fourth workload of each of the fourth parallel processing units.
4. The graphics processor according to claim 3, wherein: The target processing core includes a plurality of registers, and the target processing core executes the assigned work tasks based on the registers; the first workload, the second workload, the third workload, and the fourth workload each include a number of registers assigned to the work tasks; For any of the minimum-level parallel processing units: after the registers of the target processing core of the minimum-level parallel processing unit complete the assigned work tasks, the minimum-level parallel processing unit feeds back to the second balancing manager the first number of registers that have completed the work tasks; For each of the third parallel processing units: a second balancing manager of the parallel processing unit receives the first register quantity fed back by all the lowest-level parallel processing units in the parallel processing unit and counts the first register quantity as the second register quantity, updates the second workload based on the second register quantity, and feeds back the second register quantity to a third balancing manager of the fourth parallel processing unit where the third parallel processing unit is located; For any one of the fourth parallel processing units: the third balancing manager of the fourth parallel processing unit receives the second register quantities fed back by all third parallel processing units in the fourth parallel processing unit and counts them as the third register quantities, updates the third workload based on the third register quantities, and feeds back the third register quantity to the fourth balancing manager of the graphics processing cluster where the fourth parallel processing unit is located; For any one of the graphics processing clusters: a fourth balancing manager in the graphics processing cluster receives third register quantities fed back by all fourth parallel processing units in the graphics processing cluster and counts the third register quantities as fourth register quantities, updates the fourth workload based on the fourth register quantities, and feeds the fourth register quantity back to the first balancing manager; The first balancing manager is configured to update the first workload based on the fourth register quantity fed back by all the graphics processing clusters.
5. The graphics processor according to claim 1, wherein: The first balancing manager is configured to: determine a graphics processing cluster with the smallest load based on the first workloads of the graphics processing clusters, and allocate the current work task to the graphics processing cluster with the smallest load; The second balancing manager is configured to determine a minimum-level parallel processing unit with the smallest load based on the second workload of each minimum-level parallel processing unit in the graphics processing cluster, and allocate the current work task to the minimum-level parallel processing unit with the smallest load.
6. The graphics processor according to claim 1, wherein: The first balancing manager is further configured to: Based on the current task and the preset task execution logic, predict the derivative tasks that will be generated based on the current task; Determining a graphics processing cluster with a minimum load based on the first workload; pre-assigning the graphics processing cluster with the smallest load to execute the derived work task; The first workload is updated according to the pre-allocation result of the derived work tasks.
7. The graphics processor according to claim 1, wherein: The first balancing manager is further configured to: Based on the current task and the preset task execution logic, predict the derivative tasks that will be generated based on the current task; Calculating target resource amounts required to perform the derived work tasks; Determining a target graphics processing cluster with the largest amount of remaining resources based on the first workload of each graphics processing cluster; The target graphics processing cluster is pre-assigned to execute the derived work task, and the target graphics processing cluster and the first workload are updated.
8. The graphics processor according to claim 7, wherein: The current work task includes multiple subtasks, and the target resource amount includes the number of registers required to execute the derived work task, or the number of subtasks of the derived work task.
9. A method for load balancing a graphics processor, characterized in that: The graphics processor includes a first balancing manager and a plurality of graphics processing clusters, each of the graphics processing clusters includes a plurality of minimum-level parallel processing units, the minimum-level parallel processing units include a plurality of target processing cores, and the target processing cores include at least one of a shading processing core and a computational processing core. The load balancing method includes: Receive current work tasks; Obtaining, by the first balancing manager, a first workload of each of the graphics processing clusters; allocating, by the first balancing manager, the current work task to a target graphics processing cluster based on the first workload; When the target graphics processing cluster receives the current work task, the target graphics processing cluster obtains, through the target second balancing manager of the target graphics processing cluster, the second workload of each of the plurality of minimum-level parallel processing units in the target graphics processing cluster; The target second balancing manager redistributes the current work task to the target minimum-level parallel processing unit based on the second workload, so that the target processing core of the target minimum-level parallel processing unit executes the current work task.
10. The method according to claim 9, characterized in that The graphics processing cluster includes a third parallel processing unit; the third parallel processing unit includes at least two minimum-level parallel processing units; each of the graphics processing clusters includes a plurality of second balancing managers, each of the second balancing managers is respectively disposed in a respective third parallel processing unit, and the second balancing manager is configured to record respective second workloads of the plurality of minimum-level parallel processing units within the third parallel processing unit; Each of the graphics processing clusters further includes a third balancing manager; the third balancing manager is configured to record a third workload of each of the plurality of third parallel processing units in the graphics processing cluster; The step of obtaining, through the target second balancing manager of the target graphics processing cluster, the second workload of each of the plurality of minimum-level parallel processing units in the target graphics processing cluster includes: Obtaining, by a target third balancing manager in the target graphics processing cluster, a third workload of each of the third parallel processing units in the target graphics processing cluster; Allocating the current work task to the target third parallel processing unit based on the third workload by the target third balancing manager; The target second balancing manager in the target third parallel processing unit obtains the second workload of each of the plurality of minimum-level parallel processing units in the target graphics processing cluster.
11. The method according to claim 10, characterized in that The graphics processing cluster includes a plurality of fourth parallel processing units; Each of the fourth parallel processing units includes at least two of the third parallel processing units; each of the graphics processing clusters includes a plurality of the third balancing managers, each of the third balancing managers being respectively disposed in each of the fourth parallel processing units; Each of the graphics processing clusters further includes a fourth balancing manager; the fourth balancing manager is configured to record a fourth workload of each of the plurality of fourth parallel processing units in the graphics processing cluster; The obtaining, by the target third balancing manager in the target graphics processing cluster, the third workload of each of the third parallel processing units in the target graphics processing cluster comprises: Obtaining, by a target fourth balancing manager in the target graphics processing cluster, a fourth workload of each of the fourth parallel processing units in the target graphics processing cluster; allocating the current work task to a target fourth parallel processing unit based on the fourth workload by the target fourth balancing manager; A third workload is obtained based on a target third balancing manager in the target fourth parallel processing unit.
12. The method according to claim 11, wherein the target processing core comprises a plurality of registers, and the target processing core executes the assigned work tasks based on the registers; the first workload, the second workload, the third workload, and the fourth workload each comprise a number of registers assigned to the work tasks; The method further comprises: After the registers of the target processing core of any of the minimum-level parallel processing units complete the assigned work tasks, the minimum-level parallel processing unit feeds back the first number of registers that have completed the work tasks to the target upper-level second balancing manager; The target upper-level second balancing manager receives the first register quantity fed back by all minimum-level parallel processing units and counts it as the second register quantity; the target upper-level second balancing manager updates the second workload based on the second register quantity; the target upper-level second balancing manager feeds back the second register quantity to the target upper-level third balancing manager; The target upper level third balancing manager receives the second register quantities of all target upper level second balancing managers and counts the second register quantities as the third register quantity; The target upper level third balancing manager updates the third workload based on the third register quantity; the target upper level third balancing manager feeds back the third register quantity to the target upper level fourth balancing manager; The target upper-level fourth balancing manager receives the third register quantities of all target upper-level third balancing managers and counts them as the fourth register quantity; the target upper-level fourth balancing manager updates the fourth workload based on the fourth register quantity; the fourth balancing manager feeds back the fourth register quantity to the first balancing manager; The first balancing manager updates the first workload based on the fourth register quantities fed back by all the fourth balancing managers.
13. The method according to claim 9, characterized in that Allocating the current work task to a target graphics processing cluster based on the first workload by the first balancing manager includes: The target first balancing manager determines the graphics processing cluster with the smallest load based on the first workloads of the graphics processing clusters; and allocates the current work task to the graphics processing cluster with the smallest load; And, redistributing the target work task and the current work task to the target minimum-level parallel processing unit based on the second workload by the target second balancing manager includes: The target second balancing manager determines the minimum-level parallel processing unit with the smallest load based on the second workload of each minimum-level parallel processing unit in the graphics processing cluster, and allocates the current work task to the minimum-level parallel processing unit with the smallest load.
14. The method according to claim 9, characterized in that After receiving the current work task, the method further includes: The first balancing manager predicts the derivative work tasks that will be generated based on the current work tasks based on the current work tasks and the preset work task execution logic; determines the graphics processing cluster with the smallest load based on the first workload; pre-allocates the graphics processing cluster with the smallest load to execute the derivative work tasks; and updates the first workload according to the pre-allocation results of the derivative work tasks.
15. The method according to claim 9, characterized in that After receiving the current work task, the method further includes: The first balancing manager predicts the derivative work tasks that will be generated based on the current work tasks based on the current work tasks and the preset work task execution logic; calculates the target resource amount required to execute the derivative work tasks; determines the target graphics processing cluster with the largest remaining resources based on the first workload of each graphics processing cluster; pre-allocates the target graphics processing cluster to execute the derivative work tasks and updates the target graphics processing cluster and the first workload.
16. The method according to claim 15, characterized in that The current work task includes multiple subtasks, and the target resource amount includes the number of registers required to execute the derived work task, or the number of subtasks of the derived work task.
17. An electronic device, characterized in that: The graphics processor comprises the graphics processor according to any one of claims 1 to 8.
Citation Information
Cited By
Chip element task processing method and device, equipment, storage medium and program product
CN120823089A