Chip element task processing method and device, equipment, storage medium and program product
By dividing the tile task into evenly distributed fragment subtasks in the graphics processor and processing them in parallel, the load imbalance problem in the tile rendering architecture is solved, and rendering performance is improved.
Patent Information
- Application Number
- CN202511339931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Traditional graphics processors suffer from uneven load distribution in their tile rendering architecture, resulting in low overall rendering performance.
The task distribution unit receives the fragment task corresponding to the tile and divides it into at least two fragment subtasks. At least two processing units process these subtasks in parallel, where each subtask corresponds to a pixel sub-region of the same size, ensuring load balancing for each processing unit.
This achieves load balancing between processing units and improves overall rendering performance.
Smart Images

Figure CN120823089A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to, but is not limited to, the field of image processing technology, and in particular to a fragment task processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] A graphics processing unit (GPU) is a dedicated graphics rendering device used to process and display computerized graphics. However, GPUs typically use tile-based rendering (TBR) to render graphics. This approach divides the image into tiles, allowing each tile to fit into the on-chip cache. Consequently, the image can be rendered tile by tile to render each tile of the scene.
[0003] However, during the rendering process of the traditional TBR architecture, there is an uneven load distribution between different processing units, which results in low overall rendering performance. Summary of the Invention
[0004] In view of this, embodiments of the present application provide at least one fragment task processing method, apparatus, device, storage medium, and program product.
[0005] The technical solution of the embodiment of the present application is implemented as follows: In one aspect, an embodiment of the present application provides a fragment task processing method, the method comprising: A task distribution unit receives a fragment task corresponding to a tile; the fragment task is divided into at least two fragment subtasks; and the at least two fragment subtasks are processed in parallel by at least two processing units; wherein the fragment subtask is a pixel processing task for processing at least two pixel blocks in the tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and different groups of pixel sub-regions have the same number and correspond to different processing units.
[0006] On the other hand, an embodiment of the present application provides a fragment task processing device, the fragment task processing device includes a graphics processor, the graphics processor includes a task distribution unit and at least two processing units; wherein, The task distribution unit is configured to receive a fragment task corresponding to a tile through the task distribution unit; split the fragment task into at least two fragment subtasks; The at least two processing units are used to process the at least two fragment subtasks in parallel; wherein the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and different groups of pixel sub-regions have the same number and correspond to different processing units.
[0007] On the other hand, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.
[0008] On the other hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements some or all of the steps in the above method when executed by a processor.
[0009] On the other hand, an embodiment of the present application provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements some or all of the steps in the above method.
[0010] Based on the above embodiments provided by the present application, since the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; at least two pixel blocks are located in the same group of pixel sub-regions, and all pixel sub-regions have the same size, so that the area of each group of pixel sub-regions is the same, it is further explained that the areas of the pixel sub-regions corresponding to the at least two fragment subtasks are the same, resulting in the fragment task being divided by the task distribution unit, and the loads of the at least two fragment subtasks being substantially the same. By processing at least two fragment subtasks in parallel by at least two processing units, compared with the existing method of processing a fragment task by one processing unit, the load of one processing unit can be reduced, and load balancing between the processing units corresponding to the at least two fragment subtasks can be achieved, thereby improving the overall rendering performance.
[0011] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.
[0013] Figure 1 A typical TBR pipeline flow diagram provided in the embodiment of the present application; Figure 2 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 1 ; Figure 3 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 2 ; Figure 4 A schematic diagram of a block division provided in an embodiment of the present application; Figure 5 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 3 ; Figure 6 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 4 ; Figure 7 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 5 ; Figure 8 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 6 ; Figure 9 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 7 ; Figure 10 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 8 ; Figure 11A Schematic diagram of address remapping of a fragment task processing method provided in an embodiment of the present application Figure 1 ; Figure 11B Schematic diagram of address remapping of a fragment task processing method provided in an embodiment of the present application Figure 2 ; Figure 11C Schematic diagram of address remapping of a fragment task processing method provided in an embodiment of the present application Figure 3 ; Figure 12 A schematic diagram of a fragment task processing process provided in an embodiment of the present application; Figure 13A A schematic diagram of a rendering scene of a fragment task processing method provided in an embodiment of the present application; Figure 13B A schematic diagram of the execution time of each processing unit in a rendering scenario of a fragment task processing method provided in an embodiment of the present application; Figure 14 A schematic diagram of an improved TBR pipeline flow chart provided in an embodiment of the present application; Figure 15A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 9 ; Figure 16 A schematic diagram of the structure of a task processing unit of a fragment task processing method provided in an embodiment of the present application; Figure 17 A schematic diagram of pixel sub-region division for a fragment task processing method provided in an embodiment of the present application; Figure 18 A schematic diagram of pixel block division for a fragment task processing method provided in an embodiment of the present application; Figure 19 A schematic diagram of a pixel sub-region of a fragment task processing method provided in an embodiment of the present application; Figure 20 A schematic diagram of color remapping of a fragment task processing method provided in an embodiment of the present application; Figure 21 A schematic diagram of the structure of a fragment task processing device provided in an embodiment of the present application; Figure 22 A hardware entity diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0015] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. The terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or sequence of "first / second / third" may be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.
[0017] Tile-based rendering is the process of subdividing a computer graphics image into a regular grid in optical space and rendering each portion of the grid, or tile, separately. The advantage of this design is that it reduces memory and bandwidth consumption compared to immediate-mode rendering systems that draw the entire frame at once. This makes tile rendering systems common in low-power hardware devices. Tile rendering is sometimes also called a sort middle architecture because it sorts geometry in the middle of the graphics pipeline rather than near the end. TBR is the most common architecture used on mobile GPUs and has significant advantages in reducing power consumption.
[0018] Typical TBR pipeline process is as follows: Figure 1 As shown, the TBR pipeline process is divided into a front-end module 110 and a back-end module 120, wherein the front-end module 110 includes a vertex processing module 111, a graphics processing module 112 and a tiling module 113; the back-end module 120 includes a rasterization module 121, a hidden surface removal (HSR) module 122, a pixel shading module 123 and an output merging module 124.
[0019] Front-end module 110 performs vertex and primitive transformations (vertex processing), graphics processing (including clipping / culling, etc.), and then, during the tiling phase, completes screen segmentation, records the graphics data covering the tiles, and writes this generated information to system memory 130. This allows system memory 130 to store tile information (PrimitiveList) and vertex information (Vertex Data). The Primitive List is a fixed-length array of the same length as the tile. Each element in the array is a linked list containing pointers to all triangles intersecting the current tile, with the pointers pointing to Vertex Data. Vertex Data stores vertex and vertex attribute data.
[0020] The backend module 120 primarily performs operations such as rasterization, depth testing, and pixel shading, ultimately outputting the data to the render target. For each tile, since the data size is small, the required depth data, texture data, or color data can be loaded into the GPU's on-chip static random-access memory (SRAM), i.e., the on-chip memory 140 in the figure. For example, the hidden surface removal module 122 can store depth data in the depth buffer in the on-chip memory 140, the pixel shading module 123 can store texture data in the texture buffer in the on-chip memory 140, and the output merging module 124 can store color data in the color buffer in the on-chip memory 140.
[0021] During the rendering process, the rendering object (image) is divided into multiple tiles, so that the on-chip memory 140 can accommodate all the data for each tile. After at least one drawing instruction arrives at the GPU, the front-end module 110 processes each drawing instruction in turn and stores the corresponding tile information and vertex information in the system memory 130 until the data stored in the system memory 130 reaches a preset threshold or at least one drawing instruction has been processed. The back-end module 120 reads the corresponding vertex information from the system memory 130 in units of tiles and performs subsequent processing. In this way, by changing the back-end module 120's access to the system memory 130 to the back-end module 120's access to the on-chip memory 140, rendering efficiency can be improved.
[0022] For TBR GPUs, processing units (PUs) typically handle the fragment shading stage. Specifically, each PU is responsible for shading the fragments of a small screen tile. Each tile has a primitive list that records which primitives cover that tile's area. Therefore, the size of each tile's primitive list determines the workload of the tile's rendering task. However, within a complete screen, the primitive lists for each tile vary in size, leading to an imbalance in workload and load between PUs.
[0023] Based on this, embodiments of the present application provide a fragment task processing method that can be executed by a processor of a computer device. The computer device may refer to a server, laptop, tablet computer, desktop computer, smart TV, set-top box, mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device), or other device with data processing capabilities.
[0024] Figure 2 A schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the method includes the following steps S101 to S103: Step S101: Receive a fragment task corresponding to a tile through a task distribution unit.
[0025] In some embodiments, the task distribution unit belongs to the back-end module of the TBR architecture and is a newly added module after the rasterization unit. It is mainly used to receive the fragment task corresponding to the tile and divide the fragment task corresponding to the tile into at least two fragment sub-tasks.
[0026] In some embodiments, tiles are obtained by dividing the screen into blocks using the front-end module of the TBR architecture. A fragment task is a pixel processing task for all pixel blocks in a tile. A pixel block is a sub-region containing at least two pixels obtained by dividing the area containing the tile according to a preset size.
[0027] Exemplarily, a pixel block includes 4x4 pixels.
[0028] It should be noted that after the screen is divided into multiple blocks by the front-end module of the TBR architecture, the block range of each block on the screen is the same, and the size of the block needs to meet the storage conditions of the on-chip memory. In the embodiment of the present application, the fragment task corresponding to a block is used as an example for explanation.
[0029] In some embodiments, a fragment task refers to a pixel processing task for pixels in a tile.
[0030] Step S102: Split the fragment task into at least two fragment subtasks.
[0031] In some embodiments, a fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile.
[0032] In some embodiments, there is no inclusion relationship between at least two fragment subtasks, that is, the pixel processing tasks included in the at least two fragment subtasks are not repeated. A fragment task contains at least two fragment subtasks, that is, the pixel processing task in the fragment task includes the pixel processing tasks in at least two fragment subtasks.
[0033] In some embodiments, a fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel subregions, and all pixel subregions have the same size, i.e., all pixel subregions have the same area. Furthermore, for different groups of pixel subregions, since the number of pixel subregions in different groups is the same and the area of each pixel subregion is the same, the area of each group of pixel subregions is equal, thereby enabling the task processing unit to evenly divide the area of the tile into multiple pixel subregions. Based on the position information of the multiple pixel subregions, the multiple pixel subregions are classified to obtain at least two groups of pixel subregions. The pixel processing tasks for all pixel blocks in the tile are used as fragment tasks, and the pixel processing tasks for the pixel blocks in each of the at least two groups of pixel subregions are used as fragment subtasks, thereby obtaining at least two fragment subtasks corresponding to the at least two groups of pixel subregions.
[0034] In some embodiments, the task processing unit divides all pixel blocks in a tile and classifies the divided pixel blocks to obtain at least two categories of pixel blocks. The pixel processing tasks for all pixel blocks in the tile are used as fragment tasks, and the pixel processing tasks for each of the at least two categories of pixel blocks are used as fragment subtasks, thereby obtaining at least two fragment subtasks corresponding to the at least two categories of pixel blocks.
[0035] In some embodiments, the pixel sub-region refers to a pixel block region obtained by dividing a block into regions.
[0036] It should be noted that a group of pixel sub-regions corresponds to one fragment sub-task, that is, different groups of pixel sub-regions correspond to different fragment sub-tasks.
[0037] Exemplarily, the task processing unit evenly divides the image block into two groups of pixel sub-regions, namely a first group of pixel sub-regions and a second group of pixel sub-regions. The pixel processing task for the pixel blocks in the first group of pixel sub-regions is used as the first fragment sub-task, and the pixel processing task for the pixel blocks in the second group of pixel sub-regions is used as the second fragment sub-task. Here, the number of pixel sub-regions in the first group of pixel sub-regions is the same as the number of pixel sub-regions in the second group of pixel sub-regions. For example, the image block is divided into 6*6 pixel sub-regions, of which 18 pixel sub-regions are in the first group of pixel sub-regions and the other 18 pixel sub-regions are in the second group of pixel sub-regions.
[0038] In other embodiments, the task processing unit divides the image block into a set of pixel processing tasks (passes) to obtain multiple original pixel blocks. Each original pixel block is segmented to obtain at least two groups of pixel blocks corresponding to each original pixel block. The pixel processing tasks for the same group of pixel blocks in each original pixel block are treated as a fragment subtask, thereby obtaining at least two fragment subtasks corresponding to the at least two groups of pixel blocks.
[0039] Exemplarily, the task processing unit divides the image block by pass to obtain pass0 and pass1. pass0 and pass1 are the original pixel blocks. For pass0, the pixel block information in pass0 is segmented to obtain two groups of pixel blocks (i.e., the first group of pixel blocks and the second group of pixel blocks); for pass1, the pixel block information in pass0 is segmented to obtain two groups of pixel blocks (i.e., the first group of pixel blocks and the second group of pixel blocks). The pixel processing task of the first group of pixel blocks of pass0 and the pixel processing task of the first group of pixel blocks of pass1 are used as a fragment subtask, and the pixel processing task of the second group of pixel blocks of pass0 and the pixel processing task of the second group of pixel blocks of pass1 are used as another fragment subtask.
[0040] It should be noted that one pass consists of 32 pixels.
[0041] Step S103: Process the at least two fragment subtasks in parallel by at least two processing units; wherein the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and different groups of pixel sub-regions have the same number of pixel sub-regions corresponding to different processing units.
[0042] In some embodiments, the processing unit belongs to the back-end module of the TBR architecture, is located after the task distribution unit, and is mainly used to perform shading processing on pixel blocks.
[0043] In some embodiments, the fragment task processing method is applied to a graphics processor. The graphics processor includes multiple processing units to be selected, and at least two processing units are selected and determined by the multiple processing units to be selected. The at least two processing units can be all or some of the processing units to be selected.
[0044] For example, if the number of processing units to be selected is 16 and there are 4 tiles to be processed, the pixel block processing task for one tile can be divided among two processing units. In this case, 8 processing units are required to process the 4 tiles, and 8 processing units can be selected from the 16 processing units to complete the rendering of the tile. If the number of processing units to be selected is 16 and there are 8 tiles to be processed, the pixel block processing task for one tile can be divided among two processing units. In this case, 16 processing units are required to process the 8 tiles, that is, all the processing units to be selected are selected to complete the rendering of the tile.
[0045] In some embodiments, at least two pixel blocks are located in the same group of pixel sub-regions, that is, a group of pixel sub-regions includes at least two pixel blocks. All pixel sub-regions have the same size, that is, all pixel sub-regions have the same area. Since the number of pixel sub-regions in different groups is the same, the area of each group of pixel sub-regions is equal, which further indicates that the areas of the pixel sub-regions corresponding to at least two fragment subtasks are the same, and thus the number of primitives falling into the pixel sub-regions is substantially the same, resulting in the task distribution unit dividing the fragment task, resulting in the at least two fragment subtasks having substantially the same load.
[0046] In some embodiments, different groups of pixel sub-regions correspond to different processing units, that is, a group of pixel sub-regions corresponds to one processing unit, and one processing unit only processes one group of pixel sub-regions.
[0047] Exemplarily, the different groups of pixel sub-regions include two groups of pixel sub-regions, ie, a first pixel sub-region and a second pixel sub-region, the first pixel sub-region corresponds to the first processing unit, and the second pixel sub-region corresponds to the second processing unit.
[0048] In some embodiments, the number of processing units is consistent with the number of fragment subtasks, one processing unit processes one fragment subtask, and at least two processing units process at least two fragment subtasks simultaneously.
[0049] It should be noted that there is a mapping relationship between processing units and fragment subtasks, that is, a processing unit is bound to a fragment subtask and only processes its corresponding fragment subtask. Alternatively, a processing unit and a fragment subtask are not bound, that is, a processing unit processes any fragment subtask.
[0050] Exemplarily, the at least two processing units include two processing units (i.e., a first processing unit and a second processing unit), and the at least two fragment subtasks include two fragment subtasks (i.e., fragment subtask 1 and fragment subtask 2). When a mapping relationship exists between the processing units and the fragment subtasks, the first processing unit processes fragment subtask 1, and the second processing unit processes fragment subtask 2. When the processing units and fragment subtasks are not bound together, the first processing unit processes either fragment subtask 1 or fragment subtask 2, and the second processing unit processes all other fragment subtasks except those processed by the first processing unit. That is, the first processing unit processes fragment subtask 1, and the second processing unit processes fragment subtask 2; alternatively, the first processing unit processes fragment subtask 2, and the second processing unit processes fragment subtask 1.
[0051] In other embodiments, the number of processing units is inconsistent with the number of fragment subtasks. One processing unit may need to process multiple fragment subtasks, or one processing unit may process one fragment subtask.
[0052] For example, the at least two processing units include two processing units, and the at least two fragment subtasks include five fragment subtasks. Processing five fragment subtasks in parallel by two processing units may involve one processing unit processing two fragment subtasks and the other processing unit processing three fragment subtasks, or one processing unit processing one fragment subtask and the other processing unit processing four fragment subtasks.
[0053] In an embodiment of the present application, a fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel subregions, and all pixel subregions have the same size. Thus, the area of each group of pixel subregions is the same, further indicating that the areas of the pixel subregions corresponding to the at least two fragment subtasks are the same, resulting in the task distribution unit dividing the fragment task, resulting in the at least two fragment subtasks having substantially the same load. By having at least two processing units process the at least two fragment subtasks in parallel, the load of one processing unit can be reduced compared to the existing method of processing the fragment task by one processing unit, thereby achieving load balancing between different processing units.
[0054] Figure 3 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 2 , the method can be executed by a processor of a computer device. Figure 2 , Figure 2 Step S102 in the above example can be updated to step S201 to step S203, which will be combined with Figure 3 The steps shown are explained.
[0055] Step S201 : determining the group of pixel sub-regions corresponding to each pixel block according to the positional relationship between the pixel block in the image block and each group of pixel sub-regions in at least two groups of pixel sub-regions.
[0056] In some embodiments, the at least two groups of pixel sub-regions are obtained by dividing the image block based on the size of the pixel sub-regions. The pixel sub-region groups are obtained by classifying the pixel sub-regions into at least two categories after dividing the image block, and the number of pixel sub-region groups is the same as the number of the at least two categories of pixel sub-regions.
[0057] In some embodiments, the task distribution unit determines the group of pixel sub-regions corresponding to each pixel block according to a positional relationship between the pixel block and each of the at least two groups of pixel sub-regions in the image block.
[0058] In some embodiments, the task distribution unit determines which group of at least two groups of pixel sub-regions the pixel block is within based on the range of the pixel block in the image tile, thereby determining the group of pixel sub-regions corresponding to each pixel block.
[0059] Exemplarily, the at least two groups of pixel sub-regions include two groups of pixel sub-regions, namely, a first pixel sub-region and a second pixel sub-region, and the image block includes multiple pixel blocks, namely, pixel block 1, pixel block 2, pixel block 3, and pixel block 4. If the range of pixel block 1 and pixel block 2 is within the range of the first pixel sub-region, then the group of pixel sub-regions corresponding to pixel block 1 and pixel block 2 is the first pixel sub-region; if the range of pixel block 1 and pixel block 3 is within the range of the second pixel sub-region, then the group of pixel sub-regions corresponding to pixel block 2 and pixel block 4 is the second pixel sub-region.
[0060] Step S202: Divide the fragment task into pixel blocks with pixel blocks as the granularity, and obtain pixel block tasks corresponding to each pixel block.
[0061] In an embodiment of the present application, the fragment task includes a unit task corresponding to each of the at least two pixel groups, and the unit task includes triangle information, original pixel block information and control information of the pixel group.
[0062] In some embodiments, the segmentation granularity refers to the cutting size when segmenting the fragment task; the pixel block task is the pixel point processing task of the pixel block after the fragment task is segmented.
[0063] In some embodiments, the task distribution unit divides the fragment task based on the pixel block as the segmentation granularity to obtain the pixel block task corresponding to each pixel block.
[0064] Exemplarily, the task distribution unit divides the image block into pixel blocks as the dividing granularity to obtain multiple pixel blocks, namely pixel block 1, pixel block 2, pixel block 3 and pixel block 4, so that the fragment task corresponding to the image block is divided into pixel block tasks corresponding to pixel block 1, pixel block 2, pixel block 3 and pixel block 4 respectively.
[0065] It should be noted that if Figure 4 As shown, a tile is 32x32 pixels, so a tile can be divided into several pixel blocks (i.e. spans), and a span is 4x4 pixels.
[0066] Figure 5 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 3 , the method can be executed by a processor of a computer device. Figure 3 , Figure 3 Step S202 in the above example can be updated to step S301 to step S303, which will be combined with Figure 5 The steps shown are explained.
[0067] Step S301 : based on the relationship between the pixel block and the original pixel block, segment the original pixel block information of the pixel group to obtain the pixel block information of the pixel group.
[0068] In some embodiments, the original pixel block refers to a pixel block obtained by dividing a picture block into a preset size.
[0069] It should be noted that the preset size may be a group of pixels (ie, a pixel group), for example, a group of pixels may be 32 pixels.
[0070] In some embodiments, the task distribution unit may segment the original pixel block information of each pixel group based on the relationship between the pixel block and the original pixel block to obtain pixel block information of each pixel group.
[0071] Step S302: construct a pixel block subtask of the pixel group using the pixel block information, triangle information and control information of the pixel group.
[0072] In some embodiments, a pixel block subtask refers to a pixel processing task assigned to a pixel block within a pixel group.
[0073] In some embodiments, for each pixel group, the task processing unit may construct a pixel block subtask of the pixel group using the pixel block information, triangle information, and control information of the pixel group.
[0074] Step S303: construct a pixel block task corresponding to the pixel block based on the pixel block subtasks of at least two pixel groups.
[0075] In some embodiments, the task processing unit may construct a pixel block task corresponding to a pixel block based on pixel block subtasks of at least two pixel groups.
[0076] For example, the triangle information and control information of a pixel group will be divided into two parts, one for each processing unit, and the span information will be pre-allocated to the two processing units based on the redirection result.
[0077] In an embodiment of the present application, the original pixel block information of the pixel group is segmented based on the relationship between the pixel block and the original pixel block to obtain the pixel block information of the pixel group; the pixel block information, triangle information and control information of the pixel group are used to construct the pixel block subtask of the pixel group; based on the pixel block subtasks of at least two pixel groups, the pixel block task corresponding to the pixel block is constructed. In this way, the segmentation of the original pixel block information of the pixel group and the construction of the pixel block subtask of the pixel group can be achieved, thereby completing the construction of the pixel block task corresponding to the pixel block, which is convenient for the subsequent determination of the fragment subtask and the realization of load balancing between different processing units.
[0078] Step S203 : Based on the groups of pixel sub-regions corresponding to each pixel block and the pixel block tasks corresponding to each pixel block, determine the fragment sub-task corresponding to each group.
[0079] In some embodiments, the task processing unit may classify the pixel block tasks corresponding to each pixel block according to the group of pixel sub-regions corresponding to each pixel block, and treat the pixel block tasks of pixel blocks belonging to the same group as a type of pixel block tasks, thereby determining the fragment sub-task corresponding to each group.
[0080] Exemplarily, the groups of pixel sub-regions include two groups, a first pixel sub-region and a second pixel sub-region. The task distribution unit determines the pixel block task corresponding to the pixel block belonging to the first pixel sub-region as the first fragment sub-task, and determines the pixel block task corresponding to the pixel block belonging to the second pixel sub-region as the second fragment sub-task based on the group of pixel sub-regions corresponding to each pixel block (that is, whether each pixel block corresponds to the first pixel sub-region or the second pixel sub-region) and the pixel block task corresponding to each pixel block.
[0081] In an embodiment of the present application, the group of pixel sub-regions corresponding to each pixel block is determined based on the positional relationship between the pixel block in the image block and each group of pixel sub-regions in at least two groups of pixel sub-regions; the fragment task is divided with the pixel block as the dividing granularity to obtain the pixel block task corresponding to each pixel block; based on the group of pixel sub-regions corresponding to each pixel block and the pixel block task corresponding to each pixel block, the fragment subtask corresponding to each group is determined; in this way, the fragment task of the image block is divided to obtain the fragment subtask corresponding to each group, so that the fragment subtasks of each group can be subsequently distributed to different processing units to achieve load balancing between different processing units.
[0082] Figure 6 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 4 , the method can be executed by a processor of a computer device. Figure 2 The fragment task processing method also includes steps S401 to S402, combining Figure 6 The steps shown are explained.
[0083] Step S401: Obtain first position information of each pixel block in the image block and second position information of each pixel sub-region.
[0084] In some embodiments, the first position information is the range of each pixel block in the image block; the second position information is the range of each pixel sub-region in the image block.
[0085] In some embodiments, the task distribution unit divides the image block into multiple pixel blocks according to a preset size, and can obtain the range of each pixel block in the image block. In addition, the task distribution unit divides the image block into multiple pixel sub-regions evenly, and can obtain the range of each pixel sub-region in the image block. In some embodiments, the task distribution unit may determine the coordinates of the pixels in each pixel block in the tile according to the range of each pixel block in the tile, and determine the coordinates of the pixels in each pixel sub-region according to the range of each pixel sub-region.
[0086] Step S402: Determine the pixel sub-region corresponding to the second position information where the first position information falls as the pixel sub-region corresponding to each pixel block.
[0087] In some embodiments, the task distribution unit can determine the pixel sub-region corresponding to each pixel block based on the relationship between the first position information of each pixel block and the second position information of each pixel sub-region, that is, the pixel sub-region corresponding to the second position information where the first position information falls is determined as the pixel sub-region corresponding to each pixel block.
[0088] Exemplarily, the image block includes multiple pixel blocks, namely pixel block 1, pixel block 2, pixel block 3 and pixel block 4, and the first position information of pixel block 1, pixel block 2, pixel block 3 and pixel block 4 respectively falls within the second position information of the first pixel sub-region, the second position information of the second pixel sub-region, the second position information of the first pixel sub-region and the second position information of the second pixel sub-region, that is, pixel block 1 and pixel block 3 belong to the first pixel sub-region, and pixel block 2 and pixel block 4 belong to the second pixel sub-region.
[0089] In an embodiment of the present application, the first position information of each pixel block in the image block and the second position information of each pixel sub-region are obtained; the pixel sub-region where the first position information falls corresponding to the second position information is determined as the pixel sub-region corresponding to each pixel block; the pixel sub-region corresponding to each pixel block is determined to facilitate the division of fragment tasks and task distribution.
[0090] Figure 7 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 5 , the method can be executed by a processor of a computer device. Figure 2 The fragment task processing method further includes steps S501 to S503, combining Figure 7 The steps shown are explained.
[0091] Step S501: Determine at least two processing units from at least two to-be-selected processing units.
[0092] In some embodiments, the task distribution unit may determine at least two processing units from the at least two to-be-selected processing units based on a mapping relationship between the at least two fragment subtasks and the processing units.
[0093] In other embodiments, the task distribution unit may select processing units with smaller loads and substantially the same loads from the at least two candidate processing units based on the respective loads of the at least two candidate processing units, and determine them as the at least two processing units.
[0094] In other embodiments, the task distribution unit may randomly determine at least two processing units from at least two to-be-selected processing units.
[0095] In other embodiments, if the number of at least two processing units is consistent with the number of at least two processing units to be selected, that is, at least two processing units are all the processing units to be selected, then the number of fragment subtasks is determined based on the number of all processing units to be selected, that is, the number of fragment subtasks is equal to the number of all processing units to be selected.
[0096] Step S502: Distribute the obtained at least two fragment subtasks to the at least two processing units; wherein different processing units correspond to different groups of fragment subtasks.
[0097] In some embodiments, different processing units correspond to different groups of fragment subtasks, that is, one processing unit processes one group of fragment subtasks.
[0098] In some embodiments, the task distribution unit distributes the obtained at least two fragment subtasks to at least two processing units.
[0099] In an embodiment of the present application, at least two processing units are determined from at least two processing units to be selected; the obtained at least two fragment subtasks are distributed to the at least two processing units; and by having different processing units correspond to different groups of fragment subtasks, the fragment task processed by one processing unit can be divided into at least two processing units for processing, which can reduce the load of the processing unit, thereby achieving load balancing between different processing units.
[0100] Figure 8 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 6 , the method can be executed by a processor of a computer device. Figure 2 The fragment task processing method further includes steps S601 to S603, combining Figure 8 The steps shown are explained.
[0101] Step S601: Obtain historical load information of each processing unit.
[0102] In the embodiment of the present application, the historical load information is the number of pixels to be processed in each pixel block and the number of graphic elements corresponding to the pixels.
[0103] In some embodiments, the task processing unit obtains historical load information of each processing unit.
[0104] In the embodiment of the present application, historical load information of each processing unit is obtained to facilitate subsequent determination of the size of the pixel sub-region based on the historical load information.
[0105] Step S602: Determine the size of the pixel sub-region based on the difference between the historical load information of the processing units; wherein the size of the pixel sub-region is negatively correlated with the difference.
[0106] In some embodiments, the size of the pixel sub-region is negatively correlated with the difference, that is, the greater the difference between the historical load information of the processing units, the smaller the size of the pixel sub-region.
[0107] In some embodiments, the task processing unit may determine the size of the pixel sub-region based on the difference between historical load information of the processing units.
[0108] In some embodiments, the greater the difference between the historical load information of the processing units, the smaller the size of the pixel sub-region determined by the task processing unit; the smaller the difference between the historical load information of the processing units, the larger the size of the pixel sub-region determined by the task processing unit.
[0109] It should be noted that the size of the pixel sub-region is a multiple of the size of the pixel block.
[0110] Step S603: Divide the image block into regions according to the sizes of the pixel sub-regions to obtain the at least two groups of pixel sub-regions.
[0111] In some embodiments, the task distribution unit may divide the image block into regions according to the size of the pixel sub-region to obtain a plurality of pixel sub-regions; and determine at least two groups of pixel sub-regions based on the plurality of pixel sub-regions.
[0112] In other embodiments, based on a software identifier corresponding to a fragment task, a preset size corresponding to the software identifier is determined; the software is software that issues the fragment task. The image block is divided into regions using the preset size to obtain at least two groups of pixel sub-regions.
[0113] It should be noted that the software identifier can represent the load required for processing the fragment task, so that the preset size can be determined based on the software identifier.
[0114] For example, in the game screen, the position where the character is located contains more rendering targets.
[0115] In this embodiment of the present application, historical load information for each processing unit is obtained; the size of a pixel subregion is determined based on the difference between the historical load information of the processing units; and the image block is divided into regions based on the size of the pixel subregions to obtain at least two groups of pixel subregions. Because the size of the pixel subregions is negatively correlated with the difference, that is, the greater the difference in the historical load information of the processing units, the smaller the size of the pixel subregions, the size of the pixel subregions can be dynamically determined, making the size of the pixel subregions more accurate.
[0116] Figure 9 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 7 , the method can be executed by a processor of a computer device. Figure 7 , Figure 7 Before step S502 in step S701 to step S704 are also executed, and the Figure 9 The steps shown are explained.
[0117] Step S701: Receive triangle information, control information, and pixel block information of each group of pixel sub-regions of the image block through the task distribution unit.
[0118] In some embodiments, the triangle information, control information, and pixel block information of each group of pixel sub-regions of the tile are received by the task distribution unit.
[0119] It should be noted that the triangle information, control information and pixel block information are information output by the rasterization unit.
[0120] Step S702: storing the triangle information and control information of each group of pixel sub-regions in a cache queue of the task distribution unit.
[0121] In the embodiment of the present application, the number of the cache queues is consistent with the number of processing units; each of the cache queues includes: a first cache queue and a second cache queue; wherein, The first cache queue is used to store triangle information and control information of each group of pixel sub-regions; The second cache queue is used to store pixel block information of each group of pixel sub-regions.
[0122] In some embodiments, after receiving the triangle information, control information and pixel block information of each group of pixel sub-regions of the tile, the task dispatching unit stores the triangle information and control information of each group of pixel sub-regions in a first cache queue of the task dispatching unit.
[0123] Step S703: Determine load information of the pixel block information based on the pixel block information; wherein the load information is the number of triangles corresponding to the pixel block information.
[0124] In some embodiments, the load information is the number of triangles corresponding to the pixel block information.
[0125] In some embodiments, the task distribution unit may determine the number of triangles corresponding to the pixel block information, that is, the load information of the pixel block information, based on the pixel block information.
[0126] Step S704: Determine a processing method for the pixel block information, the triangle information, and the control information based on the load information.
[0127] In some embodiments, if the load information indicates that the number of triangles in the pixel block information is zero, the task distribution unit discards the triangle information and the control information.
[0128] In some embodiments, if the load information indicates that the number of triangles in the pixel block information is not zero, the task dispatching unit stores the pixel block information in a second cache queue of the task dispatching unit. The task dispatching unit sends the triangle information, control information, and pixel block information to the processing unit.
[0129] In an embodiment of the present application, the triangle information and control information of each group of pixel sub-regions are first stored in the cache queue of the task distribution unit, and then the processing method of the pixel block information, the triangle information and the control information is determined based on the load information of the pixel block information. This can avoid sending invalid triangle information, pixel block information and control information to the processing unit for processing, thereby improving the efficiency of the processing unit.
[0130] In the embodiment of the present application, the above-mentioned processing method of determining the pixel block information, the triangle information and the control information based on the load information can be implemented through steps S7041 to S7042.
[0131] Step S7041: If the load information represents the number of triangles in the pixel block information is zero, discard the triangle information and the control information.
[0132] In some embodiments, if the load information represents that the number of triangles in the pixel block information is zero, indicating that the pixel block has no pixels that need to be shaded, the triangle information and the control information are discarded.
[0133] Step S7042: If the load information represents the number of triangles in the pixel block information is not zero, the triangle information, the control information, and the pixel block information are sent to a processing unit.
[0134] In some embodiments, if the load information indicates that the number of triangles in the pixel block information is not zero, indicating that the pixel block has pixels that need to be shaded, the triangle information, control information, and pixel block information are sent to the processing unit.
[0135] In an embodiment of the present application, if the number of triangles of the pixel block information represented by the load information is zero, the triangle information and the control information are discarded; if the number of triangles of the pixel block information represented by the load information is not zero, the triangle information, the control information and the pixel block information are sent to the processing unit. In this way, when the number of triangles of the pixel block information is zero, it means that the pixel block is invalid information, and the triangle information and the control information are discarded. Only when the number of triangles of the pixel block information is not zero (that is, the pixel block is valid information), the triangle information, the control information and the pixel block information are sent to the processing unit, which can improve the processing efficiency of the processing unit.
[0136] Figure 10 This is a schematic diagram of the implementation process of a fragment task processing method provided in an embodiment of the present application. Figure 8 , the method can be executed by a processor of a computer device. Figure 2 The fragment task processing method also includes steps S801 to S803, combining Figure 10 The steps shown are explained.
[0137] Step S801: Receive color information output by each processing unit through the color buffer unit; wherein the color information is coloring information of pixel points.
[0138] In some embodiments, the color information is coloring information of the pixel points.
[0139] In some embodiments, the color information output by each processing unit is received via a color buffer unit.
[0140] Step S802: For each pixel point processed by the processing unit, based on the position information of each pixel point in the image block, the position information of the pixel point is updated by the remapping unit to obtain the target position information of the pixel point; wherein, the position information between the pixel blocks processed by each processing unit is discontinuous; and the target position information between the pixel blocks processed by each processing unit is continuous.
[0141] In some embodiments, the position information between the pixel blocks processed by each processing unit is discontinuous; the target position information between the pixel blocks processed by each processing unit is continuous.
[0142] In some embodiments, for each pixel point processed by each processing unit, the position information of the pixel point is updated by the remapping unit based on the position information of each pixel point in the image block to obtain the target position information of the pixel point.
[0143] For example, Figure 11A As shown, block 100 is divided into two types of pixel sub-regions, namely pixel sub-region 210 and pixel sub-region 220; wherein, the processing unit corresponding to pixel sub-region 210 is the second processing unit; and the processing unit corresponding to pixel sub-region 220 is the first processing unit. The position information of the pixel point is updated by the remapping unit to obtain the target position information of the pixel point. Specifically, the pixel sub-region (SR) to which it belongs is calculated based on the coordinates of the pixel point output by a processing unit, and then the SR is moved to the hole area covered by the SR of another processing unit to ensure the close arrangement of the SR area, and then the storage space address is calculated based on the coordinates after the move. For example, Figure 11B As shown, the pixel sub-region 220 corresponding to the first processing unit is Figure 11B The middle arrow points to move, such as Figure 11C As shown, the pixel sub-region 210 corresponding to the second processing unit is Figure 11CThe middle arrow points to move.
[0144] Address remapping can be achieved through the following code, as follows: / / SR_SIZE_X: span region width / / SR_SIZE_Y: span region height int remap_address(uint8_t fu_id, uint32_t&pixel_x, uint32_t&pixel_y) { uint32_t sr_x = pixel_x / SR_SIZE_X; uint32_t sr_y = pixel_y / SR_SIZE_Y; uint32_t sr_offset_y = (sr_y+1)>>1; pixel_x = quad_x; pixel_y = pixel_y -sr_offset_y * SR_SIZE_Y; int addr = calc_addr(pixel_x, pixel_y); return addr; } Step S803: Determine the storage space address of the pixel point based on the target position information.
[0145] In some embodiments, the storage space address of the pixel point is determined based on the target position information.
[0146] In an embodiment of the present application, the color information output by each processing unit is received through a color buffer unit, and for the pixel points processed by each processing unit, the position information of the pixel points is updated through a remapping unit based on the position information of each pixel point in the block to obtain the target position information of the pixel points; since the position information between the pixel blocks processed by each processing unit is discontinuous, the target position information between the pixel blocks processed by each processing unit is continuous, and the storage space address of the pixel point is determined based on the continuous target position information; in this way, the storage space address is also continuous, which can save storage space.
[0147] In the embodiment of the present application, the fragment task processing method further includes step S901, as follows: Step S901: Using the at least two processing units, cache the respective color information in parallel according to the storage space addresses of the pixels.
[0148] In some embodiments, at least two processing units are used to cache respective color information in parallel according to the storage space addresses of the pixels.
[0149] In the embodiment of the present application, at least two processing units are used to cache respective color information in parallel according to the storage space addresses of the pixel points. The parallel caching can improve the processing speed of the image processor.
[0150] In the embodiment of the present application, the fragment task processing method further includes step S1001, as follows: S1001: Read the color information of each of the at least two processing units from the cache in parallel through the rendering output unit, and display the color information.
[0151] In some embodiments, the color information of at least two processing units is read from the cache in parallel by the rendering output unit, and the color information is displayed.
[0152] In the embodiment of the present application, the rendering output unit reads the color information of at least two processing units in parallel from the cache and displays the color information. The parallel reading can improve the processing speed of the image processor.
[0153] The following describes the application of the fragment task processing method provided in the embodiment of the present application in actual scenarios, which mainly involves distributing the fragment task of a block to at least two processing units for processing. The following embodiments are only for the purpose of more clearly describing the implementation process of the present application.
[0154] In a traditional TBR architecture, the front-end GPU generates primitive rendering data and writes it to memory. The back-end then breaks the data into tiles based on the screen size and distributes them to different fragment units (FUs). Once a tile is sent to a FU, all pixels covered by the primitives within that tile are shaded within that FU.
[0155] See also Figure 12 , Figure 12This is a schematic diagram of the fragment task processing process in the related art. First, the screen is split into tiles, and the entire screen is split into tile11, tile12, tile13 and tile14. Then the loads of tile11, tile12, tile13 and tile14 are sent to FU15, FU16, FU17 and FU18 respectively. There are 5 pixels in tile11, 2 pixels in tile12, and 1 pixel in tile13 and tile14 respectively. Figure 12 As can be seen from the figure, if most of the pixels of a frame being rendered are concentrated in a certain tile (for example, there are 5 pixels in tile 11), the fragment load of the FU bound to the tile will also be increased, which will eventually lead to an imbalance in fragment tasks among the FUs, that is, some FUs work for a long time, while some FUs work for a short time, causing the overall GPU performance to deteriorate.
[0156] See also Figure 13A The actual rendering scene shown in the figure. There are many scenes in the benchmark where performance degradation occurs due to FU imbalance, such as Figure 13A The following figure shows a pass for rendering a head in Baldur's Gate 3. Since the scene draws hundreds of thousands of triangles, and these triangles are concentrated in the same tile, all the load is sent to the same FU, and other FUs are idle, which single-handedly prolongs the execution time of the entire frame.
[0157] See also Figure 13B The following figure shows the execution time of each processing unit in the actual rendering scene. Among them, the execution time of FU21 far exceeds the execution time of FU22, FU23, FU24, FU25, FU26 and FU27. Figure 13B As shown in the figure, FU21 bears all the loads of a header rendering Pass, and FU22, FU23, FU24, FU25, FU26 and FU27 are idle, which causes the execution time of the entire frame to be prolonged.
[0158] Based on this, this application optimizes the pipeline of the TBR architecture and adds a task dispatcher (Fragment Task Dispatcher, FTD) after the rasterization stage. The main function of FTD is to split and distribute fragment tasks. Figure 14Tile allocation 30 inputs the output of rasterization unit 31. Task dispatch unit 32 distributes the output of rasterization unit 31 to processing unit 33 (FU0) and processing unit 34 (FU1). Color buffer 35 (i.e., Color Buffer0) remaps the output data of FU0 and caches the remapped data. Color buffer 36 (i.e., Color Buffer1) remaps the output data of FU1 and caches the remapped data. Render output unit 37 (ROP) reads the remapped data from the cache for processing.
[0159] See also Figure 15 A tile 40 is divided into multiple pixel blocks, namely pixel block 401, pixel block 402, pixel block 403, pixel block 404, pixel block 405, and pixel block 406. Task dispatch unit 42 classifies the multiple pixel blocks obtained from rasterization unit 41 into two categories of pixel blocks. (One category consists of pixel blocks 401, pixel block 403, and pixel block 405, and the other category consists of pixel blocks 402, pixel block 404, and pixel block 406). Tasks are divided for each category of pixel blocks, resulting in two subtasks for each category of pixel blocks: subtasks 407 and 408 corresponding to pixel blocks 401, pixel block 403, and pixel block 405, and subtasks 409 and 410 corresponding to pixel blocks 402, pixel block 404, and pixel block 406. The processing unit 43 directly processes subtasks 407 , 408 , 409 , and 410 , and performs color buffering 45 and color buffering 46 through the rendering output unit 44 .
[0160] The Fragment Task Dispatcher (FTD) is a key component of this optimization design implementation. Figure 16 The task processing unit 51 contains three functional modules: a redirection unit 510 (Span Redirect), a judgment unit 511 (Pass Filter), and a data cache unit 512 (Data FIFO). The data cache unit 512 includes a first cache queue 5121, a second cache queue 5122, a first cache queue 5123, and a second cache queue 5124. The first cache queue 5121 and the second cache queue 5122 are used to cache data that the processing unit 52 needs to process, while the first cache queue 5123 and the second cache queue 5124 are used to cache data that the processing unit 53 needs to process.
[0161] The main functions of the task distribution unit 51 include: (1) Receive triangle information, span information and control information sent by the rasterization unit 50.
[0162] It should be noted that the control information is used downstream for downstream task assembly, and may contain some attribute information, for example, a span is a pixel block, such as 4x4 pixels.
[0163] (2) For span redirection, redirection means marking the span with the FU index according to the span region coordinates, indicating which FU the span is to be sent to.
[0164] (3) Send triangle information, span information, etc. to the FIFO of the corresponding FU.
[0165] The redirection unit 510 is the main functional module of FTD. A frame of image contains several passes. The data of a pass includes triangle information, span information and control information. The span information is generated based on the tile, for example Figure 4 A tile is 32x32 pixels, so a tile can be divided into several spans.
[0166] Redirection is to calculate which FU the span corresponding to the same area on the screen should be sent to based on the span's coordinate offset within the tile. The mapping relationship is determined by the span region pattern. Figure 17 As shown in FIG. 6 , a typical spanregion pattern is shown. Block 60 is divided into pixel subregion 61 and pixel subregion 62. The spans falling within pixel subregion 61 are sent to one processing unit, and the spans falling within pixel subregion 62 are sent to another processing unit. The user can control the size of the span region to adjust the balance between the two processing units. For example, when the load weight variation rate is large, a small pixel subregion can be set to balance the load of the two processing units; when the load weight variation rate is small, the size of the pixel subregion can be increased accordingly.
[0167] The pseudo code for redirecting span by pixel sub-region is as follows sr_x = (span_x>>sr_size_x)&0x1; sr_y = (span_y>>sr_size_y)&0x1; fu_id = sr_x ^ sr_y The information contained in the pass after being redirected, such as Figure 18As shown, the triangle information 70 and control information 72 of a pass0 will be divided into two parts, one for each processing unit, and the original pixel block 71 (i.e., pixel block group 76) will be divided into pixel block 710, pixel block 711, pixel block 712, pixel block 713, pixel block 714, pixel block 715, pixel block 716 and pixel block 717; among them, pixel block 710, pixel block 712, pixel block 714, pixel block 716 are assigned to one processing unit for processing, and pixel block 711, pixel block 713, pixel block 715, pixel block 717 are assigned to another processing unit for processing. The triangle information 73 and control information 75 of a pass 1 will be divided into two parts, one for each processing unit, and the original pixel block 74 (i.e., pixel block group 77) will be divided into pixel block 740, pixel block 741, pixel block 742, pixel block 743, pixel block 744, pixel block 745, pixel block 746 and pixel block 747; among them, pixel block 740, pixel block 742, pixel block 744, pixel block 746 is assigned to one processing unit for processing, and pixel block 741, pixel block 743, pixel block 745, pixel block 747 is assigned to another processing unit for processing.
[0168] The judgment unit 511 acts as an arbitrator to determine whether the data of a pass will actually be sent to the processing unit. This part mainly takes into account that the pixel sub-region corresponding to a processing unit may not have any valid load, such as Figure 19 As shown, tile 80 is divided into pixel sub-region 81 and pixel sub-region 82, where pixel sub-region 81 corresponds to the first processing unit and pixel sub-region 82 corresponds to the second processing unit. Since pixel sub-region 82 of the second processing unit does not have any valid spans, the triangle information and control information of this pass should be discarded to prevent invalid information from being sent to the second processing unit.
[0169] The data cache unit 512 is used to cache the data sent by the task dispatch unit to the FU, in order to avoid the pipeline stagnation caused by the blocking of a certain FU. The data cache unit 512 includes two types: one type stores the request information of the attribute; the other type stores the fragment task. Figure 16 The first cache queue 5121 and the second cache queue 5122 in.
[0170] The task distribution unit splits a tile into two subtiles and sends them to two processing units for processing. To save storage space to the greatest extent possible, the pixel colors output by the two processing units should be remapped before storage. The purpose is to arrange the spatially discrete pixel points tightly in the storage space.
[0171] like Figure 20As shown, color remapping 92 is performed on the outputs of the processing unit 90 and the processing unit 91 to obtain two color buffer queues, such as a first color buffer queue 93 and a second color buffer queue 94 .
[0172] Table 1
[0173] Based on the foregoing embodiments, an embodiment of the present application provides a fragment task processing device, which includes the various units included and the various modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0174] Figure 21 A schematic diagram of the structure of a fragment task processing device provided in an embodiment of the present application is shown in FIG. Figure 21 As shown, the fragment task processing device 200 includes: an image processor 210; the image processor 210 includes: a task distribution unit 2110 and at least two processing units 2120, wherein: The task distribution unit 2110 is configured to receive a fragment task corresponding to a tile through the task distribution unit; split the fragment task into at least two fragment subtasks; The at least two processing units 2120 are used to process the at least two fragment subtasks in parallel; wherein, the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and the number of pixel sub-regions in different groups is the same and corresponds to different processing units.
[0175] In some embodiments, the task distribution unit 2110 is further used to determine the group of pixel sub-regions corresponding to each pixel block based on the positional relationship between the pixel block in the image block and each group of pixel sub-regions in at least two groups of pixel sub-regions; divide the fragment task with the pixel block as the dividing granularity to obtain the pixel block task corresponding to each pixel block; and determine the fragment sub-task corresponding to each group based on the group of pixel sub-regions corresponding to each pixel block and the pixel block task corresponding to each pixel block.
[0176] In some embodiments, the task distribution unit 2110 is also used to obtain the first position information of each pixel block in the image block and the second position information of each pixel sub-region; and determine the pixel sub-region corresponding to the second position information where the first position information falls as the pixel sub-region corresponding to each pixel block.
[0177] In some embodiments, the task distribution unit 2110 is further used to determine the at least two processing units from the at least two to-be-selected processing units; and distribute the obtained at least two fragment subtasks to the at least two processing units; wherein different processing units correspond to different groups of fragment subtasks.
[0178] In some embodiments, the task distribution unit 2110 is further used to obtain historical load information of each of the processing units; determine the size of the pixel sub-region based on the difference between the historical load information of the processing units; wherein the size of the pixel sub-region is negatively correlated with the difference; and divide the image block into regions according to the size of the pixel sub-region to obtain the at least two groups of pixel sub-regions.
[0179] In some embodiments, the task distribution unit 2110 is further used to split the original pixel block information of the pixel group based on the relationship between the pixel block and the original pixel block to obtain the pixel block information of the pixel group; use the pixel block information, triangle information and control information of the pixel group to construct the pixel block subtask of the pixel group; based on the pixel block subtasks of at least two pixel groups, construct the pixel block task corresponding to the pixel block.
[0180] In some embodiments, the task distribution unit 2110 is also used to receive triangle information, control information and pixel block information of each group of pixel sub-regions of the image tile through the task distribution unit; store the triangle information and control information of each group of pixel sub-regions in the cache queue of the task distribution unit; determine the load information of the pixel block information based on the pixel block information; wherein the load information is the number of triangles corresponding to the pixel block information; and determine the processing method of the pixel block information, the triangle information and the control information based on the load information.
[0181] In some embodiments, the task distribution unit 2110 is also used to discard the triangle information and the control information if the number of triangles representing the pixel block information in the load information is zero; and send the triangle information, the control information and the pixel block information to the processing unit if the number of triangles representing the pixel block information in the load information is not zero.
[0182] In some embodiments, the image processor 210 further includes: a color buffer unit 2130; wherein, The color buffer unit 2130 is used to receive color information output by each of the processing units; wherein, the color information is the shading information of the pixel points; for each of the pixel points processed by the processing unit, based on the position information of each of the pixel points in the image block, the position information of the pixel points is updated by the remapping unit to obtain the target position information of the pixel points; wherein, the position information between the pixel blocks processed by each of the processing units is discontinuous; the target position information between the pixel blocks processed by each of the processing units is continuous; based on the target position information, the storage space address of the pixel points is determined.
[0183] In some embodiments, the at least two processing units 2120 are further configured to cache the respective color information in parallel according to the storage space address of the pixel point through the at least two processing units.
[0184] In some embodiments, the image processor 210 further includes a rendering output unit 2140; wherein, The rendering output unit 2140 is configured to read the color information of each of the at least two processing units from the cache in parallel and display the color information.
[0185] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0186] It should be noted that in the embodiments of the present application, if the above-mentioned fragment task processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.
[0187] An embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0188] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0189] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0190] An embodiment of the present application provides a computer program product, comprising a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, the computer program implements some or all of the steps of the above-described method. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0191] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.
[0192] Figure 22 A hardware entity diagram of a computer device provided in an embodiment of the present application is shown as follows: Figure 22 As shown, the hardware entity of the computer device 2200 includes: a processor 2201 and a memory 2202, wherein the memory 2202 stores a computer program that can be run on the processor 2201, and the processor 2201 implements the steps in the method of any of the above embodiments when executing the program.
[0193] The memory 2202 stores computer programs that can be run on the processor. The memory 2202 is configured to store instructions and applications executable by the processor 2201. It can also cache data to be processed or processed by the processor 2201 and various modules in the computer device 2200 (for example, image data, audio data, voice communication data and video communication data). It can be implemented through flash memory (FLASH) or random access memory (RAM).
[0194] When the processor 2201 executes the program, the steps of any of the above-mentioned fragment task processing methods are implemented. The processor 2201 generally controls the overall operation of the computer device 2200.
[0195] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the fragment task processing method of any of the above embodiments.
[0196] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0197] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.
[0198] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface storage device, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0199] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0200] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0201] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0202] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0203] In addition, the functional units in the various embodiments of the present application can all be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units. It can be understood by those skilled in the art that all or part of the steps of the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memories (ROMs), magnetic disks, or optical disks.
[0204] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0205] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A fragment task processing method, characterized in that: Applied to a graphics processor, the method includes: Receive the fragment task corresponding to the tile through the task distribution unit; Splitting the fragment task to obtain at least two fragment subtasks; Processing the at least two fragment subtasks in parallel by at least two processing units; Among them, the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and the number of pixel sub-regions in different groups is the same and corresponds to different processing units.
2. The method according to claim 1, characterized in that The fragment task is divided into at least two fragment subtasks, including: determining the group of pixel sub-regions corresponding to each pixel block according to a positional relationship between the pixel block and each of the at least two groups of pixel sub-regions in the image block; Using pixel blocks as the segmentation granularity, the fragment task is segmented to obtain pixel block tasks corresponding to each pixel block; Based on the groups of pixel sub-regions corresponding to each pixel block and the pixel block tasks corresponding to each pixel block, a fragment sub-task corresponding to each group is determined.
3. The method according to claim 1 or 2, characterized in that The method further comprises: Acquire first position information of each pixel block in the image block and second position information of each pixel sub-region; The pixel sub-region corresponding to the second position information and where the first position information falls is determined as the pixel sub-region corresponding to each pixel block.
4. The method according to claim 2, characterized in that The graphics processor includes at least two processing units to be selected; and the method further includes: Determining the at least two processing units from among the at least two processing units to be selected; The at least two obtained fragment subtasks are distributed to the at least two processing units, wherein different processing units correspond to different groups of fragment subtasks.
5. The method according to claim 1, wherein The method further comprises: Obtaining historical load information of each of the processing units; determining a size of a pixel sub-region based on a difference between historical load information of the processing units; wherein the size of the pixel sub-region is negatively correlated with the difference; The image block is divided into regions according to the size of the pixel sub-region to obtain the at least two groups of pixel sub-regions.
6. The method according to claim 5, characterized in that The historical load information is the number of pixels to be processed in each pixel block and the number of graphic elements corresponding to the pixels.
7. The method according to claim 2, characterized in that The fragment task includes a unit task corresponding to each of the at least two pixel groups, and the unit task includes triangle information, original pixel block information and control information of the pixel group; The fragment task is divided into pixel blocks as the segmentation granularity to obtain pixel block tasks corresponding to each pixel block, including: Based on the relationship between the pixel block and the original pixel block, segmenting the original pixel block information of the pixel group to obtain pixel block information of the pixel group; constructing a pixel block subtask of the pixel group using the pixel block information, triangle information and control information of the pixel group; Based on the pixel block subtasks of at least two pixel groups, a pixel block task corresponding to the pixel block is constructed.
8. The method according to any one of claims 1 to 7, characterized in that Before distributing the obtained at least two fragment subtasks to the at least two processing units, the method further includes: receiving triangle information, control information, and pixel block information of each group of pixel sub-regions of the image block through the task distribution unit; storing the triangle information and control information of each group of pixel sub-regions in a cache queue of the task distribution unit; Determining load information of the pixel block information based on the pixel block information; wherein the load information is the number of triangles corresponding to the pixel block information; Based on the load information, a processing manner of the pixel block information, the triangle information, and the control information is determined.
9. The method according to claim 8, characterized in that The determining, based on the load information, a processing method of the pixel block information, the triangle information, and the control information includes: If the number of triangles representing pixel block information represented by the load information is zero, discarding the triangle information and the control information; If the load information represents that the number of triangles of the pixel block information is not zero, the triangle information, the control information and the pixel block information are sent to a processing unit.
10. The method according to claim 8, characterized in that The number of the cache queues is consistent with the number of processing units; each of the cache queues includes: a first cache queue and a second cache queue; wherein, The first cache queue is used to store triangle information and control information of each group of pixel sub-regions; The second cache queue is used to store pixel block information of each group of pixel sub-regions.
11. The method according to claims 1-7, characterized in that The graphics processor includes a color buffer unit, the color buffer unit includes a remapping unit, and the method further includes: Receiving color information output by each processing unit through the color buffer unit; wherein the color information is coloring information of the pixel; For each pixel point processed by the processing unit, based on the position information of each pixel point in the image block, the remapping unit updates the position information of the pixel point to obtain the target position information of the pixel point; Wherein, the position information between the pixel blocks processed by each of the processing units is discontinuous; the target position information between the pixel blocks processed by each of the processing units is continuous; Based on the target position information, a storage space address of the pixel point is determined.
12. The method according to claim 11, characterized in that The method further comprises: The respective color information is cached in parallel according to the storage space address of the pixel point through the at least two processing units.
13. The method according to claim 12, characterized in that The graphics processor includes a rendering output unit, and the method further includes: The color information of each of the at least two processing units is read from the cache in parallel by the rendering output unit, and the color information is displayed.
14. A fragment task processing device, characterized in that: The fragment task processing device includes a graphics processor, which includes a task distribution unit and at least two processing units; wherein, The task distribution unit is configured to receive a fragment task corresponding to a tile through the task distribution unit; split the fragment task into at least two fragment subtasks; The at least two processing units are configured to process the at least two fragment subtasks in parallel; Among them, the fragment subtask is a pixel processing task for processing at least two pixel blocks in a tile; the at least two pixel blocks are located in the same group of pixel sub-regions, all pixel sub-regions have the same size, and the number of pixel sub-regions in different groups is the same and corresponds to different processing units.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 13 are implemented.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
17. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Graph processing method and device, computer equipment and storage medium
CN118397159A
Image processor, image rendering method and electronic equipment
CN120070693A
Graphics processor, load balancing method and electronic equipment
CN120500701A
Rendering task allocation method and device, electronic equipment and storage medium
CN120523610A
Z-clipping for primitive samples
US20250086882A1
Cited By
Graphics processing device and method, graphics processor and computer equipment
CN122089556A