Graphical processing method and apparatus
Patent Information
- Application Number
- CN202611170245.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-04
- Publication Date
- 2026-09-04
AI Technical Summary
[0016]In the graphics processing scheme proposed in this disclosure, the frame to be rendered is divided into multiple tiles. When the effective rendering area covered by the primitives of the frame to be rendered is relatively small, tile splitting is enabled for at least a portion of the tiles, and the tile regions obtained through splitting are allocated to multiple tile processing cores for processing. Therefore, when the primitive coverage area is relatively concentrated, tile regions can be used instead of individual tiles as the allocation unit, allowing tiles that were originally processed by a single tile processing core to be processed in parallel by two or more tile processing cores. This helps avoid the problem of one or a few cores being overloaded while other cores are idle, allowing for more full utilization of the parallel processing capabilities of multiple tile processing cores, improving load balancing, and thus helping to improve graphics processing efficiency and enhance the rendering performance of the graphics processing device.
Smart Images

Figure CN122694643A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of graphics processing technology, and more specifically, to a graphics processing method, a graphics processing device, and a computing device. Background Technology
[0002] Graphics processing technology refers to the techniques used in conjunction with hardware and software to process two-dimensional or three-dimensional visuals and images to achieve specific rendering effects. Graphics processing devices refer to devices used to perform graphics processing operations, such as graphics processing units (GPUs) or other devices with graphics processing capabilities. Generally, graphics processing devices can achieve higher processing efficiency through parallel processing cores. Therefore, how to fully utilize the characteristics and advantages of this parallel operation to further improve graphics processing efficiency has become a significant issue of interest in this field. Summary of the Invention
[0003] In view of this, this application provides a graphics processing method, a graphics processing device, and a computing device, which helps to make fuller use of the parallel operation performance of the graphics processing device and improve graphics processing efficiency.
[0004] According to one aspect of this disclosure, a graphics processing method is provided, comprising: dividing a frame to be rendered into a plurality of tiles; determining an effective rendering region of the frame to be rendered, wherein the frame to be rendered includes a plurality of primitives, and the effective rendering region includes the coverage area of the plurality of primitives; if the size of the effective rendering region is less than a threshold, enabling tile splitting for at least a portion of the plurality of tiles, wherein tile splitting includes splitting a tile into a specified number of tile regions; and allocating the plurality of tile regions obtained by tile splitting to a plurality of tile processing cores, wherein each tile processing core is configured to perform tile processing operations for the allocated tile regions.
[0005] In some embodiments, allocating multiple tile regions obtained by tile splitting to multiple tile processing cores includes: allocating multiple tiles to multiple tile processing cores, wherein each tile is allocated a number of times equal to a specified number, and sending a tile region identifier each time a tile is allocated, the tile region identifier indicating a tile region among the allocated tiles.
[0006] In some embodiments, the graphics processing method further includes: disabling tile splitting when the size of the effective rendering area is higher than a threshold; and allocating multiple tiles to multiple tile processing cores, wherein each tile processing core is configured to perform tile processing operations on the allocated tiles.
[0007] In some embodiments, a tile region identifier with a preconfigured value is sent when allocating each tile, wherein the preconfigured value indicates that tile splitting is not enabled.
[0008] In some embodiments, the effective rendering area is the smallest rectangular area that includes the coverage area of multiple primitives.
[0009] In some embodiments, the size of the effective rendering area is characterized by the number of tiles covered by the effective rendering area.
[0010] In some embodiments, the graphics processing method further includes: each tile processing core performing the following operations for each assigned tile region: obtaining a set of primitives corresponding to a tile that includes the tile region, wherein the set of primitives includes primitives that cover the tile; discarding primitives that do not cover the tile region in the set of primitives; and performing tile processing operations based on the remaining primitives in the set of primitives.
[0011] In some embodiments, discarding a primitive that does not cover the tile region in the primitive set includes: traversing the primitives in the primitive set, and, for each primitive, discarding the primitive if the primitive region corresponding to the primitive does not intersect with the tile region, wherein the primitive region corresponding to the primitive is the smallest rectangular region containing the primitive.
[0012] In some embodiments, the specified quantity is 2, 4, 8 or 16.
[0013] In some embodiments, tile processing operations include one or more of the following: rasterization, depth testing, pixel shading, and color writing.
[0014] According to another aspect of this disclosure, a graphics processing apparatus is provided, comprising: a tile partitioning module configured to: divide a frame to be rendered into a plurality of tiles, and determine an effective rendering region of the frame to be rendered, wherein the frame to be rendered includes a plurality of primitives, and the effective rendering region includes the coverage area of the plurality of primitives; an allocation module configured to: enable tile splitting for at least a portion of the plurality of tiles when the size of the effective rendering region is lower than a threshold, and allocate the plurality of tile regions obtained by tile splitting to a plurality of tile processing cores, wherein tile splitting includes splitting a tile into a specified number of tile regions; and a plurality of tile processing cores, wherein each tile processing core is configured to: perform tile processing operations on the allocated tile regions.
[0015] According to another aspect of this disclosure, a computing device is provided, including the graphics processing device described in the foregoing aspect.
[0016] In the graphics processing scheme proposed in this disclosure, the frame to be rendered is divided into multiple tiles. When the effective rendering area covered by the primitives of the frame to be rendered is relatively small, tile splitting is enabled for at least a portion of the tiles, and the tile regions obtained through splitting are allocated to multiple tile processing cores for processing. Therefore, when the primitive coverage area is relatively concentrated, tile regions can be used instead of individual tiles as the allocation unit, allowing tiles that were originally processed by a single tile processing core to be processed in parallel by two or more tile processing cores. This helps avoid the problem of one or a few cores being overloaded while other cores are idle, allowing for more full utilization of the parallel processing capabilities of multiple tile processing cores, improving load balancing, and thus helping to improve graphics processing efficiency and enhance the rendering performance of the graphics processing device.
[0017] These and other aspects of this application will be clear from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0018] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 This illustration shows an example of a tile assignment scenario; Figure 2 This illustration shows the application Figure 1 The illustrated tile allocation scheme is an example of allocation when the elements are concentrated in a single tile. Figure 3 Schematic illustration Figure 2 The load on each core is shown in the allocation example. Figure 4 An example flowchart illustrating a graphics processing method according to some embodiments of the present disclosure is shown schematically; Figure 5 The illustrations illustrate example splitting methods for dividing a tile into tile regions according to some embodiments of the present disclosure; Figure 6 This illustration schematically depicts an example process for determining an effective rendering region according to some embodiments of the present disclosure; Figure 7 This illustration schematically shows an example process for filtering elements in the core of the tile processing according to some embodiments of the present disclosure; Figure 8 This illustration schematically shows an example process of a tile processing core performing depth or color writing according to some embodiments of the present disclosure; Figure 9 The illustration schematically shows example processes of graphics processing according to some embodiments of the present disclosure; Figure 10An exemplary block diagram of a graphics processing device according to some embodiments of the present disclosure is shown schematically; Figure 11 An exemplary block diagram of a computing device according to some embodiments of the present disclosure is shown schematically. Detailed Implementation
[0019] In the following description, exemplary embodiments will be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided merely to enable those skilled in the art to clearly and fully understand this disclosure. In the drawings, the same reference numerals denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0020] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to enable those skilled in the art to fully understand the embodiments of this disclosure. However, those skilled in the art will understand that the technical solutions of this disclosure can be practiced without including one or more specific details, or other methods, components, apparatuses, steps, etc., can be used to practice the technical solutions of this disclosure. The use of the terms "one embodiment," "another embodiment," etc., refers to a specific feature, structure, material, or characteristic described in connection with that embodiment that is included in at least one embodiment of this disclosure. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this disclosure, as well as the features of different embodiments or examples.
[0021] The block diagrams shown in the accompanying drawings correspond only to functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0022] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content, operations, or steps, nor do the steps necessarily need to be executed in the described order. For example, some operations or steps can be broken down into sub-operations or sub-steps, while some operations or steps can be combined or partially combined, and some operations or steps can be executed in parallel or in reverse order. Therefore, the actual execution order of each operation or step may change depending on the actual situation.
[0023] It should be understood that although the terms first, second, third, etc., are used in this disclosure to describe some elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. Therefore, the first element mentioned below may also be referred to as the second element, which will not depart from the conception of this disclosure. In this disclosure, "a plurality of" should be understood as two or more. Furthermore, in this disclosure, the terms "and / or" and similar terms mean all combinations of any one, a plurality of, or all of the items listed therein in association.
[0024] Next, for clarity, some related concepts appearing or associated with the embodiments described in this disclosure will be briefly explained.
[0025] A frame to be rendered can refer to an image frame to be rendered, which may correspond to a part of the virtual scene to be rendered, or it may correspond to an image or video frame to be rendered.
[0026] A primitive can refer to the basic geometric shapes that make up a graphic, such as points, lines, and triangles. For example, a graphic drawn by an application or otherwise obtained can be represented in a computer by a combination of a large number of basic geometric shapes (such as triangles), where each basic geometric shape can be regarded as a primitive.
[0027] In related technologies, graphics processing can be performed based on, for example, a GPU rendering pipeline or a similar stage. The main task of the rendering pipeline is to convert input geometric primitives (such as points, lines, triangles, etc.) into pixels visible on the screen. To improve processing performance, the frame to be rendered can be divided into multiple tiles (also referred to as Tiles in this disclosure), and Tiled-Based Rendering (TBR) can be performed on the frame to be rendered, where each tile can be used to process a small portion of the content in the frame to be rendered. Specifically, the TBR rendering pipeline can include a front-end and a back-end. The front-end can perform vertex processing, primitive clipping, culling, and other operations to generate primitive rendering data and write it to memory. The back-end can perform tile partitioning operations according to the screen size and distribute it to different tile processing cores to perform rasterization, depth testing, pixel shading, and other operations, ultimately outputting the rendering result. In this disclosure, a tile processing core (also referred to as a tile processing core, or simply a core) can refer to a structure that performs related processing operations based on a tile, which may include raster units (RUs) and fragment units (FUs). After being assigned to a tile processing core, all primitives within a tile can be processed by the RU and then sent to the FU to perform pixel shading procedures, etc. Different tile processing cores are completely independent, and their internal resources cannot be shared.
[0028] by Figure 1 Taking the tile allocation scenario shown as an example, the frame 110 to be rendered can be divided into 16 tiles, namely tiles T0 to T15. The tile allocation module 120 can distribute the 16 tiles to different tile processing cores according to preset logic. For example, there can be 6 tile processing cores 130 to 135, each core including corresponding RU and FU. Among them, the currently available cores can be cores 130 to 133, while cores 134 and 135 may be unavailable due to failure, occupation, or other reasons. According to the preset logic, the tile allocation module 120 can, for example, distribute tiles T0 to T3 to core 130, tiles T4 to T7 to core 131, tiles T8 to T11 to core 132, and tiles T12 to T15 to core 133. Once a tile is distributed to a certain tile processing core, regardless of the load of that tile, it will always be handled by that tile processing core group. In the description of this disclosure, the load of a tile can be measured, for example, by the number of primitives involved.
[0029] In some rendering scenarios, the workload of different tiles can vary significantly. For example, in some game scenarios, there might be a need to render the hair on a character's head. This hair might consist of hundreds of thousands of primitives, often concentrated within a single tile area. In this case, after each tile is distributed to its corresponding tile processing core, for an extended period, only that single core might be working, while the others remain idle. Figure 2 As shown, assuming there is a heavily loaded tile T0 assigned to core 130, it's possible that core 130 will be active while other cores remain idle for extended periods. For example, as... Figure 3 As shown (where the horizontal arrows indicate the time period and the black filled area indicates the workload), FU0 is in a high-load working state during a longer time period, while other FUs are in an idle state.
[0030] Under the TBR architecture, the above situation may adversely affect the overall performance of the GPU. On the one hand, uneven tile load will lead to uneven tile processing core load, that is, some cores will work for a long time and some cores will work for a short time, which will increase the overall runtime and reduce graphics processing efficiency. On the other hand, when the primitive load in a frame is concentrated in a small number of tiles, only a few cores are working, while most cores are idle, which greatly wastes computing units and various on-chip storage resources.
[0031] To solve or at least alleviate the above problems, this disclosure proposes a novel graphics processing method. For example, Figure 4 A flowchart of a graphics processing method 400 is illustrated schematically. Exemplarily, this graphics processing method 400 may be executed by a GPU or other graphics processing device. Figure 4 As shown, the graphics processing method 400 may include steps 410 to 440.
[0032] In step 410, the frame to be rendered can be divided into multiple tiles. For example, the frame to be rendered can be divided into a certain number of tiles based on the screen size and / or pre-configuration. This disclosure does not specifically limit the number or method of tile division.
[0033] In step 420, the effective rendering area of the frame to be rendered can be determined. The frame to be rendered may include multiple primitives, and the effective rendering area includes the coverage area of these primitives. The coverage area of the multiple primitives should be understood as the total coverage of all primitives included in the frame to be rendered, i.e., the union of the coverage areas of all primitives. Optionally, the effective rendering area can be characterized in various ways, such as by a minimum region of a specified shape that contains the coverage area of the multiple primitives.
[0034] For example, the multiple primitives included in the frame to be rendered can be processed by the front end of the TBR rendering pipeline, wherein the front end can store relevant data in memory, and the back end can read the primitive data of the frame to be rendered from memory. It should be understood that the front end and back end of the rendering pipeline described in this disclosure are for ease of description only and are not intended to be limiting. The front end and back end of the rendering pipeline can be distinguished in other ways, or no distinction may be made. For example, for each tile, the primitives covering that tile can be recorded, for example, by recording the identifiers of the primitives covering each tile in a list or other manner, for use in subsequent processing stages.
[0035] In this disclosure, coverage can encompass both full coverage and partial coverage.
[0036] In step 430, if the size of the effective rendering region is below a threshold, tile splitting can be enabled for at least a portion of the tiles among multiple tiles. Tile splitting includes dividing a tile into a specified number of tile regions. Optionally, the size of the effective rendering region can be measured by the number of tiles it involves, or it can be measured in other ways (such as one or more side lengths, areas, etc. of the effective rendering region). Exemplarily, the aforementioned threshold for the size of the effective rendering region can be pre-configured, for example, through a relevant register. Alternatively, the threshold can also be determined adaptively, for example, based on one or more of the following pre-configured rules: the number of currently available tile processing cores, the number of tiles divided in the frame to be rendered, etc. Exemplarily, when tile splitting is enabled, the number of tile regions into which a tile is split (i.e., the aforementioned specified number) can be determined according to user configuration, which can be, for example, 2, 4, 8, or 16, etc. As a further example, a user can specify the number of tile regions a tile can be split into by configuring a tile splitting mode. The tile splitting mode can indicate a splitting method such as 1 to 2, 1 to 4, 1 to 8, or 1 to 16. Alternatively, the specified number may be determined adaptively, for example, based on one or more of the following pre-configured rules: the number of currently available tile processing cores, the number of tiles that need to be split. Optionally, depending on specific needs, tile splitting can be enabled for all tiles in the frame to be rendered, or only for a subset of tiles. For example, tile splitting can be enabled only for valid tiles or a portion of valid tiles, where valid tiles can refer to tiles in the frame to be rendered that are covered by at least one primitive.
[0037] Furthermore, in some examples, there may be more than one threshold, where each threshold can have a corresponding specified number. In other words, when the size of the effective rendering region is lower than the first threshold, when tile splitting is enabled, a tile can be split into a first specified number of tile regions; when the size of the effective rendering region is lower than the second threshold, when tile splitting is enabled, a tile can be split into a second specified number of tiles, and so on. In this example, as the threshold decreases, the specified number can increase, and the specified number corresponding to the smaller threshold can be preferentially adopted. That is, it can be preferentially determined whether the size of the effective rendering region is lower than the smallest threshold; if so, the corresponding specified number is adopted; otherwise, it is determined whether the size of the effective rendering region is lower than the second smallest threshold, and so on.
[0038] In step 440, multiple tile regions obtained through tile splitting can be allocated to multiple tile processing cores, wherein each tile processing core is configured to perform tile processing operations on the allocated tile regions. In other words, instead of the tile-based allocation method described above, allocation can be performed on a unit basis (tile region), allowing different tile regions within the same tile to be assigned to two or more different tile processing cores for processing. Here, multiple tile processing cores can refer to currently available tile processing cores, which can be all or some of the tile processing cores in the graphics processing device. For example, tile regions can be allocated in a round-robin manner among the multiple tile processing cores according to a preset order. Assuming there are 4 tile processing cores, the first 4 tile regions to be allocated can be sequentially allocated to these 4 cores according to the core number order or other specified order, then the 5th to 8th tile regions can be sequentially allocated to these 4 cores, and so on. Optionally, tile allocation can be performed based on a preset order of the cores, or it can be performed according to other rules or patterns. For example, when assigning a tile region to the tile processing core, the assigned tile region can be indicated to the tile processing core in an appropriate manner, such as by the identifier or coordinates of the tile region, or by the identifier or coordinates of the tile to which the tile region belongs and the sequence number or other identifier of the tile region therein. For the assigned tile region, the tile processing core can obtain the primitives associated with that tile region and perform tile processing operations on the assigned tile region based on the obtained primitive information. In some examples, the tile processing operations may include one or more of rasterization, depth testing, pixel shading, and color writing. The tile processing core may, for example, include the rasterization unit RU and fragment unit FU described above.
[0039] In the aforementioned graphics processing method 400, a tile-based allocation scheme is provided. This scheme allows for the further splitting of at least a portion of the tiles into tile regions when the effective rendering area is small (i.e., the primitive distribution is relatively concentrated). This enables different tile regions of the same tile to be allocated to two or more different tile processing cores for parallel processing. On the one hand, since different tile regions of high-load tiles can be processed in parallel by different tile processing cores, it helps improve the processing efficiency of such tiles. On the other hand, this helps improve the load balancing of each tile processing core and helps avoid situations where only a few cores are working while the rest are idle, thus making fuller use of the resources of each tile processing core. Therefore, the graphics processing method 400 can effectively improve graphics processing efficiency and optimize the overall performance of the graphics processing device. Furthermore, this scheme only requires minor modifications to the current TBR or similar tile-based graphics processing process, thus achieving the aforementioned performance optimization at a relatively low cost, while maintaining the low-power advantage of the TBR architecture or similar architecture.
[0040] In some embodiments, the method 400 may further include: if the size of the effective rendering area is higher than the aforementioned threshold, tile splitting may not be enabled, and allocation may be performed directly based on tiles, i.e., complete tiles may be directly allocated to the tile processing core. Therefore, when primitives are relatively dispersed, i.e. not concentrated in a small number of tiles, allocation and tile processing operations can be performed based on tiles to avoid introducing unnecessary additional processing in such cases, as these situations are generally less prone to severe core load imbalance.
[0041] In some embodiments, to perform tile region allocation more efficiently and minimize modifications to existing tile-based allocation methods, in step 440, multiple tile regions can be allocated to multiple tile processing cores by allocating multiple tiles to multiple tile processing cores, wherein each tile is allocated an equal number of times as specified above, and a tile region identifier is sent each time a tile is allocated, indicating a tile region within the allocated tile. In this embodiment, the number of times a tile is allocated corresponds to the granularity of tile splitting, which can be determined by user configuration. Here, the process of generating or determining the tile region identifier can be considered as enabling and splitting tiles, and the process of allocating the tile region identifier along with the tile to the tile processing core can be considered as allocating multiple tile regions to multiple tile processing cores. For example, when the specified number is n (n is an integer greater than or equal to 2), that is, when a tile is split into n tile regions, for each split tile, the tile can be allocated n times (e.g., broadcast n times), each time carrying a different tile region identifier. In this way, the allocation of tile regions obtained by splitting a tile can be conveniently implemented, where the tile region identifier can uniquely identify the allocated tile region in multiple allocations. The tile processing core can accurately determine the tile region to be processed based on the received tile region identifier and relevant information of the current tile (such as its size and the number of splits).
[0042] In some examples, when tile splitting is not enabled, the tile region identifier may not be sent to the tile processing core, or it may be set to a pre-configured value indicating that tile splitting is not currently enabled, meaning that the current allocation is based on complete tiles. Accordingly, the tile processing core can determine whether tile splitting is enabled by reading the tile region identifier or by checking its presence or absence. If enabled, the tile region to be processed is determined based on the tile region identifier; otherwise, processing is performed directly based on complete tiles.
[0043] Further exemplifying, suppose n is 4, meaning a tile can be divided into 4 tile regions, or the tile can be allocated 4 times. The first allocation of the tile can carry a tile region identifier indicating the first tile region within the tile; the second allocation can carry a tile region identifier indicating the second tile region; the third allocation can carry a tile region identifier indicating the third tile region; and the fourth allocation can carry a tile region identifier indicating the fourth tile region. Further exemplifying, with... Figure 1Taking the architecture shown as an example, when allocating the above four tile regions between cores 130 and 133 in a round-robin manner, the tile region corresponding to the first allocation can be allocated to core 130, the tile region corresponding to the second allocation can be allocated to core 131, the tile region corresponding to the third allocation can be allocated to core 132, and the tile region corresponding to the fourth allocation can be allocated to core 133. At this time, compared to... Figure 2 , 3 As shown in the diagram, when the allocation is done in units of tiles, a single heavily loaded tile can be processed by four cores in parallel instead of a single core. Each core processes only one of the four tile regions. This obviously improves the processing efficiency of the heavily loaded tile and makes fuller use of the processing resources of the four cores (130 to 133), which helps to balance the load among the cores.
[0044] In some examples, the tile splitting method can be configured as shown in the table below: Each MSAA (Multisample Anti-Aliasing) mode corresponds to a specific tile size. 1->2, 1->4, 1->8, and 1->16 correspond to tile splitting modes of 1-to-2, 1-to-4, 1-to-8, and 1-to-16, respectively. The table shows the tile region sizes for different MSAA modes (i.e., different tile sizes) and different tile splitting modes. In this type of example, the tile processing core can locate the actually assigned tile region based on the MSAA mode, the tile splitting mode, and the received tile region identifier. The MSAA mode and tile splitting mode can be obtained directly by reading configuration information (e.g., reading relevant registers).
[0045] For example, Figure 5 The illustrations depict tile splitting under different MSAA modes and different tile splitting modes according to the examples above. For example... Figure 5As shown, Figures (a) to (d) correspond to the MSAA 1X / 2X mode, and respectively to the tile splitting modes 1->2, 1->4, 1->8, and 1->16; Figures (e) to (g) correspond to the MSAA 4X mode, and respectively to the tile splitting modes 1->2, 1->4, and 1->8; Figures (h) to (i) correspond to the MSAA 8X mode, and respectively to the tile splitting modes 1->2 and 1->4. Each small square in the figure corresponds to a tile region, and the number can be regarded as the tile region identifier of the corresponding tile region. In the above example, the value of the tile region identifier can be 0 to 16, where 0 indicates that tile splitting is not enabled, that is, allocation is based on tiles, and 1 to 16 are the identifiers carried when 1->2 / 1->4 / 1->8 / 1->16, to indicate the tile regions allocated after tile splitting. Taking Figure (b) as an example, when the MSAA mode is 1X or 2X and a 1-to-4 tile splitting mode is used, a tile can be split into four tile regions as shown in the figure. When allocating these four tile regions, the tile can be allocated four times (e.g., broadcast four times), carrying tile region identifiers with values of 1, 2, 3, and 4 respectively, so that tile regions 1, 2, 3, and 4 as shown in the figure are respectively allocated to two or more tile processing cores. At this time, assuming that a tile processing core receives a tile region identifier with a value of 2, it can determine that the allocated tile region is tile region 2 shown in Figure (b) of the corresponding tile.
[0046] It should be understood that the above method of distinguishing different sized blocks using the MSAA mode is merely illustrative. Depending on actual needs, other methods can be used to determine and / or indicate block sizes. Furthermore, the above table and... Figure 5 The tile splitting method described herein is merely exemplary. Those skilled in the art can design and adopt other tile splitting methods according to actual needs.
[0047] In some embodiments, the effective rendering area determined in step 420 may be the smallest rectangular region encompassing multiple primitives, such as a bounding box (BBox). For example, as shown... Figure 6As shown, assuming the current frame to be rendered is divided into 9 tiles (i.e., T0-T8 in the figure) and includes several primitives (triangles in the figure), the determined effective rendering area can be the smallest rectangular area containing all primitives in the frame to be rendered, i.e., area 610 shown by the dashed line in the figure. For example, this effective rendering area can be determined based on the vertex coordinate information of the vertex primitives in the frame to be rendered, such as based on the maximum x-coordinate value x_max, the minimum x-coordinate value x_min, the maximum y-coordinate value y_max, and the minimum y-coordinate value y_min of all primitive vertices. These coordinates can be used to determine the coordinates of the four vertices of the rectangular effective rendering area. For example, the minimum x-coordinate, maximum x-coordinate, minimum y-coordinate, and maximum y-coordinate of the four vertices of the effective rectangular rendering area can correspond to x_min, x_max, y_min, and y_max, respectively. That is, the coordinates of the four vertices can be (x_min, y_min), (x_min, y_max), (x_max, y_min), and (x_max, y_max). Therefore, the corresponding rectangular area can be indicated by passing these four vertex coordinates or by passing the four coordinate values x_max, x_min, y_max, and y_min. By limiting the effective rendering area to the smallest rectangular area that includes the coverage area of multiple primitives, the extent of the effective rendering area can be determined and indicated more conveniently, improving overall processing efficiency.
[0048] In some embodiments, the size of the effective rendering region can be characterized by the number of tiles covered by the effective rendering region. Correspondingly, the threshold for determining whether to enable tile splitting, compared with the size of the effective rendering region, can also be characterized by the number of tiles. This representation method can more intuitively reflect the concentration of primitives in the frame to be rendered and facilitates the determination of whether primitives are concentrated in a small number of tiles. For example, when the effective rendering region is determined as the smallest rectangular region encompassing multiple primitives, the size of the effective rendering region can be the number of tiles covered by that smallest rectangular region. Figure 6 For example, the effective rendering area 610 of the rectangle can be 4 tiles. Assuming the threshold is set to 3, tile splitting can be disabled, and allocation can be performed directly based on tiles; assuming the threshold is set to 5, tile splitting can be enabled, and allocation can be performed based on tile regions.
[0049] As mentioned earlier, when dividing the frame to be rendered into multiple tiles in step 410, information about the primitives covering each tile can be obtained. However, in step 440, what is actually allocated to the multiple tile processing cores is the tile region within the tile. That is, each tile processing core does not need to process the entire tile, but rather a portion of it. In this case, the following situation may occur: for a certain tile region, one or more primitives of the tile to which the tile region belongs may not cover that tile region. For example, as... Figure 7 As shown, assuming a tile is divided into four tile regions R0 to R3, the current tile processing core needs to process tile region R1, while the primitives covering this tile are P0 to P6, of which primitives P0 to P5 do not cover tile region R1. To reduce the processing pressure in tile processing operations, such as reducing the processing pressure in the rasterization stage and subsequent stages, improving processing efficiency, and saving processing resources, the tile processing core can pre-select primitives that do not cover the tile region to be processed. Therefore, in some embodiments, the above-mentioned graphics processing method 400 may further include: each tile processing core performing the following operations for each assigned tile region: obtaining a set of primitives corresponding to the tile that includes the tile region, wherein the set of primitives includes primitives covering the tile; discarding primitives that do not cover the tile region from the obtained set of primitives; and performing tile processing operations based on the remaining primitives in the set of primitives. Continuing with... Figure 7 For example, after a tile region R1 is assigned in the shown tile, the tile processing core (such as RU) can obtain the primitive information corresponding to that tile, such as the primitive list from P0 to P6 shown, and determine one by one whether each primitive covers the tile region R1. Figure 7 In the case shown, primitives P0 to P5 that do not cover the tile area R1 will be discarded or removed, while primitive P6 that covers the tile area R1 will be retained. Subsequently, the tile processing core can perform tile processing operations based solely on primitive P6, such as rasterization, depth testing, pixel shading, and color writing.
[0050] In some embodiments, to more conveniently and efficiently determine whether a graphic element covers the current tile region and discard graphic elements that do not cover the current tile region, after obtaining the set of graphic elements covering the current tile, the graphic elements in the set can be traversed. For each graphic element, if the graphic element region corresponding to that graphic element has no intersection with the tile region, the graphic element is discarded. The graphic element region corresponding to that graphic element can, for example, be the smallest rectangular region containing that graphic element. Continuing with... Figure 7For example, for each primitive, a corresponding rectangular bounding box (as shown by the dashed line) can be determined. The determination method can be similar to the method for determining the effective rectangular rendering area described earlier. Taking primitive P5 as an example, a rectangular bounding box 710 can be determined based on the coordinates of its three vertices. The intersection of the range of the rectangular bounding box 710 and the range of the tile region R1 is then calculated. If there is no intersection, primitive P5 can be removed; otherwise, it can be retained.
[0051] Figure 8 The illustration schematically depicts an example process of a tile processing core performing depth or color write-out according to some embodiments of the present disclosure. As mentioned in the preceding embodiments, after a tile region is assigned, the tile processing core can perform tile processing operations only on that tile region (rather than the entire tile). Specifically, to avoid repeated data reading and save bandwidth, when performing depth testing, data loading can be performed only within the assigned tile region. Furthermore, depth or color write-out can be performed only on the assigned tile region to ensure the correctness of the rendering results. Figure 8 As shown, assuming that a certain tile processing core is assigned to tile region R1 in tile 810, then in the depth / color surface 820, the core can only write the depth or color within the region corresponding to tile region R1 in the region 821 corresponding to tile 810 (as shown in the shaded area) to avoid affecting the rendering results of other regions.
[0052] To further facilitate understanding, Figure 9 Taking the example of a single tile being allocated twice (i.e., split into two tile regions), the graphic processing procedure according to some embodiments of this disclosure is illustrated. First, the tile partitioning module 910 can execute the aforementioned steps 410 and 420 to divide the frame to be rendered into multiple tiles, generate corresponding tile information (such as a list of primitives covering each tile), and determine the effective rendering area in the frame to be rendered, such as a rectangular effective rendering area BBox. The tile partitioning module 910 can transmit the tile information and the effective rendering area to the allocation module 920, wherein the rectangular effective rendering area BBox can, for example, be indicated by the minimum x-coordinate value, maximum x-coordinate value, minimum y-coordinate value, and maximum y-coordinate value of its four vertices. The allocation module 920 can, for example, use the minimum x-coordinate value, maximum x-coordinate value, minimum y-coordinate value, and maximum y-coordinate value of its four vertices. Figure 1The tile allocation module 120 shown is modified to execute the aforementioned steps 430 and 440. It determines whether tile splitting should be enabled based on the size of the effective rendering area and allocates tiles or tile regions to the tile processing cores. Specifically, the allocation module 920 can determine the number of tiles covered by the rectangular effective rendering area BBox. If this number is less than a preset threshold, tile splitting is enabled; otherwise, it is disabled. As shown, when tile splitting is enabled, tile 901 can be split into two tile regions: the upper tile region 902 and the lower tile region 903 (as shown in the shaded area). Correspondingly, tile 901 can be allocated twice. For example, it can be broadcast in a polling manner according to the core order, where tile 901 can be broadcast twice and allocated to the two cores shown, the first core including RU0 and FU0, and the second core including RU1 and FU1. RU0 and FU0 can process only tile region 902, and RU1 and FU1 can process only tile region 903. After obtaining the set of primitives corresponding to tile 901, RU0 and RU1 can perform primitive filtering for their respective tile regions to discard primitives that do not cover the corresponding tile region, and perform rasterization, depth testing, and other operations based on the remaining primitives. Subsequently, FU0 and FU1 can perform pixel coloring and other operations based on the data from RU0 and RU1 respectively. OM0 and OM1 can perform color writing only for their respective tile regions. All of the above operations can be performed according to the relevant embodiments described above, and will not be elaborated further here.
[0053] Testing showed that when using the tile-based region allocation graphics processing scheme proposed in this disclosure, compared to the reference... Figures 1 to 3 The described tile-based allocation can improve performance by up to 6 times or more when rendering certain relatively concentrated scenes.
[0054] This disclosure also proposes a graphics processing device. Exemplarily, Figure 10 A block diagram of a graphics processing device 1000 according to some embodiments of the present disclosure is shown schematically.
[0055] like Figure 10As shown, the graphics processing device 1000 may include a tile partitioning module 1010, an allocation module 1020, and multiple tile processing cores 1030. Specifically, the tile partitioning module 1010 may be configured to: divide a frame to be rendered into multiple tiles, and determine the effective rendering area of the frame to be rendered, wherein the frame to be rendered includes multiple primitives, and the effective rendering area includes the coverage area of the multiple primitives; the allocation module 1020 may be configured to: enable tile splitting for at least a portion of the multiple tiles when the size of the effective rendering area is lower than a threshold, and allocate the multiple tile areas obtained by tile splitting to multiple tile processing cores, wherein tile splitting includes splitting a tile into a specified number of tile areas; the multiple tile processing cores 1030, wherein each tile processing core may be configured to: perform tile processing operations for the allocated tile areas.
[0056] The graphics processing device 1000 described above can be implemented in the form of circuits, chips, etc. For example, the graphics processing device 1000 can be a GPU or other similar device. Furthermore, the graphics processing device 1000 can be used to implement the steps in the graphics processing method 400 described above, the relevant details of which have been described in detail above and will not be repeated here for the sake of brevity. The graphics processing device 1000 can have the same features and advantages as described with respect to the foregoing method.
[0057] This disclosure also provides a computing device. Exemplarily, Figure 11 An exemplary block diagram of a computing device 1100 is shown. As illustrated, the computing device 1100 may include a graphics processing device 1000, which can be used to perform graphics processing operations and other related operations within the computing device 1100. Optionally, the computing device 1100 may also include various structures such as memory, a central processing unit (CPU), and I / O (input / output) interfaces.
[0058] By studying the accompanying drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps, and "a" or "an" does not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not imply that a combination of these measures cannot be used for profit.
Claims
1. A graphics processing method, characterized in that, The graphics processing method includes: Divide the frame to be rendered into multiple tiles; Determine the effective rendering area of the frame to be rendered, wherein the frame to be rendered includes multiple primitives, and the effective rendering area includes the coverage area of the multiple primitives; If the size of the effective rendering area is below a threshold, tile splitting is enabled for at least a portion of the plurality of tiles, wherein tile splitting includes splitting a tile into a specified number of tile regions; and, Multiple tile regions obtained by tile splitting are assigned to multiple tile processing cores, wherein each tile processing core is configured to perform tile processing operations on the assigned tile regions.
2. The graphics processing method according to claim 1, characterized in that, The step of allocating multiple tile regions obtained through tile splitting to multiple tile processing cores includes: The plurality of tiles are allocated to the plurality of tile processing cores, wherein each tile is allocated a number of times equal to the specified number, and a tile region identifier is sent each time a tile is allocated, the tile region identifier indicating a tile region among the allocated tiles.
3. The graphics processing method according to claim 1, characterized in that, The graphics processing method further includes: If the size of the effective rendering area is greater than the threshold, the tile splitting is not enabled; The plurality of tiles are assigned to the plurality of tile processing cores, wherein each tile processing core is configured to perform tile processing operations on the assigned tiles.
4. The graphics processing method according to claim 3, characterized in that, When allocating each tile, a tile region identifier with a pre-configured value is sent, wherein the pre-configured value indicates that tile splitting is not enabled.
5. The graphics processing method according to claim 1, characterized in that, The effective rendering area is the smallest rectangular area that includes the coverage area of the multiple primitives.
6. The graphics processing method according to claim 1, characterized in that, The size of the effective rendering area is characterized by the number of tiles covered by the effective rendering area.
7. The graphics processing method according to claim 1, characterized in that, The graphics processing method further includes: each tile processing core performing the following operations for each assigned tile region: Obtain a set of primitives corresponding to a tile that includes the tile region, wherein the set of primitives includes primitives that cover the tile; In the set of elements, elements that do not cover the tile area are discarded; and The block processing operation is performed based on the remaining primitives in the primitive set.
8. The graphics processing method according to claim 7, characterized in that, The step of discarding elements that do not cover the tile area in the element set includes: Traverse the primitives in the primitive set, and for each primitive, discard the primitive if the primitive region corresponding to the primitive does not intersect with the block region, wherein the primitive region corresponding to the primitive is the smallest rectangular region containing the primitive.
9. The graphic processing method according to any one of claims 1-8, characterized in that, The specified quantity is 2, 4, 8 or 16.
10. The graphic processing method according to any one of claims 1-8, characterized in that, The tile processing operations include one or more of the following: rasterization, depth testing, pixel shading, and color writing.
11. A graphics processing device, characterized in that, The graphics processing device includes: The tile division module is configured to: divide the frame to be rendered into multiple tiles, and determine the effective rendering area of the frame to be rendered, wherein the frame to be rendered includes multiple primitives, and the effective rendering area includes the coverage area of the multiple primitives; The allocation module is configured to: when the size of the effective rendering area is lower than a threshold, enable tile splitting for at least a portion of the multiple tiles, and allocate the multiple tile regions obtained by the tile splitting to multiple tile processing cores, wherein the tile splitting includes splitting a tile into a specified number of tile regions; The plurality of tile processing cores, wherein each tile processing core is configured to perform tile processing operations for an assigned tile region.
12. A computing device, characterized in that, The computing device includes the graphics processing device according to claim 11.