A method, device and storage medium for eliminating invisible pixels

By processing primitives in parallel in the rendering core of the GPU and removing obscured fragments, the resource waste caused by invisible fragments generated by rasterization is solved, and more efficient pixel culling and rendering efficiency is achieved.

CN114037795BActive Publication Date: 2025-05-16XIAN XINTONG SEMICON TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111405905.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-05-16
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

When the GPU renders the graphics, there are invisible fragments in the fragments generated by rasterization, resulting in waste of resources and affecting rendering efficiency and power consumption.

Method used

By configuring a rasterization module in the rendering core to process the primitives in parallel, the fragments are generated and the obscured fragments are eliminated based on the coordinate values ​​and depth values ​​of the fragments during the output process.

Benefits of technology

Improves the efficiency and effect of pixel culling, reduces the number and timing of clip comparisons, avoids resource waste, improves the rendering efficiency of the GPU and reduces power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114037795B_ABST
    Figure CN114037795B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a method, device and storage medium for removing invisible pixels. The method comprises: performing parallel rasterization processing on each of all the primitives covering the current tile to be processed, and obtaining the fragments corresponding to each primitive; outputting the fragments of all the primitives according to the set coordinate sequence to be processed into the fragments to be processed, and removing the blocked fragments from the fragments to be processed based on the coordinate value and depth value of the fragments during the output process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a method, device and storage medium for eliminating invisible pixels. Background Art

[0002] In general, a GPU is a specialized graphics rendering device for processing and displaying computerized graphics. The GPU is constructed with a highly parallel structure that provides more efficient processing than a typical general-purpose central processing unit (CPU) for a range of complex algorithms. For example, the complex algorithms may correspond to the representation of two-dimensional (2D) or three-dimensional (3D) computerized graphics.

[0003] However, in the process of GPU reproducing graphics, especially under the conditions of power and system bandwidth constraints, a Tile Based Rendering (TBR) solution is usually adopted. Specifically, each shader core is responsible for rendering one tile at a time, and each tile records all the primitives covering itself. The list of these primitives is the primitive list. The rasterization module in the rendering core traverses the primitive list, rasterizes the primitives one by one, and then hands the fragments generated by rasterization to the fragment shader module for fragment shading.

[0004] However, some of the fragments generated by the above rasterization will not be displayed (that is, they are invisible), and fragment shading is very time-consuming and power-consuming. In other words, shading the fragments that will not be displayed will waste time and power. If the fragments that will not be displayed can be eliminated before shading, the rendering efficiency of the GPU can be improved while reducing power consumption. Summary of the invention

[0005] In view of this, the embodiments of the present invention are intended to provide a method, device and computer storage medium for removing invisible pixels, which can achieve a good pixel removal effect and improve the efficiency of pixel removal.

[0006] The technical solution of the embodiment of the present invention is achieved as follows:

[0007] In a first aspect, an embodiment of the present invention provides a device for removing invisible pixels, comprising:

[0008] at least one rendering core and at least one rasterization module;

[0009] Each of the at least one rasterization module is configured to perform parallel rasterization processing on each of all the primitives covering the current tile to be processed, to obtain a fragment corresponding to each primitive;

[0010] Each of the at least one rendering core is configured to output the fragments of all the primitives that need to be processed by fragment shading in a set coordinate order, and to remove occluded fragments from the fragments that need to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process.

[0011] In a second aspect, an embodiment of the present invention provides a method for removing invisible pixels, comprising:

[0012] Perform parallel rasterization processing on each of the graphic elements covered by the current tile to be processed, and obtain the fragment corresponding to each graphic element;

[0013] The fragments of all the primitives are output as fragments that need to be processed by fragment shading in a set coordinate order, and blocked fragments are removed from the fragments that need to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process.

[0014] In a third aspect, an embodiment of the present invention provides a graphics processor (GPU), comprising: the invisible pixel culling device described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer storage medium, wherein the computer storage medium stores a program for eliminating invisible pixels, and when the program for eliminating invisible pixels is executed by at least one processor, the steps of the method for eliminating invisible pixels described in the second aspect are implemented.

[0016] The embodiments of the present invention provide a method, device and computer storage medium for eliminating invisible pixels, which can change the serial rasterization processing of the primitive list of Tile into parallel rasterization processing. At the same time, the position values ​​of the fragments stored in the FIFO queue of the rasterization module are changed from vertical comparison to horizontal comparison, and then the fragments of all primitives corresponding to the Tile are eliminated, thereby advancing the timing of fragment comparison, reducing the number of comparisons, and each fragment will be compared without omissions. Therefore, the pixel elimination effect is better and the efficiency is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A block diagram of a computing device capable of implementing the technical solution of an embodiment of the present invention;

[0018] Figure 2 A block diagram of a GPU that can implement the technical solution of an embodiment of the present invention;

[0019] Figure 3 Based on Figure 2 A schematic diagram of a graphics rendering pipeline formed by the structure shown;

[0020] Figure 4 An exemplary task scheduling diagram provided for an embodiment of the present invention;

[0021] Figure 5 Another exemplary task scheduling diagram provided for an embodiment of the present invention;

[0022] Figure 6 An exemplary schematic diagram of a rasterization module scanning a primitive provided by an embodiment of the present invention;

[0023] Figure 7 A further exemplary task scheduling diagram provided for an embodiment of the present invention;

[0024] Figure 8 A schematic diagram of a method for eliminating invisible pixels provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention. Figure 1 , which shows a computing device 100 capable of implementing the technical solution of the embodiment of the present invention, and the computing device 100 may include but is not limited to the following: wireless devices, mobile or cellular phones (including so-called smart phones), personal digital assistants (PDAs), video game consoles (including video displays, mobile video game devices, mobile video conferencing units), laptop computers, desktop computers, TV set-top boxes, tablet computing devices, e-book readers, fixed or mobile media players, etc. Figure 1In an example of , computing device 100 may include a central processing unit (CPU) 102 and a system memory 104 that communicates via an interconnect path of a memory bridge 105. Memory bridge 105 may be, for example, a north bridge chip, connected to an I / O (input / output) bridge 107 via a bus or other communication path 106 (e.g., a HyperTransport link). I / O bridge 107, which may be, for example, a south bridge chip, receives user input from one or more user input devices 108 (e.g., a keyboard, mouse, trackball, touch screen that can be incorporated as an integral part of a display device 110, or other type of input device) and forwards the input to CPU 102 via communication path 106 and memory bridge 105. A graphics processor (GPU) 112 is coupled to memory bridge 105 via a bus or other communication path 113 (e.g., PCI Express, Accelerated Graphics Port, or HyperTransport link); in one embodiment, GPU 112 may be a graphics subsystem that delivers pixels to display device 110 (e.g., a conventional CRT or LCD-based monitor). A system disk 114 is also connected to I / O bridge 107. Switch 116 provides connections between I / O bridge 107 and other components such as a network adapter 118 and various add-in cards 120 and 121. Other components (not explicitly shown) including USB or other port connections, CD drives, DVD drives, film recording devices, and the like may also be connected to I / O bridge 107. Figure 1 The communication paths interconnecting the various components in the system may be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol, and the connections between different devices may use different protocols known in the art.

[0026] In one embodiment, GPU 112 includes circuits optimized for graphics and video processing, including, for example, video output circuits. In another embodiment, GPU 112 includes circuits optimized for general processing while preserving the underlying computing architecture. In yet another embodiment, GPU 112 may be integrated with one or more other system elements, such as memory bridge 105, CPU 102, and I / O bridge 107, to form a system on chip (SoC).

[0027] It should be understood that the system shown herein is exemplary, and variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of CPUs 102, and the number of GPUs 112, can be modified as needed. For example, in some embodiments, the system memory 104 is directly connected to the CPU 102 instead of through a bridge, and other devices communicate with the system memory 104 via the memory bridge 105 and the CPU 102. In other alternative topologies, the GPU 112 is connected to the I / O bridge 107 or directly to the CPU 102, rather than to the memory bridge 105. In other embodiments, the I / O bridge 107 and the memory bridge 105 may be integrated into a single chip. A large number of embodiments may include two or more CPUs 102 and two or more GPUs 112. The specific components shown herein are optional; for example, any number of add-in cards or peripheral devices may be supported. In some embodiments, the switch 116 is removed, and the network adapter 118 and the add-in cards 120, 121 are directly connected to the I / O bridge 107.

[0028] based on Figure 1 The computing device 100 shown, Figure 2 A schematic block diagram of a GPU 112 that can implement one or more technical solutions of an embodiment of the present invention is shown. In an embodiment of the present invention, a graphics memory 204 may be part of the GPU 112. Therefore, the GPU 112 can read data from the graphics memory 204 and write data to the graphics memory 204 without using a bus. In other words, the GPU 112 can use a local storage device instead of an off-chip memory to process data locally. Such a graphics memory 204 may be referred to as an on-chip memory. This allows the GPU 112 to operate in a more efficient manner by eliminating the need for the GPU 112 to read and write data via a bus, where operation via the bus may experience heavy bus traffic. However, in some cases, the GPU 112 may not include a separate memory, but instead utilize the system memory 10 via the bus. The graphics memory 204 may include one or more volatile or non-volatile memories or storage devices, such as random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic data media, or optical storage media.

[0029] Based on this, GPU 112 can be configured to perform various operations related to: generating pixel data from graphics data provided by CPU 102 and / or system memory 104 via memory bridge 105 and communication path 113, interacting with local graphics memory 204 (e.g., a common frame buffer) to store and update pixel data, transferring pixel data to display device 110, etc.

[0030] In operation, CPU 102 is the main processor of computing device 100, controlling and coordinating the operation of other system components. Specifically, CPU 102 issues commands to control the operation of GPU 112. In some embodiments, CPU 102 writes a command stream to a data structure for GPU 112 (in Figure 1 or Figure 2 The data structures may be located in system memory 104, graphics memory 204, or other storage locations accessible to both CPU 102 and GPU 112. A pointer to each data structure is written to a pushbuffer to start processing the command stream in the data structure. GPU 112 reads the command stream from one or more pushbuffers and then executes the commands asynchronously with respect to the operation of CPU 102. An execution priority may be specified for each pushbuffer to control the scheduling of different pushbuffers.

[0031] Specific as Figure 2 As described in , GPU 112 includes an I / O (input / output) unit 205 that communicates with the rest of the computing device 100 via a communication path 113 connected to the memory bridge 105 (or, in an alternative embodiment, directly connected to the CPU 102). The connection of GPU 112 to the rest of the computing device 100 may also vary. In some embodiments, GPU 112 may be implemented as an add-in card that can be inserted into an expansion slot of the computer system 100. In other embodiments, GPU 112 may be integrated on a single chip with a bus bridge such as memory bridge 105 or I / O bridge 107. In still other embodiments, some or all elements of GPU 112 may be integrated on a single chip with CPU 102.

[0032] In one embodiment, communication path 113 may be a PCI-EXPRESS link in which dedicated lanes are allocated to GPU 112, as is known in the art. I / O unit 205 generates data packets (or other signals) for transmission on communication path 113, and also receives all incoming data packets (or other signals) from communication path 113, directing the incoming data packets to the appropriate components of GPU 112. For example, commands related to processing tasks may be directed to scheduler 207, while commands related to memory operations (e.g., reading or writing to graphics memory 204) may be directed to graphics memory 204.

[0033] In the GPU 112, a plurality of rendering cores may be included to form a rendering core array 230. Further, as Figure 2 As shown, the rendering core array 230 may include C general rendering cores 208, where C>1; and D fixed-function rendering cores 209. It can be understood that Figure 2 The numbers in the brackets in represent the labels of the general rendering core 208 or the fixed function rendering core 209. Based on the general rendering cores 208 in the array 230, the GPU 112 is able to concurrently execute a large number of program tasks or computing tasks. For example, each general rendering core can be programmed to be able to perform processing tasks related to a wide variety of programs, including but not limited to linear and nonlinear data transformations, video and / or audio data filtering, modeling operations (e.g., applying physical laws to determine the position, velocity and other properties of an object), graphics rendering operations (e.g., tessellation shaders, vertex shaders, geometry shaders, and / or fragment shader programs), etc.

[0034] The fixed function rendering core 209 may include hardware that is hardwired to perform certain functions. Although the fixed function hardware can be configured to perform different functions via, for example, one or more control signals, the fixed function hardware generally does not include a program memory capable of receiving user compiled programs. In some instances, the fixed function rendering core 209 may include, for example, a processing unit that performs primitive assembly, a processing unit that performs clipping and partitioning operations, a processing unit that performs rasterization operations, and a processing unit that performs fragment operations. For the processing unit that performs primitive assembly, it can restore the grid structure of the graphics, i.e., primitives, from the vertices that have been shaded by the vertex shader unit according to the original connection relationship, so as to provide processing for the subsequent fragment shader unit; the clipping and partitioning operations include clipping and culling the assembled primitives and then dividing them according to the size of the tile; the rasterization operation includes converting the primitives and outputting the fragments to the fragment shader; and the fragment operation includes, for example, depth value testing, scissor testing, alpha blending, etc. The pixel data output by the above operations can be displayed as graphics data through the display device 110.

[0035] By combining the general rendering core 208 and the fixed-function rendering core 209 in the rendering core array 230 , a complete logic model of a graphics rendering pipeline can be implemented.

[0036] In addition, the rendering core array 230 can receive processing tasks to be executed from the scheduler 207. The scheduler 207 can independently schedule the tasks to be executed by the resources of the GPU 112 (such as one or more general rendering cores 208 and fixed function rendering cores 209 in the rendering core array 230). In one example, the scheduler 207 can be a hardware processor. Figure 2 In the example shown in , the scheduler 207 may be included in the GPU 112. In other examples, the scheduler 207 may also be a unit separate from the CPU 102 and the GPU 112. The scheduler 207 may also be configured as any processor that receives a stream of commands and / or operations.

[0037] The scheduler 207 may process one or more command streams, which include scheduling operations included in the one or more command streams to be executed by the GPU 112. Specifically, the scheduler 207 may process one or more command streams and schedule operations in the one or more command streams to be executed by the rendering core array 230. In operation, the CPU 102 Figure 1 The GPU driver 103 included in the system memory 104 may send a command stream including a series of operations to be executed by the GPU 112 to the scheduler 207. The scheduler 207 may receive an operation stream including the command stream through the I / O unit 205 and may sequentially process the operations of the command stream based on the operation order in the command stream, and may schedule the operations in the command stream to be executed by one or more rendering cores in the rendering core array 230.

[0038] Also, the Tile cache 232 is a small amount of extremely high bandwidth memory located on the chip together with the GPU 112. However, the size of the Tile cache 232 is too small to hold the entire graphics data, so the rendering core array 230 must perform multiple rendering passes to reproduce the entire graphics data. For example, the rendering core array 230 may perform one rendering pass for each Tile of a frame of image. Specifically, the Tile cache 232 may include one or more volatile or non-volatile memories or storage devices, such as random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), etc. In some examples, the Tile cache 232 may be an on-chip buffer. The on-chip buffer may refer to a buffer formed, positioned and / or disposed on the same microchip, integrated circuit and / or die as the microchip, integrated circuit and / or die on which the GPU 112 is formed, positioned and / or disposed. In addition, when the Tile cache 232 is implemented on the same chip as the GPU 112, the GPU 112 does not necessarily need to access the Tile cache 232 via the communication path 113, but can access the Tile cache 232 via an internal communication interface (e.g., a bus) implemented on the same chip as the GPU 112. Because this interface is on the chip, it can operate at a higher bandwidth than the communication path 113. It can be seen that although the storage capacity of the Tile cache 232 is limited and the hardware overhead is increased, and it can only be used to cache one or a few small rectangles of data, it avoids the overhead of repeatedly accessing the video memory, reduces bandwidth, and saves power consumption.

[0039] Based on the above Figure 1 and Figure 2 Description, Figure 3 Shown Figure 2 The example of the graphics rendering pipeline 80 formed by the structure of the GPU 112 shown in FIG. 1 is that the core part of the graphics rendering pipeline 80 is a logical structure formed by cascading the universal rendering core 208 and the fixed function rendering core 209 included in the rendering core array 230. In addition, the scheduler 207, the graphics memory 204, the tile cache 232 and the I / O unit 205 included in the GPU 112 are all peripheral circuits or devices for realizing the function of the logical structure of the graphics rendering pipeline 80. Accordingly, the graphics rendering pipeline 80 generally includes a programmable level module (such as Figure 3 The rounded box in the figure) and the fixed function level modules (such as Figure 3 For example, the function of the programmable level module can be performed by the general rendering core 208 included in the rendering core array 230, and the function of the fixed function level module can be implemented by the fixed function rendering core 209 included in the rendering core array 230. Figure 3 As shown, the graphics rendering pipeline 80 includes the following stages:

[0040] Vertex grabbing module 82, in Figure 3 8 is shown as a fixed function level in the example of and is generally responsible for supplying graphics data (triangles, lines, and points) to the graphics rendering pipeline 80. For example, vertex fetch module 82 may collect vertex data for high-order surfaces, primitives, etc., and output the vertex data and attributes to vertex shader module 84.

[0041] Vertex shader module 84, in Figure 3 It is shown as a programmable stage in and is responsible for processing the received vertex data and attributes and processing the vertex data by performing a set of operations on each vertex at a time.

[0042] The primitive assembly module 86, Figure 3 8 is shown as a fixed function level, which is responsible for collecting the vertices output by the vertex shader module 84 and assembling the vertices into geometric primitives. For example, the primitive assembly module 86 can be configured to assemble every three consecutive vertices into a geometric primitive (i.e., a triangle). In some embodiments, a particular vertex can be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip can share two vertices).

[0043] The cutting and dividing module 88, in Figure 3 The fixed function level is shown in the figure, which is responsible for cutting and culling the assembled primitives and dividing them according to the size of the tile;

[0044] Rasterization module 90 is generally a fixed function level responsible for preparing primitives for fragment shader module 92. For example, rasterization module 90 may generate fragments for fragment shader module 92 to process for shading.

[0045] Fragment shader module 92, in Figure 3 , a programmable stage is shown receiving fragments from rasterization module 90 and generating per-pixel data such as color. Fragment shader module 92 may also perform per-pixel processing such as texture blending and lighting model calculations.

[0046] Output merger module 94, in Figure 3 , which is shown as a fixed function level, and is generally responsible for performing various operations on the pixel data, such as performing an alpha test, a stencil test, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the output merger module 94 has completed processing the pixel data (i.e., output data), the processed pixel data can be written to the rendering target to produce a final result.

[0047] For conventional TBR schemes, the screen area is usually divided into multiple tiles of equal size. For a frame of image, after the primitive assembly stage is completed, GPU112 calculates which tiles on the screen are covered by the primitive based on the size of the primitive, and establishes a primitive list for each tile. Once the tile is covered by the primitive, the corresponding primitive information is updated in the primitive list of the tile until all the primitives are collected. After the subsequent rasterization and other stages are completed, GPU112 will traverse the primitive list of each tile (for example, a tile may be covered by multiple primitives). After rendering each primitive in a primitive list, the data of the tile is written to the on-chip cache. The final data of the tile is not written to the video memory until all the primitives in the list are processed.

[0048] Based on the above description, it can be seen that each rendering core is processed in units of Tile, that is, the rasterization module in each rendering core will traverse the primitive list of the Tile allocated by the scheduler, rasterize each primitive one by one, and then hand over the fragments generated by the rasterization to the fragment shader module for subsequent processing.

[0049] For example, if Figure 4 As shown, the rendering scene is set to cover 4 Tiles, marked as Tile-0, Tile-1, Tile-2 and Tile-3 respectively; there are 8 primitives in total, marked as primitive 0, primitive 1, primitive 2, ..., primitive 7 respectively; the primitive list corresponding to each Tile is: Tile-0 covers primitive 0 and primitive 1; Tile-1 covers primitive 1 and primitive 2; Tile-2 covers primitive 1, primitive 2, primitive 3; Tile-3 covers primitive 2, primitive 3, primitive 4, primitive 5, primitive 6, primitive 7; continue to set the number of rendering cores to 4, marked as rendering core 0, rendering core 1, rendering core 2, rendering core 3 respectively. The scheduler allocates Tile-0 to rendering core 0, whose primitive list includes primitive 0 and primitive 1. The scheduler allocates Tile-1 to rendering core 1, whose primitive list includes primitive 1 and primitive 2. After the rasterization modules in renderer 0 and renderer 1 rasterize primitive 0, primitive 1 and primitive 1 and primitive 2 respectively, the generated fragments of the corresponding primitives are handed over to their respective fragment shader modules for subsequent processing.

[0050] However, the fragment shading process will consume a lot of time and power. If the fragments that will not be displayed in the end (that is, invisible pixels) can be eliminated in advance and then handed over to the fragment shader module for subsequent processing, the rendering efficiency can be improved and the power consumption can be reduced.

[0051] The conventional method for removing the fragments that will not be displayed in the end in advance is: first, traverse the primitive list of the current Tile, and put the fragments of the current primitive into the corresponding FIFO (First In First Out) queue. Usually, the fragments of the primitives stored in the FIFO queue can include the coordinate value and depth value of the fragment in the FIFO queue. If the fragment newly entering the FIFO queue has the same coordinate value as the fragment in the current FIFO queue, that is, there are overlapping pixels, the fragment with a large depth value will be removed from the corresponding FIFO queue, because the fragment with a large depth value (which can also be understood as far from the eye) will be blocked by the fragment with a small depth value (which can also be understood as close to the eye) and will not be displayed in the end.

[0052] It should be noted that the FIFO queue is a first-in-first-out data buffer, which is essentially a RAM. Its main functions are: caching continuous data streams to prevent data loss during machine access and storage operations; centralizing data for stacking and storage to avoid frequent bus operations. The difference between the FIFO queue and ordinary memory is that there is no external read and write address line, which is simple to use, but data can only be written and read sequentially.

[0053] Based on the above characteristics of the FIFO queue, the above conventional method has the following drawbacks: Since the size of the FIFO queue corresponding to the fragment of the current primitive is limited, and occlusion only occurs between different primitives, the occlusion fragment removal effect is not good. For example: the FIFO queue size is 20, which means that 20 fragments can be stored. The first primitive of the current Tile covers 32 fragments, then the first 12 fragments have no chance to be detected and are squeezed out of the FIFO queue. Moreover, the fragments in the primitive that have the opportunity to be detected, the later the fragments are put into the FIFO queue, the more chances they will be detected, and vice versa. It can be seen that in the above conventional method, there are large differences in the probability of different fragments in the primitive being detected, and this difference will lead to poor results using the above conventional method.

[0054] Based on this, the technical solution of the embodiment of the present invention is expected to provide a technology for eliminating invisible pixels. This technology can separate multiple rasterization modules from the rendering core to form a rasterization array, and then change the scheduling of the rasterization modules from tile-based scheduling to primitive-based scheduling, thereby changing the tile-based primitive list from serial rasterization to parallel rasterization, and changing the comparison of fragments from vertical comparison to horizontal comparison, so that the timing of fragment comparison is advanced and the number of fragment comparisons is reduced, and each fragment has the opportunity to be detected, thereby achieving better pixel culling effect and higher rendering efficiency.

[0055] An embodiment of the present application provides a device for culling invisible pixels. In some examples, the device includes: at least one rendering core and at least one rasterization module;

[0056] Each of the at least one rasterization module is configured to perform parallel rasterization processing on each of all the primitives covering the current tile to be processed, to obtain a fragment corresponding to each primitive;

[0057] Each of the at least one rendering core is configured to output the fragments of all the primitives that need to be processed by fragment shading in a set coordinate order, and to remove occluded fragments from the fragments that need to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process.

[0058] For the above example, specifically, Figure 5 As shown, first, multiple rasterization modules need to be separated from the rendering core to form a rasterization array. It can also be understood that the at least one rendering core and the at least one rasterization module are independent of each other, and the number of rasterization modules and the number of rendering cores can be the same or different, which is not limited in the embodiment of the present application.

[0059] Secondly, the scheduler should schedule the rasterization module based on primitives, so that multiple primitives of the same Tile can be rasterized separately at the same time, that is, parallel rasterization. For example, assuming that the current Tile to be processed is Tile-0, the scheduler schedules primitive 0 of Tile-0 to the idle rasterization module 0, and schedules primitive 1 of Tile-0 to the idle rasterization module 1. Correspondingly, rasterization module 0 and rasterization module 1 respectively perform rasterization on primitive 0 and primitive 1 and save the generated fragments of different primitives in the corresponding FIFO queues. Exemplarily, the fragments may include: the coordinate value and depth value of the fragment in the FIFO queue. For example, rasterization module 0 saves the fragments generated after rasterization of primitive 0 in FIFO queue 0, and rasterization module 1 saves the fragments generated after rasterization of primitive 1 in FIFO queue 1.

[0060] It should be noted that each rasterization module has a corresponding relationship with a FIFO queue, that is, the FIFO queue corresponding to the determined rasterization module can be found. The FIFO queue can be a storage space in the rasterization module or other storage space, which is not limited in the present embodiment.

[0061] Then, the scheduler schedules the rendering cores based on the tile, that is, the scheduler assigns each of the primitives in the currently processed tile to the same idle rendering core. Figure 5 As shown, assuming that the current tile to be processed is Tile-0, the scheduler assigns fragments of all primitives in the primitive list corresponding to Tile-0 to the idle rendering core 0, so that rendering core 0 can obtain fragments of corresponding primitives from FIFO queue 0 of rasterization module 0 and FIFO queue 1 of rasterization module 1 respectively.

[0062] Finally, rendering core 0 outputs the fragments of primitive 0 and primitive 1 of Tile-0 as the fragments that need to be shaded according to the set coordinate order, and during the output process, removes the occluded fragments from the fragments that need to be shaded based on the coordinate values ​​and depth values ​​of the fragments of primitive 0 and primitive 1.

[0063] In some examples, each of the at least one rendering core is further configured to compare the first fragments in the FIFO queues corresponding to all the primitives, output the fragment with the smallest coordinate value as the fragment that needs to perform fragment shading processing, and eliminate fragments with the same coordinate value and a larger depth value; update the second fragment in the FIFO queue that completes the output and / or eliminates the fragment to the first fragment of the corresponding FIFO queue; based on the updated FIFO queue, compare the first fragments in the FIFO queues corresponding to all the primitives, output the fragment with the smallest coordinate value as the fragment that needs to perform fragment shading processing, and eliminate fragments with the same coordinate value and a larger depth value until the fragments in all FIFO queues are empty.

[0064] It should be noted that if Figure 6 As shown, usually, the rasterization module scans the primitives in a row-by-row manner, that is, the primitives are scanned in order from top to bottom and from left to right. For example, the rasterization module will first save the fragment with the smallest x coordinate in the coordinate value, and for the fragments with the same x coordinate, the rasterization module will save the fragments in the order of y coordinate from small to large. In other words, the rasterization module always saves the fragment with the smaller x coordinate first, and for the fragments with the same x coordinate, the rasterization module will save the fragments in the order of y coordinate from small to large.

[0065] Based on the above description, for this example, Figure 7 As shown, it is assumed that the current scheduler assigns rasterization modules 0 to 2 to perform rasterization processing on the three primitives (i.e., primitive 1, primitive 2, and primitive 3) in the primitive list of Tile-2, and the scheduler assigns the idle rendering core 2 to Tile-2 to remove the occluded fragments from the fragments that need to perform fragment shading processing in the fragments of all primitives in the primitive list of Tile-2.

[0066] Specifically, before the rendering core 2 performs culling processing on the fragments in each FIFO queue, the fragments stored in the FIFO queues of each rasterization module corresponding to Tile-2 are as shown in Table 1:

[0067] Table 1

[0068]

[0069] For this example, the detailed description of the specific processing method is as follows:

[0070] Rendering core 2 performs the first comparison based on Table 1. Rendering core 2 obtains and compares the coordinate value of the first fragment from each FIFO queue. Since the x coordinate of the first fragment of FIFO queue 0 is 0, which is the smallest coordinate value among the three FIFO queues, rendering core 2 outputs the first fragment in FIFO queue 0 of Table 1 to the fragment shader module, and updates the second fragment in FIFO queue 0 after completing the output of the fragment to the first fragment of the corresponding FIFO queue 0. The fragments saved in each FIFO queue after processing are shown in Table 2.

[0071] Table 2

[0072]

[0073] The processing method of the following four comparisons is the same as that of the first comparison, which will not be repeated here. The fragments saved in each FIFO queue after processing are shown in Table 3.

[0074] Table 3

[0075]

[0076] Then, the rendering core 2 performs the next comparison based on Table 3. Since the x-coordinate and y-coordinate of the coordinate values ​​of the first fragment in FIFO queue 0 and FIFO queue 1 are the same, the rendering core 2 continues to compare the depth values ​​of the first fragment in FIFO queue 0 and FIFO queue 1, removes the fragment with a larger depth value (the first fragment in queue 1), outputs the fragment with a smaller coordinate value (the first fragment in queue 0) to the fragment shader module, and updates the second fragment in FIFO queue 0 that has completed the output of the fragment and the second fragment in FIFO queue 1 that has completed the removal of the fragment to the first fragment of the corresponding FIFO queue 0. The fragments stored in each FIFO queue after processing are shown in Table 4.

[0077] Table 4

[0078]

[0079] Next, the rendering core 2 performs the next two comparisons based on Table 4. Similarly, the rendering core 2 outputs the first two fragments of FIFO queue 0 in Table 4 to the fragment shader module, and updates the third fragment in FIFO queue 0 that has completed the output fragment to the first fragment of the corresponding FIFO queue 0. The fragments stored in each FIFO queue after processing are shown in Table 5.

[0080] Table 5

[0081]

[0082] Next, rendering core 2 performs the next comparison based on Table 5. Since the x-coordinate values ​​of the first fragment in each FIFO queue are the same, rendering core 2 continues to compare their y-coordinate values, among which the y-coordinate of FIFO queue 1 is the smallest. Rendering core 2 outputs the first fragment in FIFO queue 1 in Table 5 to the fragment shader module, and updates the second fragment in FIFO queue 1 that has completed the output fragment to the first fragment of the corresponding FIFO queue 1. The fragments saved in each FIFO queue after processing are shown in Table 6.

[0083] Table 6

[0084]

[0085] Next, rendering core 2 performs the next comparison based on Table 6. Since the x-coordinates and y-coordinates of the first fragments from FIFO queue 0 to FIFO queue 2 are equal, rendering core 2 continues to compare the depth values ​​of the three. Since the depth value of the first fragment in FIFO queue 0 is the smallest, rendering core 2 outputs the first fragment in FIFO queue 0 in Table 6 to the fragment shader module, and removes the first fragments in FIFO queue 1 and FIFO queue 2 with larger depth values. At this time, the fragments in FIFO queue 0 are empty (rendering core 2 outputs all fragments of primitive 0 in FIFO queue 0), rendering core 2 stops processing FIFO queue 0, and updates the second fragments in FIFO queue 1 and FIFO queue 2 that have completed the removal of fragments to the first fragment of the corresponding FIFO queue. The fragments saved in each FIFO queue after processing are shown in Table 7.

[0086] Table 7

[0087]

[0088] Next, rendering core 2 continues to process the fragments of primitive 1 in FIFO queue 1 and the fragments of primitive 2 in FIFO queue 2. The specific processing method is the same as above and will not be repeated here.

[0089] To sum up, since the rasterization module always scans the primitives from left to right and from top to bottom, the coordinate value of the first fragment in each FIFO queue is always the smallest among the coordinate values ​​of all fragments of this primitive. Therefore, the rendering core only needs to compare the coordinate value of the first fragment of each FIFO queue each time, thereby reducing the number of comparisons of the rendering core.

[0090] In some examples, the device may also include a scheduler, which is configured to access the primitive list corresponding to the tile currently to be rasterized in sequence according to a set access order, traverse all primitives in the primitive list of the current tile to be processed, and assign each traversed primitive to a currently idle rasterization module in a polling manner to perform rasterization processing.

[0091] In some possible implementations, the scheduler may access the tiles to be processed in order of the tile numbers, for example, accessing the primitive list corresponding to each tile in the order of Tile-0, Tile-1, Tile-2, and Tile-3.

[0092] In some other possible implementations, the scheduler may also access the tiles to be processed in order according to their importance. As for importance, it can be considered that the larger the tile in the primitive list, the higher its corresponding importance. Therefore, the size of the primitive list corresponding to the tile can be used as a preferred measurement indicator for importance. Alternatively, it can be considered that the closer the tile is to the center of the screen, the higher its corresponding importance. Therefore, the distance between the center of the tile and the center of the screen can be used as another preferred measurement indicator for importance. Of course, multiple importance measurement indicators can also be set according to the requirements of the specific application environment, which will not be elaborated in the embodiments of the present invention. In order to briefly explain the technical solution, the embodiments of the present invention only exemplify the access order using the tile numbering order. For example, if Figure 5 As shown, after the primitive list of Tile-0 is accessed, the scheduler will then access the primitive list of Tile-1 and traverse all primitives in the primitive list of Tile-1 (i.e., primitive 1, primitive 2). At this time, the scheduler can assign primitive 1 in the primitive list of Tile-1 to the idle rasterization module 2, assign primitive 2 in the primitive list of Tile-1 to the idle rasterization module 3, and assign Tile-1 to the idle rendering core 1.

[0093] In some examples, the device may also include a scheduler configured to assign fragments corresponding to all primitives in the primitive list corresponding to the current tile to be processed to the same currently idle rendering core to eliminate occluded fragments from the fragments that need to perform fragment shading processing.

[0094] For this example, specifically, Figure 7 As shown, if the current tile to be processed is Tile-2, the scheduler allocates the fragments of all primitives (primitive 1, primitive 2, primitive 3) in the primitive list corresponding to Tile-2 to the same idle rendering core 2 to eliminate the obstructed fragments from the fragments that need to be processed by fragment shading. As an example and not a limitation, if primitive 1 in the primitive list corresponding to Tile-2 completes the rasterization process first, the scheduler can allocate the currently idle rendering core 2 to the fragments corresponding to primitive 1, and also allocate the rendering core 2 to the fragments of primitives 2 and primitive 3 that have completed the rasterization process subsequently, so as to eliminate the obstructed fragments from the fragments that need to be processed by fragment shading corresponding to all primitives of Tile-2.

[0095] Based on the same inventive concept as the above technical solution, see Figure 8 , which shows a method for removing invisible pixels provided by an embodiment of the present invention, which can be applied to the aforementioned Figure 2 or Figure 3 In the illustrated GPU 112, the method may include:

[0096] S801: Perform parallel rasterization processing on each of all the graphic elements covered by the current tile to be processed, and obtain a fragment corresponding to each graphic element.

[0097] S802: Output the fragments of all the primitives as fragments that need to be processed by fragment shading according to the set coordinate order, and remove the blocked fragments from the fragments that need to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process.

[0098] In some examples, outputting the fragments of all the primitives according to the set coordinate order to perform fragment shading processing, and removing the blocked fragments from the fragments that need to perform fragment shading processing based on the coordinate values ​​and depth values ​​of the fragments during the output process, includes:

[0099] Compare the first fragments in the FIFO queues corresponding to all the primitives, output the fragment with the smallest coordinate value as the fragment to be processed by fragment shading, and remove the fragments with the same coordinate value and larger depth value;

[0100] Update the second fragment in the FIFO queue of completed output and / or discarded fragments to the first fragment of the corresponding FIFO queue;

[0101] Based on the updated FIFO queue, the first fragments in the FIFO queues corresponding to all graphics primitives are compared, and the fragment with the smallest output coordinate value is the fragment that needs to be shaded, and the fragments with the same coordinate value and larger depth value are eliminated until the fragments in all FIFO queues are empty.

[0102] In some examples, the method further includes:

[0103] The scheduler accesses the primitive list corresponding to the tile to be rasterized in sequence according to the set access order, traverses all primitives in the primitive list of the tile to be processed, and allocates each traversed primitive to the currently idle rasterization module in a polling manner to perform rasterization processing.

[0104] In some examples, the method further includes:

[0105] The scheduler allocates the fragments corresponding to all the primitives in the primitive list corresponding to the current tile to be processed to the same currently idle rendering core to eliminate the blocked fragments from the fragments that need to be subjected to fragment shading processing.

[0106] It can be seen that by adopting the method described in the embodiment of the present application, the serial rasterization processing of the Tile's primitive list is changed to parallel rasterization processing. At the same time, the position value of the fragment stored in the FIFO queue of the rasterization module is changed from vertical comparison to horizontal comparison, and then the primitives of the Tile are culled, so that the timing of fragment comparison is advanced, the number of comparisons is reduced, and each fragment will be compared without omission. Therefore, the method described in the embodiment of the present application has a better pixel culling effect and is more efficient.

[0107] In the above one or more instances or examples, the described functions may be implemented in, and the described functions may be implemented in hardware, software, firmware or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium. Computer-readable media may include computer data storage media or communication media, and communication media include any media that facilitates the transfer of computer programs from one place to another. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementing the technology described in the present invention. For example and without limitation, such computer-readable media may include a USB flash drive, a mobile hard disk, a RAM, a ROM, an EEPROM, a CD-ROM or other optical disk storage device, a magnetic disk storage device or other magnetic storage device, or any other media that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, server or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio and microwave are included in the definition of media. As used herein, disk and optical disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks usually reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0108] The code may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs) or other equivalent programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Therefore, the terms "processor" and "processing unit" as used herein may refer to any of the aforementioned structures or any other structures suitable for implementing the techniques described herein. In addition, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Moreover, the techniques may be fully implemented in one or more circuits or logic elements.

[0109] The techniques of the embodiments of the present invention may be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (i.e., chipsets). The various components, modules, or units described in this disclosure are intended to emphasize the functional aspects of the apparatus configured to perform the disclosed techniques, but do not necessarily need to be implemented by different hardware units. In fact, as described above, the various units may be combined in a codec hardware unit in conjunction with appropriate software and / or firmware, or provided by a collection of interoperable hardware units, including one or more processors as described above.

[0110] Various aspects of the present invention have been described. These and other embodiments are within the scope of the appended claims. It should be noted that the technical solutions described in the embodiments of the present invention can be combined arbitrarily without conflict.

[0111] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A device for removing invisible pixels, characterized in that: include: at least one rendering core and at least one rasterization module; Each of the at least one rasterization module is configured to perform parallel rasterization processing on each of all the primitives covering the current tile to be processed, to obtain a fragment corresponding to each primitive; Each of the at least one rendering core is configured to output the fragments of all the primitives according to a set coordinate order as the fragments to be processed by fragment shading, and to remove occluded fragments from the fragments to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process; Each of the at least one rendering core is further configured to compare the first fragments in the FIFO queues corresponding to all the primitives, output the fragment with the smallest coordinate value as the fragment to be processed by fragment shading, and remove the fragments with the same coordinate value and larger depth value; Update the second fragment in the FIFO queue of completed output and / or discarded fragments to the first fragment of the corresponding FIFO queue; Based on the updated FIFO queue, the first fragments in the FIFO queues corresponding to all graphics primitives are compared, and the fragment with the smallest output coordinate value is the fragment that needs to be shaded, and the fragments with the same coordinate value and larger depth value are eliminated until the fragments in all FIFO queues are empty.

2. The device according to claim 1, characterized in that The device also includes a scheduler, which is configured to access the primitive list corresponding to the tile currently to be rasterized in sequence according to a set access order, traverse all primitives in the primitive list of the current tile to be processed, and allocate each traversed primitive to a currently idle rasterization module in a polling manner to perform rasterization processing.

3. The device according to claim 1, characterized in that The device also includes a scheduler, which is configured to allocate fragments corresponding to all primitives in the primitive list corresponding to the current tile to be processed to the same currently idle rendering core to eliminate blocked fragments from the fragments that need to perform fragment shading processing.

4. A method for removing invisible pixels, characterized in that: The method comprises: Perform parallel rasterization processing on each of the graphic elements covered by the current tile to be processed, and obtain the fragment corresponding to each graphic element; Output the fragments of all the primitives according to the set coordinate order to obtain the fragments that need to be shaded, and remove the blocked fragments from the fragments that need to be shaded based on the coordinate values ​​and depth values ​​of the fragments during the output process; The step of outputting the fragments of all the primitives according to the set coordinate order to obtain the fragments to be processed by fragment shading, and removing the blocked fragments from the fragments to be processed by fragment shading based on the coordinate values ​​and depth values ​​of the fragments during the output process, includes: Compare the first fragments in the FIFO queues corresponding to all the primitives, output the fragment with the smallest coordinate value as the fragment to be processed by fragment shading, and remove the fragments with the same coordinate value and larger depth value; Update the second fragment in the FIFO queue of completed output and / or discarded fragments to the first fragment of the corresponding FIFO queue; Based on the updated FIFO queue, the first fragments in the FIFO queues corresponding to all graphics primitives are compared, and the fragment with the smallest output coordinate value is the fragment that needs to be shaded, and the fragments with the same coordinate value and larger depth value are eliminated until the fragments in all FIFO queues are empty.

5. The method according to claim 4, characterized in that The method further comprises: The scheduler accesses the primitive list corresponding to the tile to be rasterized in sequence according to the set access order, traverses all primitives in the primitive list of the tile to be processed, and allocates each traversed primitive to the currently idle rasterization module in a polling manner to perform rasterization processing.

6. The method according to claim 4, characterized in that The method further comprises: The scheduler allocates the fragments corresponding to all the primitives in the primitive list corresponding to the current tile to be processed to the same currently idle rendering core to eliminate the blocked fragments from the fragments that need to be subjected to fragment shading processing.

7. A graphics processor GPU, characterized in that: include: The invisible pixel removal device according to any one of claims 1 to 3.

8. A computer storage medium storing a program for eliminating invisible pixels, wherein the program for eliminating invisible pixels, when executed by at least one processor, implements the steps of the method for eliminating invisible pixels according to any one of claims 4 to 6.

Citation Information

Patent Citations

  • Graphics Processors

    GB202103473D0