Merging unit, method for selecting a coverage merging scheme, and depth testing system

By introducing merge units in the GPU for rough depth removal, the problem of resource waste and over-drawing of element vertices by hidden surface removal methods in the prior art is solved, and more efficient graphics processing is achieved.

CN112116518BActive Publication Date: 2025-07-08SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010547067.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-18
Filing Date
2020-06-16
Publication Date
2025-07-08
Estimated Expiration
2040-06-16

AI Technical Summary

Technical Problem

Existing graphics processors (GPUs) cannot effectively remove hidden primitive vertices and primitives in hidden surface removal (HSR) methods, resulting in waste of resources and energy, and conventional merging methods do not consider visibility, resulting in over-drawing and repeated shading.

Method used

The merging unit is used for rough depth removal, pixel coverage information and depth information are generated through the rasterizer, and visibility-based culling is performed in the local and global culling stages to reduce unnecessary primitive processing.

Benefits of technology

Reduces the workload of GPU front-end and back-end processing, reduces pixel shader calls and memory bandwidth consumption, improves rendering efficiency, and reduces over-drawing and repeated shading.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112116518B_ABST
    Figure CN112116518B_ABST
Patent Text Reader

Abstract

A merging unit, a method for selecting a coverage merging scheme, and a depth testing system are provided. An inventive aspect includes a merging unit for rough depth culling during the merging of pixel graphics. The merging unit includes a rasterizer that is configured to receive primitives and generate pixel coverage information and depth information. The merging unit includes one or more local culling stages that are configured to perform local culling within the window of a primitive. The local culling unit outputs a set of surviving coverage information and surviving depth information. The merging unit includes one or more global culling stages that are configured to perform further culling based on an overall use of previously received coverage information and depth information using the set of surviving coverage information and surviving depth information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims the benefit of U.S. Provisional Patent Application No. 62 / 864,443, filed on Jun. 20, 2019, the content of which is incorporated herein in its entirety. Technical Field

[0002] This embodiment relates to a graphics processing unit (GPU), and more particularly, to systems and methods for coarse depth cull during binning. Background Art

[0003] A GPU is a dedicated device that accelerates the processing of computer-generated graphics. GPUs are also used in a variety of modern computing environments such as neural networks, artificial intelligence (AI), high-performance systems, autonomous vehicles, mobile devices, gaming systems, etc.

[0004] Hidden surface removal (HSR) methods represent removing surfaces that are hidden or occluded from a camera by other surfaces that are closer to the camera so that the surface is not processed. Desktop GPUs maintain a depth buffer that can cull quads (i.e., 2×2 pixel blocks), where the depth of the quad indicates that the quad is occluded by other processed quads. The effectiveness of this solution depends on the degree to which the surfaces are sorted from front to back.

[0005] Existing HSR methods mainly target removing hidden quads and do not target removing the constituent vertices and primitives of hidden surfaces. Mobile GPUs can generate all output attributes (typically vertex shaders) of the front-end pipeline and read back the attributes. Considerable resources and energy are spent processing mostly fully occluded primitives and their vertices, which mostly do not end up resulting in any visible quads. GPUs typically have a limited ability to cull quads that will ultimately be occluded by later quads. A conventional method involves buffering quads before pixel shading to identify later quads that occlude earlier quads in the buffer. However, such methods are limited by the size of the buffer that is practically cost-effective.

[0006] Most tile-based deferred rendering (TBDR) GPUs run a front-end stage once per primitive per image and cache the results in an intermediate buffer, reading from the intermediate buffer once per tile to run the fragment / pixel stage. Some of these TBDR GPUs may use a similar approach for HSR. Tile-based GPUs have a merge step in which geometries are sorted by the tiles of the pixels they affect. A tile is a rectangular block of pixels. The merge unit (sometimes called a tiler) creates a list of the draws and primitives projected onto the pixels of each tile. A primitive is a geometry in a coordinate system (usually, a triangle). A tile is a group of pixels. The merge unit allows rendering to operate on a per-tile basis in the case where only those primitives affecting the tile are processed. Conventional merging is only spatial sorting and does not consider visibility. In other words, primitives occluded by other primitives within a tile are not excluded.

[0007] The lack of visibility results in overdraw or duplicate shading of some pixels in the image. Using visibility culling, the amount of duplicate pixel shading can be reduced and the corresponding pixel shader calls can also be saved. Summary of the Invention

[0008] Some embodiments include a merge unit for coarse depth culling during merging of pixel graphics. The merge unit includes a rasterizer for receiving primitives and generating pixel coverage information and depth information. The merge unit includes one or more local culling stages for performing local culling within the window of the primitive. The local culling unit outputs a set of surviving coverage information and surviving depth information. The merge unit includes one or more global culling stages for further culling based on an overall use of the previously received coverage information and depth information using the set of surviving coverage information and surviving depth information. Brief Description of the Drawings

[0009] The foregoing and additional features and advantages of the principles of the present invention will become more readily apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0010] Figure 1 is an exemplary diagram of a merge unit according to some embodiments.

[0011] Figure 2 is an exemplary diagram illustrating a hidden surface removal (HSR) technique.

[0012] Figure 3An example illustration of {primitive, tile} culling and quadrilateral culling according to some embodiments.

[0013] Figure 4 An example illustration of a depth and coverage structure stored in a memory according to some embodiments.

[0014] Figure 5 An example illustration of a situation for implementing depth and coverage merging using a local culling stage according to some embodiments.

[0015] Figure 6 Is an illustration including a legend 600 for various blocks shown in Figure 5 An illustration of a depth testing module according to some embodiments.

[0016] Figure 7 An example illustration of a depth testing module according to some embodiments.

[0017] Figure 8 Is according to some embodiments of Figure 7 An example illustration of a set tester of a depth testing module.

[0018] Figure 9 An example illustration of a depth update logic section according to some embodiments.

[0019] Figure 10 Is an example block diagram of a GPU including a merging unit according to some embodiments as disclosed herein. Figure 1 According to some embodiments as disclosed herein. Detailed Description

[0020] Now, reference will be made in detail to the embodiments disclosed herein, examples of which are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the inventive concept. However, it should be understood that those of ordinary skill in the art may practice the inventive concept without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0021] It will be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the inventive concept, a first primitive may be referred to as a second primitive, and similarly, a second primitive may be referred to as a first primitive.

[0022] The terms used in the description of the inventive concept herein are for the purpose of describing particular embodiments only and are not intended to limit the inventive concept. As used in the description of the inventive concept and the appended claims, the singular forms are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein represents and includes any and all possible combinations of one or more of the associated listed items. It will also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The components and features of the drawings are not necessarily drawn to scale.

[0023] Some embodiments include an enhanced binning unit that includes the ability to cull draws and primitives from each tile list based on visibility. The binning unit disclosed herein can create a rough approximation of the final depth representation at the granularity of pixels (or groups of pixels). The binning unit can also minimize the memory bandwidth consumed during the binning process. The binning unit can reduce work by culling primitives and draw calls from the processing performed in the tile channels. The binning unit can improve the culling performance of existing "Early-Z" hardware by preloading a rough depth representation, resulting in fewer pixels and / or fragments being shaded. "Early-Z" is a form of depth processing performed per-pixel shading.

[0024] For each primitive, the binning unit disclosed herein can rasterize the primitive at the necessary granularity (e.g., samples or pixels). When there is only one sample per pixel, the sample is equivalent to the pixel. Although not required to have only one sample per pixel, the terms "sample" and "pixel" are generally used interchangeably herein. The binning unit can compute the depth range of each primitive for each pixel block for a predefined block size. The binning unit can use this per-primitive {coverage, depth range} information to maintain an intermediate {coverage, depth range} representation of the image, culminating in a final {coverage, depth range} representation. The binning unit can maintain the {coverage, depth range} representation in a rough, compressed manner. The binning unit can use the intermediate {coverage, depth range} representation to cull primitives from one or more tiles.

[0025] In some embodiments, the merge unit may maintain the {coverage range, depth range} representation as a hierarchy. In some embodiments, the hierarchy may be maintained in a hardware circuit. Subsequent steps in the hierarchy may use the same coverage range granularity or coarsen. Each step may maintain a {coverage range, depth range} representation of one or more primitives, a window of primitives, or a subset of all primitives seen so far. Some steps may maintain this {coverage range, depth range} representation only on-chip (e.g., using a hardware circuit), while other steps may have an on-chip cache supported by memory.

[0026] An example hierarchy may include a first step that maintains a {coverage range, depth range} representation of a window of primitives on-chip, where the coverage range is maintained at a sample / pixel granularity for anti-aliased / aliased rendering, respectively. A second step may maintain a {coverage range, depth range} representation of all primitives seen so far, excluding as necessary, in memory with an on-chip cache. The coverage range may be maintained at a pixel or pixel-block granularity for anti-aliased or aliased rendering, respectively. In some embodiments, the pixel-block is a quad, i.e., a 2×2 pixel block.

[0027] The first step of the example hierarchy may cull incoming primitives based on the depth from the current window of the primitive. In some embodiments, the first step may cull the entire current window of the primitive based on the depth from the incoming primitive. The second step of the example hierarchy may cull incoming primitives or a window of primitives based on the depth from a previous primitive. In some embodiments, the second step may cull all previous primitives based on the depth from the incoming primitive or a window of primitives.

[0028] Disclosed herein is a rough depth-based hidden surface removal technique operating in a merge channel. The rough depth-based hidden surface removal technique can generate a compressed count stream representation for indicating which primitives and drawing calls affect a particular tile, and this representation does not need to include most (e.g., greater than a predetermined ratio) of the primitives and drawing calls that are invisible in the final rendered image. As a non-limiting example, the predetermined ratio can be any value between 50% and 100% (e.g., 60%, 70%, etc.). The disclosed technique can also generate an approximately compressed depth and coverage representation of the image, and this approximately compressed depth and coverage representation will be used as a preloaded depth buffer to increase pixel culling through existing depth culling hardware mandated by a graphics application programming interface (or application programming interface, API). For each pixel block, a rough coverage mask can be created by combining a dictionary of depth ranges at the granularity of pixels or pixel blocks. In some embodiments, it can be guaranteed that each covered entity (i.e., pixel or pixel block) has a depth value within a specific depth range in the dictionary.

[0029] The disclosed merge unit can cull primitives in the merge channel, thereby reducing the number of primitives processed during the color channel. This technique can reduce the per-tile processing of primitives in the front-end pipeline of the GPU during the color channel. When running the merge channel with a reduced shader that only produces vertex and primitive position information, this technique can also reduce the overall front-end shading cost. If all primitives within a drawing are culled due to depth considerations, the merge unit can cull the drawing, thereby reducing the overhead and performance impact of state management. The merge unit can use the rough depth coverage representation to cull pixels and pixel quads in the color channel, thereby generally reducing the amount of pixel shader calls and pixel processing cost.

[0030] In some embodiments, the merge unit uses a hierarchy of {coverage, depth} representations, which can be stored in a hardware circuit (such as a cache). In some embodiments, the merge unit uses the depth from earlier primitives to cull later primitives. In some embodiments, the merge unit uses the depth from later primitives to cull the range of earlier primitives.

[0031] Some embodiments described herein include a coarse visibility culling architecture for efficient 3D rendering on a tile-based deferred rendering (TBDR) GPU. At least two inefficiencies in conventional TBDR GPUs are addressed: 1) unnecessary overdraw and 2) processing unnecessary primitives during merge or tiling below the render. The methods and systems described herein add a binner or tiler that uses a coarse visibility culling step to determine a list of primitives and draw calls that affect a particular tile to minimize the amount of overdraw.

[0032] The methods and systems described herein include an added merge or tiler unit (commonly referred to herein as the "merge unit") that, in addition to generating a list of draws and primitives projected onto each tile, culls draws and primitives from such a list in cases where all fragments generated by the draw / primitive are occluded by an earlier draw / primitive. Further, the methods and systems described herein minimize overdraw by creating a coarse representation of depth at each pixel of the image during merge and preloading that representation into the depth buffer such that Early-Z hardware eliminates fragments that will be occluded by later fragments.

[0033] By using the methods and systems described herein, the GPU can minimize the amount of wasted work (i.e., wasted work in processing primitives in the front-end pipeline including vertex shaders and later shaders, and wasted work in processing pixel quads in the back-end pipeline including pixel shaders).

[0034] By processing transformed primitives during the merge pass, the methods and systems described herein create an intermediate representation of the visible depth range in the image upon receipt of each primitive, culminating in a final depth representation that can be preloaded during image rendering in the color pass. Additionally, some embodiments maintain the intermediate depth representation in a coarse, compressed representation to reduce the memory footprint of the intermediate depth representation. Further, some embodiments maintain the depth representation as multiple sets with a per-pixel or per-quad selector to select a depth set to ensure good depth resolution even when multiple surfaces are active in a block. Additionally, some embodiments use the intermediate depth representation to cull entire primitives during the merge pass itself. Further, the final depth representation can be preloaded as a starting depth buffer during the color pass to cull individual pixels and quads. The methods and systems described herein can use alternative but complementary methods that are not limited to identifying such quads within a limited window determined by a cost-effective buffer size. Instead, the methods and systems described herein can generate a coarse depth buffer during the merge pass in use.

[0035] Some embodiments disclosed herein include a rough-depth-based hidden surface removal method operating in a merge channel. The merge channel can generate a compressed count stream representation for indicating which primitives and drawing calls affect a particular tile, and excludes most primitives and drawing calls that are not visible in the final rendered image. The merge channel can generate an approximately compressed depth and coverage representation for the image, and the approximately compressed depth and coverage representation will be used as a preloaded depth buffer to increase pixel culling through existing depth culling hardware notified by the graphics API. Some embodiments can create a rough coverage mask by combining a dictionary of depth ranges at the granularity of pixels or pixel blocks for each pixel block. In some embodiments, each covered entity (pixel or pixel block) is guaranteed to have a depth value within a specific depth range in the dictionary. Some embodiments use the rough depth coverage representation in the merge channel to cull primitives, thereby reducing the number of primitives processed during the color channel. In turn, this can reduce the per-tile processing of primitives by the front-end pipeline in the color channel, and the overall front-end shading cost when running the merge channel with a reduced shader that only produces the position information of vertices and primitives.

[0036] If all primitives within a drawing are culled due to depth considerations, some embodiments cull the drawing, reducing the overhead and performance impact of state management. Some embodiments use the rough depth coverage representation to cull pixels and pixel quads in the color channel, thereby generally reducing the amount of pixel shader calls and pixel processing costs.

[0037] Some advantages of the merge unit disclosed herein are: The merge unit does not rely on the application to sort the geometry from front to back, and even when the geometry is submitted from back to front, the merge unit can successfully cull most occluded quads. Another advantage is: The merge unit described herein does not require a large buffer to hold the quads and is not affected by the latency of holding the quads to ensure culling. Yet another advantage is: Due to culling primitives and quads based on an approximate depth buffer, there are fewer pixel shader calls during rendering any image with significant depth complexity. Another advantage is: Due to culling primitives whose quads are fully occluded, the front-end shading work and associated vertex shading, setup, and rasterization are reduced. Another advantage is: Due to culling certain drawings that do not contribute to any visible quads, the state management overhead is reduced. Auxiliary benefits include fewer shader calls, and fewer shader calls include reduced memory bandwidth for textures, vertex attributes, and associated fixed-function processing. Additionally, a reduced overdraw rate during rendering is achieved, resulting in reduced wasted pixel shader work.

[0038] Figure 1FIG. 0 is an exemplary illustration of a merge unit 100 according to some embodiments. The merge unit 100 may perform some approximate but conservative HSR during merge passes to avoid the cost and complexity of running additional passes for the same operation. Thus, the additional complexity of performing this technique is confined to the merge unit 100 itself. The merge unit 100 may perform merge work in a GPU, obtain a stream of primitives and draw calls within an image, and produce a compressed count stream, specifically, one compressed count stream per tile per entity, where an entity is a primitive or a draw call. The merge may be performed at the granularity of a single tile or alternatively at the granularity of multiple tiles. The result of the merge is to produce a compressed count stream for all merged tiles in the image. The merge unit 100 may perform two types of HSR: 1) {primitive, tile} culling and 2) quad culling.

[0039] {Primitive, tile} culling involves removing primitives from the compressed count stream for a particular tile, which saves work for both front-end and back-end processing. The abbreviated term "prim" as used herein refers to one or more primitives. A tile is a rectangular block of pixels that is rendered by the GPU as a single transaction. The compressed count stream records whether an entity affects the rendering result of a tile, where an entity can be a primitive, a draw call, or something else.

[0040] Quad culling may generate an approximate z-buffer for preloading. This saves pixel shader calls for quads during the color pass. A quad is a 2×2 rectangular block of pixels that is rendered together to allow access to a texture. The disclosed technique handles cases where a quad is occluded by a later quad and thus does not rely on front-to-back sorting for occlusion.

[0041] At a higher level, the merge unit 100 may merge coverage and depth across primitives, cull {primitives, tiles} with that structure while the structure is being generated, and store the coarse depth in memory. The merge unit 100 may include a rasterizer 105 that is capable of generating coverage information at sample granularity and interpolating depth at samples within the coarse extent. The rasterizer 105 may include a first-stage coarse rasterizer 110 that may receive primitives and / or vertex data 102 and may compute coverage at the granularity of pixel blocks. This first stage may be augmented with depth interpolation logic 120 that may compute depth ranges at the corners of pixel blocks. The coarse rasterizer 110 may output intermediate rasterizer information 122 that may include values and edge equations at 2x2 tile corners or tiles and depth information at that granularity. The 2x2 tile corners or tiles may be reordered rather than running the tiles. The following discussion: This maximizes the locality of stream accumulator entries (SA entries) 135. The rasterizer 105 may also include a second-stage fine rasterizer 115 that may receive the intermediate rasterizer information 122 and compute coverage at the granularity of pixels. The coarse rasterizer 110 may compute rasterization information and depth at the granularity of blocks. The fine rasterizer 115 may compute pixel / sample coverage given the coarse rasterization result (or intermediate rasterizer information) 122 from the coarse rasterizer 110. The rasterizer 105 may output {primitive, tile, block} information 125 with depth and pixel coverage.

[0042] One or more local culling stages 130 of the merge unit 100 may perform coverage- and depth-based culling. The local culling stages 130 may perform local culling of operations within the window of primitives and draw calls using a fine-grained coverage granularity without any backing state. This stage operates on the window of primitives and draw calls within a tile and uses only the depth from those primitives to cull primitives within the window. This culling may use later primitives to cull earlier primitives or vice versa (i.e., this stage may cull in-order forward or cull in-order backward). The local culling stages 130 may include a plurality of stream accumulator entries (SA entries) 135, one or more accumulators 140, and flush control logic 145. The SA entries 135 may create an OR'ed coverage mask and maintain the depth range for each block.

[0043] The merge unit 100 can operate on pixel blocks smaller than tiles herein referred to as "blocks". {Coverage, depth range} can be referred to herein as "nodes". Nodes can define the pixel / quadrilateral coverage within a block and the depth range into which the pixels / quadrilaterals fall. The size of the block and the size of the depth dictionary can be selected at design time to minimize hardware costs. Other embodiments can choose to dynamically change the block size and depth dictionary size.

[0044] The local culling stage 130 can operate on the state local to the most recent window of primitives and is capable of culling past and current primitives. Thus, the local culling stage 130 can use the coverage and depth information from the past K primitives to cull some or all of the past K primitives within a block or to cull the current primitive. The local culling stage 130 does not need to have any knowledge of any primitives outside of that window. The size of the window can define the on-chip hardware cost and can be selected at design time. Other embodiments can choose other sizes or dynamic sizes.

[0045] The merge unit 100 can include one or more global culling stages or logic 150, which can be incorporated into at least one of one or more local culling stages 130 and rasterizers (e.g., coarse rasterizer 110 and / or fine rasterizer 115) and update the output 155 of the local culling stage 130. For example, the global culling stage 150 can cull the windows of primitives from the first-stage local culling stage 130 and use those primitives to cull the entirety of a previously seen coverage using the input coverage / depth information (i.e., output 155) from the first stage (i.e., 130), or vice versa. The global culling stage 150 can include optional extensions to improve the culling behavior. For example, the global culling stage 150 can implement context-dependent culling behavior to handle special culling behavior for performing inside-outside tests for specific geometries (such as cones or spheres) in 3D space, where a triangle can be culled if all pixels within the coverage of the triangle that is part of a cone are on one side (e.g., the side along the normal of the triangle). If the image preloading is created as a depth buffer from the output of another image, the global culling stage 150 can be used as the starting point for that subsequent image to improve culling performance. Thus, one or more custom extensions for workload-specific culling can be used, which does not need to be visibility culling or hidden surface removal. The global culling stage or local culling stage can use one or more custom extensions.

[0046] The global culling logic 150 may include a depth test module 705, which is described in detail below. In some embodiments, the global culling stage 150 includes optional components incorporated from existing merge / tile logic. For example, the global culling stage 150 may include a reorder queue 160, which may prioritize transactions for which backup data resides in on-chip memory (e.g., in on-chip buffer 165). In some embodiments, the global culling stage 150 includes merge logic 182, which may create a stream of draw calls and primitives for the coverage that will be consumed by subsequent stages rendered by the GPU. Memory for such a stream may be provided by an allocator unit 170, and data may be written into the stream by a merge logic section 175. The merge logic section 175 may be implemented on-chip. The merge logic section 175 may receive a count write request 180 from the local culling stage 130 and update the compressed count stream using the memory allocated by the allocator unit 170. In some embodiments, the global culling stage 150 includes wide and narrow on-chip networks (NOCs) 185 for communicating with the system memory cache hierarchy and / or memory subsystem (not shown).

[0047] The on-chip buffer 165 may include a prefetch queue 162, descriptor (DES) data 164, compressed count stream data or bitstream data 166, and global culling data 168 (such as rough depth information). The prefetch queue 162 may include a latency first-in first-out (FIFO) that ensures maximum utilization of the on-chip buffer 165. In other words, those transactions with on-chip data may be given a higher priority than other transactions that need to fetch data from the memory subsystem. The on-chip buffer 165 may be combined with a level 2 (L2) cache 190. The global culling data 168 may include a depth update logic section 905, which is described in detail below.

[0048] The global culling stage 150 may use rough depth information and / or fine depth information and coverage information from some or all of the primitives in past primitives to cull the current set of primitives obtained from the local culling stage 130. The global culling stage 150 is capable of culling both past and current sets of primitives.

[0049] Figure 2Is an example illustration 200 showing hidden surface removal (HSR) techniques. HSR reduces the time and resources for rendering primitives, assuming that all primitives under discussion are opaque and will ultimately be invisible. Most modern GPUs incorporate some hidden surface removal techniques. As shown in stage 210, the render queue 202 holds primitives 0, 1, 2, and 3, and the screen 205 is initially blank. At stage 215, primitives 0 and 1 are shown on the screen 205, and primitives 2 and 3 remain in the render queue 202. At stage 220, the render queue 202 is empty, and primitives 2 and 3 are occluded by primitives 0 and 1 on the screen 205. In other words, primitives 2 and 3 have a greater depth than primitives 0 and 1, and primitives 0 and 1 have a closer depth. Therefore, the surfaces of primitives 2 and 3 can be removed to reduce the time and resources for rendering those primitives.

[0050] Figure 3 Is an example illustration 300 of {primitive, tile} culling and quad culling according to some embodiments. During the merge pass, later primitives can be completely hidden by depth information from earlier primitives. The merge unit (e.g., Figure 1 100) can collect this information in a coarse manner and use this information within the merge pass to cull primitives from the tile as a whole. This culling can be represented within the compressed count stream itself, meaning that in the color pass, both front-end (VS, vertex shader) processing and back-end (PS, pixel shader) processing can be saved.

[0051] A second form of culling that can be performed by the merge unit 100 involves providing a coarse depth representation of the image to increase the effectiveness of Early-Z culling. Thus, a final or near-final version of the depth buffer can be created. The final or near-final version can be pre-loaded before running the full color pass. This form of culling saves back-end (PS) work but incurs a penalty for running the front-end (VS) for these primitives.

[0052] As Figure 3 shown, the tile 305 can consist of 16×16 pixel blocks (e.g., 310). The tile 305 can have other sizes (such as 32×16, 32×32, 64×32, 64×64, etc.). It will be understood that other suitable tile sizes can be used. As shown at stage 330, primitives 0 and 1 can be processed. The merge unit (e.g., Figure 1At stage 335, the merge unit 100 can rasterize primitive 0 and primitive 1, and accumulate the rough coverage information and depth information. The depth information can be a depth range between a predefined minimum value and a maximum value. At stage 335, the merge unit 100 can check each subsequent primitive (e.g., primitive 2 and primitive 3) for the rough coverage information and depth information. The merge unit 100 can reject primitive 2 and primitive 3 from the tile 305. This rejection can be recorded in the compressed count stream. In other words, the entire primitive 2 and primitive 3 can be culled. {Primitive, tile} culling occurs during the merge pass as shown at 315, which is one of the benefits 320.

[0053] At stage 340, the merge unit 100 can write the rough coverage information and depth information to memory. The merge unit 100 can preload the rough coverage information and depth information into the tile buffer 350 during the color pass. The tile buffer 350 is sometimes referred to as the depth buffer here. The tile buffer 350 can hold all color and depth (Z) information of the tile during the color pass. Preloading the depth buffer before the start of the color pass allows the GPU to use this depth buffer for Early-Z culling, which tests opaque objects to see if they are visible in the final image. At stage 345, the existing Early-Z logic in the tile buffer 350 can reject additional pixels or quads during the color pass. For example, multiple pixels or quads of primitive K can be Early-Z culled due to the depth information. This stage is called quad culling 325 during the color pass, which is one of the benefits 320. Two bolded pixels / quads 355 are invalidated, so primitive K can lose some pixels, thus saving pixel shading work. Three pixels / quads 360 shown in dashed lines are passed and will be rendered.

[0054] The merge unit 100 can operate in different modes. For example, the merge unit 100 can operate in a mode where the local culling stage 130 and the global culling stage 150 are enabled. In another mode, the local culling stage 130 and the global culling stage 150 can be disabled, but full rasterization can still be performed. In yet another mode, the local culling stage 130 and the global culling stage 150 can be enabled, and full rasterization can be performed. The depth used when preloading the depth buffer into the tile buffer 350 can be determined based on a predefined minimum depth and a predefined maximum depth. For example, the minimum depth can be set to 0, and the maximum depth can be set to 1. As another example, the minimum depth can be set to 0.3, the maximum depth can be set to 0.6, and anything outside this range is invisible. As yet another example, a depth range from 0.5 to 0.6 will make the processing even cheaper. The number of samples per pixel can also be predefined or set.

[0055] The merge unit 100 can internally maintain coverage at the pixel granularity, but store the coverage in the memory at the quadrilateral granularity. This can be done to reduce memory footprint. Due to the coarsening of the coverage to the quadrilateral granularity during storage, some quadrilaterals of the partial coverage may be missed. Therefore, the merge unit 100 reduces the probability of the occurrence of quadrilaterals of the partial coverage by reordering {primitives, tiles} to increase the locality of the coverage to the tiles.

[0056] Since each comparison will incur a non-negligible energy and throughput cost, efforts are made to reduce the number of deep comparisons. As a result, the depth test can be performed on clusters of primitives rather than on each primitive. Since efforts have been made to assemble complete quadrilaterals before testing, depth updates can also be implicitly performed at the cluster level. Significant efforts are made to reduce the number of depth updates going from the on-chip buffer (e.g., Figure 1 165) to the memory (e.g., Figure 1 190). Similar efforts are made to reduce the per-tile occupancy of the rough depth data of the merge unit 100 in order to minimize the increase in memory traffic.

[0057] Figure 4 FIG. 400 is an exemplary diagram of a node depth and coverage structure as stored in the memory according to some embodiments. Although the merge unit 100 can internally maintain depth information and coverage information in different formats, when written to the memory, the information can be arranged in the format as shown in Figure 4 A node may include a pad (e.g., 4 bytes), such that each node is 32 bytes in total. The depth information can be arranged in the upper 16 bytes (with 4 bytes of padding), and the coverage information can be arranged in the lower 16 bytes. Figure 4 The nodes shown in

[0058] are not necessarily drawn to scale. Multiple nodes can be arranged continuously in the memory with no empty space between any two nodes.

[0059] / / / Enable depth test and update if the state is enabled

[0060] bool DepthModeEnable = (State.Mode == ENABLE_FULLRAS_ENABLE);

[0061] / / / Control variable

[0062] / / / Whether the depth test is enabled

[0063] bool depthTestEnable = DepthModeEnable;

[0064] / / / If depth testing is disabled (i.e., depthTestEnable == false), then what does depth test resolve to.

[0065] / / / true: Always pass, false: Always discard

[0066] bool alwaysPass = true;

[0067] / / / Whether depth update is enabled, i.e., whether depth can be updated

[0068] bool depthUpdateEnable = DepthModeEnable;

[0069] / / / Step 1, controlled by State.DepthTestMode

[0070] / / / Symbolic overloading, means true if LHS equals *any* of the values in {}

[0071] depthTestEnable &= (State.DepthTestMode == {EARLYZ, LATEZ_WITH_EARLYZ_COMPARE});

[0072] depthUpdateEnable &= (State.DepthTestMode == EARLYZ);

[0073] / / / Auxiliary step, changes whether we always pass or not

[0074] alwaysPass = (State.DepthFunc != NEVER);

[0075] / / / Step 2, controlled by other state fields

[0076] / / / If any state indicates that the PS can modify the coverage, then we cannot update the depth

[0077] / / / Although we can still perform depth testing

[0078] / / / Symbolic warning:! represents boolean negation

[0079] depthUpdateEnable &= (!State.PSUsesDiscard &&!State.PSWritesCoverage &&!State.SampleAlphaToCoverage);

[0080] / / / If any state indicates that the PS writes the Z value, we cannot rely on interpolated Z,

[0081] / / / and if Z writing is disabled, we cannot update

[0082] depthTestEnable &= (!State.PSWritesZ);

[0083] depthUpdateEnable &= (!State.PSWritesZ && State.DepthWriteEnable);

[0084] / / / If any stencil test is enabled, we cannot guarantee that the coverage will ultimately be the same as the rasterized coverage.

[0085] But we can still use depth testing for culling

[0086] depthUpdateEnable &= (!State.StencilTestEnable);

[0087] / / / If blending is enabled, we can "see through" the object

[0088] / / / The DSA / local culling stage must enable depth testing iff (depthTestEnable && (State.BlendEnable == 0))

[0089] / / / The merge enables depth testing iff (depthTestEnable)

[0090] / / depthTestEnable &= (State.BlendEnable == 0);

[0091] depthUpdateEnable &= (State.BlendEnable == 0);

[0092] / / / Step 3, coverage by driver control

[0093] depthTestEnable &= (!State.SkipDepthTest);

[0094] depthUpdateEnable &= (!State.SkipDepthUpdate);

[0095] / / / Completed

[0096] Figure 5 is an example diagram 500 of a situation for implementing depth and coverage merging using a local culling stage 130 of ([[]] Figure 1 ). Figure 6 is a diagram including a legend 600 for various blocks shown in [[[]] Figure 5 . Now refer to [[[]] Figure 1 , [[[]] Figure 5 and [[[]] Figure 6 .

[0097] Several shorthand notations are used in [[[]] Figure 5 . For example, some "EX"isting coverage and depth from one or more primitives in an SA entry (e.g., 135) are affected by "IN"coming coverage and depth from one primitive. In other words, "EX" is a shorthand for existing coverage and / or depth, and "IN" is a shorthand for incoming coverage and / or depth. Multiple equivalent shorthand notations are used to describe the merge categories 505. For example, X == Y means that X and Y cover exactly the same pixels / quadrilaterals. X > Y means that the coverage of X is a strict superset of the coverage of Y (i.e., X covers all the pixels / quadrilaterals of Y and some additional pixels / quadrilaterals). X < Y means that the coverage of Y is a strict superset of the coverage of X (i.e., Y covers all the pixels / quadrilaterals of X and some additional pixels / quadrilaterals). An ALL OTHERS category is included, which is a category for those merge categories that do not belong to the ==, >, or < operators.

[0098] The behavior of the local culling stage 130 can be guided by a coverage merge scheme (or coverage merge rule) 510. The specific coverage merge rule applied can be based on the merge category 505 and the depth information 515. For example, when the input coverage == the existing coverage (IN.Cov == EX.Cov) and the existing depth is a superset of the input depth, the local culling stage 130 can apply the coverage rule 520. In this case, the local culling stage 130 will "keep both special depth", which will be described below in reference to [[[]] Figure 6will be further elaborated together with the definition of each of the other possible coverage merging rules. As another example, when the input coverage < existing coverage (IN.Cov < EX.Cov) and the existing depth is better than the input depth (i.e., relatively deeper), the local culling stage 130 may apply the coverage merging rule 525. The following refers to Figure 6 will further describe the coverage merging rule 525 of "keep both union depth".

[0099] The depth information 515 covers six columns of possibilities, and each column in the six columns of possibilities is Figure 5 shown in: 1) the input depth is strictly better (i.e., strictly deeper), 2) the input depth is better, 3) the existing depth is a superset, 4) the input depth is a superset, 5) the existing depth is better, and 6) the existing depth is strictly better. Figure 5 The illustration of each of these possibilities in

[0100] is shown relative to the depth range between the MIN (i.e., minimum) depth and the MAX (i.e., maximum) depth. Figure 6 As shown in Figure 5 the legend 600 provides additional explanations for each of the coverage merging rules 510 in Figure 6 The relevance of these rules to the culling performance is also shown. To maintain correctness, the local culling stage 130 may make a merging rule selection that maximizes the culling performance.

[0101] The rule type 605 is summarized as "discard X maintain Y.depth", where, as shown in Figure 5 X represents one of "EX" or "IN", and Y represents the other of "EX" or "IN". Similarly, the rule type 610 is summarized as "keep both maintain X.depth", where, as shown in Figure 5 X represents "EX" or "IN". The rule type 615 is "keep both special depth". The rule type 620 is "keep both union depth".

[0102] The rule type 605 means "only keep one of the input primitive coverage and the existing primitive coverage, which is the primitive coverage Y, and discard X; copy the depth from Y", where X and Y are defined as above. The rule type 610 means "keep both the input primitive coverage and the existing primitive coverage; copy the depth from one of {IN, EX}", where X and Y are defined as above.

[0103] Rule type 615 means "maintain both the input primitive coverage and the existing primitive coverage; special depth: minDepth = min(IN.minDepth, EX.minDepth); maxDepth = min(IN.maxDepth, EX.maxDepth)", where X and Y are defined as above; minDepth is the determined minimum depth, min() is the function to determine the minimum value; min(IN.minDepth, EX.minDepth) is the function to determine the minimum value between IN.minDepth and EX.minDepth; IN.minDepth is the minimum depth of the input coverage; EX.minDepth is the minimum depth of the existing coverage; maxDepth is the determined maximum depth; min(IN.maxDepth, EX.maxDepth) is the function to determine the minimum value between IN.maxDepth and EX.maxDepth; IN.maxDepth is the maximum depth of the input coverage; and EX.maxDepth is the maximum depth of the existing coverage.

[0104] Rule type 620 means "maintain both the input primitive coverage and the existing primitive coverage; combined depth: minDepth = min(IN.minDepth, EX.minDepth); maxDepth = max(IN.maxDepth, EX.maxDepth)", where X and Y are defined as above; minDepth is defined as above; min() is defined as above; min(IN.minDepth, EX.minDepth) is defined as above; IN.minDepth is defined as above; EX.minDepth is defined as above; maxDepth is the determined maximum depth; max(IN.maxDepth, EX.maxDepth) is the function to determine the maximum value between IN.maxDepth and EX.maxDepth; IN.maxDepth is defined as above; and EX.maxDepth is defined as above.

[0105] Figure 7 FIG. 700 is an exemplary diagram of a depth testing module 705 according to some embodiments. Figure 8 is according to some embodiments Figure 7 FIG. of a set tester 720 of the depth testing module 705. Now refer to Figure 7 and Figure 8 for reference.

[0106] The depth test module 705 receives an input 710 and one or more coverage sets (e.g., 718). Each coverage set (e.g., 718) may be stored in an on-chip buffer 715. The depth test module 705 may include one or more set testers (e.g., 720), and the one or more set testers may perform two separate checks for each corresponding coverage set (e.g., 718). First, the set tester 720 may use a depth tester (e.g., 740) to determine whether the depth range of the input 710 exceeds the depth range of the coverage set (e.g., 718). Second, the set tester 720 may use a coverage tester (e.g., 745) to determine whether the input 710 has any overlap with the coverage set (e.g., 718). The output (e.g., 725) of each set tester (e.g., 720) may be fed into an AND operation (e.g., 730), and the depth test module 705 may output a depth test pass signal 735.

[0107] Now referring to Figure 8 , the set tester 720 of the depth test module 705 is shown in more detail. The set tester 720 may receive the input 710 and the coverage set 718, and may perform two separate checks for each corresponding coverage set (e.g., 718). First, the set tester 720 may use the depth tester 740 to determine whether the depth range of the input 710 exceeds the depth range of the coverage set 718. Second, the set tester 720 may use the coverage tester 745 to determine whether the input 710 has any overlap with the coverage set 718.

[0108] Regarding determining first, the depth tester 740 selects the correct depth for the coverage set (e.g., 718) and the input (e.g., 710). In some embodiments, the depth tester 740 uses a look-up table (LUT). The following shows an example operation of the depth tester 740.

[0109] Table 1:

[0110]

[0111]

[0112] Therefore, the depth function 805 controls the set multiplexer 812. The set multiplexer 812 receives minDepth and maxDepth from the set 718 and outputs the output signal 815. The set depth 820 logic section sets the depth based on the output signal 815 and passes the depth to the comparison logic section 825. The depth function 805 also controls the input multiplexer 828. The input multiplexer 828 receives minDepth and maxDepth from the input 710 and outputs the output signal 830. Based on the output signal 815 and the output signal 830, the comparison logic section 825 can perform a comparison operation according to Table 1 described above. The comparison logic section 825 outputs the set depth test pass information 860.

[0113] The depth functions Never, Always, Equal, and NotEqual do not need to be recorded in Table 1 because the Never case is upstream and the remaining ones (i.e., Always, Equal, and NotEqual) always pass through the depth tester 740.

[0114] The coverage tester 745 is an overlap test. The overlap test applies an AND operation (e.g., 845) to two coverage masks (e.g., 840 and 850) to determine whether they cover the same location. The coverage tester 745 can apply an OR operation (not shown) to the result to see if there is any overlap. If the operation 855 checks that the output of the AND operation 845 is not equal to 0, then the coverage tester 745 outputs the coverage overlap information 865. For the different granularity masks being compared (e.g., 840 and 850), the set coverage mask covMask 880 is at the quadrilateral granularity, while the input coverage mask covMask 885 is at the pixel granularity. Therefore, for the coverage tester 745, the input coverage mask 885 is coarsened in the coarsening logic 875. For example, if any pixel within a 2×2 quadrilateral has coverage, the quadrilateral mask bit associated with that quadrilateral is set to 1. Therefore, the coarsening logic 875 is conservative and extends the coverage to the quadrilateral granularity. This is done to prevent any false negatives during the test.

[0115] The set depth test pass information 860 output from the depth tester 740 and the coverage overlap information 865 output from the coverage tester 745 can be used to determine the set test pass result 870. The set test pass result 870 can be determined according to the following: SetTestPass = (CoverageOverlap AND SetDepthTestPass) OR NOT(CoverageOverlap).

[0116] Figure 9 is an example diagram of the depth update logic section 905 according to some embodiments. The depth update logic section 905 receives the range of primitives within a tile (e.g., Figure 1 135 of ), and the range of the primitives may include a depth range 910 and a coverage mask 915. The depth update logic section 905 may process the range of primitives that survive the depth test (e.g., Figure 8 740 of ) within the tile. When a configuration with quadrilateral granularity is selected, the depth update logic section 905 considers the coarsened coverage, otherwise it considers the pixel granularity coverage. Compared with the depth test, the coarsening for depth update is done using the bitwise AND of the pixel coverage, i.e., it is only done when the coverage is over the entire quadrilateral. As a result, partially covered quadrilaterals may be lost during depth update. This loss of information improves the simplicity of the hardware. The depth update logic section 905 receives the range of primitives within a tile including the depth range 910 and the coverage mask 915 at quadrilateral or pixel granularity based on the selected configuration. Here it is generally assumed that the coverage is for quadrilaterals, but it will be understood that the same technique can be applied to pixel coverage.

[0117] The depth update logic section 905 performs two update levels. The first level overlaps in coverage between the range of primitives within the tile and the existing set, and decides for each quadrilateral whether the range of primitives within the tile or the set should be kept for optimal culling behavior. The second level is triggered for the remaining coverage (if any exists) to add the remaining coverage as a new set and then reduce the number of sets to the maximum allowed. If the depth test (e.g., Figure 8 740 of ) is performed, the second level of the depth update logic section 905 is guaranteed to have some coverage.

[0118] The following pseudocode covers the behavior of the first level of the depth update logic section 905. The following pseudocode relates to each node (i.e., an 8×8 or 16×16 pixel block, etc.). The following pseudocode includes the corner case definitions for other depth functions, where the coarse depth is not updated in cases such as equality and inequality.

[0119] / / / Define which depth is "better" for more culling

[0120] / / / Use two node depth copies - one from the start of processing and one the current version. The former is used for comparing and updating the set, while the latter serves as the current representation of the depth of the node.

[0121]

[0122]

[0123]

[0124] The first level of the depth update logic section 905 ensures that the extent of the primitives within the tile only has coverage for quads / pixels, where quads / pixels are the better choice. One function of the second level is to ensure that the new depth and coverage can be inserted into the coverage set while maintaining the constant maximum number of sets as elaborated by the configuration. The following pseudocode pertains to the second level of the depth update logic section 905, which involves inserting the extent of the primitives within the tile into the nodes.

[0125]

[0126]

[0127]

[0128] The guiding principle behind the second level of the depth update logic section 905 set merging is to minimize the loss in coarsening the available information, e.g., by merging information with similar "better" depths. Given a specific depth function, the best depth value can be maintained while preserving the coverage. While some information is lost, the hardware simplicity is improved.

[0129] Regarding the pseudocode for the second level, for depth functions of less than (LESS) or less than or equal to (LEQUAL), the logic attempts to minimize the maximum depth of the coverage set over time for all covered pixels. This is done to maximize culling when the depth test logic tests the minimum depth of a new primitive against the maximum depth of the set. Correspondingly, for greater than (GREATER) and greater than or equal to (GEQUAL), the logic attempts to maximize the minimum depth of the coverage set over time for all covered pixels. If the depth function changes sign within the image (i.e., a transition from {less than, less than or equal to} to {greater than, greater than or equal to} within the image), which would erode the quality of the data, then this technique may work poorly. Although expressed differently, this is the same logic as that used for merging the coverage of SA entries. Fully covered blocks / nodes can be implicitly handled in the logic.

[0130] Optional performance enhancements relate to the set merging code in the second level, which picks where to move the sets to be merged and which set to free for SA entries. For example, if the merged sets always use lower indices, set 1 might grow to be larger (e.g., in terms of coverage) and have a more diluted depth range. Thus, the following order of precedence is preferred, but any cyclic ordering is sufficient. If set 1 and set 2 are being merged, the merged set is written into set 1. If set 1 and set 3 are being merged, the merged set is written into set 3. If set 2 and set 3 are being merged, the merged set is written into set 2. Thus, the input SA entries will go into the other set indices being freed, or any other free set location, accordingly.

[0131] Figure 10 is an example block diagram of a GPU 1005 including a merge unit 100 according to some embodiments disclosed herein. The merge unit 100 may correspond to Figure 1 the merge unit. The merge unit 100 may be electrically connected to one or more processor cores 1010. The GPU 1005 may also include a memory device 1015, which may be a random access memory (RAM), flash memory, solid state drive (SSD), etc.

[0132] The various operations of the methods described above may be performed by any suitable means capable of performing the operations, such as various hardware and / or software components, circuits, and / or modules.

[0133] The blocks or steps of the methods or algorithms and functions described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination of the two. If implemented in software, the functions may be stored as one or more instructions or code on a tangible, non-transitory computer-readable medium or transmitted as one or more instructions or code on a tangible, non-transitory computer-readable medium. The software module may reside in a random access memory (RAM), flash memory, read only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, a hard disk, a removable disk, a CD ROM, or any other form of storage medium known in the art.

[0134] The following discussion is intended to provide a brief, general description of one or more suitable machines in which aspects of the inventive concept may be implemented. Generally, one or more machines include a system bus to which a processor, memory (e.g., RAM, ROM, or other state - saving media), storage device, video interface, and input / output interface ports are connected. One or more machines may be controlled at least in part by input from conventional input devices (such as a keyboard, mouse, etc.) and by instructions received from other machines, interaction with a virtual reality (VR) environment, biometric feedback, or other input signals. As used herein, the term "machine" is intended to broadly include a single machine, a virtual machine, or a system of machines, virtual machines, or devices operating together with a communication connection. Exemplary machines include computing devices (such as personal computers, workstations, servers, portable computers, handheld devices, telephones, tablets, etc.) and transportation devices (such as private or public transportation (e.g., cars, trains, taxis, etc.)).

[0135] One or more machines may include an embedded controller (such as a programmable or non - programmable logic device or array, an application - specific integrated circuit (ASIC), an embedded computer, a smart card, etc.). One or more machines may utilize one or more connections to one or more remote machines, such as via a network interface, modem, or other communication connection. Machines may be interconnected via physical and / or logical networks (such as intranets, the Internet, local area networks, wide area networks, etc.). Those skilled in the art will understand that network communication may utilize a variety of wired and / or wireless short - range or long - range carriers and protocols (including radio frequency (RF), satellite, microwave, Institute of Electrical and Electronics Engineers (IEEE) 802.11, Bluetooth, optical, infrared, cable, laser, etc.).

[0136] Embodiments of the inventive concept may be described by reference to or in conjunction with associated data including functions, procedures, data structures, applications, etc., which, when accessed by a machine, cause the machine to perform tasks or define abstract data types or low - level hardware contexts. For example, the associated data may be stored in volatile and / or non - volatile memory (e.g., RAM, ROM, etc.), or in other storage devices and their associated storage media (including hard disk drives, floppy disks, optical storage devices, magnetic tapes, flash memories, memory sticks, digital video disks, biological storage devices, etc.). The associated data may be transmitted in the form of packets, serial data, parallel data, propagating signals, etc. over a transmission environment including physical and / or logical networks and may be used in a compressed or encrypted format. The associated data may be used in a distributed environment and stored locally and / or remotely for machine access.

[0137] The principles of the inventive concept have been described and illustrated with reference to the illustrated embodiments. It will be appreciated that the illustrated embodiments may be modified in arrangement and detail and may be combined in any desired manner without departing from such principles. Also, although the foregoing discussion has focused on specific embodiments, other configurations may also be contemplated. Specifically, even though expressions such as "an embodiment according to the inventive concept" are used herein, these phrases are meant to generally refer to the possibility of embodiments and are not intended to limit the inventive concept to a particular embodiment configuration. As used herein, these terms may refer to the same or different embodiments that may be combined into other embodiments.

[0138] Embodiments of the inventive concept may include a non-transitory machine-readable medium including instructions executable by one or more processors, the instructions including instructions for performing elements of the inventive concept described herein.

[0139] The foregoing illustrative embodiments should not be construed as limiting the inventive concept. Although some embodiments have been described, those skilled in the art will readily appreciate that many modifications may be made to those embodiments without materially departing from the novel teachings and advantages of the present disclosure. Accordingly, all such modifications are intended to be included within the scope of this inventive concept as defined in the claims.

Claims

1. A merging unit for coarse depth culling during the merging of pixel geometries, the merging unit comprising: A rasterizer configured to receive one or more primitives and generate pixel coverage information and depth information; One or more local culling stages coupled to the rasterizer and configured to: perform local culling within the window of the primitive and output a set of surviving coverage information and surviving depth information; And One or more global culling stages coupled to at least one of the one or more local culling stages and the rasterizer and configured to perform further culling using the set of surviving coverage information and surviving depth information based on previously received coverage information and depth information, wherein the rasterizer, the one or more local culling stages, and the one or more global culling stages are configured to: minimize overdraw by creating a coarse representation of the depth at each pixel of the image during merging and preloading the coarse representation into the depth buffer before the full color channel.

2. The merging unit according to claim 1, wherein The one or more local culling stages are configured to: use only the depth information associated with the window of the primitive when performing local culling of the one or more primitives within the window of the primitive.

3. The merging unit according to claim 1 or claim 2, wherein The one or more global culling stages use at least one of the coarse depth information and the fine depth information from some or all of the past primitives and the coverage information to further cull the set of surviving coverage information and surviving depth information received from the one or more local culling stages.

4. The merging unit according to claim 1 or claim 2, wherein, The rasterizer, the one or more local culling stages, and the one or more global culling stages are configured to: minimize overdraw by creating a coarse representation of the depth at each pixel of the image during merging and preloading the coarse representation into the depth buffer before the full color channel, such that Early-Z hardware elimination will eliminate the fragments of the image that will be occluded by later fragments of the image.

5. The merging unit according to claim 1 or claim 2, wherein: The one or more local culling stages are configured to perform local culling within the window of the primitives within a tile; and The rasterizer, the one or more local culling stages, and the one or more global culling stages are configured to generate a representation for indicating which primitives and drawing calls affect the tile.

6. The merging unit according to claim 5, wherein The representation does not include primitives and drawing calls that are not visible in the final rendered image by more than a predetermined ratio.

7. The merging unit according to claim 1 further includes: An on-chip buffer, wherein the one or more global culling stages include a reordering queue that prioritizes transactions having backup data resident in the on-chip buffer.

8. The merging unit according to claim 7, wherein, The one or more global culling stages are configured to reorder the transactions based on the memory residency of the backup data in the on-chip buffer.

9. The merging unit according to claim 1, wherein The one or more global culling stages include: Merging logic configured to create a stream of covered drawing calls and primitives to be consumed by subsequent rendering stages of the graphics processor; and One or more custom extensions for specific workload culling.

10. The merging unit according to claim 1, wherein: The one or more local culling levels are configured to perform culling within the window of a primitive based on depth information of the input primitive; and the one or more global culling levels are configured to perform culling within the window of a primitive based on depth information of a previous primitive.

11. The merging unit according to claim 1, wherein The one or more global culling levels or the one or more local culling levels are configured to use the window of a primitive to cull previously received coverage information and depth information.

12. The merging unit according to claim 1, wherein, The one or more global culling levels or the one or more local culling levels are configured to use the previously received coverage information and depth information to cull the window of a primitive.

13. The merging unit according to claim 1 further comprises: One or more custom extensions for specific workload culling.

14. The merging unit according to claim 13, wherein, The one or more custom extensions for specific workload culling are not based on visibility culling.

15. The merging unit according to claim 14, wherein, The one or more local culling levels use the one or more custom extensions.

16. The merging unit according to claim 14, wherein, The one or more global culling levels use the one or more custom extensions.

17. A method for selecting a coverage merge rule associated with depth culling during merging of pixel geometries, the method comprising: analyzing depth information; classifying the depth information into a plurality of categories, the plurality of categories including: input depth strictly better, input depth better, existing depth superset, input depth superset, existing depth better, and existing depth strictly better; comparing input coverage information with existing coverage information, wherein the comparing step includes determining at least one of the following: whether the input coverage information is the same as the existing coverage information, whether the input coverage information is a strict superset of the existing coverage information, and whether the existing coverage information is a strict superset of the input coverage information; and selecting the coverage merge rule based on the results of the classification and comparison, wherein input depth strictly better means: the minimum depth of the input depth is greater than the maximum depth of the existing depth, wherein input depth better means: the minimum depth of the input depth is less than the maximum depth of the existing depth, and the maximum depth of the input depth is greater than the maximum depth of the existing depth, wherein existing depth superset means: the maximum depth of the input depth is less than the maximum depth of the existing depth, and the minimum depth of the input depth is greater than the minimum depth of the existing depth, wherein input depth superset means: the maximum depth of the input depth is greater than the maximum depth of the existing depth, and the minimum depth of the input depth is less than the minimum depth of the existing depth, wherein existing depth better means: the maximum depth of the input depth is greater than the minimum depth of the existing depth, and the minimum depth of the input depth is less than the minimum depth of the existing depth, wherein existing depth strictly better means: the maximum depth of the input depth is less than the minimum depth of the existing depth.

18. The method according to claim 17, wherein: the comparing step further includes further determining at least one of the following: the input coverage information is different from the existing coverage information, the input coverage information is not a strict superset of the existing coverage information, and the existing coverage information is not a strict superset of the input coverage information; and Select the coverage merging rule based on the result of further determination.

19. The method according to claim 17 or claim 18, wherein, The step of selecting the coverage merging rule further includes: selecting a first coverage merging rule from a plurality of coverage merging rules.

20. The method according to claim 19, wherein, The step of selecting the coverage merging rule further includes: selecting a second coverage merging rule from the plurality of coverage merging rules.

21. The method according to claim 20, wherein, The step of selecting the coverage merging rule further includes: selecting a third coverage merging rule from the plurality of coverage merging rules.

22. The method according to claim 21, wherein The step of selecting the coverage merging rule further includes: selecting a fourth coverage merging rule from the plurality of coverage merging rules.

23. A depth test system for depth culling during the merging of pixel geometries, the depth test system comprising: An on-chip buffer including one or more coverage sets; And A depth test module including one or more set testers, wherein each set tester of the one or more set testers is configured to receive a coverage set from the one or more coverage sets from the on-chip buffer, wherein the depth test module is configured to generate a depth test pass signal based on the results from the one or more set testers, wherein the depth tester of each set tester of the one or more set testers of the depth test module is configured to: receive the minimum depth and the maximum depth of the depth range of the corresponding coverage set from the corresponding coverage set among the one or more coverage sets to output a first output signal, receive the minimum depth and the maximum depth of the input depth range from the input to output a second output signal, and perform a comparison operation based on the first output signal and the second output signal to generate set depth test pass information.

24. The depth test system according to claim 23, wherein, Each set tester of the one or more set testers of the depth test module includes a depth tester and a coverage tester.

25. The depth test system according to claim 24, wherein, The depth tester is configured to: receive a depth function, a minimum depth, and a maximum depth, and generate set depth test pass information according to the depth function, the minimum depth, and the maximum depth.

26. The depth test system according to claim 25, wherein, The coverage tester is configured to: receive a first coverage mask and a second coverage mask, and generate coverage overlap information according to the first coverage mask and the second coverage mask.

27. The depth testing system according to claim 26, wherein, Each set tester of the one or more set testers is configured to: generate a set test pass result according to the set depth test pass information and the coverage overlap information.

Citation Information

Patent Citations

  • Tile-based rendering

    US20140267259A1

  • Updating depth related graphics data

    US20140354634A1

  • Method And Apparatus For Efficient Depth Prepass

    US20180082469A1