Conservative rasterization
By using conservative rasterization hardware to test the edge of the element and the pixel angle of the microtile in the graphics processing system, the problem of inefficiency during the rendering process is solved, and efficient and accurate coverage testing is achieved.
Patent Information
- Application Number
- CN201910562437.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-29
- Filing Date
- 2019-06-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2039-06-26
AI Technical Summary
During the rendering process, existing graphics processing systems increase the processing workload due to the increase in the number of elements, and anti-aliasing technology requires frequent sampling, which is inefficient.
Conservative rasterization hardware is used to perform parallel edge testing calculations on each edge of the primitive and every corner of each pixel in the microtile, and the external and internal coverage results are calculated by combining OR gates and AND gates to accurately perform coverage tests.
Improves the efficiency of the graphics processing system, reduces processing costs and power consumption, while ensuring the accuracy of coverage tests and reducing false positive rates.
Smart Images

Figure CN110660069B_ABST
Abstract
Description
Background Art
[0001] In computer graphics, a set of surfaces representing objects in a scene are divided into a plurality of smaller and simpler segments (referred to as primitives), typically triangles, which are easier to render. The resulting segmented surfaces are usually approximations of the original surfaces, but the accuracy of such approximations can be improved by increasing the number of generated primitives, which typically in turn results in smaller primitives. The level of subdivision is usually determined by the level of detail (LOD). Thus, a larger number of primitives are typically used where a higher level of detail is required, for example, because an object is closer to an observer and / or the object has a more complex shape. However, using a larger number of triangles increases the processing effort required to render the scene and thus increases the size of the hardware performing the processing. In addition, as the average triangle size decreases, aliasing (e.g., when angled lines appear jagged) occurs more frequently, and thus graphics processing systems employ anti-aliasing techniques, which typically involve taking several samples per pixel and then filtering the data.
[0002] As the number of generated primitives increases, the ability of a graphics processing system to process the primitives becomes more important. A known method of increasing the efficiency of a graphics processing system is to render an image in a tile-based manner. In this way, the rendering space in which the primitives are to be rendered is divided into a plurality of tiles, and then the tiles can be rendered independently of one another. A tile-based graphics system includes a tiling unit for tiling the primitives, i.e., determining for the primitives which tiles in the rendering space the primitives are in. Then, when a rendering unit renders a tile, it can be provided with information (e.g., a per-tile list) indicating which primitives should be used to render the tile.
[0003] An alternative to tile-based rendering is immediate mode rendering. In such a system, there is no tiling unit generating a per-tile list, and each primitive appears to be rendered immediately; however, even in such a system, the rendering space can still be divided into pixel tiles, and the rendering of each primitive can still be done tile by tile, advancing to the next tile only after each pixel in a tile has been processed. This is done to improve the locality of memory references.
[0004] The embodiments described below are provided only as examples and do not limit the implementation of any or all of the disadvantages of known graphics processing pipelines. Summary of the Invention
[0005] The present Summary of the Invention is provided to introduce in a simplified form some concepts that will be further described in the Detailed Description below. The present Summary of the Invention is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0006] Describes a graphics processing pipeline that includes conservative rasterization hardware. The conservative rasterization hardware includes hardware logic that is arranged to perform edge test calculations in parallel for each edge of a primitive and for each corner of each pixel in a microtile. Then, the internal and external coverage results for each pixel are calculated. For a particular pixel and a particular edge, the external coverage result is determined by combining the four corners of the pixel and the edge test result of the particular edge in an OR gate. For a particular pixel and a particular edge, the internal coverage result is determined by combining the four corners of the pixel and the edge test result of the particular edge in an AND gate. The overall external coverage result for the pixel and the primitive is calculated by combining the external coverage results for the pixel and each edge of the primitive in an AND gate. The overall internal coverage result for the pixel is calculated in a similar manner. The hardware performs the coverage test precisely.
[0007] A first aspect provides a graphics processing pipeline arranged to perform rendering in a rendering space, where the rendering space is subdivided into a plurality of tiles, each tile is subdivided into a plurality of micro-tiles, each micro-tile includes the same pixel arrangement, the graphics processing pipeline includes conservative rasterization hardware, and where the conservative rasterization hardware includes: a plurality of first hardware subunits, each first hardware subunit being arranged to calculate an external coverage result for different edges of a primitive and an internal coverage result for an edge of each pixel in a micro-tile; and a plurality of second hardware subunits, each second hardware subunit being arranged to calculate an external coverage result of a primitive and an internal coverage result of the primitive for different pixels in a micro-tile, where each first hardware subunit includes: edge test calculation hardware arranged to calculate, for each corner of a pixel in a micro-tile, a value indicating whether the pixel corner is on the left side of an edge; a plurality of OR logic blocks, each OR logic block being arranged to perform an OR operation, one OR logic block for each pixel in a micro-tile, and each OR logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and where the output of the OR logic block is the external coverage result of the pixel and the edge; and a first plurality of AND logic blocks, each AND logic block being arranged to perform an AND operation, one AND logic block for each pixel in a micro-tile, and each AND logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and where the output of the AND logic block is the internal coverage result of the pixel and the edge; and where each second hardware subunit includes: a second plurality of AND logic blocks, one AND logic block for each pixel in a micro-tile, and each AND logic block being arranged to receive the external coverage result of the pixel and each edge as inputs, one input provided by each first hardware subunit, and where the output of the AND logic block is the external coverage result of the pixel and the primitive; and a third plurality of AND logic blocks, one for each pixel in a micro-tile, and each AND logic block being arranged to receive the internal coverage result of the pixel and each edge as inputs, one input provided by each first hardware subunit, and where the output of the AND logic block is the internal coverage result of the pixel and the primitive.
[0008] A second aspect provides a method for performing conservative rasterization in a graphics pipeline arranged to perform rendering in a rendering space, where the rendering space is subdivided into a plurality of tiles, each tile is subdivided into a plurality of micro-tiles, and each micro-tile includes the same pixel arrangement. The method includes: for each edge of a primitive and each corner of a pixel in a micro-tile, calculating a value indicating whether the pixel corner is on the left side of the edge; and for each pixel, the pixel having four corners: for each edge, in an OR logic block, combining the four calculated values to generate and output an external coverage result of the pixel and the edge; for each edge, in an AND logic block, combining the four calculated values to generate and output an internal coverage result of the pixel and the edge; in an AND logic block, combining the external coverage results of the pixels of each edge of the primitive to generate and output an external coverage result of the pixel and the primitive; and in an AND logic block, combining the internal coverage results of the pixels of each edge of the primitive to generate and output an internal coverage result of the pixel and the primitive.
[0009] A graphics processing pipeline including conservative rasterization hardware can be implemented in hardware on an integrated circuit. A method for manufacturing a graphics processing pipeline including conservative rasterization hardware in an integrated circuit manufacturing system can be provided. An integrated circuit definition data set can be provided, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture a graphics processing pipeline including conservative rasterization hardware. A non-transitory computer-readable storage medium can be provided, having stored thereon a computer-readable description of an integrated circuit, which, when processed, causes a layout processing system to generate a circuit layout description for use in an integrated circuit manufacturing system to manufacture a graphics processing pipeline including conservative rasterization hardware.
[0010] An integrated circuit manufacturing system can be provided, including: a non-transitory computer-readable storage medium having stored thereon a computer-readable integrated circuit description that describes a graphics processing pipeline including conservative rasterization hardware; a layout processing system configured to process the integrated circuit description to generate a circuit layout description that implements an integrated circuit including the graphics processing pipeline with conservative rasterization hardware; and an integrated circuit generation system configured to manufacture, based on the circuit layout description, a graphics processing pipeline including conservative rasterization hardware.
[0011] Computer program code for performing any of the methods described herein can be provided. A non-transitory computer-readable storage medium having stored thereon computer-readable instructions can be provided, which, when executed at a computer system, cause the computer system to perform any of the methods described herein.
[0012] As will be apparent to those skilled in the art, the above features may be appropriately combined and may be combined with any aspect of the examples described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Examples will now be described in detail with reference to the drawings, wherein:
[0014] Figure 1A is a schematic diagram of a rendering space divided into tiles and micro - tiles;
[0015] Figure 1B is a schematic diagram showing in more detail Figure 1A a part of;
[0016] Figure 2A is a schematic diagram of an example Graphics Processing Unit (GPU) pipeline;
[0017] Figure 2B is a schematic diagram showing the edge vectors of various primitives;
[0018] Figure 3A is a schematic diagram showing in more detail Figure 2A the first part of the conservative rasterization hardware from the
[0019] Figure 3B is a schematic diagram showing in more detail Figure 2A the second part of the conservative rasterization hardware from the
[0020] Figure 4 is a flowchart of an example method for performing conservative rasterization;
[0021] Figure 5A shows in more detail Figure 3A the first example implementation of the edge test hardware from the
[0022] Figure 5B shows in more detail Figure 3A the second example implementation of the edge test hardware from the
[0023] Figure 5C shows in more detail Figure 3A the third example implementation of the edge test hardware from the
[0024] Figure 6 is a flowchart of an example method for performing edge detection;
[0025] Figure 7 shows a computer system in which a graphics processing pipeline including conservative rasterization hardware is implemented; and
[0026] Figure 8An integrated circuit manufacturing system for generating an integrated circuit embodying a graphics processing pipeline is shown, the graphics processing pipeline including the conservative rasterization hardware described herein.
[0027] The drawings illustrate various examples. Those skilled in the art will understand that the element boundaries shown in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one example of the boundaries. In some examples, one element may be designed as multiple elements, or multiple elements may be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to denote similar features. Detailed Description
[0028] The following description is presented by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0029] The embodiments are now described only by way of example.
[0030] Described herein is hardware for performing conservative rasterization. This hardware may be implemented in the rasterization stage of a graphics processing pipeline (e.g., within a graphics processing unit GPU). Conservative rasterization involves determining whether a square pixel region completely overlaps a primitive (which is referred to as "inside coverage"), partially overlaps a primitive (which is referred to as "outside coverage"), or does not overlap a primitive at all. The conservative rasterization hardware described herein provides an efficient way (e.g., in terms of physical size and power consumption) to obtain outside and inside coverage results.
[0031] The hardware described herein relies on a regular subdivision of the rendering space, as can be referenced Figure 1A and 1B as described. The rendering space 100 is divided into a plurality of tiles 102 (which may be, for example, square or rectangular), and each tile is further divided into a regular arrangement of smaller regions 104, called "micro-tiles". Within each tile 102, there is a predefined arrangement of micro-tiles 104, and in various examples, all micro-tiles 104 have the same size. Although Figure 1A shows an arrangement of 5x4 micro-tiles 104 within one tile 102, in other examples, a different number of micro-tiles 104 may be present in each tile 102. Each micro-tile 104 includes the same number (and arrangement) of pixels 106. In the examples as Figure 1A and 1B shown, each micro-tile 104 includes a 4x4 arrangement of 16 pixels 106.
[0032] As described in detail below, the conservative rasterization hardware described herein computes the edge test results for the upper left corner of each pixel (as shown by the black circle 120 in Figure 1B ) in the microtile, and additionally computes the edge test results for the remaining corners of the pixels in the microtile (as shown by the white circle 122 in Figure 1B ). For any pixel, the external coverage result for a single edge of the primitive is obtained by combining the results from all four corners of the pixel in hardware logic (e.g., using an OR gate), while the internal coverage result for a single edge of the primitive is obtained by combining the results from all four corners of the pixel in a different hardware logic (e.g., using an AND gate). In this way, after determining one coverage result (e.g., the external coverage result), the other coverage result (e.g., the internal coverage result) can be obtained at minimal additional cost (e.g., in terms of size and power consumption). The external and internal coverage results for the entire primitive (and not just a single edge of the primitive) for a particular pixel are obtained by combining the corresponding results for the pixel for each individual edge in hardware logic (e.g., using an AND gate). By using the hardware described herein, the coverage test can be performed precisely (i.e., without any uncertainty margin), although as described below, it may produce false positives.
[0033] Figure 2A FIG. 2 shows a schematic diagram of an example graphics processing unit (GPU) pipeline 200, which can be implemented in hardware within a GPU and uses a tile - based rendering method. The hardware described herein can also alternatively be used in a GPU that uses an alternative rendering method, where the rendering process processes groups of pixels (e.g., in the case of immediate - mode rendering). As shown in FIG. 2, the pipeline 200 includes a geometry processing stage 202 and a rasterization stage 204. The data generated by the geometry processing stage 202 can be passed directly to the rasterization stage 204 and / or some data can be written by the geometry processing stage 202 to a memory (e.g., parameter memory 205), and then read by the rasterization stage 204 from the memory.
[0034] The geometry processing stage 202 includes a vertex shader 206, a tessellation unit 208, and a tiling unit 210. There may be one or more optional hull shaders between the vertex shader 206 and the tessellation unit (or tessellator) 208, not shown in FIG. 2. The geometry processing stage 202 may also include other elements not shown in FIG. 2, such as memories and / or other elements.
[0035] The vertex shader 206 is responsible for performing per-vertex calculations. Different from the vertex shader, the hardware tessellation unit 208 (and any optional hull shader) operates per patch rather than per vertex. The tessellation unit 208 outputs primitives, and in a system that uses vertex indices, the output primitives take the form of a buffer of three vertex indices and vertex data (e.g., for each vertex, UV coordinates and in various examples, other parameters such as displacement factors and optionally, parent UV coordinates). In the case of not using indices, the output primitives take the form of three domain vertices, where the domain vertices can include only UV coordinates or can include UV coordinates plus other parameters (e.g., displacement factors and optionally, parent UV coordinates).
[0036] The stitching unit 210 generates per-tile display lists and outputs them to, for example, the parameter memory 205. For a specific tile, each per-tile display list identifies those primitives that are at least partially located within that tile. These display lists can be generated by the tiling unit 210 using a tiling algorithm. Then, subsequent elements within the GPU pipeline 200 (e.g., the rasterization stage 204) can read data from the parameter memory 205.
[0037] The rasterization stage 204 renders some or all of the primitives generated by the geometry processing stage 202. The rasterization stage 204 includes conservative rasterization hardware 212, a coefficient generation hardware block 214, and may include other elements not shown in FIG. 2. The coarse per-tile mask and the coefficient generation hardware block 214 generate the coefficients (e.g., A, B, and C defined as follows) used in the conservative rasterization hardware 212.
[0038] The conservative rasterization hardware 212 of the rasterization stage 204 determines, for each pixel and for each of a plurality of primitives (e.g., each primitive on a per-tile display list), whether the pixel (i.e., a square pixel region, rather than a single sample position within the pixel) overlaps the primitive partially or completely. This is referred to as external and internal coverage respectively. The rasterization hardware 212 is shown in more detail in Figure 3A and 3B and its operation can be described with reference to the flowchart in Figure 4 .
[0039] As described above and as shown in Figure 2B each primitive 21, 22, 23 has a plurality of edges (e.g., the three edges of the triangular primitive 21). Each edge is defined by an edge equation, and the edge is a vector of the following form:
[0040] f(x,y) = Ax + By + C
[0041] Where A, B, and C are constant coefficients specific to the polygon edges (and thus can be pre-computed), and C has been pre-adjusted such that the scene origin is transformed to the tile origin. The conservative rasterization hardware 212 determines whether a pixel corner (with coordinates x, y) lies to the left or right of an edge or on the edge by computing the value or sign of f(x,y) for each edge of the primitive and for each pixel corner 120, 122 in the micro-tile 104. The computation is a sum of products (SOP).
[0042] Figure 3A FIG. 4 shows a portion of the conservative rasterization hardware 212 that pertains to a single pixel of a single edge. Each edge test hardware element 302 computes whether a pixel corner lies on the edge, or to the left or right of the edge, for different pixel corners among pixel corners 120, 122, by computing the value or sign of f(x,y) for the edge (block 402). This is because:
[0043] · If f(x,y) is computed as positive (i.e., greater than zero), the pixel corner lies to the right of the edge
[0044] · If f(x,y) is computed as negative (i.e., less than zero), the pixel corner lies to the left of the edge
[0045] · If f(x,y) is computed as exactly zero, the pixel corner lies exactly on the edge
[0046] Although Figure 3A only 5 discrete edge test hardware elements 302 are shown, it should be understood that there can be more of these elements, and the number will depend on the number of pixels within the micro-tile (i.e., there can be one edge test hardware element 302 for each pixel corner in the micro-tile). For example, if the micro-tile includes a 4x4 pixel arrangement (as Figure 1B shown), there can be 25 edge test hardware elements 302, one for each pixel corner 120, 122 in the micro-tile 104. Alternatively, the edge test hardware elements 302 can be combined into edge test hardware logic that is arranged to compute multiple edge test results in parallel, e.g., to compute the edge test results for each pixel corner 120, 122 in the micro-tile 104 in parallel. By combining the edge test hardware logic, efficiency can be improved because the hardware and / or intermediate results can be reused, i.e., for computing more than one edge test result. An example of such combined hardware is described in co-pending UK application No. 1805608.5, which is also shown in Figure 5A 、 5B 、5C and 6 and described below.
[0047] After the signs (or values) of f(x,y) for each pixel corner in the microtile have been computed (in hardware element 302 and in block 402), there are four results associated with each square pixel region 106 (i.e., four computed signs or values, one for each corner of the square pixel region), where most results are associated with two or more square pixel regions (i.e., in the case where a pixel corner is a corner of two or more adjacent square pixel regions), and thus the results are reused when evaluating the external and internal coverage of different pixels (i.e., different square pixel regions).
[0048] To generate the external coverage result O for pixel i and edge n n,i (where both i and n are integers and in the Figure 1B example, i = [0,24]), an OR gate 306 is used to combine the negated signs of the four corner results from block 402 (i.e., the negated versions of all four computed signs or the signs of all four computed values) (block 404). In the Figure 3A example shown, NOT gates 305 are used to perform the negation; however, in other examples, this can be implemented using alternative hardware arrangements.
[0049] The external coverage result O for pixel i and edge n n,i is a single bit, and if it is zero, it indicates that the edge does not intersect any part of the square pixel region and the entire square pixel region is on the left side of the edge vector.
[0050] To generate the internal coverage result I for pixel i and edge n n,i , an AND gate 306 is used to combine the negated signs of the four corner results (i.e., the negated versions of all four computed signs or the signs of all four computed values from block 402) (block 406). The internal coverage result I for pixel i and edge n n,i is a single bit, and if it is 1, it indicates that no corner of the square pixel region is on the left side of the edge vector.
[0051] While Figure 3A a single OR gate 304 and a single AND gate 306 are shown, this is merely to reduce the complexity of the figure. For each edge, the conservative rasterization hardware 212 includes an OR gate 304 for each pixel in the microtile (i.e., i OR gates) and an AND gate 306 for each pixel in the microtile (i.e., i AND gates). It is also possible to duplicate for each edge Figure 3AThe hardware arrangement shown in [Figure 0] causes the conservative rasterization hardware 212 to include a total of i×n multiplexers 304 and i×n AND gates 306. The conservative rasterization hardware 212 also includes n hardware elements, one for each edge, which are arranged to determine the gradient of the edge and generate selection signals for the i multiplexers associated with that edge. Additionally, the OR gates 304 (and any other OR gates described herein) may alternatively be replaced by any logic block configured to perform an OR operation (e.g., not-AND-not, or addition and comparison, etc.). Such a logic block configured to perform an OR operation may be referred to as an OR logic block. Similarly, the AND gates 306 (and any other AND gates described herein) may alternatively be replaced by any logic block configured to perform an AND operation. Such a logic block configured to perform an AND operation may be referred to as an AND logic block.
[0052] After the external coverage results O for pixel i and each edge n have been calculated n,i then an AND gate 308 is used to combine the results for the different edges (block 408), as Figure 3B shown. This generates a single external coverage result O for pixel i i which, if zero, indicates that the primitive does not intersect any part of the square pixel region. Conservative rasterization does not allow false negatives in the external coverage result, although a small number of false positives in the external coverage result are allowed. Bounding boxes can be used to remove the false positives obtained, as described below.
[0053] After the internal coverage results I for pixel i and each edge n have been calculated n,i then an AND gate 310 is used to combine the results for the different edges (block 410), as Figure 3B shown. This generates a single internal coverage result I for pixel i i which, if zero, indicates that the primitive does not completely cover the square pixel region. The internal coverage is performed exactly and has no inherent false positives.
[0054] As described above, the external coverage results obtained using the above method include many false positives. The false positives can be eliminated by applying a bounding box and excluding any pixels outside the bounding box from the external coverage positive results. The bounding box is generated such that it encloses the primitive and can be calculated, for example, such that the vertex coordinates of the bounding box are given by the maximum and minimum x and y values of the vertices of the primitive (i.e., top - left vertex = (min x, max y), top - right vertex = (max x, max y), bottom - right vertex = (max x, min y), bottom - left vertex = (min x, min y)). For example, the application of the bounding box can be achieved by calculating (e.g., in advance) a mask corresponding to the bounding box of the primitive, where all those pixels within the bounding box have mask bit one and all those pixels outside the bounding box have mask bit zero. Then an AND logic block can be used to combine the individual external coverage result O i of pixel i and the mask bit of pixel i to generate the final external coverage result O i ′ of pixel i. The final external coverage result of the pixel has fewer false positives compared to when the bounding box is not applied.
[0055] Figure 5A and 5B illustrate Figure 3A two different example implementations of the edge - testing hardware 302 shown in Figure 5A and 5B As described above, the implementations shown in
[0056] The first example hardware arrangement 500, as shown in Figure 5A includes a single micro - tile component hardware element 502, multiple (e.g., one for each corner of the pixels in the micro - tile, and thus, for the example shown in Figure 1B 25) pixel component hardware elements 504 and multiple (e.g., at least one for each corner of the pixels in the micro - tile, and thus, for the example shown in Figure 1BIn the example shown, there are at least 25 addition and comparison elements (which can be implemented as multiple adders, for example) 508, and each addition and comparison element 508 generates an output result for different pixel corners within the same microtile. The hardware arrangement 500 may additionally include one or more multiplexers 510, which connect the pixel component hardware elements 504 and optionally the microtile component hardware elements 502 to the addition and comparison elements 508. In an example including multiplexers 510, one or more select signals (which may also be referred to as "mode signals" and may include one-hot signals encoding specific operating modes of the hardware) control the operation of the multiplexers 510, and in particular, control which combination of the hardware elements 502, 504 is connected to each specific addition and comparison element 508 (e.g., for each addition and comparison element 508, which one of the multiple pixel component hardware elements 504 is connected to the addition and comparison element 508, and each addition and comparison element 508 is also connected to a single microtile component hardware element 502).
[0057] In various examples, the hardware arrangement 500 may additionally include a subsampling component element 506, but in this case, the output of this element may be set to zero so that it does not affect the output in any way. For example, a subsampling component element 506 may be provided, where the hardware arrangement is also used for other calculations, e.g., for calculations where there are multiple samples per pixel and / or the output is not a fixed value.
[0058] As described above, if the edge test hardware 302 evaluates an SOP of the following form:
[0059] f(x,y) = Ax + By + C
[0060] In the case where the values of the coefficients A, B, C may be different for each evaluated SOP, then, the microtile component hardware element 502 evaluates:
[0061] f UT (x UT ,y UT ) = Ax UT + By UT + C
[0062] where the values of x UT and y UT (the microtile coordinates relative to the tile origin 110) are different for different microtiles. The microtile component hardware element 502 may receive the values of A, B, C, x UT as well as y UT as inputs, and this element outputs a single result f UT .
[0063] The pixel component hardware element 504 evaluates:
[0064] f P (x P ,y P ) = Ax P + By P
[0065] This evaluation is for different values of x P and y P (where these values are different for different pixel corners within a microtile). The set of values of x P and y P (i.e., the x P and y P values of all pixel corners within a microtile, as defined relative to the microtile origin) is the same for all microtiles and they can be calculated, for example, by the edge test hardware 302 or obtained from a look-up table (LUT). In various examples, the origin of a microtile can be defined as the upper left corner of each microtile, and the x P and y P values can be integers, and thus, little or no calculation is required to determine these values (thus providing an efficient implementation). Referring to the example shown in Figure 1A , each microtile includes four rows of four pixels, so there are five rows of pixel corners, with five pixel corners in each row, as Figure 1B shown. Then, the set of values of x P is {0, 1, 2, 3, 4} (which can also be written as [0, 4]), and the set of values of y P is {0, 1, 2, 3, 4} (which can also be written as [0, 4]). Each pixel component hardware element 504 receives A and B as inputs and can also receive the set of values of x P and y P (e.g., in examples where these values are not integers). Each element 504 outputs a single result f P , and thus, the calculation of f P can be combined with any calculations performed to determine x P and / or y P .
[0066] The subsample component hardware element 506 (if provided) evaluates:
[0067] f S (x S ,y S ) = Ax S + By S
[0068] Since each pixel has only one subsample location and x S and y SThere is only one value, and thus, f S There is only one value, and as described above, in various examples, f S can be set to zero.
[0069] The adder and comparator element 508 evaluates:
[0070] f(x,y) = f UT + f P
[0071] Alternatively, in the presence of the subsample component hardware element 506:
[0072] f(x,y) = f UT + f P + f S
[0073] And each adder and comparator element 508 sums different combinations of the f UT and f P values (in the case where a particular combination of values is provided as input to the adder and comparator unit 508), and the combination is either fixed (i.e., hardwired between the elements) or selected by one or more multiplexers 510 (if provided). To perform the edge test, only the MSB (or sign bit) of the output (i.e., of f(x,y)) is output, so it is not necessary for the adder and comparator element 508 to compute the full result, and the adder and comparator element 508 can perform the comparison without performing the addition (thereby reducing the overall area of the hardware). This MSB indicates the sign of the result (since a > b === sign(b - a)), and as described above, this indicates whether the pixel corner is to the left or right of the edge.
[0074] A second example hardware arrangement 520, as Figure 5B shown, is Figure 5A a variant of the hardware arrangement 500 shown in Figure 1B . This second example hardware arrangement 520 includes a single microtile component hardware element 502, a plurality (e.g., one for each corner of the pixels in the microtile, and thus, for the example shown in Figure 5A , at least 25) pixel component hardware elements 524 (although the operation of these elements is slightly different from those shown in Figure 5A and described above), and a plurality (e.g., 64) of comparator elements (e.g., which can be implemented as a plurality of adders) 528 (although the operation of these elements is slightly different from that of the adder and comparator element 508 shown in Figure 5A and described above), and each comparator element 528 generates an output result. Similar to the hardware arrangement 500 shown in Figure 5BThe hardware arrangement 520 shown may also additionally include one or more multiplexers 510 controlled by a select signal. Additionally, in various examples, the hardware arrangement 520 may also additionally include a subsampling component element 506, but in this case, the output of this element may be set to zero such that it does not affect the output in any way.
[0075] As described above, if the edge test hardware 302 evaluates an SOP of the form:
[0076] f(x,y) = Ax + By + C
[0077] where the values of the coefficients A, B, C may be different for each SOP being evaluated, then the microtile component hardware element 502 operates as described above with reference to Figure 5A ; however, in the Figure 5B arrangement 520, the output of the microtile component hardware element 502 is input to each of a plurality of pixel component hardware elements 524, rather than the output being fed directly to the comparison element 528 (as Figure 5A shown).
[0078] Figure 5B The pixel component hardware elements 524 in the Figure 5A arrangement 520 do not operate in the same manner as the UT shown manner. They receive the output f
[0079] f UT (x UT ,y UT ) + f P (x P ,y P ) = f UT (x UT ,y UT ) + Ax P + By P
[0080] This evaluation is for different values of x P and y P (where these values are different for different pixel corners within the microtile). As described above (with reference to Figure 5A ), the values of x P and y P (i.e., the values of all pixel corners x P and y P within the microtile, as defined relative to the microtile origin) can be integers, so the pixel component hardware elements 524 may include an arrangement of adders to add appropriate multiples of A and / or B to the input value f generated by the microtile component hardware element 502UT are added, and this can be achieved without using any multipliers, which reduces the size and / or power consumption of the comparison unit 528. Each element 524 outputs a single result f UT + f P , and as described above, the calculation of f P and thus the calculation of the single result can be combined with any calculations performed to determine x P and / or y P .
[0081] The comparison element 528 evaluates:
[0082] f(x,y) = f UT + f P + f S
[0083] The evaluation is similar to the above-described adder and comparator element 408; however, the inputs are different because the values of f UT and f P have already been combined in the pixel component hardware element 424. Each comparison element 528 sums different combinations of (f UT + f P ) and f S values (in the case where a particular combination of values is provided as an input to the comparison unit 528), and this combination is either fixed (i.e., hardwired) or selected by one or more multiplexers 510 (if provided). To perform the edge test, only the MSB (or sign bit) of the output (i.e., of f(x,y)) is output, so it is not necessary for the comparison element 528 to calculate the full result. This MSB indicates the sign of the result, and as described above, this indicates whether the subsample position is to the left or right of the edge.
[0084] Figure 5B The hardware arrangement 520 shown can take advantage of the fact that the value of f P can be calculated quickly, or the UTC calculation can be performed in a previous pipeline stage. By using this arrangement 520, the total area of the hardware arrangement 520 can be reduced compared to the arrangement 500 shown in Figure 5A ; however, each result output by the pixel component hardware element 524 includes more bits (e.g., approximately 15 more bits) than in the arrangement 500 shown in Figure 5A .
[0085] As described above, in various examples, there can be no subsampling component hardware element 506, and in this case, the hardware arrangement 540 shown in Figure 5C can be used. This hardware arrangement 540 is Figure 5BA variant of the hardware arrangement 520 shown in. As Figure 5C shown, the comparison operation (performed by the comparison unit 528 in Figure 5B ) is combined with the addition operation (performed by the pixel component hardware element 524 in Figure 5B ) and implemented in a single pixel component and comparison element 544. As in the hardware arrangement shown in Figure 5B , in the hardware arrangement 540 shown in Figure 5C , the output can be fixed (i.e., hardwired) or selected by one or more optional multiplexers 510.
[0086] Although Figure 5A and 5B show the hardware elements 502, 504, 506, 524 connected to a single addition and comparison element 508, 528 (optionally via a multiplexer 510), this is only to reduce the complexity of the figure. As described above, each addition and comparison element 508, 528 generates an output result, and the hardware arrangements 500, 520 are arranged in all examples to compute multiple results in parallel (e.g., one for each pixel corner in a microtile, and thus, for the example shown in Figure 1B , 25 results), and thus include multiple addition and comparison elements 508, 528 (e.g., at least 25 addition and comparison elements).
[0087] Although Figure 5A , 5B and 5C only show a single microtile component element 502, such that all the results generated in parallel by the hardware arrangements 500, 520, 540 all relate to pixel corners within the same microtile, however, in other examples, the hardware arrangement may include multiple microtile component elements 502, and in such examples, the results generated in parallel by the hardware arrangement can relate to pixel corners in more than one microtile.
[0088] In various examples, the hardware arrangements 500, 520, 540 may further include a plurality of fast decision units 530 (which may also be referred to as fast fail / pass logic elements), one for each microtile, and then conditions are applied to all the outputs (e.g., all the outputs from multiple addition and comparison elements 508, 528, 544). The fast decision unit 530 receives the output generated by the microtile component hardware element 502 and determines based on the received output whether any possible contribution from the pixel component hardware elements 504, 524, 544 will change the value of the MSB of the value output by the microtile component hardware element 502.
[0089] If the value f output by the microtile component hardware element 502 UTPositive enough such that no pixel contribution can make the result f(x,y) negative (after considering any edge rule adjustments), i.e., if:
[0090] f UT >|f Pmin |
[0091] where f Pmin is the minimum, i.e., the most negative possible value of f P then the hardware arrangements 500, 520 can determine whether the edge test passes or fails without evaluating the outputs generated by the pixel component hardware elements 504, 524, 544 (i.e., without fully evaluating the final sum).
[0092] Similarly, if the value f UT output by the microtile component hardware element 502 is negative enough such that no pixel can make the result f(x,y) positive or zero, i.e., if:
[0093] |f UT |>f Pmax
[0094] where f Pmax is the maximum, i.e., the most positive possible value of f P then the hardware arrangements 500, 520, 540 can determine whether the edge test passes or fails without evaluating the outputs generated by the pixel component hardware elements 504, 524, 544 (i.e., without fully evaluating the final sum).
[0095] The implementation of the fast decision unit 530 reduces the width of the additions performed by each addition and comparison element 508, 528 because multiple (e.g., three) MSBs in the output generated by the microtile component hardware element 502 can be omitted from the addition. The exact number of MSBs that can be omitted is determined by the number of microtiles in a tile (i.e., how many X UT bits) and the exact constraints on the coefficient C.
[0096] As described above, the hardware arrangements 500, 520, 540 are all applicable to GPUs using any rendering method in which pixel groups are processed together, and this includes tile-based rendering and immediate mode rendering. In various examples, as Figure 5B shown, the hardware 520 including the fast decision unit 530 can be particularly applicable to GPUs using immediate mode rendering. This is because immediate mode rendering results in larger UTC elements 502 than for tile-based rendering (since the coordinate range can now cover the entire screen area).
[0097] In any implementation, the choice of which hardware arrangement 500, 520, 540 to use will depend on various factors, including but not limited to the rendering method used by the GPU. Compared with Figure 5B the arrangement in the hardware 520 shown, Figure 5A the hardware arrangement 500 shown has less latency and fewer registers before the multiplexer 510 for the PPC element 504; however, Figure 5A the adder and comparator elements 508 in Figure 5B are larger and use more power than the comparator unit 528 in Figure 5B Therefore, in the case of a large number of adder and comparator elements 508 (e.g., 64 or more), using the hardware arrangement 520 shown in Figure 5B may be more appropriate. However, in the hardware arrangement 520 shown in
[0098] Figure 6 It is not possible to gate the PPC element 524 if only the micro-tile index changes, but the reduced complexity of the comparator unit 528 for 64 or more outputs can provide a significant effect in terms of the power consumption of the hardware. Figure 5A 、 5B and 5C can provide a significant effect in terms of the power consumption of the hardware.
[0099] The method includes, in a first hardware element 502, calculating a first output based on the coordinates of the micro-tile (block 602). The method further includes, in each of a plurality of second hardware elements 504, 524, 544, calculating one of a plurality of second outputs based on the coordinates of one of the plurality of pixels within the micro-tile (block 604), where each of the plurality of second hardware elements and each of the plurality of second outputs relate to different pixel corners among the plurality of pixel corners in the micro-tile. The method further includes combining different combinations of the first output and one of the second outputs by using one or more adder and / or comparator units to generate a plurality of output values (block 608), where each output value is an edge test output.
[0100] In the above method, all edges of the primitive are processed in the same way; however, if a pixel is exactly on the edge of an object, an edge rule can be applied to determine that the pixel is within only one primitive (and thus make it visible). In various examples, the edge rule can determine that a pixel located on the top or left edge is within the primitive, while if the pixel is located on another edge, it is considered to be outside the primitive. These edges can be defined based on their A and B coefficients, and an example of a triangular primitive is shown in the following table:
[0101]
[0102] For example, the edge rule can be implemented by subtracting one LSB (Least Significant Bit) in the final summation (e.g., as performed in boxes 508, 528, 544) for the right or horizontal bottom edge, and this LSB can be subtracted by subtracting one LSB from the output from the microtile component hardware element 502. This results in an effective hardware implementation because it avoids the need for a comparison element to identify the case where f(x,y) is equal to zero, but rather the comparison element only needs to determine the sign of f(x,y) and thus determine whether f(x,y) ≥ 0.
[0103] Using the above hardware arrangement and method to determine the external and internal coverage of each pixel in a microtile results in a hardware logic implementation of conservative rasterization with good utilization (e.g., because it only needs to calculate some additional SOPs, and because the calculations are performed in parallel for all pixels in the microtile, thus, the results of common pixel corners can be reused instead of being calculated separately, and in various examples, the existing hardware of the rasterization stage 204 can be reused), high performance (e.g., because it does not require any adjustment to the edge coefficients or sample positions - adjusting the edge coefficients is difficult to implement precisely, and any adjustment will introduce latency, which is worse for edge adjustment than for sample position adjustment), and is compact (in terms of physical size) and power efficient (e.g., because once the external coverage is calculated, only a small amount of additional logic is needed to calculate the internal coverage, and no adjustment to the edge coefficients is required, and because the calculations are performed in parallel for all pixels in the microtile, thus, the results of common pixel corners can be reused instead of being calculated separately). Although in Figure 1B the example shown, one microtile includes a 4x4 pixel array, however, due to the benefit of utilization from reusing the calculation results, for larger pixel arrays, the increase in utilization using the methods and hardware described herein is even more significant.
[0104] Figure 7A computer system is shown in which the graphics processing system described herein may be implemented. The computer system includes a CPU 702, a GPU 704, a memory 706, and other devices 714, such as a display 716, speakers 718, and a camera 720. The graphics processing pipeline described above, particularly the conservative rasterization hardware 212, may be implemented within the GPU 704. The components of the computer system may communicate with each other via a communication bus 722.
[0105] The hardware arrangements shown in FIGS. 2, 3A, and 3B and described above are shown as including a number of functional blocks. This is only a schematic diagram and is not intended to define a strict division between the different logical elements of these entities. Each functional block may be provided in any suitable manner. It should be understood that the intermediate values formed by any element (e.g., any element in Figure 3A and 3B need not be physically generated by the hardware arrangement at any point in time and may merely represent logical values that conveniently describe the processing performed by the hardware (e.g., the graphics processing pipeline) between its inputs and outputs.
[0106] The conservative rasterization hardware 212 described herein may be implemented in hardware on an integrated circuit. The conservative rasterization hardware 212 described herein may be configured to perform any of the methods described herein. In general, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., fixed logic circuitry), or any combination thereof. The terms "module", "function", "component", "element", "unit", "block", and "logic" may be used herein generically to represent software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, unit, block, or logic represents program code that, when executed on a processor, performs the specified task. The algorithms and methods described herein may be executed by one or more processors executing code that causes the processor to execute the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disks, flash memory, hard disk memory, and other memory devices that may use magnetic, optical, and other technologies to store instructions or other data that may be accessed by a machine.
[0107] As used herein, the terms "computer program code" and "computer-readable instructions" refer to any kind of executable code for execution by a processor, including code represented in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, bytecode, code defining an integrated circuit (e.g., a hardware description language or a netlist), and code represented in a programming language such as C, Java, or OpenCL. The executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, or compiled in a virtual machine or other software environment, causes a processor of a computer system supporting the executable code to perform the tasks specified by the code.
[0108] A processor, computer, or computer system can be any kind of device, machine, or dedicated circuit, or a collection or part thereof, that has processing capabilities such that it can execute instructions. The processor can be any kind of general-purpose or special-purpose processor, such as a CPU, GPU, system-on-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), physics processing unit (PPU), radio processing unit (RPU), digital signal processor (DSP), general-purpose processor (e.g., a general-purpose GPU), microprocessor, any processing unit designed to accelerate tasks outside of the CPU, etc. The computer or computer system can include one or more processors. Those skilled in the art will recognize that such processing capabilities are incorporated into many different devices, and thus the term "computer" includes set-top boxes, media players, digital radios, PCs, servers, mobile phones, personal digital assistants, and many other devices.
[0109] The present invention also aims to include software that defines a hardware configuration as described herein, such as HDL (hardware description language) software, for designing an integrated circuit or for configuring a programmable chip to perform the desired functions. That is, a computer-readable storage medium can be provided, encoded with computer-readable program code in the form of an integrated circuit definition data set that, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture a graphics processing pipeline configured to perform any of the methods described herein, or to manufacture a graphics processing pipeline including the conservative rasterization hardware described herein. The integrated circuit definition data set can be, for example, an integrated circuit description.
[0110] Accordingly, a method of fabricating a graphics processing pipeline including conservative rasterization hardware as described herein in an integrated circuit manufacturing system can be provided. Additionally, an integrated circuit definition dataset can be provided that, when processed in an integrated circuit manufacturing system, causes the method of fabricating a graphics processing pipeline including conservative rasterization hardware to be performed.
[0111] The integrated circuit definition dataset can be in the form of computer code, such as a netlist, code for configuring a programmable chip, a hardware description language defining an integrated circuit at any level, including as register transfer level (RTL) code, a high-level circuit representation such as Verilog or VHDL, and a low-level circuit representation such as OASIS(RTM) and GDSII. A higher-level representation (e.g., RTL) that logically defines an integrated circuit can be processed on a computer system configured to generate a manufacturing definition of the integrated circuit in the context of a software environment that includes definitions of circuit elements and rules for combining these elements to generate the manufacturing definition of the integrated circuit defined by the representation. As is typically the case where software is executed on a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system to generate the manufacturing definition of the integrated circuit, execute the code defining the integrated circuit, and generate the manufacturing definition of the integrated circuit.
[0112] Reference will now be made to Figure 8 describe an example of processing an integrated circuit definition dataset at an integrated circuit manufacturing system to configure the system to fabricate a graphics processing pipeline.
[0113] Figure 8 An example of an integrated circuit (IC) manufacturing system 802 is shown that is configured to fabricate a graphics processing pipeline including conservative rasterization hardware as described in any example herein. In particular, the IC manufacturing system 802 includes a layout processing system 804 and an integrated circuit generation system 806. The IC manufacturing system 802 is configured to receive an IC definition dataset (e.g., defining a graphics processing pipeline including conservative rasterization hardware as described in any example herein), process the IC definition dataset, and generate an IC (e.g., that implements a graphics processing pipeline including conservative rasterization hardware as described in any example herein) based on the IC definition dataset. By processing the IC definition dataset, the IC manufacturing system 802 is configured to fabricate an integrated circuit that implements a graphics processing pipeline including conservative rasterization hardware as described in any example herein.
[0114] The layout processing system 804 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout from an IC definition data set are known in the art and may include, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated, e.g., in terms of logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined from the gate-level representation of the circuit. This can be done automatically or with user participation to optimize the circuit layout. When the layout processing system 804 has determined the circuit layout, it can output a circuit layout definition to the IC generation system 806. The circuit layout definition can be, for example, a circuit layout description.
[0115] As is known in the art, the IC generation system 806 generates an IC based on the circuit layout definition. For example, the IC generation system 806 can implement a semiconductor device manufacturing process for generating an IC, which can include a multi-step sequence of lithography and chemical processing steps during which an electronic circuit is gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask, which can be used in a lithography process to generate an IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 806 can be in the form of computer-readable code, and the IC generation system 806 can use this computer-readable code to form a suitable photomask for generating the IC.
[0116] The different processes performed by the IC manufacturing system 802 can all be implemented in one location, e.g., by one party. Alternatively, the IC manufacturing system 802 can be a distributed system such that some processes can be performed at different locations and can be performed by different parties. For example, some of the following stages: (i) synthesizing RTL code representing the IC definition data set to form a gate-level representation of the circuit to be generated, (ii) generating a circuit layout based on the gate-level representation, (iii) forming a photomask according to the circuit layout, (iv) manufacturing an integrated circuit using the photomask, can be performed at different locations and / or by different parties.
[0117] In other examples, by processing an integrated circuit definition data set at an integrated circuit manufacturing system, the system can be configured to manufacture a graphics processing pipeline including conservative rasterization hardware without processing the IC definition data set to determine a circuit layout. For example, the integrated circuit definition data set can define the configuration of a reconfigurable processor (e.g., an FPGA), and by processing this data set, the IC manufacturing system can be configured to generate a reconfigurable processor with the defined configuration (e.g., by loading configuration data into the FPGA).
[0118] In some embodiments, when processed in an integrated circuit manufacturing system, an integrated circuit manufacturing definition dataset may cause the integrated circuit manufacturing system to generate a device as described herein. For example, through the integrated circuit manufacturing definition dataset, the configuration of the integrated circuit manufacturing system in the manner described above with reference to Figure 8 can manufacture a device as described herein.
[0119] In some examples, the integrated circuit definition dataset may include software that runs on the hardware defined by the dataset, or software that runs in combination with the hardware defined by the dataset. In the Figure 8 example shown, the IC generation system may also be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit when manufacturing the integrated circuit, according to program code defined in the integrated circuit definition dataset, or otherwise provide program code for use with the integrated circuit.
[0120] Those skilled in the art will recognize that storage devices for storing program instructions can be distributed over a network. For example, a remote computer may store examples of processes described as software. A local or terminal computer may access the remote computer and download some or all of the software to run the program. Alternatively, the local computer may download fragments of the software as needed, or execute some software instructions at the local terminal while executing other software instructions at the remote computer (or computer network). Those skilled in the art will also recognize that all or part of the software instructions may be executed by special-purpose circuits such as DSPs, programmable logic arrays, and the like, by using conventional techniques known to those skilled in the art.
[0121] The methods described herein may be executed by a computer configured with software in a machine-readable form stored on a tangible storage medium. For example, the software is in the form of a computer program including computer-readable program code for configuring the computer to execute components of the method, or in the form of a computer program including computer program code means that, when the program runs on a computer and in the case where the computer program can be implemented on a computer-readable storage medium, the code means is adapted to execute all steps of any method described herein. Examples of tangible (or non-transitory) storage media include magnetic disks, thumb drives, memory cards, etc., and do not include propagated signals. The software may be suitable for execution on a parallel processor or a serial processor, such that the method steps may be executed in any suitable order or simultaneously.
[0122] The hardware components described herein may be generated by a non-transitory computer-readable storage medium encoded with computer-readable program code.
[0123] The memory storing machine-executable data for implementing the disclosed aspects can be a non-transitory medium. The non-transitory medium can be volatile or non-volatile. Examples of volatile non-transitory media include semiconductor-based memories such as SRAM or DRAM. Examples of technologies that can be used to implement non-volatile memories include optical and magnetic memory technologies, flash memory, phase change memory, and resistive RAM.
[0124] A particular reference to "logic" refers to a structure that performs one or more functions. Examples of logic include circuits arranged to perform those functions. For example, such circuits can include transistors and / or other hardware elements available during a manufacturing process. Such transistors and / or other elements can be used to form circuits or structures that implement and / or contain memories such as registers, flip-flops, or latches, logic operation units such as Boolean operations, mathematical operation units such as adders, multipliers, or, shifters and interconnections, by way of example. These elements can be provided as custom circuits or in a standard cell library, macro, or at other levels of abstraction. These elements can be interconnected in a particular arrangement. Logic can include circuits with fixed functions, and the circuits can be programmed to perform one or more functions; such programming can be provided from firmware or software update or control mechanisms. Logic identified as performing one function can also include logic that implements constituent functions or sub-processes. In one example, hardware logic has circuits that implement one or more fixed-function operations, state machines, or processes.
[0125] Compared with known implementations, the implementation of the concepts set forth in this application in devices, apparatuses, modules, and / or systems (and in the methods implemented herein) can result in performance improvements. Performance improvements can include one or more of increased computing performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of such devices, apparatuses, modules, and systems (such as integrated circuits), a trade-off can be made between performance improvements and the physical implementation, thereby improving the manufacturing method. For example, a trade-off can be made between performance improvements and layout area to match the performance of known implementations but use less silicon. For example, this can be done by reusing functional blocks in a serial manner or sharing functional blocks among the elements of a device, apparatus, module, and / or system. Conversely, the concepts set forth in this application that result in improvements in the physical implementation of devices, apparatuses, modules, and systems (such as reduced silicon area) can be used to improve performance. For example, this can be done by manufacturing multiple instances of a module within a predefined area budget.
[0126] As will be apparent to those skilled in the art, any range or device value given herein can be extended or changed without losing the desired effect.
[0127] It should be understood that the above-mentioned benefits and advantages may relate to one embodiment, or may relate to multiple embodiments. Embodiments are not limited to those that solve any or all of the stated problems or have any or all of the stated benefits and advantages.
[0128] Any reference to "a" item refers to one or more of these items. The term "comprising" is used herein to mean including the identified method blocks or elements, but these blocks or elements do not comprise an exclusive list, and a device may contain additional blocks or elements, and a method may contain additional operations or elements. Further, it is not implied that the blocks, elements, and operations themselves are closed.
[0129] The steps of the methods described herein may be performed in any suitable order, or simultaneously where appropriate. The arrows between the boxes in the figures illustrate an example sequence of method steps, but are not intended to exclude other sequences or the parallel execution of multiple steps. Additionally, a single block may be removed from any method without departing from the spirit and scope of the subject matter described herein. Some aspects of any of the above examples may be combined with some aspects of any of the other examples described to form further examples without loss of the sought-after effects. Where the elements of the figures are shown connected by arrows, it should be understood that these arrows illustrate only an example flow direction of the communication (including data and control messages) between the elements. The flow direction between the elements may be in either direction or both directions.
[0130] The applicant hereby independently discloses each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the general common knowledge of a person skilled in the art, based on the whole of this specification, regardless of whether such features or combinations of features solve any of the problems disclosed herein. In view of the foregoing description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present invention.
[0131] Example Clause
[0132] A. A graphics processing pipeline (200) arranged to perform rendering in a rendering space (100), wherein the rendering space (100) is subdivided into a plurality of tiles (102), each tile is subdivided into a plurality of micro-tiles (104), and each micro-tile includes pixels (106) having the same arrangement.
[0133] The graphics processing pipeline includes conservative rasterization hardware (212), wherein the conservative rasterization hardware includes:
[0134] A plurality of first hardware subunits (300), each first hardware subunit being arranged to calculate an external coverage result of the edge and an internal coverage result of the edge for each pixel in the microtile for different edges of the primitive; and
[0135] A plurality of second hardware subunits (320), each second hardware subunit being arranged to calculate an external coverage result of the primitive and an internal coverage result of the primitive for different pixels in the microtile,
[0136] wherein each first hardware subunit (300) includes:
[0137] Edge test calculation hardware (302), which is arranged to calculate a value indicating whether the pixel corner is on the left side of the edge for each corner of the pixel in the microtile;
[0138] A plurality of OR logic blocks, each OR logic block being configured to perform an OR operation (304), one OR logic block for each pixel in the microtile, and each logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and wherein the output of the OR logic block is the external coverage result of the pixel and the edge; and
[0139] A first plurality of AND logic blocks, each AND logic block being configured to perform an AND operation (306), one AND logic block for each pixel in the microtile, and each AND logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and wherein the output of the AND logic block is the internal coverage result of the pixel and the edge;
[0140] and wherein each second hardware subunit (320) includes:
[0141] A second plurality of AND logic blocks (308), one AND logic block for each pixel in the microtile, and each AND logic block being arranged to receive the external coverage result of the pixel and each edge as inputs, each first hardware subunit providing one input, and wherein the output of the AND logic block is the external coverage result of the pixel and the primitive; and
[0142] A third plurality of AND logic blocks (308), one AND logic block for each pixel in the microtile, and each AND logic block being arranged to receive the internal coverage result of the pixel and each edge as inputs, each first hardware subunit providing one input, and wherein the output of the AND logic block is the internal coverage result of the pixel and the primitive.
[0143] B. The graphics processing pipeline according to clause A, wherein the edge test calculation hardware (302) includes one or more hardware arrangements (500, 520), each hardware arrangement being arranged to perform an edge test using a sum of products, and each hardware arrangement includes:
[0144] A micro-tile component hardware element (502), including hardware logic arranged to calculate a first output using the sum of products and the coordinates of micro-tiles within a tile in the rendering space;
[0145] A plurality of pixel component hardware elements (504, 524), each pixel component hardware element including hardware logic arranged to calculate one of a plurality of second outputs using the sum of products and the coordinates of different pixel corners defined relative to the origin of the micro-tile;
[0146] A plurality of adders arranged to generate a plurality of output results for the sum of products in parallel by combining different combinations of the first output and one of the plurality of second outputs for each output result.
[0147] C. The graphics processing pipeline according to clause B, wherein each hardware arrangement further includes:
[0148] A subsampling component hardware element (506), the subsampling component hardware element including hardware logic arranged to output a fixed third output,
[0149] and wherein the plurality of adders are arranged to generate the plurality of output results by combining the third output and different combinations of the first output and one of the plurality of second outputs for each output result.
[0150] D. The graphics processing pipeline according to clause C, wherein the fixed third output is set to zero.
[0151] E. The graphics processing pipeline according to any one of clauses B - D, wherein one or more of the hardware arrangements further includes:
[0152] A plurality of multiplexers (510) arranged to select different combinations of the first output and one of the plurality of second outputs.
[0153] F. The graphics processing pipeline according to clause B, wherein the plurality of adders includes:
[0154] A plurality of add and compare elements (508), each add and compare element being arranged to generate different results among the plurality of output results by combining different combinations of the first output and one of the plurality of second outputs.
[0155] G. The graphics processing pipeline according to clause F, wherein one or more of the hardware arrangements further include a first plurality of multiplexers (510), each multiplexer of the first plurality of multiplexers having a plurality of inputs and one output, wherein each input is arranged to receive a different output of the plurality of second outputs from the plurality of pixel component hardware elements, and the multiplexer is arranged to select one of the received second outputs and output the selected second output to the plurality of addition and comparison elements through the output terminal.
[0156] H. The graphics processing pipeline according to clause B, wherein the plurality of adders include a first subset of the plurality of adders and a second subset of the plurality of adders,
[0157] wherein each pixel component hardware element (524) further includes an input terminal for receiving the first output from the microtile component hardware element, and at least one of the first subset of the plurality of adders, which is arranged to sum the first output received from the microtile component hardware element and the second output calculated by the pixel component hardware element to generate an intermediate result, and
[0158] wherein the second subset of the plurality of adders includes:
[0159] a plurality of comparison elements (528), each comparison element being arranged to generate a different output result of the plurality of output results by evaluating different intermediate results among the intermediate results.
[0160] I. The graphics processing pipeline according to clause H, wherein one or more of the hardware arrangements further include a first plurality of multiplexers (510), each multiplexer of the first plurality of multiplexers having a plurality of inputs and one output, wherein each input is arranged to receive a different intermediate result among the intermediate results from the plurality of pixel component hardware elements, and the multiplexer is arranged to select one of the received intermediate results and output the selected intermediate result to one of the plurality of comparison elements through the output terminal.
[0161] J. The graphics processing pipeline according to clause A, wherein the edge test calculation hardware (302) includes one or more hardware arrangements (540), each hardware arrangement being arranged to perform an edge test using a sum of products, and each hardware arrangement includes:
[0162] a microtile component hardware element (502), including hardware logic, which is arranged to calculate a first output using the sum of products and the coordinates of the microtiles within the tile in the rendering space;
[0163] A plurality of pixel component hardware elements (544), each element comprising:
[0164] Hardware logic arranged to calculate one of a plurality of second outputs using the sum of products and the coordinates of different pixel corners defined relative to the origin of the microtile;
[0165] An input terminal for receiving the first output from the microtile component hardware element;
[0166] A plurality of adders arranged to sum the first output received from the microtile component hardware element and the second output calculated by the pixel component hardware element to generate an intermediate result; and
[0167] A comparison element arranged to generate one of the plurality of output results by evaluating the intermediate result.
[0168] K. A method for performing conservative rasterization in a graphics pipeline arranged to render in a rendering space (100), where the rendering space (100) is subdivided into a plurality of tiles (102), each tile is subdivided into a plurality of microtiles (104), and each microtile includes the same pixel (106) arrangement, the method comprising:
[0169] For each edge of a primitive and each corner of a pixel in the microtile, calculating a value (402) indicating whether the pixel corner is to the left of the edge; and
[0170] For each pixel, the pixel having four corners:
[0171] For each edge, in an OR logic block, combining the four calculated values to generate and output an external coverage result (404) of the pixel and the edge;
[0172] For each edge, in an AND logic block, combining the four calculated values to generate and output an internal coverage result (406) of the pixel and the edge;
[0173] Combining, in an AND logic block, the external coverage results of the pixel for each edge of the primitive to generate and output an external coverage result (408) of the pixel and the primitive; and
[0174] Combining, in an AND logic block, the internal coverage results of the pixel for each edge of the primitive to generate and output an internal coverage result (410) of the pixel and the primitive.
[0175] L. The method according to clause K, wherein calculating a value indicating whether the pixel corner is to the left of the edge comprises:
[0176] In a first hardware element, a first output (602) is calculated based on coordinates of micro - tiles within a tile;
[0177] In each of a plurality of second hardware elements, a second output (604) is calculated based on coordinates of the pixel corners within the micro - tile; and
[0178] The first output and the second output are combined (608).
[0179] M. The method according to clause L, wherein combining the first output and the second output includes:
[0180] Determining a sign of a sum of the first output and the second output.
[0181] N. The method according to clause L or M, wherein by combining different combinations of the first output and the second output in each of a plurality of adder - comparator elements (508), a plurality of values indicating whether a pixel corner is on the left side of the edge are generated in parallel for different pixel corners in the micro - tile.
[0182] O. A graphics processing pipeline configured to perform the method according to any one of clauses J - N.
[0183] P. The graphics processing pipeline according to any one of claims A - I and O, wherein the graphics processing system is implemented in hardware on an integrated circuit.
[0184] Q. A computer - readable code configured to perform the method according to any one of clauses J - N when the code is run.
[0185] R. A computer - readable storage medium having encoded thereon the computer - readable code according to clause Q.
[0186] S. A method of using an integrated circuit manufacturing system to manufacture a graphics processing pipeline according to any one of clauses A - I and O.
[0187] T. An integrated circuit definition data set which, when processed in an integrated circuit manufacturing system, configures the integrated circuit manufacturing system to manufacture a graphics processing pipeline according to any one of clauses A - I and O.
[0188] U. A computer - readable storage medium having stored thereon a computer - readable description of an integrated circuit which, when processed in an integrated circuit manufacturing system, causes the integrated circuit manufacturing system to manufacture a graphics processing pipeline according to any one of clauses A - I and O.
[0189] V. An integrated circuit manufacturing system configured to manufacture a graphics processing pipeline as described in any one of clauses A-I and O.
[0190] W. An integrated circuit manufacturing system, comprising:
[0191] A non-transitory computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, the computer-readable description describing a graphics processing pipeline as described in any one of clauses A-I and O;
[0192] A layout processing system configured to process the integrated circuit description to generate a circuit layout description of the integrated circuit embodying the graphics processing pipeline; and
[0193] An integrated circuit generation system configured to manufacture the graphics processing pipeline according to the circuit layout description.
Claims
1. A graphics processing pipeline system arranged to perform rendering in a rendering space, where the rendering space is subdivided into a plurality of tiles, each tile is subdivided into a plurality of micro-tiles, and each micro-tile includes the same pixel arrangement. The graphics processing pipeline system includes conservative rasterization hardware, where the conservative rasterization hardware includes: A plurality of first hardware sub-units, each first hardware sub-unit being arranged to calculate, for different edges of a primitive, the external coverage result of the edge and the internal coverage result of the edge for each pixel in the micro-tile; And A plurality of second hardware sub-units, each second hardware sub-unit being arranged to calculate, for different pixels in the micro-tile, the external coverage result of the primitive and the internal coverage result of the primitive, where each first hardware sub-unit includes: Edge test calculation hardware arranged to calculate, for each corner of the pixel in the micro-tile, a value indicating whether the pixel corner is on the left side of the vector of the edge; A plurality of OR logic blocks, each OR logic block being configured to perform an OR operation, one OR logic block for each pixel in the micro-tile, and each logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and where the output of the OR logic block is the external coverage result of the pixel and the edge, indicating whether the edge does not intersect any part of the square region of the pixel, and whether the entire square region of the pixel is on the left side of the vector of the edge; and A first plurality of AND logic blocks, each AND logic block being configured to perform an AND operation, one AND logic block for each pixel in the micro-tile, and each AND logic block being arranged to receive four values as inputs from the edge test calculation hardware, one value for each corner of the pixel, and where the output of the AND logic block in the first plurality of AND logic blocks is the internal coverage result of the pixel and the edge, indicating whether each corner of the square region of the pixel is not on the left side of the vector of the edge; And where each second hardware sub-unit includes: A second plurality of AND logic blocks, one AND logic block for each pixel in the micro-tile, and each AND logic block being arranged to receive as inputs the external coverage result of the pixel and each edge, one input provided by each first hardware sub-unit, and where the output of the AND logic block in the second plurality of AND logic blocks is the external coverage result of the pixel and the primitive, indicating whether the primitive does not intersect any part of the square region of the pixel; and A third plurality of AND logic blocks, one AND logic block for each pixel in the microtile, and each AND logic block is arranged to receive the pixel and the internal coverage result of each of the edges as inputs, each of the first hardware subunits provides one input, and wherein the output of the AND logic block in the third plurality of AND logic blocks is the internal coverage result of the pixel and the primitive, indicating whether the primitive does not fully cover the square area of the pixel.
2. The graphics processing pipeline system according to claim 1, wherein the edge test calculation hardware includes one or more hardware arrangements, each hardware arrangement being arranged to perform an edge test using a sum of products, and each hardware arrangement includes: A microtile component hardware element, including hardware logic, which is arranged to calculate a first output using the sum of products and the coordinates of the microtiles within the tile in the rendering space; A plurality of pixel component hardware elements, each pixel component hardware element including hardware logic, which is arranged to calculate one of a plurality of second outputs using the sum of products and the coordinates of different pixel corners defined relative to the origin of the microtile; A plurality of adders, which are arranged to generate a plurality of output results for the sum of products in parallel by combining different combinations of the first output and one of the plurality of second outputs for each output result.
3. The graphics processing pipeline system according to claim 2, wherein each hardware arrangement further includes: A subsampling component hardware element, the subsampling component hardware element including hardware logic arranged to output a fixed third output, and wherein the plurality of adders are arranged to generate the plurality of output results by combining the third output and different combinations of the first output and one of the plurality of second outputs for each output result.
4. The graphics processing pipeline system according to claim 3, wherein the fixed third output is set to zero.
5. The graphics processing pipeline system according to claim 2, wherein one or more of the hardware arrangements further includes: A plurality of multiplexers, which are arranged to select different combinations of the first output and one of the plurality of second outputs.
6. The graphics processing pipeline system according to claim 2, wherein the plurality of adders includes: A plurality of adder and comparison elements, each adder and comparison element being arranged to generate different results among the plurality of output results by combining different combinations of the first output and one of the plurality of second outputs.
7. The graphics processing pipeline system according to claim 6, wherein one or more of the hardware arrangements further includes a first plurality of multiplexers, each multiplexer in the first plurality of multiplexers having a plurality of input terminals and one output terminal, wherein each input terminal is arranged to receive different outputs among the plurality of second outputs from the plurality of pixel component hardware elements, and the multiplexer is arranged to select one of the received second outputs and output the selected second output to the plurality of adder and comparison elements through the output terminal.
8. The graphics processing pipeline system according to claim 2, wherein the plurality of adders includes a first subset of the plurality of adders and a second subset of the plurality of adders, wherein each of the pixel component hardware elements further includes an input terminal for receiving the first output from the microtile component hardware element, and at least one of the first subset of the plurality of adders, which is arranged to sum the first output received from the microtile component hardware element and the second output calculated by the pixel component hardware element to generate an intermediate result, and wherein the second subset of the plurality of adders includes: a plurality of comparison elements, each comparison element being arranged to generate a different output result among the plurality of output results by evaluating different intermediate results among the intermediate results.
9. The graphics processing pipeline system according to claim 8, wherein one or more of the hardware arrangements further includes a first plurality of multiplexers, each multiplexer of the first plurality of multiplexers having a plurality of input terminals and one output terminal, wherein each input terminal is arranged to receive a different intermediate result among the intermediate results from the plurality of pixel component hardware elements, and the multiplexer is arranged to select one of the received intermediate results and output the selected intermediate result through the output terminal to one of the plurality of comparison elements.
10. The graphics processing pipeline system according to claim 1, wherein the edge test calculation hardware includes one or more hardware arrangements, each hardware arrangement being arranged to perform an edge test using a sum of products, and each hardware arrangement includes: a microtile component hardware element, including hardware logic, which is arranged to calculate a first output using the sum of products and the coordinates of the microtiles within a tile in the rendering space; a plurality of pixel component hardware elements, each pixel component hardware element including: hardware logic, which is arranged to calculate one of a plurality of second outputs using the sum of products and the coordinates of different pixel corners defined relative to the origin of the microtile; an input terminal for receiving the first output from the microtile component hardware element; a plurality of adders, which are arranged to sum the first output received from the microtile component hardware element and the second output calculated by the pixel component hardware element to generate an intermediate result; and a comparison element, which is arranged to generate one of a plurality of output results by evaluating the intermediate result.
11. The graphics processing pipeline system according to claim 1, wherein the graphics processing pipeline system is implemented in hardware on an integrated circuit.
12. A method for performing conservative rasterization in a graphics pipeline, the graphics pipeline being arranged to perform rendering in a rendering space, wherein the rendering space is subdivided into a plurality of tiles, each tile is subdivided into a plurality of microtiles, and each microtile includes the same pixel (106) arrangement, the method comprising: For each edge of the primitive and each corner of the pixels in the microtile, calculate a value indicating whether the pixel corner is to the left of the vector of the edge; and For each pixel, the pixel has four corners: For each edge, in an OR logic block, combine the four calculated values to generate and output an external coverage result of the pixel and the edge, indicating whether the edge does not intersect any part of the square region of the pixel, and whether the entire square region of the pixel is to the left of the vector of the edge; For each edge, in an AND logic block, combine the four calculated values to generate and output an internal coverage result of the pixel and the edge, indicating whether each corner of the square region of the pixel is not to the left of the vector of the edge; Combine the external coverage results of the pixel for each edge of the primitive in an AND logic block to generate and output an external coverage result of the pixel and the primitive, indicating whether the primitive does not intersect any part of the square region of the pixel; and Combine the internal coverage results of the pixel for each edge of the primitive in an AND logic block to generate and output an internal coverage result of the pixel and the primitive, indicating whether the primitive does not completely cover the square region of the pixel.
13. The method according to claim 12, wherein calculating the value indicating whether the pixel corner is to the left of the edge comprises: In a first hardware element, calculate a first output based on the coordinates of the microtiles within the tile; In each of a plurality of second hardware elements, calculate a second output based on the coordinates of the pixel corner within the microtile; and Combine the first output and the second output.
14. The method according to claim 13, wherein combining the first output and the second output comprises: Determine the sign of the sum of the first output and the second output.
15. The method according to claim 13, wherein by combining different combinations of the first output and the second output in each of a plurality of addition and comparison elements, a plurality of values indicating whether the pixel corner is to the left of the edge are generated in parallel for different pixel corners in the microtile.
16. A computer program product comprising computer-readable code configured to cause the method according to claim 12 to be executed when the computer-readable code runs.
17. A computer-readable storage medium having stored thereon the computer program product according to claim 16.
18. A manufacturing method that uses an integrated circuit manufacturing system to manufacture the graphics processing pipeline system according to claim 1, the manufacturing method comprising: Receiving and processing an integrated circuit definition data set by a layout processing system in the integrated circuit manufacturing system to determine a circuit layout; The layout processing system outputs the circuit layout definition to an integrated circuit generation system in the integrated circuit manufacturing system; The integrated circuit generation system manufactures the graphics processing pipeline system according to the circuit layout definition.
19. An integrated circuit manufacturing system comprising: A non-transitory computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, the computer-readable description describing the graphics processing pipeline system according to claim 1, wherein the computer-readable description is computer-readable program code in the form of an integrated circuit definition dataset; A layout processing system configured to process the computer-readable description of the integrated circuit and generate a circuit layout description of the integrated circuit embodying the graphics processing pipeline system, wherein the circuit layout description is a circuit layout definition; and An integrated circuit generation system configured to fabricate the graphics processing pipeline system based on the circuit layout description generated by the layout processing system.