Graphics Processing System and Rendering Method
By introducing variable fragment shading rate and attachment FSR technology into the graphics processing system, the shading sampling points in different regions are dynamically adjusted, which solves the problems of waste of resources and inefficient processing in the existing technology, and achieves a more efficient rendering process.
Patent Information
- Application Number
- CN202211588664.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-13
- Filing Date
- 2022-12-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-12-12
AI Technical Summary
When rendering images, existing graphics processing systems are difficult to effectively manage the fragment shading rate in different regions, resulting in waste of resources and inefficient processing.
By introducing the concept of variable fragment shading rate in the graphics processing system, the number of shading sampling points is dynamically adjusted according to different primitives and regions, and a variety of fragment shading rate values (such as 1×1, 2×2, 4×4, etc.) are used and combined with the attachment FSR technology, the processing of sampler fragments during the rendering process is optimized.
It realizes dynamically adjusting the number of shading sampling points in different image areas as needed, improving rendering efficiency, and reducing unnecessary computing and resource consumption.
Smart Images

Figure CN116263973B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to UK Patent Application No. 2117998.1 filed on December 13, 2021 and UK Patent Application No. 2117999.9 also filed on December 13, 2021, which are hereby incorporated by reference in their entirety. Technical field
[0003] The present disclosure relates to graphics processing systems, and more particularly to graphics processing systems that implement variable fragment shading rates. Background art
[0004] Graphics processing systems are typically configured to receive, for example, graphics data from an application running on a computer system and render the graphics data to provide a rendered output. For example, the graphics data provided to the graphics processing system may describe geometries within a three - dimensional (3D) scene to be rendered, and the rendered output may be a rendered image of the scene. Some graphics processing systems (which may be referred to as "tile - based" graphics processing systems) use a rendering space that is subdivided into a plurality of tiles. A "tile" is a section of the rendering space and can have any suitable shape, but is typically rectangular (where the term "rectangular" includes square). As is known in the art, there are many benefits to subdividing the rendering space into tile sections. For example, subdividing the rendering space into tile sections allows the image to be rendered in a tile - by - tile manner, where the graphics data for a tile can be temporarily stored "on - chip" during the rendering of the tile, thereby reducing the amount of data transferred between the system memory and the chip of the graphics processing unit (GPU) that implements the graphics processing system.
[0005] Tile - based graphics processing systems typically operate in two stages: a geometry processing stage and a rendering stage. In the geometry processing stage, the graphics data for rendering is analyzed to determine for each tile in the tiles which graphics data items are present within that tile. Subsequently, in the rendering stage (e.g., the rasterization stage), the tile can be rendered by processing those graphics data items that are determined to be present within a particular tile (without processing the graphics data items that are determined to be absent from a particular tile during the geometry processing stage).
[0006] When rendering an image, it is well known that the rendering can use more sample points than the number of pixels by which the output image will be represented. This oversampling is useful for anti - aliasing purposes and is typically specified to the graphics processing pipeline as a constant for the entire image (i.e., a single anti - aliasing rate).
[0007] Recently, the concept of variable fragment shading rates has been considered. Here, depending on the situation, rendering can use fewer shading sample points than the number of pixels (which can be referred to as 'undersampling'), or more shading sample points than the number of pixels (which can be referred to as'multisampling'). Additionally, different parts of the same image may have different fragment shading rates. For example, a higher sampling rate may still be useful for anti-aliasing purposes in very detailed or focused parts, but a lower shading sampling rate can reduce processing in the rendering area of uniform or less important parts of the image. SUMMARY OF THE INVENTION
[0008] The present invention content is provided to introduce a series of concepts further described below in the detailed description in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0009] According to a first aspect, there is provided a method of rendering a scene formed by primitives in a rendering space in a graphics processing system, the method comprising: a rendering stage, the rendering stage including the steps of: receiving data describing one or more (optionally two or more) primitives and two or more associated fragment shading rates to be used during rendering; identifying shader fragment task instances to be shaded from the primitives, wherein the shader fragment task instances include shader fragment task instances associated with a first fragment shading rate and shader fragment task instances associated with a second fragment shading rate; combining the shader fragment task instances into a shading task, the shading task including fragment task instances associated with the first fragment shading rate and fragment task instances associated with the second fragment shading rate, and wherein the fragment task instances combined into the shading task require a common shader program; and processing the shading task.
[0010] The graphics processing system may be a tile-based graphics processing system, wherein the method includes the step of performing the rendering stage on a per-tile basis (i.e., performing the steps separately for each tile).
[0011] The range of fragment shading rates may include a 'per-pixel' fragment shading rate. However, two or more fragment shading rates associated with a primitive may each be coarser than the 'per-pixel' fragment shading rate (i.e., they may specify multiple pixels to be shaded together).
[0012] The method may include receiving data describing multiple (i.e., a plurality of) primitives. Thus, fragment task instances from multiple primitives may appear in the same shading task.
[0013] Optionally, the method further includes: using the received data describing the shading rates of the primitive and two or more associated fragments to fill a buffer indicating the sampling positions of the primitive within the region of the rendering space; and parsing the buffer to generate microtiles, each microtile corresponding to an array of sampling positions within the region and containing sampler fragments from the one or more primitives; and at least one microtile containing sampler fragments associated with the first fragment shading rate and sampler fragments associated with the second fragment shading rate; wherein the identifying step includes identifying the fragment task instances to be shaded from the microtiles.
[0014] Optionally, the method further includes: arranging the fragment task instances into blocks, wherein the fragment task instances within any given block are associated with a common fragment shading rate; and wherein the step of combining fragment task instances includes combining blocks of fragment task instances that require a common shader program into a shading task, and wherein the shading task includes a block of fragment task instances associated with the first fragment shading rate and a block of fragment task instances associated with the second fragment shading rate. Combining shader fragment task instances that require a common shader program into a shading task may include maintaining data indicating the fragment shading rate associated with each block combined into the shader task. The shader fragment task instances arranged into a given block may include adjacent shader fragments within the rendering space.
[0015] Optionally, filling the buffer includes performing hidden surface removal to identify and not store one or more sampler fragments within the rendering space that do not contribute to the scene to be rendered.
[0016] Optionally, processing the shading task includes processing the fragment task instances of the task in parallel using the common shader program. That is, processing the shading task may include processing the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate in parallel using the common shader program. Optionally, the processing is performed using a SIMD processor. Accordingly, the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate are processed simultaneously in the SIMD processor. That is, fragment task instances associated with different fragment shading rates may be present in different lanes of the SIMD processor. Processing the shading task may include using the data indicating the fragment shading rate associated with the block to perform incremental calculations on the shader fragment task instances within the block.
[0017] Optionally, the method further includes a geometry processing stage, optionally, wherein the geometry processing stage includes transforming the primitive into the rendering space, and storing data associated with the transformed primitive, and / or determining and storing control flow data indicating which primitives are associated with different regions of rendering the rendering space.
[0018] According to a second aspect, there is provided a graphics processing system configured to render a scene formed by primitives, wherein the graphics processing system includes rendering logic configured to: receive data describing one or more (optionally two or more) primitives and two or more associated fragment shading rates to be used during rendering; identify shader fragment task instances to be shaded from the primitives, wherein the shader fragment task instances include shader fragment task instances associated with a first fragment shading rate and shader fragment task instances associated with a second fragment shading rate; combine the shader fragment task instances into a shading task, the shading task including fragment task instances associated with the first fragment shading rate and fragment task instances associated with the second fragment shading rate, and wherein the fragment task instances combined into the shading task require a common shader program; and process the shading task.
[0019] The graphics processing system may be a tile-based graphics processing system, wherein the rendering logic is configured to perform the recited steps on a per-tile basis (i.e., configured to operate on each tile individually).
[0020] The range of fragment shading rates may include a 'per pixel' fragment shading rate. However, two or more fragment shading rates associated with a primitive may each be coarser than the 'per pixel' fragment shading rate (i.e., it may specify multiple pixels to be shaded together).
[0021] The rendering logic may be configured to receive data describing multiple (i.e., a plurality of) primitives. Thus, fragment task instances from multiple primitives may appear in the same shading task.
[0022] Optionally, the rendering logic is further configured to: use the received data describing the primitive and two or more associated fragment shading rates to fill a buffer indicating the sampling positions of the primitive within the region of the rendering space; and parse the buffer to generate micro-tiles, each micro-tile corresponding to an array of sampling positions within the region and containing sampler fragments from the one or more primitives; and at least one micro-tile containing sampler fragments associated with the first fragment shading rate and sampler fragments associated with the second fragment shading rate; and identify the shader fragment task instances to be shaded from the micro-tiles.
[0023] Optionally, the rendering logic is further configured to: arrange the fragment task instances into blocks, wherein the fragment task instances within any given block are associated with a common fragment shading rate; and combine shader fragment task instances by combining a block of fragment task instances associated with the first fragment shading rate and a block of fragment task instances associated with the second fragment shading rate into a shading task. The rendering logic may also be configured to combine shader fragment task instances that require a common shader program into a shading task that maintains data indicating the fragment shading rate associated with each block in the shader task. The rendering logic may also be configured to combine adjacent shader fragment task instances in the rendering space into blocks.
[0024] Optionally, the rendering logic is further configured to perform hidden surface removal when filling the buffer to identify and not store one or more sampler fragments within the rendering space that do not contribute to the scene to be rendered.
[0025] Optionally, the rendering logic is further configured to process the shading task by processing the fragment task instances of the task in parallel using the common shader program. That is, the rendering logic may be configured to process the shading task by processing the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate in parallel using the common shader program. Optionally, the processing is performed using a SIMD processor. Therefore, the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate are processed simultaneously in the SIMD processor. That is, the fragment task instances associated with different fragment shading rates may exist in different channels of the SIMD processor. The rendering logic may also be configured to process the shading task using the data indicating the fragment shading rate associated with the block so as to perform incremental calculations on the shader fragment task instances within the block.
[0026] Optionally, the graphics processing system also includes geometry processing logic, optionally wherein the geometry processing logic includes logic configured to perform the following operations: transform the primitives into the rendering space and store data associated with the transformed primitives in a memory, and / or determine and store control flow data indicating which primitives are associated with rendering different areas of the rendering space.
[0027] According to a third aspect, there is provided a graphics processing system configured to perform the method of any preceding variation of the first aspect.
[0028] The graphics processing system may be embodied in hardware on an integrated circuit. A method of manufacturing a graphics processing system at an integrated circuit manufacturing system may be provided. An integrated circuit definition data set may be provided that, when processed in an integrated circuit manufacturing system, configures the system to manufacture the graphics processing system. A non-transitory computer-readable storage medium may be provided having stored thereon a computer-readable description of the graphics processing system, the computer-readable description causing, when processed in an integrated circuit manufacturing system, the integrated circuit manufacturing system to manufacture an integrated circuit embodying the graphics processing system.
[0029] An integrated circuit manufacturing system may be provided that includes: a non-transitory computer-readable storage medium having stored thereon a computer-readable description of the graphics processing system; a layout processing system configured to process the computer-readable description to generate a circuit layout description of an integrated circuit embodying the graphics processing system; and an integrated circuit generation system configured to manufacture the graphics processing system according to the circuit layout description.
[0030] Computer program code for performing any of the methods described herein may be provided. A non-transitory computer-readable storage medium may be provided having stored thereon computer-readable instructions that, when executed at a computer system, cause the computer system to perform any of the methods described herein.
[0031] As will be apparent to those skilled in the art, the above features may be combined as appropriate and may be combined with any of the aspects of the examples described herein.
[0032] A method of rendering a scene formed by primitives in a rendering space in a graphics processing system is also provided, the method including one or more of the following: a rendering phase that includes the steps of: receiving data describing one or more primitives and one or more associated fragment shading rates to be used during rendering; storing in a buffer sampler fragments corresponding to sampling positions of the one or more primitives within a region of the rendering space of the one or more primitives; parsing the buffer to produce microtiles, each microtile corresponding to an array of sampling positions within the region and containing sampler fragments from the one or more primitives; analyzing the microtiles to identify shader fragment task instances to be shaded and arranging the shader fragment task instances into blocks, wherein at least one block of shader fragment task instances includes shader fragment task instances from more than one microtile; and shading the block of shader fragment task instances.
[0033] There is also provided a graphics processing system configured to render a scene formed by primitives, wherein the graphics processing system includes rendering logic configured to: receive data describing one or more primitives and one or more associated fragment shading rates to be used during rendering; store in a buffer sampler fragments corresponding to the sampling positions of the one or more primitives within the regions of the one or more primitives in the rendering space; parse the buffer to generate microtiles, each microtile corresponding to an array of sampling positions within the region and containing sampler fragments from the one or more primitives; analyze the microtiles to identify shader fragment task instances to be shaded and arrange the shader fragment task instances into blocks, wherein at least one block of shader fragment task instances includes shader fragment task instances from more than one microtile; and shade the blocks of shader fragment task instances. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Examples will now be described in detail with reference to the drawings, in which:
[0035] Figure 1 A graphics processing system is shown;
[0036] Figure 2 A method that can be implemented by a graphics processing system such as Figure 1 a graphics processing system is shown;
[0037] Figure 3 How a graphics processing system can process primitives for shading at a 1×1 fragment shading rate is shown;
[0038] Figure 4 How a graphics processing system can process primitives for shading at a 2×2 fragment shading rate is shown;
[0039] Figure 5A How attachment FSR values can be determined for different sampling points within a primitive is shown, and Figure 5B how those sampling points can subsequently be shaded is shown;
[0040] Figure 6 A method of how primitives can be converted into fragments and shaded is shown;
[0041] Figure 7 An example order in which microtiles can be generated from a section of a buffer is shown;
[0042] Figure 8 A method for generating microtiles is shown;
[0043] Figure 9 A method of processing microtiles to create blocks of shader fragment task instances is shown;
[0044] Figure 10 Shows a sequence of microtiles with an FSR value of 2×4;
[0045] Figure 11 Shows a sequence of microtiles with an FSR value of 1×4;
[0046] Figure 12 Shows a method of combining fragment shader instance blocks into a shader task;
[0047] Figure 13 Shows how to derive shader fragments from microtiles containing sampler fragments with different FSR values;
[0048] Figure 14 Shows a computer system implementing a graphics processing system; and
[0049] Figure 15 Shows an integrated circuit manufacturing system for generating an integrated circuit embodying a graphics processing system.
[0050] The accompanying drawings show various examples. Those skilled in the art will understand that the element boundaries shown in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one example of a boundary. In some examples, it may be the case that one element can be designed as multiple elements, or multiple elements can be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to indicate like features. Detailed Description
[0051] The following description is presented by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be apparent to those skilled in the art.
[0052] The fragment shading rate (FSR) value can be specified for a graphics processing system in a number of ways. One way is to specify the FSR value via a 'pipeline' or 'per draw' FSR technique that associates a particular fragment shading rate value with a particular draw call (and thus with the primitives associated with that draw call). Another way is to specify the FSR value via a 'primitive' or 'exciting vertex' FSR technique that sets the particular fragment shading rate value at a per-primitive granularity. A third way is to specify the FSR value via an 'attachment' or 'screen space image' which allows the fragment shading rate to be specified based on regions of the image being rendered. For example, in the attachment FSR technique, the rendering space can be divided into multiple regions, each region (or zone) being associated with a particular FSR value. Attachment information defining the texels mapping to each region in the regions of the rendering space can be used to specify the FSR value for the regions of the rendering space, each texel being associated with the FSR value of its corresponding region of the rendering space. Alternatively, a single FSR value can be set for the entire rendering space.
[0053] These three different techniques for specifying the fragment shading rate value can be used alone or in combination. Thus, in practice, having all the different techniques available forms different sources of FSR information that need to be coordinated by the graphics processing system. For example, a particular primitive can be part of a particular draw call and be rendered in a particular region of the rendering space. In this example, the particular primitive can be associated with some or all of the following: (i) a pipeline FSR value specified as part of the particular draw call, (ii) a primitive FSR value specified for the particular primitive, and (iii) an attachment FSR value specified for the particular region of the rendering space in which the primitive is rendered. In fact, the situation can be more complex than this - the primitive may fall on one or more boundaries between pixel regions mapped to different attachment FSR texels, and thus different sampling points within a single primitive may have different FSR values associated with them.
[0054] The manner in which values from different FSR sources are combined to calculate the resolved combined FSR that will be applied to a primitive (or a portion thereof) can be specified by an indication application to the graphics processing system. That is, different types of combination operations are possible. In this sense, the combination operations can be mathematical and / or logical in nature. Thus, a logical combination operation can be specified that indicates which value from a particular one of the FSR sources should be selected for use. For example, the so-called 'hold' combination operation can specify that the first of a pair of FSR values (e.g., the pipeline fragment shading rate and the primitive fragment shading rate) should be selected for use. As another example, the so-called'replace' combination operation can specify that the second of a pair of FSR values should be selected for use. Another approach may require a mathematical determination to inform the logical operation performed on different values from different FSR sources to determine the resolved combined FSR. For example, the so-called'min' combination operation can specify that the minimum FSR value of a set or subset of FSR values should be selected for use. As another example, the so-called'max' combination operation can specify that the maximum FSR value of a set or subset of FSR values should be selected for use. In these examples, a mathematical determination (i.e., establishing which value is the maximum or minimum) is used to decide which value to use. Other combination operations can be considered more 'purely' mathematical operations. For example, using the so-called'multiply' operation, which specifies that a set or subset of FSR values should be multiplied together to calculate the FSR value for use. It will be understood that in principle any other mathematical operation can be used to combine FSR values from different sources.
[0055] It will also be understood that multiple combination operations can be used to combine values from different sources - for example, a first combination operation can be used to combine the pipeline FSR value with the primitive FSR value to produce a first combined FSR value, and a second combination operation (which can be of the same type or a different type from the first combination operation) can be used to combine the attachment FSR value with the first combined FSR value to produce a second or final combined FSR value.
[0056] This disclosure presents a manner in which fragment shading rates from these different sources can be effectively processed and combined in a graphics processing system.
[0057] Embodiments are now described by way of example only.
[0058] General system
[0059] Figure 1 An example graphics processing system 100 is shown. The example graphics processing system 100 is a tile-based graphics processing system. As mentioned above, a tile-based graphics processing system uses a rendering space that is subdivided into a plurality of tiles. A tile is a section of the rendering space and can have any suitable shape, but is typically rectangular (where the term "rectangular" includes squares). The tile sections within the rendering space conventionally have the same shape and size.
[0060] System 100 includes a memory 102, geometry processing logic 104, and rendering logic 106. As is known in the art, the geometry processing logic 104 and the rendering logic 106 can be implemented on a GPU and can share some processing resources. The geometry processing logic 104 includes a geometry fetch unit 108; primitive processing logic 109, which in turn includes geometry transformation logic 110, FSR logic 111, and culling / clipping unit 112; primitive block assembly logic 113; and tiling unit 114. The rendering logic 106 includes a parameter fetch unit 116; a sampling unit 117, which includes hidden surface removal (HSR) logic 118; and a texturing / shading unit 120. The example system 100 is a so-called "deferred rendering" system because texturing / shading is performed after hidden surface removal. However, a tile-based system does not have to be a deferred rendering system, and although the present disclosure uses a tile-based deferred rendering system as an example, the ideas presented can also apply to non-deferred (referred to as immediate mode) rendering systems or non-tile-based systems. The memory 102 can be implemented as one or more physical memory blocks and includes a graphics memory 122; a transformed parameter memory 124; a control list memory 126; and a frame buffer 128.
[0061] Figure 2 A flowchart showing a method of operating a tile-based rendering system, such as Figure 1 the system shown in. The geometry processing logic 104 performs a geometry processing stage, in which the geometry fetch unit 108 fetches (e.g., previously received from an application for which rendering is being performed) geometry data from the graphics memory 122 (in step S202) and passes the fetched data to the primitive processing logic 109. The geometry data includes graphics data items (i.e., geometry items) that describe the geometry to be rendered. For example, the geometry items can represent geometries that describe structural surfaces in a scene. The geometry items can be in the form of primitives (usually triangles, but primitives can be other 2D shapes and can also be lines or points to which textures can be applied). A primitive can be defined by the vertices of the primitive, and vertex data that describes the vertices can be provided, where the combination of vertices describes the primitive (e.g., a triangle primitive is defined by the vertex data of three vertices). An object can be composed of one or more such primitives. In some examples, an object can be composed of thousands or even millions of such primitives. A scene typically contains many objects. The geometry items can also be meshes (formed by multiple primitives, such as a quadrilateral including two triangle primitives sharing an edge). The geometry items can also be patches, where a patch is described by control points and where the patch is subdivided to generate multiple subdivided primitives.
[0062] In step S204, the geometry processing logic 104 preprocesses the geometry items, for example, by transforming the geometry items into screen space, performing vertex shading, performing geometry shading, and / or performing tessellation, which is applicable to the corresponding geometry items. Specifically, the primitive processing logic 109 (and its subunits) can operate on the geometry items and can utilize the state information retrieved from the graphics memory 122 when doing so. For example, the transformation logic 110 in the primitive processing logic 109 can transform the geometry items into the rendering space and can apply lighting / attribute processing known in the art. The resulting data is passed to the culling / clipping unit 112, which culls and / or clips any geometry that falls outside the view frustum. The FSR logic 111 can also be used to determine the FSR values associated with the various primitives. The FSR value can be the result of combining some or all of the relevant values from different FSR sources. For example, the FSR logic 111 can be configured to determine the FSR value of a primitive by combining the primitive and pipeline FSR values. The remaining transformed items of the geometry (e.g., primitives) are provided from the primitive processing logic 109 to the primitive block assembly logic 113, which groups the geometry items into blocks (also referred to as "primitive blocks") for storage. A primitive block is a data structure in which data associated with one or more primitives (e.g., the transformed geometry data associated with one or more primitives) is stored together. For example, each block can include up to N primitives and up to M vertices, where the values of N and M are implementation design choices. For example, N can be 24, and M can be 16. Each block can be associated with a block ID such that the blocks can be easily identified and referenced. Primitives typically share vertices with other primitives, so storing the vertices of the primitives in the block allows the vertex data to be stored once in the block, where multiple primitives in the primitive block can reference the same vertex data in the block. The primitive block can also store the FSR information determined by the FSR logic 111. In step S206, the primitive block having the transformed geometry data items is provided to the memory 102 for storage in the transformed parameter memory 124. The transformed geometry items and information on how to pack them into the primitive block are also provided to the tiling unit 114. In step S208, the tiling unit 114 generates control flow data for each tile in the rendering space, where the control flow data for a tile includes a control list of the identifiers of the transformed primitives that will be used to render the tile, i.e., a list of the identifiers of the transformed primitives that are at least partially located within the tile. The set of control lists of the identifiers of the transformed primitives for the individual tiles can be referred to as the "control flow list" or "display list". In step S210, the control flow data for the tiles is provided to the memory 102 for storage in the control list memory 126.Accordingly, after the geometry processing stage (i.e., after step S210), the transformed primitives to be rendered are stored in the transformed parameter memory 124, and the control flow data indicating which of the transformed primitives exist in each tile in the tile is stored in the control list memory 126. In other words, for a given geometry item, before the start of the rendering stage, the geometry processing stage is completed and the results of this stage are stored in the memory.
[0063] In the rendering stage, the rendering logic 106 renders the geometry items (primitives) in a tile-by-tile manner. In step S212, the parameter acquisition unit 116 receives the control flow data of the tile, and in step S214, the parameter acquisition unit 116 fetches the indicated transformed primitives from the transformed parameter memory 124 as indicated by the control flow data of the tile. In step S216, the rendering logic 106 renders the fetched primitives by performing sampling on the primitives to determine primitive fragments representing the primitives at discrete sampling points within the tile, and then performing hidden surface removal and texturing / shading on the primitive fragments. Specifically, the fetched transformed primitives are provided to the sampling unit 117 (which can also access state information from the graphics memory or stored together with the transformed primitives), which performs sampling and determines the primitive fragments to be shaded. As part of determining the primitive fragments to be shaded, the sampling unit 117 uses the hidden surface removal (HSR) logic 118 to remove hidden (e.g., hidden by other primitive samples) primitive fragments. Methods for performing sampling and hidden surface removal are known in the art. Conventionally, the term "fragment" refers to a sample of a primitive at a sampling point, which will be shaded to assist in determining how to render the pixels of the image (note that in the case of anti-aliasing, multiple samples may be shaded to determine how to render a single pixel). However, in the case of variable FSR, there may not be a one-to-one correspondence between the fragments generated by sampling and the fragments that are shaded. Therefore, the terms "sampler fragment" (the fragment created by sampling a primitive) and "shader fragment" (the fragment on which the fragment shader program is executed) are used herein, where it is necessary to distinguish between the fragments at different units of the GPU. For example, one shader fragment can be processed to determine the color values of more than one sampler fragment. The term "sampling" is used herein to describe the process of generating discrete fragments (sampler fragments) from a geometry item (e.g., a primitive), but this process can sometimes be referred to as "rasterization" or "scan conversion". As mentioned above, Figure 1 the system 100 is a deferred rendering system and thus performs hidden surface removal before texturing / shading. However, other systems may render fragments before performing hidden surface removal to determine which fragments are visible in the scene.
[0064] Sampler fragments not removed by the HSR logic 118 are provided from the sampling unit 117 to the texturing / shading unit 120, where texturing and / or shading is applied as shader fragments. Prior to this, the accessory FSR logic 119 can be used to further determine the FSR value associated with the sample. This can be in addition to or instead of the determination performed by the FSR logic 111, depending on the system (and thus both FSR logic blocks are indicated by dashed lines to indicate that one or the other can be optional). However, if the FSR logic 111 determines a combination of the pipeline and the original FSR value, the accessory FSR logic 119 can be configured to combine the results of those first combinations with, for example, any accessory FSR values.
[0065] The texturing / shading unit 120 is generally configured to efficiently process multiple fragments in parallel. This can be done by determining the individual fragments that require the same processing (e.g., need to run the same fragment shader) and treating them as instances of the same task, which are then run in parallel in, for example, a SIMD (Single Instruction, Multiple Data) processor. To assist with this, in some specific implementations, sampler fragments from the same primitive can be provided to the texturing / shading unit 120 in the form of so-called'micro-tiles', which are groups of sampler fragments. For example, a micro-tile can correspond to a 4×4 array of sample points corresponding to a particular region of the rendering space and can thus include up to 16 sampler fragments (depending on the primitive coverage within the micro-tile), and thus if each sampler fragment is shaded as a shader fragment, the micro-tile can include up to 16 task instances. It will be understood that these micro-tiles are separate from the 'tiles' used in tile-based rendering. As explained above, tiles are sub-divisions of the entire rendering space, during which graphics data can be temporarily stored 'on-chip' during tile rendering. Thus, the sampling results for a single tile can be stored in a buffer (e.g., within the sampling unit). Micro-tiles represent the sampling (and optionally, hidden surface removal) results for part or all of a particular primitive within a particular sub-region of a tile, and the sampling results are posted from the sampling unit 117 to the texturing / shading unit 120. Thus, multiple micro-tiles can be generated from the buffer storing the sampling results for a single tile.
[0066] Although at Figure 1Not shown, but the texturing / shading unit 120 can receive texture data from the memory 102 to apply texturing to the primitive fragments, as is known in the art. As is known in the art, the texturing / shading unit 120 can apply another process (e.g., alpha blending and other processes) to the primitive fragments to determine the rendered pixel values of the image. The rendering phase is performed for each tile in the tile, such that the entire image can be rendered in the case of determining the pixel values of the entire image. In step S218, the rendered pixel values are provided to the memory 102 for storage in the frame buffer 128. Then the rendered image can be used in any suitable manner, such as being displayed on a display screen or stored in the memory or transmitted to another device, and so on.
[0067] Interaction between FSR and general system
[0068] Figure 3 and Figure 4 illustrates how different fragment shading rate values can affect the workload on the general processing system described above.
[0069] Figure 3 illustrates the simplest case using a 1×1 fragment shading rate value, where each shader fragment instance corresponds to one sampler fragment. In this example, the object 302 is formed by four right triangle primitives that meet at the center of the object. During rasterization, it is determined that the object 302 covers four microtiles 312, 314, 316, and 318 (in this example, the microtiles are a 4×4 array of sampler fragments). In this example, for ease of understanding, each primitive is within a single microtile, but this is not necessarily the case in reality. The sampler fragment coverage within each of the microtiles 312, 314, 316, and 318 is determined and indicated by the hatching. In this example, using a 1×1 FSR value, each sampler fragment corresponds to a shader fragment that is individually shaded during rasterization, and thus corresponds to one instance of a shading task. In this example, the shader fragments are grouped into instance blocks ( Figure 3In blocks 0 through 7, for parallel coloring. In this example, 2×2 instances from microtiles 312, 314, 316, and 318 are grouped into blocks (i.e., blocks 0 and 1 are derived from microtile 312, blocks 2 and 3 are derived from microtile 314, blocks 4 and 5 are derived from microtile 316, and blocks 6 and 7 are derived from microtile 318), but this depends on the configuration of the texturing / coloring unit. To emphasize that each shader fragment (regardless of block grouping) is colored separately, a dashed box is shown around each shader fragment in each block. Thus, the content of each dashed box can be considered a task instance to be processed (i.e., colored) by the texturing / coloring unit 120. After coloring, in this simple example, the coloring results can be directly combined to form output 332 (the fact that the fragments have been processed is indicated by using different cross-hatching).
[0070] In contrast, Figure 4 illustrates the use of 2×2 fragment coloring rate values, where each shader fragment corresponds to a 2×2 sampler fragment. This example starts in a similar way to Figure 3 the example where the primitives forming object 402 are determined to cover four microtiles 412, 414, 416, and 418. Again, each microtile 412, 414, 416, and 418 in the example corresponds to a 4×4 array of sample points. Again, the sampler fragment coverage within each microtile is indicated by cross-hatching. While this 4×4 sampler granularity is retained for coverage information (as will be seen later), the 2×2 fragment coloring rate values mean that shader fragments and thus the task instances for coloring are created from a set of 2×2 sampler fragments, which are then grouped into blocks ( Figure 4 blocks 0 through 3, where block 0 is derived from microtile 412, block 1 is derived from microtile 414, block 2 is derived from microtile 416, and block 3 is derived from microtile 418). As in Figure 3 , a dashed box has been shown around each shader fragment in the blocks of Figure 4 . However, compared to Figure 3In contrast, it will be seen that the content of each dashed box corresponds to four sampler fragments (i.e., 2×2) from the original sampler fragments of microtiles 412, 414, 416, and 418. A single shader task instance is run for each dashed box. In other words, shader fragments are created, where each fragment corresponds to four original sampler fragments, and a single shading task is created for each shader fragment. As shown by one of the dashed boxes in the block from block 3, this results in a single shading result 422 corresponding to the original sampler fragments for which the task instance was built. This single shading result 422 can then be recombined with coverage information (e.g., as shown in microtiles 412, 414, 416, and 418) to produce a properly shaded set of fragments 424 having the same spatial resolution as the original set of 2×2 sampler fragments (in the example shown, this results in a single shaded fragment at that resolution). After performing a similar process for each task instance, the shaded fragments can be combined to form the output 432. In other words, although in this example the 'coarser' shader fragment size causes sampler fragments to be grouped together for shading, to some extent it can cover sample points that may not actually be covered by the primitive being shaded, but the shading result 422 is only applied at known covered sample positions, which means that the outputs from Figure 3 and Figure 4 are the same in terms of spatial coverage. However, fewer task instances need to be processed to achieve the same (in terms of spatial coverage) output, resulting in higher processing efficiency. This can be seen by comparing the number of dashed boxes in the block of Figure 3 with the number of dashed boxes in Figure 4 - Figure 3 32 dashed boxes (shader task instances) are needed, while Figure 4 only needs 16. On the other hand, when determining the shading result, the processing efficiency comes at the cost of a loss in spatial resolution. That is, although the outputs 332 and 432 can have the same spatial coverage, there may be less variation in the shading result within the covered area in the output of Figure 4 . Depending on the uniformity of the covered area, there may be no difference, and thus it is up to the programmer to judge when this loss in spatial resolution in the shading result is an acceptable trade - off for increased processing efficiency.
[0071] It will be noted that in Figure 3 , there are some task instances (dashed boxes) in blocks 0 to 6 that do not include any sampler fragments and thus do not actually need to be shaded. Similarly, in Figure 4Among them, there are task instances in blocks 0 to 3 that do not contain any sampler fragments to be colored. If the system architecture expects to receive blocks containing a specific number of task instances (e.g., 2×2 instances in the presented example), such 'empty' or 'helper' instances can be created. Although systems such as SIMD systems are most efficient when each processed instance is "useful" work, the system can still operate by using helper-like instances and can still operate more efficiently (overall) than a system that does not utilize parallelism.
[0072] Figure 5A and Figure 5B illustrates (in a manner similar to how Figure 3 and Figure 4 illustrate how a single FSR value interacts with a graphics processing system) how the attached FSR technique interacts with other FSR value sources in a graphics processing system.
[0073] In Figure 5A and Figure 5B 's example, an attached texel corresponds to (or maps to) a region of 8×8 pixels, which means the rendering space is divided into a grid of 8×8 pixel regions, and an FSR value is specified for each such 8×8 pixel region. These FSR values can be different or the same (although in practice, if all texels have the same FSR value, a larger FSR texel size may be more appropriate / effective). Figure 5A illustrates a portion 502 of the rendering space containing four pixel regions 504 0至3 and, in this example, each pixel region corresponds to a different attached texel specifying a different FSR value (the attached texel corresponding to pixel region 5040 specifies a 1×1 FSR value; the attached texel corresponding to pixel region 5041 specifies a 2×2 FSR value; the attached texel corresponding to pixel region 5042 specifies a 4×4 FSR value; the attached texel corresponding to pixel region 5043 specifies a 2×1 FSR value). The pixel region corresponding to a specific FSR texel or value can also be referred to as an FSR zone.
[0074] Triangle primitive 506 overlaps four pixel regions 504 0至3 The primitive 506 is associated with its own FSR value, which in this case is a 2×2 FSR value. This can be an FSR value specified by, for example, one FSR source (e.g., a primitive FSR value), but it can also be an FSR value established based on a combination of values from different FSR sources (e.g., primitive and pipeline FSR values). In any case, during the sampling process of the primitive, the FSR value associated with the primitive is combined with the attached FSR value. In this case, as Figure 5AAs shown, although the accessory texel corresponds to a pixel area of 8×8 pixels, the microtile size is still 4×4 samples, and thus it is found that primitive 506 overlaps with eight microtiles 508 0至7 (and the sample coverage of primitive 506 within each microtile is indicated by cross-hatching). In turn, each microtile overlaps with a pixel area corresponding to a specific accessory texel associated with a specific accessory FSR value. For a sampler fragment within each microtile, the resolved combined FSR value can be calculated based on the FSR value associated with the primitive and the FSR value of the accessory texel corresponding to the pixel area that overlaps with the microtile (i.e., if the FSR value of the primitive is a first preliminary combined value derived from the combined pipeline FSR value and the primitive FSR value, the resolved combined FSR value may be a'second' combined FSR value). Thus, for the sampler fragment in microtile 5080 derived from pixel area 5040, the accessory FSR value is 1×1. In this example, the combining operation of the accessory FSR value and the FSR value associated with the primitive is a'max' combiner, so the FSR value of the sampler fragment in microtile 5080 is 2×2 (which is the FSR value associated with primitive 506, which is greater than the 1×1 FSR value of the corresponding accessory texel). Using similar reasoning, it can be understood Figure 5A how the other FSR values shown for the sampler fragments in each of the eight microtiles 508 0至7 are derived. It can be seen that for many microtiles, the resulting FSR value of the sampler fragments they contain is 2×2, but two microtiles 5084 and 5085 have a 4×4 FSR value for the sampler fragments they contain.
[0075] Note that although the previous description relates the position of the microtiles to the accessory texels to determine the relevant accessory FSR, in other specific implementations, the position of the fragment itself can be considered relative to the accessory texel.
[0076] Figure 5B Continue Figure 5A the rasterization process started in. Based on the FSR values, the sampler fragments from Figure 5A the microtiles 508 shown in 0至7 are grouped into shader fragment blocks for parallel shading. Five microtiles with an FSR value of 2×2 (microtiles 508 0至3 and 508 6至7 ) are converted into Figure 5B FSR 2×2 blocks 0 to 5 in. As in Figure 3 and Figure 4 , a dashed box is shown around each shader fragment in each block in the block. Additionally, two microtiles with a 4×4 FSR value (microtiles 5084 and 5085) are converted into Figure 5BFSR 4x4 block 0 and 1 in. Similarly, a dashed box is shown around each shader fragment within the block.
[0077] As previously referenced Figure 3 and Figure 4 discussed, a single shader task instance is run for each dashed box. As shown by one of the dashed boxes in the dashed box from the FSR 2x2 block 2, this produces a single shading result 510 corresponding to the original sampler fragment for which the task instance was constructed. This single shading result 510 can then be recombined with the coverage information from the original microtile 5082 to produce a properly shaded set of fragments 512 at the same spatial resolution as the original set of 2x2 sampler fragments. In this case, it is found that all four sample points in the sample points corresponding to the task instance are covered by the primitive, so four shaded fragments are obtained. In contrast, one of the dashed boxes in the dashed box in the FSR 4x4 block 0 is shown as producing a single shading result 514 corresponding to 16 sample points, which are not all covered by the primitive for which the task instance was constructed. In this case, the primitive only covers eight of the 16 sample points corresponding to the task instance, so when the single shading result 514 is recombined with the coverage information from the original microtile 5084, it produces eight fragments shaded according to the single shading result 514, but at the spatial resolution of the original sampler fragments. After performing a similar process for each task instance, the shaded fragments can be combined to form the output 518.
[0078] It will be understood that although the introduction of the annex FSR technology provides the possibility of different FSR values for different parts of the same primitive compared to Figure 3 and Figure 4 , once the FSR values of the individual fragments are determined, the process is very similar to that described previously. In other words, although the coarser fragment size (compared to shading each sample point individually) causes sampler fragments to be grouped together for shading, to some extent it can also cover sample points that may not actually be covered by the primitive being shaded, but the shading results 510 and 514 are only applied to the known covered sampling positions. This means that fewer task instances need to be processed to achieve the shaded output, resulting in higher processing efficiency.
[0079] However, it can be observed that in Figure 3 , eight blocks contain a total of fifteen 'empty' or 'helper' task instances out of 32 task instances, while in Figure 4 , there are eight helper task instances out of a total of 16 task instances (as indicated by the dashed boxes). In other words, the proportion of helper tasks in the task instance pool ranges from From Figure 3 to Figure 4has increased. This apparent efficiency degradation can be offset by the fact that each (useful) task instance processed contributes to coloring more than one sampler fragment, meaning that the system is still more efficient in processing incoming work faster. However, considering that the system is processing a large number of primitives, this increase in the helper task instance usage has been identified as an opportunity to further save efficiency. Specifically, this problem becomes even more important when considering scenarios using a coarser FSR rate. In Figure 5A , the microtiles 5084 and 5085 are the only microtiles associated with an FSR value of 4×4 instead of 2×2. If the FSR value of those microtiles were 2×2, each microtile in that microtile would generate four fragment task instances and thus each microtile would create a fragment task instance block without any helper task instances. In contrast, as shown in Figure 5B , each microtile in those microtiles actually creates one fragment task instance and thus creates a fragment task instance block with three helper task instances. Therefore, if these blocks were processed with an FSR value of 2×2, they would contain only a quarter of the "useful" work compared to fully "useful" work.
[0080] In other words, at an FSR value of 1×1, a 4×4 sample microtile contains up to 16 task instances from which instance blocks are constructed ('up to' because not every sampling location has to be covered by a primitive). In contrast, at an FSR value of 2×2, there are at most 4 task instances, and at an FSR value of 4×4, there can be only one task instance. Therefore, as the fragments become coarser, there are fewer fragments per microtile and thus fewer task instances aggregated from the microtiles into the instance blocks. This problem is further magnified when considering anti-aliasing. For example, at an FSR value of 2×2 and a 4x anti-aliasing rate (i.e., doubling the number of sampler fragments in both the height and width directions), the microtile will again include only one fragment and thus only one task instance. Since both the FSR value and anti-aliasing affect the number of samples required per pixel, for simplicity, it is assumed that the examples discussed below do not have anti-aliasing. Those skilled in the art will readily understand that in the absence of FSR, anti-aliasing settings do not cause the problem of reducing the number of task instances per microtile. Those skilled in the art will also readily understand how anti-aliasing can affect the number of samples per task instance and thus the number of task instances in the microtile for FSR values greater than 1×1.
[0081] To address the problem of reducing the number of task instances per microtile, it is first worthwhile to understand in more detail the formation of instance blocks from microtiles. These instance blocks may also be known or referred to as 'quads', due to the fact that they are related to four adjacent tiles - namely, the fact that they are related to a group of 2×2 adjacent tiles. In any case, forming instance blocks from a 2×2 adjacent instance group is beneficial in part because it aids in performing certain calculations during shading. For example, it is often necessary to calculate the incremental value of a parameter, i.e., the difference or gradient of a parameter at a particular location of a tile (e.g., a depth gradient when mapping a texture), and such increments cannot be determined from the information at a single location. If each task instance were processed completely independently, calculating the increment would require retrieving additional information about other locations when processing the task instance (which would cause processing delays), or including such information in the original task instance (which would increase the size of each task instance and thus the amount of information transferred through the system). However, by grouping adjacent instance blocks together and processing them together, information about surrounding locations becomes readily accessible from other instances in the block, and thus the calculation of the increment becomes more efficient. When there are not enough adjacent instances to fill an instance block, helper instances containing information can be created to assist in the calculation of the increment. Of course, SIMD processors are typically capable of processing more than four instances of a task in parallel, and thus instance blocks can be further grouped together into larger shading tasks that process multiple instance blocks in parallel (and thus process task instances from multiple instance blocks). Grouping instance blocks further into larger shading tasks can be at least partially based on whether the blocks are related to the same fragment shader (i.e., thus all instance blocks within a given larger shading task are related to the same fragment shader). Some systems may apply another criterion for grouping based on some other state associated with the instance block (e.g., possibly derived from state information associated with the primitive tile from which the task instances in the instance block are derived). In such cases, it will be understood that multiple larger tasks may be created that are related to the same fragment shader but differ in other states. In other words, while all instances within a larger shading task may be related to the same fragment shader, this does not preclude the existence of other larger shading tasks related to the same fragment shader. In any case, this additional grouping does not affect the increment calculation, since only the entire block is combined into the larger task.
[0082] To make this approach possible in a pipelined system, it is effective to provide the task instances to be processed in a way that groups instances of the same task (i.e., instances of the same fragment shader running on different shader fragments) together. This is part of the reason why it is beneficial for the sampling unit 117 to issue microtiles to the texturing / shading unit 120 - it groups the task instances together in a way that preserves the spatial arrangement of the tiles. Subsequently, this makes it possible to then group adjacent task instances together into instance blocks.
[0083] Note that although Figure 3 , Figure 4 and Figure 5A show the microtile coverage of a single primitive, in reality the microtiles can contain sampling results from different primitives (i.e., because different primitives are visible in different regions of the rendering space corresponding to the microtile). If different primitives require the same fragment shader, in some implementations, samples can be collected into the same instance block. However, if different primitives require different fragment shaders, when the microtiles are processed to create instance blocks, the instance block formation can take this into account by separating sampler fragments associated with primitives having different fragment shaders into different instance blocks. In other words, even though a microtile may contain sampler fragments associated with different fragment shaders, the instance blocks created will each be associated with a single fragment shader. This can result in more instance blocks being created from the microtiles than would be the case if, for example, all sampling positions within the microtile were covered by the same primitive having the same fragment shader.
[0084] For example, measuring a 4×4 sampling point microtile that is completely covered by a single primitive will result in 4 2×2 instance blocks (for a 1×1 FSR value). If the microtile is exactly half covered by a first primitive (e.g., the left half) and half covered by a second primitive that requires a different fragment shader (e.g., the right half), then 4 2×2 instance blocks will also be produced (since the instance blocks will correspond to the upper left, upper right, lower left, and lower right quarters of the microtile, and each quarter is covered by only one primitive). However, if the first primitive covers all but one sampling point in the microtile and the second primitive covers the last sampling point, then this will result in 5 2×2 microtiles - each quarter covered only by the first primitive will produce one instance block, and the remaining quarter will generate two instance blocks, one instance block for samples from the first primitive and one instance block for samples from the second primitive (e.g., along with three helper instances).
[0085] One aspect of this method involves determining that the system for creating instance blocks can be further extended to reduce the number of required helper instances and thus improve overall efficiency when processing coarser fragments. By creating instance blocks from instances derived not only from a single microtile, there are more opportunities to fill the instance blocks with 'useful' work.
[0086] Figure 6 shows an example method according to this approach. As discussed above, the method can be performed during the rendering phase of a tile-based system.
[0087] At step S602, data describing a primitive is received. For ease of understanding, only one primitive is considered, but it will be understood that the primitive may be received as one of a plurality of primitives, each of which may be processed in a similar manner. The data may be received by the rendering logic in the form of a primitive block and retrieved from a memory independent of the rendering logic. Data describing one or more fragment shading rates associated with the primitive is also received. As explained above, this data may be received in the same manner as the data describing the primitive (e.g., as part of the information in the primitive block) or may be provided separately (e.g., as an FSR attachment).
[0088] At step S603, shader fragment task instances to be shaded are identified from the primitive indicated by the data received at step S602. As Figure 6 shown, step S603 may be further divided into three steps S604, S606, and S608.
[0089] At step S604, the sampling results are stored in a buffer (e.g., within the sampling unit). This may occur as part of the rasterization process, in which the primitive is sampled and optionally hidden surface removal is performed. The buffer corresponds to a region of the rendering space (e.g., a tile in the tile-based method discussed above) and is used to store the sampling results before shading and texturing. The sampling results may be sampler fragments of the primitive, and each sampler fragment corresponding to a sampling position within the rendering space is overlapped by the primitive. Considering multiple primitives, each primitive will undergo the sampling process, and thus, the buffer may store the sampling results from multiple primitives.
[0090] At step S606, the buffer is parsed to generate micro-tiles. This will be discussed in more detail below, but may be part of flushing the buffer so that it can be reused. However, in general, each micro-tile output of this step corresponds to an array of sampling positions within the region. In this regard, each micro-tile contains samples to be shaded using a common fragment shading rate - that is, within one micro-tile, the same fragment shading rate will be used for all samples contained therein, but another micro-tile (possibly corresponding to the same overall region of the rendering space as the previous micro-tile) may contain all samples to be shaded using a different fragment shading rate.
[0091] At step S608, the micro-tiles are analyzed to identify shader fragment task instances to be shaded. Again, this will be discussed in more detail below. However, in summary, the number of shader fragment task instances in a micro-tile can vary. A shader fragment shaded as a single shader fragment task instance can cover multiple sampler fragments. Thus, a micro-tile covering 16 (i.e., 4×4) sampler fragments from a single primitive may contain fewer than 16 shader fragments. For example, if the FSR value of the sampler fragments in a micro-tile is 4×4, the result of analyzing the micro-tile will be to identify only one shader fragment and thus a single shader fragment task instance. In other cases, for example, if a micro-tile contains sampler fragments from different primitives, there may be multiple shader fragments and they may require different fragment shader programs.
[0092] At step S609, the fragment task instances that require a common fragment shader program are grouped into a shading task. As explained above, some systems may apply another criterion to the step of grouping task instances into a shading task, such as based on some other state associated with the instance. That is, this step does not require all instances with the same fragment shader to be grouped into the same shading task. Instead, it should be understood that all instances contained in a given shading task resulting from this step are associated with the same fragment shader, but there may be multiple shading tasks associated with the same fragment shader. In any case, as Figure 6 shown, step S609 can be performed in two steps S610 and S612.
[0093] At step S610, the shader fragment task instances are arranged into blocks. As explained above, this arrangement of instances of a particular type of task (e.g., requiring a particular fragment shader) facilitates the parallel processing of task instances, such as in a SIMD processor. The block collects instances of the same type of task, which can then be processed together. In one approach, the shader fragment task instances in a given block will be derived from the same micro-tile, but as discussed above, according to one aspect of this method, although this may still occur, the system is not limited to creating blocks in that way - in other words, at least one block of shader fragment task instances so created may include shader fragment task instances from more than one micro-tile (where these different micro-tiles are derived from the same parsing step, i.e., from different parts of the same buffer content).
[0094] At step S612, shader fragment task instance blocks can then be aggregated into shader tasks. That is, shader fragment task instance blocks related to the same fragment shader can be aggregated together in a task for parallel processing. The blocks aggregated into the same task can be derived from the same primitive or from different primitives that call the same fragment shader. As explained above, if other criteria are also used to determine which blocks to include in the same task, multiple tasks related to the same fragment shader can be created.
[0095] Finally, at step S614, the shader tasks are processed, for example, by the texturing / shading unit 120 in the example system described to produce shaded fragments. In other words, the shader fragment task instance blocks within the task are shaded in parallel to produce shaded fragments.
[0096] However, randomly combining fragment shader task instances from different micro-tiles (with the same FSR value) into blocks to form instance blocks is not the most efficient option.
[0097] Instead, to further increase the chance that fragment instances from different micro-tiles are assembled in the same block in a way that facilitates incremental computation during shading and texturing, the micro-tiles can be released in an order that preserves the locality of the micro-tiles, such as Z-order or Morton order. Figure 7 An example showing such an order.
[0098] Figure 7 A buffer 702 (e.g., maintained by the sampling unit 117) corresponding to at least a portion of the rendering space is schematically shown. For example, in the tile-based system described above, the buffer 702 can correspond to a tile. It should be noted that the buffer 702 is actually a memory region and can store values in a different way from the 2D layout shown Figure 7 which depicts the buffer in a way that aids understanding. Figure 7
[0099] The example buffer 702 is shown as an array divided into 4×4 (i.e., 16) sections 704, each section outlined in thick dashed lines in Figure 7 from which micro-tiles are created and posted to the texturing / shading unit 120. Each section 704 can store the sampling results at an array of sampling positions - i.e., sampler fragments 706 corresponding to each sampling position, which are outlined by thin continuous lines in Figure 7
[0100] In the depicted example, section 704 corresponds to an array of 4×4 (i.e., 16) sampling positions with corresponding sampler fragments 706. However, it should be noted that sampler fragments 706 do not necessarily exist for every sampling position - i.e., there may be sampling positions that are not overlapped by any primitive and thus do not generate sampler fragments. Thus, each section may contain a different number of sampler fragments 706, up to the number of sampling positions corresponding to that section (16 in the example of Figure 7 ).
[0101] Microtiles are created only for sections 704 that contain sampler fragments 706. Subsequently, according to this aspect, microtiles are issued for each FSR value associated with those sampler fragments 706. For example, if all sampler fragments 706 have the same FSR value, only one microtile will be issued, but if different sampler fragments 706 have different FSR values, multiple microtiles will be issued, one for each different FSR value. In this way, if a primitive contributes sampler fragments 706 with different FSR values to a particular section 704 of buffer 702, the sampler fragments within that section 704 will be separated into different microtiles. Thus, each microtile itself includes an array of (at most) 4×4 (i.e., 16) sampler fragments with a single FSR value (but potentially associated with different fragment shaders, as discussed above).
[0102] Figure 7 The arrows in show an example order in which microtiles generated from section 704 may be issued for shading. As mentioned above, this is of the Morton order or Z-order type, which preserves the locality of the microtiles. However, similar locality-preserving orders (e.g., using N-order) may be used in other examples. By following this order sequentially for each FSR value, the locality of microtiles with the same FSR value is maintained, thus facilitating the likelihood that task instances from different microtiles can be grouped together into instance blocks.
[0103] Figure 8 shows an exemplary method for generating microtiles. That is, Figure 8 is an example method for performing Figure 6 step S606.
[0104] Figure 8Beginning at S802, buffer 702 is refreshed at this time. Subsequently, the buffer is traversed or parsed to generate microtiles. In the tile-based system described above, an efficient way to do this in terms of memory access is to consider the buffer tiles one primitive tile at a time (although this is just an example and not necessarily the case - instead, sampler fragments can be collected into microtiles without considering tiles). Thus, for example, the first primitive tile associated with the tile is identified from the control list of the tile associated with buffer 702 (S804).
[0105] Subsequently, a first pass is begun, in which buffer 702 is traversed to identify sampler fragments corresponding to the primitives stored in the primitive tile. When the first sampler fragment corresponding to the primitive tile is identified (S806), the FSR value of the sampler fragment is determined, and the pass continues (S808), restricted to finding sampler fragments having the same first FSR value (and still from the same primitive tile), and any identified sampler fragments are published as microtiles. In some cases, corresponding to steps S806 and S808, the process of identifying sample fragments and publishing microtiles can be further divided into a two-stage process, in which first a set of relevant primitives in the buffer is identified (e.g., a set of primitives from the same primitive tile), and then the relevant (i.e., having the same FSR value) sampler fragments of the set of primitives are identified and published as microtiles. This can reduce the total number of primitives contributing to a particular microtile, which in turn can reduce the number of different fragment shaders that the content of a given microtile may involve. Thus, this can assist in later grouping of shader fragments.
[0106] Once the pass is complete, and assuming at least one sampler fragment is identified, S810 - "Yes", the buffer is scanned again to look for other sampler fragments corresponding to the primitives from the first primitive tile. In this scan, the sampler fragments from the previous scan are no longer considered 'valid' and thus will not be re-identified.
[0107] As long as there are such other sampler fragments, the process will continue to iterate through steps S806 to S810 to perform repeated scans of the buffer. In each iteration, the FSR value of the first valid sample identified from the relevant primitive tile will be used to constrain the remainder of the pass until no more samples corresponding to the primitive tile are identified.
[0108] In this case, S810 - "No", if there is another primitive tile associated with the tile, S812 - "Yes", the next primitive tile will be selected, S814, and the process will again iterate through steps S806 to S810 for the next primitive tile. Eventually, when there are no more primitive tiles remaining, S812 - "No", the refresh of the buffer is complete, S816.
[0109] As a result, the output of the refresh buffer 702 is a microtile stream. In the example described, the constrained results of performing each pass for a particular FSR value means that microtiles with the same FSR value are issued together (if the search is also constrained to be performed on a per-primitive-block basis, among the microtiles associated with each particular primitive block). These microtiles are then received by the texturing / shading unit 120 in the same order. Figure 9 An example method showing how the texturing / shading unit 120 processes these microtiles to create a shader fragment task instance block is shown in more detail in Figure 6 steps S608 and S610.
[0110] Figure 9 The method begins by receiving microtiles at S902. For the purposes of this example, it is assumed that the received microtiles have the same FSR value, and Figure 9 the method is shown as a continuous process, although it will be understood that the process will end if there are no more microtiles to be processed (e.g., due to a change in the FSR value, in which case the method can be repeated for that batch Figure 9 or because all available microtiles have been processed).
[0111] The next step is to analyze the microtiles (as in S608) to identify shader fragment task instances, S904. For the purposes of the first iteration, it is assumed that the microtiles have a different FSR value from the previously processed microtiles. Using the FSR value associated with the microtiles and, when applicable, the anti-aliasing settings, the texturing / shading unit 120 can group sampler fragments to create shader fragments of an appropriate size, from which shader fragment task instances are then generated. If a shader fragment would otherwise only cover empty sample positions (i.e., there are no sampler fragments), no shader fragment is generated.
[0112] The shader fragment task instances are then grouped (as in S610) into one or more (complete or incomplete) shader fragment instance blocks, S906. In the example, each block is a 2×2 (i.e., 4) task instance block. The number of blocks will depend on the number of shader fragment task instances identified in S904, which in turn will depend on the number of sampler fragments and the FSR value associated with them (and thus the microtiles).
[0113] Processing a single microtile may be able to generate one or more complete shader fragment instance blocks (if the microtile is completely covered by a small number of primitives and the FSR value is low - for example, an FSR value of 1×1 may generate four fragment task instances from a 4×4 microtile covered by a single primitive), and any complete blocks may then be released for processing. However, in other cases, a microtile may not be able to generate sufficient task instances arranged in a proper way to fill the block - for example, a 4×4 microtile with an FSR value of 2×4, even if completely covered by a single primitive, will only generate two task instances; and the same 4×4 microtile with an FSR value of 1×4 will generate four task instances but will not be arranged in a 2×2 block. In such cases, it may be possible to use task instances from the next (subsequent) microtile to complete the block, so it is not yet certain whether to release an incomplete block.
[0114] Figure 10 and Figure 11 Intuitively shows this. In Figure 10 it shows four microtiles 1004 with an FSR value of 2×4 being processed. The order in which the microtiles are released (i.e., read out from buffer 702) is shown by the arrows. The dashed lines are used to indicate a set of sampler fragments 1006 corresponding to shader fragments 1008. The shader fragments are filled with different patterns to indicate the different shader fragment instance blocks that can be formed. In this case, the four shader fragments from the two leftmost microtiles can form a 2×2 block (indicated by the filled pattern with lines slanting to the right), and the four shader fragments from the two rightmost microtiles can form another 2×2 block (indicated by the filled pattern with lines slanting to the right). Thus, after processing the first (i.e., top - left) microtile 1004 (and assuming all shader fragments are present and related to the same fragment shader), there will be an incomplete block; processing the second tile will start a new block while still having the possibility of completing the previous block; processing the third tile will complete the first block; and processing the fourth microtile will complete the second block.
[0115] Similarly, in Figure 11Among them, four micro tiles 1104 with an FSR value of 1×4 are being processed. Similarly, the dashed lines are used to indicate a set of sampler fragments 1106 corresponding to the shader fragment 1108. The shader fragments are filled with different patterns to indicate different shader fragment instance blocks that can be formed. In this case, two leftmost shader fragments in each of the two leftmost micro tiles can form a 2×2 block (indicated by the filled pattern with right-slanted lines), two rightmost shader fragments in each of the two leftmost micro tiles can form a 2×2 block (indicated by the filled pattern with vertical lines), two leftmost shader fragments in each of the two rightmost micro tiles can form a 2×2 block (indicated by the filled pattern with right-slanted lines), and two rightmost shader fragments in each of the two rightmost micro tiles can form another 2×2 block (indicated by the filled pattern with cross lines). Thus, after processing the first (i.e., upper left corner) micro tile 1104 (and assuming that all shader fragments are present and related to the same fragment shader), there will be two incomplete blocks; processing the second tile will start two other blocks while still having the possibility of completing the previous block; processing the third tile will complete the previous two blocks; and processing the fourth micro tile will complete the remaining two blocks.
[0116] Therefore, return Figure 9 , at step S906, any complete shader fragment instance blocks can be issued, but any incomplete blocks that can be completed by subsequent shader fragments from subsequent micro tiles remain pending. Of course, there are still cases where some blocks cannot be completed, for example, when not all possible shader fragments are present in the relevant micro tile (i.e., because the relevant sampling points are not covered by the relevant primitive). Thus, in these cases, it is still possible for the system to issue incomplete blocks (with helper instances). However, by enabling the construction of shader fragment instance blocks, the frequency of such scenarios is significantly reduced compared to the case where only task instances from a single micro tile are available to construct shader fragment instance blocks.
[0117] The method then continues to analyze the next microtile S908 and identifies shader fragment task instances from that microtile in a manner similar to step S904. The method then moves to step S910, which is similar to step S906 except that there may already be incomplete blocks. Thus, in this step, shader fragment instance blocks are created from the newly identified task instances and the incomplete blocks from the previous step. In some cases, it is possible to add the newly identified task instances to a previously incomplete block, which may or may not complete the block. For example, in the case of 4×4 microtiles and 4×4 FSR values, each microtile will create only one shader fragment and—assuming all shader fragments are associated with the same fragment shader—the first microtile in the sequence will start a new incomplete block, the second and third microtiles will each add another task instance to the incomplete block but not complete the block, and the fourth microtile will add the final fourth task instance, which completes the block. In other cases, there may be existing incomplete blocks that cannot be added and additional incomplete blocks are created (such as the case when processing the second microtile in each of Figure 10 and Figure 11 ). In still other cases, there may be no incomplete blocks from the previous step. In any case, any complete blocks are published, as are any incomplete blocks that will not be completed (as discussed above, e.g., due to empty / missing shader fragments caused by missing primitive coverage), while any incomplete blocks that may still be completed remain pending.
[0118] The method then returns to step S908, looping through steps S908 and S916 until all microtiles have been processed (at which point any remaining incomplete blocks may also be published). As mentioned above, this may be because the next microtile is associated with a different FSR value, in which case the method may start over separately for those microtiles.
[0119] Thus, the texturing / shading unit 120 can group shader fragments into shader fragment instance blocks with fewer helper instances. Since these blocks are then grouped to form larger tasks to be processed by the texturing / shading unit 120 (e.g., in a SIMD processor), this allows more useful work to be performed in parallel and thus helps with faster and more efficient processing.
[0120] In the context of FSR, another way to develop an instance block shading method is to allow instance blocks with different FSR values to be combined into the same larger task.
[0121] In a system without FSR, the SIMD shading processor may expect not only that task instances within an instance block will be related to a single fragment size, but also that each instance block will be related to fragments of the same size. That is, in the absence of FSR, the size of the fragments being rendered will be invariant (although it may vary due to rendering, e.g., due to anti-aliasing), and thus it is relatively easy to combine instance blocks into larger tasks in a way that effectively uses the entire width of the SIMD processor. In this case, shading processes that depend on the sampling pattern used to create the fragments, such as interpolation, can be performed easily and reliably. For example, this can be done by providing information that defines a single sampling pattern, which is used by all instance blocks gathered into a task as part of the overall task information. This information can then be used for all fragments processed in parallel. However, the introduction of FSR complicates this.
[0122] In one approach, only instance blocks related to the same FSR value can be combined into larger tasks and thus run in parallel. Figure 8 An example method promotes this approach by creating micro-tiles from a buffer on a per-FSR basis, which in turn results in creating shader fragment instance blocks in batches with the same FSR value, which in turn makes it easy to gather shader fragment instance blocks with the same FSR value into a task. The advantage of this is that only a relatively small change to the way tasks are submitted to the processor is required, since each fragment within the gathered blocks will have the same size, similar to a regular system. This means that operations that depend on fragment size or sampling pattern, such as interpolation, can still be handled by providing information that defines a single sampling pattern, which is used by all instance blocks gathered into a task as part of the overall task information. However, a disadvantage of this method is that it becomes more difficult to gather enough instance blocks together to keep the processor efficiently busy. That is, even Figure 8 the example results in shader fragment instance blocks being created in batches with the same FSR value, there may still be a relatively small number of shader fragment instance blocks with a particular FSR value (e.g., compared to the case where there is no variable FSR / all fragments are the same size). Thus it may be difficult to create full tasks, or at least it may be the case more frequently that tasks cannot be fully filled (i.e., to maximize processor parallelism). In other words, submitting a single task consisting of four instance blocks is more efficient than submitting two tasks each consisting of two instance blocks, but the latter case becomes more likely if there are multiple FSR values and blocks with different FSR values must remain in separate tasks.
[0123] Thus, another approach is to create tasks that include task instances related to different fragment sizes. This can be done by grouping instance chunks related to different fragment sizes into the same task for processing. However, this then affects processes such as the interpolation described above, which depends on the fragment size and / or sampling pattern, as now the instances within the task can have different sizes and different sampling patterns. Thus, the present method provides relevant information (such as the FSR value from which the sampling pattern and fragment size can be derived) in such a way that each task instance can be processed regardless of the fragment size. Specifically, the relevant information can be provided at the per-instance-chunk granularity.
[0124] It might be thought that the greatest flexibility would be achieved by allowing the instance chunks to include fragments of different sizes and then providing relevant information about the sample size and sampling pattern at the task-instance granularity. However, defining information at a higher granularity has drawbacks. Specifically, there is a data (and thus bandwidth) overhead because there is additional data to be transmitted as part of each task, and thus the size of each task is increased. For a larger SIMD width, this can be a significant overhead and may be particularly significant in devices where memory and bandwidth are very precious, such as mobile devices. Additionally, allowing different fragment sizes within the instance chunks also complicates problems such as the incremental calculations described above. For these reasons, a preferred specific implementation is to provide the sampling information at the per-instance-chunk granularity, where the task instances within a given instance chunk all have the same fragment size. This enables the creation of larger tasks (i.e., combining more instance chunks together), while minimizing the additional data overhead and avoiding the need for more extensive changes to the existing system. This enables the parallel processing of fragment task instances associated with different fragment shading rates using a common shader program, i.e., the fragment task instances can be processed simultaneously in the SIMD processor. That is, the fragment task instances associated with different fragment shading rates can be present in different lanes of the SIMD processor.
[0125] Figure 12 A method of combining fragment shader instance chunks into a task that includes task instances related to different fragment sizes, where each of the fragment shader instance chunks individually contains a task related to a specific FSR value.
[0126] Figure 12 The method starts (S1202) from the point where the microtile to be sent to the texturing / shading unit 120 is parsed from a buffer (such as buffer 704), which is equivalent to Figure 6 sub-step S606 of step S603 in Figure 8 A more detailed example for implementing this parsing is presented. For the purpose of this example, it suffices to note that, as with Figure 8Compared with the example of Figure 8 the separate passes of Figure 8 help create fragment instance blocks in batches, as explained above, but this is less important if the final task can be composed of blocks associated with different FSR values. Additionally, separating the FSR passes may cause sampler fragments associated with the same primitive to be separated in the issued microtile stream, making it less likely that shader fragments created later in the pipeline from these sampler fragments will be collected into the same task. Therefore, by parsing the buffer without considering the FSR value, sampler fragments associated with the same primitive are more likely to be grouped together in the same microtile, making it more likely that shader fragments derived from these samples will be grouped (via their corresponding shader fragment task instance blocks) into the same task. That is, this helps create better-filled tasks.
[0127] The resulting microtiles are sent to the texturing / shading unit 120, where the microtiles are analyzed to identify shader fragment task instances, which are then arranged into block S1204. This is equivalent to Figure 6 steps S608 (a sub-step of step S603) and S610 (a sub-step of step S609) in Figure 9 A specific example implementation is presented in more detail for this step. For the purposes of this example, it is sufficient to note that the effect of having multiple FSR values in the same microtile is that shader fragments of different sizes (i.e., associated with different FSR values) can be identified from the same microtile. This in turn may require more memory / buffer to be provided at this step to allow more shader fragment instance blocks to remain pending, as it is expected that the microtile will contain shader fragments that will be collected into different shader fragment instance blocks for different FSR values (while maintaining the relationship that one shader fragment instance block contains shader fragment task instances related to one fragment shader). Figure 13This is shown in the figure, which shows a microtile 1302 containing sampler fragments with an FSR value of 2×4 (shading lines slanting down to the right) and an FSR value of 1×4 (shading lines being vertical lines). The figure shows how the shader fragments obtained from the microtile 1302 are equivalent to the shader fragments obtained from two single-FSR microtiles 1304 and 1306, where it is shown how the range of the obtained shader fragments is compared with the range of the sample fragments. Thus, the resulting shader fragments are associated with different FSR values and will contribute to the parallel creation / population of the different shader fragment instance blocks that will be needed. In contrast, in the case where the incoming microtiles with different FSR values are separated into batches (i.e., in the time sense) due to buffer flushing being performed in each FSR pass, there is no need to maintain partially filled shader fragment instance blocks for different FSR values. Instead, a change in the FSR value of the incoming microtile triggers the flushing of any partially filled shader fragment instance blocks for the previous FSR value.
[0128] Thus, after step S1204, the shader fragment task instances have been collected into shader fragment instance blocks, but compared with the Figure 9 method, the blocks are not necessarily released in groups with the same FSR value, but are more mixed together. Next, at step S1206, those blocks can be collected such that the blocks (and thus the shader fragment task instances) that require a common fragment shader program are grouped into a shading task. As previously explained, other criteria (such as sharing some other state) can be additionally applied to determine which blocks are grouped into larger shading tasks. Also as mentioned above, since this method allows blocks associated with different FSR values (but the same shader) to be grouped into a shader task, the most efficient way to provide the relevant information about the sample size and sampling pattern is to maintain that information at the per-instance-block granularity within the created shader tasks. This allows the subsequent processing of the shader tasks to process the individual blocks in almost the same way as before, especially in terms of calculating the incremental values between the computational task instances, while also allowing the parallel processing of blocks associated with different FSR values.
[0129] Figure 14A computer system in which the graphics processing system described herein can be implemented is shown. The computer system includes a CPU 1402, a GPU 1404, a memory 1406, and other devices 1414, such as a display 1416, speakers 1418, and a camera 1422. One or more processing blocks 1410 (e.g., corresponding to processing blocks 104 and 106) can be implemented on the GPU 1404 as well as on a neural network accelerator (NNA) 1411. In other examples, the processing blocks 1410 can be implemented on the CPU 1402 or within the NNA 1411. The components of the computer system can communicate with each other via a communication bus 1420. A storage device 1412 (corresponding to memory 102) is implemented as part of the memory 1406.
[0130] Although Figure 14 a specific implementation of the graphics processing system is shown, it will be understood that a similar block diagram can be drawn for an artificial intelligence accelerator system—for example, by replacing the CPU 1402 or the GPU 1404 with a neural network accelerator (NNA) 1411, or by adding an NNA as a separate unit. In such cases, again, the processing blocks 1410 can be implemented in the NNA.
[0131] Figure 1 The graphics processing system is shown as including a number of functional blocks. This is merely illustrative and is not intended to define a strict partitioning between the different logical elements of such entities. Each functional block can be provided in any suitable manner. It should be understood that the intermediate values described herein that are formed by the graphics processing system do not need to be physically generated by the graphics processing system at any point in time and can merely represent logical values that conveniently describe the processing performed by the graphics processing system between its input and output.
[0132] The graphics processing units described herein may be embodied as hardware on an integrated circuit. The graphics processing units described herein may be configured to perform any of the methods described herein. In general, any of the functions, methods, techniques, or components described above may be implemented in software, firmware, hardware (e.g., fixed logic circuitry), or any combination thereof. The terms "module," "functionality," "component," "element," "unit," "block," and "logic" may be used herein generally to denote software, firmware, hardware, or any combination thereof (the term "block" is also used to refer to a group of aggregated shader fragment task instances, and the different usage is apparent from the context). In the case of a software implementation, a module, functionality, component, element, unit, block, or logic represents program code that, when executed on a processor, performs the specified task. The algorithms and methods described herein may be executed by one or more processors executing code that causes the processors to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical discs, flash memory, hard disk memory, and other memory devices that can store instructions or other data using magnetic, optical, and other technologies and that can be accessed by a machine.
[0133] As used herein, the terms computer program code and computer-readable instructions refer to any kind of executable code for a processor, including code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, bytecode, code that defines an integrated circuit (e.g., a hardware description language or netlist), and code expressed in a programming language code such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, or executed in a virtual machine or other software environment, causes a processor of a computer system that supports the executable code to perform the task specified by the code.
[0134] A processor, computer, or computer system can be any kind of device, machine, or dedicated circuit, or a collection or part thereof, that has the processing power such that instructions can be executed. A processor can be or include any kind of general-purpose or special-purpose processor, such as a CPU, GPU, NNA, system-on-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), etc. A computer or computer system can include one or more processors.
[0135] The present invention also intends to cover software that defines the configuration of hardware as described herein, such as HDL (Hardware Description Language) software, for example, for designing integrated circuits or for configuring programmable chips to implement the required functions. That is to say, a computer-readable storage medium encoded with computer-readable program code in the form of an integrated circuit definition data set can be provided, which, when processed (i.e., run) in an integrated circuit manufacturing system, configures the system to manufacture a graphics processing system configured to perform any of the methods described herein, or to manufacture a graphics processing system including any of the devices described herein. The integrated circuit definition data set can be, for example, an integrated circuit description.
[0136] Accordingly, a method of manufacturing a graphics processing system as described herein at an integrated circuit manufacturing system can be provided. In addition, an integrated circuit definition data set can be provided, which, when processed in an integrated circuit manufacturing system, causes the method of manufacturing a graphics processing system to be executed.
[0137] The integrated circuit definition data set can be in the form of computer code, such as a netlist, code for configuring a programmable chip, as a hardware description language suitable for manufacturing at any level in an integrated circuit, including as register transfer level (RTL) code, as a high-level circuit representation (such as Verilog or VHDL), and as a low-level circuit representation (such as OASIS(RTM) and GDSII). A higher-level representation (such as RTL) that logically defines the hardware suitable for manufacturing in an integrated circuit can be processed at a computer system configured to generate a manufacturing definition of an integrated circuit in the context of a software environment that includes definitions of circuit elements and rules for combining those elements to generate a manufacturing definition of an integrated circuit as so defined by the representation. As is typically the case where software is executed at a computer system to define a machine, one or more intermediate user steps (such as providing commands, variables, etc.) may be required to configure the computer system to generate a manufacturing definition of an integrated circuit to execute code that defines the integrated circuit to generate the manufacturing definition of the integrated circuit.
[0138] Now will be directed to Figure 15 Describe an example of processing an integrated circuit definition data set at an integrated circuit manufacturing system to configure the system to manufacture a graphics processing system.
[0139] Figure 15Shows an example of an integrated circuit (IC) manufacturing system 1502, which is configured to manufacture a graphics processing system as described in any of the examples herein. Specifically, the IC manufacturing system 1502 includes a layout processing system 1504 and an integrated circuit generation system 1506. The IC manufacturing system 1502 is configured to receive an IC definition data set (e.g., defining a graphics processing system as described in any of the examples herein), process the IC definition data set, and generate an IC according to the IC definition data set (e.g., which embodies a graphics processing system as described in any of the examples herein). By processing the IC definition data set, the IC manufacturing system 1502 is configured to manufacture an integrated circuit that embodies a graphics processing system as described in any of the examples herein.
[0140] The layout processing system 1504 is configured to receive and process an IC definition data set to determine a circuit layout. Methods for determining a circuit layout from an IC definition data set are known in the art and may involve, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated, e.g., in terms of logic components (such as NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined from the gate-level representation of the circuit. This can be done automatically or with user participation to optimize the circuit layout. When the layout processing system 1504 has determined the circuit layout, it can output a circuit layout definition to the IC generation system 1506. The circuit layout definition can be, for example, a circuit layout description.
[0141] As is known in the art, the IC generation system 1506 generates an IC according to the circuit layout definition. For example, the IC generation system 1506 can implement a semiconductor device manufacturing process for generating an IC, which may involve a multi-step sequence of lithography and chemical processing steps during which an electronic circuit is gradually formed on a wafer made of semiconductor material. The circuit layout definition can be in the form of a mask, which can be used in a lithography process to generate an IC according to the circuit definition. Alternatively, the circuit layout definition provided to the IC generation system 1506 can be in the form of computer-readable code, which the IC generation system 1506 can use to form a suitable mask for generating the IC.
[0142] The different processes performed by the IC fabrication system 1502 may all be implemented in one location, e.g., by one party. Alternatively, the IC fabrication system 1502 may be a distributed system such that some processes may be performed at different locations and may be performed by different parties. For example, some of the following stages may be performed at different locations and / or by different parties: (i) synthesizing RTL code representing an IC definition data set to form a gate-level representation of the circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask according to the circuit layout; and (iv) manufacturing an integrated circuit using the mask.
[0143] In other examples, processing of an integrated circuit definition data set at an integrated circuit fabrication system may configure the system to fabricate a graphics processing system without processing the IC definition data set to determine a circuit layout. For example, the integrated circuit definition data set may define the configuration of a reconfigurable processor such as an FPGA, and processing of the data set may configure the IC fabrication system to generate (e.g., by loading configuration data into the FPGA) a reconfigurable processor having the defined configuration.
[0144] In some embodiments, when processed in an integrated circuit fabrication system, an integrated circuit fabrication definition data set may cause the integrated circuit fabrication system to generate a device as described herein. For example, configuring the integrated circuit fabrication system in the manner described above for Figure 15 by the integrated circuit fabrication definition data set may cause a device as described herein to be fabricated.
[0145] In some examples, the integrated circuit definition data set may include software that runs on hardware defined at the data set, or software that runs in combination with the hardware defined at the data set. In the Figure 15 example shown, the IC generation system may also be configured by the integrated circuit definition data set to load firmware onto the integrated circuit according to program code defined at the integrated circuit definition data set when fabricating the integrated circuit, or otherwise provide program code for use with the integrated circuit.
[0146] Compared with known specific implementations, the specific implementation of the concepts set forth in the present application in devices, apparatuses, modules, and / or systems (and in the methods implemented herein) can bring about performance improvements. The performance improvements can include one or more of increased computing performance, reduced latency, increased throughput, and / or reduced power consumption. During the manufacture of such devices, apparatuses, modules, and systems (e.g., in integrated circuits), a trade-off can be made between performance improvements and physical implementation, thereby improving the manufacturing method. For example, a trade-off can be made between performance improvements and layout area to match the performance of known specific implementations but use less silicon. For example, this can be done by reusing functional blocks serially or sharing functional blocks among the elements of a device, apparatus, module, and / or system. Conversely, the concepts set forth in the present application that bring about improvements in the physical implementation of devices, apparatuses, modules, and systems (such as reduced silicon area) can be traded off against performance improvements. This can be done, for example, by manufacturing multiple instances of a module within a predefined area budget.
[0147] The applicant hereby independently discloses each individual feature described herein and any combination of two or more such features to the extent that such features or combinations can be implemented by a person of ordinary skill in the art based on the entire specification, regardless of whether such features or combinations of features solve any of the problems disclosed herein. In view of the foregoing description, it will be apparent to those skilled in the art that various modifications can be made within the scope of the present invention.
Claims
1. A method for rendering a scene formed by primitives in a rendering space of a graphics processing system, the method comprising: A rendering stage, the rendering stage comprising the steps of: Receiving (S602) data describing two or more primitives and two or more associated fragment shading rates to be used during rendering; Identifying (S603) shader fragment task instances to be shaded from the primitives, wherein the shader fragment task instances include shader fragment task instances associated with a first fragment shading rate and shader fragment task instances associated with a second fragment shading rate; Combining (S612) the shader fragment task instances into a shading task, the shading task including fragment task instances associated with the first fragment shading rate and fragment task instances associated with the second fragment shading rate, and wherein the fragment task instances combined into the shading task require a common shader program; and Processing (S614) the shading task, wherein processing the shading task includes processing the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate in parallel using the common shader program.
2. The method according to claim 1, wherein the method further comprises: Using the received data describing the primitives and two or more associated fragment shading rates to fill a buffer indicating sampling positions of the primitives within a region of the rendering space; And Parsing (S1202) the buffer to generate microtiles, each microtile corresponding to an array of sampling positions within the region and containing sampler fragments from the two or more primitives; And at least one microtile, the at least one microtile containing sampler fragments associated with the first fragment shading rate and sampler fragments associated with the second fragment shading rate; Wherein the identifying step includes identifying the fragment task instances to be shaded from the microtiles.
3. The method according to claim 1 or claim 2, wherein the method further comprises: Arranging (S1204) the fragment task instances into blocks, wherein the fragment task instances within any given block are associated with a common fragment shading rate; And wherein the step of combining the fragment task instances includes combining blocks of fragment task instances that require a common shader program into a shading task, and wherein the shading task includes a block of fragment task instances associated with the first fragment shading rate and a block of fragment task instances associated with the second fragment shading rate.
4. The method according to claim 3, wherein combining shader fragment task instances that require a common shader program into a shading task includes maintaining data indicating the fragment shading rate associated with each block combined into the shading task.
5. The method according to claim 3, wherein the shader fragment task instances arranged into a given block include adjacent shader fragments in the rendering space.
6. The method according to claim 2, wherein filling the buffer includes performing hidden surface removal to identify and not store one or more sampler fragments within the rendering space that do not contribute to the scene to be rendered.
7. The method according to claim 1 or 2, wherein processing the shading task includes using a SIMD processor.
8. The method according to claim 4, wherein processing the shading task includes using a SIMD processor, and wherein processing the shading task includes using the data indicating the fragment shading rate associated with the block to perform incremental calculations on shader fragment task instances within the block.
9. The method according to claim 1 or 2, wherein the method further includes a geometry processing stage, wherein the geometry processing stage includes transforming the primitive into the rendering space, storing data related to the transformed primitive, and / or determining and storing control flow data indicating which primitives are related to different regions of rendering the rendering space.
10. A graphics processing system configured to render a scene formed by primitives in a rendering space, wherein the graphics processing system includes rendering logic configured to perform the following operations: Receive data describing two or more primitives and two or more associated fragment shading rates to be used during rendering; Identify shader fragment task instances to be shaded from the primitives, wherein the shader fragment task instances include shader fragment task instances associated with a first fragment shading rate and shader fragment task instances associated with a second fragment shading rate; Combine the shader fragment task instances into a shading task, the shading task including fragment task instances associated with the first fragment shading rate and fragment task instances associated with the second fragment shading rate, and wherein the fragment task instances combined into the shading task require a common shader program; and Process the coloring task, where, The rendering logic is further configured to process the shading task by: using the common shader program to process in parallel the fragment task instances associated with the first fragment shading rate and the fragment task instances associated with the second fragment shading rate.
11. The graphics processing system according to claim 10, wherein the rendering logic is further configured to: Fill a buffer indicating the sampling positions of the primitive within the region of the rendering space using the received data describing the primitive and the shading rates of two or more associated fragments; And Parse (S1202) the buffer to generate microtiles, each microtile corresponding to an array of sampling positions within the region and containing sampler fragments from the two or more primitives; And at least one microtile, the at least one microtile containing sampler fragments associated with the first fragment shading rate and sampler fragments associated with the second fragment shading rate; And Identify the shader fragment task instances to be shaded from the microtiles.
12. The graphics processing system according to claim 10 or claim 11, wherein the rendering logic is further configured to: Arrange the fragment task instances into blocks, wherein the fragment task instances within any given block are associated with a common fragment shading rate; Shader fragment task instances are combined by combining fragment task instance blocks associated with the first fragment shading rate and fragment task instance blocks associated with the second fragment shading rate into a shading task.
13. The graphics processing system according to claim 12, wherein the rendering logic is further configured to: Combine shader fragment task instances that require a common shader program into a shading task, the shading task maintaining data indicating the fragment shading rate associated with each block in the shading task.
14. The graphics processing system according to claim 12, wherein the rendering logic is further configured to combine adjacent shader fragment task instances in the rendering space into blocks.
15. The graphics processing system according to claim 11, wherein the rendering logic is further configured to perform hidden surface removal when filling a buffer to identify and not store one or more sampler fragments within the rendering space that do not contribute to the scene to be rendered.
16. The graphics processing system according to claim 10 or 11, wherein The rendering logic is further configured to process the shading task using a SIMD processor.
17. The graphics processing system according to claim 13, wherein, The rendering logic is further configured to process the shading task using a SIMD processor, and wherein the rendering logic is further configured to process the shading task using the data indicating the fragment shading rate associated with a block to perform incremental calculations on shader fragment task instances within the block.
18. The graphics processing system according to claim 10 or 11, the graphics processing system further comprising geometry processing logic, wherein the geometry processing logic includes logic configured to transform the primitive into the rendering space, store data related to the transformed primitive in memory, and / or determine and store control flow data indicating which primitives are related to rendering different regions of the rendering space.
19. The graphics processing system according to claim 10 or 11, wherein the graphics processing system is embodied in hardware on an integrated circuit.
20. A method of manufacturing a graphics processing system according to any one of claims 10 to 19 using an integrated circuit manufacturing system, the method comprising: Processing a computer-readable description of the graphics processing system using a layout processing system to generate a circuit layout description of an integrated circuit including the graphics processing system; And Manufacturing the graphics processing system using an integrated circuit generation system according to the circuit layout description.
21. A computer-readable storage medium having computer-readable code stored thereon, the computer-readable code being configured to cause a method according to any one of claims 1 to 9 or claim 20 to be performed when the code is run.
22. A computer-readable storage medium having an integrated circuit definition data set stored thereon, which when processed in an integrated circuit manufacturing system configures the integrated circuit manufacturing system to manufacture a graphics processing system according to any one of claims 10 to 19.
23. An integrated circuit manufacturing system configured to manufacture a graphics processing system as described in any one of claims 10 to 19, the integrated circuit manufacturing system comprising: A computer-readable storage medium having stored thereon a computer-readable description of the graphics processing system; A layout processing system configured to process the computer-readable description to generate a circuit layout description of an integrated circuit incorporating the graphics processing system; And An integrated circuit generation system configured to manufacture the graphics processing system according to the circuit layout description.
Citation Information
Patent Citations
Hybrid hierarchy for ray tracing
CN109255828A