Use the weighted average of the properties of triangles to merge fragments of thick pixel coloring
By interpolation of overlay weighted averages of vertex attributes at the center of the thick pixels, and buffering of coarse pixel shading quaternary fragments in the graphics processing unit, the problem of limited efficiency in modern rendering workloads is solved, achieving a more efficient shading process and reducing artifacts.
Patent Information
- Application Number
- CN202210347518.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-12-04
- Filing Date
- 2016-11-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2036-11-07
AI Technical Summary
Coarse pixel shading technology is limited in efficiency in modern rendering workloads that handle small triangles and can lead to redundant pixel shading execution and instantaneous artifacts.
Combining multiple primitives by interpolation of overlay weighted averages of vertex attributes at the center of the thick pixels reduces the number of times the coarse pixel shader is executed, and buffering the coarse pixel shading quaternary fragments in the graphics processing unit to improve efficiency.
Improves the efficiency of coarse pixel shading, reduces redundant shading execution, and reduces or eliminates instantaneous artifacts.
Smart Images

Figure CN114549518B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for invention titled "Merging Fragments of Coarse Pixel Shading Using Weighted Averages of Triangle Attributes", with an international filing date of November 7, 2016, an international application number of PCT / US2016 / 060776, and a national stage entry application number in China of 201680070840.2. Background Art
[0002] In coarse pixel shading, a given frame or picture may have different shading rates. For example, certain regions of the frame or picture may have a lower shading rate, such as less than once per pixel, while in another region, the shading rate may be once per pixel, and in yet another place, the shading rate may exceed once per pixel. Examples of where the shading rate can be reduced include regions with motion and camera defocus, peripherally blurred regions, and generally any situation where the perceived visual detail is reduced in any way, or any situation where shading artifacts are generally less apparent.
[0003] Coarse Pixel Shading (CPS) is a technique that can reduce the shading rate in the rasterization pipeline. CPS merges blocks of quad-fragments from the same primitive into a coarsely shaded quad. For example, by merging a block of 4x4 fragments (4 quads) into a single shaded quad (2x2 coarse fragment), the number of shading evaluations can be reduced to 25%.
[0004] After rasterizing a triangle, coarse pixel shading reduces the shading cost. However, some factors can limit the efficiency of coarse pixel shading due to the possible generation of redundant pixel shading executions. Such efficiency limitations can occur, for example, in relation to scene depth complexity, partially covered pixels, and the quad-fragment-based scheduling of pixel shaders for finite differences.
[0005] It often happens that pixels covered by the same surface are shaded multiple times when they are covered by quad-fragments of multiple rasterized primitives. This is because coarse pixel shading only works within a single rasterized primitive. CPS uses triangles that generate several fragments by themselves. However, modern rendering workloads are typically characterized by small triangles, in which case the benefits of CPS are spread out. Brief Description of the Drawings
[0006] Some embodiments are described with reference to the following drawings:
[0007] Figure 1 is a flowchart of an embodiment;
[0008] Figure 2 is a depiction of the merging process according to an embodiment.
[0009] Figure 3 is a flowchart for a merge unit according to one embodiment;
[0010] Figure 4 is a block diagram of a processing system according to one embodiment;
[0011] Figure 5 is a block diagram of a processor according to one embodiment;
[0012] Figure 6 is a block diagram of a graphics processor according to one embodiment;
[0013] Figure 7 is a block diagram of a graphics processor engine according to one embodiment;
[0014] Figure 8 is a block diagram of another embodiment of a graphics processor;
[0015] Figure 9 is a depiction of thread execution logic according to one embodiment.
[0016] Figure 10 is a block diagram of a graphics processor instruction format according to some embodiments;
[0017] Figure 11 is a block diagram of another embodiment of a graphics processor;
[0018] Figure 12A is a block diagram of a graphics processor command format according to some embodiments;
[0019] Figure 12B is a block diagram showing a graphics processor command sequence according to some embodiments;
[0020] Figure 13 is a depiction of an exemplary graphics software architecture according to some embodiments;
[0021] Figure 14 is a block diagram showing an IP core development system according to some embodiments; and
[0022] Figure 15 is a block diagram showing an exemplary system-on-chip integrated circuit according to some embodiments. DETAILED DESCRIPTION
[0023] Two primitives can be merged by interpolating vertex attributes at the center of a thick pixel. The input attributes are calculated as a coverage weighted average of the interpolated vertex attributes. Then the resulting input attributes are used to perform thick pixel shading. After coverage weighted interpolation, the merged primitives can be discarded (they are no longer needed).
[0024] In some embodiments, the efficiency of coarse pixel shading can be improved by reducing the number of times the coarse pixel shader is executed on a set of primitive elements. Efficiency can be improved by sharing coarse pixel shader executions across multiple primitive elements in special cases where multiple primitive elements represent the same surface. This can be done by merging multiple coarse pixel shading quads if the coarse pixel shading quad fragments are from the same draw call and correspond to the same coarse pixel on the screen but their sampling coverage does not overlap. The rasterization pipeline can be extended with a small on-chip buffer within the graphics processing unit that clusters coarse pixel shading quad fragments prior to coarse pixel shading. Pixel input attributes can be pre-filtered prior to shading of the merged quad fragments.
[0025] See Figure 1 , sequence 10 can be implemented in software, firmware, and / or hardware. In software and firmware embodiments, it can be implemented by computer-executable instructions stored in one or more non-transitory computer-readable media such as magnetic, optical, or semiconductor storage.
[0026] Initially, the rasterizer 21 tests the primitive elements against a pixel region of a given size, which in this example is a 2×2 pixel region 24, called a quad. (Other sizes of pixel regions can also be used). The rasterizer traverses the quads in a space-filling order such as Morton order. If the primitive element covers any pixels or samples, in the case of multisampling, within the quad, the rasterizer sends the quad downstream to the tile buffer 16. In some embodiments, early z-culling can be done at 14.
[0027] For a given primitive element, the tile buffer can divide the screen into tiles of 2Nx2N pixel size (referred to as shading quads) and can store all rasterized quads that fall within a single tile. A screen-aligned shading grid can be evaluated for each 2Nx2N tile.
[0028] In some embodiments, the size of the grid cells or shading quads can be restricted to a power of two, measured in pixels. For example, the cell size can be 1x1, 1x2, 2x1, 4x1, 4x2, 4x4, up to NxN, including 1xN and Nx1 and all intermediate configurations.
[0029] By controlling the size of the shading grid cells, the shading rate can be controlled at block 20. That is, the larger the cell size, the lower the shading rate of the tile.
[0030] The quads stored in the tile buffer are then grouped into shaded quads 26 consisting of groups of adjacent grid cells, such as, in one embodiment, 2x2 adjacent grid cells (block 18). The shaded quads are then shaded, and the output from the shader is written back to all covered pixels in the color buffer.
[0031] The grid size is evaluated in tiles of 2N×2N pixels each (where N is the size of the largest decoupled pixel within the resized quad and is evaluated independently for each primitive). The "largest decoupled pixel" refers to the size of the pixel when the size of the shaded quad changes. In one embodiment, four pixels form a quad and each pixel forms one quarter of the quad size.
[0032] The grid size can be controlled by a property called the scaling factor, which consists of a pair of signed values - the scaling factor Sx along the X-axis and the scaling factor Sy along the Y-axis. The scaling factors can be assigned in various ways. For example, it can be interpolated from vertex attributes or calculated from screen positions.
[0033] Using signed scaling factors can be useful if a primitive crosses the focal plane, such as in the case of a defocused camera. In this case, the vertices of the primitive may be out of focus while the interior of the primitive may be in focus. One can then assign a negative scaling factor as an attribute to vertices in front of the focal plane and a positive scaling factor to vertices behind the focal plane, and vice versa. For in-focus regions of the primitive, the scaling factor interpolates to zero and thus, a high shading rate is maintained in the in-focus regions.
[0034] The scaling factor can vary within a tile, but still a single quantized grid cell size is calculated for each tile. This can result in discontinuities in the grid size moving from tile to tile and can cause visible grid transitions.
[0035] Shading of surfaces is generally the most computationally expensive part of the rendering process in terms of both memory bandwidth and power consumption. In the rasterization pipeline, when generating an image in the frame buffer, the minimum number of shader executions should match the number of pixels: the color of each visible surface needs to be calculated. However, during actual rendering, several redundant shading calculations are performed, mainly due to (1) depth complexity, (2) multiple primitives covering the same pixel, or (3) quad-based shading scheduling due to finite differences. Depth complexity can be addressed by the z-buffer algorithm. But the remaining reasons for redundant shading calculations continue to result in a significant increase in pixel shader execution.
[0036] In coarse pixel shading, the density of pixel shader execution on large triangles is reduced. Assuming a high-density display, coarse pixel shading combines pixel blocks into coarsely shaded pixels (e.g., 2×2 coarse pixel shading for pixels from a coarse pixel), thus reducing the shading cost by 25%.
[0037] However, modern rendering workloads typically generate small primitives that are comparable in size to the coarse pixel size. Then, the benefits of coarse pixel shading are amortized in such cases. If these quadtree fragments potentially belong to the same surface, this limitation can be addressed by merging coarse quadtree fragments of different triangles before coarse pixel shading.
[0038] To this end, a clustering stage or merge unit 22 can be introduced before coarse pixel shading 30, which reuses the same coarse pixel grid on the primitives. The clustering stage reuses the same coarse pixel grid on primitives that share an edge with the same vertex attributes, have the same orientation, and have mutually exclusive coverage. The clustering stage can use a small primitive merge buffer after the triangle rasterization stage and search for coarse pixel shading quadtree fragments that can be merged. The output of this stage can be referred to as a shading cluster. Assuming the above conditions are met, only one quadtree will be generated even if more triangles in the shading cluster cover the same coarsely shaded quadtree.
[0039] When performing the standard coarse pixel shading algorithm on the generated quadtree fragments, potential temporal artifacts may occur. Typically, the vertex attributes of the triangles covering the center of each coarse pixel are interpolated to generate the coarse pixel shader input attributes. If no such triangle exists, the canonical method selects a single triangle from the cluster. However, in different frames and different triangles, the same cluster can cover the sample in question, resulting in a sudden change in the pixel shader input, thus producing temporal flicker.
[0040] These artifacts can be reduced or eliminated by considering all triangles within the shading cluster that overlap the same coarse pixel, not just the center, and considering their contributions to the pixel shader input attributes that are proportional to their coverage within that coarse pixel.
[0041] For each coarse pixel, first iterate over all triangles within the cluster that cover at least one visibility sample within the coarse pixel. For these triangles, interpolate the vertex attributes at the position of the coarse pixel center, as Figure 2 shown in 50.
[0042] For this deferred attribute interpolation, retain the vertex attribute equations (i.e., the output of the triangle setup) for all triangles that cover at least one visibility sample within the coarse pixel shading grid.
[0043] The final input attributes to the coarse pixel shader are then computed as a weighted average of the interpolated vertex attributes over the coverage. Alternatively, per-vertex shader attributes along with their per-triangle coverage weights can be exposed to the coarse pixel shader, and attribute interpolation can be performed within the shader itself.
[0044] Once the coarse pixel shader input attributes are found, the coarse pixel shader quads are arranged for shading, just as in a regular coarse pixel shading operation. This interpolation method gives varying pixel shader attributes as long as the cluster holds the same triangle, even if different triangles cover the coarse shading samples.
[0045] This technique can be considered as multi-sampling anti-aliasing (MSAA) resolve on the interpolated vertex attributes within the coarse pixels prior to pixel shading. If the shader is approximated as a linear function of the shader attributes, the result of this algorithm closely matches the average of the coverage-based shader outputs.
[0046] See Figure 2 , the depiction of the merge process involves cluster formation 52, coarse pixel shading grid coverage 54, and coarse pixel shading attributes 56 for primitive A, primitive B, and the merged primitive. The effects of primitive A and primitive B are added or merged to create the final or merged cluster 60 that forms the grid coverage and shading attributes.
[0047] Thus, primitive A shown under "cluster formation" is mapped onto the grid in column 54 as shown for each primitive in column 56. For example, primitive A covers the upper right quadrant that includes the upper right quad pixel fragment and less than half of the adjacent quad pixel fragment (e.g., 6 / 16 or 6 out of 16 pixel centers). Primitive B covers the lower left quad pixel fragment and more than half of the adjacent quad pixel fragment (e.g., 10 / 16 or 10 out of 16 pixel centers). Then when merged, the entire fragment 60 is now covered in this example.
[0048] Figure 3 The sequence 70 shown in can be implemented in software, firmware, and / or hardware. In software and firmware embodiments, it can be implemented by computer-executable instructions stored in one or more non-transitory computer-readable media such as magnetic, optical, or semiconductor storage.
[0049] Sequence 70 begins by checking whether the quad fragments are from the same draw call, as indicated in diamond 72. If so, the check at diamond 74 determines whether the coverage of the samples overlaps. Only when both diamond 72 and 74 have their conditions satisfied does one reach box 76. In one embodiment, the thick pixel quad fragments in the cluster are buffered in an on-chip buffer, as indicated in box 76.
[0050] Sequence 70 continues by iterating over the triangles within the cluster, as indicated in box 78. Then the vertex attributes are interpolated at the thick pixel centers, as indicated in box 80. The input attributes are calculated as a weighted average of the interpolated vertex attributes, as indicated in box 82.
[0051] Coverage weighted attribute interpolation 58 includes 6 / 16 of one quad times the quad plus 10 / 16 of another quad times the quad region to determine the combined quad. This is because in this example, primitive A covers 6 / 16 of the quad and primitive B covers 10 / 16 of the quad.
[0052] In some cases, the limited or small capacity of the actual implementation of the cluster buffer and the cluster rasterizer may impose limitations on this technique. Different sets of triangles may be grouped together, which in turn can lead to transient artifacts. In some embodiments, the shading information across the boundaries of such clusters is not reused.
[0053] In addition, certain shader attribute extrapolation artifacts may become more severe. Attribute extrapolation occurs when a sample falls outside a triangle and the vertex attributes are extrapolated, producing values that may not correspond to the true surface. With standard thick pixel shading methods, the extrapolation artifacts only corrupt the visibility samples covered by the extrapolated triangles. The term "visibility sample" corresponds to the regular (non-thick) fragments that draw their color values from the thick shaded fragments. The problems described herein mean that the extrapolated attributes affect all fragments within the thick pixel, not just those actually covered by the primitive. In the techniques described herein, the extrapolated vertex attributes contribute to the color of the entire thick pixel partially covered by the triangle. This can be easily addressed by reducing the weight of these triangles.
[0054] Figure 4 is a block diagram of a processing system 100 according to an embodiment. In various embodiments, system 100 includes one or more processors 102 and one or more graphics processors 108, and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a large number of processors 102 or processor cores 107. In one embodiment, system 100 is a processing platform incorporated within a system-on-chip (SoC) integrated circuit for use in mobile devices, handheld devices, or embedded devices.
[0055] Embodiments of system 100 may include or incorporate a server-based gaming platform, a game console, including a game and media console, a mobile game console, a handheld game console, or an online game console. In some embodiments, system 100 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. Data processing system 100 may also include a wearable device (such as a smartwatch wearable device, a smart glasses device, an augmented reality device, or a virtual reality device), coupled to, or integrated in, the wearable device. In some embodiments, data processing system 100 is a television or set-top box device having one or more processors 102 and a graphical interface generated by one or more graphics processors 108.
[0056] In some embodiments, each of the one or more processors 102 includes one or more processor cores 107 for processing instructions that, when executed, perform the operations of system and user software. In some embodiments, each of the one or more processor cores 107 is configured to process a particular instruction set 109. In some embodiments, instruction set 109 may facilitate complex instruction set computing (CISC), reduced instruction set computing (RISC), or computing via very long instruction words (VLIW). The multiple processor cores 107 may each process a different instruction set 109, which may include instructions for facilitating the emulation of other instruction sets. The processor cores 107 may also include other processing devices, such as a digital signal processor (DSP).
[0057] In some embodiments, processor 102 includes a cache memory 104. Depending on the architecture, processor 102 may have a single internal cache or multiple levels of internal caches. In some embodiments, the cache memory is shared among the components of processor 102. In some embodiments, processor 102 also uses an external cache (e.g., a level 3 (L3) cache or a last-level cache (LLC)) (not shown), and known cache coherence techniques may be used to share the external cache among the processor cores 107. Additionally, a register file 106 is included in processor 102, and the processor may include different types of registers for storing different types of data (e.g., integer registers, floating-point registers, status registers, and instruction pointer registers). Some registers may be general-purpose registers, while other registers may be specific to the design of processor 102.
[0058] In some embodiments, the processor 102 is coupled to a processor bus 110, which is used to transmit communication signals, such as address, data, or control signals, between the processor 102 and other components within the system 100. In one embodiment, the system 100 uses an exemplary 'hub' system architecture, including a memory controller hub 116 and an input / output (I / O) controller hub 130. The memory controller hub 116 facilitates communication between the memory device and other components of the system 100, while the I / O controller hub (ICH) 130 provides a connection to I / O devices via a local I / O bus. In one embodiment, the logic of the memory controller hub 116 is integrated within the processor.
[0059] The memory device 120 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device with suitable performance for use as a processing memory. In one embodiment, the memory device 120 can operate as the system memory of the system 100 to store data 122 and instructions 121 for use when one or more processors 102 execute an application or process. The memory controller hub 116 is also coupled to an optional external graphics processor 112, which can communicate with one or more graphics processors 108 in the processor 102 to perform graphics and media operations.
[0060] In some embodiments, the ICH 130 enables peripheral components to be connected to the memory device 120 and the processor 102 via a high-speed I / O bus. The I / O peripherals include but are not limited to: an audio controller 146, a firmware interface 128, a wireless transceiver 126 (e.g., Wi-Fi, Bluetooth), a data storage device 124 (e.g., a hard disk drive, flash memory, etc.), and a legacy I / O controller 140 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system. One or more universal serial bus (USB) controllers 142 connect multiple input devices, such as a keyboard and mouse 144 combination. The network controller 134 can also be coupled to the ICH 130. In some embodiments, a high-performance network controller (not shown) is coupled to the processor bus 110. It should be understood that the system 100 shown is exemplary and not restrictive, as other types of data processing systems configured in different ways can also be used. For example, the I / O controller hub 130 can be integrated within one or more processors 102, or the memory controller hub 116 and the I / O controller hub 130 can be integrated within a discrete external graphics processor (such as the external graphics processor 112).
[0061] Figure 5is a block diagram of an embodiment of a processor 200 having one or more processor cores 202A - 202N, an integrated memory controller 214, and an integrated graphics processor 208. Figure 5 Those elements having the same reference numbers (or names) as elements in any other figure herein may operate or function in any manner similar to the ways described elsewhere herein, but are not limited to these. Processor 200 may include additional cores up to and including additional core 202N represented by the dashed box. Processor cores 202A - 202N each include one or more internal cache units 204A - 204N. In some embodiments, each processor core may also access one or more shared cache units 206.
[0062] Internal cache units 204A - 204N and shared cache unit 206 represent the cache memory hierarchy within processor 200. The cache memory hierarchy may include at least one level of instruction and data cache within each processor core and one or more levels of shared mid - level cache, such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest - level cache is classified as the LLC before external memory. In some embodiments, cache coherence logic maintains coherence between cache units 206 and 204A - 204N.
[0063] In some embodiments, processor 200 may also include a group of one or more bus controller units 216 and a system agent core 210. One or more bus controller units 216 manage a group of peripheral buses, such as one or more peripheral component interconnect buses (e.g., PCI, PCI Express). System agent core 210 provides management functions for the various processor components. In some embodiments, system agent core 210 includes one or more integrated memory controllers 214 for managing access to various external memory devices (not shown).
[0064] In some embodiments, one or more of processor cores 202A - 202N include support for simultaneous multithreading. In such embodiments, system agent core 210 includes components for coordinating and operating cores 202A - 202N during multithreaded processing. Additionally, system agent core 210 may also include a power control unit (PCU) that includes logic and components for regulating the power states of processor cores 202A - 202N and graphics processor 208.
[0065] In some embodiments, additionally, the processor 200 further includes a graphics processor 208 for performing graphics processing operations. In some embodiments, the graphics processor 208 is coupled to the shared cache unit 206 set and the system agent core 210, and the system agent core includes one or more integrated memory controllers 214. In some embodiments, the display controller 211 is coupled to the graphics processor 208 to drive the graphics processor output to one or more coupled displays. In some embodiments, the display controller 211 may be a separate module coupled to the graphics processor via at least one interconnect, or may be integrated within the graphics processor 208 or the system agent core 210.
[0066] In some embodiments, the ring-based interconnect unit 212 is used to couple the internal components of the processor 200. However, alternative interconnect units may be used, such as point-to-point interconnects, switched interconnects, or other techniques, including techniques well known in the art. In some embodiments, the graphics processor 208 is coupled to the ring interconnect 212 via an I / O link 213.
[0067] The exemplary I / O link 213 represents at least one of a plurality of varieties of I / O interconnects, including package I / O interconnects that facilitate communication between various processor components and a high-performance embedded memory module 218, such as an eDRAM module. In some embodiments, each of the processor cores 202A to 202N and the graphics processor 208 uses the embedded memory module 218 as a shared last-level cache.
[0068] In some embodiments, the processor cores 202A to 202N are homogeneous cores that execute the same instruction set architecture. In another embodiment, the processor cores 202A to 202N are heterogeneous in terms of instruction set architecture (ISA), where one or more of the processor cores 202A-N execute a first instruction set, and at least one of the other cores executes a subset or a different instruction set of the first instruction set. In one embodiment, the processor cores 202A to 202N are homogeneous in terms of microarchitecture, where one or more cores with relatively high power consumption are coupled to one or more power-efficient cores. Additionally, the processor 200 may be implemented on one or more chips or as a SoC integrated circuit having the components shown in addition to other components.
[0069] Figure 6is a block diagram of a graphics processor 300, which can be a discrete graphics processing unit or can be a graphics processor integrated with multiple processing cores. In some embodiments, the graphics processor communicates with the memory via a mapped I / O interface to registers on the graphics processor and using commands placed in the processor memory. In some embodiments, the graphics processor 300 includes a memory interface 314 for accessing the memory. The memory interface 314 can be an interface to local memory, one or more internal caches, one or more shared external caches, and / or to system memory.
[0070] In some embodiments, the graphics processor 300 further includes a display controller 302 for driving display output data to a display device 320. The display controller 302 includes hardware for one or more overlapping planes of the display and the composition of multi-layer video or user interface elements. In some embodiments, the graphics processor 300 includes a video codec engine 306 for encoding, decoding, or performing media transcoding to, from, or between one or more media coding formats, including but not limited to: Moving Picture Experts Group (MPEG) (such as MPEG-2), Advanced Video Coding (AVC) formats (such as H.264 / MPEG-4 AVC), and Society of Motion Picture and Television Engineers (SMPTE) 421M / VC-1, and Joint Photographic Experts Group (JPEG) formats (such as JPEG and Motion JPEG (MJPEG) formats).
[0071] In some embodiments, the graphics processor 300 includes a block image transfer (BLIT) engine 304 for performing two-dimensional (2D) rasterizer operations, including for example bit boundary block transfer. However, in one embodiment, one or more components of the graphics processing engine (GPE) 310 are used to perform 2D graphics operations. In some embodiments, the graphics processing engine 310 is a computing engine for performing graphics operations, including three-dimensional (3D) graphics operations and media operations.
[0072] In some embodiments, the GPE 310 includes a 3D pipeline 312 for performing 3D operations, such as rendering three-dimensional images and scenes using processing functions for 3D primitive shapes (e.g., rectangles, triangles, etc.). The 3D pipeline 312 includes programmable and fixed function elements that perform various tasks within the elements of the 3D / media subsystem 315 and / or the generated execution threads. Although the 3D pipeline 312 can be used to perform media operations, embodiments of the GPE 310 also include a media pipeline 316 specifically for performing media operations, such as video post-processing and image enhancement.
[0073] In some embodiments, media pipeline 316 includes fixed function or programmable logic units to perform one or more specialized media operations, such as video decode acceleration, video deinterlacing, and video encode acceleration, in place of, or on behalf of, video codec engine 306. In some embodiments, additionally, media pipeline 316 also includes a thread generation unit to generate threads for execution on 3D / media subsystem 315. The generated threads perform media operation computations on one or more graphics execution units included in 3D / media subsystem 315.
[0074] In some embodiments, 3D / media subsystem 315 includes logic for executing threads generated by 3D pipeline 312 and media pipeline 316. In one embodiment, the pipeline sends thread execution requests to 3D / media subsystem 315, which includes thread dispatch logic for arbitrating and dispatching requests to available thread execution resources. Execution resources include an array of graphics execution units for processing 3D and media threads. In some embodiments, 3D / media subsystem 315 includes one or more internal caches for thread instructions and data. In some embodiments, the subsystem also includes shared memory (including registers and addressable memory) for sharing data between threads and for storing output data.
[0075] Figure 7 is a block diagram of a graphics processing engine 410 of a graphics processor according to some embodiments. In one embodiment, GPE 410 is Figure 6 a version of GPE 310 shown in Figure 7 Elements with the same reference numerals (or names) as elements in any other figure herein may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto.
[0076] In some embodiments, GPE 410 is coupled to a command streamer 403 that provides a command stream to GPE 3D and media pipelines 412, 416. In some embodiments, command streamer 403 is coupled to a memory, which may be a system memory, or one or more of an internal cache memory and a shared cache memory. In some embodiments, command streamer 403 receives commands from the memory and sends these commands to 3D pipeline 412 and / or media pipeline 416. These commands are instructions fetched from a ring buffer that stores commands for 3D and media pipelines 412, 416. In one embodiment, the ring buffer may additionally include a batch command buffer that stores batches of multiple commands. 3D and media pipelines 412, 416 process commands by performing operations via logic within the respective pipelines; or dispatching one or more execution threads to execution unit array 414. In some embodiments, execution unit array 414 is scalable such that the array includes a variable number of execution units based on the target power and performance levels of GPE 410.
[0077] In some embodiments, sampling engine 430 is coupled to a memory (e.g., a cache memory or a system memory) and execution unit array 414. In some embodiments, sampling engine 430 provides a memory access mechanism for execution unit array 414 that allows the execution array 414 to read graphics and media data from the memory. In some embodiments, sampling engine 430 includes logic for performing specialized image sampling operations for media.
[0078] In some embodiments, the specialized media sampling logic in sampling engine 430 includes a denoising / deinterlacing module 432, a motion estimation module 434, and an image scaling and filtering module 436. In some embodiments, denoising / deinterlacing module 432 includes logic for performing one or more of denoising or deinterlacing on decoded video data. The deinterlacing logic combines alternating fields of interlaced video content into a single video frame. The denoising logic reduces or removes data noise from video and image data. In some embodiments, the denoising logic and the deinterlacing logic are motion adaptive and use spatial or temporal filtering based on the amount of motion detected in the video data. In some embodiments, denoising / deinterlacing module 432 includes dedicated motion detection logic (e.g., within motion estimation engine 434).
[0079] In some embodiments, the motion estimation engine 434 provides hardware acceleration for video operations by performing video acceleration functions (such as motion vector estimation and prediction) on video data. The motion estimation engine determines motion vectors that describe the transformation of image data between consecutive video frames. In some embodiments, the graphics processor media codec uses the video motion estimation engine 434 to perform operations on video at the macroblock level, which operations on video at the macroblock level may otherwise be too computationally intensive to be performed using a general-purpose processor. In some embodiments, the motion estimation engine 434 can generally be used in the graphics processor component to assist video decoding and processing functions that are sensitive to or adaptive to the direction or magnitude of motion within the video data.
[0080] In some embodiments, the image scaling and filtering module 436 performs image processing operations to enhance the visual quality of the generated images and videos. In some embodiments, the scaling and filtering module 436 processes image and video data during sampling operations before providing the data to the execution unit array 414.
[0081] In some embodiments, the GPE 410 includes a data port 444 that provides an additional mechanism for the graphics subsystem to access memory. In some embodiments, the data port 444 facilitates memory access for operations including render target writes, constant buffer reads, grab memory space reads / writes, and media surface access. In some embodiments, the data port 444 includes a cache memory space for caching accesses to memory. The cache memory can be a single data cache or can be separated into multiple caches for multiple subsystems accessing memory via the data port (e.g., render buffer cache, constant buffer cache, etc.). In some embodiments, threads executing on the execution units in the execution unit array 414 communicate with the data port by exchanging messages via the data distribution interconnect that couples each of the subsystems of the GPE 410.
[0082] Figure 8 is a block diagram of another embodiment of the graphics processor 500. Figure 8 Those elements having the same reference numbers (or names) as elements in any other figure herein may operate or function in any manner similar to the ways described elsewhere herein, but are not limited thereto.
[0083] In some embodiments, graphics processor 500 includes a ring interconnect 502, a pipeline front end 504, a media engine 537, and graphics cores 580A through 580N. In some embodiments, ring interconnect 502 couples the graphics processor to other processing units, including other graphics processors or one or more general-purpose processor cores. In some embodiments, the graphics processor is one of multiple processors integrated within a multi-core processing system.
[0084] In some embodiments, graphics processor 500 receives multiple batches of commands via ring interconnect 502. The incoming commands are interpreted by command stream converter 503 in pipeline front end 504. In some embodiments, graphics processor 500 includes scalable execution logic for performing 3D geometry processing and media processing via (multiple) graphics cores 580A through 580N. For 3D geometry processing commands, command stream converter 503 supplies the commands to geometry pipeline 536. For at least some media processing commands, command stream converter 503 supplies the commands to video front end 534, which is coupled to media engine 537. In some embodiments, media engine 537 includes a video quality engine (VQE) 530 for video and image post-processing and a multi-format encoding / decoding (MFX) 533 engine for providing hardware-accelerated encoding and decoding of media data. In some embodiments, geometry pipeline 536 and media engine 537 each generate execution threads for thread execution resources provided by at least one of graphics cores 580A.
[0085] In some embodiments, the graphics processor 500 includes scalable thread execution resource characterization module cores 580A through 580N (sometimes referred to as core shards), each of which has a plurality of sub-cores 550A through 550N, 560A through 560N (sometimes referred to as core sub-shards). In some embodiments, the graphics processor 500 can have any number of graphics cores 580A through 580N. In some embodiments, the graphics processor 500 includes graphics core 580A, which has at least a first sub-core 550A and a second core sub-core 560A. In other embodiments, the graphics processor is a low-power processor having a single sub-core (e.g., 550A). In some embodiments, the graphics processor 500 includes a plurality of graphics cores 580A through 580N, each of which includes a set of first sub-cores 550A through 550N and a set of second sub-cores 560A through 560N. Each sub-core in the set of first sub-cores 550A through 550N includes at least a first set of execution units 552A through 552N and media / texture samplers 554A through 554N. Each sub-core in the set of second sub-cores 560A through 560N includes at least a second set of execution units 562A through 562N and samplers 564A through 564N. In some embodiments, each sub-core 550A through 550N, 560A through 560N shares a set of shared resources 570A through 570N. In some embodiments, the shared resources include shared cache memory and pixel operation logic. Other shared resources may also be included in various embodiments of the graphics processor.
[0086] Figure 9 Thread execution logic 600 is shown, which includes an array of processing elements employed in some embodiments of the GPE. Figure 9 Those elements having the same reference number (or name) as an element in any other figure herein may operate or function in any manner similar to the ways described elsewhere herein, but are not limited to these.
[0087] In some embodiments, the thread execution logic 600 includes a pixel shader 602, a thread dispatcher 604, an instruction cache 606, a scalable execution unit array including a plurality of execution units 608A - 608N, a sampler 610, a data cache 612, and a data port 614. In one embodiment, the included components are interconnected via an interconnect structure that links to each of the components. In some embodiments, the thread execution logic 600 includes one or more connections to memory (such as system memory or cache memory) via the instruction cache 606, the data port 614, the sampler 610, and one or more of the execution unit arrays 608A - 608N. In some embodiments, each execution unit (e.g., 608A) is a separate vector processor capable of executing multiple simultaneous threads and processing multiple data elements in parallel for each thread. In some embodiments, the execution unit array 608A - 608N includes any number of individual execution units.
[0088] In some embodiments, the execution unit array 608A - 608N is primarily used to execute "shader" programs. In some embodiments, the execution units in the array 608A - 608N execute an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct3D and OpenGL) to be executed with minimal translation. The execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders).
[0089] Each execution unit in the execution unit array 608A through 608N operates on an array of data elements. The number of data elements is the "execution size", or the number of channels of an instruction. Execution channels are logical units that perform data - element access, masking, and flow control within an instruction. The number of channels may be independent of the number of physical arithmetic - logic units (ALUs) or floating - point units (FPUs) for a particular graphics processor. In some embodiments, the execution units 608A through 608N support integer and floating - point data types.
[0090] The execution unit instruction set includes single instruction multiple data (SIMD) instructions. Various data elements can be stored in registers as compressed data types, and the execution unit will process the various elements based on the data size of the elements. For example, when operating on a 256-bit wide vector, the 256-bit vector is stored in a register, and the execution unit operates on the vector as four separate 64-bit compressed data elements (quad word (QW) sized data elements), eight separate 32-bit compressed data elements (double word (DW) sized data elements), sixteen separate 16-bit compressed data elements (word (W) sized data elements), or thirty-two separate 8-bit data elements (byte (B) sized data elements). However, different vector widths and register sizes are possible.
[0091] One or more internal instruction caches (e.g., 606) are included in the thread execution logic 600 to cache the thread instructions of the execution unit. In some embodiments, one or more data caches (e.g., 612) are included for caching thread data during thread execution. In some embodiments, a sampler 610 is included for providing texture sampling for 3D operations and media sampling for media operations. In some embodiments, the sampler 610 includes specialized texture or media sampling functions to process texture or media data during the sampling process before providing the sampled data to the execution unit.
[0092] During execution, the graphics and media pipeline sends thread initiation requests to the thread execution logic 600 via the thread spawn and dispatch logic. In some embodiments, the thread execution logic 600 includes a local thread dispatcher 604 that arbitrates the thread initiation requests from the graphics and media pipeline and instantiates the requested threads on one or more execution units 608A - 608N. For example, the geometry pipeline (e.g., Figure 8 of 536) dispatches vertex processing, tessellation, or geometry processing threads to the thread execution logic 600 ( Figure 9 ). In some embodiments, the thread dispatcher 604 may also handle runtime thread spawn requests from executing shader programs.
[0093] Once a set of geometric objects has been processed and rasterized into pixel data, the pixel shader 602 is invoked to further compute output information and write the results to an output surface (e.g., a color buffer, a depth buffer, a stencil buffer, etc.). In some embodiments, the pixel shader 602 computes values of various vertex attributes that will be interpolated across the rasterized objects. In some embodiments, the pixel shader 602 then executes a pixel shader program supplied by an application programming interface (API). To execute the pixel shader program, the pixel shader 602 dispatches threads to execution units (e.g., 608A) via a thread dispatcher 604. In some embodiments, the pixel shader 602 uses texture sampling logic in a sampler 610 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometric data compute pixel color data for each geometric fragment, or discard one or more pixels without further processing.
[0094] In some embodiments, the data port 614 provides a memory access mechanism for the thread execution logic 600 to output processed data to memory for processing in the graphics processor output pipeline. In some embodiments, the data port 614 includes or is coupled to one or more cache memories (e.g., the data cache 612) to cache data via the data port for memory access.
[0095] Figure 10 is a block diagram showing a graphics processor instruction format 700 according to some embodiments. In one or more embodiments, the graphics processor execution units support an instruction set with multiple formats. The solid boxes show components typically included in the execution unit instructions, while the dashed boxes include optional components or components only included in a subset of the instructions. In some embodiments, the described and shown instruction format 700 is a macro-instruction as they are the instructions supplied to the execution unit, as opposed to micro-operations generated from instruction decoding (once the instruction has been processed).
[0096] In some embodiments, the graphics processor execution units natively support instructions in a 128-bit format 710. A 64-bit compact instruction format 730 can be used for some instructions based on the selected instructions, multiple instruction options, and the number of operands. The native 128-bit format 710 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 730. The native instructions available in the 64-bit format 730 vary according to the embodiment. In some embodiments, a set of index values in an index field 713 is used to partially compress the instruction. The execution unit hardware refers to a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in the 128-bit format 710.
[0097] For each format, the instruction opcode 712 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel, where the color channels represent texture elements or picture elements. By default, the execution unit executes each instruction across all data channels of the operand. In some embodiments, the instruction control field 714 enables control of certain execution options, such as channel selection (e.g., predication) and data channel ordering (e.g., mixing). For 128-bit instructions 710, the execution size field 716 limits the number of data channels that will be executed in parallel. In some embodiments, the execution size field 716 is not available for the 64-bit compact instruction format 730.
[0098] Some execution unit instructions have up to three operands, including two source operands (src0 722, src1 722) and one destination 718. In some embodiments, the execution unit supports dual-destination instructions, where one of the destinations is implicit. Data manipulation instructions may have a third source operand (e.g., SRC2 724), where the instruction opcode 712 determines the number of source operands. The last source operand of the instruction may be an immediate (e.g., hard-coded) value passed with the instruction.
[0099] In some embodiments, the 128-bit instruction format 710 includes access / address mode information 726, which, for example, defines whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are provided directly by bits in the instruction 710.
[0100] In some embodiments, the 128-bit instruction format 710 includes an access / address mode field 726, which specifies the address mode and / or access mode of the instruction. In one embodiment, the access mode is used to define the data access alignment for the instruction. Some embodiments support access modes, including a 16-byte aligned access mode and a 1-byte aligned access mode, where the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction 710 may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction 710 may use 16-byte aligned addressing for all source and destination operands.
[0101] In one embodiment, the address mode portion of the access / address mode field 726 determines whether the instruction uses direct addressing or indirect addressing. When using the direct register addressing mode, the bits in the instruction 710 directly provide the register addresses of one or more operands. When using the indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.
[0102] In some embodiments, the instructions are grouped based on the opcode 712 bit field to simplify opcode decoding 740. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of the opcode. The exact opcode grouping shown is merely exemplary. In some embodiments, the move and logic opcode group 742 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some embodiments, the move and logic group 742 shares the five most significant bits (MSBs), where the move (mov) instruction takes the form of 0000xxxxb and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 744 (e.g., call, jmp) includes instructions that take the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 746 includes a mix of instructions, including synchronization instructions (e.g., wait, send) that take the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 748 includes per-component arithmetic instructions (e.g., add, mul) that take the form of 0100xxxxb (e.g., 0x40). The parallel math group 748 performs arithmetic operations in parallel across data channels. The vector math group 750 includes arithmetic instructions (e.g., dp4) that take the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic operations on vector operands, such as dot product operations.
[0103] Figure 11 is a block diagram of another embodiment of the graphics processor 800. Figure 11 Those elements having the same reference numbers (or names) as the elements in any other figure herein may operate or function in any manner similar to the ways described elsewhere herein, but are not limited to these.
[0104] In some embodiments, graphics processor 800 includes a graphics pipeline 820, a media pipeline 830, a display engine 840, thread execution logic 850, and a render output pipeline 870. In some embodiments, graphics processor 800 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or is controlled via commands issued via a ring interconnect 802 to graphics processor 800. In some embodiments, ring interconnect 802 couples graphics processor 800 to other processing components, such as other graphics processors or general-purpose processors. Commands from ring interconnect 802 are interpreted by a command stream converter 803 that supplies instructions to individual components of graphics pipeline 820 or media pipeline 830.
[0105] In some embodiments, command stream converter 803 directs the operation of vertex fetcher 805, which reads vertex data from memory and executes vertex processing commands provided by command stream converter 803. In some embodiments, vertex fetcher 805 provides vertex data to vertex shader 807, which performs coordinate space transformation and lighting operations on each vertex. In some embodiments, vertex fetcher 805 and vertex shader 807 execute vertex processing instructions by dispatching execution threads to execution units 852A - 852B via thread dispatcher 831.
[0106] In some embodiments, execution units 852A - 852B are an array of vector processors having an instruction set for performing graphics and media operations. In some embodiments, execution units 852A - 852B have attached L1 caches 851 that are dedicated to each array or shared between the arrays. The caches can be configured as data caches, instruction caches, or a single cache that is partitioned to contain data and instructions in different partitions.
[0107] In some embodiments, graphics pipeline 820 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some embodiments, a programmable hull shader 811 configures the tessellation operation. A programmable domain shader 817 provides backend evaluation of the tessellation output. Tessellator 813 operates in the direction of hull shader 811 and includes dedicated logic for generating a detailed set of geometric objects based on a coarse geometric model that is provided as input to graphics pipeline 820. In some embodiments, if tessellation is not used, the tessellation components 811, 813, 817 can be bypassed.
[0108] In some embodiments, a complete geometric object may be processed by a geometry shader 819 via one or more threads dispatched to the execution units 852A - 852B, or may proceed directly to the clipper 829. In some embodiments, the geometry shader operates on an entire geometric object (as opposed to vertices or vertex patches as in previous stages of the graphics pipeline). If tessellation is disabled, the geometry shader 819 receives input from the vertex shader 807. In some embodiments, the geometry shader 819 may be programmed by a geometry shader program to perform geometric tessellation when the tessellation unit is disabled.
[0109] Before rasterization, the clipper 829 processes vertex data. The clipper 829 may be a fixed - function clipper or a programmable clipper with clipping and geometry shader functionality. In some embodiments, the rasterizer / depth 873 in the render output pipeline 870 dispatches pixel shaders to convert the geometric object into its per - pixel representation. In some embodiments, the pixel shader logic is included in the thread execution logic 850. In some embodiments, an application may bypass the rasterizer 873 and access the un - rasterized vertex data via the egress unit 823.
[0110] The graphics processor 800 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed among the major components of the graphics processor. In some embodiments, the execution units 852A - 852B and the associated cache(s) 851, the texture and media sampler 854, and the texture / sampler cache 858 are interconnected via a data port 856 to perform memory accesses and communicate with the render output pipeline components of the processor. In some embodiments, the sampler 854, caches 851, 858, and the execution units 852A - 852B each have separate memory access paths.
[0111] In some embodiments, the rendering output pipeline 870 includes a rasterizer and depth test component 873 that converts vertex-based objects into associated pixel-based representations. In some embodiments, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. Associated rendering cache 878 and depth cache 879 are also available in some embodiments. Pixel operation component 877 performs pixel-based operations on the data, however in some instances, pixel operations associated with 2D operations (e.g., bit-block blitting with blending) are performed by the 2D engine 841 or, at display time, by the display controller 843 using overlapping display planes instead. In some embodiments, a shared L3 cache 875 is available for all graphics components, allowing data to be shared without using the main system memory.
[0112] In some embodiments, the graphics processor media pipeline 830 includes a media engine 837 and a video front end 834. In some embodiments, the video front end 834 receives pipeline commands from the command stream converter 803. In some embodiments, the media pipeline 830 includes a separate command stream converter. In some embodiments, the video front end 834 processes media commands before sending the commands to the media engine 837. In some embodiments, the media engine 337 includes a thread generation function for generating threads for dispatch to thread execution logic 850 via a thread dispatcher 831.
[0113] In some embodiments, the graphics processor 800 includes a display engine 840. In some embodiments, the display engine 840 is external to the processor 800 and is coupled to the graphics processor via a ring interconnect 802, or some other interconnect bus or mechanism. In some embodiments, the display engine 840 includes a 2D engine 841 and a display controller 843. In some embodiments, the display engine 840 contains dedicated logic capable of operating independently of the 3D pipeline. In some embodiments, the display controller 843 is coupled to a display device (not shown), which may be a system integrated display device (such as in a laptop computer) or an external display device attached via a display device connector.
[0114] In some embodiments, the graphics pipeline 820 and the media pipeline 830 may be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). In some embodiments, the driver software of the graphics processor converts API dispatches dedicated to a particular graphics or media library into commands that can be processed by the graphics processor. In some embodiments, support is provided for the Open Graphics Library (OpenGL) and the Open Computing Language (OpenCL) from the Khronos Group, or support may be provided for both OpenGL and D3D. In some embodiments, combinations of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the pipeline of the future API to the pipeline of the graphics processor.
[0115] Figure 12A is a block diagram showing a graphics processor command format 900 according to some embodiments. Figure 12B is a block diagram showing a graphics processor command sequence 910 according to an embodiment. Figure 12A The solid boxes in show components that are typically included in a graphics command, while the dashed boxes include components that are optional or included only in a subset of the graphics commands. Figure 12A An exemplary graphics processor command format 900 includes a target client 902 for identifying the command, a command operation code (opcode) 904, and a data field 906 for the relevant data of the command. Some commands also include a sub-opcode 905 and a command size 908.
[0116] In some embodiments, the client 902 defines the client unit of the graphics device that processes the command data. In some embodiments, the graphics processor command parser examines the client field of each command to adjust further processing of the command and route the command data to the appropriate client unit. In some embodiments, the graphics processor client units include a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing the command. Once the command is received by the client unit, the client unit reads the opcode 904 and the sub-opcode 905 (if present) to determine the operation to be performed. The client unit uses the information within the data field 906 to execute the command. For some commands, an explicit command size 908 is expected to define the size of the command. In some embodiments, the command parser automatically determines the size of at least some of the commands based on the command opcode. In some embodiments, the commands are aligned by multiples of the double-word length.
[0117] Figure 12BThe flowchart in [it] shows an exemplary graphics processor command sequence 910. In some embodiments, software or firmware of a data processing system characterized by an embodiment of the graphics processor uses a version of the shown command sequence to initiate, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only, and the embodiments are not limited to these specific commands or this command sequence. Also, the commands may be issued as a batch of commands in a command sequence such that the graphics processor will process the command sequence in at least a partially simultaneous manner.
[0118] In some embodiments, the graphics processor command sequence 910 may begin with a pipeline flush clear command 912 to cause any active graphics pipeline to complete the current outstanding commands for that pipeline. In some embodiments, the 3D pipeline 922 and the media pipeline 924 do not operate simultaneously. The pipeline flush clear is executed to cause the active graphics pipeline to complete any outstanding commands. In response to the pipeline flush clear, the command parser for the graphics processor will stop command processing until the active rendering engine has completed the outstanding operations and invalidate the associated read cache. Optionally, any data marked 'dirty' in the render cache may be flushed to memory. In some embodiments, the pipeline flush clear command 912 may be used for pipeline synchronization or before putting the graphics processor into a low power state.
[0119] In some embodiments, when the command sequence requires the graphics processor to explicitly switch between pipelines, the pipeline select command 913 is used. In some embodiments, only one pipeline select command 913 is required in the execution context before issuing pipeline commands, unless the context is to issue commands for two pipelines. In some embodiments, a pipeline flush clear command 912 is exactly required before the pipeline switch via the pipeline select command 913.
[0120] In some embodiments, the pipeline control command 914 configures the graphics pipeline for operation and programs the 3D pipeline 922 and the media pipeline 924. In some embodiments, the pipeline control command 914 configures the pipeline state of the active pipeline. In one embodiment, the pipeline control command 914 is used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands.
[0121] In some embodiments, the return buffer status command 916 is used to configure a set of return buffers for corresponding pipelines to write data. Some pipeline operations require allocating, selecting, or configuring one or more return buffers, and during processing, the operations write intermediate data into the one or more return buffers. In some embodiments, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some embodiments, the return buffer status 916 includes selecting the size and number of return buffers for a set of pipeline operations.
[0122] The remaining commands in the command sequence vary based on the active pipelines for the operations. Based on the pipeline determination 920, the command sequence is customized for a 3D pipeline 922 starting with a 3D pipeline state 930, or a media pipeline 924 starting at a media pipeline state 940.
[0123] Commands for the 3D pipeline state 930 include 3D state setting commands for vertex buffer status, vertex element status, constant color status, depth buffer status, and other status variables to be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the particular 3D API in use. In some embodiments, the 3D pipeline state 930 commands can also selectively disable or bypass specific pipeline elements if those elements will not be used.
[0124] In some embodiments, the 3D primitive 932 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 932 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 932 command data to generate multiple vertex data structures. The vertex data structures are stored in one or more return buffers. In some embodiments, the 3D primitive 932 commands are used to perform vertex operations on 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 922 dispatches shader execution threads to the graphics processor execution units.
[0125] In some embodiments, the 3D pipeline 922 is triggered via the execution of a 934 command or an event. In some embodiments, a register write triggers command execution. In some embodiments, execution is triggered via a 'go' or 'kick' command in a command sequence. In one embodiment, a pipeline synchronization command is used to trigger command execution in order to clear the command sequence via a graphics pipeline dump. The 3D pipeline will perform geometric processing on 3D primitives. Once the operation is complete, the resulting geometric objects are rasterized, and the pixel engine colors the resulting pixels. For these operations, additional commands for controlling pixel coloring and pixel backend operations may also be included.
[0126] In some embodiments, when a media operation is being performed, the graphics processor command sequence 910 follows the media pipeline 924 path. Generally, the specific uses and ways of programming the media pipeline 924 depend on the media or computing operations to be performed. During media decoding, specific media decoding operations may be offloaded to the media pipeline. In some embodiments, the media pipeline may also be bypassed, and the resources provided by one or more general-purpose processing cores may be used to perform media decoding, either wholly or in part. In one embodiment, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using a compute shader program that is not explicitly related to rendering graphics primitives.
[0127] In some embodiments, the media pipeline 924 is configured in a manner similar to the 3D pipeline 922. A set of media pipeline state commands 940 are dispatched or placed into the command queue, before the media object commands 942. In some embodiments, the media pipeline state commands 940 include data for configuring the media pipeline elements that will be used to process the media object. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some embodiments, the media pipeline state commands 940 also support using one or more pointers for "indirect" state elements that contain a batch of state settings.
[0128] In some embodiments, the media object command 942 supplies a pointer to a media object for processing by a media pipeline. The media object includes a memory buffer that contains video data to be processed. In some embodiments, all media pipeline states must be valid before issuing the media object command 942. Once the pipeline state is configured and the media object command 942 is queued, the media pipeline 924 is triggered via an execute 944 command or an equivalent execution event (e.g., a register write). The output from the media pipeline 924 can then be post-processed by operations provided by the 3D pipeline 922 or the media pipeline 924. In some embodiments, GPGPU operations are configured and executed in a manner similar to media operations.
[0129] Figure 13 An exemplary graphics software architecture of a data processing system 1000 is shown in accordance with some embodiments. In some embodiments, the software architecture includes a 3D graphics application 1010, an operating system 1020, and at least one processor 1030. In some embodiments, the processor 1030 includes a graphics processor 1032 and one or more general-purpose processor cores 1034. The graphics application 1010 and the operating system 1020 each execute in the system memory 1050 of the data processing system.
[0130] In some embodiments, the 3D graphics application 1010 includes one or more shader programs that include shader instructions 1012. The shader language instructions may be in a high-level shader language, such as High-Level Shading Language (HLSL) or OpenGL Shading Language (GLSL). The application also includes executable instructions 1014 that are in a machine language suitable for execution by the general-purpose processor cores 1034. The application also includes a graphics object 1016 defined by vertex data.
[0131] In some embodiments, the operating system 1020 is an operating system from Microsoft Corporation operating system, a proprietary UNIX-like operating system, or an open-source UNIX-like operating system using a Linux kernel variant. When the Direct3D API is in use, the operating system 1020 uses a front-end shader compiler 1024 to compile any shader instructions 1012 in HLSL into a lower-level shader language. The compilation may be just-in-time (JIT) compilation, or the application may perform shader pre-compilation. In some embodiments, during the compilation of the 3D graphics application 1010, high-level shaders are compiled into low-level shaders.
[0132] In some embodiments, the user-mode graphics driver 1026 includes a backend shader compiler 1027 that is used to convert shader instructions 1012 into a hardware-specific representation. When using the OpenGL API, shader instructions 1012 in the GLSL high-level language are passed to the user-mode graphics driver 1026 for compilation. In some embodiments, the user-mode graphics driver 1026 uses operating system kernel-mode functions 1028 to communicate with the kernel-mode graphics driver 1029. In some embodiments, the kernel-mode graphics driver 1029 communicates with the graphics processor 1032 to dispatch commands and instructions.
[0133] One or more aspects of at least one embodiment may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit such as a processor. For example, the machine-readable medium may include instructions that represent the various logics within the processor. When read by a machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of the logic of an integrated circuit and may be stored on a tangible, machine-readable medium as a hardware model that describes the structure of the integrated circuit. The hardware model may be supplied to various consumers or manufacturing facilities that load the hardware model on a manufacturing machine for fabricating the integrated circuit. The integrated circuit may be fabricated such that the circuit performs operations described in association with any of the embodiments described herein.
[0134] Figure 14 is a block diagram showing an IP core development system 1100 that may be used to fabricate an integrated circuit to perform operations. The IP core development system 1100 may be used to generate modular, reusable designs that can be incorporated into a larger design or used to build an entire integrated circuit (e.g., an SOC integrated circuit). A design facility 1130 may use a high-level programming language (e.g., C / C++) to generate a software simulation 1110 of the IP core design. The software simulation 1110 may be used to design, test, and verify the behavior of the IP core using a simulation model 1112. The simulation model 1112 may include functional, behavioral, and / or timing simulations. An register transfer level (RTL) design may then be created or synthesized by the simulation model 1112. The RTL design 1115 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers (including the associated logic performed using the modeled digital signals). In addition to the RTL design 1115, lower-level designs at the logic level or transistor level may also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation may vary.
[0135] An RTL design 1115 or equivalent can be further synthesized by a design facility into a hardware model 1120, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 1140 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 1165. Alternatively, the IP core design can be transmitted (e.g., via the Internet) over a wired connection 1150 or a wireless connection 1160. The manufacturing facility 1165 can then fabricate an integrated circuit that is at least partially based on the IP core design. The fabricated integrated circuit can be configured to perform operations in accordance with at least one embodiment described herein.
[0136] Figure 15 FIG. is a block diagram showing an exemplary system-on-chip integrated circuit 1200 that can be fabricated using one or more IP cores according to an embodiment. The exemplary integrated circuit includes one or more application processors 1205 (e.g., CPUs), at least one graphics processor 1210, and may additionally include an image processor 1215 and / or a video processor 1220, any of which can be a modular IP core from the same or multiple different design facilities. The integrated circuit includes peripheral or bus logic, including a USB controller 1225, a UART controller 1230, an SPI / SDIO controller 1235, and an 2 S / I 2 2C controller 1240. Additionally, the integrated circuit may further include a display device 1245 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1250 and a mobile industry processor interface (MIPI) display interface 1255. Storage can be provided by a flash memory subsystem 1260 (including flash memory and a flash memory controller). A memory interface can be provided via a memory controller 1265 to access SDRAM or SRAM memory devices. Additionally, some integrated circuits further include an embedded security engine 1270.
[0137] In addition, other logic and circuitry can be included in the processors of the integrated circuit 1200, including additional graphics processors / kernels, peripheral interface controllers, or general-purpose processor cores.
[0138] The following clauses and / or examples relate to further embodiments:
[0139] An example embodiment can be a method that includes merging two primitives by interpolating vertex attributes at a coarse pixel center by calculating an input attribute as a coverage weighted average of interpolated vertex attributes, and performing coarse pixel shading using the merged primitives. The method can further include merging quadtree fragments. The method can further include merging only quadtree fragments from the same instance of a draw call. The method can further include merging only quadtree fragments whose sample coverage does not overlap. The method can further include buffering coarse pixel quadtree fragments in a cluster before coarse pixel shading. The method can further include using a buffer in a graphics processing unit to buffer the fragments. The method can further include retaining vertex attribute equations for primitives that cover at least one visibility sample.
[0140] Another example embodiment can be one or more non-transitory computer-readable media that store instructions for performing a sequence of operations that includes interpolating vertex attributes at a coarse pixel center by calculating an input attribute as a coverage weighted average of interpolated vertex attributes to merge two primitives, and performing coarse pixel shading using the merged primitives. The medium can further store instructions for performing a sequence of operations that includes merging quadtree fragments. The medium can further store instructions for performing a sequence of operations that includes merging only quadtree fragments from the same instance of a draw call. The medium can further store instructions for performing a sequence of operations that includes merging only quadtree fragments whose sample coverage does not overlap. The medium can further store instructions for performing a sequence of operations that includes buffering coarse pixel quadtree fragments in a cluster before coarse pixel shading. The medium can further store instructions for performing a sequence of operations that includes using a buffer in a graphics processing unit to buffer the fragments. The medium can further store instructions for performing a sequence of operations that includes retaining vertex attribute equations for primitives that cover at least one visibility sample.
[0141] Another example embodiment can be an apparatus that includes a processor and a memory coupled to the processor, the processor for: merging two primitives by interpolating vertex attributes at a coarse pixel center by calculating an input attribute as a coverage weighted average of interpolated vertex attributes, and performing coarse pixel shading using the merged primitives. The apparatus can include the processor for performing: merging quadtree fragments. The apparatus can include the processor for performing: merging only quadtree fragments from the same instance of a draw call. The apparatus can include the processor for performing: merging only quadtree fragments whose sample coverage regions do not overlap. The apparatus can include the processor for performing: buffering the coarse pixel quadtree fragments in a cluster prior to coarse pixel shading. The apparatus can include the processor for performing: using a buffer in a graphics processing unit to buffer the fragments. The apparatus can include the processor for performing: retaining vertex attribute equations for primitives that cover at least one visibility sample.
[0142] The graphics processing techniques described herein can be implemented in a variety of hardware architectures. For example, the graphics functionality can be integrated within a chipset. Alternatively, a discrete graphics processor can be used. As yet another embodiment, the graphics functionality can be implemented by a general-purpose processor that includes a multi-core processor.
[0143] References to "one embodiment" or "an embodiment" in the present specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one implementation encompassed by the present disclosure. Thus, the appearances of the phrase "one embodiment" or "in an embodiment" do not necessarily refer to the same embodiment. Additionally, the particular features, structures, or characteristics may be established in other suitable forms different from the illustrated particular embodiments, and all such forms may be covered within the scope of the claims of this application.
[0144] Although a limited number of embodiments have been described, those skilled in the art will recognize many modifications and variations therefrom. The appended claims are intended to cover all such modifications and variations that fall within the true spirit and scope of the present disclosure.
Claims
1. A method for graphics processing, comprising: Interpolating vertex attributes at the coarse pixel centers of two primitives; Calculating pixel shader input attributes as a coverage weighted average of the interpolated vertex attributes of the two primitives; If the two primitives share an edge with the same vertex attributes, have the same orientation, and have mutually exclusive coverage, then merging the quadtree fragments of the two primitives, the merging including: Checking whether the quadtree fragments are from the same draw call; Determining whether the coverage ranges of the quadtree fragments do not overlap; and If the quadtree fragments are from the same draw call and if the sample coverage ranges of the quadtree fragments do not overlap, then buffering the coarse pixel quadtree fragments in the cluster before coarse pixel shading; and Performing the coarse pixel shading using the merged two primitives.
2. The method according to claim 1, including using a buffer in a graphics processing unit to buffer the fragments.
3. The method according to claim 1, including retaining vertex attribute equations for primitives covering at least one visibility sample.
4. One or more non-transitory computer-readable media storing instructions for performing the method according to any one of the preceding claims.
5. An apparatus for graphics processing, comprising: A processor configured to: Interpolate vertex attributes at the coarse pixel centers of two primitives; Calculate pixel shader input attributes as a coverage weighted average of the interpolated vertex attributes of the two primitives; If the two primitives share an edge with the same vertex attributes, have the same orientation, and have mutually exclusive coverage, then merge the quadtree fragments of the two primitives by: Checking whether the quadtree fragments are from the same draw call; Determining whether the coverage ranges of the quadtree fragments do not overlap; and If the quadtree fragments are from the same draw call and if the sample coverage ranges of the quadtree fragments do not overlap, then buffer the coarse pixel quadtree fragments in the cluster before coarse pixel shading; and Performing the coarse pixel shading using the merged two primitives.
6. The apparatus according to claim 5, wherein the processor is configured to use a buffer in a graphics processing unit to buffer the fragments.
7. The apparatus according to claim 5, wherein the processor is configured to retain vertex attribute equations for primitives covering at least one visibility sample.
Citation Information
Patent Citations
Merge coarse pixel shaded fragments using a weighted average of the triangle's attributes
CN110136223B