Graphics discard engine
The introduction of a discard engine in graphics processing pipelines addresses inefficiencies in attribute data management by ensuring timely deallocation, reducing memory usage and improving performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ADVANCED MICRO DEVICES INC
- Filing Date
- 2022-11-23
- Publication Date
- 2026-05-25
AI Technical Summary
Existing graphics processing pipelines face inefficiencies in managing attribute data allocation and deallocation, leading to unnecessary memory usage and bandwidth consumption due to ineffective handling of data that is no longer needed.
A discard engine is introduced to collect deallocation messages from pixel shaders and determine when attribute data is no longer required, sending discard commands to caches to invalidate the data and prevent write-back, thereby optimizing memory usage.
The solution reduces memory bandwidth usage and improves efficiency by ensuring that attribute data is promptly deallocated when no longer needed, enhancing the performance of graphics processing tasks.
Smart Images

Figure 0007864835000001 
Figure 0007864835000002 
Figure 0007864835000003
Abstract
Description
Background Art
[0001] (Description of Related Art) Three-dimensional (3-D) graphics are often processed using a graphics pipeline formed by a sequence of programmable shaders and fixed-function hardware blocks. For example, a 3-D model of an object visible within a frame can be represented by a set of triangles, other polygons, or patches that are processed by the graphics pipeline to generate the pixel values for display to the user. Triangles, other polygons, and patches are collectively referred to as primitives.
[0002] In a typical graphics pipeline, a sequence of work items, which may be referred to as threads, is processed to output a final result. Each processing element executes a respective instantiation of a particular work item to process the incoming data. A work item is any one of a set of parallel executions of a kernel that is invoked on a compute unit. A work item is distinguished from other executions within the set by a global ID and a local ID. As used herein, the term "compute unit" is defined as a set of processing elements (e.g., a single-instruction, multiple-data (SIMD) unit) that perform synchronous execution of multiple work items. The number of processing elements per compute unit can vary from embodiment to embodiment. A subset of work items within a work group that are executed together simultaneously on a compute unit may be referred to as a wavefront, warp, or vector. The width of a wavefront is a characteristic of the hardware of the compute unit.
[0003] The graphics processing pipeline includes several stages that perform individual tasks such as vertex position and attribute transformation and pixel color calculation. Many of these tasks are performed in parallel by a set of processing elements for individual work items in the wavefront that traverses the pipeline. The graphics processing pipeline is constantly being updated and improved.
[0004] The advantages of the methods and mechanisms described herein can be better understood by referring to the following description in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0005] [Figure 1] This is a block diagram of one embodiment of a computing system. [Figure 2] This is a block diagram of one embodiment of a GPU. [Figure 3] This is a block diagram of one embodiment of the computing unit. [Figure 4] This is a block diagram of one embodiment of a discard engine. [Figure 5] This is a generalized flowchart illustrating one embodiment of a method for operating a waste engine. [Figure 6] This is a generalized flowchart illustrating one embodiment of a method for managing discard tables. [Figure 7] This is a generalized flowchart illustrating one embodiment of a method for generating a discard command. [Figure 8] This is a generalized flowchart illustrating another embodiment of a method for generating a discard command. [Figure 9] This is a generalized flowchart illustrating one embodiment of a method for implementing an ordered discard command generation scheme. [Modes for carrying out the invention]
[0006] The following description includes numerous specific details to provide a full understanding of the methods and mechanisms presented herein. However, those skilled in the art should recognize that various embodiments may be carried out without these specific details. In some cases, well-known structures, components, signals, computer program instructions, and techniques are not shown in detail to avoid obscuring the approaches described herein. For simplicity and clarity, it should be understood that the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to others.
[0007] Various systems, devices, and methods for implementing a discard engine in a graphics pipeline are disclosed herein. In one embodiment, the system includes a graphics pipeline having a geometry engine that invokes a shader that generates attribute data for each vertex of a set of primitives. The attribute data is consumed by pixel shaders, each pixel shader generating an attribute deallocation message when the pixel shader no longer needs the attribute data. The discard engine collects deallocations from multiple pixel shaders and determines when the attribute data is no longer needed. When a block of attributes has been consumed by all potential pixel shader consumers, the discard engine deallocates the given block of attributes. The discard engine sends a discard command to a cache so that the attribute data is invalid and cannot be written back into memory.
[0008] Referring to Figure 1, a block diagram of one embodiment of the computing system 100 is shown. In one embodiment, the computing system 100 includes at least processors 105A to 105N, an input / output (I / O) interface 120, a bus 125, a memory controller 130, a network interface 135, a memory device 140, a display controller 150, and a display 155. In other embodiments, the computing system 100 includes other components and / or the computing system 100 is arranged differently. Processors 105A to 105N represent any number of processors included in the system 100.
[0009] In one embodiment, processor 105A is a general-purpose processor such as a central processing unit (CPU). In this embodiment, processor 105A runs a driver 110 (e.g., a graphics driver) for communicating with one or more other processors in the system 100 and / or for controlling the operation of one or more of those processors. In one embodiment, processor 105N is a data-parallel processor having a highly parallel architecture such as a graphics processing unit (GPU) that processes data, performs parallel processing of workloads, renders pixels for the display controller 150 to drive the display 155 and / or performs other workloads.
[0010] GPUs can perform graphics processing tasks required by end-user applications such as video game applications. GPUs are also increasingly being used to perform other tasks unrelated to graphics. Other data parallel processors that may be included in system 100 include digital signal processors (DSPs), field programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). In some embodiments, processors 105A to 105N include multiple data parallel processors.
[0011] In some embodiments, an application running on processor 105A invokes a user-mode driver 110 (or a similar GPU driver) using a graphics application programming interface (API). In one embodiment, the user-mode driver 110 issues one or more commands to the GPU to render one or more graphics primitives into a displayable graphics image. Based on the graphics instructions issued to the user-mode driver 110 by the application, the user-mode driver 110 generates one or more graphics commands that specify one or more actions of the GPU to perform graphics rendering. In some embodiments, the user-mode driver 110 is part of an application running on the CPU. For example, the user-mode driver 110 may be part of a game application running on the CPU. In one embodiment, if the driver 110 is a kernel-mode driver, the driver 110 is part of an operating system (OS) running on the CPU.
[0012] The memory controller 130 represents any number and type of memory controllers accessible by processors 105A to 105N. While the memory controller 130 is shown as being separate from processors 105A to 105N, it should be understood that this represents only one possible embodiment. In other embodiments, the memory controller 130 can be embedded within one or more of processors 105A to 105N. The memory controller 130 is coupled to any number and type of memory devices 140.
[0013] The memory device 140 represents any number and type of devices that include memory and / or memory elements. For example, the types of memory within the memory device 140 include Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), NAND flash memory, NOR flash memory, Ferroelectric Random Access Memory (FeRAM), etc. The memory device 140 stores program instructions 145, which may include a first set of program instructions for applications, a second set of program instructions for driver components, etc. Alternatively, program instructions 145 or a portion thereof may be stored in a memory or cache device near processor 105A and / or processor 105N.
[0014] The I / O interface 120 represents any number and type of I / O interfaces (e.g., peripheral component interconnect (PCI) bus, PCI-Extended (PCI-X), PCI Express (PCIE) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)). Various types of peripheral devices (not shown) are coupled to the I / O interface 120. Such peripheral devices include, but are not limited to, displays, keyboards, mice, printers, scanners, joysticks, other types of game controllers, media recording devices, and external storage devices. The network interface 135 can receive and transmit network messages over the network.
[0015] In various embodiments, the computing system 100 is a computer, laptop, mobile device, game console, server, streaming device, wearable device, or any other type of computing system or device. Note that the number of components of the computing system 100 varies from embodiment to embodiment. For example, in other embodiments, there may be more or fewer components than those shown in Figure 1. Also note that in other embodiments, the computing system 100 includes other components not shown in Figure 1. In addition, in other embodiments, the computing system 100 is structured in ways other than those shown in Figure 1.
[0016] Looking at Figure 2, a block diagram of one embodiment of the GPU 200 is shown. In one embodiment, the command processor 210 processes commands received from the host processor (e.g., processor 105A in Figure 1). The command processor 210 also sets the GPU 200 to the correct state in order to execute the received commands. In various embodiments, the received commands are intended to cause the GPU 200 to render various scenes of a video game application, video, or other application. Based on the commands received from the command processor 210, the geometry engine 220 processes indices according to the topology (e.g., points, lines, triangles) and connectivity of the scene being rendered. For example, in one embodiment, the geometry engine 220 processes meshes based on quadrilateral or triangular primitives representing three-dimensional (3D) objects. In this example, the geometry engine 220 uses fixed-function operations to read vertices from a buffer (stored in cache / memory 275), form mesh geometry, and generate pipeline work items.
[0017] The geometry engine 220 is coupled to any number of shader processor inputs (SPIs) 230A to 230N, the number of which varies according to the embodiment. The SPIs 230A to 230N accumulate work items until enough work items are received to generate a wavefront, and then the SPIs 230A to 230N each invoke the wavefront on the compute units 240A to 240N. Depending on the embodiment, the wavefront may contain 32 work items, 64 work items, or any other number of work items. Note that the terms “work item” and “thread” may be used interchangeably herein.
[0018] Computation units 240A-240N execute shader programs to process wavefronts received from SPI230A-230N. In one embodiment, the geometry frontend includes vertex and hull shaders that operate on higher-order primitives, such as patches representing a three-dimensional (3D) model of the scene. In this embodiment, the geometry frontend provides higher-order primitives to shaders that generate lower-order primitives from higher-order primitives. The lower-order primitives are then duplicated, shaded, and / or subdivided before being processed by the pixel engine. The pixel engine performs culling, rasterization, depth testing, color blending, etc., on the primitives to generate fragments or pixels for display. In other embodiments, shaders of other types and / or sequences are used to process various wavefronts across the pipeline.
[0019] Computation units 240A to 240N read from and write to cache / memory 275 during the execution of shader programs. For example, in one embodiment, geometry engine 220 invokes shaders on computation units 240A to 240N that generate attribute data to be written to ring buffer 285. Attribute data can include any non-positional data associated with vertices. For example, attribute data can include, but is not limited to, color, texture, translucency, surface normals, etc. At a later point, pixel shaders invoked on computation units 240A to 240N consume the attribute data from ring buffer 285. There may be multiple pixels that need to access the same attribute data, and therefore, to track when attribute data can be discarded, discard engine 235 tracks deallocation from pixel shaders. When a given block of attributes has been consumed by all of its consumers, discard engine 235 sends a discard command to cache 275 with the address range of the given block of attributes. Upon receiving a discard command, cache 275 invalidates the corresponding data and prevents the dirty data from being written back to other cache levels and / or memory.
[0020] Shader export units 250A - 250N manage the outputs from compute units 240A - 240N and transfer the outputs to either primitive assemblers 260A - 260N or backend 280. For example, in one embodiment, shader export units 250A - 250N export the position of vertices after transformation. Primitive assemblers 260A - 260N accumulate and connect vertices spanning primitives and pass the primitives to scan converters 270A - 270N that perform rasterization. Primitive assemblers 260A - 260N also perform culling of non - visible primitives. Scan converters 270A - 270N determine which pixels are covered by the primitives, transfer pixel data to SPI 230A - 230N, and then SPI 210A - 230N launches pixel shader wavefronts on compute units 240A - 240N.
[0021] Referring to FIG. 3, a block diagram of one embodiment of compute unit 300 is shown. In one embodiment, compute unit 300 includes at least SIMD 310A - 310N, sequencer 305, instruction buffer 340, and local data share (LDS) 350. Note that compute unit 300 may include other components not shown in FIG. 3 to avoid obscuring the figure. In one embodiment, compute units 240A - 240N (of FIG. 2) include the components of compute unit 300.
[0022] In one embodiment, the computing unit 300 executes kernel instructions on any number of wavefronts. These instructions are stored in the instruction buffer 340 and scheduled by the sequencer 305 for execution on SIMD310A-310N. In one embodiment, the width of the wavefront corresponds to the number of lanes on lanes 315A-315N, 320A-320N, and 325A-325N within SIMD310A-310N. Each lane 315A-315N, 320A-320N, and 325A-325N of SIMD310A-310N may also be referred to as an "execution unit" or "processing element".
[0023] In one embodiment, the GPU 300 receives multiple instructions for a wavefront having several work items. When the work items are executed on SIMD 310A-310N, each work item is allocated a corresponding portion of vector general-purpose registers (VGPR) 330A-330N, scalar general-purpose registers (SGPR) 335A-335N, and local data share (LDS) 350. Note that the letter "N," when it appears next to various structures in this specification, generally refers to any number of elements of that structure (e.g., any number of SIMD 310A-310N). In addition, the different references in Figure 3 that use the letter "N" (e.g., SIMD310A~310N and lanes 315A~315N) are not intended to indicate that an equal number of different elements are provided (e.g., the number of SIMD310A~310N may differ from the number of lanes 315A~315N).
[0024] Turning to FIG. 4, a block diagram of one embodiment of the discard engine 430 is shown. As shown in FIG. 4, the discard engine 430 is coupled to a pixel shader 410 and a cache 420. The pixel shader 410 represents any number of pixel shaders. During execution, the pixel shader 410 consumes attribute data 425 from the cache 420. When a given pixel shader 410 completes consuming the corresponding attribute data, the given pixel shader 410 sends a deallocation message to the discard engine 430.
[0025] In one embodiment, the discard engine 430 uses a table 440 to track deallocation messages from the pixel shader 410. In one embodiment, each entry in the table 440 includes an attribute address range field 450, an identifier (ID) field 460 of the oldest pixel shader consumer, a received deallocation count field 470, a bin ID 480, and any number of other fields. In other embodiments, each entry in the table 440 can be structured in other suitable manners and / or include other fields. When the discard engine 430 determines that a given attribute range has been consumed by all of its consumers, the discard engine 430 sends a discard command regarding the given attribute range to the cache 420. In response to receiving the discard command, the cache 420 invalidates the corresponding data and prevents write-back of dirty data, which helps reduce memory bandwidth usage.
[0026] Referring to FIG.
[0027] The geometry engine invokes a shader that generates attribute data (block 505). The attribute data, after being generated, is stored in one or more caches (block 510). At a later point, a pixel shader that consumes the attribute data is invoked (block 515). The pixel shader sends a deallocation message to the discard engine as it consumes a portion of the attribute data (block 520). The discard engine collects the deallocation messages and tracks if a given portion of the attribute data has been consumed by all of its corresponding pixel shader consumers (block 525). The discard engine sends a discard command to one or more caches if it can discard the corresponding portion of the attribute data from the cache (block 530). Upon receiving the deallocation command, the cache invalidates the corresponding attribute data and prevents it from being written back to a lower cache level and / or memory (block 535). After block 535, method 500 terminates.
[0028] Turning to Figure 6, one embodiment of method 600 for managing the discard table is shown. A discard engine (e.g., discard engine 430 in Figure 4) receives a deallocation message from the pixel shader for a predetermined range of attribute data (block 605). In response to receiving the deallocation message, the discard engine searches the discard table (e.g., discard table 440) for entries corresponding to the predetermined range of attribute data (block 610). In one embodiment, the discard engine searches for the predetermined range of attribute data based on the memory addresses of the predetermined range. In another embodiment, the discard engine searches for the predetermined range of attribute data based on the pixel coordinates of the predetermined range of attribute data.
[0029] Next, the discard engine increments the count in the Count of Received Deallocated Fields for the matching entry (block 615). The discard engine also determines whether the pixel shader that generated the deallocated message is older than the pixel shader whose ID is stored in the entry's oldest pixel shader consumer field (block 620). In one embodiment, the discard engine determines which pixel shader is older based on the pixel shader IDs, with smaller IDs being considered older than larger IDs. In another embodiment, the discard engine uses other techniques to determine the relative age of the pixel shaders.
[0030] If the pixel shader that generated the deallocation message is older than the pixel shader whose ID is stored in the oldest pixel shader consumer field of the entry (condition block 625: "yes"), the pixel shader replaces the existing ID in the oldest pixel shader consumer field of the matching entry with the ID of the pixel shader that generated the deallocation message (block 630). Otherwise, if the pixel shader that generated the deallocation message is younger than the pixel shader whose ID is stored in the oldest pixel shader consumer field of the entry (condition block 625: "no"), the oldest pixel shader consumer field of the matching entry remains the same (block 635). After blocks 630 and 635, method 600 ends.
[0031] Referring to Figure 7, one embodiment of method 700 for generating a discard command is shown. A discard engine (e.g., discard engine 430 in Figure 4) receives a bin completion signal from a shader processor input (SPI) (e.g., SPI230A in Figure 2) for a given bin being processed (block 705). As used herein, the term “bin” is defined as a region of screen space. In one embodiment, screen space is divided into a plurality of rectangular regions or bins. The discard engine extracts a pixel shader ID from the bin completion signal, which identifies the youngest pixel shader that processed the bin (block 710). Next, the discard engine searches a discard table (e.g., discard table 440) for entries corresponding to pixel shaders with an ID higher than the pixel shader ID extracted from the bin completion signal (block 715). Then, the discard engine generates and transmits a discard command for a range of attribute data corresponding to entries with an ID higher than the pixel shader ID extracted from the bin completion signal (block 720). After block 720, method 700 terminates.
[0032] Turning to Figure 8, another embodiment of method 800 for generating a discard command is shown. A discard engine (e.g., discard engine 430 in Figure 4) receives a bin completion signal from a shader processor input (SPI) (e.g., SPI230A in Figure 2) for a given bin of a primitive (block 805). The discard engine extracts a bin ID from the bin completion signal, which identifies the bin that has just been processed by the pixel shader (block 810). Next, the discard engine searches a discard table (e.g., discard table 440) for an entry corresponding to the bin ID extracted from the bin completion signal (block 815). Then, the discard engine generates and transmits a discard command for a range of attribute data corresponding to an entry having the same bin ID as the bin ID extracted from the bin completion signal (block 820). After block 820, method 800 terminates.
[0033] Referring to Figure 9, one embodiment of method 900 for implementing an ordered discard command generation scheme is shown. A discard engine (e.g., discard engine 430 in Figure 4) maintains a deallocation counter for each pixel shader (block 905). When a deallocation message is received by the discard engine, the discard engine increments a predetermined deallocation counter corresponding to the specific pixel shader that generated the deallocation message (block 910). The discard engine monitors the counter for each pixel shader (block 915). If the counter is greater than 0 for all pixel shaders (condition block 920: "yes"), this means that all pixel shaders have completed the oldest group of attribute data, and therefore the discard engine sends a discard command for the oldest group of attribute data to one or more caches (block 925). The discard engine then decrements each counter (block 930). After block 930, method 900 returns to block 910. If any counter for any of the pixel shaders is still 0 (condition block 920: "no"), indicating that the oldest group of attribute data is still in use, then method 900 returns to block 910.
[0034] In various embodiments, the methods and / or mechanisms described herein are implemented using program instructions for a software application. For example, program instructions that can be executed by a general-purpose or dedicated processor are intended. In various embodiments, such program instructions are expressed in a high-level programming language. In other embodiments, program instructions are compiled from a high-level programming language into binary, intermediate, or other forms. Alternatively, program instructions describing the behavior or design of hardware are written. Such program instructions are expressed in a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog is used. In various embodiments, program instructions are stored in any of various non-temporary computer-readable storage media. The storage media are accessible by the computing system in use to provide the computing system with program instructions for program execution. Generally speaking, such a computing system includes at least one memory and one or more processors configured to execute program instructions.
[0035] It should be emphasized that the embodiments described above are merely non-limiting examples of embodiments. A number of variations and modifications will become apparent to those skilled in the art once the above disclosure is fully understood. The following claims are intended to be construed as encompassing all such variations and modifications.
Claims
1. It is a device, A cache configured to store attribute data for the vertices of each primitive in a set of primitives, Multiple computing units configured to execute a pixel shader to consume the aforementioned attribute data, Equipped with a discarded engine, The aforementioned discard engine is The process involves receiving a bin completion signal indicating that the binning of primitives has been processed by the pixel shaders of the aforementioned multiple computing units, The bin identifier (ID) that identifies the processed bin is obtained from the bin completion signal, In response to searching a discard table containing entries corresponding to different address ranges of attribute data, a discard command is generated and transmitted for the address range of attribute data having the same bin ID as the bin ID obtained from the bin completion signal. It is configured to do, Device.
2. The cache is configured to invalidate the attribute data within the cache in response to the receipt of the discard command. The apparatus according to claim 1.
3. The cache is configured to prevent the attribute data from being written to another level of cache or memory in response to the receipt of the discard command. The apparatus according to claim 2.
4. The aforementioned discard engine is Tracking multiple attribute deassignment messages generated by the aforementioned pixel shader, A discard command is transmitted to the cache, at least partially based on the number of deallocation messages generated by the pixel shader. It is configured to do, The apparatus according to claim 1.
5. The aforementioned discard engine is Searching the discard table for entries corresponding to pixel shaders having an ID higher than the bin ID obtained from the bin completion signal, To transmit a discard command for the address range of attribute data corresponding to the pixel shader having the aforementioned high ID, It is configured to do, The apparatus according to claim 1.
6. The discard engine is configured to maintain a discard table containing entries corresponding to different address ranges of attribute data. The apparatus according to claim 5.
7. The aforementioned discard engine is Maintain a discard table containing entries corresponding to different address ranges of attribute data, Searching the discard table for entries corresponding to the bin ID obtained from the bin completion signal, It is configured to do, The apparatus according to claim 1.
8. It is a method, The cache circuit stores attribute data for the vertices of each primitive in the set of primitives, The circuits of multiple computing units receive a bin completion signal indicating that the bins of primitives have been processed by the pixel shaders of the multiple computing units, The discard engine obtains a bin identifier (ID) that identifies the processed bin from the bin completion signal, The discard engine includes, in response to a search of a discard table containing entries corresponding to different address ranges of attribute data, generating and transmitting discard commands for address ranges of attribute data having the same bin ID as the bin ID obtained from the bin completion signal, method.
9. In response to receiving the aforementioned discard command, the following actions are taken: The method of claim 8.
10. In response to the receipt of the discard command, the following measures are taken to prevent the attribute data from being written to another level of cache or memory: The method of claim 9.
11. The discard engine tracks a plurality of attribute deallocation messages generated by the pixel shader, The discard engine transmits a discard command to the cache, at least partially based on the number of deallocation messages generated by the pixel shader. including, The method of claim 8.
12. Searching the discard table for entries corresponding to pixel shaders having an ID higher than the bin ID obtained from the bin completion signal, This includes transmitting a discard command for the address range of attribute data corresponding to the pixel shader having the aforementioned high ID, The method of claim 8.
13. Includes maintaining a discard table which includes entries corresponding to different address ranges of attribute data, The method according to claim 12.
14. Maintaining a discard table that includes entries corresponding to different address ranges of attribute data, This includes searching the discard table for entries corresponding to the bin ID obtained from the bin completion signal, The method according to claim 12.
15. It is a system, Cache and, Equipped with a discarded engine, The aforementioned discard engine is Receiving a bin completion signal indicating that the primitive bins have been processed by the pixel shaders of multiple computing units, The bin identifier (ID) that identifies the processed bin is obtained from the bin completion signal, In response to searching a discard table containing entries corresponding to different address ranges of attribute data, a discard command is generated and transmitted for the address range of attribute data having the same bin ID as the bin ID obtained from the bin completion signal. It is configured to do, system.
16. The cache is configured to invalidate the attribute data within the cache in response to the receipt of the discard command. The system according to claim 15.
17. The cache is configured to prevent the attribute data from being written to another level of cache or memory in response to the receipt of the discard command. The system according to claim 16.
18. The discard engine is configured to maintain a discard table having entries for different ranges of attribute data. The system according to claim 15.
19. The aforementioned discard engine is Tracking the oldest pixel shader that is consuming a given range of attribute data, The identifier (ID) of the oldest pixel shader is stored in a predetermined entry corresponding to a predetermined range of the attribute data, It is configured to do, The system according to claim 15.
20. The discard engine is configured to maintain a discard table having entries for different address ranges of attribute data. The system according to claim 19.