Throttle the hull shader based on the tessellation factor in the graphics pipeline
Through the throttling adjustment technology based on mosaic factor and primitive startup time interval, the resource shortage caused by too many start thread groups of shell shader circuits is solved, and the performance of the graphics pipeline is improved.
Patent Information
- Application Number
- CN202180084587.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-15
- Filing Date
- 2021-12-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-01
AI Technical Summary
When the existing graphics processing pipeline is processed, due to the excessive thread groups started by the shell shader circuit, the domain shader lacks the resources required to process mosaic elements, which affects the performance of the graphics pipeline.
The thread groups initiated from the shell shader circuit are selectively throttling based on the mosaic factor and the primitive start time interval, ensuring that the domain shader has sufficient resources to process the mosaic element.
Effectively maintain the resource balance between shell shaders and domain shaders, and improve the performance of graphics pipelines, especially in scenarios where the mosaic factor is less than or equal to one.
Smart Images

Figure CN116745801B_ABST
Abstract
Description
Background Art
[0001] A graphics processing unit (GPU) implements a graphics processing pipeline that simultaneously processes copies of commands retrieved from a command buffer. The graphics pipeline includes one or more shaders and one or more fixed-function hardware blocks that execute using the resources of the graphics pipeline. The graphics pipeline is typically divided into a geometry portion that performs geometric operations on patches or other primitives (such as triangles formed by vertices and edges and representing portions of an image). The shaders in the geometry portion can include vertex shaders, hull shaders, domain shaders, and geometry shaders. The geometry portion of the graphics pipeline is complete when the primitives generated by the geometry portion of the pipeline (e.g., by one or more scan converters) are rasterized to form a set of pixels representing a portion of the image. Subsequent processing of the pixels is called pixel processing and includes operations performed by shaders (such as pixel shaders that execute using the resources of the graphics pipeline). GPUs and other multi-threaded processing units typically implement multiple processing elements (which are also called processor cores or compute units) that execute multiple instances of a single program simultaneously as a single wave on multiple data sets. A hierarchical execution model is used to match the hierarchical structure implemented in the hardware. The execution model defines a kernel of instructions executed by all waves (also called wavefronts, threads, streams, or work items). Brief Description of the Drawings
[0002] The present disclosure can be better understood by reference to the accompanying drawings, and many of its features and advantages will be apparent to those of ordinary skill in the art. The same reference numerals are used in different drawings to denote similar or identical items.
[0003] Figure 1 is a block diagram of a processing system according to some embodiments.
[0004] Figure 2 depicts a graphics pipeline according to some embodiments capable of processing high-order geometric primitives to generate a rasterized image of a three-dimensional (3D) scene at a predetermined resolution.
[0005] Figure 3 is a block diagram of a first portion of a processing system according to some embodiments that selectively throttles a thread group launched by a hull shader circuit.
[0006] Figure 4 is a block diagram of a second portion of a processing system according to some embodiments that selectively throttles a thread group launched by a hull shader circuit.
[0007] Figure 5 is a first portion of a flowchart of a method according to some embodiments for estimating a primitive launch time interval of a domain shader using a total count and an error count.
[0008] Figure 6 Flowchart of the second part of a method for estimating the primitive launch time interval of a domain shader using total count and error count, according to some embodiments.
[0009] Figure 7 Flowchart of a method for selectively throttling the wave launch from a hull shader, according to some embodiments. DETAILED DESCRIPTION
[0010] The hull shader circuit in the geometry portion of the graphics pipeline launches waves of control points of patches processed by the hull shader. The hull shader also generates a tessellation factor indicating the subdivision of the patch. The patch and the tessellation factor processed by the hull shader are passed to the tessellator in the graphics pipeline. Before processing the tessellated primitives in the domain shader, the tessellator uses the tessellation factor to subdivide the patch into other primitives such as triangles. Thus, the domain shader typically processes a larger number of primitives than the hull shader. For example, if the tessellation factor for a quadrilateral patch processed by the hull shader is sixteen, the domain shader processes 512 triangles in response to receiving the patch from the hull shader. Patches are launched by the hull shader circuit based on a greedy algorithm that attempts to use as many resources of the graphics pipeline as possible. Launching the hull shader waves based on the greedy algorithm can leave the domain shader lacking the resources required to process the tessellated primitives. Some graphics pipelines are configured to limit the number of in-flight waves by constraining the number of compute units that can be allocated to the hull shader for processing waves. However, when the magnification of the primitives launched by the hull shader is almost non-existent (e.g., when the tessellation factor is less than or equal to one), the static limit on the number of available compute units degrades the performance of the graphics pipeline.
[0011] Figures 1 to 7Systems and techniques are disclosed for selectively launching waves from a first shader based on a measure of the graphics pipeline resources consumed by a first shader of a first type and a second shader of a second type to maintain a balance between the graphics pipeline resources consumed by the first shader and the second shader. In some embodiments, the first shader is a hull shader and the second shader is a domain shader that receives primitives from a tessellator. The hull shader generates a tessellation factor, and the tessellator subdivides (or tessellates) the primitives based on the tessellation factor to generate a plurality of higher resolution primitives. The tessellation factor for a patch launched by the hull shader circuit is saved in a buffer that supplies primitives to the domain shader. A throttling circuit uses the tessellation factor to estimate the time interval required for the domain shader to launch all the primitives from the domain shader, e.g., the number of cycles required to process the higher resolution primitives in the domain shader. This time interval is referred to herein as the "primitive launch time interval". Some embodiments of the throttling circuit include a register bank that stores information indicative of the number of higher resolution primitives (or the number of cycles required to process the higher resolution primitives) associated with a wave corresponding to an entry in the buffer. The stored information is used to set the value of a counter representing the primitive launch time interval for the domain shader. For example, a total counter is incremented by the number of cycles estimated to be required to process the higher resolution primitives in the registers associated with the buffer entry written to the tessellator for processing. In response to the domain shader launch logic completing the processing of the higher resolution primitives associated with a patch, the total counter is iteratively decremented (at the estimated primitive processing rate of the domain shader launch logic). In some embodiments, an error counter is used to modify the total counter based on a measurement of the actual time required to process a primitive in the domain shader before the domain shader launches. The value of the error counter is incremented in response to the measured latency being greater than the latency corresponding to the value of the total counter (e.g., due to backpressure on the domain shader). The value of the error counter is decremented (or set to zero) in response to the measured processing time being less than or equal to the value of the total counter. The combined total counter and error counter are then decremented based on the tessellation factor of the completed patch. Waves are selectively launched from the hull shader based on the values of the total counter and the error counter (if present).
[0012] Figure 1is a block diagram of a processing system 100 according to some embodiments. The processing system 100 includes or has access to a memory 105 implemented using a non-transitory computer-readable medium such as dynamic random access memory (DRAM) or other storage components. However, in some cases, the memory 105 is implemented using other types of memory (including static random access memory (SRAM), non-volatile RAM, etc.). The memory 105 is referred to as external memory because it is implemented external to the processing unit implemented in the processing system 100. The processing system 100 also includes a bus 110 to support communication between entities implemented in the processing system 100 such as the memory 105. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not shown in Figure 1 for clarity.
[0013] In different embodiments, the techniques described herein are used in any of a variety of parallel processors (e.g., vector processors, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), non-scalar processors, highly parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, other multi-threaded processing units, etc.). Figure 1 An example of a parallel processor, specifically a graphics processing unit (GPU) 115, is shown according to some embodiments. The graphics processing unit (GPU) 115 renders images for presentation on a display 120. For example, the GPU 115 renders objects to produce pixel values provided to the display 120, which uses the pixel values to display an image representing the rendered object. The GPU 115 implements multiple computing units (CUs) 121, 122, 123 (collectively referred to herein as "computing units 121 to 123") that execute instructions concurrently or in parallel. In some embodiments, the computing units 121 to 123 include one or more single instruction multiple data (SIMD) units and the computing units 121 to 123 are grouped into workgroup processors, shader arrays, shader engines, etc. The number of computing units 121 to 123 implemented in the GPU 115 is a matter of design choice, and some embodiments of the GPU 115 include more or fewer computing units than Figure 1 shown. The computing units 121 to 123 can be used to implement a graphics pipeline, as discussed herein. Some embodiments of the GPU 115 are used for general computing. The GPU 115 executes instructions such as program code 125 stored in the memory 105, and the GPU 115 stores information such as the results of the executed instructions in the memory 105.
[0014] Processing system 100 also includes a central processing unit (CPU) 130, which is connected to bus 110 and thus communicates with GPU 115 and memory 105 via bus 110. CPU 130 implements multiple processor cores 131, 132, 133 (collectively referred to herein as "processor cores 131 to 133") that execute instructions concurrently or in parallel. The number of processor cores 131 to 133 implemented in CPU 130 is a matter of design choice, and some embodiments include more or fewer processor cores than Figure 1 shown. Processor cores 131 to 133 execute instructions (such as program code 135 stored in memory 105), and CPU 130 stores information (such as the results of the executed instructions) in memory 105. CPU 130 is also capable of initiating graphics processing by issuing draw calls to GPU 115. Some embodiments of CPU 130 implement multiple processor cores that execute instructions simultaneously or in parallel (not shown for clarity Figure 1 in the figure).
[0015] Input / output (I / O) engine 145 processes input or output operations associated with display 120 and other elements of processing system 100 (such as, a keyboard, a mouse, a printer, an external disk, etc.). I / O engine 145 is coupled to bus 110 such that I / O engine 145 communicates with memory 105, GPU 115, or CPU 130. In the illustrated embodiment, I / O engine 145 reads information stored on external storage component 150, which is implemented using a non-transitory computer-readable medium such as a compact disc (CD), a digital versatile disc (DVD), etc. I / O engine 145 is also capable of writing information to external storage component 150, such as the processing results of GPU 115 or CPU 130.
[0016] Processing system 100 implements pipeline circuitry for executing instructions in multiple stages of a pipeline. The pipeline circuitry is implemented in some embodiments of computing units 121 - 123 or processor cores 131 - 133. In some embodiments, the pipeline circuitry is used to implement a graphics pipeline for executing different types of shaders, including but not limited to vertex shaders, hull shaders, domain shaders, geometry shaders, and pixel shaders. Some embodiments of processing system 100 include hull shader circuitry for launching a thread group including one or more primitives. For example, computing units 121 - 123 in GPU 115 can be used to implement hull shader circuitry, as well as circuitry for other shaders and throttled wave launches, as discussed herein. The hull shader circuitry also generates a tessellation factor indicative of the tessellation of the primitive. The throttling circuitry in processing system 100 estimates the primitive launch time interval of the domain shader based on the tessellation factor, and selectively throttles the launch of the thread group from the hull shader circuitry based on the latency of the domain shader and the hull shader latency. In some cases, the throttling circuitry includes a first counter that increments in response to launching a thread group from a buffer, and a second counter that modifies the first counter based on the measured latency of the domain shader.
[0017] Figure 2 Depicts a graphics pipeline 200 according to some embodiments capable of processing high - order geometric primitives to generate a rasterized image of a three - dimensional (3D) scene at a predetermined resolution. The graphics pipeline 200 is implemented in some embodiments of the processing system 100 shown in Figure 1 The illustrated embodiment of the graphics pipeline 200 is implemented according to the DX11 specification. Other embodiments of the graphics pipeline 200 are implemented according to other application programming interfaces (APIs) such as Vulkan, Metal, DX12, etc. The graphics pipeline 200 is subdivided into a geometry processing portion 201 that includes the portion of the graphics pipeline 200 before rasterization and a pixel processing portion 202 that includes the portion of the graphics pipeline 200 after rasterization.
[0018] The graphics pipeline 200 can access storage resources 205, such as a hierarchy of one or more memories or cache memories for implementing buffers and storing vertex data, texture data, etc. In the illustrated embodiment, the storage resources 205 include a load data store (LDS) 206 circuit for storing data and vector general - purpose registers (VGPRs) for storing register values used during rendering by the graphics pipeline 200. Some embodiments of the system memory 105 shown in Figure 1 are used to implement the storage resources 205.
[0019] The input assembler 210 accesses information of objects that define parts of a model representing a scene from the storage resource 205. Examples of primitives are shown as triangles 211 in Figure 2 but other types of primitives are processed in some implementations of the graphics pipeline 200. A triangle 203 includes one or more vertices 212 connected by one or more edges 214 (only one of each type is shown for clarity in Figure 2 ). The vertices 212 are shaded during the geometric processing part 201 of the graphics pipeline 200.
[0020] In the illustrated implementation, the vertex shader 215 implemented in software logically receives a single vertex 212 of a primitive as input and outputs a single vertex. Some implementations of shaders such as the vertex shader 215 implement large-scale single instruction multiple data (SIMD) processing such that multiple vertices are processed simultaneously. The graphics pipeline 200 implements a unified shader model such that all shaders included in the graphics pipeline 200 have the same execution platform on a shared large-scale SIMD computing unit. Thus, shaders including the vertex shader 215 are implemented using a common set of resources herein referred to as the unified shader pool 216.
[0021] The hull shader 218 operates on input higher-order patches or control points that define an input patch. The hull shader 218 outputs a tessellation factor and other patch data such as the control points of the patch being processed in the hull shader 218. The tessellation factor is stored in the storage resource 205 so that the tessellation factor can be accessed by other entities in the graphics pipeline 200. In some implementations, the primitives generated by the hull shader 218 are provided to the tessellator 220. The tessellator 220 receives an object (such as a patch) from the hull shader 218 and generates information identifying the primitives corresponding to the input object, for example, by tessellating the input object based on the tessellation factor generated by the hull shader 218. Tessellation subdivides an input higher-order primitive (such as a patch) into a set of lower-order output primitives that represent a finer level of detail, e.g., as indicated by the tessellation factor that specifies the primitive granularity produced by the tessellation process. Thus, the model of the scene is represented by a smaller number of higher-order primitives (to save memory or bandwidth), and additional detail is added by tessellating the higher-order primitives.
[0022] The domain shader 224 inputs the domain position and (optionally) other patch data. The domain shader 224 operates on the provided information and generates a single vertex for output based on the input domain position and other information. In the illustrated embodiment, the domain shader 224 generates the primitive 222 based on the triangle 211 and the tessellation factor. The domain shader 224 initiates the primitive 222 in response to completion of processing. The geometry shader 226 receives the input primitive from the domain shader 224 and outputs up to four primitives (per input primitive) generated by the geometry shader 226 based on the input primitive. In the illustrated embodiment, the geometry shader 226 generates the output primitive 228 based on the tessellated primitive 222.
[0023] A primitive stream is provided to one or more scan converters 230, and in some embodiments, up to four primitive streams are cascaded to a buffer in the storage resource 205. The scan converter 230 performs shading operations and other operations such as clipping, perspective division, shearing, and viewport selection. The scan converter 230 generates a set of 232 pixels, which are then processed in the pixel processing section 202 of the graphics pipeline 200.
[0024] In the illustrated embodiment, the pixel shader 234 inputs a pixel stream (e.g., including the pixel set 232) and outputs zero or another pixel stream in response to the input pixel stream. The output merger block 236 performs blending, depth, stencil, or other operations on the pixels received from the pixel shader 234.
[0025] Some or all of the shaders in the graphics pipeline 200 perform texture mapping using texture data stored in the storage resource 205. For example, the pixel shader 234 can read texture data from the storage resource 205 and use the texture data to shade one or more pixels. The shaded pixels are then provided to the display for presentation to the user.
[0026] Figure 3 is a block diagram of a first portion of a processing system 300 that selectively throttles the processing of thread groups initiated by a hull shader circuit according to some embodiments. The first portion of the processing system 300 is implemented in some embodiments of the processing system 100 shown in Figure 1 and the graphics pipeline 200 shown in Figure 2 In some embodiments.
[0027] A set of buffers 301, 302, 303, 304 (collectively referred to herein as "buffers 301 to 304") are used to store metadata associated with thread groups initiated by a hull shader circuit (such as Figure 2 the hull shader 218 shown inFigure 3 (not shown in the figure) is associated. In response to starting a thread group for execution on a computing unit or SIMD, the hull shader circuit provides metadata associated with the thread group to the corresponding buffer among buffers 301 to 304. Accordingly, each entry in buffers 301 to 304 includes metadata of the corresponding thread group.
[0028] Buffers 301 to 304 are associated with a set of counters 311, 312, 313, 314 (collectively referred to herein as "counter set 311 to 314"), and these sets of counters have values representing the measured time intervals or latencies for processing the corresponding thread groups in the hull shader. Each counter in counter set 311 to 314 is associated with an entry in the corresponding buffer among buffers 301 to 304. For example, the first counter in counter set 311 is associated with the first entry in buffer 301. When metadata is added to the corresponding entry in one of buffers 301 to 304 in response to the hull shader circuit starting a thread group, the counter starts counting (e.g., incrementing or decrementing).
[0029] Another set of buffers 321 to 324 has entries storing values indicating that the corresponding thread groups have completed processing. For example, an entry is written to buffer 321 in response to a thread group started by the corresponding hull shader circuit completing execution on the computing unit. The entries in the buffer are used to stop the counting by the corresponding counter in one of counter sets 311 to 314. Accordingly, the counter holds a value representing the measured latency of the thread group, e.g., as the number of cycles for processing the thread group. As discussed herein with respect to Figure 4 As discussed, a subset of the values of the counters in counter set 311 to 314 is provided to the second part of processing system 300 via node 1.
[0030] Arbiter 330 selects thread group metadata from buffers 301 to 304 in the order in which the hull shader circuit schedules the thread groups. For example, if the first thread group is scheduled by the hull shader circuit associated with buffer 301 and the second thread group is subsequently scheduled by the hull shader circuit associated with buffer 302, then arbiter 330 selects the thread group metadata from buffer 301 before selecting the thread group metadata from buffer 302. As discussed herein with respect to Figure 4 As discussed, arbiter 330 provides the metadata associated with the thread group to the circuit that extracts the tessellation factors of the thread group via node 2.
[0031] Figure 4 is a block diagram of the second part of processing system 300 that selectively throttles the processing of thread groups started by the hull shader circuit according to some embodiments. The second part of processing system 300 is Figure 1implemented in some implementations of the processing system 100 shown and Figure 2 the graphics pipeline 200 shown.
[0032] Figure 4 The second part of the processing system 300 shown includes circuitry 405 that extracts tessellation factors from a memory 410 and performs processing on the tessellation factors and metadata received from Figure 3 the arbiter 330 shown. Processing the metadata received from the arbiter 330 includes parsing the received thread group to identify the primitives (such as patches) included in the thread group. Then, the patches, tessellation factors, and associated metadata are provided to a buffer 415. Each entry in the buffer 415 includes a patch and its associated tessellation factor and metadata. Then, the information in the entries of the buffer 415 is provided to a patch allocator 420 that allocates the information to output buffers associated with one or more tessellators (such as Figure 2 the tessellator 220 shown) and domain shaders (such as Figure 2 the domain shader 224 shown).
[0033] The circuitry 405 also provides the tessellation factors for the primitives or patches in the thread group to a register 425 in a hull shader throttling circuit 430. Each register in the set of registers 425 stores an estimate of the number of primitives (such as, triangles) generated from a patch based on the value of the tessellation factor applied to the patch of the thread group in the corresponding entry of the buffer 415. The hull shader throttling circuit 430 also includes two counters for throttling the thread groups launched from the hull shader. The first counter 435 has a total count value representing the primitive launch time interval of the domain shader circuit (e.g., the time interval used by the domain shader to process and launch a set of primitives associated with one or more primitives provided by the hull shader). The first counter 435 is incremented in response to providing a patch (and associated tessellation factor and metadata) from the buffer 415 to the patch allocator 420. In some implementations, the first counter 435 is incremented by an amount indicated by the corresponding register in the set of registers 425. For example, the first counter 435 may be incremented by the number of primitives or patches of the patch corresponding to the entry in the buffer 415 in the register.
[0034] The second counter 440 in the hull shader throttling circuit 430 has a value representing an error count that indicates the difference between the measured downstream latency of a patch (e.g., the time interval for processing a primitive by the domain shader) and the predicted downstream primitive launch time interval indicated by the tessellation factor (e.g., the number of primitives generated from the patch based on the tessellation factor). In some embodiments, the second counter 440 increments or decrements based on whether a read enable signal associated with a thread group arrives before or after the second counter 440 has counted down to a predetermined value, such as zero. As discussed herein, the value in the second counter 440 is used to modify the first counter 435 based on the measured domain shader latency such that the value in the first counter 435 indicates the primitive launch time interval required for the domain shader to process primitives after tessellation.
[0035] The hull shader throttling circuit 430 determines the latency of the hull shader based on the values of counters that indicate the measured latency of thread groups launched from the hull shader. The values of the counters are received from registers associated with the shader engine that processes primitives in the hull shader (via node 1), e.g., Figure 3 the values of the counters in the illustrated set of counters 311 to 314. In the illustrated embodiment, the value of the counter indicates the latency as the number of clock cycles required to process the corresponding thread group. The comparison circuit 445 retrieves a predetermined number of counter values, such as eight counter values for the last eight thread groups launched by the hull shader, and uses the retrieved values to determine the average latency of the hull shader. The latency comparison circuit 445 compares the average latency of the hull shader with the primitive launch time interval indicated by the total count in the first counter 435 of the domain shader. As discussed herein, the hull shader throttling circuit 430 then selectively throttles the launching of thread groups from the hull shader circuit based on this comparison.
[0036] Figure 5 is a flowchart of a first portion of a method 500 for estimating a primitive launch time interval of a domain shader using a total count and an error count, according to some embodiments. Method 500 is implemented in some embodiments of Figure 1 the illustrated processing system 100, Figure 2 the illustrated graphics pipeline 200, and Figure 3 and Figure 4 the illustrated processing system 300. In the illustrated embodiment, the throttling circuit is used to implement method 500.
[0037] At block 505, the throttling circuit intercepts the write data of a thread group and then writes it to a FIFO buffer such as Figure 4In the buffer 415 shown. The throttling adjustment circuit uses this information to estimate the number of primitives generated based on the tessellation factors (tf1, tf2) associated with the thread group. For example, the number of primitives is equal to:
[0038] 2 * inside_tf1 * inside_tf2 (for quadrilateral patches)
[0039] floor(1.5 * inside_tf1^2) (for triangles)
[0040] factor1 * factor2 (for isolines)
[0041] Then, the number of primitives is stored in a register (e.g., Figure 4 one of the registers 425 shown) corresponding to the entry in the FIFO buffer for storing thread group data.
[0042] At block 510, in response to the corresponding thread group being written, the first counter indicating the total count is incremented by the number of primitives. At the first read operation, the second counter indicating the error count is loaded with a value equal to the number of primitives at the current position in the buffer.
[0043] At block 515, countdown (or decrement) starts for the first counter (total count) and the second counter (error count). In some embodiments, the first counter and the second counter countdown at the product of the primitive rate of the tessellator and the number of tessellators.
[0044] At decision block 520, the throttling adjustment circuit determines whether the value of the second counter (error count) has reached zero before the throttling adjustment circuit receives a read enable signal. If not, the method 500 flows to block 540. If the second counter reaches zero before receiving the read enable signal, which indicates that the primitive start time interval of the domain shader has been underestimated, the method 500 flows to block 525.
[0045] At block 525, the throttling adjustment circuit increments the second counter (error count) in each clock cycle until a read enable signal is received. If the value of the second counter reaches the maximum value, the value of the second counter is limited to the maximum value so that the second counter does not roll over. At block 530, the throttling adjustment circuit receives a read enable signal. At block 535, the throttling adjustment circuit adds the value of the second counter to the current value of the first counter. Then, the method 500 flows to block 515.
[0046] At block 540, the throttling adjustment circuit receives a read enable signal before the value of the second counter reaches zero. Then, the method 500 flows to connect block 540 to Figure 6Node 1 of decision box 605 in
[0047] Figure 6 is a flowchart of the second part of method 500 for estimating the primitive launch time interval of a domain shader using a total count and an error count according to some embodiments. Decision box 605 is connected via node 1 to Figure 5 box 540 in
[0048] At decision box 605, the throttling circuit determines whether the error count is equal to zero upon receipt of a read enable signal. If so, method 500 flows to box 610 and the next location is loaded into the second counter. Then, method 500 flows via node 2 to Figure 5 box 515 in
[0049] If the error count is not equal to zero (i.e., the value of the error count is greater than zero) upon receipt of a read enable signal, then method 500 flows to box 615. An error count greater than zero indicates that the primitive launch time interval of the domain shader has been overestimated. Thus, at box 615, the throttling circuit subtracts the value in the second counter from the value (total count) in the first counter. Then, method 500 flows via node 2 to Figure 5 box 515 in
[0050] Thus, the first counter has a value indicating the number of cycles between writing a thread group and receiving a subsequent read enable signal. Thus, the total count in the first counter indicates the total domain shader time / delay required to process the primitives generated after tessellation in the processing thread group. Thus, the total count can be used to compare the domain shader delay with the hull shader delay and selectively throttle the waves launched from the hull shader to maintain a balance between the consumption rates of the thread groups in the hull shader and the domain shader.
[0051] Figure 7 is a flowchart of method 700 for selectively throttling wave launches from a hull shader according to some embodiments. Method 700 is implemented in some embodiments of the processing system 100 shown in Figure 1 the graphics pipeline 200 shown in Figure 2 and the processing system 300 shown in Figure 3 and Figure 4 In the illustrated embodiment, the throttling circuit is used to implement method 500.
[0052] At box 705, the throttling circuit determines the total count indicated by a first counter in the throttling circuit, which total count indicates the primitive launch time interval of the domain shader. At box 710, as discussed herein, the throttling circuit determines the average hull shader delay, for example, using the value of a counter associated with the thread group processed by the shader engine.
[0053] At decision block 715, the throttling circuit compares the total count to the hull shader latency and determines whether the total count is greater than eight times the hull shader latency. If so, the comparison indicates that the hull shader is running ahead of the domain shader and throttling should occur. Accordingly, method 700 flows to block 720, and the hull shader is throttled to achieve two in-flight thread groups per shader engine. If the total count is less than or equal to eight times the hull shader latency, method 700 flows to decision block 725.
[0054] At decision block 725, the throttling circuit compares the total count to the hull shader latency and determines whether the total count is greater than four times the hull shader latency. If so, the comparison indicates that the hull shader is running ahead of the domain shader, but not as much as when the total count is greater than eight times the hull shader latency. Nevertheless, the hull shader should be throttled. Accordingly, method 700 flows to block 730, and the hull shader is throttled to achieve four in-flight thread groups per shader engine. If the total count is less than or equal to four times the hull shader latency, method 700 flows to decision block 735.
[0055] At decision block 735, the throttling circuit compares the total count to the hull shader latency and determines whether the total count is greater than two times the hull shader latency. If so, the comparison indicates that the hull shader is running ahead of the domain shader, but not as much as when the total count is greater than four times the hull shader latency. Nevertheless, the hull shader should be throttled. Accordingly, method 700 flows to block 740, and the hull shader is throttled to achieve eight in-flight thread groups per shader engine. If the total count is less than or equal to two times the hull shader latency, method 700 flows to block 745 and disables throttling of the hull shader.
[0056] A computer-readable storage medium can include any non-transitory storage medium or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media can include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, or magnetic hard disk drives), volatile memory (e.g., random access memory (RAM) or cache memory), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. The computer-readable storage medium can be embedded in the computing system (e.g., system RAM or ROM), fixedly attached to the computing system (e.g., magnetic hard disk drive), removably attached to the computing system (e.g., optical disc or universal serial bus (USB)-based flash memory), or coupled to the computer system via a wired or wireless network (e.g., network-attached storage (NAS)).
[0057] In some embodiments, certain aspects of the above techniques can be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions that are stored on or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software can include instructions and certain data that, when executed by one or more processors, manipulate the one or more processors to perform one or more aspects of the above techniques. The non-transitory computer-readable storage medium can include, for example, disk or optical storage devices, solid-state storage devices such as flash memory, cache memory, random access memory (RAM), or other one or more non-volatile memory devices. The executable instructions stored on the non-transitory computer-readable storage medium can be source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executed by one or more processors.
[0058] It should be noted that not all activities or elements described above in the general description are required, a portion of a particular activity or device may not be required, and one or more additional activities may be performed or elements may be included in addition to those described. Further, the order in which activities are listed is not necessarily the order in which they are performed. Additionally, these concepts have been described with reference to specific embodiments. However, those of ordinary skill in the art understand that various modifications and changes can be made without departing from the scope of the present disclosure. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of the present disclosure.
[0059] The foregoing has described benefits, other advantages, and solutions to problems with specific embodiments. However, the benefits, advantages, solutions to problems, and any feature that may cause any benefit, advantage, or solution to occur or become more pronounced should not be construed as a critical, required, or essential feature of any or all embodiments. Additionally, the specific embodiments disclosed above are illustrative only, as the disclosed subject matter may be modified and practiced in different but equivalent manners that are obvious to those of ordinary skill in the art who benefit from the teachings herein. The details of the construction or design shown herein are not intended to be limiting. Thus, it is evident that the specific embodiments disclosed above may be varied or modified, and all such variations are considered to be within the scope of the disclosed subject matter.
Claims
1. A device for selectively throttling the start of a thread group from a hull shader circuit, the device comprises: a processor configured to estimate a primitive launch time interval of a domain shader circuit based on a tessellation factor generated by the hull shader circuit, and throttle the start of the thread group from the hull shader circuit based on the primitive launch time interval and a latency of the hull shader circuit.
2. The device according to claim 1, further comprising the hull shader circuit configured to start a thread group including one or more primitives and generate the tessellation factor to indicate a subdivision of the one or more primitives.
3. The device according to claim 2, further comprising a tessellator configured to subdivide the one or more primitives into higher resolution primitives based on the tessellation factor.
4. The device according to claim 3, wherein the processor is configured to estimate a number of cycles for processing the higher resolution primitives at the domain shader circuit based on the tessellation factor and estimate the primitive launch time interval based on the number of cycles.
5. The device according to claim 4, further comprises: a buffer including entries configured to store the thread group started by the hull shader circuit; and a set of registers corresponding to the entries in the buffer, wherein the set of registers stores information indicating the primitive launch time interval estimated for the thread group in the entries.
6. The device according to claim 5, wherein each register in the set of registers is configured to store information indicating at least one of a number of the higher resolution primitives in the thread group associated with the register and a number of cycles required to process the higher resolution primitives in the thread group associated with the register.
7. The device according to claim 6, further comprises: a first counter that increments in response to starting a thread group from the buffer, wherein the first counter increments by an amount indicated by a corresponding register in the set of registers; and a second counter configured to modify the first counter based on a measured latency of the domain shader circuit.
8. The device according to claim 7, wherein the first counter indicates the primitive launch time interval for the domain shader circuit to process primitives after subdividing the one or more primitives into higher resolution primitives based on the tessellation factor.
9. The device according to claim 8, wherein the second counter increments or decrements based on whether a read enable signal associated with the thread group arrives before or after the second counter has counted down to zero.
10. The device according to claim 1, wherein the processor is configured to determine the latency of the circuit based on a value of a counter, the value of the counter indicating a number of primitives in the thread group started from the hull shader circuit.
11. The apparatus according to claim 10, wherein the processor is configured to determine the number of thread groups launched by the hull shader circuit based on a comparison of the primitive launch time interval of the domain shader circuit with the latency of the hull shader circuit.
12. A method for selectively throttling the launching of thread groups from a hull shader, the method comprising: estimating a primitive launch time interval of a domain shader circuit based on a thread group including one or more primitives and a tessellation factor generated by a hull shader circuit; and throttling the launching of subsequent thread groups from the hull shader circuit based on the primitive launch time interval and the latency of the hull shader circuit.
13. The method according to claim 12, further comprising subdividing the one or more primitives into higher resolution primitives based on the tessellation factor.
14. The method according to claim 13, further comprising: estimating the number of cycles for processing the higher resolution primitives at the domain shader circuit based on the tessellation factor; and estimating the primitive launch time interval based on the number of cycles.
15. The method according to claim 14, further comprising: storing thread groups launched by the hull shader circuit in entries of a buffer; and storing information indicating the primitive launch time interval estimated for the thread groups in the entries of the buffer in a set of registers corresponding to the entries of the buffer.
16. The method according to claim 15, further comprising storing in each register of the set of registers information indicating at least one of the number of higher resolution primitives in the thread group associated with the register and the number of cycles required to process the higher resolution primitives in the thread group associated with the register.
17. The method according to claim 12, further comprising determining the latency of the hull shader circuit based on a value of a counter, the value of the counter indicating the number of primitives in a thread group launched from the hull shader circuit.
18. An apparatus for selectively throttling the launching of thread groups from a hull shader, the apparatus comprising: a set of registers configured to store information indicating an estimated domain shader circuit latency for thread groups launched by the hull shader circuit and stored in a buffer, based on a tessellation factor generated by the hull shader circuit; a first counter that increments in response to launching a thread group from the buffer and has a total value representing a primitive launch time interval of the domain shader circuit; a second counter having a value based on a measured domain shader circuit latency, wherein the value of the second counter is used to modify the first counter; and a comparison circuit configured to throttle the launching of thread groups from the hull shader circuit based on the primitive launch time interval.
Citation Information
Patent Citations
Graphics processing unit for adjusting level-of-detail, method of operating the same, and devices including the same
CN105513117A
Automatic configuration of knobs to optimize performance of a graphics pipeline
US20200167985A1