Pipeline Delay Reduction for Coarse Visual Compression

By employing a pipeline latency reduction mode that utilizes on-chip memory for rendering visible primitives during visibility passes and switching to standard mode for data compression, the method addresses latency issues in graphics processing, improving rendering efficiency.

JP2025520240APending Publication Date: 2025-07-03ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024535694
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-30
Filing Date
2023-06-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing graphics processing systems experience latency due to delays in the graphics pipeline caused by waiting for compressed visibility data to be flushed from buffers during the rendering process.

Method used

Implementing a pipeline latency reduction mode that allows for the concurrent rendering of visible primitives using on-chip memory data while the visibility pass is executed, and switching to standard mode for compressing data when a binning threshold is reached, thereby reducing the need to wait for flushed compressed data.

Benefits of technology

This approach reduces the overall rendering time by minimizing latency in the graphics pipeline, enhancing the efficiency of the processing system by allowing simultaneous execution of visibility passes and data rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520240000001_ABST
    Figure 2025520240000001_ABST
Patent Text Reader

Abstract

The processing system divides the rendered image into one or more tiles and executes the visible pass of the primitives of the image. During the visible pass, the processing system generates visible data for each primitive of the draw call of the image based on the visible primitive count and the visible draw call count. In response to the primitive of the draw call being visible within the first tile, the processing system increments the visible primitive count and generates visible data indicating that the primitive of the draw call is rendered using draw call index data stored in on-chip memory. If the primitive is the first visible primitive of the draw call, the processing system increments the visible draw call count. The processing system renders the primitive of the draw call using the draw call index data stored in on-chip memory.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In a graphics processing system, a 3D scene is rendered by accelerated processing units in order to be displayed on a 2D display. To render such a scene, the graphics processing system receives a command stream from an application that represents various primitives to be rendered for the scene. The graphics processing system then renders these primitives according to a graphics pipeline that has different stages each containing instructions to be executed by the graphics processing system. The graphics processing system then displays the rendered primitives as part of the 3D scene to be displayed on the 2D display.

[0002] To help reduce the time required to render the primitives of a scene, the graphics processing system divides the scene into a plurality of tiles and performs a visibility pass of the scene to generate visibility data for each tile. Based on this visibility data, the graphics processing system generates and compresses data to be used later to render the primitives of the scene, reducing the time required to render the primitives. However, if the graphics processing system waits for the compressed data to be available, a delay occurs in the graphics pipeline, reducing the efficiency of the system.

Summary of the Invention

Means for Solving the Problems

[0003] In the embodiments described in this specification, techniques are provided for reducing latency in a graphics pipeline by coarse visibility compression. In an exemplary embodiment, the method includes executing a visibility pass of an image to determine visible primitives within a first tile of the image based on a command stream indicating a plurality of primitives, and rendering the visible primitives based on a comparison of a visible primitive count and a binning threshold.

[0004] In some embodiments, the method further includes incrementing the visible primitive count in response to determining the visible primitives within the first tile. The method may also include generating visible data indicating that the visible primitives are to be rendered using draw call index data stored in on-chip memory in response to the visible primitive count being less than the binning threshold.

[0005] In some embodiments, the visible primitives are rendered using the draw call index data simultaneously as the visibility pass is executed on the image. In some embodiments, the method includes generating visible data indicating vertex data of the visible primitives in response to the visible primitive count being greater than or equal to the binning threshold. The method may include compressing the visible data and storing the compressed visible data in a buffer associated with the first tile. In some embodiments, the method further includes flushing the compressed visible data from the buffer, and the visible primitives are rendered using the flushed visible data.

[0006] In another exemplary embodiment, the method includes generating visible data for a primitive based on a visible draw call count in response to determining that the primitive indicated in the command stream is visible within a first tile of the image. The method further includes rendering the primitive based on the visible data. In some embodiments, generating the visible data includes generating visible data indicating that the primitive should be rendered using draw call index data stored in on-chip memory in response to the visible draw count being less than a first binning threshold.

[0007] In some embodiments, the primitive is rendered using draw call index data concurrently with the visible pass of the image. In some embodiments, the method further includes incrementing the visible draw call count in response to the primitive being the first visible primitive of a draw call. The method may also include generating visible data indicating vertex data of the primitive in response to the visible draw call being greater than or equal to a first binning threshold. In some embodiments, the method further includes compressing the visible data and storing the compressed visible data in a buffer associated with the first tile. The method also includes flushing the compressed visible data from the buffer to memory, and the primitive is rendered using the flushed visible data.

[0008] In another exemplary embodiment, a processor includes one or more processing units including circuitry configured to execute a visible pass of an image based on a command stream indicating a plurality of primitives, to determine visible primitives within a first tile from the plurality of primitives, and to render the visible primitives based on a comparison of a visible primitive count and a binning threshold.

[0009] In some embodiments, visible primitives are rendered further based on a comparison of a visible draw call count with a second binning threshold. One or more processing units may include circuitry configured to generate visible data indicating that visible primitives should be rendered using draw call index data stored in on-chip memory in response to the visible primitive count being less than the binning threshold.

[0010] Visible primitives may be rendered concurrently with the visible pass of the image using draw call index data. In some embodiments, one or more processing units may include circuitry configured to generate visible data indicating vertex data of visible primitives in response to the visible primitive count being greater than or equal to the binning threshold. In some embodiments, one or more processing units include circuitry configured to compress visible data, store the compressed visible data in a buffer associated with a first tile, and flush the compressed visible data from the buffer, and visible primitives are rendered using the flushed visible data.

[0011] The present disclosure may be better understood by reference to the accompanying drawings, and its numerous features and advantages may become apparent to those skilled in the art. The use of the same reference numerals in different drawings indicates similar or identical items.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0013] To help reduce the time it takes for a processing system to render primitives of an image, the processing system first generates and compresses the visible data of each primitive of the image, and then performs coarse visible compression by rendering the primitives using the compressed visible data. For this purpose, the processing system (e.g., an accelerated processing unit (APU), a central processing unit (CPU), memory) operates in a standard mode and first divides the image into two or more tiles (e.g., bins). While in the standard mode, the processing system then performs a visible pass on the tiles of the image by determining whether each primitive of the image is visible (e.g., present) in each tile of the image. In response to a primitive not being visible within a tile, the processing system generates visible data indicating that the primitive is not visible within the tile, that the draw call associated with the primitive is not visible within the tile, or both, and that the primitive, draw call, or both should not be rendered for the tile. In response to a primitive being visible within a tile, the processing system generates visible data indicating, for example, the vertex data, shading data, positioning data of the primitive, or any combination thereof. Once the visible data is generated by the processing system, the processing system compresses the visible data and stores the compressed visible data in a buffer associated with the tile. Next, the processing system flushes the compressed visible data from the buffer and stores the flushed visible data in memory, for example, in response to the visible pass being complete (e.g., the processing system has determined whether each primitive is visible in the tiles of the image). The processing system then renders the primitives within the image using the flushed visible data. By rendering the primitives using the compressed visible data, the time required to render the primitives is reduced. However, waiting to render the primitives until after the compressed visible data has been flushed from the buffer introduces latency into the graphics pipeline used to render the primitives.

[0014] For this purpose, the techniques and systems described herein address the reduction of latency in a graphics pipeline by coarse visible compression. To reduce such latency, one or more portions of the processing system (e.g., APU, CPU) operate in a pipeline latency reduction mode. While in the pipeline latency reduction mode, the processing system maintains a visible primitive count (e.g., indicating the current number of primitives determined to be visible in the first tile) and a visible draw call count (e.g., indicating the current number of draw calls issued for the primitives determined to be visible in the first tile). Further, the processing system divides the rendered image into two or more tiles and executes the visible pass of the image by determining, for each tile of the image, whether each primitive of one or more draw calls of the image is visible (e.g., present). In response to none of the primitives indicated in the draw call being present within the tile, the processing system generates visible data indicating that the draw call, the primitives within the draw call, or both are not visible within the tile and should not be rendered for the tile. In response to the primitives of a draw call being visible within a tile (e.g., the first tile to be rendered after the visible pass), the processing system increments the visible primitive count and generates visible data including the draw call index data of the draw call (e.g., a pointer to the draw call, the number of indices within the draw call), indicating that at least a portion of the primitive, the draw call, or both are visible within the tile and should be rendered using the draw call index data. Additionally, in response to the primitives of a draw call being visible within a tile (e.g., the first tile rendered after the visible pass), the processing system increments the visible draw call count and marks the draw call as visible (e.g., generates a flag indicating a visible draw call) if a preceding primitive (e.g., a primitive for which visibility has already been determined within the tile) associated with the same draw call as the primitive (e.g., the current primitive) was not visible within the tile.That is, when the primitive is the first visible primitive within a draw call, the processing system increments the visible draw call count and marks the draw call as visible.

[0015] In addition, while in the pipeline latency reduction mode, the processing system stores the generated visible data (e.g., for the first tile rendered after the visible pass) within a data structure (e.g., an array) in memory (e.g., on-chip memory, cache). The processing system then uses the generated visible data to render primitives for one or more draw calls. For example, for each draw call marked as visible in the visible data, the processing system renders primitives based on the draw call index data. In response to the visible primitive count, the visible draw call count, or both being above a binning threshold, the processing system switches to operating in standard mode. That is, the processing system then switches to generating and compressing visible data for use in rendering the remaining primitives. In this way, the processing system does not need to wait for compressed visible data to be flushed from the buffer between when the visible pass is executed or the visible data is flushed from the buffer, or both, and reduces latency in the graphics pipeline due to coarse visible compression, enhancing the efficiency of the processing system.

[0016] FIG. 1 is a block diagram of a processing system 100 configured to reduce pipeline delay by coarse visible compression according to some embodiments. The processing system 100 includes or has access to a memory 106 or other storage component implemented using a non-transitory computer-readable storage medium such as, for example, dynamic random access memory (DRAM). However, in embodiments, the memory 106 is implemented using other types of memory including static random access memory (SRAM), non-volatile RAM, and the like. According to embodiments, the memory 106 includes external memory for being implemented external to the processing unit implemented in the processing system 100. The processing system 100 also includes a bus 112 for supporting communication between entities implemented in the processing system 100, such as the memory 106. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not shown in FIG. 1 for clarity.

[0017] The techniques described herein are employed in an Accelerated Processing Unit (APU) 114 in different embodiments. The APU 114 includes, for example, a vector processor, a coprocessor, a Graphics Processing Unit (GPU), a General-Purpose GPU (GPGPU), a non-scalar processor, a high-parallel processor, an Artificial Intelligence (AI) processor, an inference engine, a machine learning processor, other multi-threaded processing units, a scalar processor, a serial processor, or any combination thereof. The APU 114 renders images according to one or more applications 110 for presentation on a display 120. For example, the APU 114 renders objects to generate pixel values provided to the display 120, and the display 120 uses the pixel values to display an image representing the rendered objects. The APU 114 implements a plurality of processor cores 116-1 to 116-N that execute instructions simultaneously or in parallel to render objects. For example, the APU 114 uses the plurality of processor cores 116 to execute instructions from a graphics pipeline 124 to render one or more textures. According to an embodiment, one or more of the processor cores 116 operate as MD units that perform the same operations on different data sets. In the exemplary embodiment shown in FIG. 1, three cores (116-1, 116-2, 116-N) representing N cores are presented, but the number of processor cores 116 implemented within the APU 114 is a matter of design choice. Thus, in other embodiments, the APU 114 can include any number of cores 116. Some embodiments of the APU 114 are used for general-purpose computing. The APU 114 executes instructions such as program code 108 of one or more applications 110 stored in a memory 106, and the APU 114 stores information such as the results of the executed instructions within the memory 106.

[0018] The processing system 100 also includes a central processing unit (CPU) 102 connected to the bus 112 and thus communicating with the APU 114 and the memory 106 via the bus 112. The CPU 102 implements a plurality of processor cores 104-1 to 104-N that execute instructions simultaneously or in parallel. In an embodiment, one or more of the processor cores 104 operate as single instruction multiple data (SIMD) units that perform the same operation on different data sets. In the exemplary embodiment shown in FIG. 1, three cores (104-1, 104-2, 104-M) representing M cores are presented, but the number of processor cores 104 implemented within the CPU 102 is a matter of design choice. Thus, in other embodiments, the CPU 102 can include any number of cores 104. In some embodiments, the CPU 102 and the APU 114 have an equal number of cores 104, 116, but in other embodiments, the CPU 102 and the APU 114 have different numbers of cores 104, 116. The processor cores 104 execute instructions such as the program code 108 of one or more applications 110 stored in the memory 106, and the CPU 102 stores information such as the results of the executed instructions in the memory 106. Also, the CPU 102 can initiate graphics processing by issuing a draw call to the APU 114. In an embodiment, the CPU 102 implements a plurality of processor cores (not shown in FIG. 1 for clarity) that execute instructions simultaneously or in parallel.

[0019] In an embodiment, the APU 114 is configured to render one or more objects (e.g., textures) for an image to be rendered according to the graphics pipeline 124. The graphics pipeline 124 includes, for example, one or more steps, stages, or instructions executed by the APU 114 to render one or more objects for the image to be rendered. For example, the graphics pipeline 124 includes data indicating an assembler stage, a vertex shader stage, a hull shader stage, a tessellator stage, a domain shader stage, a geometry shader stage, a bin stage, a rasterizer stage, a pixel shader stage, and an output merger stage executed by the APU 114 to render one or more textures. According to an embodiment, the graphics pipeline 124 has a front end including one or more stages of the graphics pipeline 124 and a backend including one or more other stages of the graphics pipeline 124. As an example, the graphics pipeline 124 has a front end including one or more stages (e.g., an assembler stage, a vertex shader stage, a hull shader stage, a tessellator stage, a domain shader stage, a geometry shader stage, a bin stage) associated with tile-based (e.g., bin-based) rendering and a backend including one or more stages (e.g., a rasterizer stage, a pixel shader stage, an output merger stage) associated with pixel-based rendering. In an embodiment, the APU 114 is configured to execute at least a part of the front end of the graphics pipeline 124 simultaneously with at least a part of the backend of the graphics pipeline 124. For example, the APU 114 is configured to currently execute one or more stages of the front end of the graphics pipeline 124 associated with tile-based rendering with one or more stages of the backend of the graphics pipeline 124 associated with pixel-based rendering.

[0020] To render one or more objects, the APU 114 uses the original index data 168 when executing at least a portion of the graphics pipeline 124. For example, the APU 114 uses the original index data 168 when executing the front end of the graphics pipeline 124 that includes stages associated with tile-based rendering. The original index data 168 includes data representing vertices of one or more primitives of an object (e.g., a texture) to be rendered by the APU 114. In an embodiment, the APU 114 is configured to assemble, place, shade, or any combination thereof, one or more primitives according to the graphics pipeline 124 using the original index data 168. To help improve the performance of the front end of the graphics pipeline 124, the processing system 100 compresses the index data before the index data is used by the APU 114 to assemble, place, or shade one or more primitives. As an example, before the APU 114 is configured to execute at least a portion of the graphics pipeline 124, the APU 114 is configured to execute a visible pass to compress the index data of the primitives of the image. The visible pass includes, for example, first dividing the rendered image into two or more tiles (e.g., bins). Each tile includes, for example, a first number of pixels of the image rendered in a first direction (e.g., horizontal) and a second number of pixels of the image rendered in a second direction (e.g., vertical). After the image is divided into tiles, the visible pass includes determining the number of primitives to be rendered by the APU 114. For example, the APU 114 determines the number of primitives to be rendered based on a command stream indicating a batch of draw calls received by the application 110. For each primitive determined from each draw call indicated in the command stream, the APU 114 executes one or more stages of the front end of the graphics pipeline 124.As an example, the APU 114 executes an assembler stage and one or more shader stages on primitives determined from draw calls of a command stream. After one or more stages of the front end of the graphics pipeline 124 are executed on one or more primitives determined from draw calls indicated in the command stream, the APU 114 then determines whether each primitive of the draw call is present (e.g., visible) in each tile (e.g., bin) of the image and provides the visible data of the primitive to respective memory (e.g., buffer). For example, in response to determining that at least a portion of a primitive is present (e.g., visible) within a tile, the APU 114 provides visible data indicating vertex data, shading data, positioning data, or any combination thereof of the primitive to respective buffers (e.g., buffers associated with the tile). Additionally, in response to determining that a primitive is not present (e.g., visible) within a tile, the APU 114 provides visible data indicating that the primitive is not present (e.g., visible) within the tile.

[0021] According to an embodiment, the CPU 102, the APU 114, or both are configured to compress visible data before the visible data is stored in their respective buffers. For example, the CPU 102, the APU 114, or both are configured to compress data related to vertices (e.g., pointers to vertex buffers) of primitives that are visible within a tile before the data related to the vertices is stored in the buffer. In an embodiment, the CPU 102, the APU 114, or both are configured to flush visible data from the buffer in response to a threshold event. Such threshold events include, for example, the elapse of a predetermined period (e.g., nanoseconds, milliseconds, seconds, minutes), the APU 114 completing the visible pass of an image, or both. The CPU 102, the APU 114, or both flush the visible data from the buffer to the memory 106 such that the flushed visible data is available as compressed index data for the front end of the graphics pipeline 124. That is, the APU 114 is configured to use the visible data flushed from the buffer to the memory 106 as compressed index data instead of the original index data 168 when executing one or more stages of the graphics pipeline 124.

[0022] After the APU 114 completes the visible pass and visible data is flushed from one or more buffers, the APU 114 is configured to render primitives within each tile (e.g., bin) according to the graphics pipeline 124 using the compressed index data (e.g., the flushed visible data). As an example, after completing the visible pass for each tile of an image and flushing the buffer of visible data, the APU 114 uses the compressed index data to render the primitives within the first tile according to the stages of the graphics pipeline 124. When all the primitives within the first tile are rendered, the APU 114 uses the compressed index data to render the primitives within, for example, the next sequential tile (e.g., the second tile) according to the stages of the graphics pipeline 124. The APU 114 renders primitives tile (e.g., bin) by tile until the primitives within each tile are rendered. By waiting for the visible pass to complete and the visible data to be flushed from the buffer before rendering the primitives, the APU 114 helps ensure that the compressed index data from the visible pass is in memory 106 before the APU 114 begins rendering the primitives. However, waiting to render the primitives until after the visible data is flushed introduces a delay in the pipeline between the completion of the visible pass and the rendering of the primitives. To help reduce such a delay, the APU 114 is configured to operate in a pipeline delay reduction mode.

[0023] The pipeline delay reduction mode includes, for example, the APU 114 rendering one or more visible primitives within the first tile while the visible pass is being performed, while the visible data is being flushed to memory, or during both. To render one or more visible primitives within the tile, while the visible pass is being performed, while the visible data is being flushed to memory, or during both, the CPU 102, the APU 114, or both are configured to hold a visible primitive count, a visible draw call count, or both for the first tile (e.g., bin). The visible primitive count indicates, for example, the number of currently determined visible primitives within the first tile, and the visible draw call count indicates, for example, the number of draw calls including one or more currently determined visible primitives. Based on the visible primitive count and the visible draw call count, the APU 114 is configured to render a predetermined number of visible primitives, visible draw calls, or both within the first tile using, for example, draw call index data (e.g., pointers to draw calls, number of indexes within the draw call) stored in on-chip memory, and to render the remaining visible primitives within the first tile using visible data (e.g., compressed index data) flushed from the buffer. For example, the CPU 102, the APU 114, or both first receive from the application 110 a command stream indicating a batch of draw calls for the image to be rendered. Based on the draw calls, the APU 114 executes the visible pass of the image.In response to a primitive associated with a draw call not being present (e.g., not visible) within a first tile of an image, the APU 114 provides data (e.g., a flag) indicating that the draw call is not visible within the first tile and that the primitive of the draw call should not be rendered for the first tile to a data structure (e.g., an array) stored within the on-chip memory 174 (e.g., RAM, SRAM, DRAM, synchronous dynamic random access memory (SDRAM), read-only memory (ROM), programmable read-only memory (PROM), electronically erasable programmable read-only memory (EEPROM), flash memory). In response to the primitive being present within the first tile, the APU 114 stores the draw call index data associated with the primitive within the on-chip memory 174, within a buffer associated with the first tile, or both, and provides data (e.g., a flag) indicating that the primitive, the draw call associated with the primitive, or both are visible within the tile and that the draw call index data should be used for rendering. Additionally, in response to the primitive being present within the first tile, the CPU 102, the APU 114, or both increment a visible primitive count by, for example, 1. Further, the CPU 102, the APU 114, or both increment a visible draw call count by, for example, 1 if the primitive visible within the first tile is the first primitive of a draw call determined to be visible. For example, in response to the primitive being present within the first tile, the APU 114 determines whether a previous primitive of the same draw call as the primitive (e.g., a primitive for which visibility has already been determined within the tile) was visible within the first tile. As an example, the APU 114 checks a flag associated with the draw call to determine whether a previous primitive of the same draw call as the primitive was visible within the first tile. In response to determining that a previous primitive of the same draw call as the primitive was not visible within the first tile, the APU 114 increments the visible draw call count.

[0024] After the visible primitive count, the visible draw count, or both reach a threshold value, the APU 114 switches from the pipeline delay reduction mode to the standard mode and starts storing visible data (e.g., compressed index data) in each buffer as described above. That is, when the visible primitive count, the visible draw count, or both reach a predetermined value, the APU 114 switches to storing visible data in the buffer, and when the data is flushed, it uses the flushed visible data (e.g., compressed index data) to render visible primitives within a tile (e.g., a bin). In this way, the pipeline delay between the completion of the visible path and the rendering of primitives is reduced because, between when the visible path is completed, when the visible data is flushed from the buffer, or both, a predetermined number of primitives, draw calls, or both are rendered using the draw call index data stored in the on-chip memory 174. Therefore, the total time for rendering an image is shortened.

[0025] The input / output (I / O) engine 118 includes hardware and software that handle input or output operations associated with the display 120 and other elements of the processing system 100 such as a keyboard, a mouse, a printer, an external disk, etc. The I / O engine 118 is coupled to the bus 112 such that the I / O engine 118 communicates with the memory 106, the GPU 114, or the CPU 102. In the illustrated embodiment, the I / O engine 118 reads information stored in an external storage component 122 that is implemented using a non-transitory computer-readable storage medium such as a compact disk (CD), a digital versatile disk (DVD), etc. Also, the I / O engine 118 can write information such as the result of processing by the APU 114 or the CPU 102 to the external storage component 122.

[0026] Next, referring to FIG. 2, an APU 200 configured to implement a graphics pipeline 224 using coarse visual compression is shown. In an embodiment, the APU 200, similar or identical to the APU 114, is configured to render one or more textures 250 based on a command stream received from the application 110 and including data of an image to be rendered. For example, the command stream includes a batch of draw calls indicating one or more primitives to be rendered for the image. To render the image indicated in the command stream, the APU 200 is configured to render one or more primitives according to a graphics pipeline 224 similar or identical to the graphics pipeline 124. The graphics pipeline 224 includes one or more steps, stages or instructions executed by the APU 200 to render one or more objects for the image to be rendered, for example, an assembler stage 226, a vertex shader stage 228, a hull shader stage 230, a tessellator stage 232, a domain shader stage 234, a geometry shader stage 236, a binner stage 238, a rasterizer stage 240, a pixel shader stage 242, an output merger shader stage 244, or any combination thereof.

[0027] The assembler stage 226 includes data and instructions for the APU 200 to read and compile primitive data from, for example, a memory (e.g., memory 106), an application 110, a command stream, or any combination thereof, thereby generating one or more primitives to be rendered by the rest of the graphics pipeline 224. The vertex shader stage 228 includes data and instructions for the APU 200 to perform one or more operations on the primitives generated by, for example, the assembler stage 226. Such operations include, for example, transformation (e.g., coordinate transformation, modeling transformation, viewing transformation, projection transformation, viewport transformation), skinning, morphing, and lighting operations. The hull shader stage 230, the tessellator stage 232, and the domain shader stage 234 together include data and instructions for the APU 200 to perform tessellation on the primitives modified by, for example, the vertex shader stage 228. The geometry shader stage 236 includes data and instructions for the APU 200 to perform vertex operations on the tessellated primitives. Such vertex operations include, for example, point sprint expansion, dynamic particle system operations, fur-fin generation, shadow volume generation, single pass render-to-cubemap from single pass rendering, per-primitive material swapping, and per-primitive material setup. The bin stage 238 includes data and instructions for the APU 200 to perform coarse rasterization to determine, for example, whether tiles (e.g., bins) of an image overlap with one or more primitives (e.g., primitives modified by the vertex shader stage 228).That is, the bin stage 238 includes data and instructions for the APU 200 to determine which primitives (e.g., visible) are present within a tile (e.g., bin) of an image. The rasterization stage 240 includes data and instructions for determining, e.g., by the APU 200, which pixels are included in each primitive and for converting each primitive into pixels of an image. The pixel shader stage 242 includes data and instructions for the APU 200 to determine output values for the pixels determined during the rasterization stage 240, for example. The output merge stage 244 includes data and instructions for the APU 200 to merge output values of pixels using, for example, z-testing and alpha blending, for example.

[0028] According to an embodiment, each instruction of stages 226 to 244 of the graphics pipeline 224 is executed by one or more cores 248 that are similar to or the same as the core 116 of the APU 200. The exemplary embodiment shown in FIG. 2 presents an APU 200 having three cores (248-248-2, 1248-N) representing N cores, but in other embodiments, the APU 200 may have any number of cores. Each instruction of the graphics pipeline 224 is scheduled by a scheduler 246 for execution by one or more cores 248. The scheduler 246 includes, for example, hardware and software configured to schedule tasks and instructions for the cores 248 of the APU 200. In this way, two or more stages of the graphics pipeline 224 are executed simultaneously. In an embodiment, the graphics pipeline 224 includes a front end including one or more stages of the graphics pipeline 224 and a back end including one or more other stages of the graphics pipeline 224. For example, the graphics pipeline 224 includes a front end including stages related to tile-based (e.g., coarse tile-based) rendering (e.g., assembler stage 226, vertex shader stage 228, hull shader stage 230, tessellator stage 232, domain shader stage 234, geometry shader stage 236, bin stage 238) and a back end including stages related to pixel-based rendering (e.g., rasterization stage 240, pixel shader stage 242, output merger stage 244). In an embodiment, the APU 200 is configured to execute one or more stages of the front end of the graphics pipeline 224 simultaneously with one or more stages of the back end of the graphics pipeline 224.

[0029] Next, referring to FIG. 3, an APU 200 configured to reduce pipeline delay by coarse visual compression is presented. In an embodiment, the APU 200 is configured to generate one or more textures 250 according to a graphics pipeline 224. For this purpose, the APU 200 includes an assembler 354, a geometry engine 352, a shader 356, a bin 358, and an on-chip memory 374 similar to or the same as the on-chip memory 174. The assembler 354 includes, for example, hardware and software-based circuitry configured to execute one or more instructions from an assembler stage 226 of the graphics pipeline 224. That is, the assembler 354 includes hardware and software-based circuitry configured to read primitive data from a memory (e.g., memory 106), an application 110, a command stream, or any combination thereof, and compile it to generate one or more primitives to be rendered. In an embodiment, the assembler 354 includes hardware and software-based circuitry configured to read and compile data output by one or more stages of the graphics pipeline 224 so that the data is available for use by one or more other stages of the graphics pipeline 224. For example, the assembler 354 is configured to read and compile data output by a geometry shader stage 236 so that the data is available for use by a bin stage 238. The geometry engine 352 includes hardware and software-based circuitry for executing one or more instructions from one or more stages of the front end of the graphics pipeline 224, such as a vertex shader stage 228, a hull shader stage 230, a tessellator stage 232, a domain shader stage 234, and a geometry shader stage 236. As an example, the geometry engine 352 includes one or more hardware and software shaders 356 configured to execute one or more instructions from one or more stages of the front end of the graphics pipeline 224.The bin 358 includes hardware and software-based circuitry configured to execute one or more visible passes, one or more instructions from the bin stage 238, or both. For example, the bin 358 is configured to determine whether one or more primitives are visible within a tile and store visible data 360, e.g., vertex data, shading data, positioning data of the visible primitive, in respective bin buffers 364. The pixel engine 370 includes hardware and software-based circuitry configured to execute one or more instructions from one or more stages of the back end of the graphics pipeline 224, e.g., the rasterizer stage 240, the pixel shader stage 242, and the output merger stage 244.

[0030] According to an embodiment, the APU 200 is configured to simultaneously execute one or more instructions associated with the front end of the graphics pipeline 224 and one or more instructions associated with the back end of the graphics pipeline 224. For example, the assembler 354, the geometry engine 352, the bin 358, or any combination thereof is configured to execute one or more tile-based rendering instructions associated with the front end of the graphics pipeline 224 (e.g., the assembler stage 226, the vertex shader stage 228, the hull shader stage 230, the tessellator stage 232, the domain shader stage 234, the geometry shader stage 236, the bin stage 238) for primitives within a first tile (e.g., a bin), and the pixel engine 370 is configured to execute one or more pixel-based rendering instructions associated with the back end of the graphics pipeline 224 (e.g., the rasterizer stage 240, the pixel shader stage 242, the output merger stage 244) for pixels of the first tile, a different second tile, or both.

[0031] In an embodiment, the geometry engine 352 is configured to execute instructions from the front end of the graphics pipeline 224 using original index data 368 that includes data representing vertices of one or more primitives of a texture 250 rendered by, for example, the APU 114 (e.g., a pointer to a vertex buffer). To help reduce the amount of time required for the geometry engine 352 to execute instructions from the front end of the graphics pipeline 224, the APU 200 is configured to generate compressed index data 372 that includes compressed data representing vertices of one or more primitives of a texture 250 rendered by the APU 200. For this purpose, the APU 200 is configured to receive a command stream from the application 110 that indicates an image to be rendered. For example, the command stream indicates a batch of draw calls that identify one or more primitives to be rendered for the image. In response to receiving the command stream, the assembler 354, the geometry engine 352, or both are configured to execute instructions for one or more stages of the front end of the graphics pipeline 224 to generate one or more primitives. For example, the assembler 354 is configured to execute instructions from the assembler stage 226, and the geometry engine 352 is configured to execute instructions from the vertex shader stage 228, hull shader stage 230, tessellator stage 232, domain shader stage 234, geometry shader stage 236, or any combination thereof to generate one or more primitives. Next, the bin 358 is configured to divide the image into two or more tiles (e.g., bins) and execute the visible pass of the image. That is, the bin 358 determines which of the primitives generated by the assembler 354 and the geometry engine 352 are visible (e.g., present) in each tile.

[0032] Based on the visible path of the image, the bin 358 is configured to generate visible data 360 associated with the tile and store the visible data 360 in respective bin buffers 364. For example, during the visible path, in response to determining that a primitive is not visible (e.g., does not exist) within a first tile, the bin 358 provides visible data 360 (e.g., a flag) indicating that the draw call of the primitive, the primitive, or both are not visible in the first tile to respective bin buffers 364 (e.g., the bin buffer 364 associated with the first tile). Additionally, in response to determining that a primitive is visible (e.g., exists) in the first tile, the bin 358 is configured to provide visible data 360 indicating vertex data, shading data, positioning data, or any combination thereof of the primitive to respective bin buffers 364. According to an embodiment, the bin 358 is configured to compress the visible data 360 before being provided to and stored in the bin buffer 364. In an embodiment, the APU 200, the CPU 102, or both are configured to flush the compressed visible data 360 from the bin buffer 364 to the memory 106 in response to a threshold event. Such threshold events include, for example, the elapse of a predetermined period (e.g., nanoseconds, milliseconds, seconds, minutes), the APU 200 completing the visible path, or both. For example, in response to completing the visible path, the APU 200 is configured to flush the compressed visible data 360 from the bin buffer 364 associated with the first tile to the memory 106.

[0033] In an embodiment, the compressed visible data 360 flashed from the bin buffer 364 to the memory 106 can be used as compressed index data 372. That is, the assembler 354, the geometry engine 352, or both are configured to use the compressed index data 372 to render one or more primitives of the image shown in a batch of draw calls. The compressed index data 372 includes, for example, data representing vertices of one or more primitives of an image rendered by the APU 200 (e.g., pointers to vertex buffers). In an embodiment, the APU 200 is configured to render an image according to the ordering of one or more tiles and the respective visible data 360 associated with the tiles. For example, the APU 200 is configured to render each primitive visible in a first tile of the image based on the visible data 360 (e.g., based on the compressed index data 372 after the visible data 360 is flashed from the bin buffer 364). In response to rendering each primitive visible (e.g., present) in the first tile, the APU 200 is configured to render primitives visible in the next sequential tile (e.g., the first adjacent tile). According to an embodiment, the APU 200 is configured to perform tile-based rendering (e.g., the front end of the graphics pipeline 224) for primitives in a first tile while performing pixel-based rendering (e.g., the backend of the graphics pipeline 224) for primitives in a second different tile. For example, the APU 200 simultaneously performs tile-based rendering for primitives in the first tile and pixel-based rendering for primitives in a second tile where the tile-based rendering has already been completed. By simultaneously performing tile-based rendering and pixel-based rendering for primitives in different tiles, the time required to render an image is reduced.

[0034] However, waiting to execute the front end of the graphics pipeline 224 (e.g., tile-based rendering) until the visible pass is complete and the visible data is flushed from the buffer introduces latency in the graphics pipeline between the completion of the visible pass and the rendering of primitives. During such latency, the pixel engine 370 remains idle until at least a portion of the front end of the graphics pipeline 224 is complete, reducing the efficiency of the system. To help reduce such latency, the APU 200 is configured to operate in pipeline latency reduction mode.

[0035] During the pipeline delay reduction mode, the assembler 354, the geometry engine 352, the bin 358, and the pixel engine 370 are configured to render one or more visible primitives associated with one or more visible draw calls within a first tile while the visible pass is occurring, while the visible data 360 is being flushed from the bin buffer 364, or during both. For example, the assembler 354, the geometry engine 352, and the bin 358 are configured to execute one or more instructions (e.g., tile-based rendering) from the front end of the graphics pipeline 224 while the visible data 360 is being flushed from the bin buffer 364, and the pixel engine 370 is configured to execute one or more instructions (e.g., pixel-based rendering) from the back end of the graphics pipeline 224 on one or more primitives rendered during an instruction from the front end of the graphics pipeline 224. To render one or more visible primitives within a tile while the visible pass is occurring, while the visible data 360 is being flushed from the bin buffer 364, or during both, the APU 200 is configured to hold, for a first tile (e.g., bin), a visible primitive count (e.g., the number of currently determined visible primitives within the first tile), a visible draw call count (e.g., the number of draw calls including the currently determined visible primitives within the first tile), or both. The APU 200 is further configured to compare the visible primitive count, the visible draw call count, or both, with one or more binning thresholds 362 that include data representing a predetermined number (e.g., maximum number) of primitives for rendering in the pipeline delay reduction mode, a predetermined number (e.g., maximum number) of draw calls for rendering in the pipeline delay reduction mode, or both. In an embodiment, the APU 200 is configured to render one or more visible primitives within a tile based on a comparison of the visible primitive count, the visible draw call count, or both, with one or more binning thresholds 362.

[0036] As an example, in response to the APU 200 receiving a command stream from the application 110 that indicates a batch of draw calls for the image to be rendered, the assembler 354, the geometry engine 352, the bin 358, or any combination thereof executes a visible pass for the image based on one or more primitives indicated in the batch of draw calls. During the visible pass, the APU 200 compares a visible primitive count, a visible draw call count, or both with one or more binning thresholds 362 and renders visible primitives based on the comparison. For example, in response to the visible primitive count, the visible draw call count, or both being less than one or more binning thresholds 362, the APU 200 is configured to render one or more visible primitives of one or more visible draw calls in a pipeline latency reduction mode. In response to the visible primitive count, the visible draw call count, or both being greater than or equal to one or more binning thresholds 362, the APU 200 is configured to operate in a standard mode, stores visible data 360 in respective bin buffers 364, and when the visible data 360 is flushed from the bin buffer 364, renders visible primitives using compressed index data 372.

[0037] While in the pipeline delay reduction mode and in response to there being no draw call primitive (e.g., visible) within the first tile, the APU 200 generates visibility data 360 (e.g., a flag) indicating that the draw call is not visible within the first tile and that the draw call primitive, draw call, or both should not be rendered for the first tile. The APU 200 provides such visibility data to a data structure (e.g., an array) stored in the on-chip memory 374. In response to the draw call primitive being visible within the first tile, the APU 200 stores draw call index data associated with the primitive (e.g., a pointer to the draw call associated with the primitive, the number of indices within the draw call) in the on-chip memory 374 and generates visibility data 360 (e.g., a flag) indicating that the primitive, draw call associated with the primitive, or both are visible within the tile and should be rendered using the draw call index data. For example, the assembler 354, geometry engine 352, bin 358, or any combination thereof executes one or more instructions (e.g., tile-based rendering) from the front end of the graphics pipeline 224 for the visible primitives of the visible draw call using the draw call index data stored in the on-chip memory 374, and the pixel engine 370 executes one or more steps (e.g., pixel-based rendering) of the back end of the graphics pipeline 224 for the visible primitives rendered by the assembler 354, geometry engine 352, bin 358, or any combination thereof. Additionally, in response to the primitive being present within the first tile, the APU 200 increments the visible primitive count, e.g., by 1. Further, the APU 200 increments the visible draw call count, e.g., by 1, if the primitive is the first determined visible primitive of the draw call.For example, in response to a primitive being present in a first tile, the APU 200 determines whether a preceding primitive of the same draw call as the primitive (e.g., a primitive for which visibility has already been determined in the tile) was visible in the first tile. As an example, the APU 200 checks a flag associated with the draw call to determine whether a preceding primitive of the same draw call as the primitive was visible within the first tile. In response to determining that a preceding primitive of the same draw call was not visible within the first tile, the APU 200 increments a visible draw call count. In this way, the pipeline delay between the completion of the visible pass and the rendering of the primitive is reduced because a predetermined number of primitives are rendered using draw call index data stored in the on-chip memory 374 while the visible data is being flushed from the bin buffer 364. Accordingly, the time that the pixel engine 370 remains in an idle state waiting for one or more instructions from the front end of the graphics pipeline 224 to complete is also reduced, improving the efficiency of the system.

[0038] Next, referring to FIG. 4, an exemplary operation 400 for reducing pipeline delay due to a visible path in coarse visual compression is presented. In an embodiment, operation 400 includes the APU 200 receiving a command stream 405. The command stream 405 includes data generated by the application 110 that indicates, for example, a batch of draw calls for one or more primitives to be rendered for a texture, an image, or both. In response to receiving the command stream 405, the APU 200 (e.g., assembler 354) is configured to read and organize the primitive data indicated within the command stream 405 into one or more primitives to be rendered by one or more stages of the graphics pipeline 224. After reading and organizing the primitive data indicated in the command stream 405, the geometry engine 352 begins rendering one or more primitives to be rendered indicated in the command stream 405. For example, the geometry engine 352 executes one or more instructions from one or more stages (e.g., vertex shader stage 228, hull shader stage 230, tessellator stage 232, domain shader stage 234, geometry shader stage 236) associated with the front end of the graphics pipeline 224. To execute one or more instructions from one or more stages associated with the front end of the graphics pipeline 224, the geometry engine 352 is configured to use the shader 356. Operation 400 further includes providing data generated from the geometry engine 352, the shader 356, or both, which execute one or more instructions from one or more stages associated with the front end of the graphics pipeline 224, to the assembler 354, the bina 358, or both. For example, operation 400 includes the geometry engine 352, the shader 356, or both, that provide data generated from executing one or more instructions of the geometry shader stage 236 to the assembler 354.In response to assembler 354 receiving data generated from geometry engine 352, shader 356, or both, which execute one or more instructions from one or more stages associated with the front end of graphics pipeline 224, assembler 354 arranges the data so that the data is available for use by bina 358. For example, assembler 354 arranges the data into one or more primitives. As another example, operation 400 includes geometry engine 352, shader 356, or both providing data generated from executing one or more instructions of the front end of graphics pipeline 224 to bina 358. Bina 358 uses such data, for example, to perform a visible pass for two or more tiles of an image.

[0039] In response to receiving one or more primitives from assembler 354, bin 358 is configured to divide the image to be rendered into two or more tiles and execute a visible pass on the tiles of the image. In an embodiment, bin 358 is configured to execute a visible pass based on whether APU 200 is operating in a standard mode or a pipeline latency reduction mode. To determine the operating mode of APU 200, APU 200 is configured to compare a visible primitive count (e.g., indicating the current number of primitives determined to be visible within a tile), a visible draw call count (e.g., indicating the current number of draw calls issued for primitives determined to be visible within a tile), or both, with one or more binning thresholds 362. For example, APU 200 is configured to compare the visible primitive count with a predetermined visible primitive count threshold (e.g., indicating the maximum number of visible primitives), and compare the visible draw call count with a visible draw call count threshold (e.g., indicating the maximum number of draw calls having visible primitives). In response to the visible primitive count, the visible draw call count, or both being less than one or more binning thresholds 362, APU 200 is configured to operate in a pipeline latency reduction mode. For example, in response to the visible primitive count being less than the visible primitive count threshold and the visible draw call count being less than the visible draw call count threshold, the APU is configured to operate in a pipeline latency reduction mode. In response to the visible primitive count, the visible draw count, or both being greater than or equal to one or more binning thresholds 362, APU 200 is configured to operate in a standard mode. For example, in response to the visible primitive count being greater than or equal to the visible primitive count threshold, or in response to the visible draw call count being greater than or equal to the visible draw call count threshold, APU 200 is configured to operate in a standard mode.

[0040] While the APU 200 is operating in the pipeline latency reduction mode, operation 400 includes a bin 358 that generates visible data 410 similar to or the same as the visible data 360 for a first tile (e.g., a bin) of an image, based on the draw call primitives provided by the assembler 354. For example, for the first tile, the bin 358 determines whether each draw call primitive provided by the assembler 354 is visible (e.g., present) in the first tile. In response to the draw call primitive not being visible (e.g., not present) in the first tile, the bin 358 generates visible data 410 that includes data (e.g., a flag) indicating that the draw call is not visible in the first tile. Such data is stored, for example, in an array of on-chip memory 374. In response to the primitive being visible (e.g., present) in the first tile, the bin 358 stores the draw call index data associated with the primitive in the on-chip memory 374 and generates visible data 410 that includes data (e.g., a flag) indicating that the draw call, the primitive, or both are visible in the tile and are to be rendered using the draw call index data. Such data is stored, for example, in an array of on-chip memory 374. In addition, in response to the primitive being visible (e.g., present) in the first tile, the bin 358, for example, increments the visible primitive count by 1. Further, in response to the primitive being visible (e.g., present) in the first tile, the bin 358, for example, increments the visible draw call count by 1 if the primitive is the first visible primitive determined for the draw call. For example, in response to the primitive being present within the first tile, the bin 358 determines whether a preceding primitive of the same draw call as the primitive (e.g., a primitive for which visibility has already been determined within the tile) was visible within the first tile. In response to determining that a preceding primitive of the same draw call was not visible in the first tile, the bin 348 increments the visible draw call count.In an embodiment, the visible data 410 within the array of the memory 106 is provided to the APU 200, the geometry engine 352, or both, and renders one or more primitives of one or more draw calls determined to be visible within the first tile. For this purpose, for example, the APU 200, the geometry engine 352, or both are configured to render one or more primitives identified within a batch of draw calls indicated in the command stream 405 according to the visible data 410 within the on-chip memory 374. As an example, in response to the visible data 410 indicating that a draw call indicated in the command stream 405 is not visible in the first tile, the APU 200, the geometry engine 352, or both are configured to skip rendering the primitives indicated in the draw call in the first tile. In response to the visible data 410 indicating that a draw call indicated in the command stream 405 is visible in the first tile, the APU 200, the geometry engine 352, the CPU 102, or any combination thereof uses the draw call index data stored in the on-chip memory 374 to render the primitives of the draw call.

[0041] While operating in standard mode, operation 400 includes generating visible data 410 similar to or the same as visible data 360 for each tile of the image, based on each remaining primitive provided by assembler 354 (e.g., primitives not inspected during the visible pass while the APU was operating in pipeline latency reduction mode). For example, for the first tile, bin 358 determines whether each remaining primitive provided by assembler 354 is visible (e.g., present) in the first tile. In response to the remaining primitives that are not visible (e.g., not present) in the first tile, bin 358 generates visible data 410 that includes data (e.g., a flag) indicating that the primitive is not visible in the first tile. Such data is stored, for example, in each bin buffer 364 (e.g., the bin buffer associated with the first tile). In response to the remaining primitives being visible (e.g., present) in the first tile, bin 358 generates visible data 410 that includes data (e.g., a flag) indicating that the primitive is visible in the tile, and data indicating vertex data, shading data, positioning data, or any combination thereof, of the primitive. Such data is stored, for example, in each bin buffer 364. According to an embodiment, APU 200 is configured to compress visible data 410 before storing it in bin buffer 364. In an embodiment, operation 400 includes APU 200, CPU 102, or both flushing visible data 410 from each bin buffer 364 to memory 106. For example, in response to a threshold event (e.g., a predetermined period has elapsed, bin 358 has completed the visible pass for the tile, or both), APU 200 is configured to flush visible data 410 in the buffer to memory 106.After the compressed visible data 410 is flushed from the bin buffer 364 to the memory 106, the APU 200, the geometry engine 352, or both are configured to render one or more primitives identified in a batch of draw calls indicated in the command stream 405 based on the flushed visible data 410. For example, in response to the flushed visible data 410 indicating that a primitive shown in the command stream 405 is not visible in the first tile, the APU 200, the geometry engine 352, or both skip rendering of that primitive. In response to visible data 410 indicating that a primitive shown in the command stream 405 is visible in the first tile, the APU 200, the geometry engine 352, the CPU 102, or any combination thereof render the primitive using the flushed visible data 410 as compressed index data 415 that includes compressed data indicating vertex data, shading data, positioning data, or any combination thereof of the primitive. In this way, the APU 200 uses the compressed index data 415 to render the primitives of the command stream 405 and improve the rendering time. Additionally, the APU 200 reduces pipeline latency caused by waiting for the compressed index data 415 to be flushed from the bin buffer 364 by first rendering a predetermined number of primitives using draw call index data stored in the on-chip memory 374 while the APU 200 is operating in the pipeline latency reduction mode.

[0042] Next, referring to FIG. 5, an exemplary timing diagram 500 is presented that illustrates an exemplary reduction in pipeline delay in coarse visual compression. For example, the timing diagram 500 includes a first axis 505 indicating time and a second axis 540 indicating the front end 502 of the graphics pipeline 224 and the back end 504 of the graphics pipeline 224. The front end 502 includes one or more stages (e.g., assembler stage 226, vertex shader stage 228, hull shader stage 230, tessellator stage 232, domain shader stage 234, geometry shader stage 236, bin stage 238) associated with tile-based (e.g., bin-based) rendering, and the back end 504 includes one or more stages (e.g., rasterizer stage 240, pixel shader stage 242, output merger stage 244) associated with pixel-based rendering. The APU 200 is configured to perform a visible pass and tile-based rendering on one or more primitives visible within a first bin (e.g., tile) bin 0 at a first time 510, perform tile-based rendering on one or more primitives visible within a second bin (e.g., tile) bin 1 at a second time 515, and perform tile-based rendering on one or more primitives visible within a third bin (e.g., tile) bin 2 at a third time 520. Further, the APU 200 is configured to perform pixel-based rendering on one or more primitives that are visible in the first bin, bin 0, and processed by the front end 502 at a fourth time 525, perform pixel-based rendering on one or more primitives that are visible in the second bin, bin 1, and processed by the front end 502 at a fifth time 530, and perform pixel-based rendering on one or more primitives that are visible in the third bin, bin 2, and processed by the front end 502 at a sixth time 535.

[0043] In an embodiment, during at least a portion of the first time 510, the APU 200 renders a predetermined number of visible primitives of one or more visible draw calls within bin 0 using draw call index data stored in the on-chip memory 374, while the APU 200 is configured to operate in a pipeline delay reduction mode such that it simultaneously executes a visible pass, flushes visible data from the bin buffer, or both. For example, the APU 200 uses the draw call index data stored in the on-chip memory 374 to render visible primitives within bin 0 while the visible primitive count, the visible draw call count, or both are less than one or more bin thresholds 362 (e.g., a visible primitive count threshold, a visible draw call count threshold). After the visible primitive count, the visible draw call count, or both are equal to or exceed one or more binning thresholds 362, the APU 200 switches to the standard mode for the remainder of the first time 510 and the fourth time 525. In this way, the APU helps reduce the delay in the pipeline resulting from waiting for the visible data 410 to be flushed from one or more bin buffers 364 while simultaneously performing tile-based rendering of one or more primitives visible within bin 0 and pixel-based rendering of one or more primitives visible within bin 0 (as indicated by the overlap between the first time 510 and the fourth time 525), executing a visible pass, flushing visible data from the bin buffer, or both.

[0044] Next, referring to FIG. 6, an exemplary method 600 for reducing pipeline latency in coarse visible compression is presented. At step 605, an APU similar or identical to APU114, 200 receives a command stream similar or identical to command stream 405 that indicates a batch of draw calls identifying one or more primitives to be rendered for one or more textures, images, or both. For example, the APU receives a command stream from application 110 that indicates one or more primitives to be rendered for one or more textures, images, or both. At step 610, the APU executes one or more operations to at least partially render the primitives indicated in the command stream. For example, the APU executes one or more instructions from one or more stages of the front end of a graphics pipeline similar or the same as graphics pipeline 224 (e.g., assembler stage, vertex shader stage, hull shader stage, tessellator stage, domain shader stage, geometry shader stage) to at least partially render the primitives indicated in the command stream. At step 615, the APU executes the visible pass of the image indicated in the command stream. To execute the visible pass, the APU first divides the image into two or more tiles, where each tile includes a number of pixels in a first direction (e.g., horizontal) and a second number of pixels in a second direction (e.g., vertical). The APU then executes the visible pass on the tiles of the image (e.g., bins) to determine which of the primitives indicated in the command stream are visible (e.g., present) within the tile.

[0045] In steps 620 and 625, the APU determines whether to operate in pipeline delay reduction mode or standard mode while executing the visible pass. In some embodiments, the processing system 100 executes steps 620 and 625 simultaneously, while in other embodiments, the processing system 100 executes steps 620 and 625 sequentially (e.g., step 620, then step 625, step 625, then step 620). Referring to step 620, the APU determines whether the visible primitive count (e.g., a count representing the number of primitives currently determined to be visible in the first tile) is less than one or more binning thresholds similar to or the same as binning threshold 362. For example, the APU determines whether the visible primitive count is less than a visible primitive count threshold representing a predetermined number (e.g., the maximum number) of visible primitives. Referring to step 625, the APU determines whether the visible draw call count (e.g., a count representing the number of draw calls currently determined to include visible primitives) is less than one or more binning thresholds similar to or the same as binning threshold 362. For example, the APU determines whether the visible draw call count is less than a visible draw call count threshold representing a predetermined number (e.g., the maximum number) of draw calls including visible primitives. In some embodiments, the system moves to step 630 in response to both the visible primitive count being less than one or more binning thresholds (e.g., the visible primitive count threshold) and the visible draw call count being less than one or more binning thresholds (e.g., the visible draw call count threshold), while in other embodiments, the system moves to step 630 in response to either the visible primitive count being less than one or more binning thresholds (e.g., the visible primitive count threshold) or the visible draw call count being less than one or more binning thresholds (e.g., the visible draw call count threshold).Similarly, in some embodiments, the processing system 100 moves to step 645 in response to either the visible primitive count being greater than or equal to one or more binning thresholds (e.g., visible primitive count threshold) and the visible draw call count being greater than or equal to one or more binning thresholds (e.g., visible draw call count threshold). However, in other embodiments, the system moves to step 645 in response to both the visible primitive count being greater than or equal to one or more binning thresholds (e.g., visible primitive count threshold) and the visible draw call count being greater than or equal to one or more binning thresholds (e.g., visible draw call count threshold).

[0046] In step 630, while the APU is operating in the pipeline latency reduction mode, it executes a visible pass on the tiles of the image. While operating in the pipeline latency reduction mode, the APU generates visible data similar to or the same as visible data 360, 410 for a first tile (e.g., bin) of the image based on whether each primitive identified in a draw call indicated in a command stream (such as that rendered in step 610) is visible (e.g., present) in the first tile. Continuing to refer to step 630, in response to a draw call primitive not being visible (e.g., not present) in the first tile, the APU generates visible data including data (e.g., a flag) indicating that the draw call is not visible in the first tile (e.g., the draw call primitive should not be rendered in the first tile), and stores the visible data in an on-chip memory similar to or the same as on-chip memories 174, 374. In response to a draw call primitive being visible (e.g., present) in the first tile, the APU stores draw call index data associated with the primitive in the on-chip memory, and generates visible data including data (e.g., a flag) indicating that the draw call, primitive, or both should be rendered using the draw call index data stored in the on-chip memory as they are visible in the tile. In response to determining that the first primitive of a draw call is visible in the first tile, the processing system 100 moves to step 635. In step 635, the APU, CPU 102, or both increment a visible primitive count, for example, by 1. Further, in step 635, the APU, CPU 102, or both increment a visible draw call count, for example, by 1 if the first primitive is the first visible primitive determined for the draw call.For example, in response to the first primitive being visible in the first tile, the APU determines whether a primitive preceding the first primitive of the same draw call (e.g., a primitive for which visibility has already been determined in the tile) was visible in the first tile. As an example, the APU checks a flag associated with the draw call to determine whether a primitive preceding the first primitive of the same draw call was also visible within the first tile. In response to determining that a primitive preceding the same draw call was not visible within the first tile, the APU increments the visible draw call count. After incrementing the visible primitive count, the visible draw call count, or both, the system returns to steps 620 and 625, and the APU determines whether to continue operating in pipeline delay reduction mode or switch to standard mode while executing the visible pass.

[0047] In step 645, while operating in standard mode, the APU executes the visible path of the image. While operating in standard mode, for each tile of the image (e.g., bin), the APU generates visible data similar to or the same as visible data 360, 410 based on whether each primitive indicated in the command stream (e.g., as rendered in step 610) is visible (e.g., present) within the tile. Continuing to refer to step 645, in response to a primitive not being visible (e.g., not present) within the tile, the APU generates visible data that includes data (e.g., a flag) indicating that the primitive is not visible within the tile. In response to a primitive being visible (e.g., present) within the tile, the APU generates visible data that includes data (e.g., a flag) indicating that the primitive is visible within the tile and data indicating vertex data, shading data, positioning data of the primitive, or any combination thereof. While the APU is operating in standard mode, the APU stores the generated visible data in respective bin buffers similar to or the same as bin buffer 364. For example, the APU stores the generated visible data in the bin buffer associated with each tile of the image. According to an embodiment, the APU first compresses the generated visible data before storing it in respective bin buffers. In an embodiment, the APU flushes the visible data from respective primitive buffers to memory 106 in response to a threshold event (e.g., elapse of a predetermined period, completion of the visible path of the APU, or both). The visible data flushed from the bin buffer to memory 106 is then available as compressed index data similar to or the same as compressed index data 372, 415 for rendering one or more primitives determined to be visible at a first time.

[0048] In step 640, after completing the visible path of the image, the APU renders one or more primitives identified in one or more determined visible draw calls using the draw call index data stored in the on-chip memory, the visible data flushed from the bin buffer, or both. For example, the APU renders one or more primitives using the draw call index data stored in the on-chip memory while the visible data is being flushed from the bin buffer. In response to the visible data being flushed from the bin buffer, the APU renders primitives using the visible data flushed from the bin buffer.

[0049] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips) that perform operations such as those described above with reference to FIGS. 1-6 to assist in resolving pipeline delays. Electronic design automation (EDA) and computer aided design (CAD) software tools can be used to design and manufacture these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs operate a computer system to operate on code representing the circuits of the one or more IC devices to design or adapt a manufacturing system for manufacturing the circuits, or to perform at least a portion of the process for manufacturing the circuits. This code can include instructions, data, or a combination of instructions and data. Software instructions representing design tools or manufacturing tools are typically stored on a computer-readable storage medium accessible to a computing system. Similarly, code representing one or more stages of the design or manufacture of an IC device is stored on and accessed from the same or a different computer-readable storage medium.

[0050] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that is accessible by a computer system during use to provide instructions and / or data to the computer system. Such storage media include, but are not limited to, optical media (e.g., compact disc (CD), digital versatile disc (DVD), Blu-ray (registered trademark) disc), magnetic media (e.g., floppy (registered trademark) disc, magnetic tape, magnetic hard drive), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. A computer-readable storage medium (e.g., system RAM or ROM) may be built into the computing system, a computer-readable storage medium (e.g., magnetic hard drive) may be fixedly attached to the computing system, a computer-readable storage medium (e.g., optical disc or universal serial bus (USB)-based flash memory) may be removably attached to the computing system, or a computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to the computer system via a wired or wireless network.

[0051] In some embodiments, certain aspects of the above-described techniques are implemented by one or more processors of a processing system that executes software. The software includes one or more sets of executable instructions stored in a non-transitory computer-readable storage medium or otherwise tangibly embodied. The software may include instructions and certain data, and when the instructions and certain data are executed by one or more processors, the one or more processors are operated to execute one or more aspects of the above-described techniques. The non-transitory computer-readable storage medium can include, for example, magnetic or optical disk storage devices, solid-state storage devices such as flash memory, cache, random access memory (RAM), or other non-volatile memory device(s), etc. The executable instructions stored in the non-transitory computer-readable storage medium can be implemented in source code, assembly language code, object code, or other instruction formats interpretable or otherwise executable by one or more processors.

[0052] In addition to what has been described above, it should be noted that not all activities or elements described in the general description are required, some activities or parts of a particular device may not be required, one or more additional activities may be performed, and one or more additional elements may be included. Further, the order in which activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense, and all such modifications are intended to be included within the scope of the invention.

[0053] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and features that may give rise to or manifest any benefit, advantage, or solution are not to be construed as important, essential, or indispensable features of any or all of the claims. Further, the disclosed invention may be modified and practiced in different but similar ways that will be apparent to those skilled in the art having the benefit of the teachings herein, so the specific embodiments described above are merely illustrative. There is no limitation as to the details of construction or design shown herein other than as described in the appended claims. Accordingly, it is evident that the specific embodiments described above may be varied or modified and that all such variations are considered to be within the scope of the disclosed invention. Therefore, the protection sought herein is set forth in the appended claims.

Claims

1. A method comprising: executing a visible pass of an image based on a command stream indicating a plurality of primitives, and determining visible primitives within a first tile of the image from the plurality of primitives; rendering the visible primitives based on a comparison of a visible primitive count with a binning threshold. The method.

2. Including incrementing the visible primitive count in response to determining the visible primitives within the first tile. The method of Claim 1.

3. Including generating visible data indicating that the visible primitives are to be rendered using draw call index data stored in on-chip memory in response to the visible primitive count being less than the binning threshold. The method of Claim 1 or 2.

4. The visible primitives are rendered using the draw call index data simultaneously with the visible pass being executed on the image. The method of Claim 3.

5. Including generating visible data indicating vertex data of the visible primitives in response to the visible primitive count being greater than or equal to the binning threshold. The method of Claim 1.

6. Compressing the visible data; Storing the compressed visible data in a buffer associated with the first tile. The method of Claim 5.

7. Flushing the compressed visible data from the buffer, wherein the visible primitives are rendered using the flushed visible data. The method of Claim 6.

8. A method comprising: generating visible data of a primitive based on a visible draw call count in response to determining that the primitive indicated by a command stream is visible within a first tile of an image; rendering the primitive based on the visible data. The method.

9. The step of generating the visible data includes: generating visible data indicating that the primitive is to be rendered using draw call index data stored in on-chip memory in response to the visible draw call count being less than a first binning threshold. The method of claim 8.

10. wherein the primitive is rendered using the draw call index data simultaneously with the visible pass of the image The method of claim 9.

11. including incrementing the visible draw call count in response to the primitive being the first visible primitive of a draw call The method of claim 9 or 10.

12. including generating visible data indicating vertex data of the primitive in response to the visible draw call being greater than or equal to a first binning threshold The method of claim 8.

13. compressing the visible data; and storing the compressed visible data in a buffer associated with the first tile. The method of claim 12.

14. flushing the compressed visible data from the buffer to memory, wherein the primitive is rendered using the flushed visible data. The method of claim 13.

15. A processor comprising: one or more processing units including circuitry, wherein the circuitry is configured to: execute a visible pass of an image based on a command stream indicating a plurality of primitives, and determine visible primitives within a first tile from the plurality of primitives; and render the visible primitives based on a comparison of a visible primitive count with a binning threshold. Processor.

16. wherein the visible primitives are rendered based on a comparison of a visible draw call count with a second binning threshold. The processor of claim 15.

17. wherein the one or more processing units include circuitry, wherein the circuitry is configured to: generate visible data indicating that the visible primitives are rendered using draw call index data stored in on-chip memory in response to the visible primitive count being less than the binning threshold. The processor of claim 15 or 16.

18. wherein the visible primitives are rendered using the draw call index data simultaneously with the visible pass of the image. The processor of claim 17.

19. wherein the one or more processing units include circuitry, wherein the circuitry is configured to: configured to generate visible data indicating vertex data of the visible primitive in response to the visible primitive count being greater than or equal to the binning threshold The processor of claim 15 **Claim 20** the one or more processing units comprise circuitry the circuitry is compress the visible data store the compressed visible data in a buffer associated with the first tile flush the compressed visible data from the buffer, wherein the visible primitive is rendered using the flushed visible data configured to perform The processor of claim 19