Variable-ratio curved surface subdivision

By using variable-ratio surface tessellation technology, the surface tessellation factor is adjusted to adapt to the number of visible pixels and shading rate in different regions, which solves the problem of low graphics processing efficiency caused by excessive tessellation and achieves more efficient graphics rendering and lower power consumption.

CN118974779BActive Publication Date: 2026-02-17QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202380031191.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-08
Filing Date
2023-03-23
Publication Date
2026-02-17
Estimated Expiration
2043-03-23

AI Technical Summary

Technical Problem

Existing technologies suffer from excessive surface subdivision in areas or primitives with reduced detail, leading to inefficient graphics processing.

Method used

The variable ratio tessellation (VRT) technique is used to adjust the tessellation factor by identifying the number of visible pixels and the shading rate, thereby reducing unnecessary tessellation and improving rendering efficiency.

Benefits of technology

It improves the performance and efficiency of graphics processing, reduces GPU power consumption and memory bandwidth requirements, and enhances image rendering quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118974779B_ABST
    Figure CN118974779B_ABST
Patent Text Reader

Abstract

The present disclosure provides systems, devices, apparatuses, and methods, including computer programs encoded on storage media, for variable rate tessellation. A graphics processor can receive data for geometry processing of a plurality of tiles in a draw call. The graphics processor can reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a TRF. The characteristic can correspond to a shading rate or a number of visible pixels. The graphics processor can apply the TRF for each tile of the plurality of tiles. The graphics processor can render each tile of the plurality of tiles based on the applied TRF for each tile of the plurality of tiles.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims U.S. Patent Application No. 17 / 658,634, filed April 8, 2022, entitled “VARIABLE RATE TESSELLATION”, the entire contents of which are expressly incorporated herein by reference. Technical Field

[0003] In summary, this disclosure relates to processing systems, and more specifically, to one or more techniques for graphics processing. Background Technology

[0004] Computing devices typically perform graphics and / or display processing (e.g., utilizing a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices can include, for example, computer workstations, mobile phones such as smartphones, embedded systems, personal computers, tablet computers, and video game consoles. A GPU is configured to execute a graphics processing pipeline, which includes one or more processing stages that operate together to execute graphics processing commands and output frames. The CPU can control the operation of the GPU by issuing one or more graphics processing commands to it. Modern CPUs are typically capable of executing multiple applications concurrently, each of which may require the GPU to be utilized during execution. A display processor can be configured to convert digital information received from the CPU into analog values ​​and can issue commands to a display panel for displaying visual content. Devices that provide content for visual presentation on a display can utilize a CPU, GPU, and / or display processor.

[0005] Current techniques may not be able to solve the problem of excessive tessellation in regions or primitives where the level of detail (LOD) may be reduced. There is a need for improved tessellation techniques in which LOD can be taken into account. Summary of the Invention

[0006] The following provides a brief overview of one or more aspects to offer a basic understanding of such aspects. This overview is not a comprehensive summary of all anticipated aspects, nor is it intended to identify key or important elements of all aspects, nor to depict the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed descriptions that follow.

[0007] In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus can receive data for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more sub-primitives. The apparatus can reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a tessellation reduction factor (TRF). The characteristic can correspond to a shading rate or a number of visible pixels. The apparatus can apply the TRF for each tile of the plurality of tiles. The apparatus can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles.

[0008] To the accomplishment of the foregoing and related ends, one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative features of the one or more aspects. These features are indicative, however, of but a few of the various ways in which the principles of various aspects can be employed, and this description is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 is a block diagram illustrating an example content generation system in accordance with one or more techniques of this disclosure.

[0010] Figure 2 is an example image or surface in accordance with one or more techniques of this disclosure.

[0011] Figure 3 is a block diagram illustrating a graphics processing pipeline in accordance with one or more techniques of this disclosure.

[0012] Figure 4 is a block diagram illustrating an example graphics processing pipeline including a tessellation stage in accordance with one or more techniques of this disclosure.

[0013] Figure 5A is a triangle primitive including sub-primitives generated based on tessellation in accordance with one or more techniques of this disclosure.

[0014] Figure 5B is a triangle primitive including sub-primitives generated based on tessellation with a higher tessellation factor in accordance with one or more techniques of this disclosure.

[0015] Figure 6Ais a block diagram illustrating a graphics processing pipeline in which variable rate tessellation (VRT) per draw or per primitive is used, in accordance with one or more techniques of the present disclosure.

[0016] Figure 6B is a block diagram illustrating a graphics processing pipeline in which screen space image based or pixel count based VRT is used, in accordance with one or more techniques of the present disclosure.

[0017] Figure 7 is a diagram illustrating a screen space image overlaid with a frame, in accordance with one or more techniques of the present disclosure.

[0018] Figure 8 is a diagram illustrating a screen space image tile intersected with a triangle primitive, in accordance with one or more techniques of the present disclosure.

[0019] Figure 9A is a diagram illustrating a triangle primitive including triangle sub-primitives generated based on tessellation, in accordance with one or more techniques of the present disclosure.

[0020] Figure 9B is a diagram illustrating a quadrilateral primitive including triangle sub-primitives generated based on tessellation, in accordance with one or more techniques of the present disclosure.

[0021] Figure 10 is a call flow diagram illustrating example communications between a first GPU component and a second GPU component, in accordance with one or more techniques of the present disclosure.

[0022] Figure 11 is a flow diagram of an example method of graphics processing, in accordance with one or more techniques of the present disclosure.

[0023] Figure 12 is a flow diagram of an example method of graphics processing, in accordance with one or more techniques of the present disclosure. DETAILED DESCRIPTION

[0024] Various aspects of the systems, apparatuses, computer program products, and methods are more fully described below with reference to the figures. This disclosure may, however, be embodied in many different forms and should not be construed as limited to any specific structure or function presented throughout this disclosure. Rather, these aspects are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Based on the teachings herein one skilled in the art should appreciate that the scope of the disclosure is intended to cover any aspect of the systems, apparatuses, computer program products, and methods disclosed herein, whether implemented independently of, or combined with, other aspects of the disclosure. For example, an apparatus can be implemented or a method can be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method which is practiced using, in addition to or in place of the aspects set forth herein, other structures, functionalities, or structures and functions disclosed herein. Any aspect disclosed herein can be embodied by one or more elements of a claim.

[0025] Although various aspects are described herein, many variations and permutations of these aspects fall within the scope of the disclosure. While some potential benefits and advantages of aspects of the disclosure are mentioned, the scope of the disclosure is not intended to be limited to particular benefits, uses, or objectives. Rather, aspects of the disclosure are intended to be broadly applicable to different wireless technologies, system configurations, processing systems, networks, and transmission protocols, some of which are illustrated by way of example in the accompanying drawings and description below. The detailed description and drawings are merely illustrative of the disclosure and should not be taken as limiting the scope of the disclosure. The scope of the disclosure is defined by the appended claims and equivalents thereof.

[0026] Several aspects are presented with reference to various apparatus and methods. These apparatus and methods are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether such elements are implemented as hardware or software depends on the particular application and design constraints imposed on the overall system.

[0027] For example, an element, or any portion of an element, or any combination of elements can be implemented as a “processing system” that includes one or more processors. Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on a chip (SoC), baseband processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionality described throughout this disclosure. One or more processors in the processing system can execute software. Software can be interpreted and / or executable instructions that are stored in and accessed from storage. Firmware can also be used. Whether implemented as software or firmware, software can be stored in and accessed from memory. Examples of memory include random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic data media, optical data media, and / or other types of data storage media that can be used for storing and accessing instructions and / or data.

[0028] The term application can refer to software. As described herein, one or more techniques can involve an application (e.g., software) configured to perform one or more functions. In such examples, the application can be stored in memory (e.g., on-chip memory of a processor, system memory, or any other memory). Hardware, such as a processor, described herein can be configured to execute the application. For example, the application can be described as including code that, when executed by the hardware, causes the hardware to perform one or more techniques described herein. For example, the hardware can access the code from memory and execute the code accessed from memory to perform one or more techniques described herein. In some examples, components are identified in the disclosure. In such examples, a component can be hardware, software, or a combination thereof. A component can be a separate component or a subcomponent of a single component.

[0029] In one or more examples described herein, the described functions can be implemented in hardware, software, or any combination thereof. If implemented in software, the functions can be stored or encoded in one or more computer-readable media. Computer- readable media includes computer storage media. Storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, optical disk storage, magnetic disk storage, other magnetic storage devices, combinations of the foregoing, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.

[0030] As used herein, examples of the term "content" can refer to "graphics content," "images," and the like, regardless of whether the term is used as an adjective, a noun, or other part of speech. In some examples, as used herein, the term "graphics content" can refer to content produced by one or more processes of a graphics processing pipeline. In further examples, as used herein, the term "graphics content" can refer to content produced by a processing unit configured to perform graphics processing. In further examples, as used herein, the term "graphics content" can refer to content produced by a graphics processing unit.

[0031] In computer graphics, there can be a tradeoff between the level of detail (LOD) of a rendered image and the performance of a rendering GPU, which can be measured in frames per second (FPS). The LOD can be reduced in certain portions of an image without having too much impact on the perceived quality of the image. By identifying such regions (e.g., patches) or draws and reducing the detail in these regions or draws, the performance of the rendering GPU can be improved. In one or more aspects, a variable rate tessellation (VRT) technique can use shading rates from a variable rate shading (VRS) feature and / or average pixel counts per subprimitive value to identify regions or draws in which a low LOD can be used. The VRT technique can then reduce the tessellation rate for such regions or draws.

[0032] Figure 1is a block diagram illustrating an example content generation system 100 configured to implement one or more techniques of the present disclosure. The content generation system 100 includes a device 104. The device 104 can include one or more components or circuits for performing the various functions described herein. In some examples, one or more components of the device 104 can be components of a SOC. The device 104 can include one or more components configured to perform one or more techniques of the present disclosure. In the illustrated example, the device 104 can include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, the device 104 can include a number of optional components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). The display 131 can refer to one or more displays 131. For example, the display 131 can include a single display or multiple displays, which can include a first display and a second display. The first display can be a left eye display and the second display can be a right eye display. In some examples, the first display and the second display can receive different frames for presentation thereon. In other examples, the first display and the second display can receive the same frames for presentation thereon. In further examples, the results of graphics processing can not be displayed on the device, e.g., the first display and the second display can not receive any frames for presentation thereon. Instead, the frames or graphics processing results can be transmitted to another device. In some aspects, this can be referred to as split rendering.

[0033] The processing unit 120 can include internal memory 121. The processing unit 120 can be configured to perform graphics processing using the graphics processing pipeline 107. The content encoder / decoder 122 can include internal memory 123. In some examples, the device 104 can include a processor that can be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120 before the frames are displayed by one or more displays 131. While the processor in the example content generation system 100 is configured as a display processor 127, it should be understood that the display processor 127 is one example of a processor and that other types of processors, controllers, etc. can be used in place of the display processor 127. The display processor 127 can be configured to perform display processing. For example, the display processor 127 can be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120. The one or more displays 131 can be configured to display or otherwise present the frames processed by the display processor 127. In some examples, the one or more displays 131 can include one or more of a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.

[0034] Memory external to the processing unit 120 and the content encoder / decoder 122, such as system memory 124, can be accessible to the processing unit 120 and the content encoder / decoder 122. For example, the processing unit 120 and the content encoder / decoder 122 can be configured to read from and / or write to the external memory, such as system memory 124. The processing unit 120 can be communicatively coupled to the system memory 124 by a bus. In some examples, the processing unit 120 and the content encoder / decoder 122 can be communicatively coupled to the internal memory 121 by a bus or via a different connection.

[0035] The content encoder / decoder 122 can be configured to receive graphics content from any source, such as system memory 124 and / or the communication interface 126. The system memory 124 can be configured to store received encoded or decoded graphics content. The content encoder / decoder 122 can be configured to receive encoded or decoded graphics content in the form of encoded pixel data, for example, from the system memory 124 and / or the communication interface 126. The content encoder / decoder 122 can be configured to encode or decode any graphics content.

[0036] The internal memory 121 or system memory 124 can include one or more volatile or non-volatile memories or storage devices. In some examples, the internal memory 121 or system memory 124 can include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, a magnetic data medium, or an optical storage medium, or any other type of memory. According to some examples, the internal memory 121 or system memory 124 can be a non-transitory storage medium. The term "non-transitory" can indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted to mean that the internal memory 121 or system memory 124 is non-removable or that its contents are static. As one example, the system memory 124 can be removed from the device 104 and moved to another device. As another example, the system memory 124 can be non-removable from the device 104.

[0037] The processing unit 120 can be a CPU, GPU, GPGPU, or any other processing unit that can be configured to perform graphics processing. In some examples, the processing unit 120 can be integrated into a motherboard of the device 104. In further examples, the processing unit 120 can be present on a graphics card that is installed in a port of the motherboard of the device 104, or can be otherwise incorporated within a peripheral device configured to interoperate with the device 104. The processing unit 120 can include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuitry, or any combinations thereof. If the techniques are implemented partially in software, the processing unit 120 can store instructions for the software in suitable, non-transitory computer-readable storage media (e.g., the internal memory 121) and can execute the instructions in hardware to perform the techniques of this disclosure. Any of the foregoing including hardware, software, a combination of hardware and software, etc., can be considered one or more processors.

[0038] The content encoder / decoder 122 can be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 can be incorporated into a motherboard of the device 104. The content encoder / decoder 122 can include one or more processors, such as one or more microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuitry, or any combinations thereof. If the techniques are implemented partially in software, the content encoder / decoder 122 can store instructions for the software in a suitable, non-transitory computer-readable storage medium (e.g., the internal memory 123) and can execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing including hardware, software, a combination of hardware and software, etc., can be considered one or more processors.

[0039] In some aspects, the content generation system 100 can include an optional communication interface 126. The communication interface 126 can include a receiver 128 and a transmitter 130. The receiver 128 can be configured to perform any of the receiving functions described herein with respect to the device 104. In addition, the receiver 128 can be configured to receive information (e.g., eye or head position information, rendering commands, and / or position information) from another device. The transmitter 130 can be configured to perform any of the transmitting functions described herein with respect to the device 104. For example, the transmitter 130 can be configured to transmit information that can include a request for content to another device. The receiver 128 and the transmitter 130 can be combined into a transceiver 132. In such examples, the transceiver 132 can be configured to perform any of the receiving functions and / or the transmitting functions described herein with respect to the device 104.

[0040] Referring again to Figure 1In certain aspects, the display processor 127 can include a VRT unit 198 configured to receive data for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more sub-primitives. The VRT unit 198 can be configured to reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a TRF. The characteristic can correspond to a shading rate or a number of visible pixels. The VRT unit 198 can be configured to apply the TRF for each tile of the plurality of tiles. The VRT unit 198 can be configured to render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles. Although the following description can focus on graphics processing, the concepts described herein can be applicable to other similar processing techniques.

[0041] A device, such as device 104, can refer to any device, apparatus, or system configured to perform one or more techniques described herein. For example, a device can be a server, a base station, a user device, a client device, a station, an access point, a computer such as a personal computer, a desktop computer, a laptop computer, a tablet computer, a computer workstation, or a mainframe computer, an end product, an appliance, a telephone, a smartphone, a server, a video game platform or console, a handheld device such as a portable video game device or a personal digital assistant (PDA), a wearable computing device such as a smartwatch, an augmented reality device, or a virtual reality device, a non-wearable device, a display or display device, a television, a television set-top box, an intermediary network device, a digital media player, a video streaming device, a content streaming device, an in-vehicle computer, any mobile device, any device configured to generate graphics content, or any device configured to perform one or more techniques described herein. Processes herein can be described as being performed by a particular component (e.g., a GPU), but in other embodiments, can be performed using other components (e.g., a CPU) that are consistent with the disclosed embodiments.

[0042] A GPU can process multiple types of data or data packets in a GPU pipeline. For example, in some aspects, a GPU can process two types of data or data packets, such as context register packets and draw call data. A context register packet can be a collection of global state information that can manage how a graphics context will be processed, such as information about global registers, shader programs, or constant data. For example, a context register packet can include information about a color format. In some aspects of a context register packet, there can be a bit that indicates which workload belongs to the context register. Further, there can be multiple functions or programs that run simultaneously and / or in parallel. For example, a function or program can describe a particular operation, such as a color mode or color format. Thus, a context register can define multiple states of a GPU.

[0043] A context state can be used to determine how a single processing unit (e.g., a vertex fetcher (VFD), a vertex shader (VS), a shader processor, or a geometry processor) runs, and / or in what mode a processing unit runs. To do so, a GPU can use context registers and programming data. In some aspects, a GPU can generate workloads in a pipeline, such as vertex or pixel workloads, based on a context register definition of a mode or state. Certain processing units (e.g., a VFD) can use these states to determine certain functions, such as how to assemble a vertex. Since these modes or states can change, a GPU can need to change the corresponding context. Further, workloads corresponding to the modes or states can follow the changed modes or states.

[0044] A GPU can render images in a variety of different ways. In some instances, a GPU can render images using render and / or tile-based rendering. In a tile-based rendering GPU, an image can be divided or separated into different sections or tiles. After dividing the image, each section or tile can be rendered separately. A tile-based rendering GPU can divide a computer graphics image into a grid format such that each section of the grid (i.e., a tile) can be rendered separately. In some aspects, in a binning pass, an image can be divided into different bins or tiles. In some aspects, during a binning pass, a visibility stream can be constructed in which visible primitives or draw calls can be identified. In contrast to tile-based rendering, direct rendering does not divide a frame into smaller bins or tiles. Rather, in direct rendering, an entire frame is rendered at a single time. Further, some types of GPUs can allow for both tile-based rendering and direct rendering (e.g., flexible rendering).

[0045] In some aspects, a GPU can apply a draw or render process to different bins or tile regions. For example, a GPU can render for one bin and perform all of the drawing for primitives or pixels in the bin. During the process of rendering for the bin, a render target can be located in GPU internal memory (GMEM). In some instances, after rendering for one bin, the contents of the render target can be moved to system memory and the GMEM can be freed for rendering the next bin. Further, the GPU can render for another bin and perform drawing for primitives or pixels in the bin. Thus, in some aspects, there can be a small number of bins (e.g., four bins) that cover all of the drawing in one surface. Further, the GPU can loop through all of the drawing in one bin, but perform drawing for visible draw calls (i.e., draw calls that include visible geometry). In some aspects, for example in a binning pass, a visibility stream can be generated to determine visibility information for each primitive in an image or scene. For example, the visibility stream can identify whether a certain primitive is visible. In some aspects, this information can be used to remove invisible primitives, for example in a render pass. Further, at least some of the primitives that are identified as visible in the primitives can be rendered in the render pass.

[0046] In some aspects of tile-based rendering, there can be multiple processing stages or passes. For example, rendering can be performed in two passes, for example, a visibility or bin-visibility pass and a render or bin-render pass. During the visibility pass, a GPU can input a render workload, record the positions of primitives or triangles, and then determine which primitives or triangles fall into which bin or region. In some aspects of the visibility pass, the GPU can also identify or mark the visibility of each primitive or triangle in the visibility stream. During the render pass, the GPU can input the visibility stream and process one bin or region at a time. In some aspects, the visibility stream can be analyzed to determine which primitives or vertices of primitives are visible or invisible. Thus, visible primitives or vertices of primitives can be processed. By doing so, the GPU can reduce unnecessary workload of processing or rendering invisible primitives or triangles.

[0047] In some aspects, certain types of primitive geometry, such as position-only geometry, can be processed during the visibility path. Furthermore, primitives can be classified into different bins or regions based on their position or location. In some instances, classifying primitives or triangles into different bins can be performed by determining visibility information for those primitives or triangles. For example, the GPU can determine the visibility information for each primitive in each bin or region or write that visibility information to, for example, system memory. This visibility information can be used to determine or generate a visibility stream. In the rendering path, the primitives in each bin can be rendered separately. In these instances, the visibility stream can be retrieved from memory used to discard primitives that are not visible to that bin.

[0048] GPUs or GPU architectures can offer a variety of different options for rendering, such as software rendering and hardware rendering. In software rendering, the driver or CPU can copy the entire frame geometry by processing each view at a time. Furthermore, several different states can change depending on the view. Therefore, in software rendering, the software can copy the entire workload by changing some states that can be used for rendering for each viewpoint in the image. In some aspects, there can be increased overhead because the GPU may submit the same workload multiple times for each viewpoint in the image. In hardware rendering, the hardware or GPU can be responsible for copying or processing the geometry for each viewpoint in the image. Therefore, the hardware can manage the copying or processing of primitives or triangles for each viewpoint in the image.

[0049] Figure 2 An image or surface 200 according to one or more technologies of this disclosure is shown, which includes multiple elements divided into multiple compartments. For example... Figure 2 As shown, the image or surface 200 includes region 202, which includes primitives 221, 222, 223, and 224. Primitives 221, 222, 223, and 224 are divided or placed into different bins, such as bins 210, 211, 212, 213, 214, and 215. Figure 2 An example of patch rendering using multiple viewpoints is shown for primitives 221-224. For example, primitives 221-224 are located in a first viewpoint 250 and a second viewpoint 251. Therefore, a GPU that processes or renders an image or surface 200 including region 202 can utilize multiple viewpoints or multi-view rendering.

[0050] As indicated herein, a GPU or graphics processor unit can use a tiled rendering architecture to reduce power consumption or save memory bandwidth. As further stated above, such a rendering method can divide a scene into a plurality of bins, and include a visibility pass that identifies triangles that are visible in each bin. Thus, in tiled rendering, a complete screen can be divided into a plurality of bins or tiles. The scene can then be rendered multiple times, e.g., once or more times for each bin.

[0051] In aspects of graphics rendering, some graphics applications can render to a single target (i.e., render target) one or more times. For example, in graphics rendering, a frame buffer on system memory can be updated multiple times. A frame buffer can be a portion of memory or random access memory (RAM) (e.g., containing a bitmap or storage) to help store display data for a GPU. A frame buffer can also be a memory buffer containing a complete frame of data. Further, a frame buffer can be a logical buffer. In some aspects, updating a frame buffer can be performed in bin or tile rendering, where, as discussed above, a surface is divided into a plurality of bins or tiles, and each bin or tile can then be rendered separately. Further, in tiled rendering, a frame buffer can be divided into a plurality of bins or tiles.

[0052] As indicated herein, in some aspects (such as in a bin or tiled rendering architecture), a frame buffer can have data repeatedly stored or written to it, e.g., when rendering from different types of memory. This can be referred to as resolving and un-resolving a frame buffer or system memory. For example, when storing or writing to one frame buffer and then switching to another frame buffer, data or information on the frame buffer can be resolved from the GMEM at the GPU to memory in system memory, i.e., double data rate (DDR) RAM or dynamic RAM (DRAM).

[0053] In some aspects, system memory can also be a system on a chip (SoC) memory or another chip-based memory for storing data or information, e.g., on a device or smartphone. System memory can also be a physical data storage shared by a CPU and / or GPU. In some aspects, system memory can be a DRAM chip, e.g., on a device or smartphone. Thus, SoC memory can be a chip-based way of storing data.

[0054] In some aspects, the GMEM can be on-chip memory at the GPU, which can be implemented by static RAM (SRAM). Further, the GMEM can be stored on a device (e.g., a smartphone). As indicated herein, data or information can be transferred, for example, at a device, between system memory or DRAM and GMEM. In some aspects, the system memory or DRAM can be located at the CPU or GPU. Further, data can be stored at DDR or DRAM. In some aspects, such as in bin or tile-based rendering, a small portion of memory can be stored at the GPU (e.g., at the GMEM). In some instances, storing data at the GMEM can use greater processing workload and / or power consumption compared to storing data at a frame buffer or system memory.

[0055] Figure 3 is a block diagram illustrating a graphics processing pipeline 300. The pipeline 300 includes an input assembler stage 382, a vertex shader stage 384, a geometry shader stage 386, a rasterizer stage 388, a pixel shader stage 390, and an output merger stage 392. In some examples, an application program interface (API) can be configured to use each of these stages as shown in Figure 3 The graphics processing pipeline 400 is described below as being performed by the processing unit 120, but can be performed by various other graphics processors.

[0056] The graphics processing pipeline 300 generally includes programmable stages (e.g., illustrated with rounded corners) and fixed function stages (e.g., illustrated with square corners). For example, graphics rendering operations associated with certain stages of the graphics rendering pipeline 300 are generally performed by programmable shader processors, while other graphics rendering operations associated with other stages of the graphics rendering pipeline 300 are generally performed by unprogrammable, fixed function hardware units associated with the processing unit 120. Graphics rendering stages performed by shader units can generally be referred to as "programmable" stages, while stages performed by fixed function units can generally be referred to as fixed function stages. Indications herein as to whether different stages are programmable stages or fixed function stages can be exemplary, and other combinations between programmable stages or fixed function stages can also be used.

[0057] The input assembler stage 382 assembles vertices from a vertex buffer into a primitive buffer. The vertex shader stage 384 receives vertices from the input assembler stage 382 and processes the vertices. The vertex shader stage 384 can be implemented by a vertex shader unit of the processing unit 120. The geometry shader stage 386 receives vertices from the vertex shader stage 384 and processes the vertices. The geometry shader stage 386 can be implemented by a geometry shader unit of the processing unit 120. Figure 3The input assembler stage 382 is shown as a fixed function stage in the example and is generally responsible for providing graphics data (e.g., triangles, lines, and points) to the graphics processing pipeline 300. For example, the input assembler stage 382 can collect vertex data for high order surfaces, primitives, etc., and output the vertex data and attributes to the vertex shader stage 384. Thus, the input assembler stage 382 can use fixed function operations to read vertices from off-chip memory. The input assembler stage 382 can then create a pipeline work item from these vertices, while also generating a vertex identifier (“Vertex ID”), an instance identifier (“Instance ID”, which can be used by vertex shaders), and a primitive identifier (“Primitive ID”, which can be used by geometry shaders and pixel shaders). The input assembler stage 382 can automatically generate the Vertex ID, Instance ID, and Primitive ID when reading the vertices.

[0058] The vertex shader stage 384 can process the received vertex data and attributes. For example, the vertex shader stage 384 can perform per-vertex processing such as transformation, skinning, vertex displacement, and compute per-vertex material properties. In some examples, the vertex shader stage 384 can generate texture coordinates, vertex colors, vertex lighting, fog factors, etc. The vertex shader stage 384 generally takes a single input vertex and outputs a single processed output vertex.

[0059] The geometry shader stage 386 can receive a primitive defined by vertex data (e.g., three vertices for a triangle, two vertices for a line, or a single vertex for a point) and further process the primitive. For example, the geometry shader stage 386 can perform per-primitive processing such as silhouette-edge and shadow volume extrusion, among other possible processing operations. Thus, the geometry shader stage 386 can receive one primitive as input (which can include one or more vertices) and output zero, one, or more primitives (which likewise can include one or more vertices). The output primitives can contain more data than would be possible without the geometry shader stage 386. The total amount of output data can be equal to the vertex size times the vertex count, and can be limited per invocation. The stream output from the geometry shader stage 386 can allow the primitives that reach the stage to be stored to off-chip memory. The stream output is generally associated with the geometry shader stage 386, and both can be programmed together (e.g., using an API).

[0060] The rasterizer stage 388 can be a fixed function stage responsible for clipping primitives and preparing the primitives for the pixel shader stage 390. For example, the rasterizer stage 388 can perform clipping (including custom clipping bounds), perspective division, viewport / scissor selection and implementation, render target selection, and primitive setup. In this way, the rasterizer stage 388 can generate a plurality of fragments for shading by the pixel shader stage 390.

[0061] The pixel shader stage 390 can receive the fragments from the rasterizer stage 388 and generate per-pixel data such as color. The pixel shader stage 390 can also perform per-pixel processing such as texture blending and lighting model calculations. Thus, the pixel shader stage 390 can receive one pixel as input and can output one pixel (or a zero value for a pixel) at the same relative position.

[0062] The output merger stage 392 can be responsible for combining various types of output data such as pixel shader values, depth information, and stencil information to generate a final result. For example, the output merger stage 392 can perform fixed function blending, depth, and / or stencil operations for a render target (pixel location). Although the foregoing descriptions are generally described above with respect to the vertex shader stage 384, the geometry shader stage 386, and the pixel shader stage 390, each of the foregoing descriptions can refer to one or more shader units designated by the GPU to perform the respective shading operations.

[0063] Certain GPUs can not support all of the shader stages shown in Figure 3 For example, due to hardware and / or software limitations (e.g., a limited number of shader units and associated components), some GPUs can not be able to designate a shader unit to perform more than two shading operations. In an example, certain GPUs can not support the operations associated with the geometry shader stage 386. Rather, the GPU can include support for designating shader units to perform the vertex shader stage 384 and the pixel shader stage 390. Thus, the operations performed by the shader units can follow the input / output interfaces associated with the vertex shader stage 384 and the pixel shader stage 390.

[0064] Further, in some examples, the introduction of the geometry shader stage 386 into the graphics processing pipeline relative to a graphics processing pipeline that does not include the geometry shader stage 386 can result in additional reads and writes to storage units. For example, as described above, the vertex shader stage 384 can write out vertices to off-chip memory. The geometry shader stage 386 can read these vertices (output by the vertex shader stage 384) and write out new vertices that are then pixel shaded.

[0065] Figure 4 is a block diagram illustrating an example graphics processing pipeline 400 that includes a tessellation stage. For example, the pipeline 400 includes an input assembler stage 440, a vertex shader stage 442, a hull shader stage 444, a tessellator stage 446, a domain shader stage 448, a geometry shader stage 450, a rasterizer stage 452, a pixel shader stage 454, and an output merger stage 456. In some examples, an API can be configured to use each of the stages shown in Figure 4 The graphics processing pipeline 400 is described below as being performed by the processing unit 120, but can also be performed by various other graphics processors.

[0066] Certain stages shown in Figure 4 may be configured similarly or identically to the stages shown and described with respect to Figure 3 ( e.g., the assembler stage 440, the vertex shader stage 442, the geometry shader stage 450, the rasterizer stage 452, the pixel shader stage 454, and the output merger stage 456). Further, the pipeline 400 includes additional stages for hardware tessellation. For example, in addition to the stages described above with respect to Figure 3 , the graphics processing pipeline 400 also includes the hull shader stage 444, the tessellator stage 446, and the domain shader stage 448. That is, the hull shader stage 444, the tessellator stage 446, and the domain shader stage 448 are included to accommodate tessellation by the processing unit 120, rather than being performed by the software application that is being executed (e.g., by the CPU).

[0067] The hull shader stage 444 can receive primitives from the vertex shader stage 442 and can be responsible for performing at least two actions. First, the hull shader stage 444 can be responsible for determining a set of tessellation factors. The hull shader stage 444 can generate tessellation factors once per primitive. The tessellator stage 446 can use the tessellation factors to determine a level of detail at which to tessellate (e.g., split) a given primitive into smaller parts. The hull shader stage 444 can also be responsible for generating control points that will later be used by the domain shader stage 448. That is, for example, the hull shader stage 444 can be responsible for generating control points that will be used by the domain shader stage 448 to create actual tessellated vertices that are ultimately used at the time of rendering.

[0068] When the subdivision stage 446 receives data from the shell shader stage 444, it can use one of several algorithms to determine the appropriate sampling pattern for the current primitive type. For example, typically, the subdivision stage 446 can convert the requested subdivision amount (as determined by the shell shader stage 444) into a set of coordinate points within the current "domain". That is, depending on the subdivision factor from the shell shader stage 444 and the specific configuration of the subdivision stage 446, it can determine which points in the current primitive need to be sampled to subdivide the input primitive surface into smaller portions. The output of the subdivision stage 446 can be a set of domain points, which may include barycentric coordinates.

[0069] In addition to the control points generated by the shell shader stage 444, the domain shader stage 448 can also employ domain points and use them to create new vertices. The domain shader stage 448 can use a complete list of control points generated for the current primitive, texture, procedural algorithm, or any other algorithm to translate the centroid "position" for each subdivided point into the output geometry passed to the next stage in the pipeline. As mentioned above, some GPUs may not support... Figure 4 The diagram illustrates all shader stages. For example, due to hardware and / or software limitations (e.g., a limited number of shader units and associated components), some GPUs may not be able to specify shader units to perform more than two shader operations. In the example, some GPUs may not support the operations associated with geometry shader stage 450, hull shader stage 444, and domain shader stage 448. Instead, the GPU may include support for specifying shader units to perform vertex shader stage 442 and pixel shader stage 454. Therefore, the operations performed by the shader units can follow the input / output interfaces associated with vertex shader stage 442 and pixel shader stage 454.

[0070] In computer graphics, there may be a trade-off between the Level of Displacement (LOD) of a rendered image and the performance of the rendering GPU (which can be measured in frames per second, FPS). In other words, a high LOD may result in a decrease in FPS. Conversely, performance can be improved by reducing the LOD. The LOD can be reduced in certain parts of an image without significantly impacting the image's perceptual quality. The performance of the rendering GPU can be improved by identifying or drawing such regions (e.g., tiles) and reducing the detail in those regions or the drawn areas.

[0071] Examples of regions or objects where low LOD can be acceptable can include objects that are far away from the camera, regions that are not in the user’s focus (e.g., the focus of the user’s eyes), moving objects (such objects can typically be motion blurred), or objects that are in shadow or not well lit.

[0072] The number of primitives of a rendered object and the pixel shading rate of rasterized primitives can affect the LOD of a rendering operation.

[0073] The number of primitives in a draw or object can be changed in the vertex shader stage of the graphics pipeline. Changing the number of primitives specified by a developer can not be recommended because doing so can cause undesirable effects such as breaking the watertightness of an object or draw (e.g., a watertight mesh collection can refer to a mesh that consists of one closed surface). For tessellated primitives, the number of tessellated primitives on the inside (or interior) can be modified without breaking the watertightness of the draw. Sometimes, tessellation can generate a large number of sub-primitives, many of which can ultimately contribute little to the quality of an image. The tessellation rate of such primitives on the inside can be reduced without much reduction in image quality.

[0074] Figure 5A A triangle primitive 500A is shown that includes sub-primitives generated based on tessellation. The shaded sub-primitives can correspond to tessellated sub-primitives on the inside, and the unshaded sub-primitives can correspond to tessellated sub-primitives on the outside. Figure 5B A triangle primitive 500B is shown that includes sub-primitives generated based on tessellation with a higher tessellation factor. The triangle primitive 500B can be the same as the triangle primitive 500A. Due to the higher tessellation factor, the triangle primitive 500B can include many more sub-primitives than the triangle primitive 500A. Similar to Figure 5A In the Figure 5B triangle primitive 500B, the shaded sub-primitives can correspond to tessellated sub-primitives on the inside, and the unshaded sub-primitives can correspond to tessellated sub-primitives on the outside.

[0075] A tessellation rate can be specified by a tessellation factor (TessFactor). The TessFactor can indicate how many times an edge of a primitive can be tessellated. There can be separate TessFactors for outside edges and inside edges (i.e., an outside TessFactor and an inside TessFactor).

[0076] In one or more aspects, the VRT technique can use a shading rate and / or an average pixel count per subprimitive value from the VRS feature to identify areas or draws where a low LOD can be used. The VRT technique can then reduce the tessellation rate for such areas or draws.

[0077] The VRS can be an API feature and can allow a developer to specify different shading rates for different areas or draws of a frame. A shading rate can refer to a resolution at which a pixel shader (or fragment shader) is invoked. Without VRS, one pixel can be shaded per pixel / fragment shader operation. With VRS, a group of nearby pixels can be shaded together in one operation. The group of pixels can be a 1 x 2 (e.g., horizontally oriented), 2 x 1 (e.g., vertically oriented), or 4 x 4 pixel block, etc. In areas where the LOD specification is low, the developer can reduce the shading rate. In other words, more pixels can be shaded in one shader operation. In one or more aspects, based on the VRT technique, if the shading rate is low, the tessellation rate can be reduced.

[0078] In one or more aspects, in addition to or instead of using the VRS-based shading rate, the VRT technique can use a pixel count to identify a tessellation rate. Specifically, primitives with a low pixel count per subprimitive can be identified, and the tessellation rate for such primitives can be reduced.

[0079] In one or more aspects, the tessellation operation can be performed in quad and tri domains with triangles as the output primitive type.

[0080] In some examples, the inner TessFactor (rather than the outer TessFactor) can be different. In such examples, the outer TessFactor can not change because doing so can break the watertightness of the draw.

[0081] A tessellation reduction factor (TRF) can refer to a multiplicand that the inner TessFactor can be multiplied by in order to reduce the inner TessFactor. The TRF can fall within a certain range, e.g., a (0-1] range (i.e., excluding 0 and including 1). When the TRF is 1, the original TessFactor can be preserved.

[0082] When a geometry shader is enabled, the VRT technique can be omitted. As mentioned above, a geometry shader can take a single primitive as input and output multiple primitives. Because the number of input primitives for a geometry shader can be reduced based on VRT, the output image can change significantly if VRT is used in conjunction with a geometry shader. Geometry shaders and tessellation are rarely enabled together.

[0083] Per-draw and per-primitive VRT techniques can be used in both tile-based (warehouse-based) and direct rendering architectures. Screen-space image-based and pixel-count-based VRT techniques can be used in tile-based architectures. The per-draw, per-primitive, screen-space image-based, and pixel-count-based methods will be explained in further detail below.

[0084] In one or more aspects, as part of the VRT, the TRF can be identified based on the shading rate (e.g., pixel pattern) that can be received from the VRS features. According to one aspect, Table 1 below can provide TRFs corresponding to different shading rates.

[0085]

[0086]

[0087] Table 1

[0088] Here, a TRF can be chosen to maintain the ratio of the number of pixels to the number of sub-primitives ((number of pixels) / (number of sub-primitives)) (or to minimize the variation). In different configurations, the number of pixels can refer to the number of effective pixels or the number of visible pixels, as will be explained in more detail below. For example, consider the case where the shading rate is 2 x 2 pixels. Therefore, the effective pixel count can be reduced by a factor of 4 because 4 pixels can be grouped together to form 1 coarse (effective) pixel. To maintain the same ratio of the (effective) pixel count to the number of sub-primitives, the number of sub-primitives can also be reduced by a factor of 4. Since the number of sub-primitives can vary approximately by the square of the internal TessFactor, in this case, a TRF of 1 / 2 can be chosen.

[0089] In VRS features, the shading rate can be specified at different granularities. These can include per-draw granularity, per-primal (e.g., per provoking vertex) granularity, per-screen-space image granularity, or any combination of the above three granularities.

[0090] The VRT technique associated with each shading rate granularity is described in detail below. Specifically, the VRT technique can use VRS input to identify the internal TessFactor.

[0091] Figure 6A is a block diagram illustrating a graphics processing pipeline 600A in which per- draw or per-primitive VRT is used. The various stages can correspond to the same stages shown in Figure 3 and Figure 4 In the binning (visibility) pass, block 602 can include an input assembler stage and a geometry processing stage. Block 604 can include a hull shader stage, which can generate and output an unmodified TessFactor. Block 606 can include a TessFactor adjuster stage, which can adjust the TessFactor based on VRS input including a shading rate of the draw or primitive, to output an adjusted TessFactor. Block 608 can include a tessellator stage, which can tessellate based on the adjusted TessFactor. Block 610 can include a domain shader stage, a low resolution Z (LRZ) stage, and a visibility stream compressor (VSC) stage. Block 610 can output a visibility stream, which can be processed in the rendering pass. In the rendering pass, block 612 can include an input assembler stage and a geometry processing stage. Block 614 can include a hull shader stage, which can generate and output an unmodified TessFactor. Block 616 can include a TessFactor adjuster stage, which can adjust the TessFactor based on VRS input including a shading rate of the draw or primitive, to output an adjusted TessFactor. Block 618 can include a tessellator stage, which can tessellate based on the adjusted TessFactor. Block 620 can include a domain shader stage and a pixel shader stage. Block 620 can output a rendered image.

[0092] In one aspect, using the VRS feature, a developer or application can specify a per-draw shading rate in a command buffer. The LOD specified for each draw can be identified based on the shading rate. The VRT technique can receive the shading rate (e.g., at blocks 606 and 616) and can apply a corresponding TRF to all primitives in the draw call. In some examples, without directly specifying a per-draw shading rate through the API, shading rate information for the same frame can be obtained on a per-draw or per-primitive basis based on the binning pass. In further examples, the shading rate information can be obtained based on data from a previous frame.

[0093] The same TessFactor reduction technique can be applied in both the binning (visibility) pass and the rendering pass. Thus, the per-primitive VRT technique can save loops in the hardware tessellator and domain shader in both the binning pass and the rendering pass.

[0094] In one aspect, the VRS can allow a developer to specify a shading rate at a per-primitive level (e.g., per-primed vertex (one of the vertices in an output primitive can be designated as a primed vertex, and this designation can be used by other processes (e.g., using output from the primed vertex instead of output from other vertices))). The shading rate at a per-primitive level can be specified with the primed vertex in the command buffer. If there is no per-primitive specified shading rate, a shading rate of 1x1 can be assumed.

[0095] Changing the per-primitive tessellation rate can be similar to changing the per- draw tessellation rate, except for the granularity of applying the VRT technique. Here, the VRT technique can identify the TRF, and can modify the internal TessFactor for each primitive based on the shading rate of the primitive.

[0096] Similar to per-draw VRT, the per-primitive VRT technique can also save loops in the hardware tessellator and domain shader in both the binning (visibility) pass and the rendering pass.

[0097] Figure 6B is a block diagram illustrating a graphics processing pipeline 600B in which screen-space image based or pixel count based VRT is used. The various stages can correspond to the same stages shown in Figure 3 and Figure 4 In the binning (visibility) pass, block 652 can include an input assembler stage, a geometry processing stage, an hull shader stage, a tessellator stage, a domain shader stage, and an LRZ stage. A TRF for each draw or primitive in the draw or primitives can be generated at block 652 and can be stored at a TRF buffer 662. Block 654 can include a visibility stream compressor stage. In the rendering pass, block 656 can include an input assembler stage, a geometry processing stage, and a hull shader stage. The unmodified TessFactor provided by the hull shader stage can be adjusted based on the corresponding TRF value that can be fetched from the TRF buffer 662. Block 658 can include a tessellator stage and a domain shader stage. The tessellator stage can tessellate based on the adjusted TessFactor. Block 660 can include a pixel shader stage. Block 660 can output a rendered image.

[0098] In one aspect, in the VRS feature, using a screen space image, a developer can specify a shading rate for different regions of a frame. The shading rate can be specified for a block of pixel tiles (e.g., 8x 8, 16x 16, or 32x 32 pixel tiles) as indicated by the VRS tile size. Depending on the shading rate, the VRT technique can assign a TRF to each tile.

[0099] Figure 7 is a diagram 700 showing a screen space image overlapping with a frame. The origin (e.g., the top-left corner point) of the screen space image can be aligned with the origin (e.g., the top-left corner point) of the frame. If the frame is larger than the screen space image, the spill over region of the frame can have a TRF value set to “1” (i.e., no TessFactor reduction). Thus, in Figure 7 , based on different LOD specifications, the region of the frame that is within the screen space image can be associated with a low tessellation rate (for a low LOD specification), a medium tessellation rate (for a medium LOD specification), or a high tessellation rate (for a high LOD specification). The low, medium, or high tessellation rate can be associated with a low, medium, or high TRF value, respectively. The spill over region of the frame that is outside the screen space image can be associated with a default unmodified tessellation rate (i.e., the TRF value can be set to “1”).

[0100] In screen space image based VRT, the position of the primitives output from the tessellator can be available after the geometry processing stage in the binning (visibility) pass, instead of before. Thus, the VRT technique can not be applied in the binning (visibility) pass.

[0101] The LRZ stage, which can come after the geometry processing stage, can have access to both the screen space image and the position of the primitives on the frame. At this stage of the graphics pipeline, the TRF value can be computed, and the TRF value can be assigned to each tile. Among the TRFs of the tiles associated with (e.g., at least partially covered by) a primitive, the highest TRF value can be selected as the TRF of the primitive.

[0102] Figure 8 is a diagram 800 showing screen space image tiles intersecting with a triangle primitive. The triangle primitive can be distributed across 3 screen space image tiles. The numbers (“0.5”, “0.5”, “1”, and “0.25”) can be the TRF values for the respective screen space image tiles. The highest TRF value (which can correspond to the smallest TessFactor reduction) (which is “1” for the lower-left tile) can be selected as the TRF for the triangle primitive.

[0103] Referring back to Figure 6BThe selected TRF values can be stored in a TRF buffer 662. In the rendering pass, for each tessellated primitive, the TRF can be fetched from the TRF buffer 662 and the TessFactor can be modified accordingly. The adjusted TessFactor can then be sent to the hardware tessellator.

[0104] The screen-space based VRT technique can save loops in the hardware tessellator and domain shader in the rendering pass (but not in the binning pass).

[0105] In one aspect, the VRT technique can be pixel count based. A tessellator can sometimes over-tessellate a primitive. More tessellation generally results in a higher LOD. However, beyond a certain point, tessellation can provide less and less return in image quality. For small primitives, the point of diminishing returns can be reached more quickly. Primitives that are far from the camera can be associated with higher depth values, but with relatively small bounding boxes. Thus, heavy tessellation can not be desirable for such primitives.

[0106] Some of the sub-primitives can be Z culled or plane culled (i.e., farther sub-primitives can be hidden by nearer overlapping sub-primitives and thus can be removed). Culled sub-primitives can not contribute to image quality. Primitives with a large number of culled sub-primitives can be effectively over-tessellated.

[0107] The ratio of the number of pixels to the number of sub-primitives can be a good indicator that can be used to identify over-tessellation. A low value of the ratio can indicate that on average each sub-primitive holds very few pixels.

[0108] In pixel count based VRT, a developer can specify a threshold associated with the ratio of the number of visible pixels to the number of sub-primitives. For primitives with a ratio below the threshold, the number of sub-primitives can be reduced so that the ratio can become equal to the threshold.

[0109] The LRZ stage can include a counter for the pixels of each tile. Accumulating the value across all tiles of a primitive can provide the total number of visible pixels of the primitive. Using the pixel count and the threshold value, the desired new sub-primitive count can be computed as follows.

[0110] Desired new sub-primitive count N_vrt = number of visible pixels / threshold

[0111] Based on N_vrt, an iterative method can be used to calculate the TRF of a primitive (i.e., different TRF values can be tried in order to find the TRF value that produces a subprimitive count that is closest to the desired new subprimitive count N_vrt). This will be explained in further detail below. An exact TRF value can not be necessary for VRT. Thus, in some configurations, instead of using the above equation, a less accurate but simpler method can be used.

[0112] The TRF value for each primitive can be stored in the TRF buffer 662 at the end of the binning (visibility) pass. This TRF can be used to modify the TessFactor of the primitive in the rendering pass.

[0113] Pixel count based VRT techniques can save loops in the hardware tessellator and domain shader in the rendering pass.

[0114] To explain the process for finding the TRF value that produces a subprimitive count that is closest to the desired new subprimitive count N_vrt, some equations, functions, and variables can be defined as follows.

[0115] TF[x] = TessFactor[x] for edge x

[0116] P[x], the odd check for edge x = 2T[x] - P[x-1] - P[x+1] + 4P[i] - 6T[i] (T[i] - 1 - P[i])

[0117]

[0118] Table 2

[0119] For edge x, ( denotes the ceiling function), the variables denoting edges: b - bottom edge, l - left edge, t - top edge, r - right edge, i - interior edge (Tri domain), u - horizontal interior edge (Quad domain), and v - vertical interior edge (Quad domain).

[0120] Figure 9A is a schematic diagram showing a triangle primitive 900A including triangle subprimitives generated based on tessellation. In Figure 9A The left edge (l), right edge (r), bottom edge (b), and interior edge (i) are indicated in

[0121] Subprimitive count = 2T[b] + 2T[l] + 2T[r] - P[b] - P[l] - P[r] + 4P[i] + 6T[i] (T[i] - 1 - P[i]) → Equation [1]

[0122] Figure 9Bis a schematic diagram showing a quad primitive 900B comprising triangle sub-primitives generated based on tessellation. In Figure 9B The left edge (1), right edge (r), top edge (t), bottom edge (b), horizontal interior edge (u), and vertical interior edge (v) are indicated in

[0123] Sub-primitive count = 2T[b] + 2T[l] + 2T[t] + 2T[r] - P[b] - P[l] - P[t] - P[r] + 2P[u] + 2P[v] - 4T[v](P[u] + 1) - 4T[u](P[v] + 1) + 8T[u]T[v] + c → Equation [2]

[0124] To solve for the TRF value for the triangle primitive 900A, Equation [1] can be used. The total outer TessFactors can be known (e.g., they can be the same as the TessFactors of the original primitive). Thus, T[x] and P[x] for all the outer edges can be computed.

[0125] The new interior TesFactor (TF[i]) can be TF_original[i] * TRF. The values for T[i] and P[i] for the interior edges can be found based on iterations with different TRF values.

[0126] The sub-primitive count can be found based on Equation [1]. The TRF that gives the sub-primitive count closest to N_vrt can be chosen.

[0127] If N_vrt is below or the same as the minimum possible value, the TRF can be set to 1. The minimum possible value for N_vrt can be the sub-primitive count when the interior TessFactor is set to 1 (i.e., TF[i] = 1).

[0128] To solve for the TRF value for the quad primitive 900B, Equation [2] can be used. The total outer TessFactors can be known (e.g., they can be the same as the TessFactors of the original primitive). Thus, T[x] and P[x] for all the outer edges can be computed.

[0129] The new internal TessFactor (TF[u]) can be TRF*TF_original[u] and TF[v] can be TRF*TF_original[v]. The values of T[u], T[v], P[u], and P[v] for the internal edges can be found based on iterations with different TRF values.

[0130] The subpatch count can be found based on equation [2]. The TRF that gives the subpatch count closest to N_vrt can be selected.

[0131] If N_vrt is below or the same as the minimum possible value, the TRF can be set to 1. The minimum possible value for N_vrt can be the subpatch count when the internal TessFactor is set to 1 (i.e., TF[u] = 1 and TF[v] = 1).

[0132] Figure 10 is a call flow diagram 1000 illustrating example communications between a first GPU component 1002 (e.g., a processing unit, a shading unit, a tessellator, etc.) and a second GPU component 1004 (e.g., a cache, a buffer, internal memory, etc.) in accordance with one or more techniques of the present disclosure. At 1010, the first GPU component 1002 can receive, from the second GPU component 1004, data 1012 for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more subprimitives.

[0133] At 1020, the first GPU component 1002 can receive, from the second GPU component 1004, a VRS input 1022 associated with a shading rate for each tile of the plurality of tiles. A tessellation factor for each tile of the plurality of tiles can be reduced based on the VRS input.

[0134] In one configuration, the VRS input can be draw-based. The TRF for each tile of the plurality of tiles can be the same TRF. In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0135] In one configuration, the VRS input can be primitive-based. A first TRF for at least one tile of the plurality of tiles can be different than a second TRF for at least one other tile of the plurality of tiles. In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0136] In one configuration, the VRS input can be image-based. At 1030, the first GPU component 1002 can create a TRF buffer to store a TRF for each tile of the plurality of tiles based on the VRS input 1022.

[0137] In one configuration, the TRF for each tile of the plurality of tiles can be applied in a rendering pass. In one configuration, the plurality of primitives can include a first primitive associated with a plurality of patches. A first TRF for the first primitive can correspond to a maximum second TRF for the plurality of patches.

[0138] At 1040, the first GPU component 1002 can reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The characteristic can correspond to a shading rate or a number of visible pixels.

[0139] At 1050, the first GPU component 1002 can apply the TRF for each tile of the plurality of tiles.

[0140] At 1060, the first GPU component 1002 can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles.

[0141] In one configuration, the tessellation factor for each tile of the plurality of tiles can be reduced such that a ratio of a number of effective pixels to a number of one or more sub-primitives remains constant for each tile of the plurality of tiles.

[0142] In one configuration, the TRF can be dynamically identified based on a ratio of a number of visible pixels to a number of one or more sub-primitives within each tile of the plurality of tiles. The TRF can be identified based on an iterative process.

[0143] In one configuration, one or more internal tessellation factors including the tessellation factor for each tile of the plurality of tiles can be reduced. One or more external tessellation factors can not be reduced.

[0144] Figure 11 FIG. 11 is a flow diagram 1100 of an example method of graphics processing in accordance with one or more techniques of this disclosure. The method can be performed by an apparatus, such as the apparatus used in aspects combined with the Figures 1-10 apparatus, GPU, CPU, wireless communication device, etc. for graphics processing used in aspects combined with the

[0145] At 1102, the apparatus can receive data for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more sub-primitives. Reference is made to Figure 10At 1010, the first GPU component 1002 can receive data 1012 for geometry processing of a plurality of tiles in a draw call. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1102.

[0146] At 1104, the apparatus can reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a TRF. The characteristic can correspond to a shading rate or a number of visible pixels. Reference is made to Figure 10 At 1040, the first GPU component 1002 can reduce a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1104.

[0147] At 1106, the apparatus can apply a TRF for each tile of the plurality of tiles. Reference is made to Figure 10 At 1050, the first GPU component 1002 can apply a TRF for each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1106.

[0148] At 1108, the apparatus can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles. Reference is made to Figure 10 At 1060, the first GPU component 1002 can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1108.

[0149] Figure 12 is a flow diagram 1200 of an example method of graphics processing in accordance with one or more techniques of the present disclosure. The method can be performed by an apparatus, such as the apparatus used in aspects of Figures 1-10 the apparatus, GPU, CPU, wireless communication device, etc. used for graphics processing in aspects of

[0150] At 1202, the apparatus can receive data for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more sub-primitives. Reference is made to Figure 10 At 1010, the first GPU component 1002 can receive data 1012 for geometry processing of a plurality of tiles in a draw call. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1202. Furthermore,Figure 6B Block 652 in FIG. 6 can perform operation 1202.

[0151] In one configuration, at 1204, the apparatus can receive a VRS input associated with a shading rate for each tile of a plurality of tiles. A tessellation factor for each tile of the plurality of tiles can be reduced based on the VRS input. Reference is made to Figure 10 At 1020, the first GPU component 1002 can receive a VRS input 1022 associated with a shading rate for each tile of a plurality of tiles. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1204. Further, Figure 6B Block 652 in FIG. 6 can perform operation 1204.

[0152] In one configuration, the VRS input can be rendering based. A TRF for each tile of the plurality of tiles can be a same TRF. In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0153] In one configuration, the VRS input can be primitive based. A first TRF for at least one tile of the plurality of tiles can be different from a second TRF for at least one other tile of the plurality of tiles. In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0154] In one configuration, the VRS input can be image based. At 1206, the apparatus can create a TRF buffer for storing a TRF for each tile of a plurality of tiles based on the VRS input. Reference is made to Figure 10 At 1030, the first GPU component 1002 can create a TRF buffer for storing a TRF for each tile of a plurality of tiles based on the VRS input 1022. Further, Figure 1 The processing unit 120 in FIG. 12 can perform operation 1206. Further, Figure 6B Block 662 in FIG. 6 can perform operation 1206.

[0155] In one configuration, the TRF for each tile of the plurality of tiles can be applied in a rendering pass. In one configuration, the plurality of primitives can include a first primitive associated with a plurality of patches. A first TRF for the first primitive can correspond to a maximum second TRF for the plurality of patches.

[0156] At 1208, the apparatus can reduce a tessellation factor for each tile of a plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a TRF. The characteristic can correspond to a shading rate or a number of visible pixels. Reference is made toFigure 10 At 1040, the first GPU component 1002 can reduce a tessellation factor for each tile of the plurality of tiles based on a shading rate or a number of visible pixels for each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 of the device 1200 can perform operation 1208. Moreover, Figure 6A The block 606 of the device 1200 and / or Figure 6B The block 658 of the device 1200 can perform operation 1208.

[0157] At 1210, the apparatus can apply a TRF for each tile of the plurality of tiles. Reference is made to Figure 10 At 1050, the first GPU component 1002 can apply a TRF for each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 of the device 1200 can perform operation 1210. Moreover, Figure 6B The block 658 of the device 1200 can perform operation 1210.

[0158] At 1212, the apparatus can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles. Reference is made to Figure 10 At 1060, the first GPU component 1002 can render each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles. Further, Figure 1 The processing unit 120 of the device 1200 can perform operation 1212. Moreover, Figure 6B The block 660 of the device 1200 can perform operation 1212.

[0159] In one configuration, the tessellation factor for each tile of the plurality of tiles can be reduced such that a ratio of a number of effective pixels to a number of one or more sub-primitives remains constant for each tile of the plurality of tiles.

[0160] In one configuration, the TRF can be dynamically identified based on a ratio of a number of visible pixels to a number of one or more sub-primitives within each tile of the plurality of tiles. The TRF can be identified based on an iterative process.

[0161] In one configuration, one or more internal tessellation factors including the tessellation factor for each tile of the plurality of tiles can be reduced. One or more external tessellation factors can not be reduced.

[0162] In various configurations, a method or apparatus for graphics processing is provided. The apparatus can be a GPU, a CPU, or some other processor that can perform graphics processing. In various aspects, the apparatus can be a processing unit 120 within a device 104, or can be some other hardware within the device 104 or another device. The apparatus can include means for receiving data for geometry processing of a plurality of tiles in a draw call. Each tile of the plurality of tiles can include a plurality of primitives. Each primitive of the plurality of primitives in each tile of the plurality of tiles can include one or more sub-primitives. The apparatus can also include means for reducing a tessellation factor for each tile of the plurality of tiles based on a characteristic of each tile of the plurality of tiles. The reduced tessellation factor can correspond to a TRF. The characteristic can correspond to a shading rate or a number of visible pixels. The apparatus can also include means for applying the TRF for each tile of the plurality of tiles. The apparatus can also include means for rendering each tile of the plurality of tiles based on the TRF applied for each tile of the plurality of tiles.

[0163] In one configuration, the tessellation factor for each tile of the plurality of tiles can be reduced such that a ratio of a number of effective pixels to a number of one or more sub-primitives remains constant for each tile of the plurality of tiles.

[0164] In one configuration, the TRF can be dynamically identified based on a ratio of a number of visible pixels to a number of one or more sub-primitives within each tile of the plurality of tiles. The TRF can be identified based on an iterative process.

[0165] In one configuration, the apparatus can also include means for receiving a VRS input associated with a shading rate for each tile of the plurality of tiles. The tessellation factor for each tile of the plurality of tiles can be reduced based on the VRS input.

[0166] In one configuration, the VRS input can be draw-based. The TRF for each tile of the plurality of tiles can be the same TRF.

[0167] In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0168] In one configuration, the VRS input can be primitive-based. A first TRF for at least one tile of the plurality of tiles can be different than a second TRF for at least one other tile of the plurality of tiles.

[0169] In one configuration, the TRF for each tile of the plurality of tiles can be applied in a binning pass or a rendering pass.

[0170] In one configuration, the VRS input can be image-based. The apparatus can also include means for creating a TRF buffer to store a TRF for each tile of the plurality of tiles based on the VRS input.

[0171] In one configuration, the TRF for each tile of the plurality of tiles can be applied in a rendering pass.

[0172] In one configuration, the plurality of primitives can include a first primitive associated with a plurality of patches. A first TRF for the first primitive can correspond to a maximum second TRF for the plurality of patches.

[0173] In one configuration, one or more internal tessellation factors including a tessellation factor for each tile of the plurality of tiles can be reduced. One or more external tessellation factors can not be reduced.

[0174] Referring back to Figures 6A-12 In a tessellation stage, TessFactor can be reduced in the case that the LOD is lowered based on shading rate and / or pixel count. There can be no noticeable decrease in the quality of the image even if TessFactor is significantly reduced. At the same time, GPU rendering performance can be improved.

[0175] It should be understood that the particular order or hierarchy of steps in processes, flow charts, and / or calling flows disclosed herein are merely examples. It should be appreciated that the particular order or hierarchy of steps in the processes, flow charts, and / or calling flows can be re-arranged based on design. Furthermore, some steps can be deleted while others can be added in other instances. Still further, some steps can be combined and / or divided into sub-steps. Additionally, some steps can be performed concurrently. The accompanying method claims present elements of the various steps in example order. The method claims are not meant to be limited to the particular order or hierarchy presented.

[0176] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein, but is to be accorded the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean "one and only one" unless specifically so stated, but rather "one or more." The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects. Unless specifically stated otherwise, it is appreciated that throughout this specification discussions utilizing terms such as "processing," "computing," "calculating," "determining," "displaying," and / or "generating" can refer to actions or processes of a machine that manipulates or transforms data, terms

[0177] The term "some" refers to one or more, and the term "or" can be construed in either an inclusive or exclusive sense, unless it is stated otherwise, or it is clear from the context that the term is used in an otherwise manner. Combinations of components described throughout this disclosure, such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof," include any combination of A, B, and / or C, and can include multiples of A, multiples of B, or multiples of C. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof," can be A only, B only, C only, A and B, A and C, B and C, or A and B and C, where any such combination can contain one or more members of A, B, or C. Structural and functional equivalents of all elements described throughout this disclosure that serve the same, equivalent, or similar purposes as those listed and any future technical equivalents of the elements described throughout this disclosure that serve the same, equivalent, or similar purposes as those listed are expressly incorporated herein by reference and intended to be encompassed by the claims. Furthermore, no embodiment disclosed herein is intended to be dedicated to the public regardless of whether these embodiments are explicitly recited in the claims. The words "module," "mechanism," "element," "device," and the like can not be exclusive of the term "unit." Therefore, no claim element is to be construed as a means plus function unless the element is expressly recited using the phrase "means for."

[0178] In one or more examples, the functions described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage media can be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or any other medium that is suitable for transmitting software, then the coaxial cable, fiber optic cable, twisted pair, DSL, or any other medium can be properly termed a computer-readable medium. Additionally, it should be appreciated that a computer-readable medium can be implemented in a computer program product, which can be executed on a computing device.

[0179] A computer readable medium can include a computer data storage medium or a communication medium including any medium that facilitates transfer of a computer program from one place to another. In this manner, computer readable medium can be any available medium or a combination thereof that can be accessed by a one or more computers or one or more processors to implement the techniques described herein. By way of example, and not limitation, such computer readable media can comprise RAM, ROM, EEPROM, compact disk (CD) ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or any other medium that is suitable for a transmitting or otherwise providing software, then the medium is properly termed a computer readable medium. Additionally, the disclosure also contemplates computer readable media that are not tangible, such as the behaviors of wireless systems to transmit software. Accordingly, computer readable medium, as used herein includes both computer storage media and communication media.

[0180] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC) or a set of ICs (e.g., a chip set). Various components, modules, or units are described in this disclosure to emphasize that implementations of the devices have functional aspects that do not necessarily require hardware

[0181] The following aspects are illustrative only and can be combined with other aspects or teachings described herein without limitation.

[0182] Aspect 1 is an apparatus for graphics processing, comprising: at least one processor coupled to a memory and configured to: receive data for geometry processing of a plurality of tiles in a draw call, each of the plurality of tiles comprising a plurality of primitives, each of the plurality of primitives in each of the plurality of tiles comprising one or more sub-primitives; reduce a tessellation factor for each of the plurality of tiles based on a characteristic of each of the plurality of tiles, the characteristic corresponding to a shading rate or a number of visible pixels, the reduced tessellation factor corresponding to a TRF; apply the TRF for each of the plurality of tiles; and render each of the plurality of tiles based on the TRF applied for each of the plurality of tiles.

[0183] Aspect 2 is the apparatus of Aspect 1, wherein the tessellation factor for each of the plurality of tiles is reduced such that a ratio of a number of effective pixels to the number of one or more sub-primitives remains constant for each of the plurality of tiles.

[0184] Aspect 3 is the apparatus of Aspect 1, wherein the TRF is dynamically identified based on a ratio of the number of visible pixels to the number of one or more sub-primitives within each of the plurality of tiles, and the TRF is identified based on an iterative process.

[0185] Aspect 4 is the apparatus of any one of Aspects 1 and 2, the at least one processor is further configured to: receive a VRS input associated with the shading rate for each of the plurality of tiles, wherein the tessellation factor for each of the plurality of tiles is reduced based on the VRS input.

[0186] Aspect 5 is the apparatus of Aspect 4, wherein the VRS input is draw-based, and the TRF for each of the plurality of tiles is a same TRF.

[0187] Aspect 6 is the apparatus of Aspect 5, wherein the TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

[0188] Aspect 7 is the apparatus of Aspect 4, wherein the VRS input is primitive-based, and a first TRF for at least one of the plurality of tiles is different from a second TRF for at least one other of the plurality of tiles.

[0189] Aspect 8 is the apparatus of Aspect 7, wherein the TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

[0190] Aspect 9 is the apparatus of Aspect 4, wherein the VRS input is image-based, and the at least one processor is further configured to create a TRF buffer for storing the TRF for each tile of the plurality of tiles based on the VRS input.

[0191] Aspect 10 is the apparatus of Aspect 9, wherein the TRF for each tile of the plurality of tiles is applied in a rendering pass.

[0192] Aspect 11 is the apparatus of any of Aspects 9 and 10, wherein the plurality of primitives includes a first primitive associated with a plurality of patches, and a first TRF for the first primitive corresponds to a maximum second TRF for the plurality of patches.

[0193] Aspect 12 is the apparatus of any of Aspects 1-11, wherein one or more interior tessellation factors including the tessellation factors for each tile of the plurality of tiles are reduced, and one or more exterior tessellation factors are not reduced.

[0194] Aspect 13 is the apparatus of any of Aspects 1-12, wherein the apparatus is a wireless communication device.

[0195] Aspect 14 is a method for wireless communication that implements any of Aspects 1-13.

[0196] Aspect 15 is an apparatus for graphics processing, comprising means for implementing a method as in any of Aspects 1-13.

[0197] Aspect 16 is a computer-readable medium storing computer executable code, where the code when executed by at least one processor causes the at least one processor to implement a method as in any of Aspects 1-13.

[0198] Various aspects have been described herein. These and other aspects are within the scope of the following claims.

Claims

1. An apparatus for graphics processing, comprising: a memory; and at least one processor coupled to the memory and configured to: receive data for tessellation of a plurality of tiles in a draw call, each of the plurality of tiles comprising a plurality of primitives, each of the plurality of primitives in each of the plurality of tiles comprising one or more sub-primitives; reduce a tessellation factor for each of the plurality of tiles based on a characteristic of each of the plurality of tiles, the reduced tessellation factor corresponding to a tessellation reduction factor (TRF), the characteristic corresponding to a shading rate or a number of visible pixels; apply the TRF for each of the plurality of tiles; and render each of the plurality of tiles based on the TRF applied for each of the plurality of tiles.

2. The apparatus of claim 1, wherein, The tessellation factor for each of the plurality of tiles is reduced such that a ratio of a number of effective pixels to a number of the one or more sub-primitives remains constant for each of the plurality of tiles.

3. The apparatus of claim 1, wherein, The TRF is dynamically identified based on a ratio of the number of visible pixels to the number of the one or more sub-primitives within each of the plurality of tiles, and the TRF is identified based on an iterative process.

4. The apparatus of claim 1, the at least one processor further configured to: receiving a variable rate shading (VRS) input associated with the shading rate of each tile of the plurality of tiles, wherein, The tessellation factor for each of the plurality of tiles is reduced based on the VRS input.

5. The apparatus of claim 4, wherein, The VRS input is based on a draw granularity, and the TRF for each of the plurality of tiles is a same TRF.

6. The apparatus of claim 5, wherein, The TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

7. The apparatus of claim 4, wherein, The VRS input is based on a primitive granularity, and a first TRF for at least one of the plurality of tiles is different than a second TRF for at least one other of the plurality of tiles.

8. The apparatus of claim 7, wherein, The TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

9. The apparatus of claim 4, wherein, The VRS input is based on an image granularity, and the at least one processor is further configured to: create a TRF buffer for storing the TRF for each of the plurality of tiles based on the VRS input.

10. The apparatus of claim 9, wherein, The TRF for each of the plurality of tiles is applied in a rendering pass.

11. The apparatus of claim 9, wherein, The plurality of primitives comprises a first primitive associated with a plurality of patches, and a first TRF for the first primitive corresponds to a maximum second TRF for the plurality of patches.

12. The apparatus of claim 1, wherein, One or more internal tessellation factors comprising the tessellation factor for each of the plurality of tiles are reduced, and one or more external tessellation factors are not reduced.

13. The apparatus of claim 1, wherein, The apparatus is a wireless communication device.

14. A method of graphics processing, comprising: receiving data for tessellation processing of a plurality of tiles in a draw call, each of the plurality of tiles comprising a plurality of primitives, each of the plurality of primitives in each of the plurality of tiles comprising one or more sub-primitives; reducing a tessellation factor for each of the plurality of tiles based on a characteristic of each of the plurality of tiles, the characteristic corresponding to a shading rate or a number of visible pixels, the reduced tessellation factor corresponding to a tessellation reduction factor (TRF); applying the TRF for each of the plurality of tiles; and rendering each of the plurality of tiles based on the TRF applied for each of the plurality of tiles.

15. The method of claim 14, wherein, The tessellation factor for each of the plurality of tiles is reduced such that a ratio of a number of effective pixels to a number of the one or more sub-primitives remains constant for each of the plurality of tiles.

16. The method of claim 14, wherein, The TRF is dynamically identified based on a ratio of the number of visible pixels to the number of the one or more sub-primitives within each of the plurality of tiles, and the TRF is identified based on an iterative process.

17. The method of claim 14, further comprising: receiving a variable rate shading (VRS) input associated with the shading rate for each of the plurality of tiles, wherein the tessellation factor for each of the plurality of tiles is reduced based on the VRS input.

18. The method of claim 17, wherein, The VRS input is based on a draw granularity, and the TRF for each of the plurality of tiles is a same TRF.

19. The method of claim 18, wherein, The TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

20. The method of claim 17, wherein, The VRS input is based on a primitive granularity, and a first TRF for at least one of the plurality of tiles is different than a second TRF for at least one other of the plurality of tiles.

21. The method of claim 20, wherein, The TRF for each of the plurality of tiles is applied in a binning pass or a rendering pass.

22. The method of claim 17, wherein, The VRS input is based on an image granularity, and the method further comprises: creating a TRF buffer for storing the TRF for each of the plurality of tiles based on the VRS input.

23. The method of claim 22, wherein, The TRF for each of the plurality of tiles is applied in a rendering pass.

24. The method of claim 22, wherein, The plurality of primitives comprises a first primitive associated with a plurality of patches, and a first TRF for the first primitive corresponds to a maximum second TRF for the plurality of patches.

25. The method of claim 14, wherein, One or more inner tessellation factors comprising the tessellation factor for each of the plurality of tiles are reduced, and one or more outer tessellation factors are not reduced.

26. A non-transitory computer-readable medium storing computer-executable code that, when executed by at least one processor, causes the at least one processor to perform: receiving data for tessellation processing of a plurality of tiles in a draw call, each of the plurality of tiles comprising a plurality of primitives, each of the plurality of primitives in each of the plurality of tiles comprising one or more sub-primitives; reducing a tessellation factor for each of the plurality of tiles based on a characteristic of each of the plurality of tiles, the characteristic corresponding to a shading rate or a number of visible pixels, the reduced tessellation factor corresponding to a tessellation reduction factor (TRF); applying the TRF for each of the plurality of tiles; and rendering each of the plurality of tiles based on the TRF applied for each of the plurality of tiles.

27. The non-transitory computer-readable medium of claim 26, wherein, the tessellation factor for each of the plurality of tiles is reduced such that a ratio of a number of effective pixels to a number of the one or more sub-primitives remains constant for each of the plurality of tiles.

28. The non-transitory computer-readable medium of claim 26, wherein, the TRF is dynamically identified based on a ratio of the number of visible pixels to the number of the one or more sub-primitives within each of the plurality of tiles, and the TRF is identified based on an iterative process.

29. The non-transitory computer-readable medium of claim 26, the code further causing the at least one processor to: receiving a variable rate shading (VRS) input associated with the shading rate of each tile of the plurality of tiles, wherein, the tessellation factor for each of the plurality of tiles is reduced based on the VRS input.

30. The non-transitory computer-readable medium of claim 29, wherein, the VRS input is based on a draw granularity, and the TRF for each of the plurality of tiles is a same TRF.

Citation Information

Patent Citations

  • Variable rate shading

    CN110383337A

  • Graphics processing units and methods for controlling rendering complexity using cost indications for sets of tiles of a rendering space

    EP3349181A1

  • Adaptive sub-patches system, apparatus and method

    EP3396636A1

  • System And Method For Tessellation In An Improved Graphics Pipeline

    US20170358132A1