Improving visibility generation in tile-based GPU architectures
Through the two-stage binning process, the problem of low visibility testing efficiency in the existing technology is solved, and more efficient visibility generation and rendering performance improvement is achieved.
Patent Information
- Application Number
- CN202380066293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-23
- Filing Date
- 2023-09-15
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has low visibility testing efficiency during the rendering process, which leads to work waste, especially when the order dependence of the elements is strong.
Using a two-stage binning process, the first stage updates the depth buffer through vertex shading and visibility tests, and records the visibility information of the primitives. The second stage performs further visibility tests and recordings on each primitive in the set of visible primitives based on the updated depth buffer.
Improves the visibility generation efficiency during binning, reduces the workload of handling invisible primitives during rendering, thereby improving overall rendering performance.
Smart Images

Figure CN119998842A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. patent application serial number 17 / 935,031, filed on September 23, 2022, and entitled “IMPROVING VISIBILITYGENERATION IN TILE BASED GPU ARCHITECTURES,” which is expressly incorporated herein by reference in its entirety. Technical Field
[0003] The present disclosure relates generally to processing systems and, more particularly, to one or more techniques for graphics processing. Background Art
[0004] Computing devices typically perform graphics and / or display processing (e.g., using a graphics processing unit (GPU), a central processing unit (CPU), a display processor, etc.) to render and display visual content. Such computing devices may include, for example, computer workstations, mobile phones (such as smart phones), embedded systems, personal computers, tablet computers, and video game consoles. The GPU is configured to execute a graphics processing pipeline, which includes one or more processing stages that operate together to execute graphics processing commands and output frames. The central processing unit (CPU) can control the operation of the GPU by issuing one or more graphics processing commands to the GPU. Modern CPUs are typically capable of executing multiple applications concurrently, and each application may need to utilize a GPU during execution. The display processor may be configured to convert digital information received from the CPU into analog values, and may issue commands to a display panel to display visual content. A device that provides content for visual presentation on a display may utilize a CPU, a GPU, and / or a display processor.
[0005] Current graphics techniques may not account for wasted work in the rendering process when the results of the binning process may depend on the order in which primitives enter the pipeline and visibility testing may not remove all non-visible primitives. Improved visibility testing techniques are needed. Summary of the invention
[0006] The following presents a simplified summary of one or more aspects in order to provide a basic understanding of these aspects. This summary is not an extensive overview of all contemplated aspects, and is neither intended to identify key or important elements of all aspects, nor to describe the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to a more detailed description presented later.
[0007] In one aspect of the present disclosure, a method, a computer-readable medium, and an apparatus are provided. The apparatus may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set. The apparatus may update a depth buffer based on the visibility information of all of the plurality of primitives in the at least one frame. The apparatus may perform a second binning process for each of the primitives in the visible primitive set based on the updated depth buffer. The second binning process may be associated with updated visibility information of each of the primitives in the visible primitive set. The updated visibility information of each of the primitives in the visible primitive set may correspond to an updated visible indication or an updated invisible indication. The apparatus may store at least one of updated visibility information or updated position data of all of the primitives in the visible primitive set from the second binning process.
[0008] To achieve the foregoing and related ends, one or more aspects include the features fully described below and particularly pointed out in the claims. The following description and the accompanying drawings set forth in detail some illustrative features of one or more aspects. However, these features are merely indicative of some of the various ways in which the principles of the various aspects may be employed, and this specification is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is a block diagram illustrating an example content generation system in accordance with one or more techniques of this disclosure.
[0010] Figure 2 An example GPU in accordance with one or more techniques of this disclosure is illustrated.
[0011] Figure 3 Example images or surfaces are illustrated in accordance with one or more techniques of this disclosure.
[0012] Figure 4 is a diagram illustrating an example data flow associated with a procedural binning technique.
[0013] Figure 5 is a diagram illustrating an example data flow associated with a two-pass binning technique in accordance with one or more aspects.
[0014] Figure 6 is an example flow chart illustrating a first binning process in a two-process binning according to one or more aspects.
[0015] Figure 7is an example flow chart illustrating a second binning process in a two-process binning according to one or more aspects.
[0016] Figure 8 is a diagram illustrating an example data flow associated with a two-pass binning technique using on-chip storage according to one or more aspects.
[0017] Fig. 9 is an example flow chart illustrating a two-pass binning technique using on-chip storage in accordance with one or more aspects.
[0018] Fig.10 is a diagram of an example flow chart illustrating a first binning process in a two-pass binning method utilizing visibility information of previous frames according to one or more aspects.
[0019] Fig.11 is a diagram of an example flow chart illustrating a second binning process in a two-pass binning method utilizing visibility information of previous frames according to one or more aspects.
[0020] Fig.12 is an illustration of an example flow chart illustrating a bin rendering process in which a two-pass binning approach utilizing visibility information of previous frames is implemented according to one or more aspects.
[0021] Fig.13 is a call flow diagram illustrating example communications between a GPU and a storage device in accordance with one or more techniques of this disclosure.
[0022] Fig.14 is a flowchart of an example method of graphics processing in accordance with one or more techniques of this disclosure.
[0023] Fig.15 is a flowchart of an example method of graphics processing in accordance with one or more techniques of this disclosure. DETAILED DESCRIPTION
[0024] Various aspects of the system, device, computer program product and method will be described more fully below with reference to the accompanying drawings. However, the present disclosure can be embodied in many different forms and should not be interpreted as being limited to any specific structure or function presented throughout the present disclosure. On the contrary, these aspects are provided so that the present disclosure will be thorough and complete, and the scope of the present disclosure will be fully conveyed to those skilled in the art. Based on the teachings of this article, it should be understood by those skilled in the art that the scope of the present disclosure is intended to cover any aspect of the system, device, computer program product and method disclosed herein, whether it is implemented independently of other aspects of the present disclosure or implemented in combination with other aspects of the present disclosure. For example, any number of aspects set forth herein can be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods implemented using other structures, functionality, or structures and functionality other than the various aspects of the disclosure set forth herein or different from the various aspects of the disclosure set forth herein. Any aspect disclosed herein can be embodied by one or more elements of the claims.
[0025] Although various aspects are described herein, many variations and permutations of these aspects fall within the scope of the present disclosure. Although some potential benefits and advantages of various aspects of the present disclosure are mentioned, the scope of the present disclosure is not intended to be limited to specific benefits, uses or objectives. On the contrary, various aspects of the present disclosure are intended to be widely applicable to different wireless technologies, system configurations, processing systems, networks and transmission protocols, some of which are illustrated by way of example in the drawings and the description below. The specific embodiments and drawings merely illustrate the present disclosure and do not limit the present disclosure, and the scope of the present disclosure is defined by the attached claims and their equivalent technical solutions.
[0026] Several aspects are presented with reference to various devices and methods. These devices and methods are described in the following detailed description and illustrated in the accompanying drawings by various blocks, components, circuits, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, computer software, or any combination thereof. Whether these elements are implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system.
[0027] For example, an element or any part of an element or any combination of elements may be implemented as a "processing system" including one or more processors (which may also be referred to as processing units). Examples of processors include microprocessors, microcontrollers, graphics processing units (GPUs), general purpose GPUs (GPGPUs), central processing units (CPUs), application processors, digital signal processors (DSPs), reduced instruction set computing (RISC) processors, systems on chips (SOCs), baseband processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functionalities described in the present disclosure. One or more processors in a processing system may execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or other names, software can be broadly understood to mean instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software components, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, processes, functions, etc.
[0028] The term "application" may refer to software. As described herein, one or more technologies may refer to an application (e.g., software) configured to perform one or more functions. In such examples, the application may be stored in a memory (e.g., on-chip memory of a processor, system memory, or any other memory). The hardware described herein (such as a processor) may be configured to execute the application. For example, an application may be described as including code that, when executed by hardware, causes the hardware to perform one or more technologies described herein. As an example, the hardware may access code from the memory and execute the code accessed from the memory to perform one or more technologies described herein. In some examples, components are identified in the present disclosure. In such examples, the components may be hardware, software, or a combination thereof. Each component may be a separate component or a subcomponent of a single component.
[0029] In one or more examples described herein, the functions described can be implemented in hardware, software, or any combination thereof. If implemented in software, the functions can be stored or encoded on a computer-readable medium as one or more instructions or codes. Computer-readable media include computer storage media. Storage media can be any available media that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), optical disk storage, magnetic disk storage, other magnetic storage devices, a combination of computer-readable media of the above types, or any other medium that can be used to store computer executable code in the form of instructions or data structures that can be accessed by a computer.
[0030] As used herein, instances of the term "content" may refer to "graphic content," "images," and the like, regardless of whether the term is used as an adjective, noun, or other part of speech. In some examples, as used herein, the term "graphic content" may refer to content produced by one or more processes of a graphics processing pipeline. In other examples, as used herein, the term "graphic content" may refer to content produced by a processing unit configured to perform graphics processing. In yet other examples, as used herein, the term "graphic content" may refer to content produced by a graphics processing unit.
[0031] A tile-based architecture may be used for GPUs to save bandwidth and improve performance and power efficiency. A typical tile-based architecture may involve two processes: 1) a bin visibility process (hereinafter also referred to as a binning process), which includes streaming geometry positions (e.g., vertices) into the GPU, where triangle horizontal visibility tests may be performed and a low-resolution Z (LRZ) buffer and a visibility buffer may be prepared, and then triangles may be sorted into various bins with visibility information recorded, and 2) a bin rendering process, where visibility information may be used to render multiple bins by fetching one screen-space bin at a time.
[0032] The binning process may include performing a visibility test on a primitive (e.g., a triangle). An example of a visibility test may include a view frustrum test that checks whether a primitive is completely outside a viewport (a visible area expressed in rendering device specific coordinates). A further example of a visibility test may include a backface culling test that checks whether a primitive is back-facing so that it can be discarded. Many other tests may also be used to check whether a primitive generates any samples. In some configurations, a coarse or detailed Z-buffer-based test may be performed to discard a primitive (e.g., a triangle) if the primitive happens to be behind an already rendered primitive (Z-buffer-based test may also be referred to as hidden surface removal). A procedural binning architecture may be associated with various advantages. These advantages may include, for example, minimal write bandwidth for the binning process, concurrent pipelines for binning and rendering, simple visibility buffer management, and easy switching between direct mode and binning mode.
[0033] Hidden surface removal techniques may be particularly useful for generating better visibility in the binning process, since any primitives behind the visible screen area may be rejected (e.g., removed) early on. Thus, some operations that would otherwise be performed in the bin rendering process may be avoided without any negative consequences. Although these techniques may reject primitives that are otherwise hidden, they may depend heavily on the order in which the primitives are received into the GPU pipeline (e.g., front-to-back rendering or back-to-front rendering). Many workloads may include a significant proportion of primitives that, due to this limitation, are rejected in the bin rendering process (e.g., using the Z-buffer created in the binning process) rather than in the binning process (because the primitives are not actually visible). Failure to reject these primitives early in the binning process may result in wasted work in the bin rendering process (which may be associated with wasted vertex shaders or vertex bandwidth). Further, in order to scale a one-pass binning architecture, concurrency, geometry pipelines, and pixel pipelines may be considered. The one-pass binning architecture may also be associated with high software complexity.
[0034] Aspects of the present disclosure may relate to architectural techniques for improving visibility generation in the binning process so that more triangles hidden behind visible surfaces may be rejected in the binning process regardless of the order in which the triangles are received into the GPU.
[0035] In some configurations, a two-pass approach in binning may be used, where in a first pass, geometry positions may be streamed into the GPU. Vertex shading and visibility testing may then be performed on the triangles, including coarse (low resolution Z (LRZ)) / fine depth buffer updates. During the first pass, visibility of the triangles may be recorded. Further, a second pass in binning may be performed, where visibility information from the first pass may be used and visible triangles as determined in the first pass may be accessed (in other words, invisible triangles as determined in the first pass may not be accessed in the second pass). In the second pass, visible triangles from the first pass may be tested again against the already existing coarse / fine depth buffers. Any triangles that are completely behind the existing depth buffer may be rejected. During the second pass, visibility recording may be performed so that updated visibility data that removes hidden triangles may be made available to the bin rendering pass.
[0036] After the second binning process, the bin rendering process may remain the same (e.g., Figure 4 ), where a bin rendering pass may read the updated visibility data and render triangles that are considered visible after the second binning pass.
[0037] Figure 11 is a block diagram illustrating an example content generation system 100 configured to implement one or more techniques of the present disclosure. The content generation system 100 includes a device 104. The device 104 may include one or more components or circuits for performing various functions described herein. In some examples, one or more components of the device 104 may be components of a SOC. The device 104 may include one or more components configured to perform one or more techniques of the present disclosure. In the example shown, the device 104 may include a processing unit 120, a content encoder / decoder 122, and a system memory 124. In some aspects, the device 104 may include multiple components (e.g., a communication interface 126, a transceiver 132, a receiver 128, a transmitter 130, a display processor 127, and one or more displays 131). The display 131 may refer to one or more displays 131. For example, the display 131 may include a single display or multiple displays, which may include a first display and a second display. The first display may be a left-eye display, and the second display may be a right-eye display. In some examples, the first display and the second display may receive different frames for presentation thereon. In other examples, the first display and the second display may receive the same frame for presentation thereon. In another example, the result of the graphics processing may not be displayed on the device, for example, the first display and the second display may not receive any frame for presentation thereon. Instead, the frame or graphics processing result may be passed to another device. In some aspects, this situation is referred to as split rendering.
[0038] The processing unit 120 may include an internal memory 121. The processing unit 120 may be configured to perform graphics processing using the graphics processing pipeline 107. The content encoder / decoder 122 may include an internal memory 123. In some examples, the device 104 may include a processor that may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120 and then display the frames through one or more displays 131. Although the processor in the example content generation system 100 is configured as a display processor 127, it should be understood that the display processor 127 is an example of a processor and other types of processors, controllers, etc. may be used in place of the display processor 127. The display processor 127 may be configured to perform display processing. For example, the display processor 127 may be configured to perform one or more display processing techniques on one or more frames generated by the processing unit 120. The one or more displays 131 may be configured to display or otherwise present the frames processed by the display processor 127. In some examples, one or more displays 131 may include one or more of the following: a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, a projection display device, an augmented reality display device, a virtual reality display device, a head-mounted display, or any other type of display device.
[0039] Memory external to processing unit 120 and content encoder / decoder 122, such as system memory 124, may be accessible to processing unit 120 and content encoder / decoder 122. For example, processing unit 120 and content encoder / decoder 122 may be configured to read from and / or write to external memory, such as system memory 124. Processing unit 120 may be communicatively coupled to system memory 124 via a bus. In some examples, processing unit 120 and content encoder / decoder 122 may be communicatively coupled to internal memory 121 via the bus or via a different connection.
[0040] The content encoder / decoder 122 may be configured to receive graphics content from any source, such as the system memory 124 and / or the communication interface 126. The system memory 124 may be configured to store the received encoded or decoded graphics content. The content encoder / decoder 122 may be configured to receive the encoded or decoded graphics content in the form of encoded pixel data, for example, from the system memory 124 and / or the communication interface 126. The content encoder / decoder 122 may be configured to encode or decode any graphics content.
[0041] The internal memory 121 or the system memory 124 may include one or more volatile or non-volatile memories or storage devices. In some examples, the internal memory 121 or the system memory 124 may include RAM, static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable ROM (EPROM), EEPROM, flash memory, magnetic data medium or optical storage medium or any other type of memory. According to some examples, the internal memory 121 or the system memory 124 may be a non-transient storage medium. The term "non-transient" may indicate that the storage medium is not embodied in a carrier or propagation signal. However, the term "non-transient" should not be interpreted as meaning that the internal memory 121 or the system memory 124 is non-removable or that its content is static. For example, the system memory 124 may be removed from the device 104 and moved to another device. As another example, the system memory 124 may not be removable from the device 104.
[0042] The processing unit 120 may be a CPU, a GPU, a GPGPU, or any other processing unit that may be configured to perform graphics processing. In some examples, the processing unit 120 may be integrated into the motherboard of the device 104. In another example, the processing unit 120 may be present on a graphics card installed in a port of the motherboard of the device 104, or may be otherwise incorporated into a peripheral device configured to interoperate with the device 104. The processing unit 120 may include one or more processors, such as one or more microprocessors, GPUs, ASICs, FPGAs, arithmetic logic units (ALUs), DSPs, discrete logic pieces, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, the processing unit 120 may store instructions for the software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 121), and may use one or more processors to execute instructions in hardware to perform the technology of the present disclosure. Any of the above (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors.
[0043] The content encoder / decoder 122 may be any processing unit configured to perform content decoding. In some examples, the content encoder / decoder 122 may be integrated into the mainboard of the device 104. The content encoder / decoder 122 may include one or more processors, such as one or more microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), arithmetic logic units (ALUs), digital signal processors (DSPs), video processors, discrete logic, software, hardware, firmware, other equivalent integrated or discrete logic circuits, or any combination thereof. If the technology is partially implemented in software, the content encoder / decoder 122 may store instructions for the software in a suitable non-transitory computer-readable storage medium (e.g., internal memory 123), and may use one or more processors to execute instructions in hardware to perform the technology of the present disclosure. Any of the above (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors.
[0044] In some aspects, the content generation system 100 may include a communication interface 126. The communication interface 126 may include a receiver 128 and a transmitter 130. The receiver 128 may be configured to perform any receiving functions described herein with respect to the device 104. In addition, the receiver 128 may be configured to receive information from another device, such as eye or head position information, rendering commands, and / or position information. The transmitter 130 may be configured to perform any sending functions described herein with respect to the device 104. For example, the transmitter 130 may be configured to send information to another device, which may include a request for content. The receiver 128 and the transmitter 130 may be combined into a transceiver 132. In such examples, the transceiver 132 may be configured to perform any receiving functions and / or sending functions described herein with respect to the device 104.
[0045] See again Figure 1In some aspects, the processing unit 120 may include a graphics element 198 configured to perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set. The graphics element 198 may be configured to update a depth buffer based on the visibility information of all of the plurality of primitives in the at least one frame. The graphics element 198 may be configured to perform a second binning process on each of the primitives in the visible primitive set based on the updated depth buffer. The second binning process may be associated with updated visibility information of each of the primitives in the visible primitive set. The updated visibility information of each of the primitives in the visible primitive set may correspond to an updated visible indication or an updated invisible indication. Graphics element 198 may be configured to store at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process.Although the following description may focus on graphics processing, the concepts described herein may be applicable to other similar processing techniques.
[0046] Devices such as device 104 may refer to any device, apparatus or system configured to perform one or more technologies described herein. For example, a device may be a server, a base station, a user equipment, a client device, a station, an access point, a computer (such as a personal computer, a desktop computer, a laptop computer, a tablet computer, a computer workstation or a mainframe computer), a final product, an apparatus, a phone, a smart phone, a server, a video game platform or console, a handheld device (such as a portable video game device or a personal digital assistant (PDA)), a wearable computing device (such as a smart watch, an augmented reality device or a virtual reality device), a non-wearable device, a display or display device, a television, a television set-top box, an intermediate network device, a digital media player, a video streaming device, a content streaming device, a vehicle-mounted computer, any mobile device, any device configured to generate graphics content, or any device configured to perform one or more technologies described herein. The process herein may be described as being performed by a specific component (e.g., a GPU), but in other embodiments, other components (e.g., a CPU) that conform to the disclosed embodiments may be used to perform.
[0047] The GPU can process multiple types of data or data packets in the GPU pipeline. For example, in some aspects, the GPU can process two types of data or data packets, such as context register packets and draw call data. The context register packet can be a set of global state information, such as information about global registers, shaders, or constant data, which can adjust how the graphics context will be processed. For example, the context register packet can include information about the color format. In some aspects of the context register packet, there can be a bit indicating which workload belongs to the context register. In addition, multiple functions or programs can be run simultaneously and / or in parallel. For example, a function or program can describe an operation, such as a color mode or color format. Therefore, the context register can define multiple states of the GPU.
[0048] The context state may be used to determine how a single processing unit (e.g., a vertex fetcher (VFD), a vertex shader (VS), a shader processor, or a geometry processor) operates and / or in which mode the processing unit operates. To do this, the GPU may use context registers and programming data. In some aspects, the GPU may generate workloads, such as vertex or pixel workloads, in a pipeline based on context register definitions of a mode or state. Certain processing units (e.g., VFDs) may use these states to determine certain functions, such as how to aggregate vertices. Since these modes or states may change, the GPU may need to change the corresponding context. In addition, the workload corresponding to the mode or state may follow the changed mode or state.
[0049] Figure 2 2 illustrates an example GPU 200 in accordance with one or more techniques of this disclosure. Figure 2 As shown, GPU 200 includes a command processor (CP) 210, a draw call group 212, a VFD 220, a VS 222, a vertex cache (VPC) 224, a triangle setup engine (TSE) 226, a rasterizer (RAS) 228, a Z-pass engine (ZPE) 230, a pixel interpolator (PI) 232, a fragment shader (FS) 234, a rendering back end (RB) 236, an L2 cache (UCHE) 238, and a system memory 240. Although Figure 2 GPU 200 is shown to include processing units 220 to 238, but GPU 200 may include multiple additional processing units. In addition, processing units 220 to 238 are only examples, and any combination or order of processing units may be used by the GPU according to the present disclosure. GPU 200 also includes command buffer 250, context register group 260, and context state 261.
[0050] like Figure 2As shown, the GPU may utilize a CP (e.g., CP 210) or a hardware accelerator to parse the command buffer into context register packets (e.g., context register packets 260) and / or draw call data packets (e.g., draw call packets 212). The CP 210 may then transmit the context register packets 260 or the draw call data packets 212 to a processing unit or block in the GPU via a separate path. Further, the command buffer 250 may alternate different states of context registers and draw calls. For example, the command buffer may be constructed in the following manner: context registers of context N, draw calls of context N, context registers of context N+1, and draw calls of context N+1.
[0051] The GPU can render images in a variety of different ways. In some instances, the GPU can use rendering and / or tile rendering to render images. In a tile rendering GPU, the image can be divided or separated into different parts or tiles. After the image is divided, each part or tile can be rendered separately. The tile rendering GPU can divide the computer graphics image into a grid format so that each part of the grid (i.e., tile) is rendered separately. In some aspects, during the binning process (binning pass), the image can be divided into different boxes or tiles. In some aspects, during the binning process, a visibility stream can be constructed in which visible primitives or drawing calls can be identified. In contrast to tile rendering, direct rendering does not divide the frame into smaller boxes or tiles. On the contrary, in direct rendering, the entire frame is rendered at once. Additionally, some types of GPUs may allow both tile rendering and direct rendering (e.g., flexible rendering).
[0052] In some aspects, the GPU may apply a drawing or rendering process to different boxes or tiles. For example, the GPU may render for a box and perform all drawing for the primitives or pixels in the box. During the process of rendering for the box, the rendering target may be located in the GPU internal memory (GMEM). In some instances, after rendering for a box, the contents of the rendering target may be moved to the system memory, and the GMEM may be released to render the next box. Additionally, the GPU may render for another box and perform drawing for the primitives or pixels in the box. Therefore, in some aspects, there may be a small number of boxes (e.g., four boxes) that cover all the drawing in a surface. In addition, the GPU may loop through all the drawing in a box, but perform drawing of visible drawing calls, i.e., drawing calls containing visible geometry. In some aspects, a visibility stream may be generated, for example, in a binning process, to determine visibility information for each primitive in an image or scene. For example, such a visibility stream may identify whether a primitive is visible. In some aspects, this information may be used to remove invisible primitives, for example, during a rendering process. In addition, at least some of the primitives identified as visible may be rendered during a rendering process.
[0053] In some aspects of tile rendering, there may be multiple processing stages or processes. For example, rendering may be performed in two processes, such as a visibility or box visibility process and a rendering or box rendering process. During the visibility process, the GPU may input a rendering workload, record the position of primitives or triangles, and then determine which primitives or triangles fall into which box or region. In certain aspects of the visibility process, the GPU may also identify or mark the visibility of each primitive or triangle in the visibility stream. During the rendering process, the GPU may input a visibility stream and process one box or region at a time. In some aspects, the visibility stream may be analyzed to determine which primitives or primitive vertices are visible or invisible. Therefore, visible primitives or primitive vertices may be processed. By doing so, the GPU may reduce unnecessary workloads for processing or rendering invisible primitives or triangles.
[0054] In some aspects, during the visibility process, certain types of primitive geometry may be processed, such as only positioned geometry. Additionally, primitives may be classified into different bins or regions based on the position or positioning of primitives or triangles. In some instances, the classification of primitives or triangles into different bins may be performed by determining visibility information for these primitives or triangles. For example, the GPU may determine visibility information for each primitive in each bin or region or write it to, for example, system memory. The visibility information may be used to determine or generate a visibility stream. During the rendering process, the primitives in each bin may be rendered separately. In these cases, the visibility stream may be extracted from a memory for discarding primitives that are not visible to the bin.
[0055] Some aspects of the GPU or GPU architecture may provide a number of different options for rendering (e.g., software rendering and hardware rendering). In software rendering, the driver or CPU may process each view. Figure 1 In hardware rendering, the hardware or GPU may be responsible for copying or processing the geometry of each viewpoint in the image. Therefore, the hardware can manage the copying or processing of primitives or triangles for each viewpoint in the image.
[0056] Figure 3 An image or surface 300 is illustrated, including a plurality of primitives divided into a plurality of bins, in accordance with one or more techniques of this disclosure. Figure 3 As shown, image or surface 300 includes region 302, which includes primitives 321, 322, 323, and 324. Primitives 321, 322, 323, and 324 are divided or placed into different bins, such as bins 310, 311, 312, 313, 314, and 315. Figure 3 An example of using multiple viewpoints for sub-tile rendering for primitives 321-324 is illustrated. For example, primitives 321-324 are in a first viewpoint 350 and a second viewpoint 351. Therefore, GPU processing or rendering of an image or surface 300 including region 302 may utilize multi-viewpoint or multi-view rendering.
[0057] As indicated herein, a GPU or graphics processor unit may use a tile rendering architecture to reduce power consumption or save memory bandwidth. As further described above, this rendering method may divide a scene into a plurality of bins, and include a visibility pass that identifies visible triangles in each bin. Thus, in tile rendering, a full screen may be divided into a plurality of bins or tiles. The scene may then be rendered multiple times, for example, once or multiple times for each bin.
[0058] In various aspects of graphics rendering, some graphics applications may render to a single target (i.e., a render target) one or more times. For example, in graphics rendering, a frame buffer on system memory may be updated multiple times. A frame buffer may be a portion of a memory or random access memory (RAM) (e.g., containing a bitmap or storage) to help store display data for a GPU. A frame buffer may also be a memory buffer containing a complete frame of data. Additionally, a frame buffer may be a logical buffer. In some aspects, updating a frame buffer may be performed in box or tile rendering, where, as described above, a surface is divided into multiple boxes or tiles, and then each box or tile may be rendered separately. Furthermore, in tile rendering, a frame buffer may be divided into multiple boxes or tiles.
[0059] As noted herein, in some aspects, such as in a bin or tile rendering architecture, frame buffers may have data stored or written to them repeatedly, for example, when rendering from different types of memory. This may be referred to as resolving and unresolving frame buffers or system memory. For example, when storing or writing to one frame buffer and then switching to another frame buffer, data or information on the frame buffer may be resolved from the GMEM at the GPU to system memory, i.e., memory in double data rate (DDR) RAM or dynamic RAM (DRAM).
[0060] In some aspects, system memory may also be a system on chip (SoC) memory or another chip-based memory for storing data or information, such as on a device or smartphone. System memory may also be a physical data storage shared by the CPU and / or GPU. In some aspects, system memory may be a DRAM chip, such as on a device or smartphone. Thus, SoC memory may be a chip-based way to store data.
[0061] In some aspects, GMEM can be an on-chip memory at the GPU, which can be implemented by static RAM (SRAM). In addition, GMEM can be stored on a device (e.g., a smart phone). As indicated herein, data or information can be transferred between system memory or DRAM and GMEM, for example, at the device. In some aspects, the system memory or DRAM can be located at the CPU or GPU. In addition, the data can be stored in DDR or DRAM. In some aspects, such as in box or tile rendering, a small portion of the memory can be stored at the GPU, for example, in GMEM. In some cases, storing data at GMEM may consume a larger processing workload and / or power consumption than storing data at a frame buffer or system memory.
[0062] Figure 44 is a diagram illustrating an example data flow 400 associated with a procedural binning technique. The operations of a binning pipeline 430 and a rendering pipeline 432 may be performed at a GPU (e.g., processing unit 120). In binning pipeline 430, position pipeline 402 may read positions 412 associated with triangles. Position pipeline 402 may then perform a visibility check 404 (e.g., a visibility test as described above) to determine whether each of the triangles is visible. The result of visibility check 404 may be visibility data / LRZ buffer 414. Positions 412 and visibility data / LRZ buffer 414 may be stored in memory 450 (e.g., DDR / main memory).
[0063] In the rendering pipeline 432, the GPU may render bins. Specifically, the visibility data / LRZ buffer 416 and the position / attribute 418, which may be stored in the memory 450, may correspond to the visibility data / LRZ buffer 414 and the position 412 in the binning pipeline 430. The geometry pipeline, which may be involved in vertex shading and / or per-primitive shading and culling, may read the visibility data / LRZ buffer 416 and the position / attribute 418. The output of the geometry pipeline 406 may be fed into the pixel pipeline. The pixel pipeline 408 may perform per-pixel tests and operations (e.g., depth testing, pixel blending, etc.) as well as programmable shading and lighting before outputting the final color. The pixel pipeline 408 may read data from or store data to the GMEM 410. Further, the GMEM 410 may exchange data with the memory 450 via the render target (RT) 420.
[0064] Aspects of the present disclosure may relate to architectural techniques for improving visibility generation in the binning process so that more triangles hidden behind visible surfaces may be rejected in the binning process regardless of the order in which the triangles are received into the GPU.
[0065] In some configurations, a two-pass approach in binning may be used, wherein in a first pass, geometry positions may be streamed into the GPU. Vertex shading and visibility testing may then be performed on the triangles, including coarse (LRZ) / fine depth buffer updates. In this context, the depth buffer may be a memory-resident buffer in which the latest depth values may be stored on a per-pixel basis or on a pixel block basis. During the first pass, visibility of the triangles may be recorded. Further, a second pass in binning may be performed, wherein visibility information from the first pass may be used and visible triangles as determined in the first pass may be accessed (in other words, invisible triangles as determined in the first pass may not be accessed in the second pass). In the second pass, visible triangles from the first pass may be tested again against the already existing coarse / fine depth buffers. Any triangles that are completely behind the existing depth buffer may be rejected. During the second pass, visibility recording may be performed so that updated visibility data that removes hidden triangles may be made available to the bin rendering pass.
[0066] After the second binning process, the bin rendering process may remain the same (e.g., Figure 4 ), where a bin rendering pass can read the updated visibility data and render triangles that are considered visible after the second binning pass. The software-based two-pass binning technique can be associated with minimal hardware changes as long as the binning pipeline supports visibility checking. Further, because the binning pipeline can be concurrent with the rendering pipeline, the time associated with the additional binning pass can be hidden behind a rendering pass (e.g., a rendering pass for a previous frame).
[0067] Figure 5 is a diagram illustrating an example data flow 500 associated with a two-pass binning technique in accordance with one or more aspects. Figure 6 is an example flow chart 600 illustrating a first binning process in a two-process binning according to one or more aspects. Figure 7700 is an example flow chart illustrating a second binning process in a two-pass binning process according to one or more aspects. These operations may be performed at a GPU (e.g., processing unit 120). As shown, in a first binning process 540 in a binning pipeline 530, at 602, the GPU may read a binning command stream. At 604, the GPU (e.g., position pipeline 502a) may read position data 512a (e.g., positions associated with one or more triangles). At 606, the GPU (e.g., position pipeline 502a) may perform triangle visibility tests (visibility checks 504a) (e.g., frustum tests, backface culling, clipping, and zero pixel tests (i.e., no screen space pixels are covered by a triangle), etc.). At 608, the GPU may determine whether each of the triangles is visible. The GPU may mark the invisible triangles as such in a visibility stream (e.g., visibility data / LRZ buffer 514a) at 618. For the visible triangles identified at 608, at 610, the GPU may rasterize the visible triangles and may perform a coarse horizontal depth test. At 612, the GPU may update a coarse depth buffer (e.g., an LRZ buffer) (e.g., visibility data / LRZ buffer 514a). At 614, the GPU may determine whether each of the remaining triangles is still visible after the depth test at 610. The GPU may mark the triangles that are not visible after the depth test at 610 as such in the visibility stream at 618. For the triangles that are still visible after the depth test at 610, at 616, the GPU may perform a bin level visibility test (e.g., the GPU may classify the visible triangles into different bins and identify the triangles in the bins that are still visible). At 618, the GPU may mark the remaining visible triangles as such in the visibility stream. The position data 512a and the visibility data / LRZ buffer 514a may be stored in a memory 550 (e.g., DDR / main memory).
[0068] In the second binning process 542 in the binning pipeline 530, at 702, the GPU may read the binning command stream and may read visibility data (e.g., visibility data / LRZ buffer 514a) from the first binning process 540 (e.g., at the position pipeline 502b). At 704, the GPU may perform a visibility check 504b and may determine, for each of the visible triangles from the first binning process, whether the triangle is visible. The GPU may mark the invisible triangles identified at 704 as such in the visibility stream at 714. For the visible triangles identified at 704, at 706, the GPU may read the geometry position data 512b. At 708, the GPU may rasterize the visible triangles and may perform a coarse horizontal depth test. At 710, the GPU may update the coarse depth buffer (e.g., LRZ buffer) from the first binning process 540 (e.g., visibility data / LRZ buffer 514b). At 712, the GPU may determine whether each of the remaining triangles is still visible after the depth test at 708. The GPU may mark the triangles that are not visible after the depth test at 708 in the visibility stream, as shown at 714. The GPU may mark the remaining visible triangles as such in the visibility stream, at 714. The position data 512b and the visibility data / LRZ buffer 514b may be stored in the memory 550.
[0069] The rendering pipeline 532 may be similar to Figure 4 For example, the geometry pipeline 506, the pixel pipeline 508, the GMEM 510, the visibility data / LRZ buffer 516, the position / attribute 518, and the RT 520 may be similar to Figure 4 4. The geometry pipeline 406, pixel pipeline 408, GMEM 410, visibility data / LRZ buffer 416, position / attribute 418, and RT 420 in FIG. 4 may be used to generate the pixel pipeline 406, pixel pipeline 408, GMEM 410, visibility data / LRZ buffer 416, position / attribute 418, and RT 420. In other words, the bin rendering process may be performed as usual based on the output from the second binning process 542. The visibility data / LRZ buffer 416 and the position / attribute 418 may be stored in the memory 550 and may correspond to the visibility data / LRZ buffer 514 b and the position data 512 b, respectively.
[0070] In some configurations, an intermediate storage structure may be implemented within the GPU. After the visibility check is performed in the first binning process, the storage structure may store the position data. The geometry buffer may keep accumulating all visible triangle information until all geometry is streamed and the coarse / fine depth buffer is updated. Once all geometry of the surface is processed or if the geometry buffer becomes full, the stored visible geometry may be replayed (e.g., processed again in the binning pipeline) to check visibility again based on the updated coarse / fine depth buffer. Any triangle that was previously visible and has since been overwritten by a later geometry may now fail the depth visibility test. The visibility stream may be updated accordingly. Because the position data associated with the visible triangles identified as in the first binning process may be stored in a hardware intermediate storage device, vertex fetching and vertex shading operations may be bypassed in the second binning process, and the second binning process may directly use the position data stored in the hardware intermediate storage device. This hardware solution may be appropriate or desirable when the GMEM at the GPU is large enough (because a portion of the GMEM can be used as an intermediate storage structure).
[0071] The bin rendering process may use the visibility stream from the second binning process so that those triangles that pass the visibility test in both binning processes may be taken for the bin rendering process. Triangles that fail the visibility test in the second binning process may not be taken for the bin rendering process.
[0072] Figure 8 is a diagram illustrating an example data flow 800 associated with a two-pass binning technique using on-chip storage in accordance with one or more aspects. Fig. 9900 is an example flow chart illustrating a two-pass binning technique using on-chip storage according to one or more aspects. These operations may be performed at a GPU (e.g., processing unit 120). As shown, in a first binning process 840 in a binning pipeline 830, at 902, the GPU may read a binning command stream. At 904, the GPU (e.g., at a position pipeline 802a) may read geometry position data 812a (e.g., geometry position data associated with one or more triangles). At 906, the GPU may perform triangle visibility tests (e.g., visibility check 804a) (e.g., frustum test, backface culling, clipping and zero pixel test, etc.). At 908, the GPU may determine whether each of the triangles is visible. The GPU may mark the invisible triangles as such in the visibility stream at 924. For the visible triangles identified at 908, at 910, the GPU may store position data 812b (e.g., the positions of these triangles) in a storage device (e.g., an on-chip storage device such as GMEM 852). At 912, the GPU may rasterize the visible triangles and may perform a coarse horizontal depth test. At 914, the GPU may update the coarse depth buffer (e.g., LRZ buffer) (e.g., visibility data / LRZ buffer 814a). At 916, the GPU may determine whether each of the remaining triangles is still visible after the depth test at 912. The GPU may mark the triangles that are not visible after the depth test at 912 in the visibility stream at 924. For triangles that are still visible after the depth test at 912 (position data 812b associated with these triangles may be stored in an intermediate on-chip storage device (such as GMEM852)), in a second binning process 842, at 918, the GPU (e.g., position pipeline 802b) may replay the stored geometry to check visibility again based on the updated depth buffer (e.g., visibility check 804b). For triangles that have passed the replay check at 918, at 922, the GPU may perform a bin level visibility test. At 924, the GPU may mark the remaining visible triangles as such in the visibility stream (e.g., visibility data / LRZ buffer 814b). Once all triangles for a particular surface have been processed, as determined at 920, the process may return to 912. Position data 812a and visibility data / LRZ buffers 814a and 814b may be stored in memory 850 (e.g., DDR / main memory).
[0073] The rendering pipeline 832 may be similar to Figure 4 For example, the geometry pipeline 806, the pixel pipeline 808, the GMEM 810, the visibility data / LRZ buffer 816, the position / attribute 818, and the RT 820 may be similar to Figure 4 4, the geometry pipeline 406, the pixel pipeline 408, the GMEM 410, the visibility data / LRZ buffer 416, the position / attribute 418, and the RT 420 in FIG. In other words, the bin rendering process may be performed as usual based on the output from the second binning process 842. The visibility data / LRZ buffer 816 and the position / attribute 818 may correspond to the visibility data / LRZ buffer 814 b and the position data 812 b, respectively, and may be stored in the memory 850.
[0074] In some aspects described above, two binning passes can be used: a first binning pass can be used to generate a coarse / fine depth buffer, and a second binning pass can include a visibility check using an updated depth buffer. Using the above aspects, visible geometry can be accessed twice to obtain a better visibility flow. In some further aspects, visibility information of a previous frame can be used to obtain a better visibility flow for the current frame.
[0075] In most games and benchmarks, scenes may show minimal changes in dynamic geometry from frame to frame. A small portion of geometry that was previously visible (e.g., in a previous frame) may become invisible in the next frame, and vice versa.
[0076] Therefore, in some configurations, first, the binning pipeline can be enhanced to add the ability to accept and process visibility data. In the first binning pass, the visibility of triangles marked as visible in the previous frame can be tested using the camera / viewport parameters of the current frame. The coarse / fine depth buffer can be updated accordingly. It is possible that some triangles that were previously visible may become invisible in the current frame. Once the perceived visible triangles have been processed using the camera / viewport positions of the current frame to update the coarse depth buffer, the second binning pass can process the triangles marked as invisible in the previous frame.
[0077] When the camera parameters of the current frame are used, some of the triangles marked as invisible in the previous frame may become visible, but most of the triangles marked as invisible may remain invisible in the current frame as in the previous frame. The visibility stream may be recorded for both passes. Since the depth buffer may be updated with triangles that are likely to be visible, triangles behind the visible triangles may be rejected. Therefore, valid visibility information may be obtained at the end of the second pass.
[0078] The bin rendering pass can use data from both visibility streams to process triangles designated to be rendered for the final render target. The scheme that utilizes visibility information from the previous frame can be beneficial when the visibility frame of the previous frame is the best. Therefore, in some configurations, an initial two-pass binning mechanism can be used for the first frame after a scene change (e.g., as described above with respect to Figures 5 to 9 described in the technique).
[0079] Fig.10 is a diagram of an example flowchart 1000 illustrating a first binning process in a two-pass binning method that utilizes visibility information of previous frames according to one or more aspects. Fig.11 is a diagram of an example flowchart 1100 illustrating a second binning process in a two-pass binning method that utilizes visibility information of previous frames according to one or more aspects. Fig.12 1 is a diagram of an example flow chart 1200 illustrating a bin rendering process, wherein a two-pass binning method utilizing visibility information of a previous frame is implemented according to one or more aspects. These operations may be performed at a GPU (e.g., processing unit 120). As shown, at 1002, the GPU may read a binning command stream. At 1004, the GPU may retrieve a visibility stream of a previous frame. At 1006, the GPU may determine whether each of the triangles is visible in the previous frame. For triangles visible in the previous frame, at 1008, the GPU may read geometry position data. At 1010, the GPU may perform triangle visibility tests (e.g., frustum tests, backface culling, clipping and zero pixel tests, etc.) using the camera / viewport of the current frame. At 1012, the GPU may determine whether each of the triangles visible in the previous frame is visible in the current frame. The GPU may mark triangles that have become invisible in the current frame as such at 1022, and may update the first visibility stream accordingly. For triangles that are still visible in the current frame, as identified at 1012, at 1014, the GPU may rasterize the visible triangles and may perform a coarse horizontal depth test. At 1016, the GPU may update a coarse depth buffer (e.g., an LRZ buffer). At 1018, the GPU may determine whether each of the triangles is still visible after the depth test at 1014. For triangles that are not visible after the depth test at 1014, the GPU may mark these invisible triangles in the current frame as such at 1022, and may update the first visibility stream accordingly. For triangles that remain visible after the depth test at 1014, at 1020, the GPU may perform a bin horizontal visibility test. At 1022, the GPU may mark the remaining visible triangles in the current frame as such, and may update the first visibility stream accordingly.
[0080] At 1102, the GPU may read the binning command stream. At 1104, the GPU may retrieve the visibility stream of the previous frame. At 1106, the GPU may determine whether each of the triangles is visible in the previous frame. For triangles that are not visible in the previous frame, at 1108, the GPU may read the geometry position data. At 1110, the GPU may perform triangle visibility tests (e.g., frustum tests, backface culling, clipping, and zero pixel tests, etc.) using the camera / viewport of the current frame. At 1112, the GPU may determine whether each of the triangles that were not visible in the previous frame has become visible in the current frame. The GPU may mark the triangles that remain invisible in the current frame as such at 1122, and may update the second visibility stream accordingly. For triangles that have become visible in the current frame, as identified at 1112, at 1114, the GPU may rasterize the visible triangles and may perform a coarse horizontal depth test. At 1116, the GPU may update a coarse depth buffer (e.g., LRZ buffer) from the first binning process. At 1118, the GPU may determine whether each of the triangles is still visible after the depth test at 1114. For triangles that are not visible after the depth test at 1114, the GPU may mark these invisible triangles in the current frame as such at 1122, and may update the second visibility stream accordingly. For triangles that remain visible after the depth test at 1114, the GPU may perform a bin level visibility test at 1120. At 1122, the GPU may mark the remaining visible triangles in the current frame as such, and may update the second visibility stream accordingly.
[0081] To perform a bin rendering process, at 1202, the GPU may read a binning command stream. At 1204, the GPU may read a first visibility stream (e.g., as updated at 1022), and may read a second visibility stream (e.g., as updated at 1122). At 1206, the GPU may determine whether each of the triangles marked as visible in the first visibility stream or the second visibility stream is actually visible. At 1208, the GPU may read geometry attribute data for the remaining visible triangles. At 1210, the GPU may render the bins, and may update the RT of each bin.
[0082] Therefore, based on the various aspects of the present disclosure described above, workload can be reduced in terms of bandwidth as well as ALU instructions.The performance benefits may be particularly significant for geometry-heavy workloads.
[0083] Fig.131300 is a call flow diagram illustrating example communications between a GPU and a storage device in accordance with one or more techniques of the present disclosure. At 1306, the GPU 1302 may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set.
[0084] In one configuration, each primitive in the plurality of primitives may correspond to a triangle.
[0085] At 1308 , the GPU 1302 may store at least one of visibility information or position data for all primitives in the plurality of primitives from the first binning process in the storage device 1304 before updating the depth buffer at 1310 .
[0086] At 1310 , GPU 1302 may update a depth buffer based on visibility information of all primitives in the plurality of primitives in the at least one frame.
[0087] In one configuration, at least one of visibility information or position data for all primitives in the plurality of primitives from the first binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory at 1310 .
[0088] At 1312, GPU 1302 may perform a second binning process on each primitive in the set of visible primitives based on the updated depth buffer. The second binning process may be associated with updated visibility information for each primitive in the set of visible primitives. The updated visibility information for each primitive in the set of visible primitives may correspond to an updated visible indication or an updated invisible indication.
[0089] In one configuration, the first binning process at 1306 may include generating a first visibility test of visibility information for each primitive in the plurality of primitives. The second binning process at 1312 may include generating a second visibility test of updated visibility information for each primitive in the set of visible primitives.
[0090] At 1314 , the GPU 1302 may store at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process in the storage device 1304 .
[0091] In one configuration, at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory.
[0092] At 1316, GPU 1302 may perform a bin rendering process for each primitive in the set of visible primitives including the updated visible indication. The bin rendering process may be performed after storing at least one of the updated visibility information or the updated position data.
[0093] In one configuration, the at least one frame may include one or more of a current frame or a previous frame in a set of frames of a scene associated with the graphics processing.
[0094] In one configuration, the first binning pass at 1306 and the second binning pass at 1312 may be based on visibility information of each primitive in the plurality of primitives in the previous frame.
[0095] In one configuration, the at least one frame may be included in a scene associated with graphics processing at GPU 1302. The first binning process at 1306 and the second binning process at 1312 may be performed by GPU 1302.
[0096] Fig.14 1400 is a flowchart of an example method of graphics processing according to one or more techniques of the present disclosure. The method may be performed by a device such as: a device for graphics processing, a GPU, a CPU, a wireless communication device, etc., such as in combination with Figures 1 to 13 used in all aspects.
[0097] At 1402, the device may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set. For example, referring to Fig.13 At 1306, GPU 1302 may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. Further, 1402 may be performed by Figure 1 The processing unit 120 in is executed.
[0098] At 1404, the device may update the depth buffer based on visibility information of all of the plurality of primitives in the at least one frame. Fig.13At 1310, GPU 1302 may update the depth buffer based on visibility information of all the primitives in the plurality of primitives in the at least one frame. Further, 1404 may be performed by Figure 1 The processing unit 120 in is executed.
[0099] At 1406, the device may perform a second binning process on each primitive in the set of visible primitives based on the updated depth buffer. The second binning process may be associated with updated visibility information for each primitive in the set of visible primitives. The updated visibility information for each primitive in the set of visible primitives may correspond to an updated visible indication or an updated invisible indication. For example, referring to Fig.13 At 1312, GPU 1302 may perform a second binning process on each primitive in the visible primitive set based on the updated depth buffer. Further, 1406 may be performed by Figure 1 The processing unit 120 in is executed.
[0100] At 1408, the device may store at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process. Fig.13 At 1314, GPU 1302 may store at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process (e.g., in storage device 1304). Further, 1408 may be performed by Figure 1 The processing unit 120 in the embodiment executes.
[0101] Fig.15 1500 is a flowchart of an example method of graphics processing according to one or more techniques of the present disclosure. The method may be performed by a device such as: a device for graphics processing, a GPU, a CPU, a wireless communication device, etc., such as in combination with Figures 1 to 13 used in all aspects.
[0102] At 1502, the device may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set. For example, referring to Fig.13 At 1306, GPU 1302 may perform a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. Further, 1502 may be performed by Figure 1 The processing unit 120 in is executed.
[0103] At 1506, the device may update the depth buffer based on visibility information of all of the plurality of primitives in the at least one frame. Fig.13 At 1310, GPU 1302 may update the depth buffer based on visibility information of all the primitives in the plurality of primitives in the at least one frame. Further, 1506 may be performed by Figure 1 The processing unit 120 in is executed.
[0104] At 1508, the device may perform a second binning process on each primitive in the set of visible primitives based on the updated depth buffer. The second binning process may be associated with updated visibility information for each primitive in the set of visible primitives. The updated visibility information for each primitive in the set of visible primitives may correspond to an updated visible indication or an updated invisible indication. For example, referring to Fig.13 At 1312, GPU 1302 may perform a second binning process on each primitive in the visible primitive set based on the updated depth buffer. Further, 1508 may be performed by Figure 1 The processing unit 120 in is executed.
[0105] At 1510, the device may store at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process. Fig.13 At 1314, GPU 1302 may store at least one of updated visibility information or updated position data of all primitives in the visible primitive set from the second binning process (e.g., in storage device 1304). Further, 1510 may be performed by Figure 1 The processing unit 120 in is executed.
[0106] In one configuration, at 1504, before updating the depth buffer, the device may store at least one of visibility information or position data for all primitives in the plurality of primitives from the first binning process. Fig.13 At 1308, GPU 1302 may store at least one of visibility information or position data of all primitives in the plurality of primitives from the first binning process (e.g., in storage device 1304) before updating the depth buffer at 1310. Further, 1504 may be performed by Figure 1 The processing unit 120 in is executed.
[0107] In one configuration, refer to Fig.13At least one of visibility information or position data for all primitives in the plurality of primitives from the first binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory at 1310 .
[0108] In one configuration, at 1512, the device may perform a box rendering process for each primitive in the set of visible primitives that includes the updated visibility indication. The box rendering process may be performed after storing at least one of the updated visibility information or the updated position data. For example, referring to Fig.13 At 1308, before updating at 1316, GPU 1302 may store, GPU 1302 may perform a bin rendering process on each primitive in the visible primitive set including the updated visible indication. Further, 1512 may be performed by Figure 1 The processing unit 120 in is executed.
[0109] In one configuration, each primitive in the plurality of primitives may correspond to a triangle.
[0110] In one configuration, refer to Fig.13 The first binning process at 1306 may include generating a first visibility test of visibility information for each primitive in the plurality of primitives. The second binning process at 1312 may include generating a second visibility test of updated visibility information for each primitive in the set of visible primitives.
[0111] In one configuration, at least one of updated visibility information or updated position data for all primitives in the visible primitive set from the second binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory.
[0112] In one configuration, the at least one frame may include one or more of a current frame or a previous frame in a set of frames of a scene associated with the graphics processing.
[0113] In one configuration, refer to Fig.13 , the first binning process at 1306 and the second binning process at 1312 may be based on visibility information of each of the plurality of primitives in the previous frame.
[0114] In one configuration, refer to Fig.13 , the at least one frame may be included in a scene associated with graphics processing at GPU 1302. The first binning process at 1306 and the second binning process at 1312 may be performed by GPU 1302.
[0115] In a configuration, a method or apparatus for graphics processing is provided. The apparatus may be a GPU, a CPU, or some other processor that can perform graphics processing. In various aspects, the apparatus may be a processing unit 120 within a device 104, or may be some other hardware within a device 104 or another device. The apparatus may include a component for performing a first binning process associated with visibility information of each of a plurality of primitives in at least one frame. The visibility information of each of the plurality of primitives may correspond to a visible indication or an invisible indication. Each of the plurality of primitives having the visible indication may correspond to a visible primitive set. Each of the plurality of primitives having the invisible indication may correspond to an invisible primitive set. The apparatus may further include a component for updating a depth buffer based on the visibility information of all primitives in the plurality of primitives in the at least one frame. The apparatus may further include a component for performing a second binning process for each primitive in the visible primitive set based on the updated depth buffer. The second binning process may be associated with the updated visibility information of each primitive in the visible primitive set. The updated visibility information for each primitive in the set of visible primitives may correspond to an updated visible indication or an updated invisible indication. The apparatus may further include means for storing at least one of updated visibility information or updated position data for all primitives in the set of visible primitives from the second binning process.
[0116] In one configuration, the device may further include means for storing at least one of visibility information or position data for all of the plurality of primitives from the first binning process before updating the depth buffer. In one configuration, at least one of visibility information or position data for all of the plurality of primitives from the first binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory. In one configuration, the device may further include means for performing a bin rendering process for each primitive in the set of visible primitives including the updated visible indication. The bin rendering process may be performed after storing at least one of the updated visibility information or the updated position data. In one configuration, each of the plurality of primitives may correspond to a triangle. In one configuration, the first binning process may include a first visibility test that generates visibility information for each of the plurality of primitives. The second binning process may include a second visibility test that generates updated visibility information for each of the primitives in the set of visible primitives. In one configuration, at least one of updated visibility information or updated position data for all primitives in the set of visible primitives from the second binning process may be stored in at least one of a GPU buffer, an on-chip buffer, or a system memory. In one configuration, the at least one frame may include one or more of a current frame or a previous frame in a set of frames of a scene associated with graphics processing. In one configuration, the first binning process and the second binning process may be based on visibility information for each of the plurality of primitives in the previous frame. In one configuration, the at least one frame may be included in a scene associated with graphics processing at a GPU. The first binning process and the second binning process may be performed by a GPU.
[0117] It should be understood that the specific order or hierarchy of the boxes / steps in the processes, flow charts and / or call flow charts disclosed herein is only an illustration of the exemplary method. It should be understood that the specific order or hierarchy of the boxes / steps in these processes, flow charts and / or call flow charts can be rearranged based on design preferences. In addition, some boxes / steps can be combined or omitted. Other boxes / steps can also be added. The attached method claims provide the elements of various boxes / steps in a sample order, but are not meant to be limited to the specific order or hierarchy provided.
[0118] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the various aspects shown herein, but should be given the full scope consistent with the language of the claims, wherein unless otherwise specified, reference to an element in the singular is not intended to represent "one and only one", but to represent "one or more". The word "exemplary" is used herein to mean "used as an example, instance, or illustration". Any aspect described herein as "exemplary" is not necessarily to be interpreted as being preferred or having advantages over other aspects.
[0119] Unless specifically stated otherwise, the term "some" refers to one or more, and the term "or" may be interpreted as "and / or" if the context does not dictate otherwise. Combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof," include any combination of A, B, and / or C, which may include multiple A, multiple B, or multiple C. Specifically, combinations such as "at least one of A, B, or C," "one or more of A, B, or C," "at least one of A, B, and C," "one or more of A, B, and C," and "A, B, C, or any combination thereof" may be only A, only B, only C, A and B, A and C, B and C, or A and B and C, where any such combination may include one or more members of A, B, or C. All structural and functional equivalents of the elements of various aspects described throughout this disclosure that are known or will later become known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly stated in the claims. Words such as "module", "mechanism", "element", "device", etc. cannot replace the word "component". Therefore, no claim element will be understood as a part plus function unless the element is explicitly stated using the phrase "component for..."
[0120] In one or more examples, the functionality described herein may be implemented in hardware, software, firmware, or any combination thereof. For example, although the term "processing unit" is used throughout this disclosure, such processing unit may be implemented in hardware, software, firmware, or any combination thereof. If any functionality, processing unit, technique, or other module described herein is implemented in software, the functionality, processing unit, technique, or other module described herein may be stored or sent as one or more instructions or codes on a computer-readable medium.
[0121] Computer-readable media may include computer data storage media and communication media, including any media that facilitate the transfer of computer programs from one place to another. In this manner, computer-readable media may generally correspond to: (1) a non-transitory tangible computer-readable storage medium; or (2) a communication medium, such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to extract instructions, codes, and / or data structures for implementing the techniques described in this disclosure. As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, or other magnetic storage devices. As used herein, magnetic disks and optical disks include compact discs (CDs), laser discs, optical disks, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where magnetic disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. A computer program product may include computer-readable media.
[0122] The technology of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs) or IC groups (e.g., chipsets). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of devices configured to perform the disclosed technology, but they do not necessarily need to be implemented by different hardware units. On the contrary, as described above, the various units can be combined in any hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with suitable software and / or firmware. Therefore, the term "processor" as used herein can refer to any of the above structures or any other structure suitable for implementing the technology described herein. Similarly, these technologies can be fully implemented in one or more circuits or logic elements.
[0123] The following aspects are merely illustrative and may be combined with other aspects or teachings described herein without limitation.
[0124] Aspect 1 is a graphics processing method, the method comprising: performing a first binning process associated with visibility information of each of a plurality of primitives in at least one frame, wherein the visibility information of each of the plurality of primitives corresponds to a visible indication or an invisible indication, wherein each of the plurality of primitives having the visible indication corresponds to a visible primitive set, wherein each of the plurality of primitives having the invisible indication corresponds to an invisible primitive set; updating a depth buffer based on the visibility information of all of the plurality of primitives in the at least one frame; performing a second binning process on each of the primitives in the visible primitive set based on the updated depth buffer, wherein the second binning process is associated with updated visibility information of each of the primitives in the visible primitive set, wherein the updated visibility information of each of the primitives in the visible primitive set corresponds to an updated visible indication or an updated invisible indication; and storing at least one of the updated visibility information or updated position data of all primitives in the visible primitive set from the second binning process.
[0125] Aspect 2 may be combined with Aspect 1, and further comprises: before updating the depth buffer, storing at least one of the visibility information or the position data of all primitives in the plurality of primitives from the first binning process.
[0126] Aspect 3 may be combined with Aspect 2 and includes storing at least one of the visibility information or the position data of all primitives in the plurality of primitives from the first binning process in at least one of a GPU buffer, an on-chip buffer, or a system memory.
[0127] Aspect 4 can be combined with any one of Aspects 1 to 3, and further includes: performing a box rendering process for each primitive in the set of visible primitives including the updated visible indication, wherein the box rendering process is performed after storing at least one of the updated visibility information or the updated position data.
[0128] Aspect 5 may be combined with any one of Aspects 1 to 4, and includes: each primitive of the plurality of primitives corresponds to a triangle.
[0129] Aspect 6 can be combined with any of Aspects 1 to 5, and includes: the first binning process includes generating a first visibility test of the visibility information for each of the multiple primitives, and wherein the second binning process includes generating a second visibility test of the updated visibility information for each primitive in the visible primitive set.
[0130] Aspect 7 can be combined with any one of Aspects 1 to 6, and includes: at least one of the updated visibility information or the updated position data of all primitives in the visible primitive set from the second binning process is stored in at least one of a GPU buffer, an on-chip buffer, or a system memory.
[0131] Aspect 8 may be combined with any one of aspects 1 to 7, and includes: the at least one frame comprising one or more of a current frame or a previous frame in a set of frames of a scene associated with the graphics process.
[0132] Aspect 9 may be combined with aspect 8, and comprises that the first binning process and the second binning process are based on the visibility information of each primitive of the plurality of primitives in the previous frame.
[0133] Aspect 10 may be combined with any one of aspects 1 to 9, and includes: the at least one frame is included in a scene associated with the graphics processing at a GPU, wherein the first binning process and the second binning process are performed by the GPU.
[0134] Aspect 11 is a device for graphics processing, the device comprising: at least one processor, the at least one processor is coupled to a memory, and based at least in part on information stored in the memory, the at least one processor is configured to implement the method according to any one of Aspects 1 to 10.
[0135] Aspect 12 may be combined with aspect 11 and further comprises a transceiver, wherein the apparatus is a wireless communication device.
[0136] Aspect 13 is a device for graphics processing, the device comprising means for implementing the method according to any one of aspects 1 to 10.
[0137] Aspect 14 is a computer-readable medium (eg, a non-transitory computer-readable medium) storing computer-executable code, which, when executed by at least one processor, causes the at least one processor to implement the method according to any one of aspects 1 to 10.
[0138] Various aspects have been described herein. These and other aspects are within the scope of the following claims.
Claims
1. A device for graphics processing, the device comprising: Memory; and at least one processor coupled to the memory and based at least in part on the information stored in the memory, the at least one processor configured to: performing a first binning process associated with visibility information of each of a plurality of primitives in at least one frame, wherein the visibility information of each of the plurality of primitives corresponds to a visible indication or an invisible indication, wherein each of the plurality of primitives having the visible indication corresponds to a visible primitive set, and wherein each of the plurality of primitives having the invisible indication corresponds to an invisible primitive set; updating a depth buffer based on the visibility information for all primitives in the plurality of primitives in the at least one frame; performing a second binning process on each primitive in the visible primitive set based on the updated depth buffer, wherein the second binning process is associated with updated visibility information for each primitive in the visible primitive set, wherein the updated visibility information for each primitive in the visible primitive set corresponds to an updated visible indication or an updated non-visible indication; as well as At least one of the updated visibility information or updated position data for all primitives in the set of visible primitives from the second binning process is stored.
2. The apparatus of claim 1 , wherein the at least one processor is further configured to: Prior to updating the depth buffer, at least one of the visibility information or position data for all primitives in the plurality of primitives from the first binning process is stored.
3. The apparatus of claim 2 , wherein at least one of the visibility information or the position data for all primitives in the plurality of primitives from the first binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
4. The apparatus of claim 1 , wherein the at least one processor is further configured to: A bin rendering process is performed on each primitive in the set of visible primitives that includes the updated visible indication, wherein the bin rendering process is performed after storing at least one of the updated visibility information or the updated position data. The apparatus of claim 1 , wherein each primitive of the plurality of primitives corresponds to a triangle.
6. The apparatus of claim 1 , wherein the first binning process comprises generating a first visibility test of the visibility information for each primitive in the plurality of primitives, and wherein the second binning process comprises generating a second visibility test of the updated visibility information for each primitive in the set of visible primitives.
7. The apparatus of claim 1 , wherein at least one of the updated visibility information or the updated position data for all primitives in the visible primitive set from the second binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
8. The device of claim 1, wherein the at least one frame comprises one or more of a current frame or a previous frame in a set of frames of a scene associated with the graphics process. 9 . The apparatus of claim 8 , wherein the first binning process and the second binning process are based on the visibility information of each primitive in the plurality of primitives in the previous frame.
10. The apparatus of claim 1, wherein the at least one frame is included in a scene associated with the graphics processing at a graphics processing unit (GPU), wherein the first binning process and the second binning process are performed by the GPU.
11. The apparatus of claim 1, wherein the apparatus is a wireless communication device.
12. A graphics processing method, the method comprising: performing a first binning process associated with visibility information of each of a plurality of primitives in at least one frame, wherein the visibility information of each of the plurality of primitives corresponds to a visible indication or an invisible indication, wherein each of the plurality of primitives having the visible indication corresponds to a visible primitive set, and wherein each of the plurality of primitives having the invisible indication corresponds to an invisible primitive set; updating a depth buffer based on the visibility information for all primitives in the plurality of primitives in the at least one frame; performing a second binning process on each primitive in the visible primitive set based on the updated depth buffer, wherein the second binning process is associated with updated visibility information for each primitive in the visible primitive set, wherein the updated visibility information for each primitive in the visible primitive set corresponds to an updated visible indication or an updated non-visible indication; as well as At least one of the updated visibility information or updated position data for all primitives in the set of visible primitives from the second binning process is stored.
13. The method according to claim 12, further comprising: Prior to updating the depth buffer, at least one of the visibility information or position data for all primitives in the plurality of primitives from the first binning process is stored.
14. The method of claim 13, wherein at least one of the visibility information or the position data for all primitives in the plurality of primitives from the first binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
15. The method according to claim 12, further comprising: A bin rendering process is performed on each primitive in the set of visible primitives that includes the updated visible indication, wherein the bin rendering process is performed after storing at least one of the updated visibility information or the updated position data. The method of claim 12 , wherein each primitive of the plurality of primitives corresponds to a triangle.
17. The method of claim 12, wherein the first binning process comprises generating a first visibility test of the visibility information for each primitive in the plurality of primitives, and wherein the second binning process comprises generating a second visibility test of the updated visibility information for each primitive in the set of visible primitives.
18. The method of claim 12, wherein at least one of the updated visibility information or the updated position data for all primitives in the visible primitive set from the second binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
19. The method of claim 12, wherein the at least one frame comprises one or more of a current frame or a previous frame in a set of frames of a scene associated with the graphics process.
20. The method of claim 19, wherein the first binning process and the second binning process are based on the visibility information of each primitive in the plurality of primitives in the previous frame.
21. The method of claim 12, wherein the at least one frame is included in a scene associated with the graphics processing at a graphics processing unit (GPU), wherein the first binning process and the second binning process are performed by the GPU.
22. A computer readable medium storing computer executable code which, when executed by at least one processor, causes the at least one processor to: performing a first binning process associated with visibility information of each of a plurality of primitives in at least one frame, wherein the visibility information of each of the plurality of primitives corresponds to a visible indication or an invisible indication, wherein each of the plurality of primitives having the visible indication corresponds to a visible primitive set, and wherein each of the plurality of primitives having the invisible indication corresponds to an invisible primitive set; updating a depth buffer based on the visibility information for all primitives in the plurality of primitives in the at least one frame; performing a second binning process on each primitive in the visible primitive set based on the updated depth buffer, wherein the second binning process is associated with updated visibility information for each primitive in the visible primitive set, wherein the updated visibility information for each primitive in the visible primitive set corresponds to an updated visible indication or an updated non-visible indication; as well as At least one of the updated visibility information or updated position data for all primitives in the set of visible primitives from the second binning process is stored.
23. The computer readable medium of claim 22, said code further causing said at least one processor to: Prior to updating the depth buffer, at least one of the visibility information or position data for all primitives in the plurality of primitives from the first binning process is stored.
24. The computer-readable medium of claim 23, wherein at least one of the visibility information or the position data for all primitives in the plurality of primitives from the first binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
25. The computer readable medium of claim 22, said code further causing said at least one processor to: A bin rendering process is performed on each primitive in the set of visible primitives that includes the updated visible indication, wherein the bin rendering process is performed after storing at least one of the updated visibility information or the updated position data.
26. The computer-readable medium of claim 22, wherein each primitive of the plurality of primitives corresponds to a triangle.
27. The computer-readable medium of claim 22, wherein the first binning process comprises generating a first visibility test of the visibility information for each primitive in the plurality of primitives, and wherein the second binning process comprises generating a second visibility test of the updated visibility information for each primitive in the set of visible primitives.
28. The computer-readable medium of claim 22, wherein at least one of the updated visibility information or the updated position data for all primitives in the visible primitive set from the second binning process is stored in at least one of a graphics processing unit (GPU) buffer, an on-chip buffer, or a system memory.
29. The computer-readable medium of claim 22, wherein the at least one frame comprises one or more of a current frame or a previous frame in a set of frames of a scene associated with graphics processing.
30. The computer readable medium of claim 29, wherein the first binning process and the second binning process are based on the visibility information of each primitive in the plurality of primitives in the previous frame.