Systems and methods for efficient access to memory and avoidance of unnecessary computations
By compressing the surface on the GPU and building a dedicated execution path, leveraging the value locality in the texture, solving the challenges of modern graphics applications for memory bandwidth and computing requirements, achieving more efficient graphics rendering performance.
Patent Information
- Application Number
- CN201911201912.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-24
- Filing Date
- 2019-11-29
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-12-23
AI Technical Summary
Modern graphics applications have extremely high demands on GPU memory bandwidth and computing, and it is difficult for the prior art to effectively utilize the locality of values in textures to reduce storage and computing load.
By checking and building dedicated execution paths for compressed surfaces, leveraging value localization in textures, such as UniformTexOpti software optimization and dirty tile mapping techniques, avoiding dynamic computation redundancy.
Improves GPU efficiency, reduces memory access and computing requirements, and improves graphics rendering performance and frame rate.
Smart Images

Figure CN111986279B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] None.
[0003] Statement Regarding Federally Sponsored Research or Development
[0004] None. Technical Field
[0005] The present technology relates to techniques for efficiently processing surfaces (such as textures) to exploit value locality. More specifically, the techniques herein relate to runtime checks for compressing surfaces, and specialized execution paths that reduce the need for memory load and / or computation. Background Art
[0006] The never-ending pursuit of photo-realistic real-time rendering and increased display resolution means that graphics-intensive applications continue to place high memory bandwidth and computational demands on modern graphics processing units (GPUs). GPU manufacturers have historically addressed these challenges by leveraging technology scaling and building GPUs with higher processing capabilities and supplying more memory bandwidth. However, as technology scaling approaches its end, it becomes important to adopt a first-principles approach to improving GPU efficiency to find alternative ways to meet the demands of modern graphics applications.
[0007] Texture mapping is a ubiquitous technique for efficiently achieving various effects in computer-generated images, such as realistic modeling of rough surfaces (e.g., brick or stone walls), fabric patterns, the texture of a desktop, the leaves of a tree, or any complex image feature that does not require 3D detail. Texture mapping typically involves defining a texture map, most commonly an array of texture elements (or "texels"). Simply put, a texture is a one-dimensional, two-dimensional, or three-dimensional array of such floating-point or integer texel values. Texel values typically represent color or other visualization parameters. Thus, in most textures, each texel has a unique coordinate (e.g., in one, two, or three dimensions, such as coordinates u, v, w), color, and in some cases, other attributes (such as surface normals).
[0008] Textures can be static, or they can be dynamically generated. Static textures are typically stored in mass storage and provided by application developers as part of an application. Dynamic textures are generated within a frame and can take various forms, such as shadow mapping, light mapping, reflection mapping, etc. Shaders can combine dynamic textures with other scene effects (e.g., other textures, geometric rendering, etc.) to produce an image for storage and / or display. As an example, real-time ray tracing can be used to generate maps of shadows and / or reflections, which can then be blended by a shader with the scene image generated from geometry to produce a real-time display.
[0009] There is significant value locality in the textures used in modern graphics applications. In the case of textures, value locality means that the values of texels are very similar or even identical to the values of other texels in the same texture map. Value locality can manifest itself spatially locally or globally on the texture surface. For example, some dynamic surfaces are cleared to a background color (e.g., black for night scenes, or sky blue for day scenes) and then conditionally rendered such that when they are read in as textures, most of them return the same background color.
[0010] Value Locality in Example Textures - A Case Study
[0011] Figures 1A - 1F Example textures with high value locality (also referred to herein as partially uniform textures) from various real-world applications are shown. Such textures can be static or dynamic and can be predetermined or generated at runtime. For example, Figures 1A - 1C is a light map, Figure 1D and Figure 1E is a reflection map generated at runtime, and Figure 1F can be a static texture predetermined by the application developer and provided with the application.
[0012] "High value locality" means that several texels have exactly the same color value such that texture mapping or other operations on these texels will result in common (identical) texture map values. This value locality typically manifests as one or more spatially contiguous regions of texels. It can be seen that dynamically rendered textures have a non-trivial amount of value locality. On average for an exemplary sampling, 38% of the textures have > 30% of their values exactly the same.
[0013] As a non-limiting example, Figure 2 shows the normalized breakdown of the total count of static (not generated within a frame) and dynamic (generated within a frame) textures and the proportion of partially uniform textures in each category for various different applications. Figure 2 is obtained by taking a histogram of the values in all textures in the frames of 12 different applications. For this particular non-limiting test, on average 38% of all textures in the frame tend to show partial uniformity, where in one example, the non-limiting context can be arbitrarily defined as 30% or more of the texels in a texture having exactly the same color value. In this context, a "texel" represents a cell of a multi-dimensional texture array. Textures typically store floating-point color values. Although adjacent texels may visually appear the same, their actual floating-point values may have slight differences. Thus, the fact that 38% of the textures show partial uniformity is significant. Additionally, in a particular Figure 2 sampling, dynamic textures exhibit much higher partial uniformity than static textures.
[0014] The memory systems of modern GPUs have been designed to operate efficiently for graphics applications, where memory access patterns exhibit a large amount of value locality. For example, some current GPUs identify and exploit this value locality in texture surfaces to effectively compress textures and save memory bandwidth (but these GPUs do not necessarily do much else with such value locality). The sizes of textures and other image surfaces, especially high-resolution images, can be large. Loading them from main memory may require many processor cycles. To reduce storage size and loading time, textures are typically compressed as much as possible. In particular, modern GPUs identify and exploit value locality in partially uniform textures by compressing them to save memory bandwidth. See, for example, Brennan, C., "Delta Color Compression Overview" (March 14, 2016); Smith, R., "NVIDIA Geforce GXT 980 Review: Maxwell Mark 2" (September 18, 2014); and U.S. Patent No. US8330766B1.
[0015] Commonly used texture compression / decompression CODECs include DXT, ETC, and ASTC. Texture compression techniques typically provide different modes that exploit redundancy in texture data to increase the compression ratio. For example, some modes (sometimes called "reduction compression" modes) reduce the overall texture data size based on redundancy. For example, if a texture includes many identical color values (e.g., many adjacent texture pixels have the same blue or black sky shade), the texture compression CODEC can store the color value once and repeat that value when the texture is decompressed. Other modes (sometimes called "differential compression") determine the differences of texture values relative to one or more baseline colors and encode the differences. This is a bit like writing down the height of the center of a basketball team and then measuring the heights of everyone else relative to the center ("Jane is 2 inches shorter than Alyssa, and Katie is 1 inch taller than Alyssa"). Each texel can be recovered by, for example, adding or subtracting the difference from one or more baseline colors. See, for example, USP 8,594,441.
[0016] Texture compression techniques typically generate and store metadata associated with the compressed texture, sometimes referred to as "compression state" information. The compression state information typically describes the compression mode and some characteristics of the compression result. The CODEC uses this compression state information as a guide for decompressing the texture. In some cases, the compression state information is stored in a memory table separate from the compressed texture, so that it can be accessed more conveniently (e.g., from the on-chip cache rather than the main memory). Some proprietary texture compression formats operate entirely within the GPU and can only be accessed through kernel authorization.
[0017] Another known technique for more efficiently storing and processing textures is texture tiling. See, e.g., Wei, Tile-Based Texture Mapping on Graphics Hardware (Graphics Hardware 2004). In some cases, a larger virtual texture can be generated by repeating a smaller number of texture tiles. In other cases, the entire texture or other surface is explicitly stored, but is still divided into sub-regions or "tiles" (somewhat like the tiles on a kitchen floor) to assist with storage and memory management. For example, since Maxwell, NVIDIA GPUs have supported tile caches, taking advantage of locality and the L2 cache by processing geometry and textures in small enough chunks so that both the input and output can reside in the on-chip cache. This tiling can also be used to assist with parallel processing. See generally, e.g., McCormack et al., "Neon: A Single-Chip 3D Workstation Graphics Accelerator", p123, Proceedings of the ACM / SIGGRAPH / EUROGRAPHICS Workshop on Graphics Hardware (Association for Computing Machinery, August 1998); Akenine-Moeller et al., Real-Time Rendering, especially Chapters 6 and 23 (4th Edition, CRC Press 2018).
[0018] It would be useful to take advantage of texture and other surface value locality to reduce or eliminate dynamic computational redundancy (e.g., by computing value reuse). BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Read the following detailed description of exemplary non-limiting illustrative embodiments in conjunction with the accompanying drawings:
[0020] Figures 1A - 1F is an image of an example of a partially uniform texture.
[0021] Figure 2Shows an exemplary non - restrictive characterization of partial uniformity in static and dynamic textures for multiple different examples (description of partial uniformity in the texture) (in this non - restrictive example, "partial uniformity" means that >= 30% of the texels in the texture have the same color).
[0022] Figure 3 Shows an example non - restrictive system process.
[0023] Figure 3A Shows an example non - restrictive system.
[0024] Figure 3B Shows an example non - restrictive process.
[0025] Figures 4A - 4D Shows a dirty tile mapping example.
[0026] Figure 5 Shows an example dirty tile mapping construction pre - pass.
[0027] Figure 6A 、 6B Shows an example texture with many uniform tiles.
[0028] Figure 7 Shows an example valid dirty tile mapping construction.
[0029] Figure 7A Shows an example non - restrictive code snippet.
[0030] Figure 8 Shows an example query tile and dirty tile mapping tile operation.
[0031] Figure 9 Shows before and after the fragment, showing how UniformTexOpti transforms the API code sequence and a single texture lookup in the shader program.
[0032] Figure 10 Shows an example partial evaluation and reuse.
[0033] Figure 11 Shows an example partial evaluation and reuse with leaders and non - leaders depending on the compression mode.
[0034] Figure 12A Shows an example non - restrictive unique FP color value seen across frames.
[0035] Figure 12B Shows an example non - restrictive performance boost of simple UniformTexOpti.
[0036] Figure 12CShows texture lookup reduction from UniformTexOpti.
[0037] Figure 13 Shows a conceptual diagram of an example graphics processing pipeline implemented by a PPU according to an embodiment.
[0038] Detailed Description of Exemplary Non - Limiting Embodiments
[0039] The exemplary non - restrictive techniques herein describe ways to utilize surface / texture value locality to avoid dynamic computational redundancy. For example, assume that it is possible to avoid the memory fetches for texture lookups in the black regions of an Figures 1A - 1F image and directly write the value 0.0 (for black) to the destination register for such lookups. Then, imagine being able to specialize the dependent code for these textures with substantial value locality.
[0040] Example non - restrictive embodiments provide a software optimization, referred to as "UniformTexOpti", which utilizes surface memory compression information to efficiently construct a coarse - grained representation of a surface, called a value locality map or "dirty tile map" ("DTM"), and then uses these DTMs (and / or other pre - information mechanisms) to avoid dynamic computational redundancy and optimize the runtime performance of shader programs through non - speculative software optimizations (e.g., avoid redundant memory lookups and / or mathematical operations).
[0041] The example non - restrictive techniques herein can be applied to any compressed data (images, videos, audio, disk files, etc.) processed on a CPU and / or GPU. Example optimizations (for a CPU or GPU system processing a compressed data array) include:
[0042] · Uniform texture optimization - Benefiting from the extensive repetition of a small number of values in an image texture (or more generally, a data array)
[0043] · Partial evaluation and reuse - Benefiting from local uniformity (repeated values or low value variability) in an image texture (or more generally, a data array).
[0044] · Can operate "on - the - fly"
[0045] · Applicable to tiled videos, video analysis, etc. images
[0046] · Applicable to any compression.
[0047] · Applicable to deep learning.
[0048] Deep learning:
[0049] Weight vectors in deep learning systems typically have most weights as 0. By querying compression information and knowing in advance which parts of the weights are 0, not only can memory lookups be avoided, but many related mathematical codes can also be specialized and optimized. Both training and inference can benefit.
[0050] High-value locality textures (“global partial uniformity”):
[0051] -> Well compressed with uniform compression mode
[0052] -> Compression information can be used for efficient dirty tile map construction
[0053] -> Determine version control conditions for uniformTexOpti
[0054] -> Create fast path and default slow path for version optimization
[0055] -> For data with high value locality, the fast path is adopted more frequently, which provides performance benefits.
[0056] This optimization has shown promise in real-time graphics applications, including but not limited to virtual reality, gaming, head-up displays, and any other environment that utilizes textures or other surfaces. It works by selectively avoiding memory fetches for partially uniform textures in shader programs and instead using program paths dedicated to one or more frequently occurring values. The decision to use the dedicated fast path is made dynamically by consulting a coarse-grained representation of the partially uniform texture, called the dirty tile map (“DTM”). These techniques can accelerate the processing of partially uniform textures, increase frame rates, and may also save energy (which is especially important for mobile chips and real-time graphics systems).
[0057] Applications include real-time graphics and deep learning. In many cases, machine learning is implemented as brute-force technology where the designer doesn't know how to design an elegant way to solve a difficult problem and thus turns to machine learning. Often, the designer may not know which features are important and which are not. When training is complete, it can be determined that only some features are actually important and contribute to the final result. If certain features are determined to be irrelevant, the weights corresponding to those features (e.g., in a neural network) can be set to zero. It is not uncommon to see a non-trivial proportion of weights in a machine learning system ultimately being set to zero. However, currently, GPUs and other processors still perform matrix multiplication operations on each individual value (zero or non-zero) loaded from machine learning vectors. If the processor knew in advance that a weight portion was zero, the number of computations the processor needs to perform could be reduced. Similarly, during the training of a deep neural network (DNN), the output of an intermediate layer called activation (which can be thought of as weights for combinations of features) also tends to exhibit partial uniformity and this knowledge can be beneficially exploited in subsequent DNN layers.
[0058] Example non-limiting overall system
[0059] Figure 3 An example non-limiting system 200 for avoiding dynamic computational redundancy is shown. In this example embodiment, surfaces (such as textures) are stored in memory in a compressed form, e.g., sub-sampled compression 204, differential compression 206, and a technique called "zero-bandwidth clearing" 202 is used. Metadata associated with such compression (e.g., compression state) can also be stored in memory - typically in a processor chip memory (such as an L2 cache). Information stored in this way is accessed (208) to construct an abbreviated or simplified (e.g., coarse-grained) data structure called a "dirty tile map" (DTM) (201), and expressions are also reconstructed during shader execution (212). The combination of the DTM and expression reconstruction allows the shader to use / access such surfaces to avoid dynamic computational redundancy (e.g., by skipping memory fetches, providing specialized code, etc.) (214).
[0060] Example non-limiting real-time graphics system
[0061] As a non-limiting example, Figure 3A An exemplary non-limiting real-time graphics system 50 is shown, which includes using such DTM and techniques to avoid dynamic computational redundancy.
[0062] In response to real-time input from an input device 54, a CPU and / or GPU 56 that executes one or more shaders 58 accesses graphic information (such as geometry and texture arrays 64) stored in a DRAM 62 to generate an image for display on a display 60.
[0063] In the illustrated embodiment, tile textures or other surfaces 64 are stored in DRAM 62, and corresponding DTMs 64' are stored in the L2 cache memory 66 on the same chip as the CPU and / or GPU 56. In a simple case, the DTM 64' uses one or a small number of bits in each tile of the texture or other surface 64 to represent or indicate whether all positions in each texture tile of the texture / surface have a uniform value. Although other regions can be used, in this embodiment, the granularity is at the texture tile level, and the DTM 64' indicates for each tile whether all the texels in the tile have the same value. If all positions in a tile do not have the same value, the tile is considered "dirty". The DTM 64' can be flexibly defined by software (SW) as needed. They can be defined, for example, simply as a one-dimensional array of, for example, 32b words, where each word holds the dirty or non-dirty state of the corresponding texture 64 tile.
[0064] The granularity of the tiles themselves can be defined flexibly again. For example, the tiles herein can refer to ROP (raster operation) tiles or higher granularity regions. Similarly, the number of bits per tile in the DTM can be defined flexibly based on usage. In their simplest form, the DTM uses a 1-bit representation to simply convey whether the value of the corresponding tile is in the cleared (i.e., initial) state or not cleared (i.e., dirty). More generally, a DTM implementation can use n bits per tile and use 2 n to the 2 n minus 1 possible states to convey which of the 2 n minus 1 unique uniform values the tile has and the remaining one-bit pattern to convey that none of the other 2 n minus 1 values apply.
[0065] The DTM 64' construction can be very efficient such that the gain in avoiding memory fetches far outweighs the cost of the DTM construction. In terms of performance, a simple DTM 64' implementation for retrieving and analyzing the values of all texels in a texture can be very expensive. A faster alternative is to directly or indirectly use the memory compression state (simply referred to as compstatus) of, for example, 256B tiles. The compression state can reflect one of several compression modes. In some example non-limiting embodiments, of particular interest are the color and depth zero bandwidth clear (ZBC) compression state (see, for example, U.S. Patent No. US8330766B1) and the 8:1 downsampling compression mode.
[0066] Therefore, Figure 3AShows the compressed tile surface / texture 64 stored in the main memory DRAM 62. The compressed surface / texture 64 typically includes metadata (compression status information) indicating the compression status of each tile in the surface / texture 64. In one example non-limiting embodiment, driver software (which can also be executed by the CPU and / or GPU 56) can analyze this compression-related metadata to generate DTM 64'. Hardware support can be added to assist in the efficient construction of DTM 64'. Thus, the above functions do not depend on any specific hardware features or implementations.
[0067] Example of the entire process
[0068] In Figure 3B In the embodiment shown, the process analyzes the uniformity and / or interest value (82) of a texture or other surface, and based on the observed uniformity and / or any identified interest value, compiles shader code 58 (84) using a dedicated execution path. For example, an optimization compiler 52 (which can run on a development computer 52, although it is also possible to optimize the interpreter executed on the CPU and / or GPU 56) creates an executable shader target code that has a "clean" dedicated execution path in addition to the "dirty" (normal or default) execution path. The optimized shader target code executable 58 determines whether to execute the dedicated path or the normal path based on the content of DTM 64' (which has not been created yet, but whose format and specification are pre-specified).
[0069] At runtime, the system uses a driver pre-pass process to inspect the surface and create DTM 64' (86). When DTM 64' indicates uniformity and / or interest value (88), the CPU / GPU 58 executes the compiled shader code and invokes one or more dedicated execution paths. As described in detail below, executing the dedicated execution path can save the need to load the surface / texture 64 from DRAM 62 and / or perform calculations on the loaded surface / texture data.
[0070] Thus, the exemplary non-limiting technique can assist GPU applications by detecting and optimizing value locality in dynamic textures. It can assist deep learning applications by detecting and optimizing sparse weight and activation matrices. For example, its benefits stem from avoiding TEX traffic (reducing memory congestion and increasing the effective L1 cache capacity) and enabling code specialization.
[0071] The example non-limiting technique improves GPU efficiency by eliminating dynamic computational redundancy originating from textures / surfaces with high value locality. This improves performance and thus can even save energy.
[0072] Some example non - limiting embodiments of the present disclosure provide software improvements and / or optimizations referred to as “UniformTexOpti” and enabling techniques to read memory compression states to utilize memory compression information already available in modern GPUs, thereby avoiding dynamic computational redundancy in graphics - intensive applications. In one example embodiment, “UniformTexOpti” works by selectively avoiding parts of uniform textures in a memory fetch shader program and instead using a program path dedicated to statically - known values. In some non - limiting embodiments, during program execution, a decision on whether to use the dedicated fast path is made dynamically by consulting a coarser - grained representation of such parts of uniform textures stored previously (i.e., a “dirty” tile map (DTM) or other high - level information that indicates whether a given tile (e.g., 8x8 or 16x16 texels) is “dirty” (i.e., has a value different from the value assumed to be specialized for the program)).
[0073] The use of the term “dirty” is somewhat different from the conventional usage of the term in the context of caches, where “dirty” typically means “written to” (a data block in the cache that the processor has modified since reading from main memory and thus needs to be written back to main memory before releasing the associated cache line). In certain contexts such as zero - bandwidth clear (ZBC), the term “dirty” does mean writing data after an area of a cleared surface / texture. But in other cases, “dirty” simply means that the surface / texture area is non - uniform. Although texture memory access for texture mapping purposes is read - only in some contexts because texture mapping generally does not change the texture and thus texture mapping operations generally do not “dirty” the texture in memory by writing, in other contexts, e.g., a dynamic texture processor will write to the texture after it has been cleared to a uniform color. Some exemplary non - limiting embodiments may claim a texture tile as “dirty” when its texels (by any mechanism) are found to be dissimilar or non - identical. Example non - limiting embodiments may use effective techniques (such as Figure 3B the driver pre - pass shown in) to determine when a texture tile is “dirty”.
[0074] For example, in an example non - limiting embodiment, the aforementioned DTM 64' can be constructed dynamically within a frame by an explicit driver - introduced pre - pass executed before a draw call for “UniformTexOpti”. Figure 3AOne example non - limiting embodiment for assisting in fast and effective DTM construction as shown involves using vanilla memory load instructions in DTM construction code to directly read the compressed state of a tile from a virtual memory system (including cache and main memory) or from a dedicated hardware structure (where the system may choose to keep or cache the compressed state in a dedicated storage structure). This embodiment requires that the compressed state storage be directly addressable by the driver software. In an alternative embodiment, where the compressed state storage cannot be directly accessed by the user - mode driver (UMD), simple hardware enhancements and a suitably enhanced style of memory load instructions can be used to map the data address of a tile to the corresponding compressed state address. A prototype of the above - mentioned simple, non - invasive form of "UniformTexOpti" is able to reduce the total number of memory lookups by, for example, up to 16.5% and is implemented through proof - of - concept software on a modern high - performance GPU 56, achieving an average of 2.5% and up to 6.5% frame - time acceleration in a set of modern graphics applications.
[0075] Other exemplary non - limiting embodiments of this document provide the following non - limiting features and / or advantages:
[0076] 1) A software optimization called "UniformTexOpti" to avoid memory lookups and dependent computations for partially uniform textures by consulting a pre - built coarse - grained representation called the dirty tile map (DTM) 64'.
[0077] 2) A way to read memory compression information through user - mode driver (UMD) software to assist in fast and effective DTM 64' construction.
[0078] In the following description, the first part provides a high - level background on various prominent aspects of 3D graphics programming in an API - agnostic way and introduces some terms. The next part describes non - limiting embodiments for developing the dirty texture map (DTM). The next part describes non - limiting embodiments for providing shader execution efficiency and specialized execution to avoid redundant computations and exploit value uniformity. The last part presents quantitative results.
[0079] High - level background on various prominent aspects of 3D graphics programming
[0080] At a high level, one can Figure 3AFrames of real-time 3D graphics applications (such as virtual reality, augmented reality, head-up displays, games, etc.) executed on a system take (e.g., virtual) eye positions, the level at which a viewer is located in a 3D scene, and various static textures as their inputs to produce a final image output to a display. From a software perspective, it is very useful to view these applications as a two-level hierarchy of API calls and shader programs. A frame makes one or more subordinate API calls (in modern applications, this can be up to, for example, 5000 calls). The calls can be graphics draw calls, compute dispatches, clears, copies, and other calls that manipulate the API state. Depending on resource availability, multiple API calls can be made simultaneously in the GPU 56.
[0081] A draw or dispatch call consumes zero or more input textures 64 and produces one or more output textures or other surfaces. During a draw or dispatch call, a shader program 58 is typically used to read the input textures at the desired locations, perform mathematical transformations on the read values, and produce location-specific output values in the output surface / texture of that draw call. Some high-performance GPU implementations perform such operations in a massively parallel manner. For more details on example GPU architectures and their use in real-time graphics, deep learning, and other environments, see Figure 13 and the content that follows.
[0082] Example construction of a dirty tile map
[0083] One exemplary non-limiting technique for providing the above "advanced knowledge" is to use a "dirty tile map" (DTM) 64'.
[0084] More specifically, one example non-limiting embodiment provides UniformTexOpti, which is a driver-managed computational value reuse (CVR) technique that shows promise in graphics applications. It works by selectively avoiding memory fetches of portions of uniform textures in a shader program and instead using program paths dedicated to the most frequently occurring values. In some non-limiting embodiments, the decision to use a dedicated fast path can be made dynamically by consulting a coarse-grained representation of a portion of the uniform texture, called a dirty tile map (DTM) 64'.
[0085] In one example non - limiting embodiment, each tile of the DTM 64' uses one or a small number of bits to convey whether all positions in a given tile have a uniform value. The DTM 64' can be flexibly defined by software (SW) according to its needs. For example, the tiles here can refer to raster operation (ROP) granularity tiles or higher - granularity regions. Similarly, the number of bits for each tile in the DTM 64' can be defined flexibly based on usage. In their simplest form, the DTM 64' can use a 1 - bit representation to simply convey whether the value of the corresponding block is in a cleared (i.e., initial) state or an uncleared state (i.e., dirty). More generally, the DTM 64' implementation can use n bits per tile and use 2 n of the 2 n - 1 possible states to convey which of the 2 n - 1 unique uniform values the tile has and the remaining one - bit pattern to convey that none of the other 2 n - 1 values apply.
[0086] Thus, an example surface (such as a texture) is divided into coarse - granularity regions, e.g., tile regions. Some example non - limiting embodiments use a pre - pass to determine whether all the texels in the region have the same value. If they have the same value, the metadata is set to "clean". If the texels do not have the same value, the metadata for the region is set to "dirty". In some example embodiments, a single bit for each region can be used to indicate whether the region is "clean" or "dirty". In other embodiments, multiple bits can be used to indicate whether a tile is "dirty" or "clean". For example, in a texture that repeatedly uses two different colors such that all the texels in multiple tiles are the first color and all the texels in multiple other tiles are a second color different from the first color, it can be used to represent a "clean" state with two different bit patterns depending on the color (e.g., clean state "01" represents all black texture pixels, and clean state "10" represents all white texture pixels, and the state "11" represents a dirty tile that does not belong to either of the above types). In general, N bits can be used to represent 2 N possibilities.
[0087] Software can define the number of texels represented by a single DTM tile and the number of bits used to represent the tiles in the DTM, e.g.:
[0088] · Unit per - tile DTM - If the values of all texels in the tile are equal to a specific (global) value, the bit of the tile is transmitted. 0 represents clean (i.e., expected value), and 1 represents dirty.
[0089] · Multi - bit per - tile DTM - If the values of all texels in the tile are equal to at most 2 n- One of a specific (global) value or not present at all (for a total of 2 n possibilities), then transmit the n bits of the tile.
[0090] A version control transformation can be designed (discussed below) to handle 2 n cases.
[0091] Example DTM
[0092] Figures 4A - 4D Shows an example of a "dirty" DTM 64' constructed using an example non - restrictive automated infrastructure:
[0093] · Figure 4B Shows Figure 4A A texture having a "dirty" (non - uniform) region d and a "clean" region of uniform color c0.
[0094] · Figure 4D Shows Figure 4C A texture containing a "clean" region of uniform color c0, a "clean" region of a second uniform color c1, and a "dirty" region d. (In this particular example, c1 is the color of the grass in the foreground of this particular intermediate texture, so the grass is not actually green, but some shade of blue).
[0095] Figure 4B 、 4D The dirty tile mapping shown in can be much smaller than the corresponding texture. For example, Figure 4B DTM 64' can be 8192x smaller than a 1 - bit DTM (a single bit indicates "clean" or "dirty"), and Figure 4D DTM can be 4096x smaller than a 2 - bit DTM (2 bits are used to represent three states: dirty, "clean" with color c0; "clean" with color c1).
[0096] An example detailed process for creating such a dirty tile mapping can include:
[0097] · Divide the surface into appropriately sized coarse - grained tiles (see of the original texture Figure 6A and of the segmented texture Figure 6B )
[0098] · Determine whether all the texels in the tile have a given value of interest (in the case of Figure 6B , 0.0 represents black)
[0099] · If so, mark the tile as clean in DTM 64'; otherwise, dirty
[0100] · Thus, in DTM 64', tiles with flares or other features will be marked as dirty.
[0101] In some exemplary embodiments, the DTM 64' can be constructed by software in an explicit pre-pass that runs before a draw call that requires their output. In some example non-limiting embodiments, tests can be performed based on the identified interest values. Such interest values can be identified based on analysis, programmer knowledge, heuristics, artificial intelligence / machine learning, or any other technique. Then, the driver software can introduce specialized channels to test these interest values and use the results of these tests to construct such a dirty tile map. Thus, in some example non-limiting embodiments, the driver pre-pass not only identifies tiles whose all texels have the same color; it can also adjust the identification of whether the texels of a tile have one of a small set of predetermined colors. Then, the DTM can be used by compiled programs that are used to process raw partially uniform textures. In another exemplary non-limiting embodiment, the driver pre-pass can not only identify uniform tiles whose all constituent texture pixels have the same color, but also dynamically discover the 2 n-1 most popular colors in the texture and then appropriately encode them with n bits for each DTM block. In such an embodiment, the pre-pass not only passes the DTM to the subsequent program that is the target of UniformTexOpti, but also passes an auxiliary array of the 2 n-1 most popular colors.
[0102] In one example non-limiting embodiment, black can be the only color detected in the Figure 6B example pre-pass. Thus, the only time the pre-pass operation marks a tile as "clean" is when the pre-pass detects (a) that the texel value in the tile is black and (b) that the high-level information of the tile indicates that the texels in all tiles have the same color (more details below). When all texels have the same value, the high-level information will indicate.
[0103] The compiler and the driver software work together to provide version control transformation. In an example non-limiting embodiment, the driver uses the DTM information to convey the results of its pre-pass checks to the executing shader or other application process. Thus, the compiler that compiles the shader can also know which specific interest values the driver pre-pass is checking and can generate specialized execution paths based on those specific interest values (e.g., skip blending operations, blend into black to produce black, while performing block blending operations on different colors such as pink or blue).
[0104] Overall transformation using the DTM
[0105] As described above, the exemplary non-limiting embodiments use the DTM 64' to avoid doing unnecessary work and / or memory accesses. Figure 5 Shows how to use Figure 4BExample of DTM choosing between default execution path and faster dedicated execution path:
[0106] Example baseline program;
[0107]
[0108] Note that in the code above, the instruction "DTM[TileID] tests the DTM 64' of the tile to determine if the tile is "dirty". If the tile is not dirty ("!="), the fast path is taken. Otherwise, the slow (default) path is taken. Thus, when the tile is not dirty, the DTM 64' enables the code to avoid executing expensiveWork().
[0109] Efficient DTM construction leveraging memory compression information
[0110] The DTM 64' construction should preferably be very efficient such that the gain from avoiding memory fetches far outweighs the cost of the DTM construction. One exemplary non - limiting way to do this is to leverage compression information that may already be available for surface / texture tiles without adding any further overhead.
[0111] In many existing systems, value locality has been captured in the form of memory compression information. Historically, compression has been used to save bandwidth and optionally save storage. The GPU compresses uniform dynamic textures to minimize storage requirements. The compression information is stored as compressed state (metadata) and compressed data. The specific case depends on the compression type. Programs such as shaders can read the compressed state and compressed data in some way.
[0112] Reading each texel in a region to determine if all texels in the region have the same value is expensive. Thus, some example non - limiting embodiments leverage compression information to infer high - level information about whether all texels in a given region have the same value (and in some cases, which value). In some example embodiments, a pre - pass can be performed based on compression information associated with a texture tile (e.g., an 8×8 or other - sized region) that already resides in memory.
[0113] Conventional texture compression is a valuable tool for reducing the size of textures stored in memory. For example, in some texture compression arrangements, the texture tiles themselves reside in (texture) memory, while the compressed information for the tiles (e.g., compression state) resides in L2 or other caches. In the case where all the texels in a given tile have the same value, the corresponding index value can be stored in a table that also resides in the L2 cache. Thus, the processor does not need to leave the chip to create the DTM 64' and / or otherwise determine whether the tile should be processed by the default execution path or a dedicated execution path - the processor can determine this at runtime by examining the contents of its on-chip L2 cache. Additionally, in an example embodiment, if a dedicated path is employed, there is no need to access the original texture tile in memory at all - because the processor only takes the dedicated path when it determines that all the texels in the tile have the same predetermined known value (e.g., black), such that the dedicated path can bypass / eliminate per-texel operations on the texels themselves.
[0114] Accordingly, one example non-limiting embodiment uses the compressed information to infer value locality and eliminate value-dependent dynamic computation redundancy and improve program performance. Thus, such an example embodiment can use the decompression information that is already available to understand value locality. For example, tiles can be utilized to determine whether all the values in a tile are the same. Some examples are dynamically generated, but the technique also applies to static or "fixed" versions.
[0115] In one example context, conventional software and / or hardware memory compression is available and deployed on a system that is optimized. Certain fixed-granularity data (e.g., 1KB DRAM blocks) are compressed to a smaller number of bytes to save memory bandwidth and optionally save memory storage. One example non-limiting embodiment uses color and depth zero-bandwidth clear (ZBC) compression (see, e.g., U.S. Patent No. US8330766B1 titled "Zero-Bandwidth Clears"); the compression ratio is 8:1 (see Smith, R, "NVIDIA Geforce GTX 980 Review: Maxwell Mark 2" (September 18, 2014)).
[0116] When a texture is compressed, metadata about the compression type and other parameters related to the compression are typically stored in the DRAM 62 and / or the cache 66 as the compression state and the compressed data, respectively, referred to as compstatus and compdata. The compression state is the metadata about the compression, while the compressed data is the compressed data itself. The system uses the compression state metadata to decompress the compressed data. The compression state and the compressed data will vary according to the compression type, e.g., compression types: zero bandwidth clear, reduction, differential compression, etc. One example embodiment uses a memory compression state (simply referred to as the compression state) of tiles of, e.g., 256B or other sizes to infer the value locality characteristics of the underlying texture without reading or analyzing the compressed data. The compression state can, in some cases, constitute a DTM or an equivalent of a DTM, and thus can be directly used by a shader program to determine when to follow a dedicated execution path, or it can be used to generate a DTM. Thus, the shader program can use the compression state as "advanced information" that the shader program can use to infer the value locality characteristics of the surface without the shader program having to read or access the surface itself.
[0117] For example, in some embodiments, the driver checks known values of information stored in the memory to determine the presence of known values used in a "clear" API function call. Such a "clear" operation can be used to initialize most of a texture or other surface. If only some tiles are later changed dynamically, the remaining tiles will retain the color values initialized by the "clear", and the driver can test for this in a pre-pass. As an example, assume that the "clear" function is used to clear the entire screen to draw the screen sky blue (or black for a night sky). Then, assume that a dynamic process adds clouds, a moon, and a rocket, but most of the sky remains blue or black. The example non-limiting techniques herein can be used to identify when most tiles or other screen areas retain their initialized values and use specialized path execution to avoid the need to spend processing time and memory accesses to retrieve and process redundant values.
[0118] Figure 7 An example scenario showing the use of two different example uniform compression modes that can be used to compress textures with high value locality, where at least one tracks such a "clear":
[0119] · Zero bandwidth clear (ZBC) (see U.S. Patent 8330766B1)
[0120] · Reduction ("red") compression.
[0121] In this example, the texture is stored in DRAM 224 in a compressed or compact form. For example, texture 1 (with tiles 232a, 232b, 232c, 232d) is stored in DRAM, and texture 2 (with tiles 234a, 234b, 234c, 234d) is also stored in DRAM (here, "red" does not refer to the color red, but rather the fact that the tiles are compressed using a certain type of reduction compression as described above).
[0122] Regardless of the compression mode, in an exemplary non - limiting embodiment of encoding multiple colors, actual data values are typically required to construct the DTM 64' (e.g., Figure 4D ). For the ZBC (Zero Bandwidth Clear) compression state, knowledge of the data values is embedded in the compression state itself.
[0123] In Figure 7 , the "compression state" metadata associated with the compression type / characteristics is stored in the L2 cache 222. The data structure ZBC compression state 226 conveys all positions in a block that contain the same value, and a pointer to the value is encoded in the compression state block. It is well known that ZBC compression can be triggered in response to an API - level clear() call (i.e., an initialization call) that assigns an initial value to a texture surface. In one example non - limiting embodiment, in response to such a clear(0) call, the software driver programs the clear value table (or "ZBC table") in the L2 cache 222 with the desired clear value, such that all tiles of the cleared surface will have their compression state "ZBC" and contain a pointer to the relevant ZBC table entry 230, which contains the value for the region that was cleared. Subsequent writes to an individual tile can change the compression state, but on a later read, if a tile is found to have the ZBC state, then given that all of its constituent pixels have not been modified since being cleared (i.e., "not dirty"). Thus, for ZBC - based DTM 64' construction, software or other query functions only need to access the compression state - rather than the texture 232 itself stored in DRAM 224 or any more detailed compression - related information - at least in cases where the region is marked as ZBC - cleared.
[0124] For 8:1 reduced tiles, the DTM 64' construction in an example non - limiting embodiment accesses the compression state 228 and 8:1 reduced data 234 from the memory 224. In this case, the DTM construction shader may need to issue a plurality of loads (e.g., two subsequent regular 16B loads) to extract the reduced values for a large (e.g., 256B) tile. Even in this case of accessing compressed data 234 in DRAM 224, the DTM 64' structure for 8:1 reduced tiles is 8 times faster than simply reading each texel.
[0125] Query tiles and DTM tiles
[0126] A number of surfaces are defined as multiple n-dimensional arrays (data matrices) of different resolutions that can be compressed differently. In one non-limiting exemplary embodiment, the granularity of a query can be determined by the compression granularity in the system, referred to as a "query tile".
[0127] One or more query tiles can form a DTM tile, the state of which will be equal to the logical OR of the "dirty" states of all the constituent query tiles. The number of texels in a query tile will depend on the size of the texel (written as bits per pixel or bpp). The granularity of the DTM tile is determined by the overall DTM size budget and the size of the input surface.
[0128] Figure 8 Shows how the size and number of query tiles for each DTM tile can vary according to bits per pixel. This example shows three different query tiles:
[0129] 128-bpp 4x4 query tile 252;
[0130] 32-bpp 8x8 query tile 254;
[0131] 8-bpp 16x16 query tile 256.
[0132] A common DTM tile 250 is constructed by performing a logical OR operation on the "dirty" results of three query tiles 252, 254, 256 with different sizes and resolutions.
[0133] Efficient DTM construction
[0134] The following is a code snippet for 1b / tile DTM to capture 1 out of 2 possibilities:
[0135]
[0136] The following is an example code snippet for 2b / tile DTM to capture 1 out of 4 possibilities;
[0137]
[0138]
[0139] Hierarchical DTM
[0140] In another non - limiting embodiment, the thus - constructed "DTM" can be further compressed hierarchically to convey information at an even coarser granularity. For example, the first - level DTM conveys whether a group of 8x8 texels has the same color. The next - level DTM constructed from the first - level DTM can represent the information of a 16×16 region of the first - level DTM, actually representing 128×128 texels (16×16 times 8×8) of the original surface. Due to its very coarse granularity, this second - level DTM can represent the entire original texture surface very concisely in a few bytes. In addition, this second - level DTM can convey at least three possible values, with 2 bits representing each 16x16 region. These values will convey whether all the first - level DTM bits in the 16x16 region are "clean" (represented by the 2 - bit pattern "00"), or whether they are all "dirty" (represented by the 2 - bit pattern "11"), or whether they are a mixture of clean and dirty tiles (represented by the 2 - bit pattern "10"). For the first - level DTM, a 3840×2160 texture will require (3840×2160) / (8×8) / 8 = 16,200 bytes, and for the second - level DTM, it will require (3840×2160)×2 / (128×128) / 8 = 127 bytes.
[0141] The consumer shader will first access the second - level DTM and only access the first - level DTM when the second - level DTM indicates that a particular texel belongs to a 16×16 region with a mixture of clean and dirty. If most lookups can be serviced outside the second - level (and even higher - level) DTMs, the dynamic working set of bytes due to DTM lookups and thus the resulting runtime overhead can be kept very low.
[0142] Profiling to provide higher efficiency
[0143] There are other ways to determine high-level information about the uniformity of values for textures or other surfaces. For example, profiling can be performed to gather information about texture value uniformity. For example, profiling can be used to determine which textures and which shader programs / processes are of interest in terms of potential increases in efficiency, and which localized texel values are produced by processing one or more such textures. Such profiling can be performed offline (based on results logged from program execution) or online (while the application is running in real time). For example, online profiling can run as a background process to analyze one or more specific scenarios. Profiling can consider a single user running a specific application on a particular system and / or multiple users running the application across multiple systems, or can use deep learning to perform the application and / or analyze surface / texture tasks that result in the greatest latency when the application is executed in various different ways. Surface / texture access and computation information can be gathered through such profiling and used to derive values of interest. In other contexts, the developer (the person designing the texture) can provide information identifying the most common texel values for a particular texture. Or deep learning can be used to analyze all the surfaces / textures of an application to identify one or more of the most popular values.
[0144] In such an analysis-feedback-based system, for example, some or all of the 8:1 reduced tiles described below can be harvested as ZBC tiles. By detecting a priori (e.g., through online or offline profiling) that a texture exhibits high value locality and its tiles are compressed with 8:1 reduction, the driver software can introduce explicit clear() calls with an analysis-determined clear value and modify the writers for all such surfaces to delete their writes if the value they are about to write is equal to the clear value. In this way, the compressed state remains as ZBC, which in turn will enable fast DTM construction. The DTM can be built online or offline, depending on when the compressed data is available for aggregation.
[0145] Example Execution Specialization
[0146] As described above, high-level knowledge of value locality can be exploited to eliminate dynamic computational redundancy ("computation value reuse"). Suppose a program is to perform the following operation on all texels of a Figure 6A / 6B texture, which can be very large (e.g., 16,777,216 texels):
[0147] result = tex(u, v) x expensiveWork()
[0148] Today, the processor would simply extract and process each individual texel as if the value of each texel were unique. However, all texel values in, for example, a black area are the same, i.e., 0.0. Redoing the same thing over and over again is wasteful.
[0149] The advanced knowledge of such only black regions through DTM 64' can not only help avoid memory fetches, but also help avoid expensiveWork() through code specialization (compiler constant folding), such as:
[0150]
[0151] The above code snippet provides version control transformation with optimized fast and default slow paths. The default slow path states "result = tex(u, v) x expensiveWork()". However, the code also takes into account that if the advanced knowledge of the texture element at coordinates u, v indicates that the texture element color value is black, then the product of the texture element and any result of the "expensive work" will be zero, meaning the time spent by the processor doing the "expensive work" will be wasted. Therefore, the code adds more fast (specialized) instructions that, based on the prior knowledge that the product will be zero, set the result to zero without doing the "expensive work". This is a bit like a restaurant refusing to prepare the meal ordered by a customer when it is known in advance that the customer has no money.
[0152] The more clean tiles there are, the higher the frequency of dynamically adopting the fast path and the higher the performance gain. The software system can choose to perform further code specialization in the "fast" path based on the statically available knowledge about the texture values in that path.
[0153] Knowledge about specific values will help us avoid irrelevant computations. For example, if we know that all tiles are black, then we can avoid memory lookups and set the output to zero. Different code snippets can be executed based on the context (e.g., providing a specialized path based on a specific identification value). Knowledge about specific values can help us avoid unnecessary memory lookups and other computations that would otherwise need to be performed on these specific values.
[0154] It is useful to ensure that the resulting program is functionally correct. Therefore, we introduce the normal (slow) path. But most of the time we want to execute in the optimized fast path.
[0155] Thus, version control relies on the value at a specific location. If a specific location contains a specific value, the code will execute in one way (e.g., in a specialized optimization path), and if that specific location contains a different value, it will execute in a different way (e.g., in a non-specific default path). In some of the example non-limiting embodiments herein, this is achieved through compiler optimizations. Thus, the compiler can avoid any computations in the fast path because the compiler knows that for that path, the result of the computation will be zero. In some cases, the optimizations performed by the compiler can be dramatic, e.g., eliminating the need for memory loads and any related computations on data that the load would have retrieved anyway. In other cases, the optimizations performed by the compiler may be less dramatic (e.g., providing "shortcuts" or other more efficient computations on values that are still loaded from memory).
[0156] Figure 7A A sample DirectX assembly code snippet from an application is shown to illustrate the advantages of specialization in a shader program. In this example, the 4-vector register r1.xyzw loaded from texture t47 is zero (0) 99% of the time. When the value of texture t47 is zero, the common 99% case, specialization through code replication and version control can help eliminate the lookup of texture t47 (in combination with texture t47). This example illustrates how specialization can further be used to reduce memory fetches and reduce related computations. These optimizations can significantly improve GPU efficiency by reducing or avoiding dynamic computational redundancy.
[0157] "DTM" can provide high-level knowledge for performing specialization
[0158] A coarse-grained DTM representation that conveys tile-granularity value locality is used to inform Figure 9 the execution specialization in the example, which is also elaborated below with additional annotations:
[0159] Baseline code:
[0160]
[0161] which can be converted to the following modified code:
[0162]
[0163] In the italicized code lines above (some of which are elaborated in the block in Figure 9 , one or more pre-passes are introduced to create a dirty tile map (DTM), which is bound as a read-only resource for Draw3() (e.g., as a constant buffer). Then the shader for Draw3() is modified to look up the DTMs for A and B and jump to an optimized fast path or a default slow path based on the results of the DTM lookups.
[0164] Figure 9 Shows another baseline consumer shader fragment, which shows an example of shader code specialization with an optimized "clean" path that avoids memory fetches and instead directly provides the result from the DTM:
[0165] Uniform texel optimization shader code fragment:
[0166]
[0167]
[0168] Partial evaluation and expression result reuse
[0169] In some cases, there is no extensive overall uniformity within a surface such as a texture, but there is local uniformity. Imagine a checkerboard where each tile is unique. In such a scenario, there may be no good or efficient way to construct a DTM with one or two bits per tile. For example, the tiles may be mostly uniform but show minute variations (e.g., such that the tiles are suitable for compression using a differential compression scheme). For example, assume that the texels do not have exactly the same color but only differ by a small amount relative to a base value. To compress such tiles using differential compression, the base value can be stored, and then each texel can be encoded with a value indicating the difference (e.g., magnitude and sign) between the value of the texel and the base value.
[0170] Instead of or in addition to using a pre-computed DTM to avoid redundant work, another way to exploit value locality is through expression result reuse, where the work is done in a leader thread and non-leader threads simply reuse the results from the leader thread. This is a useful strategy when there is no global uniformity (and thus the DTM is not valid), but there is local value locality (e.g., a checkerboard with uniquely colored tiles). Thus, surfaces with global variability are generally not suitable for DTMs, but they may be suitable for leader / non-leader work partitioning. In such cases, the compiler performs the partitioning and leader / non-leader communication setup. The leader reads the compression state and compressed data and computes its result. The non-leader can directly reuse the leader's result or reuse the leader's result along with a small amount of additional reconstruction work, depending on the compression mode (the compiler creates versions for different compression modes and reuse possibilities). Memory fetches and mathematical operations can be reduced in non-leader threads, thus saving energy and potentially improving performance.
[0171] Local value locality can be of two broad types: 1) repetitive, or 2) showing minute variations relative to a base value. Thus, the compression machine can compress them with different algorithms (e.g., run-length reduction for repetitive values or differential compression for values with mild variations).
[0172] For duplicate values, any applicable expressions can be directly reused from the leader thread. For non-duplicate value locality, understanding the underlying compression techniques is useful for code restructuring, and only certain expressions can be applied to restructuring and partial result reuse.
[0173] Toy example: Partial evaluation and reuse
[0174] Figure 10 An example is shown. Suppose a 4x4 matrix where each 2x2 has the same data. Suppose a program wants to add a constant to each element of the matrix, i.e., matrix[i][j]+K.
[0175] If 16 threads work on this 4×4 matrix (one thread per block), then only 4 threads compute the unique output. The remaining 12 threads can simply reuse the results of the 4 unique threads.
[0176] Direct reuse of duplicate value locality
[0177] Suppose there are multiple tiles in a tiled matrix and each tile is processed in the same way as all other tiles (e.g., blended with a solid color, etc.). In this case, the processing can be performed on the "leader" tile and then these results can be used to process all other tiles (considering the changes between each other tile and the "leader" tile). Thus, the hard work performed by the processor on the "leader" tile can then be reused for some or all of the other tiles. Suppose thread 0 and thread 1 are running on two adjacent positions (positions 0 and 1 respectively) of a 4B element array. Without loss of generality, assume the values at positions 0 and 1 are the same and have been compressed with a 2 to 1 reduction (i.e., storing the 2:1 reduction compression state + one 4B compressed data).
[0178] Also assume that thread 0 is the leader thread in a 2-thread group (thread 0, thread 1). Before reuse, the situation would be:
[0179] Thread 0
[0180] A = value[loc0] x K1
[0181] B = A + K2
[0182] Thread 1
[0183] A = value[loc1] x K1
[0184] B = A + K2.
[0185] In this "before" case, the threads work independently and are unaware of value locality, which results in redundant work.
[0186] As shown above, the original value lookup thread 0 multiplies the value stored at location "Location 0" by the constant K1 and adds the second constant K2 to the product to provide the result "B". Note that thread 1 performs the same operation on the value stored at location "Location 1". If the entire process can determine that the value stored at location "Location 1" is the same as the value stored at location "Location 0", the process can perform the calculation on the value stored at location "Location 0" in thread 0 and simply pass the result to thread 1 so that thread 1 does not need to access the value [Location 1] or perform the calculation again. Although thread 0 and thread 1 are independent threads that can execute in parallel, the resulting efficiency can reduce the memory load by half and reduce the mathematical processing overhead. A potential drawback is that thread 1 is now dependent on thread 0, which may or may not be tolerable depending on the situation.
[0187] Conversely, in the "post" scenario, the following operations can alternatively be performed;
[0188] Thread 0:
[0189] (leader_data) = LOAD.CD(loc0)
[0190] A = leader_data x K1
[0191] B = A + K2
[0192] SEND(B)
[0193] Thread 1:
[0194] (leader_B) = RECEIVE()
[0195] B = leader_B
[0196] In this "post" scenario, the memory load is halved and the mathematical overhead is reduced. The threads are now dependent threads with thread 0 messaging thread 1.
[0197] Code refactoring for non-repeating value locality
[0198] As another example, assume differential compression is used to compress the values of a particular tile or other region. In this case, all the texels in the tile have substantially the same value, and only a few (e.g., the last few least significant bits) are different. In the example shown, thread 0 can send the result of a computation based on a "leader" texel, along with a "delta" value indicating the difference between value[position 0] and value[position 1], to thread 1. Thread 1 can now reuse the result of value[position 0] computed by thread 0 and compute a correction factor (e.g., Δx K1) that corrects the result of the difference between the leader thread's value[position 0] and value[position 1]. This reuse saves the potentially time-consuming multiplication of value x K1, but still requires a multiplication (Δx K1) and an addition ("leader_B + [result of the multiplication]). However, there are ways to exploit specific values to provide specialized execution that offers further optimization.
[0199] Intuition: If f(x) = x.K1 + K2, then f(x + d) = (x + d).K1 + K2 = (x.K1 + K2) + d.K1 = f(x) + d.K1.
[0200] This is similar to the example above. Without loss of generality, now assume that the array has been successfully compressed using differential compression such that value[position 1] is guaranteed to be within a small constrained increment of value[position 0]:
[0201] Thread 0
[0202] (leader_data,delta) = LOAD.CD(loc0)
[0203] A = leader_data x K1
[0204] B = A + K2
[0205] SEND(B,delta)
[0206] Thread 1
[0207] (leader_B,delta) = RECEIVE()
[0208] B = leader_B + delta x K1
[0209] If the range of delta is known, the process can be optimized statically.
[0210] Code refactoring for non-repeating values locally
[0211] The example non-limiting embodiment uses a compiler to refactor certain expressions on non-leader threads and express them as simpler functions of partial results of leader texels and statically evaluable constants.
[0212] The refactoring has the following benefits:
[0213] · Memory loads are only performed in the leader thread,
[0214] · Expression evaluation in non-leader threads is simplified through partial evaluation and reuse;
[0215] · Helps save memory system bandwidth and energy;
[0216] · Fewer operations are performed on the core again, thus saving energy
[0217] The following example assumes that 2x4B is compressed into 1x4B + 4b delta. This example is generally written with f(), f'() and g(), and shows how to specialize f'() for a small delta range:
[0218]
[0219]
[0220] More specifically, the following table gives some examples of how the compiler refactors expressions and uses differential compression to combine the partial evaluation results from the leader thread and compile-time evaluation expressions to derive the final results for non-leaders. Topic:
[0221]
[0222] K1, K2, K3 in the above expressions are uniform in the threads of interest. The expressions d.K1, d.K2 and K1^d highlight the f'() function evaluated on the delta values.
[0223] Leaders and non-leaders depend on the compression mode
[0224] Assume that an application is running to create a shadow map. When creating the shadow map, it stores it in the L2 cache (and thus may also be stored in the main memory). When calculating the shadow map, the GPU detects the value uniformity in coarse-grained regions (e.g., cache lines, ROP tiles, etc.) and compresses (e.g., using reduction compression) the shadow map for storage. The GPU also stores compression state values and / or index the shadow map, where the compression state values indicate the reduction compression. When the shader process then wishes to utilize the shadow map to render an image to the display, the shader process (which may run in multiple threads and / or warps) reads the previously stored compression state values and discovers that the compression of multiple texels for multiple tiles of the shadow map is the same. At this stage, one or more leader threads retrieve / decompress the texels, calculate results based on those texels, and use the results in the rendering. One or more leader threads also send messages to other (follower) threads, sending the calculated results to the other (follower) threads. The other (follower) threads independently identify, based on the compression state values they read, that the leader threads are calculating values that the other (follower) threads can reuse. Thus, the other (follower) threads wait for the leader threads to send their calculated results to the other (follower) threads. When the leader threads send their results, the other (follower) threads reuse the results, avoiding the need to retrieve and decompress the texture and also the need to recalculate the same values that the leader threads have already calculated.
[0225] In the above variation, assume that the shadow map is differentially compressed. The other (follower) threads can calculate a result difference (Δs) based on the difference between the texel values they are processing and the texel values the leader threads are processing (the differential result is calculated based on the difference values provided in the differential compression map or other data structures). The other (follower) threads use their respective calculated Δs to correct the calculated values sent by the leader threads. The calculation for each thread is specified by the compiler at compile time, expecting that the shadow map can be differentially compressed. Thus, the other (follower) threads still need to do some work on their own, but not as much as when the leader threads do not share their calculated results with the other (follower) threads. Additionally, the other (follower) threads do not need to perform memory loads; when the other (follower) threads start accessing their respective texels, they determine that the texels are differentially compressed, so they read the difference values corresponding to the texels and then wait for the leader threads to send the calculated values instead of loading and decompressing them themselves. The compression state generally remains resident on the chip (e.g., in the L2 cache), so loading the compression state for a single tile is generally cheaper than loading all the texels of a tile (e.g., from the texture, L2, and / or main memory).
[0226] Figure 11Additional non - restrictive examples are shown. As before, we first create a LOAD.CS to understand the compression pattern for each memory lookup. Based on the pattern, the leader and non - leader threads can be flexibly determined. Suppose in Figure 11 (position 0, position 1), (position 2, position 3) in Figure 11 has a 2:1 reduction compression (thread 0 (since it reads position 0) is the leader of thread 1 (reads position 1), and similarly thread 2 is the leader of thread 3). For the next lookup, if (position 4, position 5, position 6, position 7) in Figure 11 has a 4:1 reduction compression, then thread 0 (since it reads position 4) will be the leader of threads 1, 2, and 3.
[0227] The high - level content is that the compiler creates versions for each lookup based on the initial compression state. Each version will know whether the current thread is the leader for that particular lookup or otherwise and react appropriately. For example:
[0228]
[0229]
[0230] In some applications, the system may encounter some tiles that undergo reduction compression, some that undergo differential compression, and some that are not compressed at all. In this case, the compiler can create multiple versions. One version can be customized for tiles with reduction compression, another for tiles with differential compression, and another for tiles not compressed using either method.
[0231] Example Performance Statistics / Results
[0232] Figure 12A Shows the number of unique FP color values seen in sampling per single frame and across multiple frames where possible. For most applications, an example non - restrictive experimental framework uses only a single - frame APIC for each application. For some applications, we are able to capture and use multiple single - frame APICs, and for those we have presented the number of unique values seen in all the study frames.
[0233] From Figure 12AWe observe that, on average, a single frame may require the driver to track and set the ZBC table to maintain up to 20 unique colors. For applications where we have multiple APICs, we see that the total number of unique colors seen across frames is typically 3 or 4 more than the average number of colors required for a single frame. In any case, the total number of unique colors for a typical application seems to be less than 32. Aside from the exact details of how large the GPU's ZBC table is, it doesn't seem unreasonable to have 32 ZBC tables. See, for example, NVIDIA Geforce GTX 1080 (2016). This means that all value locality in the frames we studied could in theory be harvested as ZBC.
[0234] Figure 12B Shows the performance improvement of clear value optimization from a conservative evaluation of an example non - restrictive GPU. The performance acceleration plotted against the primary y - axis is conservative because one example non - restrictive prototype uses knowledge about partial uniformity to avoid only texture lookups, but does not perform any code specialization based on the clear value. Even so, this example non - restrictive prototype shows an average increase of 2.5% through uniform texel optimization. As Figure 12C shown, the texture lookup count itself drops by 8.5%.
[0235] The reduction in texture access does not result in as much performance improvement because the performance of different regions of the frame / draw calls tends to be limited by different GPU bottlenecks, so the reduction in texture lookups does not directly translate into equivalent performance improvements. However, this reduction in workload is expected to translate into some energy savings.
[0236] Conclusion
[0237] Value locality is inherent in many real - time graphics applications. Uniform Texel Optimization (UniformTexOpti) exploits memory compression information to eliminate dynamic computational redundancy, thus improving GPU efficiency. Software optimizations utilize existing and future compression capabilities to build a coarse - grained representation of textures called the dirty tile map.
[0238] Graphics Processing Pipeline
[0239] In one embodiment, the PPU 300 is configured to receive commands specifying a shader program for processing graphics data. The graphics data can be defined as a set of primitives, such as points, lines, triangles, quads, triangle strips, etc. Typically, a primitive includes data specifying multiple vertices of the primitive (e.g., in a model - space coordinate system) and attributes associated with each vertex of the primitive. The PPU can be configured to process the primitives to generate a frame buffer (e.g., pixel data for each pixel of a display).
[0240] The application writes the model data of the scene (e.g., a collection of vertices and attributes) to a memory (such as system memory or a memory). The model data defines each of the objects that may be visible on the display. The application then makes an API call to the driver kernel, which requests the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. These commands can refer to different shader programs to be implemented on the SMs of the PPU, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, one or more of the SMs can be configured to execute a vertex shader program that processes multiple vertices defined by the model data. In one embodiment, different SMs can be configured to concurrently execute different shader programs. For example, a first subset of the SMs can be configured to execute a vertex shader program, while a second subset of the SMs can be configured to execute a pixel shader program. The first subset of the SMs processes the vertex data to produce processed vertex data and writes the processed vertex data to the L2 cache and / or memory. After the processed vertex data is rasterized (e.g., converted from three-dimensional data to two-dimensional data in screen space) to produce fragment data, the second subset of the SMs performs pixel shading to produce processed fragment data, which is then blended with other processed fragment data and written to the frame buffer in memory. The vertex shader program and the pixel shader program can be executed concurrently to pipeline-process different data from the same scene until all the model data of the scene has been rendered to the frame buffer. Then, the content of the frame buffer is transferred to the display controller for display on the display device.
[0241] Figure 13 is a conceptual diagram of a graphics processing pipeline 600 implemented by a PPU according to one embodiment. The graphics processing pipeline 600 is an abstract flowchart of processing steps implemented to generate a 2D computer-generated image from 3D geometry data. As is well known, pipeline architectures can execute long-latency operations more efficiently by dividing the operations into multiple stages, where the output of each stage is coupled to the input of the next consecutive stage. Thus, the graphics processing pipeline 600 receives input data 601 that is passed from one stage of the graphics processing pipeline 600 to the next stage to generate output data 602. In one embodiment, the graphics processing pipeline 600 can represent a graphics processing pipeline defined by the API. Alternatively, the graphics processing pipeline 600 can be implemented in the context of the functionality and architecture of the previous figures and / or any one or more of the subsequent figures.
[0242] As Figure 13As shown, the graphics processing pipeline 600 includes a pipeline architecture that includes multiple stages. These stages include, but are not limited to, a data assembly stage 610, a vertex shading stage 620, a primitive assembly stage 630, a geometry shading stage 640, a viewport scale, cull, and clip (VSCC) stage 650, a rasterization stage 660, a fragment shading stage 670, and a raster operations stage 680. As described above, software shading algorithms that work in conjunction with such shading hardware can be optimized to reduce computation time.
[0243] In one embodiment, the input data 601 includes commands that configure the processing unit to implement the stages of the graphics processing pipeline 600 and configure geometric primitives (e.g., points, lines, triangles, quads, triangle strips, or fans, etc.) to be processed by these stages. The output data 602 can include pixel data (e.g., color data) that is copied into a frame buffer or other type of surface data structure in memory.
[0244] The data assembly stage 610 receives the input data 601 that specifies vertex data for high-order surfaces, primitives, etc. The data assembly stage 610 collects the vertex data in temporary storage or a queue, such as by receiving a command from a host processor that includes a pointer to a buffer in memory and reading the vertex data from that buffer. The vertex data is then transmitted to the vertex shading stage 620 for processing.
[0245] The vertex shading stage 620 processes the vertex data by performing a set of operations (e.g., a vertex shader or program) once for each vertex. A vertex can be specified, for example, as a 4 - coordinate vector (e.g., <x, y, z, w>) associated with one or more vertex attributes (e.g., color, texture coordinates, surface normal, etc.). The vertex shading stage 620 can manipulate the individual vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading stage 620 performs operations on the vertex coordinates or other vertex attributes associated with the vertex. These operations typically include lighting operations (e.g., modifying the color attribute of the vertex) and transformation operations (e.g., modifying the coordinate space of the vertex). For example, a vertex can be specified using coordinates in an object coordinate space, which are transformed by multiplying the coordinates by a matrix that transforms the coordinates from the object coordinate space to the world space or the normalized - device - coordinate (NCD) space. The vertex shading stage 620 generates the transformed vertex data that is transmitted to the primitive assembly stage 630.
[0246] The primitive assembly stage 630 collects the vertices output by the vertex shading stage 620 and groups the vertices into geometric primitives for processing by the geometry shading stage 640. For example, the primitive assembly stage 630 may be configured to group every three consecutive vertices into a geometric primitive (e.g., a triangle) for transmission to the geometry shading stage 640. In some embodiments, a particular vertex may be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip may share two vertices). The primitive assembly stage 630 transmits the geometric primitives (e.g., a set of associated vertices) to the geometry shading stage 640.
[0247] The geometry shading stage 640 processes the geometric primitives by performing a set of operations (e.g., a geometry shader or program) on the geometric primitives. Tessellation operations may generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading stage 640 may subdivide each geometric primitive into a finer mesh of two or more geometric primitives for processing by the remainder of the graphics processing pipeline 600. The geometry shading stage 640 transmits the geometric primitives to the viewport SCC stage 650.
[0248] In one embodiment, the graphics processing pipeline 600 may operate within the streaming multiprocessors and the vertex shading stage 620, the primitive assembly stage 630, the geometry shading stage 640, the fragment shading stage 670, and / or associated hardware / software, and may perform processing operations sequentially. Once the sequential processing operations are complete, in one embodiment, the viewport SCC stage 650 may utilize the data. In one embodiment, the primitive data processed by one or more of the stages in the graphics processing pipeline 600 may be written into a cache (e.g., an L1 cache, a vertex cache, etc.). In such a case, in one embodiment, the viewport SCC stage 650 may access the data in the cache. In one embodiment, the viewport SCC stage 650 and the rasterization stage 660 are implemented as fixed function circuitry.
[0249] The viewport SCC stage 650 performs viewport scaling, culling, and clipping of the geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of the viewer watching the scene and defines a viewing frustum that encloses the objects in the scene. The viewing frustum may include a viewing plane, a back plane, and four clipping planes. Any geometric primitive that is completely outside the viewing frustum may be culled (e.g., discarded) because these geometric primitives will not contribute to the final rendered scene. Any geometric primitive that is partially inside and partially outside the viewing frustum may be clipped (e.g., transformed into a new geometric primitive that is enclosed within the viewing frustum). Additionally, each geometric primitive may be scaled based on the depth of the viewing frustum. Then all potentially visible geometric primitives are transmitted to the rasterization stage 660.
[0250] The rasterization stage 660 converts 3D geometric primitives into 2D fragments (e.g., can be used for display, etc.). The rasterization stage 660 can be configured to set a set of plane equations using the vertices of the geometric primitives, from which various attributes can be interpolated. The rasterization stage 660 can also calculate a coverage mask for multiple pixels, which indicates whether one or more sample positions of the pixels intercept the geometric primitive. In one embodiment, a z-test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have already been rasterized. The rasterization stage 660 generates fragment data (e.g., interpolated vertex attributes associated with specific sample positions of each covered pixel), which is sent to the fragment shading stage 670.
[0251] The fragment shading stage 670 processes the fragment data by performing a set of operations (i.e., fragment shader or program) on each of the fragments. The fragment shading stage 670 can generate pixel data (i.e., color values) for the fragments, such as by performing lighting operations or sampling texture maps using the interpolated texture coordinates of the fragments. The fragment shading stage 670 generates pixel data, which is sent to the raster operations stage 680.
[0252] The raster operations stage 680 can perform various operations on the pixel data, such as performing an alpha test, a stencil test, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the raster operations stage 680 has completed processing the pixel data (i.e., the output data 602), the pixel data can be written to a render target, such as a frame buffer, a color buffer, etc. The raster engine includes multiple fixed-function hardware units configured to perform various raster operations. In one embodiment, the raster engine includes a setup engine, a coarse raster engine, a culling engine, a clipping engine, a fine raster engine, and a tile aggregation engine. The setup engine receives the transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices. The plane equations are sent to the coarse raster engine to generate coverage information for the primitives (e.g., x, y coverage masks for tiles). The output of the coarse raster engine is sent to the culling engine, where fragments associated with primitives that fail the z-test are culled, and the non-culled fragments are sent to the clipping engine, where fragments located outside the view frustum are clipped off. Those fragments that remain after clipping and culling can be passed to the fine raster engine to generate the attributes of the pixel fragments based on the plane equations generated by the setup engine. The output of the raster engine includes, for example, fragments to be processed by the fragment shader implemented within the DPC.
[0253] It should be appreciated that one or more additional stages may be included in the graphics processing pipeline 600 in addition to or instead of one or more of the above stages. Various implementations of the abstract graphics processing pipeline may implement different stages. Additionally, in some embodiments, one or more of the above stages may be excluded from the graphics processing pipeline (such as the geometry shader stage 640). Other types of graphics processing pipelines are considered to be contemplated within the scope of the present disclosure. Additionally, any stage of the graphics processing pipeline 600 may be implemented by one or more dedicated hardware units within a graphics processor (such as a PPU). Other stages of the graphics processing pipeline 600 may be implemented by programmable hardware units (such as the SMs of a PPU).
[0254] The graphics processing pipeline 600 may be implemented via an application executed by a host processor (such as a CPU). In one embodiment, a device driver may implement an application programming interface (API) that defines various functions that may be utilized by an application to generate graphics data for display. A device driver is a software program that includes a plurality of instructions for controlling the operation of a PPU. The API provides an abstraction to a programmer that allows the programmer to utilize dedicated graphics hardware (such as a PPU) to generate graphics data without requiring the programmer to utilize the specific instruction set of the PPU. An application may include API calls that are routed to the device driver of the PPU. The device driver interprets the API calls and performs various operations in response to the API calls. In some cases, the device driver may perform operations by executing instructions on the CPU. In other cases, the device driver may perform operations by at least partially initiating operations on the PPU by utilizing an input / output interface between the CPU and the PPU. In one embodiment, the device driver is configured to utilize the hardware of the PPU to implement the graphics processing pipeline 600.
[0255] Various programs may be executed within the PPU to implement the various stages of the graphics processing pipeline 600. For example, the device driver may initiate a kernel on the PPU to execute the vertex shader stage 620 on one SM (or multiple SMs). The device driver (or an initial kernel executed by the PPU) may also initiate other kernels on the PPU to execute other stages of the graphics processing pipeline 600, such as the geometry shader stage 640 and the fragment shader stage 670. Additionally, some of the stages of the graphics processing pipeline 600 may be implemented on fixed unit hardware (such as a rasterizer or data assembler implemented within the PPU). It should be appreciated that the results from one kernel may be processed by one or more intermediate fixed function hardware units before being processed by subsequent kernels on the SMs.
[0256] The SM includes programmable streaming processors configured to process tasks represented by multiple threads. Each SM is multi-threaded and configured to concurrently execute multiple threads (e.g., 32 threads) from a particular thread group. In one embodiment, the SM implements a SIMD (Single Instruction, Multiple Data) architecture, where each thread in a thread group (e.g., a warp) is configured to process different data sets based on the same instruction set. All threads in the thread group execute the same instruction. In another embodiment, the SM implements a SIMT (Single Instruction, Multiple Threads) architecture, where each thread in a thread group is configured to process different data sets based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each warp, enabling concurrency between the warp and serial execution within the warp when threads within the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, enabling equal concurrency among all threads within and between warps. When the execution state is maintained for each individual thread, threads executing the same instruction can be coalesced and executed in parallel for maximum efficiency. A multi-level memory hierarchy is implemented. In one embodiment, a memory partitioning unit supports unified memory to provide a single unified virtual address space for CPU and PPU memories, enabling data sharing between virtual memory systems. In one embodiment, the access frequency of memory located on other processors by the PPU is tracked to ensure that memory pages are moved to the physical memory of the PPU that accesses the page more frequently. In one embodiment, NVLink supports address translation services, which allow the PPU to directly access the CPU's page table and provide full access by the PPU to CPU memory.
[0257] In one embodiment, a copy engine transfers data between multiple PPUs or between a PPU and a CPU. The copy engine can generate a page fault for an address not mapped to a page table. Then, the memory partitioning unit can service the page fault, map the address into the page table, after which the copy engine can perform the transfer. In a conventional system, for multiple copy engine operations between multiple processors, memory is fixed (e.g., non-pageable), which significantly reduces the available memory. Due to hardware page faults, an address can be passed to the copy engine without worrying about whether the memory page is resident and whether the copy process is transparent.
[0258] Data from the memory 62 or other system memory can be retrieved by the memory partitioning unit and stored in the L2 cache 66, which is on-chip and shared among the various GPCs. Each memory partitioning unit includes a portion of the L2 cache 66 associated with a corresponding memory device. Then, lower-level caches can be implemented in multiple units within the GPC. For example, each SM can implement a level-1 (L1) cache. The L1 cache is dedicated memory for a particular SM. Data from the L2 cache 66 can be fetched and stored in each L1 cache for processing in the functional units of the SM. The L2 cache 66 is coupled to the memory interface and the XBar.
[0259] The ROP unit performs graphics raster operations related to pixel colors, such as color compression, pixel blending, etc. The ROP unit also implements depth testing together with the raster engine, receiving the depth of the sample position associated with the pixel fragment from the culling engine of the raster engine. Test the depth of the sample position associated with the fragment relative to the corresponding depth in the depth buffer. If the fragment passes the depth test of the sample position, the ROP unit updates the depth buffer and sends the result of the depth test to the raster engine. It will be understood that the number of partitioning units can be different from the number of GPCs, and thus each ROP unit can be coupled to each GPC. The ROP unit tracks packets received from different GPCs and determines which GPC the results generated by the ROP unit 450 are routed to via the Xbar. Although the ROP unit is included within the memory partitioning unit, in other embodiments, the ROP unit can be outside the memory partitioning unit. For example, the ROP unit can reside in a GPC or another unit.
[0260] Each SM includes L processing cores. In one embodiment, the SM includes a large number (e.g., 128, etc.) of different processing cores. Each core can include fully pipelined, single-precision, double-precision, and / or mixed-precision processing units, which include a floating-point arithmetic logic unit and an integer arithmetic logic unit. In one embodiment, the floating-point arithmetic logic unit implements the IEEE
[0261] 754-2008 standard for floating-point operations. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0262] The tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included in a core. Specifically, the tensor cores are configured to perform deep learning matrix operations, such as convolutional operations for neural network training and inference. In one embodiment, each tensor core operates on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A·B + C, where A, B, C, and D are 4×4 matrices.
[0263] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating-point matrices, while the accumulation matrices C and D can be 16-bit floating-point or 32-bit floating-point matrices. The tensor cores operate on 16-bit floating-point input data and 32-bit floating-point accumulations. The 16-bit floating-point multiplication requires 64 operations to produce a full-precision product, which is then accumulated by adding with other intermediate products of the 4×4×4 matrix multiplication using 32-bit floating-point. In practice, the tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations built from these smaller elements. APIs (such as the CUDA9 C++ API) expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to effectively use the tensor cores from a CUDA-C++ program. At the CUDA level, the warp-level interface assumes a 16×16 size matrix spanning all 32 threads of a warp.
[0264] In some embodiments, transpose hardware is included in a processing core or another functional unit and is configured to generate matrix data stored diagonally and / or generate an original matrix and / or a transposed matrix from matrix data stored diagonally. The transpose hardware can be provided inside the shared memory to register the file load path of the SM.
[0265] In one example, matrix data stored diagonally can be fetched from DRAM and stored in the shared memory. When processing an instruction that performs processing using matrix data stored diagonally, the transpose hardware and register file set in the path of the shared memory can provide the original matrix, the transposed matrix, the compressed original matrix, and / or the compressed transposed matrix. Until the last store before the instruction, a single matrix data stored diagonally can be maintained, and the matrix types specified by the instruction can be generated in the register file as needed.
[0266] Each SM also includes M SFUs that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, an SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, an SFU may include a texture unit configured to perform texture mapping filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texels) from memory and sample the texture map to produce a sampled texture value for use in a shader program executed by the SM. In one embodiment, the texture map is stored in shared memory / L1 cache. The texture unit implements texture operations, such as filtering operations using mip mapping (e.g., texture maps at different levels of detail). In one embodiment, each SM includes two texture units.
[0267] Each SM also includes N LSUs that implement load and store operations between the shared memory / L1 cache and the register file. Each SM includes an interconnect network that connects each functional unit to the register file and connects the LSUs to the register file and the shared memory / L1 cache. In one embodiment, the interconnect network is a crossbar switch that can be configured to connect any functional unit to any register in the register file and connect the LSUs to memory locations in the register file and the shared memory / L1 cache.
[0268] The shared memory / L1 cache is an on-chip memory array that allows data storage and communication between the SM and the primitive engine and between threads in the SM. In one embodiment, the shared memory / L1 cache includes a storage capacity of 128 KB and is in the path from the SM to the partition unit. The shared memory / L1 cache 570 can be used for cache reads and writes. One or more of the shared memory / L1 cache, the L2 cache, and the memory are backup storage.
[0269] Combining the data cache and shared memory functions into a single memory block provides optimal overall performance for both types of memory access. This capacity can be used by the program as a cache when not using shared memory. For example, if the shared memory is configured to use half of its capacity, texture and load / store operations can use the remaining capacity. The integration within the shared memory / L1 cache enables the shared memory / L1 cache to act as a high-throughput pipeline for streaming data and at the same time provides high bandwidth and low latency access to frequently reused data.
[0270] The PPU can be included in a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In one embodiment, the PPU is included on a single semiconductor substrate. In another embodiment, the PPU is included on a system-on-chip (SoC) together with one or more other devices such as an additional PPU, a memory 62, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.
[0271] In one embodiment, the PPU can be included on a graphics card that includes one or more memory devices 62. The graphics card can be configured to interface with a PCIe slot on the motherboard of a desktop computer. In yet another embodiment, the PPU can be an integrated graphics processing unit (iGPU) or a parallel processor included in a chipset of the motherboard.
[0272] Example computing systems
[0273] Systems with multiple GPUs and CPUs are used in various industries because developers expose and utilize more parallelism in applications such as artificial intelligence computing. High-performance GPU-accelerated systems with dozens to thousands of computing nodes are deployed in data centers, research institutions, and supercomputers to solve larger problems. As the number of processing devices within a high-performance system increases, the communication and data transfer mechanisms need to scale to support the increased bandwidth.
[0274] In the context of this specification, a single semiconductor platform can refer to a unique single semiconductor-based integrated circuit fabricated on a die or chip. It should be noted that the term single semiconductor platform can also refer to a multi-chip module with increased connectivity that emulates on-chip operation and makes substantial improvements by leveraging conventional bus implementation methods. Of course, depending on the user's needs, various circuits or devices can also be placed separately or in various combinations of semiconductor platforms. Optionally, the parallel processing module can be implemented as a circuit board substrate, and each of the PPU and / or memory can be a packaged device. In one embodiment, the CPU, switch, and parallel processing module are located on a single semiconductor platform.
[0275] As Figure 3AAs shown, a system 50 is provided, which includes at least one central processing unit 56 connected to a communication bus. The communication bus can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or one or more point-to-point communication protocols. The system 50 also includes a main memory 62. Control logic (software) and data are stored in the main memory 62, and the main memory 62 can take the form of a random access memory (RAM).
[0276] The system 50 also includes an input device 54, a parallel processing system 56, and a display device 60, such as a conventional CRT (Cathode Ray Tube), LCD (Liquid Crystal Display), LED (Light Emitting Diode), plasma display, etc. User input can be received from the input device 54 (such as a keyboard, mouse, touchpad, microphone, etc.). Each of the foregoing modules and / or devices can even be located on a single semiconductor platform to form the system 50. Optionally, according to the user's needs, the individual modules can also be placed separately or in various combinations of semiconductor platforms.
[0277] In addition, the system 50 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) (such as the Internet), a peer-to-peer network, a cable network, etc.) for communication purposes through a network interface.
[0278] The system 50 can also include auxiliary storage (not shown). The auxiliary storage 610 includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a tape drive, an optical disk drive, a digital versatile disk (DVD) drive, a recording device, a universal serial bus (USB) flash drive. The removable storage drive reads and / or writes to the removable storage unit in a well-known manner.
[0279] A computer program or a computer control logic algorithm can be stored in the main memory 62 and / or the auxiliary storage. When these computer programs are executed, they enable the system 50 to perform various functions. The memory 62, the storage, and / or any other storage are possible examples of computer-readable media.
[0280] The architectures and / or functionality of the various prior figures can be implemented in the context of a general purpose computer system, a circuit board system, a gaming console system dedicated to entertainment purposes, a special purpose system, and / or any other desired system. For example, system 565 can take the form of a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, a cellular phone device, a television, a workstation, a gaming console, an embedded system, and / or any other type of logic.
[0281] All patents and printed publications mentioned above are hereby incorporated by reference as if set forth in full.
[0282] While the invention has been described in connection with what are presently considered to be the most practical and preferred embodiments, it is to be understood that the invention is not limited to the disclosed embodiments, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A system comprising: at least one memory storing (a) image element values representing a surface, and (b) surface memory compression information indicating a compression state of the stored image element values representing the surface; a processor operatively coupled to the at least one memory, the processor configured to read the surface memory compression information and use the read surface memory compression information to construct a value locality map of the surface, the value locality map indicating image element values that are similar or identical to each other; and a shader operatively coupled to the at least one memory and the processor, the shader configured to process the image element values representing the surface, the shader further configured to selectively process different image element values representing the surface in different ways based on the value locality map to reduce dynamic calculation redundancy when the shader processes the image element values.
2. The system of claim 1, wherein in response to the value locality map indicating image element values that are similar or identical to each other, the shader includes a dedicated execution path.
3. The system of claim 1, wherein the value locality map provides at least one bit for each tile of the surface, the at least one bit indicating whether a given surface tile has image element values that are similar or identical to other image element values.
4. The system of claim 1, wherein the value locality map (a) provides multiple bits for each region of the surface, the multiple bits indicating whether a given surface tile has image element values that are similar or identical to one of multiple possible image element values of the surface, and (b) provides at least one bit pattern value associated with each type of value locality and a bit pattern to convey a lack of locality of any of these values.
5. The system of claim 4, wherein the value locality map includes a coarse-grained value locality map that conveys whether a coarse-grained region of the surface contains image element values that are similar or identical to one or more values, or contains no values at all, or the coarse-grained region contains a mixture of tiles with and without similar or identical image element values.
6. The system of claim 1, wherein the shader selectively performs expression reconstruction based on the value locality map indicating similar or identical image element values.
7. The system of claim 1, wherein the surface memory compression information is stored in the L2 cache of the processor.
8. The system of claim 1, wherein the surface memory compression information includes zero-bandwidth clear data, and the shader is configured to selectively process image element values in response to the zero-bandwidth clear data.
9. The system of claim 1, wherein the surface memory compression information includes reduced compression data, and the shader is configured to selectively process image element values in response to the reduced compression data.
10. The system of claim 1, wherein The surface memory compressed information includes differential compression, and the shader is configured to selectively process image element values in response to the differential compression.
11. The system according to claim 1, wherein, the processor uses a driver to read the surface memory compressed information and construct the value locality map.
12. The system according to claim 1, wherein the shader is compiled to selectively trigger specialized execution for display using the value locality map.
13. The system according to claim 1, wherein, the processor causes the surface memory compressed information to be read by user mode driver software using a memory load operation to assist in constructing the value locality map of the surface.
14. The system according to claim 1, wherein, tile query processing is performed to combine value localities of surfaces from different sizes and / or resolutions into a common value locality map.
15. The system according to claim 1, wherein, the shader is multi-threaded, and based on the value locality map, multiple threads share and reuse computations.
16. A method, comprising: reading surface memory compressed information from at least one memory that stores (a) image element values representing a surface, and (b) surface memory compressed information indicating a compression state of the stored image element values representing the surface; using the read surface memory compressed information to construct a value locality map of the surface, the value locality map indicating image element values that are similar or identical to each other; and processing the image element values representing the surface using a shader, the processing including selectively processing different image element values representing the surface in different ways according to the value locality map, thereby reducing shader dynamic computation redundancy.
17. The method according to claim 16, wherein, the shader includes a dedicated execution path activated in response to the value locality map indicating an image element value that is similar or identical to other image element values of the surface.
18. The method according to claim 16, wherein, the value locality map provides at least one bit for each surface tile of the surface, the at least one bit indicating whether the surface tile has image element values that are similar or identical to each other.
19. The method according to claim 16, wherein, the value locality map: (a) provides multiple bits for each region of the surface, the multiple bits indicating that one of a plurality of possible values has a plurality of similar or identical occurrence rates; and (b) provides at least one bit pattern value associated with each type of value locality and a bit pattern to convey a lack of locality.
20. The method according to claim 19, wherein, the value locality map includes a coarse-grained value locality map that conveys whether a coarse-grained region of the surface contains image element values that are similar or identical to one or more values, or contains no values at all, or the coarse-grained region contains a mixture of tiles with and without similar or identical image element values.
21. The method according to claim 16, wherein, the shader selectively performs expression reconstruction based on the value locality mapping indicating similar or identical image element values.
22. The method according to claim 16, further comprising: reading the surface memory compression information from the L2 cache of the processor.
23. The method according to claim 16, wherein, the surface memory compression information includes zero-bandwidth clear data, and the processing includes selectively processing the image element values in response to the zero-bandwidth clear data.
24. The method according to claim 16, wherein, the surface memory compression information includes reduced compression data, and the processing includes selectively processing the image element values in response to the reduced compression data.
25. The method according to claim 16, wherein, the surface memory compression information includes differential compression, and the processing includes selectively processing the image element values in response to the differential compression.
26. The method according to claim 16, further comprising: using a driver to read the surface memory compression information and construct the value locality mapping.
27. The method according to claim 16, further comprising: compiling the shader to selectively trigger specialized execution using the value locality mapping.
28. The method according to claim 16, further comprising: reading the surface memory compression information by a user-mode driver software using a memory load operation to assist in constructing the value locality mapping of the surface.
Citation Information
Patent Citations
Zero-bandwidth clears
US8330766B1
Efficient line and page organization for compression status bit caching
US20110087840A1
Caching of adaptively sized cache tiles in a unified l2 cache with surface compression
US20140118379A1
Method and apparatus for processing compressed texture
US20170084055A1
Per-Sample MSAA Rendering Using Comprehension Data
US20170287209A1