Delta Triplet Index Compression

The method compresses index streams in computer graphics by using delta value triplets and a lookup table to identify common patterns, thereby reducing storage and bandwidth needs while maintaining efficient rendering performance.

JP7676447B2Active Publication Date: 2025-05-14ADVANCED MICRO DEVICES INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022577276
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-26
Filing Date
2021-06-01
Publication Date
2025-05-14
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

Existing methods for compressing index streams in computer graphics are inefficient, particularly in reducing storage and bandwidth requirements while maintaining effective rendering performance.

Method used

The proposed solution involves a method for compressing index streams using delta value triplets, where the delta values are compared to a lookup table to determine if they match common patterns. If a match is found, the index triplet is compressed using a pattern ID; otherwise, variable length encoding is applied to the delta values.

Benefits of technology

This approach effectively reduces the storage and bandwidth requirements for index streams by leveraging temporal locality in index streams, allowing for efficient compression and decompression of index data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676447000002
    Figure 0007676447000002
  • Figure 0007676447000003
    Figure 0007676447000003
  • Figure 0007676447000004
    Figure 0007676447000004
Patent Text Reader

Abstract

A method, device, and system are provided for compressing and decompressing an index stream associated with a graphics primitive. A delta value group is determined based on an index group of the index stream. The delta value group is compared to delta values ​​in a lookup table. The index group is compressed based on an entry in the lookup table if the delta value group matches all delta values ​​of the entry, and is otherwise compressed based on variable length encoding.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 042,384, filed June 22, 2020, and U.S. Patent Application No. 17 / 187,625, filed February 26, 2021, which are incorporated by reference as if fully set forth herein. [Background technology]

[0002] In computer graphics, objects are typically represented in three-dimensional (3D) space as a group of two-dimensional (2D) polygons, which are typically referred to as primitives in this context. The polygons are typically triangles, each with three vertices. Other types of polygon primitives may be used, but triangles are typically the most common. Each vertex contains information that defines its position in 3D space, and typically includes other information, such as color, normal vectors, and / or texture information.

[0003] A vertex can be part of more than one triangle (or other primitive). For example, two triangles with a common edge share two vertices. Multiple triangles sharing various common edges to describe an object are in some cases referred to as a mesh or triangle mesh. To reduce the amount of information needed to represent the vertices, it is typically useful to reference each vertex by an index value rather than referencing its entire 3D coordinates, color and / or other information. Index data is widely used to specify primitive connectivity.

[0004] In one example, the first triangle in the stream contains three vertices referenced by indices 0, 1, and 2. The next triangle in the stream adjacent to the first triangle contains three vertices referenced by indices 1, 2, and 3. Here, the triangles share a common edge, which is defined by the vertices corresponding to indices 1 and 2. Thus, the two triangles are represented by only four unique vertices, and the vertices indexed by 1 and 2 are reused. The indices are always integer values.

[0005] A more detailed understanding may be had from the following description, given by way of example in conjunction with the accompanying drawings, in which: [Brief description of the drawings]

[0006] [Figure 1] 1 is a block diagram of an example device in which one or more features of the present disclosure may be implemented. [Diagram 2] 2 is a block diagram of the device of FIG. 1 showing additional details. [Diagram 3] FIG. 2 is a block diagram illustrating a graphics processing pipeline, according to an example. [Figure 4] FIG. 2 is a flow diagram illustrating an example triplet-based delta compression of an example index stream. [Diagram 5] FIG. 13 is a flow diagram illustrating an example triplet-based delta compression of an example index stream for an example case. [Figure 6] FIG. 11 is a flow diagram illustrating an example triplet-based delta compression of an example index stream for another example case. [Figure 7] FIG. 11 is a flow diagram illustrating an example triplet-based delta compression of an example index stream for another example case. [Figure 8] 4 is a flowchart illustrating an example process for compressing an example index stream. [Figure 9] 4 is a flowchart illustrating an example process for decompressing an example compressed index stream. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0007] Some embodiments provide a method for compressing an index stream associated with a graphics primitive. A delta value group is determined based on an index group of the index stream. The delta value group is compared to delta values ​​in a lookup table. The index group is compressed based on an entry in the lookup table if the delta value group matches all delta values ​​in the entry, otherwise the index group is compressed based on a variable length encoding.

[0008] In some embodiments, if a delta value group includes at least one value that matches a delta value in an entry in the lookup table and includes at least one value that does not match any delta value in the entry, and the at least one value that does not match any delta value in the entry corresponds to a don't-care value in the entry, the index group is compressed based on the entry and based on a variable length encoding of the at least one value that does not match any delta value in the entry.

[0009] In some embodiments, if the delta value group does not include at least one value that matches a delta value in an entry in the lookup table, the index group is compressed based on an entry in the lookup table that includes only don't care values ​​and based on a variable length encoding of each of the delta values. In some embodiments, the index group is a triplet and the graphics primitive is a triangle. In some embodiments, the index stream indexes the vertices of a mesh.

[0010] Some embodiments provide a method for decompressing a compressed index stream associated with a graphics primitive. Compressed index groups of the compressed index stream are compared to entries of a lookup table. If the compressed index groups match lookup table entries that include delta values ​​corresponding to each of the compressed index groups, the compressed index groups are decompressed based on the delta values, otherwise the compressed index groups are decompressed based on variable length decoding.

[0011] In some embodiments, if a compressed index group includes at least one delta value corresponding to at least one of the compressed index groups and matches a lookup table entry that includes at least one don't care value corresponding to at least one of the compressed index groups, the compressed index group is decompressed based on the at least one delta value and based on variable length decoding of the delta value corresponding to the at least one don't care value.

[0012] In some embodiments, if the compressed index group matches a lookup table entry that contains only don't care values, the compressed index group is decompressed based on a variable length decoding of the delta values ​​corresponding to each of the compressed index groups. In some embodiments, the compressed index group is a triplet and the graphics primitive is a triangle. In some embodiments, the compressed index stream indexes the vertices of a mesh.

[0013] Some embodiments provide a compressor configured to compress an index stream associated with a graphics primitive, the compressor including circuitry configured to determine a delta value group based on an index group of the index stream, the compressor including circuitry configured to compare the delta value group to delta values ​​in a lookup table, and the compressor including circuitry configured to compress the index group based on the entry if the delta value group matches all delta values ​​in an entry in the lookup table, and otherwise compress the index group based on a variable length encoding.

[0014] In some embodiments, the compressor includes circuitry configured to compress the index group based on an entry in the lookup table, where the delta value group includes at least one value that matches a delta value in an entry in the lookup table and includes at least one value that does not match any delta value in the entry, and where the at least one value that does not match any delta value in the entry corresponds to a don't care value in the entry, and based on a variable length encoding of the at least one value that does not match any delta value in the entry.

[0015] In some embodiments, the compressor includes circuitry configured to compress the index group based on an entry in the lookup table that includes only don't care values ​​if the delta value group does not include at least one value that matches a delta value in an entry in the lookup table, and based on a variable length encoding of each of the delta values. In some embodiments, the index group is a triplet and the graphics primitive is a triangle. In some embodiments, the index stream indexes vertices of a mesh.

[0016] Some embodiments provide a decompressor configured to decompress a compressed index stream associated with a graphics primitive. The decompressor includes circuitry configured to compare compressed index groups of the compressed index stream with entries of a lookup table. The decompressor also includes circuitry configured to decompress the compressed index groups based on the delta values ​​if the compressed index groups match lookup table entries that include delta values ​​corresponding to each of the compressed index groups, and to decompress the compressed index groups based on variable length decoding if not.

[0017] In some embodiments, the decompressor includes circuitry configured to decompress the compressed index group if the compressed index group matches a lookup table entry that includes at least one delta value corresponding to at least one of the compressed index groups and that includes at least one don't care value corresponding to at least one of the compressed index groups based on the at least one delta value and based on variable length decoding of the delta value corresponding to the at least one don't care value.

[0018] In some embodiments, the decompressor includes circuitry configured to decompress the compressed index groups based on variable length decoding of delta values ​​corresponding to each of the compressed index groups if the compressed index groups match lookup table entries that include only don't care values. In some embodiments, the compressed index groups are triplets and the graphics primitives are triangles. In some embodiments, the compressed index stream indexes vertices of a mesh.

[0019] Some embodiments provide a system, device, and method for compressing indices in an index stream representing vertices in a triangular mesh. The index stream is input, and a delta value triplet is calculated based on index value triplets from the index stream. If the delta value triplet corresponds to a pattern identifier (ID) in a common pattern lookup table, the index value triplet is represented by the pattern ID. If the delta value triplet does not correspond to any pattern ID in the common pattern lookup table, the index value triplet is represented by the delta value triplet.

[0020] In some embodiments, if the delta value triplet does not correspond to any pattern ID in the common pattern lookup table, the delta triplet is encoded and the index value triplet is represented by the encoded delta triplet. In some embodiments, the encoded delta value triplet is encoded using variable length coding.

[0021] 1 is a block diagram of an example device 100 capable of implementing one or more features of the present disclosure. Device 100 may include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. Device 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Device 100 may also optionally include an input driver 112 and an output driver 114. It should be understood that device 100 may include additional components not shown in FIG. 1.

[0022] In various alternatives, the processor 102 may include a central processing unit (CPU), a graphics processing unit (GPU), a CPU and a GPU located on the same die, or one or more processor cores, each of which may be a CPU or a GPU. In various alternatives, the memory 104 may be located on the same die as the processor 102 or may be located separately from the processor 102. The memory 104 may include volatile or non-volatile memory (e.g., random access memory (RAM), dynamic RAM, cache).

[0023] Storage devices 106 include fixed or removable storage devices (e.g., hard disk drives, solid state drives, optical disks, flash drives). Input devices 108 include, but are not limited to, a keyboard, a keypad, a touch screen, a touch pad, a detector, a microphone, an accelerometer, a gyroscope, a biometric scanner, or a network connection (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals). Output devices 110 include, but are not limited to, a display, a speaker, a printer, a haptic feedback device, one or more optical, antennae, or network connections (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0024] The input driver 112 communicates with the processor 102 and the input device 108, allowing the processor 102 to receive input from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110, allowing the processor 102 to send output to the output device 110. It should be noted that the input driver 112 and the output driver 114 are optional components, and that the device 100 can operate in the same manner without the input driver 112 and the output driver 114 being present. The output driver 114 includes an accelerated processing device ("APD") 116 coupled to a display device 118. The APD accepts computational and graphics rendering commands from the processor 102, processes the computational and graphics rendering commands, and provides pixel output to the display device 118 for display. As described in more detail below, the APD 116 includes one or more parallel processing units that perform computations according to a single-instruction-multiple-data ("SIMD") paradigm. Thus, although various functions are described herein as being performed by or in conjunction with APD 116, in various alternatives, the functions described as being performed by APD 116 may additionally or alternatively be performed by other computing devices having similar capabilities that are not driven by a host processor (e.g., processor 102) to provide graphical output to display device 118. For example, it is contemplated that any processing system that performs processing tasks according to the SIMD paradigm may perform the functions described herein. Alternatively, it is contemplated that computing systems that do not perform processing tasks according to the SIMD paradigm may perform the functions described herein.

[0025] FIG. 2 is a block diagram of device 100 showing additional details regarding the execution of processing tasks on APD 116. Processor 102 maintains, in system memory 104, one or more control logic modules for execution by processor 102. The control logic modules include operating system 120, kernel mode driver 122, and applications 126. These control logic modules control various aspects of the operation of processor 102 and APD 116. For example, operating system 120 communicates directly with hardware and provides an interface to the hardware for other software executing on processor 102. Kernel mode driver 122 controls the operation of APD 116, for example, by providing an application programming interface ("API") to software executing on processor 102 (e.g., applications 126) to access various features of APD 116. Kernel mode driver 122 also includes a just-in-time compiler that compiles programs for execution by processing components of APD 116 (such as SIMD unit 138, described in more detail below).

[0026] APD 116 executes commands and programs for selected functions, such as graphic and non-graphic operations that may be suitable for parallel processing. APD 116 may be used to perform graphics pipeline operations, such as pixel operations, geometry calculations, and rendering of images to display device 118, based on commands received from processor 102. APD 116 also performs computational operations not directly related to graphic operations, such as operations related to video, physics simulations, computational fluid dynamics, or other tasks, based on commands received from processor 102.

[0027] The APD 116 includes a computation unit 132 that includes one or more SIMD units 138 that perform operations in a parallel manner according to the SIMD paradigm at the request of the processor 102. The SIMD paradigm is one in which multiple processing elements share a single program control flow unit and program counter, and thus execute the same program, but can execute the program on different data. In one example, each SIMD unit 138 includes 16 lanes, each lane executes the same instruction simultaneously with other lanes in the SIMD unit 138, but can execute the instruction on different data. Lanes can be switched off with predication when not all lanes need to execute a given instruction. Predication can also be used to execute programs with branching control flows. More specifically, for programs with conditional branches or other instructions where the control flow is based on calculations made by individual lanes, the predication of lanes corresponding to the control flow paths that are not currently executed, and the serial execution of the different control flow paths allows for arbitrary control flow.

[0028] The basic unit of execution within the compute unit 132 is the work item. Each work item represents a single instantiation of a program that executes in parallel on a particular lane. Work items can execute simultaneously as a "wavefront" on a single SIMD unit 138. One or more wavefronts are included in a "workgroup," which includes a collection of work items designated to execute the same program. A workgroup can be executed by executing each of the wavefronts that make up the workgroup. In the alternative, a wavefront executes sequentially on a single SIMD unit 138, or partially or fully in parallel on different SIMD units 138. A wavefront can be thought of as the largest collection of work items that can execute simultaneously on a single SIMD unit 138. Thus, if a command received from the processor 102 indicates that a particular program is parallelized to an extent that it cannot be run simultaneously on a single SIMD unit 138, the program is split into wavefronts that are either parallelized on two or more SIMD units 138, or serialized on the same SIMD unit 138 (or both parallelized and serialized, as appropriate). The scheduler 136 performs operations related to scheduling the various wavefronts on the different compute units 132 and SIMD units 138 .

[0029] The parallel processing provided by the computation units 132 is well suited to graphics-related operations, such as pixel value calculations, vertex transformations, and other graphics operations. Thus, in some cases, the graphics processing pipeline 134, which accepts graphics processing commands from the processor 102, provides computation tasks to the computation units 132 for execution in parallel.

[0030] Computation units 132 are also used to perform computational tasks that are not related to graphics or that are not performed as part of the "normal" operation of graphics processing pipeline 134 (e.g., custom operations performed to supplement the operations performed on graphics processing pipeline 134). Applications 126 or other software executing on processor 102 send programs defining such computational tasks to APD 116 for execution.

[0031] Figure 3 is a block diagram illustrating additional details of the graphics processing pipeline 134 shown in Figure 2. The graphics processing pipeline 134 includes stages, each of which performs a particular function. The stages represent subdivisions of the functionality of the graphics processing pipeline 134. Each stage is implemented partially or fully as a shader program executing within the programmable processing unit 202, or partially or fully as fixed-function, non-programmable hardware external to the programmable processing unit 202.

[0032] The input assembler stage 302 reads user-filled buffers (e.g., buffers filled with requests from software executed by the processor 102, such as applications 126) and assembles the data into primitives to be used by the rest of the pipeline. The input assembler stage 302 can generate different types of primitives based on the primitive data contained in the user-filled buffers. The input assembler stage 302 formats the assembled primitives for use by the rest of the pipeline.

[0033] The vertex shader stage 304 processes the vertices of the primitives assembled by the input assembler stage 302. The vertex shader stage 304 performs various per-vertex operations such as transformation, skinning, morphing, and per-vertex lighting. Transformation operations include various operations for transforming the coordinates of vertices. These operations include one or more of modeling transformations, view transformations, projection transformations, perspective division, and viewport transformations. As used herein, such transformations are considered to change the coordinates or "position" of the vertices at which the transformation occurs. Other operations of the vertex shader stage 304 modify attributes other than coordinates.

[0034] The vertex shader stage 304 is implemented partially or completely as a vertex shader program that executes on one or more compute units 132. The vertex shader program is provided by the processor 102 and is based on a program pre-written by a computer programmer. The driver 122 compiles such a computer program to generate a vertex shader program having a format suitable for execution within the compute units 132.

[0035] The hull shader stage 306, the tessellator stage 308, and the domain shader stage 310 work together to implement tessellation, which converts simple primitives into more complex primitives by subdividing the primitives. The hull shader stage 306 generates a patch for tessellation based on the input primitive. The tessellator stage 308 generates a sample set for the patch. The domain shader stage 310 calculates vertex positions for vertices that correspond to the samples of the patch. The hull shader stage 306 and the domain shader stage 310 can be implemented as shader programs executing on the programmable processing unit 202.

[0036] The geometry shader stage 312 performs vertex operations on a primitive basis. A variety of different types of operations can be performed by the geometry shader stage 312, including operations such as point sprint expansion, dynamic particle system operations, fur-fin generation, shadow volume generation, single pass render-to-cubemap, per-primitive material swapping, and per-primitive material setup. In some cases, a shader program executing on the programmable processing unit 202 performs the operations of the geometry shader stage 312.

[0037] The Rasterizer stage 314 accepts and rasterizes the simple primitives generated upstream. Rasterization consists of determining which screen pixels (or sub-pixel samples) are covered by a particular primitive. Rasterization is performed by fixed function hardware.

[0038] The pixel shader stage 316 calculates output values ​​for screen pixels based on the upstream generated primitives and the results of rasterization. The pixel shader stage 316 can apply textures from texture memory. The operations of the pixel shader stage 316 are performed by shader programs executing on the programmable processing unit 202.

[0039] The output merge stage 318 accepts the outputs from the pixel shader stage 316, merges them, and performs operations such as z-testing and alpha blending to determine the final color of the screen pixel.

[0040] Texture data defining textures is stored and / or accessed by texture unit 320. Textures are bitmap images that are used at various points in graphics processing pipeline 134. For example, in some cases, pixel shader stage 316 applies textures to pixels to improve the apparent rendering complexity (e.g., to provide a more "photorealistic" look) without increasing the number of vertices rendered.

[0041] In some cases, the vertex shader stage 304 uses texture data from the texture unit 320 to modify primitives to increase complexity, for example, by generating or modifying vertices for improved aesthetics. In one example, the vertex shader stage 304 uses a height map stored in the texture unit 320 to modify the displacement of vertices. This type of technique can be used to generate more realistic looking water, for example, by modifying the position and number of vertices used to render the water, as compared to textures used only by the pixel shader stage 316. In some cases, the geometry shader stage 312 accesses texture data from the texture unit 320.

[0042] Index data is widely used in computer graphics to specify primitive connectivity. For example, an object is typically specified as a stream of indexes, starting with a starting vertex, then proceeding to other vertices of the first polygon, and so on through the indices of adjacent polygons that share a common vertex, until the object is fully defined.

[0043] In some cases, it may be desirable to compress the index stream. For example, in "traditional" dual-pass rendering (e.g., separate z-pass and rendering passes controlled by software), it may be advantageous to store primitives that survive the z-pass as a compressed index stream. In some such cases, the compressed index stream can be used as index data during the rendering pass instead of the original index buffer. Exemplary cases where this technique may be applicable include when the geometry of the z-pass and the rendering pass are the same.

[0044] In some cases, it is desirable to compress the index stream to reduce storage and bandwidth used in processing the indexes. An index buffer may be reused without modification. However, in some cases, several versions of an index buffer are used. In a landscape rendering example, multiple index buffers access landscape vertices at different levels of detail (LOD). In another example, multiple index buffers are used for different viewing angles, and landscape features that cannot be seen from certain angles are removed from some of the index buffers. Therefore, in some such cases, it is advantageous to compress multiple index buffers to save storage space and / or memory bandwidth in some situations.

[0045] In some cases, index data is streamed in the graphics pipeline from the start of geometry processing to later hardware units in the pipeline. In some such cases, a normal buffer for indexes (e.g., a first-in-first out (FIFO) buffer) would require a significant amount of hardware to store the index data. As such, in some such cases, it is desirable to compress the index stream (e.g., using triplet compression as discussed herein) to reduce the amount of storage used to store the index stream. In some such cases, a parallel variant of compression can be used (i.e., compressing or decompressing multiple primitives in parallel for greater primitive throughput).

[0046] Because the numerical values ​​of successive indices in a stream are typically fairly close in number due to spatial locality in the way they are typically constructed, index streams typically exhibit substantial temporal locality (i.e., the numbers are close in value, such as within a threshold of each other). For example, considering one index in a stream, the surrounding indices (e.g., the previous and / or subsequent indexes, such as the immediately preceding and / or immediately following indexes, etc.) typically have close (e.g., within a threshold number) values. In some embodiments, this property provides compressibility that is exploited to reduce geometry data size. In some embodiments, it is possible to further compress the index stream by compressing the index patterns in the stream.

[0047] In computer graphics, a variety of primitives are used. For example, it is possible to render a line defined by two vertices, or a quadrilateral defined by four vertices. In some embodiments, the most common primitive is a triangle, which is represented by three vertices. Therefore, for compression, it may be desirable to split the index stream representing the vertices of a triangle primitive into triplets (i.e., groups or indices of three vertices). It is noted that the techniques, devices, and systems herein are applicable to other primitives, e.g., by splitting the index stream into groups of other numbers of vertices. For example, it may be desirable to decompose the index stream representing the vertices of a quadrilateral primitive into quadruplets (i.e., groups or indices of four vertices) for compression. It is noted that while the examples herein are described with respect to the triangle / triplet case, accommodation for different numbers of vertices is contemplated, e.g., by splitting the index stream into groups of other numbers of vertices. Furthermore, it is noted that special cases arise in the index stream. For example, a triangle mesh is typically aligned as a triplet. However, some embodiments include patch primitives. In some cases, patch primitives include more than three indices per primitive. Additionally, some embodiments include point primitives (one index per primitive) and line primitives (two indices per primitive) that are not aligned by triplets. So, in an example where a mesh includes 7 primitives with 5 indices each, the total number of indices is 35. The 35 indices can be represented as 11 triplets (=33 indices) and 2 index entries at the end to match up to 35. Alternatively, the data is represented such that each primitive includes a single triplet and another 2 index entries, so that the data is aligned by 5. Now, each primitive has 5 indices and there are 7 primitives.So the data can be stored as 7x(3+2), i.e. each primitive is stored as a triplet and a group containing only two indices. Alternatively, the data can be stored as a long sequence of triplets, with the last entry matching the amount of data with the number of entries stored (e.g. 11x3+2, 11 triplets and one extra slot at the end).

[0048] Because each vertex is indexed by a different value, numbers typically are not repeated in the index stream. However, the underlying patterns in the indices of a primitive are comparable across primitives. For example, delta values ​​(e.g., difference values) between indices and / or between indices in a primitive may be comparable to, and similar to, delta values ​​between indices and / or between indices in another primitive, even if the index values ​​are different. In this way, in some embodiments, common patterns are identifiable in the index stream.

[0049] Therefore, in some embodiments, the index stream is compressed based on delta values ​​within and / or between index triplets (in the case of triangle primitives or in the case of quad primitives, within and / or between index quadruplets).

[0050] In some embodiments, the compressed delta triplet stream is used in various operations as a compressed version of the index stream. In some embodiments, the compressed delta triplet stream is used as a form of visibility. For example, in some embodiments, the compressed delta triplet values ​​are used in place of index values ​​in a buffer (e.g., a "visibility buffer") that tracks which triangles are visible in a bin, tile, or other portion of the rendered screen. In some embodiments, using compressed index data is advantageous for per-bin visibility information. For example, in some cases, it may be advantageous to store per-primitive information so that when rendering a bin, primitives that do not contribute to the bin in question do not need to be processed. It is not necessary to fetch visibility flags by building a compressed index stream and then sparse fetch to the index buffer based on the fetched visibility flags, because the compressed index stream provides the necessary information directly. This can have the advantage of reducing bandwidth and latency when rendering primitives. One way to store visibility information is to compute a bit for each primitive indicating whether the primitive is visible and store this information in memory in either compressed or uncompressed form. Using this approach, in some cases the process for fetching an index requires that the visibility information is fetched first, followed by a dependent fetch to index data based on the visibility bit. In some embodiments, this approach doubles the overall fetch latency due to the requirement of two data fetch passes. In some embodiments, there will also be over-fetching of index data since memory fetches are typically done in blocks of several hundred bits, while the dependent fetch requires only a subset of the data in the index buffer, such that some portions of the fetched data are not used as indexes for visible primitives.Therefore, some embodiments generate compressed index data in memory that contains only indexes of visible primitives, which results in all fetched data being used. In some such embodiments, more data is generated for the compressed index stream when writing visibility information, but less information is read using the compressed index stream, using more favorable access patterns.

[0051] FIG. 4 is a flow diagram 400 illustrating an exemplary triplet-based delta compression of an exemplary index stream. Flow diagram 400 illustrates a technique for compressing an index stream that is triplet-based. The technique is optimized for the most common primitive type, which is a triangle having three vertices. As such, in this technique, delta values ​​(e.g., difference values) are calculated within triplets of index values ​​from the index stream to generate delta triplets for the triplets of index values, and a lookup table captures the most common delta triplets of the stream. The index triplets that correspond to the delta triplets captured in the lookup table (i.e., the most common patterns or delta triplets in the stream) are represented by a pattern identifier (ID) associated with an entry in the lookup table. The index triplets that correspond to delta triplets that are not captured in the lookup table are stored or encoded in another manner, such as by applying a variable length coding to the delta triplet and representing the index triplet by the coded value.

[0052] In some embodiments, this technique results in an advantageously compressed representation of the index stream, which in some embodiments is due to temporal locality in the index stream that results in a relatively small number (e.g., below a threshold number or percentage) of patterns occurring in the delta stream, such that a relatively large number (e.g., above a threshold number or percentage) of delta triplets match entries in the common pattern lookup table.

[0053] More specifically, flow diagram 400 shows an index stream that includes six indexes split into two index triplets 402, 404. Index triplet 402 includes indexes 406, 408, 410, which occupy positions N-3, N-2, and N-1, respectively. Index triplet 404 includes indexes 412, 414, 416, which occupy positions N, N+1, and N+2, respectively.

[0054] In some embodiments, to compress index triplet 404, three delta values ​​418, 420, 422 corresponding to indexes 412, 414, 416, respectively, are calculated based on the values ​​of indexes 412, 414, 416. For example, delta value 418 is determined as (index 412)-(index 410), i.e., the difference between the value of index 412 and the value of index 410. Delta value 420 is calculated as (index 414)-(index 412), i.e., the difference between the value of index 414 and the value of index 412. Delta value 422 is calculated as (index 416)-(index 412), i.e., the difference between the value of index 416 and the value of index 412. Note that one of the three delta values ​​(delta value 418 in this example) is based on an index value from the previous triplet, while the other two delta values ​​(delta value 420 and delta value 422 in this example) are calculated from index values ​​within the triplet. The three delta values ​​418, 420, 422 are referred to as delta triplet 424 for convenience.

[0055] After the value of the delta triplet 424 is determined, the delta triplet 424 is compressed. To compress the delta triplet, the three delta values ​​418, 420, 422 are compared to values ​​in a common pattern lookup table 426. The common pattern lookup table 426 may be implemented in any suitable manner, such as a dedicated register in the graphics hardware, and may be of any suitable size. In this example, the common pattern lookup table 426 stores delta values ​​for 15 patterns. In some embodiments, the number of entries in the common pattern lookup table is based on a desired lookup speed, die area, and / or design complexity. Each of the patterns in the common pattern lookup table is associated with a pattern identifier (ID).

[0056] In some embodiments, if the delta values ​​418, 420, 422 in the delta triplet 424 (in order) match a common pattern in the common pattern lookup table 426, then the index triplet 404 is compressed into a compressed representation 428 based on the pattern ID 430 that corresponds to the common pattern that matches the delta triplet 424. In some embodiments, the delta values ​​418, 420, 422 are not part of the compressed representation 428 of the index triplet 404 (apart from their representation in the pattern ID 430).

[0057] In some embodiments, if the delta values ​​418, 420, 422 in delta triplet 404 (in order) do not match any common patterns in common pattern lookup table 426, then index triplet 424 is represented in another suitable manner compressed representation 428. For example, in some embodiments, delta values ​​418, 420, 422 in delta triplet 424 are encoded in any suitable manner, such as variable length coding, and index triplet 404 is compressed into compressed representation 428 based on the encoded delta values ​​432.

[0058] In some embodiments, in this case, delta values ​​that do not match any common pattern in common pattern lookup table 426 are handled by matching them to entries in common pattern lookup table 426 where each delta value is a "don't care" value (i.e., matches any delta value). In some embodiments, index triplet 404 is compressed into a compressed representation 428 based on a combination of the pattern ID 430 and encoded delta value 432 that corresponds to this particular entry in common pattern lookup table 426.

[0059] In some embodiments, if at least one but less than all (i.e., one or two in the case of a triplet) of delta values ​​418, 420, 422 in delta triplet 424 (in order) match a common pattern in common pattern lookup table 426, index triplet 404 is compressed into a compressed representation 428 in another suitable manner, such as based on a combination of the pattern ID and the encoded delta value.

[0060] For example, in some embodiments, the common pattern lookup table 426 includes at least one common pattern, where the triplet includes one or more particular delta values ​​and one or more "don't care" values. In such cases, one or more delta values ​​of the index triplet 404 that match a specified value can be represented in the compressed representation 428 based on a pattern ID 430 corresponding to the common pattern, and one or more delta values ​​of the index triplet 404 that correspond to the one or more "don't care" values ​​are encoded in any suitable manner, such as variable length coding, such that the index triplet is represented by a combination of the pattern ID 430 corresponding to the common pattern and one or more encoded delta values ​​432.

[0061] In some embodiments, the common pattern lookup table 426 is pre-populated (e.g., statically). In some embodiments, the common pattern lookup table 426 is fixed or configurable. In some embodiments, the pattern lookup table 426 is pre-populated with values ​​determined by analyzing a set of training data (e.g., a set of data collected from application traces (e.g., empirically, i.e., from "real-world" application traces)).

[0062] In some embodiments, the common pattern lookup table 426 is populated during compression of the index stream by tracking the frequency of observed delta triplet values. In some embodiments, after the lookup table becomes full (e.g., after 14 delta triplets or patterns are stored in a common pattern lookup table with 14 entries), the least frequent delta triplet values ​​or patterns are evicted from the common pattern lookup table if a different delta triplet value or pattern is observed by the common pattern lookup table hardware and / or software and occurs more frequently in the delta stream.

[0063] FIG. 5 is a flow diagram 500 illustrating an example triplet-based delta compression of an example index stream for an example case in which all of the delta values ​​of an example index triplet match a particular delta value of an entry in a common pattern lookup table.

[0064] More specifically, flow diagram 500 shows an index stream that includes six indexes split into two index triplets 502, 504. Index triplet 502 includes indexes 506, 508, and 510. Index 510, in this example, has a value of 999. The values ​​of indexes 506 and 508 are not used in this exemplary compression and are simply shown in the figure as "D." Index triplet 504 includes indexes 512, 514, and 516, which have values ​​of 1000, 1001, and 1002, respectively.

[0065] In this example, to compress index triplet 504, three delta values ​​518, 520, 522 corresponding to indexes 512, 514, 516, respectively, are calculated based on the values ​​of indexes 512, 514, 516. For example, delta value 518 is determined as (index 512)-(index 510), i.e., (1000)-(999)=1. Delta value 520 is calculated as (index 514)-(index 512), i.e., (1001)-(1000)=1. Delta value 522 is calculated as (index 516)-(index 512), i.e., (1002)-(1000)=2. The three delta values ​​518, 520, 522 are conveniently referred to as delta triplet 524.

[0066] After the value of the delta triplet 524 is determined, the delta triplet 524 is compressed. To compress the delta triplet, the three delta values ​​518, 520, 522 are compared to values ​​in a common pattern lookup table 526. The common pattern lookup table 426 may be implemented in any suitable manner, such as dedicated registers in the graphics hardware, and may be of any suitable size. In this example, the common pattern lookup table 526 stores delta values ​​of 15 patterns. Table 1 shows the delta values ​​of the 15 patterns stored in the exemplary common pattern lookup table 526 and the corresponding pattern ID for each entry.

[0067] [Table 1]

[0068] In this example, delta values ​​518, 520, 522, in that order, match a pattern in common pattern lookup table 526 that corresponds to pattern ID 3 (binary 0011). In this case, index triplet 504 is compressed into compressed representation 528 based on pattern ID 530, which has a value of 3 (b0011). In this example, since each of delta values ​​518, 520, 522 is captured by a pattern ID, there are no "mismatched" deltas, and therefore encoded delta value 532 is not used to generate compressed representation 528. In this example, locality among index values ​​advantageously facilitates compression without the need for variable length coding.

[0069] FIG. 6 is a flow diagram 600 illustrating exemplary triplet-based delta compression of an exemplary index stream for an exemplary case in which a subset of the delta values ​​of an exemplary index triplet match particular delta values ​​of an entry in a common pattern lookup table, while other delta values ​​match “don't care” values ​​of the entry.

[0070] More specifically, flow diagram 600 shows an index stream that includes six indexes split into two index triplets 602, 604. Index triplet 602 includes indexes 606, 608, and 610. Index 610 has a value of 0 in this example. The values ​​of indexes 606 and 608 are not used in this exemplary compression and are simply shown as "D" in the figure. Index triplet 604 includes indexes 612, 614, and 616 that have values ​​of 1000, 1001, and 1002, respectively.

[0071] In this example, to compress index triplet 604, three delta values ​​618, 620, 622 corresponding to indexes 612, 614, 616, respectively, are calculated based on the values ​​of indexes 612, 614, 616. For example, delta value 618 is determined as (index 612)-(index 610), i.e., (1000)-(0)=1000. Delta value 620 is calculated as (index 614)-(index 612), i.e., (1001)-(1000)=1. Delta value 622 is calculated as (index 616)-(index 612), i.e., (1002)-(1000)=2. The three delta values ​​618, 620, 622 are conveniently referred to as delta triplet 624.

[0072] After the value of the delta triplet 624 is determined, the delta triplet 624 is compressed. To compress the delta triplet, the three delta values ​​618, 620, 622 are compared to values ​​in a common pattern lookup table 626. The common pattern lookup table 626 may be implemented in any suitable manner, such as a dedicated register in the graphics hardware, and may be of any suitable size. In this example, the common pattern lookup table 626 stores delta values ​​of 15 patterns. Table 1 shows the delta values ​​of the 15 patterns stored in the exemplary common pattern lookup table 626 and the corresponding pattern ID for each entry.

[0073] In this example, delta values ​​618, 620, 622, in that order, match a pattern in common pattern lookup table 626 that corresponds to pattern ID 4 (binary 0100). In this case, index triplet 604 is compressed into a compressed representation 628 based on pattern ID 630, which has a value of 4 (b0100). In this example, two of the specific delta values ​​620, 622 are captured by the pattern IDs, but delta value 618 is represented by a "don't care" entry that matches any delta value. Therefore, the value of delta value 618 (1000, or b01111101000 including the sign bit) is "mismatched," and therefore these bits are encoded as encoded data value 632, for example, using variable length encoding. The combination of pattern ID 630 and encoded data value 632 is used to generate compressed representation 628. In this example, locality between some of the index values ​​advantageously facilitates compression without requiring variable length encoding of all of the delta values, although some encoding compensates for compression due to the large difference in values ​​between index 610 and index 612.

[0074] FIG. 7 is a flow diagram 700 illustrating exemplary triplet-based delta compression of an exemplary index stream for an exemplary case in which none of the delta values ​​of an exemplary index triplet match a particular delta value of an entry in a common pattern lookup table, but rather match an entry having three “don't care” values.

[0075] More specifically, flow diagram 700 shows an index stream that includes six indexes split into two index triplets 702, 704. Index triplet 702 includes indexes 706, 708, and 710. Index 710 has a value of 0 in this example. The values ​​of indexes 706 and 708 are not used in this exemplary compression and are simply shown as "D" in the figure. Index triplet 704 includes indexes 712, 714, and 716 that have values ​​of 1000, 2000, and 2001, respectively.

[0076] In this example, to compress index triplet 704, three delta values ​​718, 720, 722 corresponding to indexes 712, 714, 716, respectively, are calculated based on the values ​​of indexes 712, 714, 716. For example, delta value 718 is determined as (index 712)-(index 710), i.e., (1000)-(0)=1000. Delta value 720 is calculated as (index 714)-(index 712), i.e., (2000)-(1000)=1000. Delta value 722 is calculated as (index 716)-(index 712), i.e., (2001)-(1000)=1001. The three delta values ​​718, 720, 722 are conveniently referred to as delta triplet 724.

[0077] After the value of the delta triplet 724 is determined, the delta triplet 724 is compressed. To compress the delta triplet, the three delta values ​​718, 720, 722 are compared to values ​​in a common pattern lookup table 726. The common pattern lookup table 726 may be implemented in any suitable manner, such as dedicated registers in the graphics hardware, and may be of any suitable size. In this example, the common pattern lookup table 726 stores delta values ​​of 15 patterns. Table 1 shows the delta values ​​of the 15 patterns stored in the exemplary common pattern lookup table 726 and the corresponding pattern ID for each entry.

[0078] In this example, delta values ​​718, 720, 722, in that order, match a pattern in common pattern lookup table 726 that corresponds to pattern ID 14 (binary 1110). In this case, index triplet 704 is compressed into a compressed representation 728 based on pattern ID 730, which has a value of 14 (b1110). In this example, none of the specific delta values ​​718, 720, 722 are captured by a pattern ID, but all of the delta values ​​718, 720, 722 are represented by "don't care" entries that match any delta value. Thus, the values ​​of delta values ​​718, 720, 722 (1000, 2000, 2001) are all "mismatched," and therefore all of them are encoded as encoded data values ​​732, e.g., using variable length encoding. The combination of pattern ID 730 and encoded data values ​​732 is used to generate compressed representation 728. In this example, variable length encoding of all of the delta values ​​is required due to the relative lack of locality between some of the index values.

[0079] FIG. 8 is a flow chart illustrating an example process 800 for compressing an example index stream.

[0080] In step 802, an index triplet is input from the primitive index stream to encoding hardware (e.g., part of the rasterizer stage 314 as shown and described with respect to FIG. 3). In step 804, a delta value triplet is calculated based on the index triplet. On condition 806 that the delta values ​​of the index triplet all match particular values ​​of entries in the common pattern lookup table, the index triplet is compressed in step 808 based on the pattern ID that corresponds to the entry in the common pattern lookup table (e.g., its index value).

[0081] On condition 806 that none of the delta values ​​of the index triplet match a particular value of an entry in the common pattern lookup table, the index triplet is compressed based on an encoding (e.g., variable length encoding) of the delta values ​​in step 810. In some embodiments, the index triplet is compressed based on this encoding and a pattern ID (e.g., index value) that corresponds to an entry in the common pattern lookup table for which each delta value is a "don't care" value (i.e., matches any delta value).

[0082] On condition 806 that a subset (e.g., one or two) of the delta values ​​of the index triplet match a particular value of an entry in the common pattern lookup table and the remainder do not match, the index triplet is compressed in step 812 based on a combination of an encoding (e.g., variable length encoding) of one or more of the mismatched delta values ​​and a pattern ID (e.g., index value) corresponding to the entry in the common pattern lookup table that matches the particular delta and "don't care" values. On condition 814 that the index stream has more index triplets for compression, flow returns to step 802. Otherwise, the process ends.

[0083] (Decompress) The techniques described above with respect to compression are applicable to decompression. For example, the compressed representation of each triplet (or other grouping of primitive information in the case of non-triangular primitives) includes or is based on a pattern ID. In some embodiments, the decoder decodes the pattern ID to determine how the triplet is compressed. Depending on the pattern ID, the triplet is compressed based on the pattern ID and previous index, based on the pattern ID, previous index and some variable length coded delta value, or based entirely on the previous index and variable length coded delta value, as indicated by the pattern ID.

[0084] FIG. 9 is a flow chart illustrating an example process 900 for decompressing an example compressed index stream, for example, compressed according to the techniques described above.

[0085] In step 902, a compressed index triplet is input from the compressed stream of primitives to a decoding hardware (e.g., part of the assembler stage 302 shown and described with respect to FIG. 3). In step 904, a pattern ID is extracted (e.g., parsed) from the compressed index triplet and decoded. On condition 906 that the decoded pattern ID indicates that all of the delta values ​​of the index triplet match a particular value of an entry in the common pattern lookup table corresponding to the pattern ID, the index triplet is decompressed in step 908 by decoding the delta value in the common pattern lookup table corresponding to the pattern ID based on the previous triplet index in the stream. In certain cases, such as the first triplet in the stream, the value of the previous triplet index is set to 0 or a different suitable value in some embodiments.

[0086] If the condition 906 indicates that the decoded pattern ID indicates that the delta values ​​of the index triplet do not match a particular value of the entry of the common pattern lookup table corresponding to the pattern ID, then all of the delta values ​​of the compressed index triplet are decompressed in step 908 based on the encoded (e.g., variable length encoded) delta values ​​in the compressed index triplet.

[0087] On condition 906 that the decoded pattern ID indicates that a subset (e.g., one or two) of the delta values ​​of the delta triplet match a particular value in the common pattern lookup table entry corresponding to the pattern ID, but the rest do not match, the index triplet is decompressed in step 912 based on a combination of decoding (e.g., variable length decoding) the non-matching delta values ​​and decoding the matching delta values ​​in the common pattern lookup table corresponding to the pattern ID based on the previous triplet index in the stream. On condition 914 that the index stream has more index triplets for compression, flow returns to step 902. Otherwise, the process ends.

[0088] It should be noted that information about primitives other than vertex index values ​​may, in some embodiments, be compressed in accordance with the techniques described herein.

[0089] For example, in some embodiments, each primitive is associated with a compressible instance identifier (ID). Instancing is a feature in 3D graphics APIs that allows the same mesh to be rendered multiple times with a single API call. Each instance of a mesh has an instance ID (which increments by one for each instance), which is used, for example, to select a different transformation matrix or animation frame for each instance of the same mesh. In some cases, such instancing is used to render scenes that contain the same object repeated multiple times (i.e., multiple instances). An example includes instances of people in a large crowd, instances of trees in a forest, etc., and in some cases includes some variation between objects. Since each instance is typically rendered with a different transformation, primitive visibility may vary from instance to instance. Thus, visibility data for each instance is stored separately in some embodiments, even though the source index data for each instance is the same. Thus, in some embodiments, instance ID changes are embedded in the visibility data. In some embodiments, this has the advantage of facilitating straightforward reconstruction of the original instance ID without, for example, splitting each instantiated draw call into multiple draws. Instance ID changes typically do not occur for every primitive (i.e., they are typically sparser than every primitive), but in some cases instanced rendering is used and each rendered instance contains very simple geometry (e.g., only one or two primitives for a particle system). Because some instances may be completely invisible, the stored instance ID delta between successive instances is not always 1 (i.e., it can jump forward by a larger amount).

[0090] Referring to FIG. 7, index triplet 702 points to the vertices of a first primitive, and index triplet 704 points to the vertices of a second primitive. In some embodiments, each of the first and second primitives is associated with an instance ID. Similar to the vertex indexes described above, the instance IDs may exhibit substantial time locality in some cases. For example, the instance ID associated with the first primitive may have a value of 100000, and the instance ID associated with the second primitive may have a value of 100001. Although each of these instance IDs may require multiple binary bits to represent, the difference between these particular example values ​​is representable using only a single bit (i.e., in this case, the difference is 1 or b1). As such, the instance IDs are compressed in the compressed representation (compressed representation 728 using the example of FIG. 7) based on the delta of the instance ID values.

[0091] In some embodiments, the delta between the instance ID of the first primitive and the instance ID of the second primitive is encoded along with any mismatch deltas (e.g., encoded along with delta values ​​718, 720, 722 in the variable length coding of FIG. 7). In some embodiments, if there are no mismatch deltas (or other information for the variable length coding), then only the instance ID delta is encoded as the coded delta value 732.

[0092] In some embodiments, the common pattern lookup table is configured to accommodate the instance ID delta. For example, in some embodiments, the common pattern lookup table (e.g., common pattern lookup table 726 in the example of FIG. 7) includes a fourth index that matches the combined pattern of the instance ID delta and the vertex index delta. In such a case, the instance ID is compressed into a compressed representation (e.g., compressed representation 728 in the example of FIG. 7) by the pattern ID (e.g., pattern ID 730 in the example of FIG. 7) along with any matching vertex index deltas.

[0093] Note that beyond instance IDs, this technique can be extended to any one or more information items associated with a primitive, such as a primitive ID. A primitive ID, in some embodiments, identifies each primitive individually. In some embodiments, a primitive ID identifies a primitive within a mesh (e.g., as a "running number"). For example, in some embodiments, a primitive ID starts at 0 at the beginning of a mesh and increments by 1 for each primitive. Because the visibility stream contains only some of the primitives, a primitive ID may jump forward several values ​​between primitives, rather than a continuously incrementing value after the data is compressed and decompressed. In some embodiments, this is captured in the compressed index stream so that the data can be reconstructed using the original primitive ID. In some embodiments, such items include integer values, such as integer values ​​with a threshold degree of temporal locality, where numbers associated with primitives adjacent to each other in the stream are close in value, such as within a threshold.

[0094] It should be noted that in some embodiments, primitives in a stream that have a different number of vertices than the base primitive may also be compressed according to the techniques described herein.

[0095] For example, examples herein are presented with respect to a triangle primitive having three vertices and therefore three vertex indices. Note, however, that some embodiments also accommodate primitives in the index stream that are defined with respect to a different number of vertex indices. For example, in some embodiments, primitives in the triangle stream include only one or two vertex indices (e.g., point and line primitives) or include four or more vertex indices (e.g., rectangle primitives, etc.).

[0096] A patch primitive is an exemplary use case of a primitive with more than three indices. Rectangles are typically used for 2D operations and in some embodiments are defined with three vertices, with a fourth vertex calculated from the other vertices. A patch primitive contains any number of source vertices (e.g., up to 32 vertices per primitive, depending on the API). Patch primitives are typically tessellated into triangles in the geometry processing pipeline, but the original indices are used, for example, to store index data.

[0097] In some embodiments, this situation is matched by a common pattern lookup table to one or more "special" pattern IDs (e.g., specific pattern IDs indicating one vertex index, two vertex indexes, or four vertex indexes, etc.), and the delta values ​​of the index vertices are compressed into a compressed representation using variable length coding, as described above with respect to mismatch deltas. In some embodiments, the common pattern LUT is configured to accommodate pattern compression of two or more primitive sizes. In some such instances, the vertex indices of two or more types of primitives (i.e., primitives having different numbers of vertex indexes) can be compressed into a compressed representation based on the corresponding pattern IDs (and variable length coding, e.g., if there are mismatch deltas).

[0098] In some embodiments, primitives having different numbers of vertices are combined to form triplets (or vertices of other cardinalities). For example, in some embodiments, one vertex index from a point primitive is combined with two vertex indexes of a line primitive to form a triplet, and the triplet is compressed as described above for delta compression of triplets of triangle primitives.

[0099] Typically, it is not necessary to combine point and line primitives into triplets. Each draw call uses only primitives of a single type, so there are typically long sequences of a single type. However, this can be implemented if, for example, only some visible primitives are visible from such a sequence, followed by only some visible primitives from another sequence having a different type.

[0100] Some embodiments compress point and line primitives by storing each primitive as a separate "special case" primitive (with only one or two index values) instead of using triplets, or by forming several consecutive triplets, where each triplet consists of three points (for point primitives), or where every two triplets consists of two lines (for line primitives). In some such cases, a single 1-index or 2-index "special case" data entry is used (e.g., at the end of the data) for consistency with the total index count of the draw call.

[0101] In some embodiments, multiple streams are compressed and / or decompressed simultaneously. For example, if the same primitive streams are compressed simultaneously for different purposes (e.g., for different bins), the delta calculations within each triplet (e.g., deltas 720, 722 in the example of FIG. 7) will be the same in some embodiments. Therefore, in such cases, only the delta from the previously stored index (e.g., delta 718 calculated from index 710 in the example of FIG. 7) is calculated for each additional stream. In some embodiments, the deltas to the previously stored index are different for each stream, but the triplet is the same for each stream. In some embodiments, each triangle is represented by a triplet, and the first delta stored is the difference from the last vertex of the last stored triangle. Since each bin may contain a different set of triangles, the deltas to the last stored primitive are different for each bin. However, the deltas within a triangle (the delta between index 0 and index 1, and the delta between index 0 and index 2) are the same for all bins.

[0102] It should be understood that many variations are possible based on the disclosure herein, and although features and elements have been described above in particular combinations, each feature or element can be used alone without the other features and elements, or in various combinations with or without the other features and elements.

[0103] The various functional units illustrated in the figures and / or described herein (including, but not limited to, the processor 102, the input driver 112, the input device 108, the output driver 114, the output device 110, the acceleration processing device 116, the scheduler 136, the graphics processing pipeline 134, the computation unit 132, the SIMD unit 138) may be implemented as a general purpose computer, processor or processor core, or as a program, software or firmware stored in a non-transitory computer readable storage medium or another medium executable by the general purpose computer, processor or processor core. The methods provided may be implemented in a general purpose computer, processor or processor core. Suitable processors include, by way of example, general purpose processors, special purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines. Such processors may be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data such as netlists (such instructions may be stored on a computer readable medium). The result of such processing may be a mask work that is used in a subsequent semiconductor manufacturing process to manufacture a processor implementing features of the present disclosure.

[0104] The methods or flow diagrams provided herein may be implemented in a computer program, software, or firmware embodied in a non-transitory computer-readable storage medium for execution by a general purpose computer or processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (e.g., internal hard disks and removable disks), magneto-optical media, and optical media (e.g., CD-ROM disks and digital versatile disks (DVDs)).

Claims

1. 1. A method for compressing an index stream associated with a graphics primitive, comprising: determining a group of delta values ​​based on the group of indexes of the index stream; comparing the group of delta values ​​to delta values ​​in a lookup table; compressing the index group based on an entry in the lookup table if the delta value group matches all delta values ​​of the entry, and compressing the index group based on variable length coding otherwise. method.

2. and compressing the index group based on the entry and based on a variable length encoding of the at least one value that does not match any delta value of the entry, if the delta value group includes at least one value that matches a delta value of an entry in the lookup table and includes at least one value that does not match any delta value of the entry, and the at least one value that does not match any delta value of the entry corresponds to a don't care value of the entry.

2. The method of claim 1.

3. if the delta value group does not include at least one value that matches a delta value of an entry in the lookup table, compressing the index group based on entries in the lookup table that include only don't care values ​​and based on a variable length encoding of each of the delta values.

2. The method of claim 1.

4. the index group is a triplet and the graphics primitive is a triangle; 2. The method of claim 1.

5. the index stream indexes the vertices of a mesh; 2. The method of claim 1.

6. 1. A method for decompressing a compressed index stream associated with a graphics primitive, comprising: comparing compressed index groups of the compressed index stream with entries of a lookup table; if the compressed index group matches a lookup table entry including a delta value corresponding to each of the compressed index groups, decompressing the compressed index group based on the delta value, and otherwise decompressing the compressed index group based on variable length decoding. method.

7. If the compressed index group includes at least one delta value corresponding to at least one of the compressed index groups and matches a lookup table entry including at least one don't care value corresponding to at least one of the compressed index groups, decompressing the compressed index group based on the at least one delta value and based on variable length decoding of the delta value corresponding to the at least one don't care value. The method of claim 6.

8. and decompressing the compressed index groups based on variable length decoding of delta values ​​corresponding to each of the compressed index groups when the compressed index groups match a lookup table entry that includes only don't care values. The method of claim 6.

9. the compressed index group is a triplet and the graphics primitive is a triangle; The method of claim 6.

10. the compressed index stream indexing the vertices of a mesh; The method of claim 6.

11. 1. A compressor configured to compress an index stream associated with a graphics primitive, the compressor comprising: a circuit configured to determine a group of delta values ​​based on a group of indexes of the index stream; a circuit configured to compare the group of delta values ​​to delta values ​​in a lookup table; and a circuit configured to compress the index group based on an entry in the lookup table if the delta value group matches all delta values ​​of the entry, and to compress the index group based on a variable length coding if the delta value group matches all delta values ​​of the entry in the lookup table. Compressor.

12. the delta value group includes at least one value that matches a delta value of an entry in the lookup table and includes at least one value that does not match any of the delta values ​​of the entries, and where the at least one value that does not match any of the delta values ​​of the entries corresponds to a don't care value of the entry, further comprising a circuit configured to compress the index group based on the entry and based on a variable length encoding of the at least one value that does not match any of the delta values ​​of the entries. The compressor of claim 11.

13. and a circuit configured to compress the index group based on entries in the lookup table that include only don't care values ​​and based on a variable length encoding of each of the delta values ​​if the delta value group does not include at least one value that matches a delta value of an entry in the lookup table. The compressor of claim 11.

14. the index group is a triplet and the graphics primitive is a triangle; The compressor of claim 11.

15. the index stream indexes the vertices of a mesh; The compressor of claim 11.

16. 1. A decompressor configured to decompress a compressed index stream associated with a graphics primitive, the decompressor comprising: a circuit configured to compare compressed index groups of the compressed index stream with entries of a lookup table; and a circuit configured to, if the compressed index group matches a lookup table entry including a delta value corresponding to each of the compressed index groups, decompress the compressed index group based on the delta value, and otherwise decompress the compressed index group based on variable length decoding. Decompressor.

17. and a circuit configured to decompress the compressed index group based on the at least one delta value and based on variable length decoding of the delta value corresponding to the at least one don't care value if the compressed index group matches a lookup table entry including at least one delta value corresponding to at least one of the compressed index groups and including at least one don't care value corresponding to at least one of the compressed index groups.

17. The decompressor of claim 16.

18. and a circuit configured to decompress the compressed index groups based on variable length decoding of delta values ​​corresponding to each of the compressed index groups when the compressed index groups match a lookup table entry that includes only don't care values.

17. The decompressor of claim 16.

19. the compressed index group is a triplet and the graphics primitive is a triangle; 17. The decompressor of claim 16.

20. the compressed index stream indexing the vertices of a mesh; 17. The decompressor of claim 16.

Citation Information

Patent Citations

  • Memory for fast table lookup and low power mechanism

    JP2007508653A

  • Memory efficient ray tracing with hierarchical mesh quantization

    JP2011081788A

  • Randomly accessible, lossless parameterized data compression for tile-based 3D computer graphics systems

    JP2013541755A

  • Vertex parameter data compression

    US20140354666A1

  • Buffer index format and compression

    US20180232849A1