Data compression and packing
Patent Information
- Application Number
- CN202111440952.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-01
- Filing Date
- 2021-11-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-11-30
AI Technical Summary
与此相悖的是,希望在更快的GPU上使用更高质量的渲染算法,因而对相对有限的资源(存储器带宽)造成了压力
Smart Images

Figure CN114584778B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to data compression and packaging. Background Technology
[0002] Data compression, whether lossless or lossy, is desirable in many applications where data will be stored in and / or read from memory. By compressing data before storing it in memory, the amount of data transferred to memory can be reduced. Examples of data for which data compression is particularly useful are image data, such as depth data to be stored in a depth buffer, pixel data to be stored in a frame buffer, and texture data to be stored in a texture buffer. These buffers can be any suitable type of memory, such as cache memory, a separate memory subsystem, a storage area in a shared memory system, or some combination thereof.
[0003] A graphics processing unit (GPU) processes image data to determine the pixel values of an image to be stored in a frame buffer for output to a display. GPUs typically have a highly parallelized architecture for processing large blocks of data in parallel. There is significant commercial pressure to keep GPUs (especially those intended for mobile devices) running at low power levels. Conversely, there is a desire to use higher-quality rendering algorithms on faster GPUs, thus putting pressure on relatively limited resources (memory bandwidth). However, increasing the bandwidth of the memory subsystem may not be an attractive solution, as moving data in and out of the GPU, and even within the GPU itself, consumes a significant portion of the GPU's power budget. The same problem may exist for other processing units, such as the central processing unit (CPU), as well as the GPU.
[0004] The implementation schemes described below are provided by way of example only and do not limit the ways in which any or all of the shortcomings of known data compression and decompression methods can be addressed. Summary of the Invention
[0005] This summary is provided to introduce, in a simplified form, a series of concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
[0006] A data compression method is described. The method includes receiving input pixel data blocks in raster scan order, wherein the pixel data includes at least red channel data, green channel data, and blue channel data for each pixel. The pixel data is compressed using a block-based coding scheme, and then the compressed pixel data is output substantially in raster scan order.
[0007] A data decompression method is described. The method includes receiving compressed pixel data substantially in raster scan order, and determining the number of bits in the compressed data corresponding to a row of pixels. Then, for each group of pixels in the row, the method identifies a block-based decoding scheme for the pixel group and decodes the pixel group using the identified scheme.
[0008] A first aspect provides a data compression method, comprising: receiving input pixel data of a data block in raster scan order, the pixel data including at least a first channel data, a second channel data, and a third channel data for each pixel; compressing the pixel data substantially in raster scan order using a block-based coding scheme; and outputting the compressed pixel data substantially in raster scan order, wherein compressing the pixel data substantially in raster scan order using a block-based coding scheme comprises: subdividing the input block of pixels into partial blocks, each partial block including more than one row of pixels; and compressing the pixel data substantially in raster scan order using the block-based coding scheme and the subdivision.
[0009] Compressing the pixel data using a block-based coding scheme and the subdivision substantially in raster scan order may include: for each sub-block, analyzing the pixel data of pixels in the sub-block to identify the coding pattern used on the pixels in the sub-block; and encoding the pixel data substantially in raster scan order using the selected pattern.
[0010] A partial block may include 2×2 pixels, and analyzing pixel data to identify the encoding pattern for pixels in the partial block may include, for each partial block: calculating the difference between each pair of pixels in the partial block; in response to determining that the minimum difference does not exceed a predefined threshold, selecting a ternary encoding pattern from a set of ternary encoding patterns based on the pixel pairs having the minimum difference, and calculating the average pixel data of the pixels in the pixel pairs having the minimum difference and storing it in a buffer; and in response to determining that the minimum difference exceeds a predefined threshold, selecting a quaternary encoding pattern.
[0011] Analyzing pixel data to identify the encoding pattern used on pixels in the partial block may further include, for each partial block: in response to determining that the minimum difference does not exceed a predefined threshold and that the pixel pair having the minimum difference is a pixel pair in the first row of the partial block, storing the pixel data of the first pixel in the second row of the partial block in the buffer.
[0012] Encoding the pixel data substantially in raster scan order using the selected type may include: encoding the pixel data substantially in raster scan order using the selected type such that each pixel or adjacent pixel pair from the same partial block is compressed into a sequence of multiples of P bits, where P is an integer.
[0013] Encoding the pixel data substantially in raster scan order using the selected pattern, such that a sequence of pixels or adjacent pixel pairs from the same partial block compressed to multiples of P bits may include: encoding the pixel data substantially in raster scan order using the selected pattern; and wherein a sequence of pixels or adjacent pixel pairs from the same partial block compressed to multiples of P bits, embedding one or more control and / or padding bits to increase the sequence length to multiples of P bits.
[0014] A partial block may include 2×2 pixels, and the pixel data encoded substantially in raster scan order using a selected pattern may include: for adjacent pixel pairs from the first row of the partial block and, in the case of using a ternary encoding pattern, outputting a sequence of two encoded pixel values; and for adjacent pixel pairs from the second row of the partial block and, in the case of using a ternary encoding pattern, outputting a sequence of one encoded pixel value.
[0015] A sequence comprising two encoded pixel values may include either: average pixel data of pixel pairs in the partial block read from a buffer, and converted pixel data of one pixel from an adjacent pixel pair in the first row of the partial block; or converted pixel data of each pixel in an adjacent pixel pair in the first row of the partial block; and a sequence comprising one encoded pixel value may include either: average pixel data of pixel pairs in the partial block read from a buffer; or converted pixel data of one pixel from an adjacent pixel pair in the second row of the partial block.
[0016] Outputting compressed pixel data in essentially raster scan order can include: packing the compressed pixel data into a data structure in essentially raster scan order; and outputting the data structure.
[0017] Packing the compressed pixel data into a data structure substantially in raster scan order may include: concatenating the compressed pixel data substantially in raster scan order to form a compressed pixel data block; and concatenating the compressed pixel data block and a control data block.
[0018] Connecting the compressed pixel data block and the control data block may include: appending or prepending the control data block to the compressed pixel data block.
[0019] The method may further include embedding one or more bits of control data into the pixel data before concatenating the compressed pixel data.
[0020] The first channel data, the second channel data, and the third channel data for each pixel may include red channel data, green channel data, and blue channel data for each channel.
[0021] A second aspect provides data compression hardware, comprising: an input terminal for receiving input pixel data of a data block in raster scan order, the pixel data including at least a first channel data, a second channel data, and a third channel data for each pixel; hardware logic arranged to compress the pixel data substantially in raster scan order using a block-based encoding scheme; and an output terminal for outputting the compressed pixel data substantially in raster scan order, wherein the hardware logic includes: an analysis pipeline arranged to subdivide the input block of pixels into partial blocks, each partial block including more than one row of pixels; and encoding hardware arranged to compress the pixel data substantially in raster scan order using a block-based encoding scheme and the subdivision.
[0022] The analysis pipeline can also be arranged such that, for each sub-block, pixel data of pixels in the sub-block is analyzed to identify the encoding pattern used on the pixels in the sub-block, and the encoding hardware is arranged to encode the pixel data substantially in raster scan order using the selected pattern.
[0023] The data compression hardware may further include a buffer, wherein the analysis pipeline is further arranged to store data in the buffer for use by the encoding hardware, and wherein the encoding hardware is arranged to read the stored data from the buffer when encoding the pixel data.
[0024] The data compression hardware may further include packing hardware, which is arranged to pack the compressed pixel data into a data structure substantially in raster scan order and output the data structure.
[0025] The packaging hardware may include: partial block packaging hardware arranged to connect the compressed pixel data substantially in raster scan order to form compressed pixel data blocks; and tile block packaging hardware arranged to connect the compressed pixel data blocks and control data blocks.
[0026] A third aspect provides a graphics processing system that includes data compression hardware as detailed above.
[0027] The fourth aspect provides a graphics processing system configured to perform the methods detailed above.
[0028] The fifth aspect provides computer-readable code configured to cause the methods detailed above to be executed when the code is run.
[0029] A sixth aspect provides a computer-readable storage medium having a computer-readable description of an integrated circuit stored thereon, wherein when the computer-readable description is processed in an integrated circuit manufacturing system, the integrated circuit manufacturing system causes the integrated circuit manufacturing system to manufacture data compression hardware as detailed above.
[0030] A seventh aspect provides an integrated circuit manufacturing system comprising: a computer-readable storage medium storing a computer-readable description of an integrated circuit, the computer-readable description describing data compression hardware as detailed above; a layout processing system configured to process the integrated circuit description to generate a circuit layout description of the integrated circuit embodying the data compression hardware; and an integrated circuit manufacturing system configured to manufacture the data compression hardware according to the circuit layout description.
[0031] The eighth aspect provides a data decompression method comprising: receiving compressed pixel data substantially in raster scan order; determining the number of bits of compressed data corresponding to a row of pixels; and for each group of pixels in the row, identifying a block-based decoding scheme for the group of pixels, and decoding the group of pixels using the identified scheme.
[0032] A ninth aspect provides data decompression hardware comprising: an input terminal for receiving compressed image data substantially in raster scan order; control hardware arranged to determine the number of bits of compressed data corresponding to a row of pixels; and decoding hardware arranged to decode groups of pixels, wherein the control hardware or the decoding hardware is further arranged to identify a block-based decoding scheme for each group of pixels in the row, and wherein the decoding hardware is arranged to decode the group of pixels using the identified scheme.
[0033] The data compression and / or decompression units described herein can be implemented in hardware on an integrated circuit. A method for manufacturing the data compression and / or decompression units as described herein in an integrated circuit manufacturing system can be provided. An integrated circuit definition dataset can be provided, which, when processed in an integrated circuit manufacturing system, configures the system to manufacture the data compression and / or decompression units as described herein. A non-transitory computer-readable storage medium storing a computer-readable description of an integrated circuit can be provided, which, when processed, causes a layout processing system to generate a circuit layout description used in the integrated circuit manufacturing system to manufacture the data compression and / or decompression units as described herein.
[0034] An integrated circuit manufacturing system may be provided, comprising: a non-transitory computer-readable storage medium storing a computer-readable integrated circuit description describing data compression and / or decompression units as described herein; a layout processing system configured to process the integrated circuit description to generate a circuit layout description of an integrated circuit embodying the data compression and / or decompression units as described herein; and an integrated circuit generation system configured to manufacture the data compression and / or decompression units as described herein based on the circuit layout description.
[0035] Computer program code for performing any of the methods described herein may be provided. A non-transitory computer-readable storage medium may be provided storing computer-readable instructions that, when executed at a computer system, cause the computer system to perform any of the methods described herein.
[0036] As will be apparent to those skilled in the art, the above features can be appropriately combined, and can be combined with any aspect of the examples described herein. Attached Figure Description
[0037] The example will now be described in detail with reference to the accompanying drawings, in which:
[0038] Figure 1 The graphics rendering system is shown;
[0039] Figure 2 This is a flowchart of an exemplary data compression method;
[0040] Figure 3 This is a diagram illustrating two input blocks with different formats;
[0041] Figure 4 It is configured to implement Figure 2 A schematic diagram of exemplary compression hardware for the method;
[0042] Figure 5 It is configured for decompression. Figure 2 A schematic diagram of exemplary hardware for decompressing compressed data generated by the method;
[0043] Figure 6 It is possible to pass Figure 5 A flowchart illustrating an example of hardware-based data decompression;
[0044] Figure 7 This is a schematic diagram of an example compressed data block;
[0045] Figure 8 This is a flowchart of an exemplary method for compressing sub-blocks;
[0046] Figure 9 It shows that it can be done Figure 8 A diagram illustrating the coding style used in the method;
[0047] Figure 10 It is possible Figure 2 A flowchart of the encoding method used in the method;
[0048] Figure 11 This is a flowchart of a first exemplary method for converting 10-bit data to 8-bit data;
[0049] Figure 12 This is a flowchart of an exemplary method for converting a-digit number to b-digit number, where a > b;
[0050] Figure 13 Here is a flowchart of another exemplary data compression method;
[0051] Figure 14 and Figure 15 Two different exemplary implementations are shown for determining whether to use a constant α mode or a variable α mode;
[0052] Figure 16 A computer system implementing data compression and / or decompression units is shown; and
[0053] Figure 17 An integrated circuit manufacturing system for generating integrated circuits embodying data compression and / or decompression units as described herein is illustrated.
[0054] The accompanying drawings illustrate various examples. Those skilled in the art will understand that the element boundaries (e.g., boxes, groups of boxes, or other shapes) shown in the drawings represent one example of a boundary. In some examples, it may be that one element can be designed as multiple elements, or multiple elements can be designed as one element. Where appropriate, common reference numerals are used throughout the drawings to indicate similar features. Detailed Implementation
[0055] The following description is given by way of example to enable those skilled in the art to make and use the invention. The invention is not limited to the embodiments described herein, and various modifications to the disclosed embodiments will be readily apparent to those skilled in the art.
[0056] The embodiments will now be described by way of example only.
[0057] As mentioned above, memory bandwidth is a relatively limited resource within a processing unit (e.g., a CPU or GPU), and similarly, memory space is also a limited resource, as increasing it impacts both the physical size and power consumption of the device. By using data compression before storing data in memory, memory bandwidth and memory space can be saved.
[0058] Data compression techniques can be lossless or lossy. Lossless compression allows for the perfect reconstruction of the original data from the compressed data. Conversely, lossy compression does not perfectly reconstruct the data; instead, the decompressed data is only an approximation of the original. Lossy compression typically compresses data to a greater extent than lossless compression (i.e., achieves a lower compression ratio). The amount of compression achievable using lossless compression depends on the nature of the data being compressed, some of which is easier to compress than others.
[0059] The amount of compression achieved through compression techniques (whether lossless or lossy) can be expressed as a percentage, referred to here as the compression ratio, and is given by the following formula:
[0060]
[0061] This means that a compression ratio of 100% means no compression has been achieved, a compression ratio of 50% means the data has been compressed to half its original, uncompressed size, and a compression ratio of 25% means the data has been compressed to one-quarter of its original, uncompressed size.
[0062] It should be understood that there are other ways to limit the compression ratio, and there are other ways to express the amount of compression, which can be achieved by data compression methods, such as by referring to the size of the compressed data block (e.g., in terms of the number of bytes).
[0063] The variability in the amount of compression that can be achieved (depending on the characteristics of the actual data being compressed) affects both memory bandwidth and memory space, and may mean that the full benefits of the achieved compression will not be realized in one or both of these aspects, as described below.
[0064] In many use cases, random access to raw data is required. Typically, for image data, this is achieved by dividing the image data into independent, non-overlapping rectangular blocks before compression. If the size of each compressed block varies depending on the nature of the data within the block (e.g., a block containing only the same color might compress much more than a block containing a lot of detail), such that in some cases the block might not be compressed at all, then to maintain the ability to randomly access compressed data blocks, memory space could be allocated as if the data wasn't compressed at all. Alternatively, it's necessary to maintain an index with entries for each block identifying where the compressed data for that block resides in memory. This requires memory space to store the index (which can be relatively large), and memory access (performing lookups in the index) increases system latency. For example, in systems where random access to each compressed data block is important and an index isn't used, even achieving an average compression ratio of 50% (for all data blocks) still requires allocating memory space assuming a 100% compression ratio, because for some blocks, lossless compression techniques might not be able to achieve any compression.
[0065] Furthermore, because data transfer to memory occurs in fixed-size bursts (e.g., in 64-byte bursts), for any given block, only one set of discrete effective compression ratios applies to transferring data to memory. For example, if a data block contains 256 bytes and data transfer occurs in 64-byte bursts, the effective compression ratios for data transfer are 25% (if the block is compressed from 256 bytes to no more than 64 bytes, so only one burst is needed), 50% (if the block is compressed to 65-128 bytes, so two bursts are needed), 75% (if the block is compressed to 129-192 bytes, so three bursts are needed), and 100% (if the block is not compressed at all or is compressed to 193 or more bytes, so four bursts are needed). This means that if a data block comprising 256 bytes is compressed to any value in the range of 129-192 bytes, the compressed block requires three bursts, while the uncompressed block requires four bursts. This results in an effective compression ratio of 75% for memory transfers, while the actual data compression achieved can be much lower (e.g., as low as 50.4% if compressed to 129 bytes). Similarly, if compression can only reduce the block to 193 bytes, the benefits of using data compression are not seen in memory transfers because four bursts are still required to transfer the compressed data block to memory. In other examples, data blocks may include different numbers of bytes, and the bursts to memory may include different numbers of bytes.
[0066] Compressed data can be packaged into data structures (which may be referred to as compressed data blocks) before being transferred to memory or another device. As part of the data decompression process, the compressed data is unpacked (e.g., extracted from the data structure) before being decompressed. The way compressed data is packaged into data structures (i.e., the order in which bits are placed into the data structure) can affect the operation of compression and decompression hardware (especially the amount of buffering required), and thus can affect the size of the hardware and the efficiency of its operation (e.g., in terms of speed, processing power, power consumption, etc.).
[0067] This paper describes methods for performing data compression and data packing of pixel data (e.g., RGB / RGBA / ARGB pixel data), and corresponding data unpacking and decompression methods that provide improved efficiency for compression and decompression operations, such that the compression and decompression operations can be implemented efficiently in hardware (e.g., with a small hardware size). For example, compression hardware can be implemented efficiently in hardware because the amount of multiplexing hardware required is reduced, and decompression hardware can be implemented efficiently in hardware for the same reason and also because the amount of buffering required is reduced. This can be particularly useful in space-constrained implementations (e.g., mobile / portable / handheld computing devices, such as smartphones, tablets, portable gaming devices, smartwatches, etc.). The methods described herein can be used to provide a guaranteed compression ratio (e.g., 50%) or a fixed compressed block size (where this can be defined in bits per byte, e.g., 128 bytes, or bits per pixel, e.g., 16 bits). In various examples, the method described herein provides a packaging scheme for compressed data that has easily determined the boundaries between data of different pixels (or groups of pixels, such as pixel pairs), which leads to further improvements in the efficiency of unpacking and decoding operations in the decompression process (and thus the decompression hardware is less complex and more efficient).
[0068] As described in more detail below, the method for performing data compression and data packing of pixel data involves the subdivision of an input block of uncompressed image data. Subdivision can be a one-step process (which divides a block into multiple sub-blocks), a two-step process (which divides the block into multiple sub-blocks and then each sub-block into multiple mini-blocks), or it may include more than two steps (where the resulting block is still referred to as a mini-block). Regardless of the number of subdivision steps, the input block is subdivided into multiple smaller blocks, which may be collectively referred to as "partial blocks." Therefore, the term "partial block" can refer to either a sub-block or a mini-block.
[0069] Figure 1A graphics rendering system 100 that can be implemented in an electronic device such as a mobile device is shown. The graphics rendering system 100 includes a host CPU 102, a GPU 104, and memory 106 (e.g., graphics memory). The CPU 102 is arranged to communicate with the GPU 104. Data, which may be compressed data, can be transferred between the GPU 104 and the memory 106 in either direction.
[0070] GPU 104 includes a rendering unit 110, a compression / decompression unit 112, a memory interface 114, and a display interface 116. System 100 is arranged such that data can be transferred in any direction: (i) between CPU 102 and rendering unit 110; (ii) between CPU 102 and memory interface 114; (iii) between rendering unit 110 and memory interface 114; (iv) between memory interface 114 and memory 106; (v) between rendering unit 110 and compression / decompression unit 112; (vi) between compression / decompression unit 112 and memory interface 114; and (vii) between memory interface 114 and display interface. System 100 is also arranged such that data can be transferred from compression / decompression unit 112 to display interface 116. An image rendered by GPU 104 can be sent from display interface 116 to a display for display thereon.
[0071] In operation, GPU 104 processes image data. For example, rendering unit 110 may use known techniques such as depth testing (e.g., for hiding surface removal) and texturing and / or shading to perform scan transformations of graphics primitives such as triangles and lines. Rendering unit 110 may include a cache unit to reduce memory traffic. Some data is read from or written to memory 106 by rendering unit 110 via memory interface unit 114 (which may include a cache), but for other data, such as data to be stored in a frame buffer, data is preferably transferred from rendering unit 110 to memory interface 114 via compression / decompression unit 112. Compression / decompression unit 112 reduces the amount of data to be transferred to memory 106 via an external memory bus by compressing the data, as described in more detail below.
[0072] Display interface 116 sends the completed image data to the display. Uncompressed images can be accessed directly from memory interface unit 114. Compressed data can be accessed via compression / decompression unit 112 and sent as uncompressed data to display 108. In an alternative example, compressed data can be sent directly to display 108, and display 108 may include logic for decompressing compressed data in a manner equivalent to decompression by compression / decompression unit 112. Although shown as a single entity, compression / decompression unit 112 may comprise multiple parallel compression and / or decompression units for performance enhancement.
[0073] Figure 2 A first exemplary compression method is shown that can be performed by the compression / decompression unit 112. For example... Figure 2 As shown, the method takes uncompressed image data, i.e., blocks of concatenated pixel values, as input. This image data can be RGB or RGBA data, or in a corresponding format with channels in different orders (e.g., ARGB data). Exemplary formats include RGB888, RGBA8888, ARGB8888, RGBA1010102, and ARGB2101010 formats. The method can also be used for other types of image data (e.g., YUV data); however, when the method utilizes color correlation between channels, it may introduce some errors for formats with less correlation. Each received block of image data is associated with an N×M pixel block (where N and M are integers), and this block can also be referred to as a tiled block of pixels (e.g., in a graphics processing device using a tiled block-based rendering technique), for example, a block / tile 302 comprising 8×8 pixels (N=M=8) or a block / tile 304 comprising 16×4 pixels (N=16, M=4), as shown. Figure 3 As shown, and in the following description, the terms "tile" and "block" may be used synonymously. Image data is received in raster scan order, i.e., pixel values are received line by line, starting with the least significant bit (LSB) in the first line (line 0) and ending with the most significant bit (MSB) in the last line (tile block 302 in line 7 and tile block 304 in line 3), as shown. Figure 3 As shown by the arrow in the image.
[0074] like Figure 2As shown, the input block of data is subdivided into sub-blocks (block 202), each sub-block comprising n x m pixels, where n and m are integers greater than one, and n and m can be different or the same. For input blocks (or tiled blocks) 302, 304 comprising 8 x 8 pixels or 16 x 4 pixels, the block can be subdivided into sub-blocks, each sub-block comprising 4 x 4 pixels (n = m = 4). In various examples, these sub-blocks can be further subdivided into mini-blocks, each mini-block comprising n' x m' pixels, where n' and m' are integers, m' is greater than one, and n' and m' can be different or the same. In various examples, n' can also be greater than one. For example, a 4 x 4 sub-block can be subdivided into four mini-blocks, each mini-block comprising 2 x 2 pixels (n' = m' = 2).
[0075] After the data block is divided into sub-blocks (in box 202), compression is then performed using a block-based encoding scheme (in box 204). Compression (in box 204) involves some analysis performed on the sub-blocks to identify the specific encoding scheme (or pattern) to be used for each sub-block; however, compression is generally performed in raster scan order. When the sub-blocks are further subdivided into mini-blocks, the analysis can be performed mini-block by mini-block rather than sub-block by sub-block. Any suitable block-based compression scheme can be used to perform compression of the input image data in box 204.
[0076] The term "basically in raster scan order" is used in this document to refer to the ordering of pixel data, which is precisely the raster scan order for most encoding schemes (where the raster scan order is determined by...). Figure 3 (The arrows in the two examples indicate this); however, in a small subset of encoding schemes (e.g., only one out of all used encoding schemes), isolated pixel values can be output before their proper output positions according to the exact raster scan order. This sorting is very different from sorting by sub-blocks or mini-blocks, as... Figure 3 As you can see in [the image / video]. (Reference) Figure 3 In the second example 304, where the block comprises four blocks, it can be seen that the pixels from the first row, row 0 (starting with LSB) of each sub-block are output in raster scan order, followed by the second row, row 1, and so on for each sub-block. In contrast, sorting by sub-blocks would result in the following order: traversing the first row, row 0 of the first sub-block, followed by the second, third, and fourth rows of the first sub-block, then the first row, row 0 of the second sub-block, followed by the second, third, and fourth rows of the second sub-block, and so on.
[0077] In various examples, compression (in box 204) includes analyzing the pixel data of each sub-block or mini-block to select an encoding scheme and / or pattern to be used for that sub-block or mini-block (box 204A), and then compressing the pixel data in that sub-block / mini-block using the selected encoding scheme and / or pattern (box 204B). In such examples, the analysis (in box 204A) can be performed sub-block by sub-block (or mini-block by mini-block), while the encoding of the pixel data (in box 204B) can be performed on the input data substantially in raster scan order using the selected encoding scheme and / or pattern determined in the analysis phase. Different encoding schemes and / or patterns may be referred to as different compression modes. References are made below, for example. Figures 8 to 10 Provide a detailed description of various examples of type-based compression methods.
[0078] After compression is performed (in box 204), the compressed data is then packaged into compressed blocks of image data, wherein the compressed pixel data is not arranged in the order of individual sub-blocks (or individual mini-blocks), but is arranged in essentially raster scan order (in box 206), that is, the compressed pixel values are output line by line, starting from the first line and LSB, in essentially the same order as they were received in the original uncompressed data.
[0079] Packing data to form a compressed data block (in box 206) may include concatenating compressed pixel data in substantially raster scan order (box 206B), and then concatenating the compressed pixel data with one or more bits of control data (box 206C). This concatenation with the control data (in box 206C) may include one or more bits of additional control data (e.g., as an LSB) or one or more bits of preceding control data (e.g., as an MSB). In various examples, packing may also include embedding one or more bits of control data within the pixel data prior to concatenation (in box 206B) (box 206A). Whether any control bits are embedded (e.g., whether box 206A is omitted) and exactly which control bits are added and at what position within the compressed pixel data (in box 206A) can depend on the specific encoding scheme and / or type and / or format of the input data for a particular sub-block or mini-block, examples of which will be described in more detail below.
[0080] In various examples, the compressed data can be output exactly in raster scan order (compressed data blocks packed into each input block); however, in some examples, there may be very few compression patterns (e.g., encoding schemes and / or types) that result in only slight modifications to the raster scan order for a small number of pixels (e.g., one pixel per mini-block). In such examples, isolated pixel values may be output before the positions that should be output according to the exact raster scan order, and in these cases, the compressed pixel data is still output substantially in raster scan order, not in a sub-block (or mini-block) order. Exemplary compression patterns that result in minor differences from the exact raster scan order are described in detail below. In particular, using the compression and packing methods described herein, the decompression operation always has enough data to decompress (e.g., decode) the next pixel in raster scan order, and there is no need to store compressed values to enable decoding while waiting for data to be received (e.g., in subsequent rows).
[0081] When compressed data is output in essentially raster scan order rather than in sub-blocks (compressed data blocks packed into each input block in that order), this reduces the amount of buffering required as part of the decompression operation.
[0082] Figure 4 It shows the configuration to implement Figure 2 The method is illustrated in the schematic diagram of an exemplary compression hardware 400 that can be implemented within the compression / decompression unit 112. The compression hardware 400 includes an analysis pipeline 402, a buffer 404 (which may be a FIFO), encoding hardware 406, sub-block packing hardware 408, and tile packing hardware 410.
[0083] Analysis pipeline 402 includes an input terminal configured to receive uncompressed input image data in raster scan order, and the pipeline is arranged to accommodate pixel data of an entire input block (e.g., Figure 3 The example shown contains 8×8 or 16×4 pixel data. Input data can be input to the analysis pipeline 402 at a rate of, for example, 8 pixels per clock cycle, and if a stage in the analysis pipeline holds fewer pixels than the entire input block, the analysis pipeline 402 can include an output buffer (e.g., in the form of a FIFO with read and write pointers) at its output. In other examples, a buffer can be additionally or alternatively incorporated into another location of the analysis pipeline 402 (e.g., at its input) such that pixel data for the entire input block (e.g., the entire tile block) can be contained within the analysis pipeline 402.
[0084] Analysis pipeline 402 is configured to subdivide the input block of pixels into sub-blocks, and in some cases, to further subdivide the sub-blocks into mini-blocks (box 202). As described above, each sub-block and each mini-block includes more than one row of pixels (i.e., m>1 and m'>1). Analysis pipeline 402 is further configured to analyze the pixel data of pixels in the sub-blocks or mini-blocks, and to identify the encoding scheme and / or type of each sub-block or each mini-block within each sub-block (box 204A). Analysis pipeline 402 does not reorder the pixel data; therefore, after analysis, it is configured to output the pixel data to encoding hardware 406 in raster scan order.
[0085] In addition to outputting pixel data, the analysis pipeline 402 is also arranged to generate and output some control data used by the encoding hardware 406. This control data can be directly output to the encoding hardware 406, or the analysis pipeline 402 can be configured to write control data into the buffer 404. The control data includes data identifying the encoding scheme and / or type to be used for each sub-block or mini-block, and may also include other parameters used in the encoding operation. In various examples, the analysis pipeline 402 is arranged to calculate any parameters used for encoding data but based on pixels from different rows in the input data block or pixels that are not adjacent in the raster scan order. This is because only the analysis pipeline 402 has this data available simultaneously. In contrast, the encoding hardware 406 processes pixels in raster scan order. The generated parameters form part of the control data output to the encoding hardware 406 or written to the buffer 404.
[0086] In various examples, the analysis pipe 402 can be implemented as a large shift register, where function blocks are arranged to perform comparisons (e.g., in...). Figure 8 (in box 804) and averaging (e.g., in...) Figure 8 (in box 808).
[0087] Encoding hardware 406 includes inputs for receiving input pixel data from analysis pipeline 402 in raster scan order and inputs for receiving control data from buffer 404. The control data read by encoding hardware 406 from buffer 404 may include data identifying an encoding scheme and / or type to be used for each sub-block or mini-block, and may include other control data, such as palette color data (described below). Encoding hardware 406 is arranged to encode pixel data according to the identified encoding scheme and / or type of a particular sub-block or mini-block, where pixels are part of the generated compressed pixel data (box 204B). Encoding hardware 406 operates in raster scan order rather than sub-block or mini-block by mini-block. Encoding hardware 406 includes an output for outputting compressed pixel data to sub-block packing hardware 408 substantially in raster scan order, and may also include an output for outputting control data to tile packing hardware 410.
[0088] Encoding hardware 406, sub-block packing hardware 408, and tile packing hardware 410 can be arranged to collectively pack compressed pixel data (generated in encoding hardware 406) into a data structure (box 206), and in various examples, these individual elements can be combined into a single encoding and packing hardware element 412. In various examples, encoding hardware 406 can also be arranged to embed one or more control bits within the pixel data before outputting the compressed pixel data (with any embedded control bits) substantially in raster scan order (box 206A). In various examples, one or more padding bits may also be embedded.
[0089] Sub-block padding hardware 408 includes inputs for receiving compressed pixel data (which may include some embedded control and / or padding bits) output by encoding hardware 406. As detailed above, the compressed pixel data is received from encoding hardware 406 substantially in raster scan order; however, as encoding hardware 406 reduces the size (i.e., bit width) of the data for each pixel, "gaps" (e.g., unused bits) exist in the data words between the compressed pixel data. Sub-block packing hardware 408 is arranged to concatenate the compressed pixel data (box 206B) such that the data substantially maintains the raster scan order; that is, the packing hardware is configured to perform packing without reordering the compressed pixel data. This packing operation includes removing the "gaps" and packing all used bits (i.e., bits of the compressed pixel data) together without changing the order of the compressed pixel data.
[0090] Tiling block filling hardware 410 includes one or more input terminals (such as buffer 404 or encoding hardware 406) for receiving compressed pixel data blocks generated by sub-block filling hardware 408 and control data that can be read from buffer 404 or received from encoding hardware 406. Figure 4(As shown by the dashed arrow in the image). The tile fill hardware 410 is arranged to assemble control data block 704 and attach it to compressed pixel data block 702 to form compressed image data block 700 (box 206C). Unlike compressed pixel data block 702, control data block 704 can be arranged sub-blocks, such as... Figure 7 As shown (e.g., having separate portions 706-712 of control data block 704 associated with each sub-block in the input block). The tile block filling hardware 410 also includes an output for outputting compressed image data blocks. The control data in control block 704 includes data different from any embedded control bits within pixel data block 702.
[0091] Figure 5 A schematic diagram of an exemplary decompression hardware 500 configured for decompression is shown. Figure 2 The data is compressed using this method, and this compression / decompression can be implemented within the compression / decompression unit 112. The decompression hardware 500 is configured to implement... Figure 6 The decompression method shown includes unpacking hardware 502, control hardware 504, and decoding hardware 506.
[0092] The unpacking hardware 502 includes an input for receiving compressed image data blocks 700. The compressed data blocks may not be received as a single block, but may be received incrementally over several clock cycles (e.g., depending on the burst size). The unpacking hardware 502 is configured to pass control data blocks 704 (which are the first portion of the compressed data blocks to be received) to the control hardware 504, which can be implemented as a state machine. The control hardware 504 is configured to analyze the control data and, based on the analysis, determine the number of bits in the pixel data block 702 corresponding to each row of pixels in the input block (box 602). The control hardware 504 is configured to relay this information back to the unpacking hardware 502, and in response, the unpacking pipeline is configured to then pass pixel data corresponding to each row of the input block (e.g., input tile blocks) to the decoding hardware 506.
[0093] Control hardware 504 and / or decoding hardware 506 are arranged to determine the encoding scheme and / or type used for each pixel group in each row, and therefore the decoding scheme and / or type needed to decode those pixels (box 604), where a pixel group in a row includes those pixels in rows belonging to the same sub-block (where encoding is performed sub-block by sub-block) or mini-block (where encoding is performed mini-block by mini-block). Identification of the decoding scheme and / or type can be performed by control hardware 504 based on control data block 704, or by decoding hardware 506 based on control bits within embedded compressed pixel data 702 and optionally based on data provided by control hardware 504. In the case where identification is performed by control hardware 504, control hardware 504 is configured to provide this information to decoding hardware 506, and in this example or other examples, control hardware 504 can be configured to provide decoding hardware 506 with additional data it has extracted from the control data block. When the identification is performed by the decoding hardware 506, the control hardware 504 can be configured to provide the decoding hardware 506 with information identifying the location of any embedded control bits within the compressed pixel data.
[0094] Decoding hardware 506 is arranged to decode the corresponding pixel data in raster scan order (box 606) using the identified decoding scheme and / or type and output it as decompressed pixel data. In various examples, decoding hardware 506 is arranged to additionally use other data provided by control hardware 504 to perform decoding. For a particular encoding scheme and / or type (as described above) that causes a deviation from a strict raster scan order, decoding hardware 506 is arranged to store the pixel data received before it is in its correct position in raster scan order, so that it can subsequently be output in the exact raster scan order. In such examples, decoding hardware 506 may be arranged to store encoded pixel data such that it is decoded in raster scan order, or decoding hardware 506 may be arranged to perform decoding in the order in which the encoded pixel data is received, and then store the decoded pixel data before outputting it in the exact raster scan order. As described above, the compression method described herein ensures that decoding hardware 506 always has enough data to decode the next pixel in raster scan order and does not need to wait for data to be provided out of order.
[0095] By using the methods described herein, in various examples, the decompression hardware 500 can be implemented with only pipeline-level control timing, without buffering the compressed pixel data. The control hardware 504 may include a buffer to store control data used by the decoding hardware 506; however, the size of the control data is relatively small compared to the amount of compressed pixel data.
[0096] In various examples, the decoding hardware 506 described above may include multiple identical hardware units operating in parallel on different bits within a pixel row. Where the encoding scheme and / or typology is identified based on each mini-block (where each mini-block comprises n' x m' pixels), there may be N / n' decoding hardware units (e.g., four decoding hardware units for an 8×8 tile block 302 and eight decoding units for a 16×4 tile block 304, such that a decoding hardware unit exists for each mini-block in a row of the input tile block), each configured to decode those n' pixels within a row from a different mini-block. Similarly, where the sub-blocks are not subdivided into mini-blocks, there may be N / n decoding hardware units (e.g., two decoding hardware units for an 8×8 tile block 302 and four decoding units for a 16×4 tile block 304), each configured to decode those n pixels within a row from a different sub-block.
[0097] To further improve the efficiency of the compression / decompression unit 112, compressed pixel data of pixels or pixel pairs (i.e., adjacent pixel pairs from the same row and the same sub-block / mini-block) can be arranged to always be a multiple of P bits (where P is an integer, e.g., P = 10). This can be implemented as part of an encoding scheme and / or format (e.g., by the encoding hardware 406 in box 204B) and / or as part of a packing process (in box 206). In various examples, this can be achieved by embedding control bits (e.g., by the encoding hardware 406 in box 206A) and / or one or more padding bits (e.g., by the encoding hardware 406 before, after, or replacing box 206A).
[0098] By compressing each pixel or pixel pair into multiples of P bits, the hardware used to implement the packing operation (in box 206, and particularly box 206B) is simplified because there is less variability (i.e., fewer possible positions for each pixel or pixel pair to be written), and therefore it is more efficient (e.g., it can be implemented in smaller hardware because the number of multiplexing hardware required is reduced). Similarly, the unpacking operation (in box 602) is simplified because the number of boundary positions between the compressed data of each pixel is significantly reduced (e.g., by a factor of P or P / 2). In particular, this reduces the complexity of the hardware that distributes the compressed pixel data to each decoding hardware unit when multiple identical decoding hardware units exist (e.g., when multiples of P bits are always input to each decoding hardware unit). Reference Figure 4 The data transmitted from the encoding hardware 406 to the sub-block packing hardware 408 can always be a multiple of P bits, and reference... Figure 5 The data transmitted from the unpacking hardware 502 to the decoding hardware 506 can always be a multiple of P bits.
[0099] As described above, the compression of pixel data (in box 204) can use one of the predefined encoding patterns in a set of encoding patterns, and the encoding pattern can be determined per sub-block or per mini-block (in box 204A, and by analysis pipeline 402). Figure 8 This is a flowchart of an exemplary method for analyzing pixel data to select an encoding pattern for use on a 2×2 miniblock. This method can be implemented by analysis pipeline 402 and can alternatively be modified for use on larger miniblocks or subblocks instead of 2×2 miniblocks, but this could result in a significantly larger number of encoding patterns from a predefined set of encoding patterns from which to select an encoding pattern. As described above, when selecting an encoding pattern per subblock / miniblock, the actual encoding is performed in raster scan order.
[0100] like Figure 8 As shown, the method includes calculating the color difference of each pixel pair in a 2×2 mini-block (box 802), where the most similar pixel has the smallest color difference with a given pixel, and where the color difference between a pixel and its neighboring pixels (i.e., between pixel pairs) can be calculated as follows:
[0101] |Red Difference|+|Green Difference|+|Blue Difference|+|α Difference|
[0102] After calculating the three color differences (one color difference per pixel pair) (in box 802), the minimum color difference between any pixel pair in the miniblock is used to determine the miniblock encoding mode to be used (box 802). In this example, two different types of miniblock encoding modes are used depending on whether the minimum color difference (between any pixel pair in the miniblock) exceeds a threshold (e.g., a value that can be set in the range of 0-50, such as 40). If the minimum color difference does not exceed the threshold ("Yes" in box 804), one of six predefined three-color encoding patterns 902-912 is selected (box 806), and an average palette color is calculated (box 808). However, if the minimum color difference exceeds the threshold ("No" in box 804), a four-color mode 914 is selected (box 810). These different miniblock modes 902-914 are... Figure 9 As shown in the diagram. The encoding pattern relies on the assumption that in most mini-blocks, there are no more than three different colors, and in this case, the mini-block can be represented by three palette colors and pixel-to-palette entry assignments. A four-color mode exists to handle other exceptions, and a threshold (used in box 804) provides an approximation to select the mode with the least error.
[0103] As mentioned above, the threshold value can be in the range of 0-50. In various examples, it can be a fixed value set at design time. Alternatively, it can be a variable stored in a global register and changed each time execution... Figure 8The method reads it. This allows the threshold to change dynamically or at least periodically. In various examples, the threshold value can be set based on the results from the training phase. The training phase can use image quality metrics, such as Peak Signal-to-Noise Ratio (PSNR) or Structural Similarity metric (SSIM), to evaluate a set of images compressed using each of the three-color modes 902-912 or the four-color mode 914, and then select the threshold such that the highest image quality metric is obtained overall. It should be understood that in other examples, box 804 can be replaced by different tests.
[0104] like Figure 8 As shown, if the minimum color difference is less than or equal to a threshold ("Yes" in box 804), allowing the use of a three-color coding scheme, then based on the pixel pair with the minimum color difference (in the mini-block), from... Figure 9 Choose the specific assignment type from the set of six shown (box 806). Figure 9 In each pattern, the two pixels shown as shaded are the two pixels with the smallest color difference. If pixels AB have the smallest color difference, the selected pattern is pattern 902, called "Top". If pixels CD have the smallest color difference, the selected pattern is pattern 904, called "Bottom". If pixels AC have the smallest color difference, the selected pattern is pattern 906, called "Left". If pixels BD have the smallest color difference, the selected pattern is pattern 908, called "Right". If pixels AD have the smallest color difference, the selected pattern is pattern 910, called "Diag1". If pixels BC have the smallest color difference, the selected pattern is pattern 912, called "Diag2".
[0105] After selecting the tri-color encoding mode (box 806), the new pixel color is calculated by averaging the source pixel colors of the two pixels in the pixel pair with the minimum color difference (box 808). In various examples, the new pixel color can be calculated in 5555 format, which provides a compression ratio of 50%. If other compression ratios are required, the average pixel data can be calculated with different bit depths. The following pseudocode illustrates an exemplary implementation of this calculation, where the two pixels in the pixel pair are represented as Pix1 and Pix2, and the RGBA values are represented as .red, .grn, .blu, and .alp.
[0106] PIXEL Average(PIXEL Pix1,PIXEL Pix2)
[0107] {
[0108] PIXEL Result;
[0109] Result.red := (Pix1.red + Pix2.red + 1) >> 1;
[0110] Result.grn:=(Pix1.grn+Pix2.grn+1)>>1;
[0111] Result.blu:=(Pix1.blu+Pix2.blu+1)>>1;
[0112] Result.alp:=(Pix1.alp+Pix2.alp+1)>>1;
[0113] Return Result;
[0114] }
[0115] While the pseudocode above includes α values, where the α mode has a constant value or the compression mode used involves a constant α value, the same technique can be used but the α channel can be omitted. The use of constant α values is described in more detail below. Similarly, in the absence of an α channel (e.g., for RGB data), this is omitted from the calculation.
[0116] The average pixel data calculated when the three-color mode is selected (in box 810) is generated by the analysis pipeline 402 and can then be written to the buffer 404 so that it can be subsequently read by the encoding hardware 406 and used when encoding pixel data (in box 204B). The average pixel data is calculated by the analysis pipeline 402 (in box 810) as part of the selected encoding pattern (in box 204A), rather than by the encoding hardware 406, because the encoding hardware 406 processes pixel data essentially in raster scan order. This means that when the encoding hardware 406 encodes pixels from one row, it has no access to pixel data from subsequent rows and therefore cannot calculate the average pixel data for four of the six-color modes 906-912. Therefore, the average pixel data forms part of the control data generated by the analysis pipeline 402 (as described above).
[0117] Similarly, when "top" mode 902 is selected (in box 806), the value of input pixel C is written to buffer 404 so that it can subsequently be read by encoding hardware 406 and used when encoding pixel data (box 204B). As mentioned above, encoding hardware 406 processes pixel data substantially in raster scan order, and therefore, while encoding hardware 406 is encoding pixels from one row, it has no access to pixel data from subsequent rows. If "top" mode is selected, the value of input pixel C thus forms part of the control data generated by analysis pipeline 402 (as described above). When the value of input pixel C is written to the buffer, it can be converted to a different format (e.g., from 8888 format to 5555 format) before being written, or alternatively, any format conversion can be performed by encoding hardware 406 after the value has been read from the buffer.
[0118] In selecting a predefined encoding mode (box 204A, for example, ...), Figure 8 Following the encoding operation (shown in the diagram), the encoding process (in box 204B) will depend on the selected mode, the input data format, and, in various examples, other factors such as the desired compression ratio. Examples of RGBA8888 or ARGB8888 format data are shown in... Figure 10 As shown in the diagram and listed in the table below, it provides a 50% compression ratio. In examples requiring different compression ratios, the conversion stages (in boxes 1004 and 1012) can be modified accordingly. The table below (Table 1) illustrates what the encoded pixel data (which can be represented as output pixel AD) includes for each encoding mode and each input pixel. As can be seen in Table 1, for the tri-color mode, output pixel C or output pixel D contains the data (not both).
[0119] Table 1:
[0120]
[0121] As shown in the table above and Figure 10 As shown, the method takes pixel data of a single pixel as input. This pixel data is received in raster scan order, and each pixel belongs to a mini-block and is pixel A, B, C, or D within that mini-block, and its position within the mini-block (i.e., as pixel A, B, C, or D) affects how the pixel data is encoded. Pixels A and B form part of even-numbered rows (row 0, row 2, ...), where pixel A precedes pixel B in the raster scan order, and pixels C and D form part of odd-numbered rows (row 1, row 3, ...), where pixel C precedes pixel D in the raster scan order.
[0122] If the selected mode is a four-color mode ("Yes" in box 1002) where the pixel is part of a mini-block, the encoding operation (in box 204A) is performed in the same way (box 1004) by converting 8888 format data to 4434 format data, regardless of whether the input pixel is located at position A, B, C, or D within the mini-block. However, if the selected mode is a three-color mode, the encoding operation will differ based on the pixel's position within the mini-block and the selected encoding mode, and for each odd-numbered row, only a single pixel will have encoded data output (e.g., output pixel C or output pixel D). This is because only three colors are output in such modes.
[0123] If the pixel is at position A within the mini-block ("Yes" in box 1006), then if the selected mode is top 902, left 906, or diag1 910 ("Yes" in box 1008), the encoding operation involves reading the average pixel value from buffer 404 (box 1010). Depending on the mode, different average values will be calculated by the analysis pipeline 402 and written to buffer 404 as detailed above. Otherwise, the encoding operation involves converting the pixel data from 8888 format to 5555 format (box 1012).
[0124] If the pixel is at position B in the mini-block (No in box 1006, Yes in box 1014), then if the selected mode is right 908 or diag2 912 (Yes in box 1016), the encoding operation includes reading the average pixel value from buffer 404 (box 1010). Depending on the mode, different average values will be calculated by the analysis pipeline 402 and written to buffer 404 as detailed above. If the selected mode is bottom 904, left 906, or diag1 910 (Yes in box 1018), the encoding operation includes converting the pixel data from 8888 format to 5555 format (box 1012). Otherwise, the encoding operation includes reading the input pixel C data from buffer 404 (box 1019), and if this is not completed before the data is stored in the buffer, the pixel data is converted from 8888 format to 5555 format (box 1012).
[0125] If the pixel is at position C in the mini-block (No in boxes 1006 and 1014, Yes in box 1020), then if the selected mode is bottom 904 (Yes in box 1022), the encoding operation includes reading the average pixel value from buffer 404 (box 1010). Depending on the mode, different average values will be calculated by the analysis pipeline 402 and written to buffer 404 as detailed above. If the selected mode is right 908 or diag1 910 (Yes in box 1024), the encoding operation includes converting the pixel data from 8888 format to 5555 format (box 1012). Otherwise, no encoding operation is performed on the pixel data, and the pixel data is discarded (box 1026).
[0126] Finally, if the pixel is at position D within the mini-block ("No" in boxes 1006, 1014, and 1020), then if the selected mode is top 902, left 906, or diag2 912 ("Yes" in box 1028), the encoding operation includes converting the pixel data from 8888 format to 5555 format (box 1012). Otherwise, no encoding operation is performed on the pixel data, and the pixel data is discarded (box 1026).
[0127] It should be understood that Figure 10 Decisions can be organized in different ways while still achieving the same results (e.g., the results shown in Table 1 above).
[0128] Figure 10 The method compresses each pixel or pixel pair into multiples of P bits, where P = 10. In four-color mode, each compressed pixel pair comprises 30 bits, in three-color mode, each compressed pixel comprises 20 bits, and in two-color mode, each compressed pixel comprises 30 bits. As described below, by ensuring that each pixel or pixel pair is compressed into multiples of P bits, the packing operation is simplified, and therefore the packing hardware is simplified.
[0129] As described above, there are seven different encoding types or patterns 902-914 (six three-color patterns and one four-color pattern), such as Figure 9 As shown. In various examples, an additional two-color mode can exist, where pixels A and B share a palette color that is the average of pixels A and B, and pixels C and D share a palette color that is the average of pixels C and D. Since only two palette colors are stored (compared to three or four in other encoding modes), palette colors can be stored with higher precision (e.g., RGBA7878). This two-palette color mode can be used, for example, if the color difference between pixels A and B and the color difference between pixels C and D are both less than a predefined threshold.
[0130] When using this two-palette color mode, you can Figure 8 An additional test for a predefined threshold (which may differ from the threshold used in box 804) is inserted between boxes 802 and 804. If the test passes, a two-color mode is selected, and the two average pixel colors are calculated and written to a buffer (e.g., in a manner similar to box 808). Alternatively, averaging can be performed in the encoding hardware 406 while averaging two adjacent pixels in the same row; however, this would require averaging both the encoding hardware 406 and the hardware within the analysis pipeline 402, and therefore may be less efficient in terms of hardware size. It should be understood that in other examples, additional two-color modes may exist (e.g., where different pixel pairs share a palette color, such as pixels A and C and pixels B and D).
[0131] In the encoding operation (box 204B) and when using two-color mode, the additional rows provided below in Table 2 have been added to Table 1 above. Similarly, modifications can be made. Figure 10 The system includes an additional decision block such that if a two-color mode has been selected for a pixel at position A or C, the average value is read from the buffer (box 1010), and for a pixel at position B or D, no encoding operation is performed on the pixel data and the pixel data is discarded (box 1026). That is, for each row (whether even or odd), only a single encoded data pixel is output for each pair of input pixels from the same mini-block (e.g., pixel A is output for even rows and pixel C is output for odd rows, where in each row, the output pixel is the average of the two input pixels in that row).
[0132] Table 2:
[0133]
[0134] In various examples, each encoding mode can be identified by a 3-bit encoded value detailed in the table below (Table 3), and these three bits can be written to buffer 404 by analysis pipe 402 and used by encoding hardware 406 when encoding input pixel data. In various examples, these values can be included within control data block 704 (in box 206C) of the encoded pixel data added (e.g., appended or prepended) by tile block filling hardware 410, and in other examples, these values can be embedded in the pixel data (in box 206A).
[0135] Table 3:
[0136] Four colors 101 Two colors 100 top 000 bottom 001 Left side 010 right side 011 Diag1 110 Diag2 111
[0137] like Figure 7As shown, in various examples, the control data 730 of sub-blocks 706-712 may include the coded values 714-720 of the connections of each of the four mini-blocks in the sub-block. For example, if the control data of sub-block 0 includes the coded value 111011000010, then mini-block P has been encoded using the left-hand encoding mode, mini-block Q has been encoded using the top-hand encoding mode, mini-block R has been encoded using the right-hand encoding mode, and mini-block S has been encoded using the Diag2 encoding mode. One or more additional bits 722 (e.g., 2 additional bits) may be present within the control data of the sub-block. Another example of the control data 740 of sub-blocks 706-712 is described below.
[0138] The above describes the data in 8888 format. Figure 8 and Figure 10 The method can also be used for RGBA1010102 or ARGB2101010 format data in the following manner: First, a preprocessing step is applied, which converts the source data into 888Z or Z888 format data with a flag bit, where Z is an integer not greater than 8. Then, boxes 808, 1004, and 1012 are adjusted. In box 808, before averaging, if the flag is set, the 8-bit value is converted to a value of up to 9 bits, and then the new pixel color is calculated in RGBA6652 or ARGB2665 format instead of 5555 format. In box 1004, instead of converting the pixel data to RGBA4434 or ARGB4443 format, the pixel data is converted to RGBA4442 or ARGB2444 format. In box 1012, instead of converting the pixel data to 5555 format, the pixel data is converted to RGBA6652 or ARGB2665 format.
[0139] An exemplary preprocessing method is described in GB2575436, such as... Figure 11 As shown. Figure 11 The method converts pixel data from RGBA1010102 format to RGBA8883 or from ARGB2101010 format to ARGB3888 format (i.e., Z=3), and sets the flag value. For example... Figure 11As shown, the MSB (Maximum Segment Bus) of each RGB channel is checked (box 1102), and if one or more of the three MSBs are equal to one ("Yes" in box 1102), a flag is set (box 1104); otherwise, the flag is not set. This flag can be called the High Dynamic Range (HDR) flag because if at least one MSB is equal to one, then the pixel data is likely HDR data. HDR images can represent a larger range of brightness levels than non-HDR images, and HDR images are typically created by merging multiple low or standard dynamic range (LDR or SDR) photos or by using a special image sensor. The mixing logarithm γ is the HDR standard, which defines a non-linear transfer function where the lower half of the signal value (which is the SDR part of the range) is multiplied by x. 2 The curve, the upper half of the signal value (this is the HDR portion of the range), uses a logarithmic curve, and the white signal standard level is set to 0.5 of the signal value. In the 10 bits of the R / G / B data, the most significant bit indicates whether the value is in the lower half of the range (SDR portion) or the upper half (HDR portion).
[0140] Aside from setting a flag, pixel data is reduced from 10 bits to 8 bits in different ways depending on whether one or more MSBs of the RGB channel are one. If none of the three MSBs are equal to one ("No" in box 1102), each 10-bit value of the RGB channel is truncated by removing both the MSB (known to be zero) and the LSB (box 1110). If any of the three MSBs are equal to one ("Yes" in box 1102), there are two different ways to reduce the 10-bit value to 8 bits (box 1106). In the first example, two LSBs can be removed from each 10-bit value, and in the second example, methods such as... Figure 12 The method shown.
[0141] Figure 12 This is an exemplary method for converting an input of 'a' digits to 'b' digits, where a and b are integers and a > b. For example... Figure 12 As shown, the method includes receiving an input of a number of bits A and truncating the number from a bits to b bits (box 1202). An adjustment value is then determined based on the input of a number of bits A (box 1204), and this can be implemented using multiple AND and OR gates. These AND and OR logic gates (or alternative, functionally equivalent alternative logic arrangements) compare multiple predetermined subsets of the bits of the input of a number of bits with predetermined values in a fixed-function circuit, and determine the adjustment value based on the comparison results, which is then added to the truncated value from block 1202 (box 1206). The value of the adjustment value is zero, one, or negative one.
[0142] In use Figure 12In the case of preprocessing RGBA1010102 or ARGB2101010 format data, a = 10 and b = 8, and the following exemplary VHDL shows how to calculate the adjustment value (in box 1204):
[0143] function CorrectionsFor10to8(i:std_logic_vector(9downto 0))
[0144] return std_logic_vector is
[0145] variable results:std_logic_vector(1downto 0);
[0146] begin
[0147] results: = (others => '0');
[0148] if std_match(i,"00------11")then results:="01";end if;
[0149] if std_match(i,"0-00----11")then results:=results OR"01";end if;
[0150] if std_match(i,"0-0-00--11")then results:=results OR"01";end if;
[0151] if std_match(i,"0-0-0-0011")then results:=results OR"01";end if;
[0152] if std_match(i,"1-1-1-1100")then results:=results OR"11";end if;
[0153] if std_match(i,"1-1-11--00")then results:=results OR"11";end if;
[0154] if std_match(i,"1-11----00")then results:=results OR"11";end if;
[0155] if std_match(i,"11------00")then results:=results OR"11";end if;
[0156] Return results;
[0157] end function CorrectionsFor10to8;
[0158] It is estimated that approximately 25 AND / OR gates are needed to determine the values of the adjustment values for a=10 and b=8 (in box 1204).
[0159] Regardless of the values of the three MSBs of the pixel's RGB channels, modify the 2-bit alpha channel value in the same way. For example... Figure 11 As shown, the HDR flag is appended to the existing 2-bit value (box 1108), making the output α channel value 3 bits.
[0160] Figure 11 The method can be implemented pixel by pixel, but in a variation of the method, the decision that causes the setting of HDR blocks (in box 1102) can be performed less frequently, for example, pixel by sub-block or pixel by mini-block.
[0161] In various examples, to compress each pixel or pixel pair into a multiple of P bits, where P = 10 for RGBA1010102 or ARGB2101010 format data, the encoding and / or packing operation may include adding one or more padding bits. In four-color mode, each compressed pixel pair includes 30 bits, and therefore no padding bits are needed. In three-color mode, each compressed pixel includes 20 bits, and again no padding bits are needed; however, in two-color mode, each compressed pixel includes 27 bits, so three padding bits (e.g., 000) can be added, making each compressed pixel include 30 bits. As described below, by ensuring that each pixel or pixel pair is compressed into a multiple of P bits, the packing operation is simplified, and therefore the packing hardware is simplified. The unpacking operation (and hardware) is also simplified.
[0162] In all the examples above, the pixel data includes four channels: red, green, blue, and α (although they can be in different orders), and these compression methods can be collectively described as variable α modes. This does not mean that α must vary between pixels, but rather that α is specified individually for each pixel and therefore can vary. However, in other examples, one of several constant α compression modes can be used instead, and a corresponding constant α mode can exist for each of the above variable α modes. In the constant α mode, the pixel data compressed using the above methods includes only three channels: red, green, and blue, and the α channel values are processed separately.
[0163] In various examples where both variable and constant α modes are available for selection (in box 204), additional stages may be included, such as Figure 13 As shown. Figure 13 As shown, before selecting an encoding mode for a sub-block or mini-block (in box 204A), it is determined on a per-sub-block basis whether the sub-block is compressed using a variable alpha mode or a constant alpha mode (box 1302). In the case of using both constant alpha and variable alpha modes, control block 704 may include a data field 722 (e.g., at the beginning of the data for each sub-block 706-712 within control block 704, such as...). Figure 7 As shown), it indicates whether a variable or constant α pattern is being used for each sub-block, and in various examples, this can be a 2-bit field for each sub-block.
[0164] exist Figure 14 and Figure 15 Two exemplary methods for determining whether to use a constant α or a variable α are shown and described below (in box 1302). In the first exemplary method, as Figure 14 As shown, the α value of each pixel within the sub-block is analyzed, and two parameters, minalpha and maxalpha, are calculated, which are the minimum and maximum values of α for all pixels in the sub-block (box 1402). These can be determined in any way, including, for example, using loops (as shown in the pseudocode example below, or its functional equivalent) or using a test tree, where a first step determines the maximum and minimum α values for pixel pairs, and a second step determines the maximum and minimum α values for the output pairs from the first step, etc. These two parameters (minalpha and maxalpha) are then used in the subsequent decision-making process (box 1404), and, if it is determined that a constant α mode should be selected, the value of α that should be used for the sub-block is determined (boxes 1406-1408).
[0165] The decision regarding whether to use a constant or variable alpha involves evaluating the range of alpha values over sub-blocks against a threshold alphadifftol (box 1404). This test determines whether the range is greater than the error introduced by using a (best-case) variable alpha mode (e.g., due to the additional compression applied to pixel data to achieve the same compression ratio), and the magnitude of these errors is represented as alphadifftol and can be predetermined. The value of alphadifftol can be determined during training by comparing the quality loss caused by different methods in the variable alpha mode (i.e., 4-color encoding with 4-bit alpha or 3-color encoding with 5-bit alpha and two pixels sharing the same color) (hence the use of the above term "best-case"). Alternatively, the value of alphadifftol can be determined (also during training) by evaluating different candidate values over a large set of test images to find the candidate value that provides the best results using visual comparison or image difference metrics. The value of alphadifftol can be fixed or can be programmable.
[0166] In response to determining that the range is greater than the error introduced by using the (best case) variable alpha mode ("Yes" in box 1404), a variable alpha compression mode is applied to the sub-block. However, in response to determining that the range is not greater than the error introduced by using the (best case) variable alpha mode ("Yes" in box 1404), a constant alpha compression mode is applied to the sub-block. In the latter case, two additional decision operations (boxes 1406, 1408) can be used to determine the value of alpha for the entire sub-block. If the value of maxalpha is the maximum possible value of alpha (e.g., 0xFF, "Yes" in box 1406), then the alpha value (constalphaval) used in the constant alpha mode is set to that maximum possible value (box 1410). This ensures that if any pixels are completely opaque, they remain completely opaque after the data is compressed and subsequently decompressed. If the value of minalpha is zero (e.g., 0x00, "Yes" in box 1408), then the alpha value (constalphaval) used in the constant alpha mode is set to zero (box 1412). This ensures that if any pixels are completely transparent, they remain completely transparent after the data is compressed and subsequently decompressed. If none of these conditions are met ("No" in boxes 1406 and 1408), the average value of α is calculated for the pixels in the sub-block (box 1414), and this average value is used in constant α mode.
[0167] For example, the following pseudocode (or its functional equivalent) can be used to implement... Figure 14 The analysis shown in this code, where P.alp is the α value of the pixel P being considered:
[0168] CONST Alphadifftol = 4;
[0169] U8 Minalpha = 0xFF;
[0170] U8 Maxalpha = 0x00;
[0171] U12 AlphaSum = 0;
[0172] FOREACH Pixel, P, in the 4×4 block
[0173] Minalpha = MIN(P.alp, Minalpha);
[0174] Maxalpha = MAX(P.alp, Maxalpha);
[0175] AlphaSum += P.alp;
[0176] ENDFOR
[0177] IF ((Maxalpha–Minalpha) > Alphadifftol) THEN
[0178] Mode = VariableAlphaMode;
[0179] ELSE IF (Maxalpha == 0xFF)
[0180] Mode = ConstAlphaMode;
[0181] Constalphaval = 0xFF;
[0182] ELSE IF (Minalpha == 0x00)
[0183] Mode = ConstAlphaMode;
[0184] Constalphaval = 0x00;
[0185] ELSE
[0186] Mode = ConstAlphaMode;
[0187] Constalphaval = (AlphaSum + 8) >> 4;
[0188] ENDIF
[0189] It should be understood that although the decision process is in Figure 14 The test is shown in a specific order, but in other examples, the same test can be applied in a different order (e.g., assuming alphadifftol < 254, boxes 1406 and 1408 can be interchanged). Furthermore, it should be understood that the test in box 1404 can alternatively be maxalpha > (minalpha + alphadifftol).
[0190] Figure 15 An alternative example implementation of the analysis phase (box 1302) is shown. In this example, the parameter `constalphaval` is initially set to the α value of the pixel at a predefined location within the sub-block (box 1502). For example, `constalphaval` could be set to the α value of the pixel at the top-left corner of the sub-block (i.e., the first pixel in the sub-block). Then, all α values of the other pixels in the sub-block are compared to this `constalphaval` (box 1504). If all α values are very similar to `constalphaval` (e.g., within ±5, marked "yes" in box 1504), a constant α mode is used; otherwise, if their variation is greater than this (marked "no" in box 1504), a variable α mode is used. Then, for the constant α mode, the values are compared with... Figure 14 The parameter constalphaval is calculated in a similar way. Figure 15 This includes setting it to zero (in box 1508) or the maximum value (in box 1512), where all pixels are almost completely transparent (constalphaval < 5, "Yes" in box 1506) or almost completely opaque (constalphaval > 250, "Yes" in box 1510). It should be understood that... Figure 15 The specific values used as part of the analysis (e.g., in boxes 1504, 1506, and 1510) are provided only as examples, and these values may be slightly different in other examples.
[0191] and Figure 14 Compared to the methods, Figure 15 This method does not require determining minalpha and maxalpha, which reduces the computational workload required to perform the analysis. However, Figure 15 The method may produce some visible artifacts (such as aliasing), especially when the object moves slowly on the screen and "constant α" tiles are unlikely to be detected because a predefined position is used as the center of the α value.
[0192] Regardless of how the constant α value is calculated (by analysis pipeline 402), the value is then written to buffer 404 so that it can subsequently be included in control data block 704 (in box 206C) added (e.g., appended or pre-loaded) by tile-filled hardware 410.
[0193] After determining whether to use a constant or variable α (box 1302), the method for compressing the sub-box data continues, as follows: Figure 13 As shown in the diagram. For each sub-block or mini-block, the pixel data continues to be analyzed to select the encoding pattern (box 204A), and this can be referenced as above. Figure 8 The implementation described above differs only in that the calculated color difference and average pixel data (in boxes 802 and 808) do not include any consideration for the alpha channel, and the average pixel data is determined at a higher resolution, such as RGB676 format instead of RGBA5555. When still using a two-color mode, the average pixel data is determined at a higher resolution and compared to a variable alpha equivalent, such as RGB888 format instead of RGBA8787.
[0194] In selecting a mode from the predefined encoding modes (box 204A, for example, such as...) Figure 8 Following the sequence shown in Figure 204B, the encoding operation will depend on the selected mode, the input data format, and, in various examples, other factors such as the desired compression ratio. Figure 10 The RGBA8888 or ARGB8888 format data shown below differs only slightly from the variable α example when using a constant α mode instead of a variable α mode. This is described below specifically for RGB888 data and is also listed in Table 4 below (listed for each encoding mode and each input pixel; the encoded pixel data includes these and also includes...). Figure 10 (Two color modes not shown).
[0195] As shown in Table 4 below, when the four-color mode is selected, in the variation of box 1004, the pixel data is converted from RGB888 to RGB554 format. And when the three-color mode is selected and the pixel data is converted, in the variation of box 1012, the pixel data is converted from RGB888 to RGB676 format.
[0196] Table 4:
[0197]
[0198] When using the constant alpha mode, the way data is packaged into a data structure (in box 206) can differ from the way it is used with the variable alpha mode, and the nature of the added (e.g., additional or pre-added) control data 704 (in box 206C) can also differ. Figure 7 As shown, in various examples, the control data 740 for sub-blocks 706-712 may include a field 722 that identifies the alpha mode for sub-block 724 (as described above), a constant alpha value (constalphaval), and a set of fields 726 (e.g., four 1-bit fields) indicating whether a two-color or four-color mode has been selected for each mini-block in the sub-block. This, in turn, provides an indication of the number of bits per row of mini-blocks. In various examples, this field may include a 1 if a two-color or four-color mode has been used, and a zero if a three-color mode has been used. In examples using the encoded values provided above and shown in Table 3, this bit value may be determined as follows:
[0199] EV[2] AND NOT EV[1]
[0200] The three bits of the encoded value are represented as EV[2], EV[1] and EV[0], where EV[2] is the MSB and EV[0] is the LSB.
[0201] The example of control data 740 can be used for RGB888 and RGB8888 format data, where a constant α exists. In contrast, for RGB888 and RGB8888 format data with variable α, and for RGBA1010102 and ARGB2101010 format data, the example of control data 730 (previously described) can be used regardless of whether α is constant or variable.
[0202] As described above, to further improve the efficiency of the compression / decompression unit 112, the compressed pixel data of pixels or pixel pairs (i.e., a pair of adjacent pixels from the same row and the same sub-block / mini-block) can be arranged to always be a multiple of P bits (where P is an integer, e.g., P = 10). In the case of using the constant α mode, Figure 10 The method does not compress each pixel or pixel pair to a multiple of P bits. In four-color mode, each compressed pixel includes 14 bits, in three-color mode, each compressed pixel includes 19 bits, and in two-color mode, each compressed pixel includes 24 bits. To improve efficiency, one or more control bits, and in some cases padding bits, can be embedded within the pixel data (box 206A) so that the compressed data of the pixel or pixel pair is a multiple of P bits.
[0203] In the example, with P=10, one control / padding bit is added to each compressed pixel, so that in four-color mode, a pair of compressed pixels has a total of 30 bits, and in three-color mode, each compressed pixel includes 20 bits. By adding one control / padding bit to each pixel, the total number of bits per pixel in two-color mode increases from 24 to 25, thus adding another five bits (e.g., one control bit and four or five padding bits) to make each compressed pixel include 30 bits.
[0204] In addition to improving efficiency, the embedded control bits also facilitate the decompression operation by ensuring that each pixel or pixel pair is a multiple of P. The added control bits can be the encoded value, EV, or bits, where two bits of the encoded value are embedded in the pixel data of each even-numbered row (row 0, row 2, ...), and the remaining bits of the encoded value are embedded in the pixel data of the subsequent odd-numbered rows (row 1, row 3, ...), as shown in Table 5 below.
[0205] Table 5:
[0206]
[0207]
[0208] By embedding these encoded bits, they can be used by the decompression hardware 502 in conjunction with a single field in the control data when decompressing the data to distinguish the encoding mode (in box 604), as shown in Table 6 below.
[0209] Table 6:
[0210]
[0211] In both scenarios, it is impossible to distinguish between the two different tri-color coding modes based on the bits in the control data combined with the control bits embedded in the even-numbered rows. This will not affect the decompression operation because for both the left and Diag1 modes (embedded bits = 10), the even-numbered rows are decompressed in the same way (e.g., compressed pixel A corresponds to the average of the two used for decompressing pixel A and a pixel in the odd-numbered row, and compressed pixel B corresponds to decompressed pixel B), and similarly, for both the right and Diag2 modes (embedded bits = 11), the even-numbered rows are decompressed in the same way (e.g., compressed pixel A corresponds to decompressed pixel A, and compressed pixel B corresponds to the average of the two used for decompressing pixel B and a pixel in the odd-numbered row). When decompressing pixels in the odd-numbered rows, the decompression operation uses the control bits (EV[2]) embedded in the odd-numbered rows to distinguish between the left and Diag1 modes and the right and Diag2 modes.
[0212] The table below (Table 7) shows examples of the positions of the embedded control bits for various packing modes and for even and odd rows. As shown in the table below, the positions of the embedded control bits are the same for two-color mode and four-color mode. This is to assist in the decompression operation, since the bits in control block 704 only indicate whether two-color mode or four-color mode has been used, and the first embedded control bit EV[0] indicates which of the two modes has been used (i.e., if EV[0] is one, four-color mode is used to compress the pixel data, and if EV[0] is zero, two-color mode is used).
[0213] Table 7:
[0214]
[0215] Table 7 also illustrates that the bits included in the control data block 704, which identifies when a two-color / four-color mode is used, also indicate the compression bit depth corresponding to a row of pixels (in box 602). In this example, for two-color / four-color mode, each mini-block row includes 30 bits, and for three-color mode, even-numbered rows include 50 bits and odd-numbered rows include 20 bits.
[0216] Although Table 7 above shows how compressed pixel data for mini-blocks is packed into pixel data blocks (box 206), as detailed above, the compressed pixel data is packed primarily in raster scan order rather than mini-block order. For an 8×8 tile block 302, pixel data is packed using the symbols [SB,MB,R] as follows, where SB identifies sub-blocks, where SB = [0,1,2,3], MB identifies mini-blocks, where MB = [P,Q,R,S], and R identifies rows, where R = [E,O], E = even, O = odd:
[0217] [0,P,E],[0,Q,E],[1,P,E],[1,Q,E],
[0218] [0,P,O],[0,Q,O],[1,P,O],[1,Q,O],
[0219] [0,R,E],[0,S,E],[1,R,E],[1,S,E],
[0220] [0,R,O],[0,S,O],[1,R,O],[1,S,O],
[0221] [2,P,E],[2,Q,E],[3,P,E],[3,Q,E],
[0222] [2,P,O],[2,Q,O],[3,P,O],[3,Q,O],
[0223] [2,R,E],[2,S,E],[3,R,E],[3,S,E],
[0224] [2,R,O],[2,S,O],[3,R,O],[3,S,O].
[0225] Similarly, for a 16×8 tile 304, the pixel data is packaged as follows:
[0226] [0,P,E],[0,Q,E],[1,P,E],[1,Q,E],[2,P,E],[2,Q,E],[3,P,E],[3,Q,E],
[0227] [0,P,O],[0,Q,O],[1,P,O],[1,Q,O],[2,P,O],[2,Q,O],[3,P,O],[3,Q,O],
[0228] [0,R,E],[0,S,E],[1,R,E],[1,S,E],[2,R,E],[2,S,E],[3,R,E],[3,S,E],
[0229] [0,R,O],[0,S,O],[1,R,O],[1,S,O],[2,R,O],[2,S,O],[3,R,O],[3,S,O].
[0230] When the input data is in RGBA1010102 format or ARGB2101010 and a constant alpha mode is used, the same method as for RGBA8888 or ARGB8888 format data can be used, but in implementation... Figure 8 and Figure 10 Prior to this method, a preprocessing step (as described above) can be used to convert the source data into 888 format data with flag bits. Additionally, the averaging function changes: as mentioned above, if a flag is set (to realign values according to HDR values) before averaging, the 8-bit value is converted to a maximum of 9-bit values. For RGBA8888 or ARGB8888 format data, no further changes to the compression method described above are required.
[0231] In various examples, to compress each pixel or pixel pair into a multiple of P bits, where P = 10 for RGBA1010102 or ARGB2101010 format data, and in the case of using constant α mode, the encoding and / or packing operations may include adding one or more padding bits. In four-color mode, each compressed pixel pair comprises 30 bits, and no padding bits are required. In three-color mode, each compressed pixel comprises 20 bits, and again no padding bits are required; however, in two-color mode, each compressed pixel comprises 25 bits, and therefore five padding bits (e.g., 00000) are added, making each compressed pixel comprise 30 bits.
[0232] In all the examples above, P = 10. In other examples, this can be achieved by changing the bit depth used (e.g., Figure 10 Different values of P can be used by adding different numbers of padding bits (in the middle) and / or adding different numbers of padding bits. Similarly, the above method can be applied to other formats of RGBA / RGB / ARGB data by modifying the bit depth used.
[0233] All the examples above involve RGBA or ARGB data. If the input data is RGB (i.e., there is no α channel), the constant α mode described above can be used (decision points are omitted in box 1302) without calculating or storing a constant α value.
[0234] The above description outlines various compression methods, and the inverse of these methods is used to decompress the data (e.g., the relevant encoding modes are identified in box 606 and box 604). (Return to reference shown) Figure 10 Tables 1, 2, and 4 show the implementation methods and variations of the encoding. For the four-color mode 914, decompression involves converting the reduced bit depth data of each of the four colors back to the desired format, and any suitable method can be used. For the two-color mode, an average encoded pixel color is received, and this pixel color is used for two adjacent pixels, which can then be converted to the data format again. For the three-color mode, for each pair of pixels, two encoded pixel values are received in the even-numbered rows and one encoded pixel value is received in the odd-numbered rows. Depending on the specific encoding mode, these pixel colors are assigned to the decompressed pixels in different ways and can be converted to the data format again.
[0235] Table 8 below shows which decompressed pixel values are used for each of the decompressed output pixels A, B, C, and D in the mini-block. Similar to the previous tables, although this is shown per mini-block, the pixels are decompressed and output in raster scan order, such that pixels A and B are decompressed and output in even-numbered rows, and pixels C and D are decompressed and output in odd-numbered rows. As mentioned above, even-numbered rows always include both compressed pixel values for each mini-block, while odd-numbered rows may include one or both compressed pixel values for each mini-block. This means that when decompression is performed, the pixel data in even-numbered rows always includes all the pixel data required to output the two decompressed output pixels in that row, and may also include some advance data for pixel pairs in subsequent rows used to decompress the same mini-pixel. This advance data can be stored within the decoding hardware 506. There is no situation where decompression cannot occur in raster scan order because the data required for decompression is later stored in the data stream.
[0236] Table 8:
[0237]
[0238] Figure 16 A computer system is illustrated in which the data compression and decompression methods and apparatus described herein can be implemented. The computer system includes a CPU 1602, a GPU 1604, a memory 1606, and other devices 1614, such as a display 1616, a speaker 1618, and a camera 1620. A data compression and / or decompression block 1621 (which can implement any of the methods described herein) is implemented on the GPU 1604. In other examples, the data compression and / or decompression block 1621 may be implemented on the CPU 1602. Components of the computer system can communicate with each other via a communication bus 1622.
[0239] Figure 4 and Figure 5 The data compression hardware is shown as comprising multiple functional blocks. This is merely illustrative and not intended to define a strict division between different logical elements of such an entity. Each functional block can be provided in any suitable manner. It should be understood that the intermediate values formed by the data compression hardware described herein do not need to be physically generated by the data compression hardware at any point in time, and may only represent logical values that conveniently describe the processing performed by the data compression hardware between its inputs and outputs.
[0240] The data compression and decompression hardware described herein (including any hardware arranged to implement any of the methods described above) can be implemented in hardware on an integrated circuit. The data compression and decompression hardware described herein can be configured to perform any of the methods described herein. Generally, any of the functions, methods, techniques, or components described above can be implemented in software, firmware, hardware (e.g., a fixed logic circuit system), or any combination thereof. The terms “module,” “function,” “component,” “element,” “cell,” “block,” and “logic” are used herein to generally denote software, firmware, hardware, or any combination thereof. In the case of a software implementation, a module, function, component, element, cell, block, or logic represents program code that, when executed on a processor, performs a specified task. The algorithms and methods described herein can be executed by one or more processors that execute code that causes the processor to perform the algorithm / method. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), optical disk, flash memory, hard disk storage, and other memory devices that can use magnetic, optical, and other techniques to store instructions or other data and can be accessed by a machine.
[0241] As used herein, the terms computer program code and computer-readable instructions refer to any kind of executable code for a processor, comprising code expressed in machine language, interpreted language, or scripting language. Executable code includes binary code, machine code, bytecode, code defining integrated circuits (e.g., hardware description languages or netlists), and code expressed in programming languages such as C, Java, or OpenCL. Executable code can be, for example, any kind of software, firmware, script, module, or library that, when properly executed, processed, interpreted, compiled, or executed in a virtual machine or other software environment, causes the processor of a computer system that supports the executable code to perform tasks specified by said code.
[0242] A processor, computer, or computer system can be any kind of device, machine, or special-purpose circuit, or a collection or portion thereof, having processing power that enables it to execute instructions. A processor can be any kind of general-purpose or special-purpose processor, such as a CPU, GPU, system-on-a-chip, state machine, media processor, application-specific integrated circuit (ASIC), programmable logic array, field-programmable gate array (FPGA), physical processing unit (PPU), radio processing unit (RPU), digital signal processor (DSP), general-purpose processor (e.g., general-purpose GPU), microprocessor, any processing unit designed to accelerate tasks other than those performed by the CPU, etc. A computer or computer system may include one or more processors. Those skilled in the art will recognize that this processing power is incorporated into many different devices; therefore, the term "computer" includes set-top boxes, media players, digital radios, PCs, servers, mobile phones, personal digital assistants, and many other devices.
[0243] This invention also intends to cover software, such as HDL (Hardware Description Language) software, that defines the configuration of hardware as described herein for designing integrated circuits or configuring programmable chips to perform desired functions. That is, a computer-readable storage medium may be provided on which computer-readable program code in the form of an integrated circuit definition dataset is encoded, which, when processed (i.e., executed) in an integrated circuit manufacturing system, configures the system to manufacture data compression and / or decompression hardware configured to perform any of the methods described herein, or to manufacture data compression and / or decompression hardware including any of the means described herein. The integrated circuit definition dataset may, for example, be an integrated circuit description.
[0244] Therefore, a method for manufacturing data compression and / or decompression hardware as described herein can be provided in an integrated circuit manufacturing system. Furthermore, an integrated circuit definition dataset can be provided, which, when processed in the integrated circuit manufacturing system, enables the method for manufacturing the data compression and / or decompression hardware to be executed.
[0245] Integrated circuit definition datasets can be in the form of computer code, such as as a netlist, code for configuring programmable chips, or a hardware description language for defining integrated circuits at any level, including register-transfer level (RTL) code, high-level circuit representations such as Verilog or VHDL, and low-level circuit representations such as OASIS (RTM) and GDSII. Higher-level logical representations of integrated circuits (such as RTL) can be processed at a computer system configured to generate manufacturing definitions of integrated circuits within a software environment that includes definitions of circuit elements and rules for combining those elements to generate manufacturing definitions representing such integrated circuits. As is typically the case where software executes at a computer system to define a machine, one or more intermediate user steps (e.g., providing commands, variables, etc.) may be required to configure the computer system to generate manufacturing definitions of integrated circuits, executing the code that defines the integrated circuits to generate the manufacturing definitions of the integrated circuits.
[0246] Now refer to Figure 17 Describe an example of processing integrated circuit definition datasets at an integrated circuit manufacturing system in order to configure the system as hardware for manufacturing data compression and / or decompression.
[0247] Figure 17 An example of an integrated circuit (IC) manufacturing system 1702 is shown, configured to manufacture data compression and / or decompression hardware as described in any of the examples herein. Specifically, the IC manufacturing system 1702 includes a layout processing system 1704 and an integrated circuit generation system 1706. The IC manufacturing system 1702 is configured to receive an IC definition dataset (e.g., defining data compression and / or decompression hardware as described in any of the examples herein), process the IC definition dataset, and generate an IC (e.g., embodying the data compression and / or decompression hardware as described in any of the examples herein) based on the IC definition dataset. Through the processing of the IC definition dataset, the IC manufacturing system 1702 is configured to manufacture integrated circuits embodying the data compression and / or decompression hardware as described in any of the examples herein.
[0248] The layout processing system 1704 is configured to receive and process an IC definition dataset to determine a circuit layout. Methods for determining a circuit layout based on an IC definition dataset are known in the art and may involve, for example, synthesizing RTL code to determine a gate-level representation of the circuit to be generated, for example, in relation to logic components (e.g., NAND, NOR, AND, OR, MUX, and FLIP-FLOP components). By determining the location information of the logic components, the circuit layout can be determined based on the gate-level representation of the circuit. This can be done automatically or with user intervention to optimize the circuit layout. Once the layout processing system 1704 has determined the circuit layout, it can output the circuit layout definition to the IC generation system 1706. The circuit layout definition may be, for example, a circuit layout description.
[0249] As is known in the art, IC generation system 1706 generates ICs according to a circuit layout definition. For example, IC generation system 1706 can implement a semiconductor device manufacturing process for generating ICs, which may involve a multi-step sequence of photolithography and chemical processing steps, during which electronic circuits are gradually formed on a wafer made of semiconductor material. The circuit layout definition may be in the form of a mask, which can be used in the photolithography process to generate ICs according to the circuit definition. Alternatively, the circuit layout definition provided to IC generation system 1706 may be in the form of computer-readable code, which IC generation system 1706 can use to form a suitable mask for generating ICs.
[0250] The various processes performed by the IC manufacturing system 1702 can all be implemented in one location, for example, by one party. Alternatively, the IC manufacturing system 1702 can be a distributed system, allowing some processes to be performed in different locations and by different parties. For example, some of the following stages can be performed in different locations and / or by different parties: (i) synthesizing RTL codes representing an IC definition dataset to form a gate-level representation of the circuit to be generated; (ii) generating a circuit layout based on the gate-level representation; (iii) forming a mask based on the circuit layout; and (iv) using the mask to manufacture the integrated circuit.
[0251] In other examples, by processing the integrated circuit definition dataset at the integrated circuit manufacturing system, the system can be configured to manufacture data compression and / or decompression hardware without processing the IC definition dataset to determine circuit layout. For example, the integrated circuit definition dataset can define the configuration of a reconfigurable processor such as an FPGA, and processing of that dataset can configure the IC manufacturing system to (e.g., by loading configuration data into the FPGA) generate a reconfigurable processor with that defined configuration.
[0252] In some implementations, when processed in an integrated circuit manufacturing system, an integrated circuit manufacturing definition dataset can enable the integrated circuit manufacturing system to generate devices as described herein. For example, using an integrated circuit manufacturing definition dataset, as described above regarding... Figure 17 The configuration of the integrated circuit manufacturing system described herein enables the production of equipment as described in this article.
[0253] In some examples, an integrated circuit definition dataset may include software running on hardware defined at the dataset, or software running in combination with hardware defined at the dataset. Figure 17 In the example shown, the IC generation system can be further configured by the integrated circuit definition dataset to load firmware onto the integrated circuit according to the program code defined at the integrated circuit definition dataset during the manufacturing of the integrated circuit, or otherwise provide the integrated circuit with program code for use with the integrated circuit.
[0254] Those skilled in the art will recognize that the storage devices used to store program instructions can be distributed across a network. For example, a remote computer can store examples of the described process as software. A local or terminal computer can access the remote computer and download part or all of the software to run the program. Alternatively, a local computer can download fragments of software as needed, or execute some software instructions at a local terminal while executing others at a remote computer (or computer network). Those skilled in the art will also recognize that, by utilizing conventional techniques known to them, all or part of the software instructions can be executed by dedicated circuitry such as DSPs, programmable logic arrays, etc.
[0255] The applicant has independently disclosed each individual feature described herein, as well as any combination of two or more such features, to the extent that such features or combinations can be implemented based on the specification as a whole, in accordance with the common knowledge of those skilled in the art, regardless of whether such features or combinations of features solve any problem disclosed herein. In view of the foregoing description, those skilled in the art will understand that various modifications can be made within the scope of this invention.
Claims
1. A data compression method, the data compression method comprising: The input pixel data of the data block is received in raster scan order, and the pixel data includes at least a first channel data, a second channel data and a third channel data for each pixel; The pixel data is compressed essentially in raster scan order using a block-based coding scheme; as well as Basically, compressed pixel data is output in raster scan order. The block-based coding scheme, which compresses the pixel data essentially in raster scan order, includes: The input block of pixels is subdivided into partial blocks, each partial block comprising pixels in more than one row; and The pixel data is compressed using a block-based coding scheme and the subdivisions essentially in raster scan order. Basically, the raster scan order refers to the sorting of pixel data, which is exactly the raster scan order for most encoding schemes, and in a small number of encoding schemes, isolated pixel values are output before their proper output positions according to the exact raster scan order.
2. The method of claim 1, wherein compressing the pixel data using a block-based coding scheme and the subdivision substantially in raster scan order comprises: For each partial block, the pixel data of the pixels in the partial block is analyzed to identify the coding pattern used on the pixels in the partial block; as well as The selected method encodes the pixel data essentially in raster scan order.
3. The method of claim 2, wherein a partial block comprises 2×2 pixels, and wherein analyzing pixel data to identify the encoding pattern on pixels within the partial block comprises, for each partial block: Calculate the difference between each pair of pixels in the partial block; In response to determining that the minimum difference does not exceed a predefined threshold, a ternary encoding pattern is selected from a set of ternary encoding patterns based on the pixel pair having the minimum difference, and the average pixel data of the pixels in the pixel pair having the minimum difference is calculated and stored in a buffer; as well as In response to determining that the minimum difference exceeds a predefined threshold, a four-value encoding pattern is selected.
4. The method of claim 3, wherein analyzing pixel data to identify the coding pattern on pixels in the partial block further comprises, for each partial block: In response to determining that the minimum difference does not exceed a predefined threshold and that the pixel pair having the minimum difference is a pixel pair in the first row of the partial block, the pixel data of the first pixel in the second row of the partial block is stored in the buffer.
5. The method of claim 2, wherein encoding the pixel data using the selected type substantially in raster scan order comprises: The pixel data is encoded essentially in raster scan order using the selected pattern, such that each pixel or adjacent pixel pair from the same partial block is compressed into a sequence of multiples of P bits, where P is an integer.
6. The method of claim 5, wherein the pixel data is encoded substantially in raster scan order using the selected pattern, such that each pixel or adjacent pixel pair from the same partial block is compressed into a sequence of multiples of P bits, comprising: The selected method encodes the pixel data essentially in raster scan order; and Pixels or adjacent pixel pairs from the same partial block are compressed into a sequence of bits less than a multiple of P, and one or more control and / or padding bits are embedded to increase the sequence length to a multiple of P.
7. The method of claim 6, wherein a portion of the block comprises 2×2 pixels, and wherein encoding the pixel data substantially in raster scan order using a selected pattern comprises: For adjacent pixel pairs from the first row of a partial block, and in the case of using the ternary encoding format, the output is a sequence of pixel values including the two encoded values; as well as For adjacent pixel pairs from the second row of a partial block, and in the case of using ternary encoding, the output consists of a sequence of encoded pixel values.
8. The method according to claim 7, wherein: A sequence of two encoded pixel values includes any of the following: The average pixel data of the pixel pairs in the partial block read from the buffer, and the converted pixel data of one pixel from the adjacent pixel pair in the first row of the partial block; or Transformed pixel data for each pixel in the adjacent pixel pairs from the first row of the partial block; and A sequence of encoded pixel values includes any of the following: Average pixel data of pixel pairs in the partial block read from the buffer; or Transformed pixel data from one of the adjacent pixel pairs in the second row of the partial block.
9. The method of claim 1, wherein outputting compressed pixel data substantially in raster scan order comprises: The compressed pixel data is packaged into a data structure in essentially raster scan order; as well as Output the data structure.
10. The method of claim 9, wherein packing the compressed pixel data into a data structure substantially in raster scan order comprises: The compressed pixel data is basically connected in raster scan order to form compressed pixel data blocks; as well as Connect the compressed pixel data block and the control data block.
11. The method of claim 10, wherein connecting the compressed pixel data block and the control data block comprises: The control data block is appended to or prepended to the compressed pixel data block.
12. The method of claim 10, further comprising, before concatenating the compressed pixel data: One or more bits of the control data are embedded within the pixel data.
13. The method of claim 1, wherein the first channel data, the second channel data and the third channel data of each pixel include red channel data, green channel data and blue channel data of each channel.
14. A data compression hardware, the data compression hardware comprising: The input terminal is used to receive input pixel data of data blocks in raster scan order, wherein the pixel data includes at least a first channel data, a second channel data and a third channel data for each pixel; Hardware logic, which is arranged to compress the pixel data substantially in raster scan order using a block-based coding scheme; as well as The output terminal is used to output compressed pixel data substantially in raster scan order. The hardware logic mentioned above includes: An analysis pipeline, arranged to subdivide an input block of pixels into partial blocks, each partial block comprising more than one row of pixels; and Encoding hardware, configured to compress the pixel data using a block-based coding scheme and the subdivision substantially in raster scan order. Basically, the raster scan order refers to the sorting of pixel data, which is exactly the raster scan order for most encoding schemes, and in a small number of encoding schemes, isolated pixel values are output before their proper output positions according to the exact raster scan order.
15. The data compression hardware according to claim 14, wherein: The analysis pipeline is further arranged such that, for each partial block, pixel data of pixels in the partial block is analyzed to identify the encoding pattern used on the pixels in the partial block, and The encoding hardware is configured to encode the pixel data substantially in raster scan order using a selected method.
16. The data compression hardware of claim 14, further comprising a buffer, wherein the analysis pipeline is further arranged to store data in the buffer for use by the encoding hardware, and wherein the encoding hardware is arranged to read the stored data from the buffer when encoding the pixel data.
17. The data compression hardware of claim 14, further comprising packing hardware arranged to pack the compressed pixel data substantially in raster scan order into a data structure and output the data structure.
18. The data compression hardware of claim 17, wherein the packing hardware comprises: Partial block packing hardware, the partial block packing hardware being arranged to connect the compressed pixel data substantially in raster scan order to form compressed pixel data blocks; as well as Tiled block packaging hardware, the tiled block packaging hardware being arranged to connect the compressed pixel data block and the control data block.
19. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed at a computer system, cause the computer system to perform the method according to any one of claims 1 to 13.
20. A computer-readable storage medium having stored thereon a computer-readable description of an integrated circuit, wherein when the computer-readable description is processed in an integrated circuit manufacturing system, the computer-readable description causes the integrated circuit manufacturing system to manufacture data compression hardware according to any one of claims 14-18.
Citation Information
Patent Citations
Guaranteed data compression
GB2575436A
Guaranteed data compression
CN110662049A