Data compression method and device, electronic equipment and readable storage medium

By preprocessing and combining coding schemes on the graphics processor data, differential coding and Huffman coding are optimized, the problems of waste of hardware area and high encoding failure rates in graphics processor data compression are solved, and efficient data compression and hardware savings are achieved.

CN120034664APending Publication Date: 2025-05-23LOONGSON TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411977055.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When processing frame buffer and depth buffer data, the graphics processor uses different compression algorithms and structures to cause waste of hardware area, and traditional differential coding and Huffman coding have high failure rates and low adaptability in some cases.

Method used

A data compression method is proposed. By pre-processing the frame buffer and depth buffer data, the combination scheme of differential encoding, general encoding and entropy encoding is adopted to optimize the differential encoding and Huffman encoding algorithms to improve the compression rate and success rate.

Benefits of technology

It realizes general lossless compression of graphics processor data, saves hardware area, and improves the success rate and adaptability of differential coding and Huffman coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034664A_ABST
    Figure CN120034664A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data compression method and device, electronic equipment and a readable storage medium, and the method comprises the steps: carrying out the preprocessing operation of to-be-processed original data, and obtaining to-be-processed target data; the original data is a Tile block in a frame buffer area or a depth buffer area; inputting the target data into a first coding module for differential coding to obtain a first data set and a second data set; the first data set comprises incomplete difference data and deviation points obtained after differential coding; the second data set comprises complete difference data obtained after differential coding and coordinate data of a deviation point; inputting the first data set into a second coding module for universal coding to obtain a second coding result; inputting the second data set into a third coding module for entropy coding to obtain a third coding result; and splicing the second coding result and the third coding result to obtain compressed data. The hardware area can be saved, and the compression rate can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a data compression method, device, electronic equipment and readable storage medium. Background Art

[0002] In the field of graphics technology, Graphics Processing Unit (GPU) occupies an increasingly important position in computing technology with its high throughput and high bandwidth characteristics. Data compression plays an important role in improving the memory access performance of GPU.

[0003] The graphics data that the graphics processor needs to process mainly comes from the frame buffer and the depth buffer. In order to match different data characteristics, the two often need to use different compression algorithms and structures. For example, in the depth compression direction, the compression algorithm generally tends to have good linear characteristics of the data; while in the frame buffer compression direction, the compression algorithm generally tends to have characteristics such as the sequence of data within a certain range. Summary of the invention

[0004] In view of the above problems, an embodiment of the present invention is proposed to provide a data compression method that overcomes the above problems or at least partially solves the above problems. The method is applicable to general data, can save hardware area, and improve compression rate.

[0005] Correspondingly, an embodiment of the present invention further provides a data compression device, an electronic device, and a computer program product to ensure the implementation and application of the above method.

[0006] In a first aspect, an embodiment of the present invention discloses a data compression method, the method comprising:

[0007] Performing a preprocessing operation on the raw data to be processed to obtain target data to be processed; the raw data to be processed is a Tile block in the frame buffer or the depth buffer;

[0008] The target data is input into a first encoding module for differential encoding to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold;

[0009] Inputting the first data set into a second encoding module for universal encoding to obtain a second encoding result; and inputting the second data set into a third encoding module for entropy encoding to obtain a third encoding result;

[0010] The second encoding result and the third encoding result are concatenated to obtain compressed data of the original data.

[0011] In a second aspect, an embodiment of the present invention discloses a data compression device, the device comprising:

[0012] A preprocessing module, used for performing preprocessing operations on the raw data to be processed to obtain target data to be processed; the raw data to be processed is a Tile block in the frame buffer or the depth buffer;

[0013] A first encoding module is used to input the target data into the first encoding module for differential encoding to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold;

[0014] A second encoding module, used for performing universal encoding on the first data set to obtain a second encoding result;

[0015] A third encoding module, used for performing entropy encoding on the second data set to obtain a third encoding result;

[0016] The merging output module is used to concatenate the second encoding result and the third encoding result to obtain compressed data of the original data.

[0017] In the third aspect, an embodiment of the present invention discloses an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of any of the data compression methods described above.

[0018] In a fourth aspect, an embodiment of the present invention discloses a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the data compression method as described in any of the above can be implemented.

[0019] In a fifth aspect, an embodiment of the present invention discloses a computer program product, including a computer program, which, when executed by a processor, performs the steps of any of the data compression methods described above.

[0020] The embodiments of the present invention include the following advantages:

[0021] The data compression method provided by the embodiment of the present invention is a universal lossless compression solution for image data suitable for an image processor, which can process data compression in a frame buffer and a depth buffer at the same time, solves the problem of area overhead caused by using different compression algorithms and processing structures in the frame buffer and the depth buffer, and can save hardware area.

[0022] In addition, the data compression method provided in the embodiment of the present invention optimizes differential coding. A preprocessing operation based on breakpoint detection is performed before differential coding, and the compression environment of the differential coding algorithm is improved by adding deviation point data in differential coding, and the tolerance for non-compliant pixel points (deviation points) is increased, and the coordinate data format of the additionally stored deviation points also maintains the characteristic of being able to be compressed twice. Therefore, compared with traditional differential coding, the optimization strategy of the optimized differential coding scheme of the present invention has low cost and high efficiency, improves the adaptability of the differential coding algorithm to the graphics processor, and improves the success rate of differential coding.

[0023] Furthermore, the data compression method provided by the embodiment of the present invention optimizes Huffman coding. By constructing a symbol set containing all possible members and recording the information of each parent node of the Huffman tree, the symbol correspondence table and the codeword correspondence table that the Huffman tree originally needs to transmit are combined into one, and only the symbol-codeword correspondence table (i.e., the third encoding result) needs to be transmitted. The data width required by the symbol-codeword correspondence table is smaller, thereby improving the compression rate of Huffman coding. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a flowchart of DDPCM encoding in an example;

[0025] Figure 2 This is a flowchart of Huffman coding in an example;

[0026] Figure 3 is a flow chart of steps of an embodiment of a data compression method of the present invention;

[0027] Figure 4 is a schematic diagram of the architecture of the universal lossless compression scheme of the present invention;

[0028] Figure 5 It is a schematic diagram of adding deviation points to DDPCM coding in the present invention;

[0029] Figure 6 This is a schematic diagram of the existence of primitive boundaries in a Tile block;

[0030] Figure 7 It is a schematic diagram of the process of Huffman coding in an example of the present invention;

[0031] Figure 8It is a schematic diagram of combining and encoding each level of parent node information to obtain a symbol-codeword correspondence table in an example of the present invention;

[0032] Fig. 9 is a structural block diagram of an embodiment of a data compression device of the present invention;

[0033] Fig.10 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable when appropriate, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of the present invention, the term "multiple" refers to two or more, and other quantifiers are similar.

[0036] DDPCM (Differential-Differential Pulse Coded Modulation) is a compression algorithm that uses double row difference and double column difference. DDPCM encoding performs the following differential operation on the tile block:

[0037] 1. First-order (or first-order) sequence difference;

[0038] 2. Double (or second-order) column difference is performed synchronously with the first and second row differences;

[0039] 3. Second-order difference of the first row.

[0040] Reference Figure 1 , showing a schematic diagram of the DDPCM encoding process in an example. In the graphic image processing technology, the image to be processed is usually divided into smaller, regular-shaped (usually square or rectangular) areas, and these small areas are called tiles. Figure 1The Tile block size shown is 4×4 (indicating that the Tile block has 4 pixels per row, 4 pixels per column, and a total of 16 pixels). In practice, Tile blocks of sizes such as 4×8 (4 rows and 8 columns, a total of 32 pixels) or 8×8 (8 rows and 8 columns, a total of 64 pixels) can also be used.

[0041] like Figure 1 As shown, first, a first-order column difference operation is performed on the Tile block (mat_i) to obtain a first-order column difference matrix mat_1c, and then a second-order column difference operation is performed on mat_1c to obtain a second-order column difference matrix mat_2c, and then a first-order row difference operation is performed on the first two rows of mat_1c to obtain first-order row difference matrices mat_1r1 and mat_1r2, and finally a second-order row difference operation is performed on the first row of data of mat_1r1 to obtain a second-order row difference matrix mat_2r. Based on the above matrices, the compressed data of the Tile block can be obtained as mat_o.

[0042] The success of DDPCM encoding depends on the linear characteristics of the input image data, which means that the change amplitude of the image data cannot be too large and lacks anti-interference ability. For frame buffer images, their data can be attributed to the combination of texture information and interpolation in a certain proportion, so frame buffer images usually do not have good linear characteristics, which will lead to DDPCM encoding failure. In addition, in image data with good linear characteristics, data with a second-order difference transform amplitude exceeding the expected threshold will directly lead to DDPCM encoding failure, thereby affecting the success rate of DDPCM encoding.

[0043] Huffman Coding, also known as Huffman coding, is a variable length coding (VLC). Huffman coding represents characters with higher frequencies with shorter codes and characters with lower frequencies with longer codes, thereby achieving efficient data compression.

[0044] Reference Figure 2 , shows a flow chart of Huffman coding in an example. Figure 2 As shown, assume that the data to be compressed is: ABACAA. First, count the frequency of each character in the data to be compressed. Then construct a Huffman tree based on the frequency, where characters with higher frequencies are located on shorter paths, and characters with lower frequencies are located on longer paths. Finally, generate the code corresponding to each character based on the Huffman tree, thereby encoding characters with higher frequencies into shorter bit strings, and characters with lower frequencies into longer bit strings.

[0045] Reference Figure 3, shows a flow chart of steps of an embodiment of a data compression method of the present invention, the method may include the following steps:

[0046] Step 101, preprocessing the original data to be processed to obtain target data to be processed; the original data to be processed is a Tile block in the frame buffer or the depth buffer;

[0047] Step 102: input the target data into a first encoding module for differential encoding to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold;

[0048] Step 103: input the first data set into a second encoding module for general encoding to obtain a second encoding result; and input the second data set into a third encoding module for entropy encoding to obtain a third encoding result;

[0049] Step 104: concatenate the second encoding result and the third encoding result to obtain compressed data of the original data.

[0050] The data compression method provided by the embodiment of the present invention is a universal lossless compression solution for image data suitable for an image processor GPU. It can process data compression in a frame buffer and a depth buffer at the same time, solves the problem of area overhead caused by using different compression algorithms and processing structures in the frame buffer and the depth buffer, and can save hardware area.

[0051] Reference Figure 4 , shows a schematic diagram of the architecture of the universal lossless compression solution of the present invention. Figure 4 As shown, the architecture includes a preprocessing module, a first encoding module (differential encoding module), a second encoding module (universal encoding module), a third encoding module (entropy encoding module), and a merging output module.

[0052] The preprocessing module is used to perform preprocessing operations on the raw data to be processed to obtain the target data to be processed. The preprocessing operations may include but are not limited to preprocessing operations based on breakpoint detection and / or preprocessing operations based on MSAA (MulitiSampling Anti-Aliasing) detection. The preprocessing operations can be used to improve the efficiency and success rate of differential encoding.

[0053] The first encoding module is also called a differential encoding module, which is used to perform differential encoding on the received data to obtain a first encoding result.

[0054] Differential coding is to realize compression coding by using the difference between pixel values. The embodiment of the present invention does not limit the type of differential coding. For example, the differential coding may include DDPCM or DPCM (Differential Pulse Code Modulation).

[0055] The first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold.

[0056] Incomplete difference data includes original pixel points that do not participate in the difference calculation during the difference coding process or difference data obtained by incomplete difference calculation. For example, for DDPCM coding, since DDPCM coding requires double (i.e., second-order) difference calculations of row difference and column difference. Therefore, after DDPCM coding, incomplete difference data includes original pixel points that do not participate in the difference calculation and first-order difference data. Since DDPCM coding requires second-order difference calculation, the first-order difference data obtained by only performing first-order difference calculation is called difference data obtained by incomplete difference calculation, and the second-order difference data is called complete difference data. That is, for DDPCM coding, the first data set includes the original pixel points, first-order difference data and deviation points obtained after DDPCM coding; the second data set includes the second-order difference data obtained after DDPCM coding and the coordinate data of the deviation points; the deviation point is the original pixel point where the second-order difference data exceeds the target difference threshold. For DDPCM coding, the target difference threshold refers to the second-order difference threshold.

[0057] For DPCM encoding, since DPCM encoding only needs to perform row difference or column difference calculation, whether row difference or column difference is set according to actual needs, DPCM only needs to perform first-order difference calculation. Therefore, after DPCM encoding, the incomplete difference data only includes the original pixel points that are not involved in the difference calculation, and the complete difference data includes the first-order difference data. That is, for DPCM encoding, the first data set includes the original pixel points and deviation points obtained after DPCM encoding; the second data set includes the first-order difference data obtained after DPCM encoding and the coordinate data of the deviation points; the deviation point is the original pixel point where the first-order difference data exceeds the target difference threshold. For DPCM encoding, the target difference threshold refers to the first-order difference threshold.

[0058] It should be noted that, since the compression rate of DDPCM coding is greater than that of DPCM coding, and the failure rate of DDPCM coding is less than that of DPCM coding, the embodiment of the present invention mainly uses DDPCM coding as an example for explanation. The difference between DPCM coding and DDPCM coding is that the first-order difference calculation is less performed, and the other steps can be referred to each other.

[0059] The second encoding module is used to perform universal encoding on the received data to obtain a second encoding result. Furthermore, the embodiment of the present invention does not limit the universal encoding algorithm adopted by the second encoding module. Exemplarily, a dictionary-based compression algorithm (Dictionary-based compression) can be adopted, or other encoding compression algorithms can be adopted. The second encoding module only needs to maintain a stable compression rate for the image data with non-linear features. Preferably, the second encoding module of the embodiment of the present invention adopts a dictionary-based compression algorithm, and the second encoding module is also called a dictionary encoding module, which is used to perform dictionary encoding on the received data to obtain a second encoding result.

[0060] The third encoding module is also called an entropy encoding module, which is used to perform entropy encoding on the received data to obtain a third encoding result.

[0061] Entropy coding is to reduce the redundant information in the data and make the average code length after coding as close to the information entropy of the source as possible, so as to achieve data compression. The embodiment of the present invention does not limit the entropy coding algorithm used by the second coding module. Exemplarily, the embodiment of the present invention is mainly described by taking the entropy coding as Huffman coding as an example. Huffman coding is a VLC type entropy coding algorithm. Of course, in the specific implementation, other VLC type entropy coding algorithms, such as arithmetic coding (Arithmetic Coding), etc. can also be used.

[0062] The merging output module is used to perform splicing processing on the second encoding result and the third encoding result to obtain compressed data after compressing the target data.

[0063] Further, Figure 4 The illustrated architecture may further include a post-processing module, which is a hardware module. The data input to the post-processing module may include multiple types of data, such as data in RGBA8 format or RGB10A2 format, etc., divided by color format; such as data pre-processed based on breakpoint detection, or pre-processed based on MSAA detection, etc., divided by pre-processing effect. The post-processing module identifies the type of received data and inputs the data to be input to the Huffman coding module into the Huffman coding module.

[0064] In the related art, when a Tile block is encoded with DDPCM, as long as there is a data exceeding the target difference threshold (second-order difference threshold), the DDPCM encoding will fail directly, greatly affecting the success rate of DDPCM encoding. Similarly, when a Tile block is encoded with DPCM, as long as there is a data exceeding the target difference threshold (first-order difference threshold), the DPCM encoding will fail directly.

[0065] The embodiment of the present invention optimizes differential coding. For DDPCM coding, certain tolerance measures are provided for data whose second-order difference variation is greater than the expected threshold (second-order difference threshold) to improve the success rate of DDPCM coding. Similarly, for DPCM coding, certain tolerance measures are provided for data whose first-order difference variation is greater than the expected threshold (first-order difference threshold) to improve the success rate of DPCM coding. Specifically, the embodiment of the present invention records additional deviation points by adding a coding suffix to improve the tolerance of DDPCM coding to individual data exceeding the second-order difference threshold, or to improve the tolerance of DPCM coding to individual data exceeding the first-order difference threshold.

[0066] Taking DDPCM encoding as an example, when a Tile block is DDPCM encoded, when a data point exceeding the second-order difference threshold appears, the judgment of encoding failure will not be triggered immediately, and the information of the deviation point will be recorded additionally. The embodiment of the present invention does not limit the way of recording the deviation point. For example, a bit can be added at the end of the original DDPCM encoding result to record the information of the deviation point. Exemplarily, the recording method can adopt the format of "row coordinates + column coordinates + original data of the point", and the row coordinates and column coordinates need to fill the width to the width represented by the second-order difference threshold. After such processing, the coordinate data of the deviation point can be treated as an "extra second-order difference member" after a successful DDPCM encoding, so that the coordinates of the deviation point can be compressed using Huffman coding in the next stage. When the number of deviation points accumulates to a certain number (such as exceeding a preset number), the judgment of DDPCM encoding failure is triggered, thereby greatly improving the tolerance of DDPCM encoding to deviation points, thereby improving the success rate of DDPCM encoding.

[0067] It should be noted that the width described in the embodiments of the present invention, such as the width of the row coordinates and the width of the column coordinates, refers to the bits occupied by the data. For example, for a 4×4 Tile, the Tile includes 4 rows and 4 columns, a total of 16 pixels. (i, j) represents the coordinates of the pixel in the i-th row and j-th column, where i is the row coordinate and j is the column coordinate. If the row coordinate and the column coordinate each occupy 3 bits, the width of the row coordinate is called 3 bits, and the width of the column coordinate is called 3 bits.

[0068] Reference Figure 5, which shows a schematic diagram of adding deviation points to DDPCM coding according to an embodiment of the present invention. Figure 5 The Tile block shown has two deviation points, point0 and point1, after DDPCM encoding. Record the encoding results and information about the deviation points. The information about the deviation points includes: the original pixel point of point0 (the value of the pixel point) and the column coordinate data tile_col and row coordinate data tile_row of point0; the original pixel point of point1 (the value of the pixel point) and the column coordinate data tile_col and row coordinate data tile_row of point1.

[0069] like Figure 5 As shown, although adding the coding suffix to record the information of the deviation point brings additional data width, it greatly improves the success rate of DDPCM encoding. Moreover, since the coordinate data of the deviation point also has good statistical characteristics, it can be compressed at a secondary level in the subsequent stage to reduce the loss of compression rate, thereby improving the overall compression rate.

[0070] In an example, assuming that a Tile block is in RGBA8 format, the Tile block is split according to the color channel, and DDPCM encoding is performed on each color channel. The Tile block data of a color channel is 8 bits. For example, if the first-order difference threshold is configured to 8 bits (the same as the original data width, the first-order difference actually has no compression benefit, so the first-order difference cannot exceed the first-order difference threshold, and the first-order difference will not be judged as encoding failure). Assuming that the second-order difference threshold is configured to 3 bits (the data is cropped from 8-bit width to 3-bit width, and there is compression benefit), if the second-order difference exceeds 3 bits, it is recorded as a deviation point. If the number of deviation points does not exceed the preset number, it is considered that the DDPCM encoding of the Tile block is successful; if the number of deviation points exceeds the preset number, it is considered that the DDPCM encoding of the Tile block has failed.

[0071] It should be noted that, in a specific implementation, if DDPCM encoding is used, the width of the first-order difference threshold can be set to be consistent with the width of the original pixel, that is, the first-order difference data does not generate compression benefits; the width of the second-order difference threshold is set to be smaller than the width of the original pixel, that is, the second-order difference data generates compression benefits, so as to improve the compression rate and reduce the failure rate of encoding. At this time, the deviation point refers to the original pixel point whose second-order difference exceeds the second-order difference threshold.

[0072] If DPCM encoding is used, since only the first-order difference calculation is required, the width of the first-order difference threshold can be set smaller than the width of the original pixel, that is, the first-order difference data generates compression benefits. At this time, the deviation point refers to the original pixel whose first-order difference exceeds the first-order difference threshold.

[0073] Furthermore, the embodiment of the present invention further inputs the data obtained after DDPCM encoding into the Huffman encoding module for secondary compression. Since Huffman encoding is based on members (members refer to data participating in Huffman encoding, such as each data in the second data set input to the third module), each member has a fixed data width (referring to the bits occupied by the data), therefore, in order to allow the coordinate data of these deviation points to also participate in Huffman encoding, the embodiment of the present invention adjusts the width of the coordinate data of the deviation points to be the same as the second-order difference threshold, that is, the same as the width of other successfully encoded second-order difference data, thereby the coordinate data of these deviation points can be used together with the second-order difference data as members of Huffman encoding for Huffman encoding.

[0074] For example, assuming that the second-order difference threshold is 3 bits, if the coordinate width of the deviation point is less than 3 bits, the coordinate width of the deviation point is padded to 3 bits. If the coordinate width of the deviation point exceeds 3 bits, it means that the Tile block granularity has become larger (because the 3-bit data width can already express all the coordinates of the 4×4 Tile block), for example, the Tile block changes from 4×4 to 8×8 or 4×8. At this time, in order to ensure the success rate of compression, the second-order difference threshold can be adjusted, such as adjusting the second-order difference threshold from the original 3 bits to 4 bits or 5 bits, so that the second-order difference threshold can express all the coordinates of the enlarged Tile block. In actual use, the second-order difference threshold can be configured according to the size of the Tile block so that it can express all the coordinates of the current Tile block.

[0075] After the target data is DDPCM-encoded by the first encoding module, if the DDPCM encoding is successful, a first encoding result is obtained; the first encoding result includes a first data set and a second data set; the first data set includes original pixel points obtained after DDPCM encoding (1, such as Figure 1 The data A in mat_o in the target data set is included, first-order difference data (2) and deviation points; the second data set includes second-order difference data ((number of pixels in the Tile block - 3)) obtained after DDPCM encoding and coordinate data of the deviation points; the deviation points are original pixel points in the target data whose second-order difference exceeds the second-order difference threshold.

[0076] In this embodiment of the present invention, if the target data is DDPCM encoded successfully, the length of the data after DDPCM encoding is composed of: original pixel width × 1 + first-order difference data width × 2 + second-order difference data width × (number of pixels in the Tile block - 3) + deviation point data row and column coordinate width × number of deviation points × 2 + deviation point data × number of deviation points.

[0077] In the embodiment of the present invention, the width of the first-order difference data is consistent with the width of the original pixel point, and the row and column coordinate width of the deviation point data is consistent with the width of the second-order difference data. After the target data DDPCM encoding is successful, the first encoding result obtained can be divided into two groups of data: the first data set and the second data set. The first data set consists of 1 original pixel point, 2 first-order difference data, and deviation point data. The width of the data members in the first data set is the original pixel point width of the target data. The first data set can also be called the original width data set. The second data set consists of (the number of pixels in the Tile block-3) second-order difference data and (the number of deviation points × 2) deviation point row and column coordinate data. The width of the data members in the second data set is the width represented by the second-order difference threshold. The second data set can also be called the difference width data set. The data width of the second data set output by the first encoding module is consistent with the member width specified by the Huffman coding used in the secondary compression, and has the conditions for using Huffman coding for secondary compression.

[0078] The data after DDPCM encoding is a fixed-length data set, most of which is second-order difference data, and still has certain statistical characteristics. The embodiment of the present invention uses Huffman coding to perform secondary compression on the data after DDPCM encoding, which can further improve the compression rate. In addition, after DDPCM encoding, the original width data set and the data that failed DDPCM encoding can be input into the second encoding module for dictionary encoding, and the difference width data set can be input into the third encoding module for Huffman coding for secondary compression. The target data often has good statistical characteristics after the second-order difference operation of DDPCM encoding, and the majority of pixels maintain good linear characteristics. Therefore, compared with directly performing Huffman coding on the original data, the compression rate of Huffman coding can be improved to a greater extent.

[0079] Finally, the second encoding result and the third encoding result are concatenated to obtain compressed data of the original data.

[0080] The data compression method of the embodiment of the present invention uses Tile blocks as the granularity unit, but the data in a Tile block may be split, and some data in a Tile block will enter the Huffman coding module, and some data will enter the dictionary coding module. Since the data length of different coding algorithms may be different, the data length of the second coding result and the third coding result may be different. The embodiment of the present invention splices the second coding result and the third coding result through a merging output module to obtain the compressed data of the Tile block.

[0081] The splicing process refers to tight splicing, that is, splicing the valid data in the results output by the second encoding module and the third encoding module. Since the width of the valid data in different registers may be different in hardware implementation, invalid data may be introduced if the data in the registers are directly spliced. For example, the data width of register data_a is 1024 bits, and the data width of register data_b is 1024 bits, but the valid data in data_a actually has 512 bits, and the valid data in data_b actually has 10 bits. If the data in these two registers are directly spliced, a large amount of invalid data will be introduced. To avoid this problem, the splicing process of an embodiment of the present invention may include: splicing the valid data in the register storing the second encoding result with the register storing the third encoding result.

[0082] In an optional embodiment of the present invention, the preprocessing operation of the raw data to be processed may include:

[0083] Step S11, performing breakpoint detection on the Tile block to determine whether there is a primitive boundary in the Tile block;

[0084] Step S12: If there is a primitive boundary in the Tile block, determine the Tile corner containing the most repeated data in the Tile block according to the primitive boundary;

[0085] Step S13, performing a preset operation on the Tile block so that the Tile corner is located at the target position of the Tile block to obtain target data to be processed; the target position is determined according to the order of difference operations in the differential encoding process, and the target position is a position in the Tile block that does not participate in the difference operation; the preset operation includes a pixel mapping operation or a Tile block flipping operation.

[0086] In a specific implementation, for a Tile block, the location where the linear feature of the data is broken is usually located at the primitive boundary (such as the boundary between the primitive and the background). If there is a primitive boundary in the Tile block, then within the Tile block, the primitive boundary divides the Tile block into two blocks of data with different features. The primitive boundary contains primitive data with good linear features, and the primitive boundary contains data without good linear features, such as background (single color) data, data of another primitive, or texture data.

[0087] Reference Figure 6 , which shows a schematic diagram of the existence of primitive boundaries in a Tile block. Figure 6 As shown in (a), the pixel boundary in the tile block divides the tile block into two parts. The pixel boundary consists of the detected breakpoints. Figure 6In (a), the detected breakpoints include A, F, K, and P, and the primitive boundary consists of points A, F, K, and P. Inside the primitive boundary are primitive data with good linear features, such as A, B, C, D, F, G, H, K, L, and P. Outside the primitive boundary are data without good linear features, such as E, I, J, M, N, and O.

[0088] The data outside the primitive boundary usually contains a large amount of repeated data. If there is a primitive boundary in the Tile block, the Tile corner containing the most repeated data in the Tile block can be determined based on the primitive boundary, such as Figure 6 (a) shows the lower left corner (shaded part) of the Tile block. This Tile corner contains repeated data, and for this Tile block, this Tile corner contains the most repeated data.

[0089] Since differential coding (such as DDPCM coding) performs difference operations on Tile blocks by row and column, Figure 1 As shown, the difference operation of DDPCM encoding has the following characteristics: the difference operation for columns is from top to bottom in the Tile block, and the difference operation for rows is from left to right in the Tile block. Therefore, the data in the lower right corner of the Tile block will not participate in the difference operation.

[0090] The embodiment of the present invention is based on the characteristics of the difference operation of DDPCM encoding. If a primitive boundary is detected in a Tile block, the Tile corner containing the most repeated data in the Tile block is determined according to the primitive boundary; and a preset operation is performed on the Tile block so that the Tile corner is located at the target position of the Tile block (such as the lower right corner of the Tile block) to obtain the target data to be processed. For example, for Figure 6 The Tile block shown in (a) is subjected to a preset operation to obtain the Tile block shown in 6(b), and the Tile block shown in 6(b) is input into the DDPCM encoding module as the target data. Therefore, the data in the shaded part that does not have good linear characteristics will not participate in the difference operation of the DDPCM encoding, thereby improving the success rate of the DDPCM encoding.

[0091] Furthermore, the target position can be determined according to the order of difference calculation in the DDPCM encoding process, and the target position is the position in the Tile block that does not participate in the difference calculation. For example, if the order of difference calculation in the DDPCM encoding is as follows: Figure 1 As shown, the target position is the lower right corner of the Tile block. Depending on the order of the difference operation in DDPCM encoding, the target position may also be the lower left corner of the Tile block.

[0092] Optionally, a preset operation is performed on the Tile block, and the preset operation may include but is not limited to a pixel mapping operation or a Tile block flipping operation. The embodiment of the present invention moves the data that does not have good linear characteristics in the Tile block to the target position of the Tile block (such as the lower right corner) through a preprocessing operation based on breakpoint detection. Due to the sequential characteristics of the difference operation of the DDPCM encoding, this part of the data will not participate in the difference operation, and therefore will not affect the DDPCM encoding process of the data that has good linear characteristics. In addition, since after the preset operation, the data that does not have good linear characteristics will not participate in the difference operation, the success rate of the DDPCM encoding can be further improved.

[0093] It should be noted that the preset operation is not limited to the pixel mapping operation or the Tile block flipping operation. For example, it may also include the Tile block rotation operation, or the use of neighboring points to complete the deviation points. Among them, operations such as pixel mapping, Tile block flipping or rotation are achieved by guiding the deviation points in the Tile block and moving them to the target position of the Tile block, so as to improve the encoding success rate of the DDPCM encoding for the Tile block with some nonlinear feature data. The method of using neighboring points to complete the deviation points does not require guiding the deviation points. Since the neighboring points are pixel points that are close to the deviation point position and the DDPCM encoding is successful, there is no need to change the position of the pixel points in the Tile block, and the effect of improving the encoding success rate of the DDPCM encoding for the Tile block with some nonlinear feature data can also be achieved.

[0094] In an optional embodiment of the present invention, performing breakpoint detection on the Tile block to determine whether there is a primitive boundary in the Tile block may include:

[0095] Step S21, using the four corners of the Tile block as starting points, respectively using a preset algorithm to perform duplicate data detection;

[0096] Step S22, scoring each corner according to the number of repeated data contained in each detected corner;

[0097] Step S23: If the maximum score of the four corners of the Tile block exceeds a preset score threshold, it is determined that there is a primitive boundary in the Tile block; otherwise, it is determined that there is no primitive boundary in the Tile block.

[0098] In an embodiment of the present invention, for a Tile block, the four corners of the Tile block are used as starting points, and duplicate data detection is performed separately using a preset algorithm. After the detection is completed, the four corners will each hold a set of detection data (such as bp_data), and the Tile corner containing the most duplicate data is determined based on the detection data of each corner. The embodiment of the present invention does not limit the preset algorithm. Exemplarily, the preset algorithm can be a flood fill algorithm, or an improved algorithm of the flood fill algorithm, etc.

[0099] In an example, a 4×4 Tile block is taken as an example, and Table 1 shows the value of each pixel in the Tile block. The four corners of the Tile block are the upper right corner (RU), the lower right corner (RD), the upper left corner (LU), and the lower left corner (LD).

[0100] Table 1

[0101] 1(LU) 1 3 4(RU) 2 1 4 4 2 2 6 7 2(LD) 2 2 7(RD)

[0102] First, taking the four corners as starting points, we can use algorithms such as flood fill to perform duplicate data detection respectively, and obtain the detection data bp_data corresponding to the four corners respectively.

[0103] Furthermore, the present invention improves the flood fill algorithm and adds a range limit for detecting duplicate data as follows. Specifically, for a certain starting point, the range for detecting duplicate data may include: the maximum length and width of the duplicate data coordinate point does not exceed the duplicate value of the Tile edge whose corresponding starting point is zero. That is, among the duplicate data coordinate points (x, y) that can be detected, taking LU as an example, x-1, x-2, etc., must be the same value until the edge, and the same applies to y. Taking RU as the starting point as an example, the value of RU is 4, and the point with a duplicate value of 4 is found. The point in the 4th column of the 2nd row is 4, which is a duplicate value. Although the value of the 3rd column of the 2nd row is also 4, since the value of the 3rd column of the 1st row is not 4, the point in the 3rd column of the 2nd row cannot be considered as duplicate data of RU. Taking LD as an example again, the value of LD is 2, and the point with a duplicate value of 2 is found. The value of the 2nd column of the 3rd row is 2, and the data from the point to the left to the edge is 2, and the value from the point to the edge is also 2, so this point is the duplicate data of LD. The value of the 3rd column of the 4th row is also 2, and the value of this point is 2 all the way to the left to the edge, and the point downward is the edge, so this point is also the duplicate data of LD. The purpose of setting the range of detecting duplicate data in the embodiment of the present invention is to make the target Tile block finally obtained after mapping or flipping, and the duplicate points must be continuously distributed in the target position (such as the lower right corner) in the Tile block, to ensure that the point that can be discarded by the DDPCM encoding must be the last item of the difference operation of the row and column of the DDPCM encoding, so as to improve the success rate of subsequent DDPCM encoding.

[0104] The embodiment of the present invention uses four corners as starting points, uses a flood fill algorithm to detect duplicate data respectively, and scores each corner according to the detected duplicate data, and determines that the corner with the highest score is the tile corner containing the most duplicate data. If the scores of the four corners are the same, LU is selected by default.

[0105] Taking Table 1 as an example, RU is scored as 2, RD is scored as 2, LU is scored as 2, and LD is scored as 6. For RU, its own value is 4, and the value of the 4th column of the 2nd row is 4. There are 2 duplicate data in total, and the score is 2. For RD, its own value is 7, and the value of the 4th column of the 3rd row is 7. There are 2 duplicate data in total, and the score is 2. For LU, its own value is 1, and the value of the 2nd column of the 1st row is 1. There are 2 duplicate data in total, and the score is 2. For LD, its own value is 2, and the values ​​of the 1st column of the 2nd row, the 1st column of the 3rd row, the 2nd column of the 3rd row, the 2nd column of the 4th row, and the 3rd column of the 4th row are all 2. There are 6 duplicate data in total, and the score is 6. Therefore, it is determined that the Tile corner containing the most duplicate data is LD. Perform a preset operation on the Tile block shown in Table 1, such as flipping the Tile block. The flipped Tile block is shown in Table 2, and Table 2 is the target data to be processed.

[0106] Table 2

[0107] 4 3 1 1 4 4 1 2 7 6 2 2 7 2 2 2

[0108] The target data shown in Table 2 is input into the first encoding module for DDPCM encoding, and the data points with a value of 2 will not participate in the difference operation of the DDPCM encoding.

[0109] Furthermore, the embodiment of the present invention may also include the following steps:

[0110] After the breakpoint detection is performed on the Tile block, the breakpoint detection related data of the Tile block is recorded; the breakpoint detection related data includes: the mapping mode bp_mode of the Tile block, the breakpoint data bp_msk of the Tile block, the repeated data bp_data, and the breakpoint valid signal bp_valid.

[0111] Among them, the mapping mode bp_mode of the Tile block is used to indicate the transformation base point of the Tile block, such as which corner of the Tile block is used for mapping or flipping. The breakpoint data bp_msk includes the pixel points corresponding to each breakpoint detected in the Tile block. The repeated data bp_data refers to the specific value of the repeated data contained in the detected Tile corner. The breakpoint valid signal bp_valid is used to indicate whether the preprocessing based on breakpoint detection (hereinafter referred to as breakpoint preprocessing) is successful. If the breakpoint preprocessing is successful, it means that the Tile block has undergone a transformation (such as mapping or flipping, etc.); if the breakpoint preprocessing is unsuccessful, it means that the Tile block has not undergone a transformation.

[0112] Taking Table 2 as an example, the breakpoint detection related data that can be recorded include: the mapping mode bp_mode is recorded as LD, indicating that it is flipped based on LD. The breakpoint data bp_msk is recorded as {000, 001, 010, 011}. The embodiment of the present invention uses a binary system to record the breakpoint data. Specifically, the number of rows of each breakpoint is recorded in binary from left to right and from bottom to top. The repeated data bp_data is recorded as 2. The breakpoint valid signal bp_valid is recorded as 1, indicating that the breakpoint preprocessing is successful and the original Tile block is transformed (the preset operation is performed).

[0113] The recorded breakpoint detection related data can be passed to the lower-level modules, such as the DDPCM encoding module, so that the lower-level modules can know whether the breakpoint preprocessing of the Tile block is successful according to the breakpoint valid signal, and know which corner the Tile block is flipped based on according to the mapping mode bp_mode. In addition, according to the breakpoint data bp_msk and the repeated data bp_data, the decompression module (decoding module) can restore the Tile block. Furthermore, if there is a primitive boundary in the Tile block, bp_valid will be pulled high (the value is 1); for the lower-level module, if the bp_valid signal is detected to be high, it can be determined that the related data with bp as the prefix is ​​valid, and the related data with bp as the prefix can be further read. If the primitive boundary is not detected in the Tile block, bp_valid is 0, and the lower-level module will not continue to read the related data with bp as the prefix, which can reduce the number of data readings and improve data processing efficiency.

[0114] Furthermore, an optional interval can be set for the control mechanism of bp_valid. For example, a scoring threshold can be set. If the largest score of the four corners of the Tile block does not exceed the scoring threshold, it is considered that there is no primitive boundary in the Tile block; if the largest score of the four corners exceeds the scoring threshold, it is considered that there is a primitive boundary in the Tile block, and the angle with the largest score is determined to be the Tile corner containing the most repeated data. Through the embodiment of the present invention, data outside the primitive boundary will be discarded (not involved in the difference calculation of DDPCM encoding), and will not be included in the consideration of the successful status of DDPCM encoding, thereby expanding the applicable occasions of DDPCM encoding, ensuring that the algorithm can still complete the encoding task in some special occasions, such as achieving successful compression of DDPCM encoding in the presence of primitive boundaries and solid color backgrounds, and the implementation cost is low.

[0115] In an optional embodiment of the present invention, the preprocessing operation of the raw data to be processed may include:

[0116] Step S31, detecting whether the Tile block has been subjected to multi-sampling anti-aliasing processing;

[0117] Step S32: If it is detected that the Tile block has been subjected to multi-sampling anti-aliasing processing, the repeated data in the Tile block are merged, and the vacant positions after the merging are filled with 0 to obtain the target data to be processed.

[0118] When the image processor performs drawing tasks, in order to achieve the purposes of edge softening, anti-aliasing, and anti-image folding loss, the image data will be processed with MSAA (multi-sampling anti-aliasing). After MSAA processing, the data will lose its original linear characteristics.

[0119] Table 3 shows an example of a 4×4 Tile block processed by MSAA. Each data in Table 3 represents the value of a pixel.

[0120] Table 3

[0121] A1B1C1FF A1B1C1FF A2B2C2FF A2B2C2FF A1B1C1FF A1B1C1FF A2B2C2FF A2B2C2FF A3B3C3FF A3B3C3FF A4B4C4FF A4B4C4FF A3B3C3FF A3B3C3FF A4B4C4FF A4B4C4FF

[0122] As shown in Table 3, after the MSAA processing, the values ​​of the four pixels in each 2×2 block in the Tile block are the same, which means that the Tile block has the MSAA property, that is, the Tile block has been processed by MSAA.

[0123] In a specific implementation, for a Tile block, the difference between adjacent rows and adjacent columns in the Tile block can be used to detect whether the Tile block has been processed by MSAA. For example, for a Tile block, subtract the first row from the second row, the third row from the fourth row, the first column from the second column, and the third column from the fourth column. If each result is 0, it is determined that the Tile block has been processed by MSAA; otherwise, it is determined that the Tile block has not been processed by MSAA. If it is determined that the Tile block has been processed by MSAA, the repeated data in each 2×2 block in the Tile block is merged, and the vacant positions are filled with 0 to obtain the target data to be processed.

[0124] By executing the process of step S32 on the Tile block shown in Table 3, the target data shown in Table 4 is obtained.

[0125] Table 4

[0126] A1B1C1FF A2B2C2FF 0 0 A3B3C3FF A4B4C4FF 0 0 0 0 0 0 0 0 0 0

[0127] In this way, the size of the Tile block can be reduced to one-fourth of its original size, resulting in a compression ratio of about 4:1, while also improving the success rate of differential encoding.

[0128] Furthermore, the step of detecting whether the Tile block has been subjected to multi-sampling anti-aliasing processing may be performed in a pre-processing stage, or may be performed in a first-order difference operation process of differentially encoding the Tile block.

[0129] For example, during the first-order difference operation of DDPCM encoding of a Tile block, it is possible to detect whether the Tile block has been subjected to multi-sampling anti-aliasing processing. Specifically, the first-order row difference operation can be extended from the original first two rows to all rows in the Tile block, and every two columns of the first-order row difference matrix are compared with every two rows of the first-order column difference matrix. If all are consistent, it is determined that the Tile block has been subjected to MSAA processing. Therefore, MSAA detection can be performed simultaneously during the first-order difference operation of DDPCM encoding of the Tile block, which can improve data compression efficiency.

[0130] Furthermore, the preprocessing based on MSAA detection and the preprocessing based on breakpoint detection are independent of each other. The embodiment of the present invention does not limit the execution order of the two preprocessings, and they can be processed simultaneously or sequentially. Both preprocessings can be implemented in the preprocessing module. Preferably, the preprocessing based on MSAA detection can be implemented inside the differential coding because the first-order difference generated in the differential coding operation process can be used for MSAA detection, which can reduce the operation overhead.

[0131] In the specific implementation, for a tile block processed by MSAA, each pixel will be sampled 4 times, which means that one pixel is expanded to 4 pixels. Therefore, in a tile block processed by MSAA, the values ​​of the 4 pixels in each 2×2 block must be the same, so each corner contains at least 2×2=4 repeated data (as shown in Table 3), and the score of each corner is at least 4. Therefore, the following two situations may occur when performing breakpoint detection on a tile block:

[0132] Case 1: The tile block has been processed by MSAA, and there is a primitive boundary in the tile block. In this case, the result of the MSAA detection processing does not affect the breakpoint detection.

[0133] Case 2: The Tile block is processed by MSAA, but there is no primitive boundary in the Tile block. At this time, the score of each corner is 4, which will cause bp_valid to be set to 1 and bp_mode to select LU (the scores of the four corners are the same, and LU will be selected by default). However, this situation is undesirable and will bring redundant data to the compression process. Therefore, the embodiment of the present invention can set the score threshold to be greater than 4 to avoid interference between the preprocessing of MSAA detection and the preprocessing of breakpoint detection, which affects the compression efficiency.

[0134] In an optional embodiment of the present invention, the method may further include:

[0135] Step S41: if it is detected that the Tile block has been processed by multi-sampling anti-aliasing, the multi-sampling flag of the Tile block is set to the first flag bit; if it is detected that the Tile block has not been processed by multi-sampling anti-aliasing, the multi-sampling flag of the Tile block is set to the second flag bit; the multi-sampling flag being the first flag bit indicates that the valid data in the Tile block is 1 / 4 of the Tile block; the multi-sampling flag being the second flag bit indicates that the valid data in the Tile block is the entire Tile block;

[0136] Step S42: input the multi-sampling identifier and the Tile block into the first encoding module, so that the first encoding module determines the valid data in the Tile block according to the multi-sampling identifier and performs differential encoding on the valid data.

[0137] In the embodiment of the present invention, a multi-sampling flag (such as msaa_valid) may be set for each Tile block to indicate whether the Tile block has been processed by MSAA. msaa_valid is the first flag bit (such as 1) indicating that the Tile block has been processed by MSAA and has the MSAA property, that is, the values ​​of the pixels in each 2×2 block in the Tile block are consistent; msaa_valid is the second flag bit (such as 0) indicating that the Tile block has not been processed by MSAA and does not have the MSAA property.

[0138] The multi-sampling identifier msaa_valid and the Tile block are input into the next level module (such as the first encoding module) together, so that the first encoding module can determine the valid data in the Tile block according to msaa_valid, and then perform differential encoding (such as DDPCM encoding) on ​​the valid data. For example, if msaa_valid is 0, it means that the Tile block does not have the MSAA property, and the valid data is the complete Tile block. If msaa_valid is 1, it means that the Tile block has the MSAA property, then after the Tile block is pre-processed by MSAA detection, duplicate data will be removed, and the width of the valid data is the original width of the Tile block divided by 2, and the length of the valid data is the original length of the Tile block divided by 2, that is, the valid data is 1 / 4 of the Tile block.

[0139] In an optional embodiment of the present invention, the entropy coding is Huffman coding, and the step of inputting the second data set into a third coding module for entropy coding to obtain a third coding result may include:

[0140] Step S51, determining the upper limit of the number of members of Huffman coding according to the target difference threshold;

[0141] Step S52: construct a symbol set including all possible members according to the upper limit of the number of members;

[0142] Step S53, constructing a Huffman tree based on the order of each member in the symbol set and the frequency of each member in the second data set, and recording the parent node information of each level in the process of constructing the Huffman tree;

[0143] Step S54: Combination-encode the recorded parent node information of each level, and remove the redundant items therein to obtain a third encoding result.

[0144] In the embodiment of the present invention, the entropy coding used by the third coding module can be Huffman coding. Further, the embodiment of the present invention optimizes Huffman coding by constructing a symbol set containing all possible members and recording the parent node information of each level of the Huffman tree, so that the symbol correspondence table and the codeword correspondence table that the Huffman tree originally needs to transmit are combined into one, and only the symbol-codeword correspondence table (i.e., the third coding result) needs to be transmitted, and the data width required by the symbol-codeword correspondence table is smaller, thereby improving the compression rate of Huffman coding.

[0145] The data compression method of the embodiment of the present invention is mainly aimed at compressing image data by an image processor, and the data input to the third encoding module (Huffman encoding module) is a member set with a fixed bit width, that is, the second data set. Taking the differential encoding as DDPCM encoding as an example, the fixed bit width is the target difference threshold, that is, the second-order difference threshold (such as 3 bits). For the second data set, since the bit width of a single member is fixed, the upper limit of the number of members that the second data set can contain is known in advance. For example, if the second-order difference threshold is 3 bits, the upper limit of the number of members that the second data set can contain is 2. 3 = 8. Based on this, the embodiment of the present invention proposes a Huffman coding optimization scheme for reducing the width of the symbol-codeword correspondence table. Specifically, all possible members are arranged in a fixed order, and the root node and each level of parent nodes are recorded in this order during the Huffman coding process. Finally, the recorded data are merged and encoded to obtain a symbol-codeword correspondence table with root node information.

[0146] Taking the second-order difference threshold of 3 bits as an example, the Huffman coding module can construct a symbol set containing all possible members according to the second-order difference threshold as follows: 000, 001, 010, 011, 100, 101, 110, 111. That is, when the second-order difference threshold is 3 bits, the members participating in Huffman coding can include the above 8 types at most. Therefore, when the Huffman coding module has not yet received the second data set transmitted by the upper module, it can be judged that the second data set transmitted by the upper module cannot include more than the above 8 types of members only through the second-order difference threshold uniformly configured by the pipeline.

[0147] After receiving the second data set, the Huffman coding module performs Huffman coding according to the order of each member in the constructed symbol set and the frequency of each member in the second data set, constructs a Huffman tree, and in the process of constructing the Huffman tree, records the parent node information of each level (such as member_msk) according to the upper limit of the number of members (such as 8). The parent node information of each level is arranged in a fixed order by default. When the members in the second data set complete the first sorting and select the two members with the smallest frequency, the positions of the two members in the order are recorded, that is, the corresponding positions of the two members with the smallest frequency in the member_msk of this level are set high (set to 1), so as to record the parent node information obtained after this merger. Because of the merger of members, the member_msk of the next level will lose an effective width of one bit, and this is repeated until the construction of the Huffman tree is completed.

[0148] Reference Figure 7 , shows a flow chart of Huffman coding in an example of the present invention. Assuming that the second-order difference threshold is 3 bits, the upper limit of the number of members is 8, and the recording process of each level member_msk is as follows Figure 7 As shown. First, arrange the 8 possible members in order, such as: 000, 001, 010, 011, 100, 101, 110, 111. For the convenience of description, the 8 members are named A to H in sequence, that is, A to H represent a member, such as A represents member 000, B represents member 001, C represents member 010, D represents member 011, E represents member 100, F represents member 101, G represents member 110, and H represents member 111. The process of constructing the Huffman tree is as follows Figure 7 As shown by the arrows in , the parent node information of each level is recorded through a set of member_msk signals.

[0149] Figure 7The first row in represents the order in which all members contained in the constructed symbol set are arranged (the first row is the root node). Since the symbol set contains all possible members, such as A to H, in practice, the second data set input to the Huffman coding module may not contain all members. For example, members B, G, and H do not appear in the second data set. In this case, the symbol set is said to contain invalid members B, G, and H. Invalid members in the symbol set do not participate in the construction of the Huffman tree. Figure 7 As shown, the valid under members B, G, and H indicates that the member is invalid and does not participate in the construction of the Huffman tree.

[0150] In an embodiment of the present invention, a Huffman tree is constructed based on the order of each member in the symbol set and the frequency of occurrence of each member in the second data set, and the parent node information of each level is recorded in the process of constructing the Huffman tree. For example, starting from the root node, it is counted that the frequency of occurrence of members C and D in the second data set is the smallest, then the positions corresponding to members C and D in the parent node information (member_msk_0) of this level are set high, and therefore, the parent node information of this level is recorded as member_msk_0=00110000. The two 1s therein represent the positions corresponding to members C and D, respectively. Similarly, the parent node information of each level is recorded in the process of constructing the Huffman tree.

[0151] After the Huffman tree is constructed, the parent node information of each level is obtained and combined and encoded, such as discarding the redundant items in the parent node information of each level by shifting right (or left) to obtain the Huffman encoding result (i.e., the symbol-codeword correspondence table), such as Figure 8 As shown, it is a schematic diagram of combining and encoding the parent node information of each level to obtain a symbol-codeword correspondence table. Among them, discarding redundant items refers to removing the data 0 in member_msk from left to right or from right to left until the first 1 is encountered. For example, taking from right to left as an example, member_msk_0=00110000, 4 0s can be removed from right to left to get 0011, and the redundant item is 0000. member_msk_1=1010000, 4 0s can be removed from right to left to get 101, and the redundant item is 0000. And so on. In this way, the redundant 0s in each level of member_msk can be removed to obtain a symbol-codeword correspondence table. In a specific implementation, whether to remove redundant items from left to right or from right to left can be determined according to the hardware configuration. The redundant 0s in member_msk are generated due to the existence of invalid members in the symbol set.

[0152] In the related art, Huffman coding requires the transmission of a symbol correspondence table. If the shortest length scheme is considered, the variable-length codes generated by the Huffman tree need to be tightly spliced, which requires additional computing resources. Figure 7Taking the data shown as an example, traditional Huffman coding needs to transmit the following symbol correspondence table: A000 C0010D0011 E01 F1, the format is: member + code; each member is 3 bits, totaling 29 bits.

[0153] The embodiment of the present invention optimizes Huffman coding, and the symbol codeword correspondence table (such as Figure 8 The length of the Huffman coding is only 16 bits (the full signal does not need to be encoded). Compared with the traditional Huffman coding, the data width required for the transmitted symbol-codeword correspondence table is smaller, thereby improving the compression rate of the Huffman coding.

[0154] In an optional embodiment of the present invention, the method may further include:

[0155] Step S61: setting a redundant item flag according to whether there is an invalid member in the symbol set; the redundant item flag is used to indicate whether the redundant item in the last-level parent node information is allowed to be discarded;

[0156] Step S62: discard or retain the redundant items in the last-level parent node information according to the redundant item identifier, and transfer the redundant item identifier to the next-level module.

[0157] The embodiment of the present invention optimizes Huffman coding. Specifically, by constructing a symbol set containing all possible members and recording the parent node information of each level of the Huffman tree, only a symbol-codeword correspondence table with a smaller required width needs to be transmitted.

[0158] Furthermore, since the constructed symbol set may contain invalid members that do not appear in the second data set, and in the Huffman encoding process, the redundant items in the parent node information generated by the invalid members will be discarded, therefore, when there is invalid data in the symbol set and the redundant items in the last-level parent node information are discarded, it is difficult for the decoding module to determine whether the member information is lost based on the symbol-codeword correspondence table, and it is impossible to know who the lost member is, which will make it difficult to correctly restore the Huffman tree due to incomplete information.

[0159] To avoid this problem, an embodiment of the present invention sets a redundant item identifier according to whether there are invalid members in the symbol set; the redundant item identifier is used to indicate whether the redundant items in the last-level parent node information are allowed to be discarded. For example, if there are no invalid members in the symbol set, the redundant item identifier is set to the first flag bit. Since there are no invalid members in the symbol set, the symbol-codeword correspondence table finally obtained itself contains the correct encoding information, and the Huffman tree can be correctly restored according to the symbol-codeword correspondence table. Therefore, if the redundant item identifier is the first flag bit, the redundant items in the last-level parent node information can be discarded. If there are invalid members in the symbol set, the redundant item identifier is set to the second flag bit to indicate that the redundant items in the last-level parent node information are not allowed to be discarded, so that the auxiliary decoding module can correctly restore the Huffman tree by retaining the redundant items in the last-level parent node information.

[0160] The embodiment of the present invention adds a full signal (redundancy item identifier) ​​to indicate whether there are invalid members in this Huffman coding. If full is 0, it means that the last level member_msk does not allow the discarding of redundant items. Figure 8 As shown, member_msk_3=10100, member_msk_3 (the last level parent node information) retains the redundant item 00. Since the decoding module decoding is the opposite of Huffman coding, the first level parent node information read by the decoding module is the last level parent node information output by the coding module. Therefore, the first level parent node information read by the decoding module is 10100, indicating that there are 3 invalid members (3 zeros), so the Huffman tree can be accurately restored.

[0161] In summary, the data compression method provided in the embodiment of the present invention is a general lossless compression solution for image data suitable for an image processor, which can process data compression in a frame buffer and a depth buffer at the same time, solves the problem of area overhead caused by using different compression algorithms and processing structures in the frame buffer and the depth buffer, and can save hardware area.

[0162] In addition, the data compression method provided in the embodiment of the present invention optimizes differential coding. A preprocessing operation based on breakpoint detection is performed before differential coding, and the compression environment of the differential coding algorithm is improved by adding deviation point data in differential coding, and the tolerance for non-compliant pixel points (deviation points) is increased, and the coordinate data format of the additionally stored deviation points also maintains the characteristic of being able to be compressed twice. Therefore, compared with traditional differential coding, the optimization strategy of the optimized differential coding scheme of the present invention has low cost and high efficiency, improves the adaptability of the differential coding algorithm to the graphics processor, and improves the success rate of differential coding.

[0163] Furthermore, the data compression method provided by the embodiment of the present invention optimizes Huffman coding. By constructing a symbol set containing all possible members and recording the information of each parent node of the Huffman tree, the symbol correspondence table and the codeword correspondence table that the Huffman tree originally needs to transmit are combined into one, and only the symbol-codeword correspondence table (i.e., the third encoding result) needs to be transmitted. The data width required by the symbol-codeword correspondence table is smaller, thereby improving the compression rate of Huffman coding.

[0164] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0165] Reference Fig. 9 , shows a structural block diagram of an embodiment of a data compression device of the present invention, the device may include:

[0166] A preprocessing module 201 is used to perform a preprocessing operation on the raw data to be processed to obtain target data to be processed; the raw data to be processed is a Tile block in a frame buffer or a depth buffer;

[0167] The first encoding module 202 is used to perform differential encoding on the target data to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold;

[0168] A second encoding module 203, configured to perform universal encoding on the first data set to obtain a second encoding result;

[0169] A third encoding module 204, configured to perform entropy encoding on the second data set to obtain a third encoding result;

[0170] The merging output module 205 is used to concatenate the second encoding result and the third encoding result to obtain compressed data of the original data.

[0171] Optionally, the preprocessing module includes a first preprocessing submodule, including:

[0172] A breakpoint detection unit, used to perform breakpoint detection on the Tile block to determine whether there is a primitive boundary in the Tile block;

[0173] A Tile corner determination unit, configured to determine a Tile corner containing the most repeated data in the Tile block according to a primitive boundary if there is one in the Tile block;

[0174] A preset operation unit is used to perform a preset operation on the Tile block so that the Tile corner is located at a target position of the Tile block to obtain target data to be processed; the target position is determined according to the order of difference operations in the differential encoding process, and the target position is a position in the Tile block that does not participate in the difference operation; the preset operation includes a pixel mapping operation or a Tile block flipping operation.

[0175] Optionally, the breakpoint detection unit includes:

[0176] A duplicate detection subunit, used to use the four corners of the Tile block as starting points and use a preset algorithm to perform duplicate data detection respectively;

[0177] A scoring subunit, used to score each corner according to the number of repeated data contained in each detected corner;

[0178] The boundary determination subunit is used for determining that there is a primitive boundary in the Tile block if the maximum score among the four corners of the Tile block exceeds a preset score threshold; otherwise, it is determined that there is no primitive boundary in the Tile block.

[0179] Optionally, the device further comprises:

[0180] A data recording and transmission module is used to record the breakpoint detection related data of the Tile block after performing breakpoint detection on the Tile block; the breakpoint detection related data includes: the mapping mode of the Tile block, the breakpoint data of the Tile block, repeated data, and the breakpoint valid signal; and transmit the breakpoint detection related data to the lower-level module.

[0181] Optionally, the preprocessing module includes a second preprocessing submodule, including:

[0182] A multi-sampling detection unit, used to detect whether the Tile block has been processed by multi-sampling anti-aliasing;

[0183] The deduplication processing unit is used to merge the duplicate data in the Tile block if it is detected that the Tile block has been processed by multi-sampling anti-aliasing, and fill the vacant positions after the merging with 0 to obtain the target data to be processed.

[0184] Optionally, the device further comprises:

[0185] A multi-sampling flag setting module is used to set the multi-sampling flag of the Tile block to a first flag bit if it is detected that the Tile block has been processed by multi-sampling anti-aliasing; the multi-sampling flag being the first flag bit indicates that the valid data in the Tile block is 1 / 4 of the Tile block; if it is detected that the Tile block has not been processed by multi-sampling anti-aliasing, set the multi-sampling flag of the Tile block to a second flag bit; the multi-sampling flag being the second flag bit indicates that the valid data in the Tile block is the entire Tile block;

[0186] The multi-sampling identifier processing module inputs the multi-sampling identifier and the Tile block into the first encoding module, so that the first encoding module determines the valid data in the Tile block according to the multi-sampling identifier and performs differential encoding on the valid data.

[0187] Optionally, the entropy coding is Huffman coding, and the third coding module includes:

[0188] An upper limit determination submodule, used for determining an upper limit of the number of members of Huffman coding according to the target difference threshold;

[0189] A symbol set construction submodule, used to construct a symbol set containing all possible members according to the upper limit of the number of members;

[0190] An information recording submodule, used for constructing a Huffman tree based on the order of each member in the symbol set and the frequency of each member in the second data set, and recording the parent node information of each level in the process of constructing the Huffman tree;

[0191] The combined coding submodule is used to perform combined coding on the recorded parent node information of each level and remove the redundant items therein to obtain a third coding result.

[0192] Optionally, the device further comprises:

[0193] A redundant item identification setting module, used to set a redundant item identification according to whether there is an invalid member in the symbol set; the redundant item identification is used to indicate whether the redundant item in the last-level parent node information is allowed to be discarded;

[0194] The redundant item identifier processing module is used to discard or retain the redundant items in the last level parent node information according to the redundant item identifier, and transmit the redundant item identifier to the next level module.

[0195] Optionally, the data width of the first data set is the width of original pixels of the target data, and the data width of the second data set is the target difference threshold of the differential encoding.

[0196] Optionally, the device further comprises:

[0197] The failure trigger module is used to trigger the failure of the DDPCM encoding if the number of deviation points in the Tile block exceeds a preset number.

[0198] In summary, the data compression device provided in the embodiment of the present invention adopts a general lossless compression scheme suitable for image data of an image processor, and can process data compression in the frame buffer and the depth buffer at the same time, thereby solving the problem of area overhead caused by different compression algorithms and processing structures used in the frame buffer and the depth buffer, and can save hardware area.

[0199] In addition, the data compression device provided by the embodiment of the present invention optimizes differential coding. A preprocessing operation based on breakpoint detection is performed before differential coding, and the compression environment of the differential coding algorithm is improved by adding deviation point data in differential coding, and the tolerance for non-compliant pixel points (deviation points) is increased, and the coordinate data format of the additionally stored deviation points also maintains the characteristic of being able to be compressed twice. Therefore, compared with traditional differential coding, the optimization strategy of the optimized differential coding scheme of the present invention has low cost and high efficiency, improves the adaptability of the differential coding algorithm to the graphics processor, and improves the success rate of differential coding.

[0200] Furthermore, the data compression device provided by the embodiment of the present invention optimizes Huffman coding. By constructing a symbol set containing all possible members and recording the information of each parent node of the Huffman tree, the symbol correspondence table and the codeword correspondence table that the Huffman tree originally needs to transmit are combined into one, and only the symbol-codeword correspondence table (i.e., the third encoding result) needs to be transmitted. The data width required by the symbol-codeword correspondence table is smaller, thereby improving the compression rate of Huffman coding.

[0201] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0202] Reference Fig.10 , is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Fig.10 As shown, the electronic device includes: a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the data compression method of the aforementioned embodiment.

[0203] An embodiment of the present invention provides a non-transitory computer-readable storage medium. When instructions in the storage medium are executed by a program or a processor of a terminal, the terminal is enabled to execute the steps of the data compression method of the aforementioned embodiment.

[0204] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0205] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0206] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0207] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0208] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0209] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0210] Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A data compression method, characterized in that: The method comprises: Performing a preprocessing operation on the raw data to be processed to obtain target data to be processed; the raw data to be processed is a Tile block in the frame buffer or the depth buffer; The target data is input into a first encoding module for differential encoding to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold; Inputting the first data set into a second encoding module for universal encoding to obtain a second encoding result; and inputting the second data set into a third encoding module for entropy encoding to obtain a third encoding result; The second encoding result and the third encoding result are concatenated to obtain compressed data of the original data.

2. The method according to claim 1, characterized in that The preprocessing operation of the raw data to be processed includes: Performing breakpoint detection on the Tile block to determine whether there is a primitive boundary in the Tile block; If there is a primitive boundary in the Tile block, determining the Tile corner containing the most repeated data in the Tile block according to the primitive boundary; A preset operation is performed on the Tile block so that the Tile corner is located at a target position of the Tile block to obtain target data to be processed; the target position is determined according to the order of difference operations in the differential encoding process, and the target position is a position in the Tile block that does not participate in the difference operation; the preset operation includes a pixel mapping operation or a Tile block flipping operation.

3. The method according to claim 2, characterized in that The performing breakpoint detection on the Tile block to determine whether there is a primitive boundary in the Tile block includes: Taking the four corners of the Tile block as starting points, duplicate data detection is performed using a preset algorithm respectively; Each corner is scored according to the number of repeated data contained in each detected corner; If the maximum score of the four corners of the Tile block exceeds a preset score threshold, it is determined that there is a primitive boundary in the Tile block; otherwise, it is determined that there is no primitive boundary in the Tile block.

4. The method according to claim 2, characterized in that: The method further comprises: After performing breakpoint detection on the Tile block, record breakpoint detection related data of the Tile block; the breakpoint detection related data includes: mapping mode of the Tile block, breakpoint data of the Tile block, repeated data, and breakpoint valid signal; The breakpoint detection related data is transferred to the lower-level module.

5. The method according to claim 1, characterized in that The preprocessing operation of the raw data to be processed includes: Detect whether the Tile block has been processed by multi-sampling anti-aliasing; If it is detected that the Tile block has been subjected to multi-sampling anti-aliasing processing, the repeated data in the Tile block are merged, and the vacant positions after the merging are filled with 0 to obtain the target data to be processed.

6. The method according to claim 5, characterized in that The method further comprises: If it is detected that the Tile block has been subjected to multi-sampling anti-aliasing processing, the multi-sampling flag of the Tile block is set to the first flag bit; the multi-sampling flag being the first flag bit indicates that the valid data in the Tile block is 1 / 4 of the Tile block; If it is detected that the Tile block has not been subjected to multi-sampling anti-aliasing processing, the multi-sampling flag of the Tile block is set to the second flag bit; the multi-sampling flag being the second flag bit indicates that the valid data in the Tile block is the entire Tile block; The multiple sampling identifier and the Tile block are input into the first encoding module together, so that the first encoding module determines the valid data in the Tile block according to the multiple sampling identifier and performs differential encoding on the valid data.

7. The method according to claim 1, characterized in that The entropy coding is Huffman coding, and the inputting the second data set into a third coding module for entropy coding to obtain a third coding result includes: Determining an upper limit on the number of members of the Huffman code according to the target difference threshold; According to the upper limit of the number of members, a symbol set including all possible members is constructed; Constructing a Huffman tree based on the order of each member in the symbol set and the frequency of each member in the second data set, and recording the parent node information of each level in the process of constructing the Huffman tree; The recorded parent node information of each level is combined and encoded, and the redundant items therein are removed to obtain a third encoding result.

8. The method according to claim 7, characterized in that The method further comprises: According to whether there is an invalid member in the symbol set, a redundant item identifier is set; the redundant item identifier is used to indicate whether the redundant item in the last-level parent node information is allowed to be discarded; According to the redundant item identifier, the redundant item in the last level parent node information is discarded or retained, and the redundant item identifier is transmitted to the next level module.

9. The method according to any one of claims 1 to 8, characterized in that: The data width of the first data set is the width of original pixels of the target data, and the data width of the second data set is the target difference threshold of the differential encoding.

10. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: If the number of deviation points in the Tile block exceeds a preset number, the differential encoding failure is triggered.

11. A data compression device, characterized in that: The device comprises: A preprocessing module, used for performing preprocessing operations on the raw data to be processed to obtain target data to be processed; the raw data to be processed is a Tile block in the frame buffer or the depth buffer; A first encoding module is used to input the target data into the first encoding module for differential encoding to obtain a first encoding result; the first encoding result includes a first data set and a second data set; the first data set includes incomplete difference data and deviation points obtained after differential encoding; the second data set includes complete difference data obtained after differential encoding and coordinate data of the deviation points; the deviation points are original pixel points where the complete difference data exceeds the target difference threshold; A second encoding module, used for performing universal encoding on the first data set to obtain a second encoding result; A third encoding module, used for performing entropy encoding on the second data set to obtain a third encoding result; The merging output module is used to concatenate the second encoding result and the third encoding result to obtain compressed data of the original data.

12. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute the steps of the data compression method according to any one of claims 1 to 10.

13. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the data compression method according to any one of claims 1 to 10 are implemented.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data compression method according to any one of claims 1 to 10 are implemented.