Lossless compression and decompression method and system based on probability model
Through the lossless compression method based on the probability model, using probability inference and asymmetric digital system entropy coding, the problems of compression loss and slow speed in the existing lossless compression algorithm are solved, and efficient lossless compression and high-speed decompression are achieved.
Patent Information
- Application Number
- CN202510333077.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
Existing lossless compression algorithms have problems such as compression loss, slow speed and low efficiency, especially the compression algorithms based on neural networks do not perform well in compression speed.
The lossless compression method based on the probability model is adopted, and the input data is divided into multiple input sequences, the probability model is used for probability reasoning, and the entropy encoding of the asymmetric digital system is used for compression and decompression, and the data processing is performed using probability circuits, neural networks or traditional probability distribution models.
It realizes lossless, high-speed, and high compression ratio compression effects, and is suitable for different types of input data, such as pictures, audio and video, and further improves the compression speed through the selection or design of hardware modules.
Smart Images

Figure CN120263194A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the design of lossless compression in compression algorithms, and particularly to the design of lossless compression and decompression related to probability model algorithms. Background Art
[0002] Compression algorithms are techniques used to reduce data storage space or transmission time by converting the original data into a more compact form through specific mathematical or coding methods. As Figure 1 shown, according to whether the data can be completely restored to its original state after decompression, compression algorithms can be divided into two categories: lossy compression and lossless compression. Lossy compression algorithms, such as JPEG and MP3, achieve high compression ratios by removing information that is insensitive to human perception or redundant in the data. This method is particularly effective in compressing multimedia data such as images, audio, and video, and can significantly reduce the storage space while ensuring a certain quality. However, the disadvantage is that the decompressed data cannot be completely restored to its original state, resulting in some information loss.
[0003] Lossless compression algorithms can compress data without losing any original data information. This compression method relies on redundant information in the data, such as repeated characters, patterns, or predictable numerical sequences. By identifying and eliminating this redundancy, lossless compression algorithms can significantly reduce the size of the data while ensuring that the decompressed data is exactly the same as the original data. Common lossless compression algorithms include Huffman Coding, RLE (Run-Length Encoding), and Lempel-Ziv-Welch (LZW) algorithms, etc. Lossless compression is widely used in scenarios where data integrity needs to be maintained, such as text files, source code, database backups, and certain types of image files (such as PNG format). However, the compression ratio of lossless compression is usually limited by the theoretical redundancy of the data statistics and it is difficult to achieve a high compression ratio like lossy compression.
[0004] Lossless compression algorithms based on neural networks combine the powerful feature extraction and expression capabilities of neural networks to achieve a higher level of compression effect. By learning complex patterns and structures in the data, this algorithm can more precisely identify and eliminate redundant information while maintaining data integrity. Lossless compression algorithms based on neural networks not only have the advantages of lossless compression, that is, the decompressed data is exactly the same as the original data, but also can achieve a higher compression ratio than conventional lossless compression in some cases. In addition, the neural network can automatically adjust the compression parameters according to different scenarios and contents, so as to obtain a better compression effect in different situations.
[0005] Ordinary compression algorithms are fast and have a high compression ratio, but there is compression loss; conventional lossless compression algorithms avoid compression loss, but sacrifice the compression ratio; existing neural network-based compression algorithms improve the compression ratio, but have high computational requirements and very slow compression speed. Summary of the Invention
[0006] In view of the problems of compression loss, slow compression speed, and low efficiency in the above-mentioned existing technologies, the present invention proposes an efficient lossless compression and decompression method and system based on a probabilistic model. For the compression of different types of data, through the technique of segmentation, probability inference is performed using a probabilistic model, and the asymmetric numeral system entropy coding compression / decompression algorithm is used for complete lossless compression / decompression.
[0007] The technical solution of the present invention is as follows:
[0008] A lossless compression and decompression method based on a probabilistic model, characterized by comprising the following steps:
[0009] (1) Divide the input compressed data into multiple input sequences, and each input sequence includes multiple basic units;
[0010] (2) Based on the trained probabilistic model, perform probability inference on each basic unit in each input sequence in turn to generate corresponding conditional probabilities, forming a probability sequence;
[0011] (3) Perform compression encoding operations of asymmetric numeral system entropy coding on all conditional probabilities in the probability sequence generated by the probability inference operation in step (2), and convert the compressed data into a binary coding format for storage;
[0012] (4) When decoding the compressed bitstream, perform decompression operations of asymmetric numeral system entropy coding to obtain conditional probabilities;
[0013] (5) According to the decompressed conditional probabilities, use the dichotomy method to reconstruct all basic units of the input sequence in turn, and calculate the original values of all input sequences.
[0014] Further, the probabilistic model is a probabilistic circuit, a neural network, or a traditional probability distribution model.
[0015] Further, in step (2), when performing probability inference on each basic unit in the input sequence, the generated corresponding conditional probability is the conditional probability of each basic unit based on all previous units in the input sequence:
[0016] Conditional probability of basic unit x0: P(x0 = X0), P(x0 < X0)
[0017] Conditional probabilities of the basic unit x1: (P(x0 = X0, x1 = X1), P(x0 = X0, |x1 < X1)) ....
[0019] The basic unit x i 's conditional probabilities: P(x0 = X0, …, x i = X i ), P(x0 = X0, …|x i < X i )
[0020] Furthermore, in step (3), a compression operation of asymmetric numeral system entropy coding is performed on all conditional probabilities in the probability sequence. Specifically, when using asymmetric numeral system entropy coding for compression, a normalization precision p and n need to be selected, generally not exceeding 10. Then, the operations are carried out according to the following steps:
[0021] (3.1) Normalize the conditional probabilities and convert them into integer frequencies f(x i ), C(x i ):
[0022] f(x i ) = P(x i = X i |x {i-1} = X {i-1} ) × 2 p
[0023] C(x i ) = P(x i < X i |x {i-1} = X {i-1} ) × 2 p
[0024] (3.2) Convert the frequencies and states into two state variables a1, a2, and the complete state s(x {i} ):
[0025]
[0026] a2 = s(x {i-1} ) mod (f(x i )) + C(x i )
[0027] s(x {i} ) = a1 + a2
[0028] Store the finally obtained a1 and a2 in binary form, and the compression process of a complete input sequence is completed. After completing the compression of one input sequence, repeat the above steps until the compression of all input sequences is completed.
[0029] Further, step (5) is specifically as follows:
[0030] For an input sequence, for the basic unit x0 of the input sequence, according to the conditional probability of this basic unit, the specific data of x0 is reconstructed by the dichotomy method: there is a restricted interval 0 < x0 ≤ X for the basic unit x0 lim_0 , determine the initial hypothesis value x0 = X lim_0 / 2, calculate the corresponding probability, and judge whether the interval of x0 is 0 < x0 ≤ X lim_0 / 2 or X lim_0 / 2 < x0 ≤ X lim_0 , by iterating this process, adjust the hypothesis value x0 until the specific data of x0 is calculated; for subsequent inputs of the probability sequence, repeat this process until the original value of a single input sequence is calculated; for the remaining input sequences, repeat this process until the original values of all input sequences are calculated.
[0031] The present invention also provides a system for implementing the above-mentioned lossless compression and decompression method based on a probability model, which is characterized in that it includes a software module and a hardware module. The software module includes:
[0032] A probability model training part, which is used to perform probability distribution fitting according to the target data set and construct a probability model describing its distribution frequency;
[0033] A probability inference part, which is used to perform probability inference on each basic unit of the input sequence by using the trained probability model;
[0034] An asymmetric digital system entropy coding compression / decompression part, which is used to perform compression coding on the probability sequence generated by the probability inference and decode the compression coding;
[0035] The hardware module includes:
[0036] A configurable probability model operation module, which is used to deploy the trained probability model and execute the operations of the probability inference part;
[0037] A compression / decompression module, which is used to deploy the asymmetric digital system entropy coding compression / decompression process and execute compression coding / decoding operations;
[0038] Among them, the probability model operation module dynamically selects an operation platform according to the type of the probability model, including FPGA, GPU, CPU or a custom chip ASIC.
[0039] Further, when the probability model training is constructed using a probabilistic circuit, the probability model operation module is implemented by an FPGA or a custom chip ASIC; when the probability model training is constructed using a neural network, the probability model operation module is implemented by a GPU; when the probability model training is constructed using a simple distribution, the probability model operation module is implemented by a CPU.
[0040] Further, the compression / decompression module is implemented by a CPU and supports the use of Verilog to construct a specific encoding module to implement random asymmetric digital system entropy encoding.
[0041] Further, the software module is implemented using software programming codes Julia, C++ and Python.
[0042] The technical effects of the present invention are as follows:
[0043] The present invention provides an efficient lossless compression and decompression method and system based on a probability model, which can perform lossless compression with high speed and high compression ratio. By segmenting and parameterizing the input sequence, it can compress input sequences of different sizes and types, such as pictures, audio, video, etc.; through the combined functions of probability model inference and asymmetric digital system entropy encoding, it realizes lossless compression with a high compression ratio; through the independent selection or design of the hardware module, it realizes high-speed lossless compression. Description of the Drawings
[0044] Figure 1 Shows the difference between lossy compression and lossless compression;
[0045] Figure 2 Is a schematic diagram of the working process of the lossless compression and decompression method based on the probability model of the present invention;
[0046] Figure 3 Is a schematic diagram of the working process of the lossless compression and decompression method based on the probability model of the present invention on the hardware module;
[0047] Figure 4 Is a schematic diagram of the "image compression" task in the embodiment of the present invention. Detailed Embodiments
[0048] The following further clearly and completely elaborates the present invention with reference to the drawings through specific embodiments.
[0049] An efficient lossless compression and decompression method based on a probability model of the present invention Figure 2 For the schematic diagram of the working process of this method, it includes the following steps:
[0050] (1) Divide the input compressed data into multiple input sequences, and each input sequence includes multiple basic units;
[0051] (2) Based on the trained probability model, perform probability inference on each basic unit in each input sequence successively to generate corresponding conditional probabilities, forming a probability sequence;
[0052] (3) Perform lossless compression encoding operation of asymmetric numeral system entropy encoding on all conditional probabilities in the probability sequence generated by the probability inference operation in step (2), and convert the compressed data into binary encoding format for storage;
[0053] (4) When decoding the compressed bitstream, perform decompression operation of asymmetric numeral system entropy encoding to obtain conditional probabilities;
[0054] (5) According to the decompressed conditional probabilities, reconstruct all basic units of the input sequence successively by using the dichotomy method, and calculate the original values of all input sequences.
[0055] Implement the system design of the lossless compression and decompression method based on the probability model of the present invention. The system includes a software module and a hardware module. The software module includes three parts: probability model training, probability inference, and asymmetric numeral system entropy encoding compression / decompression. The hardware module includes two parts: a probability model operation module and a compression / decompression module.
[0056] The software module includes three parts:
[0057] The probability model training part uses the target compression data training set to fit the probability distribution. After the user selects the target compression task, the probability model used in the present invention is trained on the data set corresponding to the task, and a probability model describing its distribution frequency is constructed on the original data. The training process can use a traditional simple probability model to describe the probability distribution, or use a new probability model, such as a probability circuit, etc. to describe it.
[0058] For different compression tasks, the specific compression examples are segmented into smaller input sequences to facilitate the construction of the probability model. The probability inference part uses the trained probability model to perform probability inference on the content to be compressed specifically to obtain the probabilities corresponding to each input sequence (x0 = X0, x1 = X1, …, x n = X n ) of the target compression task, where x i is the basic unit of the input sequence, and X0 is the specific value of the basic unit. After the probability model training is completed, the hardware module uses the trained probability model to perform probability inference on each basic unit x i of the input sequence successively, and obtains the conditional probability of each basic unit based on all the previous units of the input sequence (P(x0 = X0, …, x i = X i ), P(x0 = X0, …|x i <Xi ). The conditional probability of the operation completion forms a probability sequence, which is fed into the subsequent part for operation as the compressed data of the asymmetric digital system entropy coding.
[0059] The compression / decompression part of the asymmetric digital system entropy coding compresses / decompresses the probability sequence generated by the probability inference using the asymmetric digital system entropy coding. The asymmetric digital system entropy coding is a predictive coding algorithm based on the statistical probability of data. It treats the entire input sequence as a symbol stream and encodes according to the probability of each symbol occurrence. The basic principle of the asymmetric digital system entropy coding is to map the input symbol to an interval, each interval representing a probability range, and then scale and update the interval according to the probability of the input symbol. Finally, the encoder maps the input sequence to the final coding interval. The compression part of the asymmetric digital system entropy coding converts the probability generated by the probability inference into a binary coding format through such a mapping process, realizing a highly efficient compression process.
[0060] The decompression process reverses the above compression process, decompresses the encoding back into the corresponding probability. Then, based on the corresponding probability, the input sequence is reconstructed using the bisection method. Specifically, the process of the bisection method is as follows. If there is a restricted interval 0 < x0 ≤ X for the basic unit x0 of the input sequence lim_0 , the decompression process first assumes x0 = X lim_0 / 2, calculates the corresponding probability in this case, and judges whether the interval of x0 is 0 < x0 ≤ X lim_0 / 2 or X lim_0 / 2 < x0 ≤ X lim_0 . And so on, until the specific value of x0 is calculated. For the subsequent input of the probability sequence, repeat this process until the original value of a single input sequence is calculated. After that, the decompression process repeats the above process for the remaining data until the original values of all input sequences are obtained.
[0061] The hardware module includes two modules:
[0062] The probability model operation module performs efficient probability inference. After the probability model training is completed, it is deployed on the probability model operation module for high-speed operation. Depending on the probability model used, the operation platform used by the probability model operation module can be different. When using a simple distribution to construct the probability model, the CPU can be used as the probability model operation module; when using a neural network to construct the probability model, the GPU can be used as the probability model operation module. In particular, when using a probability circuit to construct the probability model, a targeted circuit design, such as a hardware programmable array FPGA or a custom chip design ASIC, can be used for probability inference.
[0063] The compression / decompression module performs the asymmetric numeral system entropy encoding process of the probability sequence and the dichotomy reconstruction process of the decompression process. Specifically, for most asymmetric numeral system entropy encoding methods, the CPU can be used as the compression / decompression module. In particular, when using the random asymmetric numeral system entropy encoding method for compression / decompression, a specific encoding module can be constructed using Verilog code.
[0064] Figure 3 This is a schematic diagram of the working process of the lossless compression and decompression method based on the probability model of the present invention on the hardware module; below, an efficient lossless compression method implemented on the system of the present invention will be introduced with "image compression" as an example. In this embodiment, the software module is implemented using software programming codes Julia, C++, and Python, and the hardware module is implemented using the CPU and the hardware programmable array FPGA, and can also be implemented using the GPU.
[0065] "Image compression" is a classic probability inference problem, and its tasks are as Figure 4 shown. In a 32x32x3 RGB image, each color channel of each pixel is regarded as a basic unit of the input sequence. In the RGB representation, each channel is an integer from 0 to 255. To reduce the model overhead, it is split into 16x16x3 sequences for separate compression, and each input sequence has 768 basic units. Here, a probability circuit is used to construct the probability model. After the probability model training is completed, the probability model is deployed to the probability model operation module built by the FPGA or GPU on the hardware module for efficient probability inference, and the asymmetric numeral system entropy encoding process is deployed to the hardware module, and the compression / decompression module built by the CPU to complete the complete compression or decompression process. The compression process goes through the following steps:
[0066] 1) Sequentially perform probability inference on all basic units x 767 =X 767 ) of the entire input sequence (x0 = X0, x1 = X1,..., x i ). Based on the trained probability model, through the calculation of the probability model operation module, the conditional probabilities P(x0 = X0), P(x0 < X0), P(x0 = X0, x1 = X1), P(x0 = X0|x1 < X1),... (P(x0 = X0,..., x 767 =X 767 ), P(x0 = X0,...|x 767 <X 767 )) are successively obtained through probability inference.
[0067] 2) Perform asymmetric numeral system entropy encoding on all the generated conditional probabilities. During the compression process, the conditional probabilities P(x0 = X0), P(x0 < X0), P(x0 = X0, x1 = X1), P(x0 = X0|x1 < X1), … (P(x0 = X0, …, x 767 = X 767 ), P(x0 = X0, …|x 767 < X 767 )) obtained based on the probability model operation module are input into the compression / decompression module for asymmetric numeral system entropy encoding compression operations.
[0068] Specifically, when using asymmetric numeral system entropy encoding in the asymmetric numeral system entropy encoding method for compression, a normalized precision p and n are selected. The selection process is determined according to the probability magnitude. Here, p = n = 8 is selected. Then, the operations are carried out according to the following steps:
[0069] Normalize the conditional probabilities and convert them into integer frequencies f(x i ), C(x i ):
[0070] f(x i ) = P(x i = X i |x {i-1} = X {i-1} ) × 2 p
[0071] C(x i ) = P(x i < X i |x {i-1} = X {i-1} ) × 2 p
[0072] Convert the frequency and state into two state variables a1, a2, and the complete state s(x {i} ):
[0073]
[0074] a2 = s(x {i-1} ) mod (f(x i )) + C(x i )
[0075] s(x {i} ) = a1 + a2
[0076] Store the finally obtained a1 and a2 in binary form, that is, complete the compression process of a complete input sequence. After completing the compression of one input sequence, repeat the above steps three times until the compression of all four input sequences is completed.
[0077] The decompression of the compressed binary stream (bit stream) requires the following steps:
[0078] 1) Decode the compressed bit stream. During decompression, after the binary stream is converted back to a1, a2, it is input into the compression / decompression module for asymmetric digital system entropy coding decompression operation to obtain the conditional probabilities P(x0 = X0), P(x0 < X0), P(x0 = X0, x1 = X1), P(x0 = X0|x1 < X1),... (P(x0 = X0,..., x 767 = X 767 ), P(x0 = X0,...|x 767 < X 767 ))). In terms of the process, it is the reverse process of the second step of compression to obtain the conditional probabilities.
[0079] 2) Reconstruct all the basic units of the input sequence in turn according to the decompressed conditional probabilities. Specifically, the decompressed probabilities are stored locally in the compression / decompression module to provide the dichotomy operation standard for the probability model operation module. According to P(x0 = X0), P(x0 < X0), x0 = X0 is reconstructed in a dichotomy manner first.
[0080] Specifically, the process of the dichotomy is as follows. If there is a restricted interval 0 ≤ x0 < 256 for the basic unit x0 of the input sequence, during decompression, it is first assumed that x0 = 128, and the corresponding probability in this case is calculated, and then it is judged whether the interval of x0 is 0 ≤ x0 < 128 or 128 ≤ x0 < 256 according to the size of the calculated probability. And so on until the specific data of x0 is calculated.
[0081] On this basis, x1 = X1 is reconstructed according to P(x0 = X0, x1 = X1), P(x0 = X0|x1 < X1), and so on until an input sequence (X0, X1,..., X 767 ) is reconstructed. After that, the decompression process repeats the above process for the remaining data until the original values of all four input sequences are obtained.
[0082] In the embodiments of the present invention, the task of image compression is taken as an example, but the present invention is not limited to a specific lossless compression task and is also applicable to lossless compression of other input types. Although the present invention emphasizes lossless compression, in essence, it can replace the application scenarios of lossy compression and can be applied to application scenarios involving compression.
[0083] Finally, it should be noted that the purpose of disclosing the embodiments is to help further understand the present invention, enabling those skilled in the art to understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.
Claims
1. A lossless compression and decompression method based on a probability model, characterized in that, It includes the following steps: (1) Split the input compressed data into multiple input sequences, where each input sequence includes multiple basic units; (2) Based on the trained probability model, perform probability inference on each basic unit in each input sequence successively to generate corresponding conditional probabilities, forming a probability sequence; (3) Perform a compression encoding operation of asymmetric numeral system entropy encoding on all conditional probabilities in the probability sequence generated by the probability inference operation in step (2), and convert the compressed data into a binary encoding format for storage; (4) When decoding the compressed bitstream, perform a decompression operation of asymmetric numeral system entropy encoding to obtain conditional probabilities; (5) According to the decompressed conditional probabilities, use the dichotomy method to reconstruct all basic units of the input sequence successively, and calculate the original values of all input sequences.
2. The lossless compression and decompression method based on a probability model according to claim 1, wherein The probability model is a probability circuit, a neural network, or a traditional probability distribution model.
3. The lossless compression and decompression method based on a probability model according to claim 1, characterized in that, In step (2), when performing probability inference on each basic unit in the input sequence, the generated corresponding conditional probability is the conditional probability of each basic unit based on all previous units in the input sequence: Conditional probability of basic unit x0: P(x0 = X0), P(x0 < X0) Conditional probability of basic unit x1: (P(x0 = X0, x1 = X1), P(x0 = X0, |x1 < X1)) .... Basic unit x i Conditional probability of: P(x0 = X0,…,x i = X i ), P(x0 = X0,…|x i < X i ).
4. The lossless compression and decompression method based on a probability model according to claim 1, characterized in that, In step (3), when performing a compression operation of asymmetric numeral system entropy encoding on all conditional probabilities in the probability sequence, select a normalization precision p and n, where p and n do not exceed 10, and perform operations according to the following steps: (3.1) Normalize the conditional probability and convert it into an integer frequency f(x i ), C(x i ): f(x i ) = P(x i = X i |x {i-1} = X {i-1} ) × 2 p C(x i ) = P(x i <X i |x {i-1} = X {i-1} ) × 2 p (3.2) Convert the frequency and status into two state variables a1 and a2, and the complete state s(x {i} ): a2 = s(x {i-1} ) mod (f(x i )) + C(x i ) s(x {i} ) = a1 + a2 Store the finally obtained a1 and a2 in binary form, that is, complete the compression process of a complete input sequence; After completing the compression of one input sequence, repeat the above steps until the compression of all input sequences is completed.
5. The lossless compression and decompression method based on a probability model according to claim 1, characterized in that Step (5) is specifically: For an input sequence, the basic unit x0 of the input sequence is reconstructed into specific data of x0 by dichotomy according to the conditional probability of the basic unit: there is a restricted interval 0 < x0 ≤ X for the basic unit x0 lim_0 , determine the initial hypothesis value x0 = X lim_0 / 2, calculate the corresponding probability, and judge whether the interval of x0 is 0 < x0 ≤ X lim_0 / 2 or X lim_0 / 2 < x0 ≤ X lim_0 , by iterating this process, adjust the hypothesis value x0 until the specific data of x0 is calculated; For subsequent inputs of the probability sequence, repeat this process until the original value of a single input sequence is calculated; for the remaining input sequences, repeat this process until the original values of all input sequences are calculated.
6. A lossless compression and decompression system based on a probability model, characterized in that, It includes a software module and a hardware module. The software module includes: A probability model training part, which is used to perform probability distribution fitting according to the target data set and construct a probability model describing its distribution frequency; A probability inference part, which is used to perform probability inference on each basic unit of the input sequence by using the trained probability model; An asymmetric numeral system entropy encoding compression / decompression part, which is used to perform compression encoding on the probability sequence generated by probability inference and decode the compression encoding; The hardware module includes: A configurable probability model operation module, which is used to deploy the trained probability model and execute the operations of the probability inference part; A compression / decompression module, which is used to deploy the asymmetric numeral system entropy encoding compression / decompression process and execute compression encoding / decompression operations; Among them, the probability model operation module dynamically selects an operation platform according to the probability model type, including FPGA, GPU, CPU, or a custom chip ASIC.
7. The lossless compression and decompression system based on a probability model according to claim 6, characterized in that, When the probability model training is constructed using a probabilistic circuit, the probability model operation module is implemented by an FPGA or a custom chip ASIC; when the probability model training is constructed using a neural network, the probability model operation module is implemented by a GPU; when the probability model training is constructed using a simple distribution, the probability model operation module is implemented by a CPU.
8. The lossless compression and decompression system based on a probability model according to claim 6, characterized in that, The compression / decompression module is implemented by a CPU and supports the use of Verilog to construct a specific coding module to implement random asymmetric digital system entropy coding.
9. The lossless compression and decompression system based on a probability model according to claim 6, wherein The software module is implemented using software programming codes Julia, C++, and Python.
Citation Information
Cited By
Data compression method and system
CN121143732A
A data compression method and system
CN121143732B