An image encoding method and device based on deep learning
By introducing causal context adjustment loss function and uneven channel dimension scheduling, the problem of difficulty in generalizing the causal context model of the entropy model in the prior art is solved, and the rate distortion balance and compression performance of image compression is improved.
Patent Information
- Application Number
- CN202411263786.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-09-10
AI Technical Summary
In existing deep learning-based image compression methods, the performance of the entropy model is highly dependent on the hand-designed causal context model, making it difficult to generalize to all types of natural images, and lacking a more widely applicable context model structure.
The causal context adjustment loss function (CCA-loss) is introduced, and the causal context information is clearly adjusted through multi-stage autoregression processing capabilities and uneven channel dimension scheduling, and the network is forced to encode important information as soon as possible, thereby improving the accuracy of autoregression prediction of the entropy model.
The autoregressive prediction accuracy of the entropy model is improved, and better rate distortion balance and compression performance are achieved, while reducing the computational burden in the encoding and decoding process.
Smart Images

Figure CN119135910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image coding, and particularly to an image coding method and device based on deep learning. Background Art
[0002] The continuous growth of contemporary high-quality, high-resolution photos has driven the increasing demand for advanced image storage and transmission technologies. Therefore, lossy image compression technologies have developed rapidly in recent years. In parallel with traditional coding technologies (such as JPEG, BPG, WebP, VVC), a series of deep learning-based image compression methods (LIC) have emerged, which achieve high peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MS-SSIM) while maintaining a relatively fast running speed. Their excellent compression effect compared to VVC indicates that LIC technology is expected to keep pace with traditional technologies in the near future.
[0003] The deep learning-based lossy image compression method is built on the variational autoencoder (VAE) framework proposed by Ballé et al. The VAE-based LIC framework mainly includes an autoencoder and an entropy model. The autoencoder performs a non-linear transformation between the image space and the latent space; while the entropy model minimizes the coding length by estimating the probability distribution of the latent variables. Compared with the autoencoder, the entropy model is a unique and important part of LIC and has an important impact on the final compression result.
[0004] In the literature of LIC, the entropy model usually refers to a parameterized distribution model. In some pioneering work, Ballé et al. established an end-to-end rate-distortion optimization framework and proposed that the average coding length of the optimized latent variables is given by the Shannon cross-entropy between the actual marginal distribution and the learned entropy model. Since then, many types of entropy models have been studied. One type of research explores advanced network architectures for accurately predicting the distribution of latent representations. At the same time, another type of research studies the entropy model from a more fundamental perspective, that is, conditional distribution modeling, in order to pursue a better rate-distortion balance. Conditioning on auxiliary information (also called hyperprior) and the decoded latent representation (also called causal context) has become a mainstream strategy in modern LIC models.
[0005] Existing work usually trains the LIC network through a combination of rate loss and distortion loss. The conditional predictability of the representation is indirectly optimized, and the performance of the entropy model highly depends on hand-designed causal context models, such as channel grouping or checkerboard context models. It is difficult for a fixed context model to generalize to all types of natural images, and an improved method for obtaining a more widely applicable context model structure is urgently needed. Summary of the Invention
[0006] To address this issue, the present invention provides a Causal Context Adjustment Loss function (CCA-loss) to advance conditional distribution modeling in the entropy model. The CCA loss function is the first attempt to explicitly adjust the causal context, enabling subsequent representations to be more accurately predicted by previously decoded representations. Specifically, the autoregressive entropy model has the ability of multi-stage autoregressive processing. By dividing the latent variables into two stages, the information from the previous stage is processed to obtain the information of the current stage. In particular, considering a two-stage autoregressive context model with a hyperprior z, the latent representations to be decoded in the first stage and the second stage are denoted as y1 and y2 respectively. In addition to minimizing the cross-entropy loss by reducing the bitstream, the present invention introduces an auxiliary entropy model and a customized causal context adjustment loss, which enables y2 to be accurately estimated by y1 and z, while making it impossible for y2 to be accurately estimated only by z. Therefore, the CCA loss function explicitly guides the encoder to adjust important information to the early stage of the autoregressive entropy model, providing a more reasonable causal context sequence for entropy coding. Since the coding in the early stage is enhanced by the CCA loss function, and the present invention provides an uneven channel dimension scheduling to obtain a better rate-distortion balance. The uneven channel scheduling also helps to reduce the computational burden in the encoding and decoding processes, enabling the model to achieve state-of-the-art compression performance in less running time.
[0007] The specific steps of an image coding method based on deep learning provided by the present invention are as follows:
[0008] Step S1: Input the picture to be compressed into the encoder to obtain image latent variables in the high-dimensional latent space;
[0009] Step S2: Input the image latent variables into the hyperprior encoder and the auxiliary entropy model respectively to obtain hyperprior latent variables and auxiliary estimated probabilities;
[0010] Step S3: Quantize the hyperprior latent variables to obtain hyperprior quantized latent variables, and then through the entropy coding process to become a bitstream, which is transmitted to the decoding party and restored to the hyperprior quantized latent variables through entropy decoding, and then input into the hyperprior decoder to obtain prior information;
[0011] Step S4: The prior information is input into the auxiliary entropy model to help estimate the auxiliary probability, and input into the entropy model to estimate the probability distribution of the prior information. The entropy model is specifically a causal context entropy model based on the autoregressive framework;
[0012] Step S5: Use the probability distribution estimated by the entropy model to perform entropy coding on the image latent variables to obtain a bitstream, which is restored to the image reconstruction latent variables after compression transmission and entropy decoding;
[0013] Step S6: Finally, the image reconstruction latent variables are restored through the decoder to obtain the reconstructed image;
[0014] During the training process, the following loss functions are used to optimize the model: the rate-distortion loss function optimized by end-to-end backpropagation adjusts the compression quality and the compression size ratio; the causal context adjustment loss function adjusts the distribution of information in the latent variables estimated by the entropy model;
[0015] The following loss functions are also introduced to enhance the predictability of the causal context entropy model based on the autoregressive framework:
[0016]
[0017] Among them, CCA-loss represents the CCA loss function, represents the i-th part of the quantized latent variable. After the image latent variable is quantized, it is evenly divided into n latent variable parts in the channel dimension represents the decoded part of the i-th part of the image quantized latent variable, represents excluding the decoded part of the previous part; represents the hyperprior quantized latent variable after quantization of the hyperprior latent variable; and are the estimated distributions of the auxiliary entropy model and the entropy model respectively.
[0018] The present invention also provides an electronic device, including:
[0019] a memory; a processor; and a computer program;
[0020] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method as described above.
[0021] The present invention also provides a computer-readable storage device, characterized in that a computer program is stored thereon; the computer program is executed by a processor to implement the method as described above.
[0022] The advantages of the technical solution provided by this application are as follows:
[0023] The present invention introduces a causal context adjustment loss to explicitly adjust causal context information, forcing the network to encode important information earlier, thereby improving the autoregressive prediction accuracy of the entropy model. The present invention adopts an uneven autoregressive causal context scheduling and a convolutional autoencoder architecture, providing an efficient compression network that is easy to implement on modern deep learning platforms. It has extremely strong practicability and high performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 It is a schematic flowchart of an image coding method based on deep learning provided by an embodiment of the present invention;
[0026] Figure 2 It is a schematic diagram of the entropy model network structure of an image coding method based on deep learning provided by an embodiment of the present invention;
[0027] Figure 3 It is a schematic diagram of the autoencoder network structure of an image coding method based on deep learning provided by an embodiment of the present invention. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0029] Embodiment 1
[0030] An image coding method based on deep learning provided in this embodiment includes:
[0031] The encoder module is a non-linear transformation structure for the mutual transformation of images and latent variables;
[0032] The entropy model module is a causal context and non-linear transformation structure for entropy coding;
[0033] A rate-distortion loss function and a causal context adjustment loss function for end-to-end backpropagation optimization.
[0034] The specific steps of the image coding method based on deep learning provided in this embodiment are as follows:
[0035] Step S1: Input the picture to be compressed into the encoder to obtain the image latent variable in the high-dimensional latent space;
[0036] Step S2: Input the image latent variable into the hyperprior encoder and the auxiliary entropy model respectively to obtain the hyperprior latent variable and the auxiliary estimated probability;
[0037] Step S3: The hyperprior latent variable is quantized to become a hyperprior quantized latent variable, and then through an entropy encoding process to become a bitstream. After being transmitted to the decoding side, it is restored to the hyperprior quantized latent variable through entropy decoding, and then input into the hyperprior decoder to obtain prior information;
[0038] Step S4: The prior information is input into the auxiliary entropy model to help estimate the auxiliary probability, and input into the entropy model to estimate the probability distribution of the prior information. The entropy model is specifically a causal context entropy model based on an autoregressive framework;
[0039] Step S5: The probability distribution estimated by the entropy model is used to perform entropy encoding on the image latent variable to obtain a bitstream, which is restored to the image reconstruction latent variable after compression transmission and then entropy decoding;
[0040] Step S6: Finally, the image reconstruction latent variable is restored through the decoder to obtain the reconstructed image.
[0041] Specifically, this embodiment involves an encoder, a decoder, a hyperprior encoder, a hyperprior decoder, an entropy model, and an auxiliary entropy model; as Figure 3 shown, the encoder specifically includes a downsampling module, a residual module, and a module without a non-linear activation function; further, after the input image passes through the downsampling module, it is processed by the three modules of the residual module, 4 stacked modules without a non-linear activation function, and the downsampling module for 3 times to obtain the image latent variable; furthermore, the residual module is composed of 3 stacked residual blocks, and the residual block includes: a convolutional layer and its non-linear activation function included in the identity mapping.
[0042] The decoder specifically includes an upsampling module, a residual module, and a module without a non-linear activation function; further, after the image reconstruction latent variable passes through the upsampling module, it is processed by the three modules of the module without a non-linear activation function, the residual module, and the upsampling module for 3 times to obtain the reconstructed image. In the up / downsampling module, an up / downsampling operation is performed using a convolutional kernel.
[0043] The hyperprior encoder specifically includes a downsampling module and a Gaussian error linear unit; further, the latent variable sequentially passes through the downsampling module, the Gaussian error linear unit, the downsampling module, the Gaussian error linear unit, and the downsampling module to obtain the prior variable. The Gaussian error linear unit is an activation function with Gaussian error in the negative domain;
[0044] The hyperprior decoder specifically includes an upsampling module and a Gaussian error linear unit; further, after the hyperprior quantized latent variable passes through the hyperprior entropy model (i.e., entropy encoding), it sequentially passes through the upsampling module, the Gaussian error linear unit, the upsampling module, the Gaussian error linear unit, and the upsampling module to obtain the prior information.
[0045] Specifically, as Figure 2As shown in the figure, the entropy model of this embodiment is specifically a causal context entropy model based on an autoregressive framework, and its structure specifically includes a module without a non-linear activation function, a Gaussian error linear unit, and a convolutional layer; further, prior information is input into the entropy model, and after passing through 4 stacked modules without a non-linear activation function, a convolutional layer, a Gaussian error linear unit, a convolutional layer, a Gaussian error linear unit, and a convolutional layer in sequence, an entropy estimation probability distribution is obtained, and finally, the supplementary residual obtained through the residual prediction module is added to recover the error caused by quantization. The latent variable residual prediction module includes: 3 convolutional layers and their non-linear activation functions.
[0046] Further, the structure of the module without a non-linear activation function is specifically that the input data passes through layer normalization, a convolutional layer, a depth convolutional layer, a simple gating, a simple channel attention, a convolutional layer included by the identity mapping, and layer normalization, a convolutional layer, a simple gating, and a convolutional layer included by the identity mapping in sequence. Furthermore, layer normalization normalizes the parameters of the entire layer, depth convolution performs convolution operations on each channel of the variable respectively, and simple gating multiplies the variable after dividing it into two parts along the channel dimension.
[0047] The training method of the method provided in this example before inference use includes: a rate distortion loss function and a causal context adjustment loss function for end-to-end backpropagation optimization; specifically including: adjusting the compression quality and compression size ratio through the rate distortion loss function; adjusting the distribution of information in the latent variables estimated by the entropy model through the causal context adjustment loss function acting on the probability distributions estimated by the auxiliary entropy model and the entropy model.
[0048] The expression of the rate distortion loss function for end-to-end backpropagation optimization is:
[0049]
[0050] The expression of the causal context adjustment loss function is:
[0051]
[0052] Among them, represents the rate distortion loss function, represents the causal context adjustment loss function;
[0053] represents the reconstructed image, and x represents the original input image; represents the image quantization latent variable obtained through entropy model estimation and entropy coding, represents the auxiliary latent variable used for loss function calculation estimated by the auxiliary entropy model, represents the hyperprior quantization latent variable, represents the corresponding variable probability distribution; represents the cross - entropy of its estimated true probability distribution.
[0054] The main path of this embodiment includes: First, the original image is non - linearly transformed into an image latent variable y by an encoder, and the image latent variable is quantized into an image - quantized latent variable. After that, it is evenly divided into multiple latent variable parts in the channel dimension. Then, the transmitted image - quantized latent variable is autoregressively decoded through the estimated probability distribution of the entropy model. Autoregressive decoding is a step - by - step decoding process that uses causal context, where the causal context includes the already decoded parts. Specifically, considering a decoding process of multiple latent variable parts, first the first part is decoded, and then the second part is decoded by using the already decoded first part... and so on until all parts are decoded. Finally, the decoded quantized latent variable is reconstructed into the original image.
[0055] In addition, this embodiment uses the CCA loss to explicitly encourage encoding important information of the image into earlier causal contexts to enhance the predictability of the autoregressive entropy model. Specifically, in addition to the original autoregressive entropy model, the current - stage latent variable is estimated through the already decoded parts and the hyper - prior quantized latent variable as follows: In addition, the present invention introduces an auxiliary entropy model that takes the hyper - prior latent variable and the already decoded parts excluding the previous part as inputs: to obtain the auxiliary latent variable where represents the already decoded parts, is the quantized latent variable part, represents that the entropy model F uses the already decoded parts and the hyper - prior quantized latent variable to obtain the latent variable part at the current stage; represents excluding the already decoded parts of the previous part, represents obtaining the auxiliary latent variable by using the already decoded parts excluding the previous part and the hyper - prior latent variable.
[0056] By introducing the auxiliary entropy model, a two - stage CCA loss can be defined as follows:
[0057]
[0058] measures the information gain between latent variables, that is, the entropy gap of the auxiliary latent variable with respect to the quantized latent variable, where and are the estimated distributions of the auxiliary entropy model and the entropy model respectively. The auxiliary entropy model and the entropy model are parameterized by two networks, namely the auxiliary entropy model and the entropy model It should be noted that the auxiliary entropy model is only introduced in the training stage to better optimize the causal context; in the testing stage, the model still uses the entropy model to compress the latent variables, and the CCA loss function does not introduce additional computational burden for image compression.
[0059] A schematic diagram of the action of the CCA loss function provided by the present invention can be found in Figure 1 . By substituting the information entropy of the latent variable predicted by inputting the entropy model into the hyperprior latent variable and the auxiliary information entropy of the auxiliary latent variable predicted by inputting the auxiliary entropy model into the hyperprior latent variable into the CCA loss function for gradient descent optimization. For an autoregressive model with more than two stages, the equation can be easily extended. The multi-stage CCA loss can be defined as follows:
[0060]
[0061] where CCA-loss represents the CCA loss function, represents the i-th part of the image quantization latent variable, and the quantization latent variable is evenly divided into n latent variable parts in the channel dimension represents the decoded part of the i-th part of the quantization latent variable; and are the estimated distributions of the auxiliary entropy model and the entropy model respectively. Through the CCA loss provided by the present invention, the image compression model of this embodiment can automatically adjust the causal context, thereby improving the rate-distortion performance. This embodiment uses a network based on a convolutional neural network (CNN). Due to the adoption of the convolutional structure, the running efficiency of this embodiment is high and it is easy to be deployed on lightweight devices. The detailed structure of the autoencoder of this embodiment can be found in Figure 3 . This embodiment adopts a non-uniform channel grouping autoregressive architecture to design the entropy model. The detailed network architectures of the entropy model and the auxiliary entropy model of this embodiment can be found in Figure 2 .
[0062] In this embodiment, applying the causal context adjustment loss (CCA-loss) on the CNN-based model achieves advanced rate-distortion performance. Thanks to the advantages of the convolutional neural network, the non-uniform grouping scheme, and the proposed training method of CCA-loss, this embodiment maintains a good balance between compression latency and rate-distortion performance.
[0063] In addition, the embodiment of the present application also provides an electronic device, including:
[0064] a memory; a processor; and a computer program;
[0065] Wherein, the computer program is stored in the memory and configured to be executed by the processor to implement the method as described above.
[0066] In addition, an embodiment of the present application further provides a computer-readable storage device, on which a computer program is stored; the computer program is executed by a processor to implement the method as described above.
[0067] In the present application, the method and the device are based on the same inventive concept. Since the principles of the method and the device for solving problems are similar, the implementation of the method and the device can be referred to each other, and the repeated parts will not be described again.
[0068] The memory may be a computer-readable storage medium, which may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which can store program codes.
[0069] The above embodiments only represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. An image coding method based on deep learning, characterized in that: The following steps are involved: Step S1: Input the image to be compressed into the encoder to obtain the image latent variables in the high-dimensional latent space; Step S2: input the image latent variables into the super prior encoder and the auxiliary entropy model respectively, and obtain the super prior latent variables and the auxiliary estimated probability respectively; Step S3: quantize the super-a priori latent variable to become a super-a priori quantized latent variable, then go through an entropy coding process to become a bit stream, transmit it to the decoder, and then recover it to become a super-a priori quantized latent variable after entropy decoding, and then input it into a super-a priori decoder to obtain a priori information; Step S4: the prior information is input into the auxiliary entropy model to help estimate the auxiliary probability, and the input entropy model estimates the probability distribution of the prior information. The entropy model is specifically a causal context entropy model based on an autoregressive framework; Step S5: entropy encoding the image latent variables using the probability distribution estimated by the entropy model to obtain a bit stream, which is then restored to the image reconstruction latent variables through entropy decoding after compression transmission; Step S6: Finally, the image reconstruction latent variables are restored through the decoder to obtain the reconstructed image; During the training process, the model is optimized using the following loss functions: the rate-distortion loss function optimized by end-to-end backward gradient propagation adjusts the compression quality and compression size ratio; the causal context adjustment loss function adjusts the distribution of information in the latent variables estimated by the entropy model; The following loss function is introduced to enhance the predictability of the causal context entropy model based on the autoregressive framework: Among them, CCA-loss represents the causal context adjustment loss function, Represents the i-th part of the quantized latent variable. After the image latent variable is quantized, it is divided into n latent variable parts in the channel dimension. represents the decoded part of the i-th part of the image quantization latent variable, Indicates excluding the decoded part of the previous part; represents the super-prior quantized latent variable after the super-prior latent variable is quantized; and are the estimated distributions of the auxiliary entropy model and the entropy model respectively, and log is the cross entropy of the corresponding estimated distribution.
2. The image coding method based on deep learning according to claim 1, characterized in that: The image quantization latent variable is divided into n latent variable parts in the channel dimension Then, the transmitted latent variable part is autoregressively decoded by using the estimated probability distribution of the entropy model. The autoregressive decoding is a part-by-part decoding process using the causal context, and the causal context includes the decoded part. Specifically, the autoregressive decoding is as follows: consider a decoding process of n latent variable parts, first decode the first part, then decode the second part by using the decoded first part... and so on to complete the decoding of all parts; The auxiliary entropy model is only introduced during training. During testing and use, only the entropy model is used to compress latent variables.
3. The image coding method based on deep learning according to claim 1, characterized in that: The expression of the rate-distortion loss function of the end-to-end reverse gradient propagation optimization is: in, represents the rate-distortion loss function; represents the reconstructed image, and x represents the original input image; represents the reconstructed image quantization latent variable, represents the probability distribution of the corresponding variable; represents the cross entropy of its estimated true probability distribution.
4. The image coding method based on deep learning according to claim 3, characterized in that: The encoder specifically includes a downsampling module, a residual module, and a non-linear activation function module; further, after the input image passes through the downsampling module, it passes through the residual module, 4 stacked non-linear activation function modules, and the downsampling module to process the sequence three times to obtain the image latent variable; The decoder specifically includes an upsampling module, a residual module, and a non-linear activation function module; further, after the reconstructed image quantization latent variable passes through the upsampling module, it passes through the non-linear activation function module, the residual module, and the upsampling module for three times to obtain a reconstructed image; The structure of the module without nonlinear activation function is specifically that the input data passes through layer standardization, convolution layer, deep convolution layer, simple gating, simple channel attention, convolution layer and layer standardization, convolution layer, simple gating and convolution layer included in the identity mapping in sequence; the layer standardization standardizes the parameters of the entire layer, the deep convolution performs convolution operations on each channel of the variable respectively, and the simple gating divides the variable into two parts by the channel dimension and then multiplies them.
5. The image coding method based on deep learning according to claim 4, characterized in that: The super-a priori encoder specifically includes a downsampling module and a Gaussian error linear unit; the image latent variable is sequentially passed through the downsampling module, the Gaussian error linear unit, the downsampling module, the Gaussian error linear unit, and the downsampling module to obtain a super-a priori latent variable; The super-prior decoder specifically includes an upsampling module and a Gaussian error linear unit; further, after the super-prior quantized latent variables are entropy encoded, they are sequentially passed through an upsampling module, a Gaussian error linear unit, an upsampling module, a Gaussian error linear unit, and an upsampling module to obtain prior information.
6. The image coding method based on deep learning according to claim 5, characterized in that: The structure of the entropy model specifically includes a non-linear activation function module, a Gaussian error linear unit, and a convolutional layer; further, prior information is input into the entropy model, and the entropy estimation probability distribution is obtained by sequentially passing through 4 stacked non-linear activation function modules, convolutional layers, Gaussian error linear units, convolutional layers, Gaussian error linear units, and convolutional layers, and finally the supplementary residual obtained by the residual prediction module is added to restore the error caused by quantization.
7. The image coding method based on deep learning according to claim 6, characterized in that: The residual module is composed of three stacked residual blocks, and the residual block includes a convolution layer and a nonlinear activation function including an identity mapping; the Gaussian error linear unit is an activation function with a Gaussian error in the negative domain.
8. The image coding method based on deep learning according to claim 7, characterized in that: The latent variable residual prediction module includes: 3 convolutional layers and their nonlinear activation functions.
9. An electronic device comprising a memory; a processor; and a computer program, characterized in that: The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.
10. A computer readable storage device having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image compression method based on multi-scale space and context information fusion
CN114792347A
Image compression method and apparatus, electronic device, computer program product, and storage medium
WO2024164694A1