Bridging of gaps between diffusion models and uniform quantization for image compression
By using a forward diffusion process based on quantization and a diffusion model trained with uniform noise, the problems of noise type, noise level and discretization gap in neural image compression are solved, and high-quality image reconstruction at low bit rates is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2026-03-27
AI Technical Summary
Existing neural image compression methods tend to produce blurry and unrealistic images at low bit rates, and cannot effectively remove quantization errors, resulting in visual interference and loss of detail.
A quantization-based forward diffusion process is adopted, using a diffusion model trained with uniform noise. The discretization gap is eliminated through quantization scheduling and general quantization to ensure signal-to-noise ratio matching. Combined with the denoising capability of the diffusion model, the quantization error is corrected.
Generating realistic and detailed image reconstructions at low bit rates improves image quality, addresses differences in noise type, noise level, and discretization, and achieves more realistic image reconstruction.
Smart Images

Figure CN121750862A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] Pursuant to 35 U.SC §119(e), this application has the right to and claims the filing date benefit of U.S. Provisional Application No. 63 / 700,489, filed on September 27, 2024, entitled "A diffusion model for image compression and a linking of uniform quantization", the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] The embodiments generally relate to the field of compression. Background Technology
[0004] Multimedia content is transmitted globally via the internet and constitutes a large portion of network traffic. Developing efficient compression algorithms is crucial for the efficient transmission of multimedia content over networks.
[0005] Traditional encoder-decoder systems (CODECSs) use user-manufactured transforms, which may not perform as well as data-driven neural image compression (NIC) methods optimized for both bitrate and distortion. However, NIC methods can still produce blurry and unrealistic images, for example, at low bitrate settings. This is because these methods may be optimized for bitrate distortion, which is measured by pixel-level metrics such as mean squared error. Optimizing for low distortion, such as pixel-level error, can result in unrealistic images. This may be because emphasizing pixel-level accuracy or similarity to the original image can lead to overly smoothed or blurry outputs. Summary of the Invention
[0006] In some embodiments, a method receives an image and encodes the image into a latent representation in a latent space. A quantization process is performed on the latent representation to generate a quantized latent representation. The quantization process is based on uniform noise. The method transmits the quantized latent representation to a receiver. An inverse quantization process is performed using a diffusion model to generate a reconstructed latent representation, the diffusion model performing a denoising process iteratively multiple times based on time step t to remove noise from the reconstructed latent representation. The diffusion model is trained to perform denoising using the uniform noise. Attached Figure Description
[0007] The accompanying drawings are for illustrative purposes only, and are intended to provide examples of possible structures and operations of the disclosed inventive systems, apparatuses, methods, and computer program products. These drawings are in no way intended to limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.
[0008] Figure 1A simplified system for performing compression is drawn according to some embodiments.
[0009] Figure 2 A simplified flowchart of a method for determining quantization scheduling is drawn according to some embodiments.
[0010] Figure 3 A simplified flowchart of a method for performing a diffusion process is drawn according to some embodiments.
[0011] Figure 4 A simplified flowchart of a method for performing a training process is drawn according to some embodiments.
[0012] Figure 5 An example of a computing device is shown according to some embodiments. DETAILED DESCRIPTION
[0013] Techniques for content processing systems are described herein. In the following description, for the purposes of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Particular embodiments defined by the claims can include some or all of the examples, and can further include modifications and equivalents of the features and concepts described herein.
[0014] System overview
[0015] Generative neural image compression supports very low bitrates for data compression, enabling receivers (e.g., client devices) to synthesize details and consistently produce highly realistic images. By exploiting the similarity between quantization error and additive noise, diffusion-based generative image compression codecs use a latent diffusion model to denoise artifacts introduced by quantization. An image compression pipeline can use the diffusion model to synthesize details lost in the compression process. For example, the diffusion model can be used to correct for quantization error that can arise when using a quantization process. The error introduced in the quantization process can be similar to adding uniform noise. In fact, adding uniform noise is commonly used as a differentiable substitute for the quantization operation during the training process of a neural codec. Since diffusion models are denoising models by nature, a diffusion model can be used to counteract the quantization error introduced in the encoding process. By exploiting the similarity between quantization error and noise, the diffusion model can perform a partial denoising step that corresponds to the noise level (e.g., quantization error) of the quantized latent representation. The resulting output of the diffusion model can correct for the quantization error from the quantization process. This can improve the resulting decoded image, making it more realistic, especially at low bitrates, but this improvement can occur at all bitrates.
[0016] Previous approaches following this paradigm can have three gaps (i.e., noise type gap, discretization gap, and noise level gap) that cause the quantized data to deviate from the data distribution known by the diffusion model. When this happens, the diffusion model can not optimally denoise the artifacts introduced by quantization. However, the present system overcomes all three gaps using a quantization-based forward diffusion process.
[0017] The system addresses the three gaps of noise type gap, noise level gap, and discretization gap. The noise type gap represents the distribution difference between quantization error (e.g., uniform noise) and the Gaussian diffusion model. The noise level gap refers to the potential mismatch between the expected signal-to-noise ratio of the partially noisy data and the actual ratio. The discretization gap stems from passing discrete data to a continuous diffusion model. If not addressed, these gaps can cause the data to deviate from the distribution of the diffusion model, negatively impacting the final reconstruction quality. The system includes a quantization-based forward diffusion process to eliminate the discretization and noise level gaps and uses a uniform noise diffusion model to eliminate the noise type gap. The system consistently produces realistic and detailed reconstructed images even at very low bitrates.
[0018] In some embodiments, the system uses a quantization-based forward diffusion process that places the quantized data on the diffusion trajectory. The forward process uses a universal quantization to eliminate the discretization gap and introduces a quantization schedule to specify the signal-to-noise ratio of the quantized data. Finally, the system uses a diffusion model trained on uniform noise, which matches the distribution of the quantization error to address the noise type gap. Additionally, in some embodiments, the uniform noise diffusion model can be efficiently obtained by fine-tuning an existing Gaussian diffusion model. The system provides an image codec that produces more realistic and detailed reconstructions than previous approaches while being able to operate over a wider range of target bitrates.
[0019] The system includes a pipeline that incorporates non-integer universal quantization and a fine-tuned uniform noise diffusion model. The system selects the quantization bin width to ensure that the signal-to-noise ratio (SNR) of the quantized variable matches the expected signal-to-noise ratio at each timestep while adopting universal quantization to eliminate the discretization gap and mitigate the discretization gap between the expected Gaussian noise and the actual uniform noise in the quantization error.
[0020] In view of the quantized noise-like characteristics, the system exploits the denoising capabilities of diffusion models to develop a codec that explicitly removes the errors introduced by the quantization process using these models. The system exploits the connection between the quantization errors and the diffusion process. In view of the iterative nature of diffusion models, instead of sampling new images from the generative model, the system can take existing images and obtain a partially noisy sample at some arbitrary iteration, which can be reconstructed to the original image by performing partial diffusion steps. The system uses a general quantization and chooses a "quantization schedule" such that the quantized variables lie on the diffusion trajectory, enabling detailed and realistic reconstruction through the denoising diffusion process.
[0021] The system can address the distribution mismatch between Gaussian and uniform noise. Instead of trying to warp the uniform noise to a normal distribution, the system can replace the Gaussian diffusion model with a model that operates on uniform noise to achieve the same effect. Since the uniform distribution satisfies this requirement, diffusion models that denoise uniform noise can be used.
[0022] System
[0023] Figure 1 A simplified system 100 for performing compression is drawn in accordance with some embodiments. The system 100 includes a server system 102 and a receiver 104. The server system 102 can encode content and the receiver 104 can decode content. In some embodiments, the server system 102 can transmit encoded content to the receiver 104 over a network. However, encoding and decoding can also be performed on a single system. In some embodiments, the server system 102 can encode a video and transmit it over a network to a client device that is the receiver 104. The client device can decode the video and display it using a media player on an interface. In this case, the client device can be a smartphone, a living room device, a television, a personal computer, a laptop, a tablet device, etc. Other system configurations are also contemplated.
[0024] A diffusion model can be a class of generative models that define an iterative process that gradually corrupts an input signal by adding noise as the time step t increases (e.g., a forward diffusion process), and then attempts to simulate the reverse process by denoising the noisy image (e.g., a reverse diffusion process). Empirically, the forward process is performed by adding noise (e.g., uniform noise) to the signal. Thus, the reverse process is a denoising process that removes noise from the input. The diffusion model approximates the reverse process by estimating the noise level of the image and using it to predict the previous step of the forward process. This can remove an amount of noise from the image. This can be performed for multiple time steps. To fully denoise the image, the diffusion model can iteratively perform a full set of time steps in the reverse process.
[0025] A latent diffusion model can provide improved memory and computational efficiency by moving the diffusion process to a latent space that is lower in spatial dimensions compared to the image space (e.g., pixel space). The latent space can provide similar performance compared to a corresponding image space diffusion model while requiring fewer parameters and memory. Here, the latent diffusion model can be trained in the latent space, where an encoder can encode an image into a latent representation in the latent space. The latent representation is then processed by the latent diffusion model to denoise the latent representation through a time step t. The denoised latent representation can be decoded back into a decoded image in the image space.
[0026] The system addresses the noise type gap. In many domains (such as latent domains), quantization error can be approximated by uniform noise. However, some diffusion models assume a Gaussian noise structure because it is consistent with natural data distribution assumptions and helps facilitate easy modeling. This leads to a noise type gap - a difference between the quantization error (which can be well approximated by uniform noise) and the Gaussian noise used in the diffusion process. This inconsistency means that when uniform quantization noise interacts with a Gaussian denoising diffusion model, the model cannot correctly predict the actual noise characteristics to denoise, resulting in generated artifacts. Specifically, this mismatch can lead to visual interference effects such as unnatural color shifts, texture inconsistencies, and artificial patterns that degrade the realism and fidelity of the generated images.
[0027] Furthermore, the system addresses the discretization gap. While the neural decoder is a continuous model, it operates on the discrete representation extracted from the transmitted bitstream; most methods construct robust decoders to minimize the negative effects that arise from this. However, constructing similarly robust diffusion models in this case can be infeasible, as they model transitions between continuous states, which inherently cannot handle discrete inputs. This leads to the discretization gap - the incompatibility between using discrete input data and continuous diffusion models. Under the discretization gap, small changes in the input data are eliminated, which leads to flat textures and loss of detail, while using large quantization bin sizes results in blocky artifacts and color shifts due to low palette resolution.
[0028] Finally, the system addresses the noise level gap. Diffusion image generation assumes a fixed process of variance scheduling, which dictates the noise level at each time step t. Thus, it is crucial to ensure that the noise levels match between the forward diffusion process and the backward diffusion process (e.g., the noise at each corresponding time step t should be the same in the forward process and the reverse process; failing to do so violates the theoretical foundation of the diffusion model). However, when a different forward process is used (e.g., when quantization replaces the forward process), the noise levels in the forward and reverse processes can not be consistent. This is the noise level gap - the difference between the actual noise level of the diffusion variable and the noise level expected at any time step. Intuitively, the diffusion model either overestimates or underestimates the noise of the variable in the time step throughout the diffusion process, which leads to image reconstruction that is either noisy or overly smoothed.
[0029] The following pipeline addresses the three gaps described above - the noise type gap, the noise level gap, and the discretization gap. In this pipeline, the server system 102 can receive an image x. For example, the image x can be an image in a video being encoded. The following process can be performed on each image of the video. The encoder 106 can encode the image into a latent representation y in a latent space. The latent space can be a lower dimensional space compared to the image space. That is, the latent space can represent a compressed version of the input that captures important features. In some embodiments, the encoder 106 can be a variational autoencoder (VAE), which can be a neural network or machine learning model trained to represent images in the latent space. The encoder 106 can be considered part of the diffusion model 116 or can be independent. In some embodiments, the latent representation y can be a latent vector that captures key features of the input image in the latent space. The latent representation y can be mapped from the input image to a distribution in the latent space, which can be parameterized by a mean and a variance. While a variational autoencoder is described, other encoders that are capable of mapping input images to a latent space can also be used.
[0030] The quantization process 108 can quantize the latent representation y using a time step t-based quantization schedule to generate a quantized latent representation The quantization process reduces the precision of an image by representing it using a finite number of discrete values. The quantization process 108 can convert a continuous-valued signal into a digital signal with a finite range of values. This is done by mapping the continuous signal to a set of discrete values, called quantization levels or bins. The quantization process can be an affine transformation T applied to the latent representation y before applying integer quantization. The affine transformation can be a linear mapping used to convert floating-point values to fixed-point representations, such as integers. These channels can be channels of a latent encoding or an image, such as color, intensity, etc.
[0031] The quantized latent representation The quantized latent representation can be entropy encoded by entropy encoding 110. The quantized latent representation can be encoded into a bitstream containing the quantized latent representation using an entropy model The quantized latent representation can be entropy encoded into a bitstream using different entropy models. Entropy encoding can use entropy encoding methods, including Huffman coding and arithmetic coding, to reduce the average number of bits required to represent the quantized latent representation. Entropy encoding can reduce the average length of the quantized latent representation by assigning shorter encodings for more frequent symbols and longer encodings for less frequent symbols.
[0032] The server system 102 can transmit the bitstream to the receiver 104. The bitstream can be entropy decoded by entropy decoding 112 to reconstruct the quantized latent representation Entropy decoding is the inverse process of entropy encoding, used to reconstruct a bitstream into a quantized latent representation
[0033] The renoising process 114 can perform partial dequantization using a time step t-based quantization schedule to generate a reconstructed latent representation The renoising process dequantizes a discrete representation as input and outputs a representation in the continuous domain. The reconstructed latent representation The reconstructed latent representation can have quantization errors due to loss of information. This quantization error can be similar to noise. For example, in the process of converting continuous (high-precision) data, such as floating-point numbers, to discrete (low-precision) values, such as integers, a quantization error can introduce random variations in the latent representation. Therefore, the quantization process is split into two stages, with one half performed on the server system 102 (sender) and the other half performed on the receiver 104.
[0034] The reconstructed latent representation is input into a diffusion model 116 to perform a dequantization process. The diffusion model 116 can denoise the reconstructed latent representation to remove the noise, which can remove the introduced quantization error, outputting a denoised latent representation In other words, quantization error can be analogous to adding noise to an image, and the reconstructed latent representation can be denoised using a diffusion model 116. This process will be described in more detail below.
[0035] Decoder 118 can denoise the latent representation Decode into decoded image In some embodiments, decoder 118 may be a variational self-decoder, but other decoders may also be used. Decoder 118 can reconstruct the denoised latent representation from the latent space into the decoded image in the image space. The decoded image may be improved because the diffusion model 116 may have removed at least some of the introduced quantization errors. Removing these errors may produce a more realistic reconstructed image.
[0036] Diffusion process
[0037] The diffusion model defines the process of modeling the transformation between random noise and structured data. As the forward process (e.g., from data to noise) and the reverse process (e.g., from noise to data) are broken down into smaller steps, the transformation between each step is the addition or removal of Gaussian noise samples. Therefore, the complete diffusion process is a traversal across a series of time steps t∈[N,0]. While this process is iterative, the noisy diffusion variable y at any given time step t is partially diffused. t This can be represented by the original data y0 and the noise sample ∈:
[0038]
[0039] Where α t (referred to as "variance scheduling") defines y at each time step t. t The signal-to-noise ratio (SNR) increases with t→0. Noise samples ∈ [0, t] represent the forward propagation of the diffusion model. Variance scheduling controls the diffusion variable y at each time step t. t The trade-off between signal and noise components. The backdiffusion process is difficult to handle, therefore it is parameterized by a diffusion model that learns iteratively to y through steps t = {N,...,1,0}. t Denoising is performed. Partially denoised data y t-1 From noisy data y t The calculation yielded:
[0040]
[0041] Where ∈ θ (y t ,t) is the output of the diffusion model, which uses y t And the current time step t as input, while is the fully denoised data estimate, which is computed from the output of the diffusion model and the current time step. The fully denoised data is produced by successively performing equation (2) to produce data with slightly reduced noise until the output fully denoised data.
[0042] To use the diffusion model as an image compression codec, the forward noise process can be replaced with quantization (which is similar to adding uniform noise to the original signal) and use the diffusion model to denoise the quantization error. As mentioned above, this can lead to three shortcomings: noise type disparity, noise level disparity, and discretization disparity.
[0043] The system 100 provides an improved forward process. The system 100 uses quantization as the forward process to replace the diffusion model forward process. In the forward process using quantization, the discrete variable can be encoded into a bitstream while preserving the noise characteristics of the original diffusion variance schedule. The standard forward process (slightly reorganized from equation 1) is:
[0044]
[0045] To address the discretization disparity, the quantization process 108 can use a universal quantization as the forward process to add noise to obtain a discrete variable for entropy coding. The universal quantization can be a hard quantization dithered by a uniform random variable. This has the unique property of being equivalent in distribution to simply adding another sample (from the same random variable) to the original unquantized variable:
[0046]
[0047] where denotes rounding to an interval of width Δ. The interval of width Δ can be a quantization step or interval, which refers to the range of values that are mapped to a single quantized value. Hard quantization refers to the process of rounding a continuous-valued signal to the nearest quantization level. The term denotes a hard quantization operation that is rounded to an interval of width Δ and then dithered by a uniform random variable u to randomize the quantization error. The uniform random variable is a uniform noise sample u, u' that is uniformly distributed between -Δ / 2 and Δ / 2. The value of u, u' can take any value in this range with equal probability. The quantization process 108 and the re-noise process 114 introduce error to the input signal y, but in a way that preserves the statistical properties of the signal degradation expected by the diffusion process.
[0048] Combining equations (3) and (4), the forward noise process performed by the quantization process 108 and the re-noise process 114 becomes:
[0049]
[0050]
[0051]
[0052] Reconstructing the latent representation again becomes a continuous variable; vectorizing the latent representation adding uniform noise samples u, reconstructing the latent representation back into continuous space. For example, after rounding in the quantization process 108 (Eq. 6), the denoising process 114 outputs the reconstructed latent representation (Eq. 7) in continuous space. Eq. 6 of quantization is performed by the quantization process 108 on the server system 102, and Eq. 7 is performed by the denoising process 114 on the receiver 104.
[0053] To address the noise level disparity, the system 100 ensures that the latent representation y t and the reconstructed latent representation have the same signal-to-noise ratio at all time steps t. The signal-to-noise ratio of the reconstructed latent representation can be controlled by adjusting the quantization bin width and the uniform noise support (defined with Δ). Thus, to eliminate the noise level disparity, the system 100 matches the noise levels of the latent representation y t and the reconstructed latent representation (Eq. 7):
[0054]
[0055] The quantization process 108 uses a quantization schedule that changes the quantization bin width according to the time step t. The uniform variable support is an interval or range of values that the random variable u is uniformly distributed over. This means that each value in the range has an equal probability of being selected. The added noise u is uniformly distributed between -Δ / 2 and Δ / 2. In some embodiments, the quantization process 108 sets the bin width in Eqs. 5 and 6 can be determined by substituting Eqs. 1 and 5 into Eq. 7 and solving for the bin width Δ. Following the quantization schedule and changing the bin size resolves the signal-to-noise ratio disparity by maintaining a consistent signal-to-noise ratio between the diffuse variable of the reconstructed latent representation and the quantized variable of the latent representation y t
[0056] The quantization-based forward process simultaneously removes both discretization and noise level disparity through universal quantization and quantization scheduling, respectively. An additional benefit of quantization scheduling is that it also becomes a rate-distortion trade-off parameter, as quantization bin width directly impacts the final size of the compressed bitstream. Furthermore, since the diffusion model can denoise data at any time step, the system 100 performs compression by taking time step t as input at inference, thereby supporting multi-bitrate compression using a single model.
[0057] In some embodiments, the system 100 can train the uniform diffusion model by starting from a pre-trained Gaussian diffusion model. In some embodiments, the system 100 fine-tunes the base diffusion model and replaces the Gaussian noise with uniform noise. Since diffusion models are sensitive to changes in variance scheduling, the system 100 can keep the variance schedule constant despite the change in distribution. To adapt the Gaussian diffusion model to uniform noise, this is done by drawing the noise of the forward process to a uniform distribution between :∈~U during training in equation (1).
[0058] Quantization scheduling settings
[0059] Figure 2 A simplified flowchart 200 of a method for determining quantization scheduling is depicted in accordance with some embodiments. At 202, the system 100 receives a time step t. The time step can be received from user input, automatically generated, or dynamically determined based on the input image. The value of the time step t is a rate-distortion trade-off parameter, as quantization bin width directly impacts the final size of the compressed bitstream. The rate-distortion trade-off can improve the balance between compression rate and distortion. The compression rate can be the number of bits used to encode the data, while the distortion can be the difference between the reconstructed image and the original image. The input settings can trade-off between high bitrate low distortion or low bitrate high distortion. When quantization results in a higher compression rate, a higher bitrate is used, which can result in a more accurate reconstruction and less distortion. When quantization uses fewer bits, a lower bitrate is used, which can result in a less accurate reconstruction and more distortion. When the number of bits is reduced, a lower rate can be achieved, but can result in higher quantization error and greater distortion in the reconstructed image. However, the pipeline can use the diffusion model 116 to compensate for the higher quantization error and greater distortion. The time step t can be the optimal number of denoising steps that the diffusion model 116 should perform, which results in a true image. The quantization process 108 can add a certain amount of quantization error (e.g., noise). There can be a certain number of time steps to remove that amount of noise. When the diffusion model 116 performs this number of time steps, a true image is produced.
[0060] At 204, the system 100 determines a quantization schedule that changes the quantization bin width as a function of the time step t.
[0061] As described above, the bin width can be
[0062] At 206, the system 100 sets the quantization schedule in the quantization process 108 and the re-noising process 114. In some embodiments, the server system 102 can send the quantization schedule to the receiver 104 over a network.
[0063] At 208, the system 100 sends the time step t to the receiver 104. The diffusion model 116 can perform a denoising step on the reconstructed latent representation using the value of the time step t for a number of iterations. For example, if the value 50 is received, the diffusion model 116 can perform a denoising for 50 iterations on the reconstructed latent representation.
[0064] Diffusion model
[0065] Figure 3 A simplified flowchart 300 of a method for performing a diffusion process is drawn in accordance with some embodiments. At 302, the diffusion model 116 receives the time step t and the reconstructed latent representation The time step t can be specified by an input.
[0066] At 304, the diffusion model 116 performs a denoising operation on the reconstructed latent representation As described above, the diffusion model 116 can receive the reconstructed latent representation as input and the time step t. The diffusion model 116 can then estimate the noise level of the reconstructed latent representation and predict a previous step of the forward process to remove some noise from the reconstructed latent representation At 306, the diffusion model 116 outputs the denoised reconstructed latent representation
[0067] At 308, it is determined whether the time step t is reached. For example, if there are 50 time steps, the diffusion model 116 can compare the time step of 50 to the current number of time steps. If the time step t is not reached, the process repeats to 304, where another denoising process is performed on the output of the diffusion model 116. Here, the denoised reconstructed latent representation is denoised again using the same process as described above. This process continues until the time step t is reached. When the time step t is reached, at 310, the diffusion model 116 outputs the final denoised reconstructed latent representation Here, the diffusion model 116 can have removed some of the noise introduced by the quantization process as a quantization error.
[0068] Training
[0069] In some embodiments, the diffusion model 116 can be used as a base model. The system can use a stable diffusion as a base model, but other models can also be used. The base model (e.g., diffusion model) is trained, e.g., using a large amount of data, and thus has excellent generation capabilities to denoise the latent representation of an image. In some embodiments, the system 100 fine-tunes a pre-trained Gaussian diffusion model to enable it to handle uniform noise.
[0070] Figure 4 A simplified flowchart 400 of a method for performing a training process is drawn in accordance with some embodiments. At 402, images of a training dataset can be input into a pipeline. The images of the training dataset can be ground truths. At 404, the pipeline outputs decoded images. The pipeline can process the images in the manner described above. At 406, the decoded images can be compared to the ground truths of the original images to determine the differences between the decoded images and the original images. At 408, based on these differences, the parameters of the denoising model 116 can be adjusted to minimize the differences, e.g., using a loss function. For example, the noise ∈ is sampled from a uniform distribution of unit variance instead of a standard normal Gaussian distribution (typically used in diffusion models). This noise is applied to a source tensor (e.g., latent representation or image). The denoising model is then trained to predict the noise added to the source tensor when given a noisy tensor.
[0071] At the second stage, the process of 402-406 can be performed again. Then, at 410, the parameters of the encoder 106, the quantization process 108, the re-noising process 114, the diffusion model 116, and the decoder 118 can be frozen at the second stage. At 412, the system 100 trains the entropy encoding 110 and the entropy decoding 112 to effectively encode the quantized latent representation into a bitstream and return. Notably, since all image transformation modules are frozen, and the entropy encoding stage is lossless, the system 100 can only optimize for the rate target. Moreover, the system 100 samples the time step t during training - the function of the entropy model is an accurate probability model of the quantized data, and since the distribution of the reconstructed quantized latent representation z ^ varies depending on the input parameter t, the system 100 varies it to reflect the operating conditions at inference. The range of t can vary.
[0072] Conclusion
[0073] Accordingly, a lossy image compression codec based on a latent diffusion model can be provided to produce realistic image reconstructions at low to very low bitrates. By combining the denoising capabilities of the diffusion model with the inherent characteristics of quantization noise, the system produces perceptually satisfying reconstructions at a range of bitrates. By allowing the quantization error to use fewer bits in the quantization process, lower bitrates can be achieved. This error can be corrected by using the diffusion model 116 to remove noise. The system 100 minimizes the noise level gap, the noise type gap, and the discretization gap. Furthermore, the diffusion model can be trained to denoise uniform noise.
[0074] System
[0075] Figure 5 One example of a computing device is shown in accordance with some embodiments. According to various embodiments, a system 500 suitable for implementing the embodiments described herein includes a processor 501, a memory 503, a storage device 505, an interface 511, and a bus 515 (e.g., a PCI bus or other interconnect) in accordance with various embodiments. The system 500 can operate as a variety of devices or as any other device or service described herein. Although a particular configuration has been described, various alternative configurations can be employed. The processor 501 can perform the operations described herein. Instructions for carrying out such operations can be embodied in the memory 503, on one or more non-transitory computer-readable media, or on some other storage medium. Various specially configured devices can also be used in place of, or in addition to, the processor 501. The memory 503 can be random access memory (RAM) or another dynamic storage device. The storage device 505 can include a non-transitory computer-readable storage medium that stores information, instructions, or some combination thereof. For example, the storage device 505 can store instructions that, when executed by the processor 501, cause the processor 501 to be configured or capable of performing one or more operations of a method as described herein. The bus 515 or other communication component can support communication within the system 500. The interface 511 can connect to the bus 515 and be configured to send and receive data packets over a network. Examples of supported interfaces include, but are not limited to: Ethernet, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High- Speed Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces can include ports suitable for communication with the appropriate media. They can also include independent processors and / or volatile RAM. A computer system or computing device can include or be in communication with a monitor, printer, or other suitable display for providing any results mentioned herein to a user.
[0076] Any of the disclosed embodiments can be implemented in various types of hardware, software, firmware, computer readable media, and combinations thereof. For example, some of the techniques disclosed herein can be implemented, at least in part, through non-transitory computer readable media comprising program instructions, state information, and the like for configuring computing systems to perform various services and operations described herein. Examples of program instructions include machine code, such as produced by a compiler, and higher level code that can be executed by an interpreter. The instructions can be embodied in any suitable language, such as Java, Python, C++, C, HTML, any other markup language, JavaScript, ActiveX, VBScript, or Perl. Examples of non-transitory computer readable media include, but are not limited to: magnetic media such as hard disks and tapes; optical media such as flash memory, compact disks (CDs) or digital versatile disks (DVDs); magneto-optical media; and other hardware devices such as read-only memory ("ROM") devices and random access memory ("RAM") devices. The non-transitory computer readable media can be any combination of such storage devices.
[0077] In the foregoing specification, various technical and mechanisms can be described in the singular form for clarity. However, it should be noted that some embodiments include multiple iterations of a technique or multiple instances of a mechanism, unless otherwise specified. For example, a system uses a processor in various contexts, but can use multiple processors while remaining within the scope of the present disclosure, unless otherwise specified. Similarly, various techniques and mechanisms can be described as including a connection between two entities. However, the connection does not necessarily imply a direct, unimpeded connection, as various other entities (e.g., bridges, controllers, gateways, etc.) can exist between the two entities.
[0078] Some embodiments can be implemented in non-transitory computer readable storage media for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer readable storage media contains instructions for controlling a computer system to perform a method described by some embodiments. The computer system can include one or more computing devices. The instructions, when executed by one or more computer processors, can be configured or able to perform what is described in some embodiments.
[0079] As used in this specification and the appended claims, "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Additionally, as used in this specification and the appended claims, the term "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0080] The above description sets forth numerous specific details regarding various embodiments and aspects of some embodiments. However, these embodiments are illustrative only and not restrictive; they do not delimit the scope of the protection, which is defined only by the claims. Other arrangements, embodiments, implementations and equivalents can be employed without departing from the scope of the claims.
Claims
1. A method comprising: Receive image; The image is encoded into a latent representation in a latent space; A quantization process is performed on the latent representation to generate a quantized latent representation; as well as The quantized latent representation is transmitted to a receiver, wherein a re-noising process is performed to add uniform noise to generate a reconstructed latent representation, and a diffusion model performs a denoising process based on time step t for multiple iterations to remove noise from the reconstructed latent representation to generate a denoised reconstructed latent representation, wherein the diffusion model is trained to perform denoising using the uniform noise.
2. The method according to claim 1, further comprising: Entropy encoding is performed on the quantized latent representation, wherein the receiver performs entropy decoding on the entropy-encoded quantized latent representation.
3. The method of claim 1, wherein performing quantization on the latent representation comprises: A quantization schedule is selected, wherein the quantization schedule changes the quantization interval width according to the time step t used by the diffusion model to denoise the reconstructed latent representation.
4. The method of claim 3, wherein performing quantization on the latent representation comprises: Select the interval width to match the signal-to-noise ratio at each time step t in the time step t used by the diffusion model to denoise the reconstructed potential representation.
5. The method according to claim 4, wherein: The interval width is equal to Where α t It is variance scheduling, which defines the signal-to-noise ratio at each time step t in the denoising process.
6. The method according to claim 3, further comprising: Receive the time step t; as well as The time step t is used to determine the quantization schedule.
7. The method of claim 1, wherein performing quantization on the latent representation comprises: The latent representation is jittered using uniformly distributed random variables.
8. The method of claim 1, wherein the denoising process is performed using variance scheduling, the variance scheduling defining the signal-to-noise ratio at each time step t and increasing as the time step approaches zero.
9. The method of claim 1, wherein the re-noising process is performed on the latent representation, and uniform noise is added to the quantized latent representation during the re-noising process.
10. The method of claim 9, wherein the reconstructed latent representation following the re-noise process is a continuous variable.
11. The method according to claim 1, wherein: The diffusion model is trained to denoise the first type of noise, and The diffusion model was adjusted to denoise the uniform noise.
12. The method of claim 1, wherein the diffusion model reconstructs the information lost in the quantization of the latent representation.
13. The method according to claim 1, wherein: The quantization error from generating the quantized latent representation adds quantization noise to the reconstructed latent representation, and The diffusion model denoises the reconstructed latent representation to remove the quantization noise from the reconstructed latent representation.
14. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a computing device, enable the computing device to: Receive image; The image is encoded into a latent representation in a latent space; A quantization process is performed on the latent representation to generate a quantized latent representation; as well as The quantized latent representation is transmitted to a receiver, wherein a re-noising process is performed to add uniform noise to generate a reconstructed latent representation, and a diffusion model performs a denoising process based on time step t for multiple iterations to remove noise from the reconstructed latent representation to generate a denoised reconstructed latent representation, wherein the diffusion model is trained to perform denoising using the uniform noise.
15. A method comprising: A quantized latent representation in the latent space of a received image, wherein the image is encoded into a latent representation in the latent space and quantized to generate the quantized latent representation; A re-noise process is performed to add uniform noise to generate a reconstructed latent representation; A denoising reconstruction latent representation is generated by performing a denoising process at multiple time steps based on a diffusion model at time step t to remove noise from the reconstructed latent representation, wherein the diffusion model is trained to perform denoising using the uniform noise. as well as The denoised reconstruction latent representation is decoded into a reconstructed image.
16. The method of claim 15, wherein: The quantized latent representation is entropy encoded, and The quantized latent representation is entropy decoded before the re-noise process is performed.
17. The method of claim 15, wherein the quantized latent representation is generated by: The quantization schedule is selected, and the quantization schedule changes the interval width according to each time step t in the time step t used by the diffusion model to denoise the quantized latent representation.
18. The method of claim 15, wherein the denoising process is performed using variance scheduling, the variance scheduling defining the signal-to-noise ratio at each time step t and increasing as the time step approaches zero.
19. The method of claim 18, wherein: Select the interval width to match the signal-to-noise ratio at each time step t in the time step t used by the diffusion model to denoise the quantized latent representation.
20. The method of claim 15, wherein: The diffusion model is trained to denoise the first type of noise, and The diffusion model was adjusted to denoise the uniform noise.