Lossy image compression using diffusion model

By using diffusion model for inverse quantization and denoising processing in image compression, the problem of traditional methods generating blurred images at low bit rate is solved, and the effect of generating real images at low bit rate is achieved.

CN120017831APending Publication Date: 2025-05-16DISNEY ENTERPRISES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411549132.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-18
Filing Date
2024-11-01
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional image compression methods can produce blurry and unreal images at low bit rate settings, making it difficult to maintain similarity to the original image in terms of pixels.

Method used

Using a lossy image compression pipeline based on diffusion model, the quantization error is corrected and the real image reconstruction is generated by performing inverse quantization and denoising processes in the latent space.

Benefits of technology

Generate highly realistic and detailed image reconstruction at low bit rate to improve the authenticity of decoded images, especially at low bit rate settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017831A_ABST
    Figure CN120017831A_ABST
Patent Text Reader

Abstract

In some embodiments, a method receives a quantized potential representation of an image in a potential space. The image is encoded as a representation in potential space and quantized to generate the quantized potential representation. A time step parameter generated based on the representation is received. The method performs an inverse quantization process to generate a reconstructed representation. A diffusion model performs a de-noising process of multiple iterations based on the time step parameter to remove noise from the reconstructed representation, generating a de-noised reconstructed representation. The denoised reconstructed representation is decoded as a reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] Pursuant to 35 U.S.C. §119(e), this application is entitled to and claims the benefit of the filing date of U.S. Provisional Application No. 63 / 599,541, filed on November 15, 2023, entitled “LOSSY IMAGE COMPRESSION WITH FOUNDATION DIFFUSION MODELS,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] The present application relates to a content processing system. Background Art

[0004] Multimedia content is transmitted globally over the Internet and constitutes a large portion of network traffic. Developing efficient compression algorithms is crucial for efficient transmission of multimedia content over the Internet.

[0005] Traditional encoder-decoders (CODECS) using user-crafted transformations may be outperformed by data-driven neural image compression (NIC) methods that optimize both rate and distortion. However, neural image compression methods may still produce blurry and unrealistic images in situations such as low bitrate settings. This is because these methods may optimize for rate distortion, where distortion is measured by pixel-wise matrices such as mean squared error. Optimizing for low distortion (e.g., pixel-wise error) can lead to unrealistic images. This may be because emphasizing pixel-wise accuracy or similarity to the original image may lead to overly smoothed or blurry outputs. Summary of the invention

[0006] In some embodiments, a method receives a quantized latent representation of an image in a latent space. The image is encoded as a representation in the latent space and quantized to generate the quantized latent representation. A time step parameter generated based on the representation is received. The method performs a dequantization process to generate a reconstructed representation. A diffusion model performs a denoising process for multiple iterations based on the time step parameter to remove noise from the reconstructed representation to generate a denoised reconstructed representation. The denoised reconstructed representation is decoded into a reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The included drawings are for illustrative purposes only and are only used to provide examples of possible structures and operations of the disclosed inventive systems, devices, methods, and computer program products. These drawings in no way limit any changes in form and detail that may be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.

[0008] Figure 1 A simplified system performing compression according to some embodiments is depicted.

[0009] Figure 2 A simplified flow chart of a method of performing parameter estimation according to some embodiments is depicted.

[0010] Figure 3 A simplified flow chart of a method of performing a diffusion process according to some embodiments is depicted.

[0011] Figure 4 A simplified flow chart of a method of performing a training process according to some embodiments is depicted.

[0012] Figure 5 An example of a computing device in accordance with some embodiments is illustrated. DETAILED DESCRIPTION

[0013] The technology of content processing system is described herein. In the following description, for the purpose of explanation, many examples and specific details are listed to provide a thorough understanding of some embodiments. Some embodiments as defined by the claims may include some or all of the features in these examples, or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.

[0014] System Overview

[0015] Both traditional and neural network-based image compression methods can produce unrealistic artifacts (e.g., blocking or ringing), especially in extreme compression settings (e.g., low bitrates). These artifacts appear unnatural to the human eye, and thus the decoded image may not look realistic. However, in certain bitrate scenarios, such as low bitrate scenarios, it may be preferable to decode realistic images that are more perceptually pleasing to the viewer, even if this means lower performance in terms of pixel metrics. In such scenarios, it may make more sense to allow the generation of textures and other difficult-to-encode content, even if they are not exactly the same as the source image, as long as they look realistic and resemble the original content. In some embodiments, the system improves image compression by using a generative model (e.g., a diffusion model) at the decoder end to synthesize realistic details while using as few bits as possible. The system may use a parameter estimation network that optimizes a rate-distortion target between an input image and a reconstructed image to generate realistic images.

[0016] Image compression pipelines can use diffusion models that can synthesize details lost during compression. For example, diffusion models can be used to correct quantization errors that may be produced by the quantization process. Errors introduced during the quantization process can be similar to adding noise. In fact, adding uniform noise is often used as a differentiable surrogate for the quantization operation during the training of neural codecs. Since diffusion models are essentially denoising models, diffusion models can be used to counteract quantization errors introduced during the encoding process. In addition, the system predicts the ideal number of denoising diffusion steps to compensate for the information lost during the quantization process and continue to generate realistic images. Exploiting the similarity between quantization error and noise, the diffusion model can perform a subset of denoising steps that corresponds to the noise level of the quantized latent representation (e.g., quantization error). The resulting output of the diffusion model can correct the quantization error during the quantization process. This can improve the resulting decoded image to make it a more realistic image, especially at low bitrates, but this improvement can occur at all bitrates.

[0017] In some embodiments, the system can be improved by using a basic diffusion model. The basic diffusion model may not require training to perform denoising operations. Instead, the parameter estimation process can predict the ideal denoising timestep for the diffusion model, which allows a balance between transmission cost and reconstruction quality. The diffusion model can then synthesize information lost in the quantization process based on the predicted timestep to correct quantization errors. Therefore, a lossy image compression pipeline based on potential diffusion is able to produce highly realistic and detailed image reconstructions at potentially lower bitrates. In addition to the ideal number of denoising diffusion timesteps, the parameter estimation process can also learn adaptive quantization parameters to allow realistic reconstruction for a target bitrate range.

[0018] system

[0019] Figure 1 A simplified system 100 for performing compression according to some embodiments is depicted. The system 100 includes a server system 102 and a recipient 104. The server system 102 can encode content and the recipient 104 can decode the content. In some embodiments, the server system 102 can transmit the encoded content to the recipient 104 over a network. However, the encoding and decoding can be performed on a single system. In some embodiments, the server system 102 can be encoding a video that is transmitted over a network to a client device that is the recipient 104. The client device can decode the video and display the video using a media player on an interface. In this case, the client device can be a smartphone, a living room device, a television, a personal computer, a laptop, a tablet device, etc. Other system configurations are also conceivable.

[0020] Diffusion models can be a class of generative models that define an iterative process in which the input signal is gradually corrupted by adding noise with increasing time steps t, and then attempts to model the inverse process. As a rule of thumb, the forward process is performed by adding noise, such as Gaussian noise, to the signal. Therefore, the reverse process is the denoising process of removing noise from the input. Diffusion models approximate the inverse process by estimating the noise level of the image and using it to predict the previous step of the forward process. This removes a certain amount of noise from the image. This can be performed for multiple time steps. To completely denoise the image, the diffusion model can iteratively perform a full set of time steps in the inverse process.

[0021] The latent diffusion model can provide improved memory and computational efficiency by moving the diffusion process to a latent space that is spatially lower dimensional than the image space (e.g., pixel space). The latent space can provide similar performance to the corresponding image space diffusion model while requiring fewer parameters and memory. Here, the latent diffusion model can be trained in the latent space, where the encoder can encode the image into a latent representation in the latent space. The latent representation is then processed by the latent diffusion model to denoise the latent representation through time step t. The denoised latent representation can be decoded back to a decoded image in the image space.

[0022] Since latent diffusion models are based on encoders, they encode images into a latent space that maps the image into a lower dimensional space. Therefore, latent diffusion models can also be considered a compression method. In some cases, latent diffusion models may not produce realistic image reconstructions. However, system 100 can control parameters in the pipeline to improve the performance of the pipeline and produce realistic images even at low bit rates.

[0023] In the pipeline, the server system 102 may receive an image x. For example, the image x may be an image from a video to be encoded. The following process may be performed for each image of the video. The encoder 106 may encode the image into a latent representation y in a latent space. The latent space may be a space of lower dimensionality than the image space. That is, the latent space may represent a compressed version of the input that captures important features. In some embodiments, the encoder 106 may be a variational autoencoder (VAE), which may be a neural network or a machine learning model that is trained to represent images in a latent space. The encoder 106 may be considered as part of the diffusion model 116, or may be independent. In some embodiments, the latent representation y may be a latent vector that captures the key features of the input image in the latent space. The latent representation y may be mapped from the input image to a distribution in the latent space, which may be parameterized by a mean and a variance. Although a variational autoencoder is described, other encoders that can map an input image to a latent space may also be used.

[0024] The quantization process 108 may quantize the latent representation y into a quantized latent representation The quantization process may be to apply an affine transformation T to each channel of the latent representation y, the transformation being parameterized by the quantization setting γ, and then applying integer quantization. The affine transformation may be a linear mapping for converting a floating point value to a fixed point representation (e.g., an integer). The channel may be a channel of an image, such as color, brightness, etc.

[0025] Quantization of the potential representation y can generate a finite set of discrete values ​​or codes based on the values ​​in the potential representation y. The quantization setting Y can balance between the compression used and the quality. The quantization setting Y can represent a rate-distortion balance or trade-off, which will be described in detail below. In some embodiments, the quantization setting scales the value range of the potential representation y before rounding the potential representation y to an integer. The quantization setting can control an affine transformation, and then the quantization setting is just rounded. For example, in some embodiments, the quantization parameter Y can be a number that is multiplied by the value of the potential representation y. Therefore, for example, if the potential representation y is [-3, -2.9, -2.8, ..., -2.1, -2.0, -1.9, ...2.9, 3.0], and Y is 1, the rounded quantized values ​​are [-3, -3, -3, ..., -2, -2, -2, ..., 3, 3]. If γ is larger, such as 10, the quantized values ​​will be [-30,-29,-28,…,-21,-20,-19,…,29,30]. This has more values ​​than the previous case, requiring more bits for transmission, but also introduces less quantization error. In contrast, if the quantization setting Υ is 1 / 3, the quantized values ​​become [-1,-1,-1,…,-1,-1,-1,…,1,1,1], which are very few values, requiring only a few bits, but introducing a large error. Although one value is described, the quantization setting Υ can have multiple values.

[0026] Quantized Latent Representation Entropy encoding may be performed by entropy encoding 110. The quantized latent representation may be encoded into a bitstream containing the quantized latent representation using an entropy model. The quantized latent representation can be entropy encoded into a bitstream using different entropy models. Entropy coding can reduce the average number of bits required to represent the quantized latent representation using entropy coding methods, including Huffman coding and arithmetic coding. Entropy coding can reduce the average length of the quantized latent representation by assigning shorter codes to more frequent symbols and longer codes to less frequent symbols.

[0027] The server system 102 may transmit the bitstream to the recipient 104. The entropy decoding 112 may entropy decode the bitstream to reconstruct the quantized latent representation Entropy decoding is the inverse process of entropy coding, which is used to reconstruct the bitstream into a quantized latent representation.

[0028] The inverse quantization process 114 may perform partial inverse quantization (e.g., the inverse of an affine transformation of multiplying the latent representation by the quantization setting Y) using the quantization setting Y to generate a reconstructed latent representation Dequantization is the inverse of the quantization process 108, which can be used to obtain the quantized potential representation. The original value is approximately reconstructed in . The reconstructed latent representation Quantization error may be added due to information loss. Such quantization error may be similar to noise. For example, quantization error may introduce random variations into the latent representation in the process of converting continuous (high precision) data (such as floating point numbers) to discrete (low precision) values ​​(such as integers). The quantization setting may balance the compression used and the quality. In some embodiments, using the example of multiplying the latent representation y by the quantization setting γ before rounding, the process is then divided by the quantization setting Y. Dividing by the quantization setting Y provides the required input value range for the diffusion model 116.

[0029] The reconstructed latent representation is input into the diffusion model 116 to perform a partial inverse quantization process. This partial inverse quantization of the quantized latent representation may map a finite set of discrete values ​​or codes to the original continuous values ​​found in the latent representation. The diffusion model 116 may denoise the reconstructed latent representation to remove noise, which may remove the introduced quantization error. That is, the quantization error may be similar to adding noise to an image, and the diffusion model 116 may be used to denoise the reconstructed latent representation.

[0030] As described above, the diffusion model 116 is trained to remove noise in an iterative step. The diffusion model 116 may use a time step parameter t (where t is a number) to determine the time step for denoising the reconstructed latent representation into a denoised reconstructed latent representation For example, one time step may be one iteration, in which the reconstructed latent representation is input into the diffusion model 116 and the denoised latent representation is output. An additional t-1 time steps are required to determine This denoised latent representation can be re-input into the diffusion model 116 in another iteration of time steps. This iteration performs another denoising process. The denoising process is performed as described above, where the diffusion model approximates the inverse process by estimating the noise level of the image and using it to predict the previous step of the forward process. In this process, denoising starts from a quantized latent representation that already contains structural and semantic information from the image. In this case, performing the entire range of denoising steps during the decoding process (which the diffusion model 116 can be trained to perform to produce a fully denoised image) may be wasteful and may result in an overly smoothed image. If from pure noise to the complete image, the entire range of denoising steps is used. Since the system does not start from complete noise, but from a noisy version of the latent representation y, the system only needs to run a portion of the full range of steps. Running too many steps may result in a smoothed image because the image may be denoised using too many steps. That is, during the training process, noise is added to the signal, and then the diffusion model is trained to completely denoise the signal back to the original signal. However, all time steps may not be needed because the diffusion model can perform denoising in a subset of time steps using some structural and semantic information. Thus, time step t may be a subset of the denoising diffusion steps required from the complete set of denoising steps. The output of the diffusion model 116 produces a denoised reconstructed latent representation Noise may have been removed to correct for quantization errors.

[0031] The decoder 118 can convert the denoised reconstructed latent representation Decode to decode image In some embodiments, decoder 118 may be a variational self-decoder, although other decoders may also be used. Decoder 118 may reconstruct the denoised latent representation from the latent space into a decoded image in the image space. The decoded image may be improved because at least some of the introduced quantization error may have been removed by the diffusion model 116. Such error removal may result in a more realistic reconstructed image.

[0032] Parameter estimation

[0033] As described above, the time step t and the quantization setting γ can be used as parameters. The parameter estimation network 120 can estimate the parameters used in the pipeline. For example, the parameter estimation network 120 can receive an input setting λ and output a time step t and a quantization setting Υ. The input setting λ can specify a setting for controlling a rate distortion tradeoff that balances improved compression rate and distortion. The compression rate can be the number of bits used to encode the data, and the distortion can be the difference between the reconstructed image and the original image. The input setting can be a tradeoff between high bit rate and low distortion or low bit rate and high distortion. When quantization results in a higher compression rate, a higher bit rate is used, which can result in more accurate reconstruction and less distortion. When quantization uses fewer bits, a lower bit rate is used, which can result in less accurate reconstruction and more distortion.

[0034] When the number of bits is reduced, a lower rate can be achieved, but a higher quantization error and greater distortion may be generated in the reconstructed image. However, the pipeline can use the diffusion model 116 to compensate for the higher quantization error and greater distortion. The parameter estimation network 120 can be trained to discard information that can be synthesized using the diffusion model 116 through the quantization process. That is, if the diffusion model 116 can be used to remove the quantization error introduced by the quantization process 108, the parameter estimation network 120 can output a quantization setting that provides a lower bit rate but has a higher quantization error (e.g., higher distortion). Then, even in the presence of higher distortion, the diffusion model 116 can be used to generate the lost information and reduce the distortion. This allows the pipeline to use a lower bit rate but still produce a high-quality, realistic reconstructed image. However, if the diffusion model 116 cannot remove certain quantization errors (e.g., distortion), then the parameter estimation network 120 can output a quantization setting with lower distortion using a higher bit rate.

[0035] Time step t can be the optimal number of denoising steps that the diffusion model 116 should perform. The parameter estimation network 120 can be trained to predict a subset of the entire denoising step range that can be performed to produce the best decoded image. Time step t can be used to produce a realistic image. For a given quantization setting Y, a certain amount of quantization error (e.g., noise) is added. There can be a specific number of time steps to remove this certain amount of noise. When the diffusion model 116 executes this number of time steps, a realistic image will be generated, but executing other numbers of time steps may not generate a realistic image. For example, executing too many steps will result in a smooth image, while too few steps will result in a noisy image. Therefore, by predicting the time step t to be executed and learning this prediction, the parameter estimation network 120 learns how to generate a realistic image.

[0036] In some embodiments, the parameter estimation network 120 may receive a latent representation y and an input setting λ. Based on the latent representation y and the input setting λ, the parameter estimation network 120 outputs a time step t and a quantization setting γ. In some embodiments, the input setting λ is a number, such as 5. The output time step t may also be a number, i.e., the number of diffusion iterations to be performed. The quantization setting γ is a set of numbers that parameterize the transformation. Because the optimal denoising time step depends on the amount of noise in the latent representation and therefore on the severity of the quantization, and vice versa, the parameter estimation network 120 jointly predicts the time step t and the quantization setting γ. The parameter estimation network 120 can be trained to map the amount of noise in the latent representation y and the input setting λ to the optimal number of time steps t and the quantization setting γ.

[0037] Figure 2 A simplified flow chart 200 of a method for performing parameter estimation according to some embodiments is depicted. At 202, the parameter estimation network 120 receives input settings λ and a latent representation y. The input settings λ may be set on a per-content basis, such as for a particular video, or on a per-recipient basis, or a combination of per-content and per-recipient. Fixed input settings λ may also be received for multiple videos or recipients. The latent representation y may be received from the output of the encoder 106 for a video.

[0038] At 204, the parameter estimation network 120 outputs a quantization setting γ and a time step t based on the latent representation y and the input setting λ. The parameter estimation network 120 may generate the quantization setting γ and the time step t based on training to optimize the parameters of the parameter estimation network 120 to map the input setting λ and the latent representation y to the quantization setting γ and the time step t that can reduce the bit rate while producing a realistic image. As described above, the quantization setting γ and the time step t may balance between reducing the bit rate while being able to remove noise generated by the quantization process.

[0039] At 206, the parameter estimation network 120 sends the quantization setting γ to the quantization process 108 and the inverse quantization process 114. In some embodiments, the parameter estimation network 120 may send the quantization setting γ over the network to the recipient 104. Additionally, if the parameter estimation network 120 is not local to the encoder 106 and the quantization process 108, the parameter estimation network 120 may send the quantization setting Y to the quantization process 108 over the network.

[0040] At 208, parameter estimation network 120 sends time step t and quantization setting γ to receiver 104. Diffusion model 116 may use time step t to perform a number of denoising step iterations on the reconstructed latent representation based on the value of time step t. For example, if a value of 50 is received, diffusion model 116 may perform 50 denoising iterations. Inverse quantization process 114 may use quantization setting γ.

[0041] Diffusion Model

[0042] Figure 3 A simplified flowchart 300 of a method of performing a diffusion process according to some embodiments is depicted. At 302, the diffusion model 116 receives a time step t and a reconstructed latent representation Time step t is estimated by parameter estimation network 120.

[0043] At 304, the diffusion model 116 reconstructs the latent representation As described above, the diffusion model 116 may receive the reconstructed latent representation As input and time step t. Then, the diffusion model 116 can estimate the reconstructed latent representation The noise level and predict the previous step of the forward process to reconstruct the latent representation At 306, the diffusion model 116 outputs the denoised reconstructed latent representation

[0044] At 308, a determination is made as to whether time step t has been reached. For example, if there are 50 time steps, the diffusion model 116 may compare time step 50 with the current time step. If time step t has not been reached, the process repeats to 304, where another denoising process is performed on the output of the diffusion model 116. Here, the denoised reconstructed latent representation is again denoised using the same process described above. This process will continue until time step t is reached. When time step t is reached, at 310, the diffusion model 116 outputs the final denoised reconstructed latent representation Here, the diffusion model 116 may have removed some of the noise (quantization error) introduced by the quantization process.

[0045] train

[0046] In some embodiments, the diffusion model 116 may be used as a base model. The system may use stable diffusion as a base model, but other models may also be used. These base models (e.g., diffusion models) are trained, for example, using a large amount of data, and therefore have excellent generative capabilities to denoise the potential representation of the image. Here, the parameters of the base diffusion model may not be trained in the training process of the parameter estimation network 120. Instead, the parameters of the diffusion model 116 may remain fixed during the training process. The training focuses on providing optimization for the rate distortion of further quantizing the potential representation produced by the quantization process.

[0047] Figure 4A simplified flow chart 400 of a method for performing a training process according to some embodiments is depicted. At 402, a training data set of images may be input into the pipeline. The training data set of images may be ground truth. At 404, the pipeline outputs a decoded image. The pipeline may process the image as described above. At 406, the decoded image may be compared with the ground truth of the original image to determine the difference between the decoded image and the original image. At 408, based on the difference, the parameters of the parameter estimation network 120 may be adjusted, for example, using a loss function to minimize the difference. Here, the time step t and the quantization settings may be adjusted to minimize the loss. By adjusting the parameters of the parameter estimation network, the pipeline may be trained to produce realistic images at low bit rates. This may improve the use of the system because expensive training of the diffusion model 116 may not be required. In addition, the system may adjust the parameters of the entropy model of the entropy encoding 110 or the entropy decoding 112. Here, the system learns the best entropy model for the quantized data so that it can be more efficiently encoded into a bitstream.

[0048] in conclusion

[0049] Thus, a lossy image compression codec based on a latent diffusion model can be provided to produce realistic image reconstructions at low to very low bit rates. By combining the denoising capabilities of the diffusion model with the inherent characteristics of quantization noise, the system predicts the ideal number of denoising steps to produce perceptually pleasing reconstructions at a range of bit rates. Lower bit rates can be achieved by allowing quantization errors in the quantization process to use fewer bits. This error can be corrected by removing noise using the diffusion model 116. Training of the parameter estimation network 120 can generate quantization settings and time steps t that can produce realistic images at lower bit rates.

[0050] The system can achieve faster decoding time than other diffusion codecs due to the reuse of the underlying diffusion model and a lower training budget. In addition, the system provides control over the rate-distortion trade-off by using the input settings of the parameter estimation network.

[0051] system

[0052] Figure 5An example of a computing device according to some embodiments is illustrated. According to various embodiments, a system 500 suitable for implementing the embodiments described herein includes a processor 501, a memory module 503, a storage device 505, an interface 511, and a bus 515 (e.g., a PCI bus or other interconnect structure). The system 500 can operate as various devices, or as any other device or service described herein. Although a specific configuration is described, various alternative configurations are possible. The processor 501 can perform the operations described herein. Instructions for performing such operations can be embodied in the memory 503, on one or more non-transitory computer-readable media, or on other storage devices. Various specially configured devices can also be used in place of or in conjunction with the processor 501. The memory 503 can be a random access memory (RAM) or other dynamic storage device. The storage device 505 can include a non-transitory computer-readable medium that stores information, instructions, or a combination thereof, for example, when executed by the processor 501, instructions that enable the processor 501 to be configured or capable of performing one or more operations of the method described herein. The bus 515 or other communication components can support information communication within the system 500. Interface 511 can be connected to bus 515 and configured to send and receive data packets through the network. Examples of supported interfaces include, but are not limited to, Ethernet, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High Speed ​​Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces may include ports suitable for communicating with corresponding media. They may also include independent processors and / or volatile RAM. The computer system or computing device may include or communicate with a monitor, printer, or other suitable display device to provide any results mentioned herein to the user.

[0053] Any of the disclosed embodiments may be embodied in various types of hardware, software, firmware, computer-readable media, and combinations thereof. For example, some of the techniques disclosed herein may be implemented at least in part by non-transitory computer-readable media including program instructions, state information, and the like, for configuring a computing system to perform the various services and operations described herein. Examples of program instructions include machine code generated by a compiler, and higher-level code that can be executed by an interpreter. The instructions may be embodied in any suitable language, such as Java, Python, C++, C, HTML, any other markup language, JavaScript, ActiveX, VBScript, or Perl. Examples of non-transitory computer-readable media include, but are not limited to: magnetic media, such as hard disks and tapes; optical media, such as flash memory, compact disks (CDs), or digital versatile disks (DVDs); magneto-optical media; and other hardware devices, such as read-only memory ("ROM") devices and random access memory ("RAM") devices. Non-transitory computer-readable media may be a combination of any such storage devices.

[0054] In the foregoing description, various techniques and mechanisms may be described in singular form for clarity. However, it should be noted that some embodiments include multiple iterations of a technique or multiple instances of a mechanism, unless otherwise specified. For example, a system uses a processor in various contexts, but multiple processors may be used while still within the scope of the present disclosure, unless otherwise specified. Similarly, various techniques and mechanisms may be described as including a connection between two entities. However, a connection does not necessarily mean a direct, unimpeded connection, as various other entities (e.g., bridges, controllers, gateways, etc.) may be located between the two entities.

[0055] Some embodiments may be implemented in a non-transitory computer readable medium for use with or in connection with an instruction execution system, device, system, or machine. The computer readable storage medium contains instructions for controlling a computer system to perform the methods described in some embodiments. The computer system may include one or more computing devices. These instructions, when executed by one or more computer processors, may be configured or capable of performing the content described in some embodiments.

[0056] As used in this specification and the claims that follow, “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. In addition, as used in this specification and the claims that follow, the meaning of “in” includes “in” and “on” unless the context clearly dictates otherwise.

[0057] The above description illustrates various embodiments, as well as examples of how to implement aspects of some embodiments. The above examples and embodiments should not be considered as the only embodiments, but are intended to illustrate the flexibility and advantages of some embodiments, which are defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations and equivalents may be adopted without departing from the scope defined by the claims.

Claims

1. A method comprising: receiving a quantized latent representation of an image in a latent space, wherein the image is encoded into the latent representation in the latent space and quantized to generate the quantized latent representation; receiving a time step parameter generated based on the latent representation; performing an inverse quantization process to generate a reconstructed latent representation; performing a plurality of iterative denoising processes based on the time step parameter using a diffusion model to remove noise from the reconstructed latent representation and generate a denoised reconstructed latent representation; as well as The denoised reconstructed latent representation is decoded into a reconstructed image.

2. The method of claim 1, wherein: The quantized latent representation is entropy encoded, and The quantized latent representation is entropy decoded before performing an inverse quantization process.

3. The method of claim 1 , further comprising: receiving a quantization setting generated based on the latent representation in the latent space; as well as The inverse quantization process is performed using the quantization settings.

4. The method of claim 3, wherein the inverse quantization process uses the quantization setting to adjust the number of bits used for quantization to generate the reconstructed latent representation.

5. The method of claim 3, wherein: The quantization settings are generated based on the latent representation of the image and input parameters, wherein the input parameters are based on a rate-distortion balance.

6. The method of claim 1, wherein: The time step parameters are generated based on the latent representation of the image and input parameters, wherein the input parameters are based on a rate-distortion balance.

7. The method of claim 1, wherein: The quantization and dequantization processes add quantization errors to the reconstructed latent representation compared to the latent representation of the image, and The diffusion model is configured to remove noise associated with the quantization error.

8. The method of claim 1, wherein: quantization error generated by generating the quantized latent representation adds noise to the reconstructed latent representation, and The diffusion model denoises the reconstructed latent representation to remove noise from the reconstructed latent representation.

9. The method of claim 1, wherein: A network is trained to generate the time step parameters.

10. The method of claim 9, wherein: During training of the network, the diffusion model is not trained.

11. The method of claim 9, wherein: The network is trained to generate a quantization setting that is used by the inverse quantization process to adjust the number of bits used for quantization to generate the reconstructed latent representation.

12. A non-transitory computer-readable medium having computer-executable instructions stored thereon that, when executed by a computing device, cause the computing device to: receiving a quantized latent representation of an image in a latent space, wherein the image is encoded into the latent representation in the latent space and quantized to generate the quantized latent representation; receiving a time step parameter generated based on the latent representation; performing an inverse quantization process to generate a reconstructed latent representation; performing a plurality of iterative denoising processes based on the time step parameter using a diffusion model to remove noise from the reconstructed latent representation and generate a denoised reconstructed latent representation; as well as The denoised reconstructed latent representation is decoded into a reconstructed image.

13. A method comprising: receiving an image; encoding the image into a latent representation in a latent space; estimating time step parameters based on the latent representation; performing a quantization process on the latent representation to generate a quantized latent representation; as well as The quantized latent representation is transmitted to a receiver, where an inverse quantization process is performed to generate a reconstructed latent representation, and a diffusion model performs a denoising process for multiple iterations based on the time step parameter to remove noise from the reconstructed latent representation.

14. The method of claim 13, wherein: The quantized latent representation is entropy encoded, and The quantized latent representation is entropy decoded before performing an inverse quantization process.

15. The method of claim 13, further comprising: Determine the quantization settings based on the generation of latent representations in the latent space; as well as The quantization process is performed using the described quantization settings.

16. The method of claim 15, wherein: The quantization settings are generated based on a latent representation of the image and input parameters, wherein the input parameters are based on a rate-distortion tradeoff.

17. The method of claim 13, wherein estimating the time step parameter comprises: The time step parameters are estimated based on a latent representation of the image and input parameters, wherein the input parameters are based on a rate-distortion tradeoff.

18. The method of claim 13, wherein: The quantization error from the quantization process adds noise to the reconstructed latent representation, and The diffusion model denoises the reconstructed latent representation to remove noise from the reconstructed latent representation.

19. The method of claim 13, wherein: A network is trained to generate the time step parameters.

20. The method of claim 9, wherein: The network is trained to generate a quantization setting that is used by the quantization process to adjust a number of bits used for quantization to generate the quantized latent representation.