Lossy image compression with diffusion model
By using a diffusion model to synthesize realistic details and correct quantization errors, the system addresses the issue of unrealistic artifacts in low-bitrate image compression, achieving a balance between efficiency and quality.
Patent Information
- Application Number
- JP2024198019
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-18
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-27
AI Technical Summary
Conventional image compression methods, both traditional and neural network-based, often produce unrealistic artifacts at extreme compression settings, such as low bitrates, leading to blurry or artificial-looking images.
The system employs a generative diffusion model on the decoder side to synthesize realistic image details while minimizing bit consumption. This involves a parameter estimation network that optimizes the rate-distortion tradeoff, and a diffusion model that corrects quantization errors by performing a subset of noise removal steps.
The approach results in more realistic and perceptually comfortable images at low bitrates, effectively balancing compression efficiency with image quality, and can produce high-quality reconstructions across a range of bitrates.
Smart Images

Figure 2025081265000001_ABST
Abstract
Description
Cross - reference to related applications
[0001]
[0001] According to 35 U.S.C. § 119(e), this application claims the benefit of, and claims priority to, the filing date of U.S. Provisional Patent Application No. 63 / 599,541, entitled "LOSSY IMAGE COMPRESSION WITH FOUNDATION DIFFUSION MODELS," filed on November 15, 2023, the content of which is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0002]
[0002] Multimedia content is distributed globally over networks and occupies a large portion of the traffic. The development of efficient compression algorithms is important for the efficient distribution of multimedia content over networks.
[0003]
[0003] Conventional encoders - decoders (CODECS) that use hand - crafted transformations by users can be outperformed by data - driven neural image compression (NIC) methods that optimize for both rate and distortion. However, neural image compression methods can still produce blurry and unrealistic images, for example, in low - bitrate settings. This is because these methods can be optimized for rate - distortion, where distortion is measured using a pixel - level metric such as mean squared error. Optimization for low distortion, such as pixel - level error, can result in unrealistic images. This is because emphasizing pixel - level accuracy or similarity to the original image can lead to overly smooth or blurry outputs.
[0004]
[0004] The accompanying drawings are for illustrative purposes only and are useful only to provide examples of possible structures and operations for the disclosed system, apparatus, method, and computer program product. These drawings do not limit in any way the forms and details that can be made by those skilled in the art without departing from the spirit and scope of the disclosed embodiments.
Brief Description of the Drawings
[0005]
Figure 1
[0005] FIG. 1 illustrates a simplified system for performing compression according to some embodiments.
Figure 2
[0006] FIG. 2 illustrates a simplified flowchart of a method for performing parameter estimation according to some embodiments.
Figure 3
[0007] FIG. 3 illustrates a simplified flowchart of a method for performing a diffusion process according to some embodiments.
Figure 4
[0008] FIG. 4 illustrates a simplified flowchart of a method for performing a training process according to some embodiments.
Figure 5
[0009] FIG. 5 illustrates an example of a computing device according to some embodiments.
Modes for Carrying Out the Invention
[0006]
[0010] In this specification, techniques for a content processing system are described. In the following description, for the purpose of explanation, numerous examples and specific details are set forth in order to provide a thorough understanding of some embodiments. Some embodiments defined by the claims may include some or all of the features in these examples, either alone or in combination with other features described below, and may further include modifications and equivalents of the features and concepts described herein.
[0007]
[0011] <Overview of the System>
[0008]
[0012] Both conventional image compression methods and neural network-based image compression methods can produce unrealistic artifacts (e.g., blocking or ringing) at extreme compression settings (e.g., low bitrates). These artifacts look artificial to the human eye, and the decoded image can look unrealistic. However, in some bitrate scenarios, such as low bitrate scenarios, it may be preferable to decode a realistic image that is perceptually comfortable for the viewer, even if it means lower performance in terms of pixel-level metrics. In such scenarios, it may be more interesting to enable the generation of textures and other difficult-to-encode content, even if they are not exactly the same as the source image, as long as they look realistic and resemble the original content. In some embodiments, the system improves image compression by using a generative model (e.g., a diffusion model) on the decoder side to synthesize realistic details while consuming as few bits as possible. The system may use a parameter estimation network that can optimize the rate-distortion target between the input image and the reconstructed image to produce a realistic image.
[0009]
[0013] The image compression pipeline may use a diffusion model that can synthesize details lost from the compression process. For example, the diffusion model may be used to correct quantization errors that may result as a consequence of using a quantization process. The errors introduced during quantization may be similar to adding noise. In fact, adding uniform noise is often used as a differentiable surrogate for the quantization operation during the training of neural codecs. Since the diffusion model is essentially a noise removal model, in that case, the diffusion model may be used to cancel out the quantization errors introduced during encoding. Additionally, the system predicts the ideal number of noise removal diffusion steps to compensate for the information lost during quantization and consistently produces realistic images. Leveraging the similarity between quantization error and noise, the diffusion model may perform a subset of the noise removal steps corresponding to the noise level (e.g., quantization error) of the quantized latent representation. The output resulting from the diffusion model may correct the quantization errors from the quantization process. This can improve the resulting decoded image with more realistic images, especially at low bitrates, but this improvement can occur at all bitrates.
[0010]
[0014] In some embodiments, the system can be improved by using a base diffusion model. The base diffusion model may not need to be trained to perform denoising operations. Rather, a parameter estimation process can predict the ideal denoising time steps for the diffusion model, which enables balancing between transmission cost and reconstruction quality. The diffusion model can then synthesize the information lost during quantization to correct the quantization error based on the predicted number of time steps. Thus, a latent diffusion-based irreversible image compression pipeline can create a very realistic and detailed image reconstruction at a potentially low bitrate. The parameter estimation process can learn adaptive quantization parameters in addition to the ideal number of denoising diffusion time steps to enable realistic reconstructions for a range of target bitrates.
[0011]
[0015] <System>
[0012]
[0016] FIG. 1 illustrates a simplified system 100 for performing compression according to some embodiments. System 100 includes a server system 102 and a receiver 104. Server system 102 can encode content, and receiver 104 can decode the content. In some embodiments, server system 102 can transmit the encoded content to receiver 104 over a network. However, encoding and decoding can be performed on a single system. In some embodiments, server system 102 can encode video transmitted over a network to a client device acting as receiver 104. The client device can decode the video and display the video using a media player on the interface. The client device in this case can be a smartphone, a living room device, a television, a personal computer, a laptop, a tablet device, etc. Other system configurations can also be recognized.
[0013]
[0017] A diffusion model can be a class of generative models that define an iterative process that gradually destroys the input signal by adding noise as the time step t increases and then attempts to model the reverse process. Empirically, the forward process is executed by adding noise such as Gaussian noise to the signal. Therefore, the reverse process is a noise removal process that removes noise from the input. The diffusion model approximates the reverse process by estimating the noise level of the image and using it to predict the step before the forward process. This can remove a certain amount of noise from the image. This can be done over several time steps. To completely remove the noise from the image, the diffusion model can iteratively execute the full set of time steps in the reverse process.
[0014]
[0018] The latent diffusion model can provide improved memory and computational efficiency by moving the diffusion process to a spatially lower-dimensional latent space compared to the image space (e.g., pixel space). The latent space can provide similar performance to the corresponding image space diffusion model while requiring fewer parameters and less memory. Here, the latent diffusion model can be trained in the latent space where the encoder can encode the image into a latent representation in the latent space. Then, the latent representation is processed by the latent diffusion model to remove noise from the latent representation by the time step t. The denoised latent representation can be decoded back into the decoded image in the image space.
[0015]
[0019] The latent diffusion model is encoder-based because it encodes the image into a latent space that maps the image to a lower-dimensional space. Therefore, the latent diffusion model can also be considered a type of compression method. In some cases, the latent diffusion model may not produce a realistic reconstruction of the image. However, the system 100 can control the parameters in the pipeline to improve the performance of the pipeline to produce a realistic image even at a low bit rate.
[0016]
[0020] In a pipeline, the server system 102 may receive an image x. For example, the image x may be an image from an encoded video. The following process may be executed for each image of the video. The encoder 106 may encode the image into a latent representation y in a latent space. The latent space may be a lower-dimensional space compared to the image space. That is, the latent space may represent a compressed version of the input that captures important features. In some embodiments, the encoder 106 may be a variational autoencoder (VAE), which may be a neural network or a machine learning model trained to represent an image in a latent space. The encoder 106 may be considered part of the diffusion model 116 or may be separate. In some embodiments, the latent representation y may be a latent vector that captures important features of the input image in the latent space. The latent representation y may be mapped from the input image to a distribution in the latent space parameterized by a mean and a variance. Although a variational autoencoder is described, other encoders that can map an input image to a latent space may also be used.
[0017]
[0021] The quantization process 108 may quantize the latent representation y into a quantized latent representation
[0018]
Number
[0019] The quantization process may be an affine transformation T for each channel of the latent representation y parameterized by quantization settings
[0020]
Number
[0021] before applying integer quantization. The affine transformation may be a linear mapping used to convert floating-point values to a fixed-point representation such as an integer. The channels may be channels of an image such as color, intensity, etc.
[0022]
[0022] The quantization of the latent representation y can generate a finite set of discrete values or codes based on the values found in the latent representation y. Quantization settings
[0023]
Number
[0024] can balance the compression used and the quality. Quantization settings
[0025]
Number
[0026] can represent a rate-distortion balance or trade-off, which will be described in more detail below. In some embodiments, the quantization settings scale the range of values of the latent representation y before rounding the latent representation y to an integer. The quantization settings can control an affine transformation, in which case the quantization settings are simply rounding. For example, in some embodiments, the quantization parameter
[0027]
Number
[0028] is a single number that is multiplied by the value of the latent representation y. Thus, for example, if the latent representation y is [-3, -2.9, -2.8, ..., -2.1, -2.0, -1.9, ..., 2.9, 3.0] and gamma is 1, then the quantized values will be [-3, -3, -3, ..., -2, -2, -2, ..., 3, 3] when rounding is performed. If gamma is larger, for example 10, the quantized values will be [-30, -29, -28, ..., -21, -20, -19, ..., 29, 30]. This is many more values than in the previous case, which requires more bits to transmit but also introduces less quantization error. On the other hand, the quantization settings
[0029]
Number
[0030] When it is 1 / 3, the quantized value becomes a very small value [-1, -1, -1, ..., -1, -1, -1, ..., 1, 1, 1], which consumes only a few bits but introduces a large error. Although one value is described, the quantization setting
[0031]
Number
[0032] can be multiple values.
[0033]
[0023] Quantized latent representation
[0034]
Number
[0035] can be entropy encoded by entropy encoding 110. The entropy model converts the quantized latent representation into a bit stream containing the quantized latent representation
[0036]
Number
[0037] can be used to encode. Different entropy models can be used to entropy encode the quantized latent representation into a bitstream. Entropy coding can reduce the average number of bits required to represent the quantized latent representation using an entropy coding method including Huffman coding and arithmetic coding. Entropy coding can reduce the average length of the quantized latent representation by assigning shorter codes to symbols with higher frequencies and longer codes to symbols with lower frequencies.
[0038]
[0024] The server system 102 can transmit the bitstream to the receiver 104. The entropy decoder 112 can entropy decode the bitstream to reconstruct the quantized latent representation
[0039]
Number
[0040] Entropy decoding is the reverse process of entropy coding to reconstruct it into the quantized latent representation
[0041]
Number
[0042] is the reverse process of entropy coding to reconstruct it into the quantized latent representation.
[0043]
[0025] The inverse quantization process 114 uses the quantization settings to generate the reconstructed latent representation
[0044]
Number
[0045] using part of the inverse quantization (e.g., the quantization settings
[0046]
Mathematics
[0047] can perform the inverse of the affine transformation that multiplies the latent representation by (). Inverse quantization is the reverse process of the quantization process 108, and the quantized latent representation
[0048]
Mathematics
[0049] can approximately reconstruct the original value from. The reconstructed latent representation
[0050]
Mathematics
[0051] may have quantization errors added due to information loss. This quantization error may be similar to noise. For example, the quantization error may introduce random fluctuations into the latent representation during the process of converting continuous (high-precision) data such as floating-point numbers into discrete (low-precision) values such as integers. The quantization settings can balance the compression used and the quality. In some embodiments, before rounding, the quantization settings
[0052]
Mathematics
[0053] Using the example of multiplying the latent representation y by, in this case, this process is
[0054]
Mathematics
[0055] will divide by. The quantization settings
[0056]
Mathematics
[0057] Dividing by provides the diffusion model 116 with the range of values as the required input.
[0058]
[0026] The reconstructed latent representation is input into the diffusion model 116 to perform a part of the inverse quantization process. This part of the inverse quantization of the quantized latent representation can map a finite set of discrete values or codes to the original continuous values found in the latent representation. The diffusion model 116 can denoise the reconstructed latent representation to remove the introduced quantization error. That is, the quantization error can be similar to adding noise to the image, and the diffusion model 116 can be used to denoise the reconstructed latent representation.
[0059]
[0027] As described above, the diffusion model 116 is trained to remove noise in iterative steps. The diffusion model 116 denoises the reconstructed latent representation to obtain the denoised reconstructed latent representation
[0060]
Number
[0061] The parameter of the time step t can be used to determine the number of iterations to be used to obtain, where t is a number. For example, the time step can be one iteration in which the reconstructed latent representation is input into the diffusion model 116 and the denoised latent representation
[0062]
Number
[0063] is output. t - 1 additional time steps are
[0064]
Number
[0065] is required to determine. This noise-removed latent representation can be re-input into the diffusion model 116 in another iteration for the time step. This iteration performs another noise removal process. The noise removal process is performed as described above, where the diffusion model estimates the noise level of the image and approximates the reverse process by using it to predict the step before the forward process. In this process, noise removal starts from the quantized latent representation that already contains the structural and semantic information from the image. In this scenario, it may be wasteful to perform the full range of noise removal steps during the decoding process in which the diffusion model 116 is trained to produce a completely noise-removed image, and it may result in an overly smoothed image. The full range of noise removal steps is used when going from pure noise to a complete image. Since the system starts from a noisy version of the latent representation y rather than from complete noise, in this case, the system only needs to perform a portion of the full range of steps. Executing too many steps can cause the image to be noise-removed at too many steps, resulting in a smooth image. That is, in the training process, noise is added to the signal, and then the diffusion model is trained to completely remove the noise from the signal and return it to the original signal. However, since the diffusion model can utilize some structural and semantic information to perform noise removal in a subset of the time steps, not all time steps may be required. Thus, the time step t can be a subset of the required noise removal diffusion steps from the full set of noise removal steps. The output of the diffusion model 116 is the noise-removed reconstructed latent representation that may have been noise-removed to correct the quantization error.
[0066]
Number
[0067] to create.
[0068]
[0028] Decoder 118 decodes the noise-removed latent representation
[0069]
Number
[0070] into the decoded image
[0071]
Number
[0072] and can decode it into the decoded image in the image space. In some embodiments, decoder 118 can be a variational auto-decoder, but other decoders can also be used. Decoder 118 reconstructs the noise-removed latent representation from the latent space into the decoded image
[0073]
Number
[0074] in the image space. The decoded image can be improved in that at least part of the introduced quantization error may have been removed by diffusion model 116. The removal of this error can result in a more realistic reconstructed image.
[0075]
[0029] <Parameter Estimation>
[0076]
[0030] As described above, time step t and quantization settings
[0077]
Number
[0078] can be used as a parameter. The parameter estimation network 120 can estimate the parameters used in the pipeline. For example, the parameter estimation network 120 receives an input setting λ and outputs a time step t and a quantization setting
[0079]
Number
[0080] can be output. The input setting λ can specify a setting that can control the rate-distortion trade-off when improving the balance between the compression rate and distortion. The compression rate can be the number of bits used to encode the data, and the distortion can be the difference between the reconstructed image and the original image. The input setting can trade off a high bit rate for low distortion or a low bit rate for high distortion. When quantization results in a higher compression rate, a higher bit rate is used, which can lead to a more accurate reconstruction and less distortion. When fewer bits are used for quantization, a lower bit rate is used, which can lead to a less accurate reconstruction and more distortion.
[0081]
[0031] When the number of bits is reduced, a lower rate is achieved, but higher quantization error and greater distortion can occur in the reconstructed image. However, the pipeline can compensate for the higher quantization error and greater distortion using the diffusion model 116. The parameter estimation network 120 can be trained to discard the information that can be synthesized using the diffusion model 116 through the quantization process. That is, if the diffusion model 116 can be used to remove the quantization error introduced by the quantization process 108, the parameter estimation network 120 can output a quantization setting that provides a lower bit rate but with higher quantization error (e.g., higher distortion). In this case, even if there is higher distortion, the diffusion model 116 can be used to generate the lost information and reduce the distortion. This enables the pipeline to use a lower bit rate but still produce a high-quality realistic reconstructed image. However, if the diffusion model 116 cannot remove some quantization error (e.g., distortion), the parameter estimation network 120 can output a quantization setting that can use a higher bit rate with lower distortion.
[0082]
[0032] The time step t can be the optimal number of noise removal steps for the diffusion model 116 to execute. The parameter estimation network 120 can be trained to predict a subset of the full range of noise removal steps that can be executed to produce an optimal decoded image. The time step t can be used to produce a realistic image. For a given quantization setting
[0083]
Number
[0084] For a certain amount of quantization error (e.g., noise) is added. There may be a specific number of time steps to remove that certain amount of noise. When this number of time steps is executed by the diffusion model 116, a realistic image is generated, but performing other numbers of time steps may not result in a realistic image. For example, if too many steps are performed, a smooth image is obtained, and if too few steps are performed, a noisy image is obtained. Therefore, by predicting the number of time steps t to be executed and learning this prediction, the parameter estimation network 120 learns how to generate a realistic image.
[0085]
[0033] In some embodiments, the parameter estimation network 120 may receive a latent representation y and an input setting λ. Based on the latent representation y and the input setting λ, the parameter estimation network 120 determines the time step t and the quantization setting
[0086]
Number
[0087] and outputs them. In some embodiments, the input setting λ is a number such as 5. The output time step t may also be a number, which is the number of diffusion iterations to be executed. The quantization setting
[0088]
Number
[0089] is a set of numbers that parameterize the transformation. The optimal number of noise removal time steps depends on the amount of noise in the latent representation and thus the severity of the quantization, and vice versa. Therefore, the parameter estimation network 120 determines the time step t and the quantization setting
[0090]
Number
[0091] Predict them together. The parameter estimation network 120 determines the amount of noise in the latent representation y and the input setting λ, the optimal number of time steps, and the quantization setting
[0092]
Number
[0093] so as to be trained to map to.
[0094]
[0034] FIG. 2 illustrates a simplified flowchart 200 of a method for performing parameter estimation according to some embodiments. At 202, the parameter estimation network 120 receives an input setting λ and a latent representation y. The input setting λ can be set for each receiver, for each content such as a specific video, or a combination of each content and each receiver. A fixed input setting λ can also be received for several videos or receivers. The latent representation y can be received from the output of the encoder 106 for the video.
[0095]
[0035] At 204, the parameter estimation network 120 outputs a quantization setting
[0096]
Number
[0097] and a time step t based on the latent representation y and the input setting λ. The parameter estimation network 120 maps the input setting λ and the latent representation y to a quantization setting that reduces the bitrate while generating realistic images
[0098]
Number
[0099] And, based on training to optimize the parameters of the parameter estimation network 120 for mapping to the quantization settings and time step t
[0100]
Number
[0101] And time step t can be generated. As described above, the quantization settings
[0102]
Number
[0103] And time step t can balance reducing the bit rate while being able to remove noise resulting from the quantization process.
[0104]
[0036] In 206, the parameter estimation network 120 sends the quantization settings
[0105]
Number
[0106] to the quantization process 108 and the inverse quantization process 114. In some embodiments, the parameter estimation network 120 can send the quantization settings
[0107]
Number
[0108] to the receiver 104 via the network. Also, when the parameter estimation network 120 is not local to the encoder 106 and the quantization process 108, the parameter estimation network 120 sends the quantization settings
[0109]
Number
[0110] can be sent to the quantization process 108 via the network.
[0111]
[0037] In 208, the parameter estimation network 120 sends the time step t and the quantization setting
[0112]
Number
[0113] to the receiver 104. The diffusion model 116 can use the time step t to perform the number of iterations of the noise removal step on the reconstructed latent representation based on the value of the time step t. For example, if a value of 50 is received, the diffusion model 116 can perform 50 iterations of noise removal. The inverse quantization process 114 can use the quantization setting
[0114]
Number
[0115] as well.
[0116]
[0038] <Diffusion Model>
[0117]
[0039] FIG. 3 illustrates a simplified flowchart 300 of a method for performing a diffusion process according to some embodiments. At 302, the diffusion model 116 receives the time step t and the reconstructed latent representation
[0118]
Number
[0119] The time step t is estimated by the parameter estimation network 120.
[0120]
[0040] In 304, the diffusion model 116 performs a noise removal operation on the reconstructed latent representation
[0121]
Number
[0122] and executes a noise removal operation on it. As described above, the diffusion model 116 takes the reconstructed latent representation as input
[0123]
Number
[0124] and may receive the time step t. Next, the diffusion model 116 estimates the noise level of the reconstructed latent representation
[0125]
Number
[0126] and may predict the step before the forward process to remove some noise from the reconstructed latent representation
[0127]
Number
[0128] In 306, the diffusion model 116 outputs the denoised reconstructed latent representation
[0129]
Number
[0130] .
[0131]
[0041] In 308, it is determined whether the time step t is satisfied. For example, if there are 50 time steps, the diffusion model 116 may compare the 50 time steps with the current time step. If the time step t is not satisfied, the process iterates to 304, where another noise removal process is performed on the output of the diffusion model 116. Here, the noise-removed and reconstructed latent representation
[0132]
Number
[0133] is denoised again using the same process as described above. This process will continue until the time step t is satisfied. When the time step t is satisfied, in 310, the diffusion model 116 outputs the final noise-removed and reconstructed latent representation
[0134]
Number
[0135] . Here, the diffusion model 116 may have removed some of the noise introduced as quantization error by the quantization process.
[0136]
[0042] <Training>
[0137]
[0043] In some embodiments, diffusion model 116 can be used as a base model. The system can use stable diffusion as the base model, although other models can also be used. A base model (e.g., a diffusion model) trained with a large amount of data, etc., thus has excellent generative power for denoising the latent representation of an image. Here, the parameters of the base diffusion model may not be trained during the training process of parameter estimation network 120. Rather, the parameters of diffusion model 116 can be fixed during the training process. The training focuses on providing rate distortion optimization when further quantizing the latent representation created by the quantization process.
[0138]
[0044] FIG. 4 illustrates a simplified flowchart 400 of a method for performing a training process according to some embodiments. At 402, a training dataset of images can be input into the pipeline. The training dataset of images can be ground truth. At 404, the pipeline outputs a decoded image. The pipeline can process the image as described above. At 406, the decoded image can be compared to the ground truth of the original image to determine the difference between the decoded image and the original image. At 408, based on the difference, the parameters of parameter estimation network 120 can be adjusted to minimize the difference, such as by using a loss function. Here, the time step t and the quantization settings can be adjusted to minimize the loss. The pipeline can be trained to generate realistic images at a low bit rate by adjusting the parameters of the parameter estimation network. This can improve the use of the system since expensive training of diffusion model 116 may not be required. Also, the system can adjust the parameters of the entropy model of entropy encoding 110 or entropy decoding 112. Here, the system learns the best entropy model for the quantized data so that the quantized data can be more efficiently encoded into a bitstream.
[0139]
[0045] <Conclusion>
[0140]
[0046] Therefore, an irreversible image compression codec based on a latent diffusion model can be provided to create realistic image reconstructions from low bitrates to ultra-low bitrates. By combining the noise removal ability of the diffusion model with the inherent characteristics of quantization noise, the system predicts the ideal number of noise removal steps to create perceptually satisfactory reconstructions over a range of bitrates. By enabling the use of fewer bits with less quantization error in the quantization process, lower bitrates can be achieved. The error can be corrected by removing noise using the diffusion model 116. Training of the parameter estimation network 120 can generate quantization settings and time steps t that can create realistic images at lower bitrates.
[0141]
[0047] The system can also have a faster decoding time than other diffusion codecs by reusing the underlying diffusion model and having a lower training budget. Additionally, the system provides control over the rate-distortion tradeoff using the input settings to the parameter estimation network.
[0142]
[0048] <System>
[0143]
[0049] Figure 5 illustrates an example of a computing device according to some embodiments. According to various embodiments, a system 500 suitable for implementing the embodiments described herein includes a processor 501, a memory module 503, a storage device 505, an interface 511, and a bus 515 (e.g., a PCI bus or other interconnect fabric). System 500 may operate as various devices described herein, or any other device or service. Although a particular configuration is described, various alternative configurations are possible. Processor 501 may execute operations such as those described herein. Instructions for performing such operations may be embodied in memory 503, on one or more non-transitory computer-readable media, or on some other storage device. Various specially configured devices may also be used in place of, or in addition to, processor 501. Memory 503 may be a random access memory (RAM) or other dynamic storage device. Storage device 505 may include a non-transitory computer-readable storage medium that holds information, instructions, or some combination thereof, such as instructions that, when executed by processor 501, cause processor 501 to be configured to or capable of performing one or more operations of the methods described herein. Bus 515 or other communication component may support the communication of information within system 500. Interface 511 may be connected to bus 515 and may be configured to transmit and receive data packets over a network. Examples of supported interfaces include, but are not limited to: Ethernet®, Fast Ethernet, Gigabit Ethernet, Frame Relay, Cable, Digital Subscriber Line (DSL), Token Ring, Asynchronous Transfer Mode (ATM), High-Speed Serial Interface (HSSI), and Fiber Distributed Data Interface (FDDI). These interfaces may include ports suitable for communication with appropriate media. They may also include independent processors and / or volatile RAM.A computer system or computing device may include, or communicate with, a monitor, printer, or other suitable display for providing any of the results referred to herein to a user.
[0144]
[0050] Any of the disclosed implementations may be embodied in various types of hardware, software, firmware, computer-readable media, and combinations thereof. For example, some of the techniques disclosed herein may be implemented, at least in part, by a non-transitory computer-readable medium that includes program instructions, state information, etc. for configuring a computing system to perform the various services and operations described herein. Examples of program instructions include both machine code such as that generated by a compiler and higher-level code that may be executed via an interpreter. The instructions may be embodied in any suitable language, such as Java (registered trademark), Python, C++, C, HTML, any other markup language, JavaScript (registered trademark), ActiveX, VBScript, or Perl. Examples of non-transitory computer-readable media include, but are not limited to, magnetic media such as hard disks and magnetic tapes, optical media such as flash memory, compact discs (CDs) or digital versatile discs (DVDs), magneto-optical media, and other hardware devices such as read-only memory ("ROM") devices and random access memory ("RAM") devices. The non-transitory computer-readable media may be any combination of such storage devices.
[0145]
[0051] In the foregoing specification, various techniques and mechanisms may be described in the singular for clarity. However, it should be noted that some embodiments include multiple repetitions of a technique or multiple instantiations of a mechanism unless otherwise stated. For example, a system may use a processor in various contexts and, unless otherwise stated, may use multiple processors while remaining within the scope of this disclosure. Similarly, various techniques and mechanisms may be described as including a connection between two entities. However, since there may be various other entities (e.g., bridges, controllers, gateways, etc.) between two entities, a connection does not necessarily mean a direct, unobstructed connection.
[0146]
[0052] Some embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform the methods described by some embodiments. The computer system may include one or more computing devices. The instructions, when executed by one or more computer processors, may be configured to perform or be operable to perform what is described in some embodiments.
[0147]
[0053] As used in the description herein and throughout the following claims, "a", "an", and "the" include references to the plural unless the context clearly dictates otherwise. Also, as used in the description herein and throughout the following claims, the meaning of "in" includes "in" and "on" unless the context clearly dictates otherwise.
[0148] [
[0054] ]The above description has illustrated various embodiments, along with examples of how aspects of some embodiments can be implemented. The above examples and embodiments should not be regarded as the only embodiments, but are presented to illustrate the flexibility and advantages of some embodiments as defined by the following claims. Based on the above disclosure and the following claims, other arrangements, embodiments, implementations, and equivalents can be used without departing from the scope of the invention as defined by the claims.
Claims
1. receiving a quantized latent representation of an image in a latent space, where the image is encoded into a latent representation in the latent space and quantized to generate the quantized latent representation; receiving time step parameters generated based on the latent representation; performing an inverse quantization process to generate a reconstructed latent representation; performing a denoising process for a number of iterations based on the time step parameter to remove noise from the reconstructed latent representation using a diffusion model to generate a denoised reconstructed latent representation; decoding the denoised reconstructed latent representation into a reconstructed image; and A method for providing the above.
2. The quantized latent representation is entropy coded; the quantized latent representation is entropy decoded prior to performing the inverse quantization process. The method of claim 1.
3. receiving a quantization setting generated based on the latent representation in the latent space; performing the inverse quantization process using the quantization settings; The method of claim 1 further comprising:
4. The method of claim 3 , wherein the quantization settings are used by the inverse quantization process to adjust a number of bits used for quantization to generate the reconstructed latent representation.
5. The method of claim 3 , wherein the quantization settings are generated based on the latent representation of the image and input parameters, the input parameters being based on a rate-distortion balance.
6. The method of claim 1 , wherein the time step parameters are generated based on the latent representation of the image and input parameters, the input parameters being based on a rate-distortion balance.
7. a quantization process and a dequantization process add a quantization error to the reconstructed latent representation compared to the latent representation of the image; the diffusion model is configured to remove noise associated with the quantization error. The method of claim 1.
8. quantization error from generating the quantized latent representation adds noise to the reconstructed latent representation; The diffusion model denoises the reconstructed latent representation to remove noise from the reconstructed latent representation. The method of claim 1.
9. The method of claim 1 , wherein a network is trained to generate the time-step parameters.
10. The method of claim 9 , wherein the diffusion model is not trained during training of the network.
11. 10. The method of claim 9, wherein the network is trained to generate quantization settings used by the inverse quantization process to adjust a number of bits used for quantization to generate the reconstructed latent representation.
12. A non-transitory computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a computing device, cause the computing device to: receiving a quantized latent representation of an image in a latent space, where the image is encoded into a latent representation in the latent space and quantized to generate the quantized latent representation; receiving time step parameters generated based on the latent representation; performing an inverse quantization process to generate a reconstructed latent representation; performing a denoising process for a number of iterations based on the time step parameter to remove noise from the reconstructed latent representation using a diffusion model to generate a denoised reconstructed latent representation; decoding the denoised reconstructed latent representation into a reconstructed image; and 23. A non-transitory computer-readable storage medium operable to:
13. Receiving an image; and encoding the image into a latent representation in a latent space; estimating time step parameters based on the latent representation; and performing a quantization process on the latent representation to generate a quantized latent representation; transmitting the quantized latent representation to a receiver, where an inverse quantization process is performed to generate a reconstructed latent representation, and a diffusion model performs a denoising process over a number of iterations based on the time step parameter to remove noise from the reconstructed latent representation; A method for providing the above.
14. The quantized latent representation is entropy coded; the quantized latent representation is entropy decoded prior to performing the inverse quantization process. The method of claim 13.
15. determining a quantization setting based on the latent representation in the latent space; and performing the quantization process using the quantization settings; and The method of claim 13 further comprising:
16. The method of claim 15 , wherein the quantization settings are generated based on the latent representation of the image and input parameters, the input parameters being based on a rate-distortion tradeoff.
17. estimating the time step parameters estimating the time step parameters based on the latent representation of the image and input parameters, where the input parameters are based on a rate-distortion tradeoff. The method of claim 13.
18. quantization error from the quantization process adds noise to the reconstructed latent representation; The diffusion model denoises the reconstructed latent representation to remove noise from the reconstructed latent representation. The method of claim 13.
19. The method of claim 13 , wherein a network is trained to generate the time-step parameters.
20. 10. The method of claim 9, wherein the network is trained to generate quantization settings used by the quantization process to adjust a number of bits used for quantization to generate the quantized latent representation.
Citation Information
Patent Citations
Training method, image coding method, image decoding method, and device
JP2021150955A
Method, apparatus and computer program for video encoding
JP2023528176A
Removing spatially varying noise from images using diffusion
JP2025525780A
Method and data processing system for lossy image or video encoding, transmission and decoding
WO2023118317A1
Cited By
Communication device and method in communication device
JP7874810B1