Method for training diffusion model and recovering image quality
By using the visual coding features of noise as a generation condition, the diffusion model is trained to restore image quality, solving the problem of low efficiency in the prior art, and achieving an efficient and high-quality image denoising effect.
Patent Information
- Application Number
- CN202510347528.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, denoising models need to be trained separately for different types of noise, resulting in low efficiency and poor effect of image quality recovery. Especially when the noise type is unknown in low-quality images, it is difficult to efficiently restore image quality.
By training the diffusion model, the visual encoding features of noise are used as a generation condition, and the diffusion model is guided to denoise. The pre-trained encoder is used to obtain the visual encoding features of noise, and image quality recovery is performed using this as a priori knowledge to avoid training a separate model for each type of noise.
Efficient and high-quality image quality recovery in low-quality images is achieved, improving training efficiency and recovery effects without multiple denoising and training models for each noise.
Smart Images

Figure CN120374424A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to a method for training a diffusion model and image quality restoration. Background Art
[0002] In today's era of information explosion, images have become an important carrier for information transmission. However, due to limitations in shooting equipment, transmission processes, or storage conditions, image quality often faces degradation problems. Image quality degradation is ultimately due to the noise added to the original image, and there are various types of noise, such as rain, fog, night, motion blur, white noise, etc. Image quality restoration is essentially a process of denoising the image.
[0003] However, the noise contained in an image often consists of a superposition of multiple types of noise. For different types of noise, different denoising methods are often required for denoising. In the prior art, for different types of noise, different denoising models also need to be trained separately, which results in low training efficiency for the model used for image quality restoration and low efficiency of image quality restoration. Summary of the Invention
[0004] Embodiments of this specification provide a method, device, storage medium, and electronic device for training a diffusion model and image quality restoration to partially solve the problems existing in the above prior art.
[0005] Embodiments of this specification adopt the following technical solutions:
[0006] A method for training a diffusion model provided in this specification, the method comprising:
[0007] Determine a first sample noise according to each preset type of noise;
[0008] Obtain a visual coding feature corresponding to the first sample noise through a pre-trained encoder, and obtain a first sample image. Add the first sample noise to the first sample image to obtain a first noisy image;
[0009] Input a standard noise image into the diffusion model, and inject the visual coding feature and the first noisy image as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image;
[0010] Determine a noise prediction loss according to the first sample image and the restored image;
[0011] Taking reducing the noise prediction loss as a training objective, adjust the model parameters of the diffusion model.
[0012] A method for image quality restoration provided in this specification, the method comprising:
[0013] Obtain the image to be processed;
[0014] According to the image to be processed, select at least one type of noise from each preset type of noise;
[0015] Determine the fourth sample noise according to the at least one type of selected noise;
[0016] Obtain the visual coding feature corresponding to the fourth sample noise;
[0017] Input the standard noise image into a pre-trained diffusion model, and inject the visual coding feature and the image to be processed as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image, wherein the diffusion model is pre-trained by the method of training the diffusion model described above.
[0018] An apparatus for training a diffusion model provided in this specification, the apparatus comprising:
[0019] A preprocessing module, configured to determine the first sample noise according to each preset type of noise; obtain the visual coding feature corresponding to the first sample noise through a pre-trained encoder, and obtain a first sample image, and add the first sample noise to the first sample image to obtain a first noise-added image;
[0020] A denoising module, configured to input the standard noise image into the diffusion model, and inject the visual coding feature and the first noise-added image as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image;
[0021] A loss determination module, configured to determine a noise prediction loss according to the first sample image and the restored image;
[0022] A parameter adjustment module, configured to adjust the model parameters of the diffusion model with the goal of reducing the noise prediction loss.
[0023] An apparatus for image quality restoration provided in this specification, the apparatus comprising:
[0024] An acquisition module, configured to acquire the image to be processed;
[0025] A selection module, configured to select at least one type of noise from each preset type of noise according to the image to be processed;
[0026] A noise determination module, configured to determine a fourth sample noise according to at least one selected type of noise;
[0027] A condition acquisition module, configured to obtain a visual coding feature corresponding to the fourth sample noise;
[0028] A denoising module, configured to input a standard noise image into a pre-trained diffusion model, and inject the visual coding feature and the image to be processed as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image, wherein the diffusion model is pre-trained by the method of training the diffusion model described above.
[0029] A computer-readable storage medium provided in this specification, where the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method of training the diffusion model and / or the method of image quality restoration described above.
[0030] An electronic device provided in this specification, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method of training the diffusion model and / or the method of image quality restoration described above.
[0031] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0032] The embodiments of this specification disclose a method for training a diffusion model. The method injects a first noise-added image obtained by adding a first sample noise to a first sample image and a visual coding feature of the first sample noise into the diffusion model as generation conditions, so that the diffusion model learns that the noise in the first noise-added image is the noise corresponding to the visual coding feature, and uses this as prior knowledge. Under the guidance of this prior knowledge, the diffusion model learns how to denoise a standard noise image to obtain the first sample image before noise addition. Through the diffusion model trained by this method, even if the noise in the low-quality image is composed of a superposition of multiple types of noise, the diffusion model can still guide the diffusion model to denoise according to the visual coding feature of the noise as one of the generation conditions, without the need to use multiple denoising models to denoise the low-quality image multiple times, nor to train a denoising model for each type of noise, which can effectively improve the efficiency of training the denoising model and image quality restoration of the low-quality image. Description of the Drawings
[0033] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0034] Figure 1 A flow chart of a method for training a diffusion model provided in an embodiment of this specification;
[0035] Figure 2 A schematic diagram of the training process of the diffusion model provided in the embodiment of this specification;
[0036] Figure 3 A flow chart of a method for image quality restoration provided in an embodiment of this specification;
[0037] Figure 4 A schematic diagram of a device for training a diffusion model provided in an embodiment of this specification;
[0038] Figure 5 A schematic diagram of a device for restoring image quality provided in an embodiment of this specification;
[0039] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0040] Since the prior art often requires training corresponding denoising models for different types of noise such as rain, fog, night, motion blur, white noise, etc., when restoring the image quality of a low-quality image, it is often necessary to use a denoising model corresponding to each type of noise to perform multiple denoising on the low-quality image. Moreover, it is usually not known in advance which types of noise the noise in the low-quality image is superimposed on, so the image quality restoration method in the prior art is not only inefficient, but also has poor results.
[0041] The embodiments of this specification are intended to train a diffusion model for image quality restoration, and guide the diffusion model to denoise by taking the visual encoding features of noise as at least one of the generation conditions, thereby achieving efficient and high-quality image quality restoration.
[0042] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0043] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0044] Figure 1 The method flow chart of training diffusion model provided in the embodiment of this specification includes the following steps:
[0045] S100: Determine the first sample noise according to each preset type of noise.
[0046] In the embodiments of this specification, the diffusion model to be trained includes a denoising network and an adaptation network. As the core, the denoising network is used to denoise the subsequent input standard noise image, and the detailed process will be described later. The adaptation network is used to adapt to different noises, and based on different noises, different guiding controls are performed on the denoising network to enable the denoising network to perform different denoising.
[0047] It should be emphasized that different types of noises and different noises described in this specification are not the same concept. Different types of noises refer to different types of basic noises such as rain, fog, night, motion blur, white noise, etc. Different noises can include, in addition to different types of basic noises: basic noises of the same type but different intensities, noises with incomplete identical types of included basic noises, and noises with the same type of included basic noises but incomplete identical intensities.
[0048] That is to say, the adaptation network included in the diffusion model in the embodiments of this specification can adapt to different types of single basic noises, or can also adapt to noises formed by superimposing multiple different intensities and different types of basic noises.
[0049] When training this diffusion model, the first sample noise needs to be obtained first. Since when using the trained diffusion model for denoising, the noise in the low-quality image often consists of multiple different types of basic noises superimposed, therefore, the first sample noise in the embodiments of this specification can be a single type of basic noise, or can also be formed by superimposing multiple different types of basic noises. That is, at least one type of noise can be selected from each preset type of noise (basic noise), and then the selected types of noises are superimposed to obtain the first sample noise. When superimposing, the intensity of each selected type of noise can be set.
[0050] S102: Obtain the visual coding feature corresponding to the first sample noise through a pre-trained encoder, obtain a first sample image, and add the first sample noise to the first sample image to obtain a first noise-added image.
[0051] In the embodiments of this specification, Figure 1 The steps S102 shown can synchronously execute the process of obtaining the visual coding feature corresponding to the first sample noise and the process of obtaining the first noise-added image, and the execution order of the two is not sequential.
[0052] When obtaining the visual coding features corresponding to the first sample noise, the noise map corresponding to the first sample noise can be directly visually coded by a pre-trained encoder to obtain the visual coding features corresponding to the first sample noise. In order to further improve the accuracy of the visual coding features corresponding to the obtained first sample noise, in the embodiments of this specification, a second sample image can be first obtained, and the first sample noise can also be added to the second sample image to obtain a second noisy image. Then, the second noisy image is input into the pre-trained encoder to obtain the image features of the second noisy image output by the encoder. Finally, the visual coding features corresponding to the first sample noise are extracted from the image features of the second noisy image.
[0053] Since the image features of the second noisy image contain both the image features of the second sample image and the visual coding features of the first sample noise, when extracting the visual coding features corresponding to the first sample noise from the image features of the second noisy image, the image features of the second sample image can be removed from the image features of the second noisy image to obtain the visual coding features corresponding to the first sample noise. Specifically, the second sample image can also be input into the encoder to obtain the image features of the second sample image output by the encoder, and then the image features of the second sample image are removed from the image features of the second noisy image to obtain the visual coding features corresponding to the first sample noise.
[0054] As for the training method of the encoder, it will be described in detail later.
[0055] S104: Input the standard noise image into the diffusion model, and inject the visual coding features and the first noisy image as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image.
[0056] In the embodiments of this specification, the diffusion model can select a diffusion model whose core denoising network has been trained, such as a diffusion model based on U-Net, including DDPM or DDIM, etc., all of which are diffusion models with U-Net as the denoising network.
[0057] The output of the denoising network of the diffusion model can be used as the input of the adaptation network. Then, the standard noise image can be directly input into the denoising network, and at the same time, the visual coding features of the first sample noise and the first noisy image are injected into the diffusion model as generation conditions. Among them, the standard noise image described in the embodiments of this specification can be a random noise image.
[0058] According to the above generation conditions, the diffusion model can obtain the following two kinds of information as prior knowledge:
[0059] 1. The task of the diffusion model is to denoise a standard noise image to obtain the image before the first noisy image was noise-added (i.e., the first sample image), which is used as the restored image.
[0060] 2. The noise added to the first noisy image is the noise corresponding to the injected visual coding features mentioned above.
[0061] Based on the above prior knowledge, the diffusion model can denoise the standard noise image through the denoising network, and then input the output result of the denoising network into the adaptation network to achieve a single denoising process for the standard noise image.
[0062] In practical applications, the diffusion model needs to perform the above denoising process multiple times to obtain the restored image. The specific number of times to perform the above denoising process can be set as needed.
[0063] However, experimental tests have found that directly injecting the first noisy image as a generation condition into the denoising network does not achieve good denoising effects. Therefore, in the embodiments of this specification, in addition to the denoising network and the adaptation network, the diffusion model may also include an image restoration bridging network. The input of the image restoration bridging network is the first noisy image.
[0064] Then, in step S104, the first noisy image can be input into the image restoration bridging network to obtain the image features corresponding to the first noisy image output by the image restoration bridging network. Inject the image features corresponding to the first noisy image into the denoising network, and at the same time inject the visual coding features of the first sample noise as a generation condition into the adaptation network, so that the denoising network of the diffusion model denoises the standard noise image according to the injected image features corresponding to the first noisy image, and then input the output result of the denoising network into the adaptation network. The adaptation network then processes the output result of the denoising network according to the injected visual coding features of the first sample noise to achieve a single denoising process for the standard noise image. Similarly, in practical applications, the diffusion model also needs to perform this denoising process multiple times to obtain the restored image.
[0065] When the denoising network is a U-Net, since the U-Net contains denoising processing layers corresponding to multiple resolutions, after the first noisy image is input into the image restoration bridging network, the image restoration bridging network needs to encode the first noisy image into image features of multiple different resolutions. Then, for each image feature of a different resolution corresponding to the first noisy image, inject the image feature of this resolution into the denoising processing layer corresponding to this resolution in the U-Net. Finally, input the output results of the denoising processing layers corresponding to each resolution in the U-Net into the adaptation network.
[0066] Further, when the denoising network is U-Net, to improve the accuracy of training the diffusion model, the image features of each resolution corresponding to the first denoised image output by the image restoration bridging network can be injected into the denoising processing layer of the corresponding resolution in the U-Net through the AdaLN-Zero input mechanism.
[0067] Among them, AdaLN-Zero is a conditional injection mechanism for neural networks, which combines the techniques of Adaptive Layer Normalization (AdaLN) and Zero Initialization. Layer Normalization (LN) is a normalization technique used to stabilize the training process of neural networks. It normalizes the input of each layer so that its mean is 0 and variance is 1. Adaptive Layer Normalization is an extended version of LN, which dynamically adjusts the normalization parameters (such as scaling factors and offsets) by introducing additional information (such as feature vectors). This enables the model to flexibly adjust the feature distribution of each layer according to the injected generation conditions. Zero Initialization means that during model initialization, some parameters (such as the weights corresponding to conditional injection) are initialized to zero. The purpose of this is to reduce the impact of the injected conditions on the model in the initial stage of training, thereby avoiding the model prematurely relying on the injected conditions and resulting in unstable training. That is to say, in the embodiments of this specification, first, the image features of multiple different resolutions corresponding to the first denoised image output by the image restoration bridging network are normalized through Adaptive Layer Normalization, then the injection weights corresponding to the multiple normalized image features of different resolutions are set to 0 through Zero Initialization respectively, then they are injected into the denoising processing layer corresponding to the corresponding resolution in the denoising network respectively, and finally the output results of the denoising processing layer corresponding to each resolution are input into the adaptation network, and the adaptation network processes the output results of the denoising network according to the visual coding features of the injected first sample noise to complete a denoising process. After repeating this denoising process multiple times, the restored image is obtained.
[0068] It should be noted that in the embodiments of this specification, the visual coding features of the first sample noise are used as one of the generation conditions and injected into the diffusion model because the object processed by the diffusion model is an image. If the text features of the text used to describe the first sample noise are used as the generation conditions, it will affect the accuracy of the diffusion model in restoring the image quality, and the effect is poor. For example, assume that the first sample noise is a noise composed of the superposition of motion blur noise and white noise. Then, indeed, the text features of the prompt text "motion blur noise superposed with white noise" can also be used as one of the generation conditions and injected into the diffusion model. However, after all, there are differences between such text features and the visual coding features of the first sample noise. Therefore, for the diffusion model used to process images, using the visual coding features of the first sample noise as the generation condition is more accurate and direct, and can improve the accuracy of image quality restoration.
[0069] S106: Determine the noise prediction loss according to the first sample image and the restored image.
[0070] In the embodiments of this specification, a standard noise prediction loss can be used to train the diffusion model. Specifically, the process of adding noise (the noise here is not the first sample noise) to the first sample image to obtain a standard noise image can be regarded as a process of adding noise to the first sample image multiple times (the number of times of adding noise is the same as the number of times of repeatedly performing the denoising process above) based on a Markov model to obtain a standard noise image. Therefore, according to the first sample image and the standard noise image, a Markov model can be used to fit the noise addition process to estimate the noise distribution of the noise added to the first sample image each time of adding noise, and then according to the standard noise image and the restored image, estimate the noise distribution of the noise removed from the standard noise image each time of performing the denoising process. According to the noise distribution of the noise added to the first sample image each time of adding noise and the noise distribution of the noise removed from the standard noise image each time of performing the denoising process, determine the noise prediction loss.
[0071] Furthermore, since the denoising process of the diffusion model for the standard noise image is the inverse process of the noise addition process fitted by the Markov model, assume that a total of N times of noise addition and N times of denoising are required. Then, the noise prediction loss can be determined according to the estimated noise distribution of the noise added to the first sample image at the nth time of adding noise and the noise distribution of the noise removed from the standard noise image at the (N - n + 1)th time of denoising, where 1 ≤ n ≤ N. The greater the difference between the noise distribution of the noise added to the first sample image at the nth time of adding noise and the noise distribution of the noise removed from the standard noise image at the (N - n + 1)th time of denoising, the greater the noise prediction loss, and vice versa, the smaller the noise prediction loss.
[0072] S108: Adjust the model parameters of the diffusion model with the goal of reducing the noise prediction loss.
[0073] Since in the embodiments of this specification, the diffusion model needs to learn how to control the denoising network of the guided diffusion model through the adaptation network to denoise the standard noise image so as to obtain a restored image as identical as possible to the original first sample image, therefore, in the embodiments of this specification, it is necessary to use reducing the noise prediction loss determined in step S106 as the training objective to adjust the model parameters of the adaptation network.
[0074] Furthermore, when the diffusion model further includes an image restoration bridging network, it is also necessary to adjust the model parameters of the image restoration bridging network. For the already trained denoising network, the model parameters of the denoising network remain unchanged.
[0075] Figure 2 is a schematic diagram of the training process of the diffusion model provided by the embodiments of this specification. In Figure 2 , a first noise-added image is obtained by adding first sample noise to the first sample image and is injected as one of the generation conditions into the image restoration bridging network of the diffusion model. The visual coding feature of the first sample noise is injected as the second generation condition into the adaptation network. The image feature output by the image restoration bridging network is injected into the core denoising network. The denoising network denoises the standard noise image, and the output result of the denoising network is input into the adaptation network for processing. After multiple denoising operations, a restored image is obtained.
[0076] Furthermore, since in the above step S102, it is necessary to use a pre-trained encoder to encode the second noise-added image with the first sample noise added thereto to obtain the image feature of the second noise-added image, and it is also necessary to use this encoder to encode the second sample image without the first sample noise added thereto to obtain the image feature of the second sample image, and the visual coding feature corresponding to the first sample noise is obtained by subtracting the image feature of the second sample image from the image feature of the second noise-added image, therefore, in the embodiments of this specification, it is necessary to train an encoder that can accurately encode both noise and the original image.
[0077] Specifically, when training the encoder, the codec model to be trained and the third sample image can be obtained. The codec model includes an encoder and a decoder. Then, according to each type of preset noise, the second sample noise is determined, and the second sample noise is added to the third sample image to obtain a third noisy image. The third noisy image is input into the encoder in the codec model to be trained, and the image features of the third noisy image output by the encoder in the codec model to be trained are obtained. The image features of the third noisy image are input into the decoder in the codec model to be trained, and the reconstructed noise output by the decoder in the codec model to be trained is obtained. Finally, according to the second sample noise and the reconstructed noise, the reconstruction loss is determined, and according to the reconstruction loss, at least the model parameters of the encoder in the codec model to be trained are adjusted, and the encoder after adjusting the model parameters is used as the pre-trained encoder.
[0078] Among them, the decoder in the above codec model is used to reconstruct the noise. In an ideal situation, since the image features obtained by encoding the third noisy image by the encoder should include both the image features of the third sample image and the visual coding features of the second sample noise, the encoder can recover the second sample noise, that is, the reconstructed noise, according to the visual coding features therein. Thus, in a supervised learning manner, using the second sample noise as a supervision signal, the difference between the second sample noise and the reconstructed noise can be determined. The greater this difference, the greater the reconstruction loss, and vice versa, the smaller the reconstruction loss. With the minimization of the reconstruction loss as the training objective, at least the model parameters of the encoder are adjusted, and the trained encoder can be obtained.
[0079] The diffusion model trained by the above method can adapt to various different noises through its adaptation network. As long as the low-quality image and the visual coding features of the noise contained in the low-quality image are used as generation conditions and injected into the diffusion model, the diffusion model can know that the high-quality image corresponding to the low-quality image needs to be generated, and the noise contained in the low-quality image is the noise corresponding to the visual coding features, and thus generate the high-quality image corresponding to the low-quality image to achieve image quality restoration of the low-quality image, without training a denoising model for each type of basic noise. Furthermore, the denoising model can be trained efficiently and with high quality, and the image quality of the low-quality image can also be restored efficiently and with high quality.
[0080] However, since the diffusion model trained by the above method needs to inject the visual coding features of the noise contained in the low-quality image as one of the generation conditions, but in practical applications, it is often unknown which types of basic noises are superimposed to form the noise contained in the low-quality image. Therefore, after training the above diffusion model in the embodiments of this specification, the following can be adopted Figure 3The method shown performs image quality restoration on low-quality images.
[0081] Figure 3 The following is a flowchart of a method for image quality restoration provided by an embodiment of this specification, including the following steps:
[0082] S300: Obtain the image to be processed.
[0083] In the embodiment of this specification, the image to be processed is a low-quality image containing noise.
[0084] S302: According to the image to be processed, select at least one type of noise from each type of preset noise.
[0085] S304: Determine the fourth sample noise according to the at least one type of selected noise.
[0086] Since the components of the noise contained in the low-quality image in practical applications are unknown, in step S302, a standard image without any noise can be obtained, and according to the noise visual effect of the image to be processed, at least one type of noise is selected from each type of preset noise, and then the selected types of noise are added to the standard image. As long as the noise visual effect of the standard image after adding noise is similar to the noise visual effect of the image to be processed, it means that the noise contained in the standard image after adding noise at this time is similar to the noise contained in the image to be processed. Therefore, the visual coding features of the noise added to the standard image can be directly used as generation conditions and injected into the diffusion model.
[0087] Therefore, after adding the selected types of noise to the standard image to obtain the standard image after adding noise, the similarity between the noise visual effect of the image to be processed and the noise visual effect of the standard image after adding noise can be evaluated. If the similarity is higher than the set similarity, the types of noise currently added to the standard image are used as the selected types of noise, and the selected types of noise are superimposed accordingly to determine the fourth sample noise, that is, step S304 is executed. Otherwise, continue to select at least one type of noise from each type of preset noise, and continue to add the selected noise to the standard image until the similarity between the noise visual effect of the standard image after adding noise and the noise visual effect of the image to be processed is higher than the set similarity.
[0088] Specifically, when evaluating the similarity between the noise visual effect of the image to be processed and the noise visual effect of the standard image after adding noise, a pre-trained machine learning model can be used for evaluation. For example, the above-mentioned encoding and decoding model can be directly used. The image to be processed and the standard image after adding noise are respectively input into the encoding and decoding model to obtain the noise reconstructed by the encoding and decoding model for the image to be processed and the standard image after adding noise respectively, and then the similarity between the two noises is determined.
[0089] Of course, it can also be directly evaluated manually by observation. Specifically, after obtaining the standard image, at least one type of noise can be manually selected from each preset type of noise and added to the standard image, and then the visual effect of the noise on the standard image after adding noise and the visual effect of the noise on the image to be processed are observed manually to see if they are similar. If they are similar, it can be considered that the similarity of the noise visual effects of the two is higher than the set similarity. If they are not similar, continue to manually select and add noise.
[0090] S306: Obtain the visual coding features corresponding to the fourth sample noise.
[0091] In the embodiments of this specification, the visual coding features corresponding to the fourth sample noise can also be obtained by using the above-mentioned pre-trained encoder. The process is exactly the same as the method for obtaining the visual coding features corresponding to the first sample noise by this encoder in the Figure 1 shown training process, and will not be elaborated here one by one.
[0092] S308: Input the standard noise image into the pre-trained diffusion model, and inject the visual coding features and the image to be processed as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image.
[0093] Similarly, in step S308, the image to be processed can be used as one of the generation conditions and injected into the image restoration bridging network of the diffusion model. The image features of the image to be processed are injected into the denoising network through the image restoration bridging network. The visual coding features of the fourth sample noise are used as the second generation condition and injected into the adaptation network of the diffusion model, so that the diffusion model denoises the standard noise image through the denoising network that injects the image features of the image to be processed, and inputs the output result after denoising into the adaptation network that injects the visual coding features of the fourth sample noise for processing, completing a denoising process for the standard noise image. Repeatedly execute this denoising process to obtain a restored image.
[0094] In addition, in the embodiments of this specification, the diffusion model may also include multiple adaptation networks corresponding to different noises. The components (i.e., the types of basic noises included) and intensities of the noises corresponding to different adaptation networks are not exactly the same. Therefore, the diffusion model may also include model parameters for selecting adaptation networks.
[0095] Correspondingly, when training the diffusion model, the diffusion model can, based on the visual coding features of the first noise-added image injected as the first generation condition and the first sample noise injected as the second generation condition, select at least one adaptation network from multiple adaptation networks according to the model parameters for selecting the adaptation network, and inject the visual coding features of the first sample noise into the at least one selected adaptation network. After injecting the visual coding features of the first sample noise, the selected adaptation network can process the output result input to the denoising network, and fuse the processing results of the selected adaptation network to obtain a restored image. Then, according to the first sample image and the restored image, a noise prediction loss is determined, and with the goal of reducing the noise prediction loss, the model parameters of the image restoration bridging network, the model parameters of the selected adaptation network, and the model parameters for selecting the adaptation network in the diffusion model are adjusted, so that the diffusion model can not only learn how to control and guide the denoising network of the diffusion model to denoise a standard noise image through the adaptation network to obtain a restored image as similar as possible to the original first sample image, but also learn how to select a suitable adaptation network.
[0096] The above is a method for training a diffusion model and a method for image quality restoration provided by an embodiment of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.
[0097] Figure 4 The following is a schematic diagram of a device for training a diffusion model provided by an embodiment of this specification. The device includes:
[0098] A preprocessing module 401, configured to determine first sample noise according to each preset type of noise; obtain the visual coding features corresponding to the first sample noise through a pre-trained encoder, and acquire a first sample image, and add the first sample noise to the first sample image to obtain a first noise-added image;
[0099] A denoising module 402, configured to input a standard noise image into the diffusion model, and inject the visual coding features and the first noise-added image into the diffusion model as generation conditions, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image;
[0100] A loss determination module 403, configured to determine a noise prediction loss according to the first sample image and the restored image;
[0101] A parameter adjustment module 404, configured to adjust the model parameters of the diffusion model with the goal of reducing the noise prediction loss.
[0102] Optionally, the preprocessing module 401 is specifically configured to select at least one type of noise from each preset type of noise; superimpose the selected types of noise to obtain a first sample noise.
[0103] Optionally, the preprocessing module 401 is specifically configured to obtain a second sample image; add the first sample noise to the second sample image to obtain a second noise-added image; input the second noise-added image into a pre-trained encoder to obtain the image features of the second noise-added image output by the encoder; extract the visual coding features corresponding to the first sample noise from the image features of the second noise-added image.
[0104] Optionally, the preprocessing module 401 is specifically configured to input the second sample image into a pre-trained encoder to obtain the image features of the second sample image output by the encoder; remove the image features of the second sample image from the image features of the second noise-added image to obtain the visual coding features corresponding to the first sample noise.
[0105] Optionally, the device further includes:
[0106] A training module 405, specifically configured to obtain a codec model to be trained and a third sample image, where the codec model includes an encoder and a decoder; determine a second sample noise according to each preset type of noise; add the second sample noise to the third sample image to obtain a third noise-added image; input the third noise-added image into the encoder in the codec model to be trained to obtain the image features of the third noise-added image output by the encoder in the codec model to be trained; input the image features of the third noise-added image into the decoder in the codec model to be trained to obtain the reconstructed noise output by the decoder in the codec model to be trained; determine a reconstruction loss according to the second sample noise and the reconstructed noise; at least adjust the model parameters of the encoder in the codec model to be trained according to the reconstruction loss, and use the encoder with adjusted model parameters as the pre-trained encoder.
[0107] Optionally, the diffusion model includes an image restoration bridging network, a denoising network, and an adaptation network;
[0108] The denoising module 402 is specifically configured to input the first noise-added image into the image restoration bridging network to obtain multiple image features with different resolutions corresponding to the first noise-added image output by the image restoration bridging network. For each image feature with a different resolution corresponding to the first noise-added image, inject the image feature at this resolution into the denoising processing layer corresponding to this resolution in the denoising network; and inject the visual coding feature into the adaptation network to enable the adaptation network to process the output results of the denoising processing layers corresponding to each resolution in the denoising network.
[0109] Optionally, the parameter adjustment module 404 is specifically configured to adjust the model parameters of the image restoration bridging network and the adaptation network with the goal of reducing the noise prediction loss.
[0110] Figure 5 The following is a schematic diagram of an apparatus for image quality restoration provided by an embodiment of this specification. The apparatus includes:
[0111] An acquisition module 501, configured to acquire an image to be processed;
[0112] A selection module 502, configured to select at least one type of noise from each preset type of noise according to the image to be processed;
[0113] A noise determination module 503, configured to determine a fourth sample noise according to the at least one type of noise selected;
[0114] A condition acquisition module 504, configured to obtain a visual coding feature corresponding to the fourth sample noise;
[0115] A denoising module 505, configured to input a standard noise image into a pre-trained diffusion model, and inject the visual coding feature and the image to be processed as generation conditions into the diffusion model, so that the diffusion model performs denoising on the standard noise image based on the generation conditions to obtain a restored image, where the diffusion model is pre-trained by the method of training the diffusion model described above.
[0116] Optionally, the selection module 502 is configured to obtain a standard image; select at least one type of noise from each type of preset noise according to the noise visual effect of the image to be processed; add the selected types of noise to the standard image; evaluate the similarity between the noise visual effect of the image to be processed and the noise visual effect of the standard image after adding noise; if the similarity is higher than the set similarity, then use the types of noise currently added to the standard image as the selected types of noise, otherwise, continue to select at least one type of noise from each type of preset noise, and continue to add the selected noise to the standard image until the similarity between the noise visual effect of the standard image after adding noise and the noise visual effect of the image to be processed is higher than the set similarity.
[0117] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to execute the above-provided method for training a diffusion model and / or the method for image quality restoration.
[0118] Based on Figure 1 and Figure 3 the method for training a diffusion model and the method for image quality restoration shown, the embodiments of this specification also provide Figure 6 the structural schematic diagram of the electronic device shown. As Figure 6 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above method for training a diffusion model and / or the method for image quality restoration.
[0119] The above are only the embodiments of this specification and are not used to limit this specification. For those skilled in the art, this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for training a diffusion model, the method comprising: Determining first sample noise according to each preset type of noise; Obtaining visual encoding features corresponding to the first sample noise through a pre-trained encoder, acquiring a first sample image, adding the first sample noise to the first sample image to obtain a first noise-added image; Inputting a standard noise image into the diffusion model, and injecting the visual encoding features and the first noise-added image as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image; Determining a noise prediction loss according to the first sample image and the restored image; Adjusting model parameters of the diffusion model with the goal of reducing the noise prediction loss.
2. The method according to claim 1, wherein determining first sample noise according to each preset type of noise specifically comprises: Selecting at least one type of noise from each preset type of noise; Superimposing the selected types of noise to obtain first sample noise.
3. The method according to claim 1, wherein obtaining visual encoding features corresponding to the first sample noise through a pre-trained encoder specifically comprises: Acquiring a second sample image; Adding the first sample noise to the second sample image to obtain a second noise-added image; Inputting the second noise-added image into the pre-trained encoder to obtain image features of the second noise-added image output by the encoder; Extracting visual encoding features corresponding to the first sample noise from the image features of the second noise-added image.
4. The method according to claim 3, wherein extracting visual encoding features corresponding to the first sample noise from the image features of the second noise-added image specifically comprises: Inputting the second sample image into the pre-trained encoder to obtain image features of the second sample image output by the encoder; Removing the image features of the second sample image from the image features of the second noise-added image to obtain visual encoding features corresponding to the first sample noise.
5. The method according to any one of claims 2 to 4, wherein pre-training the encoder specifically comprises: Acquiring a to-be-trained encoding-decoding model and a third sample image, wherein the encoding-decoding model comprises an encoder and a decoder; Determining second sample noise according to each preset type of noise; Adding the second sample noise to the third sample image to obtain a third noise-added image; Inputting the third noise-added image into the encoder in the to-be-trained encoding-decoding model to obtain image features of the third noise-added image output by the encoder in the to-be-trained encoding-decoding model; Inputting the image features of the third noise-added image into the decoder in the to-be-trained encoding-decoding model to obtain a reconstructed noise output by the decoder in the to-be-trained encoding-decoding model; Determining a reconstruction loss according to the second sample noise and the reconstructed noise; Adjusting at least model parameters of the encoder in the to-be-trained encoding-decoding model according to the reconstruction loss, and using the encoder with adjusted model parameters as the pre-trained encoder.
6. The method according to claim 1, wherein the diffusion model comprises an image restoration bridging network, a denoising network, and an adaptation network; Injecting the visual encoding features and the first noisy image as generation conditions into the diffusion model specifically includes: Inputting the first noisy image into the image restoration bridging network to obtain multiple image features with different resolutions corresponding to the first noisy image output by the image restoration bridging network. For each image feature with a different resolution corresponding to the first noisy image, injecting the image feature at this resolution into the denoising processing layer corresponding to this resolution in the denoising network; and injecting the visual encoding features into the adaptation network to enable the adaptation network to process the output results of the denoising processing layers corresponding to each resolution in the denoising network.
7. The method according to claim 6, wherein the model parameters of the diffusion model are adjusted with the training objective of reducing the noise prediction loss, specifically including: Adjusting the model parameters of the image restoration bridging network and the adaptation network with the training objective of reducing the noise prediction loss.
8. A method for image quality restoration, the method comprising: Obtaining an image to be processed; Selecting at least one type of noise from each type of preset noise according to the image to be processed; Determining a fourth sample noise according to the at least one type of selected noise; Obtaining visual encoding features corresponding to the fourth sample noise; Inputting a standard noise image into a pre-trained diffusion model, and injecting the visual encoding features and the image to be processed as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image, wherein the diffusion model is pre-trained by the method according to any one of claims 1 to 7.
9. The method according to claim 8, wherein selecting at least one type of noise from each type of preset noise according to the image to be processed specifically includes: Obtaining a standard image; Selecting at least one type of noise from each type of preset noise according to the noise visual effect of the image to be processed; Adding the selected types of noise to the standard image; Evaluating the similarity between the noise visual effect of the image to be processed and the noise visual effect of the standard image after adding noise; If the similarity is higher than the set similarity, taking the types of noise currently added to the standard image as the selected types of noise; otherwise, continuing to select at least one type of noise from each type of preset noise, and continuing to add the selected noise to the standard image until the similarity between the noise visual effect of the standard image after adding noise and the noise visual effect of the image to be processed is higher than the set similarity.
10. An apparatus for training a diffusion model, the apparatus comprising: A preprocessing module for determining a first sample noise according to each type of preset noise; Obtaining visual encoding features corresponding to the first sample noise through a pre-trained encoder, and obtaining a first sample image, adding the first sample noise to the first sample image to obtain a first noisy image; A denoising module, configured to input a standard noise image into the diffusion model, and inject the visual encoding feature and the first noise-added image as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image; A loss determination module, configured to determine a noise prediction loss according to the first sample image and the restored image; A parameter adjustment module, configured to adjust the model parameters of the diffusion model with the goal of reducing the noise prediction loss during training.
11. An image quality restoration device, the device comprising: An acquisition module, configured to acquire an image to be processed; A selection module, configured to select at least one type of noise from each preset type of noise according to the image to be processed; A noise determination module, configured to determine a fourth sample noise according to the at least one type of selected noise; A condition acquisition module, configured to obtain a visual encoding feature corresponding to the fourth sample noise; A denoising module, configured to input a standard noise image into a pre-trained diffusion model, and inject the visual encoding feature and the image to be processed as generation conditions into the diffusion model, so that the diffusion model denoises the standard noise image based on the generation conditions to obtain a restored image, wherein the diffusion model is pre-trained by using the method according to any one of claims 1 to 7.
12. A computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 above is implemented.
13. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method according to any one of claims 1 to 9 above is implemented.
Citation Information
Cited By
Image processing method and device, equipment and storage medium
CN121563835A