Diffusion model processing method and device for unified image restoration, electronic equipment and storage medium
By introducing virtual consistency models and lightweight noise correction networks into image restoration technology, the performance limitation problem of the prior art in complex scenarios is solved, and a more efficient and robust image restoration effect is achieved.
Patent Information
- Application Number
- CN202510006142.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing image restoration technology shows the problems of performance limitation and complex model in low light, haze, rain, snow, sports and other scenarios.
A diffusion model-based processing method is adopted to establish a virtual consistency model through virtual consistency functions, and combine lightweight noise correction and prompt word refinement network to perform noise correction and refinement of images.
Without retraining the pre-trained diffusion model, the performance and fidelity of image restoration are significantly improved, and different types of degradation can be adaptively handled, improving the efficiency and robustness of image restoration.
Smart Images

Figure CN119941536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a diffusion model processing method, device, electronic equipment and storage medium for unified image restoration. Background Art
[0002] In low light, fog, rain, snow, moving objects, or mixed scenes, the captured photos may have problems such as blurred subjects or serious noise interference. Such images are called degraded images. Image restoration refers to the process of processing degraded images to obtain clearer images with the same image content. Currently, related technologies such as airnet, IDR, Painter, ProRes, and DA-CLIP can be applied to image restoration, but these related technologies generally have shortcomings. For example, airnet uses a module to map different degradation distributions to the same distribution through contrastive learning, but faces challenges in training and performance; IDR observes that degradation types can be segmented by singular value decomposition, and reconstructs clear images by reformulating singular values and vectors. Painter, ProRes, and DA-CLIP use dynamic learning to utilize the potential of large models to restore images. However, due to the multi-part mapping strategy, methods such as IDR, Painter, ProRes, and DA-CLIP can only achieve moderate performance improvements and require a lot of model complexity. Summary of the invention
[0003] In view of the technical problems of current image restoration technology such as limited performance and complex models, the purpose of the present invention is to provide a diffusion model processing method, device, electronic device and storage medium for unified image restoration.
[0004] In one aspect, an embodiment of the present invention includes a diffusion model processing method for unified image restoration, the diffusion model processing method for unified image restoration comprising the following steps:
[0005] Perform at least one iteration process; each iteration process includes the following steps:
[0006] Get the input image and prompt word;
[0007] Performing noise processing on the input image to obtain a noisy image;
[0008] Processing the noisy image using a noise correction network to obtain noise spatial features;
[0009] Using a pre-trained diffusion model to process the noise spatial features and the prompt word to obtain a noise residual;
[0010] Performing noise reconstruction processing on the input image to obtain reconstructed noise;
[0011] The noise spatial feature, the noise residual and the reconstruction noise are processed using a virtual consistency function to obtain a restored image of this round of the iterative process.
[0012] Furthermore, the use of a pre-trained diffusion model to process the noise spatial features and the prompt word to obtain a noise residual includes:
[0013] Obtaining a refined embedding feature corresponding to the prompt word;
[0014] Inputting the noise spatial features and the refined embedded features into the pre-trained diffusion model for processing;
[0015] Acquire first noise prediction information and second noise prediction information obtained by processing the pre-trained diffusion model;
[0016] A difference process is performed according to the first noise prediction information and the second noise prediction information to obtain the noise residual.
[0017] Furthermore, the obtaining of the refined embedding feature corresponding to the prompt word includes:
[0018] Processing the prompt word using a text encoder to obtain a prompt word embedding feature;
[0019] Processing the input image using a CLIP network to obtain a CLIP embedding feature;
[0020] The prompt word embedding feature and the CLIP embedding feature are subjected to prompt word refinement processing to obtain the refined embedding feature.
[0021] Furthermore, the using of a virtual consistency function to process the noise spatial feature, the noise residual and the reconstruction noise to obtain a restored image of this round of the iterative process includes:
[0022] According to the formula
[0023]
[0024] Perform calculations to obtain the restored image of this round of iterative process in, is the noise spatial feature, t represents the round order of the current iterative process, is the reconstruction noise, △∈ is the noise residual, Represents a low-quality image, f() satisfies
[0025]
[0026] in
[0027]
[0028] z0 represents the initial sampling data, z t represents the result after z0 is t times of noise addition, ∈(z t ,t;z0) represents the noise predictor, α t is a hyperparameter that satisfies the variance of the Gaussian distribution.
[0029] Furthermore, when the current iteration process is the first iteration process, the input image is an image to be processed;
[0030] When the current iteration process is an iteration process other than the first iteration process, the input image is the restored image obtained by processing the previous iteration process.
[0031] Furthermore, when the diffusion model processing method for unified image restoration is executed in the training phase, any round of iteration process further includes the following steps:
[0032] Acquire a high-quality image corresponding to the input image;
[0033] The noise correction network is supervisedly trained based on the high-quality image and the restored image.
[0034] Furthermore, when the diffusion model processing method for unified image restoration is executed in the inference stage, the diffusion model processing method for unified image restoration further includes the following steps:
[0035] Get the image enhancement level;
[0036] Determining the total number of rounds according to the image enhancement level;
[0037] After executing any round of the iterative process, the number of executed rounds of the iterative process is counted, and when the number of executed rounds is less than the total number of rounds, the next round of the iterative process is executed, otherwise, the execution of all the iterative processes is terminated;
[0038] The restored image obtained by the last round of the iterative process is obtained.
[0039] On the other hand, an embodiment of the present invention also includes a computer device including a memory and a processor, the memory is used to store at least one program, and the processor is used to load at least one program to execute the diffusion model processing method for unified image restoration in the embodiment.
[0040] On the other hand, an embodiment of the present invention also includes an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the diffusion model processing method for unified image restoration in the embodiment.
[0041] On the other hand, an embodiment of the present invention further includes a computer-readable storage medium storing a program executable by a processor, and the program executable by the processor is used to execute the diffusion model processing method for unified image restoration in the embodiment when executed by the processor.
[0042] The beneficial effects of the present invention are: the diffusion model processing method for unified image restoration in the embodiment establishes a virtual consistency model based on a virtual consistency function. Specifically, a lightweight noise correction and cue word refinement network is introduced, which can be used for noise and cue embedding modification, and the image restoration performance is enhanced without retraining the pre-trained diffusion model; by using the virtual consistency function, the virtual consistency model VCMUIR utilizes the uniform degradation characteristics of the high noise space of the diffusion model, which can effectively solve the general image restoration problem and significantly improve the fidelity; moreover, the virtual consistency model VCMUIR is a unified model as a whole, and the noise correction network therein can be trained through the training stage, so that it can adaptively handle different degradation types to enhance robustness, and can handle different degradation tasks at the same time. Therefore, it is a general image restoration model, which eliminates the need for specialized models for specific tasks, and thus can improve image restoration efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of the principle of a virtual consistency model in an embodiment;
[0044] Figure 2 Schematic diagram of the overall steps of the diffusion model processing method for unified image restoration in the embodiment;
[0045] Figure 3 It is a structural diagram of a virtual consistency model in an embodiment;
[0046] Figure 4 It is a structural diagram of the prompt word refinement network in the embodiment;
[0047] Figure 5 A schematic diagram of the steps of a diffusion model processing method for unified image restoration during the training phase in the embodiment;
[0048] Figure 6Schematic diagram of the steps of the diffusion model processing method for unified image restoration during the inference stage in the embodiment. DETAILED DESCRIPTION
[0049] In response to the technical problems of current image restoration technologies such as limited performance and complex models, DiffUIR proposes a selective hourglass mapping strategy based on a residual diffusion model, which simultaneously achieves shared distribution mapping and strong conditional guidance. However, DiffUIR relies heavily on the residual diffusion model, which makes it difficult to exploit the powerful priors in the community-based diffusion model, thus hindering further improvements in general image restoration. In addition, the original training objective of the diffusion model is not to directly solve the image restoration problem, which may lead to significant fidelity loss.
[0050] It can be seen that the key challenge facing image restoration technology is how to effectively learn multiple degradation distributions simultaneously to enhance the performance of image restoration tasks with multiple degradation types. The standard diffusion model exhibits degradation uniformity and obvious content in high-noise environments, making it very suitable for general image restoration.
[0051] The forward process and denoising process of the standard diffusion model are as follows:
[0052]
[0053] Based on the standard diffusion model, the key consistency function f(z t ,t), the key consistency function is used to ensure that the output of all points on the probability flow ordinary differential equation is consistent, and a consistency model can be obtained. Compared with the standard diffusion model, the consistency model greatly speeds up the generation process. In the consistency model, sampling is performed in a series of time steps T 1:n ∈[t0,T]. With initial noise and Initially, at each time step τ i Sampling ∈~N(0,1), iteratively updating the multi-step consistency sampling process is:
[0054]
[0055] Consider a special case of formula (2), when σ t Take at all time steps t When , formula (2) can be rewritten as
[0056]
[0057] Let ∈ represent a more general noise predictor, we can get
[0058]
[0059] Combining formula (5) and formula (6), a sampling process similar to multi-step consistency sampling [formula (3) and formula (4)] is as follows:
[0060]
[0061] In order to ensure that f(z t ,t;z0) is self-constrained and can be regarded as a consistency function. The goal of DDCM is f(z t ,t;z0)=z0. This involves solving the equations directly to compute ∈ without parameterization, as follows:
[0062]
[0063] Where ∈ rec is the reconstructed noise z0.
[0064] The consistency function f(z t ,t; z0) is not derived from the probability flow ordinary differential equation, and the consistency function f(z t ,t;z0) is called a virtual consistency function, and the model using it is called a virtual consistency model (VCMs). Based on formula (7) and formula (8), a new virtual consistency model for general image restoration (essentially a diffusion model) can be designed, called a virtual consistency model for image restoration (VCMUIR). In this embodiment, the principle of processing data by the virtual consistency model (VCMUIR) is as follows Figure 1 As shown, the consistency function used by the virtual consistency model (VCMUIR) is as follows:
[0065]
[0066] Where ∈ θ (z t ,t,c * ) is a pre-trained diffusion model; c * is the input condition (specifically, the refined embedding feature c HQ and c LQ ); represents the predicted high-quality image; Represents a low-quality image, which is equivalent to the high-quality image output by the previous step, and is initially set to According to formula (1), The result of adding noise. When t = 0, due to and Can cancel each other out, substituting into the formula to get Therefore, formula (9.1) satisfies the boundary conditions of the consistency model. With this goal, the sampling process of VCMUIR can be Figure 1 The noise correction network is shown in part (b) of φ It is actually designed to convert low-quality images Re-identify as high-quality low-components in a high-noise space.
[0067] Based on the above principle, in this embodiment, a diffusion model processing method for unified image restoration is provided. Figure 2 ,The diffusion model processing method for unified image restoration includes the following steps:
[0068] S1. Obtain the image to be processed and the prompt word;
[0069] Perform at least one iteration process; each iteration process includes the following steps:
[0070] S201. Get input image and prompt word;
[0071] S202. Performing noise processing on the input image to obtain a noisy image;
[0072] S203. Process the noisy image using a noise correction network to obtain noise spatial features;
[0073] S204. Process the noise spatial features and the prompt word using a pre-trained diffusion model to obtain a noise residual;
[0074] S205. Perform noise reconstruction processing on the input image to obtain reconstructed noise;
[0075] S206. Use a virtual consistency function to process the noise spatial features, noise residuals and reconstruction noise to obtain a restored image of this round of iterative process.
[0076] In this embodiment, the image to be processed obtained in step S1 may be a low-quality image (degraded image) or a high-quality image. Among them, a low-quality image is an image of a certain object to be photographed in scenes such as low light, fog, rain, snow, and the object to be photographed is moving, or an image obtained by artificially adding noise, and such an image has low-quality characteristics such as unclearness. A high-quality image is an image with less noise and clearer content than a low-quality image. The prompt word is a text used to describe the quality of the image to be processed. For example, if there is a low-quality area in the image to be processed, the prompt word may include a low-quality image prompt word, and its content may be "haze", "rain", "snow", "dark light", "blur", "motion blur", "low resolution", "unnatural", "low quality", etc. If there is a high-quality area in the image to be processed, the prompt word may include a high-quality image prompt word, and its content may be "clean", "real", "high resolution", "natural", "high quality", "fine details", etc.
[0077] Steps S201-S206 constitute one round of iteration. Generally, multiple rounds of iteration are required. The tth round of iteration is taken as an example for explanation. The processing flow executed by the virtual consistency model in the tth round of iteration is as follows: Figure 3 Specifically, Figure 3 The encoder, CLIP, cue word refinement network, noise correction network and diffusion model (main network) in the virtual consistency model. Figure 3 ,Networks such as decoders and LoRA adapters can also be included in the virtual consistency model.
[0078] If the current iteration process, i.e., the t-th iteration process, is the first iteration process, then the input image to be obtained in step S201 is the image to be processed obtained in step S1; if the current iteration process, i.e., the t-th iteration process, is an iteration process after the first iteration process, then the input image to be obtained in step S201 is the restored image obtained by the previous iteration process. The prompt word obtained in step S201 can be the same as the prompt word obtained in step S1.
[0079] In step S202, refer to Figure 3 , the input image can be Input into the encoder for encoding, so as to obtain the encoded image in the tth round of iteration Next, the principle of formula (1) is used to encode the image Perform noise processing to obtain the noise image in the tth round of iteration Specifically, for formula (1), that is,
[0080]
[0081] Can make it Execute formula (1) to calculate z t Afterwards Thus, we can obtain a noisy image
[0082] In step S203, refer to Figure 3 , the noisy image Input into the noise correction network for processing to obtain the noise spatial characteristics Specifically, in this embodiment, a UNet network can be used as the noise correction network φ The noise correction network specifically applies the first formula in formula (9.2), that is,
[0083]
[0084] For noisy images Processing to obtain noise spatial characteristics Among them, r φ () represents the processing process of the UNet network.
[0085] In this embodiment, when executing step S204, that is, using the pre-trained diffusion model to process the noise spatial features and the prompt word to obtain the noise residual, the following steps may be specifically performed:
[0086] S20401. Obtain the refined embedding features corresponding to the prompt word;
[0087] S20402. Inputting the noise spatial features and the refined embedded features into the pre-trained diffusion model for processing;
[0088] S20403. Obtaining first noise prediction information and second noise prediction information obtained through pre-trained diffusion model processing;
[0089] S20404. Perform difference processing based on the first noise prediction information and the second noise prediction information to obtain a noise residual.
[0090] In step S20401, refer to Figure 3 , the prompt word (specifically the low-quality image prompt word or the high-quality image prompt word) can be input into the text encoder for encoding processing to obtain the prompt word embedding features output by the text encoder; on the other hand, the input image is input into the CLIP network for feature extraction processing to obtain the CLIP embedding features; the prompt word embedding features and the CLIP embedding features are input into the prompt word refinement network for processing.
[0091] In this embodiment, the structure of the prompt word refinement network is as follows: Figure 4 As shown in Figure 1, it includes a cross-attention module, a CLIP feature mapping module, and a Transformer decoder. The structures of the cross-attention module and the CLIP feature mapping module are also shown in Figure 1. Figure 4 The cross attention module processes the cue word embedding features, and the CLIP feature mapping module processes the CLIP embedding features. The final results are added to obtain the refined embedding feature c. * (Specifically, for low-quality image prompt words, the refined embedding feature obtained can be expressed as c LQ , for high-quality image prompt words, the refined embedding features obtained can be expressed as c HQ ). The CLIP network extracts image features and initial quality descriptive cue embeddings of the input image, which are processed by the cross-attention and transformer decoder in the cue word refinement network to generate refined cue embeddings, i.e., refined embedding features c LQ and c HQ .
[0092] In step S20402, the noise spatial feature obtained in step S203 is executed and the refined embedding feature c obtained by executing step SS20401 * Combined, they are input into the pre-trained diffusion model for processing. The pre-trained diffusion model specifically applies the second formula in formula (9.2), that is,
[0093]
[0094] Specifically, ∈ θ () represents the processing of the pre-trained diffusion model. In step S20403, the pre-trained diffusion model predicts the high-quality image prompt word c HQ The first noise prediction information obtained by processing, In step S20403, the pre-trained diffusion model predicts the low-quality image prompt word c LQ The second noise prediction information obtained by processing. △∈ is the first noise prediction information obtained in step S20404. and the second noise prediction information The noise residual obtained by taking the difference.
[0095] In step S205, refer to Figure 4 , for the input image (corresponding encoded image ) is used to perform noise reconstruction processing to obtain the reconstructed noise. Specifically, using the third formula in formula (9.2), that is,
[0096]
[0097] According to the input image (corresponding coded image ) and noise spatial characteristics Calculate and obtain the reconstruction noise in the tth round of iteration
[0098] In this embodiment, when executing step S206, that is, using a virtual consistency function to process the noise spatial features, the noise residual and the reconstruction noise to obtain the restored image of this round of iterative process, the following steps may be specifically performed:
[0099] According to formula 9.1, that is
[0100]
[0101] Execute the calculation to obtain the restored image of this round of iteration, that is, the tth round of iteration in, is the noise spatial feature obtained by executing step S203, is the reconstructed noise obtained by executing step S205, and Δ∈ is the noise residual obtained by executing step S204.
[0102] In this embodiment, f() in Formula 9.1 satisfies
[0103]
[0104] in
[0105]
[0106] z0 represents the initial sampling data, z t represents the result after z0 is t times of noise addition, ∈(z t ,t;z0) represents the noise predictor, α t is a hyperparameter that conforms to the variance of the Gaussian distribution. Specifically, when performing calculations, we can let Calculate to obtain the restored image of the tth round of iteration process
[0107] If the iteration end condition is not met after the tth iteration, the t+1th iteration will be executed. The restored image of the tth iteration is (relatively high-quality image) will become the input image (relatively low-quality image) of the t+1th round of iteration process and be processed by a new round of iteration process.
[0108] In this embodiment, Figure 2 and Figure 3The processing process of the diffusion model for general image restoration on the processed image and the prompt word is shown. Before the diffusion model for general image restoration is applied to reasoning (restoring the degraded image that actually needs to be restored), a training phase can be performed to train the diffusion model for general image restoration.
[0109] During the training phase, refer to Figure 5 The image to be processed obtained in step S1 is specifically a sample image, which can be an image taken in a specially set specific scene (e.g., low light), or an image obtained by deliberately adding interference factors such as noise. When executing step S1, a high-quality image corresponding to the sample image is obtained. The high-quality image has the same content (photographing subject) as the sample image, and is taken in a better shooting scene (e.g., sufficient light, clear weather, and the photographing subject remains still) than the sample image, or has been subjected to noise reduction processing.
[0110] In the training phase, after each iteration (e.g., the tth iteration), the noise correction network r is trained based on the high-quality image corresponding to the sample image and the restored image obtained in the tth iteration. φ Conduct supervised training.
[0111] In this embodiment, the noise correction network r φ The training of θ needs to meet two main requirements: (1) to achieve the recovery function r in the high noise space φ , and (2) provide image restoration supervision in the image space.
[0112] For the first requirement, since the noisy image is the noise correction network r φ When t is large enough, r φ The range of is naturally in the high noise space. Taking advantage of uniform degradation, a simple and lightweight UNet network can be used to correct the input noise. For the second condition, first define the optimization objective. By expanding the consistency function [Formula (9)], we can get
[0113]
[0114] Formula (10) shows that the noise correction network r φ The main purpose is to fit the noise residuals. When the pre-trained diffusion model is a consistency model, it can be assumed that these residuals are approximately consistent at different time steps. Therefore, a fixed total step size (the maximum value of t) can be selected, and the perceptual loss function is determined by L1 loss and LPIPS as follows
[0115]
[0116] in is the restored image obtained in the tth round of iteration Represents the high-quality image corresponding to the sample image. λ is the balancing coefficient. The time step is set to 999 (i.e., the maximum value of t, that is, the total number of rounds of the iterative process performed is 999) to ensure maximum sub-uniform degradation. Although only one step of optimization is taken, it is sufficient to generalize to a multi-step denoising process. Obviously, formula (11) satisfies the conditions for supervision in the image space.
[0117] Reference Figure 5 In the tth round of iteration, after executing steps S201-S206, the loss function value can be calculated by formula (11). φ Perform supervised training. Since the noise correction network r in this embodiment φ It is a Unet network, so the noise correction network r can be trained using the Unet network training method. φ Conduct training.
[0118] Reference Figure 3 In this embodiment, the encoder, CLIP, text encoder, pre-trained diffusion network and decoder are frozen, so according to Figure 5 When supervised training is performed using the process of φ To train, or to correct the noise network r φ And the prompt word refinement network is trained.
[0119] pass Figure 5 As shown in the steps, the diffusion model processing method for unified image restoration can train VCMUIR.
[0120] After the training of the diffusion model for general image restoration is completed, the diffusion model for general image restoration can be applied to the inference stage, and the diffusion model for general image restoration is used to perform image restoration.
[0121] In the inference phase, refer to Figure 6 The image to be processed obtained in step S1 is specifically a degraded image. The degraded image may be an image taken under poor photographic conditions, or an image affected by interference factors such as noise.
[0122] Formula (10) proves that VCMUIR actually learns the noise residual between the degraded image and the high-quality image in the high-noise space. Therefore, through multi-step reasoning, the image enhancement level can be adjusted to meet the preferences of different users. In the reasoning stage, the image enhancement level can be determined according to the user's preference, where the image enhancement level indicates the degree to which the user wants to enhance the degraded image. The greater the degree of enhancement the user wants to perform on the degraded image, the higher the image enhancement level, and the greater the total number of rounds determined accordingly.
[0123] Reference Figure 6 After each round of iteration, check whether the total number of iterative processes that have been executed has reached the total number of rounds; if it has not reached the total number of rounds, the restored image obtained by the iterative process that has just been executed will be used as the input image of the next round of iterative process, and the next round of iterative process will be executed; if the total number of rounds has been reached, then all iterative processes will be terminated, and the next round of iterative process will not be executed. The restored image obtained by the last round of iterative process will be output, refer to Figure 3 , the restored image is input into the decoder for decoding, thereby obtaining the restoration result of the degraded image.
[0124] Reference Figure 3 , trainable LoRA and skip connections can also be added to the decoder that processes the restored image, thereby enhancing the detail restoration of the restoration result.
[0125] pass Figure 6 The steps shown in the figure, the diffusion model processing method for unified image restoration can be applied to the following scenarios:
[0126] 1. Enhanced image brightness in low-light scenes
[0127] 2. Image defogging in foggy scenes
[0128] 3. Deraining images in rainy scenes
[0129] 4. Remove snow from images in snowy scenes
[0130] 5. Image deblurring in motion blur scenes
[0131] 6. Image enhancement in mixed degradation scenarios
[0132] The present invention constructs a method VCMUIR for restoring images in a high-noise space and proposes a virtual consistency model for general image restoration. The model effectively realizes general image restoration and ensures accurate supervision of images in the image space through the proposed consistency function. The degraded image exhibits uniform degradation characteristics in the high-noise space of the diffusion model, so the degradation conflict problem can be solved. At the same time, ensuring accurate supervision data during training is crucial for image restoration, and the image space is the optimal space to achieve this goal. A large number of experiments show that the present invention outperforms existing methods in 22 different degrees of degradation, including general restoration and challenging real-world restoration tasks.
[0133] The diffusion model processing method for unified image restoration in this embodiment reformulates the inverse denoising process in the diffusion model by designing a virtual consistency function, thereby establishing a virtual consistency model VCMUIR; the virtual consistency model VCMUIR naturally unifies different degradations through a high-noise spatial standard diffusion model, while retaining the core image content, and solves the technical problems of current image restoration technology such as limited performance and complex models while retaining the pre-trained diffusion model prior knowledge to the greatest extent; specifically, in the virtual consistency model VCMUIR in this embodiment, a lightweight noise correction and cue word refinement network are introduced, which can be used for noise and cue embedding modification, and the performance of image restoration is enhanced without retraining the pre-trained diffusion model; by using the virtual consistency function, the virtual consistency model VCM UIR utilizes the uniform degradation characteristics of the high-noise space of the diffusion model, which can effectively solve the general image restoration problem and significantly improve the fidelity; the virtual consistency model VCMUIR is a unified development model as a whole. The noise correction network in it can be trained through the training stage, so that it can adaptively handle different degradation types to enhance robustness, and can handle different degradation tasks at the same time. Therefore, it is a general image restoration model, eliminating the need for specialized models for specific tasks, and thus can improve image restoration efficiency; the image restoration function and supervision function of the virtual consistency model VCMUIR are decoupled, and it can enhance detail restoration by freezing the pre-trained diffusion model, introducing a lightweight noise correction network, a cue word refinement network, and adding trainable LoRA and jump connections in the decoder, thereby achieving image restoration of multiple degradation types.
[0134] A computer program for executing the diffusion model processing method for unified image restoration in this embodiment can be written and written into a computer device or a storage medium. When the computer program is read out and run, the diffusion model processing method for unified image restoration in this embodiment is executed, thereby achieving the same technical effect as the diffusion model processing method for unified image restoration in the embodiment.
[0135] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. In addition, the descriptions of up, down, left, right, etc. used in the present disclosure are only relative to the relative positional relationship of the components of the present disclosure in the accompanying drawings. The singular forms of "a", "" and "the" used in the present disclosure are also intended to include the plural forms, unless the context clearly indicates other meanings. In addition, unless otherwise defined, all technical and scientific terms used in this embodiment have the same meaning as those generally understood by those skilled in the art. The terms used in the specification of this embodiment are only for describing specific embodiments and are not intended to limit the present invention. The term "and / or" used in this embodiment includes any combination of one or more related listed items.
[0136] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided by the present embodiment is only intended to better illustrate embodiments of the present invention, and unless otherwise required, the scope of the present invention will not be limited.
[0137] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed dedicated integrated circuit for this purpose.
[0138] In addition, the operations of the process described in this embodiment may be performed in any suitable order, unless otherwise indicated in this embodiment or otherwise clearly contradicted by the context. The process described in this embodiment (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as a code (e.g., executable instructions, one or more computer programs, or one or more applications) executed on one or more processors in common, by hardware or a combination thereof. A computer program includes a plurality of instructions that may be executed by one or more processors.
[0139] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or part thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the above steps in conjunction with a microprocessor or other data processor, the invention of this embodiment includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.
[0140] The computer program can be applied to input data to perform the functions of the present embodiment, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.
[0141] The above are only preferred embodiments of the present invention. The present invention is not limited to the above embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical solutions and / or implementation methods may have various modifications and changes.
Claims
1. A diffusion model processing method for unified image restoration, characterized in that: The diffusion model processing method for unified image restoration includes: Perform at least one iteration process; each iteration process includes the following steps: Get the input image and prompt word; Performing noise processing on the input image to obtain a noisy image; Processing the noisy image using a noise correction network to obtain noise spatial features; Using a pre-trained diffusion model to process the noise spatial features and the prompt word to obtain a noise residual; Performing noise reconstruction processing on the input image to obtain reconstructed noise; The noise spatial feature, the noise residual and the reconstruction noise are processed using a virtual consistency function to obtain a restored image of this round of the iterative process.
2. The diffusion model processing method for unified image restoration according to claim 1, characterized in that: The using of the pre-trained diffusion model to process the noise spatial feature and the prompt word to obtain the noise residual includes: Obtaining a refined embedding feature corresponding to the prompt word; Inputting the noise spatial features and the refined embedded features into the pre-trained diffusion model for processing; Acquire first noise prediction information and second noise prediction information obtained by processing the pre-trained diffusion model; A difference process is performed according to the first noise prediction information and the second noise prediction information to obtain the noise residual.
3. The diffusion model processing method for unified image restoration according to claim 2, characterized in that: The obtaining of the refined embedding feature corresponding to the prompt word includes: Processing the prompt word using a text encoder to obtain a prompt word embedding feature; Processing the input image using a CLIP network to obtain a CLIP embedding feature; The prompt word embedding feature and the CLIP embedding feature are subjected to prompt word refinement processing to obtain the refined embedding feature.
4. The diffusion model processing method for unified image restoration according to claim 1, characterized in that: The process of using a virtual consistency function to process the noise spatial feature, the noise residual and the reconstruction noise to obtain a restored image of this round of the iterative process includes: According to the formula Perform calculations to obtain the restored image of this round of iterative process in, is the noise spatial feature, t represents the round order of the current iterative process, is the reconstruction noise, △∈ is the noise residual, Represents a low-quality image, f() satisfies in z0 represents the initial sampling data, z t represents the result after z0 is t times of noise addition, ∈(z t ,t;z0) represents the noise predictor, α t is a hyperparameter that satisfies the variance of the Gaussian distribution.
5. The diffusion model processing method for unified image restoration according to any one of claims 1 to 4, characterized in that: When the current iteration process is the first iteration process, the input image is an image to be processed; When the current iteration process is an iteration process other than the first iteration process, the input image is the restored image obtained by processing the previous iteration process.
6. The diffusion model processing method for unified image restoration according to any one of claims 1 to 4, characterized in that: When the diffusion model processing method for unified image restoration is executed in the training phase, any round of iteration process further includes the following steps: Acquire a high-quality image corresponding to the input image; The noise correction network is supervisedly trained based on the high-quality image and the restored image.
7. The diffusion model processing method for unified image restoration according to any one of claims 1 to 4, characterized in that: When the diffusion model processing method for unified image restoration is executed in the inference stage, the diffusion model processing method for unified image restoration further includes the following steps: Get the image enhancement level; Determining the total number of rounds according to the image enhancement level; After executing any round of the iterative process, the number of executed rounds of the iterative process is counted, and when the number of executed rounds is less than the total number of rounds, the next round of the iterative process is executed, otherwise, the execution of all the iterative processes is terminated; The restored image obtained by the last round of the iterative process is obtained.
8. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory is used to store at least one program, and the processor is used to load at least one program to execute the diffusion model processing method for unified image restoration as described in any one of claims 1 to 7.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to execute the diffusion model processing method for unified image restoration as described in any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Image super-resolution reconstruction method and system based on diffusion model
CN117522694A
Diffusion models having continuous scaling through patch-wise image generation
US20240161327A1
Cited By
Unified image restoration method and system based on task-driven semantic adjustment and color gamut correction
CN121169758A