Arbitrary multiplying power super-resolution enhancement system and method based on diffusion model
Through an arbitrary magnification super-resolution enhancement system based on the diffusion model, combined with a dual-branch network and a symmetric diffusion encoder and decoder, the problem of high-magnification scene blur artifacts and generation results in the prior art do not conform to the natural image distribution, and image generation with high fidelity and high conform to the natural image distribution is achieved.
Patent Information
- Application Number
- CN202510067217.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing arbitrary super-resolution enhancement method exhibits blur artifacts in high-magnification scenarios, reducing visual quality, and the generation results are inconsistent with the original image, and the supplementary pixels do not conform to objective logic.
A random magnification super-resolution enhancement system based on diffusion model is adopted, and images are combined with dual-branch networks and symmetric diffusion encoder and decoder to produce images that meet the target super-resolution enhancement magnification.
Real high-resolution image generation with high fidelity and highly consistent with natural image distribution is achieved, super-resolution enhancement of arbitrary resolution magnifications, and high-quality images can be generated even with low resolution image resolution.
Smart Images

Figure CN119991483A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a super-resolution enhancement method based on deep learning, and in particular to an arbitrary-magnification super-resolution enhancement system and method based on a diffusion model. Background Art
[0002] In the 5G era, the increase in network bandwidth has promoted the development of high-definition video streaming, virtual reality (VR) and augmented reality (AR), and users have higher and higher requirements for clear and delicate visual experience. With the development of technology and the expansion of application scenarios, especially in the context of the popularization of ultra-high-definition display devices such as 4K and 8K, traditional low-resolution images cannot meet the growing visual quality requirements, and super-resolution technology has become the key to solving this bottleneck. Super-resolution technology can effectively enhance low-resolution images to high resolution, enhance visual details, reduce data transmission pressure, and ensure viewing experience at the same time, which provides technical support for major Internet platforms, games, and film and television production companies.
[0003] Traditional super-resolution enhancement methods are usually only applicable to fixed magnifications (such as 2×, 3×, 4×). Different magnifications require different super-resolution enhancement models. Arbitrary magnification super-resolution uses one model to flexibly generate high-resolution images of different magnifications according to needs. It also supports fractional super-resolution, which greatly improves the flexibility and practicality of the application of super-resolution enhancement technology.
[0004] Most current arbitrary-rate super-resolution enhancement methods use regression-based optimization methods, which usually exhibit pattern averaging behavior in high-rate super-resolution enhancement scenarios, resulting in blurring artifacts and significantly reducing the visual quality of human observers. In recent years, researchers have improved the optimization model by introducing generative models GAN and Diffusion, which has enhanced the authenticity of the image, but has also introduced new problems, namely, the generated results are inconsistent with the original image, or the supplemented pixels do not conform to objective logic. Therefore, how to achieve high fidelity and high compliance with natural image distribution have become two main problems in current arbitrary-rate super-resolution enhancement methods. Summary of the invention
[0005] In view of the defects in the prior art, the object of the present invention is to provide an arbitrary magnification super-resolution enhancement system and method based on a diffusion model.
[0006] According to one aspect of the present invention, there is provided an arbitrary magnification super-resolution enhancement system based on a diffusion model, comprising:
[0007] Input module, given a low-resolution image of any size and any target super-resolution enhancement factor;
[0008] An upsampling unit, which improves the low-resolution image to a target resolution by bicubic interpolation to obtain an intermediate target resolution image;
[0009] An image encoder is used to encode the intermediate target resolution image to obtain latent space input features;
[0010] A dual semantic guidance module, collecting text descriptions and semantic features of the intermediate target resolution image;
[0011] A latent space module, which denoises the latent space input features based on the target super-resolution enhancement factor, the text description and the semantic features to obtain intermediate features;
[0012] An image decoder decodes the intermediate features to obtain a final image that meets the target super-resolution enhancement ratio.
[0013] Preferably, the latent space module adopts a two-branch network, including a low-resolution image branch and a diffusion model main branch;
[0014] The low-resolution image branch extracts multi-scale features of the latent space input features by combining the text description, the semantic features and the target super-resolution enhancement factor;
[0015] Among them, the main branch of the diffusion model combines the text description, the semantic features, the target super-resolution enhancement factor and random noise to extract the latent space input feature image features, and fuses them with the multi-scale features to obtain the fused intermediate features.
[0016] Preferably, the low-resolution image branch includes a plurality of small groups, each of which includes a convolution layer, a text attention layer and a feature attention layer in sequence; the convolution kernel size of the convolution layer of each small group is different; wherein the convolution layer injects the target super-resolution enhancement factor into the latent space input feature or the output of the previous small group; the text attention layer injects the text description into the latent space input feature or the output of the previous small group; the feature attention layer injects the semantic feature into the latent space input feature or the output of the previous small group; and the multi-scale feature of the latent space input feature is obtained through the outputs of the plurality of small groups;
[0017] The main branch of the diffusion model includes a symmetrical diffusion encoder and a diffusion decoder, and the diffusion encoder and the diffusion decoder respectively include multiple groups, each of which includes a convolution layer, a text attention layer and a feature attention layer in sequence; the convolution kernel size of the convolution layer of each group is different; the diffusion encoder takes the noise, the text description, the semantic feature and the target super-resolution enhancement ratio as input, gradually reduces the width and height of the noise feature, and increases the number of channels; the decoder takes the text description, the semantic feature, the target super-resolution enhancement ratio and the multi-scale feature as input, gradually restores the original width and height of the noise feature, and reduces the number of channels to obtain the fused intermediate feature.
[0018] Preferably, it further comprises a global rate modulation unit, which modulates the time step characteristics of the main branch of the diffusion model based on the target super-resolution enhancement rate;
[0019] The global rate modulation unit comprises:
[0020] A first encoder is used to perform position encoding on the target super-resolution enhancement factor;
[0021] A first linear layer performs linear processing on the position code to obtain a linear processing result;
[0022] The first activation layer performs activation processing on the linear processing result to obtain a super-resolution magnification feature;
[0023] The updating unit adds the super-resolution ratio feature to the time step feature of the main branch of the diffusion model to update the original time step feature.
[0024] Preferably, it further comprises a local rate modulation unit, which modulates the multi-scale low-resolution image features extracted by the low-resolution image conditional branch based on the target super-resolution enhancement rate;
[0025] The local rate modulation unit comprises:
[0026] A second encoder performs position encoding on the target super-resolution enhancement factor;
[0027] A second linear layer performs linear processing on the position code to obtain a linear processing result;
[0028] The second activation layer performs activation processing on the linear processing result to obtain a super-resolution magnification feature;
[0029] The third linear layer performs linear processing on the super-resolution magnification feature to obtain a scale parameter and a bias parameter;
[0030] The modulation unit multiplies the multi-scale low-resolution image features by the scale parameters correspondingly, and then adds the multi-scale low-resolution image features by the bias parameters correspondingly to obtain the modulated features.
[0031] Preferably, it comprises a first training module, which trains the main branch network of the diffusion model;
[0032] The first training module includes:
[0033] The data submodule, given a high-resolution image, obtains the latent space representation z0 after passing through the main branch network of the diffusion model;
[0034] The noise adding submodule gradually adds noise to the latent space representation z0 to obtain the noisy latent space representation z t ;
[0035] Training submodule, combined with low-resolution image I lr , target super-resolution enhancement factor s, diffusion model time step t, text description c, semantic feature p, training diffusion model main branch network ∈ θ To predict the latent space representation z added to the noise t The noise in the training submodule; the loss function used in the training submodule is:
[0036]
[0037] in Represents the calculation prediction noise ∈ θ The L2 loss of the real noise ∈ is used as the diffusion loss, N represents the uniformly distributed random noise distribution, and ε~N represents sampling a noise sample from the noise distribution N.
[0038] Preferably, a second training module is included, which trains the image encoder; the loss function of the second training module is a reconstruction loss function, and the reconstruction loss preliminarily reconstructs the structural detail information by constraining the intermediate features of the image encoder, and its specific expression is:
[0039]
[0040] in Represents the nth multi-scale feature extracted by the image encoder, toRGB means using convolution to convert the feature into an RGB image, It represents downsampling the real high-resolution image to the resolution of the nth feature.
[0041] According to a second aspect of the present invention, there is provided an arbitrary magnification super-resolution enhancement method based on a diffusion model, comprising:
[0042] Given a low-resolution image of any size and any target super-resolution enhancement factor;
[0043] The low-resolution image is enhanced to a target resolution by bicubic interpolation to obtain an intermediate target resolution image;
[0044] Encoding the intermediate target resolution image to obtain latent space input features;
[0045] Collecting text description and semantic features of the intermediate target resolution image;
[0046] Based on the target super-resolution enhancement factor, the text description and the semantic feature, denoising the latent space input feature to obtain an intermediate feature;
[0047] The intermediate features are decoded to obtain a final image that meets the target super-resolution enhancement ratio.
[0048] According to a third aspect of the present invention, there is provided a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor can be used to execute the system or the method when executing the program.
[0049] According to a fourth aspect of the present invention, a computer-readable storage medium stores a computer program, which, when executed by a processor, can be used to run the system or execute the method.
[0050] Compared with the prior art, the embodiments of the present invention have at least one of the following beneficial effects:
[0051] An arbitrary magnification super-resolution enhancement system and method based on a diffusion model is proposed in an embodiment of the present invention. LR and the expected super-resolution enhancement factor s (or the resolution of the target image), a unified latent space module is used to enhance the given low-resolution image to a high-resolution image of the expected resolution I HR , expand image resolution, enhance visual details, and generate real high-resolution images that conform to natural image distribution while fully maintaining the original image information.
[0052] An arbitrary-resolution-ratio super-resolution enhancement system and method based on a diffusion model proposed in an embodiment of the present invention supports super-resolution enhancement of arbitrary resolution ratio. Through super-resolution ratio modulation, the super-resolution ratio perception capability is injected into the generation process of the diffusion model, so that the diffusion model can adaptively process image features according to the super-resolution enhancement ratio.
[0053] An arbitrary-magnification super-resolution enhancement system and method based on a diffusion model is proposed in an embodiment of the present invention. Through dual semantic guidance, the semantic preservation capability in the generation process is enhanced, so that the image generated by the diffusion model has high fidelity and highly consistent with the distribution of real images.
[0054] An arbitrary-magnification super-resolution enhancement system and method based on a diffusion model proposed in an embodiment of the present invention can still generate realistic and faithful images even when the resolution of the low-resolution image is as low as 32×32 and the target super-resolution enhancement ratio exceeds 20×, which is superior to other arbitrary-magnification super-resolution enhancement methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0056] Figure 1 This is a structural diagram of an arbitrary-magnification super-resolution enhancement system based on a diffusion model according to an embodiment of the present invention;
[0057] Figure 2 A schematic diagram of a training process of a system and method for super-resolution enhancement at any magnification based on a diffusion model according to an embodiment of the present invention;
[0058] Figure 3 The results of comparison between the arbitrary-magnification super-resolution enhancement system based on the diffusion model in the embodiment of the present invention and other baseline models;
[0059] Figure 4 The visualization result of arbitrary magnification enhancement of a specific embodiment of the present invention and the result compared with the baseline method;
[0060] Figure 5 The figure is a visualization result of comparing a specific embodiment of the present invention with more baseline methods. DETAILED DESCRIPTION
[0061] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0062] In one embodiment of the present invention, Figure 1 As shown, an arbitrary magnification super-resolution enhancement system based on a diffusion model comprises:
[0063] Input module, given a low-resolution image of any size and any target super-resolution enhancement factor;
[0064] The upsampling unit uses bicubic interpolation to increase the low-resolution image to the target resolution and obtain an intermediate target resolution image.
[0065] Image encoder, which encodes the intermediate target resolution image to obtain latent space input features;
[0066] The dual semantic guidance module collects text descriptions and semantic features of intermediate target resolution images;
[0067] The latent space module denoises the latent space input features based on the target super-resolution enhancement ratio, text description and semantic features to obtain intermediate features;
[0068] The image decoder decodes the intermediate features to obtain the final image that meets the target super-resolution enhancement ratio.
[0069] In the above embodiment, by giving a low-resolution image I LR As well as the expected super-resolution enhancement factor s (or the resolution of the target image), a unified latent space module is used to upgrade the specified low-resolution image to a high-resolution image I with the expected resolution HR , thereby expanding the resolution of the image and enhancing the visual details. In this process, the information of the original image is fully preserved, and a true high-resolution image that conforms to the distribution of natural images is generated. Among them, the intermediate target resolution image is a rough image, which is simply a low-resolution image I LR Resized to the target resolution, the image is a blurry, low-quality target resolution image lacking texture details.
[0070] In a preferred embodiment of the present invention, an upsampling unit and an image encoder are designed. The image encoder maps low-resolution image features to a latent space module, and before passing through the image encoder, the low-resolution image is first upsampled to a target resolution by bicubic interpolation, and then the latent space representation of the low-resolution image is obtained by the image encoder implemented by a convolutional neural network as the input of the latent space diffusion model (low-resolution branch).
[0071] In order to improve the high fidelity of the final generated image and conform to the distribution characteristics of natural images, in a preferred embodiment of the present invention, the latent space module adopts a dual-branch design, one is a low-resolution image branch, and the other is a diffusion model main branch. The low-resolution image branch mainly targets the latent space input features mapped by the image encoder, and further extracts multi-scale features from the latent space input features. The main branch of the diffusion model iteratively denoises the input random noise mainly based on text descriptions, semantic features, and target super-resolution enhancement ratios. In the process of iterative denoising, image features are generated and fused with multi-scale features to obtain fused intermediate features.
[0072] In order to avoid problems such as the result of decoding the intermediate features of the latent space module being inconsistent with the original image, or the supplemented pixels not conforming to objective logic, a dual semantic guidance module is designed in a preferred embodiment of the present invention to guide the latent space module to generate a high-resolution image that conforms to the content, structure, and texture of the original image, thereby enhancing the fidelity of the low-resolution input image. Specifically, the dual semantic guidance module includes a text guidance unit and a semantic guidance unit. The text guidance unit first generates a text description for the low-resolution image, and then uses the generated text description to guide the generation of the high-resolution image. The semantic guidance unit extracts the semantic features of the low-resolution image, and calculates cross-attention with the features of the diffusion model, so that the semantic information of the low-resolution image is fully perceived during the generation of the high-resolution image.
[0073] In a specific embodiment, the low-resolution image branch includes a plurality of groups, each of which includes a convolution layer, a text attention layer, and a feature attention layer in sequence; the convolution kernel size of each group is different; wherein the convolution layer injects a target super-resolution enhancement factor into the latent space input feature or the output of the previous group, the text attention layer injects a text description into the latent space input feature or the output of the previous group; the feature attention layer injects a semantic feature into the latent space input feature or the output of the previous group; and the multi-scale feature of the latent space input feature is obtained through the outputs of the plurality of groups;
[0074] The main branch of the diffusion model includes a symmetrical diffusion encoder and a diffusion decoder. The diffusion encoder and the diffusion decoder each include multiple groups, each of which includes a convolution layer, a text attention layer, and a feature attention layer in sequence; the convolution kernel size of each group is different; the diffusion encoder takes noise, text description, semantic features, and target super-resolution enhancement ratio as input, gradually reduces the width and height of the image features, and increases the number of channels; the decoder takes text description, semantic features, target super-resolution enhancement ratio, and multi-scale features as input, gradually restores the original width and height of the image features, and reduces the number of channels to obtain fused intermediate features. The fusion process is that the multi-scale features output by the local rate modulation unit of the decoder are fused with the corresponding scales in the decoder in the main branch, and finally output a feature with the same dimension as the encoder input feature.
[0075] In the above embodiment, the abstract semantic information of the text description and the spatial semantic information of the semantic features promote each other, thereby improving the fidelity of the high-resolution image to the low-resolution input image.
[0076] After solving the above-mentioned fidelity problem, it is necessary to consider the adaptive problem of the super-resolution enhancement factor s. In order to enable the entire latent space module to adaptively process image features according to the super-resolution enhancement factor, in a preferred embodiment, a global factor modulation unit and a local factor modulation unit are designed.
[0077] The global rate modulation unit uses the time-stepping feature of the super-resolution enhanced rate modulation diffusion model.
[0078] In the global rate modulation unit, there are encoders, linear layers, activation layers, and update layers. First, the encoder is implemented to positionally encode the super-resolution rate s, then two linear layers activated by the activation layer SiLU are used to obtain the super-resolution rate feature, and finally the update layer adds it to the time step feature E t To update the original time step features, the global information can be modulated. Specifically, the formula for modulating the time step features based on the super-resolution ratio is as follows:
[0079] E p =Sinusoidal(s)
[0080] E s =Linear(SiLU(Linear(E p )))
[0081] E=E t +E s
[0082] Where s represents the super-resolution ratio, Sinusoidal represents sinusoidal encoding, Linear represents linear layer, and E t represents the original time step, and SiLU represents the activation layer function.
[0083] The local rate modulation unit extracts the multi-scale low-resolution image features F from the low-resolution image conditional branch, modulates the multi-scale features F through the super-resolution rate, and then adds the modulated features to the features of the corresponding resolution of the Unet decoder of the diffusion model.
[0084] The local rate modulation unit is a channel-level super-resolution enhancement rate adaptive feature transformation unit. It first obtains a rate feature for local feature modulation by using a method similar to global rate modulation:
[0085] E ls =Linear(SiLU(Linear(E p )))
[0086] Then, the scale feature α and the bias feature β are obtained through the linear layer, and the multi-scale low-resolution features are modulated by these two parameter features. The formula for modulating the low-resolution feature by super-resolution magnification is as follows:
[0087] α,β=Linear α (E ls ),Linear β (E ls )
[0088] F=α⊙F+β
[0089] Linear α and Linear β Represents a linear layer.
[0090] The designed super-resolution rate modulation unit in the above embodiment injects the super-resolution rate perception capability into the generation process of the diffusion model, so that the diffusion model can adaptively process image features according to the super-resolution enhancement rate, and adaptively adjust the generation capability according to the target enhancement rate: when the rate is high, the generation capability is fully utilized to expand the missing pixels, and when the rate is low, more emphasis is placed on the fidelity capability.
[0091] The above-mentioned embodiment uses the generation process of the target super-resolution modulated high-resolution image to adaptively adjust the generation capability of the generation model, thereby achieving super-resolution enhancement supporting any resolution ratio.
[0092] In order to improve the accuracy of the entire arbitrary-magnification super-resolution enhancement system based on the diffusion model, training is required. In a preferred embodiment, a training module is used, such as Figure 2 As shown, the training module mainly includes:
[0093] Data pair construction submodule: Construct arbitrary-scale super-resolution enhanced training data pairs. Randomly crop 512×512 images from high-quality images as high-resolution images, and then downsample the high-resolution images with a downsampling factor of any floating point number between [1,16] as low-resolution images. In order to support high-quality high-resolution image output of any target resolution (including resolution lower than 512×512), randomly transform the resolution of data pairs during subsequent training. Generate text descriptions based on high-resolution images for subsequent training. At this time, paired high-resolution-low-resolution image pairs and corresponding text descriptions are obtained as training data for subsequent training. Data enhancement methods can be used during training.
[0094] System framework construction sub-module: Construct an arbitrary-magnification super-resolution enhancement system based on a diffusion model.
[0095] Using an architecture based on the latent space diffusion model, the VAE codec is used to map the low-resolution image to the latent space module, and the diffusion model is applied to the latent space module. The VAE codec uses pre-trained weights and does not update parameters during the entire training process. The latent space module contains two branches, one is the diffusion model main branch, and the other is the low-resolution image branch. The low-resolution image conditional branch first upsamples the low-resolution image obtained by downsampling to the target resolution, then encodes the low-resolution image to the latent space module through an image encoder, and then extracts multi-scale features and integrates the features into the diffusion model main branch.
[0096] For the entire diffusion model main branch, the super-resolution rate modulation module (including global rate modulation unit and local rate modulation unit) and the dual semantic information guidance module (including text guidance unit and semantic guidance unit) mentioned in the above embodiment are used to enhance the fidelity and quality of the generated image, and obtain a high-resolution image that is closer to the real image distribution and has better visual effects without damaging the original image information. Among them, the semantic features in the semantic guidance unit are obtained through the latent space features of an object recognition network.
[0097] The above-mentioned construction submodule applies the diffusion model to the latent space module, and encodes the image to the VAE codec required by the latent space module using a pre-trained model. The latent space diffusion model uses the super-resolution magnification modulation and dual semantic information guidance designed in the embodiment of the present invention.
[0098] Loss function submodule: Use diffusion loss and reconstruction loss to train the latent space module.
[0099] Given a high-resolution image, we first map it to a latent space module through a VAE encoder to obtain the latent space feature z0, and then gradually add noise to z0 to obtain the noisy latent space representation z t . Using low-resolution images I LR , super-resolution enhancement factor s, diffusion time step t, text description c and semantic feature p, train the diffusion network ε θ To predict the latent space representation z added to the noise t The noise in . The loss function is:
[0100]
[0101] in Represents the calculation prediction noise ∈ θ The L2 loss of the real noise ∈ is used as the diffusion loss, N represents the uniformly distributed random noise distribution, and ε~N represents sampling a noise sample from the noise distribution N.
[0102] While training the diffusion model, the low-resolution image encoder is jointly trained. The multi-scale features extracted by the encoder are converted into RGB images of corresponding resolution through convolution, and then the L1 loss is calculated with the RGB images of corresponding resolution:
[0103]
[0104] Among them, F e n Represents the multi-scale features of the low-resolution image encoder, n represents a certain dimension. toRGB means converting the features into an RGB image of the corresponding resolution.
[0105] The complete training loss is the weighted sum of diffusion loss and reconstruction loss:
[0106] L total =L diff +α rec L rec
[0107] where α rec is a hyperparameter to balance the diffusion loss and reconstruction loss.
[0108] The above loss function sub-module jointly optimizes the image encoder and the latent space diffusion model that maps the low-resolution image to the latent space module.
[0109] Application module: When applying inference, given the input low-resolution image and the target super-resolution enhancement ratio, the resolution of the low-resolution image is first increased to the target resolution through bicubic interpolation, and then the low-resolution image is encoded into the latent space module using the image encoder and sent to the low-resolution image branch network. For the high-resolution rough image, its text and semantic features are extracted and sent to the low-resolution image branch and the main branch as conditions;
[0110] The main branch network randomly initializes noise, uses the DDPM sampling method, iteratively denoises, and sets the number of steps to 20. After obtaining the denoised latent space representation, the final high-resolution output image is obtained through the pre-trained VAE decoder.
[0111] The reasoning process of the above application module is implemented after all training is completed, and the input low-resolution image can be enhanced into a high-resolution image of any resolution according to the expected super-resolution ratio.
[0112] Based on the same inventive concept, in other embodiments of the present invention, an arbitrary magnification super-resolution enhancement method based on a diffusion model is also provided, and the steps are as follows:
[0113] Step 1: Given a low-resolution image of any size and an arbitrary target super-resolution enhancement factor;
[0114] Step 2: The low-resolution image is upgraded to the target resolution by bicubic interpolation to obtain a rough target resolution image;
[0115] Step 3, encode the coarse target resolution image to obtain latent space input features;
[0116] Step 4, collecting text descriptions and semantic features of the rough target resolution image;
[0117] Step 5: Based on the target super-resolution enhancement ratio, text description and semantic features, the latent space input features are denoised to obtain intermediate features;
[0118] Step 6: decode the intermediate features to obtain the final image that meets the target super-resolution enhancement ratio.
[0119] The various steps in the above examples of the present invention may refer to the implementation techniques of the corresponding units / modules of the arbitrary magnification super-resolution enhancement system based on the diffusion model in the above embodiments, which will not be described in detail here.
[0120] In order to verify the feasibility and effect of the arbitrary-rate super-resolution enhancement system and method based on the diffusion model in the above embodiment, a specific embodiment of the present invention selects the verification set of the DIV2K dataset as the test image dataset. The arbitrary-rate super-resolution enhancement methods LIIF, CiaoSR, IDM and the real blind super-resolution enhancement methods StableSR, PASD, and SeeSR are used as baseline models.
[0121] PSNR, LPIPS, FID, and CLIPIQA are used as evaluation indicators, where PSNR is used to measure pixel similarity, LPIPS is used to evaluate perceptual similarity, FID is the distance between the generated image distribution and the true distribution, and CLIPIQA is used to evaluate image quality.
[0122] Figure 3 The results of the comparison between the arbitrary-magnification super-resolution enhancement system and method based on the diffusion model of the embodiment of the present invention and other baseline models are shown, where 4×, 8×, and 16× are the results within the training distribution, and 18.2× is the case outside the training distribution. CIQA is the abbreviation of CLIPIQA. Figure 3 The best and second best results are marked in bold and underlined respectively. Figure 3It can be seen that the system\method of the embodiment of the present invention is significantly better than other methods in the case of large super-resolution ratios. The two regression-based methods LIIF and CiaoSR achieved the highest PSNR values, but their LPIPS, FID and CLIPIQA were far inferior to the method of the embodiment of the present invention, which shows that they cannot generate high-quality images that conform to the real image distribution and human eye perception. Most diffusion-based blind super-resolution methods perform well at 4× super-resolution, but as the super-resolution ratio gradually increases, the quality of images generated by IDM, StableSR and PASD decreases significantly, which is reflected in the obvious deterioration of LPIPS, FID and CLIPIQA indicators.
[0123] Figure 4 It is visualized that the system and method of the embodiment of the present invention have the ability to achieve super-resolution enhancement at any magnification. The top and bottom images show two arbitrary super-resolution scenes, respectively. The top image shows the ability to achieve super-resolution enhancement at any magnification with a fixed low-resolution image. The bottom image shows the ability to upsample different low-resolution images to a fixed-resolution high-resolution image, where the top row shows the result of interpolation and the bottom row shows the result of the system and method of the embodiment of the present invention. The figure in the upper right corner visualizes the comparison results between the system and method of the embodiment of the present invention and IDM. Obviously, it is significantly better than IDM in terms of the real naturalness of structure and texture. The image in the lower right corner shows the comparison results between the system and method of the embodiment of the present invention and SeeSR. Although SeeSR can generate rich details, it is significantly better than SeeSR in terms of semantic rationality and image fidelity. Figure 5 The comparison results with other baseline models at 16× super-resolution enhancement are shown. CiaoSR and StableSR cannot generate results that conform to the natural image distribution in large-scale super-resolution enhancement scenarios. The images generated by SeeSR have higher realism but lower fidelity. The system and method of the embodiments of the present invention can generate high-quality and high-fidelity images.
[0124] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0126] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0128] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0129] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. An arbitrary-magnification super-resolution enhancement system based on a diffusion model, characterized in that: include: Input module, given a low-resolution image of any size and any target super-resolution enhancement factor; An upsampling unit, which improves the low-resolution image to a target resolution by bicubic interpolation to obtain an intermediate target resolution image; An image encoder is used to encode the intermediate target resolution image to obtain latent space input features; A dual semantic guidance module, collecting text descriptions and semantic features of the intermediate target resolution image; A latent space module, which denoises the latent space input features based on the target super-resolution enhancement factor, the text description and the semantic features to obtain intermediate features; An image decoder decodes the intermediate features to obtain a final image that meets the target super-resolution enhancement ratio.
2. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 1, characterized in that: The latent space module adopts a two-branch network, including a low-resolution image branch and a diffusion model main branch; The low-resolution image branch extracts multi-scale features of the latent space input features by combining the text description, the semantic features and the target super-resolution enhancement factor; Among them, the main branch of the diffusion model combines the text description, the semantic features, the target super-resolution enhancement factor and random noise to extract the latent space input feature image features, and fuses them with the multi-scale features to obtain the fused intermediate features.
3. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 2, characterized in that: The low-resolution image branch includes a plurality of small groups, each of which includes a convolution layer, a text attention layer and a feature attention layer in sequence; the convolution kernel size of the convolution layer of each small group is different; wherein the convolution layer injects the target super-resolution enhancement factor into the latent space input feature or the output of the previous small group; the text attention layer injects the text description into the latent space input feature or the output of the previous small group; the feature attention layer injects the semantic feature into the latent space input feature or the output of the previous small group; and the multi-scale feature of the latent space input feature is obtained through the outputs of the plurality of small groups; The main branch of the diffusion model includes a symmetrical diffusion encoder and a diffusion decoder, and the diffusion encoder and the diffusion decoder respectively include multiple groups, each of which includes a convolution layer, a text attention layer and a feature attention layer in sequence; the convolution kernel size of the convolution layer of each group is different; the diffusion encoder takes the noise, the text description, the semantic feature and the target super-resolution enhancement ratio as input, gradually reduces the width and height of the noise feature, and increases the number of channels; the decoder takes the text description, the semantic feature, the target super-resolution enhancement ratio and the multi-scale feature as input, gradually restores the original width and height of the noise feature, and reduces the number of channels to obtain the fused intermediate feature.
4. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 2, characterized in that: It also includes a global rate modulation unit, which modulates the time step characteristics of the main branch of the diffusion model based on the target super-resolution enhancement rate; The global rate modulation unit comprises: A first encoder is used to perform position encoding on the target super-resolution enhancement factor; A first linear layer performs linear processing on the position code to obtain a linear processing result; The first activation layer performs activation processing on the linear processing result to obtain a super-resolution magnification feature; The updating unit adds the super-resolution ratio feature to the time step feature of the main branch of the diffusion model to update the original time step feature.
5. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 2, characterized in that: It also includes a local rate modulation unit, which modulates the multi-scale low-resolution image features extracted by the low-resolution image conditional branch based on the target super-resolution enhancement rate; The local rate modulation unit comprises: A second encoder performs position encoding on the target super-resolution enhancement factor; A second linear layer performs linear processing on the position code to obtain a linear processing result; The second activation layer performs activation processing on the linear processing result to obtain a super-resolution magnification feature; The third linear layer performs linear processing on the super-resolution magnification feature to obtain a scale parameter and a bias parameter; The modulation unit multiplies the multi-scale low-resolution image features by the scale parameters correspondingly, and then adds the multi-scale low-resolution image features by the bias parameters correspondingly to obtain the modulated features.
6. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 2, characterized in that: It includes a first training module, which trains the diffusion model main branch network; The first training module includes: The data submodule, given a high-resolution image, obtains the latent space representation z0 after passing through the main branch network of the diffusion model; The noise adding submodule gradually adds noise to the latent space representation z0 to obtain the noisy latent space representation z t ; Training submodule, combined with low-resolution image I lr , target super-resolution enhancement factor s, diffusion model time step t, text description c, semantic feature p, training diffusion model main branch network ∈ θ To predict the latent space representation z added to the noise t The noise in the training submodule is: in Represents the calculation prediction noise ∈ θ The L2 loss of the real noise ∈ is used as the diffusion loss, N represents the uniformly distributed random noise distribution, and ε~N represents sampling a noise sample from the noise distribution N.
7. The arbitrary magnification super-resolution enhancement system based on a diffusion model according to claim 6, characterized in that: A second training module is included, which trains the image encoder; the loss function of the second training module is a reconstruction loss function, and the reconstruction loss preliminarily reconstructs the structural detail information by constraining the intermediate features of the image encoder, and its specific expression is: in Represents the nth multi-scale feature extracted by the image encoder, toRGB means using convolution to convert the feature into an RGB image, It represents downsampling the real high-resolution image to the resolution of the nth feature.
8. A method for arbitrary magnification super-resolution enhancement based on a diffusion model, characterized in that: include: Given a low-resolution image of any size and any target super-resolution enhancement factor; The low-resolution image is enhanced to a target resolution by bicubic interpolation to obtain an intermediate target resolution image; Encoding the intermediate target resolution image to obtain latent space input features; Collecting text description and semantic features of the intermediate target resolution image; Based on the target super-resolution enhancement factor, the text description and the semantic feature, denoising the latent space input feature to obtain an intermediate feature; The intermediate features are decoded to obtain a final image that meets the target super-resolution enhancement ratio.
9. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it can be used to run the system described in any one of claims 1 to 7, or to execute the method described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it can be used to run the system described in any one of claims 1 to 7, or to execute the method described in claim 8.
Citation Information
Patent Citations
Method and device for determining a high resolution output image
CN105793891A
Image super-resolution method, electronic equipment and storage medium
CN118521476A
Real world image super-resolution method based on stable diffusion
CN118918009A
Hyperspectral image super-resolution method based on global guide condition diffusion model
CN119027317A
Image super-resolution generation model construction method and image super-resolution generation method and system
CN119205515A
Cited By
Image preprocessing method, system and equipment based on OCR (Optical Character Recognition), medium and product
CN120236285A