Image Super-Resolution Enhancement Diffusion Method and System Based on a Reference Information Database

Through the image super-resolution enhanced diffusion method based on the reference information database, using patch matching and multi-scale feature alignment technology, combined with the pre-trained diffusion model, the problems of poor image super-resolution generalization effect and difficulty in recovering real image details in the prior art are solved, and efficient and high-speed image super-resolution recovery is achieved.

CN120031722BActive Publication Date: 2025-07-01XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510495947.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-01
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing image super-resolution technology has poor generalization effect in complex and diverse degraded real low-resolution images, and the existing RefSR methods are difficult to recover real image details, and the acquisition of related reference images is high.

Method used

Using the image super-resolution enhanced diffusion method based on the reference information database, high-resolution image patches are obtained from the reference information database through a patch matching pipeline, and the multi-scale feature transfer and aggregation control module are used to align and fuse the features, and image recovery is carried out in combination with a pre-trained diffusion model.

Benefits of technology

It achieves advanced performance in quantitative and qualitative evaluation on multiple benchmarks, and can recover high-resolution images from low-resolution images with high speed and high accuracy, improving the visual quality and detail recovery capabilities of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031722B_ABST
    Figure CN120031722B_ABST
Patent Text Reader

Abstract

The present invention discloses an image super-resolution enhanced diffusion method and system based on a reference information database, belonging to the technical field of digital image processing. The method includes: creating a reference information database and a patch matching pipeline to obtain corresponding high-resolution image patches; using a multi-scale feature transfer and aggregation control module to control the generation ability of the diffusion process and help align the input features and reference features; adopting pre-trained stable diffusion, and under the guidance of fused multi-scale features, restoring the noise input into a high-resolution image. This method obtains a series of highly relevant reference patches from the reference information database, then aligns and fuses them through the multi-scale feature transfer and aggregation control module. The output features are used as guidance in the subsequent diffusion restoration process, achieving advanced performance in quantitative and qualitative evaluations on multiple benchmarks.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Super-Resolution reconstruction is a process of improving the resolution of original image data through hardware or software methods, and obtaining a high-resolution image data from a series of low-resolution image data. Single Image Super Resolution (SISR) aims to recover a high-resolution (HR) image from a low-resolution (LR) image, which has attracted extensive attention from all walks of life due to its great practical value. However, due to the unknown degradation in real scenes, SISR is a highly ill-posed problem. Convolutional Neural Networks (CNNs) and transformers are two main research methods, which have made remarkable progress in recent years. Although these methods have achieved remarkable performance on synthetic datasets by training on predefined degraded data, their generalization effect is very poor in real low-resolution images with complex and diverse degradations.

[0003] In addition, in real-world scenarios, there is a weak correlation between the two pixelization accuracy metrics, Peak Signal-to-Noise Ratio (PSNR) and the evaluation metric designed based on the sensitivity of the human eye to image structure information, Structural Similarity Index Measure (SSIM), and human perception of image quality. It should be noted that restoring the high fidelity (i.e., pixelization accuracy) of images is of great significance in research fields such as medical imaging processes or remote sensing. However, for other cases, such as personal use and daily applications, the visual quality perceived by humans has a higher priority than the accuracy perceived by pixels. To obtain better visual effects with rich textures, pre-trained Generative Adversarial Networks (GANs) can be used to improve the visual quality of super-resolution, but due to the inherent limitations of GANs, these methods have unrealistic textures. In recent years, diffusion-based models, which have the ability to recover images from noise, have shown great potential in various content creation and recovery tasks such as image synthesis, image editing, and image super-resolution.

[0004] On the other hand, to construct an effective reconstruction process from low-resolution (LR) images to high-resolution (HR) images, a series of reference-based super-resolution (RefSR) methods have been proposed, and high-quality reference images and corresponding texture transfer schemes have been introduced. Due to the additional high-frequency information, RefSR shows better results than single-image super-resolution (SISR). However, there are still some difficulties and problems that limit its performance and application. First, most existing RefSR methods are oriented towards peak signal-to-noise ratio (PSNR) and cannot recover real image details. In addition, few works provide methods to obtain paired reference images with the input, which limits their application value. Most existing methods only focus on transferring high-frequency information from well-prepared reference images to low-resolution images, but hardly mention how to obtain relevant reference images. Moreover, in the latest RefSR dataset (Large-scale Multi-Reference, LMR), the transformed high-resolution reference images (HR-Refs) are collected from exactly the same scene under different lighting conditions, weather, or shooting angles. In the practical application of RefSR, it is difficult and very costly to find a reference image with exactly the same scene as the LR image input. The most economical and practical method is to obtain reference images with similar semantic information. In addition, different regions of the low-resolution image require different high-frequency information, and one or a few reference images cannot meet this requirement. However, existing RefSR methods only consider obtaining relevant textures and content from multiple images, rather than obtaining a series of patch-level references from a large database. From another perspective, directly adding a large number of reference images to the model during the inference process brings a huge computational cost.

[0005] In recent years, diffusion models have developed rapidly due to their impressive content generation capabilities. Diffusion-based image super-resolution methods have significantly improved the quality of low-resolution inputs by generating realistic photo details. However, without detailed information guidance, key semantic details may be missed or inaccurate textures may be introduced during the diffusion process. Summary of the Invention

[0006] The present invention aims at the above deficiencies in the prior art and provides an image super-resolution enhanced diffusion method and system based on a reference information database, which can recover high-resolution images from low-resolution images at high speed and with high precision, and is used to solve the technical problems of image super-resolution.

[0007] The present invention adopts the following technical solutions:

[0008] An image super-resolution enhanced diffusion method based on a reference information database, comprising the following steps:

[0009] Construct a reference information database, and obtain corresponding high-resolution image patches from the reference information database using a patch matching pipeline;

[0010] Build a multi-scale feature transfer and aggregation control module, align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patches, and obtain a low-resolution image with aligned reference features;

[0011] Based on the overall learning objective L, use a pre-trained diffusion model to restore the low-resolution image with aligned reference features to a high-resolution image, and complete the image super-resolution enhanced diffusion.

[0012] Preferably, the constructing a reference information database and obtaining corresponding high-resolution image patches from the reference information database using a patch matching pipeline is specifically as follows:

[0013] Randomly select several images from each category of the dataset, and then form a reference information database through cropping and padding. Crop the low-resolution images in the reference information database into first patch vectors of a set size;

[0014] Use a pre-trained residual network to transform the cropped first patches into the vector space;

[0015] Calculate the vector similarity matrix between the patch vectors of the low-resolution image and the first patch vectors of the reference images in the reference information database based on the obtained vector space;

[0016] Locate the corresponding second patch vectors in the vector latent space using the vector similarity matrix, and use the second patches as the high-resolution image patches.

[0017] Preferably, the vector similarity matrix between the patch vectors of the low-resolution image and the first patch vectors of the reference images in the reference information database is calculated as follows:

[0018]

[0019] where, represents the patch vector of the low-resolution image, represents the first patch vector, ,[[]]END]] are both constants.

[0020] Preferably, the dataset uses the ImageNet dataset.

[0021] Preferably, the randomly selecting several images from each category of the dataset is specifically as follows:

[0022] Randomly select several low-resolution images with a size of 512×512 pixels from each category of the ImageNet dataset.

[0023] Preferably, a multi-scale feature transfer and aggregation control module is constructed to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch. Specifically:

[0024] Use two trainable multi-scale encoders to extract multi-scale features from the low-resolution image and the reference information in the reference information database respectively; align and fuse the obtained multi-scale features with the intermediate diffusion result at time step T to generate aligned multi-scale features of the low-resolution image as the low-resolution image with reference features aligned.

[0025] Preferably, based on the overall learning objective L, a pre-trained diffusion model is used to restore the low-resolution image with reference features aligned to a high-resolution image, specifically as follows:

[0026] In the forward process of the pre-trained diffusion model, Gaussian noise is added to the low-resolution image with reference features aligned using the pre-trained diffusion model; in the reverse process of the pre-trained diffusion model, the low-resolution image with reference features aligned is restored to a high-resolution image through the inference model.

[0027] Preferably, the forward process of the pre-trained diffusion model is specifically as follows:

[0028]

[0029] where t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, is the input, is the initial value, is the identity matrix;

[0030] The reverse process of the pre-trained diffusion model is specifically as follows:

[0031]

[0032] where, is the reverse process of the diffusion model, is the mean, is the variance, is the input, is the input at the previous moment, is the input The mean of the Gaussian distribution.

[0033] Preferably, the overall learning objective is as follows:

[0034]

[0035] Wherein, is the mathematical expectation, is the target image of a latent space, is the number of times of adding Gaussian noise, indicates that the cropped text prompt is set to empty, and are the low-resolution image and the reference image input respectively, is the conditional denoising network in the diffusion model, is the noise that follows the Gaussian distribution sampled from the Gaussian noise, is the normal distribution.

[0036] In a second aspect, an image super-resolution enhanced diffusion system based on a reference information database provided by an embodiment of the present invention includes:

[0037] A data module that makes a reference information database and obtains corresponding high-resolution image patches from the reference information database using a patch matching pipeline;

[0038] An alignment module that constructs a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patches to obtain a low-resolution image with aligned reference features;

[0039] An output module that, based on the overall learning objective L, uses a pre-trained diffusion model to restore the low-resolution image with aligned reference features to a high-resolution image, completing image super-resolution enhanced diffusion.

[0040] Compared with the prior art, the present invention has at least the following beneficial effects:

[0041] An image super-resolution enhanced diffusion method based on a reference information database, which adopts a patch matching (CMPM, Class-based Multi-Reference Patch Matching) pipeline to obtain a series of highly relevant image patches from a reference information database (Routing Information Base, RIB). The reference information database contains millions of high-quality image patches, input image patches and low-resolution images, and then aligns and fuses them through a well-designed multi-scale feature transfer and aggregation control module (MFTA, Multi-Scale Feature Transfer and Aggregation) for guidance during the diffusion recovery process; the method of the present invention has achieved advanced performance in terms of quantitative and qualitative evaluations on multiple benchmarks.

[0042] Further, the reference information database aligns the image scales by randomly selecting several images from each category of the ImageNet dataset and then applying cropping and padding to obtain an image size of 512×512 pixels for convenient further processing.

[0043] Further, the patch matching pipeline obtains corresponding high-resolution image patches from the reference information database. Under the condition of low computational cost and time consumption, it obtains corresponding high-resolution image patches from the reference information database to prepare for the next multi-scale feature transfer and aggregation.

[0044] Further, two trainable multi-scale encoders extract multi-scale features from the low-resolution image and the reference information; the extracted multi-scale features are aligned and fused with the intermediate diffusion result at time step T using multi-scale alignment, and aligned multi-scale features are generated of the low-resolution image to guide the next stable diffusion process.

[0045] Further, based on the overall learning objective L, the time for training the neural network is reduced, accelerating the entire image super-resolution process.

[0046] It can be understood that the beneficial effects of the second aspect above can be referred to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0047] In summary, the present invention improves the quality of super-resolution images by using a reference information database and has achieved advanced performance in terms of quantitative and qualitative evaluations on multiple benchmarks.

[0048] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments of the present application. Obviously, the accompanying drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart for the embodiment of the present invention to restore an image according to the image super-resolution enhancement diffusion method based on a reference information database;

[0051] Figure 2 It is an effect diagram for the embodiment of the present invention to restore an image according to the image super-resolution enhancement diffusion method based on a reference information database. Among them, (a) is the LR image before the experiment, and (b) is the HR image after the experiment. Detailed implementation manners

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0053] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0054] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0055] It should be further understood that the term " / and" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present invention generally represents an "or" relationship between the preceding and following related objects.

[0056] It should be understood that although terms such as first, second, and third may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0057] Depending on the context, as used herein, the word "if" can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0058] Schematic diagrams of various structures according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. And those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0059] The present invention provides an image super-resolution enhanced diffusion method based on a reference information database, making a large-scale reference information database and an effective patch matching pipeline to obtain corresponding high-quality image patches; using a multi-scale feature transfer and aggregation module to control the generation ability of the diffusion process and help align the input features and reference features; adopting pre-trained stable diffusion, under the guidance of fusing multi-scale features, restoring the noise input into a high-quality image. The method of the present invention achieves advanced performance in quantitative and qualitative evaluations on multiple benchmarks, aiming to improve the quality of super-resolution images by using auxiliary reference information and solve the problem of image distortion.

[0060] Embodiment 1

[0061] Please refer to Figure 1 , an image super-resolution enhanced diffusion method based on a reference information database of the present invention, comprising the following steps:

[0062] S1. Make a large-scale reference information database and an effective patch matching pipeline to obtain corresponding high-resolution image patches;

[0063] The reference information database randomly selects 100 images from each category of the ImageNet dataset, a total of 100K, and then applies cropping and padding to obtain images with a size of 512×512 pixels.

[0064] The effective patch matching pipeline is a class-based multi-reference patch matching pipeline that obtains corresponding high-resolution image patches from a large-scale reference information database with low computational cost and time consumption.

[0065] The steps of obtaining corresponding high-resolution image patches from the reference information database using the patch matching pipeline are as follows:

[0066] S101. Crop the low-resolution image into a first patch vector with a size of 128×128 pixels;

[0067] S102. Use a pre-trained ResNet (Residual Network) to transform the first patch vector into a vector space;

[0068] S103. Use the cosine formula to calculate the vector similarity matrix between the patch vector of the input low-resolution image and the first patch vector of the reference image in the reference information database ;

[0069] Vector similarity matrix The calculation is as follows:

[0070]

[0071] Among them, represents the input patch vector of the low-resolution image, represents the reference patch vector, , are all constants.

[0072] S104. Use the vector similarity matrix to locate the corresponding second patch in the vector space , and use the second patch as the high-resolution image patch.

[0073] S2. Use the multi-scale feature transfer and aggregation control module to control the diffusion process, help align the input features of the low-resolution image and the reference features in the high-resolution image patch, and generate aligned multi-scale features of the low-resolution image;

[0074] The multi-scale feature transfer and aggregation control module includes two parts:

[0075] Two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information respectively;

[0076] A multi-scale alignment module for aligning and fusing the extracted multi-scale features with the intermediate diffusion results at time step T, and generating aligned multi-scale features 。

[0077] Under normal circumstances, generating features at four scales makes it easier to control the diffusion process.

[0078] Since the goal of the MFTA module is to generate features for controlling the diffusion process, according to the Neural Network ControlNet settings, the trainable multi-scale encoder is copied from the diffusion UNet encoder in Stable Diffusion and loaded with pre-trained corresponding weights for fast training.

[0079] ‌The Neural Network ControlNet‌ is an extended model of Stable Diffusion. Stable Diffusion is a diffusion model based on the latent variable model, which intervenes in the image generation process by introducing additional conditions. For example, Canny edge detection can be used to control the contour of the image, or OpenPose can be used to transfer human pose features. These conditions can extract the features of the reference image through a pre-processor and act on the target image generation process together with the text conditions.

[0080] The Neural Network ControlNet aims to add spatial condition control to large pre-trained text-to-image diffusion models. Its core idea is to copy the frozen original network blocks and connect them to the trainable copies using zero convolutional layers. The task goal is to enable users to precisely control the spatial layout, pose, shape, and form of the generated image by providing additional images (such as edge maps, pose skeletons, segmentation maps, depth maps, etc.), so as to express their visual inspiration more accurately.

[0081] S3. Adopt a pre-trained diffusion model to restore the low-resolution image with aligned reference features to a high-resolution image, completing image super-resolution enhanced diffusion.

[0082] The diffusion model is the Stable Diffusion 2.1 model‌; the Stable Diffusion 2.1 model‌ is a deep learning technology based on the diffusion model, mainly used to generate high-quality images. It generates images through a process of gradually adding and removing noise, combining a variational autoencoder and a U-Net network to achieve the mapping from the latent space to the image space.

[0083] The working principle of the Stable Diffusion 2.1 model is divided into two processes: the forward process and the reverse process

[0084] Forward process: Gradually reduce the high-dimensional image information to the low-dimensional latent space, and destroy the original image by adding noise.

[0085] Reverse process: Gradually restore the image from the latent space, and generate the final image by removing noise, as Figure 2 shown.

[0086] The core components of the Stable Diffusion 2.1 model include the variational autoencoder and U-Net; the variational autoencoder is used to encode the image into a latent representation; U-Net is used to decode from the latent space to the image space to achieve high-quality image generation.

[0087] The pre-trained diffusion model includes two processes:

[0088] During the forward process, noise is gradually added to the data, and during the reverse process, the data is denoised to generate the data.

[0089] The forward process is to iteratively add Gaussian noise to the input data, that is:

[0090]

[0091] where is a scalar function, by making close to zero, converges to .

[0092] The reverse process uses the standard Gaussian prior to gradually denoise the data from the Gaussian noise distribution to the target distribution. The reverse process of the diffusion model is:

[0093]

[0094] where is the mean value, is the variance.

[0095] Overall learning objective is as follows:

[0096]

[0097] where is the mathematical expectation, is a target image in the latent space, t represents the number of times noise is added, indicates that the cropped text prompt is set to empty, and are the low-resolution image and the reference image input (in the vector latent space) respectively, is the conditional denoising network in the diffusion model, is noise that is sampled from Gaussian noise and follows a Gaussian distribution.

[0098] Those skilled in the art can understand that various aspects of the present invention can be implemented as a system, method, or program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "platform" here.

[0099] Embodiment 2

[0100] The present invention provides an image super-resolution enhanced diffusion system based on a reference information database, which can be used to implement the above-mentioned image super-resolution enhanced diffusion method based on a reference information database. Specifically, the image super-resolution enhanced diffusion system based on a reference information database includes a data module, an alignment module, and an output module.

[0101] Among them, the data module creates a reference information database and obtains corresponding high-resolution image patches from the reference information database using a patch matching pipeline;

[0102] Randomly select several images from each category of the dataset, and then form a reference information database through cropping and padding. Crop the low-resolution images in the reference information database into first patch vectors of a set size;

[0103] Use a pre-trained residual network to transform the cropped first patch vectors into a vector space;

[0104] Calculate a vector similarity matrix between the patch vectors of the low-resolution image and the first patch vectors of the reference images in the reference information database based on the obtained vector space;

[0105] Locate the corresponding second patch vectors in the vector space using the vector similarity matrix, and use the second patch vectors as high-resolution image patches.

[0106] The vector similarity matrix between the patch vectors of the low-resolution image and the first patch vectors of the reference images in the reference information database is calculated as follows:

[0107]

[0108] Among them, represents the patch vector of the low-resolution image, represents the first patch vector, , are both constants.

[0109] Among them, the dataset uses the ImageNet dataset, and several low-resolution images with a size of 512×512 pixels are randomly selected from each category of the ImageNet dataset.

[0110] Alignment module, constructing a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch, obtaining a low-resolution image with aligned reference features;

[0111] Use two trainable multi-scale encoders to extract multi-scale features from the low-resolution image and the reference information in the reference information database respectively; align and fuse the obtained multi-scale features with the intermediate diffusion result at time step T to generate aligned multi-scale features .

[0112] Output module, based on the overall learning objective L, using a pre-trained diffusion model to restore the low-resolution image with aligned reference features to a high-resolution image, completing image super-resolution enhanced diffusion.

[0113] In the forward process of the pre-trained diffusion model, add Gaussian noise to the low-resolution image with aligned reference features using the pre-trained diffusion model; in the reverse process of the pre-trained diffusion model, restore the low-resolution image with aligned reference features to a high-resolution image through the inference model.

[0114] The forward process of the pre-trained diffusion model is specifically as follows:

[0115]

[0116] Among them, t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, is the input, is the initial value, is the identity matrix;

[0117] The reverse process of the pre-trained diffusion model is specifically as follows:

[0118]

[0119] Among them, is the mean value, is the variance, is the input, is the input at the previous moment, is the input the mean of the Gaussian distribution of.

[0120] Overall learning objective The details are as follows:

[0121]

[0122] Among them, is the mathematical expectation, is the target image of a latent space, is the number of times of adding Gaussian noise, indicates that the cropped text prompt is set to be empty, and are the low-resolution image and the reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and following a Gaussian distribution, is the normal distribution.

[0123] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0124] The proposed reference information database enhanced diffusion model is compared with the existing state-of-the-art SISR methods, namely the real super-resolution algorithm RealSR (ICCV conference paper in 2019), the real enhanced super-resolution generative adversarial network Real-ESRGAN+ (ICCV conference paper in 2021), the blind image super-resolution generative adversarial network BSRGAN (ICCV conference paper in 2021), the adaptive degradation super-resolution network DASR (ECCV conference paper in 2022), the feature matching super-resolution algorithm FeMaSR (ACCM conference paper in 2022), the image inpainting model SwinIR based on the sliding window transformer, the latent space diffusion model LDM (ICCV conference paper in 2021), and the stable super-resolution algorithm StableSR (IJCV conference paper in 2024).

[0125] To further verify the effectiveness of the model, a series of reference-based SR methods were also used to conduct extensive experiments, namely the image super-resolution algorithm SRNTT based on neural texture transfer (a paper published in the ICCV conference in 2019), the texture transformer network TTSR for image super-resolution (a paper published in the CVPR conference in 2020), the reference super-resolution algorithm MASA that matches acceleration and spatial adaptability (a paper published in the CVPR conference in 2021), the super-resolution algorithm DATSR of the reference image combined with deformable attention transformers (a paper published in the ECCV conference in 2022), the multi-reference attention super-resolution algorithm MRefSR (a paper published in the ICCV conference in 2023), and the reference image and video super-resolution algorithm c2-matching based on c2 matching (a paper published in the TPAMI conference in 2023).

[0126] Among them, PSNR: Peak Signal-to-Noise Ratio; SSIM: Structural Similarity; FID: Frechet Distance; LPIPS: Learned Perceptual Image Patch Similarity; MUSIQ: Multi-Scale Unified Image Quality Assessment; CLIP-IAQ: CLIP-based Image Quality Assessment; MANIQA: Multi-Modal Adaptive Image Quality Assessment; Ours: The method of the present invention.

[0127] Table 1 Quantitative results of SISR methods on synthetic datasets

[0128]

[0129] Table 2 Quantitative results of SISR methods on real datasets (RealSR, DRealSR, and RealSet300)

[0130]

[0131] To evaluate the performance of the method, extensive experiments were conducted on four synthetic datasets (ImageNet Test2000, Set5, Set14, and BSD100) and three real-world benchmarks (RealSR, DRealSR, and RealSet300). The numerical results are shown in Table 1 and Table 2. The best and second-best results in Table 1 are highlighted in bold and underlined respectively, and the best and second-best results in Table 2 are highlighted in bold and underlined respectively. It can be seen that the model proposed by the present invention is superior to the SISR method in multiple prediction metrics, demonstrating the effectiveness of the reference information database enhanced diffusion algorithm.

[0132] Specifically, the method of the present invention obtained the highest scores on all perceptual metrics on the synthetic datasets. On the DRealSR and RealSet300 datasets, the model of the present invention has superior performance on all reference-free perceptual metrics.

[0133] Table 3 Quantitative Results of the RefSR Method

[0134]

[0135]

[0136]

[0137] Extensive experiments were conducted on four datasets, namely WR-SR, Urban100, Manga109, and LMR, as shown in Table 3. The best and second-best results in Table 3 are highlighted in bold and underlined respectively. For each dataset, two degradation methods were used to obtain low-resolution image - high-resolution image pairs (Bicubic: bilateral downsampling, Real-ESRGAN: pipeline of enhanced super-resolution generative adversarial network). The reference information database enhanced diffusion proposed in the present invention always shows superior performance in almost all datasets and perceptual metrics, highlighting its robustness and superiority. Although some RefSR methods such as MASA, DATSR, and c2-matching have several good scores on the bilateral downsampling dataset, however, when the degradation of the low-resolution image is turned off and in the case of unknown degradation in real scenarios (downsampling pipeline Real-ESRGAN), their performance drops significantly.

[0138] A large number of experiments show that the performance of the method proposed in the present invention is superior to the current state-of-the-art SISR methods and RefSR methods. In addition, the success of the system and its components of the present invention also demonstrates the effectiveness of this reinforcement strategy, which has the potential to adapt to other downstream tasks.

[0139] In summary, an image super-resolution enhanced diffusion method and system based on a reference information database according to the present invention utilizes a class-based multi-reference patch matching pipeline and a reference information database to obtain a series of high-quality reference patches, which contain necessary high-frequency information and contribute to restoring the low-resolution image input. This system is the first method that can obtain relevant reference information from millions of high-quality reference patches. At the same time, a multi-scale feature transfer and aggregation control module is developed to align and fuse the information of the low-resolution image input, reference patches, and intermediate diffusion results at the current time step, and the obtained features are used as guiding information in the subsequent diffusion restoration process. Ablation studies verify the effectiveness of each proposed module.

[0140] The above content is only to illustrate the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. Image super-resolution enhancement diffusion method based on reference information database, characterized in that: The following steps are involved: Create a reference information database and use a patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; Construct a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features. Specifically, two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information in the reference information database respectively; the obtained multi-scale features are aligned and fused with the intermediate diffusion results at the time step T to generate aligned multi-scale features. The low-resolution image of is used as the low-resolution image for reference feature alignment; Based on the overall learning goal L, the pre-trained diffusion model is used to restore the low-resolution image with reference feature alignment to a high-resolution image to complete the image super-resolution enhancement diffusion, specifically: In the forward process of the pre-trained diffusion model, the pre-trained diffusion model is used to add Gaussian noise to the low-resolution image aligned with the reference features; in the reverse process of the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image through the inference model. The overall learning goal The details are as follows: in, is the mathematical expectation, is a target image in a latent space, is the number of times Gaussian noise is added, Indicates that the crop text hint is set to empty. and are low-resolution image and reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and follows Gaussian distribution, is a normal distribution.

2. The image super-resolution enhanced diffusion method based on a reference information database according to claim 1, characterized in that: The reference information database is prepared, and the corresponding high-resolution image patches are obtained from the reference information database using a patch matching pipeline, specifically: A number of images are randomly selected from each category of the data set, and then a reference information database is formed by cropping and padding, and the low-resolution images in the reference information database are cropped into a first patch vector of a set size; Convert the cropped first patch vector to vector space using a pre-trained residual network; Calculate a vector similarity matrix between a patch vector of the low-resolution image and a first patch vector of a reference image in a reference information database based on the obtained vector space; The vector similarity matrix is ​​used to locate the corresponding second patch vector in the vector space, and the second patch vector is used as the high-resolution image patch.

3. The image super-resolution enhanced diffusion method based on a reference information database according to claim 2, characterized in that: Vector similarity matrix between the patch vector of the low-resolution image and the first patch vector of the reference image in the reference information database The calculation is as follows: in, represents the patch vector of the low-resolution image, represents the first patch vector, , are all constants.

4. The image super-resolution enhanced diffusion method based on a reference information database according to claim 2, characterized in that: The dataset uses the ImageNet dataset.

5. The image super-resolution enhanced diffusion method based on a reference information database according to claim 4 is characterized in that: The method randomly selects several images from each category of the data set, specifically: Randomly select several low-resolution images of size 512×512 pixels from each class of the ImageNet dataset.

6. The image super-resolution enhanced diffusion method based on a reference information database according to claim 1, characterized in that: The forward process of the pre-trained diffusion model is as follows: Where t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, For input, is the initial value, is the identity matrix; The reverse process of the pre-trained diffusion model is as follows: in, is the reverse process of the diffusion model, is the average value, is the variance, For input, is the input of the previous moment, For input The mean of the Gaussian distribution.

7. An image super-resolution enhanced diffusion system based on a reference information database, characterized in that: include: Data module, making a reference information database and using patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; The alignment module constructs a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features. Specifically, two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information in the reference information database respectively; the obtained multi-scale features are aligned and fused with the intermediate diffusion results at the time step T to generate aligned multi-scale features. The low-resolution image of is used as the low-resolution image for reference feature alignment; The output module, based on the overall learning goal L, uses the pre-trained diffusion model to restore the low-resolution image aligned with the reference features to a high-resolution image, completing the image super-resolution enhancement diffusion, specifically: In the forward process of the pre-trained diffusion model, the pre-trained diffusion model is used to add Gaussian noise to the low-resolution image aligned with the reference features; in the reverse process of the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image through the inference model. The overall learning goal The details are as follows: in, is the mathematical expectation, is a target image in a latent space, is the number of times Gaussian noise is added, Indicates that the crop text hint is set to empty. and are low-resolution image and reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and follows Gaussian distribution, is a normal distribution.