Image super-resolution enhanced diffusion method and system based on reference information database

Through the image super-resolution enhancement diffusion method based on the reference information database, the problem of poor generalization effect in real scenes and difficulty in recovery of real image details by the prior art is solved, and high-speed and high-precision image super-resolution recovery is achieved, which improves visual quality and reduces calculation costs.

CN120031722AActive Publication Date: 2025-05-23XI AN JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510495947.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art has poor generalization effect in complex and diverse degraded low-resolution images in real scenes, and the existing RefSR methods are difficult to recover real image details, and the cost of obtaining related reference images is high, and different regions of low-resolution images require different high-frequency information, but the existing methods are difficult to meet this requirement.

Method used

Using the image super-resolution enhanced diffusion method based on the reference information database, acquiring high-resolution image patches by making a reference information database and using a patch matching pipeline, a multi-scale feature transfer and aggregation control module is constructed, the input features of the low-resolution image are aligned with the reference features in the high-resolution image patch, and image recovery is performed using a pre-trained diffusion model.

Benefits of technology

It realizes high-speed and high-precision recovery of high-resolution images from low-resolution images, improves the visual quality of image super-resolution, reduces computing costs, and can effectively meet the high-frequency information needs of different regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031722A_ABST
    Figure CN120031722A_ABST
Patent Text Reader

Abstract

The invention discloses an image super-resolution enhancement diffusion method and system based on a reference information database, and belongs to the technical field of digital image processing, and the method comprises the steps: making a reference information database and a patch matching pipeline, and obtaining a corresponding high-resolution image patch; a multi-scale feature transfer and aggregation control module is utilized to control the generation capability of the diffusion process, and the alignment of input features and reference features is facilitated; pre-trained stable diffusion is adopted, and noise input is recovered into a high-resolution image under the guidance of fusion of multi-scale features. According to the method, a series of highly related reference patches are obtained from a reference information database, then alignment and fusion are carried out through a multi-scale feature transfer and aggregation control module, output features serve as guidance in the subsequent diffusion recovery process, and advanced performance is achieved in the aspect of quantitative and qualitative evaluation of multiple standards.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Super-resolution reconstruction is the process of improving the resolution of the original image data through hardware or software methods, and obtaining a high-resolution image data through a series of low-resolution image data. Single Image Super Resolution (SISR) aims to restore high-resolution (HR) images from low-resolution (LR) images. Due to its great practical value, it has attracted widespread attention from all walks of life. However, due to the existence of unknown degradation in real scenes, SISR is a highly ill-posed problem. Convolutional Neural Networks (CNNs) and transformers are two major research methods that have made significant progress in recent years. Although these methods have achieved remarkable performance on synthetic datasets by training on predefined degraded data, their generalization effect is very poor in real low-resolution images with complex and diverse degradation.

[0003] In addition, in real-world scenarios, two pixelation accuracy metrics, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM), which is an evaluation metric designed based on the sensitivity of the human eye to image structural information, have a weak correlation with human perception of image quality. It should be noted that high fidelity of restored images (i.e., pixelation accuracy) is of great significance in research fields such as medical image processing or remote sensing. But for other situations, such as personal use and daily applications, human-perceived visual quality has a higher priority than pixel-perceived accuracy. In order to obtain better visual results with rich textures, pre-trained generative adversarial networks (GANs) can be used to improve the visual quality of super-resolution, but due to the inherent limitations of GANs, these methods have unrealistic textures. In recent years, diffusion-based models have the ability to recover images from noise and have shown great potential in various content creation and restoration tasks such as image synthesis, image editing, and image super-resolution.

[0004] On the other hand, in order to construct an effective reconstruction process from LR images to HR images, a series of reference-based image super-resolution (RefSR) methods have been proposed, and high-quality reference images and corresponding texture transfer schemes have been introduced. Due to the additional high-frequency information, RefSR shows good results compared to SISR. However, there are still some difficulties and problems that limit its performance and application. First, most of the existing RefSR methods are PSNR-oriented and cannot recover the real image details. In addition, few works provide methods to obtain paired reference images with inputs, which limits their application value. Most of the existing methods only focus on transferring high-frequency information from well-prepared reference images to low-resolution images, but almost no mention is made on how to obtain relevant reference images. In addition, in the latest RefSR dataset (Large-scale Multi-Reference, LMR), the transformed high-resolution reference image (HR-Ref) is collected from exactly the same scene and under different lighting conditions, weather or shooting angles. In the practical application of RefSR, it is difficult and very costly to find a reference image with exactly the same scene as the LR image input. The most economical and practical method is to obtain reference images with similar semantic information. In addition, different regions of low-resolution images require different high-frequency information, which cannot be met by one or two reference images. However, existing RefSR methods only consider obtaining relevant textures and content in multiple images, rather than obtaining a series of patch-level references from a large database. From another perspective, directly adding a large number of reference images to the model during inference brings huge computational costs.

[0005] In recent years, diffusion models have gained rapid development due to their impressive content generation capabilities. Diffusion-based image super-resolution methods significantly improve the quality of low-resolution inputs by generating photo-realistic details. However, without detailed information guidance, key semantic details may be missed or inaccurate textures may be introduced during the diffusion process. Summary of the invention

[0006] In view of the deficiencies in the above-mentioned prior art, the present invention provides an image super-resolution enhancement diffusion method and system based on a reference information database, which can restore a high-resolution image from a low-resolution image at high speed and high precision, and is used to solve the technical problem of image super-resolution.

[0007] The present invention adopts the following technical solutions: An image super-resolution enhancement diffusion method based on a reference information database comprises the following steps: Create a reference information database and use a patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; Constructing a multi-scale feature transfer and aggregation control module to align input features of the low-resolution image to be processed with reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features; Based on the overall learning objective L, a pre-trained diffusion model is used to restore the low-resolution image with reference feature alignment to a high-resolution image, completing image super-resolution enhancement diffusion.

[0008] Preferably, the making of the reference information database uses a patch matching pipeline to obtain corresponding high-resolution image patches from the reference information database, specifically: A number of images are randomly selected from each category of the data set, and then a reference information database is formed by cropping and padding, and the low-resolution images in the reference information database are cropped into a first patch vector of a set size; Convert the cropped first patch to vector space using a pre-trained residual network; Calculate a vector similarity matrix between a patch vector of the low-resolution image and a first patch vector of a reference image in a reference information database based on the obtained vector space; The vector similarity matrix is ​​used to locate the corresponding second patch vector in the vector latent space, and the second patch is used as the high-resolution image patch.

[0009] Preferably, the vector similarity matrix between the patch vector of the low-resolution image and the first patch vector of the reference image in the reference information database is The calculation is as follows:

[0010] in, represents the patch vector of the low-resolution image, represents the first patch vector, , are all constants.

[0011] Preferably, the dataset uses the ImageNet dataset.

[0012] Preferably, the randomly selecting a number of images from each category of the data set is specifically: Randomly select several low-resolution images of size 512×512 pixels from each class of the ImageNet dataset.

[0013] Preferably, a multi-scale feature transfer and aggregation control module is constructed to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch, specifically: Two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information of the reference information database respectively; the obtained multi-scale features are aligned and fused with the intermediate diffusion results at time step T to generate aligned multi-scale features The low-resolution image of is used as the low-resolution image for reference feature alignment.

[0014] Preferably, based on the overall learning goal L, a pre-trained diffusion model is used to restore the low-resolution image aligned with the reference features to a high-resolution image, as follows: In the forward process of the pre-trained diffusion model, the pre-trained diffusion model is used to add Gaussian noise to the low-resolution image aligned with the reference features; in the reverse process of the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image through the inference model.

[0015] Preferably, the forward process of the pre-trained diffusion model is as follows:

[0016] Where t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, For input, is the initial value, is the identity matrix; The reverse process of the pre-trained diffusion model is as follows:

[0017] in, is the reverse process of the diffusion model, is the average value, is the variance, For input, is the input of the previous moment, For input The mean of the Gaussian distribution.

[0018] Preferably, the overall learning objectives The details are as follows:

[0019] in, is the mathematical expectation, is a target image in a latent space, is the number of times Gaussian noise is added, Indicates that the crop text hint is set to empty. and are low-resolution image and reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and follows Gaussian distribution, is a normal distribution.

[0020] In a second aspect, an embodiment of the present invention provides an image super-resolution enhanced diffusion system based on a reference information database, comprising: Data module, making a reference information database and using patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; An alignment module constructs a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features; The output module, based on the overall learning objective L, uses the pre-trained diffusion model to restore the low-resolution image with reference feature alignment to a high-resolution image, completing the image super-resolution enhancement diffusion.

[0021] Compared with the prior art, the present invention has at least the following beneficial effects: An image super-resolution enhanced diffusion method based on a reference information database adopts a patch matching (CMPM, Class-based Multi-Reference Patch Matching) pipeline to obtain a series of highly correlated image patches from a reference information database (RoutingInformation Base, RIB), which contains millions of high-quality image patches. The input image patches and low-resolution images are then aligned and fused through a well-designed multi-scale feature transfer and aggregation control module (MFTA, Multi-Scale Feature Transfer and Aggregation) for guidance in the diffusion restoration process. The method of the invention achieves advanced performance in quantitative and qualitative evaluations of multiple benchmarks.

[0022] Furthermore, the reference information database aligns the image scale by randomly selecting several images from each category of the ImageNet dataset and then applying cropping and padding to obtain an image size of 512×512 pixels for further processing.

[0023] Furthermore, the patch matching pipeline obtains the corresponding high-resolution image patches from the reference information database with low computational cost and time consumption, preparing for the next step of multi-scale feature transfer and aggregation.

[0024] Furthermore, two trainable multi-scale encoders extract multi-scale features from the low-resolution image and the reference information; the extracted multi-scale features are aligned and fused with the intermediate diffusion results at the time step T using multi-scale alignment, and aligned multi-scale features are generated. The low-resolution image is used to guide the next step of stable diffusion process.

[0025] Furthermore, based on the overall learning objective L, the time for training the neural network is reduced, accelerating the entire image super-resolution process.

[0026] It can be understood that the beneficial effects of the second aspect mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0027] In summary, the present invention improves the quality of super-resolution images by utilizing a reference information database, achieving advanced performance in quantitative and qualitative evaluations of multiple benchmarks.

[0028] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1 This is a flow chart of restoring an image according to an image super-resolution enhancement diffusion method based on a reference information database according to an embodiment of the present invention; Figure 2 This is a diagram showing the image restoration effect according to the image super-resolution enhancement diffusion method based on a reference information database in an embodiment of the present invention, wherein (a) is the LR image before the experiment, and (b) is the HR image after the experiment. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.

[0033] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0034] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.

[0035] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0036] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0037] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0038] The present invention provides an image super-resolution enhancement diffusion method based on a reference information database, which makes a large-scale reference information database and an effective patch matching pipeline to obtain corresponding high-quality image patches; uses a multi-scale feature transfer and aggregation module to control the generation capacity of the diffusion process to help align input features with reference features; uses pre-trained stable diffusion to restore noisy input to a high-quality image under the guidance of fused multi-scale features. The method of the present invention achieves advanced performance in quantitative and qualitative evaluation of multiple benchmarks, and aims to improve the quality of super-resolution images and solve the problem of image distortion by using auxiliary reference information.

[0039] Example 1 See also Figure 1 The present invention provides an image super-resolution enhancement diffusion method based on a reference information database, comprising the following steps: S1. Create a large-scale reference information database and an efficient patch matching pipeline to obtain corresponding high-resolution image patches. The reference information database is constructed by randomly selecting 100 images from each class of the ImageNet dataset, a total of 100K images, and then applying cropping and padding to obtain images of size 512×512 pixels.

[0040] The effective patch matching pipeline is a class-based multi-reference patch matching pipeline that obtains corresponding high-resolution image patches from a large-scale reference information database with low computational cost and time consumption.

[0041] Use the patch matching pipeline to obtain the corresponding high-resolution image patch from the reference information database; the steps are as follows: S101, cropping the low-resolution image into a first patch vector of 128×128 pixels; S102, using a pre-trained ResNet (Residual Network) to convert the first patch vector into a vector space; S103, using the cosine formula to calculate the vector similarity matrix between the patch vector of the input low-resolution image and the first patch vector of the reference image in the reference information database ; Vector Similarity Matrix The calculation is as follows:

[0042] in, represents the low-resolution image input patch vector, denotes the reference patch vector, , are all constants.

[0043] S104, using vector similarity matrix Locate the corresponding second patch in vector space , the second patch As high-resolution image patches.

[0044] S2. Use the multi-scale feature transfer and aggregation control module to control the diffusion process, help align the low-resolution image input features and the reference features in the high-resolution image patch, and generate aligned multi-scale features. Low-resolution images; The multi-scale feature transfer and aggregation control module consists of two parts: Two trainable multi-scale encoders for extracting multi-scale features from low-resolution images and reference information, respectively; A multi-scale alignment module is used to align and fuse the extracted multi-scale features with the intermediate diffusion results at time step T and generate aligned multi-scale features .

[0045] Typically, generating features at four scales makes it easier to control the diffusion process.

[0046] Since the goal of the MFTA module is to generate features that control the diffusion process, according to the neural network ControlNet setting, the trainable multi-scale encoder is copied from the diffusion UNet encoder in stable diffusion and loaded with pre-trained corresponding weights for fast training.

[0047] ‌Neural Network ControlNet‌ is an extended model of Stable Diffusion, which is a diffusion model based on latent variable model. It intervenes in the image generation process by introducing additional conditions. For example, you can use Canny edge detection to control the outline of the image, or use OpenPose to transfer the character pose features. These conditions can extract the features of the reference image through the preprocessor and work together with the text conditions in the target image generation process.

[0048] The neural network ControlNet aims to add spatial conditioning control to large pre-trained text-to-image diffusion models. The core idea is to copy the frozen original network block and connect it with the trainable copy using zero convolutional layers. The task goal is to enable users to precisely control the spatial layout, pose, shape and form of the generated images by providing additional images (such as edge maps, pose skeletons, segmentation maps, depth maps, etc.), so as to more accurately express their visual inspiration.

[0049] S3. Using the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image to complete the image super-resolution enhancement diffusion.

[0050] The diffusion model is the Stable Diffusion 2.1 model; the Stable Diffusion 2.1 model is a deep learning technology based on the diffusion model, which is mainly used to generate high-quality images. It generates images by gradually adding and removing noise, combining the variational autoencoder and the U-Net network to achieve mapping from the latent space to the image space.

[0051] The working principle of the Stable Diffusion 2.1 model is divided into two processes: forward process and reverse process: Forward process: gradually reduce the high-dimensional image information to a low-dimensional latent space and destroy the original image by adding noise.

[0052] Reverse process: gradually recover the image from the latent space and generate the final image by removing noise, such as Figure 2 shown.

[0053] The core components of the Stable Diffusion 2.1 model include a variational autoencoder and U-Net; the variational autoencoder is used to encode images into latent representations; U-Net is used to decode from the latent space to the image space to achieve high-quality image generation.

[0054] The pre-trained diffusion model consists of two processes: In the forward process, noise is gradually added to the data, and in the backward process, the generated data is denoised.

[0055] The forward process is to iteratively add Gaussian noise to the input data, that is:

[0056] in, is a scalar function, by letting Close to zero, Converge to .

[0057] The reverse process is to gradually denoise the data from the Gaussian noise distribution to the target distribution by using the standard Gaussian prior, the reverse process of the diffusion model for:

[0058] in, is the average value, is the variance.

[0059] Overall learning objectives as follows:

[0060] in, is the mathematical expectation, is a target image in a latent space, t represents the number of times noise is added, Indicates that the crop text hint is set to empty. and are the low-resolution image and reference image input (in vector latent space), respectively. is the conditional denoising network in the diffusion model, is a Gaussian-distributed noise sampled from Gaussian noise.

[0061] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".

[0062] Example 2 The present invention provides an image super-resolution enhancement diffusion system based on a reference information database, which can be used to implement the above-mentioned image super-resolution enhancement diffusion method based on a reference information database. Specifically, the image super-resolution enhancement diffusion system based on a reference information database includes a data module, an alignment module and an output module.

[0063] Among them, the data module makes a reference information database and uses a patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; A number of images are randomly selected from each category of the data set, and then a reference information database is formed by cropping and padding, and the low-resolution images in the reference information database are cropped into a first patch vector of a set size; Convert the cropped first patch vector to vector space using a pre-trained residual network; Calculate a vector similarity matrix between a patch vector of the low-resolution image and a first patch vector of a reference image in a reference information database based on the obtained vector space; The vector similarity matrix is ​​used to locate the corresponding second patch vector in the vector space, and the second patch vector is used as the high-resolution image patch.

[0064] Vector similarity matrix between the patch vector of the low-resolution image and the first patch vector of the reference image in the reference information database The calculation is as follows:

[0065] in, represents the patch vector of the low-resolution image, represents the first patch vector, , are all constants.

[0066] The dataset uses the ImageNet dataset, and several low-resolution images of 512×512 pixels are randomly selected from each category of the ImageNet dataset.

[0067] An alignment module constructs a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features; Two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information of the reference information database respectively; the obtained multi-scale features are aligned and fused with the intermediate diffusion results at time step T to generate aligned multi-scale features .

[0068] The output module, based on the overall learning objective L, uses the pre-trained diffusion model to restore the low-resolution image with reference feature alignment to a high-resolution image, completing the image super-resolution enhancement diffusion.

[0069] In the forward process of the pre-trained diffusion model, the pre-trained diffusion model is used to add Gaussian noise to the low-resolution image aligned with the reference features; in the reverse process of the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image through the inference model.

[0070] The forward process of the pre-trained diffusion model is as follows:

[0071] Where t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, For input, is the initial value, is the identity matrix; The reverse process of the pre-trained diffusion model is as follows:

[0072] in, is the average value, is the variance, For input, is the input of the previous moment, For input The mean of the Gaussian distribution.

[0073] Overall learning objectives The details are as follows:

[0074] in, is the mathematical expectation, is a target image in a latent space, is the number of times Gaussian noise is added, Indicates that the crop text hint is set to empty. and are low-resolution image and reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and follows Gaussian distribution, is a normal distribution.

[0075] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can usually be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0076] The proposed reference information database enhanced diffusion model is compared with existing state-of-the-art SISR methods, namely, the real super-resolution algorithm RealSR (ICCV conference paper in 2019), the real enhanced super-resolution generative adversarial network Real-ESRGAN+ (ICCV conference paper in 2021), the blind image super-resolution generative adversarial network BSRGAN (ICCV conference paper in 2021), the adaptive degradation super-resolution network DASR (ECCV conference paper in 2022), the feature matching super-resolution algorithm FeMaSR (ACCM conference paper in 2022), the image inpainting model SwinIR based on the sliding window converter, the latent space diffusion model LDM (ICCV conference paper in 2021) and the stable super-resolution algorithm StableSR (IJCV conference paper in 2024).

[0077] To further verify the effectiveness of the model, a series of reference-based SR methods were also used for extensive experiments, namely, the image super-resolution algorithm SRNTT based on neural texture transfer (ICCV conference paper in 2019), the texture transformer network TTSR for image super-resolution (CVPR conference paper in 2020), the reference super-resolution algorithm MASA with matching acceleration and spatial adaptation (CVPR conference paper in 2021), the reference image super-resolution algorithm DATSR combined with deformable attention transformer (ECCV conference paper in 2022), the multi-reference attention super-resolution algorithm MRefSR (ICCV conference paper in 2023), and the reference image and video super-resolution algorithm c2-matching based on c2 matching (TPAMI conference paper in 2023).

[0078] Among them, PSNR: peak signal-to-noise ratio; SSIM: structural similarity; FID: Fréchet distance; LPIPS: learning perceptual image block similarity; MUSIQ: multi-scale unified image quality assessment; CLIP-IAQ: CLIP-based image quality assessment; MANIQA: multi-modal adaptive image quality assessment; Ours: the method of the present invention.

[0079] Table 1 Quantitative results of SISR methods on synthetic datasets

[0080] Table 2 Quantitative results of SISR methods on real datasets (RealSR, DRealSR and RealSet300)

[0081] To evaluate the performance of the method, extensive experiments are conducted on four synthetic datasets (ImageNet Test2000, Set5, Set14 and BSD100 and three real benchmarks (RealSR, DRealSR and RealSet300). The numerical results are shown in Tables 1 and 2. The best and second best results are highlighted in bold and underlined in Table 1, respectively, and the best and second best results are highlighted in bold and underlined in Table 2. It can be seen that the proposed model outperforms the SISR method in multiple prediction metrics, which proves the effectiveness of the reference information database enhanced diffusion algorithm.

[0082] Specifically, our method achieves the highest scores on all perceptual metrics on synthetic datasets. On DRealSR and RealSet300 datasets, our model has superior performance on all no-reference perceptual metrics.

[0083] Table 3 Quantitative results of RefSR method

[0084]

[0085]

[0086] Extensive experiments are conducted on four datasets, WR-SR, Urban100, Manga109 and LMR, as shown in Table 3. The best and second best results in Table 3 are highlighted in bold and underline respectively. For each dataset, two degradation methods are used to obtain low-resolution image-high-resolution image pairs (Bicubic: bilateral downsampling, Real-ESRGAN: pipeline of enhanced super-resolution adversarial generative network). The proposed reference information database enhanced diffusion consistently shows superior performance on almost all datasets and perceptual indicators, highlighting its robustness and superiority. Although some RefSR methods such as MASA, DATSR and c2-matching perform well on the dual-edge downsampled dataset with several scores, their performance drops significantly when the degradation of low-resolution images is turned off and when the degradation is unknown in real scenes (downsampling pipeline Real-ESRGAN).

[0087] Extensive experiments show that the proposed method outperforms the most advanced SISR and RefSR methods. In addition, the success of the proposed system and its components also demonstrates the effectiveness of this enhancement strategy and has the potential to be adapted to other downstream tasks.

[0088] In summary, the present invention provides an image super-resolution enhanced diffusion method and system based on a reference information database, which utilizes a class-based multi-reference patch matching pipeline and a reference information database to obtain a series of high-quality reference patches, which contain necessary high-frequency information and help to restore low-resolution image inputs. This system is the first method that can obtain relevant reference information from millions of high-quality reference patches. At the same time, a multi-scale feature transfer and aggregation control module is developed to align and fuse the information of the low-resolution image input, reference patches, and intermediate diffusion results of the current time step. The obtained features are used as guidance information in the subsequent diffusion restoration process. The ablation study verifies the effectiveness of each proposed module.

[0089] The above contents are only for explaining the technical idea of ​​the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.

Claims

1. Image super-resolution enhancement diffusion method based on reference information database, characterized in that: The following steps are involved: Create a reference information database and use a patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; Constructing a multi-scale feature transfer and aggregation control module to align input features of the low-resolution image to be processed with reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features; Based on the overall learning objective L, a pre-trained diffusion model is used to restore the low-resolution image with reference feature alignment to a high-resolution image, completing image super-resolution enhancement diffusion.

2. The image super-resolution enhancement diffusion method based on a reference information database according to claim 1, characterized in that: The reference information database is prepared, and the corresponding high-resolution image patches are obtained from the reference information database using a patch matching pipeline, specifically: A number of images are randomly selected from each category of the data set, and then a reference information database is formed by cropping and padding, and the low-resolution images in the reference information database are cropped into a first patch vector of a set size; Convert the cropped first patch vector to vector space using a pre-trained residual network; Calculate a vector similarity matrix between a patch vector of the low-resolution image and a first patch vector of a reference image in a reference information database based on the obtained vector space; The vector similarity matrix is ​​used to locate the corresponding second patch vector in the vector space, and the second patch vector is used as the high-resolution image patch.

3. The image super-resolution enhanced diffusion method based on a reference information database according to claim 2, characterized in that: Vector similarity matrix between the patch vector of the low-resolution image and the first patch vector of the reference image in the reference information database The calculation is as follows: in, represents the patch vector of the low-resolution image, represents the first patch vector, , are all constants.

4. The image super-resolution enhanced diffusion method based on a reference information database according to claim 2, characterized in that: The dataset uses the ImageNet dataset.

5. The image super-resolution enhanced diffusion method based on a reference information database according to claim 4 is characterized in that: The method randomly selects several images from each category of the data set, specifically: Randomly select several low-resolution images of size 512×512 pixels from each class of the ImageNet dataset.

6. The image super-resolution enhanced diffusion method based on a reference information database according to claim 1, characterized in that: Construct a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch, specifically: Two trainable multi-scale encoders are used to extract multi-scale features from the low-resolution image and the reference information of the reference information database respectively; the obtained multi-scale features are aligned and fused with the intermediate diffusion results at time step T to generate aligned multi-scale features The low-resolution image of is used as the low-resolution image for reference feature alignment.

7. The image super-resolution enhanced diffusion method based on a reference information database according to claim 1, characterized in that: Based on the overall learning objective L, a pre-trained diffusion model is used to restore the low-resolution image aligned with the reference features to a high-resolution image as follows: In the forward process of the pre-trained diffusion model, the pre-trained diffusion model is used to add Gaussian noise to the low-resolution image aligned with the reference features; in the reverse process of the pre-trained diffusion model, the low-resolution image aligned with the reference features is restored to a high-resolution image through the inference model.

8. The image super-resolution enhanced diffusion method based on a reference information database according to claim 7, characterized in that: The forward process of the pre-trained diffusion model is as follows: Where t represents the number of times Gaussian noise is added, is the forward process of the diffusion model, is a scalar function, is a Gaussian distribution, For input, is the initial value, is the identity matrix; The reverse process of the pre-trained diffusion model is as follows: in, is the reverse process of the diffusion model, is the average value, is the variance, For input, is the input of the previous moment, For input The mean of the Gaussian distribution.

9. The image super-resolution enhanced diffusion method based on a reference information database according to claim 7, characterized in that: Overall learning objectives The details are as follows: in, is the mathematical expectation, is a target image in a latent space, is the number of times Gaussian noise is added, Indicates that the crop text hint is set to empty. and are low-resolution image and reference image input respectively, is the conditional denoising network in the diffusion model, is the noise sampled from Gaussian noise and follows Gaussian distribution, is a normal distribution.

10. An image super-resolution enhanced diffusion system based on a reference information database, characterized in that: include: Data module, making a reference information database and using patch matching pipeline to obtain the corresponding high-resolution image patches from the reference information database; An alignment module constructs a multi-scale feature transfer and aggregation control module to align the input features of the low-resolution image to be processed with the reference features in the high-resolution image patch to obtain a low-resolution image with aligned reference features; The output module, based on the overall learning objective L, uses the pre-trained diffusion model to restore the low-resolution image with reference feature alignment to a high-resolution image, completing the image super-resolution enhancement diffusion.

Citation Information

Patent Citations

  • Nuclear magnetic resonance image super-resolution recovery method and model construction method

    CN117611453A

  • Face recognition method in low resolution image and face recognition device in low resolution image

    KR101382892B1

  • Improved noise schedules, losses, and architectures for generation of high-resolution imagery with diffusion models

    WO2024159002A2