An image blind super-resolution reconstruction method for complex degradation scenes

CN116433483BActive Publication Date: 2026-09-11CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310203802.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-06
Publication Date
2026-09-11
Estimated Expiration
2043-03-06

AI Technical Summary

Technical Problem

面对场景复杂的退化环境,现有的盲超分辨率算法估计的退化信息仍有不足,尤其是存在噪声环境下,模糊核的估计也未能充分覆盖

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116433483B_ABST
    Figure CN116433483B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image processing, and relates to computer vision, in particular to a blind super-resolution reconstruction method for complex degradation scenes; the method comprises the following steps: obtaining a low-resolution image under a complex degradation scene; estimating degradation prior information by using a degradation estimation network with a double-branch prediction structure, wherein the degradation prior information comprises a blur kernel and a noise map; processing the low-resolution image and the degradation prior information by using a super-resolution reconstruction network to obtain a high-resolution image; the present application fully estimates degradation information such as a blur kernel and noise under a complex degradation scene, and uses a denoising module and a deblurring module in a super-resolution network to guide image reconstruction by using estimated degradation characteristic information; the present application fully considers the degradation characteristic information of a complex degradation scene to realize the function of blind super-resolution reconstruction of an image, thereby obtaining a high-quality high-resolution image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing technology, specifically computer vision, and more specifically to a blind super-resolution reconstruction method for complex degraded scenes. Background Technology

[0002] Image super-resolution technology aims to reconstruct high-resolution (HR) images from one or more low-resolution (LR) images, and it is widely used in fields such as medical imaging, satellite remote sensing, and security monitoring. Due to hardware resource limitations, the resolution of images directly captured by sensors often falls short of practical requirements. Therefore, using software algorithms to reconstruct image super-resolution is a current focus in the field of computer vision.

[0003] Existing research methods, particularly deep learning-based image super-resolution, have achieved significant visual results. However, these methods rely on a degradation setting of bicubic downsampling for image super-resolution reconstruction. While the network performance is quite good under this degradation setting, real-world image degradation scenarios are far more complex than this. Therefore, how to use degradation prior information to guide the super-resolution network to accurately reconstruct high-resolution images is a current research focus.

[0004] Image super-resolution, a typical ill-posed problem in computer vision, merely establishes a mapping relationship between low-resolution and high-resolution images without considering the blur kernel as prior information to guide the network's reconstruction process. Therefore, when the blur kernel information is biased, the network performance is suboptimal. To address this issue, IKC proposed a strategy of alternating iterative correction between blur kernel estimation and image reconstruction as two sub-networks. This correction yields an accurate blur kernel to guide the super-resolution network in recovering a clearer high-resolution image. Similarly, DAN, another blind super-resolution method using alternating iterative optimization, integrates the blur kernel estimation module and the reconstruction module into a unified super-resolution network. This allows the low-resolution image to participate in all iterative sub-modules of the network, fully exploring the mapping relationship between low-resolution and high-resolution images. However, both of these iterative optimization methods are prone to small deviations or errors in blur kernel estimation, affecting the reconstruction effect of subsequent super-resolution networks. Therefore, subsequent research has decomposed this type of super-resolution problem into two sub-tasks: accurately estimating the blur kernel information using the low-resolution image, and then using the estimated blur kernel to reconstruct the super-resolution image. MANet utilizes affine convolution modules to fully extract degradation information from low-resolution images and accurately estimate the blur kernel. This blur kernel information is then integrated into a non-blind super-resolution network to improve the robustness of network reconstruction. However, existing blind super-resolution algorithms still fall short in estimating degradation information in complex degraded environments, especially in noisy environments where blur kernel estimation fails to adequately cover the data. Summary of the Invention

[0005] To further improve existing blind super-resolution algorithms, this invention provides a blind super-resolution reconstruction method for images in complex degraded scenes. This method more accurately estimates image degradation information in noisy and complex degraded environments to guide the super-resolution network in reconstructing higher-quality images.

[0006] To achieve the above objectives, the present invention proposes the following technical solution:

[0007] A blind super-resolution method for images in complex degradation scenarios includes the following steps:

[0008] Acquire low-resolution images in complex degradation scenarios;

[0009] A degradation estimation network with a dual-branch prediction structure is used to estimate degradation prior information, which includes a fuzzy kernel and a noise map.

[0010] A super-resolution reconstruction network is used to process the low-resolution image and the degradation prior information to reconstruct a high-resolution image.

[0011] The beneficial effects of this invention are as follows:

[0012] Compared to existing blind super-resolution methods, this invention fully estimates degradation information such as blur kernels and noise in complex degradation scenarios, and uses the estimated degradation feature information to guide image reconstruction based on the denoising and deblurring modules in the super-resolution network. This method fully utilizes the degradation feature information of complex degradation scenarios to achieve blind super-resolution image reconstruction, thereby obtaining high-quality, high-resolution images. Attached Figure Description

[0013] Figure 1 This is a flowchart of the image blind super-resolution reconstruction method according to an embodiment of the present invention;

[0014] Figure 2 This is a schematic diagram of the degradation estimation network in an embodiment of the present invention;

[0015] Figure 3 This is a schematic diagram of the structure of the super-resolution reconstruction network in an embodiment of the present invention;

[0016] Figure 4 This is a schematic diagram of the noise reduction module in an embodiment of the present invention.

[0017] Figure 5 This is a schematic diagram of the deblurring module in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] This invention provides a blind super-resolution method for images in complex degradation scenarios. This method uses a combination of two sub-networks: a low-resolution image as input to the degradation estimation network, which outputs a corresponding estimated blur kernel and noise; and a super-resolution reconstruction network that uses the low-resolution image as basic input and the estimated degradation information as conditional input to form a corresponding feature map in the network, enhancing the role of degradation information in the reconstruction process, and finally outputting a corresponding high-resolution image.

[0020] Figure 1 This is a flowchart of an image blind super-resolution method for complex degradation scenarios according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes:

[0021] 101. Acquire low-resolution images in complex degradation scenes;

[0022] In this embodiment of the invention, the low-resolution image can be obtained from existing public datasets or captured by devices such as cameras, video cameras, and mobile phones. The invention does not impose specific limitations on this.

[0023] As used herein, the term "low-resolution image" for a sample is generally defined as an image in which all patterned features in the region of the sample that produces the image are not resolvable in the image. For example, some of the patterned features in the region of the sample that produces the low-resolution image may be resolvable in the low-resolution image, provided that they are large enough to be resolvable. However, the low-resolution image is not produced at a resolution that would make all the patterned features in the image resolvable. In this way, the term "low-resolution image" as used herein does not contain information about the patterned features on the sample sufficient to make the low-resolution image usable for applications such as defect review, which may include defect classification and / or verification, as well as metrology. Furthermore, the term "low-resolution image" as used herein generally refers to an image produced by an inspection system that typically has a relatively low resolution (e.g., lower than that of a defect review and / or metrology system) in order to allow for relatively fast processing.

[0024] "Low-resolution images" can also mean "low-resolution" because they have a lower resolution compared to the "high-resolution images" described herein. The term "high-resolution image," as used herein, can generally be defined as an image that resolves all patterned features of a sample with relatively high accuracy. In this way, all patterned features in the region of the sample that produces the high-resolution image are resolved in the high-resolution image regardless of their size. Similarly, the term "high-resolution image," as used herein, contains information about the patterned features of a sample sufficient to make the high-resolution image usable for applications such as defect review, which may include defect classification and / or verification, as well as metrology. Furthermore, the term "high-resolution image," as used herein, generally refers to an image that cannot be generated by an inspection system during routine operations, configured to sacrifice resolution capability for increased processing power. In this way, "high-resolution image" can also be referred to herein and in the relevant field as "high-sensitivity image," another term for "high-quality image." Therefore, the terms low-resolution image and high-resolution image are used relative to each other and are not intended to indicate any particular image resolution.

[0025] In this embodiment of the invention, if the low-resolution image in the complex degradation scene is obtained from a high-resolution image in the dataset, this embodiment obtains the training low-resolution image from the high-resolution image in the DIV2K dataset through the following complex degradation processing. A 21×21 isotropic Gaussian kernel and an anisotropic Gaussian kernel are used as the blur kernels, with the standard deviations of both kernels uniformly sampled from [0.2, 4.0]. Additive white Gaussian noise and JPEG compressed noise are used as references for comparison. The level range of the additive white Gaussian noise is set in [0, 50], and the compression quality of the JPEG compressed noise is selected from [0, 100]. If the low-resolution image in the complex degradation scene is a real image or other unprocessed low-resolution image, then no noise processing is required; the blur kernel and noise map can be directly estimated using the network model of this invention.

[0026] 102. A degradation estimation network with a dual-branch prediction structure is used to estimate degradation prior information, wherein the degradation prior information includes a fuzzy kernel and a noise map;

[0027] In this embodiment of the invention, the low-resolution image I LR As input to the degradation estimation network, it estimates the blur kernel information and additive noise information corresponding to the low-resolution image.

[0028] In this embodiment of the invention, the degradation estimation network extracts shallow features from a convolutional layer, then passes through three residual blocks formed by stacking mutually affine transformation layers in MANet, and performs downsampling and upsampling operations through a convolutional layer and a transposed convolutional layer, ultimately forming a network structure similar to U-Net. The degradation information reconstruction branch is divided into two branches to estimate blur kernel information and noise information respectively. The estimated degradation information will further help the super-resolution reconstruction network generate more accurate and clearer high-resolution images.

[0029] Specifically, such as Figure 2As shown, the degradation estimation network includes a MANet backbone network, and a first branch and a second branch connected to the MANet backbone network. The MANet backbone network includes three alternating convolutional layers and three residual blocks, with the last convolutional layer being a transposed convolutional layer. The MANet backbone network is used to input a low-resolution image and output a processed intermediate feature map. The first convolutional layer extracts the shallow features of the low-resolution image, the first residual block extracts the deep features of the shallow features, the downsampling convolutional layer downsamples the deep features, the second residual block extracts even deeper features from the downsampled deep features, a transposed convolutional layer upsamples the even deeper features, and the third residual block extracts even deeper features from the even deeper features, which is the intermediate feature map. The first branch includes a convolutional layer and is used to process the input intermediate feature map, outputting a noise map. The second branch includes two convolutional layers, a pooling layer, and a softmax layer and is used to process the input intermediate feature map, outputting a blur kernel.

[0030] 103. A super-resolution reconstruction network is used to process the low-resolution image and the degradation prior information to reconstruct a high-resolution image.

[0031] In this embodiment of the invention, the low-resolution image I LR As input to the degradation estimation network, it estimates the blur kernel information and additive noise information corresponding to the low-resolution image.

[0032] The super-resolution reconstruction network consists of a denoising module and multiple deblurring modules that process the estimated degradation information and the low-resolution image, respectively. The denoising module associates the feature information of the low-resolution image with the noise information to generate a corresponding feature map to guide the network in denoising. The deblurring module uses a residual structure of stacked spatial feature transformation layers and convolutional layers to fuse image features and blur kernel information to enhance the influence of the blur kernel degradation features in the reconstruction process. The network uses long skip connections to make full use of the low-frequency information in the low-resolution image for image reconstruction.

[0033] Specifically, such as Figure 3 As shown, the super-resolution reconstruction network includes a denoising module, several deblurring modules, a single convolutional layer, and a transposed convolutional layer. The denoising module denoises the input low-resolution image and the noise map, outputting a denoised feature map. The deblurring module deblurs the input blur kernel and the denoised feature map, outputting a deblurred feature map. The single convolutional layer processes the input low-resolution image and the deblurred feature map, outputting a reconstructed feature map. The transposed convolutional layer processes the input reconstructed feature map, outputting a high-resolution image.

[0034] In embodiments of the present invention, such as Figure 4 As shown, the denoising module includes a convolutional layer, a fully connected layer, and an activation function. The low-resolution image is input from the convolutional layer of the denoising module, and a first feature information is output. The noise map is input from the fully connected layer, passes through the activation function, and outputs a second feature information. The first feature information and the second feature information are multiplied to obtain a denoised feature map. The noise map, combined with the low-resolution image by the denoising module, generates a corresponding feature map, which can be used to guide the network in denoising.

[0035] In embodiments of the present invention, such as Figure 5 As shown, the deblurring module includes alternating spatial transformation feature layers and convolutional layers. The spatial transformation feature layer includes a reshaping processing module, a feature map connection module, a sigmoid activation function, and two inner convolutional layers. The deblurring kernel input from the reshaping processing module in the first spatial transformation feature layer is processed to output a reshaping blur kernel. The feature map connection module in the first spatial transformation feature layer fuses and concatenates the denoised feature map with the reshaping blur kernel to obtain a reshaping feature map. One inner convolutional layer in the first spatial transformation feature layer processes the reshaping feature map using the sigmoid activation function, and then convolves it with the denoised feature map to obtain a first intermediate blur map. Another inner convolutional layer in the first spatial transformation feature layer adds the reshaping feature map and the first intermediate blur map to obtain a second intermediate blur map. The second intermediate blur map is then passed through the convolutional layer of the spatial transformation feature layer to output a deblurred feature map.

[0036] In this embodiment of the invention, the size of each convolutional layer can be the same, for example, its size can be 3×3, and the invention does not make a specific limitation on this.

[0037] In the super-resolution network of this invention, a linear combination structure of a denoising module and a deblurring module is adopted. The deblurring module serves as the basic module of the network for basic input, and the denoising module serves as the conditional module for conditional input. The inputs from two paths are processed separately, with the basic input and conditional input defined as the low-resolution image and degradation information, respectively. For the conditional input branch, the feature information of different levels of the basic input branch is processed through a mutual affine transformation layer and a convolutional layer to guide the reconstruction of the super-resolution network. Subsequent convolutional layers link the degradation information with the feature information of the low-resolution image. Unlike previous blind super-resolution algorithms, this invention adds a denoising module for noise information to the super-resolution reconstruction network, combining the feature information of the low-resolution image with the noise map information in subsequent convolutions for feature mapping.

[0038] In a preferred embodiment of the present invention, the super-resolution dataset DIV2K undergoes preprocessing for complex degradation scenarios, specifically dynamic synthetic blurring and noise addition, which is then applied to the network for training and inference. Simultaneously, Set5, Set14, Urban100, BSD100, and Manga109 are used as test sets to evaluate the network's performance. Training is completed by setting the initial learning rate, optimizer, loss function, and number of iterations, and the final network reconstruction effect is obtained through testing and comparison.

[0039] During the training of the network model, this embodiment simulates the degradation of complex degradation scenarios in the image. Therefore, this invention adds noise to the DIV2K dataset used for training and inference, obtaining a noisy low-resolution image as the network input. This invention uses two sub-networks to achieve image super-resolution. The input noisy low-resolution image first passes through a degradation estimation network to estimate the corresponding blur kernel information and noise map information. The network adopts a MANet-like network structure, consisting of residual blocks, upsampling, and downsampling. This structure allows for better extraction of deep features from low-resolution images, and the network fully learns the corresponding degradation feature information. Finally, two branches are extended to estimate the blur kernel information and noise map information, respectively. The blur kernel estimation branch uses two convolutional layers and one pooling layer for prediction and uses Softmax to constrain the blur kernel to generate the final blur kernel result; the noise map estimation branch directly uses a single convolutional layer to predict the corresponding noise result.

[0040] The degradation estimation network and the super-resolution reconstruction network are constructed using degradation reconstruction loss and image reconstruction loss, respectively. Stable training performance is achieved through the convergence model of these two types of losses. The network's loss function is expressed as:

[0041] L = w dr ×L DR +w1×L1

[0042]

[0043]

[0044] Where L represents the loss function, w dr L represents the weight of the degradation and reconstruction loss. DR Let w1 represent the weights of the image reconstruction loss, and L1 represent the image reconstruction loss; K represents the image reconstruction loss. GT and K est N represents the actual fuzzy kernel and the estimated fuzzy kernel, respectively; GT and N estThese represent the actual noise information and the estimated noise information, respectively; C, H, and W represent the number of channels and the image dimensions (length and width) of the high-resolution image, respectively; I HR and I SR These are represented as the real high-resolution image and the reconstructed high-resolution image, respectively. represents the L2 norm; |||1 represents the L1 norm.

[0045] Through the above iterative loop, the trained degradation estimation network and super-resolution reconstruction network can be obtained, which can be used to perform blind super-resolution reconstruction on the low-resolution image to be processed, thereby obtaining a high-quality high-resolution image.

[0046] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.

[0047] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for blind super-resolution image reconstruction in complex degraded scenes, characterized in that, The method includes: Acquire low-resolution images in complex degradation scenarios; A degradation estimation network employing a dual-branch prediction structure estimates degradation prior information, which includes a blur kernel and a noise map. The degradation estimation network comprises a MANet backbone network, and a first branch and a second branch connected to the MANet backbone network. The MANet backbone network includes three alternating convolutional layers and three residual blocks, with the last convolutional layer being a transposed convolutional layer. The MANet backbone network is used to input a low-resolution image and output a processed intermediate feature map. The first convolutional layer extracts shallow features from the low-resolution image, and the first residual block extracts... The deep features of the shallow features are taken, and the deep features are downsampled using a second convolutional layer. A second residual block is used to extract even deeper features from the downsampled deep features. The even deeper features are upsampled using a transposed convolutional layer, and the feature information of the even deeper features is extracted using a third residual block, which forms the intermediate feature map. The first branch includes a convolutional layer and is used to process the input intermediate feature map, outputting a noise map. The second branch includes two convolutional layers, a pooling layer, and a softmax layer. The second branch is used to process the input intermediate feature map and output a blur kernel. A super-resolution reconstruction network is used to process the low-resolution image and the degradation prior information to reconstruct a high-resolution image. The super-resolution reconstruction network includes a denoising module, several deblurring modules, a single convolutional layer, and a transposed convolutional layer. The denoising module denoises the input low-resolution image and the noise map, outputting a denoised feature map. The deblurring module deblurs the input blur kernel and the denoised feature map, outputting a deblurred feature map. The single convolutional layer processes the input low-resolution image and the deblurred feature map, outputting a reconstructed feature map. The transposed convolutional layer processes the input reconstructed feature map, outputting a high-resolution image.

2. The image blind super-resolution reconstruction method for complex degradation scenes according to claim 1, characterized in that, The denoising module includes a convolutional layer, a fully connected layer, and an activation function. The low-resolution image is input from the convolutional layer of the denoising module, and a first feature information is output. The noise map is input from the fully connected layer, passes through the activation function, and outputs a second feature information. The first feature information and the second feature information are multiplied to obtain a denoised feature map.

3. The image blind super-resolution reconstruction method for complex degradation scenes according to claim 1, characterized in that, The deblurring module includes alternating spatial transformation feature layers and convolutional layers. The spatial transformation feature layer includes a reshaping processing module, a feature map connection module, a sigmoid activation function, and two inner convolutional layers. The deblurring kernel input from the reshaping processing module in the first spatial transformation feature layer is processed to output a reshaping blur kernel. The feature map connection module in the first spatial transformation feature layer fuses and concatenates the denoised feature map with the reshaping blur kernel to obtain a reshaping feature map. One inner convolutional layer in the first spatial transformation feature layer processes the reshaping feature map using the sigmoid activation function, and then convolves it with the denoised feature map to obtain a first intermediate blur map. Another inner convolutional layer in the first spatial transformation feature layer adds the reshaping feature map and the first intermediate blur map to obtain a second intermediate blur map. The second intermediate blur map is then passed through the convolutional layers of the spatial transformation feature layer to output a deblurred feature map.

4. The image blind super-resolution reconstruction method for complex degradation scenes according to claim 1, characterized in that, The degradation estimation network and the super-resolution reconstruction network are constructed using degradation reconstruction loss and image reconstruction loss, respectively, and are expressed as follows: in, Represents the loss function. The weights representing the degradation and reconstruction loss, Indicates the degradation and reconstruction loss. The weights represent the image reconstruction loss. K represents the image reconstruction loss; GT and K est N represents the actual fuzzy kernel and the estimated fuzzy kernel, respectively; GT and N est These represent the actual noise information and the estimated noise information, respectively; C, H, and W represent the number of channels and the image dimensions (length and width) of the high-resolution image, respectively; I HR and I SR These are represented as the real high-resolution image and the reconstructed high-resolution image, respectively. express Norm; express Norm.

Citation Information

Patent Citations

  • Image restoration method based on multi-scale self-similarity and conformal constraint

    CN109859131A

  • Image blind super-resolution method and system

    CN113139904A