Reference image feature dynamic correction and space-channel fusion image reconstruction method

By using the dynamic correction of reference image features and spatial-channel fusion techniques in the super-resolution reconstruction method, the existing methods have solved the problem of low image reconstruction quality and poor model robustness when processing complex texture information or reference images with little correlation, thus achieving higher quality image reconstruction and stronger model robustness.

CN120235758APending Publication Date: 2025-07-01BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510351995.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing super-resolution reconstruction method based on reference images cannot accurately obtain information related to low-resolution images when processing complex texture information or reference images with little correlation with low-resolution images, resulting in low image reconstruction quality and poor model robustness.

Method used

The reference image features are dynamically corrected and spatial-channel fusion methods are adopted to improve the accuracy of image reconstruction through the combination processing of shallow feature extraction, matching and extraction modules, feature fusion parts and image reconstruction parts.

Benefits of technology

Improve image reconstruction quality, enhance model robustness, and more efficiently process complex texture information and reference images that are not closely related to low-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235758A_ABST
    Figure CN120235758A_ABST
Patent Text Reader

Abstract

The invention discloses an image reconstruction method for reference image feature dynamic correction and space-channel fusion, and the method comprises the steps: taking a low-resolution image and a reference image as input, firstly carrying out the shallow feature extraction of the low-resolution image, and obtaining the features of the low-resolution image; meanwhile, the low-resolution image and the reference image are input into a matching and extracting module to be processed, and reference image features of the * 1 scale and the * 2 scale related to the low-resolution image are obtained; the low-resolution image features and the * 1 scale reference image features are spliced according to channels to serve as input of a feature fusion part, the * 1 scale reference image features and the * 2 scale reference image features serve as the other two inputs of the feature fusion part, and reconstruction features are obtained; and the feature fusion part is repeatedly executed for three times to obtain final reconstructed features. And finally, inputting the final reconstruction feature into an image reconstruction part for image reconstruction to obtain a super-resolution reconstruction result, and training the model by using a loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image super - resolution reconstruction, aiming to reconstruct and enlarge a low - resolution image into a super - resolution image through a model. Specifically, it is an image super - resolution reconstruction method with dynamic correction of reference - image features and spatial - channel fusion. Background Art

[0002] Image super - resolution reconstruction is to reconstruct a low - resolution image into a high - resolution image through a network model, which has wide applications in fields such as transportation, medicine, and monitoring, and has strong practicality. With the development of deep - learning technology, image super - resolution reconstruction technology has developed rapidly. However, due to its ill - posed problem, that is, for a low - resolution image, there are multiple possible high - resolution images during the super - resolution reconstruction process, which has hindered the development of image super - resolution reconstruction.

[0003] The super - resolution reconstruction technology based on a reference image is based on traditional image super - resolution reconstruction, providing an additional high - resolution reference image to assist in completing super - resolution reconstruction. The reference image generally has similar semantic and texture information to the high - resolution image corresponding to the low - resolution image. By extracting relevant information from the reference image, super - resolution reconstruction is carried out.

[0004] Currently, although the super - resolution reconstruction technology based on a reference image has made significant progress, after extracting relevant information from the reference image, most previous models use convolution to fuse features with the low - resolution image. Convolution can extract local features, ignoring global information to a certain extent, and convolution does not specially process the channel relationship of features. In addition, after extracting relevant information from the reference image, most previous methods directly fuse the extracted relevant information with the low - resolution image features, lacking dynamic correction of the reference - image features according to the reconstructed features during the reconstruction process to obtain more relevant reference - image features. Therefore, there is a need for a new super - resolution reconstruction method based on a reference image to solve the above problems. Summary of the Invention

[0005] The problem to be solved by the present invention is that in the super - resolution reconstruction method based on a reference image, although existing reconstruction methods have made great progress in both reconstruction effect and running speed, when faced with a reference image containing complex texture information or having little correlation with the low - resolution image, they cannot accurately obtain the information related to the low - resolution image in the reference image, resulting in low image - reconstruction quality and poor model robustness.

[0006] To solve the above problems, the present invention provides an image super - resolution reconstruction method with dynamic correction of reference - image features and spatial - channel fusion. The method includes the following steps:

[0007] 1) This method takes a low-resolution image and a reference image as inputs;

[0008] 2) First, perform shallow feature extraction on the low-resolution image to obtain low-resolution image features; meanwhile, input the low-resolution image and the reference image into a matching and extraction module for processing to obtain reference image features of ×1 scale and ×2 scale related to the low-resolution image;

[0009] 3) Concatenate the low-resolution image features and the ×1 scale reference image features by channel as the input to the feature fusion part;

[0010] 4) Use the ×1 scale and ×2 scale reference image features as the other two inputs to the feature fusion part. The feature fusion part processes based on the three input features to obtain reconstructed features;

[0011] 5) Repeat step 4) three times to obtain the final reconstructed features;

[0012] 6) Input the final reconstructed features into the image reconstruction part for image reconstruction to obtain the super-resolution reconstruction result. The image reconstruction part consists of a sub-pixel convolutional layer and two 3×3 convolutional kernels, and the model is trained using reconstruction loss, perceptual loss, and adversarial loss. Description of the Drawings

[0013] Figure 1 is a flowchart of the image super-resolution reconstruction method for dynamic correction of reference image features and spatial-channel fusion in the present invention.

[0014] Figure 2 is a network model diagram.

[0015] Figure 3 is a residual fusion module diagram.

[0016] Figure 4 is a spatial Transformer diagram.

[0017] Figure 5 is a channel Transformer diagram.

[0018] Figure 6 is a Window Mixformer diagram. Detailed Embodiments

[0019] The present invention provides an image super-resolution reconstruction method for dynamically correcting reference image features and spatial-channel fusion. This method takes a low-resolution image and a reference image as inputs. First, shallow feature extraction is performed on the low-resolution image to obtain low-resolution image features. At the same time, the low-resolution image and the reference image are input into a matching and extraction module for processing to obtain reference image features of ×1 scale and ×2 scale related to the low-resolution image. Then, the low-resolution image features and the ×1 scale reference image features are concatenated by channel as the input of the feature fusion part, and the ×1 scale and ×2 scale reference image features are used as the other two inputs of the feature fusion part. The feature fusion part processes based on the three input features to obtain reconstructed features. The feature fusion part is repeatedly executed 3 times to obtain the final reconstructed features. Finally, the final reconstructed features are input into the image reconstruction part for image reconstruction to obtain the super-resolution reconstruction result, and the model is trained using a loss function.

[0020] The present invention includes the following steps:

[0021] 1) This method takes a low-resolution image and a reference image as inputs;

[0022] 2) First, the low-resolution image is input into a shallow feature extraction module for shallow feature extraction to obtain low-resolution image features. At the same time, the low-resolution image and the reference image are input into a matching and extraction module for processing to obtain reference image features of ×1 scale and ×2 scale related to the low-resolution image. The shallow feature extraction module consists of a convolution with a convolution kernel of 3×3.

[0023] 3) The low-resolution image features and the ×1 scale reference image features are concatenated by channel as the input of the feature fusion part;

[0024] 4) The ×1 scale and ×2 scale reference image features are used as the other two inputs of the feature fusion part. The feature fusion part processes based on the three input features to obtain reconstructed features;

[0025] 5) Step 4) is repeatedly executed 3 times to obtain the final reconstructed features;

[0026] 6) The final reconstructed features are input into the image reconstruction part for image reconstruction to obtain the super-resolution reconstruction result. The image reconstruction part consists of a sub-pixel convolution layer and two convolutions with a convolution kernel of 3×3, and the model is trained using a reconstruction loss, a perceptual loss, and an adversarial loss.

[0027] Furthermore, the matching and extraction module in step 2) is specifically:

[0028] 2.1) First, the reference image Ref is bilinearly downsampled by a factor of four to obtain Ref↓, which has the same size as the low-resolution image LR.

[0029] 2.2) Then, Ref, Ref↓, and LR are respectively input into the encoder for feature extraction. The encoder consists of three residual blocks. The feature maps output by the second and third residual blocks are respectively 1 / 2 and 1 / 4 of the size of the input feature map. The residual block consists of a convolution with a 3×3 convolutional kernel, an activation function, and a skip connection.

[0030] For the reference image Ref, the output results of the second and third residual blocks are taken as the reference image features at the ×2 scale and the reference image features at the ×1 scale respectively. For LR and Ref↓, only the output feature F LR and F Ref↓ are taken, and then three-stage processing is performed.

[0031] 2.3) In the first stage, F LR and F Ref↓ are divided into non-overlapping blocks. By calculating the cosine similarity, the most similar block in F LR to each block in F Ref↓ is determined.

[0032] 2.4) The second stage is to divide non-overlapping patches within each block of F LR and F Ref↓ . According to the similar corresponding relationship between the blocks in F LR and F Ref↓ obtained in step 2.3), the similarity between each pair of patches in the similar corresponding blocks is also calculated by the cosine similarity, and the patch with the largest similarity in each patch of each block in F LR in the corresponding block in F Ref↓ is taken to obtain the index map and the similarity map.

[0033] 2.5) Finally, according to the index map, the reference image features at the ×2 scale and the reference image features at the ×1 scale are recombined, and a dot product operation is performed with the similarity map to obtain the reference image features at the ×1 scale and the ×2 scale with spatially position-related information aligned.

[0034] Furthermore, the feature fusion part in step 4) is specifically:

[0035] The feature fusion part consists of a Residual Fusion Module (RFM), upsampling, a residual fusion module, and downsampling in sequence to achieve feature fusion at multiple scales. Among them, the residual fusion module consists of three Feature Aggregation Blocks (FABs), three Window Mixformers, a 3×3 convolution, and a skip connection in sequence. The input of the first Window Mixformer is the feature output by the third feature aggregation block and the reference image feature.

[0036] 4.1) The feature aggregation block consists of a Spatial Transformer (ST) and a Channel Transformer (CT) in sequence to achieve feature fusion in the spatial and channel dimensions.

[0037] 4.1.1) Specifically, the spatial transformer is as follows:

[0038] 4.1.1.1) First, after passing the input through a layer normalization layer, it is divided into two sub-vector sequences according to the channels: the global feature sequence X1 and the local feature sequence

[0039] 4.1.1.2) Secondly, X1 is input into the sliding window multi-head self-attention (SW-MSA) to extract global features, X2 is input into the input-dependent depthwise convolution (IDConv) to extract local features, and the global features and local features are concatenated in channels to obtain:

[0040] X' = Concat(SW-MSA(X1), IDConv(X2)) (1)

[0041] Among them, Concat(·) represents the concatenation operation.

[0042] 4.1.1.3) Then, to effectively fuse global features and local features, achieve cross-channel information fusion and control the computational complexity, 3×3 depthwise convolution DWConv is used to enhance local relationships, and two 1×1 convolutions are used to compress and expand channels to reduce the computational complexity. At the same time, the fused feature X” is obtained through a residual connection:

[0043]

[0044] Among them, r is the channel reduction multiple.

[0045] 4.1.1.4) Finally, the fused feature X” passes through the LayerNorm and MLP layers respectively to obtain the final result X s 。

[0046] 4.1.2) The channel Transformer is specifically as follows:

[0047] 4.1.2.1) First, the feature sequence X s is normalized through a layer to obtain X' s , which are respectively used as the inputs of the Channel-Wise SelfAttention (CW-SA) and the Depthwise Separable Convolution (DS-Conv);

[0048] 4.1.2.2) In CW-SA, the query Q c , key K c and value V c are successively generated, their dimensions are reshaped to and self-attention calculation is performed to obtain the channel self-attention Y c :

[0049]

[0050] where α is a learnable scaling parameter.

[0051] 4.1.2.3) In the depthwise separable convolution, X' s is first converted into the form of a feature map, and after passing through the depthwise convolution and the pointwise convolution respectively, it is then converted back into the form of a feature sequence:

[0052] V = Reshape((Reshape(X′ S ) * W d ) * W p ) (4)

[0053] where W d represents the convolution kernel of the depthwise convolution, and W p represents the convolution kernel of the pointwise convolution.

[0054] 4.1.2.4) Then, the outputs of CW-SA and DS-Conv are added together to obtain the feature Y:

[0055] Y = Y c + V (5)

[0056] 4.1.2.5) Finally, the feature Y passes through skip connection, layer normalization, and MLP to obtain the feature X sc 。

[0057] 4.2) The Window Mixformer further integrates the reference image features into the low-resolution image features and corrects the reference image features according to the low-resolution image features to obtain more relevant reference image features.

[0058] The first Window Mixformer in each residual fusion block has two inputs, namely the features output by its previous feature aggregation block and the reference image features output by the matching and extraction module. The inputs of the remaining two Window Mixformers are both the output of its previous Window Mixformer. Only the LR Token part of the output of the last Window Mixformer is retained.

[0059] The processing process of the first Window Mixformer is as follows:

[0060] 4.2.1) First, perform window partitioning operations on X sc and X Ref respectively to obtain the low-resolution image features LRFeature and the reference image features RefFeature, and perform Flatten transformation into sequences X' sc and X' Ref , and respectively obtain through linear operations: Q LR , K LR , V LR , Q Ref , K Ref and V Ref .

[0061] 4.2.2) Then, concatenate K LR and K Ref to obtain K, concatenate V LR and V Ref to obtain V, and perform hybrid attention calculations respectively. For the low-resolution image features, through hybrid attention calculation, fuse the reference image features to obtain the low-resolution image feature attention Attention LR . For the reference image features, through hybrid attention calculation, fuse the low-resolution image-related features to obtain the reference image feature attention Attnetion Ref :

[0062]

[0063] where d represents the dimension of the key K.

[0064] 4.2.3) Finally, combine the low-resolution image feature attention Attention LR with X' scAdd, reference image feature attention Ref Add with X' Ref Add them together, perform a concatenation operation on the two results, obtain the result after layer normalization and MLP processing respectively, and split the result to get the LR Token and RefToken as the input for the next WindowMixformer.

[0065] The loss function in step 6) is specifically:

[0066] The loss function includes reconstruction loss, perceptual loss, and adversarial loss. The loss value is calculated as follows:

[0067]

[0068] Among them, represents the reconstruction loss, represents the perceptual loss, represents the adversarial loss, and λ1 and λ2 are hyperparameters.

[0069] 6.1) To calculate the pixel-level difference between the generated super-resolution image and the original image, and at the same time reduce the influence of outliers on the result, the L1 norm is selected to calculate the reconstruction loss:

[0070]

[0071] Among them, X HR is the original high-resolution image, and X SR is the generated super-resolution image, and ||·|| is the L1 norm.

[0072] 6.2) To improve the visual quality of the super-resolution image, calculate the perceptual loss:

[0073]

[0074] Among them, ||·|| F is the Frobenius norm, V and C are the number of volume and channels of the feature map respectively. Φ i (·) is the i-th channel extracted by the ReLU5_1 activation function in VGG19.

[0075] 6.3) The present invention uses the Wasserstein Generative Adversarial Network (WGAN) loss to calculate the adversarial loss:

[0076]

[0077] Among them, D(·) is the discriminator, and PSR is the distribution of the generated super-resolution image, P HR is the distribution of the original high-resolution image, is the mathematical expectation.

[0078] The present invention has wide applications in the field of super-resolution reconstruction based on reference images, such as: medical imaging, public safety and surveillance, etc. The following will refer to the attached Figure 1 to describe the present invention in detail.

[0079] (1) This method takes a low-resolution image and a reference image as inputs.

[0080] (2) First, the low-resolution image is input into the shallow feature extraction module for shallow feature extraction to obtain low-resolution image features; at the same time, the low-resolution image and the reference image are input into the matching and extraction module for processing to obtain reference image features of ×1 scale and ×2 scale related to the low-resolution image.

[0081] (3) The low-resolution image features and the ×1 scale reference image features are concatenated by channel as the input of the feature fusion part.

[0082] (4) The ×1 scale and ×2 scale reference image features are used as the other two inputs of the feature fusion part. The feature fusion part processes based on the three input features to obtain reconstructed features;

[0083] (5) Step (4) is repeated 3 times to obtain the final reconstructed features.

[0084] (6) The final reconstructed features are input into the image reconstruction part for image reconstruction to obtain the super-resolution reconstruction result, and the model is trained using reconstruction loss, perceptual loss, and adversarial loss.

[0085] The present invention provides a reference super-resolution reconstruction method with dynamic correction of reference image features and spatial-channel fusion, which is applicable to super-resolution reconstruction tasks based on reference images, has high image reconstruction quality and good robustness. Experiments show that the invention can quickly and effectively perform super-resolution reconstruction based on reference images.

Claims

1. An image reconstruction method based on dynamic correction of reference image features and space-channel fusion, characterized in that: For a given low-resolution image and reference image, perform the following operations: Step 1), taking a low-resolution image and a reference image as input; Step 2), firstly input the low-resolution image into the shallow feature extraction module, perform shallow feature extraction, and obtain low-resolution image features; at the same time, input the low-resolution image and the reference image into the matching and extraction module for processing, and obtain the reference image features of ×1 scale and ×2 scale related to the low-resolution image; The shallow feature extraction module consists of a convolution with a kernel of 3×3; Step 3), the low-resolution image features and the ×1-scale reference image features are concatenated by channel as the input of the feature fusion part; Step 4), the ×1 scale and ×2 scale reference image features are used as the other two inputs of the feature fusion part; The feature fusion part processes the three input features to obtain the reconstructed features; Step 5), repeat step 4) three times to obtain the final reconstructed features; Step 6), input the final reconstructed features into the image reconstruction part, perform image reconstruction, and obtain the super-resolution reconstruction result; the image reconstruction part consists of a sub-pixel convolution layer and two convolutions with a convolution kernel of 3×3, and the model is trained using reconstruction loss, perceptual loss, and adversarial loss.

2. The image super-resolution reconstruction method of dynamic correction of reference image features and space-channel fusion according to claim 1 is characterized in that: The feature fusion part in step 4) is specifically: The feature fusion part is composed of the residual fusion module RFM, upsampling, residual fusion module and downsampling sequence to achieve feature fusion at multiple scales; the residual fusion module is composed of three feature aggregation blocks FAB, three Window Mixformers, a 3×3 convolution and a jump connection sequence; the input of the first Window Mixformer is the features output by the third feature aggregation block and the reference image features; Step 4.1), the feature aggregation block is composed of spatial Transformer and channel Transformer in sequence to achieve spatial and channel dimension feature fusion; Step 4.2), Window Mixformer further integrates the reference image features into the low-resolution image features, and modifies the reference image features according to the low-resolution image features to obtain more relevant reference image features; The first Window Mixformer in each residual fusion block has two inputs, namely the features output by the previous feature aggregation block and the reference image features output by the matching and extraction module; the inputs of the other two Window Mixformers are the outputs of the previous Window Mixformer; the output of the last Window Mixformer only retains its LRToken part.

3. The image super-resolution reconstruction method of dynamic correction of reference image features and space-channel fusion according to claim 2 is characterized in that: Step 4.1) includes, 4.1.1) spatial Transformer specifically: Step 4.1.1.1), after the input passes through the layer normalization layer, it is divided into two sub-vector sequences according to the channel: the global feature sequence X1 and the local feature sequence X2, X1, Step 4.1.1.2), input X1 into the sliding window multi-head self-attention SW-MSA to extract global features, input X2 into the input dependent deep convolution IDConv to extract local features, and channel-join the global features with the local features to obtain: X'=Concat(SW-MSA(X1),IDConvX2)) (1) Among them, Concat(·) represents the concatenation operation; Step 4.1.1.3), in order to effectively fuse global features with local features, realize cross-channel information fusion and control the amount of calculation, 3×3 deep convolution DWConv is used to enhance local relations, and two 1×1 convolution compression and expansion channels are used to reduce the amount of calculation, and the fused feature X" is obtained through residual connection: Among them, r is the channel reduction factor; Step 4.1.1.4), the fused feature X" passes through the LayerNorm and MLP layers respectively to obtain the final result X s .

4. The image super-resolution reconstruction method of dynamic correction of reference image features and space-channel fusion according to claim 2 is characterized in that: The channel Transformer is specifically: Step 4.1.2.1) First, transform the feature sequence X s After layer normalization, we get X' s , respectively as the input of channel-level self-attention CW-SA and depth-separable convolution DS-Conv; Step 4.1.2.2) In CW-SA, generate query Q in sequence c , key K c Sum value V c , reshape it into And perform self-attention calculation to obtain channel self-attention Y c : Where α is a learnable scaling parameter; Step 4.1.2.3) In the depthwise separable convolution, X' s First convert it into a feature map, then go through depth convolution and point-by-point convolution, and then convert it into a feature sequence: V=Reshape((Reshape(X' S )·W d )·W p ) (4) Among them, W d represents the convolution kernel of the depthwise convolution, W p Represents the convolution kernel of point-by-point convolution; Step 4.1.2.4) Then, add the outputs of CW-SA and DS-Conv to get feature Y: Y=Y c +V (5) Step 4.1.2.5) Finally, feature Y undergoes skip connection, layer normalization and MLP to obtain feature X sc .

5. The image super-resolution reconstruction method of dynamic correction of reference image features and space-channel fusion according to claim 2, characterized in that: The processing of Window Mixformer is as follows: Step 4.2.1) First, sc With X Ref Perform window division operations respectively to obtain low-resolution image features LRFeature and reference image features Ref Feature, and convert them into sequence X' by Flatten sc and X' Ref , and through linear operations, we can get: Q LR , K LR 、V LR , Q Ref , K Ref and V Ref ; Step 4.2.2) Then, K LR and K Ref Splice to get K, and V LR and V Ref The splicing results in V, and mixed attention calculations are performed separately; for low-resolution image features, mixed attention calculations are performed to fuse the reference image features to obtain low-resolution image feature attention LR ; For the reference image features, the hybrid attention calculation is used to fuse the low-resolution image related features to obtain the reference image feature attention. Ref : Where d represents the dimension of key K; Step 4.2.3) Attention is given to the low-resolution image features LR With X' sc Add, refer to image feature attention Ref With X' Ref Add the two results together, concatenate them, and obtain the results after layer normalization and MLP processing respectively. Then split the results to obtain LR Token and Ref Token as the input of the next Window Mixformer.