Super-resolution method

CN120339070APending Publication Date: 2025-07-18HEBEI UNIV OF ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510429291.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18

Smart Images

  • Figure CN120339070A_ABST
    Figure CN120339070A_ABST
Patent Text Reader

Abstract

The invention provides a super-resolution method, and aims to improve the spatial resolution and detail recovery capability of an image. According to the method, a super-resolution convolutional neural network is used as a baseline model, and a multi-head self-attention architecture is used for improving and optimizing the baseline model. In the improved network, a dense residual connection module in an original network is replaced by a visual multi-head self-attention module, so that the network can dynamically adjust the feature weight according to the input content, local and global features are modeled at the same time, information loss is avoided, and the feature extraction capability is enhanced. The method further introduces a frequency domain filtering residual module, further optimizes the structure of the super-resolution network, and improves the performance of the network when generating a high-resolution image. Paired image data sets are adopted for training, the model obtained after improved network training can effectively reconstruct high-quality images, the definition and detail performance of the images are remarkably improved, and good adaptability and wide application prospects in the fields of remote sensing, monitoring, microscopy and the like are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and deep learning, and particularly to a super-resolution method. This method aims to improve the spatial resolution and detail restoration ability of images. Background Art

[0002] Image super-resolution technology aims to recover high-resolution images from low-resolution images to enhance the clarity, detail expressiveness, and spatial information of images. This technology has wide application value in fields such as medical imaging, satellite remote sensing, security monitoring, and digital entertainment. However, due to factors such as the physical limitations of imaging devices, environmental noise interference, and transmission compression losses, the acquired images often have problems of insufficient resolution, resulting in the loss of key details, blurred edges, and degraded textures, thus affecting subsequent visual analysis and intelligent decision-making. Traditional super-resolution methods mainly include methods based on interpolation (such as bicubic interpolation), reconstruction (such as iterative back projection), and learning (such as sparse coding). Although these methods have certain effects in specific scenarios, they generally have problems such as poor adaptability, high computational complexity, and limited high-frequency detail restoration ability. For example, interpolation methods are prone to edge smoothing and jagged effects, while sparse representation-based methods rely on artificially designed features and are difficult to adapt to complex natural image degradation models. In recent years, the rapid development of deep learning technology has provided new solutions for image super-resolution. Convolutional neural networks have made remarkable progress in single-image super-resolution tasks due to their powerful feature extraction ability. The rise of the multi-head self-attention mechanism has brought new breakthroughs to image super-resolution. Different from convolutional neural networks, the multi-head self-attention mechanism can dynamically model long-range dependencies in images through the self-attention mechanism, thus more effectively restoring high-frequency details and structural information. However, directly applying the multi-head self-attention mechanism to the super-resolution task still faces challenges, such as high computational complexity, insufficient sensitivity to local details, and the risk of overfitting on small-sample data. Therefore, there is an urgent need to study an efficient and robust image super-resolution method that combines the local feature extraction ability of convolutional neural networks and the global modeling advantages of the multi-head self-attention mechanism to achieve higher-quality reconstruction effects under complex degradation conditions. At the same time, for different application scenarios (such as the restoration of fine structures in medical images, real-time super-resolution of surveillance videos, etc.), it is necessary to further optimize the network architecture and improve the generalization ability and computational efficiency of the algorithm to meet actual needs. Summary of the Invention

[0003] Object of the Invention. In order to improve the performance of the super-resolution neural network and solve the problem of poor super-resolution image generation effect existing in the prior art, the present invention provides a method for image super-resolution reconstruction using an image super-resolution network.

[0004] Technical solution: To solve the above technical problems, the present invention proposes an image super-resolution network. The network introduces a visual multi-head self-attention mechanism and a frequency-domain filtering residual module into the super-resolution neural network to enhance the network's expressive ability and image reconstruction quality. The specific implementation steps are as follows: Step 1: Obtain a high-resolution image dataset, preprocess and degrade the images to generate their corresponding low-resolution images; Step 2: Construct an image super-resolution network. The network uses a super-resolution convolutional neural network as the baseline model and improves it using a multi-head self-attention mechanism and a frequency-domain filtering residual module; Step 3: Use the images as high-resolution images and the corresponding degraded blurred images as low-resolution images to construct a data pair set, and input it into the improved image super-resolution network for training until the network model converges to obtain a trained image super-resolution network model; Step 4: Apply the trained image super-resolution network model, input the image to be super-resolved into the test network for super-resolution processing to generate a higher-resolution image.

[0005] Further, in Step 1, the preprocessing of the images includes: random cropping, pixel value normalization processing, and converting the processed images into PyTorch tensors for easy input into the neural network. The degradation processing uses bicubic interpolation to downsample the images to low resolution to obtain the low-resolution images required for training.

[0006] Further, in Step 2, in the construction of the image super-resolution network, the dense residual connection module in the original network is replaced by a new visual multi-head self-attention mechanism, which models local and global features simultaneously to avoid information loss, thereby enhancing the feature extraction ability.

[0007] Further, in Step 2, in the construction of the image super-resolution network, a frequency-domain filtering residual module is introduced into the network. The frequency-domain filtering residual module includes the following parts: Frequency-domain transformation layer: Use the fast Fourier transform (FFT) to transform the input sequence along the time dimension into the frequency-domain space to achieve the conversion from time-domain signal to frequency-domain representation; Complex linear transformation layer: Use linear transformation (Linear) to perform the same linear transformation operation on the real and imaginary parts of the frequency-domain signal respectively, keeping the dimension of the frequency-domain features unchanged; Complex activation layer: Use an improved complex ReLU activation function (ModReLU) to perform non-linear activation on the transformed frequency-domain signal. This activation function enhances the model's expressive ability through a learnable phase bias; Time-domain inverse transformation layer: Use the inverse fast Fourier transform (iFFT) to convert the processed frequency-domain signal back to the time-domain space, and only retain the real part as the final output.

[0008] Further, in step three, the training of the image super-resolution network includes: inputting the images and related information in the data set into the improved network; during the training process, calculating the loss function according to the difference between the generated image and the real image, and updating the parameters of the generator and discriminator according to backpropagation, so as to optimize the network parameters to reduce the reconstruction error.

[0009] Improvement effect: Compared with the prior art, the present invention combines the feature extraction ability of the multi-head self-attention mechanism with the advantages of the frequency-domain filtering residual module to achieve effective reconstruction of image super-resolution, and has the following technical advantages: 1) Higher reconstruction quality: By introducing a new convolutional network architecture, the optimized network structure significantly improves the detail recovery ability of the image, and the reconstructed image has significant improvements in clarity and details.

[0010] 2) Strong feature expression ability: The multi-head self-attention mechanism enhances the model's ability to capture complex textures and edges, which helps to recover key information in blurred images.

[0011] 3) Efficient training process: The improved network structure improves the training efficiency. Compared with the traditional super-resolution convolutional neural network, it can converge to the ideal state faster and reduces the consumption of computing resources.

[0012] 4) Strong adaptability: The present invention shows good adaptability in various image types and is suitable for different application scenarios. Description of the Drawings

[0013] The drawings of the present invention form a part of the specification and are intended to provide readers with an in-depth understanding of the present invention. These schematic embodiments and descriptions do not constitute a limitation to the present invention.

[0014] Figure 1 is the step flow chart of the method of the present invention; Figure 2 is the structural diagram improved by using the multi-head self-attention mechanism; Figure 3 is the structural diagram of the frequency-domain filtering residual module; Figure 4 is the comparison diagram of the implementation effects. Detailed Embodiments

[0015] The present invention will be further described below through embodiments. It should be noted that the purpose of disclosing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.

[0016] Embodiment: Perform super-resolution processing on low-resolution images.

[0017] As Figure 1 shown, the present invention proposes a super-resolution method. The method mainly includes the following steps: Step 1: Obtain a high-resolution image dataset, and perform preprocessing and degradation processing on the images to generate their corresponding low-resolution images; Step 2: Construct an image super-resolution network, which uses a super-resolution convolutional neural network as the baseline model and improves it using a multi-head self-attention mechanism and a frequency-domain filtering residual module; Step 3: Use the images as high-resolution images and the corresponding degraded blurred images as low-resolution images to construct a data pair set, and input it into the improved image super-resolution network for training until the network model converges to obtain a trained image super-resolution network model; Step 4: Apply the trained image super-resolution network model, input the image to be super-resolved into the test network for super-resolution processing to generate a higher-resolution image.

[0018] In Step 1, preprocess the image dataset, randomly crop the high-resolution images to a size of 256×256. Convert the images to tensors, and normalize the numerical range of the images from [0, 255] to [0, 1]. After conversion to tensors, the dimension of the images will change from (H, W, C) to (C, H, W). Degrade each image in the image dataset, and use bicubic interpolation to reduce the high-resolution images to 1 / 4 of their original size, that is, adjust the image size to 64×64 pixels to generate low-resolution images.

[0019] In Step 2, construct an image super-resolution network, where a super-resolution convolutional neural network is used as the baseline model, and the improvements include: First, use a visual multi-head self-attention mechanism to replace the original dense residual connection. The structure is as Figure 2 shown. The visual multi-head self-attention mechanism directly calculates the global relationship between image patches and can model long-range dependencies without stacking multiple layers. Among them, for the input feature map X, first, it is divided into non - overlapping image patches of a fixed size, and each patch is flattened into a vector through linear projection. Assume the shape of the input feature map is N , C, H , W (batch size, number of channels, height, width). It is divided into P ´ P image patches, then the size of each patch is N , C , H / p , W / p . Through linear transformation, each patch is mapped to an embedding vector with a dimension of D . Finally, the serialized image patch size is N , L , D , where L is the sequence length, and the calculation formula is as follows:

[0020] To retain the spatial position information of the image patches, before inputting into the Transformer encoder, a learnable position encoding E pos :

[0021] Subsequently, the enhanced patch sequence is input into the Transformer layer composed of multi - head self - attention (MSA) and feed - forward network (FFN). In the self - attention mechanism, each image patch dynamically aggregates global information (Attention) by calculating the similarity of query (Query), key (Key), and value (Value):

[0022] Among them, Q, K, V are obtained by linear transformation of the input X ′ respectively. Multi - head self - attention (MSA) further enhances the model's expressive power by calculating multiple groups of attention heads ([[]] head h ) in parallel and concatenating the results:

[0023] The output dimension of each attention head is D / h , W o is the output projection matrix.

[0024] The self-attention layer is followed by a feed-forward network (FFN), which consists of two fully-connected layers and the activation function GELU, and is used for non-linear feature transformation:

[0025] Among them, X’’ is the output of the self-attention layer. W 1 , W 2 both represent learnable weight matrices. b 1 , b 2 both represent learnable bias terms.

[0026] To stabilize the training, each sub-layer (MSA and FFN) adopts residual connection and layer normalization:

[0027]

[0028] Finally, after being processed by multiple Transformer layers, the model can efficiently capture the global dependencies between image patches.

[0029] In this embodiment, for the input image, it is first converted into a tensor, and then the input channel number is changed from 3 to 64 using the initial convolutional layer Conv2d to obtain the input tensor X , the input tensor X is of the shape of N , C, H , W , and several Transformer layers are used for feature extraction. PixelShuffle is performed on the high-order feature map after feature extraction to double its spatial resolution. Finally, the tensor is converted back to twice the scale of the original shape N , C, 2 *H ,2* W .

[0030] Secondly, a frequency-domain filtering residual module is introduced, which enhances the feature expression ability through time-frequency conversion and frequency-domain non-linear transformation. This module includes: a frequency-domain transformation layer, a complex linear transformation layer, a complex activation layer, and a time-domain inverse transformation layer. For the input feature map X 0 , assuming the shape of the input feature map is N , C, H , W(Batch size, number of channels, height, width), first flatten it to obtain a tensor with a shape of N , H´W, C X 1 , perform a fast Fourier transform (FFT) along the sequence dimension (dim = 1) to convert the time-domain features into a frequency-domain representation:

[0031] Among them, the complex features contain two components: the real part (Real) and the imaginary part (Imaginary).

[0032] Apply the same linear transformation (weight sharing) to the real and imaginary parts of the frequency-domain features respectively:

[0033]

[0034] Among them, W is the learnable weight matrix, b is the bias term. The transformed complex features are:

[0035] Use the ModReLU activation function to process the complex features, and its calculation process is: Calculate the magnitude (modulus):

[0036] Calculate the phase (angle):

[0037] Apply the phase bias and ReLU activation:

[0038] Among them, b is the learnable phase bias parameter.

[0039] Restore the frequency-domain features to the time domain through the inverse Fourier transform (IFFT):

[0040] In this embodiment, the input tensor X is rearranged into N , H*W , C shape, use a complex linear transformation layer, a complex activation layer, and a time-domain inverse transformation layer for feature extraction, and finally convert the tensor back to its original shape N , C, H , W .

[0041] In step three, the training of the image super-resolution network includes: inputting the images and related information in the data set into the improved network, using the improved network to generate high-resolution images, updating the parameters in the network according to backpropagation, so as to optimize the network parameters to reduce the reconstruction error, and finally obtaining a trained image super-resolution network.

[0042] In this embodiment, the batch size is set to 1, the learning rate is 1e-4, and the number of training epochs is set to 100 during the training process. The L1 norm loss is used as the loss function to effectively evaluate the difference between the generated image and the real image. The Adam optimizer is used to update the parameters of the network to improve the model performance.

[0043] Finally, the model of the trained image super-resolution network is obtained, and a low-resolution image can be input into the model for super-resolution to obtain a high-resolution image.

[0044] In summary, the above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A super-resolution method, characterized in that, This method uses an image super-resolution network based on the multi-head self-attention mechanism combined with the frequency-domain filtering residual module.

2. The super-resolution method according to claim 1, characterized in that This method is implemented by the following steps: Step 1: Obtain a high-resolution image dataset and generate corresponding low-resolution images; Step 2: Construct an image super-resolution network based on the combination of the frequency-domain filtering residual module and the multi-head self-attention mechanism; Step 3: Use the high-resolution images and low-resolution images to train the image super-resolution network; Step 4: Use the trained network to perform super-resolution reconstruction on the input images.

3. The super-resolution method according to claim 1, characterized in that During the construction of the adopted image super-resolution network, the dense residual connection module in the original network is replaced by the visual multi-head self-attention mechanism, which models local and global features simultaneously to avoid information loss, thereby enhancing the feature extraction ability.

4. The super-resolution method according to claim 1, characterized in that During the construction of the adopted image super-resolution network, the introduced frequency-domain filtering residual module enhances the model's expressive ability through time-frequency conversion and frequency-domain feature transformation.

5. The super-resolution method according to claim 2, characterized in that In Step 2, in the construction of the improved image super-resolution network, the improvement of the network structure includes: the dense residual connection module in the original network is replaced by the visual multi-head self-attention mechanism. The specific structure of the multi-head self-attention mechanism includes: first, through a normalization layer to adjust the mean and variance of the features, enhance the training stability, and reduce the internal covariate shift; then, through the multi-head self-attention module, enable the model to simultaneously focus on different positions of the input sequence, capture long-range dependencies, and enhance the feature representation ability; introduce a residual at the end of the multi-head self-attention to alleviate the gradient disappearance problem in deep networks; normalize the features again to stabilize the training of the subsequent multi-layer perceptron module; perform non-linear transformation on the features through at least one fully connected layer to enhance the model's fitting ability; add the input and output of the multi-layer perceptron to further retain the original information, prevent information loss, and promote gradient backpropagation to make the optimization of the deep network more stable.