Single Image Super-Resolution Reconstruction Method Based on Lightweight Neural Network and Transformer
By combining lightweight neural networks and Transformer models, low-frequency and high-frequency features of the image are extracted and deep-separable convolutional fusion is performed to achieve efficient super-resolution reconstruction of the image, solving the problems of large amount of calculation and blurring of details in the prior art, and high-resolution and rich details are obtained.
Patent Information
- Application Number
- CN202210463442.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The existing super-resolution reconstruction technology has problems such as large calculation volume, many parameters and high computational complexity, making it difficult to effectively improve image resolution and maintain clear details and textures.
A single image super-resolution reconstruction method based on lightweight neural networks and Transformer is adopted to extract low-frequency features through three-layer convolutional neural networks, high-frequency features are extracted using lightweight Transformer modules, and fusion features can be separated by deep convolutional networks, and finally high-resolution images are reconstructed through jump connections.
Efficient super-resolution reconstruction of low-resolution images is realized, and the resulting image is rich in details, clear texture, high spatial resolution, small calculation amount, few parameters, and low calculation complexity.
Smart Images

Figure CN114841859B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a single-image super-resolution reconstruction method based on a lightweight neural network and Transformer. Background Art
[0002] In recent years, with the vigorous development of Internet technology, images have become one of the important sources for people to obtain information. However, due to the influence of hardware performance, environmental noise, transmission methods, and storage methods, most images will undergo a degradation process, resulting in a decrease in image quality. People often cannot obtain high-resolution images. Therefore, reconstructing the image resolution is a research focus and difficulty in the field of image processing. Improving the resolution of the original image at low cost through various software algorithms is super-resolution (SR) reconstruction. The super-resolution reconstruction technology breaks through the constraints of device performance and environmental factors, saves costs and resources, and has important application value and broad application prospects in many fields.
[0003] So far, a large number of super-resolution reconstruction methods have been proposed, which are mainly divided into the following three categories: (1) interpolation-based methods; (2) reconstruction-based methods; (3) learning-based methods. Interpolation-based methods have small computational complexity, low computational cost, and fast reconstruction speed. However, the reconstructed images are relatively smooth, with blurred details and ringing artifacts. Interpolation methods mainly utilize the correlation between image pixels and obtain the pixel values of the estimated points through function formulas based on the local pixel values in the neighborhood. Common interpolation methods include: nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Reconstruction-based methods mainly utilize the image degradation model to reconstruct a single high-resolution image from the information of multiple low-resolution images. The basic idea of learning-based methods is to study the mapping relationship between low-resolution images and corresponding high-resolution images, and then use this relationship to reconstruct the input images.
[0004] With the development of computer software and hardware technologies, deep learning technology has gradually emerged and been widely applied in various fields, especially in the field of computer vision. With the proposal of various deep neural network models, the Transformer model has stood out among these models. The self-attention mechanism in the Transformer module can effectively overcome the limitations brought by convolutional inductive bias. However, it has more parameters and a larger amount of computation. Therefore, the present invention combines a lightweight network and the Transformer model to propose a single-image super-resolution reconstruction method based on a lightweight neural network and Transformer. Summary of the Invention
[0005] The object of the present invention is to overcome the deficiencies existing in the prior art. The present invention provides a single-image super-resolution reconstruction method based on a lightweight neural network and Transformer, which has a small number of parameters and a small amount of computation.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A single-image super-resolution reconstruction method based on a lightweight neural network and Transformer, characterized by comprising the following steps:
[0008] Step 1: Downsample the original high-resolution image X HR to obtain a low-resolution image X by bicubic interpolation. LR ;
[0009] Step 2: Use a three-layer convolutional neural network to extract the low-frequency feature map X of the low-resolution image X LR ;
[0010] Step 3: Feed the low-resolution image feature map X obtained in Step 2 into the main network to obtain a high-frequency feature map. Among them, the main network is composed of N cascaded lightweight Transformer modules, and each lightweight Transformer module is cascaded by a first block module, a first linear transformation module, a flattening module, a Transformer module, a second linear transformation module, and a depthwise separable convolutional neural network;
[0011] Step 4: Fuse the low-frequency feature map extracted in Step 2 with the high-frequency feature map obtained in Step 3 to obtain a reconstructed high-resolution image.
[0012] Further, in the low-frequency feature extraction of Step 2, three-layer standard convolutional neural networks are used to extract the low-frequency features of the image, which specifically include the following steps:
[0013] Step 2.1: Use a three-layer convolutional neural network to perform feature extraction on the low-resolution image X LR to obtain a feature map where 3 represents the three channels of R, G, and B, and H×W represents the size of the feature map;
[0014] Step 2.2: Uniformly divide the feature map X into N image blocks, and the i-th image block i = 1, 2,..., N; where h = H / N, w = W / N.
[0015] Further, the processing flow of the lightweight Transformer module in step 3 is as follows: The first chunking module divides the received feature map into N image chunks of the same size. Each image chunk is flattened into a one-dimensional sequence after linear transformation. The N one-dimensional sequences are output as the feature maps of N image chunks after passing through the Transformer module. The feature maps of the N image chunks are recombined into a complete feature map after linear transformation. The complete feature map formed by recombination passes through a depthwise separable convolutional neural network to generate a high-frequency feature map.
[0016] Further, the encoder and decoder in the Transformer module in step 3 are specifically as follows:
[0017] The encoder includes a first self-attention layer and a first feed-forward neural network MLP. A residual module and a normalization operation module are provided after both the first self-attention layer and the first MLP.
[0018] The decoder includes a second self-attention layer and a second feed-forward neural network. An encoder-decoder attention layer is also provided between the second self-attention layer and the second feed-forward neural network.
[0019] Further, the depthwise separable convolutional neural network in step 3 includes a pointwise convolutional layer and a depth convolutional layer. A batch normalization layer and a ReLU activation function are provided after both the pointwise convolutional layer and the depth convolutional layer.
[0020] Further, in step 4, skip connections are used to fuse the low-resolution image feature map extracted in step 2 with the high-resolution feature map obtained in step 3, and a reconstructed high-resolution image is obtained through a convolutional layer.
[0021] The beneficial effects of the present invention are as follows:
[0022] The present invention uses depthwise separable convolution to extract the fine texture features of images. Compared with ordinary convolution operations, the depthwise separable convolutional network has fewer parameters, less computational volume, and lower computational complexity. At the same time, the encoder and decoder modules of the Transformer are used, and the self-attention mechanism in the modules can effectively overcome the limitations brought by the convolutional inductive bias. This method can achieve super-resolution reconstruction of low-resolution images and obtain high-resolution images with rich details, clear textures, and high spatial resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic flow chart of the single-image super-resolution reconstruction method based on a lightweight neural network and Transformer proposed by the present invention;
[0024] Figure 2It is a schematic diagram of the single-image super-resolution reconstruction method model based on lightweight neural network and Transformer proposed by the present invention;
[0025] Figure 3 It is a schematic diagram of the Mobile-T architecture of the main network proposed by the present invention. Detailed implementation manners
[0026] The following elaborates on this solution in conjunction with the accompanying drawings and specific implementation manners.
[0027] As Figure 1 shown, the present invention discloses a single-image super-resolution reconstruction method based on lightweight neural network and Transformer. First, the present invention downsamples the original high-resolution image using bicubic downsampling to obtain a low-resolution image, and uses the obtained low-resolution image as the input of the network, and the original high-resolution image as the true annotation data during the training of the network (including a three-layer convolutional neural network, a main network, and a convolutional layer for final fusion). Secondly, a low-frequency feature extraction module (a three-layer convolutional neural network) is used to extract the spatial structure features of the low-resolution image. Then, the main network (multiple Mobile-T models) is used to extract the high-frequency information of the image, output the feature maps of each image block, and finally fuse the high-frequency and low-frequency feature maps (a convolutional layer) through skip connections to reconstruct a high-resolution image. This method can achieve super-resolution reconstruction of low-resolution images and obtain high-resolution images with rich details, clear textures, and high spatial resolution.
[0028] In one embodiment, as Figure 1 shown, the single-image super-resolution reconstruction method based on lightweight neural network and Transformer includes the following steps:
[0029] Step 1: Downsample the original high-resolution RGB image by bicubic method to obtain a low-resolution image X LR . Use the obtained low-resolution image as the input of the network, and the original high-resolution image as the true annotation data during the training of the network.
[0030] Step 2: Use a three-layer standard convolutional neural network to extract the low-frequency feature map X of the low-resolution image X LR .
[0031] Step 3: Feed the low-resolution image feature map X obtained in Step 2 into the main network to obtain a high-frequency feature map; wherein, the main network is composed of N lightweight Transformer modules cascaded, and each lightweight Transformer module is cascaded by a first block module, a first linear transformation module, a flattening module, a Transformer module, a second linear transformation module, and a depthwise separable convolutional neural network.
[0032] Step 4: Fuse the low-frequency feature map extracted in Step 2 with the high-frequency feature map obtained in Step 3 to obtain a reconstructed high-resolution image.
[0033] Step 1 includes constructing the dataset required for the network. In the present invention, we use the DIV2K dataset as the training dataset for the experiment. First, perform bicubic downsampling preprocessing on the images in the dataset, and then use the preprocessed low-resolution images as the input of the network model, and the original high-quality images as the true annotation data to make the training dataset.
[0034] The specific downsampling operation is as follows:
[0035] X LR = f(X HR )
[0036] where X HR is the original high-resolution image, f(·) represents the bicubic downsampling operation, and X LR is the low-resolution image obtained after downsampling.
[0037] As Figure 2 shown, the network model includes a low-frequency feature extraction module, a main network, and a reconstruction module. Among them, the main network is composed of N lightweight Transformer modules cascaded. As Figure 3 shown, each lightweight Transformer module is cascaded by a first block module, a first linear transformation module, a flattening module, a Transformer module, a second linear transformation module, and a depthwise separable convolutional neural network; the reconstruction module fuses the low-frequency feature map extracted by the low-frequency feature extraction module with the high-frequency feature map output by the main network, and obtains a reconstructed high-resolution image through a convolutional layer. The main network can realize the extraction of fine features of the image.
[0038] The low-frequency feature extraction module is composed of three layers of standard convolutional neural networks cascaded. The specific method for extracting the low-frequency features of the image is as follows:
[0039] The convolutional kernel size of the three layers of standard convolutional neural networks is 3×3:
[0040] L j = Max(0, W j * X LR + B j ), j = 1, 2, 3
[0041] In the formula, X LR is the input low-resolution image, W j is the convolutional kernel of the j-th layer of the convolutional neural network layer, B j is the bias of the j-th layer of the convolutional neural network layer, and * represents the convolutional operation.
[0042] The first and second chunking modules chunk the received feature maps: The feature map X is evenly divided into N image chunks X of a fixed size i , i = 1, 2, …, N. Among them, Here, 3 refers to the three channels of R, G, and B, and H×W represents the size of the feature map. Each segmented feature image chunk i = 1, 2, …, N, where N is the number of image chunks, h = H / N, w = W / N.
[0043] The encoder and decoder in the Transformer module are as follows:
[0044] Encoder: Using the standard Transformer encoder architecture, for each encoder, it mainly contains two layers of structures. One is the self-attention layer (Self-Attention), which is used to obtain context semantic information. The other is the multi-layer perceptron (MLP). In addition, there is a residual module and a normalization operation after each self-attention layer and multi-layer perceptron layer, and their role is to accelerate the convergence of the model, prevent gradient vanishing or gradient explosion. Among them, the specific calculation formula for the self-attention layer to output attention values is as follows:
[0045]
[0046] Among them, Q refers to the query matrix, K refers to the key matrix, V refers to the value matrix, and d k refers to the input vector of the main network, that is, the dimension of the image chunk. After the three matrices are calculated, the attention values are output through the softmax function.
[0047] Decoder: The decoder uses the standard Transformer decoder architecture. It mainly includes three modules: the self-attention layer, the multi-layer perceptron, and the encoder-decoder attention layer. The decoder outputs the feature maps of N image chunks, and then the image chunk features are recombined into a complete feature map through a linear transformation.
[0048] The depthwise separable convolution module is specifically as follows: Using depthwise separable convolution to further extract the high-frequency features of the image. Different from the traditional depthwise separable convolution, it first uses pointwise convolution to deepen the dimension of the feature channels, and then uses depthwise convolution to achieve feature fusion. After each convolutional layer (depthwise convolution, pointwise convolution), there are batch normalization and ReLU activation functions. After N Mobile-T modules, the high-frequency feature map X′ is finally generated. The specific convolution operation is as follows:
[0049] L point = ReLU(W4 * L last + B4)
[0050] L dep = ReLU(W5 * L point + B5)
[0051] where L point represents a pointwise convolutional layer, W4 and B4 respectively represent the convolutional kernel and bias of the pointwise convolutional layer, L dep represents a depth convolutional layer, W5 and B5 respectively represent the convolutional kernel and bias of the depth convolutional layer, L last is the output of the previous layer, and * represents the convolution operation.
[0052] The reconstruction module is specifically as follows:
[0053] The low-resolution image feature map extracted by the low-frequency feature extraction module is added to the high-resolution feature map output by the main network using a skip connection, and a convolutional layer is used to obtain a high-resolution image. The specific fusion formula is as follows:
[0054] HR = X' * W + X
[0055] In the formula, HR is the finally reconstructed high-resolution image, X' represents the high-resolution feature map output by the main network, W represents the convolutional kernel of this convolutional layer, X represents the low-resolution image feature map, and * represents the convolution operation.
[0056] The above method is only a relatively reasonable implementation manner of the present invention. The protection scope of the present invention is not limited to the above implementation method. Any equivalent modifications or changes made by those of ordinary skill in the art according to the content disclosed by the present invention should be included in the protection scope recorded in the claims.
Claims
1. A single-image super-resolution reconstruction method based on a lightweight neural network and Transformer, characterized in that, It includes the following steps: Step 1: Downsample the original high-resolution image X HR to obtain the low-resolution image X by bicubic interpolation LR ; Step 2: Use a three-layer convolutional neural network to extract the low-frequency feature map X of the low-resolution image X LR ; Step 3: Feed the low-resolution image feature map X obtained in Step 2 into the main network to obtain a high-frequency feature map. The main network is composed of N lightweight Transformer modules cascaded. Each lightweight Transformer module is composed of a first block module, a first linear transformation module, a flattening module, a Transformer module, a second linear transformation module, and a depthwise separable convolutional neural network cascaded. Step 4: Fuse the low-frequency feature map extracted in Step 2 with the high-frequency feature map obtained in Step 3 to obtain a reconstructed high-resolution image. The processing flow of the lightweight Transformer module in Step 3 is as follows: The first block module divides the received feature map into N image blocks of the same size. Each image block is flattened into a one-dimensional sequence after linear transformation. The N one-dimensional sequences output the feature maps of N image blocks after passing through the Transformer module. The feature maps of the N image blocks are recombined into a complete feature map after linear transformation. The complete feature map formed by recombination passes through the depthwise separable convolutional neural network to generate a high-frequency feature map. The depthwise separable convolutional neural network in Step 3 includes a pointwise convolutional layer and a depth convolutional layer. After the pointwise convolutional layer and the depth convolutional layer, there are a batch normalization layer and a ReLU activation function respectively.
2. The single-image super-resolution reconstruction method based on lightweight neural network and Transformer according to claim 1, characterized in that In the low-frequency feature extraction of Step 2, a three-layer standard convolutional neural network is used to extract the low-frequency features of the image, which specifically includes the following steps: Step 2.1: Use a three-layer convolutional neural network to extract features from the low-resolution image X LR to obtain a feature map where 3 represents the three channels of R, G, and B, and H×W represents the size of the feature map; Step 2.2: Uniformly divide the feature map X into N image patches, and the i-th image patch where h = H / N, w = W / N.
3. The single-image super-resolution reconstruction method based on a lightweight neural network and Transformer according to claim 1, characterized in that The encoder and decoder in the Transformer module in Step 3 are specifically as follows: The encoder includes a first self-attention layer and a first feedforward neural network MLP. After the first self-attention layer and the first MLP, there is a residual module and a normalization operation module respectively. The decoder includes a second self-attention layer and a second feedforward neural network. There is also an encoder-decoder attention layer between the second self-attention layer and the second feedforward neural network.
4. The single-image super-resolution reconstruction method based on a lightweight neural network and Transformer according to claim 1, wherein In Step 4, skip connections are used to fuse the low-frequency feature map extracted in Step 2 with the high-frequency feature map obtained in Step 3, and a reconstructed high-resolution image is obtained through a convolutional layer.
Citation Information
Patent Citations
Image processing method and device, storage-calculation integrated chip and electronic equipment
CN114049255A
Image super-resolution enhancement method based on bidirectional recursion convolution neural network
WO2017219263A1