Blind Compressed Image Super-Resolution Reconstruction Based on Multi-Scale Channel Pyramid Residual Attention

By using the multi-scale channel pyramid residual attention mechanism and the predictive quality factor of the classified convolution neural network in the super-resolution reconstruction of compressed images, the problem of artifact amplification in super-resolution reconstruction of compressed images is solved, and efficient image high-frequency information recovery and quality improvement are achieved.

CN115496652BActive Publication Date: 2025-06-24SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110679946.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-18
Publication Date
2025-06-24
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle super-resolution reconstruction of compressed images, especially in image recovery under unknown quality factors, which are prone to amplification of compressed block artifacts and have problems with discomfort.

Method used

The blind compressed image super-resolution method based on multi-scale channel pyramid residual attention is adopted to predict the quality factor of the compressed image by classifying convolutional neural networks, and the multi-scale channel pyramid residual attention extracts and fuses image information of different channels and depths to build an end-to-end compressed image super-resolution network.

Benefits of technology

It effectively solves the problem of super-resolution reconstruction of compressed images under unknown quality factors, improves the high-frequency information recovery ability of the image, significantly improves the quality of the reconstruction results, and obtains high PSNR and SSIM values.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496652B_ABST
    Figure CN115496652B_ABST
Patent Text Reader

Abstract

The present invention discloses a blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention, which mainly includes the following steps: separately training a QF value segmented prediction sub-network for compressed images, a QF segmented decompression effect sub-network, and an image super-resolution sub-network, and using a cascaded three-sub-network to form an end-to-end network for training. Taking the image downsampled by the JPEG algorithm as the input, and through the network model trained above, the finally super-resolved image is obtained. In the image super-resolution feature extraction stage, efficient multi-scale channel pyramid residual attention is added to fuse different channel information and image features of different depths in the same channel, so as to restore more high-frequency information. The method described in the present invention can specifically suppress the blocking effect of JPEG compressed images, reconstruct high-resolution images, and the obtained subjective visual effects and objective evaluation indicators show that the present invention is an effective compressed image super-resolution restoration method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of compressed image super-resolution reconstruction, and particularly to a blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention, belonging to the image restoration direction in the field of digital image processing. Background Art

[0002] With the development of the Internet and video technology, people have higher requirements for the quality of the collected images. However, since the image degradation process is unknown, there are many similar low-resolution images for different high-resolution image blocks. The super-resolution problem of obtaining a high-resolution image from a single low-resolution image is a challenge.

[0003] In order to save bandwidth and data storage space, most of the images we obtain are compressed, which inevitably introduces image block effects. If traditional super-resolution processing is directly performed on the image, the compression block artifacts will inevitably be amplified. Since the super-resolution reconstruction of a single image has an ill-posed problem, certain prior knowledge or a cleverly designed algorithm (such as predicting image degradation information) is required to improve the result of image super-resolution. We divide the compressed image super-resolution algorithm into compressed image QF value segmented prediction, QF segmented de-compression effect, and image super-resolution sub-processes, train them separately, and finally form an end-to-end network. By adding efficient multi-scale channel pyramid residual attention in the image super-resolution feature extraction stage, different channel information is fused, and the image feature relationship of different depths in the same channel is studied to recover more high-frequency information. Summary of the Invention

[0004] The purpose of the present invention is to use a classification convolutional neural network to predict the quality factor of a compressed image, and use efficient multi-scale channel pyramid residual attention to extract and fuse image information of different channels and different depths in the same channel of a low-resolution image, so as to construct an effective compressed image super-resolution method based on quality factor prediction.

[0005] The blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention proposed by the present invention mainly includes the following operation steps:

[0006] (1) For a low-resolution compressed image with an unknown quality factor, a blind compressed image super-resolution model is proposed, the compressed image super-resolution problem is decomposed to obtain a compressed image QF value segmented prediction sub-problem, a QF segmented de-compression effect sub-problem, and an image super-resolution sub-problem, and these problems are solved separately;

[0007] (2) For the compressed image QF value segmented prediction sub-problem described in step (1), compressed images of different QF segments generated by the JPEG compression algorithm are used as the training set, and a QF fuzzy classification prediction network is designed and built;

[0008] (3) For the QF segmented decompression effector problem described in step (1), design and build a QF segmented decompression effect network, and train different decompression effect models for the compressed images of different QF segments;

[0009] (4) Incorporate the image downsampling constraint, and design and build an image super-resolution convolutional neural network with multi-scale channel pyramid residual attention;

[0010] (5) Cascade the networks designed in step (3) and step (4) to form an end-to-end compressed image super-resolution network for joint training;

[0011] (6) Use the compressed low-quality image with an unknown QF value as the input, and utilize the model trained in step (2) to obtain a rough range of the QF value of the compressed image;

[0012] (7) Use the JPEG compressed image as the input, and utilize the corresponding network model trained in step (5), combined with the optimized reconstruction cost function, to obtain a high-resolution image. Description of the Drawings

[0013] Figure 1 is a block diagram of the blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention of the present invention.

[0014] Figure 2 is the composition of the effective dense connection channel attention module in the QF segmented decompression stage of the compressed image.

[0015] Figure 3 is the composition of the multi-scale channel pyramid residual attention module in the image super-resolution stage.

[0016] Figure 4 is the composition of the multi-scale channel pyramid attention module in the multi-scale channel pyramid residual attention module.

[0017] Figure 5 (a) is the wide activation residual block (WARB), Figure 5 (b) is the enhanced wide activation residual block (EWARB).

[0018] Figure 6 is a comparison chart of the reconstruction results of the test image "Bikes" by the present invention and eight methods (the super-resolution reconstruction factor is 2, and the JPEG compression quality factor is 10): Among them, Figure 6 (a) is the test image, Figure 6(b), (c), (d), (e), (f), (g), (h), (i), and (j) are the results of bicubic interpolation, comparison method 1, comparison method 2, comparison method 3, comparison method 4, comparison method 5, comparison method 6, comparison method 7, comparison method 8, and the reconstruction result of the present invention, respectively.

[0019] Figure 7 This is a comparison chart of the reconstruction results of the present invention and eight methods for the test image "Statue" (the super-resolution reconstruction factor is 2, and the JPEG compression quality factor is 20): Among them, Figure 7 (a) is the test image, Figure 7 (b), (c), (d), (e), (f), (g), (h), (i), and (j) are the results of bicubic interpolation, comparison method 1, comparison method 2, comparison method 3, comparison method 4, comparison method 5, comparison method 6, comparison method 7, comparison method 8, and the reconstruction result of the present invention, respectively.

[0020] Figure 8 This is a comparison chart of the reconstruction results of the present invention and eight methods for the test image "Monarch" (the super-resolution reconstruction factor is 2, and the JPEG compression quality factor is 30): Among them, Figure 8 (a) is the test image, Figure 8 (b), (c), (d), (e), (f), (g), (h), (i), and (j) are the results of bicubic interpolation, comparison method 1, comparison method 2, comparison method 3, comparison method 4, comparison method 5, comparison method 6, comparison method 7, comparison method 8, and the reconstruction result of the present invention, respectively. Detailed implementation manners

[0021] The present invention will be further described below with reference to the accompanying drawings:

[0022] Figure 1 In the blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention, the following steps are included:

[0023] (1) For a low-resolution compressed image with an unknown quality factor, a blind compressed image super-resolution model is proposed, and the compressed image super-resolution problem is decomposed into a compressed image QF value segmented prediction sub-problem, a QF segmented decompression effect sub-problem, and an image super-resolution sub-problem, and these problems are solved separately;

[0024] (2) For the compressed image QF value segmented prediction sub-problem described in step (1), compressed images with different QF segments generated by the JPEG compression algorithm are used as the training set, and a QF fuzzy classification prediction network is designed and built;

[0025] (3) Designing and building a QF segmented decompression effect network for the QF segmented decompression effect sub-problem described in step (1), and training different decompression effect models for compressed images of different QF segments;

[0026] (4) Integrate image downsampling constraints, design and build a multi-scale channel pyramid residual attention image super-resolution convolutional neural network;

[0027] (5) cascading the networks designed in step (3) and step (4) to form an end-to-end compressed image super-resolution network for joint training;

[0028] (6) Taking the compressed low-quality image with unknown QF value as input, using the model trained in step (2), obtain the approximate range of the QF value of the compressed image;

[0029] (7) Taking the JPEG compressed image as input, using the corresponding network model trained in step (5) combined with the optimized reconstruction cost function, a high-resolution image is obtained.

[0030] Specifically, in step (1), the convolutional neural network model structure is as follows: Figure 1 As shown in the figure, the constructed model consists of a compressed image QF value segmentation prediction network, a QF segmentation decompression effect network, and an image super-resolution network.

[0031] In the step (2), in the QF value segmented prediction network, the network is composed of 8 convolutional layers and a fully connected layer. The convolution kernel of each convolution layer is 3×3, and the number of output channels of the convolutional layer is doubled every two layers. The image block size is reduced on the convolutional layer with a moving step of 2. A random cropping method is used to extract 9 image blocks of a fixed size of 128×128 from the compressed image to be tested, and the image blocks are input into the QF prediction network. After each two layers of convolution, the image block size is reduced to half of the original size. The feature matrix output by the Conv8 convolution is flattened into a one-dimensional feature vector and input into the fully connected layer, and finally output after being processed by the Softmax activation function. Specifically, the parameters of the network are uniformly set to η, and N network training samples are represented as {(X (1) ,q (1) ),...,(X (N) ,q (N) )}, where X (i) represents the grayscale image sample block of the i-th input, q (i) Indicates the true QF value corresponding to the sample. represents the nth activation value of the kth input image block in the fully connected layer FC2. The activation function of the image block after Softmax can be expressed as follows:

[0032]

[0033] In the formula is the 7-dimensional probability vector of the k-th input image, which represents the probability of the image at seven specified QF nodes. The cross-entropy loss function is used to optimize and train the network, and this process can be expressed as:

[0034]

[0035] Finally, after processing the one-dimensional feature vector by the Softmax activation function, the final QF value predicted by the network is output to approximate the true QF result of the image to be measured.

[0036] In the step (3), the QF segmented decompression effect sub-problem can be divided into three parts: shallow feature extraction, cross-channel information fusion and channel attention enhancement model, and image reconstruction process. The initial image feature extraction H FE1 is used to generate shallow features for deep learning and provide a wider receptive field for the network. Dense connection channel attention H DCEA fully utilizes the initial features and effectively reduces compression noise. Finally, the reconstruction module H REC reconstructs the image vector after removing compression artifacts into an image. Specifically, the extraction of shallow features of the image can be expressed by the following formula:

[0037] F0 = H FE1 (X)

[0038] where H FE1 is the initial feature extraction, which is achieved by overlapping and taking blocks of the JPEG compressed image with a convolution stride of 2, and then taking F0 as the input of the subsequent layer to generate deep features for image restoration.

[0039] Fr = H DCEA (F0)

[0040] where H DCEA is the dense connection channel attention module, and Fr is the output of DCEA. Finally, the image vector passes through the overlapping block reconstruction module to obtain the LR image without compression artifacts, which is expressed as:

[0041] Z = H REC (F r )

[0042] The dense connection channel attention (DCEA) module takes the output fWR of the enhanced wide activation residual block (EWARB) as the input, and the structure of EWARB is as Figure 5As shown in (b), the number of input and output channels of convolutional layers C1 and C4 is 64, the number of input channels of convolutional layer C2 is 64, and the number of output channels is 128. The number of input channels of convolutional layer C3 is 128, and the number of output channels is 64. EWARB improves the information transmission efficiency while keeping the number of network parameters and computational complexity unchanged. DCEA cascades an effective channel attention (ECA) and dilated convolution with dense connections to make full use of feature information. The ECA module adaptively selects the size of the one-dimensional convolutional kernel to determine the coverage range of local cross-channel interaction of information. The output of the dense connection channel attention is expressed as follows:

[0043] f DCEA =H DCEA (f WR )+f WR

[0044] In the ECA module of the present invention, the size of the convolutional kernel is taken as 5, and then information exchange and fusion are carried out between the high-dimensional channel and the low-dimensional channel. Finally, dense connection information extraction and enhancement are carried out by the following method:

[0045] v=F fir (f ECA )+W 5×5 (W 3×3 (Re(F fir (f ECA ))))+Re(W 5×5 (W 1×1 (F IFEB )))

[0046] Among them, F IFEB =W 5×5 (W 3×3 (Re(F fir (f ECA )))) is the information exchange and fusion module, f ECA represents the output of the ECA module, W 3×3 represents the weight of the dilated convolution with a 3×3 convolutional kernel, W 5×5 represents the weight of the dilated convolution with a 5×5 convolutional kernel, W 1×1 represents the weight of the ordinary convolution with a 1×1 convolutional kernel.

[0047] In step 4, the image super-resolution sub-process is divided into three parts: initial feature extraction, multi-scale channel pyramid residual attention feature extraction, and sub-pixel convolutional layer reconstruction. The purpose of feature extraction is to form the initial features of super-resolution. The multi-scale channel pyramid residual attention is used to predict high-frequency details and improve the visual quality. Finally, the sub-pixel convolutional layer generates the final super-resolution reconstructed image as the output. We use a convolutional layer to extract the shallow features of the image:

[0048] F SR = H FE2 (Z)

[0049] In the formula, H FE2 represents the convolutional operation for extracting shallow features.

[0050] Then we extract and enhance the deep features from the shallow features through cascaded wide activation residual blocks (WARB) and multi-channel pyramid residual attention. H MCPRA represents the multi-scale channel pyramid residual attention (MCPRA) function. f WR represents the input of MCPRA and the output of WARB, which directly forms a residual structure in the network. H MCPRA is composed of the cascading of channel attention and pyramid residual attention, and finally added to the shallow features of the low-resolution image to obtain the final output f CCPRA :

[0051] f CCPRA = H MCPRA (f WR ) + f WR

[0052] In the channel attention, first, a global average pooling layer (GAP) is introduced to pay more attention to the most valuable information of the low-resolution image by using the global information of the feature map. Let p be the output of GAP. Its input f WR has a size of H×W, and the output of the channel attention module can be expressed as the following formula:

[0053] u = s(W u Re(W d p))

[0054] where s() and Re() represent the sigmoid and ReLU functions, and W u and W d represent the weights of the upsampling and downsampling channel convolutions. We can obtain f CA after passing through the channel attention module:

[0055] f CA = u × f WR

[0056] Although both the channel attention and pyramid residual attention mechanisms have their respective advantages, they also have their own disadvantages. The channel attention can explore the relationships between different channels, but in the same channel, the information at different depths has different dependencies. Therefore, we cascade a pyramid attention after the channel attention, and finally multiply the output of the pyramid attention by the input. This multi-scale channel pyramid residual attention module can be expressed by the following formula as

[0057] H(x) = (1 + M(f CA ) + P1(P2(f CA ))) × V(f CA )

[0058] Among them, H(x) represents the output feature, x represents the original feature, V represents the convolution operation, M is the parameter obtained by training the attention mask, P1 and P2 are the parameters of the two-layer pyramid network. After cascading the channel attention and pyramid residual attention mechanisms, MCPRA can extract more information than a single attention, thereby improving the performance of the network.

[0059] In the image super-resolution process, given the training set D i , H i represent the low-resolution and high-resolution images respectively. The loss function of image super-resolution can be expressed as:

[0060]

[0061] Among them, Θ ISR represents all the trainable parameters in the image super-resolution network, and H ISR represents the mapping function of the network.

[0062] In step (5), the three networks are trained separately. The image with an unknown quality factor is input into the QF fuzzy classification prediction network model to obtain the QF value of the compressed image. The QF decompression network and the image super-resolution network are cascaded to form an end-to-end network. The compressed image is input into the corresponding QF value decompression end-to-end network to obtain the final image blind super-resolution reconstruction result.

[0063] To better illustrate the effectiveness of the present invention, we selected 10 test images (including Bikes, Circuit, House, Leaves, Monarch, Parrots, Peppers, Statue, Woman, Zebra) from the public datasets Set12 and Set14. Simulate the generation method of low-resolution images compressed by JPEG. Use the bicubic interpolation method to downsample the high-resolution test images by a factor of 2, and then compress the sampled images with JPEG at different compression quality factors. The compressed images are the images to be reconstructed.

[0064] In the experiment, the compared compressed image super-resolution reconstruction methods are:

[0065] Method 1: FSRCNN: The method proposed by Chao et al., reference "D. Chao, C. L. Chen, and X. Tang. Accelerating the super-resolution convolutional neural network. In European Conference on Computer Vision, 2016."

[0066] Method 2: SRCDFOE: The method proposed by Jia et al., reference "X. Jia, W. Chen, and X. Hu. Single image super-resolution in compressed domain based on field of expert prior. In Image and Signal Processing (CISP), 2012 5th International Congress on, 2012."

[0067] Method 3: ICDBSR: The method proposed by Li et al., reference "T. Li, X. He, L. Qing, Q. Teng, and H. Chen. An iterative framework of cascaded deblocking and super-resolution for compressed images. IEEE Transactions on Multimedia, 2017."

[0068] Method 4: VDSR: The method proposed by Kim et al., reference "J. Kim, J. K. Lee, and K. M. Lee. Accurate image super-resolution using very deep convolutional networks. In IEEE Conference on Computer Vision & Pattern Recognition, 2016."

[0069] Method 5: DNCNN-3 + VDSR: The method DNCNN-3 proposed by Kai et al., reference "K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. IEEE Transactions on Image Processing, 26(7): 3142–3155, 2017." The method VDSR proposed by Kim et al., reference "J. Kim, J.K. Lee, and K.M. Lee. Accurate image super-resolution using very deep convolutional networks. In IEEE Conference on Computer Vision & Pattern Recognition, 2016."

[0070] Method 6: CISRDCNN: The method proposed by Chen et al., reference "H. Chen, X. He, C. Ren, L. Qing, and Q. Teng. Cisrdcnn: Super-resolution of compressed images using deep convolutional neural networks. NEUROCOMPUTING, 285: 204–219, 2018."

[0071] Method 7: MEMNET+RCAN: The MEMNET method proposed by Tai et al., see reference "Y. Tai, J. Yang, X. Liu, and C. Xu. Memnet: A persistent memory network for image restoration. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 4549–4557, Los Alamitos, CA, USA, oct 2017. IEEE Computer Society." The RCAN method proposed by Zhang et al., see reference "Y. Zhang, K. Li, K. Li, L. Wang, B. Zhong, and Y. Fu. Image super-resolution using very deep residual channel attention networks. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss, editors, Computer Vision–ECCV 2018, pages 294–310, Cham, 2018. Springer International Publishing."

[0072] The content of the comparative experiment is as follows:

[0073] Table 1, Table 2, and Table 3 are respectively the low-resolution test image sets generated by degrading 10 test images with 2x downsampling and JPEG algorithms with quality factors (QFs) of 10, 20, 30, and 40, and then performing 2x decompression super-resolution reconstruction. The super-resolution reconstruction results of some images are as Figure 6 , Figure 7 , Figure 8 shown. The objective evaluation results of the reconstructed images are shown in Tables 1 to 3. PSNR (Peak Signal to Noise Ratio, unit: dB) and SSIM (Structure Similarity Index) are used to evaluate the reconstruction effect respectively. The higher the values of PSNR / SSIM, the better the reconstruction effect.

[0074] Table 1 Comparison of PSNR values of 10 test images (QF = 10 or QF = 20, super-resolution ×2)

[0075]

[0076] Comparison of PSNR values of 10 test images (QF = 30 or QF = 40, super-resolution × 2)

[0077]

[0078] Comparison of average SSIM values of 10 test images (QF = 10, 20, 30, 40, super-resolution × 2)

[0079]

[0080] It can be seen from Table 1, Table 2 and Table 3 that the present invention has achieved higher PSNR and SSIM. In Figure 5 the "Bikes" test image, the present invention can reconstruct clearer text "21" compared with the comparative method; in Figure 6 the "Statue" test image, the statue reconstructed by the present invention is more realistic in terms of head texture and details; in Figure 7 the "Monarch" test image, the reconstructed petal contours and texture details of the present invention are clearer.

[0081] In summary, compared with the comparative method, the reconstruction result of the present invention has great advantages in both subjective and objective evaluations. Therefore, the present invention is an effective method for super-resolution reconstruction of compressed images.

Claims

1. Blind compressed image super-resolution reconstruction based on multi-scale channel pyramid residual attention, characterized in that Including the following steps: Step 1: For a low-resolution compressed image with an unknown quality factor, a blind compressed image super-resolution model is proposed. The super-resolution problem of the compressed image is decomposed to obtain a sub-problem of segmental prediction of the QF value of the compressed image, a sub-problem of QF segment-based decompression effect, and a sub-problem of image super-resolution, and these problems are solved separately; Step 2: For the sub-problem of segmental prediction of the QF value of the compressed image described in Step 1, compressed images of different QF segments generated by the JPEG compression algorithm are used as the training set, and a QF fuzzy classification prediction network is designed and built; Step 3: For the sub-problem of QF segment-based decompression effect described in Step 1, a QF segment-based decompression effect network is designed and built, and different decompression effect models are trained for compressed images of different QF segments to effectively suppress the compression effect for compressed images within a wide QF range; Step 4: Incorporating the image downsampling constraint, a multi-scale channel pyramid residual attention-based image super-resolution convolutional neural network is designed and built; Step 5: The networks designed in Step 3 and Step 4 are cascaded to form an end-to-end compressed image super-resolution network for joint training; Step 6: Using the compressed low-quality image with an unknown QF value as the input, and utilizing the model trained in Step 2, the approximate range of the QF value of the compressed image is obtained; Step 7: Using the JPEG compressed image as the input, and utilizing the corresponding network model trained in Step 5, combined with the optimized reconstruction cost function, a high-resolution image is obtained; It is characterized by the QF segmented prediction classification network described in step 2: Different from the traditional deep learning-based decompression algorithm for compressed images that trains a single network to map compressed images to images without compression effects, the output of the QF prediction network of the present invention does not use the prediction result of a single image block as the result of the network. Instead, it extracts the QF value range with the highest frequency of occurrence in the QF nodes from 9 different blocks as the result approximating the true QF of the image. This sub-network finally uses the sigmoid activation function and finally selects the decompression effect model corresponding to the node to decompress the image; It is characterized by the QF segmented decompression network described in step 3: Different from the single QF value decompression model, the QF segmented decompression network divides the image into seven sub-segments with QF values of {1-10, 11-20, …, 51-60} and {61-100} in the JPEG codec to generate a mixed compression sample set for different segment generalization networks, and trains a decompression network with dense connection and effective attention for the corresponding QF segments. In the network construction, a wide activation residual block and effective channel dense connection attention are introduced. The wide activation residual block improves the information transmission efficiency while keeping the number of network parameters and computational complexity unchanged. The effective channel dense connection attention captures local cross-channel interactions to obtain the most valuable image information while reducing the computational complexity and improves the information utilization rate; It is characterized by the construction of an image super-resolution convolutional neural network based on multi-scale channel pyramid residual attention described in step 4: The present invention constructs a multi-scale channel pyramid residual attention network by studying the image degradation process. It is composed of a channel attention cascaded with a multi-scale pyramid residual attention. In this way, the network not only pays attention to the relationship between different channel information, but also pays attention to the dependence of different depth information in the same channel. Finally, the information extracted from different pyramid layers is feature fused to reconstruct a high-resolution image.

Citation Information

Patent Citations

  • JPEG compressed image super-resolution reconstruction method based on convolutional neural network

    CN107563965A

  • Compressed multi-scale feature fusion network-based image super-resolution reconstruction method

    CN108537731A