High-definition image real-time lightweight compressed sensing method based on convolutional neural network
By using initial linear reconstruction branches, DAP modules and adaptive weighted merge reconstruction modules of different reconstruction scales in high-definition image compression perception, combined with L1 loss function training, the problem of high computing resource consumption in high-definition image compression perception is solved, and the effect of real-time lightweight compression perception is achieved.
Patent Information
- Application Number
- CN202510350127.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-08
AI Technical Summary
The existing image compression perception method based on convolutional neural networks has the problem that computing resource consumption is high or difficult to meet real-time requirements in high-definition image applications.
The initial linear reconstruction branch, DAP module and adaptive weighted merge reconstruction module of different reconstruction scales are used, and the L1 loss function is trained to achieve real-time lightweight compression perception of high-definition images.
It effectively reduces the computing resource usage of high-definition image compression perception, improves computing speed and reconstruction accuracy, and meets the real-time compression perception needs of high-definition images.
Smart Images

Figure CN120455691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image perception method, and more particularly to a real-time lightweight compressed perception method for high-definition images based on a convolutional neural network. Background Art
[0002] Image compression sensing is an emerging image signal processing technology that recovers the original signal through a small number of linear observations, thereby breaking the limitations of the Nyquist sampling theorem. It is an important application of compressed sensing theory in the field of image processing, which can effectively improve the transmission efficiency of images and reduce the amount of stored data. In recent years, with the widespread application of deep learning technology, more and more researchers have combined deep convolutional neural networks with image compression sensing and proposed methods such as ISTA-Net. + OPINE-Net + Advanced image compression sensing methods such as AMP-Net and CSNet have emerged. Compared to traditional image compression sensing methods (such as orthogonal matching pursuit, fast iterative shrinkage thresholding, and approximate message passing), image compression sensing methods based on convolutional neural networks (CNNs) offer higher computational efficiency and reconstruction accuracy, gradually replacing them. However, the test image resolutions used to evaluate the performance of CNN-based ... Summary of the Invention
[0003] In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a real-time lightweight compressed sensing method for high-definition images based on convolutional neural networks to solve the above technical problems.
[0004] To achieve the above object, the present invention provides the following technical solution: a real-time lightweight compressed sensing method for high-definition images based on convolutional neural networks, characterized in that it includes the following steps:
[0005] Step 1: The input HD image size is determined to be C×W×H, where C is the number of HD image channels, W and H are the width and height of the HD image respectively. Then, a convolution layer with a convolution kernel of B×B and a stride of B is used to perform adaptive sampling of the HD image, and a size of rCB is obtained.2 ×W / B×H / B compressed sensing measurement value y;
[0006] Step 2: The high-definition image compressed sensing measurement y sampled in step 1 is first passed through a 1×1 convolutional layer to expand the number of channels of the measurement value y, and then two initial linear reconstruction branches with different reconstruction scales are set for upsampling;
[0007] Step 3: Downsampling. In the downsampling deep nonlinear reconstruction, the number of channels is first expanded through a 3×3 convolution layer, and then a nonlinear activation is performed using the PReLU function. The network is then fed into three series-connected DAP modules to complete a more downscaled deep nonlinear reconstruction.
[0008] Step 4: The feature map fed into the DAP module is convolutionally downsampled with a step size of 2 or 4 times its size to obtain a feature map of size W / 4 and H / 4 or W / 8 and H / 8. After batch normalization, it is nonlinearly processed using the PReLU activation function and then fed into two series-connected residual modules improved by the SE attention module.
[0009] Step 5: The feature map of the improved residual module passes through a 3×3 convolution layer, a batch normalization layer, a PReLU activation layer, and then another 3×3 convolution layer and batch normalization layer before being sent to the SE attention module.
[0010] Step 6: After the feature map processed by the two improved Residual modules in series, it is reconstructed to the size of the DAP module input feature map by double or quadruple upsampling through Pixelshuffle. After passing through a 1×1 convolution layer, it is skip-connected with the feature map of the input DAP module and fused as the output of the DAP module.
[0011] Step 7: After executing three rounds of DAP module processing in series from step 4 to step 6, the feature map is first resized to 4C×W / 2×H / 2 through a 3×3 convolutional layer. Then, it is resized to the original image size of C×W×H by using Pixelshuffle and then aligned with the feature map of the branch that was initially linearly reconstructed to the original image size.
[0012] Step eight, adaptively weighting and reconstructing the reconstructed image obtained by nonlinearly reconstructing the downsampled depth to the size of the original image and the reconstructed image obtained by initial linear reconstruction to the size of the original image.
[0013] As a further improvement of the present invention, the two branches in step 2 are specifically as follows: one branch implements B-fold upsampling through Pixelshuffle, directly reconstructs the expanded compressed sensing measurement feature map into a feature map of the original image size of C×W×H, and completes the initial linear reconstruction to the original image size; the other branch uses Pixelshuffle to only complete B / 2-fold upsampling, reconstructs the feature map to a size of 4C×W / 2×H / 2, and only restores to 1 / 4 of the original image size, and based on this, implements downsampling deep nonlinear reconstruction.
[0014] As a further improvement of the present invention, the specific method of sending the input to the SE attention module in step five is: the input feature map first passes through an adaptive average pooling layer to compress the 2D feature map of each channel into a 1×1 value, and then passes through a 1×1 convolution layer, a ReLU activation layer, a 1×1 convolution layer and a Sigmoid layer in sequence to obtain the weight value of each channel, and then the feature map input to the SE attention module is weighted and output.
[0015] As a further improvement of the present invention, the specific method of adaptively weighted merging and reconstructing the reconstructed image of the downsampled depth nonlinear reconstruction to the size of the original image and the reconstructed image of the initial linear reconstruction to the size of the original image in step eight is as follows: first, the reconstructed images output by the two branches are connected in series and fused into a feature map of size 2C×W×H; secondly, a dual-channel feature map of size W×H is obtained through a 1×1 convolutional layer, and after processing by the Softmax function, weights corresponding to the reconstructed image sizes W×H of the two branches are generated, and then the reconstructed images output by the two branches are adaptively weighted merged; finally, the merged image is processed through a 3×3 convolutional layer, and the final merged reconstruction is completed at the original image size, and then the reconstructed high-definition image is output.
[0016] As a further improvement of the present invention, a training step is further included, wherein the L1 loss function is used in the training step to complete the end-to-end training of the network, specifically as follows:
[0017]
[0018] Where x i and They represent the i-th original image and the corresponding reconstructed image in the training set respectively.
[0019] The beneficial effects of the present invention are as follows: a real-time lightweight compressed sensing method for high-definition images based on convolutional neural networks of the present invention aims at the problems of difficulty in real-time calculation or large computational space resource consumption in the reconstruction process of high-definition images by mainstream image compressed sensing methods. Firstly, by setting different initial linear reconstruction scale branches, the high-definition image compressed sensing measurement feature map is directly subjected to initial linear reconstruction of different reconstruction scales, and then only the initial linear reconstruction image reconstructed to 1 / 4 of the original image size is subjected to deep nonlinear reconstruction, which greatly reduces the computational resource occupation of high-definition image compressed sensing reconstruction; secondly, the convolution downsampling technology is introduced in the deep nonlinear reconstruction process, and the DAP module is established, so that the deep nonlinear reconstruction process is further improved. The method is completed on the downscaled feature map, thereby further reducing the computing resource consumption required for the deep nonlinear reconstruction of high-definition images and improving the computing speed; thirdly, an adaptive weighted merging reconstruction module of the initial linear reconstructed image and the deep nonlinear reconstructed image is proposed, which realizes adaptive fusion reconstruction of the decoupled initial linear reconstruction process and the deep nonlinear reconstruction process, and fully improves the reconstruction accuracy of the compressed perception of high-definition images compared with the undecoupled mode or the residual connection mode; finally, the L1 loss function is used for training, which further improves the quality of high-definition image reconstruction under the condition of reducing computing resource occupancy of the method of the present invention, thereby effectively improving the accuracy, efficiency and applicability of compressed perception of high-definition images. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a network structure diagram of the real-time lightweight compressed sensing method for high-definition images based on convolutional neural networks of the present invention;
[0021] Figure 2 This is a network structure diagram of the SE attention module used in the method of the present invention;
[0022] Figure 3 Schematic diagram showing the comparison of high-definition image reconstruction results between the method of the present invention and the mainstream advanced image compression sensing method based on convolutional neural network when the compression ratio is 0.01;
[0023] Figure 4 Schematic diagram comparing the high-definition image reconstruction results of the method of the present invention and the mainstream advanced convolutional neural network-based image compression sensing method when the compression ratio is 0.1. DETAILED DESCRIPTION
[0024] The present invention will be further described below with reference to the embodiments shown in the accompanying drawings.
[0025] Reference Figure 1 As shown, a real-time lightweight compressed sensing method for high-definition images based on a convolutional neural network in this embodiment includes the following steps:
[0026] Step 1: Assume that the input HD image size is C×W×H, where C is the number of HD image channels and is 3 for RGB images. W and H are the width and height of the HD image, respectively. Use a convolution layer with a convolution kernel of B×B and a stride of B to complete the adaptive sampling of the HD image and obtain a size of rCB. 2 The compressed sensing measurement value y is calculated as ×W / B×H / B. r is the compression ratio, which can be 0.01, 0.04, 0.1, etc. for high-definition images, and B is generally 32 or 33.
[0027] In step 2, the sampled high-definition image compressed sensing measurement y is first passed through a 1×1 convolutional layer to expand the number of channels of the measurement value y. Then, two initial linear reconstruction branches with different reconstruction scales are set. One branch implements B-fold upsampling through Pixelshuffle, and directly reconstructs the expanded compressed sensing measurement feature map into a feature map of the original image size of C×W×H, completing the initial linear reconstruction to restore it to the original image size; the other branch uses Pixelshuffle to only complete B / 2-fold upsampling, reconstructing it into a feature map of size 4C×W / 2×H / 2, restoring it to only 1 / 4 of the original image size, and based on this, implements downsampling deep nonlinear reconstruction.
[0028] In step 3, in the downsampling deep nonlinear reconstruction, the number of channels is first expanded through a 3×3 convolution layer, and then a nonlinear activation is performed using the PReLU function, and then sent to three series-connected DAP modules to complete a more downscaled deep nonlinear reconstruction.
[0029] In step 4, the feature map fed into the DAP module is convolutionally downsampled with a step size of 2 or 4 according to its size, and a feature map of size W / 4 and H / 4 (corresponding to 2 times the step size convolution downsampling) or size W / 8 and H / 8 (corresponding to 4 times the step size convolution downsampling) is obtained. After batch normalization, the PReLU activation function is used for nonlinear processing, and then the feature map is fed into two series-connected residual modules improved by the SE attention module.
[0030] Step 5: The feature map of the improved residual module is fed into a 3×3 convolution layer, a batch normalization layer, a PReLU activation layer, and then another 3×3 convolution layer and batch normalization layer before being fed into the SE attention module. Figure 2 The figure shows the network structure of the SE attention module used in the present invention. The input feature map first passes through an adaptive average pooling layer to compress the 2D feature map of each channel into a 1×1 value, and then passes through a 1×1 convolution layer, a ReLU activation layer, a 1×1 convolution layer and a Sigmoid layer in sequence to obtain the weight value of each channel. The feature map of the input SE attention module is then weighted and output.
[0031] In step 6, the feature map processed in series by the two improved Residual modules is reconstructed to the size of the DAP module input feature map by double or quadruple upsampling through Pixelshuffle, and then skip-connected with the feature map of the input DAP module after passing through a 1×1 convolution layer. After fusion, it is output as the DAP module.
[0032] Step 7: After three rounds of DAP module serial processing from step 4 to step 6, the feature map is first resized to 4C×W / 2×H / 2 through a 3×3 convolutional layer, and then upsampled by Pixelshuffle to reconstruct the feature map of the original image size C×W×H. The feature map is aligned with the feature map of the branch that was initially linearly reconstructed to the original image size.
[0033] In step 8, the reconstructed image obtained by nonlinearly reconstructing the downsampled depth to the size of the original image and the reconstructed image obtained by linearly reconstructing the original image are adaptively weighted and merged. First, the reconstructed images output by the two branches are connected in series and fused into a feature map of size 2C×W×H. Secondly, a dual-channel feature map of size W×H is obtained through a 1×1 convolutional layer. After processing with the Softmax function, weights corresponding to the size W×H of the reconstructed images of the two branches are generated. The reconstructed images output by the two branches are then adaptively weighted and merged. Finally, the merged image is processed through a 3×3 convolutional layer to complete the final merge reconstruction at the original image size, and the reconstructed high-definition image is output.
[0034] Step 9: To further improve the reconstruction accuracy of the high-definition image compression sensing method of the present invention, the L1 loss function is used to complete the end-to-end training of the network during the training phase. The calculation formula is as follows:
[0035]
[0036] Where x i and Respectively, they represent the i-th original image and the corresponding reconstructed image in the training set. Compared to the L2 loss function used in existing methods, the method of the present invention also uses the L2 loss function for training to compare the results of training using the L1 loss function. As shown in Table 1, at several common compression ratios of 0.01, 0.04, and 0.1, the PSNR and SSIM test accuracies obtained by training the method of the present invention using the L1 loss function are superior to those obtained by training using the L2 loss function, verifying the superiority of the method of the present invention using the L1 loss function for training.
[0037] This embodiment provides the following example, using the DIV2K high-definition image dataset for training and testing, selecting several commonly used compression ratios of 0.01, 0.04, and 0.1, and comparing the method of the present invention with the mainstream image compression sensing method. Figure 3 and 4 The comparison results of the reconstruction images of the method of the present invention and the mainstream advanced image compression sensing method based on convolutional neural network when the compression ratio is 0.01 and 0.1 respectively. To ensure the fairness of the comparison, the brightness component of the image is used to calculate the evaluation index, and the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are selected as the evaluation indicators of the reconstructed image quality accuracy. Figure 3 and Figure 4 As shown, the corresponding PSNR and SSIM values are also noted below the reconstructed image of each method to help illustrate the reconstruction effect of each method. Figure 3 and Figure 4 From the enlarged image of the reconstructed image details, it can be seen that the detailed features reconstructed by the method of the present invention are clearer and more detailed than those of other mainstream methods, and its PSNR and SSIM results are also the highest among all methods. Tables 2 and 3 respectively give the average PSNR results and average SSIM results of different methods on the test set under three compression ratios, and the method of the present invention has achieved the best. At the same time, the experiment configured an NVIDIA A40 graphics card to test the computing time and video memory occupancy of different methods. As shown in Table 4, the average computing time and GPU video memory occupancy of the mainstream methods for achieving compressed sensing of a high-definition image under a compression ratio of 0.1 are compared. It can be seen from the table that although the video memory occupancy and computing speed of the method of the present invention are not optimal, the method of the present invention has the best comprehensive space occupancy and time computing efficiency among all methods. The computing time of the method of the present invention is almost the same as that of the fastest method, but the video memory space required is close to 1 / 3 of the space occupied by the fastest method. At the same time, the method with less video memory occupancy than the method of the present invention requires nearly 27 times the computing time required by the method of the present invention. Therefore, compared with mainstream methods for processing high-definition image compression sensing tasks, the method of the present invention is the best method to balance time and space complexity, and it is also the best method to reconstruct image quality.
[0038] It should be noted that the different reconstruction scales and network parameter selections in the above embodiment are only one of the best embodiments of the present invention, and the patent scope of the present invention cannot be limited by this embodiment alone.
[0039] Table 1 Comparison of the average PSNR (dB) results of the proposed method and mainstream image compression sensing methods on the DIV2K high-definition image test set under different compression ratios
[0040]
[0041] Table 2 Comparison of the average SSIM results of the proposed method and mainstream image compression sensing methods on the DIV2K high-definition image test set under different compression ratios
[0042]
[0043] Table 3 Performance comparison of the proposed method and mainstream image compression sensing methods in reconstructing a single high-definition image (1800×1196) at a compression ratio of 0.1
[0044]
[0045] Table 4 Comparison of reconstruction quality of the proposed method using L1 and L2 loss functions on the DIV2K high-definition image test set
[0046]
[0047] In summary, the present invention proposes a real-time lightweight method for high-definition image compressed sensing tasks based on convolutional neural networks;
[0048] This lightweight method proposes initial linear reconstruction branches at different reconstruction scales, directly performing initial linear reconstruction of the HD image compressed sensing measurement feature map at different reconstruction scales. This includes a Pixelshuffle linear reconstruction branch that restores the image to the original size and a Pixelshuffle linear reconstruction branch that restores the image to 1 / 4 the original size. Compared to mainstream image compressed sensing methods that perform initial linear reconstruction to the original image size and then directly perform deep nonlinear reconstruction, which results in high computational resource consumption for HD image compressed sensing, this method only performs deep nonlinear reconstruction on the initial linearly reconstructed image at 1 / 4 the original image size, effectively reducing the computational resources required for HD image compressed sensing and improving computational speed.
[0049] A DAP module is proposed to achieve downsampled deep nonlinear reconstruction. The module includes a double or quadruple convolution downsampling layer, a batch normalization layer, a nonlinear activation layer, two series-connected residual modules improved by the SE attention module, a corresponding double or quadruple pixelshuffle layer, a convolution layer, and a skip connection. The improved residual module of the SE attention module specifically consists of a convolution layer, a batch normalization layer, a PReLU activation function, a convolution layer, a batch normalization layer, and an SE attention module connected in series. The DAP module completes the deep nonlinear reconstruction process on the downscaled feature map by utilizing downsampling technology and pixelshuffle technology, thereby improving the computational speed of deep reconstruction and fully reducing the space resource usage of the deep nonlinear reconstruction process.
[0050] An adaptive weighted merging reconstruction module is proposed. The module includes first connecting the reconstructed image that is initially linearly reconstructed to the size of the original image in series with the reconstructed image that is output after alignment of the downsampled deep nonlinear reconstruction, then compressing it into a feature map of two channels through convolution, generating the corresponding weight feature map using Softmax, and then weighted fusion of the above-mentioned initial linear reconstructed image and the deep nonlinear reconstructed image, and finally processing it through a convolution layer to output the final reconstructed high-definition image. Compared with the residual connection mode in the invention [2], the method of the present invention proposes a mode of adaptive weighted merging reconstruction of the initial linear reconstructed image and the deep nonlinear reconstructed image, realizing the adaptive weighted fusion of the corresponding pixels of the initial linear reconstructed image and the deep nonlinear reconstructed image through the Softmax function, and performing only one convolution transformation on the merged reconstructed image at the size of the original image, thereby further improving the quality of the reconstructed high-definition image on the basis of reducing the computing resource occupation.
[0051] It is proposed to use the L1 loss function to replace the L2 loss function used in mainstream image compression sensing methods (including the invention [1,2]). Experiments have verified that the method of the present invention can more effectively improve the image reconstruction quality of high-definition image compression sensing using the L1 loss function than the L2 loss function.
[0052] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A real-time lightweight compressed sensing method for high-definition images based on convolutional neural networks, characterized by: The steps include: Step 1: The input HD image size is determined to be C×W×H, where C is the number of HD image channels, W and H are the width and height of the HD image respectively. Then, a convolution layer with a convolution kernel of B×B and a stride of B is used to perform adaptive sampling of the HD image, and a size of rCB is obtained. 2 ×W / B×H / B compressed sensing measurement value y; Step 2: The high-definition image compressed sensing measurement y sampled in step 1 is first passed through a 1×1 convolutional layer to expand the number of channels of the measurement value y, and then two initial linear reconstruction branches with different reconstruction scales are set for upsampling; Step 3: Downsampling. In the downsampling deep nonlinear reconstruction, the number of channels is first expanded through a 3×3 convolution layer, and then a nonlinear activation is performed using the PReLU function. The network is then fed into three series-connected DAP modules to complete a more downscaled deep nonlinear reconstruction. Step 4: The feature map fed into the DAP module is convolutionally downsampled with a step size of 2 or 4 times its size to obtain a feature map of size W / 4 and H / 4 or W / 8 and H / 8. After batch normalization, it is nonlinearly processed using the PReLU activation function and then fed into two series-connected residual modules improved by the SE attention module. Step 5: The feature map of the improved residual module passes through a 3×3 convolution layer, a batch normalization layer, a PReLU activation layer, and then another 3×3 convolution layer and batch normalization layer before being sent to the SE attention module. Step 6: After the feature map processed by the two improved Residual modules in series, it is reconstructed to the size of the DAP module input feature map by double or quadruple upsampling through Pixelshuffle. After passing through a 1×1 convolution layer, it is skip-connected with the feature map of the input DAP module and fused as the output of the DAP module. Step 7: After executing three rounds of DAP module processing in series from step 4 to step 6, the feature map is first resized to 4C×W / 2×H / 2 through a 3×3 convolutional layer. Then, it is resized to the original image size of C×W×H by using Pixelshuffle and then aligned with the feature map of the branch that was initially linearly reconstructed to the original image size. Step eight, adaptively weighting and reconstructing the reconstructed image obtained by nonlinearly reconstructing the downsampled depth to the size of the original image and the reconstructed image obtained by initial linear reconstruction to the size of the original image.
2. The method for real-time lightweight compressed sensing of high-definition images based on convolutional neural networks according to claim 1 is characterized in that: The two branches in step 2 are specifically as follows: one branch implements B-fold upsampling through Pixelshuffle, directly reconstructs the expanded compressed sensing measurement feature map into a feature map of the original image size of C×W×H, and completes the initial linear reconstruction to restore to the original image size; the other branch uses Pixelshuffle to only complete B / 2-fold upsampling, reconstructs the feature map into a size of 4C×W / 2×H / 2, and only restores to 1 / 4 of the original image size, and based on this, implements downsampling deep nonlinear reconstruction.
3. The method for real-time lightweight compressed sensing of high-definition images based on convolutional neural networks according to claim 1 or 2, characterized in that: The specific method of feeding into the SE attention module in step 5 is as follows: the input feature map first passes through an adaptive average pooling layer to compress the 2D feature map of each channel into a 1×1 value, and then passes through a 1×1 convolution layer, a ReLU activation layer, a 1×1 convolution layer and a Sigmoid layer in sequence to obtain the weight value of each channel, and then the feature map input to the SE attention module is weighted and output.
4. The method for real-time lightweight compressed sensing of high-definition images based on convolutional neural networks according to claim 1 or 2, characterized in that: The specific method of adaptively weighting and merging the reconstructed image obtained by nonlinearly reconstructing the downsampled depth to the size of the original image and the reconstructed image obtained by initial linear reconstruction to the size of the original image in step eight is as follows: first, the reconstructed images output by the two branches are connected in series and fused into a feature map of size 2C×W×H; second, a dual-channel feature map of size W×H is obtained through a 1×1 convolutional layer, and after processing by a Softmax function, weights corresponding to the size W×H of the reconstructed images of the two branches are generated, and then the reconstructed images output by the two branches are adaptively weighted and merged; finally, the merged image is processed through a 3×3 convolutional layer, and the final merged reconstruction is completed at the original image size, and then the reconstructed high-definition image is output.
5. The method for real-time lightweight compressed sensing of high-definition images based on convolutional neural networks according to claim 1 or 2, characterized in that: It also includes a training step, in which the L1 loss function is used to complete the end-to-end training of the network, as follows: Where x i and They represent the i-th original image and the corresponding reconstructed image in the training set respectively.
Citation Information
Cited By
Image depth compressed sensing method and system oriented to multi-level privacy protection
CN121357293A