Image super-resolution model based on performance and efficiency compromise consideration

Through multi-path decomposition and hourglass type variable depth distillation block design and residual learning module, the existing lightweight model has solved the problems of high computational complexity and insufficient reconstruction quality, and achieved efficient and high-quality image super-resolution reconstruction at extremely low parameters, which is suitable for mobile and embedded devices.

CN120355574APending Publication Date: 2025-07-22HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510468206.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing lightweight image super-resolution model has high computational complexity and insufficient reconstruction quality on resource-constrained devices, insufficient feature extraction, low parameter efficiency, and poor parallelism, making it difficult to maintain high-quality image reconstruction at extremely low parameters.

Method used

Multipath decomposition and hourglass-type variable depth distillation block design are adopted, combining residual learning and contrast-perceptual attention modules, feature extraction is performed by gradually combining channels and variable depth distillation blocks, reducing parameters and improving feature utilization, and retaining contrast-perceptual attention modules to enhance detail recovery.

Benefits of technology

The image reconstruction quality is significantly improved under extremely low parameters, suitable for mobile and embedded devices, achieving efficient computing and high-quality image reconstruction, especially in complex texture scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355574A_ABST
    Figure CN120355574A_ABST
Patent Text Reader

Abstract

The invention discloses an image super-resolution model based on compromise consideration of performance and efficiency, and relates to the technical field of image processing, the method adopts a multi-path structure to decompose an input channel into four paths for parallel processing, and efficient feature extraction is realized by gradually combining the channel and a variable depth distillation block; sand clock type IMDB-H3d and IMDB-2d distillation blocks are designed, and the parameter quantity is reduced by combining a structure with two small sides and a large middle part; a contrast perception attention module is reserved to enhance detail recovery; a traditional ICC module is replaced with residual learning, and cross-layer feature fusion is achieved. According to the method, the parameter quantity is only 32K under a double amplification factor and is 5% of that of an original IMDN, compared with an IMDN-RTC model, the PSNR is increased by 0.32 dB at most through 12K parameter quantity increase, and part of indexes on a test set such as Set5 exceed that of a non-lightweight model. The method is suitable for resource-limited scenes such as mobile terminals and medical images, and high-quality real-time super-resolution reconstruction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to an image super-resolution model based on a trade-off between performance and efficiency. Background Art

[0002] Image super-resolution (SR) technology aims to restore or reconstruct high-resolution (HR) images from low-resolution (LR) images to enhance the details and clarity of images. This technology has important application values in fields such as medical imaging, satellite remote sensing, and video surveillance. With the development of deep learning, super-resolution methods based on convolutional neural networks (CNNs) have become mainstream, but their computational complexity and the number of parameters are often high, making it difficult to run in real time on resource-constrained devices (such as mobile terminals and embedded systems).

[0003] The Information Multiple Distillation Network (IMDN) is an efficient lightweight super-resolution model that achieves a good balance between performance and the number of parameters through multiple information distillation blocks (IMDBs) and channel attention mechanisms (CCAs). However, the number of parameters of the standard IMDN is still as high as 694K at a 2x magnification factor, making it difficult to meet the requirements of ultra-lightweight applications. For this reason, researchers proposed the IMDN-RTC model, which compressed the number of parameters to 20K by removing the CCA module, reducing the number of channels and distillation blocks. However, this radical compression led to a significant decline in the reconstruction quality, especially in images with complex textures.

[0004] Existing lightweight solutions mainly face the following problems: 1. Insufficient feature extraction: Excessively reducing the number of channels in the feature extraction layer (such as only 12 channels in IMDN-RTC) results in a lack of input features for subsequent distillation modules, affecting the ability to recover details.

[0005] 2. Low parameter efficiency: The number of parameters in the distillation structure of traditional IMDB blocks increases quadratically with the number of input channels, making it difficult to maintain performance under the premise of lightweight.

[0006] 3. Poor parallelism: After removing the intermediate information collection (ICC) module, the model cannot fully utilize multi-level features, and the calculation is difficult to parallelize.

[0007] Therefore, there is an urgent need for a new lightweight solution to maintain or even improve the image reconstruction quality while having an extremely low number of parameters. Summary of the Invention

[0008] Aiming at the deficiencies of the prior art, the present invention provides an image super-resolution model based on a trade-off between performance and efficiency, solving the problems raised in the above background art.

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: An image super-resolution model based on a trade-off between performance and efficiency, comprising the following steps: Perform preliminary feature extraction on the input low-resolution image through a feature extraction module; Input the extracted features into a multi-path distillation part for multi-level feature distillation processing, and the multi-path distillation part adopts a structural design of channel decomposition and step-by-step merging; During the distillation processing, use a variable-depth distillation block with a hourglass structure for feature extraction; Connect the distillation outputs at all levels through a residual learning module; Finally, output a high-resolution image through an upsampling module.

[0010] Furthermore, the feature extraction module uses a 3×3 convolutional kernel, and the output channel number is set to 24.

[0011] Furthermore, the multi-path distillation part includes: Average split the input channels into 4 paths for parallel processing, with each path processing 6 channels; Gradually merge the channels, with the number of paths changing from 4 to 2 and 1, and the number of channels processed by each path increasing from 6 to 12 and 24; Set 4-path processing in the first distillation layer, with multiple distillation blocks set in the 1st and 3rd paths, and a single distillation block set in the 2nd and 4th paths; Set 2-path and 1-path processing in the second and third distillation layers respectively, with a single distillation block set in each layer.

[0012] Furthermore, the variable-depth distillation block includes two structures: IMDB-H3d block: A hourglass structure with a distillation depth of 3; IMDB-2d block: A hourglass structure with a distillation depth of 2.

[0013] Furthermore, the hourglass structure adopts a design of "small on both sides and large in the middle", with a larger convolutional kernel used in the middle layer for feature extraction, and smaller convolutional kernels used in the two side layers for dimension adjustment.

[0014] Furthermore, the residual learning module connects the distillation outputs at all levels to realize the shortcut function of the distillation part, enabling the output of each distillation block to directly reach the end of the distillation part.

[0015] Furthermore, a 3×3 convolutional layer is set after the third distillation layer for feature integration of the merged channels.

[0016] Furthermore, the upsampling module includes a 3×3 convolutional layer and a sub-pixel convolutional layer.

[0017] Furthermore, a Contrastive Perception Attention module (CCA) is retained in the distillation block to enhance image detail restoration.

[0018] The present invention provides an image super-resolution model based on a trade-off between performance and efficiency. Compared with the prior art, it has the following beneficial effects: 1. Extremely low number of parameters and efficient computation At a magnification factor of 2, the number of model parameters is only 32K, which is about 5% of the original IMDN (694K), and only 12K more parameters than IMDN-RTC (20K).

[0019] Adopting a multi-path decomposition and hourglass-shaped distillation block design significantly reduces the computational complexity and is suitable for deployment on mobile and embedded devices. 2. Multi-path distillation structure improves feature utilization The 24-channel input is split into 4 parallel paths (6 channels per path), gradually merged into 2 paths (12 channels) and a single path (24 channels), maintaining the overall feature expression ability while reducing the computational amount of a single path.

[0020] Dynamically adjust the distillation depth (IMDB-H3d and IMDB-2d), adopt deeper distillation (3 layers) for the low-channel number path and shallow distillation (2 layers) for the high-channel number path to optimize the allocation of computing resources.

[0021] 3. Hourglass-shaped distillation block enhances feature extraction efficiency Introduce an "hourglass structure with small sides and large middle" (1×1 convolution for dimensionality reduction → 3×3 convolution for feature extraction → 1×1 convolution for dimensionality increase) in the distillation block, reducing the number of parameters by about 15% compared with the traditional design.

[0022] Retain the Contrastive Perception Attention (CCA) module (only accounting for 2% of the number of parameters), and enhance the detail restoration ability through the fusion of standard deviation and mean.

[0023] 4. Residual learning replaces the ICC module Implement the "direct connection" function similar to the ICC module through cross-layer residual connections, allowing simple image features to be directly transmitted to the output layer, reducing the parameter overhead caused by splicing operations by about 30%.

[0024] Alleviate the degradation problem of deep networks and improve training stability.

[0025] 5. The reconstruction quality is significantly better than that of similar lightweight models On test sets such as Set5 and Urban100, the PSNR index comprehensively exceeds IMDN-RTC (the highest increase is 0.32dB), and in some scenarios (such as Manga109), it is even better than the VDSR model (665K) with 20 times more parameters.

[0026] It is particularly prominent in restoring complex textures (such as urban buildings and comic lines), and the edge sharpness is improved by about 15%.

[0027] 6. Wide application adaptability It can be adapted to real-time super-resolution on mobile devices (30fps@1080p, Snapdragon 865 platform).

[0028] It supports medical image enhancement (such as low-resolution CT image reconstruction) and satellite image optimization (the power consumption is reduced by 40% when the resolution is doubled). Description of the drawings

[0029] Figure 1 It is the structural diagram of IMDN-MVDND in the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Please refer to Figure 1 , the present invention provides a technical solution: an image super-resolution model based on a trade-off between performance and efficiency, including the following steps: Perform preliminary feature extraction on the input low-resolution image through a feature extraction module; Input the extracted features into a multi-path distillation part for multi-level feature distillation processing, and the multi-path distillation part adopts a structural design of channel decomposition and step-by-step merging; Adopt a variable-depth distillation block with a hourglass structure for feature extraction during the distillation processing; Connect the distillation outputs at all levels through a residual learning module; Finally, output a high-resolution image through an upsampling module.

[0032] Specifically, the lightweight image super-resolution network (IMDN-MVDND) mainly includes the following modules: (a), Feature extraction module; (b), Multi-path distillation part (including variable-depth distillation block); (c), Residual learning module; (d), Upsampling module; Among them, I. Feature extraction module Input: Low-resolution (LR) image (size: H×W×3, RGB three channels).

[0033] Processing: A 3×3 convolution kernel is used for preliminary feature extraction.

[0034] The number of output channels is set to 24 (double the 12 channels of IMDN-RTC, but much lower than the 64 channels of the original IMDN).

[0035] Output feature map size: H×W×24.

[0036] Function: While reducing the number of parameters, it provides sufficient initial feature information to prevent the subsequent distillation from affecting the reconstruction quality due to too few input features; II. Multi-path distillation section The multi-path distillation part adopts the design of channel decomposition and gradual merging to reduce the amount of calculation and maintain the feature expression ability. The specific structure is as follows: (1) First distillation layer (4-way parallel processing) Input: 24-channel feature map → split into 4 paths, each with 6 channels.

[0037] deal with: Pass 1 and Pass 3: Two IMDB-H3d blocks (distillation depth = 3, hourglass structure) are set in each.

[0038] Path 2 and Path 4: Set one IMDB-H3d block each (to reduce the amount of calculation).

[0039] Output: 4 6-channel feature maps → concatenated into 24 channels (4×6=24).

[0040] (2) Second distillation layer (2-way parallel processing) Input: 24-channel feature map → split into 2 paths, each with 12 channels.

[0041] deal with: One IMDB-2d block (distillation depth = 2, hourglass structure) is set for each channel.

[0042] Output: 2 12-channel feature maps → concatenated into 24 channels (2×12=24).

[0043] (3) Third distillation layer (single-pass processing) Input: 24-channel feature map (no longer split).

[0044] deal with: Set up 1 IMDB-2d block (distillation depth = 2).

[0045] Output: 24-channel feature map.

[0046] (4) Feature integration Add a 3×3 convolutional layer after the third distillation layer to integrate the features after multi-channel distillation. III. Variable-depth distillation block design The distillation block adopts a hourglass structure to reduce the number of parameters and improve the feature extraction efficiency.

[0047] (1) IMDB-H3d block (distillation depth = 3) Structure: Input → 1×1 convolution (dimensionality reduction) → 3×3 convolution (feature extraction) → 1×1 convolution (dimensionality increase).

[0048] Perform distillation three times, retaining some channels each time (such as 6→4→3→2).

[0049] Function: While reducing the number of parameters, maintain the feature expression ability.

[0050] (2) IMDB-2d block (distillation depth = 2) Structure: Input → 1×1 convolution (dimensionality reduction) → 3×3 convolution (feature extraction) → 1×1 convolution (dimensionality increase).

[0051] Perform distillation twice, retaining some channels each time (such as 12→8→6).

[0052] Function: Further reduce the computational load and is suitable for the subsequent distillation layer with a high number of channels.

[0053] IV. Residual learning module Design: Introduce a residual connection (Shortcut) between each level of the distillation layer.

[0054] The output of each distillation block can be directly passed to the end of the distillation part.

[0055] Function: Alleviate the degradation problem of deep networks.

[0056] Replace the ICC module of the original IMDN and reduce the number of parameters.

[0057] V. Upsampling module Structure: 3×3 convolutional layer (adjust the number of channels).

[0058] Sub-pixel convolutional layer (realize 2-fold upsampling).

[0059] Output: High-resolution (HR) image (size: 2H×2W×3).

[0060] VI. Training method (1) Dataset Training set: DIV2K (800 high - definition images).

[0061] Test sets: Set5, Set14, BSD100, Urban100, Manga109.

[0062] (2) Training parameters Optimizer: Adam (learning rate = 2×10 -4 )

[0063] Batch Size: 32.

[0064] Number of training epochs: 1000.

[0065] Hardware: NVIDIA GeForce GTX 1080 Ti.

[0066] VII. Specific parameters and results (a) Specific parameters for model training The model is trained using the training set in the DIV2K

[11] dataset. DIV2K is a high - quality single - image super - resolution dataset. There are a total of 1000 images, 800 as the training set and 200 as the test set. The characteristics of the DIV2K dataset are high resolution and covering a variety of image types.

[0067] The model in this paper is trained on NVIDIA GeForce GTX 1080 Ti. For images with a ×2 magnification factor, the Adam optimizer is used during model training: lr (learning rate) is set to 2e - 4. The model is trained for 1000 epochs. The batch_size of the model is set to 32.

[0068] (b) Test results Use Set5

[12] , Set14

[13] , BSD100

[14] , Urban100

[15] , Manga109

[16] as the test sets for evaluation.

[0069] Set5 and Set14 are small - sized datasets. They have diverse image types and are suitable for quickly judging the super - resolution performance of the model. BSD100 is 100 images selected from the Berkeley Segmentation Dataset and also has high diversity. Urban100 contains 100 urban landscape images, and the images in Urban100 have rich details. Manga109 contains 109 comic images.

[0070] The test results are as follows in the table

[0071] Table 1 Test Results of Different Methods Compared with the original IMDN-RTC model, the model in this paper, with the number of parameters between 20K and 35K, becomes a good compromise. Only by increasing the number of parameters by 12K, the super-resolution performance of the model has been greatly improved. On all 5 datasets, it outperforms the original IMDN-RTC model. The PSNR values of IMDN-MVDND on the 4 test sets of Set5, Set14, BSD100, Urban100, and Manga109 are respectively 0.13dB, 0.01dB, 0.008dB (not reflected in the table due to rounding), 0.06dB, and 0.32dB higher than those of the original IMDN-RTC model.

[0072] The super-resolution effect of this technology is better than that of all models with the number of parameters less than 32K, and it is also better than the FSRCNN model with the number of parameters of 57K. The super-resolution effect on some datasets is even higher than that of some non-ultra-lightweight models. Among the compared models, the PSNR values of the model LapSRN

[17] with the number of parameters up to 813K on Urban100 and Manga109 are 0.18dB and 0.02dB lower than those of the model in this paper. Also, the PSNR value of the model VDSR with the number of parameters up to 665K on Manga109 is also 0.02dB lower than that of the model in this paper. However, the number of parameters is about 5% and 4% of these two models.

Claims

1. An image super-resolution model based on a trade-off between performance and efficiency, characterized in that It includes the following steps: The input low-resolution image is preliminarily feature-extracted by the feature extraction module; The extracted features are input into the multi-path distillation part for multi-level feature distillation processing, and the multi-path distillation part adopts a structural design of channel decomposition and step-by-step merging; During the distillation processing, a variable-depth distillation block with a hourglass structure is used for feature extraction; The residual learning module is used to connect the distillation outputs at all levels; Finally, a high-resolution image is output through the upsampling module.

2. The image super-resolution model based on a trade-off between performance and efficiency according to claim 1, wherein The feature extraction module uses a 3×3 convolution kernel, and the number of output channels is set to 24.

3. An image super-resolution model based on a trade-off between performance and efficiency according to claim 1, characterized in that, The multi-path distillation part includes: The input channels are evenly split into 4 paths for parallel processing, and each path processes 6 channels; The channels are gradually merged, and the number of paths changes from 4 to 2 and 1, and the number of channels processed by each path increases from 6 to 12 and 24; 4-path processing is set in the first distillation layer, and multiple distillation blocks are set in the 1st and 3rd paths, and a single distillation block is set in the 2nd and 4th paths; 2-path and 1-path processing are respectively set in the second and third distillation layers, and a single distillation block is set in each layer.

4. A super-resolution image model based on a trade-off between performance and efficiency according to claim 1, characterized in that, The variable-depth distillation block includes two structures: IMDB-H3d block: A hourglass structure with a distillation depth of 3; IMDB-2d block: A hourglass structure with a distillation depth of 2.

5. An image super-resolution model based on a trade-off between performance and efficiency according to claim 4, characterized in that, The hourglass structure adopts a "small on both sides and large in the middle" design, and a larger convolution kernel is used in the middle layer to extract features, and smaller convolution kernels are used in the two side layers to adjust the dimensions.

6. An image super-resolution model based on a trade-off between performance and efficiency according to claim 1, characterized in that, The residual learning module connects the distillation outputs at all levels to implement the shortcut function of the distillation part, so that the output of each distillation block can directly reach the end of the distillation part.

7. An image super-resolution model based on a trade-off between performance and efficiency according to claim 1, characterized in that, A 3×3 convolution layer is set after the third distillation layer for feature integration of the merged channels.

8. An image super-resolution model based on a trade-off between performance and efficiency according to claim 1, characterized in that The upsampling module includes a 3×3 convolution layer and a sub-pixel convolution layer.

9. An image super-resolution model based on a trade-off between performance and efficiency according to claim 1, characterized in that, The contrast-aware attention module (CCA) is retained in the distillation block for enhancing image detail restoration.