Hybrid adaptive perception super-resolution reconstruction model and reconstruction method for complex scene
Through the hybrid adaptive perceptual super-resolution reconstruction model, the problems of inflexible resource allocation and limited feature extraction capabilities in complex scenes are solved, and efficient image super-resolution reconstruction is achieved. It is suitable for portable devices and improves image quality and computational efficiency.
Patent Information
- Application Number
- CN202510515272.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing single-image super-resolution reconstruction methods based on deep learning have inflexible resource allocation strategies and limited local feature extraction capabilities when processing complex scenes, resulting in difficulty in recovering detailed information and low efficiency in global information modeling, making it difficult to efficiently deploy them on portable devices while maintaining high resolution.
A hybrid adaptive perceptual super-resolution reconstruction model for complex scenes is adopted, including a shallow feature extraction module, a deep feature extraction module and a reconstruction module. Through a dynamic adaptive hybrid module, a pixel fusion convolutional feedforward network and an enhanced spatial channel perception attention module, computing resources are dynamically allocated to improve the quality of feature extraction and reconstruction.
It achieves high-quality and efficient image super-resolution reconstruction in complex scenes, taking into account both local detail preservation and global consistency modeling. It is suitable for deployment on mobile devices, improving visual effects and model deployment friendliness.
Smart Images

Figure CN120672570A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image reconstruction and relates to a hybrid adaptive perceptual super-resolution reconstruction model, method, device and storage medium for complex scenes. Background Art
[0002] With the widespread adoption of image acquisition devices such as smartphones and digital cameras, images have become an essential medium for daily communication and information transfer. Image quality directly impacts the visual experience, and high-resolution, high-definition images are becoming increasingly necessary. However, due to hardware limitations and environmental interference, captured images often suffer from low resolution and blurred details, making them difficult to meet practical requirements. Therefore, efficient and feasible technical means to improve image quality and enhance visual effects are urgently needed, which has important research value and application significance.
[0003] Traditional image super-resolution reconstruction methods fall into three main categories: interpolation-based, reconstruction-based, and learning-based. Interpolation methods, such as nearest neighbor, bilinear, and bicubic interpolation, are computationally simple but have poor detail recovery capabilities and can easily blur images. Reconstruction methods, such as maximum a posteriori estimation, iterative backprojection, and convex set projection, can recover some high-frequency information but are computationally complex and rely on degradation models. Learning methods, which effectively reconstruct image details by constructing mappings between low- and high-resolution images, are the current mainstream technology.
[0004] Against this backdrop, the Transformer architecture has gradually attracted the attention of researchers due to its global modeling capabilities enabled by its self-attention mechanism. It can exploit internal self-similarity in images to capture long-range dependencies, demonstrating superior performance in recovering global texture details. However, existing Transformer and CNN architectures still have numerous limitations, such as uneven distribution of computing resources, limited feature extraction capabilities, and performance degradation when handling complex scenes (such as those requiring high-frequency textures or large receptive fields).
[0005] Existing lightweight models generally adopt a uniform computation strategy, which makes them difficult to adapt flexibly when dealing with complex scenes. ClassSR reduces computational complexity through classification strategies, but its pre-classification mechanism limits flexibility, especially in areas with complex textures. For details, see "Kong X, Zhao H, Qiao Y, et al. Classsr: A general framework to accelerate super-resolution networks by data characteristic[C] / / Proceedingsof the IEEE / CVF conference on computer vision and pattern recognition. 2021: 12016-12025."; ARM improves the model's adaptability through meta-training, but still struggles to recover details in extremely complex scenes. The complexity of flat areas (such as the sky) and high-frequency areas (such as building textures) differs significantly. Fixed computational paths can easily lead to wasted resources in simple areas and insufficient reconstruction of complex areas, causing artifacts in areas with overlapping textures. For details, see "Chen B, Lin M, Sheng K, et al. Arm: Any-time super-resolution method [C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022: 254-270."; PAN (pixel attention network) enhances the representation of detailed areas through pixel-level attention. For details, see "Zhao H, Kong X, He J, et al. Efficient image super-resolution using pixel attention [C] / / Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16. Springer International Publishing, 2020: 56-72."; However, these methods still find it difficult to strike a balance between model lightweighting and reconstruction quality in practical applications.
[0006] The inventors have discovered that existing deep learning-based single-image super-resolution reconstruction methods have poor reconstruction results, mainly due to the following shortcomings: (1) insufficient flexibility in resource allocation strategies, making it difficult to cope with the diverse needs of complex scenes; (2) limited local feature extraction capabilities, resulting in difficulty in recovering detailed information; and (3) low efficiency in global information modeling, which affects overall image quality. The current problem to be solved is how to maintain high resolution while ensuring that no information is lost, and how to enable efficient deployment on portable devices, balancing performance and model complexity. Summary of the Invention
[0007] In order to solve the above problems, the present invention proposes a hybrid adaptive perceptual super-resolution reconstruction model and reconstruction method for complex scenes, which solves the problems of excessive computational complexity, serious loss of high-frequency information and poor visual effect of reconstructed images in super-resolution reconstruction tasks.
[0008] The second object of the present invention is to provide a hybrid adaptive perceptual super-resolution reconstruction model and reconstruction method for complex scenes.
[0009] A third object of the present invention is to provide an electronic device.
[0010] A fourth object of the present invention is to provide a computer storage medium.
[0011] The technical solution adopted by the present invention is a hybrid adaptive perceptual super-resolution reconstruction model for complex scenes, comprising a shallow feature extraction module, a deep feature extraction module and a reconstruction module;
[0012] The shallow feature extraction module uses a standard 3×3 convolution to expand the input image channel and map it to a high-dimensional feature space to extract shallow feature information, which is used to extract shallow features of low-resolution images and obtain effective information;
[0013] The deep feature extraction module is composed of 24 dynamic adaptive hybrid modules (DAFMixer) connected in series, which is used to extract deep high-frequency detail information;
[0014] The DAFMixer module includes a dynamic adaptive perceptual mixing module (DAPMixer), a pixel fusion convolutional feedforward network (PFCFN) and enhanced spatial channel-aware attention (SCEA).
[0015] The reconstruction module adds the input low-resolution image through a 3×3 convolution layer and a sub-pixel convolution module to the image upsampled by bicubic interpolation to obtain a clear reconstructed image.
[0016] Furthermore, the deep feature extraction module consists of 24 dynamic adaptive perceptual hybrid modules (DAPMixer) modules;
[0017] The DAPMixer module included in the DAFMixer first obtains the value V through point-wise convolution (PWConv) of the input features, the global predictor obtains the offset map, the mixer mask and simple spatial / channel attention, obtains the attention area ratio through the attention branch, and then passes through the convolution branch to achieve a more sophisticated self-attention mechanism in complex areas and use basic convolution operations in simple areas.
[0018] The global predictor included in DAPMixer, based on the fusion of local features and global conditions, first generates an intermediate feature map for estimating offsets and attention weights. The offset guides flexible sampling of window positions, and the mask is used to filter valid regions. Simultaneously, spatial attention and channel attention branches are introduced in parallel, focusing on key locations and channel importance, respectively. The final fusion generates a joint channel-spatial attention map, which performs weighted feature enhancement and improves subsequent reconstruction quality.
[0019] The attention branch included in the DAPMixer first performs bilinear interpolation on the input feature map under the guidance of an offset matrix to obtain aligned sparse window features. During the training phase, Gumbel Softmax is used to achieve binary sampling. During the inference phase, a top-M filtering method is used to obtain the position index of the high-attention region. The sampling window is then linearly mapped to generate query, key, and value vectors. Attention weights are calculated through scaled dot products and weighted fusion is performed to obtain a sparse attention feature map.
[0020] The convolution branch contained in the DAPMixer is composed of a longitudinal depth-wise separable convolution DWConv of size 1×(2d-1), a transverse depth-wise separable convolution DWConv of size (2d–1)×1, a longitudinal depth-wise separable convolution DWConv of size 1×d / 2, and a transverse depth-wise separable convolution DWConv of size d / 2×1.
[0021] The PFCFN module included in the DAFMixer consists of a LayerNorm layer, a pixel attention module, two fully connected layers (FC1 and FC2), and a nonlinear activation function module (Act). The input and output are added through a residual connection structure to enhance the stability of feature transfer.
[0022] The SCEA module included in the DAFMixer is composed of an ECA attention branch and a backbone feature extraction branch in parallel, which is used to fuse channel relationships and spatial structure information. The ECA branch first performs average pooling on the input feature map to extract channel information, then captures local channel dependencies through one-dimensional convolution, and then upsamples the channel attention features to restore the original size. The backbone branch extracts multi-scale spatial information through strided convolution and maximum pooling, then introduces depthwise separable convolution (DWConv3) for lightweight feature extraction, and then restores spatial resolution through upsampling. The features output by the two branches are spliced and fused in the Concat module, and the channel information is further integrated through a 1×1 convolution, and then the fused attention map is generated using the Sigmoid function. Finally, the attention map is element-wise multiplied with the original input features of the main branch to output the fused feature map.
[0023] A hybrid adaptive perceptual super-resolution reconstruction method for complex scenes is characterized by following the steps below:
[0024] S1: low-resolution input image I LR Perform shallow feature extraction processing, the shallow feature extraction function is recorded as H SFE (·), the extraction result is the shallow feature F0, and the calculation formula is:
[0025] F0=H SFE (I LR ) (6)
[0026] Among them, I LR is the input low-quality image, F0 is the shallow feature of the extracted LR image, H SFE It is a convolution layer operation with a convolution kernel of 3×3.
[0027] S2: The deep feature extraction module consists of 24 dynamic adaptive hybrid modules (DAFMixer) connected in series. Each DAFMixer module includes a dynamic adaptive perception module (DAPMixer), a pixel fusion convolutional feedforward network (PFCFN), and an enhanced spatial channel perception module (SCEA). The modules use a residual connection structure to achieve feature preservation and enhancement. The shallow feature map F0 is used as input and is sequentially input into each DAFMixer module for processing to achieve multi-level deep feature extraction. The processing process is as follows:
[0028]
[0029] Among them, F0 represents the output feature map of the shallow feature extraction module, F n (n=1···24) represents the output of the n-th DAFMixer, The function representing the nth DAFMixer, F DAFMixer It indicates the final fusion feature after this process.
[0030] The processing in each DAFMixer module can be further expanded into the following composite structure:
[0031]
[0032] Among them, F input The feature map of the DAFMixer module representing the input, F output Describe the output feature map, Indicates that after n SCEA modules are operated, represents n PFCFN operations, Indicates n DAPMixer module operations, n is 24;
[0033] S3: The output feature F of the last DAFMixer module n Input to a 3×3 convolution layer for feature integration to obtain the fusion feature F DAFMixer , and its calculation formula is:
[0034] F DAFMixer =Conv3(F n ) (9)
[0035] Among them, Conv3(·) represents the convolution operation with a 3×3 convolution kernel, F DAFMixer is the final fusion feature.
[0036] S4: The fused deep feature F DAFMixer Concatenate (Concat) or element-wise weighted fusion with the shallow feature F0 in step S1 and input to the image reconstruction module H IRB (·), and obtain the final reconstructed image I SR , whose expression is:
[0037]
[0038] Among them, F0 is the shallow feature of the extracted LR image, F DAFMixer Indicates the final fusion feature after this process, H IRB represents the image reconstruction function, which usually includes sub-pixel convolution operations to achieve upsampling, represents the feature fusion operation, I SR Represents the final output high-resolution image.
[0039] An electronic device adopts the above method to realize super-resolution reconstruction of a single image.
[0040] A computer storage medium stores at least one program instruction, which is loaded and executed by a processor to achieve super-resolution reconstruction of the above-mentioned complex image.
[0041] The beneficial effects of the present invention are:
[0042] In response to the problems of insufficient edge detail recovery, limited global modeling capabilities, and difficulty in efficiently deploying models to mobile devices in existing super-resolution reconstruction methods, the present invention proposes a lightweight dynamic adaptive hybrid reconstruction method DAFMixer. This method introduces a dynamic adaptive perceptual hybrid module (DAPMixer) to dynamically allocate computing resources according to the content of the image region, thereby achieving focused modeling of complex texture areas; constructs a pixel fusion convolutional feedforward network (PFCFN) to improve pixel-level feature expression capabilities and effectively eliminate redundant information; and designs an enhanced spatial channel attention module (SCEA) to parallel optimize spatial and channel information and dynamic position encoding strategies. Each module is collaboratively optimized in terms of feature extraction, attention guidance, and context modeling, taking into account both local detail preservation and global consistency modeling, ultimately achieving high-quality and efficient super-resolution reconstruction of images with good visual effects and model deployment friendliness. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 2 is a schematic structural diagram of a reconstruction model according to an embodiment of the present invention.
[0045] Figure 2 It is a structural diagram of the dynamic adaptive perception mixing module DAPMixer in the reconstruction model of an embodiment of the present invention.
[0046] Figure 3 It is a schematic diagram of the pixel fusion convolutional feedforward network PFCFN structure in the reconstruction model of an embodiment of the present invention.
[0047] Figure 4 It is a schematic diagram of the enhanced spatial channel perception attention SCEA structure in the reconstruction model of an embodiment of the present invention.
[0048] Figure 5This is a comparison of the subjective visual effects of the reconstruction method of the embodiment of the present invention and other algorithms for the reconstruction images of the "img092 and img008" in Urban100 with a magnification of ×4; the reconstruction image of the "barbara" image in Set14 with a magnification of ×2; and the reconstruction image of the "img86000" image in B100 with a magnification of ×3. DETAILED DESCRIPTION
[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0050] Example 1,
[0051] A hybrid adaptive perceptual super-resolution reconstruction model for complex scenes, whose structure is as follows Figure 1 As shown, it includes a shallow feature extraction module, a deep feature extraction module and a reconstruction module;
[0052] The shallow feature extraction module uses a standard 3×3 convolution to expand the input image channel and map it to the high-dimensional feature space to extract shallow feature information, which is used to extract shallow features of low-resolution images and obtain effective information;
[0053] like Figure 2The DAFMixer module shown here includes a dynamic adaptive perceptual mixing module (DAPMixer), a pixel-fused convolutional feedforward network (PFCFN), and enhanced spatial channel-aware attention (SCEA). The DAPMixer module first applies point-wise convolution (PWConv) to the input features to obtain a value V. The global predictor obtains an offset map, a mixer mask, and simple spatial / channel attention. The attention branch obtains the proportion of the attention region, and then passes it through the convolution branch. This enables a more refined self-attention mechanism in complex regions, while using basic convolution operations in simple regions. The global predictor in DAPMixer, based on the fusion of local features and global conditions, first generates an intermediate feature map for estimating offsets and attention weights. The offset is used to guide the flexible sampling of window positions, and the mask is used to select valid regions. Simultaneously, spatial attention and channel attention branches are introduced in parallel, focusing on key positions and channel importance, respectively. Finally, they are fused to generate a joint channel-spatial attention map, which performs weighted feature enhancement and improves subsequent reconstruction quality. After passing through the attention branch, the input feature map is first adjusted by bilinear interpolation under the guidance of the offset matrix to obtain aligned sparse window features. During the training phase, Gumbel Softmax is used to achieve binary sampling. During the inference phase, a top-M filtering method is used to obtain the position index of the high-attention region. The sampling window is then linearly mapped to generate query, key, and value vectors. Attention weights are calculated through scaled dot products and weighted fusion is performed to obtain a sparse attention feature map. Finally, the convolution branch is used, which consists of a vertical depthwise separable convolution DWConv of size 1×(2d-1), a horizontal depthwise separable convolution DWConv of size (2d-1)×1, a vertical depthwise separable convolution DWConv of size 1×d / 2, and a horizontal depthwise separable convolution DWConv of size d / 2×1.
[0054] like Figure 3 As shown in the figure, the PFCFN module consists of a LayerNorm layer, a pixel attention module, two fully connected layers (FC1 and FC2), and a nonlinear activation function module (Act). The input and output are added through a residual connection structure to enhance the stability of feature transfer.
[0055] like Figure 4As shown in the figure, the SCEA module consists of an ECA attention branch and a backbone feature extraction branch in parallel, which is used to fuse channel relationships and spatial structure information. The ECA branch first performs average pooling on the input feature map to extract channel information, then captures local channel dependencies through one-dimensional convolution, and then upsamples the channel attention features to restore the original size. The backbone branch extracts multi-scale spatial information through strided convolution and maximum pooling, then introduces depthwise separable convolution (DWConv3) for lightweight feature extraction, and then restores the spatial resolution through upsampling. The features output by the two branches are spliced and fused in the Concat module, and the channel information is further integrated through a 1×1 convolution. The fused attention map is then generated using the Sigmoid function. Finally, the attention map is element-wise multiplied with the original input features of the main branch to output the fused feature map.
[0056] The deep learning-based network model design is based on the application context. Image super-resolution reconstruction involves fitting the nonlinear mapping relationship between low-resolution images and high-resolution images. A convolutional neural network is used to learn the degradation relationship between low-resolution and high-resolution images, restoring and reconstructing a high-resolution image rich in high-frequency detail. To enable practical deployment in embedded devices, the present invention proposes a dynamic adaptive hybrid super-resolution reconstruction method to fully extract feature information from images, minimize the loss of high-frequency detail, and reduce network parameters and computational overhead. This method aims to adaptively extract image features and restore details. The method uses attention region ratios for dynamic perception. For simple regions, the network uses convolution operations for efficient processing. For complex regions, it combines the LSKA module and the attention mechanism. The LSKA mechanism replaces the standard K×K convolution operation with a K×1 longitudinal one-dimensional convolution and a 1×K horizontal one-dimensional convolution. In addition, the introduced efficient predictor can generate guidance signals based on rich input information. The PFCFN module uses two fully connected layers and a nonlinear activation function module, and adds the input and output through a residual connection structure. The SCEA module effectively compensates for the lack of local information by fusing channel and spatial domain information, further improving the network's representational capabilities. Finally, the reconstruction module recovers the high-resolution reconstructed image, completing the entire super-resolution reconstruction process.
[0057] Example 2,
[0058] A hybrid adaptive perceptual super-resolution reconstruction method for complex scenes is characterized by following the steps below:
[0059] S1: low-resolution input image I LR Perform shallow feature extraction processing, the shallow feature extraction function is recorded as H SFE (·), the extraction result is the shallow feature F0, and the calculation formula is:
[0060] F0=H SFE (I LR ) (11)
[0061] Among them, I LR is the input low-quality image, F0 is the shallow feature of the extracted LR image, H SFE It is a convolution layer operation with a convolution kernel of 3×3.
[0062] S2: The deep feature extraction module consists of 24 dynamic adaptive hybrid modules (DAFMixer) connected in series. Each DAFMixer module includes a dynamic adaptive perception module (DAPMixer), a pixel fusion convolutional feedforward network (PFCFN), and an enhanced spatial channel perception module (SCEA). The modules use a residual connection structure to achieve feature preservation and enhancement. The shallow feature map F0 is used as input and is sequentially input into each DAFMixer module for processing to achieve multi-level deep feature extraction. The processing process is as follows:
[0063]
[0064] Among them, F0 represents the output feature map of the shallow feature extraction module, F n (n=1···24) represents the output of the n-th DAFMixer, The function representing the nth DAFMixer, F DAFMixer It indicates the final fusion feature after this process.
[0065] The processing in each DAFMixer module can be further expanded into the following composite structure:
[0066]
[0067] Among them, F input The feature map of the DAFMixer module representing the input, F output Describe the output feature map, Indicates that after n SCEA modules are operated, represents n PFCFN operations, Indicates n DAPMixer module operations, n is 24;
[0068] S3: The output feature F of the last DAFMixer module n Input to a 3×3 convolution layer for feature integration to obtain the fusion feature F DAFMixer , and its calculation formula is:
[0069] F DAFMixer =Conv3(F n ) (14)
[0070] Among them, Conv3(·) represents the convolution operation with a 3×3 convolution kernel, F DAFMixer is the final fusion feature.
[0071] S4: The fused deep feature F DAFMixer Concatenate (Concat) or element-wise weighted fusion with the shallow feature F0 in step S1 and input to the image reconstruction module H IRB (·), and obtain the final reconstructed image I SR , whose expression is:
[0072]
[0073] Among them, F0 is the shallow feature of the extracted LR image, F DAFMixer Indicates the final fusion feature after this process, H IRB represents the image reconstruction function, which usually includes sub-pixel convolution operations to achieve upsampling, represents the feature fusion operation, I SR Represents the final output high-resolution image.
[0074] To test the effectiveness of the model in this embodiment, a comparative experiment was conducted. Quantitative evaluation was performed using two objective metrics: Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM). The validation sets consisted of four datasets: Set5, Set 14, B100, and Urban100, containing images of different types from the training set. Reconstruction and restoration were performed on these four datasets at three different magnifications: 2x, 3x, and 4x. The super-resolution networks compared with the method of the present invention are Dong's algorithm (Dong C, Loy CC, He K, et al. Learning a deep convolutional network for image super-resolution [C] / / European conference on computer vision (CVPR). 2014: 184-199.); Kim's algorithm (Kim J, Lee JK, Lee K M. Accurate image super-resolution using very deep convolutional networks [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 1646-1654.); Ahn's algorithm (Ahn N, Kang B, Sohn K A. Fast, accurate, and lightweight super-resolution with cascading residual network [C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 252-268.); Lim's algorithm (Lim B, Son S, Kim H, et al. Enhanced deepresidual networks for single image super-resolution[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition workshops.2017:136-144.); Hui’s algorithm (Hui Z, Gao X, Yang Y, et al.Lightweight image super-resolution with information multi-distillation network[C] / / Proceedings of the 27th ACM international conference on multimedia. 2019: 2024-2032.); Zhao's algorithm (Zhao H, Kong X, He J, et al. Efficient image super-resolution using pixel attention[C] / / Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16. Springer International Publishing, 2020: 56-72.); Li's algorithm (Li W, Zhou K, Qi L, et al. Lapar: Linearly-assembled pixel-adaptive regression network for single image super-resolution and beyond[J]. Advances in Neural Information Processing Systems, 2020, 33: 20343-20355.); Luo's algorithm (Luo X, Xie Y, Zhang Y, et al. Lattice net: Towards lightweight image super-resolution with lattice block[C] / / Computer vision–ECCV 2020: 16th European conference, Glasgow, UK, August 23–28, 2020, proceedings, part XXII 16. Springer International Publishing, 2020: 272-289.); Luo's algorithm (Luo X, Qu Y, Xie Y, et al. Lattice network for lightweight image restoration[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(4):4826-4842.); Park's algorithm (Park K, Soh JW, ChoN IA dynamic residual self-attention network for lightweight single image super-resolution[J]. IEEE Transactions on Multimedia, 2021, 25:907-918.); ZHAOXQ's algorithm (ZHAOXQ, LIXY, SONGZY.Lightweight Inverse Separable Residual IntormationDistillation Network for Image Super-Resolntion Reconstruction.PatternRecognition and Artificial Intelligence, 2023, 36(5):419-432); Gao's algorithm (Gao G, Wang Z, Li J, et al.Lightweight bimodal network for single-image super-resolution via symmetric CNN and recursive transformer[J].arXiv preprintarXiv:2204.13286, 2022.); Liu's algorithm (Liu Y, Jia Q, Zhang J, et al. Hierarchical similarity learning for aliasing suppression image super-resolution[J]. IEEE Transactions on Neural Networks and Learning Systems, 2024, 35(2):2759-2771.) The PSNR / SSIM results of different methods on different data sets are compared in Table 1.
[0075] Table 1 Comparison of PSNR / SSIM results of different methods on different datasets at different magnifications
[0076]
[0077]
[0078] The results in Table 1 show that the embodiments of the present invention are moderate in terms of parameter and computational complexity, and perform well across different magnification factors and datasets. At a magnification factor of 2, the image textures in datasets such as Set5 and Set14 are relatively simple, dominated by low-frequency information, and the model relies primarily on its ability to integrate global information. DAFMixerSR utilizes a dynamic adaptive hybrid module, employing convolution in simple regions to reduce computational cost. In complex regions, it combines a self-attention mechanism with multi-scale separable large kernel convolutions, achieving a balanced recovery of global information and local details. It achieves PSNRs of 38.15 and 33.78 on Set5 and Set14, respectively, outperforming lightweight methods such as IMDN and PAN at low parameter counts. However, at magnification factors of 3 and 4, image detail loss is exacerbated, particularly on the BSD100 and Urban100 datasets, which have complex high-frequency textures. The Urban100 dataset achieves a PSNR of 32.62, surpassing HSRNet's 32.53 and LIRIDN's 32.43, demonstrating superior detail recovery capabilities.
[0079] In summary, the method of the present invention achieves a good balance between image reconstruction effect and parameter quantity, and can meet the application requirements of mobile devices in terms of model lightweighting.
[0080] like Figure 5 Figure 2 shows a subjective visual comparison of the single-image hybrid adaptive perceptual super-resolution reconstruction method provided by the present invention with other algorithms for reconstructions of images "img092" and "img008" from Urban100 at a magnification of ×4; image "barbara" from Set14 at a magnification of ×2; and image "img86000" from B100 at a magnification of ×3. The method of the present invention excels in recovering high-frequency stripes in complex architectural scenes. While other methods suffer from blurred stripes and directional distortion during reconstruction, the present invention accurately restores the stripe structure, maintaining clear lines and correct orientation. Furthermore, observations of images "barbara" and "img86000" show that the present invention more accurately restores the line structure than other methods, avoiding geometric distortion and blurring artifacts. These reconstructions of images with other methods exhibit chaotic, distorted textures and overall blur. Therefore, the super-resolution images obtained by the method of the present invention are closer to true high-resolution images.
[0081] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0082] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A hybrid adaptive perceptual super-resolution reconstruction method for complex scenes, characterized by: The following steps are involved: S1: Build dataset images; S2: Build a super-resolution reconstruction model; The S2 includes: S21: Construct shallow feature extraction module SFE; S22: Construct the deep feature extraction module DAFMixer module; S23: constructing an image reconstruction module IRB; S3: training super-resolution reconstruction model; S4: Test the trained super-resolution reconstruction model.
2. The hybrid adaptive perceptual super-resolution reconstruction method for complex scenes according to claim 1, characterized in that: Said S1 comprises: The dataset was selected. The training dataset was augmented by random rotations of 90°, 180°, and 270° and horizontal flipping. Low-resolution and high-resolution image training pairs were constructed by bicubic downsampling (sampling multiples were '2, '3, and '4 respectively). The test dataset used 5 common benchmark datasets.
3. The hybrid adaptive perceptual super-resolution reconstruction method for complex scenes according to claim 1, characterized in that: In said S21: The shallow feature extraction module is used to map the input low-resolution image to a high-dimensional feature space through a 3×3 standard convolution. The process of extracting shallow feature information is as follows: F0=H SFE (I LR ) (1) Among them, I LR is the input low-quality image, F0 is the shallow feature of the extracted LR image, H SFE It is a convolution layer operation with a convolution kernel of 3×3.
4. The hybrid adaptive perceptual super-resolution reconstruction method for complex scenes according to claim 1, characterized in that: In said S22: The S22 consists of 24 dynamic adaptive hybrid modules (DAFMixer) connected in series. Each DAFMixer module consists of a dynamic adaptive perception module (DAPMixer), a pixel fusion convolutional feedforward network (PFCFN), and an enhanced spatial channel perception module (SCEA). The modules use a residual connection structure to achieve feature preservation and enhancement. The shallow feature map F0 is used as input and is sequentially input into each DAFMixer module for processing to achieve multi-level deep feature extraction. The processing process is as follows: Among them, F0 represents the output feature map of the shallow feature extraction module, F n (n=1···24) represents the output of the n-th DAFMixer, The function representing the nth DAFMixer, F DAFMixer Indicates the final fusion features after this process; The processing in each DAFMixer module can be further expanded into the following composite structure: Among them, F input The feature map of the DAFMixer module representing the input, F output Describe the output feature map, Indicates that after n SCEA modules are operated, represents n PFCFN operations, Indicates n DAPMixer module operations, n is 24; The output feature F of the last DAFMixer module n Input to a 3×3 convolution layer for feature integration to obtain the fusion feature F DAFMixer , and its calculation formula is: F DAFMixer =Conv3(F n ) (4) Among them, Conv3(·) represents the convolution operation with a 3×3 convolution kernel, F DAFMixer is the final fusion feature.
5. The hybrid adaptive perceptual super-resolution reconstruction method for complex scenes according to claim 1, characterized in that: In said S23: The S23 includes a reconstruction module to transform the fused deep features F DAFMixer Perform element-wise weighted fusion with the shallow feature F0 in step S1 and input it to the image reconstruction module H IRB (·), and obtain the final reconstructed image I SR , whose expression is: Among them, F0 is the shallow feature of the extracted LR image, F DAFMixer Indicates the final fusion feature after this process, H IRB represents the image reconstruction function, which usually includes sub-pixel convolution operations to achieve upsampling, represents the feature fusion operation, I SR Represents the final output high-resolution image.
6. The DAFMixer module according to claim 4, comprising: The DAPMixer first obtains the value V through 1×1 point-wise convolution (PWConv), and generates the offset map offset, mixer mask mask and preliminary spatial / channel attention factors ca / pa through the global predictor; wherein, the global predictor is used to fuse local and global context information to predict the offset information and mask information for feature sampling adjustment and screening; the DAPMixer module includes an attention branch, which uses bilinear interpolation to sample and align the input feature map under the guidance of the offset to construct sparse window features, and uses Gumbel in the training stage. The Softmax technique implements binary sampling of the attention region, and a Top-M screening strategy is used to determine the key position index during the inference phase. The selected sparse window features are then linearly mapped to generate three sets of vectors: query Q, key K, and value V. The attention weights are calculated using a scaled dot-product attention mechanism to obtain a weighted fused sparse attention feature map. The DAPMixer module also includes a convolution branch, which is sequentially composed of a 1×(2d–1) vertical depth-wise separable convolution, a (2d–1)×1 horizontal depth-wise separable convolution, a 1×(d / 2) vertical depth-wise separable convolution, and a (d / 2)×1 horizontal depth-wise separable convolution. The PFCFN module includes a LayerNorm layer, a pixel attention PA module, a first fully connected layer FC1, a nonlinear activation function module Act and a second fully connected layer FC2, which are connected in sequence, and a residual connection path is provided for element-by-element addition of the module input and the final output; The SCEA module first passes through a 1×1 standard convolution operation and a strided convolution layer, and then is divided into two parallel branches, consisting of a channel attention branch and a trunk spatial feature extraction branch; among them, the channel attention branch is used to perform Avgpooling on the input feature map to extract channel statistical information, and then model the local dependency between channels through 1DConv, and restore the obtained channel attention feature map to the original spatial size through Upsampling; the trunk spatial feature extraction branch passes through a Maxpooling in sequence, and then passes it to a 3×3 DWConv, and then restores it to the spatial size through Upsampling; the output feature maps of the two branches are spliced and fused in the channel dimension through the Concat module, and the channel dimension of the fused feature map is compressed through a 1×1 convolution layer, and then the fused attention map is generated through the Sigmoid activation function.
7. An electronic device, characterized in that: The method according to any one of claims 2 to 6 is used to realize super-resolution reconstruction of a single image.
8. A computer storage medium, characterized in that The storage medium stores at least one program instruction, and the at least one program instruction is loaded and executed by a processor to implement super-resolution reconstruction of a single image as claimed in any one of claims 2 to 6.
Citation Information
Cited By
Road surface defect detection method based on improved YOLOv8 model
CN121527082A