Image super-resolution reconstruction method based on channel perception aggregation Transform
By using a channel-aware aggregation Transformer approach, and leveraging dual-aggregation interactive attention and hierarchical slice attention mechanisms, the high computational complexity and weak local detail modeling capabilities of existing Transformer models are addressed, achieving efficient and high-quality image super-resolution reconstruction.
Patent Information
- Application Number
- CN202511499896.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing Transformer models have high computational complexity and resource consumption in image super-resolution reconstruction tasks, making them difficult to deploy on resource-constrained devices. Furthermore, they have weak local detail modeling capabilities and struggle to recover high-frequency texture information.
We employ a channel-aware aggregation Transformer-based approach, combining convolution operations with subpixel reconstruction through dual-aggregation interactive attention and hierarchical slice attention mechanisms, to achieve high-resolution image restoration.
It effectively captures global dependencies and local details of images, reduces computational complexity, improves reconstruction quality and efficiency, and enhances the model's adaptability and stability to different types of images.
Smart Images

Figure CN120976024A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image super-resolution, and particularly relates to an image super-resolution reconstruction method based on a channel perception aggregation Transformer. BACKGROUND
[0002] The purpose of image super-resolution technology is to recover a high-quality high-resolution image from a given low-resolution image. Early methods are mostly based on interpolation algorithms or reconstruction strategies of sparse representation, but these methods often have difficulty in recovering the details and textures in the image.
[0003] In recent years, convolutional neural networks and Transformers have been widely used in image reconstruction tasks, and significant progress has been made in reconstruction quality. However, the existing Transformer model has a computational complexity that grows quadratically with the input resolution, consumes a lot of resources, and is difficult to deploy on resource-constrained devices.
[0004] Before the widespread application of deep learning, image super-resolution mainly relied on interpolation methods and prior-based image modeling methods. For example, classic algorithms such as bilinear interpolation and bicubic interpolation achieve pixel interpolation through simple mathematical rules, which are fast but have limited recovery accuracy, often causing image blurring. Sparse representation and dictionary learning methods learn the mapping relationship between low-resolution pictures (LR) and high-resolution images (HR) to recover, but the overall effect still has deficiencies in detail restoration, especially in recovering high-frequency texture information.
[0005] With the development of deep learning technology, convolutional neural networks have been widely used in super-resolution tasks, significantly improving image recovery quality. Early representative methods such as SRCNN (Super-Resolution Convolutional Neural Network) first used a three-layer CNN to learn the mapping relationship from LR images to HR images. Subsequently, VDSR (Very Deep Super-Resolution) introduced a residual learning structure to further deepen the network layers to enhance modeling capabilities. RCAN (Residual Channel Attention Network) introduced a channel attention mechanism, effectively improving the network's ability to perceive key information. Although CNN methods have made significant achievements in SR tasks, their receptive field is limited by the size of the convolution kernel, making it difficult to capture global dependencies between distant pixels.
[0006] In recent years, inspired by the natural language processing Transformer architecture, visual Transformers have been widely applied to image tasks. Its multi-head self-attention mechanism has strong global modeling capability and can effectively capture long-range image dependency relationships, which is suitable for texture complex and high-frequency detail rich SR tasks, and has achieved leading performance on multiple benchmark datasets. However, due to the quadratic growth of the Transformer computation complexity with the input image size, the model inference speed is slow, the resource consumption is large, and the local detail modeling capability is weak, which may lose the local texture of the image. SUMMARY
[0007] In order to overcome the problems in the prior art, the present application provides an image super-resolution reconstruction method based on a channel perception aggregation Transformer.
[0008] The technical solution of the present application to solve the above technical problems is as follows: The present application provides an image super-resolution reconstruction method based on a channel perception aggregation Transformer, comprising: Step 100: shallow feature extraction is performed on the low-resolution image to obtain a shallow feature map; Step 200: input the shallow feature map into a deep feature extraction module for deep feature extraction to obtain a deep feature map, the deep feature extraction module comprising a plurality of residual groups, each residual group comprising at least one double-aggregation interactive attention and a hierarchical slice attention; Step 300: combine the convolution operation and sub-pixel reorganization to realize the preliminary recovery of the high-resolution image; and superimpose and fuse the global residual information extracted from the preliminary recovered high-resolution image and the original low-resolution image after global upsampling to generate a high-resolution reconstructed image.
[0009] Further, the step 100 specifically comprises: using a convolution operation to map the original low-resolution image from the original pixel space to a high-dimensional feature space, thereby obtaining a shallow feature map.
[0010] Further, the step 200 specifically comprises: The shallow feature map passes through a convolution layer to extract initial features of the image and perform double-aggregation interaction on the initial features to output double-aggregation interactive image features; the double-aggregation interactive image features are subjected to hierarchical slice self-attention processing to obtain hierarchical slice output features. The hierarchical slice output features are connected in residual to the initial features to obtain output features of the residual group, i.e. a deep feature map.
[0011] Further, the double-aggregation interaction on the initial features to output double-aggregation interactive image features comprises: The initial features are subjected to layer normalization operation, and via double-flow feature interaction, double-flow feature interaction output features are obtained; the double-flow feature interaction output features are divided into a first feature flow and a second feature flow through a channel segmentation operation; The first feature flow is subjected to multi-scale spatial compression attention processing, and the second feature flow is subjected to depth separable convolution processing; the processed first feature flow and the processed second feature flow are fused to output double-aggregated interactive image features.
[0012] Further, via double-flow feature interaction, double-flow feature interaction output features are obtained, including: the features after the layer normalization operation are subjected to double-flow feature interaction, the double-flow feature interaction adopts a parallel structure, combines a convolution branch and a linear branch, jointly models spatial structure information and channel dependency of the features, and element-wise multiplies output features of the two branches to obtain double-flow feature interaction output features.
[0013] Further, the first feature flow is subjected to multi-scale spatial compression attention processing, and the second feature flow is subjected to depth separable convolution processing; the processed first feature flow and the processed second feature flow are fused to output double-aggregated interactive image features, including: The first feature flow is sent into a multi-scale spatial compression attention module for global dependency modeling to obtain the processed first feature flow; at the same time, the second feature flow is sent into a depth separable convolution module for local context modeling to obtain the processed second feature flow; The processed first feature flow, the processed second feature flow, and the initial features are element-wise added to obtain first fusion features; the first fusion features are sent into a feedforward network for nonlinear enhancement, and the enhanced first fusion features are residual connected with the first fusion features to obtain double-aggregated interactive image features.
[0014] Further, the first feature flow is sent into a multi-scale spatial compression attention module for global dependency modeling, including: A linear transformation is applied to the first feature flow to generate a query vector, and the feature dimension thereof is compressed; A depth separable convolution is performed on the first input feature flow to obtain a spatially compressed feature map; a one-time depth separable convolution and a one-time 1x1 convolution are applied to the spatially compressed feature map in sequence to further transform the features, thereby obtaining a feature-transformed feature map; different linear transformations are then applied to the feature-transformed feature map to respectively generate a key vector and a value vector; An attention score is calculated by dot product of the key vector and the query vector, multiplied by a scaling factor, and then normalized by a Softmax function to calculate a weighted sum, thereby obtaining an output of the multi-scale spatial compression attention module.
[0015] Further, the double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including: The double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including: The double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including:
[0016] Further, the double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including: The double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including: The double-aggregated interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including:
[0017] Further, in the step 300, the deep feature map is subjected to convolution operation and sub-pixel reorganization to realize preliminary recovery of the high-resolution image, including: The deep feature map is subjected to convolution operation and sub-pixel reorganization to sample it to a high-resolution space to obtain a preliminary recovered high-resolution image.
[0018] Compared with the prior art, the present application has the following technical effects: (1) The present application proposes double-aggregated interactive attention and hierarchical slicing attention, which can effectively capture the global dependency relationship and local detail information of the image. The multi-scale spatial compression attention processing and depth separable convolution processing in the double-aggregated interactive attention model the features from different angles, enhancing the network's perception ability of key information of the image. The hierarchical slicing attention reduces the computational complexity through slicing processing while improving the local representation ability. These mechanisms work together to make the model more accurately recover the details and textures in the image, significantly improving the quality of the reconstructed image.
[0019] (2) The layered slicing attention mechanism of the present application divides the feature map into local regions and independently performs attention calculation within the regions, greatly reducing the computational load and reducing the computational complexity. This enables the model to operate more efficiently while maintaining good reconstruction performance, reducing the dependence on computing resources.
[0020] (3) The method of the present application adopts multiple residual group structures in the deep feature extraction module, and fuses the features of different stages through residual connection and other means. This design not only helps the effective propagation of gradients, avoids the problem of gradient vanishing or explosion, makes the model easier to train, but also enhances the adaptability of the model to different types of images. The model can stably learn effective feature representation, thereby generating high-quality high-resolution reconstructed images, improving the reliability and stability of the model in practical applications. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the description of the embodiments or the prior art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0022] Figure 1 The flowchart of the image super-resolution reconstruction method based on channel perception aggregation Transformer of the present application; Figure 2 The structure diagram of the image super-resolution reconstruction method based on channel perception aggregation Transformer of the present application; Figure 3 The structure diagram of the double-flow feature interaction module of the present application; Figure 4 The structure diagram of the multi-scale spatial compression attention of the present application; Figure 5 The selected first original image; Figure 6 In (a) of the figure, the low-resolution image corresponding to the detail area intercepted in the figure is shown. Figure 5 In (b) of the figure, the original high-resolution image corresponding to the detail area intercepted in the figure is shown. Figure 5 In (c) of the figure, the super-resolution image obtained after the low-resolution image is subjected to image super-resolution reconstruction is shown. Figure 7 The selected second original image; Figure 8 In (a) of the figure, the low-resolution image corresponding to the detail area intercepted in the figure is shown. Figure 7 In (b) of the figure, the original high-resolution image corresponding to the detail area intercepted in the figure is shown. Figure 7The original high-resolution image corresponding to the detail area intercepted in (a), (c) is the super-resolution image obtained after the low-resolution image is reconstructed by image super-resolution; Figure 9 The third original image selected; Figure 10 The original high-resolution image corresponding to the detail area intercepted in (a), Figure 9 The low-resolution image corresponding to the detail area intercepted in (a), (b) is Figure 9 The original high-resolution image corresponding to the detail area intercepted in (a), (c) is the super-resolution image obtained after the low-resolution image is reconstructed by image super-resolution; Figure 11 The slice self-attention structure diagram of the present application; Figure 12 The hierarchical window design diagram in the slice self-attention. DETAILED DESCRIPTION
[0023] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined object, the specific implementation, structure, features and effects of the technical solutions proposed according to the present application are described in detail below in combination with the drawings and preferred embodiments. The specific features, structures or characteristics in one or more embodiments can be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as understood by those skilled in the art of the technology to which the present application belongs.
[0024] In the present embodiment, reference is made to Figures 1-12 The present application provides an image super-resolution reconstruction method based on channel perception aggregation Transformer, comprising: Step 100: shallow feature extraction is performed on the low-resolution image to obtain a shallow feature map; Step 200: the shallow feature map is input into a deep feature extraction module for deep feature extraction to obtain a deep feature map, the deep feature extraction module comprising a plurality of residual groups, each residual group comprising at least one double-aggregation interactive attention and hierarchical slice attention; Step 300: the deep feature map is combined with sub-pixel reorganization through convolution operation to realize preliminary recovery of the high-resolution image; the preliminary recovered high-resolution image is superimposed and fused with global residual information extracted after global up-sampling of the original low-resolution image to generate a high-resolution reconstructed image.
[0025] The above steps are described in detail as follows: Step 100: shallow feature extraction is performed on the low-resolution image to obtain a shallow feature map.
[0026] The public data set DIV2K is selected as the training data set. The input low-resolution image is cropped to a size of 128x128 , the high-resolution image is cropped to , where r is a scaling factor.
[0027] For the input low-resolution image , a convolution operation is used to map the original low-resolution image from the original pixel space to a high-dimensional feature space, thereby obtaining a shallow feature map with C channels , where, represents the spatial dimension, and C represents the number of channels: ; In the above formula, is a 3x3 convolution operation.
[0028] Step 200: input the shallow feature map into a deep feature extraction module for deep feature extraction to obtain a deep feature map; wherein the deep feature extraction module includes multiple residual groups, and each residual group includes at least one double-aggregated interactive attention and hierarchical slice attention.
[0029] The shallow feature map is input into a deep feature extraction module composed of N residual groups (RG) in cascade for processing, and a convolution is used after the Nth residual group to refine the features and improve the reconstruction quality.
[0030] The i th residual group processes the output of the first i- 1 residual group, then passes through a 3x3 convolution to further adjust the features, and finally adds the processed result to the output of the first i- 1 residual group , and the output after the N residual groups, i.e., the deep feature map. F
[0031] ; In the above formula, represents the i th residual group, represents the output of the i th residual group; represents the output of the first i- 1 residual group. Through residual connection, the gradient can be effectively backpropagated, avoiding the problem of gradient vanishing or explosion, making the network easier to train, and better preserving and utilizing the feature information of the previous stage.
[0032] In this implementation, the residual set is mainly divided into three parts: Dual Convergence Interaction Transformer (DAIT), Hierarchical Slicing Transformer (HST), and convolutional layers. The convolutional layers are used to optimize and enhance the feature representation within the current module.
[0033] Step 210: Shallow Feature Map After passing through the first convolutional layer (Conv), the initial features of the image are extracted.
[0034] The first convolutional layer, Conv, extracts the initial features of the image by capturing local feature information through the sliding operation of the convolutional kernel across the image. Specifically, the first convolutional layer, Conv, uses a single 3x3 convolution.
[0035] Step 220: Perform dual-aggregation interaction on the initial features and output dual-aggregation interaction image features.
[0036] As an example, step 220 specifically includes the following sub-steps: Step 2201: Perform LayerNorm operation on the initial features.
[0037] The LayerNorm operation normalizes the initial feature maps output by convolutional layers, which helps to accelerate the training convergence speed of the model and improve the stability of the model.
[0038] Step 2202: After performing layer normalization on the initial features, the dual-stream feature interaction (DSFI) output features are obtained through dual-stream feature interaction.
[0039] The features after LayerNorm operation are subjected to dual-stream feature interaction (DSFI). The dual-stream feature interaction (DSFI) adopts a parallel structure, combining convolutional branches and linear branches to jointly model the spatial structure information and channel dependencies of the features. The output features of the two branches are multiplied element-wise to obtain the output features of dual-stream feature interaction (DSFI).
[0040] Among them, reference Figure 3 Convolutional branches Employing a Bottleneck-like design, it includes sequentially concatenated 1×1 convolutions (Conv), ReLU activation, 3×3 convolutions (Conv), ReLU activation, and then another 1×1 convolution (Conv), demonstrating strong local feature extraction capabilities; linear branches Then the input features are directly mapped to the target dimension by a single-layer linear transformation Liner to capture the global dependency of channel hierarchy. The outputs of the two branches are fused by element-wise multiplication, which introduces a channel-selective feature enhancement mechanism while maintaining the spatial structure perception ability. Specifically, let the input features be X after the layer normalization operation, and the fused output is ; wherein, represents the double-stream feature interaction DSFI output feature, represents the element-wise multiplication operation. This fusion strategy effectively integrates multi-scale information at the spatial and channel levels, providing rich semantic support for subsequent high-quality reconstruction.
[0041] Step 2203: The double-stream feature interaction DSFI output feature is divided into a first feature stream and a second feature stream by a channel segmentation operation.
[0042] The double-stream feature interaction DSFI output feature is expanded by a dilated convolution (i.e., a 1x1 convolution) to expand the channel number to twice the original dimension, and then the expanded double-stream feature interaction DSFI output feature is uniformly divided into a first feature stream split and a second feature stream by a channel dimension operation, i.e., ; In the above formula, Split represents the channel segmentation operation.
[0043] Step 2204: The first feature stream is processed by multi-scale spatial compression attention, the second feature stream is processed by depth separable convolution, and the processed first feature stream and the processed second feature stream are fused to output a double-aggregated interactive image feature.
[0044] The first feature stream is input into a multi-scale spatial compression attention (MSCA) module for global dependency modeling to obtain the processed first feature stream; at the same time, the second feature stream is input into a depth separable convolution (DWConv) module for local context modeling to obtain the processed second feature stream.
[0045] Wherein, the multi-scale spatial compression attention module effectively reduces the computational cost of the traditional self-attention mechanism at high-resolution input by introducing an iterative depth convolution downsampling and channel dimension compression mechanism, and its architecture is shown in Figure 4 For the first feature stream , wherein B represents the batch size, where C is the number of channels. The query vector Q is generated by applying a channel compression operation to the first feature stream and then applying a linear transformation to generate the query vector Q and compress its feature dimension to 1.
[0046] The key vector and the value vector are generated by performing a depth separable convolution on the first feature stream, sequentially performing spatial downsampling by four times of depth separable convolution with a stride of 2, each operation reducing the spatial resolution of the feature map by half, thereby reducing the original spatial dimension to a smaller spatial dimension to obtain a spatially compressed feature map; sequentially applying a depth separable convolution and a 1x1 convolution to the spatially compressed feature map to further transform the features to obtain a feature-transformed feature map; and then applying a linear transformation to the feature-transformed feature map to generate the key vector K and the value vector V, wherein the feature dimension of K is compressed to 1 and the feature dimension of V remains unchanged.
[0047] In the attention calculation process, when constructing the key vector and the query vector, the channel dimension is reduced from C to 1 through the channel compression operation, effectively reducing the parameter quantity and computational overhead in the dot product operation. The dot product of the channel-compressed query vector and the transpose of the key vector is performed, and a scaling operation (Scale) is performed by dividing by a scaling factor (the square root of the dimension of the key vector) to stabilize the numerical distribution and prevent the Softmax gradient from being too small, followed by a weighted operation (Weighted), and then normalized by the Softmax function to obtain the weight coefficient. Finally, the weight coefficient and the value vector V are multiplied by a matrix to obtain the attention output A: ; In the above formula, represents the dimension of the key vector.
[0048] Finally, the attention output is projected on the original spatial dimension to ensure effective utilization and preservation of global context information. The multi-scale spatial compression attention module further reduces the computational burden by compressing the channel dimension, achieving the purpose of significantly improving the computational efficiency while maintaining the ability to model global dependencies.
[0049] The processed first feature stream (output of the multi-scale spatial compression attention module), the processed second feature stream (output of the depth separable convolution module), and the double-aggregated interactive image feature X0 (initial feature) are added element by element to obtain a first fusion feature; the fusion feature is input into a feedforward network (FFN) for nonlinear enhancement, and is connected in residual connection with the first fusion feature to obtain a double-aggregated interactive image feature: ; wherein, represents a first fusion feature, represents a deep convolution, is a double-aggregated interactive image feature.
[0050] Step 230: performing hierarchical slice self-attention processing on the double-aggregated interactive image feature to obtain a hierarchical slice output feature.
[0051] As an example, the step 230 specifically includes the following sub-steps: Step 2301: performing a LayerNorm operation on the double-aggregated interactive image feature again.
[0052] Step 2302: performing weighting processing on the LayerNorm processed double-aggregated interactive image feature by a slice attention mechanism to obtain a feature processed by the slice attention mechanism.
[0053] By introducing a slice self-attention module (HSA), the calculation complexity is reduced and the local representation ability is improved, and the architecture thereof is shown in Figure 11 The method comprises: performing an image division operation (Image divide) on the LayerNorm processed double-aggregated interactive image feature According to a set region size, the region size P×P is divided into Figure 12 local regions, and the total number of local regions is In each local region (the size of the local region is determined by the network depth), the calculation of the attention mechanism is independently performed, including mapping the features in the region to query (Q), key (K) and value (V) respectively; performing the attention mechanism based on the Q, K and V inside each local region to obtain a local attention result; restoring and splicing all local region attention results according to the spatial position to obtain a complete feature map, i.e., a feature processed by the slice attention mechanism.
[0054] Step 2303: performing residual connection on the feature processed by the slice attention mechanism and the image feature obtained by the double aggregation to obtain a second fusion feature; sending the second fusion feature into a feedforward network (FFN) to perform nonlinear enhancement to obtain an enhanced second fusion feature; performing residual connection on the enhanced second fusion feature and the second fusion feature, and passing through two 3×3 convolution layers to obtain a layer slice output feature.
[0055] Step 240: performing residual connection on the hierarchical slice output feature and the initial feature to obtain an output feature map of a residual group, i.e., a deep feature map.
[0056] Step 300: The deep feature map is combined with sub-pixel reconstruction through convolution operation to realize the preliminary recovery of the high-resolution image; the preliminary recovered high-resolution image is superimposed and fused with the global residual information extracted after the global up-sampling of the original low-resolution image to generate a high-resolution reconstructed image.
[0057] Output features of residual groups F 3x3 convolution is performed, and then sub-pixel reconstruction is performed Pixel The operation is up-sampled to a high-resolution space to obtain a preliminary recovered high-resolution image. At the same time, the original low-resolution image is up-sampled Up ; finally, the preliminary recovered high-resolution image is superimposed and fused with the global residual information extracted after the global up-sampling of the original low-resolution image, so as to fully utilize the information of the original low-resolution image and the high-frequency detail information learned from the deep features, thereby generating a high-quality high-resolution reconstructed image : ; wherein, indicates the reconstructed high-resolution image, Pixel indicates sub-pixel reconstruction, Up indicates up-sampling operation, which can be bilinear interpolation operation; F indicates the output of N residual groups.
[0058] The application also includes: step 400: using an L1 loss function to supervise the learning and training of the reconstructed image and the real image, and using an Adam optimizer to iteratively update.
[0059] ; wherein, n is the total number of pixels, is the predicted value, is the true value.
[0060] The number of training times is set to 500000 iter, the training is performed on a standard SISR data set (such as DIV2K), and the evaluation is performed on benchmark data sets such as Set5, Set14, BSD100, Urban100 and Manga109. The model is optimized using an Adam optimizer, the learning rate is set to , the loss function is L1 loss, and the pixel difference between the reconstructed image and the real HR image is minimized. In the test stage, a self-integrated inference strategy is adopted, the input image is enhanced, and the output is averaged to further improve the reconstruction quality.
[0061] During the testing phase, this invention evaluates super-resolution images using two metrics based on the Y channel of the YCbCr color space. Specifically, PSNR (Peak Signal-to-Noise Ratio) is used to measure image reconstruction error, and SSIM (Structural Similarity Score) is used to measure image structural fidelity.
[0062] The experimental results of image reconstruction using the proposed method on the Set14 and Urban100 benchmark datasets show a reconstruction scaling factor of 4. Figure 5 The original image is presented, and to explore the detailed characteristics of the image in depth, from... Figure 5 The image shows a portion of representative details (highlighted in red). Figure 6 (a) shows the low-resolution image corresponding to the cropped detail area, where the details are blurred; Figure 6 (b) shows the original high-resolution image corresponding to the cropped detail area. Figure 6 Figure (c) shows the image result obtained after reconstructing a low-resolution image using the method of the present invention. Figure 7 Another original image was presented, from Figure 7 The image precisely captures some representative details (red box). Figure 8 (a) shows the low-resolution image corresponding to the cropped detail area. Figure 8 (b) shows the original high-resolution image corresponding to the cropped detail area. Figure 8 Figure (c) shows the image result obtained after reconstructing a low-resolution image using the method of the present invention. Figure 9 Another original image was presented, from Figure 9 The image shows a portion of representative details (highlighted in red). Figure 10 (a) shows the low-resolution image corresponding to the cropped detail area. Figure 10 (b) shows the original high-resolution image corresponding to the cropped detail area. Figure 10 Figure (c) shows the image result obtained after reconstructing a low-resolution image using the method of the present invention.
[0063] The image comparison results clearly demonstrate that the algorithm proposed in this invention has significant advantages in image edge sharpness, texture restoration capability, and preservation of structural details. The reconstructed image visually closely resembles the original high-resolution image, accurately restoring complex texture regions and fine-grained structures, exhibiting excellent image perceptual quality. These results fully demonstrate that the super-resolution reconstruction method proposed in this invention can significantly improve the subjective visual quality of images, and it has broad application prospects and practical value in areas such as image magnification and enhancement.
[0064] The above examples are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing examples, those ordinarily skilled in the art should understand: the technical solutions recorded in the foregoing examples can still be modified, or some technical features can be replaced equivalently; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An image super-resolution reconstruction method based on channel-aware aggregation Transformer, characterized in that, Includes the following steps: Step 100: Perform shallow feature extraction on the original low-resolution image to obtain a shallow feature map; Step 200: Input the shallow feature map into the deep feature extraction module for deep feature extraction to obtain a deep feature map; the deep feature extraction module includes multiple residual groups, and each residual group includes at least one dual-convergence interactive attention and a hierarchical slice attention; Step 300: The deep feature map is combined with convolution operation and subpixel reconstruction to achieve preliminary restoration of the high-resolution image; the preliminary restored high-resolution image is superimposed and fused with the global residual information extracted after global upsampling of the original low-resolution image to generate a high-resolution reconstructed image.
2. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 1, characterized in that, Step 100 specifically includes: using convolution operations to map the original low-resolution image from the original pixel space to a high-dimensional feature space, thereby obtaining a shallow feature map.
3. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 1, characterized in that, Step 200 specifically includes: The shallow feature map is passed through a convolutional layer to extract the initial features of the image, and the initial features are subjected to dual-aggregation interaction to output dual-aggregation interaction image features; the dual-aggregation interaction image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features. The output features of the layered slices are residually concatenated with the initial features to obtain the output features of the residual group, which is the deep feature map.
4. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 3, characterized in that, Perform a dual-convergence interaction on the initial features to output dual-convergence interaction image features, including: The initial features are normalized by layer, and the dual-stream feature interaction output features are obtained through dual-stream feature interaction. The dual-stream feature interaction output features are divided into a first feature stream and a second feature stream by channel segmentation. The first feature stream is subjected to multi-scale spatial compression attention processing, and the second feature stream is subjected to depthwise separable convolution processing. The processed first feature stream and the processed second feature stream are fused to output dual-aggregated interactive image features.
5. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 4, characterized in that, The dual-stream feature interaction output features are obtained through dual-stream feature interaction, including: performing dual-stream feature interaction on the features after layer normalization operation. The dual-stream feature interaction adopts a parallel structure, combining convolutional branches and linear branches, jointly modeling the spatial structure information and channel dependency of the features, and multiplying the output features of the two branches element by element to obtain the dual-stream feature interaction output features.
6. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 5, characterized in that, The first feature stream undergoes multi-scale spatial compression attention processing, and the second feature stream undergoes depthwise separable convolution processing. The processed first and second feature streams are then fused to output dual-aggregated interactive image features, including: The first feature stream is fed into a multi-scale spatial compression attention module for global dependency modeling to obtain the processed first feature stream; at the same time, the second feature stream is fed into a depthwise separable convolution module for local context modeling to obtain the processed second feature stream. The processed first feature stream, the processed second feature stream, and the initial features are added element-wise to obtain the first fused feature; the first fused feature is fed into a feedforward network for nonlinear enhancement, and the enhanced first fused feature is residually connected with the first fused feature to obtain the dual-aggregated interactive image feature.
7. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 6, characterized in that, The first feature stream is fed into a multi-scale spatial compression attention module for global dependency modeling, including: Apply a linear transformation to the first feature stream to generate a query vector and compress its feature dimensions; A depthwise separable convolution is performed on the first input feature stream to obtain a spatially compressed feature map. A depthwise separable convolution and a 1×1 convolution are then applied to the spatially compressed feature map to perform further feature transformation, resulting in a transformed feature map. Different linear transformations are then applied to the transformed feature map to generate key vectors and value vectors, respectively. The attention score is calculated by performing a dot product between the key vector and the query vector, multiplied by a scaling factor, and then normalized using the Softmax function to calculate a weighted sum, resulting in the output of the multi-scale spatially compressed attention module.
8. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 3, characterized in that, The dual-aggregate interactive image features are subjected to hierarchical slicing self-attention processing to obtain hierarchical slicing output features, including: The dual-aggregate interactive image features are then subjected to layer normalization again; the layer-normalized dual-aggregate interactive image features are then weighted using a piecewise attention mechanism to obtain the features processed by the piecewise attention mechanism. The features processed by the slice attention mechanism are residually connected with the image features obtained by the double aggregation interaction to obtain the second fused feature. The second fused feature is then fed into a feedforward network for nonlinear enhancement. The enhanced second fused feature is residually connected with the second fused feature and passed through a convolutional layer to obtain the slice output feature.
9. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 8, characterized in that, The dual-aggregate interactive image features after layer normalization are weighted using a piecewise attention mechanism to obtain the features processed by the piecewise attention mechanism, including: The double-aggregated interactive image features after layer normalization are divided into several local regions. In each local region, the attention mechanism is calculated independently. The attention mechanism calculation includes mapping the features in the local region into query vector, key vector and value vector respectively. An attention mechanism is executed based on the query vector, key vector, and value vector within each local region to obtain local attention results. All local region attention results are then restored according to their spatial location and stitched together to form a complete feature map, which is the feature map processed by the piecewise attention mechanism.
10. The image super-resolution reconstruction method based on channel-aware aggregation Transformer according to claim 1, characterized in that, In step 300, the preliminary restoration of the high-resolution image by combining convolution operations and sub-pixel reconstruction of the deep feature map includes: The deep feature map is convolved, and then upsampled to a high-resolution space through subpixel recombination to obtain a preliminary high-resolution image.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on fused attention mechanism residual network
CN111192200A
Image super-resolution reconstruction model and method based on residual mixed attention network
CN115222601A
Remote sensing image super-resolution reconstruction method based on MT-SRGAN
CN117196950A
CAD (Computer Aided Design) image super-resolution enhancement method and system based on dual-polymerization Transformer
CN120339074A
Channel attention-based swin-transformer image denoising method and system
US20240193723A1
Cited By
Stable feature enhancement method, device and equipment for super-resolution of single image
CN121981892A
Stable feature enhancement methods, apparatus and devices for single image super-resolution
CN121981892B
Image super-resolution reconstruction method
CN122115217A