Image super-resolution reconstruction method based on hierarchical directional attention fusion

The image super-resolution reconstruction method using hierarchical directional attention fusion, by utilizing bifurcation attention residual blocks and nonlocal coordinate attention, reduces computational complexity and improves image reconstruction quality, solving the problem of high computational resource efficiency in existing technologies and achieving efficient image super-resolution reconstruction.

CN121504722APending Publication Date: 2026-02-10ANHUI QIURUI INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411677800.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing image super-resolution reconstruction methods suffer from high computational complexity, large number of parameters, and high computational resource requirements while maintaining performance, and are particularly inefficient when processing large images.

Method used

A hierarchical directional attention fusion-based image super-resolution reconstruction method is adopted, which designs a bifurcated attention residual block, a non-local coordinate attention and directional attention fusion module. Through the hierarchical attention fusion group and reconstruction module, the computational complexity is reduced and the feature extraction efficiency is improved.

Benefits of technology

Without increasing the number of additional parameters, it significantly improves image reconstruction quality, reduces computation time, and enhances model performance, especially in large-size image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504722A_ABST
    Figure CN121504722A_ABST
Patent Text Reader

Abstract

The invention discloses an image super-resolution reconstruction method based on hierarchical directional attention fusion, and relates to the field of image enhancement. According to the invention, a bifurcated attention residual block is provided, attention features are separated in the residual block, and residual information can be transmitted to different hierarchies independently. Non-local coordinate attention is designed, and average pooling operation is carried out on different directions of a one-dimensional space of input features, so that the complexity is reduced, and meanwhile, the performance is kept. A direction attention fusion module is introduced, and attention features and residual features are fused by using information of different spatial dimensions. Through the attention fusion layer and the residual fusion layer, the correlation between the features is captured in a global range, so that the model is helped to obtain more accurate features, and the image reconstruction quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image enhancement, and in particular to an image super-resolution reconstruction method based on hierarchical directional attention fusion. Background Technology

[0002] Super-resolution (SR) image reconstruction refers to the technique of reconstructing high-resolution (HR) images from low-resolution (LR) images. This technique offers an opportunity to address resolution limitations in various computer vision applications, such as medical imaging and security monitoring. SR is a polysodic process because one LR image may correspond to multiple different HR images. To overcome this uncertainty, researchers have proposed various solutions, ranging from interpolation-based methods to reconstruction-based methods, and more recently, learning-based methods.

[0003] Convolutional Neural Networks (CNNs), with their powerful representation capabilities and end-to-end training paradigm, have played a crucial role in the development of SR tasks (Super-Resolution Convolutional Neural Network, SRCNN). Most CNN-based SR methods stack numerous convolutional layers to deepen the network and achieve higher performance; however, this approach leads to the vanishing gradient problem during network training. Subsequently, the emergence of residual learning (Very Deep Convolutional Network for Single Image Super-Resolution, VDSR) effectively alleviated the vanishing gradient problem, but residual features require long propagation paths, which may lead to information loss in residual connections. Furthermore, increasing network depth leads to an increase in the number of parameters and computational cost, limiting its application on mobile devices. While some lightweight SR methods (Information Multi-Distillation Network, IMDN) have addressed the network complexity issue, simply reducing the number of parameters does not sufficiently improve the network's sensitivity to key information. Research (RCAN) has attempted to reweight feature channels by directly introducing channel attention within residual blocks to achieve more effective feature representation. However, introducing attention mechanisms within residual blocks increases model complexity, potentially interfering with the original residual features and even discarding some deeper information. On the other hand, simply introducing attention mechanisms cannot meet the model's need to capture and integrate global information across the entire input space. Multi-Scale Residual Network (MSRN) uses a simple hierarchical feature fusion structure, which helps fuse global information into feature representations, thereby reducing redundant information and improving model efficiency and generalization ability. However, deep residual blocks can only receive complex fused features, ignoring clearer residual features. Residual Feature Aggregation Network (RFANet) aggregates local features within residual blocks to obtain more powerful feature representations, but it is not conducive to information transfer and interaction between features, weakening the model's perception of the global structure. In recent years, Non-local attention (NLA) has become an important tool in computer vision and natural language processing. In the field of super-resolution reconstruction, this attention mechanism has demonstrated powerful performance. NLA calculates the correlation between each location and all other locations in the image. This consideration of global correlation allows the model to more effectively obtain correlation information between distant regions, which helps to preserve image details and structural information.While NLA offers significant advantages in improving model performance, its computational complexity is high, requiring more computing resources and memory, especially for SR tasks with large input feature sizes. Therefore, reducing computational cost while maintaining performance is a challenge for nonlocal attention research in the SR field.

[0004] For example, patent application CN116228542 A discloses an image super-resolution reconstruction method and system based on a cross-scale nonlocal attention mechanism. This method introduces a cross-scale nonlocal attention mechanism into the image super-resolution reconstruction model to learn and mine the relationship between LR features and large-scale HR patches within the same feature map. A dual regression network is introduced to provide an additional constraint; however, its computational complexity is high, requiring more computational resources and memory, especially for SR tasks with large input feature sizes. Therefore, there is an urgent need to propose an image super-resolution reconstruction method based on hierarchical directional attention fusion to solve the above problems. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an image super-resolution reconstruction method based on hierarchical directional attention fusion, which addresses the shortcomings of the prior art.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0007] A hierarchical directional attention fusion-based image super-resolution reconstruction method includes the following steps:

[0008] S1. Design an image super-resolution reconstruction model framework based on hierarchical directional attention fusion: The model framework includes shallow feature extraction, three hierarchical attention fusion groups, an attention fusion layer, a residual fusion layer, a directional attention fusion module, and a reconstruction module; and the hierarchical attention fusion group includes three bifurcated attention residual blocks;

[0009] S2, using a 3×3 single-layer convolution from the input image I LR Extract shallow features F0 from them;

[0010] (1)

[0011] Where: H SFE This is a shallow feature extraction function; It is a 3×3 convolutional layer;

[0012] S3. Input the shallow features F0 into the hierarchical attention fusion group for nonlinear mapping learning. Each hierarchical attention fusion group produces two output paths, namely the attention fusion layer and the residual fusion layer.

[0013] S4. The features from the attention fusion layer and the residual fusion layer are fused using the directional attention fusion module, i.e.:

[0014] (2)

[0015] In the formula: F DF For nonlinear mapping feature output; F DF For F G Element-wise addition of F0; F G The output of the final directional attention fusion module is DAFM(); DAFM() is the directional attention fusion function; AFP() represents the attention fusion feature; RFP() represents the residual fusion feature.

[0016] S5. Finally, the reconstruction module uses the sub-pixel shuffling convolutional layer in ESPCN to process F. DF Perform upsampling reconstruction, i.e.:

[0017] (3)

[0018] In the formula: I SR For output image; H RM For reconstruction function; H HDAFN This refers to the HDAFN network function.

[0019] As a further preferred embodiment of the present invention, it also includes S6, optimizing the image super-resolution reconstruction model framework based on hierarchical directional attention fusion: given a training set of N input images and their corresponding output images, the parameter set is... The network L1 loss function is:

[0020] (4)

[0021] As a further preferred embodiment of the present invention, the bifurcated attention residual block obtains the attention feature map directly using the Sigmoid activation function before the residual features are added element by element.

[0022] As a further preferred embodiment of the present invention, the residual features and attention features are expressed as follows:

[0023] (5)

[0024] (6)

[0025] in, This indicates a convolution operation with a kernel size of 3; Represents the ReLU activation function; X is the input feature; represents the Sigmoid activation function; residual represents the residual feature output; Attention represents the attention feature output.

[0026] As a further preferred embodiment of the present invention, the attention fusion layer and the residual fusion layer are subjected to 1×1 convolution after feature splicing for dimensionality reduction filtering.

[0027] As a further preferred embodiment of the present invention, the attention fusion layer incorporates nonlocal coordinate attention (NLCA).

[0028] As a further preferred embodiment of the present invention, the algorithm of the directional attention fusion module includes: given input residual group features X and attention group features Y, concatenating and fusing the two sets of features, and then reducing the dimensionality of their 1×1 convolutional channels to obtain the fused feature Z:

[0029] (7)

[0030] In the formula, Concat() is the concatenation and merging operation. This represents a convolution with a kernel size of 1; then, feature extraction is performed on Z using both horizontal attention XAttn and vertical attention YAttn:

[0031] (8)

[0032] (9)

[0033] In the formula, Let Xpool() represent the Sigmoid activation function, Xpool() represent one-dimensional horizontal global pooling, and Ypool() represent one-dimensional vertical global pooling.

[0034] (10)

[0035] (11)

[0036] Finally, the residuals of the lateral attention and vertical attention features are added and fused to obtain the output feature Out:

[0037] (12)

[0038] In the formula, This indicates that the broadcast class performs element-wise multiplication.

[0039] The present invention has the following beneficial effects:

[0040] 1. This invention captures the correlation between features globally through an attention fusion layer and a residual fusion layer, thereby helping the model obtain more accurate features and improve the image reconstruction quality. (1) A bifurcated attention residual block is proposed, which separates attention features within the residual block to ensure that residual information can be propagated to different levels independently; the attention feature is separated from the residual block into an independent output path, which runs parallel to the residual feature path to the fusion layer, obtaining attention features without increasing the number of additional parameters, and retaining the original residual features. (2) A non-local coordinate attention is designed to perform average pooling operations on different directions in the one-dimensional space of the input features to reduce complexity while maintaining performance; its complexity only increases linearly with the width and height of the input feature map, reducing the computational burden of self-attention features. (3) A directional attention fusion module is introduced to fuse attention features and residual features using information from different spatial dimensions.

[0041] 2. The parameterless attentional residual block FARB can focus on useful information through attention paths, improving feature fluidity. Compared to the basic residual block RB, the PSNR of the Set5 dataset is improved by 0.13dB with the addition of a fusion bottleneck convolution of about 60K parameters.

[0042] 3. Introducing Nonlocal Coordinate Attention (NLCA) only increased the reconstruction time by 23ms, far lower than that of Nonlocal Attention, while maintaining performance. Introducing Directional Attention Fusion (DAFM) improved the PSNR of the Set5 dataset by 0.05dB. This module effectively fuses attention features and residual features.

[0043] 4. Extensive experimental results demonstrate that HDAFN exhibits superior performance in peak signal-to-noise ratio (PSNR) and structural similarity in quantization tests across four benchmark datasets. Specifically, by adding the proposed modules to the residual blocks, HDAFN increases the number of parameters by only 149K, achieving an average improvement of 0.23 dB in PSNR across four times the benchmark datasets. Attached Figure Description

[0044] Figure 1 This is a flowchart of the image super-resolution reconstruction method based on hierarchical directional attention fusion according to the present invention;

[0045] Figure 2 This is a model structure diagram of the image super-resolution reconstruction method based on hierarchical directional attention fusion of the present invention;

[0046] Figure 3 It is the bifurcated attention residual block in the method model of this invention;

[0047] Figure 4 This refers to the nonlocal coordinate attention in the method model of this invention;

[0048] Figure 5 It is the directional attention fusion module in the method model of this invention;

[0049] Figure 6 Visualizations of the reconstruction results of different 4x SR models on the B100 dataset are shown;

[0050] Figure 7 The visualization results show the reconstruction results of different 4x SR models on the Urban100 dataset. Detailed Implementation

[0051] The present invention will now be described in further detail with reference to the accompanying drawings and specific preferred embodiments.

[0052] In the description of this invention, it should be understood that the terms "left side," "right side," "upper part," "lower part," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. "First," "second," etc., do not indicate the importance of the components, and therefore should not be construed as a limitation of this invention. The specific dimensions used in this embodiment are only for illustrating the technical solution and do not limit the scope of protection of this invention.

[0053] like Figure 1-7 As shown, an image super-resolution reconstruction method based on hierarchical directional attention fusion includes the following steps:

[0054] S1. Design an image super-resolution reconstruction model framework based on hierarchical directional attention fusion: The model framework includes shallow feature extraction, three hierarchical attention fusion groups, an attention fusion layer, a residual fusion layer, a directional attention fusion module, and a reconstruction module; and the hierarchical attention fusion group includes three bifurcated attention residual blocks;

[0055] S2, using a 3×3 single-layer convolution from the input image I LR Extract shallow features F0 from them;

[0056] (1)

[0057] Where: H SFE This is a shallow feature extraction function; It is a 3×3 convolutional layer;

[0058] S3. Input the shallow features F0 into the hierarchical attention fusion group for nonlinear mapping learning. Each hierarchical attention fusion group produces two output paths, namely the attention fusion layer and the residual fusion layer.

[0059] S4. The features from the attention fusion layer and the residual fusion layer are fused using the directional attention fusion module, i.e.:

[0060] (2)

[0061] In the formula: F DF For nonlinear mapping feature output; F DF For F G Element-wise addition of F0; F G The output of the final directional attention fusion module is DAFM(); DAFM() is the directional attention fusion function; AFP() represents the attention fusion feature; RFP() represents the residual fusion feature.

[0062] S5. Finally, the reconstruction module uses the sub-pixel shuffling convolutional layer in ESPCN to process F. DF Perform upsampling reconstruction, i.e.:

[0063] (3)

[0064] In the formula: I SR For output image; H RM For reconstruction function; H HDAFN This refers to the HDAFN network function.

[0065] This invention continues the use of residual blocks (RBs) in the classic EDSR method. While attention mechanisms have achieved significant results in SR, they require additional parameter computation, leading to an increase in model parameters and slower inference speed. Therefore, this invention designs a forked attention residual block (FARB), which focuses on learning information-rich regions without requiring additional parameters, such as... Figure 3 As shown, FARB directly uses the sigmoid activation function to obtain the attention feature map before element-wise summing of the residual features. This design not only maintains model complexity but also introduces an attention output path. Furthermore, preserving the output path of the original residual features helps address the information loss problem that attention may cause. The residual features and attention features can be represented as:

[0066] (5)

[0067] (6)

[0068] in, This indicates a convolution operation with a kernel size of 3; Represents the ReLU activation function; X is the input feature; represents the Sigmoid activation function; residual represents the residual feature output; Attention represents the attention feature output.

[0069] HAFG mainly consists of three branched attention residual blocks FARB, such as Figure 1 As shown, the attention features and residual features output by FARB are fused through two separate paths. After feature concatenation, a 1×1 convolution is added for dimensionality reduction filtering. This convolution may cause information loss during dimensionality reduction. Therefore, the attention fusion layer AFL also introduces an additional non-local coordinate attention (NLCA) to learn and adjust the most important features in the feature hierarchy. Subsequently, the two feature paths enter the directional attention fusion module DAFM, which aims to reduce feature degradation from three consecutive residual learning layers, improve the utilization rate of residual features in each layer, and pass the fused key detail features to the subsequent HAFG. Similarly, HAFG also introduces higher-level attention feature paths and residual feature paths.

[0070] Traditional non-local attention (NLA) involves a large number of matrix multiplications, which is costly in SR tasks dealing with large image inputs. For some large-scale computational operations, redundancy may exist in SR tasks. Therefore, this invention proposes a lightweight non-local coordinate attention (NLCA), such as... Figure 4 As shown, given a C×H×W feature map input, a 1×1 convolution is used to linearly map the input features to θ, Φ, and g. Then, average pooling is performed on θ and g in the one-dimensional horizontal direction to change the feature matrix structure. Matrix multiplication is then performed on θ and Φ to obtain autocorrelation features. Normalized attention coefficients are obtained through the Softmax activation function, and finally, these are multiplied back into the feature matrix g to restore its corresponding dimensions before outputting the features. Similarly, another attention parameter comes directly from global average pooling in the one-dimensional vertical direction, preserving accurate vertical positional information. Finally, the positional information sets from both directions are summed to strengthen the focus on the region of interest.

[0071] like Figure 5As shown, this invention addresses the potential for information redundancy after feature fusion by introducing a Directed Attention Fusion Module (DAFM). DAFM utilizes lateral and longitudinal attention across different spatial dimensions to extract and fuse features from the Attention Fusion Layer (AFL) and the Residual Fusion Layer (RFL), suppressing useless information and thus enhancing information representation across different channels and spatial regions. This achieves linear fusion of attention features and residual features at different levels. The following is a detailed description of DAFM:

[0072] Given input residual features X and attention features Y, the two sets of features are concatenated and fused, and then their dimensionality is reduced by 1×1 convolutional channels to obtain the fused feature Z:

[0073] (7)

[0074] In the formula, Concat() is the concatenation and merging operation. This represents a convolution with a kernel size of 1. Next, feature extraction is performed on Z using both horizontal attention (XAttn) and vertical attention (YAttn).

[0075] (8)

[0076] (9)

[0077] In the formula, Let Xpool() represent the Sigmoid activation function, Xpool() represent one-dimensional horizontal global pooling, and Ypool() represent one-dimensional vertical global pooling.

[0078] (10)

[0079] (11)

[0080] Finally, the residuals of the lateral attention and vertical attention features are added and fused to obtain the output feature Out:

[0081] (12)

[0082] In the formula, This indicates that the broadcast class performs element-wise multiplication.

[0083] Figure 6The visualization results show the reconstruction results of different 4x SR models on the B100 dataset. By comparing the SR images, it was found that the "zebra stripes" recovered by models such as NGSwin in image "253027" have incorrect textures and do not match the HR image, resulting in texture blurring. Although VapSR recovers the stripe direction correctly, it exhibits distortion. The HDAFN model provided by this invention performs more reasonably in the recovered stripes, ensuring texture clarity and accuracy, and is closest to the HR image. In image "148026", the HDAFN model recovers the "wooden plank" lines more accurately, without producing obvious incorrect directions. These results further demonstrate the superior performance of the HDAFN model of this invention.

[0084] Figure 7 The images presented are a series of images with frequently repeating regions. The comparison results fully demonstrate the superior performance of the HDAFN model provided by this invention in reconstructing geometric structures. Specifically, HDAFN successfully restored the square structure of the floor grid in the "Img061" image, significantly suppressed blur artifacts, and achieved clearer edge reconstruction. In contrast, the visualization effects of Transformer-based methods such as ESRT and NGSwin are significantly lower than those of HDAFN, especially when dealing with geometric structures and high-frequency information, where HDAFN performs better. HDAFN's effective noise handling and more accurate restoration of image details and textures enable it to better preserve image details in various scenarios.

[0085] The software used in the experiments of this invention includes an Ubuntu 18.04 operating system, Python 3.6 programming language, and a PyTorch 1.7 deep learning framework, accelerated using CUDA version 11.0. For hardware, it uses an Intel(R) Core(TM) i9-10980XE CPU @ 3.00 GHz, 32GB of memory, with 18 cores and 36 threads, and an NVIDIA RTX 3090 24G graphics card. The training dataset is DIV2K, and the training images undergo data augmentation techniques such as horizontal flipping. Each batch randomly crops 64 low-resolution image patches of 64×64 pixels as input, for a total of 600 training rounds, using a learning rate of 10. -3The ADAM optimizer was trained with the learning rate halved every 150 epochs. During the testing phase, four benchmark datasets were used: Set5, Set14, B100, and Urban100. Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) were used as performance metrics, calculated on the Y-channel of the image. FLOPs were calculated for an output image size of 1280×720 at the corresponding scale. The HDAF of this invention consists of three HAFGs, each with NLCA, DAFM, and three FARBs, and all convolutional layers use 64 channels.

[0086] To verify the effectiveness of the HDAFN method provided in this invention, its performance was compared with lightweight state-of-the-art methods such as BSRN and VapSR. Table 1 lists the quantitative results of PSNR and SSIM evaluation using different algorithms on four benchmark datasets with scale factors of ×2, ×3, and ×4. The table also includes the number of parameters and FLOPs for a more comprehensive comparative analysis. Experimental results show that, compared with other SR methods, the HDAFN method of this invention achieves the best PSNR and SSIM scores on all datasets. Especially in methods with relatively similar parameter counts, such as NGSwin and A2N, HDAFN demonstrates competitive or even superior results. It also exhibits excellent performance in recent Transformer-based methods such as NGSwin and ESRT. The experimental data fully demonstrate that HDAFN can effectively fuse, select, and preserve important relevant information throughout the network.

[0087] Table 1. Comparison of model performance metrics on benchmark datasets with different scale factors.

[0088] Note: The best performing result is highlighted in bold.

[0089] This invention proposes an image super-resolution reconstruction network based on hierarchical directional attention fusion. By using attention fusion layers and residual fusion layers, this invention globally captures the correlations between features, thereby helping the model obtain more accurate features and improving image reconstruction quality.

[0090] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all fall within the protection scope of the present invention.

Claims

1. An image super-resolution reconstruction method based on hierarchical directional attention fusion, characterized in that: Includes the following steps: S1. Design an image super-resolution reconstruction model framework based on hierarchical directional attention fusion; the image super-resolution reconstruction model framework based on hierarchical directional attention fusion includes shallow feature extraction, three hierarchical attention fusion groups, attention fusion layer, residual fusion layer, directional attention fusion module and reconstruction module; and the hierarchical attention fusion group includes three bifurcated attention residual blocks; S2, using a 3×3 single-layer convolution from the input image I LR Extract shallow features F0 from them; F0=H SFE (I LR )=Conv 3×3 (I LR ) (1) Where: H SFE Conv is a shallow feature extraction function; 3×3 (·) represents a 3×3 convolutional layer; S3. Input the shallow features F0 into the hierarchical attention fusion group for nonlinear mapping learning. Each hierarchical attention fusion group produces two output paths, namely the attention fusion layer and the residual fusion layer. S4. The features from the attention fusion layer and the residual fusion layer are fused using the directional attention fusion module, i.e.: F DF =F0+F G =F0+DAFM(AFP(·),RFP(·)) (2) In the formula: F DF For nonlinear mapping feature output; F DF For F G Element-wise addition of F0; F G This is the output of the final directional attention fusion module; DAFM() is the directional attention fusion function; AFP() represents the attention fusion feature; RFP() represents residual fusion features; S5. Finally, the reconstruction module uses the sub-pixel shuffling convolutional layer in ESPCN to process F. DF Perform upsampling reconstruction, i.e.: I SR =H RM (F DF )=H HDAFN (I LR ) (3) In the formula: I SR For output image; H RM For reconstruction function; H HDAFN This refers to the HDAFN network function.

2. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 1, characterized in that: This also includes S6, optimizing the image super-resolution reconstruction model framework based on hierarchical directional attention fusion: Given a training set of N input images and their corresponding output images, the network L1 loss function with parameter set θ is:

3. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 1, characterized in that: The bifurcated attention residual block obtains the attention feature map directly using the Sigmoid activation function before the residual features are added element by element.

4. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 3, characterized in that: The residual features and attention features are represented as follows: residual=X+Conv 3×3 (δ(Conv 3×3 (X))) (5) Attention=σ(Conv 3×3 (δ(Conv 3×3 (X)))) (6) Among them, Conv 3×3 (·) represents a convolution operation with a kernel size of 3; δ(·) represents the ReLU activation function; X is the input feature; σ(·) represents the Sigmoid activation function; residual represents the residual feature output; Attention represents the attention feature output.

5. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 1, Its features are: The attention fusion layer and residual fusion layer incorporate 1×1 convolutions for dimensionality reduction filtering after feature concatenation.

6. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 5, Its features are: The attention fusion layer incorporates nonlocal coordinate attention (NLCA).

7. The image super-resolution reconstruction method based on hierarchical directional attention fusion according to claim 1, Its features are: The algorithm of the directional attention fusion module includes: given input residual group features X and attention group features Y, concatenating and fusing the two sets of features, and then reducing the dimensionality of their 1×1 convolutional channels to obtain the fused feature Z: Z=Conv 1×1 (Concat(X,Y)) (7) In the formula, Concat() is the concatenation and merging operation, and Conv... 1×1 (·) represents a convolution with a kernel size of 1; then, feature extraction is performed on Z using both horizontal attention XAttn and vertical attention YAttn: XAttn(Z)=σ(Conv 1×1 (Xpool(Conv 1×1 (Z)))) (8) YAttn(Z)=σ(Conv 1×1 (Ypool(Conv 1×1 (Z)))) (9) In the formula, σ(·) represents the Sigmoid activation function, Xpool() represents one-dimensional horizontal global pooling, and Ypool() represents one-dimensional vertical global pooling, that is: Finally, the residuals of the lateral attention and vertical attention features are added and fused to obtain the output feature Out: In the formula, This indicates that the broadcast class performs element-wise multiplication.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method based on cross-scale non-local attention mechanism

    CN116228542A