A lightweight multi-scale global attention enhanced network for image super-resolution
The problem of high computing resources and difficulty in global information capture is solved through the lightweight multi-scale global attention enhancement network (LMGAE-Net), and efficient image super-resolution performance improvement is achieved, especially in image detail and edge recovery.
Patent Information
- Application Number
- CN202510008551.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing image super-resolution methods have problems such as high computing resource consumption and long computing time, and CNN-based methods are difficult to effectively capture global information.
The lightweight multi-scale global attention enhancement network (LMGAE-Net) is adopted to reduce parameter and calculation costs through the multi-scale global attention module (MGAB) and local information fusion module (LIFB), while improving image detail recovery capabilities. Multiple sets of shift fusion modules (MSFBs) are used to expand the receptive field and capture local information.
While reducing network parameters and computing costs, it significantly improves image super-resolution performance, which can effectively restore the details and edges of high-resolution images, and performs better than existing methods.
Smart Images

Figure CN119887524B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and analysis, and particularly relates to a lightweight multi-scale global attention enhanced network for image super-resolution. Background Art
[0002] Single-image super-resolution refers to the image reconstruction process of restoring a low-resolution image (LR) to a high-resolution image (HR). This image reconstruction technology has been widely applied in fields such as medical imaging, remote sensing images, and surveillance systems. In the process of generating high-resolution images, traditional methods mainly rely on interpolation techniques, such as nearest-neighbor interpolation, bilinear interpolation, and bicubic interpolation, to restore pixel information around the image.
[0003] Subsequently, the emerging convolutional neural network (CNN) has surpassed traditional interpolation-based methods in terms of the generated image quality. To further improve the performance, efforts have been made to train models using larger-scale and deeper-structured networks to achieve precise restoration of more details. However, these large models often require huge computational resources, rely on expensive hardware devices, and need longer computation time. For application scenarios that require obtaining high-resolution images, this may face problems such as high costs or long waiting times. Therefore, it is particularly urgent and necessary to construct a lightweight and high-performance model to effectively address such problems.
[0004] Although CNN-based methods have made significant progress, limited by the size of the convolutional kernel, CNN can only process local information. Introducing the self-attention mechanism (SA) in Transformer to process non-local regions of features, it adopts a redundant attention mechanism, resulting in a quadratic relationship between the computational complexity and the image size, and thus requires high computational resources. Nevertheless, this discontinuous window shifting method cannot integrate information outside the window, so there are certain limitations in capturing global information. Summary of the Invention
[0005] To solve the above problems, the present application proposes a lightweight multi-scale global attention enhanced network (LMGAE-Net) for image super-resolution. The influence brought by window discontinuity is alleviated through a multi-scale global attention module (MGAB). At the same time, in order to learn local fine features more deeply, a local information fusion module (LIFB) is designed, and multiple groups of shifted fusion modules (MSFB) are used to capture local information, which not only reduces the number of network parameters and computational costs, but also achieves excellent performance in super-resolution (SR).
[0006] To solve the above technical problems, the present invention is implemented in the following ways:
[0007] A lightweight multi-scale global attention enhanced network for image super-resolution, including a shallow feature learning module, an LMGAE deep feature extraction module, a multi-layer feature fusion module (MFF), and an image reconstruction module;
[0008] Given a degraded LR image I LR ∈R 3×H×W , where H and W represent the height and width of the LR image respectively, use a shallow feature learning module f SF (·) composed of a single 3×3 convolution to extract local feature F0∈R C×H×W , where C represents the number of channels of the intermediate feature, and its specific expression is as follows:
[0009] F0 = f SF (I LR )
[0010] Send F0 into N cascaded LMGAE deep feature extraction modules for feature extraction. The output f n (1≤n≤N) of the nth LMGAE deep feature extraction module is expressed as follows:
[0011]
[0012] Among them, represents the function of the nth LMGAE deep feature extraction module;
[0013] The LMGAE deep feature extraction module is jointly composed of a multi-scale global attention module (MGAB) and a local information fusion module (LIFB), and both modules adopt a residual learning strategy; the output feature map of the (n - 1)th LMGAE module is directly input into the nth LMGAE module, and the output function expression of the nth LMGAE module is as follows:
[0014]
[0015] Among them, X = (f MGAB (F n-1 ) + F n-1 );
[0016] All LMGAE deep feature extraction modules are input into the multi-layer feature fusion module (MFF module). The multi-layer feature fusion module fuses the output features of all LMGAE deep feature extraction modules through a 1×1 convolution, and further extracts features with the proposed MSFB module. To assist network learning, a global residual learning mechanism is also introduced; the output expression of the multi-layer feature fusion module is as follows:
[0017] F MFF = f MFF(Concat[F n , F n-1 ,..., F1]) + F0
[0018] where F n , F n-1 , …, F1 represent the outputs of all previous LMGAE modules;
[0019] Then, through a 3×3 convolutional layer and a pixel - shuffle layer in the image reconstruction module, perform upsampling to reconstruct the SR image I SR , and its specific expression is as follows:
[0020] I SR = f UP (W * F MFF )
[0021] where W represents the weight of the 3×3 convolutional layer, and f UP represents the pixel - shuffle operation.
[0022] Furthermore, the multi - scale global attention module (MGAB) divides a set of input features X ∈ R C×H×W into G groups, denoted as and calculates the self - attention of each group of features with different window sizes M g . Suppose the G groups of features are separated by channels, then the computational complexity of the G - group self - attention (SA) is expressed as
[0023] After completing the self - attention, the scattered channels are re - combined, and a 1×1 convolutional layer is used to integrate the feature information from different groups. Then the overall process of the multi - scale global attention module can be expressed as follows:
[0024] [X0, X1,..., X g = Split(X)
[0025] X = Conv 1×1 (Concat[X0, X1,..., X g )
[0026] where Split(·) represents the channel splitting operation, SA(·) represents the self - attention calculation process, Concat(·) represents the channel concatenation, and Conv 1×1 (·) represents the 1×1 convolutional operation, and X represents the input feature.
[0027] Furthermore, the local information fusion block (LIFB) consists of an enhanced spatial attention block (ESA) and a multi-group shifted fusion block (MSFB) in sequence, and the expression of the local information fusion block processing flow is as follows:
[0028] F LIFB = F MSFB (F ESA (X))
[0029] The enhanced spatial attention block reduces the channel number of the input features through a 1×1 convolutional layer, reduces the spatial size of the features by using strided convolution and max pooling layers, extracts features through a group of convolutional layers, and restores the spatial size of the features through an upsampling operation based on interpolation. The output F of the enhanced spatial attention block ESA The expression is as follows:
[0030]
[0031] where f sigmoid represents the sigmoid function, represents the weight of the 1×1 convolutional layer, represents the weight of the 3×3 convolutional layer with a stride of 2, represents the weight of the 1×1 convolutional layer that restores the channel dimension, represents the output of the first layer of ESA, represents the output of the 2nd to 5th layers of ESA, f pool represents the max pooling operation, f up represents the upsampling function implemented by bilinear interpolation, Conv g represents the Group Conv layer;
[0032] The multi-group shifted fusion block evenly divides the output feature F ESA into nine groups, and moves these features along nine different spatial directions (down, right, up, left, unchanged, bottom-right, top-right, bottom-left, top-left), further extracts features by using 1×1 convolution, and comprehensively combines the information of the surrounding nine pixels to effectively expand the receptive field without adding additional learnable parameters and significant computational overhead, and maintains an arithmetic complexity comparable to that of a single 1×1 convolution; given the input feature Y, the process of the multi-group shifted fusion block (MSFB) is expressed as follows:
[0033] F MSFB = Conv 1×1 (Concat(Y0, Y1,..., Y9))
[0034] [Y0, Y1,..., Y9] = Split(Y)
[0035] Y i = fshift (Y i ), i ∈ {0, 1, ..., 9}
[0036] Among them, f shift (·) represents a shift operation.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] The multi-scale global attention module (MGAB) proposed in the present invention divides features into different groups through a novel grouping strategy, and independently calculates self-attention using windows of different sizes, effectively expanding the receptive field and enhancing the ability to capture long-range information; the proposed multi-group shift fusion module (MSFB) captures local information, and achieves the receptive field range of a 3×3 convolution with the computing resources of a 1×1 convolution. While maintaining a low computational cost, it effectively learns and fuses local features, enhancing the network's ability to restore image details. The proposed lightweight multi-scale global attention enhanced network (LMGAE-Net) for image super-resolution reduces the number of network parameters and computational cost, effectively solving the single-image super-resolution problem; it has been rigorously qualitatively and quantitatively evaluated on common benchmark datasets, and has achieved good performance in terms of super-resolution performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic structural diagram of the lightweight multi-scale global attention enhanced network (LMGAE-Net) for image super-resolution of the present invention;
[0040] Figure 2 It is a schematic structural diagram of the lightweight multi-scale global attention module (LMGAE) of the present invention;
[0041] Figure 3 It is a schematic structural diagram of the multi-scale global attention module (MGAB) of the present invention;
[0042] Figure 4 It is a schematic structural diagram of the local information fusion module (LIFB) of the present invention;
[0043] Figure 5 It is a schematic structural diagram of the enhanced spatial attention module (ESA) of the present invention;
[0044] Figure 6 It is a schematic structural diagram of the multi-group shift fusion module (MSFB) of the present invention;
[0045] Figure 7 It is a schematic diagram for comparing the visual effects of ×4SR on the Urban100 and BSD100 datasets of the present invention;
[0046] Figure 8Schematic diagram of the influence of the adjusted number of groups on the test performance of the present invention. Detailed implementation manners
[0047] The following further elaborates on the detailed implementation manners of the present invention in conjunction with the accompanying drawings and specific embodiments.
[0048] As Figure 1 shown, a lightweight multi-scale global attention enhanced network for image super-resolution includes a shallow feature learning module, an LMGAE deep feature extraction module, a multi-layer feature fusion module (MFF), and an image reconstruction module;
[0049] Given a degraded LR image ILR ∈ R 3×H×W , where H and W respectively represent the height and width of the LR image, a shallow feature learning module f SF (·) composed of a single 3×3 convolution is used to extract local features F0 ∈ R C×H×W , where C represents the number of channels of the intermediate features, and its specific expression is as follows:
[0050] F0 = f SF (I LR )
[0051] F0 is fed into N cascaded LMGAE deep feature extraction modules for feature extraction. Assuming the model has N LMGAE modules, the output f n (1 ≤ n ≤ N) of the nth LMGAE deep feature extraction module is expressed as follows:
[0052]
[0053] Among them, represents the function of the nth LMGAE deep feature extraction module, and F n represents the output of this module.
[0054] As Figure 2 shown, the LMGAE deep feature extraction module is jointly composed of a multi-scale global attention module (MGAB) and a local information fusion module (LIFB), and both modules adopt a residual learning strategy; the output feature map F n-1 of the (n - 1)th LMGAE module is directly input into the nth LMGAE module, and the output function expression of the nth LMGAE module is as follows:
[0055]
[0056] Among them, X = (f MGAB (F n-1 ) + F n-1 );
[0057] All the deep feature extraction modules of LMGAE are input into the multi-layer feature fusion module (MFF module). The multi-layer feature fusion module fuses the output features of all the deep feature extraction modules of LMGAE through a 1×1 convolution, and further extracts features with the proposed MSFB module. To assist the network in learning, a global residual learning mechanism is also introduced. The MFF module fuses the features of all previous LMGAE modules. The output expression of the multi-layer feature fusion module is as follows:
[0058] F MFF = f MFF (Concat[F n ,F n-1 ,...,F1]) + F0
[0059] where F n ,F n-1 ,…,F1 represent the outputs of all previous LMGAE modules;
[0060] Then, it is upsampled through a 3×3 convolutional layer and a pixel-shuffle layer in the image reconstruction module to reconstruct the SR image I SR , and its specific expression is as follows:
[0061] I SR = f UP (W * F MFF )
[0062] where W represents the weight of the 3×3 convolutional layer, and f UP represents the pixel-shuffle operation.
[0063] The window-based self-attention mechanism plays an important role in reducing the huge computational overhead of the Transformer model. Assuming the size of the feature map is C×H×W, when using non-overlapping windows of size M×M, its computational complexity is reduced to 2M 2 HWC; where the window size M plays a decisive role in the calculation range of self-attention. Selecting a larger M value helps the network capture global features more effectively and deeply mine self-similar information. Simply increasing M will lead to a quadratic growth in computational cost and resource requirements.
[0064] To learn long-range features more efficiently, a multi-scale global attention module (MGAB) is proposed. As Figure 3 shown, the multi-scale global attention module (MGAB) divides a set of input features X ∈ R C×H×W into G groups, denoted as and adopts different window sizes M for each group of features gCalculate the self-attention of Group G. This method not only balances the computational efficiency and feature representation ability, but also enables the network to capture rich context information at multiple scales, thereby improving the overall performance of the model. Set that the G groups of features are separated by channels, then the computational complexity of the G groups of self-attention (SA) is expressed as
[0065] After completing the self-attention, the scattered channels are remerged, and the 1×1 convolutional layer is used to integrate the feature information from different groups. Then the overall process of the multi-scale global attention module can be expressed as follows:
[0066] [X0, X1,..., X g = Split(X)
[0067] X i = SA(X i ), i ∈ {0, 1,..., g}
[0068] X = Conv 1×1 (Concat[X o , X1,..., X g )
[0069] where Split(·) represents the channel splitting operation, SA(·) represents the self-attention calculation process, Concat(·) represents the channel concatenation, Conv 1×1 (·) represents the 1×1 convolutional operation, and X represents the input feature.
[0070] As Figure 4 shown, the local information fusion module (LIFB) is composed of a lightweight enhanced spatial attention module (ESA) and a multi-group shift fusion module (MSFB) in sequence, and the expression of the processing flow of the local information fusion module is as follows:
[0071] F LIFB = F MSFB (F ESA (X))
[0072] To ensure efficiency while enhancing the expression ability of the model, we adopt the lightweight enhanced spatial attention module (ESA), which can effectively enhance the performance of the model in the spatial dimension. As Figure 5 shown, the enhanced spatial attention module reduces the number of channels of the input feature through the 1×1 convolutional layer, uses the strided convolution and the max pooling layer to reduce the spatial size of the feature, extracts features through a group of convolutional layers, and restores the spatial size of the feature through the upsampling operation based on interpolation. The output F ESA of the lightweight enhanced spatial attention module is expressed as follows:
[0073]
[0074] Among them, f sigmoid represents the sigmoid function, represents the weight of the 1×1 convolutional layer, represents the weight of the 3×3 convolution with a stride of 2, represents the weight of the 1×1 convolutional layer for restoring the channel dimension, represents the output of the first layer of ESA, represents the output of the 2nd to 5th layers of ESA, f pool represents the max pooling operation, f up represents the upsampling function implemented by bilinear interpolation, Conv g represents the Group Conv layer, and X represents the input feature of LIFB;
[0075] To deeply capture local features, 1×1 or 3×3 convolutional layers have usually been relied on in the past. However, the 1×1 convolution is not sufficient to capture enough context information due to its limited receptive field, and although the 3×3 convolution can expand the receptive field, it comes at the cost of consuming a large amount of computing resources. Therefore, in order to expand the receptive field without sacrificing computational efficiency in this application, a multi-group shift fusion module (MSFB) is proposed. As Figure 6 shown, the multi-group shift fusion module divides the output feature F ESA evenly into nine groups and moves these features along nine different spatial directions (down, right, up, left, unchanged, bottom-right, top-right, bottom-left, top-left), and uses 1×1 convolution to further extract features. This design enables the 1×1 convolution to synthesize the information of the surrounding nine pixels, effectively expanding the receptive field without adding additional learnable parameters and significant computational overhead, and maintaining an arithmetic complexity comparable to that of a single 1×1 convolution; given the input feature Y, the process of the multi-group shift fusion module (MSFB) is as follows:
[0076] F MSFB = Conv 1×1 (Concat(Y0, Y1,..., Y9))
[0077] [Y0, Y1,..., Y9] = Split(Y)
[0078] Y i = f shift (Y i ), i ∈ {0, 1,..., 9}
[0079] Among them, f shift (·) represents the shift operation.
[0080] In the embodiments of the present application, through quantitative and qualitative evaluation methods, the effectiveness of the proposed LMGAE-Net is verified on five commonly used SR benchmark datasets.
[0081] Datasets and evaluation metrics: 800 images from the DIV2K dataset are selected as the training set, and the corresponding low-resolution (LR) images are created by applying bicubic downsampling technology to the HR images. To verify the performance of the model, it is tested on five commonly used benchmark datasets, including Set5, Set14, BSD100, Urban100, and Manga109; to evaluate the quality of the reconstructed images, the peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) are used. Among them, the PSNR and SSIM values are calculated on the Y channel after converting the images from the RGB format to the YCbCr color space.
[0082] Experimental details: In the training stage, data augmentation methods such as random rotation by 90°, 180°, 270°, and horizontal flipping are introduced. The LMGAE-Net architecture contains 16 LMGAE modules equipped with 60 channels. These modules calculate the MGAB using window sizes of 4×4, 8×8, 16×16, and 32×32 on four equally divided channels. 16 image patches of size 64×64 are randomly selected from the low-resolution (LR) images as the basic input units for training. The model training uses the Adam optimizer with parameter settings of β1 = 0.9 and β1 = 0.99, and the initial learning rate is set to 1×10 -3 , and the total number of iterations is set to 300,000 times. To guide the model training, a composite metric of the mean absolute error (MAE) loss and the frequency loss function based on the fast Fourier transform (FFT) is used. All experimental operations are completed using the Pytorch framework in an environment equipped with an NVIDIA A100 Tensor Core GPU.
[0083] To evaluate the performance of the LMGAE-Net model, it is compared with the state-of-the-art lightweight SR at different scaling factors for quantitative result analysis.
[0084] Table 1 Performance comparison of different lightweight SR models on five benchmark models
[0085]
[0086]
[0087] Table 1 shows the quantitative comparison of various lightweight methods on five benchmark datasets. The results show that LMGAE-Net achieved the best performance on all datasets. Although LMGAE-Net has more parameters than RepRFN, its performance has been significantly improved. Taking the ×3 scale as an example, LMGAE-Net obtained a high PSNR of 28.65 on the Urban100 dataset, which is better than 28.06 of RepRFN.
[0088] Qualitative comparison: The SR results of LMGAE-Net and six representative models (including CARN, IMDN, RFDN, ShuffleMixer, SAFMN, and RepRFN) in the ×4 upscaling task were compared in terms of visual quality. As Figure 7 shown, it was observed that most of the comparison methods would produce blurred or inaccurate edges and textures during image restoration, while the method of this application was able to restore more accurate and clear edge details. This advantage of LMGAE-Net was particularly prominent in images containing repetitive patterns and clear edges. During the reconstruction process, most methods would introduce artifacts and deformations, while in contrast, LMGAE-Net was able to restore clearer patterns and edges.
[0089] To deeply understand the working principle of LMGAE-Net, a series of ablation experiments were conducted to analyze the contributions of each component in the method. To ensure a fair comparison with the designed baseline model, all experiments were carried out at the ×4 upscaling scale and with consistent training configurations. The experiment started with a streamlined baseline model from which the MGAB module and the LIFE module were removed (Model ①). Then, the MGAB module (Model ②) and the MSFB module (Model ③) were gradually introduced into the baseline model, and the effect after removing the ESA module was shown (Model ④). Finally, all these modules together constituted the complete version of the method (Model ⑤), and the specific results are shown in Table 2.
[0090] Table 2 PSNR / SSIM comparison of the ×4 models under different settings
[0091]
[0092] As the core component of the LMGAE-Net architecture, MGAB performs excellently in obtaining global information. By comparing Model ① and Model ②, we found that the introduction of MGAB led to a stable improvement in PSNR and SSIM on all five datasets, with the maximum increase in PSNR being 0.51 dB. To deeply understand its mechanism of action, MGAB was further explored by adjusting the number of groups N to test the effect of the grouping strategy. As Figure 8As shown, with the increase in the number of groups, the performance of the model has been steadily improved, and this result confirms the effectiveness of the grouping strategy.
[0093] Effectiveness of MSFB: The MSFB module can efficiently capture local features while reducing the computational cost. Comparing Model ① and Model ③, it can be clearly observed that after introducing the MSFB module, the number of parameters of the model is reduced by approximately 38.8%, and the performance has also been significantly improved. To further verify the effectiveness of the MSFB module, a series of ablation experiments were carried out. As shown in Table 3, on the premise of keeping the number of parameters constant, only learning the left-middle-right (①) or up-middle-down (②) positions of pixels will lead to a significant decline in the model performance. And only moving the modules up, down, left, and right (③) can only capture the information of 4 surrounding pixels and itself, and its performance is still inferior to that of the module moving in nine different directions (④). Generally speaking, the MSFB module performs well in effectively encoding local information and reducing the number of parameters.
[0094] Table 3 Influence of Different Moving Groups on Experimental Results
[0095]
[0096]
[0097] To confirm the actual effect of this attention module, ablation experiments were conducted. As shown by Model ④ in Table 2, compared with the model without ESA, the complete LMGAE-Net has achieved performance improvement on five commonly used datasets. This result indicates that ESA can effectively enhance the overall expression ability of the model.
[0098] The above are only the implementation manners of the present invention. Once again, it is stated that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements can still be made to the present invention, and these improvements are also included in the protection scope of the claims of the present invention.
Claims
1. A lightweight multi-scale global attention enhanced network for image super-resolution, characterized in that: It includes a shallow feature learning module, an LMGAE deep feature extraction module, a multi-layer feature fusion module, and an image reconstruction module; Given the degenerate LR image I LR ∈R 3×H×W , where H and W respectively represent the height and width of the LR image, a shallow feature learning module f SF (·) is used to extract the local feature F0 ∈ R C×H×W , where C represents the number of channels of the intermediate feature, and its specific expression is as follows: F0 = f SF (I LR ) Send F0 to N cascaded LMGAE deep feature extraction modules for feature extraction. The output F of the nth LMGAE deep feature extraction module n The expression is as follows: Among them, represents the nth LMGAE deep feature extraction module function, where 1 ≤ n ≤ N; The LMGAE deep feature extraction module is jointly composed of a multi-scale global attention module and a local information fusion module, and both modules adopt a residual learning strategy; the output feature map of the (n-1)th LMGAE module is directly input into the nth LMGAE module, and the output function expression of the nth LMGAE module is as follows: where X = (f MGAB (F n-1 ) + F n-1 ); All LMGAE deep feature extraction modules are input into the multi-layer feature fusion module. The multi-layer feature fusion module fuses the output features of all LMGAE deep feature extraction modules through a 1×1 convolution, and further extracts features by means of the proposed MSFB module. The output expression of the multi-layer feature fusion module is as follows: F MFF = f MFF (Concat[F n , F n-1 ,..., F1) + F0 Among them, F n , F n-1 ,..., F1 represents the outputs of all previous LMGAE modules; Then, upsampling is performed through a 3×3 convolutional layer and a pixel-shuffle layer in the image reconstruction module to reconstruct the SR image I SR , and its specific expression is as follows: I SR = f UP (W * F MFF ) Among them, W represents the weight of the 3×3 convolutional layer, and f UP represents the pixel-shuffle operation; The local information fusion module is sequentially composed of an enhanced spatial attention module and a multi-group shift fusion module, and the expression of the processing flow of the local information fusion module is as follows: F LIFB = F MSFB (F ESA (X)) The enhanced spatial attention module reduces the number of channels of the input features through a 1×1 convolutional layer, uses strided convolution and max pooling layers to reduce the spatial size of the features, extracts features through a set of convolutional layers, and restores the spatial size of the features through an interpolation-based upsampling operation. The output F of the lightweight enhanced spatial attention module ESA The expression is as follows: F1 esa = w1 esa * Conv 1×1 (X) Among them, f sigmoid represents the sigmoid function, W1 esa represents the weight of the 1×1 convolutional layer, represents the weight of the 3×3 convolution with a stride of 2, represents the weight of the 1×1 convolutional layer that restores the channel dimension, F1 esa represents the output of the first layer of ESA, represents the output of the 2nd to 5th layers of ESA, f pool represents the max pooling operation, f up represents the upsampling function implemented by bilinear interpolation, Conv g represents the Group Conv layer; The multiple groups of shift fusion modules will output the feature F ESA Evenly divide it into nine groups, move these features along nine different spatial directions, further extract features using 1×1 convolution, and maintain an arithmetic complexity comparable to that of a single 1×1 convolution; given the input feature Y, the process of the multiple groups of shift fusion modules is expressed as follows: F MSFB = Conv 1×1 (Concat(Y0, Y1, …, Y9)) [Y0, Y1,..., Y9] = Split(Y) Y i = f shift (Y i ), i ∈ {0, 1,..., 9} where f shift (·) represents a shift operation.
2. A lightweight multi-scale global attention enhanced network for image super-resolution according to claim 1, characterized in that: The multi-scale global attention module divides a set of input features \(X\in\mathbb{R}\) C×H×W into \(G\) groups, denoted as and calculates the self-attention of each group of features with different window sizes \(M\) g Suppose the \(G\) groups of features are separated by channels, then the computational complexity of the \(G\) groups of self-attention is expressed as After self-attention is completed, the scattered channels are recombined, and the feature information from different groups is integrated by a 1×1 convolutional layer. Then the overall process of the multi-scale global attention module can be expressed as follows: [X0, X1,...., X g = Split(X) X i = SA(X i ), i ∈ {0, 1, ..., g} X = Conv 1×1 (Concat[X0, X1, …, X g ) Among them, Split(·) represents the channel splitting operation, SA(·) represents the self-attention calculation process, Concat(·) represents channel concatenation, and Conv 1×1 (·) represents the 1×1 convolution operation, and X represents the input feature.
Citation Information
Patent Citations
Global-local cooperation lightweight image super-resolution method based on semantic guidance
CN117314753A
Image super-resolution reconstruction method, terminal equipment and storage medium
CN117575915A