Lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel cavity coordinate attention

A lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention solves the problems of high-frequency information loss and high computational complexity in remote sensing image super-resolution reconstruction, achieving high-quality image reconstruction and improved computational efficiency.

CN120876231APending Publication Date: 2025-10-31INST OF EARTH ENVIRONMENT CHINESE ACAD OF SCI

Patent Information

Application Number
CN202511054087.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods for super-resolution reconstruction of remote sensing images suffer from problems such as loss of high-frequency information leading to loss of texture details and blurred edges, and increased network depth resulting in increased parameter and computational costs, which reduce the practicality and efficiency of the model.

Method used

A lightweight super-resolution reconstruction method employing multi-scale feature extraction and parallel dilated coordinate attention is proposed. Through depthwise separable convolution and parallel dilated convolution attention mechanism, multi-scale features are extracted and detailed information is captured, reducing computational complexity.

Benefits of technology

It effectively preserves high-frequency details, avoids detail loss, reduces computational complexity, improves image quality and practicality, and is suitable for low-quality and noisy image reconstruction tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876231A_ABST
    Figure CN120876231A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel cavity coordinate attention. The lightweight super-resolution reconstruction method comprises the following steps: constructing a shallow feature extraction sub-network for remote sensing image super-resolution reconstruction; constructing a deep feature extraction sub-network for remote sensing image super-resolution reconstruction; constructing an image restoration and reconstruction self-sub-network for remote sensing image super-resolution reconstruction; constructing a remote sensing image super-resolution reconstruction network; generating a training set, a verification set and a test set; training a remote sensing image super-resolution reconstruction network; reconstructing remote sensing image resolution; according to the method, multi-scale feature extraction, a parallel cavity coordinate attention mechanism and depth separable convolution are combined, so that the effect and the quality of super-resolution reconstruction are improved while relatively low calculation cost is kept; the problems of detail loss, high calculation complexity, insufficient utilization of multi-scale features and the like when complex textures and high-frequency details are processed in an existing reconstruction technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image enhancement technology, and specifically relates to a lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention. Background Technology

[0002] Super-resolution reconstruction of remote sensing images is an image processing technique that uses computers to process a low-resolution (LR) remote sensing image or image sequence to recover a high-resolution (HR) remote sensing image. HR means that the image has a high pixel density and can provide more details. Due to hardware limitations, the quality of remote sensing images acquired from satellites may not meet the needs of multiple tasks such as ground target recognition, semantic segmentation, and change detection. Using remote sensing image super-resolution reconstruction technology can improve the resolution of remote sensing images, improve the clarity of remote sensing images at a relatively low economic cost, and increase the utilization rate of remote sensing images.

[0003] Traditional super-resolution reconstruction methods are mainly based on theories such as interpolation or sparse representation. However, these methods have limited effectiveness when dealing with complex textures and edges. With the rise of deep learning, CNN-based super-resolution reconstruction methods have gradually become the mainstream methods. CNNs can automatically learn the feature representation of images and restore image details through multi-layer convolution operations. However, existing CNN-based super-resolution reconstruction methods still have problems such as loss of high-frequency information, high network complexity, and insufficient utilization of multi-scale features.

[0004] Inventors Hu Shufeng and Chen Xiaoxuan, in their patent "Remote Sensing Image Super-Resolution Reconstruction Method Based on Hybrid Multi-Scale Attention Group" (application filed on February 27, 2025, CN202510227286.5), proposed a remote sensing image super-resolution reconstruction method based on hybrid multi-scale attention group. This method effectively combines global and local information of remote sensing images, meeting the requirements for breadth and fine granularity in remote sensing image reconstruction. The specific steps of this method are as follows: shallow feature map extraction is performed on the input low-resolution remote sensing image; deep feature map extraction is performed on the shallow feature map using a residual hybrid multi-scale attention network; the shallow feature map and the deep feature map are aggregated and then reconstructed to obtain a super-resolution remote sensing image. The shortcomings of this method are that remote sensing images have characteristics such as multiple bands, angular distortion, and large differences in ground object scale, and the method requires angular invariance design and physical model constraints.

[0005] In their patent "A Deep Learning-Based Single Remote Sensing Image Super-Resolution Method" (application dated April 3, 2023, CN202310344467.7), inventors Jiao Jizong and Pei Xiudong proposed a deep learning-based single remote sensing image super-resolution method. This method improves the spatial resolution of low-resolution satellite images, providing an effective solution for image processing and denoising applications. The specific steps of this method are: downsampling high-resolution remote sensing sample images using a bicubic kernel function to generate corresponding low-resolution images; using the original high-resolution images as corresponding label images; using high- and low-resolution remote sensing image pairs as training data; dividing the high- and low-resolution image pairs into training and validation sets proportionally, and performing data augmentation on the training set; feeding the training set into a deep learning-based super-resolution model for training, and using the validation set to evaluate the super-resolution model to obtain the optimal network parameter model; the drawback of this method is that pre-training on ImageNet (Natural RGB) and direct transfer to remote sensing images can lead to feature distribution misalignment.

[0006] In their paper "Image Super-Resolution Reconstruction Method Based on Lightweight Symmetric CNN-Transformer" (Pattern Recognition and Artificial Intelligence, Vol. 37, No. 7, 2024, pp. 626-637), Wang Tingwei, Zhao Jianwei, and other authors proposed an image super-resolution reconstruction method combining a lightweight symmetric convolutional neural network (CNN) and a Transformer structure. This method, while maintaining low computational cost, fully utilizes the local feature extraction capabilities of CNNs and the global information modeling capabilities of Transformers to improve the super-resolution reconstruction effect. The specific steps of this method are as follows: A lightweight symmetric CNN module is used to extract local features of the low-resolution image, reducing computational complexity while preserving effective information; a symmetric Transformer module is used, employing a self-attention mechanism to enhance the global feature modeling capability and improve the image structure restoration effect; a feature fusion mechanism is used to integrate the features extracted by CNNs and Transformers to achieve efficient information interaction; and an upsampling operation is used to generate a super-resolution image, improving image clarity and detail representation. The drawback of this method is that, although significant progress has been made in terms of lightweighting and performance improvement, some detail loss may still exist when processing complex textures and high-frequency details, and real-time performance on large-scale datasets still needs further optimization.

[0007] In their paper "RDNet: Lightweight Residual and Detail Self-Attention Network for Infrared Image Super-Resolution" (Infrared Physics and Technology, 2024, Vol. 141, No. 105480), Feiyang Chen, Detian Huan, and other authors proposed a lightweight residual and detail self-attention network (RDNet) for infrared image super-resolution reconstruction. This method reduces computational complexity and improves the efficiency and practicality of the model while maintaining the quality of high-resolution infrared images. The specific steps of this method are as follows: Lightweight Residual Blocks (LRBs) are used to extract local features of the infrared image to enhance the preservation of texture information; a detail self-attention module is then used... The method employs a self-attention mechanism to capture long-range dependencies and optimizes the detail recovery of infrared images. It integrates features at different levels using a multi-scale feature fusion strategy to improve the reconstruction accuracy of the model. Finally, an upsampling module is used to convert low-resolution infrared images into high-resolution output to achieve super-resolution reconstruction. The drawback of this method is that although it achieves significant performance improvement in infrared image super-resolution reconstruction, it still has certain limitations when dealing with complex multi-scale details. Although the lightweight design improves computational efficiency, it sacrifices the generalization ability of the model to some extent.

[0008] In their paper "Image Super-Resolution Reconstruction Network Based on Transposed Attention and CNN" (Journal of Image and Graphics, 2025, Vol. 46, No. 1, pp. 35-46), Chen Guanhao, Xu Dan, and other authors proposed a method combining transposed attention... This image super-resolution reconstruction method utilizes a combination of attention mechanisms and convolutional neural networks (CNNs) to improve the restoration quality of low-resolution images by effectively modeling long-range dependencies and enhancing local feature extraction capabilities. The specific steps are as follows: First, a convolutional neural network is used to extract local features from the low-resolution image to obtain basic information. Second, a transposed attention mechanism is introduced to enhance feature interaction through rearranged attention mapping, better modeling long-range dependencies and improving the recovery of high-frequency details. Third, a feature fusion module is used to integrate local and global information to improve the reconstruction effect. Finally, an upsampling layer is used to generate a super-resolution image, achieving enhanced sharpness and detail. The limitations of this method are that, although the transposed attention mechanism helps capture global information, its computational complexity is higher than traditional convolutional methods, which may lead to a decrease in inference speed. When processing noisy, low-quality images, this method may suffer from artifacts and detail loss, and its ability to restore complex texture regions still needs further optimization.

[0009] Existing methods for super-resolution reconstruction of remote sensing images have the following drawbacks:

[0010] First, the loss of high-frequency information leads to the loss of texture details and blurred edges. Some super-resolution reconstruction methods are prone to losing some high-frequency information during image reconstruction, resulting in problems such as loss of texture details and blurred edges in the reconstructed image, which affects the reconstruction effect and practicality of the super-resolution reconstruction method.

[0011] Second, increasing network depth leads to an increase in the number of parameters and computational cost. Some super-resolution reconstruction methods achieve better image reconstruction results by increasing network depth, but the number of parameters and computational cost of the method will also increase, which is not conducive to building practical network structures. Summary of the Invention

[0012] To address the shortcomings of existing technologies, this invention designs a lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel dilated coordinate attention. By introducing multi-scale feature extraction and parallel dilated coordinate attention mechanisms, this method solves the problems of existing super-resolution reconstruction methods often resulting in blurred texture details and unclear edges due to the loss of high-frequency information during image reconstruction, and the problem that some methods increase network depth to improve reconstruction results, leading to a significant increase in the number of parameters and computation, thus reducing the practicality and efficiency of the model.

[0013] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0014] A lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention includes the following steps:

[0015] Step 1: Construct a shallow feature extraction subnetwork for super-resolution reconstruction of remote sensing images to extract shallow features from the original remote sensing images.

[0016] Step 2: Construct a deep feature extraction subnetwork for super-resolution reconstruction of remote sensing images, and extract deep features from the shallow feature map obtained in Step 1.

[0017] Step 3: Construct an image restoration and reconstruction sub-network for super-resolution reconstruction of remote sensing images, restore the deep feature map obtained in Step 2 to the original feature map size, and obtain the final super-resolution reconstructed image;

[0018] Step 4: Based on the shallow feature extraction subnetwork constructed in Step 1, the deep feature extraction subnetwork constructed in Step 2, and the image restoration and reconstruction subnetwork constructed in Step 3, construct a remote sensing image super-resolution reconstruction network.

[0019] Step 5: Generate the training set, validation set, and test set;

[0020] Step 6: Use the training set and validation set generated in Step 5 to train the remote sensing image super-resolution reconstruction network constructed in Step 4, and obtain the model weight file.

[0021] Step 7: Using the model weight file obtained from training in Step 6, test the remote sensing images in the test set generated in Step 5 to obtain the super-resolution reconstruction result image of the remote sensing images.

[0022] The specific method for step 1 is as follows:

[0023] A shallow feature extraction module is constructed for raw remote sensing images. The input is a remote sensing image at its original resolution. The structure of the shallow feature extraction module is: one depthwise separable convolutional module (DSC). q Extracting shallow feature maps, such as local edge and texture features; one LeakyReLU activation function is used to capture basic and intuitive features of the image, in DSC. q Based on the output, the non-linear response of edges and textures is enhanced to obtain a more expressive shallow feature map, which serves as the input to the deep module;

[0024] Step 1.1, Construct a depthwise separable convolutional module (DSC) q ;

[0025] The shallow feature extraction module includes Depthwise Separable Convolution (DSC). DSC decomposes the standard convolution operation into depthwise convolution and pointwise convolution, employing a channel-adaptive grouping strategy to group the input feature map F... o The channels are dynamically divided into several groups based on feature similarity. Each group shares a learnable sparse mask to reduce unnecessary computation. Each input channel is convolved separately, and then the channel information is integrated through pointwise convolution to extract the basic features of the image. A channel weight matrix is ​​generated through 1×1 convolution to suppress redundant channel interactions, such as deep convolutional layers for local edge and texture features.

[0026] Step 1.2 Construct the LeakyReLU activation function;

[0027] The LeakyReLU activation function performs a non-linear mapping on the shallow features extracted in step 1.1, enhancing the high-frequency contrast of edges and textures while preserving the weak response of the negative half-axis. This is achieved through DSC. q Based on the output, the nonlinear response of edges and textures is enhanced to obtain a more expressive shallow feature map F. q As the input to the deep feature extraction module, its expression is as follows:

[0028]

[0029] Where x is the gray value of the feature map, and α is a positive number less than 1, 0.01. When the input x is positive, the function output is equal to x; when the input x is negative, the function output is α times x.

[0030] The specific method for step 2 is as follows:

[0031] A deep feature extraction subnetwork with 16 feature extraction modules (FE) is constructed. The FE module structure is as follows: depthwise separable convolution DSC-1, depthwise separable convolution DSC-2, hybrid feature extraction HSEM, depthwise separable convolution DSC-3, and depthwise separable convolution DSC-4. The output of this subnetwork and the output of the shallow feature extraction subnetwork are merged through skip links before output.

[0032] Step 2.1: Construct the depthwise separable convolutional module DSC-1 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0033] The depthwise separable convolution module DSC-1 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map extracted by the shallow feature extraction module, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map.

[0034] Step 2.2: Construct the depthwise separable convolutional module DSC-2 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel dilated coordinate attention;

[0035] The depthwise separable convolution module DSC-2 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the depthwise separable convolution DSC-1, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map;

[0036] Step 2.3: Construct the Hybrid Feature Extraction Module (HSEM) in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention. The feature map is obtained by fusing the feature map output by DSC-2 through four branch paths.

[0037] Step 2.3.1: The feature map output by DSC-2 is directly passed through the single feature extraction module SFE to obtain the first intermediate feature map F1;

[0038] The Single Feature Extraction (SFE) module processes the input feature map F. i-1 Perform 3×3 convolution to extract the primary feature map F. i-2 The extracted primary feature map F i-2 Input a parallel hole coordinate attention module (AB) to perform inter-channel feature enhancement, resulting in feature map F. ia The enhanced feature map F ia Perform downsampling (3×3 convolution) and output feature map F. a To reduce spatial dimensionality; the extracted primary feature map F i-2 Perform two 3×3 convolutions to extract deep semantic features, resulting in a deep semantic feature map F. i-3 The extracted feature map F a Perform Sigmoid activation function processing and combine it with deep semantic features F i-3 Multiplying the results yields the deep feature map F. i-4; to extract deep feature maps F i-4 Perform a 1×1 convolution and then concatenate it with the input feature map F. i-1 The features are added together to restore the spatial resolution, resulting in the output feature map F. i ;

[0039] The parallel hole coordinate attention module (AB) extracts features from the input feature map F from the two branches of X and Y respectively;

[0040] In the X branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps X_avg and X_max. After merging X_avg and X_max, 3×3 dilated convolutions with dilation rates r=1, 2, 3, and 4 are performed. After merging the output features (concat), 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the X branch. x ;

[0041] In the Y branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps Y_avg and Y_max. After merging Y_avg and Y_max, 3×3 dilated convolutions with dilation rates r=1, 2, 3, and 4 are performed. After merging the output features (Concat), 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the Y branch. x ;

[0042] The feature maps Fx and Fy extracted from the X and Y branches are concatted along the channel dimension to form an attention map A. The attention map is then multiplied channel by channel with the original input F to achieve feature recalibration, resulting in the enhanced feature map F'.

[0043] Step 2.3.2: Input the feature map output by DSC-2 into the downsampling module Dwon-1, which outputs a feature map F with half the spatial resolution. d1 (Dimensions are H / 2 × W / 2 × C1); The feature map F d1 The input is fed into the Single Feature Extraction (SFE) module, and the output is the second intermediate feature map F2.

[0044] Step 2.3.3, transfer feature map F d1 The input is fed into the Dwon-2 downsampling module to obtain a feature map F with half the spatial resolution. d2 (Dimensions are H / 4 × W / 4 × C2); The feature map F d2 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F3.

[0045] Step 2.3.4, transfer the feature map F d2 The input is fed into the Dwon-3 downsampling module to obtain a feature map F with half the spatial resolution. d3 (Dimensions are H / 8×W / 8×C2); The feature map F d3 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F4.

[0046] Step 2.3.5: Input feature maps F3 and F4 into the Non-local Block (NLB), and use bilinear interpolation to restore the spatial resolution to H / 2×W / 2, obtaining feature map F. 34 ; the feature map F 34 Adding F2 to obtain the feature map F 234 ; the feature map F 234 The F1 input is fed into the NLB module, and bilinear interpolation is used to restore the spatial resolution to H×W to obtain the feature map F; the feature map F is convolved with the feature map F' by 1×1, and then added to the original feature map X to obtain the mixed feature map X';

[0047] Step 2.4: Construct the depthwise separable convolutional module DSC-4 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0048] The depthwise separable convolution module DSC-3 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map X' output by the hybrid feature extraction module HSEM, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map;

[0049] Step 2.5: Construct the depthwise separable convolutional module DSC-4 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0050] The depthwise separable convolutional module DSC-4 uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the depthwise separable convolutional module DSC-3, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map;

[0051] Step 2.6: Construct the skip link module in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0052] The skip link module uses the feature map extracted by the previous FE module as the input feature map of the next FE module; it adds the feature map output by the shallow feature extraction sub-network to the output feature map of the 16th FE module to output a deep feature map, which is then used as the input feature map of the image restoration and reconstruction sub-network.

[0053] The specific method for step 3 is as follows:

[0054] Construct an image restoration and reconstruction subnetwork, including two identical image restoration submodules and one depthwise separable convolutional submodule;

[0055] Step 3.1, construct the image restoration submodule 1 of the image restoration and reconstruction subnetwork;

[0056] The image restoration submodule 1 includes the following structure: a depthwise separable convolutional module DSC-h, two PixelShufflers, and one LeakyReLU activation function;

[0057] The depthwise separable convolutional module DSC-h uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the deep feature extraction module in step 2, extracting spatial features from different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused, reducing the number of parameters and computational cost, resulting in a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the depthwise separable convolutional module by a factor of 1 each time. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to the image restoration submodule 2.

[0058] Step 3.2: Construct the image restoration submodule 2 of the image restoration and reconstruction subnetwork;

[0059] The image restoration submodule 2 includes the following structure: a depthwise separable convolutional module DSC-h, two PixelShufflers, and one LeakyReLU activation function;

[0060] The depthwise separable convolutional module DSC-h uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the image restoration submodule 1, extracting spatial features from different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused, reducing the number of parameters and computational cost, resulting in a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the depthwise separable convolutional module by a factor of 1 in each stage. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to the depthwise separable convolutional module DSC-hc.

[0061] Step 3.3: Construct the depthwise separable convolutional module DSC-hc of the image restoration and reconstruction subnetwork;

[0062] The depth-separable convolution module DSC-hc uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the image restoration submodule 2, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the reconstructed remote sensing image.

[0063] The specific method for step 4 is as follows:

[0064] Based on the shallow feature extraction subnetwork constructed in step 1, the deep feature extraction subnetwork constructed in step 2, and the image restoration and reconstruction subnetwork constructed in step 3, a remote sensing image super-resolution reconstruction network is constructed.

[0065] Step 4.1: Concatenate the shallow feature map extracted in Step 1 and the deep feature map extracted in Step 2.

[0066] Step 4.2: Concatenate the feature map obtained by concatenating steps 1 and 2 with the feature map recovered from the image in step 3.

[0067] The specific method for step 5 is as follows:

[0068] Step 5.1: Obtain the publicly available dataset UC Merced Land-Use Dataset for super-resolution reconstruction of remote sensing images;

[0069] Step 5.2: Divide the dataset from Step 5.1 into training set, validation set and test set according to the proportions.

[0070] The specific method for step 6 is as follows:

[0071] Step 6.1: Set training parameters. Each time, randomly and non-repeating remote sensing images are selected from the training set and input into the network. The loss function L uses a content-based loss function (Li). Con ) and perception-based adversarial loss function (L Adv ) and use the sum to optimize the network;

[0072] The content-based loss function (L) Con )as follows:

[0073]

[0074] Where, φ i,j W represents the feature map of the j-th convolution before the i-th max pooling and after passing through the activation layer; i,j H i,j These are the width and height of the feature map, respectively; I LR For low-resolution remote sensing imagery, G θG : Generator, I HR High-resolution remote sensing imagery;

[0075] The perception-based adversarial loss function (L) Adv )as follows:

[0076]

[0077] Among them, D θD (G θG (I LR )) represents the probability that the generated image is real;

[0078] The super-resolution reconstruction loss function (L) is as follows:

[0079]

[0080] Step 6.2: Input the training set and validation set generated in Step 5 into the shallow feature extraction subnetwork of the remote sensing image super-resolution reconstruction method constructed in Step 4 to perform shallow feature extraction and obtain a shallow feature map; then, input the shallow feature map output by the shallow feature extraction subnetwork into the deep feature extraction subnetwork to perform deep feature extraction and obtain a deep feature map; finally, input the deep feature map output by the deep feature extraction subnetwork into the image restoration and reconstruction subnetwork to obtain the restored remote sensing image.

[0081] The specific method for step 7 is as follows:

[0082] Step 7.1: Input the test set generated in Step 5 into the remote sensing image super-resolution reconstruction network trained in Step 6, and load the remote sensing image super-resolution reconstruction network parameter file "best_loss.pth" trained in Step 7.1.

[0083] Step 7.2: Perform remote sensing image super-resolution reconstruction on the images in the test set to obtain high-resolution remote sensing images;

[0084] Step 7.3: Output the detection results of the remote sensing image super-resolution reconstruction network and save it as an image file in ".png" format.

[0085] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0086] (1) The lightweight super-resolution reconstruction network based on multi-scale feature extraction and parallel dilated coordinate attention proposed in this invention introduces a deep feature extraction module, which can process image features of different scales at the same time. Through the parallel dilated convolutional attention mechanism, the receptive field is effectively expanded and more detailed information is captured. When processing complex textures and edges, this invention can better preserve high-frequency details and avoid the problem of detail loss. In addition, the fusion design of multi-scale feature extraction and parallel dilated coordinate attention can improve image quality while avoiding the problem of reduced inference speed caused by complex attention mechanism. This makes the method perform well in low-quality and noisy image reconstruction tasks, reduce artifacts and enhance the restoration ability of complex texture regions.

[0087] (2) This invention introduces a depth-separable convolution module into a lightweight super-resolution reconstruction network based on multi-scale feature extraction and parallel dilated coordinate attention. This module splits the traditional convolution operation into two independent processes: depthwise convolution and pointwise convolution. This significantly reduces the number of model parameters and computational complexity, improves training speed, and saves storage space. This lightweight design ensures computational efficiency while avoiding the problem of decreased model generalization ability, making it more real-time on large-scale datasets and giving it broader practical application value.

[0088] In summary, compared with the best existing techniques for super-resolution reconstruction of remote sensing images, this invention innovatively combines multi-scale feature extraction, parallel hole coordinate attention mechanism, and depthwise separable convolution. It not only improves the effect and quality of super-resolution reconstruction while maintaining low computational cost, but also solves the problems of detail loss, high computational complexity, and insufficient utilization of multi-scale features in existing techniques when dealing with complex textures and high-frequency details. It has higher practicality and innovation. Attached Figure Description

[0089] Figure 1 This is a schematic diagram of the method flow in a specific embodiment of the present invention;

[0090] Figure 2 This is a schematic diagram of the network model structure in a specific embodiment of the present invention;

[0091] Figure 3 This is a schematic diagram of the shallow feature extraction subnetwork structure in a specific embodiment of the present invention;

[0092] Figure 4 This is a schematic diagram of the structure of the depth-separable convolution module (DSC) in a specific embodiment of the present invention;

[0093] Figure 5 This is a schematic diagram of the deep feature extraction subnetwork structure in a specific embodiment of the present invention;

[0094] Figure 6 This is a schematic diagram of the hybrid feature extraction module HSEM in a specific embodiment of the present invention;

[0095] Figure 7 This is a schematic diagram of the parallel dilated convolution coordinate attention AB module in a specific embodiment of the present invention;

[0096] Figure 8 This is a schematic diagram of the image restoration and reconstruction sub-network structure in a specific embodiment of the present invention; Figure 9 This is a visualization of the super-resolution reconstruction results on the UC Merced Land Use Dataset in a specific embodiment of the present invention. Detailed Implementation

[0097] The technical solution and effects of the present invention will be further described in detail below with reference to the accompanying drawings.

[0098] Reference Figure 1 The implementation steps of the embodiments of the present invention will be further described below;

[0099] The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention constructed in this invention comprises three main parts: a shallow feature extraction subnetwork, a deep feature extraction subnetwork, and an image restoration and reconstruction subnetwork. (Refer to...) Figure 2 The shallow feature extraction subnetwork is responsible for extracting shallow features from remote sensing images; the deep feature extraction subnetwork is responsible for extracting deep features from remote sensing images; the image restoration and reconstruction subnetwork is responsible for restoring the image and restoring the feature map of the remote sensing image to its original resolution; among them, the skip link module is responsible for fusing the feature map output by the shallow feature extraction subnetwork with the feature map output by the deep feature extraction subnetwork.

[0100] Step 1: Construct a shallow feature extraction subnetwork for super-resolution reconstruction of remote sensing images to extract shallow features from the original remote sensing images.

[0101] Reference Figure 3The following is a further description of the shallow feature extraction subnetwork for constructing super-resolution reconstruction of remote sensing images in this invention;

[0102] Figure 3 This diagram illustrates the shallow feature extraction subnetwork structure. A shallow feature extraction module with one depthwise separable convolution and a LeakyReLU activation function is constructed to extract shallow features from remote sensing images. Its structure consists of a depthwise separable convolution module (DSC-1) and a LeakyReLU activation function module. The DSC-1 module extracts shallow feature maps, while the LeakyReLU activation function captures basic and intuitive features of the image, enhancing the non-linear response of edges and textures based on the DSC-1 output to obtain a more expressive shallow feature map, which serves as the input to the deep modules.

[0103] Step 1.1: Construct the depthwise separable convolutional module DSC-1;

[0104] Reference Figure 4 The depthwise separable convolutional module DSC-1 in the feature extraction subnetwork structure for building changes in remote sensing images constructed in this invention will be further described.

[0105] Figure 4 This is a schematic diagram of the depthwise separable convolutional module structure. The input is the original remote sensing image, used to extract basic features of the input remote sensing image data, such as edges and textures. A channel-adaptive grouping strategy is adopted, which dynamically divides the input channels into several groups according to feature similarity. Each group shares a learnable sparse mask to reduce invalid computation. Convolution is performed on each input channel separately, and then the channel information is integrated through pointwise convolution to extract the basic features of the image. A channel weight matrix is ​​generated through 1×1 convolution to suppress redundant channel interactions, such as the depthwise convolutional layer for local edge and texture features.

[0106] Step 1.2 Construct the LeakyReLU activation function;

[0107] Based on the feature map output by DSC-1, the nonlinear response of edges and textures is enhanced to obtain a more expressive shallow feature map, which is used as the input to the deep feature extraction module. Its expression is as follows:

[0108]

[0109] Where x is the gray value of the feature map, and α is a positive number less than 1, 0.01; when the input x is positive, the function output is equal to x; when the input x is negative, the function output is α times x.

[0110] Step 2: Construct a deep feature extraction subnetwork with 16 feature extraction modules (FE). The output of this subnetwork and the output of the shallow feature extraction subnetwork are merged through skip links before output.

[0111] Reference Figure 5 The feature extraction module FE in the deep feature extraction subnetwork for super-resolution reconstruction of remote sensing images constructed in this invention will be further described.

[0112] Figure 5 This is a schematic diagram of the feature extraction module structure, which mainly includes 5 parts: depthwise separable convolution DSC-2, depthwise separable convolution DSC-3, hybrid feature extraction HSEM, depthwise separable convolution DSC-4, and depthwise separable convolution DSC-5; the input is the feature map output by the shallow feature extraction module or the feature map output by the previous FE;

[0113] Step 2.1: Construct the depthwise separable convolutional module DSC-2 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0114] Multi-scale convolution kernels (such as 3×3, 5×5, 7×7) are used to independently convolve each channel of the feature map extracted by the shallow feature extraction module to extract spatial features of different receptive fields and generate multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused to reduce the number of parameters and computational cost, and the fused feature map is obtained.

[0115] Step 2.2: Construct the depthwise separable convolutional module DSC-3 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0116] Multi-scale convolution kernels (such as 3×3, 5×5, 7×7) are used to independently convolve each channel of the feature map generated by depthwise separable convolution DSC-2 to extract spatial features of different receptive fields and generate multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused to reduce the number of parameters and computational cost, resulting in a fused feature map.

[0117] Step 2.3: Construct the Hybrid Feature Extraction Module (HSEM) in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention. The feature map is obtained by fusing the feature map output by DSC-3 through four branch paths.

[0118] Step 2.3.1: The feature map output by DSC-3 is directly passed through the single feature extraction module SFE to obtain the first intermediate feature map F1;

[0119] Reference Figure 6 The hybrid feature extraction in the deep feature extraction subnetwork structure for super-resolution reconstruction of remote sensing images constructed in this invention is further described.

[0120] Figure 6 This is a schematic diagram of the hybrid feature extraction module structure, which processes the input feature map F. i-1 Perform 3×3 convolution to extract the primary feature map F. i-2 The extracted primary feature map F i-2 Input a parallel hole coordinate attention module (AB) to perform inter-channel feature enhancement, resulting in feature map F. ia The enhanced feature map F ia Perform downsampling (3×3 convolution) and output feature map F. a To reduce spatial dimensionality; the extracted primary feature map F i-2 Perform two 3×3 convolutions to extract deep semantic features, resulting in a deep semantic feature map F. i-3 The extracted feature map F a Perform Sigmoid activation function processing and combine it with deep semantic features F i-3 Multiplying the results yields the deep feature map F. i-4 ; to extract deep feature maps F i-4 Perform a 1×1 convolution and then concatenate it with the input feature map F. i-1 The features are added together to restore the spatial resolution, resulting in the output feature map F. i ;

[0121] Reference Figure 7 The parallel hole coordinate attention module (AB) in the deep feature extraction sub-network structure for super-resolution reconstruction of remote sensing images constructed in this invention will be further explained.

[0122] Figure 7 This is a schematic diagram of the parallel hole coordinate attention module structure, which extracts features from the input feature map F from the X and Y branches respectively;

[0123] In the X branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps X_avg and X_max. After merging X_avg and X_max, 3×3 dilated convolutions with dilation rates r=1, 2, 3, and 4 are performed. After merging the output features (concat), 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the X branch. x ;

[0124] In the Y branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps Y_avg and Y_max. After merging Y_avg and Y_max, 3×3 dilated convolutions with dilation rates r=1, 2, 3, and 4 are performed. After merging the output features (Concat), 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the Y branch. x ;

[0125] The feature maps F extracted from the X branch and Y branch x and F y Perform channel dimension merging (Concat) to form attention map A; multiply the attention map with the original input F channel by channel to achieve feature recalibration and obtain the enhanced feature map F';

[0126] Step 2.3.2: Input the feature map output by DSC-3 into the downsampling module Dwon-1, and output a feature map F with half the spatial resolution. d1 (Dimensions are H / 2 × W / 2 × C1); The feature map F d1 The input is fed into the Single Feature Extraction (SFE) module, and the output is the second intermediate feature map F2.

[0127] Step 2.3.3, transfer feature map F d1 The input is fed into the Dwon-2 downsampling module to obtain a feature map F with half the spatial resolution. d2 (Dimensions are H / 4 × W / 4 × C2); The feature map F d2 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F3.

[0128] Step 2.3.4, transfer the feature map F d2 The input is fed into the Dwon-3 downsampling module to obtain a feature map F with half the spatial resolution. d3 (Dimensions are H / 8×W / 8×C2); The feature map F d3 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F4.

[0129] Step 2.3.5: Input feature maps F3 and F4 into the nonlocal module NLB, and use bilinear interpolation to restore the spatial resolution to H / 2×W / 2, obtaining feature map F. 34 ; the feature map F 34 Adding F2 to obtain the feature map F 234 ; the feature map F 234The F1 input is fed into the nonlocal module NLB, and bilinear interpolation is used to restore the spatial resolution to H×W to obtain the feature map F; the feature map F is convolved with the feature map F' by 1×1, and then added to the original feature map X to obtain the mixed feature map X';

[0130] Step 2.4: Construct the depthwise separable convolutional module DSC-4 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0131] Multi-scale convolution kernels (such as 3×3, 5×5, 7×7) are used to independently convolve each channel of the feature map X' output by the hybrid feature extraction module HSEM to extract spatial features of different receptive fields and generate multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused to reduce the number of parameters and computational cost, and the fused feature map is obtained.

[0132] Step 2.5: Construct the depthwise separable convolutional module DSC-5 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel dilated coordinate attention;

[0133] Multi-scale convolution kernels (such as 3×3, 5×5, 7×7) are used to independently convolve each channel of the feature map output by the depthwise separable convolution module DSC-4 to extract spatial features of different receptive fields and generate multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused to reduce the number of parameters and computational cost, resulting in a fused feature map.

[0134] Step 2.6: Construct the skip link module in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention;

[0135] The feature map extracted by the previous FE module is used as the input feature map of the next FE module; the feature map output by the shallow feature extraction sub-network is added to the output feature map of the 16th FE module to output the deep feature map, which is used as the input feature map of the image restoration and reconstruction sub-network.

[0136] Step 3: Construct an image restoration and reconstruction sub-network for super-resolution reconstruction of remote sensing images, restore the deep feature map obtained in Step 2 to the original feature map size, and obtain the final super-resolution reconstructed image;

[0137] Reference Figure 8 The image restoration and reconstruction subnetwork for remote sensing image super-resolution reconstruction constructed in this invention will be further described.

[0138] Figure 8This is a schematic diagram of the image restoration and reconstruction subnetwork, which includes two identical image restoration submodules and one depthwise separable convolutional submodule, DSC-hc.

[0139] Step 3.1, construct the image restoration submodule 1 of the image restoration and reconstruction subnetwork;

[0140] Image restoration submodule 1, whose structure includes: a depthwise separable convolutional module DSC-6, two PixelShufflers, and one LeakyReLU activation function;

[0141] Multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) are used to independently convolve each channel of the feature map output by the deep feature extraction module in step 2, extracting spatial features of different receptive fields and generating multi-scale spatial features. Multi-scale spatial features are then fused using 1×1 pointwise convolutions, reducing the number of parameters and computational cost to obtain a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the separable convolutional module by a factor of 1 each time. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to image restoration submodule 2.

[0142] Step 3.2: Construct the image restoration submodule 2 of the image restoration and reconstruction subnetwork;

[0143] Image restoration submodule 2, whose structure includes: a depthwise separable convolutional module DSC-7, two PixelShufflers, and one LeakyReLU activation function;

[0144] Multi-scale convolutional kernels (such as 3×3, 5×5, and 7×7) are used to independently convolve each channel of the feature map output by the image restoration submodule 1 to extract spatial features from different receptive fields and generate multi-scale spatial features. These multi-scale spatial features are then fused using 1×1 pointwise convolutions, reducing the number of parameters and computational cost to obtain a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the depthwise separable convolutional module by a factor of 1 each time. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to the depthwise separable convolutional module DSC-hc.

[0145] Step 3.3: Construct the depthwise separable convolutional module DSC-hc of the image restoration and reconstruction subnetwork;

[0146] The depthwise separable convolution module DSC-hc uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the image restoration submodule 2, extracting spatial features of different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the reconstructed remote sensing image.

[0147] Step 4: Based on the shallow feature extraction subnetwork constructed in Step 1, the deep feature extraction subnetwork constructed in Step 2, and the image restoration and reconstruction subnetwork constructed in Step 3, construct a remote sensing image super-resolution reconstruction network.

[0148] Step 4.1: Concatenate the shallow feature map extracted in Step 1 and the deep feature map extracted in Step 2.

[0149] Step 4.2: Concatenate the feature map obtained by concatenating steps 1 and 2 with the feature map recovered from the image in step 3.

[0150] Step 5: Generate the training set, validation set, and test set;

[0151] Step 5.1: Obtain the publicly available dataset UC Merced Land-Use Dataset for super-resolution reconstruction of remote sensing images;

[0152] Step 5.2: Divide the dataset from Step 5.1 into training set, validation set and test set according to the proportions.

[0153] Step 6: Use the training set and validation set generated in Step 5 to train the remote sensing image super-resolution reconstruction network constructed in Step 4, and obtain the model weight file.

[0154] Step 6.1: Set training parameters. Each time, randomly and non-repeating remote sensing images are selected from the training set and input into the network. The loss function L uses a content-based loss function (Li). Con ) and perception-based adversarial loss function (L Adv ) and use the sum to optimize the network;

[0155] The content-based loss function (L) Con )as follows:

[0156]

[0157] Where, φ i,j W represents the feature map of the j-th convolution before the i-th max pooling and after passing through the activation layer. i,j H i,j These are the width and height of the feature map, respectively; I LR For low-resolution remote sensing imagery, G θG : Generator, IHR High-resolution remote sensing imagery;

[0158] The perception-based adversarial loss function (L) Adv )as follows:

[0159]

[0160] Among them, D θD (G θG (I LR )) represents the probability that the generated image is real;

[0161] The super-resolution reconstruction loss function (L) is as follows:

[0162]

[0163] Step 6.2: Input the training set and validation set generated in step 5 into step 4 in sequence;

[0164] In the constructed remote sensing image super-resolution reconstruction method, shallow feature extraction is performed in the shallow feature extraction subnetwork to obtain a shallow feature map. Next, the shallow feature map output by the shallow feature extraction subnetwork is input into the deep feature extraction subnetwork to perform deep feature extraction to obtain a deep feature map. Finally, the deep feature map output by the deep feature extraction subnetwork is input into the image restoration and reconstruction subnetwork to obtain the restored remote sensing image.

[0165] Step 7: Using the model weight file obtained from step 6, test the remote sensing images in the test set generated in step 5 to obtain the super-resolution reconstruction result image of the remote sensing images.

[0166] Step 7.1: Input the test set generated in Step 5 into the remote sensing image super-resolution reconstruction network trained in Step 6, and load the remote sensing image super-resolution reconstruction network parameter file "best_loss.pth" trained in Step 7.1.

[0167] Step 7.2: Perform remote sensing image super-resolution reconstruction on the images in the test set to obtain high-resolution remote sensing images;

[0168] Step 7.3: Output the reconstruction results of the remote sensing image super-resolution reconstruction network and save it as an image file in ".png" format.

[0169] Reference Figure 9 The reconstruction results of the remote sensing image super-resolution reconstruction network constructed in this invention will be further described.

[0170] like Figure 9 As shown, where, Figure 9The first and third rows are the reconstructed remote sensing images, while the second and fourth rows are the local magnification effects of different models on the reconstruction results.

[0171] From left to right, the models are: low-resolution remote sensing imagery LR, SRCNN, FSRCNN, SubPixel, VDSR, DRCN, SRGAN, LGCNet, EDSR, ESRGAN, DCM, CTNET, FENET, SRDD, HSENET, TRANSENET, OmniSR, HAUNet, and the proposed method. It can be seen that due to the multi-scale feature extraction module, the model focuses more on the details of ground features. The parallel hole coordinate attention module improves image quality while avoiding the inference speed reduction problem caused by complex attention mechanisms.

Claims

1. A lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention, characterized in that, Includes the following steps: Step 1: Construct a shallow feature extraction subnetwork for super-resolution reconstruction of remote sensing images to extract shallow features from the original remote sensing images; Step 2: Construct a deep feature extraction subnetwork for super-resolution reconstruction of remote sensing images to extract deep features from shallow features; Step 3: Construct an image restoration and reconstruction subnetwork for super-resolution reconstruction of remote sensing images, restore the deep feature map to the original feature map size, and obtain the final super-resolution reconstructed image; Step 4: Construct a super-resolution reconstruction network for remote sensing images; Step 5: Use the publicly available super-resolution reconstruction dataset to generate training, validation, and test sets; Step 6: Train the constructed lightweight super-resolution reconstruction network based on multi-scale feature extraction and parallel hole coordinate attention to obtain the model weight file; Step 7: Using the model weight file obtained from training, test the remote sensing images in the generated test set to obtain the super-resolution reconstruction result image of the remote sensing images.

2. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 1 is as follows: A shallow feature extraction module for raw remote sensing images is constructed. The input is a remote sensing image at the original resolution. The structure of the shallow feature extraction module is: a depthwise separable convolutional module DSC-1, which extracts shallow feature maps, such as local edge and texture features. One LeakyReLU activation function is used to capture the basic and intuitive features of the image, and enhances the non-linear response of edges and textures on the basis of DSC-1 output to obtain a more expressive shallow feature map as input to the deep module; Step 1.1: Construct the depthwise separable convolutional module DSC-1; The shallow feature extraction module includes Depthwise Separable Convolution (DSC). DSC decomposes the standard convolution operation into depthwise convolution and pointwise convolution. It adopts a channel adaptive grouping strategy to dynamically divide the input channels into several groups according to feature similarity. Each group shares a learnable sparse mask to reduce invalid computation. Each input channel is convolved separately, and then the channel information is integrated through pointwise convolution to extract the basic features of the image. The inter-channel weight matrix is ​​generated by 1×1 convolution to suppress redundant channel interactions, such as deep convolutional layers for local edges and texture features; Step 1.2 Construct the LeakyReLU activation function; The LeakyReLU activation function performs a non-linear mapping on the shallow features extracted in step 1.

1. While preserving the weak response of the negative half-axis, it enhances the high-frequency contrast of edges and textures. Based on the DSC-1 output, it enhances the non-linear response of edges and textures, resulting in a more expressive shallow feature map, which serves as the input to the deep feature extraction module. Its expression is as follows: LeakyReLU(x)=max(0,x)+α*min(0,x); Where x is the gray value of the feature map, and α is a positive number less than 1, 0.

01. When the input x is positive, the function output is equal to x; when the input x is negative, the function output is α times x.

3. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 2 is as follows: A deep feature extraction subnetwork with 16 feature extraction modules (FE) is constructed. The FE module structure is as follows: depthwise separable convolution DSC-2, depthwise separable convolution DSC-3, hybrid feature extraction HSEM, depthwise separable convolution DSC-4, and depthwise separable convolution DSC-5. The output of this subnetwork and the output of the shallow feature extraction subnetwork are merged through skip links before output. Step 2.1: Construct the depthwise separable convolutional module DSC-2 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention; The depth-separable convolution module DSC-2 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map extracted by the shallow feature extraction module, extracting spatial features of different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, multi-scale spatial features are fused, reducing the number of parameters and computational cost, and obtaining a fused feature map. Step 2.2: Construct the depthwise separable convolutional module DSC-3 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention; The depthwise separable convolution module DSC-3 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the depthwise separable convolution DSC-2, extracting spatial features of different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, multi-scale spatial features are fused, reducing the number of parameters and computational cost, and obtaining a fused feature map. Step 2.3: Construct the Hybrid Feature Extraction Module (HSEM) in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention. The feature map is obtained by fusing the feature map output by DSC-3 through four branch paths. Step 2.3.1: The feature map output by DSC-3 is directly passed through the single feature extraction module SFE to obtain the first intermediate feature map F1; The Single Feature Extraction (SFE) module processes the input feature map F. i-1 Perform 3×3 convolution to extract the primary feature map F. i-2 The extracted primary feature map F i-2 Input a parallel hole coordinate attention module (AB) to perform inter-channel feature enhancement, resulting in feature map F. ia The enhanced feature map F ia Perform downsampling (3×3 convolution) and output feature map F. a To reduce spatial dimensionality; the extracted primary feature map F i-2 Perform two 3×3 convolutions to extract deep semantic features, resulting in a deep semantic feature map F. i-3 The extracted feature map F a Perform Sigmoid activation function processing and combine it with deep semantic features F i-3 Multiplying the results yields the deep feature map F. i-4 ; to extract deep feature maps F i-4 Perform a 1×1 convolution and then concatenate it with the input feature map F. i-1 The features are added together to restore the spatial resolution, resulting in the output feature map F. i ; The parallel hole coordinate attention module (AB) extracts features from the input feature map F from the X and Y branches respectively; In the X branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps X_avg and X_max. After concatenating X_avg and X_max, 3×3 dilated convolutions with dilation rates r = 1, 2, 3, and 4 are performed. After concatenating the output features, 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the X branch. x ; In the Y branch, global average pooling (AvgPool) and global max pooling (MaxPool) are performed on the input feature map F to obtain feature maps Y_avg and Y_max. After merging Y_avg and Y_max, 3×3 dilated convolutions with dilation rates r = 1, 2, 3, and 4 are performed. After merging the output features (concat), 1×1 convolutions and sigmoid activation functions are performed to obtain the feature map F of the Y branch. x ; The feature maps Fx and Fy extracted from the X and Y branches are concatted along the channel dimension to form an attention map A. The attention map is then multiplied channel by channel with the original input F to achieve feature recalibration, resulting in the enhanced feature map F'. Step 2.3.2: Input the feature map output by DSC-3 into the downsampling module Dwon-1, and output a feature map F with half the spatial resolution. d1 (Dimensions are H / 2×W / 2×C1); Feature map F d1 The input is fed into the Single Feature Extraction (SFE) module, and the output is the second intermediate feature map F2. Step 2.3.3, transfer feature map F d1 The input is fed into the Dwon-2 downsampling module to obtain a feature map F with half the spatial resolution. d2 (Dimensions are H / 4×W / 4×C2); Feature map F d2 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F3. Step 2.3.4, transfer the feature map F d2 The input is fed into the Dwon-3 downsampling module to obtain a feature map F with half the spatial resolution. d3 (Dimensions are H / 8×W / 8×C2); Feature map F d3 The input is fed into the Single Feature Extraction (SFE) module, and the output is the third intermediate feature map F4. Step 2.3.5: Input feature maps F3 and F4 into the Non-local Block (NLB), and use bilinear interpolation to restore the spatial resolution to H / 2×W / 2, obtaining feature map F. 34 ; the feature map F 34 Adding F2 to obtain the feature map F 234 ; the feature map F 234 The F1 input is fed into the nonlocal module NLB, and bilinear interpolation is used to restore the spatial resolution to H×W to obtain the feature map F; the feature map F is convolved with the feature map F' by 1×1, and then added to the original feature map X to obtain the mixed feature map X'; Step 2.4: Construct the depthwise separable convolutional module DSC-4 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention; The depthwise separable convolution module DSC-4 uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map X' output by the hybrid feature extraction module HSEM, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map; Step 2.5: Construct the depthwise separable convolutional module DSC-5 in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel dilated coordinate attention; The depthwise separable convolutional module DSC-5 uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the depthwise separable convolutional module DSC-4, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the fused feature map; Step 2.6: Construct the skip link module in the deep feature extraction subnetwork of the lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention; The skip link module uses the feature map extracted by the previous FE module as the input feature map of the next FE module; it adds the feature map output by the shallow feature extraction sub-network to the output feature map of the 16th FE module to output a deep feature map, which is then used as the input feature map of the image restoration and reconstruction sub-network.

4. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 3 is as follows: Construct an image restoration and reconstruction subnetwork, including two identical image restoration submodules and one depthwise separable convolutional submodule; Step 3.1, construct the image restoration submodule 1 of the image restoration and reconstruction subnetwork; The image restoration submodule 1 includes the following structure: a depthwise separable convolutional module DSC-h, two PixelShufflers, and one LeakyReLU activation function; The depthwise separable convolutional module DSC-h uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the deep feature extraction module in step 2, extracting spatial features from different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused, reducing the number of parameters and computational cost, resulting in a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the depthwise separable convolutional module by a factor of 1 each time. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to the image restoration submodule 2. Step 3.2: Construct the image restoration submodule 2 of the image restoration and reconstruction subnetwork; The image restoration submodule 2 includes the following structure: a depthwise separable convolutional module DSC-h, two PixelShufflers, and one LeakyReLU activation function; The depthwise separable convolutional module DSC-h uses multi-scale convolutional kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the image restoration submodule 1, extracting spatial features from different receptive fields and generating multi-scale spatial features. By using 1×1 pointwise convolution, the multi-scale spatial features are fused, reducing the number of parameters and computational cost, resulting in a fused feature map. Two consecutive PixelShuffle layers are used for progressive upsampling, increasing the image resolution of the feature map output by the depthwise separable convolutional module by a factor of 1 in each stage. The LeakyReLU activation function performs a non-linear mapping on the feature map output after PixelShuffle upsampling, serving as the input to the depthwise separable convolutional module DSC-hc. Step 3.3: Construct the depthwise separable convolutional module DSC-hc of the image restoration and reconstruction subnetwork; The depth-separable convolution module DSC-hc uses multi-scale convolution kernels (such as 3×3, 5×5, 7×7) to independently convolve each channel of the feature map output by the image restoration submodule 2, extracting spatial features of different receptive fields and generating multi-scale spatial features; by using 1×1 pointwise convolution, the multi-scale spatial features are fused, thereby reducing the number of parameters and computational cost, and obtaining the reconstructed remote sensing image.

5. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 4 is as follows: Step 4.1: Concatenate the shallow feature map extracted in Step 1 and the deep feature map extracted in Step 2. Step 4.2: Concatenate the feature map obtained by concatenating steps 1 and 2 with the feature map recovered from the image in step 3.

6. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 5 is as follows: Step 5.1: Obtain the publicly available dataset UC Merced Land-Use Dataset for super-resolution reconstruction of remote sensing images; Step 5.2: Divide the dataset from Step 5.1 into training set, validation set and test set according to the proportions.

7. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 6 is as follows: Step 6.1: Set training parameters. Each time, randomly and non-repeating remote sensing images are selected from the training set and input into the network. The loss function L uses a content-based loss function (Li). Con ) and perception-based adversarial loss function (L Adv ) and use the sum to optimize the network; The content-based loss function (L) Con )as follows: in, W represents the feature map of the j-th convolution before the i-th max pooling and after passing through the activation layer; i,j H i,j These are the width and height of the feature map, respectively; I LR For low-resolution remote sensing imagery, G θG : Generator, I HR High-resolution remote sensing imagery; The perception-based adversarial loss function (L) Adv )as follows: in, This represents the probability that the generated image is real; The super-resolution reconstruction loss function (L) is as follows: L=L Com +L Adv ; Step 6.2: Input the training set and validation set generated in Step 5 into the shallow feature extraction subnetwork of the remote sensing image super-resolution reconstruction method constructed in Step 4 to perform shallow feature extraction and obtain a shallow feature map; then, input the shallow feature map output by the shallow feature extraction subnetwork into the deep feature extraction subnetwork to perform deep feature extraction and obtain a deep feature map; finally, input the deep feature map output by the deep feature extraction subnetwork into the image restoration and reconstruction subnetwork to obtain the restored remote sensing image.

8. The lightweight super-resolution reconstruction method based on multi-scale feature extraction and parallel hole coordinate attention as described in claim 1, characterized in that, The specific method for step 7 is as follows: Step 7.1: Input the test set generated in Step 5 into the remote sensing image super-resolution reconstruction network trained in Step 6, and load the remote sensing image super-resolution reconstruction network parameter file "best_loss.pth" trained in Step 7.

1. Step 7.2: Perform remote sensing image super-resolution reconstruction on the images in the test set to obtain high-resolution remote sensing images; Step 7.3: Output the detection results of the remote sensing image super-resolution reconstruction network and save them as an image file in ".png" format.

Citation Information

Patent Citations

  • A Single Remote Sensing Image Super-Resolution Method Based on Deep Learning

    CN116342392B

  • Remote sensing image super-resolution reconstruction method based on mixed multi-scale attention group

    CN120259076A

Cited By

  • Remote sensing image deep learning super-resolution reconstruction method

    CN121073781A

  • Multi-scale progressive surface temperature fusion downscaling method and device

    CN121121533A

  • Multi-scale progressive land surface temperature fusion downscaling method and device

    CN121121533B

  • Super-resolution remote sensing image target detection method based on multi-modal fusion

    CN121459181A

  • A multi-modal fusion-based super-resolution remote sensing image target detection method

    CN121459181B