Double-domain lightweight reconstruction method for remote sensing image super-resolution reconstruction
Through the dual-domain lightweight reconstruction method, the dual-domain reconstruction network DDRN combined with spatial domain and frequency domain features, the problems of high model complexity and feature redundancy in remote sensing image super-resolution reconstruction are solved, and efficient image reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510116635.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing remote sensing image super-resolution reconstruction methods have problems such as excessive model parameter quantity and computational complexity, insufficient context information capture, and feature redundancy, making it difficult to effectively capture image details and global information.
A two-domain lightweight reconstruction method is proposed. By constructing a two-domain reconstruction network DDRN, the blueprint can be used to separate convolution and multiple two-domain recovery modules to deeply extract features, combining spatial domain and frequency domain features to reduce the amount of model parameters and calculation complexity.
It significantly reduces the complexity of the model and inference time, maintains subjective visual effects and quantitative indicators, and improves the accuracy and efficiency of image reconstruction.
Smart Images

Figure CN120070178A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image super-resolution reconstruction, and relates to a dual-domain lightweight reconstruction method for remote sensing image super-resolution reconstruction. Background Art
[0002] High-resolution remote sensing images play a crucial role in multiple fields, including land use monitoring, environmental change detection, post-disaster assessment, and target recognition. Due to the limitations of remote sensing equipment, directly obtaining high-resolution images is often restricted by hardware performance and cost. Therefore, super-resolution (SR) technology for reconstructing high-resolution (HR) images from low-resolution (LR) images has become an important solution. SR technology can not only enhance the details and textures of images but also reduce the cost of obtaining high-resolution images, thereby improving the analysis efficiency of remote sensing images. Traditional image super-resolution methods can generally be divided into three categories: interpolation methods, reconstruction methods, and learning methods. Interpolation methods such as bicubic interpolation generate high-resolution images from low-resolution images through simple mathematical formulas. Although they are computationally simple and easy to implement, since they rely on local information, they cannot effectively recover the high-frequency information of images, resulting in blurred image edges and lost details. Reconstruction methods estimate high-resolution images by optimizing problems and often use iterative solution strategies to recover image details, but they have a large computational overhead and are difficult to handle complex degradation models. In recent years, with the rapid development of deep learning technology, learning-based image super-resolution methods have gradually become mainstream. Among them, methods based on deep neural networks (DNNs) have demonstrated excellent image reconstruction capabilities. In particular, through multi-level feature learning, they can capture more high-frequency information and significantly improve the super-resolution reconstruction effect.
[0003] In recent years, deep convolutional neural networks (DCNNs) have been widely applied to image super-resolution reconstruction tasks. Most DCNN-based methods use convolutional layers, attention mechanisms, and residual connections to extract local information in the spatial dimension and global information in the channel dimension in the spatial domain. There are also several methods based on large kernel convolution operations and Transformers to capture long-range spatial dependencies. At the same time, methods based on generative adversarial networks are also applied to image super-resolution reconstruction tasks, but the problem of image distortion during recovery also occurs. In addition to recovering images in the spatial domain, there are also a small number of methods for image super-resolution reconstruction in the frequency domain. However, these methods do not fully consider the combination of frequency domain features and spatial features, resulting in poor recovery effects.
[0004] Guo et al. proposed a novel dense super resolution generative adversarial network (NDSRGAN) for real aerial image super-resolution reconstruction, using a multi-level dense network in the generative network to connect the dense connection blocks in the residual dense blocks, and using a matrix mean discriminator in the discriminative network to locally discriminate the generated images. Huan et al. designed a residual multi-scale block (RMSB) and a residual multi-scale dilation block (RMSDB) to extract shallow features and deep features respectively with fewer parameters, and proposed a feature-refinement fusion (FRF) module to update the shallow features with the deep features to achieve better reconstruction results. Haut et al. combined deep residual learning and channel attention mechanism to propose deep residual channel attention (DRCA) for learning the weights between different channels, thus enhancing the flexibility of the network in processing different features. Meng et al. proposed a super-resolution dense-sampling residual attention network (SRDSRGAN), whose encoder-decoder modules of the generator and discriminator adopt dense residual blocks and convolutional layers respectively, and both use serial channels and spatial attention in the image reconstruction process to capture global and local information. A chained training method was introduced during the training process for model training of large-scale super-resolution tasks. Traditional meteorological data is usually limited by spatial resolution. Cheng et al. addressed the spatial downsampling problem of downward surface shortwave radiation, and combined residual blocks and channel attention to propose a deep neural network that can learn the mapping relationship between low-resolution downward surface shortwave radiation (DSSR) data and high-resolution DSSR data. Zhang et al. proposed a frequency distribution integration module to obtain local-global frequency distribution information, thereby promoting the recognition of remote sensing scenes, and at the same time introduced an adaptive feature refinement module to filter redundant features generated by domain differences.
[0005] The above-mentioned existing technologies face some challenges. First, the number of model parameters and computational complexity are too large: Many methods improve performance by deepening the number of network layers or expanding the number of channels, but this increases the difficulty of network training and deployment, especially it is difficult to apply on resource-constrained remote sensing devices. Second, the capture of context information is insufficient: Existing methods often ignore the complementary relationship between spatial domain and frequency domain features, and it is difficult to effectively extract and fuse global and local information, thus restricting the improvement of the reconstruction effect. Finally, the problem of feature redundancy: Traditional deep networks often improve performance by stacking features of different layers, but do not fully consider the differences between features at different levels, which easily leads to information redundancy, thereby affecting the reconstruction effect and efficiency. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to solve the problems of effectively capturing the detailed information in remote sensing images, fully mining the global information and local details in the images, and reducing the number of parameters and computational complexity in the super-resolution reconstruction of remote sensing images, and provide a dual-domain lightweight reconstruction method for the super-resolution reconstruction of remote sensing images.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A dual-domain lightweight reconstruction method for the super-resolution reconstruction of remote sensing images, comprising the following steps:
[0009] Construct a dual-domain reconstruction network DDRN. In the DDRN network, first use a Blueprint Separation Convolution (BSConv) to obtain shallow features F LR ∈R 3×H×W from an input remote sensing image I in ∈R C×H×W , then deeply extract the shallow features through multiple Dual-domain Restoration Modules (DRMs), splice the output of each DRM in the channel dimension and input it into a convolutional layer composed of a 1×1 ordinary convolution, a GELU activation function, and a BSConv for feature fusion to obtain F out ; finally add F out and F in , and obtain a high-resolution image I HR ∈R 3×sH×sW through upsampling, where s is the upsampling factor;
[0010] Train and test the DDRN network using remote sensing images to obtain a dual-domain lightweight reconstruction model for the super-resolution reconstruction of remote sensing images;
[0011] Use the dual-domain lightweight reconstruction model for super-resolution reconstruction of remote sensing images to perform super-resolution reconstruction on remote sensing images.
[0012] Furthermore, the overall calculation formula of the DDRN is as follows:
[0013] F in = BSConv I LR
[0014]
[0015] I HR = Upsample F in + F out
[0016] The BSConv is composed of a 1×1 ordinary convolution and a 3×3 depthwise separable convolution. DRM(·) refers to the output operation of the input features through the dual-domain restoration module. is the concatenation operation along the channel dimension. Conv 1×1 is an ordinary convolution with a convolution kernel size of 1. Upsample(·) is composed of PixcleShuffle and a 3×3 ordinary convolution.
[0017] Furthermore, the dual-domain restoration module DRM is composed of a dual-domain feature extraction module DFEM and a dual-domain pixel attention module DPAM connected. The dual-domain feature extraction module DFEM is used to co-extract and fuse multi-level features in the spatial domain and the frequency domain to capture the detailed information in the remote sensing image. The dual-domain pixel attention module DPAM is used to mine the global information and local details in the image and re-distribute the weights of the pixels.
[0018] Furthermore, the dual-domain feature extraction module DFEM includes an input layer, a convolution layer, a receptive field expansion block RFEB, a frequency domain residual block FRB, and BSConv;
[0019] The input layer is used to input the input data of the dual-domain feature extraction module
[0020] The convolution layer is composed of a 1×1 ordinary convolution and is responsible for extracting local features in the space;
[0021] The RFEB is responsible for expanding the receptive field to obtain richer image information. By splitting the input features into two equal parts along the channel dimension, where half of the features pass through a 5×5 depthwise separable convolution, and the other half of the features are retained. Then, they are concatenated along the channel dimension and fused through a 1×1 ordinary convolution. Finally, the features are output through the GELU activation function The definition formula of RFEB is as follows:
[0022]
[0023] Among them, Split(·) is an operation that divides the feature into two parts along the channel dimension, and DWConv 5×5 is a depthwise separable convolution with a convolution kernel size of 5, and Identity(·) is an operation that retains the original feature;
[0024] The frequency-domain residual block FRB is responsible for extracting global information in the frequency domain; the definition formula of the FRB is as follows:
[0025]
[0026] where FFT(·) and IFFT(·) represent the Fourier transform and the inverse Fourier transform respectively;
[0027] After each expansion of the receptive field, the RFEB outputs deeper features Then further are respectively input into the convolutional layer, RFEB, and FRB; after three expansions of the receptive field, for the outputs of all previous convolutional layers and FRBs, as well as the output after passing through BSConv is concatenated along the channel dimension, and feature fusion is performed through a 1×1 ordinary convolution to obtain the output feature
[0028] Furthermore, the dual-domain pixel attention module DPAM includes a spatial attention block SAB and a frequency-domain pixel block FPB;
[0029] The input feature is first input into the SAB. First, extracts the maximum and average values of the pixels, and then the two are concatenated along the channel dimension and successively input into a 7×7 ordinary convolution and a Sigmoid activation function to obtain the spatial weight W SAB ∈R 1×H×W , and finally through and W SAB are multiplied element-wise to obtain the output of the SAB The definition formula of the SAB is as follows:
[0030]
[0031] where Sigmoid(·) is the Sigmoid activation function, Conv 7×7 is an ordinary convolution with a convolution kernel size of 7, Max(·) and Mean(·) are max pooling and average pooling in the channel dimension, and ⊙ is the element-wise multiplication operation;
[0032] After passing through the SAB, then Input into the FPB, and perform global pixel weight assignment in the frequency domain. First, Convert it to the frequency domain through Fourier transform, and obtain the frequency domain weight W through the frequency domain residual convolution layer FPB ∈R C×H×W , and finally through Perform element-wise multiplication with W FPB to obtain the output of SAB The definition formula of FPB is as follows:
[0033]
[0034] Among them, FreConv(·) includes three 1×1 ordinary convolutions and two LeakyRelu activation functions.
[0035] The beneficial effects of the present invention are as follows: The method proposed by the present invention significantly reduces the model complexity and inference time while maintaining the subjective visual effect and quantization index close to the current mainstream algorithms, demonstrating superior lightweight advantages and application potential.
[0036] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0037] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0038] Figure 1 is the DDRN network structure;
[0039] Figure 2 is the structure diagram of the dual-domain feature extraction module;
[0040] Figure 3 In (a) is the structure diagram of the receptive field expansion block, and (b) is the structure diagram of the frequency domain residual block;
[0041] Figure 4 is the structure diagram of the dual-domain pixel attention module;
[0042] Figure 5 is the subjective visual graph of the comparative experiment. Detailed Embodiment
[0043] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0044] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0045] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0046] The present invention provides a lightweight reconstruction method in the dual domain for remote sensing image super-resolution reconstruction. Due to the limitations of imaging devices and communication bandwidth, and being easily interfered by external noise and other factors, the restoration of remote sensing images is very challenging. Therefore, for effective remote sensing image restoration, the feature extraction in the frequency domain and the spatial domain is crucial. Therefore, the present invention simultaneously considers the local and global feature information in the spatial domain and the frequency domain, and proposes DDRN, as Figure 1 shown.
[0047] DDRN consists of a shallow feature extraction layer, a deep feature extraction layer, and a graph reconstruction layer. Specifically, the present invention first uses a Blueprint Separation Convolution (BSConv) to obtain shallow features F LR ∈R 3×H×W (where H and W are the height and width of the input LR remote sensing image respectively) from an input remote sensing image I in ∈R C×H×W (C represents the number of channels of the current feature). Then, the shallow features are deeply extracted through multiple Dual-domain Restoration Modules (DRMs), and the output of each DRM Concatenate in the channel dimension and input it into a convolutional layer composed of a 1×1 ordinary convolution, a GELU activation function, and a BSConv for feature fusion to obtain F out Finally, add it to F in and perform upsampling to obtain the high-resolution image I HR ∈R 3×sH×sW (where s is the upsampling factor). The overall network definition is shown by the following formula:
[0048] F in = BSConv I LR (1)
[0049]
[0050] I HR = Upsample F in + F out (5)
[0051] where BSConv(·) is composed of a 1×1 ordinary convolution and a 3×3 depthwise separable convolution, DRM(·) refers to the operation of the input features output through the dual-domain restoration module, is the concatenation operation in the channel dimension, Conv 1×1 is an ordinary convolution with a convolution kernel size of 1, and Upsample(·) is composed of PixcleShuffle and a 3×3 ordinary convolution.
[0052] The dual-domain restoration module (DRM) includes two major modules: the dual-domain feature extraction module and the dual-domain pixel attention module.
[0053] Comprehensive and in-depth feature extraction is crucial for restoring the texture details of local regions and global structures. Most of the previous feature extraction processes for remote sensing image super-resolution tasks are only performed in the spatial domain. However, the feature extraction method only in the spatial domain cannot accurately mine the potential local and global information in remote sensing images, and the features in the frequency domain can be used as an effective supplement. And in the feature extraction process, the size of the receptive field directly affects the network's ability to capture local and global information, which is crucial for the performance and application effect of the model. Therefore, in the proposed DDRN network, the present invention first proposes a dual-domain feature extraction module (DFEM), as Figure 2 shown.
[0054] Specifically, the present invention inputs into a convolutional layer, a receptive field expansion block (RFEB), and a frequency domain residual block (FRB) respectively. The convolutional layer is composed of a 1×1 ordinary convolution and is responsible for extracting local features in space. As Figure 3The RFEB shown in (a) is responsible for expanding the receptive field to obtain richer image information by splitting the input features into two equal parts along the channel dimension. Half of the features are passed through a 5×5 depthwise separable convolution, and the other half of the features are retained. Then, they are concatenated along the channel dimension and fused through a 1×1 ordinary convolution. Finally, the features are output through the GELU activation function The definition formula of RFEB is as follows:
[0055]
[0056] Among them, Split(·) is the operation of splitting the features into two equal parts along the channel dimension, DWConv 5×5 is a depthwise separable convolution with a kernel size of 5, and Identity(·) is the operation of retaining the original features.
[0057] As Figure 3 shown in (b), the FRB is responsible for extracting global information in the frequency domain. After each expansion of the receptive field, the RFEB will output deeper features which are then further input into the convolutional layer, RFEB, and FRB respectively. After three expansions of the receptive field, the present invention will concatenate the outputs of all previous convolutional layers and FRBs, as well as the output after passing through BSConv along the channel dimension, and perform feature fusion through a 1×1 ordinary convolution to obtain the output features The definition formula of FRB is as follows:
[0058]
[0059] Among them, FFT(·) and IFFT(·) represent the Fourier transform and the inverse Fourier transform respectively.
[0060] After deep feature extraction, the present invention needs to correct the obtained features and assign more weights to the information that needs to be enhanced. Therefore, the present invention designs a dual-domain pixel attention module (DPAM) for redistributing the weights of pixels. As Figure 4 shown, the DPAM includes a spatial attention block (SAB) and a frequency-domain pixel block (FPB). Specifically, the input features are first input into the SAB. First, the maximum and average values of the pixels are extracted, and then they are concatenated along the channel dimension and successively input into a 7×7 ordinary convolution and a Sigmoid activation function to obtain the spatial weight W SAB ∈R 1×H×W . Finally, through element-wise multiplication with W SAB , the output of the SAB is obtained The SAB definition formula is as follows:
[0061]
[0062] Where Sigmoid(·) is the Sigmoid activation function, Conv 7×7 is a normal convolution with a convolution kernel size of 7, Max(·) and Mean(·) are max pooling and average pooling in the channel dimension, and ⊙ is an element-wise multiplication operation. After passing through SAB, the is input into the FPB to perform global pixel weight allocation in the frequency domain. First, the present invention is transformed into the frequency domain through Fourier transform, and the frequency domain weight W FPB ∈R C×H×W is obtained through the frequency domain residual convolution layer. Finally, through and W FPB are multiplied element-wise to obtain the output of SAB The FPB definition formula is as follows:
[0063]
[0064]
[0065] For fair comparison, the present invention is trained and tested on the UC Merced dataset and the AID dataset respectively. Among them, UC Merced is a popular land use dataset proposed by the National Atlas of the United States Geological Survey in 2010. The dataset includes a total of 21 categories, such as agriculture, aircraft, beach, buildings, bushes, runway and tennis court. Each category has 100 images with a size of 256×256. The AID dataset is a remote sensing image dataset, containing 30 types of scene images. Each type has about 220 to 420 images with a size of 600×600, totaling 10,000 images. And 1000 images are randomly selected from the AID dataset as the AID test set.
[0066] The present invention uses Peak Signal to Noise Ratio (PSNR), Structural Similarity Index (SSIM), number of model parameters (Parameters, Params) and number of floating point operations (Floating Point of Operations, FLOPs) as evaluation metrics to evaluate the performance of DDRN.
[0067] DDRN contains 8 DRMs, its feature channels are 56, and it is trained by the Adam optimizer (β 1 =0.9, β 2= 0.9). The initial learning rate is set to 5×10 -3 , and then it is decreased to 1×10 -7 through the cosine annealing strategy. The training image patch size is 64×64, and the batch size is 16. The training images are augmented by random rotation and flipping. The loss function for training the DDRN is:
[0068] L total = L content + L fft (9)
[0069] where L content is the L1 loss between the high-resolution image I HR reconstructed by the network and the reference clear image I GT . L fft is the L1 loss in the frequency domain between the high-resolution image I HR reconstructed by the network and the reference clear image I GT , and λ = 0.1. The present invention realizes the training and testing of the model on a dual-card Nvidia RTX 2080Ti and PyTorch platform.
[0070] The present invention statistically analyzes the quantitative analysis of the latest state-of-the-art remote sensing image SR methods on the UC Merced dataset and the AID dataset in terms of average PSNR, average SSIM, Parms, FLOPs, and inference time. As can be seen from Table 1, when the model is trained on the UC Merced dataset and the AID dataset, in most cases, the proposed DDRN of the present invention, in the tests with scaling factors of ×2 and ×4, with similar performance, ensures a smaller number of model parameters and a faster inference time. Specifically, when the scaling factor is ×2, the Parms of DDRN is 4.6M lower than that of the sub-optimal SRAGAN, the inference speed is 8ms faster, the number of floating-point operations is 19G lower, the PSNR and SSIM on the UC Merced dataset are increased by 0.06dB and 0.002 respectively, and the PSNR and SSIM on the AID dataset are increased by 0.04dB and 0.002 respectively; its inference speed is 1ms faster than that of the sub-optimal NDSRGAN, the number of model parameters is 17.2M lower, the number of floating-point operations is 71G lower, the PSNR and SSIM on the UC Merced dataset are increased by 4.9dB and 0.127 respectively, and the PSNR and SSIM on the AID dataset are increased by 3.92dB and 0.109 respectively. When the scaling factor is ×4, the Parms of DDRN is 4.8M lower than that of the sub-optimal SRAGAN, the inference speed is 9ms faster, and the PSNR and SSIM on the AID dataset are increased by 1.87dB and 0.022 respectively; its inference speed is only 4ms slower than that of the optimal NDSRGAN, the PSNR and SSIM on the UC Merced dataset are increased by 2.75dB and 0.026 respectively, and the PSNR and SSIM on the AID dataset are increased by 3.19dB and 0.035 respectively.
[0071] Table 1
[0072]
[0073] Figure 5 shows the comparison of the subjective visual effects between the proposed DDRN of the present invention and the latest state-of-the-art remote sensing image SR methods on the UC Merced dataset and the AID dataset. As Figure 5As shown, the images reconstructed by the DDRN proposed in the present invention contain clearer and more accurate edges and details. For example, in the "Airplane" image, the DDRN proposed in the present invention can accurately reconstruct the tail part of the airplane, while the tail reconstructed by the method carries a certain degree of high-frequency noise and blurriness. At the same time, the reconstruction results of other methods are also too smooth, losing a lot of texture details. In the "River" image, the DDRN proposed in the present invention can reconstruct the texture of the woods on both sides of the riverbank, while other methods cannot remove the noise and blurriness in the woods part or cause distortion in the woods part. In the "Bridge" image, the DDRN proposed in the present invention can accurately restore the cars and dividing lines in the bridge. However, there is still obvious noise in the road restored by other methods. At the same time, the cars restored by other methods are also too smooth, losing the detailed parts of the cars. In the "Farmland" image, the DDRN of the present invention can reconstruct the houses and roads near the farmland, while other methods cannot remove the noise in the farmland or lose the details of the houses. In summary, the DDRN proposed in the present invention can reconstruct more accurate structures and clearer textures (see Figure 5 and the corresponding enlarged areas). Table 1 and Figure 5 demonstrate the advantages of the DDRN proposed in the present invention from the aspects of qualitative evaluation and quantitative indicators respectively.
[0074] In the above embodiments, the reference to "this embodiment" in the specification means that the specific features, structures or characteristics described in connection with the embodiments are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.
[0075] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many substitutions, modifications and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) can be used in the embodiments discussed. The embodiments of the present invention are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims.
[0076] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the methods in this embodiment.
[0077] This embodiment also provides an electronic terminal, including: a processor and a memory;
[0078] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any one of the methods in this embodiment.
[0079] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to a computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0080] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program so that the electronic terminal executes each step of the above method.
[0081] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0082] The above-mentioned processor can be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it can also be a digital signal processor (Digital Signal Processing, abbreviated as DSP), an application specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field-programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0083] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on.
[0084] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A dual-domain lightweight reconstruction method for super-resolution reconstruction of remote sensing images, characterized in that: The following steps are involved: Construct a dual domain reconstruction network DDRN. In the DDRN network, a blueprint separable convolution BSConv is first used to obtain an input remote sensing image I LR ∈R 3×H×W Get the shallow feature F in ∈R C×H×W Then, multiple dual-domain restoration modules DRM are used to deeply extract shallow features, and each DRM output The concatenation is performed in the channel dimension and input into the convolution layer consisting of 1×1 ordinary convolution, GELU activation function and BSConv for feature fusion to obtain F out ; Finally, F out With F in Add and upsample to obtain a high-resolution image I HR ∈R 3×sH×sW , where s is the upsampling factor; Using remote sensing images to train and test the DDRN network, a dual-domain lightweight reconstruction model for super-resolution reconstruction of remote sensing images is obtained; The remote sensing image is super-resolution reconstructed using the dual-domain lightweight reconstruction model for super-resolution reconstruction of remote sensing images.
2. The dual-domain lightweight reconstruction method for super-resolution reconstruction of remote sensing images according to claim 1, characterized in that: The overall calculation formula of DDRN is as follows: The BSConv consists of a 1×1 normal convolution and a 3×3 depthwise separable convolution. DRM(·) refers to the operation of input features being output through a dual domain restoration module. It is a concatenation operation based on the channel dimension. Conv 1×1 It is a normal convolution with a kernel size of 1, and Upsample(·) is composed of PixcleShuffle and 3×3 normal convolution.
3. The dual-domain lightweight reconstruction method for super-resolution reconstruction of remote sensing images according to claim 1, characterized in that: The dual-domain restoration module DRM is composed of a dual-domain feature extraction module DFEM and a dual-domain pixel attention module DPAM; the dual-domain feature extraction module DFEM is used to collaboratively extract and fuse multi-level features in the spatial domain and frequency domain to capture detailed information in the remote sensing image; the dual-domain pixel attention module DPAM is used to mine the global information and local details in the image and redistribute the weights of pixels.
4. The dual-domain lightweight reconstruction method for super-resolution reconstruction of remote sensing images according to claim 1, characterized in that: The dual-domain feature extraction module DFEM includes an input layer, a convolutional layer, a receptive field expansion block RFEB, a frequency domain residual block FRB and BSConv; The input layer is used to input the input data of the dual-domain feature extraction module The convolution layer consists of a 1×1 ordinary convolution, which is responsible for extracting local features in space; The RFEB is responsible for expanding the receptive field to obtain richer image information by The channel dimension is divided into two parts, half of which is separated by 5×5 depth-wise convolution, and the other half is retained. Then they are concatenated according to the channel dimension and fused by 1×1 ordinary convolution, and finally the features are output through the GELU activation function. The RFEB definition formula is as follows: Among them, Split(·) is the operation of splitting the feature into two according to the channel dimension, and DWConv 5×5 It is a depth-wise separable convolution with a kernel size of 5, and Identity(·) is an operation that preserves the original features; The frequency domain residual block FRB is responsible for extracting global information in the frequency domain; the FRB definition formula is as follows: Where FFT(·) and IFFT(·) represent Fourier transform and inverse Fourier transform, respectively; After each expansion of the receptive field, RFEB outputs deeper features Then further are input into the convolutional layer, RFEB and FRB respectively; after three receptive field expansions, the outputs of all previous convolutional layers and FRB, and The output of BSConv is concatenated according to the channel dimension, and the output features are obtained by feature fusion through 1×1 ordinary convolution.
5. The dual-domain lightweight reconstruction method for super-resolution reconstruction of remote sensing images according to claim 1, characterized in that: The dual-domain pixel attention module DPAM includes a spatial attention block SAB and a frequency domain pixel block FPB; Input Features First input into SAB, first The maximum and average values of pixels are extracted from the network, and then the two are concatenated according to the channel dimension and successively input into the 7×7 ordinary convolution and Sigmoid activation function to obtain the spatial weight W. SAB ∈R 1 ×H×W , and finally through With W SAB Perform element-by-element multiplication to get the output of SAB The SAB definition formula is as follows: Where Sigmoid(·) is the Sigmoid activation function, Conv 7×7 is a normal convolution with a kernel size of 7, Max(·) and Mean(·) are the maximum pooling and average pooling in the channel dimension, and ⊙ is an element-by-element multiplication operation; After SAB, Input to FPB, and perform global pixel weight distribution in the frequency domain. First, It is converted to the frequency domain through Fourier transform, and the frequency domain weight W is obtained through the frequency domain residual convolution layer. FPB ∈R C×H×W , and finally through With W FPB Perform element-by-element multiplication to get the output of SAB The FPB definition formula is as follows: FreConv(·) contains three 1×1 ordinary convolutions and two LeakyRelu activation functions.
Citation Information
Cited By
Deep learning model remote sensing image fusion method based on dense residual error
CN120410890A
Cross-domain applicable lightweight super-resolution reconstruction method
CN121235904A
A cross-domain applicable lightweight super-resolution reconstruction method
CN121235904B
Remote sensing image super-resolution reconstruction method, computer device and readable storage medium
CN122550363A