A method and device for remote sensing image super-resolution based on a generative adversarial network with fusion attention and frequency domain enhancement
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-10
- Publication Date
- 2026-08-11
AI Technical Summary
本发明针对遥感图像中多尺度、多方向特征提取不平衡及空间-频域信息利用不足等问题,提出一种基于融合注意力与频域增强的生成对抗网络遥感图像超分辨率方法及装置
1、本发明提出了一种空间域特征增强模块。该网络集成了动态卷积模块、多尺度残差模块、Shuffle Attention注意力模块以及通道-空间协同注意力模块。它能够自适应地捕获遥感图像中从宏观结构到微观纹理的多尺度特征,并通过注意力机制实现跨通道与跨空间维度的信息融合,有效解决了多尺度、多方向特征提取不平衡及空间-频域信息利用不足等问题,有效增强了网络对高频边缘和复杂纹理细节的特征表征与恢复能力。实现了遥感图像从宏观结构到微观纹理的多尺度特征自适应捕获及跨通道与跨空间维度的信息融合。
Smart Images

Figure CN122550358A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image super-resolution reconstruction, specifically a method and apparatus for super-resolution of remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement. Background Technology
[0002] Super-resolution reconstruction of remote sensing images, a popular research direction at the intersection of computer vision and remote sensing applications, aims to reconstruct high-resolution images with rich details from low-resolution remote sensing images using image processing and deep learning techniques, providing support for tasks such as remote sensing interpretation. The application of deep convolutional neural networks has significantly improved performance in this field compared to traditional methods, but core challenges remain in practical applications: imbalanced multi-scale and multi-directional feature extraction, making it difficult for existing methods to achieve a balanced capture of both macroscopic structure and microscopic texture, resulting in lost details and insufficient representation of directional features in the reconstructed image; and insufficient utilization of spatial and frequency domain information, with most methods limited to single spatial domain processing and lacking an effective synergistic mechanism between the two, failing to fully exploit deep image features and causing loss of high-frequency details and structural blurring. Existing technologies have failed to address these problems specifically, making it difficult to meet the demands of high-precision applications.
[0003] The Chinese patent announcement number is "CN120339072B", entitled "A Low-Quality Image Super-Resolution Reconstruction Method Based on Generative Adversarial Networks". This method first preprocesses the remote sensing image, fixing the input image size to meet network input requirements. Next, it extracts spatial domain features from the low-resolution image using a single-scale or simple multi-scale convolution module. Then, the extracted features are processed by dimensionality reduction or cascading before being fed into the reconstruction module. Finally, an upsampling module amplifies the feature map, outputting a high-resolution reconstructed image. However, this method is limited to single-space domain feature processing, lacking a balanced capture mechanism for multi-scale and multi-directional features, and failing to establish an effective collaborative link between the spatial and frequency domains. This results in an inability to fully extract deep high-frequency information from the image, leading to problems such as missing details in the reconstructed image, insufficient directional feature representation, and structural ambiguity. Consequently, it struggles to meet the practical needs of high-precision remote sensing interpretation. Therefore, addressing the imbalance in multi-scale and multi-directional feature extraction and the insufficient utilization of spatial-frequency domain information is a key technical bottleneck that this invention urgently needs to overcome. Summary of the Invention
[0004] (a) Technical problems to be solved This invention addresses the imbalance in multi-scale and multi-directional feature extraction and insufficient utilization of spatial-frequency domain information in remote sensing images. It proposes a generative adversarial network-based super-resolution method and apparatus for remote sensing images, incorporating attention and frequency domain enhancement. This method integrates a dynamic convolution module, a multi-scale residual module, a shuffle attention module, a channel-spatial collaborative attention module, and a frequency-gated feedforward network. Through multi-scale feature extraction, lightweight attention interaction, and frequency domain feature enhancement, it achieves collaborative representation in both spatial and frequency domains, improving the accuracy and efficiency of subsequent remote sensing interpretation tasks.
[0005] (II) Technical Solution To achieve the above objectives, the present invention specifically adopts the following technical solution: A generative adversarial network (GAN) method for super-resolution of remote sensing images based on the fusion of attention and frequency domain enhancement includes the following steps: Step 1, Constructing the network model: Construct an adaptive feature enhancement network, which includes a spatial domain feature enhancement module and a frequency domain enhancement module; Step 2, Prepare the dataset: Divide the training set and test set using the first remote sensing image dataset and preprocess them; fine-tune the model using the second remote sensing image dataset; Step 3, train the network model: input the dataset prepared in step 2 into the network model built in step 1 for training; Step 4, fine-tuning the model: retrain and fine-tune the network model using the second remote sensing image dataset to obtain the final model; Step 5, Save the model: Solidify the parameters of the final model and save the model.
[0006] The above-mentioned method for super-resolution remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement is characterized by: In step 1, the spatial domain feature enhancement module includes a shallow feature extraction module and a deep feature extraction module; the frequency domain enhancement module includes a frequency-gated feedforward network.
[0007] In step 1, the adaptive feature enhancement network is composed of a spatial domain feature enhancement module and a frequency domain enhancement module. The spatial domain feature enhancement module includes a dynamic convolution module, a multi-scale residual module, a ShuffleAttention module, and a channel-space collaborative attention module. Specifically: the dynamic convolution module adaptively adjusts the convolution kernel parameters using dynamic convolution components to enhance the multi-directional representation capability of features; the multi-scale residual module captures multi-scale features of the data using multi-scale convolution; the ShuffleAttention module achieves cross-channel information interaction and feature weighting using channel shuffling and attention calculation components; the channel-space collaborative attention module utilizes the attention mechanism of channel and spatial dimensions to complete cross-dimensional feature information fusion; and the frequency domain enhancement module, with a frequency-gated feedforward network at its core, effectively mines and enhances the frequency domain features of the data, assisting in improving the performance of super-resolution tasks.
[0008] In step 2, preprocessing is performed on the prepared dataset, and the size of each image is adjusted to ensure that the image size input to the network is fixed.
[0009] In step 3, a composite loss function is selected, which includes pixel loss, artifact loss, and adversarial loss. The choice of loss function affects the quality of the model, accurately reflecting the difference between the predicted and true values, and correctly reflecting the model's quality.
[0010] In step 3, the training of the network model also includes evaluating the quality and degree of image distortion of the remote sensing image reconstruction results through evaluation metrics, and measuring the role of the super-resolution reconstruction network.
[0011] In step 4, the network is trained using the Gaofen Image Dataset to enhance its robustness.
[0012] A super-resolution device for remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement, comprising: Image acquisition module: used to load datasets, the loading items being remote sensing images to be preprocessed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded and processed remote sensing images, including spatial domain feature enhancement module and frequency domain enhancement module, and obtains the final super-resolution model through iteration; Image output module: Used to display the reconstructed super-resolution image and output the reconstructed high-resolution image using electronic devices.
[0013] The image acquisition module is the starting point of the entire process. It is responsible for loading datasets containing remote sensing images to be processed. The goal of this module is to collect and prepare image data for subsequent processing. Next is the image processing module. Once the images are loaded, this module's task is to preprocess the remote sensing images, resizing them to 256×256 pixels. This ensures the consistency and standardization of the input images, providing a consistent benchmark for subsequent processing. The third is the image reconstruction module, which is the core of the entire process. It includes two main modules: a spatial domain feature enhancement module and a frequency domain enhancement module. These modules process the loaded and preprocessed images through iterative training. The spatial domain feature enhancement module is responsible for adaptively capturing multi-scale features of remote sensing images and achieving cross-channel and cross-spatial dimensional information fusion, while the frequency domain enhancement module focuses on mining and enhancing the frequency domain features of remote sensing images, working in conjunction with the spatial domain feature enhancement module to strengthen the high-frequency details and structural information of the images. Finally, there is the image output module, which is responsible for displaying the reconstructed super-resolution image. This module uses electronic devices to output the processed high-resolution image, allowing users to observe and evaluate the final result. These four modules work together to form a process that, through preprocessing, training, and output, ultimately achieves super-resolution processing of remote sensing images.
[0014] The aforementioned electronic device includes an input / output unit, a central processing unit, a memory, and a display. The computer program is stored in the memory and, when executed by the processor, implements the various steps of the aforementioned generative adversarial network remote sensing image super-resolution method based on fusion attention and frequency domain enhancement.
[0015] (III) Beneficial Effects Compared with existing technologies, this invention provides a method and apparatus for super-resolution of remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement, which has the following beneficial effects: 1. This invention proposes a spatial domain feature enhancement module. This network integrates a dynamic convolution module, a multi-scale residual module, a Shuffle Attention module, and a channel-spatial collaborative attention module. It adaptively captures multi-scale features from macroscopic structure to microscopic texture in remote sensing images and achieves cross-channel and cross-spatial dimension information fusion through an attention mechanism. This effectively solves problems such as imbalance in multi-scale and multi-directional feature extraction and insufficient utilization of spatial-frequency domain information, effectively enhancing the network's feature representation and recovery capabilities for high-frequency edges and complex texture details. It achieves adaptive capture of multi-scale features from macroscopic structure to microscopic texture in remote sensing images and cross-channel and cross-spatial dimension information fusion.
[0016] 2. This invention designs a frequency domain enhancement module, the core of which is a frequency-gated feedforward network. This module adaptively enhances features in the frequency domain. Through collaboration with a spatial domain feature enhancement network, it effectively strengthens the high-frequency details and structural information of the image. This effectively solves the problem that existing super-resolution reconstruction techniques are mostly limited to spatial domain processing, lacking an effective collaborative mechanism between the spatial and frequency domains, making it difficult to fully extract high-frequency detail information. This results in reconstructed images with structural blurring, insufficient detail representation, and limited overall feature expression capabilities. It significantly improves the overall feature expression capability and reconstruction quality.
[0017] 3. This invention designs a composite loss function comprising an improved pixel loss, an optimized artifact loss, and an adversarial loss. The pixel loss achieves pixel-level consistency supervision between the super-resolution image and the real high-resolution image; the artifact loss, through multi-step recognition logic combined with EMA technology, accurately locates artifact regions while avoiding excessive penalty for real textures in the early stages of adversarial training, effectively suppressing artifact generation; the adversarial loss guides the generator network output to match the distribution of the real high-resolution image, improving visual realism; finally, by balancing the contributions of the three through weight coefficients, a synergistic optimization of pixel fidelity, artifact suppression, and visual realism is achieved. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a schematic diagram of the adaptive feature enhancement network of the present invention; Figure 3 This is a schematic diagram of the dynamic convolution module structure of the present invention; Figure 4 This is a schematic diagram of the multi-scale residual module structure of the present invention; Figure 5 This is a schematic diagram of the Shuffle Attention module structure of the present invention; Figure 6 This is a schematic diagram of the dual-branch attention calculation module structure of the present invention; Figure 7 This is a schematic diagram of the channel-space collaborative attention module structure of the present invention; Figure 8 This is a schematic diagram of the frequency-gated feedforward network structure of the present invention; Figure 9 This is a schematic diagram of the discriminator structure of the present invention; Figure 10 A schematic diagram of the image super-resolution device provided by the present invention; Figure 11 A schematic diagram comparing relevant indicators of existing technologies and the method proposed in this invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1 like Figure 1 As shown in the figure, Implementation Example 1 of the invention provides a flowchart of a method for super-resolution of remote sensing images based on generative adversarial networks that fuses attention and frequency domain enhancement. The method specifically includes the following steps: Step 1: Construct the network model: Construct an adaptive feature enhancement network. This network consists of a spatial domain feature enhancement module and a frequency domain enhancement module. The spatial domain feature enhancement module includes a dynamic convolution module, a multi-scale residual module, a Shuffle Attention module, and a channel-spatial collaborative attention module; the frequency domain enhancement module uses a frequency-gated feedforward network as its core.
[0021] Step 2: Prepare the dataset: Use the NWPU-RESISC45 remote sensing scene classification dataset as both training and testing data. Use the Gaofen Image Dataset to train and fine-tune the model.
[0022] Step 3: Train the network model: Train the adaptive feature enhancement network model. Preprocess the dataset prepared in Step 2, adjust the size of each image in the dataset, fix the size of the input image, and input the processed dataset into the network model built in Step 1 for training.
[0023] During training, model parameters are optimized by minimizing the composite loss function between the network output features and the real target. This process continues until the number of training iterations reaches a set threshold or the loss function value converges to a stable range. At this point, model pre-training is considered complete, and the parameters from this stage are saved. The composite loss function includes pixel loss, artifact loss, and adversarial loss to comprehensively supervise the quality of feature enhancement. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used as evaluation metrics to objectively measure the enhancement results in terms of pixel fidelity, structural preservation, and feature layer quality.
[0024] Step 4: Fine-tuning the Model: To further improve the model's robustness and cross-scene adaptability, the Gaofen ImageDataset dataset was used for training and fine-tuning. This stage involves adjusting model parameters to optimize its feature enhancement stability and generalization ability under complex terrain and different imaging conditions. Ultimately, this results in better image reconstruction quality. Step 5: Save the model: Solidify the finalized model parameters.
[0025] In step 1, the adaptive feature enhancement network is a generator model, such as... Figure 2 As shown.
[0026] Adaptive feature enhancement networks such as Figure 2 As shown, the system includes a shallow feature extraction module, a deep feature extraction module, a feature enhancement module, and an upsampling module. The deep feature extraction module consists of a dynamic convolution module, a multi-scale residual module, a ShuffleAttention module, and a channel-space co-attention module. The feature enhancement module consists of a frequency-gated feedforward network.
[0027] The shallow feature extraction module consists of convolutional layers with a kernel size of 3×3 and a stride of 1.
[0028] Dynamic convolution modules, such as Figure 3 As shown, the module consists of a global average pooling unit, a first fully connected layer (feature compression and transformation), a second fully connected layer (dynamic parameter generation), a predefined multi-directional convolutional kernel group, and a dynamic weight normalization unit (Softmax). Input features are first fed into the global average pooling unit; then, this global information is input into the first fully connected layer; the processed features are fed into the second fully connected layer; the generated parameters are simultaneously applied to the predefined multi-directional convolutional kernel group and the dynamic weight normalization (Softmax) component, which uses the Softmax function; finally, the output of the multi-directional convolutional kernel group and the normalized weights are processed through an element-wise operation component to output the final processed features of this module.
[0029] Multi-scale residual modules such as Figure 4 As shown, the module consists of three deep convolutional layers: Layer 1, Layer 2, Layer 3, and Layer 4, along with activation functions. Layer 1 has a kernel size of k=3×3 and a stride of s=1; Layer 2 has a kernel size of k=5×5 and a stride of s=1; Layer 3 has a kernel size of k=7×7 and a stride of s=1; and Layer 4 has a kernel size of k=1×1 and a stride of s=1. Input features are simultaneously input into these three layers simultaneously. The output of each layer is fed into a ReLU activation function. The activated features from these three layers are then combined and input into Layer 4. Finally, the output of Layer 4 is processed with the original input features using an element-wise arithmetic component to output the final processed features of this module. The formula for the ReLU activation function is as follows: The Shuffle Attention module consists of a feature grouping unit, a channel shuffling unit, a dual-branch attention calculation unit, an attention fusion component, a feature weighting component, and a grouping and concatenation unit. Figure 5 As shown. Where the number of feature groups is g, and the input feature dimension is... Dimension is The features to be processed are first input into feature grouping units, and the features are split along the channel dimension according to the number of groups g. After splitting, the dimension of a single feature group is... The split feature input channel shuffling unit shuffles and rearranges the feature order along dimension C. The shuffled features are then input into a dual-branch attention calculation unit to learn the attention weights corresponding to different dimensions in parallel. Next, the attention fusion component performs element-wise multiplication on the attention weights output from the dual branches to fuse multi-dimensional attention information. The fused result is then multiplied element-wise with the original features by a feature weighting component to perform weighted enhancement. Finally, the weighted features are input into a grouping and concatenation unit to concatenate and integrate g groups of features along the channel dimension, restoring them to dimension C. Output the processing results of this module.
[0030] The dual-branch attention mechanism in the Shuffle Attention module consists of two parallel feature processing branches and an element-wise operation component, such as... Figure 6 As shown. Each branch contains a global average pooling unit, a fully connected (FC) layer dimensionality reduction unit, an activation function unit, and a dimensionality increase unit. Activation functions one and two are ReLU activation functions, while activation functions three and four are Sigmoid activation functions. The formula for the Sigmoid activation function is as follows: Dimensions The input features are synchronously input into two parallel branches. Each branch first compresses the H×W spatial dimension to 1×1 through a global average pooling unit; then, the global information is input into the FC layer dimensionality reduction unit to reduce the feature dimension from C to C / r; finally, one of the branches uses a dimensionality upscaling unit to upscal the features to the required spatial dimension. Another branch uses a dimensionality-upgrading unit to upgrade the features to the channel dimension. The two upscaled features are mapped to the 0-1 interval using a Sigmoid activation component, yielding attention weights for the spatial and channel dimensions. Finally, the two weights are multiplied using an element-wise operation component, resulting in an output dimension of... The weighted results are fused using a dual-branch attention mechanism.
[0031] Channel-space collaborative attention module such as Figure 7As shown. In the channel-space collaborative attention module, the self-attention mechanism is executed separately for the channel and spatial dimensions, as follows. Figure 7 As shown. For channel self-attention, the input feature map Then, a query, key, and value matrix is generated first through linear projection, i.e.: in Let be an unbiased linear projection matrix, and reshape it as... of The channel is then divided into h heads (each head has a dimension of h). ),Right now The calculation process for the i-th head is as follows: in To obtain learnable temperature parameters, the outputs of all heads are concatenated and reconstructed to obtain the channel self-attention features. Following that, regarding the Global average pooling is applied, followed by processing through a convolutional layer with kernel size k=3×3 and stride s=1. The result is then input into activation function one, the GeLU activation function, and subsequently processed through another convolutional layer with kernel size k=3×3 and stride s=1. Finally, the transformed features are obtained through activation function two, the Sigmoid activation function. .
[0032] Parallel execution of spatial self-attention (corresponding to the spatial dimension processing part in the block diagram) is to... Divide and flatten into containing Non-overlapping windows of pixels, reshaped into of Similarly, after dividing into h heads, the calculation process for the i-th head is as follows: Where B is a learnable relative position code, which is then reshaped and concatenated to obtain spatial self-attention features. (The default operation uses a shift window to enhance spatial information capture).
[0033] Finally, the channel self-attention process obtained Output characteristics of spatial self-attention Element-wise multiplication is performed to fuse channel and spatial attention information, collaboratively capture global context information and optimize feature interaction, thereby improving the model's understanding of the global structure of the image and the efficiency and accuracy of feature extraction.
[0034] Frequency-gated feedforward networks such as Figure 8As shown, it consists of fully connected layers, activation functions, channel grouping units, frequency domain processing units, splitting units, depthwise convolutional layers, element-wise operation components, and convolutional layers, as follows. Figure 8 As shown. The frequency domain processing unit consists of an integrated 2D Fast Fourier Transform, a learnable filter M, and a 2D Inverse Fast Fourier Transform. The input features are first fed into the fully connected layer of the network, and then pass through the GeLU activation function to obtain the features. Then on Features are split along the channel dimension by a channel grouping unit. The split features are then input into a frequency domain processing unit: first, they undergo an integrated 2D Fast Fourier Transform, then multiply with the learnable filter M within the unit, and finally undergo a 2D Inverse Fast Fourier Transform to obtain the frequency domain processed features. ; then will and Feature information is aggregated by an element-wise addition component, and then the aggregated features are split into two parts in the channel dimension by a splitting unit. One part of the features is input into a deep convolutional layer with a kernel size of k=3×3 and a stride of s=1. Then, it interacts with the other part of the features split by the element-wise multiplication component. Finally, the element-wise multiplication result is input into a convolutional layer with a kernel size of k=3×3 and a stride of s=1 to output the final features.
[0035] Discriminator structure as follows Figure 9 As shown, the system consists of convolutional layers 1 to 7, batch normalization layers 1 to 6, and activation functions 1 to 7. Specifically, the kernel size of convolutional layer 1 is 4×4, and the stride is s=2; the kernel size of convolutional layer 2 is 4×4, and the stride is s=2; the kernel size of convolutional layer 3 is 4×4, and the stride is s=2; the kernel size of convolutional layer 4 is 4×4, and the stride is s=2; the kernel size of convolutional layer 5 is 4×4, and the stride is s=1; the kernel size of convolutional layer 6 is 4×4, and the stride is s=1; and the kernel size of convolutional layer 7 is 4×4, and the stride is s=1. Activation functions 1 to 6 are LeakyReLU activation functions, and activation function 7 is a Sigmoid activation function. The formula for the LeakyReLU activation function is as follows: in =0.01.
[0036] In step 3, the appropriate evaluation metrics are Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity. PSNR is an image quality evaluation metric based on pixel-level errors, used to measure image sharpness and accuracy. Structural Similarity, on the other hand, considers aspects such as brightness, contrast, and structure of an image, used to measure the similarity between images. PSNR and Structural Similarity are defined as follows: in , Let x and y represent the mean and variance of the image, respectively. and Let x and y represent the standard deviations of the image, respectively. Represents the covariance of the x and y graphs. and It is a constant.
[0037] The training iterations are set to 200. The learning rate for the first 100 training iterations is set to 0.0001, and the learning rate for the next 100 training iterations gradually decreases from 0.0001 to 0. The upper limit of the number of images input to the network each time is mainly determined by the performance of the computer's graphics processor. Generally, the number of images input to the network each time is in the range of 10-20, which can make the network training more stable and the training results better, and can ensure that the network fits quickly.
[0038] In step 3, the network output and the ground truth labels are compared using a loss function. Minimizing this loss function optimizes model performance and improves the quality of the output. During training, a composite loss function is used: the image super-resolution reconstruction network employs an adversarial loss function to enhance the visual realism of the output image; the adaptive feature enhancement network uses a multi-domain composite loss function, including pixel loss and artifact loss, to supervise the quality of feature enhancement from two dimensions: pixel fidelity and artifact suppression.
[0039] Pixel loss calculates the difference between two images pixel by pixel. Therefore, image SR reconstruction methods can use pixel loss to determine the pixel-level consistency between SR and HR images. The formula for pixel loss is as follows: In the formula for pixel loss, and It is the feature image of paired LR and HR images in the dataset, and M is the logarithm of the dataset. In generative networks and The mapping relationship between them.
[0040] This invention uses an artifact loss function to address artifacts that often appear in high-frequency feature regions of images in adversarial networks.
[0041] First, the present invention calculates and The remaining as follows: Subsequently, local statistics are calculated to determine pixel differences within a specific region. This formula can be expressed as: In the formula for calculating local statistics This represents the statistical area; var(·) calculates the variance of the pixels in that region. Local statistics can better identify texture details with regular edges, but they are less effective at identifying randomly distributed texture details. Therefore, this invention uses a global patch. The formula for solving this problem is as follows: In calculation In the formula, This is a global patch parameter. This invention uses... This is used to identify artifact regions. However, in the early stages of adversarial training, some identification errors still exist, leading to over-penalization of realistic texture details. To address this issue, this invention employs the EMA technique, as shown in the following formula: In calculation In the formula, These are weight parameters; It represents the number of network iterations. Through the first The model obtained from the training; It is a model obtained through EMA technology. (And...) compared to, It is more reliable and can reduce the generation of random artifacts.
[0042] This invention uses EMA technology for optimization. In each iteration, obtain and Then, these two models used generate and Then calculate the residuals to obtain and . By comparison and Determine the penalty location, and finally multiply Mr by The final artifact loss is obtained. The calculation process can be expressed as the formula: The core idea of adversarial loss is to introduce a discriminative network that guides the generative network to produce reconstruction results that are difficult to distinguish from real high-resolution images in terms of data distribution. Therefore, image SR reconstruction methods can use adversarial loss to determine whether the SR image lies on the manifold of the natural image, thereby improving its visual realism. The formula for adversarial loss is as follows: In the formula for adversarial loss, G(⋅) is the generator network; D(⋅) is the discriminator network; and E[⋅] represents the mathematical expectation of the data distribution.
[0043] Therefore, the loss function is defined as: in =0.3, =0.4, =0.3.
[0044] In step 4, the model is trained and fine-tuned using the second remote sensing image dataset. The Gaofen Image Dataset is used during the fine-tuning of model parameters. For the Gaofen Image Dataset, 300 images are used for training, and 100 images are used for testing.
[0045] In step 5, after the network training is completed, all parameters in the network are saved. The remote sensing image to be super-resolution is then input into the network to obtain the reconstructed image.
[0046] Example 2 like Figure 10 As shown, this embodiment provides a remote sensing image super-resolution device based on generative adversarial networks that fuse attention and frequency domain enhancement, capable of executing the above-described method. The device includes: Image acquisition module: Used to load datasets, responsible for loading the remote sensing image datasets to be processed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded and processed remote sensing images, including spatial domain feature enhancement module and frequency domain enhancement module. The spatial domain feature enhancement module includes shallow feature extraction module, deep feature extraction module and feature enhancement module; the frequency domain enhancement module includes frequency-gated feedforward network, and obtains the final super-resolution model through iteration. Image output module: Used to display the reconstructed super-resolution image and output the reconstructed super-resolution image using electronic devices.
[0047] The entire process begins with the image acquisition module, whose task is to load the dataset containing the remote sensing images to be processed. The goal of this module is to prepare the image data for subsequent processing. Next is the image processing module, which, once the images are loaded, preprocesses them, resizing them to 256×256 pixels to ensure consistency and standardization of the input images. This standardized size provides a consistent benchmark for subsequent processing. The third module is the image reconstruction module, which is the core of the entire process. It comprises two main networks: a spatial domain feature enhancement module and a frequency domain enhancement module. These networks are iteratively trained to process the loaded and preprocessed images. The spatial domain feature enhancement module adaptively captures multi-scale features of the remote sensing images and achieves cross-channel and cross-spatial dimensional information fusion, while the frequency domain enhancement module focuses on mining and enhancing the frequency domain features of the remote sensing images, working in conjunction with the spatial domain feature enhancement module to strengthen high-frequency details and structural information. Finally, the image output module is responsible for displaying the reconstructed super-resolution image. This module uses electronic devices to output the processed high-resolution image, allowing users to observe and evaluate the final result. These four modules work together to form a process that, through preprocessing, training, and output, ultimately achieves super-resolution processing of remote sensing images.
[0048] This invention constructs a method and apparatus for super-resolution remote sensing images based on generative adversarial networks (GANs) that fuses attention and frequency domain enhancement. This method can directly generate high-resolution remote sensing images from low-resolution images without intermediate steps. Under the same conditions, the feasibility and superiority of this proposed method are further verified by calculating and comparing the image correlation metrics obtained from existing methods. A comparison of the correlation metrics between existing technologies and the proposed method is provided below. Figure 11 As shown in the table, the method proposed in this invention has a higher peak signal-to-noise ratio and structural similarity than existing methods. These indicators further demonstrate that the method proposed in this invention has better image super-resolution reconstruction quality.
[0049] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A super-resolution method for remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement, characterized in that: Includes the following steps: Step 1: Construct an adaptive feature enhancement network, which includes a spatial domain feature enhancement module and a frequency domain enhancement module; Step 2: Using NWPU-RESISC45 as the first remote sensing image dataset, divide it into training and test sets and preprocess it; The Gaofen Image Dataset was used as a second remote sensing image dataset to fine-tune the model; Step 3: Input the dataset prepared in Step 2 into the network model built in Step 1 for training until the number of training iterations reaches the initial threshold or the value of the loss function reaches the preset range. At this point, the network model is considered to have been trained and the network model parameters are saved. Step 4: Use the second remote sensing image dataset to retrain and fine-tune the network model, optimize the network model parameters, further improve the image super-resolution reconstruction performance, and obtain a network model that can achieve the best results. Step 5: Solidify the finalized model parameters and save the model. The model can output a corresponding high-resolution remote sensing image after inputting a low-resolution remote sensing image.
2. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: In step 1, the spatial domain feature enhancement module includes a shallow feature extraction module and a deep feature extraction module; the frequency domain enhancement module includes a frequency-gated feedforward network.
3. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: The spatial domain feature enhancement module includes a dynamic convolution module, a multi-scale residual module, a shuffle attention module, and a channel-space collaborative attention module. Specifically: the dynamic convolution module adaptively adjusts the convolution kernel parameters using dynamic convolution components to enhance the multi-directional representation capability of features; the multi-scale residual module captures multi-scale features of the data using multi-scale convolution; the shuffle attention module achieves cross-channel information interaction and feature weighting using channel shuffling and attention calculation components; the channel-space collaborative attention module utilizes the attention mechanism of channel and spatial dimensions to complete cross-dimensional feature information fusion; and the frequency domain enhancement module, with a frequency-gated feedforward network at its core, effectively mines and enhances the frequency domain features of the data, assisting in improving the performance of super-resolution tasks.
4. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: In step 2, the prepared dataset is preprocessed, the size of each image in the dataset is adjusted, and the size of the input image is fixed as the network input.
5. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: In step 3, a composite loss function is selected, which includes pixel loss, artifact loss and adversarial loss.
6. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: In step 3, the training of the network model also includes evaluating the quality and degree of image distortion of the remote sensing image reconstruction results through evaluation metrics, and measuring the role of the super-resolution reconstruction network.
7. The method for super-resolution of remote sensing images based on generative adversarial networks with fused attention and frequency domain enhancement as described in claim 1, characterized in that: In step 4, the network is trained using the Gaofen Image Dataset to enhance its robustness.
8. A super-resolution device for remote sensing images based on generative adversarial networks that fuse attention and frequency domain enhancement, characterized in that: The device includes: Image acquisition module: used to load datasets, the loading items being remote sensing images to be preprocessed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded and processed remote sensing images, including spatial domain feature enhancement module and frequency domain enhancement module, and obtains the final super-resolution model through iteration; Image output module: Used to display the reconstructed super-resolution image and output the reconstructed high-resolution image using electronic devices.
Citation Information
Patent Citations
A low-quality image super-resolution reconstruction method based on a generative adversarial network
CN120339072B