A real scene-oriented remote sensing image super-resolution reconstruction method and device

CN122529971APending Publication Date: 2026-08-07CHANGCHUN UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGCHUN UNIV OF SCI & TECH
Filing Date
2026-05-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本发明旨在解决缺乏配对训练数据集的条件下难以提升遥感图像清晰度的问题,提出了一种面向真实场景的遥感图像超分辨率重建方法及装置

Benefits of technology

1、本发明提出了一种更加接近实际情况的模糊核估计网络,该模型仅以真实低分辨率遥感图像作为输入,通过引入频域特征提取模块生成伪OTF(光学传递函数)并进一步估计实际模糊核。所生成的模糊核能够有效反映真实成像系统与场景退化过程,从而克服合成退化模型与真实退化之间的分布差异问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122529971A_ABST
    Figure CN122529971A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image processing, and is especially a remote sensing image super-resolution reconstruction method and device for real scenes, and specifically comprises the following steps: step 1, constructing a network model: constructing a network model comprising a blur kernel estimation network and an image super-resolution reconstruction network; step 2, preparing a data set: using a DIV2K remote sensing image data set to divide a training set and a test set and to perform preprocessing; using a UCMerced data set to fine-tune the model; step 3, training the network model: inputting the data set prepared in step 2 into the network model constructed in step 1 to perform training; step 4, fine-tuning the model: using the UCMerced data set to perform retraining and fine-tuning of the network model, to obtain a final model; step 5, saving the model: solidifying the parameters of the final model obtained, to save the model. The present method utilizes the information and features of low-resolution images themselves, and does not rely on corresponding high-resolution images, and finally improves the definition and resolution of remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image super-resolution reconstruction, specifically a method and apparatus for super-resolution reconstruction of remote sensing images for real-world scenes. Background Technology

[0002] Super-resolution reconstruction of remote sensing images has always been an important research direction in the field of computer vision. Its core objective is to reconstruct clear images with higher resolution and richer details such as edge textures and high-frequency structures from single or multiple low-resolution images using image processing techniques and deep learning models. Since the introduction of convolutional neural networks into image super-resolution tasks, their reconstruction performance has significantly improved compared to traditional reconstruction-based or example-learning methods. However, the performance of current mainstream super-resolution methods in real-world remote sensing scenarios has not yet reached ideal levels. Deep learning models still have significant shortcomings in feature representation, degradation modeling, and high-frequency information recovery. Therefore, improving the super-resolution reconstruction capability of remote sensing images for real-world scenarios remains a key issue that requires in-depth research and optimization.

[0003] The Chinese patent publication number is "CN119494781B," entitled "A Method for Super-Resolution Reconstruction of Remote Sensing Images Based on Degradation Mechanisms." This method first processes high-resolution images in both the spatial and frequency domains using a degradation model comprised of spatial domain filtering, frequency domain filtering, and random degradation control. The two types of results are then fused, downsampled, and noise-added to generate a degraded image that approximates the real imaging conditions. Subsequently, this degraded image is input into a super-resolution network composed of spatial and frequency domain branches to extract local multi-scale features and global frequency features, which are then fused to generate a reconstructed image. This method can achieve high-quality super-resolution reconstruction of degraded remote sensing images. While this method has achieved some improvement in experimental scenarios, this type of degradation mechanism based on artificial design and random control still has limitations in complex remote sensing scenarios: its degradation form relies on preset filtering operators, making it difficult to accurately match the degradation patterns formed by multiple factors such as optical blurring, platform motion, and sensor characteristics in real remote sensing imaging systems. Degradation deviations are particularly prone to occur under different sensors, different atmospheric conditions, or complex texture regions. This leads to a significant decrease in reconstruction quality in practical applications. Therefore, constructing a mechanism that can more realistically simulate the degradation of remote sensing images has become a key issue that needs to be addressed. Summary of the Invention

[0004] This invention aims to address the challenge of improving the sharpness of remote sensing images when paired training datasets are lacking. It proposes a method and apparatus for super-resolution reconstruction of remote sensing images in real-world scenarios. This method addresses the complex degradation characteristics of remote sensing images in real-world environments by constructing a blur kernel estimation network to obtain degradation information that more closely reflects the actual imaging process. By adaptively estimating scene-related blur kernels from real degraded images, it replaces the fixed or random filtering process in traditional degradation models. This makes the degraded images generated during the training phase more closely resemble the real imaging degradation process, thereby significantly improving the generalization ability and reconstruction accuracy of the super-resolution network.

[0005] To achieve the above objectives, the present invention specifically adopts the following technical solution: A method for super-resolution reconstruction of remote sensing images for real-world scenes includes the following steps: Step 1, Construct the network model: Construct a network including a fuzzy kernel estimation network and an image super-resolution reconstruction network; Step 2, Prepare the dataset: Split the training and test sets using the DIV2K dataset and preprocess them; fine-tune the model using the UCMerced dataset; Step 3, train the network model: input the dataset prepared in step 2 into the network model built in step 1 for training; Step 4, fine-tune the model: retrain and fine-tune the network model using the UCMerced dataset to obtain the final model; Step 5, Save the model: Solidify the parameters of the final model and save the model.

[0006] The above-mentioned method for super-resolution reconstruction of remote sensing images for real-world scenes is characterized by: In step 1, the fuzzy kernel estimation module includes a frequency domain feature extraction network and a fuzzy kernel generation network; the image super-resolution reconstruction network consists of an attention-guided multi-feature fusion network and a dynamic attention channel feature enhancement network.

[0007] In step 1, the fuzzy kernel estimation module includes a frequency domain feature extraction module and a fuzzy kernel generation module. The frequency domain feature extraction module utilizes frequency domain-spatial domain transformation and optical transfer function to extract frequency domain information. It first maps the low-resolution image to the frequency domain space to capture degradation priors, then performs an inverse transformation back to the spatial domain while retaining key frequency features, fully simulating the real degradation patterns of the imaging system. The fuzzy kernel generation module uses max pooling and convolutional layers to simplify the dimensionality of the frequency domain degradation representation, and then obtains the degraded fuzzy kernel through upsampling and double-layer convolutional blocks. The image super-resolution reconstruction network uses a multi-stage channel separation feature fusion module (MCSF) to perform layered extraction and fusion of input features along the channel dimension. It combines dynamic convolution to achieve adaptive feature transformation, and then uses a channel feature enhancement module (CFE) to strengthen the effective information of the channel dimension. Finally, through layer-by-layer processing with convolutional layers and activation functions, the remote sensing image is reconstructed. The adjuster network uses a dual-path feature mapping structure to refine the input fuzzy kernel, and then outputs two feature paths through residual connections.

[0008] In step 2, preprocessing is performed on the prepared dataset, and the size of each image is adjusted to ensure that the image size input to the network is fixed.

[0009] In step 3, a composite loss function is selected for the loss function selection. The fuzzy kernel estimation network adopts frequency domain consistency loss; the image super-resolution reconstruction network adopts total variation loss, texture consistency loss and local self-similarity loss. The selection of the loss function affects the quality of the model, can truly reflect the difference between the predicted value and the true value, and can correctly reflect the quality of the model.

[0010] In step 3, the training of the network model also includes evaluating the quality and degree of image distortion of the remote sensing image reconstruction results through evaluation metrics, and measuring the role of the super-resolution reconstruction network.

[0011] In step 4, the network is trained using the DIV2K dataset to enhance its robustness.

[0012] A method and apparatus for super-resolution reconstruction of remote sensing images for real-world scenarios, comprising: Image acquisition module: used to load datasets, the loading items being remote sensing images to be preprocessed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded and processed remote sensing images, including a blur kernel estimation network and an image super-resolution reconstruction network, and obtains the final super-resolution model through iteration; Image output module: Used to display the reconstructed super-resolution image and output the reconstructed high-resolution image using electronic devices.

[0013] The image acquisition module is the starting point of the entire process. It is responsible for loading the dataset, which contains remote sensing images to be processed. The goal of this module is to collect and prepare image data for subsequent processing. Next is the image processing module. Once the images are loaded, this module's task is to preprocess the remote sensing images, resizing them to 256×256 pixels. This ensures the consistency and standardization of the input images, providing a consistent benchmark for subsequent processing. The third module is the image reconstruction module, which is the core of the entire process. It includes two main networks: an attention-guided multi-feature fusion network and a dynamic attention channel feature augmentation network. Strong networks, trained iteratively, process loaded and preprocessed images; fuzzy kernel estimation networks generate degraded representations of low-resolution images; and super-resolution networks reconstruct high-resolution images from low-resolution images, with the goal of generating an optimized super-resolution model. Finally, the image output module displays the reconstructed super-resolution image, using electronic devices to output the processed high-resolution image, allowing users to observe and evaluate the final result. These four modules work together to form a process that, through preprocessing, training, and output, ultimately achieves blind super-resolution processing of remote sensing images.

[0014] The aforementioned electronic device includes an input / output unit, a central processing unit, a memory, and a display, etc. The computer program is stored in the memory and, when executed by the processor, implements the various steps of the aforementioned method for super-resolution reconstruction of remote sensing images for real-world scenarios.

[0015] Compared with existing technologies, the present invention provides a method and apparatus for super-resolution reconstruction of remote sensing images for real-world scenarios, which has the following advantages: 1. This invention proposes a fuzzy kernel estimation network that more closely approximates reality. This model uses only real low-resolution remote sensing images as input, and generates a pseudo-OTF (Optical Transfer Function) by introducing a frequency domain feature extraction module to further estimate the actual fuzzy kernel. The generated fuzzy kernel can effectively reflect the real imaging system and scene degradation process, thereby overcoming the distribution difference problem between synthetic degradation models and real degradation.

[0016] 2. This invention proposes a super-resolution reconstruction network to enhance the recovery of high-frequency details during the reconstruction process. The network consists of an attention-guided multi-feature fusion network and a dynamic attention channel feature enhancement network. The attention-guided multi-feature fusion network solves the problem of insufficient feature expression in low-resolution images through attention mechanisms and bi-branch pooling, and realizes refined selection and fusion of multi-branch features. The dynamic attention channel feature enhancement network combines dynamic convolution and channel attention to capture the complex degradation characteristics of low-resolution images, enhance the channel dimension expression of features, and provide support for the recovery of high-frequency details.

[0017] 3. This invention proposes a composite loss function consisting of frequency domain consistency loss, local self-similarity loss, texture consistency loss, and total variation loss, which optimizes edge structure and visual perception, effectively improving the quality of image super-resolution and enhancing the realism of the image. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the method flow of the present invention; Figure 2 This is a diagram of the overall network structure described in this invention; Figure 3 This is a schematic diagram of the frequency domain feature extraction module of the present invention; Figure 4 This is a schematic diagram of the attention-guided multi-feature fusion network module of the present invention; Figure 5 This is a schematic diagram of the multi-stage channel separation feature fusion module of the present invention; Figure 6 This is a schematic diagram of the channel feature enhancement module of the present invention; Figure 7 This is a schematic diagram of the dual-path feature mapping network module of the present invention; Figure 8 This is a schematic diagram of the dynamic attention channel feature enhancement network module of the present invention; Figure 9 A schematic diagram of the image super-resolution device provided by the present invention; Figure 10 This is a schematic diagram comparing relevant indicators of the method proposed in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Example 1 like Figure 1 As shown in the flowchart of Embodiment 1, a method for super-resolution reconstruction of remote sensing images for real-world scenes is provided. This method specifically includes the following steps: Step 1: Construct the network model: Construct a fuzzy kernel estimation network and an image super-resolution reconstruction network; wherein, the fuzzy kernel estimation network includes a frequency domain feature extraction network and a fuzzy kernel generation network; the image super-resolution reconstruction network consists of an attention-guided multi-feature fusion network and a dynamic attention channel feature enhancement network.

[0021] Step 2: Prepare the dataset: Divide the DIV2K remote sensing image dataset into training and test sets.

[0022] Step 3: Train the network model: Train the remote sensing image super-resolution network model. Preprocess the dataset prepared in Step 2, adjust the size of each image in the dataset, fix the size of the input image, and input the processed dataset into the network model built in Step 1 for training.

[0023] During training, the optimal evaluation metric and the minimum loss function value are selected. Specifically, the loss function between the network output image and the label is minimized until the number of training iterations reaches a set threshold or the loss function value falls within a set range. The model parameters are then considered pre-trained and saved. Simultaneously, the optimal evaluation metric is selected to measure the algorithm's accuracy and evaluate the system's performance. During training, a composite loss function is used: frequency domain consistency loss is employed for the fuzzy kernel estimation network, while local similarity loss, texture consistency loss, and total variation loss are used for the image super-resolution reconstruction network. The choice of loss function significantly impacts the model's performance, accurately reflecting the difference between predicted and true values ​​and providing correct feedback on model quality. Appropriate evaluation metrics include the Learning Perceptual Patch Similarity Index (LPISP) ​​and the Natural Image Quality Evaluation Index (NIQE), which effectively assess the quality of remote sensing image reconstruction results and the degree of image distortion, measuring the effectiveness of the super-resolution reconstruction network.

[0024] Step 4: Fine-tune the model: Train and fine-tune the model using the UCMerced dataset to obtain stable and usable model parameters, further improving the model's remote sensing image super-resolution capability; ultimately resulting in better image quality reconstructed by the model. Step 5: Save the model: Solidify the finalized model parameters.

[0025] The overall network model in step 1 is as follows: Figure 2 As shown, the processing flow of the frequency domain feature extraction module is as follows: Figure 3As shown, this module addresses the frequency domain information requirements of low-resolution remote sensing images by constructing a complete processing chain of "frequency domain transformation - frequency domain prior extraction - inverse frequency domain transformation": first, the spatial domain image is mapped to the frequency domain to analyze its frequency distribution characteristics; then, the optical transfer function (OTF) of the imaging system is used to perform prior enhancement of the frequency domain information; finally, the processed frequency domain features are restored to the spatial domain, providing richer frequency domain detail support for subsequent image reconstruction. The proposed frequency domain feature extraction model can be described as follows: (1) Frequency domain transformation: To analyze the frequency components of low-resolution images (such as high-frequency edges and low-frequency smooth regions), the Discrete Fourier Transform (DFT) is used to transform the spatial domain image to the frequency domain. Its mathematical expression is as follows: in, Low-resolution remote sensing images representing the spatial domain are used for Spatial domain pixel coordinates; This is the transformed frequency domain representation. is the frequency domain coordinate; M and N correspond to the width and height dimensions of the image; j is the imaginary unit, and a linear transformation from the spatial domain to the frequency domain is achieved through a complex exponent term.

[0026] (2) Frequency Domain Prior Extraction: To simulate the frequency domain response characteristics of the imaging system and enhance the effective details in the frequency domain, the optical transfer function (OTF) and frequency domain weighting operators are introduced to perform prior processing on the frequency domain information. Its definition is expressed as: in, It is the optical transfer function of the imaging system, obtained by Fourier transform of the system's point spread function (PSF), i.e. The PSF is used to describe the diffusion effect of the imaging system on the "point light source". It corresponds to the amplitude attenuation and phase shift of the OTF in the frequency domain and can accurately simulate the optical blurring characteristics in the imaging process. The frequency domain weighting operator uses a Gaussian high-pass filter to enhance high-frequency details; its expression is: Where σ is the standard deviation of the Gaussian function, and random sampling is performed from a uniform distribution U[0.1,0.5] in the experiment to balance the intensity of high-frequency enhancement with the noise suppression effect.

[0027] (3) Inverse Frequency Domain Transform: The frequency domain information after prior processing is transformed back to the spatial domain to obtain an image containing frequency domain enhancement features. This is achieved using the Inverse Discrete Fourier Transform (IDFT), and the formula is as follows: in, The final output spatial domain image retains the detailed features after frequency domain processing and can be directly used for subsequent image super-resolution, enhancement and other tasks.

[0028] By processing the entire process of the frequency domain feature extraction model, the frequency domain distribution of low-resolution remote sensing images can be obtained. At the same time, by simulating the frequency domain response of the imaging system using the optical transfer function, high-frequency details and frequency domain degradation characteristics can be accurately captured, providing prior support in the frequency domain dimension for subsequent blur kernel generation. These frequency domain features can help the blur kernel more accurately match the optical blur rules of the real imaging system, further improving the generation accuracy of the blur kernel and providing samples that are closer to real scenes for the training of the super-resolution network.

[0029] In step 1, the attention-guided multi-feature fusion network (AGMF) is as follows: Figure 4 As shown, the network targets four sets of features from different sources in the input. Grouping: First, group content-related features... Structure-related features Preliminary feature fusion is achieved by concatenating the features along the channel dimension. Then, local feature information is extracted through the local convolutional mapping module to obtain two sets of intermediate fused features with unified dimensions.

[0030] The two sets of intermediate fused features are concatenated again and input into the attention guidance module. Global information is extracted through max pooling and average pooling, and attention weights are generated through a multilayer perceptron to enhance key texture and structural regions. The weighted features are then remapped by convolution and residually connected to the initial input features to output the final fused features. .

[0031] The convolution kernel size is 3×3 with a stride of 1. Combined with the sigmoid non-linear activation function, the 3×3 convolution ensures reasonable coverage of the local receptive field, fully capturing the detailed features of remote sensing images and enhancing feature discrimination capabilities. The AGMF module enhances texture and edge structures through attention enhancement and residual mechanisms, providing high-quality features for deep reconstruction networks, thereby improving the detail restoration and edge sharpness of super-resolution images.

[0032] In step 1, the Multi-Stage Channel Separation Feature Fusion Network (MCSF) is as follows: Figure 5 As shown, this module addresses the issues of redundant feature channel information and insufficient contrast in detailed regions in remote sensing image super-resolution tasks by introducing multi-stage channel separation and branch feature attention enhancement operations.

[0033] First, the input features are processed through a multi-stage process involving four sets of convolutional layers and channel separation layers. In each set, the input features are first locally mapped by a convolutional layer (3×3 kernel size, stride of 1), and then feature decomposition is performed in the channel separation layer, which decomposes the 64-channel features into 48-channel main features and 16-channel detail features, achieving channel-level separation of features with different attributes. The channel detail features processed in each stage are input into the fusion module for channel concatenation and channel fusion with the channel main features from the last stage.

[0034] Subsequently, the enhanced features are input into parallel convolutional layers in three directions, all of which use 3×3 kernels. The three outputs are then concatenated and processed again by convolution. Finally, the multi-channel fused features are weighted by attention and fused to obtain the final fused features.

[0035] The MCSF module achieves feature decoupling through multi-stage channel separation and, combined with a dynamic attention mechanism, effectively improves the ability to express the details of features, providing more accurate structural and texture priors for super-resolution variable reconstruction.

[0036] In step 1, the Channel Feature Enhancement Network (CFE) is as follows: Figure 6 As shown, this module addresses the issues of redundant feature information and insufficient weighting of key details in super-resolution tasks by introducing channel attention and gating mechanisms to adaptively filter and enhance effective information. First, the input features are processed by a channel attention (CA) unit to extract initial weights for each channel dimension. After preliminary weighting, the features are split into two paths into a gating branch: one path directly inputs to convolutional layer one for feature mapping; the other path is processed by convolutional layer two and an activation function to generate gating weights. Subsequently, the output of convolutional layer one is multiplied element-wise with the gating weights, achieving selective feature enhancement through the gating mechanism. By dynamically adjusting the gating weights, detailed features can be specifically enhanced while suppressing redundant information.

[0037] The gated enhanced features are sequentially passed through convolutional layer three and channel attention (CA) units for secondary feature mapping and channel weight optimization. Finally, a residual connection is performed with the initial input features to output the enhanced features. This module, through the combination of channel attention and gating mechanisms, achieves global information mining along the channel dimension and precisely enhances key details through gating branches, effectively improving the expressive power of the features.

[0038] In step 1, the dual-path feature mapping network is as follows: Figure 7 As shown, this module addresses the issues of low efficiency in low-resolution feature mapping and insufficient fusion of global and local information in super-resolution tasks by constructing a structure for dual-path parallel processing and complementary feature fusion: First, the input feature k is processed in two paths: one path upsamples through pixel shuffling to expand the feature dimension, and then extracts local detail features through a 3×3 convolutional layer; the other path captures global statistical information through global average pooling, and then transforms the dimensionality of the global features through a fully connected layer. The output of the local path is then passed through a 3×3 depthwise convolutional layer and a 3×3 convolutional layer to enhance the detail representation, while the output of the global path is stacked through fully connected layers to deepen the global information representation. The features processed by the two paths are then residually connected, and finally fused element-wise to integrate local details and global information.

[0039] Finally, the fused features are activated to output the mapping result. This module uses a dual-path parallel design to simultaneously cover the local details and global distribution information of the features, improving the completeness and accuracy of the feature mapping.

[0040] In step 1, the Dynamic Attention Channel Feature Enhancement Network (DACE) is as follows: Figure 8 As shown, this module includes a Multi-Stage Channel Separation Feature Fusion (MCSF) module and a Channel Attention (ECA) module. This module addresses the issues of insufficient channel-dimensional information and low discriminative power of detailed features in the output features of the preceding AGMF module by constructing a feature enhancement structure.

[0041] First, the input features are processed by convolutional layer one and then passed through two paths: one path inputs to the MCSF module to extract multi-scale features, and the other path inputs to the dynamic convolutional module to capture dynamic features. The two paths are then weighted and fused, and residual connections are performed with the initial input features. Finally, convolutional layer two unifies the feature dimensions. The features are then input to the ECA module to extract channel weights, achieving global information enhancement across the channel dimensions. The enhanced features are processed through two paths: one path uses dynamic convolution, an L-shaped activation function, and convolutional layer three to enhance dynamic adaptability; the other path uses continuous feature enhancement (CFE) modules to deepen detailed representation. The outputs from both paths are activated and convolved, and then pass through the ECA module again for secondary channel weight optimization. The number of parameters used in the dynamic convolution is approximately equal to that of a regular 2D convolution with the same parameters (input channels, output channels, kernel size, etc.), and the computational cost is only slightly larger than that of a regular convolution. All convolutional layers used in this network are 3×3 in size.

[0042] This module achieves effective interaction of multi-dimensional features through multi-scale feature fusion of the MCSF module and channel attention mechanism of the ECA module, and also strengthens channel-level key information in a targeted manner, effectively improving the expressive power of features and providing more accurate detailed and global feature support for super-resolution reconstruction.

[0043] To ensure network robustness, retain more structural information, and fully extract image features, this invention uses three activation functions: T-function, L-function, and S-function, defined as follows: In step 3, the appropriate evaluation metrics selected are the Learned Perceptual Patch Similarity Index (LPISP) ​​and the Natural Image Quality Evaluation Model (NIQE). The LPISP is a perceptual similarity evaluation metric based on deep features. It extracts image features through a pre-trained deep neural network and calculates the distance between image patches in the feature space to measure the perceptual similarity between images. NIQE uses a multivariate Gaussian model to establish a probability distribution of natural image features on a distortion-free training set and uses this distribution to evaluate the image quality score, assessing the quality stability of the reconstructed image in a referenceless manner. The LPISP and NIQE are defined as follows: in, Features of the l-th layer of the pre-trained network; To perform channel normalization on the features; The channel weights are learned.

[0044] The training iterations are set to 500. The learning rate for the first 300 training iterations is set to 0.0001, and the learning rate for the next 200 training iterations is gradually reduced from 0.0001 to 0. The upper limit of the number of images input to the network each time is mainly determined by the performance of the computer's graphics processor. Generally, the number of images input to the network each time is in the range of 10-20, which can make the network training more stable and the training results better, and can ensure that the network fits quickly.

[0045] In step 3, the loss function is calculated based on the network output and labels. Minimizing this loss function achieves better results. During training, a composite loss function is used: frequency domain consistency loss is employed for the fuzzy kernel estimation network, while total variation loss, texture consistency loss, and local self-similarity loss are used for the image super-resolution reconstruction network.

[0046] Total variation loss Introduced as a regularization constraint, it achieves noise reduction while preserving image edge information. Specifically, it minimizes the variation corresponding to the noise by calculating the differences between image pixels in the horizontal and vertical directions. Its formula is defined as: in, The feature map represents the pixel location. eigenvalues ​​at that location The amount is small; the introduction of this loss can help improve the clarity of the final HR image while avoiding blurring of edge information.

[0047] Frequency consistency loss This is used to constrain the model to maintain consistency of structural details in the frequency domain. By transforming the image to the frequency domain, the model maintains spectral consistency at low frequencies and preserves texture details at high frequencies, thereby improving reconstruction accuracy and perceptual quality. Its formula is defined as: in, This indicates the Fourier transform operation. It is the feature map corresponding to the LR image. These are intermediate feature maps from the reconstruction process. The L1 norm is used to introduce this loss, which enhances the consistency between the super-resolution results and the original LR image at the frequency level, and improves the realism of texture details and the sharpness of edges.

[0048] Local self-similarity loss Initial features used to constrain the output of the feature extraction module The structural similarity between the initial feature and the intermediate feature output by each stage module in the local region is calculated; local neighborhood blocks of the initial feature and intermediate feature are selected respectively, and the similarity difference between the blocks is calculated, with the formula defined as: in, This indicates selecting the i-th local neighborhood block from the feature map. For local feature embedding mapping (such as convolution operation), N is the number of local blocks; this loss enhances the consistency of repetitive textures in remote sensing images by maintaining the self-similarity of local regions, while improving the detail fidelity of the super-resolution results.

[0049] Texture consistency loss Initial features used to constrain the output of the feature extraction module The matching of intermediate features output by the module in terms of texture distribution. Using the initial features as the texture baseline, texture descriptors are extracted from the intermediate features at each stage, and the differences between the descriptors are calculated. The formula is defined as: in, This indicates the extraction of the j-th texture region block from the feature map. For texture descriptor extraction, M is the number of texture regions; this loss is introduced to compare features at each stage in pairs, thereby ensuring the consistency of remote sensing image texture information throughout the process and improving the texture realism of the super-resolution results.

[0050] In step 4, the model is trained and fine-tuned using the UCMerced dataset. For the UCMerced dataset, 1000 images are used for training, and 100 images are used for testing.

[0051] In step 5, after the network training is completed, all parameters in the network are saved. The remote sensing image to be super-resolution is then input into the network to obtain the reconstructed image.

[0052] Example 2 like Figure 9 As shown, this embodiment provides a remote sensing image super-resolution reconstruction device for real-world scenes, capable of executing the above-described method. The device includes: Image acquisition module: Used to load datasets, responsible for loading the remote sensing image datasets to be processed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded remote sensing images, including a fuzzy kernel estimation network and an image super-resolution reconstruction network. The fuzzy kernel estimation network includes a frequency domain feature extraction module and a fuzzy kernel generation module; the image super-resolution reconstruction network consists of an attention-guided multi-feature fusion module and a dynamic attention channel feature enhancement module, and obtains the final super-resolution model through iteration. Image output module: Used to display the reconstructed super-resolution image and output the reconstructed super-resolution image using electronic devices.

[0053] The entire process begins with the image acquisition module, whose task is to load the dataset containing the remote sensing images to be processed. The goal of this module is to prepare the image data for subsequent processing. Next is the image processing module, which, once the images are loaded, preprocesses them, resizing them to 256×256 pixels to ensure consistency and standardization of the input images. This standard size provides a consistent benchmark for subsequent processing. The third module is the image reconstruction module, which is the core of the entire process. It includes two main networks: a blur kernel estimation network and an image super-resolution reconstruction network. These networks process the loaded and preprocessed images through iterative training. The blur kernel estimation network is responsible for generating a degraded representation of the low-resolution image, while the super-resolution network focuses on reconstructing a high-resolution image from the low-resolution image. The goal of this module is to generate an optimized super-resolution model. Finally, the image output module is responsible for displaying the reconstructed super-resolution image. This module uses electronic devices to output the processed high-resolution image, allowing users to observe and evaluate the final result. These four modules work together to form a process that, through preprocessing, training, and output, ultimately achieves blind super-resolution processing of remote sensing images.

[0054] This invention constructs a method and apparatus for super-resolution reconstruction of remote sensing images for real-world scenarios. It can directly generate high-resolution remote sensing images from low-resolution images without intermediate steps. Under the same conditions, the feasibility and superiority of this method are further verified by calculating relevant indices of the images obtained using existing methods. A comparison of relevant indices between existing technologies and the method proposed in this invention is provided below. Figure 10 As shown.

[0055] from Figure 10 It can be seen that the method proposed in this invention has a higher perceptual similarity index and perceptual index than existing methods. These indicators further demonstrate that the method proposed in this invention has better image super-resolution reconstruction quality.

[0056] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for super-resolution reconstruction of remote sensing images for real-world scenes, characterized in that: Includes the following steps: Step 1: Construct a network including a fuzzy kernel estimation network and an image super-resolution reconstruction network; Step 2: Split the training and test sets using the DIV2K dataset and preprocess them; fine-tune the model using the UCMerced dataset; Step 3: Input the dataset prepared in Step 2 into the network model built in Step 1 for training until the number of training iterations reaches the initial threshold or the value of the loss function reaches the preset range. At this point, the network model is considered to have been trained and the network model parameters are saved. Step 4: Retrain and fine-tune the network model using the UCMerced dataset to optimize the network model parameters, further improve the image super-resolution reconstruction performance, and obtain a network model that can achieve the best results. Step 5: Solidify the parameters of the final model, save the model, and the model can output a corresponding high-resolution remote sensing image after inputting a low-resolution remote sensing image.

2. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 1, the fuzzy kernel estimation network includes a frequency domain feature extraction module and a fuzzy kernel estimation module; the super-resolution reconstruction network consists of an attention-guided multi-feature fusion network and a dynamic attention channel feature enhancement network.

3. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 1, the fuzzy kernel estimation module includes a frequency domain feature extraction module and a fuzzy kernel generation module; The frequency domain feature extraction module utilizes frequency domain transformation and optical transfer function to extract frequency domain information, then performs inverse transformation back to the spatial domain while retaining key frequency features, fully simulating the real degradation law of the imaging system; the blur kernel generation module uses max pooling and convolutional layers to reduce the dimensionality of the frequency domain degradation representation, and then obtains the degradation blur kernel through upsampling and double-layer convolutional blocks; the feature extraction module uses convolutional layers and residual blocks to effectively learn and represent image features; The reconstruction module utilizes a multi-stage channel separation feature fusion module (MCSF) to perform layered extraction and fusion of input features along the channel dimension. It combines dynamic convolution to achieve adaptive transformation of features, and then uses a channel feature enhancement module (CFE) to strengthen the effective information of the channel dimension. Finally, the remote sensing image is reconstructed through layer-by-layer processing of convolutional layers and activation functions. The adjuster network uses a dual-path feature mapping structure to refine the input blur kernel, and then outputs two-way features through residual connections.

4. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 2, the prepared dataset is preprocessed, the size of each image in the dataset is adjusted, and the size of the input image is fixed as the network input.

5. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 3, a composite loss function is selected for the loss function selection. The fuzzy kernel estimation network adopts frequency consistency loss; the image super-resolution reconstruction network adopts total variation loss, local self-similarity loss and texture consistency loss. The selection of loss function affects the quality of the model, can truly reflect the difference between the predicted value and the true value, and can correctly reflect the quality of the model.

6. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 3, the training of the network model also includes evaluating the quality and degree of image distortion of the remote sensing image reconstruction results through evaluation metrics, and measuring the role of the super-resolution reconstruction network.

7. The method for super-resolution reconstruction of remote sensing images for real-world scenes according to claim 1, characterized in that: In step 4, the network is trained using the DIV2K dataset to enhance its robustness.

8. A remote sensing image super-resolution reconstruction device for real-world scenes, characterized in that: The device includes: Image acquisition module: used to load datasets, the loading items being remote sensing images to be preprocessed; Image processing module: used to preprocess the loaded remote sensing images, adjusting the size of each image to 256×256 pixels to ensure that the size of the input image remains constant; Image reconstruction module: used to train the loaded and processed remote sensing images, including a blur kernel estimation network and an image super-resolution reconstruction network, and obtains the final super-resolution model through iteration; Image output module: Used to display the reconstructed super-resolution image and output the reconstructed high-resolution image using electronic devices.

Citation Information

Patent Citations

  • Remote Sensing Image Super-Resolution Reconstruction Method Based on Degradation Mechanism

    CN119494781B