Double-domain remote sensing image super-resolution method

By combining the dual-domain strategy with the SwinTransformer architecture, the problems of insufficient remote sensing image resolution and frequency domain processing distortion are solved, efficient super-resolution reconstruction of remote sensing images is achieved, and image quality and robustness are improved.

CN120833260APending Publication Date: 2025-10-24TAIYUAN UNIVERSITY OF TECHNOLOGY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510958520.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies have problems with insufficient resolution and poor quality of remote sensing images, especially the inability to accurately restore the edge and texture features of remote sensing images, and deep learning network models have data anomaly distortion problems in frequency domain processing.

Method used

A dual-domain remote sensing image super-resolution method is adopted. The remote sensing image is converted from the image domain to the frequency domain through the domain conversion operator. The frequency domain features are optimized using a complex neural network. Combined with the SwinTransformer architecture, super-resolution reconstruction is performed in the image domain. The dual-domain data consistency constraint is introduced to optimize the training process.

Benefits of technology

It achieves efficient super-resolution reconstruction of remote sensing images, preserves the high-level semantic information and low-level statistical features of the image, improves the fidelity of image details and reconstruction quality, shows robustness and generalization ability, and is applicable to a variety of downsampling factors and datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833260A_ABST
    Figure CN120833260A_ABST
Patent Text Reader

Abstract

The invention provides a double-domain remote sensing image super-resolution method, and belongs to the technical field of super-resolution algorithms in remote sensing image processing. According to the method, a remote sensing image is converted from an image domain to a frequency domain through a domain conversion operator module, phase and amplitude information is reserved, and feature optimization is conducted on frequency domain data through a complex value neural network module. And the optimized frequency domain data is converted back to an image domain through an inverse domain conversion operator, and meanwhile, super-resolution reconstruction is carried out in the image domain in combination with a SwinTransform architecture. And the double-domain data consistency constraint module is used for constraining the training optimization direction through the loss function design of Fourier transform. By adopting the double-domain remote sensing image super-resolution method, remote sensing image super-resolution reconstruction under a double-domain strategy is realized, the problem of low availability of low-resolution remote sensing satellite image data is effectively solved, and the application scene of the low-resolution remote sensing image data is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of super-resolution algorithm in remote sensing image processing, and particularly relates to a dual-domain remote sensing image super-resolution method. BACKGROUND

[0002] Geographical information provides support for human activities in the fields of natural resources, agriculture and industry, and the like, and remote sensing image data, as common geographical information, is of great importance to remote sensing imaging applications in terms of high quality and high resolution. However, remote sensing images are acquired by imaging sensors carried by remote sensing platforms, and the differences in sensor types and imaging principles determine that the imaging schemes basically include photographic imaging, scanning imaging and radar imaging. Complex natural factors and limitations of sensors make remote sensing applications face great challenges in spatial resolution, resulting in data defects such as poor quality and low resolution of remote sensing images, which need to be solved urgently.

[0003] In order to overcome the negative effects of insufficient resolution of remote sensing images on human activities, traditional algorithms based on interpolation are proposed to improve the blurring caused by image enlargement. The interpolation-based method has significant advantages in terms of image processing speed and algorithm complexity, but the defects of this method are also obvious. When the image pixel details change greatly, the processing scheme cannot maintain the original edges and texture features of the image, and cannot accurately restore the original high-frequency information of the image.

[0004] Image reconstruction is another method for low-resolution image restoration, and remote sensing image super-resolution reconstruction is a common application field thereof. Good reconstruction effects are usually obtained by increasing constraint conditions and introducing prior information, but the convergence of the model combined with prior knowledge is not ideal.

[0005] In recent years, thanks to the improvement of the underlying visual capabilities of computers by convolution operation, the research on remote sensing image super-resolution reconstruction based on deep learning has begun to rapidly advance. However, the work of exploring the potential of frequency domain features of remote sensing images has not been thoroughly studied, and the scheme of exploring the potential of frequency domain features of remote sensing images and improving the data abnormal distortion of the deep learning network model needs to be proposed and verified. SUMMARY

[0006] The purpose of the present application is to provide a dual-domain remote sensing image super-resolution method, which realizes remote sensing image super-resolution reconstruction under a dual-domain strategy, effectively solves the problem of low availability of low-resolution remote sensing satellite image data, and expands the application scenarios of low-resolution remote sensing image data.

[0007] To achieve the above purpose, the present application provides a dual-domain remote sensing image super-resolution method, comprising:

[0008] A domain conversion operator module for converting a remote sensing image from an image domain to a frequency domain while preserving phase and amplitude information, comprising:

[0009] a two-dimensional discrete Fourier transform unit, performing two-dimensional discrete Fourier transform on the input remote sensing image to generate frequency domain complex data;

[0010] a zero frequency component shift unit, performing zero frequency component shift and spectrum visualization on the frequency domain complex data;

[0011] a complex-valued neural network module for optimizing the phase and amplitude characteristics of the frequency domain complex data, comprising:

[0012] a batch normalization unit for covariance matrix normalization of the complex data;

[0013] a complex convolution unit for complex feature extraction through a complex convolution kernel;

[0014] a complex activation unit for activation processing of the complex convolution result for inputting the next layer of the model for continuous training;

[0015] a complex linear connection layer for adjusting the output size;

[0016] a SwinTransformer architecture for image domain super-resolution reconstruction, comprising:

[0017] a shallow feature extraction unit for extracting shallow features of the input image through 3x3 convolution;

[0018] a plurality of residual SwinTransformer sub-modules, each residual SwinTransformer sub-module comprising six SwinTransformer layers and a convolution unit, each SwinTransformer layer comprising a multi-head attention unit, a multi-layer perceptron unit, and a layer normalization unit; the multi-layer perceptron unit comprising two fully connected layers and a GELU layer in between;

[0019] a dual-domain data consistency constraint module based on a Fourier transform loss function design, which constrains and optimizes the training direction through bidirectional conversion between the frequency domain and the image domain.

[0020] Preferably, the two-dimensional discrete Fourier transform unit independently performs the following operations on each channel of the RGB three-channel remote sensing image:

[0021]

[0022] wherein, represents the complex value of the corresponding channel in the frequency domain, (u, v) represents the frequency in the horizontal and vertical directions, represents the pixel value of the coordinate variable (x, y) in a channel, and j represents the imaginary unit, represents the image frequency domain representation after Fourier transform, fft represents Fourier transform, and I represents original image data.

[0023] Preferably, the zero frequency component shifting unit realizes spectrum centering by the following operation:

[0024]

[0025] wherein, represents the frequency domain image data after the shift operation, and Shift represents the shift operation.

[0026] Preferably, the complex convolution unit performs the following operation on the input complex number z = a + bi by using a complex convolution kernel ω = A + iB:

[0027] ω * z = (A * a - B * b) + i (B * a + A * b) ;

[0028] and expressed in matrix form as:

[0029]

[0030] wherein, A and a both represent real parts, B and b both represent imaginary parts, i represents an imaginary unit, (ω * z) (r) represents the real part of the complex convolution result, and is marked by r in the matrix form, (ω * z) (i) represents the imaginary part of the complex convolution result, and is marked by i in the matrix form, and * represents the convolution operation.

[0031] Preferably, the activation function of the complex activation unit is as follows:

[0032]

[0033] wherein, zReLU represents the activation function, and θ z represents the phase angle of the complex number z.

[0034] Preferably, the batch normalization unit normalizes the complex data by the following operation:

[0035]

[0036] wherein, represents the normalized complex number, E(z) represents the average value of the complex number, and C represents a 2x2 covariance matrix;

[0037] The normalized complex number is processed using a learnable shift parameter β and a scale parameter γ, and the complex value batch normalization is defined as:

[0038]

[0039] Preferably, the shallow feature extraction layer unit is composed of a 3x3 convolutional layer, given an image I of insufficient resolution LR , it is processed by a shallow feature extraction operation LFE to obtain shallow features f

[0040] f (0) = LFE(I LR );

[0041] where f (0) represents the shallow features used to input the first residual SwinTransformer sub-module.

[0042] Preferably, the multi-head attention unit works as follows:

[0043] Given the size of the input f (i) is HxWxC', the model first reshapes the input into a feature;

[0044] The self-attention calculation scheme of each window is as follows:

[0045] Generate query Q, key K and value V through linear transformation;

[0046]

[0047] Calculate the self-attention matrix;

[0048]

[0049] Parallelly execute h times of attention, and finally concatenate the multiple attention results;

[0050] where H, W represent the height and width of the feature respectively, C' represents the number of channels, M 2 represents the number of windows, P Q , P K and P V represent the shared projection matrices between different windows, represents the local window feature, SA(Q, K, V) represents the self-attention matrix of the local window feature, B represents the relative position encoding, and d represents the dimension of the query and key.

[0051] Preferably, the layer normalization unit adds a residual jump connection, denoted as:

[0052] f (i)′ = MSA(LN(f (i) ))+f (i) ;

[0053] f (i)″ = MLP(LN(f (i)′ ))+f(i)′ ;

[0054] wherein, f (i) denotes the original input feature, LN denotes layer normalization, MSA denotes multi-head self-attention mechanism, f (i)′ denotes the feature after multi-head self-attention mechanism processing and adding residual connection, f (i)″ denotes the feature after multi-layer perception processing and adding residual connection, and MLP denotes multi-layer perception.

[0055] Preferably, the loss function design based on Fourier transform constrains the training optimization direction through bidirectional conversion between the frequency domain and the image domain, and the specific operation is as follows:

[0056]

[0057] wherein, epoch denotes the range of training rounds, from the 1st round to the nth round, respectively denote Fourier transform operator and inverse operator, I (g) denotes the image generated in the gth round, and epsilon g denotes the learning rate, denotes the real image, denotes the square of L2 norm, denotes the loss value.

[0058] Therefore, the application adopts the above-mentioned dual-domain remote sensing image super-resolution method, and has the beneficial technical effects as follows:

[0059] (1) Synergistic advantage of dual-domain strategy:

[0060] Fusion of frequency domain and image domain: the application realizes seamless conversion between image domain and frequency domain through discrete Fourier transform operator, and innovatively expands the super-resolution reconstruction task of remote sensing image to the frequency domain. This cross-domain processing method not only retains the high-level semantic information and low-level statistical characteristics of the image, but also provides a new perspective and method for the enhancement of low-resolution remote sensing image data.

[0061] Deep mining of frequency domain features: in the frequency domain task, a deep complex network is used to recover the spectrum and phase information, fully exploiting the frequency domain feature potential of remote sensing images, effectively improving the data abnormal distortion problem of the deep learning network model in the frequency domain processing, and improving the fidelity of image details and reconstruction quality.

[0062] (2) Efficient reconstruction of image domain.

[0063] Efficiency of SwinTransformer architecture: In image domain tasks, the SwinTransformer network with residual is adopted, which utilizes its multi-head attention mechanism and local window feature extraction capability to significantly improve the performance of image resolution enhancement. When processing low-resolution remote sensing images, this architecture can better capture the local features and details of the image, generating high-quality high-resolution images.

[0064] Optimization of dual-domain data consistency constraints: For image domain models, dual-domain data consistency constraints are introduced to optimize the training of image domain models through frequency domain loss. This cross-domain constraint mechanism not only improves the distortion of image detail reconstruction, but also guides the model to maintain positive during the training optimization process from the cross-domain dimension, ensuring the convergence and stability of the model.

[0065] (3) Robustness and generalization ability:

[0066] Robustness verification: Experiments show that the present invention performs excellently in different data sets and various down-sampling factors (such as 2 times, 4 times, and 8 times), and the qualitative and quantitative results are better than existing technical solutions. This fully proves the robustness and reliability of the present invention method in processing low-resolution remote sensing images.

[0067] Generalization ability: The present invention also exhibits good reconstruction effect on data sets not involved in training, verifying its strong generalization ability. This means that the method is not only suitable for specific data sets, but also can play a role in a wider range of remote sensing image data, providing a general solution for remote sensing image super-resolution reconstruction. BRIEF DESCRIPTION OF DRAWINGS

[0068] Figure 1 is a structural diagram of the present invention, a dual-domain remote sensing image super-resolution method;

[0069] Figure 2 is a comparison experiment effect display diagram under the NWPU-RESISC45 data set, the illustration category is airplane, and the scale factor is 2, wherein, Figure 2 (a) in is the Bicubic scheme, Figure 2 (b) in is the SRCNN scheme, Figure 2 (c) in is the VDSR scheme, Figure 2 (d) in is the WTCRR scheme, Figure 2 (e) in is the DSSR scheme, Figure 2 (f) in is the PCRBPN scheme, Figure 2 (g) in is the WINet scheme, Figure 2 (h) in is the present invention method, Figure 2 (i) in is the reference image;

[0070] Figure 3 Figure 1 is a comparison experiment effect display diagram under the NWPU-RESISC45 dataset, the illustration category is airport runway, the scale factor is 4, wherein, Figure 3 (a) in the figure is a Bicubic scheme, Figure 3 (b) in the figure is an SRCNN scheme, Figure 3 (c) in the figure is a VDSR scheme, Figure 3 (d) in the figure is a WTCRR scheme, Figure 3 (e) in the figure is a DSSR scheme, Figure 3 (f) in the figure is a PCRBPN scheme, Figure 3 (g) in the figure is a WINet scheme, Figure 3 (h) in the figure is the method of the present application, Figure 3 (i) in the figure is a reference image;

[0071] Figure 4 Figure 2 is a comparison experiment effect display diagram under the NWPU-RESISC45 dataset, the illustration is storage tank, the scale factor is 8, wherein, Figure 4 (a) in the figure is a Bicubic scheme, Figure 4 (b) in the figure is an SRCNN scheme, Figure 4 (c) in the figure is a VDSR scheme, Figure 4 (d) in the figure is a WTCRR scheme, Figure 4 (e) in the figure is a DSSR scheme, Figure 4 (f) in the figure is a PCRBPN scheme, Figure 4 (g) in the figure is a WINet scheme, Figure 4 (h) in the figure is the method of the present application, Figure 4 (i) in the figure is a reference image;

[0072] Figure 5 Figure 3 is a robustness verification experiment effect display diagram under the UCMerced dataset, the illustration category is shrub, the scale factor is 2, wherein, Figure 5 (a) in the figure is a Bicubic scheme, Figure 5 (b) in the figure is an SRCNN scheme, Figure 5 (c) in the figure is a VDSR scheme, Figure 5 (d) in the figure is a WTCRR scheme, Figure 5 (e) in the figure is a DSSR scheme, Figure 5 (f) in the figure is a PCRBPN scheme, Figure 5 (g) in the figure is a WINet scheme, Figure 5 (h) in the figure is the method of the present application, Figure 5 (i) in the figure is a reference image;

[0073] Figure 6 Figure 6 is a diagram showing the results of a robustness verification experiment performed on the UCMerced dataset, with a scale factor of 4, and an illustration category of overpass, wherein, Figure 6 (a) in Figure 6 is a Bicubic scheme, Figure 6 (b) in Figure 6 is an SRCNN scheme, Figure 6 (c) in Figure 6 is a VDSR scheme, Figure 6 (d) in Figure 6 is a WTCRR scheme, Figure 6 (e) in Figure 6 is a DSSR scheme, Figure 6 (f) in Figure 6 is a PCRBPN scheme, Figure 6 (g) in Figure 6 is a WINet scheme, Figure 6 (h) in Figure 6 is the method of the present application, Figure 6 (i) in Figure 6 is a reference image;

[0074] Figure 7 Figure 7 is a diagram showing the results of a robustness verification experiment performed on the UCMerced dataset, with a scale factor of 8, and an illustration category of golf course, wherein, Figure 7 (a) in Figure 7 is a Bicubic scheme, Figure 7 (b) in Figure 7 is an SRCNN scheme, Figure 7 (c) in Figure 7 is a VDSR scheme, Figure 7 (d) in Figure 7 is a WTCRR scheme, Figure 7 (e) in Figure 7 is a DSSR scheme, Figure 7 (f) in Figure 7 is a PCRBPN scheme, Figure 7 (g) in Figure 7 is a WINet scheme, Figure 7 (h) in Figure 7 is the method of the present application, Figure 7 (i) in Figure 7 is a reference image;

[0075] Figure 8 Figure 8 is a diagram showing the results of a consistency constraint module verification experiment performed on the NWPU-RESISC45 dataset, with a scale factor of 2, and an illustration category of traffic trunk, wherein, Figure 8 (a) in Figure 8 is the method of the present application without a consistency constraint, Figure 8 (b) in Figure 8 is the method of the present application with a consistency constraint, Figure 8 (c) in Figure 8 is a reference image;

[0076] Figure 9 Figure 9 is a diagram showing the loss results of a consistency constraint module verification experiment performed on the NWPU-RESISC45 dataset, wherein, Figure 9 (a) in Figure 9 is the method of the present application without a consistency constraint, Figure 9 (b) in Figure 9 is the method of the present application with a consistency constraint;

[0077] Figure 10 Effect verification of the influence of the number of residual SwinTransformer blocks on the PSNR of the model reconstructed image, wherein, Figure 10 (a) in the above is an experimental result based on the NWPU-RESISC45 dataset, Figure 10 (b) in the above is an experimental result based on the UCMerced dataset;

[0078] Figure 11 Effect verification of the influence of the number of residual SwinTransformer blocks on the SSIM of the model reconstructed image, wherein, Figure 11 (a) in the above is an experimental result based on the NWPU-RESISC45 dataset, Figure 11 (b) in the above is an experimental result based on the UCMerced dataset;

[0079] Figure 12 From left to right, the low-resolution image, the image processed by the CertSR scheme and the processing, the reference image, and the results after the semantic segmentation processing of these images are sequentially arranged, and the data source is a remote sensing image of a city in a certain city. DETAILED DESCRIPTION

[0080] The technical solutions of the present application are further described below by means of the accompanying drawings and examples.

[0081] Unless otherwise defined, the technical terms or scientific terms used in the present application shall have the usual meanings understood by persons having ordinary skills in the art to which the present application belongs.

[0082] Example 1

[0083] A dual-domain remote sensing image super-resolution method, comprising a domain conversion operator module, a complex neural network module, a SwinTransformer architecture, and a dual-domain data consistency constraint module. The overall process diagram is shown in the first part of Figure 1 , wherein the domain conversion operator module is shown in the second and third parts of Figure 1 , the complex neural network module is shown in the fourth part of Figure 1 , the SwinTransformer architecture and the dual-domain data consistency constraint module are shown in the fifth part of Figure 1 , and the implementation process comprises the following steps:

[0084] A, remote sensing image downsampling strategy setting and data processing.

[0085] Remote sensing image downsampling strategy and data processing is a key step in remote sensing data analysis, especially in the case of large amount of data, downsampling can reduce the redundancy of data, reduce the computational burden, while keeping the key information of the image. But downsampling will lead to the loss of part of the high frequency information, so we need to pay attention to the balance between data compression and information preservation when downsampling, and often give up the requirement of data compression in the actual work due to the pursuit of information preservation. In order to simulate low resolution remote sensing image data, the original high resolution data is preprocessed and the data set is constructed, and bicubic interpolation is used to resample the remote sensing data as follows:

[0086] I LR = Bicubic(I, Scale), Scale ∈ [2, 4, 8]

[0087] where I LR is the image data after reducing the resolution, I is the original image data, and Scale is the resampling factor, which affects the resolution of the resampled image. Here, the resampling factor is set to 2, 4 and 8. After downsampling the original remote sensing image, the resolution of the original remote sensing image is reduced to 1 / 2, 1 / 4 and 1 / 8, the size is reduced, and the pixel information contained in the image is greatly reduced. In order to construct the same size input data and label data for model training, bicubic interpolation is used to upsample the downsampled remote sensing image as follows:

[0088] I' = Bicubic(I LR , Scale), Scale ∈ [0.5, 0.25, 0.125]

[0089] where I' is the image data with restored size, and Scale is still the resampling factor here. Here, the resampling factor is set to 0.5, 0.25 and 0.125 to restore the size of the downsampled image to the original size. Finally, Gaussian noise is added to the image to simulate random noise in low resolution imaging:

[0090] I * = I' + ∈, ∈ ~ N(0, σ 2 = 0.01)

[0091] where I * is the final image data, ∈ is the random added Gaussian noise, and the mean is 0, and the variance σ 2 is 0.01. The final image is a low resolution, size invariant simulated low resolution remote sensing image. The original remote sensing image and the processed remote sensing image are one-to-one corresponding in the data set, which constitutes the label and model input data, so as to construct the data set.

[0092] B, based on the domain conversion operator module to convert data.

[0093] a domain conversion operator module for converting a remote sensing image from an image domain to a frequency domain while preserving phase and amplitude information, comprising:

[0094] B1, a two-dimensional discrete Fourier transform unit, performing a two-dimensional discrete Fourier transform on an input remote sensing image to separate the RGB three channels and generate frequency domain complex data;

[0095]

[0096] wherein, represents the complex value of the corresponding channel in the frequency domain, (u, v) represents the frequency in the horizontal and vertical directions, represents the pixel value of the coordinate variable (x, y) in a channel, j represents an imaginary unit, represents the frequency domain representation of the image after Fourier transform, fft represents the Fourier transform, and I represents the original image data.

[0097] B2, a zero frequency component shift unit, shifting the frequency domain complex data and visualizing the spectrum.

[0098]

[0099] wherein, represents the frequency domain image data after the shift operation, and Shift represents the shift operation.

[0100] C, a complex-valued neural network module for optimizing amplitude and phase data.

[0101] The complex-valued neural network module is used to optimize the phase and amplitude characteristics of the frequency domain complex data, comprising:

[0102] A batch normalization unit normalizes the covariance matrix of the complex data;

[0103] The batch normalization unit normalizes the complex data by the following operation:

[0104]

[0105] wherein, represents the normalized complex number, E(z) represents the mean value of the complex number, and C represents a 2x2 covariance matrix;

[0106] The normalized complex number is processed using a learnable shift parameter β and a scale parameter γ, and the complex batch normalization is defined as:

[0107]

[0108] The complex convolution unit realizes complex feature extraction through a complex convolution kernel.

[0109] The complex convolution unit performs the following operation on the input complex number z = a + bi by the complex convolution kernel ω = A + iB:

[0110] ω * z = (A * a - B * b) + i (B * a + A * b) ;

[0111] And expressed in matrix form:

[0112]

[0113] Wherein, A, a represent real part, B, b represent imaginary part, i represents imaginary unit, (ω * z) (r) The real part of the complex convolution result is represented by r in the matrix form, (ω * z) (i) The imaginary part of the complex convolution result is represented by i in the matrix form, * represents convolution operation.

[0114] The complex activation unit activates the complex convolution result for the next layer of the input model to continue training.

[0115] The activation function of the complex activation unit is as follows:

[0116]

[0117] Wherein, zReLU represents the activation function, θ z The phase angle of the complex number z.

[0118] The complex linear connection layer is used to adjust the output size.

[0119] The complex-valued neural network is trained and parameter archived using the Fourier domain data set obtained in step B, and the input of the Fourier domain data set is processed again by reading the parameters, and the output data is the input of the new data set, and the corresponding label is still the original Fourier domain label.

[0120] D. Convert the new Fourier domain data set to the real number domain data set based on the domain conversion operator.

[0121] Based on the inverse process of the domain conversion operator, the new Fourier domain data set is converted to the image domain (which can also be called the real number domain in the current specific task scenario) data. This facilitates the second step of the entire method framework. After the inverse process of the domain conversion operator, the new Fourier domain data set is converted to the new real number domain data set, and the input and label correspond respectively.

[0122] E. Construct the double-domain consistency constraint of the image domain subnetwork model.

[0123] The image domain subnetwork model is a SwinTransformer architecture model. The conventional target setting for training the model is generally to minimize the mean square error or cumulative average error between the output and the label, which is essentially an implementation process based on a gradient descent algorithm. Unlike gradient descent, the method of the present application contains an innovative loss design, which is based on the principle of calculating the loss of the output data and the label during the training process, converting the obtained loss data to the Fourier domain through a domain conversion operator, minimizing the target in the Fourier domain, and simultaneously performing loss backpropagation during the training process based on the principle of minimizing the Fourier domain loss target:

[0124]

[0125] wherein, is the loss value, and are Fourier conversion operators and inverse conversion operators, I (g) represents the image generated in the gth round, represents the real image, represents the square of the L2 norm. This is used as a loss function for constraining the training of the image domain subnetwork module. A large number of experimental results show that the implementability and effectiveness of this scheme far exceed the benefits brought by conventional loss functions.

[0126] The optimization target is achieved under the dual-domain data consistency constraint algorithm. The dual-domain data consistency constraint algorithm for remote sensing images is as follows:

[0127]

[0128] wherein, epoch represents the range of training rounds, from the 1st round to the nth round, respectively represent the Fourier conversion operator and the inverse operator, I (g) represents the image generated in the gth round, ε g represents the learning rate, represents the real image.

[0129] F, training and testing of the SwinTransformer architecture model based on a new real number domain data set.

[0130] The SwinTransformer architecture model is trained and tested by using a new real number domain data set, and the model is specifically trained and parameter-archived by combining the dual-domain consistency constraint in the method of the present application. In subsequent testing, the archived parameters are read and the images with insufficient resolution are input into the model, and a high-resolution image almost identical to the original image can be output.

[0131] SwinTransformer architecture for image domain super-resolution reconstruction, comprising:

[0132] The shallow feature extraction unit extracts shallow features of the input image through a 3x3 convolution.

[0133] The plurality of residual SwinTransformer sub-modules each include six SwinTransformer layers and a convolution unit, and each SwinTransformer layer includes a multi-head attention unit, a multi-layer perception unit, and a layer normalization unit. The multi-layer perception unit includes two fully connected layers and a GELU layer in between.

[0134] The shallow feature extraction layer unit is composed of a 3x3 convolution layer. Given an image I LR , the shallow feature f (0) is obtained by processing the image I LR through a shallow feature extraction operation LFE.

[0135] f (0) = LFE(I LR );

[0136] wherein f (0) represents the shallow feature used to input the first residual SwinTransformer sub-module.

[0137] The multi-head attention unit works as follows:

[0138] Given the size of the input f (i) as HxWxC', the model first reshapes the input into a feature;

[0139] The self-attention calculation scheme of each window is as follows:

[0140] Generate query Q, key K, and value V through linear transformation;

[0141]

[0142] Calculate the self-attention matrix;

[0143]

[0144] Parallelly execute h times of attention, and finally concatenate the multiple attention results;

[0145] wherein H and W represent the height and width of the feature respectively, C' represents the number of channels, M 2 represents the number of windows, P Q , P K , and P V represent shared projection matrices between different windows, denotes the local window feature, SA(Q, K, V) denotes the self-attention matrix of the local window feature, B denotes the relative position encoding, and d denotes the dimension of the query and the key.

[0146] Layer normalization unit adds a residual skip connection, denoted as:

[0147] f (i)′ = MSA(LN(f (i) ))+f (i) ;

[0148] f (i)″ = MLP(LN(f (i)′ ))+f (i)′ ;

[0149] wherein f (i) denotes the original input feature, LN denotes layer normalization, MSA denotes multi-head self-attention mechanism, f (i)′ denotes the feature processed by the multi-head self-attention mechanism and added with the residual connection, f (i)″ denotes the feature processed by the multi-layer perception and added with the residual connection, and MLP denotes multi-layer perception.

[0150] G, verification work after all training work is completed.

[0151] Taking the simulation data set as an example, read and load the complex-valued neural network model, input the simulation low-resolution image for verification into the trained complex-valued neural network after processing by the domain conversion operator, and output the Fourier domain intermediate data for further processing. The Fourier domain intermediate data is processed by the inverse process of the Fourier conversion operator to obtain the image domain intermediate data and input the SwinTransformer architecture model under the constraint of dual-domain consistency, and output the final result image with high resolution and high quality. In practical applications, the input data in the simulation data set is replaced by the low-resolution remote sensing image collected by a satellite, and a high-quality image with high resolution can be obtained.

[0152] In order to further illustrate the effectiveness of the method of the present application, experiments were conducted on two remote sensing image data sets, NWPU-RESISC45 and UCMerced, and the specific experimental methods and results are as follows:

[0153] Experimental method.

[0154] 1. Data set preparation: adopt NWPU-RESISC45 and UCMerced data sets, and perform 2 times, 4 times and 8 times downsampling processing, respectively, to construct a low-resolution remote sensing image data set.

[0155] 2. Experimental setup: The low-resolution remote sensing images are super-resolution reconstructed using the above method, and compared with various existing methods, including Bicubic, SRCNN, VDSR, WTCRR, DSSR, PCRBPN and WINet, etc.

[0156] 3. Evaluation index: The reconstruction results are evaluated by visual effect and quantitative index (such as PSNR, SSIM).

[0157] Experimental results.

[0158] NWPU-RESISC45 dataset experimental results:

[0159] Aircraft category, Scale factor is 2: Figure 2 The reconstruction effect comparison of different methods is shown. It can be seen that the method of the application is superior to other methods in detail recovery and edge fidelity.

[0160] Airport runway category, Scale factor is 4: Figure 3 The performance of the method of the application in the 4-fold downsampling task is shown. Compared with other methods, the image generated by the method of the application is clearer and more detailed.

[0161] Storage Tank category, Scale factor is 8: Figure 4 It is shown that the method of the application can still effectively recover image details in high-multiplying downsampling tasks and is superior to other comparison methods.

[0162] UCMerced dataset experimental results:

[0163] Shrub category, Scale factor is 2: Figure 5 It is used to prove the robustness of the method of the application in the 2-fold downsampling task, and the generated image is superior to other methods in visual effect.

[0164] Interchange category, Scale factor is 4: Figure 6 The robustness of the method of the application in the 4-fold downsampling task is shown, and the image details and structural integrity are significantly improved.

[0165] Golf course category, Scale factor is 8: Figure 7 The robustness of the method of the application in the 8-fold downsampling task is verified, and the generated image is superior to other methods in visual effect.

[0166] Consistency constraint module verification.

[0167] Traffic trunk category, Scale factor is 2: Figure 8The image restoration effects of the method of the present application with and without consistency constraint are compared. The results show that the consistency constraint module significantly improves the image detail restoration capability.

[0168] Loss effect diagram: Figure 9 The loss changes of the method of the present application with and without consistency constraint are shown. The consistency constraint module effectively reduces semantic distortion and improves image quality.

[0169] Residual SwinTransformer block quantity verification:

[0170] PSNR influence effect: Figure 10 The influence of the number of residual SwinTransformer blocks on the PSNR of the reconstructed image of the model is shown. The results show that the current block quantity can achieve optimal performance.

[0171] SSIM influence effect: Figure 11 The influence of the number of residual SwinTransformer blocks on the SSIM of the reconstructed image of the model is shown. The results show that the current block quantity performs best in visual quality.

[0172] Downstream task application verification:

[0173] Semantic segmentation application: Figure 12 The application effect of the low-resolution remote sensing image processed by the method of the present application in the semantic segmentation task is shown. The results show that the high-resolution image generated by the method of the present application can effectively improve the performance of the downstream task, and help target recognition and classification of remote sensing images.

[0174] From the above experimental results, it can be seen that the method of the present application performs excellent performance in different data sets and various down-sampling factor task scenarios, and the qualitative and quantitative results are better than the prior art. This fully proves the robustness and reliability of the method of the present application in processing low-resolution remote sensing images, and verifies its strong generalization ability.

[0175] It is worth noting that the contents not elaborated in the present application are all prior art and are well known to those skilled in the art.

[0176] Therefore, the above-mentioned dual-domain remote sensing image super-resolution method is adopted, the remote sensing image super-resolution reconstruction under the dual-domain strategy is realized, the low usability problem of low-resolution remote sensing satellite image data is effectively solved, and the application scenarios of low-resolution remote sensing image data are expanded.

[0177] It should be pointed out finally that the above examples are only used to illustrate the technical solutions of the present application but not to limit it, and although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can still be modified or replaced equivalently, and these modifications or equivalent replacements should not make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.

Claims

1. A dual-domain remote sensing image super-resolution method, characterized in that, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture.

2. The dual-domain remote sensing image super-resolution method of claim 1, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. wherein, represents the complex value of the corresponding channel lower frequency domain, (u, v) represents the frequency in horizontal and vertical directions, represents the pixel value on the coordinate variable (x, y) in one channel, j represents the imaginary unit, represents the image frequency domain representation after Fourier transform, fft represents Fourier transform, and I represents the original image data.

3. The dual-domain remote sensing image super-resolution method of claim 2, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. wherein represents the frequency domain image data after the shift operation, and Shift represents the bit shift operation.

4. The dual-domain remote sensing image super-resolution method of claim 3, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. where A, a represent real parts, B, b represent imaginary parts, i represents an imaginary unit, (ω * z) (r) represents a real part of a complex convolution result, identified by r in a matrix form, (i) represents an imaginary part of a complex convolution result, identified by i in a matrix form, * represents a convolution operation.

5. The dual-domain remote sensing image super-resolution method of claim 4, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. where zReLU denotes an activation function, denotes the phase angle of a complex number z.

6. The dual-domain remote sensing image super-resolution method of claim 5, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. wherein denotes the normalized complex, E(z) denotes the mean value of the complex, and C denotes a 2x2 covariance matrix; The standardized complex is processed using a learnable shift parameter β and scale parameter γ, complex batch normalization is defined as:

7. The dual-domain remote sensing image super-resolution method of claim 6, wherein, The shallow feature extraction layer unit is composed of a 3x3 convolutional layer, given an image I of insufficient resolution LR is processed by a shallow feature extraction operation LFE to obtain shallow features: f (0) = LFE(I LR ); wherein f (0) represents a shallow feature for inputting the first residual SwinTransformer submodule.

8. The dual-domain remote sensing image super-resolution method of claim 7, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. Given input f (i) of size H x W x C', the model first reshapes the input into a of size 1 x H x W x C' The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. where H, W represent the height and width of the feature respectively, C' represents the number of channels, M 2 represents the number of windows, P Q , P K and P V represent the shared projection matrix between different windows respectively, represents the local window feature, SA(Q, K, V) represents the self-attention matrix of the local window feature, B represents the relative position encoding, and d represents the dimension of the query and the key.

9. The dual-domain remote sensing image super-resolution method of claim 8, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. f (i)' = MSA(LN(f (i) ))+ f (i) ; f (i)″ = MLP(LN(f (i)′ ))+f (i)′ ; wherein f (i) denotes the original input feature, LN denotes layer normalization, MSA denotes multi-head self-attention mechanism, f (i)′ denotes the feature processed by the multi-head self-attention mechanism and added with the residual connection, f (i)″ denotes the feature processed by the multi-layer perceptron and added with the residual connection, and MLP denotes the multi-layer perceptron.

10. The dual-domain remote sensing image super-resolution method of claim 9, wherein, The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and a SwinTransformer architecture. The application relates to a remote sensing image super-resolution method based on a complex-valued neural network and wherein epoch denotes a range of rounds of training, from the 1st round to the nth round, denote Fourier transform operator and inverse operator, respectively, I (g) denote the image generated in the gth round, ε g denote learning rate, denote real image, denote square of L2 norm, denote loss value.

Citation Information

Patent Citations

  • Deep learning-based cross-dataset magnetic resonance multi-modal super-resolution image synthesis method

    CN119205527A