Remote sensing image reconstruction method and device cooperating with space channel attention and dynamic frequency domain multi-scale fusion

Through the remote sensing image reconstruction method that integrates the attention of the space channel and the dynamic frequency domain multi-scale fusion, the data loss problem caused by cloud occlusion in Sentinel-2 images is solved, and high-precision image reconstruction and resolution improvement are achieved, meeting the needs of agricultural and ecological monitoring.

CN120339436AActive Publication Date: 2025-07-18ZHEJIANG UNIV

Patent Information

Application Number
CN202510463260.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively fill the lack of Sentinel-2 image data caused by cloud occlusion and other factors, especially in the key growth stage in the agricultural field, which affects continuous time series monitoring, and heterologous image fusion has problems with acquisition geometry, spectral characteristics and spatial resolution differences.

Method used

The remote sensing image reconstruction method that integrates the attention and dynamic frequency domain multi-scale fusion of collaborative spatial channel attention and multi-scale fusion of dynamic frequency domains is adopted. By constructing a training model, the time embedding layer, multi-scale convolution module, enhanced attention module and upsampling module of Landsat-8 images and Sentinel-2 images are used to reconstruct Sentinel-2 images, especially the red edge band, taking into account dynamic changes in the surface and spatial heterogeneity.

Benefits of technology

High-precision reconstruction of Sentinel-2 images is realized, which fills in data loss, improves the complete timing and resolution of the images, and provides more accurate data support, especially in the fields of agricultural and ecological monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339436A_ABST
    Figure CN120339436A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image reconstruction method and device cooperating with spatial channel attention and dynamic frequency domain multi-scale fusion, and the method comprises the steps: carrying out the decomposition, fusion and reconstruction of a frequency spectrum of a down-sampling feature in an attention enhancement module, thereby enhancing the expression of high-frequency information through the extraction of a frequency domain feature; according to the method, the Landsat-8 image slice sample set with month information is taken as the training sample set, the month information in the training sample is embedded into the Landsat-8 image data through the time embedding layer to obtain a comprehensive feature map including time and space for training, and thus the influence of dynamic change of the earth surface is considered. Through the two points, image data missing caused by factors such as cloud layer shielding can be filled, meanwhile, the red edge wave band of Sentinel-2 is reconstructed, and more accurate and comprehensive data support is provided for the fields of agriculture, ecological monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a remote sensing image reconstruction method and device that synergistically integrate spatial channel attention and dynamic frequency domain multi-scale fusion. Background Art

[0002] Satellite remote sensing technology acquires the geometric and physical characteristics of the earth's surface by collecting data in different bands of the electromagnetic spectrum, and is widely used in fields such as land use, agriculture, forestry, water resources, and disaster emergency. With the development of global earth observation technology, multiple satellite systems such as the Landsat and Sentinel series have been successively launched. However, relying solely on single-satellite data for large-scale and long-term surface monitoring still faces many challenges. Although Landsat-8 / 9 has an 8-day revisit cycle and the Sentinel-2A / B combination can provide 5-day global coverage, the working conditions of satellite sensors and atmospheric factors often result in missing remote sensing image data.

[0003] Especially in the agricultural field, the critical growth stages of most crops are short and usually coincide with the rainy season. If satellite images are missing due to cloud cover or other reasons, the critical observation periods are often missed, affecting continuous time series monitoring. Therefore, filling data gaps and ensuring data continuity are crucial for long-term and accurate surface change monitoring.

[0004] Current remote sensing image filling techniques are mainly divided into spatial-based methods, spectral-based methods, time-based methods, and spatio-temporal spectral methods. However, these techniques still target single data sources, restricting the requirements for high-frequency and continuous monitoring. Another effective solution is to fuse different satellite images to increase the monitoring frequency and the probability of obtaining cloud-free data. However, heterogenous image data often has differences in acquisition geometry, spectral characteristics, and spatial resolution, and simple image fusion will have significant uncertainties.

[0005] The Harmonized Landsat Sentinel (HLS) products (HLSL30 and HLSS30) use the Hyperion hyperspectral dataset to simulate the surface reflectance of Landsat-8 and Sentinel-2, and correct the spectral differences between different bands by constructing a fixed linear regression model, solving the problem of differences in heterogenous remote sensing images. However, Landsat-8 lacks the three additional red-edge bands provided by Sentinel-2. Therefore, when there is no Sentinel-2 data, the HLS images still lack red-edge bands and cannot completely solve the data missing problem. In addition, the spatial resolution of the HLS dataset is 30m.

[0006] The invention patent application with the publication number CN118691496A discloses an image fusion method and system based on a multi-scale smoothing and sharpening filter, belonging to the technical field of remote sensing image processing. The method includes: obtaining Landsat-8 images at the target time and Sentinel-2 images at the reference time, and performing preprocessing on them respectively; extracting high-frequency components from the preprocessed images through a smoothing and sharpening filter; using a multi-scale smoothing and sharpening filter to extract multi-scale detail images from the high-frequency components of the Sentinel-2 images; summing the multi-scale detail images of the Sentinel-2 images and adding them to the preprocessed Landsat-8 images to obtain an image fusion result. This invention patent application improves the accuracy of the fusion of Landsat-8 images and Sentinel-2 images by extracting and transferring detail information from the high-frequency components of the Sentinel-2 images, and can accurately capture phenological and land cover changes. However, it is unable to complement the Sentinel-2 images with high quality for areas with large differences in full-coverage heterogeneity.

[0007] Therefore, there is an urgent need to develop a reconstruction framework for remote sensing images to reconstruct missing images, while improving the resolution and ensuring the complete timeliness of the images. Summary of the Invention

[0008] The present invention provides a remote sensing image reconstruction method that combines spatial-channel attention and dynamic frequency-domain multi-scale fusion. This method can effectively fill the data gap of Sentinel-2 images caused by factors such as cloud occlusion and reconstruct the red-edge band of Sentinel-2.

[0009] The present invention provides a remote sensing image reconstruction method that combines spatial-channel attention and dynamic frequency-domain multi-scale fusion, including:

[0010] Performing preprocessing and slicing on the selected Sentinel-2 and Landsat-8 images to obtain Landsat-8 and Sentinel-2 image pair slice samples that match in time and space and cover the whole year's monthly time series. Using the Landsat-8 image slice sample set with month information as the training sample set and the corresponding Sentinel-2 image slice samples as labels;

[0011] Build a training model, where the training model includes a time embedding layer, a multi-scale convolution module, a downsampling module, an enhanced attention module, an upsampling module, and an activation function. Among them, the month information in the training samples is embedded into the Landsat-8 image data through the time embedding layer to obtain a comprehensive feature map. The comprehensive feature map is sequentially passed through the multi-scale convolution module and the downsampling module to obtain downsampled features. The enhanced attention module extracts the frequency-domain reconstruction features and self-attention features of the downsampled features, and then fuses them to obtain fused features. The fused features are sequentially passed through the upsampling module and the activation function to obtain the predicted Sentinel-2 image;

[0012] Construct a loss function through the predicted Sentinel-2 image and the label, and train the training model based on the training sample set through the loss function to obtain a remote sensing image reconstruction model;

[0013] During application, the Landsat-8 image corresponding to the occluded Sentinel-2 image and the corresponding month information are input into the remote sensing image reconstruction model to obtain the reconstructed Sentinel-2 image.

[0014] Preferably, extracting the frequency-domain reconstruction features of the downsampled features by the enhanced attention module includes:

[0015] The enhanced attention module converts the downsampled features from the spatial domain to the frequency domain, then performs spectral decomposition to obtain high-frequency components and low-frequency components, and after converting the high-frequency components and low-frequency components back to the spatial domain and fusing them, the frequency-domain reconstruction features are obtained.

[0016] Preferably, the method for obtaining the frequency-domain reconstruction features includes:

[0017] Perform a two-dimensional fast Fourier transform on the downsampled features to convert the spatial domain to the frequency domain to obtain the spectrum of the downsampled features. Use a differentiable band separation algorithm to decompose the spectrum into high-frequency components and low-frequency components. After performing an inverse two-dimensional fast Fourier transform on the high-frequency components and low-frequency components to convert them back to the spatial domain and fusing them, the frequency-domain reconstruction features are obtained.

[0018] Preferably, fusing the frequency-domain features and self-attention features of the downsampled features by the enhanced attention module to obtain fused features includes:

[0019] The enhanced attention module includes a first branch, a second branch, and a fusion module;

[0020] The first branch is a spatial channel attention module. Through the spatial channel attention module, spatial self-attention is performed on the downsampled features, and the spatial self-attention result is connected to the downsampled features by residual connection to obtain self-attention features;

[0021] The second branch is a frequency domain layer, and the frequency domain reconstruction features of the downsampled features are extracted through the frequency domain layer;

[0022] The fusion module performs pixel-by-pixel weighted fusion on the self-attention features and the frequency domain reconstruction features to obtain fusion features.

[0023] Preferably, the month information in the training samples is embedded into the Landsat-8 image information through the time embedding layer to obtain a spatio-temporal comprehensive feature map, including:

[0024] The time embedding layer includes an embedding layer, which maps the month information into a high-dimensional feature vector, and after normalizing the high-dimensional feature vector, it is element-wise added to the corresponding Landsat-8 image data along the channel dimension to obtain a comprehensive feature map containing time and space information.

[0025] Preferably, a method for obtaining Landsat-8 and Sentinel-2 image pair slice samples that are monthly in chronological order throughout the year and match in time and space includes:

[0026] The area of the remote sensing image is sampled by a grid to generate multiple random points, and then the points covered by both Landsat-8 and Sentinel-2 images in each month are screened out. Taking the screened points as the center, image patches with a set area are extracted, and the corresponding image patches of the closest dates of Landsat-8 and Sentinel-2 images at each point are extracted, so as to preliminarily screen and obtain Landsat-8 and Sentinel-2 image pair slice samples that are monthly in chronological order throughout the year and match in time and space;

[0027] Then, after selecting the image pair slice samples with the number of null value pixels accounting for less than 1% of the total number of pixels in the slice, the missing pixels are compensated by bicubic interpolation to obtain Landsat-8 and Sentinel-2 image pair slice samples that are monthly in chronological order throughout the year and match in time and space.

[0028] Preferably, the preprocessing of the screened Sentinel-2 and Landsat-8 images includes:

[0029] Screen Sentinel-2 images with cloud cover less than 10%, screen Sentinel-2 images with cloud cover less than 15%, and perform radiometric calibration, atmospheric correction, cloud masking, BRDF correction, topographic correction, and resampling preprocessing on the screened Sentinel-2 and Landsat-8 images in sequence.

[0030] Preferably, the multi-scale convolution module includes a plurality of convolutional layers and activation layers with different numbers of convolutional kernel layers. The comprehensive feature map is respectively subjected to feature extraction through the convolutional layers with different numbers of convolutional kernel layers to obtain feature maps of different sizes. The feature maps of different sizes are spliced and then activated through the activation layer, and the activation result is input into the downsampling module.

[0031] The present invention also provides a remote sensing image reconstruction device for synergistic spatial-channel attention and dynamic frequency-domain multi-scale fusion, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the remote sensing image reconstruction of synergistic spatial-channel attention and dynamic frequency-domain multi-scale fusion.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] Since the area of the remote sensing image is a region with complex terrain covering the entire surface of the earth, that is, a region with large differences in the spatial heterogeneity of different surface cover environments, the present invention decomposes, fuses, and reconstructs the spectrum of the downsampled features in the enhanced attention module. By extracting frequency-domain features, the surface detail information with different-scale spatial heterogeneity can be accurately reconstructed, so that the missing Sentinel-2 images can be reconstructed more accurately.

[0034] At the same time, since the surface impact changes over time, the present invention uses the Landsat-8 image slice sample set with month information as the training sample set, and embeds the month information in the training sample into the Landsat-8 image data through the time embedding layer to obtain a comprehensive feature map including time and space for training, thereby considering the impact of time on the Sentinel-2 image.

[0035] Through the above two points, the missing image data caused by factors such as cloud occlusion can be filled, and at the same time, the red-edge band of Sentinel-2 is reconstructed, providing more accurate and comprehensive data support for fields such as agricultural and ecological monitoring. Description of the Drawings

[0036] Figure 1 It is a flowchart of the remote sensing image reconstruction method for synergistic spatial-channel attention and dynamic frequency-domain multi-scale fusion according to a specific embodiment of the present invention;

[0037] Figure 2 It is a structural diagram of the remote sensing image reconstruction method for synergistic spatial-channel attention and dynamic frequency-domain multi-scale fusion according to a specific embodiment of the present invention;

[0038] Figure 3 It is a structural diagram of the residual connection according to a specific embodiment of the present invention;

[0039] Figure 4This is a comparison chart of the reconstruction results of specific embodiments of the present invention and other methods. Specific Embodiments

[0040] In order to make the objectives, technical solutions, and technical effects of the present invention clearer and more understandable, the following further elaborates on the present invention in conjunction with the attached drawings of the specification.

[0041] Since most areas of full-coverage remote sensing images include multiple land use types, and there are significant differences in surface features of different types, and the prior art does not consider the impact of surface dynamic changes on the quality of Sentinel-2 image completion when performing Sentinel-2 image completion, specific embodiments of the present invention extract frequency-domain reconstructed features of downsampled features by introducing dynamic frequencies to achieve the reconstruction of Sentinel-2 images.

[0042] Specific embodiments of the present invention provide a remote sensing image reconstruction method that synergistically combines spatial channel attention and dynamic frequency domain multi-scale fusion, as Figure 1 shown, including:

[0043] (1) Construct a training sample set and labels: Preprocess and slice the selected Sentinel-2 and Landsat-8 images to obtain Landsat-8 and Sentinel-2 image pair slices that are time-sequential throughout the year and match in time and space. Use the Landsat-8 image slice sample set with month information as the training sample set, and use the corresponding Sentinel-2 image slice samples as labels.

[0044] In this embodiment, a typical cloudy and rainy area in southern China - Zhejiang Province is selected as an example to construct a complete time-sequential remote sensing image set of Zhejiang Province in 2019. Before making slice sample pairs, first obtain multi-source remote sensing data covering Zhejiang Province and preprocess it. The specific steps are as follows:

[0045] The specific steps for obtaining regional remote sensing images provided by specific embodiments of the present invention are as follows: Obtain Sentinel-2 and Landsat-8 images of Zhejiang Province throughout 2019 on the Google Earth (GEE) engine. Through data screening, for Sentinel-2, select images with cloud cover below 10% (except for images in June with cloud cover below 50% and images in July with cloud cover below 40%), and a total of 760 scenes are obtained. For Landsat-8 images, select images with cloud cover below 15% (except for images in January and June with cloud cover below 50% and images in July with cloud cover below 40%), and 103 images are obtained.

[0046] After obtaining the Sentinel-2 and Landsat-8 images after cloud screening, some preprocessing operations are performed on the multi-source remote sensing data, including radiometric calibration, atmospheric correction, cloud masking, BRDF correction, terrain correction, and resampling. First, to reduce the residual errors caused by using different atmospheric correction methods, this study applied the Py6S atmospheric correction model to all top-of-atmosphere (TOA) images of Landsat-8 and Sentinel-2. In this model, the zenith angle is encoded as "0". For the cloud masking of Sentinel-2 data, it is based on the cloud probability data provided by GEE. By selecting the cloud cover threshold, the corresponding cloud masking data is obtained and applied to the Sentinel-2 images after atmospheric correction to effectively remove the influence of clouds. For Landsat-8 images, the cloud and cloud shadow are masked through the quality assessment band (QA_PIXEL) provided by GEE. Next, the bidirectional reflectance distribution function (BRDF) model is used to process the reflectance changes of Landsat-8 and Sentinel-2 images caused by the observation geometry and atmospheric effects. The viewing angle is set to the positive viewing angle, and the illumination is set according to the central latitude of the image. Subsequently, the reflectance is adjusted by the improved sun-canopy-sensor terrain correction method due to the influence of slope, aspect, and elevation. Finally, the Sentinel-2 images are resampled to a resolution of 10 meters, reprojected into the Mercator projection coordinate system, and uniformly converted to the World Geodetic System 1984 (WGS1984). The georegistration of these two data types is performed using the Sentinel-2 global reference image data provided by the European Space Agency.

[0047] The specific steps for obtaining the slice samples provided by the specific embodiments of the present invention are as follows: After completing the preprocessing of the data, 1000 random points are generated by grid sampling in the entire Zhejiang Province area. Then, the points covered by Landsat-8 and Sentinel-2 images every month in 2019 are screened out. Taking the points as the center, image blocks of 2.4 km × 2.4 km are extracted. By matching the closest dates of the two types of data for each point, the image pairs are screened to obtain the monthly time-series Landsat-8 and Sentinel-2 image pair slice samples throughout the year, a total of 9238 pairs. On this basis, after screening out the image pairs with the number of null pixels accounting for less than 1% of the total number of pixels in the slice, the missing pixels are filled as much as possible by bicubic interpolation. After several rounds of manual screening, finally, approximately 6725 pairs of high-quality Landsat-8 and Sentinel-2 image pairs in 2019 are retained. The Landsat-8 image slice sample set with month information is used as the training sample set, and the corresponding Sentinel-2 image slice samples are used as labels.

[0048] (2) Construct a training model, that is, for the construction of a collaborative spatial channel attention and dynamic frequency domain multi-scale fusion network, which specifically includes the following steps:

[0049] The training model provided by the specific embodiment of the present invention includes a time embedding layer, a multi-scale convolution module, a downsampling module, an enhanced attention module, an upsampling module, and an activation function.

[0050] The time embedding layer provided by the specific embodiment of the present invention includes an embedding layer that maps month information (represented in digital form) to a high-dimensional feature vector. After normalizing it, the embedded time features are concatenated with the Landsat-8 image data in the channel dimension to form a comprehensive feature map X1 that contains both time and space information.

[0051] The interpolation module provided by the specific embodiment of the present invention uses interpolation technology to improve the spatial resolution of the image and obtains X2.

[0052] The specific design steps of the multi-scale convolution module provided by the specific embodiment of the present invention are as follows: This module includes three convolution operations, as Figure 2 shown. The first convolution operation contains 1 layer of convolution kernels, aiming to capture the details of small features and local changes, as well as the local relationships in the image, and obtains X3; the second convolution operation increases the number of convolution kernel layers to 2, expanding the receptive field, which helps to capture the ground object structures and global context information in a larger range, and obtains X4; the last convolution operation increases the number of convolution layers to 3, improving the network's perception ability of larger-scale features and the overall scene, and obtains X5. Finally, the three feature maps are concatenated to obtain X6, that is, X3 + X4 + X5 = X6. This module consists of a convolution Conv layer and a (ReLU) layer.

[0053] The specific design steps of the downsampling module provided by the specific embodiment of the present invention are as follows: The downsampling process is divided into three layers, as Figure 2 shown. Each layer contains 2 convolution layers, a Pixel-Unshuffle layer, and a ReLU activation layer. This method of layer-by-layer downsampling can extract higher-level and more abstract semantic features, increase the receptive field, enable each feature to capture the context information in a larger range, while reducing the consumption of computing resources and alleviating data redundancy. The first downsampling layer outputs X7, the second downsampling layer outputs X8, and the last downsampling layer outputs X9.

[0054] The specific design steps of the enhanced attention mechanism module provided by the specific embodiment of the present invention are as follows: The enhanced attention mechanism module introduces a frequency domain layer on the basis of the collaborative spatial channel attention module, as Figure 2As shown. X9 is transformed from the spatial domain to the frequency domain through two-dimensional fast Fourier transform (FFT2), and the spectrum is decomposed into high-frequency components X high and low-frequency components X low ; secondly, a trainable weight coefficient β∈[0,1] is introduced to establish a dynamic fusion equation, and the fused frequency-domain feature X fused =β 2 ·X high +X low is obtained, where X high and X low represent high / low-frequency features respectively. In the specific embodiment of the present invention, the fusion of high / low-frequency features can reduce noise and retain high-frequency information to a large extent; in addition, the frequency-domain feature is transformed back to the spatial domain through inverse FFT2 to obtain the frequency-domain reconstructed feature X 10 , X9 is processed through collaborative spatial channel attention to obtain X 11 , and X 11 is combined with X9 through a residual connection method to obtain the collaborative spatial channel attention feature X 12 ; an adaptive fusion weight α∈R^(1×C×1×1) is used to perform channel-wise weighted fusion on X 10 and X 12 to obtain the fused feature X 13 =α⊙X 10 +(1-α)⊙X 12 , where ⊙ represents the channel-wise multiplication operation; finally, a multi-layer perceptron (MLP) including a gated linear unit (GELU) is used to perform a non-linear transformation on the fused feature. Each layer used for the non-linear transformation includes a linear transformation layer, a GELU layer, a linear transformation layer, and a Dropout layer, and the output X 14 processed by the enhanced attention mechanism is obtained. Its mathematical expression is: X 14 =MLP(X 13 )+X 13 , realizing the deep optimization of the feature space.

[0055] The specific steps of the enhanced upsampling module design provided by the specific embodiment of the present invention are as follows: The upsampling module corresponds to the downsampling process, as Figure 2As shown, the upsampling module consists of three layers. The first layer sequentially includes a convolutional layer, a Pixel-shuffle layer, and a ReLU activation layer. The second and third layers each independently and sequentially contain a convolutional layer, a ReLU activation layer, a residual connection layer, a convolutional layer, a Pixel-shuffle layer, and a ReLU activation layer. The last layer sequentially includes a convolutional layer and a ReLU activation layer. The outputs of the first three upsampling layers are skip-connected to the corresponding inputs of the downsampling. The spatial resolution of the feature map is restored through three successive upsampling and convolution operations. After each upsampling, convolution operations and residual block processing are performed. First, for the input feature map X 14 is first convolved through a convolutional layer to double the number of channels; then, through a depth-to-space conversion operation (Pixel shuffle layer), the feature map is rearranged from the depth dimension to the spatial dimension to achieve upsampling and double the spatial resolution; then the result of the upsampling is concatenated with the downsampled feature map of the previous stage through a skip connection and convolved to fuse multi-scale information, enhancing the feature representation during the reconstruction process and avoiding or reducing the problem of gradient disappearance. After three upsamplings, X 15 、X 16 、X 17 are obtained in sequence. The introduced residual connection mechanism enables the network to learn the residual mapping between the input and the output, thereby accelerating the training process and improving the convergence and generalization ability of the network. The residual block structure includes Conv, ReLU, residual scaling (ResScale), and element-wise summation operations. In the ResScal operation, a constant scalar is used to deepen the network and slow down its convergence speed, thereby stabilizing the training without increasing the number of parameters, as Figure 3 shown. Finally, the processed feature map X 17 is mapped to the final output result, i.e., the predicted Sentinel-2 image through a convolutional layer and a ReLU layer.

[0056] (3) A loss function is constructed by predicting Sentinel-2 images and labels, and a remote sensing image reconstruction model is obtained by training the training model based on the training sample set through the loss function:

[0057] The specific steps for establishing the loss function provided by the specific embodiments of the present invention are as follows: Assume that L is the input of the network and S is the desired output (or known reference data). Network training is to learn a mapping function S = f(L; w), where S is the predicted output and w is the set of all parameters including filter weights and biases. The mean absolute error MAE is used as the loss function: where L represents Landsat-8, S represents Sentinel-2, f(L) is the predicted output, i.e., the reconstructed Sentinel-2, and n represents the number of training samples.

[0058] The specific steps for network parameter setting provided by the specific embodiments of the present invention are as follows: Using the deep learning framework pytorch 1.9.0, set the number of epochs to 200 rounds, the batch size to 64, the division ratio of the training set, validation set, and test set to 8:1:1, and set the initial learning rate to 1×10 -4 , and use the Adam optimizer to update the model parameters until the loss no longer decreases.

[0059] (4) Use the remote sensing image reconstruction model to complete the occluded Sentinel-2 image, and evaluate the completed Sentinel-2 image, including:

[0060] During application, input the Landsat-8 corresponding to the occluded Sentinel-2 image and the corresponding month information into the remote sensing image reconstruction model to obtain the reconstructed Sentinel-2 image.

[0061] The specific steps for remote sensing image reconstruction provided by the specific embodiments of the present invention are as follows: Input the Landsat-8 data in the test set into the trained remote sensing image reconstruction model to reconstruct the corresponding Sentinel-2 data, and obtain the reconstructed image.

[0062] The specific steps for evaluating the accuracy of the reconstructed image provided by the specific embodiments of the present invention are as follows: Compare the pixel values and band values of the reconstructed image and the real image, generally using the mean absolute error (MAE) and peak signal-to-noise ratio (PSNR). The formulas are as follows:

[0063]

[0064] Among them, Y and respectively represent the real image and the reconstructed image, and m and n respectively represent the height and width of the image. The lower the MAE value, the better the reconstruction performance.

[0065]

[0066] Among them, max(Y) represents the maximum pixel value of the image.

[0067] In this embodiment, when comparing the reconstructed image with the real image in the established network, the evaluation index MAE is 0.0165 and the PSNR is 34.162. Compared with the traditional interpolation method Bicubic (MAE: 0.0308, PSNR: 27.232) and other models (such as AUNet: MAE: 0.0185, PSNR: 33.291; DCMNET: MAE: 0.0182, PSNR: 33.424; Transenet: MAE: 0.0189, PSNR: 33.164; RCAN: MAE: 0.0202, PSNR: 32.634), the network proposed by the present invention has a significant improvement in the image reconstruction accuracy. Visually, the reconstructed image obtained by the network provided by the present invention shows clearer and more continuous edges, and the tone is closer to the real Sentinel-2, and the visual effect is better than other models.

[0068] As Figure 4 shown, the first column is the original Sentinel-2 image, and the second, third, fourth, fifth, and sixth columns are the Sentinel-2 images reconstructed based on Landsat-8 using the Bicubic, AUNet, Transenet, RCAN, and DCMNE methods respectively. The last column is the reconstructed image obtained based on the algorithm proposed by the present invention, and the fourth row is a partial enlarged view of the third row image.

[0069] In summary, in the specific embodiment of the present invention, by collaborating the spatial channel attention and the dynamic frequency domain multi-scale fusion network, feature information is extracted at multiple scales and the key features are strengthened by combining a strong attention mechanism, thereby effectively improving the accuracy and robustness of image analysis. In the specific implementation process, first, multi-scale convolutional layers are used to extract multi-level and different-scale feature maps from the input image. On this basis, a strong attention mechanism is used to perform weighted fusion on the feature maps of different scales, so as to more accurately focus on the important regions or information in the image.

[0070] Compared with the prior art based on traditional image interpolation methods and using traditional convolutional neural networks or simple multi-scale networks, the framework provided by the present invention has significant advantages in reconstruction accuracy. Compared with the reconstruction methods using a single image or a global model, the enhanced attention mechanism can effectively suppress the influence of noise, with higher accuracy and better effect.

[0071] The present invention also provides a remote sensing image reconstruction device that collaborates spatial channel attention and dynamic frequency domain multi-scale fusion, which is characterized by including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the remote sensing image reconstruction that collaborates spatial channel attention and dynamic frequency domain multi-scale fusion.

Claims

1. A remote sensing image reconstruction method that combines spatial channel attention and dynamic frequency domain multi-scale fusion, characterized in that, Including: Preprocess and slice the selected Sentinel-2 and Landsat-8 images to obtain Landsat-8 and Sentinel-2 image pairs sliced samples that match in time and space for the whole year's monthly time series. Use the Landsat-8 image sliced sample set with month information as the training sample set, and use the corresponding Sentinel-2 image sliced samples as labels; Construct a training model. The training model includes a time embedding layer, a multi-scale convolution module, a downsampling module, an enhanced attention module, an upsampling module, and an activation function. Among them, the month information in the training samples is embedded into the Landsat-8 image data through the time embedding layer to obtain a comprehensive feature map. The comprehensive feature map is sequentially passed through the multi-scale convolution module and the downsampling module to obtain downsampled features. The enhanced attention module extracts the frequency-domain reconstruction features and self-attention features of the downsampled features, and then fuses them to obtain fused features. The fused features are sequentially passed through the upsampling module and the activation function to obtain the predicted Sentinel-2 image; Construct a loss function through the predicted Sentinel-2 image and the label, and train the training model based on the training sample set through the loss function to obtain a remote sensing image reconstruction model; During application, input the Landsat-8 image and the corresponding month information of the adjacent date of the occluded Sentinel-2 image into the remote sensing image reconstruction model to obtain the reconstructed Sentinel-2 image.

2. The remote sensing image reconstruction method based on collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, characterized in that Extracting the frequency-domain reconstruction features of the downsampled features by the enhanced attention module includes: The enhanced attention module converts the downsampled features from the spatial domain to the frequency domain, and then performs spectral decomposition to obtain high-frequency components and low-frequency components. After converting the high-frequency components and low-frequency components back to the spatial domain, they are fused to obtain frequency-domain reconstruction features.

3. The remote sensing image reconstruction method based on collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 2, characterized in that The method for obtaining the frequency-domain reconstruction features includes: Perform a two-dimensional fast Fourier transform on the downsampled features to convert the spatial domain to the frequency domain to obtain the spectrum of the downsampled features. Use a differentiable frequency band separation algorithm to decompose the spectrum to obtain high-frequency components and low-frequency components. After performing an inverse two-dimensional fast Fourier transform on the high-frequency components and low-frequency components to convert them back to the spatial domain and then fusing them, frequency-domain reconstruction features are obtained.

4. The remote sensing image reconstruction method combining collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, characterized in that Fusing the frequency-domain features and self-attention features of the downsampled features by the enhanced attention module to obtain fused features includes: The enhanced attention module includes a first branch, a second branch, and a fusion module; The first branch is a spatial channel attention module. Through the spatial channel attention module, spatial self-attention is performed on the downsampled features, and the spatial self-attention result is connected with the downsampled features through a residual connection to obtain self-attention features; The second branch is a frequency-domain layer, which extracts the frequency-domain reconstruction features of the downsampled features; The fusion module performs pixel-by-pixel weighted fusion on the self-attention features and the frequency-domain reconstruction features to obtain fused features.

5. The remote sensing image reconstruction method for collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, wherein Embedding the month information in the training samples into the Landsat-8 image information through the time embedding layer to obtain a spatio-temporal comprehensive feature map includes: The time embedding layer includes an embedding layer. The month information is mapped into a high-dimensional feature vector through the embedding layer. After normalizing the high-dimensional feature vector, it is element-wise added to the corresponding Landsat-8 image data along the channel dimension to obtain a comprehensive feature map containing temporal and spatial information.

6. The remote sensing image reconstruction method combining collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, characterized in that, A method for obtaining slice samples of Landsat-8 and Sentinel-2 image pairs that match in time and space over the entire year's monthly time series, including: Generating multiple random points by grid sampling in the area of the remote sensing image, then screening out the points covered by both Landsat-8 and Sentinel-2 images each month. Taking the screened points as the center, extracting image patches with a set area, and extracting the corresponding image patches of the Landsat-8 and Sentinel-2 images with the closest dates for each point, so as to preliminarily screen and obtain slice samples of Landsat-8 and Sentinel-2 image pairs that match in time and space over the entire year's monthly time series; Then, after selecting the image pair slice samples with the number of null value pixels accounting for less than 1% of the total number of pixels in the slice, the missing pixels are filled by bicubic interpolation to obtain slice samples of Landsat-8 and Sentinel-2 image pairs that match in time and space over the entire year's time series.

7. The remote sensing image reconstruction method combining collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, wherein Preprocessing the selected Sentinel-2 and Landsat-8 images, including: Selecting Sentinel-2 images with cloud cover less than 10%, selecting Sentinel-2 images with cloud cover less than 15%, and successively performing radiometric calibration, atmospheric correction, cloud masking, BRDF correction, terrain correction, and resampling preprocessing on the selected Sentinel-2 and Landsat-8 images.

8. The remote sensing image reconstruction method combining collaborative spatial channel attention and dynamic frequency domain multi-scale fusion according to claim 1, wherein, The multi-scale convolution module includes multiple convolutional layers and activation layers with different numbers of convolutional kernel layers. The comprehensive feature map is respectively subjected to feature extraction through multiple convolutional layers with different numbers of convolutional kernel layers to obtain feature maps of different sizes. The feature maps of different sizes are concatenated and then activated through the activation layer, and the activation result is input into the downsampling module.

9. A remote sensing image reconstruction device that collaboratively integrates spatial channel attention and dynamic frequency-domain multi-scale fusion, characterized in that, It includes a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the remote sensing image reconstruction of collaborative spatial-channel attention and dynamic frequency-domain multi-scale fusion as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Image fusion method and system based on multi-scale smooth sharpening filter

    CN118691496A

  • Target detection method and device based on multi-scale image

    CN113221925A

  • Remote sensing image space-time fusion method based on multi-scale and hybrid convolutional network

    CN117876240A

  • Method for extracting building change area in double-time-phase remote sensing image based on twinborn mixed attention mechanism and multi-scale feature fusion

    CN118212532A

  • Sentinel-2 remote sensing image super-resolution reconstruction method based on attention mechanism

    CN119205503A

Cited By

  • Landsat-MODIS remote sensing image space-time fusion method based on double-layer cross attention mechanism

    CN121010871A