A multi-view and multi-scale based remote sensing image spatio-temporal fusion method
The spatiotemporal fusion method for remote sensing images, which employs multi-view, multi-scale feature extraction and edge attention mechanisms, addresses the issue of insufficient attention to spectral quality and channel relationships in existing technologies. This method achieves higher-precision image fusion and enhances the application effect of remote sensing images.
Patent Information
- Application Number
- CN202310246762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-15
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2043-03-15
AI Technical Summary
Existing spatiotemporal fusion methods for remote sensing images fail to adequately consider inter-channel relationships and spectral quality, resulting in limited image fusion accuracy.
A spatiotemporal fusion method based on multi-view and multi-scale remote sensing images is adopted. The feature information of different viewpoints and scales is extracted through a multi-view and multi-scale feature extraction model, and the image fusion is performed by using an edge attention mechanism and a decoding module to reduce spectral distortion and preserve structural texture information.
It improves image fusion quality, reduces spectral distortion, and more fully preserves structural texture information, thus enhancing the image fusion effect. It is of great significance in fields such as land use, natural disaster monitoring, urban planning, agricultural management, and environmental monitoring.
Smart Images

Figure CN116229284B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and specifically to a spatiotemporal fusion method for remote sensing images based on multiple perspectives and scales. Background Technology
[0002] In recent years, the spatiotemporal fusion technology of remote sensing images has been widely used in many fields, such as land use and legend cover classification, natural disaster monitoring, urban planning and construction, agricultural management, and environmental monitoring. It is of great significance for realizing resource management, environmental protection, and economic development. Therefore, high spatial resolution high density time series remote sensing images (HSHT) are crucial for timely and accurate detection of ecology and environment.
[0003] Spatiotemporal fusion algorithms refer to the techniques of using two or more data with similar spectral ranges to fuse richer data than a single data source through specific algorithms. In other words, spatiotemporal fusion technology combines different information by retrieving temporal ground changes from high temporal and low spatial resolution images and extracting detailed ground textures from low temporal and high spatial resolution images to simultaneously fuse them into an image with both high spatial and high temporal resolution.
[0004] Existing spatiotemporal fusion algorithms can be broadly categorized into four types: weighted function-based, decomposition-based, Bayesian-based, and learning-based. Weighted function-based spatiotemporal fusion algorithms primarily calculate changes in image information through HSLT images and assign these changes to the HSLT images according to weights to obtain HSHT images. The most classic algorithm is the Spatial and Temporal Adaptive Reflection Fusion Model (STARM). A drawback of weighted function-based algorithms is their reliance on empirical assumptions and insufficient sensitivity to dataset requirements. Decomposition-based methods mainly involve: endmember extraction, abundance estimation, decomposition within a sliding window, and high-resolution image reconstruction. Bayesian-based methods are based on Bayesian statistics, assuming the spatiotemporal fusion problem is a maximum a posteriori problem, attempting to integrate spatial-temporal-spectral fusion into a unified framework. Their drawback is uncertainty, significantly influenced by variations in the geographic region and land type of the image. Learning-based methods use machine learning to model the relationship between low-temporal, high-spatial-resolution images and high-temporal, low-spatial-resolution images, then predict the temporal high-resolution image. For example, AMNet directly uses the residual image obtained by subtracting the MODIS image from two imaging sessions to train the network, and employs two special structures, a multi-scale mechanism and an attention mechanism, to improve fusion accuracy. However, even though there are many current methods for spatiotemporal fusion of remote sensing images, they do not pay sufficient attention to inter-channel relationships and spectral quality, which means they cannot perform feature extraction, greatly limiting their image fusion accuracy. Summary of the Invention
[0005] To address the problems existing in the background technology, this invention provides a spatiotemporal fusion method for remote sensing images based on multiple views and scales, which reduces spectral distortion to a greater extent, preserves structural texture information more fully, and improves image fusion quality, including:
[0006] S1: Obtain high temporal and low spatial resolution images at times t1 and t0, low temporal and high spatial resolution image at time t0, and low temporal and high spatial resolution reference image at time t1;
[0007] S2: Upsample the high temporal and low spatial resolution images at times t0 and t1 respectively to make them the same size as the low temporal and high spatial resolution images at time t0, thus obtaining two first intermediate images;
[0008] S3: The two first intermediate images are stitched together with the low temporal and high spatial resolution image at time t0 to obtain the second intermediate image; and the first encoder is used to encode the low temporal and high spatial resolution image at time t0 to generate the first intermediate feature map.
[0009] S4: Encode the second intermediate image using the second encoder to generate a second intermediate feature map; add the first intermediate feature map and the second intermediate feature map along the feature dimension to obtain a third intermediate feature map;
[0010] S5: Input the third intermediate feature map into the first multi-view multi-scale feature extraction model and the second multi-view multi-scale feature extraction model respectively, and perform feature extraction on the third intermediate feature map at multiple scales and from multiple perspectives to obtain the fourth intermediate feature map and the fifth intermediate feature map respectively;
[0011] S6: Add the fourth and fifth intermediate feature maps along the feature dimension to obtain the sixth intermediate feature map; input the sixth intermediate feature map into the attention mechanism model to weight the sixth intermediate feature map to obtain the seventh intermediate feature map; concatenate the first and seventh intermediate feature maps to obtain the eighth intermediate feature map;
[0012] S7: Input the eighth intermediate feature map into the decoder to obtain a high temporal and high spatial resolution image that integrates spatiotemporal information at time t1. Then, use the Adam algorithm to supervise the training of the model based on the error between the high temporal and high spatial resolution image that integrates spatiotemporal information and the reference image.
[0013] This invention constructs a multi-view, multi-scale feature extraction model to extract and fuse feature information from different perspectives and scales, improving the network's feature extraction capabilities. It employs an edge attention mechanism to obtain the correlation of edge information and assigns different weights to edge details, enhancing the restoration of texture details during edge feature fusion. Simultaneously, it decodes the extracted deep features and structural information through a fusion image decoding module, using its own internal convolution to obtain related feature information, resulting in better image fusion quality. This significantly reduces spectral distortion and better preserves structural and texture information, leading to improved image fusion quality. The enhanced fusion effect is of great significance for land use and legend cover classification, natural disaster monitoring, urban planning and construction, agricultural management, environmental monitoring, and for achieving resource management, environmental protection, and economic development. Attached Figure Description
[0014] Figure 1 This is a flowchart of the method of the present invention;
[0015] Figure 2 This is a schematic diagram of the method flow structure of the present invention;
[0016] Figure 3 This is a schematic diagram of the structure of a multi-view, multi-scale feature extraction model;
[0017] Figure 4 This is a schematic diagram of the ZBlock layer structure;
[0018] Figure 5 This is a schematic diagram of the ResZNet residual network extracting features along the height, width, and channels of an image;
[0019] Figure 6 This is a comparison chart of the experimental results. Detailed Implementation
[0020] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0021] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0022] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0023] Please see Figure 1 and Figure 2 This invention provides a spatiotemporal fusion method for remote sensing images based on multiple perspectives and scales, comprising:
[0024] S1: Obtain high temporal and low spatial resolution images at times t1 and t0, low temporal and high spatial resolution image at time t0, and low temporal and high spatial resolution reference image at time t1;
[0025] Preferably, an implementation method for acquiring high temporal and low spatial resolution images includes: acquiring high temporal and low spatial resolution images in real time using MODIS satellite sensors; the Moderate-resolution Imaging Spectroradiometer (MODIS) can acquire images of the same area every day or half a day, but its spatial resolution ranges from 250 meters to 1000 meters.
[0026] Preferably, one implementation method for acquiring low temporal, high spatial resolution images includes: acquiring low temporal, high spatial resolution images in real time using Landsat satellite sensors. The spatial resolution of the images acquired by Landsat is 10 to 30 meters, and the revisit period is 16 days.
[0027] The high spatial resolution of low temporal, high spatial resolution images acquired by Landsat satellite sensors is 16a×16a; the high spatial resolution of high temporal, low spatial resolution images acquired by MODIS satellite sensors is a×a.
[0028] Preferably, an implementation method for acquiring low-temporal, high-spatial-resolution images and high-temporal, low-spatial-resolution images includes: obtaining low-temporal, high-spatial-resolution images and high-temporal, low-spatial-resolution images by downloading existing Guangdong and Shandong datasets from the internet; the satellite remote sensing image dataset mainly selects representative satellite data images (Landsat-8 and MOD09A1) from Guangdong and Shandong datasets, each containing a pair of Landsat and MODIS images; Landsat has a spatial resolution of 30 meters and a 16-day revisit time, while MODIS can cover most parts of the Earth daily, but can only acquire data at a spatial resolution of 250 to 1000 meters. In this embodiment, the LEVEL2 product of Landsat8 OLI (which has undergone preliminary radiometric calibration and atmospheric correction) and the 8-day composite data MOD09A1 from MODIS are used, and fusion is performed using four bands: blue, green, red, and near-infrared (NIR). To verify the generality of the proposed model, we selected specific areas in Shandong and Guangdong provinces for testing. In Shandong, the Landsat image with WRS coordinates P122R034 and the corresponding MODIS sinusoidal grid was h27v05. In Guangdong, the Landsat image with WRS coordinates P123R043 and the corresponding MODIS sinusoidal grid was h28v06. The imagery period was from January 1, 2013 to December 31, 2017. Landsat8 images were required to have a cloud cover of less than 5%, and each scene was cropped to 4800x4800 pixels (excluding edges without data). The corresponding MODIS image was reprojected at a spatial resolution of 480 meters, and then cropped to the same area as the Landsat image, resulting in an image size of 300x300. The cropped Landsat image and its corresponding MODIS image constitute a single data pair. Finally, the Landsat and MODIS data pairs were grouped, with each group consisting of two Landsat and two MODIS images. Fourteen groups were selected for each region, and then the groups were randomly divided into training and test sets, with 10 data sets for the training set and 4 data sets for the test set.
[0029] Preferably, in this embodiment, the low temporal and high spatial resolution images and the high temporal and low spatial resolution images are images of the same region at different resolutions.
[0030] Preferably, the reference image is a low temporal and high spatial resolution image at time t1, that is, the low temporal and high spatial resolution image that should be acquired at time t1 in one of the above embodiments. However, due to the long sampling period of the Landsat satellite sensor, the Landsat satellite sensor cannot acquire a low temporal and high spatial resolution image at time t1. Therefore, the present invention generates a low temporal and high spatial resolution image at time t1 by using the high temporal and low spatial resolution images at times t1 and t0, and the low temporal and high spatial resolution image at time t0.
[0031] S2: Upsample the high temporal and low spatial resolution images at times t0 and t1 respectively to make them the same size as the low temporal and high spatial resolution images at time t0, thus obtaining two first intermediate images;
[0032] Preferably, an implementation method for upsampling high temporal resolution and low spatial resolution images at times t0 and t1 includes: interpolating the low spatial resolution image using bilinear interpolation.
[0033] Preferably, an implementation method for upsampling high temporal resolution, low spatial resolution images at times t0 and t1 includes: upsampling using the interpolate() function provided by PyTorch, with a specific upsampling factor of 16.
[0034] Preferably, an implementation method for upsampling high temporal resolution, low spatial resolution images at times t0 and t1 includes: upsampling the images using a convolutional neural network, such as a U-net neural network.
[0035] S3: The two first intermediate images are stitched together with the low temporal and high spatial resolution image at time t0 to obtain the second intermediate image; and the first encoder is used to encode the low temporal and high spatial resolution image at time t0 to generate the first intermediate feature map.
[0036] Preferably, a method for stitching two first intermediate images with a low temporal, high spatial resolution image at time t0 includes: using the CONCAT function to stitch the two first intermediate images with the low temporal, high spatial resolution image at time t0.
[0037] Preferably, the first encoder includes: a first convolutional layer, a first ReLU activation function, a second convolutional layer, a second ReLU activation function, a third convolutional layer, an inner convolutional layer, and a third ReLU activation function connected in sequence; wherein the first convolutional layer, the second convolutional layer, and the third convolutional layer are 3×3 convolutional layers.
[0038] S4: Encode the second intermediate image using the second encoder to generate a second intermediate feature map; add the first intermediate feature map and the second intermediate feature map along the feature dimension to obtain a third intermediate feature map;
[0039] Preferably, the second encoder includes: a first convolutional layer, an inner convolutional layer, and a second convolutional layer connected in sequence; wherein, the first convolutional layer and the second convolutional layer are 3×3 convolutional layers; it should be noted that the first convolutional layer in this invention represents two different convolutional layers in different encoders.
[0040] Involution is a novel atomic operation based on deep neural networks that reverses the convolution design and incorporates a complex instance of the inner convolution operator. Compared to traditional convolution operators with spatially undistorted and channel-specific kernels, involution involves both channel-undistorted and spatially specific operations. Furthermore, involution significantly reduces computational costs in spatial modeling and network architecture design.
[0041] Preferably, the inner convolution includes:
[0042]
[0043] in, It is the size of the involution kernel, Y i,j,k This is the final output, where H, W, K, and G represent height, width, kernel size, and number of channels, respectively, and u and v represent the values in Δ... k Points within the region, X i,j,k This represents the input pixel value. This represents the value of the convolution kernel.
[0044] Please see Figure 3 S5: Input the third intermediate feature map into the first multi-view multi-scale feature extraction model and the second multi-view multi-scale feature extraction model respectively, and perform feature extraction on the third intermediate feature map at multiple scales and multiple views to obtain the fourth intermediate feature map and the fifth intermediate feature map respectively;
[0045] Preferably, both the first and second multi-view multi-scale feature extraction models include: two ZBlock modules and an edge attention mechanism module; the ZBlock module is composed of multiple ResZNet residual networks, which extract features from the third intermediate feature map from multiple views and multiple scales based on the length, width, and channels of the image, respectively; the edge attention mechanism module is used to weight the extracted features; wherein, in the first multi-view multi-scale feature extraction model, the first ZBlock module extracts features from the third intermediate feature map by increasing channel scale and the second ZBlock module extracts features from the third intermediate feature map by decreasing channel scale; in the second multi-view multi-scale feature extraction model, the first ZBlock module extracts features from the third intermediate feature map by decreasing channel scale and the second ZBlock module extracts features from the third intermediate feature map by increasing channel scale.
[0046] The first multi-view, multi-scale extraction module and the second multi-view, multi-scale extraction model extract features from different multi-scale perspectives. The first multi-view, multi-scale extraction model first enlarges the channel scale and then reduces it to extract features from multiple perspectives. For example, in the first ZBlock module, the input image has 24 channels and the output image has 30 channels. In the second ZBlock module, the input image has 30 channels and the output image has 24 channels, i.e., the channel variation parameters are [24, 30, 24]. The second multi-view, multi-scale extraction model first reduces the channel scale and then enlarges it to extract features from multiple perspectives. For example, in the first ZBlock module, the input image has 24 channels and the output image has 12 channels. In the second ZBlock module, the input image has 12 channels and the output image has 24 channels, i.e., the channel variation parameters are [24, 12, 24].
[0047] Please see Figure 4 Preferably, the ZBlock module includes three parallel ZBlock layers: a first ZBlock layer, a second ZBlock layer, and a third ZBlock layer.
[0048] The first ZBlock layer includes: ResZNetC1 residual network, ResZNetH1 residual network, and ResZNetW1 residual network;
[0049] The second ZBlock layer includes: ResZNetC2 residual network, ResZNetH2 residual network, and ResZNetW2 residual network;
[0050] The third ZBlock layer includes: ResZNetC3 residual network, ResZNetH3 residual network, and ResZNetW3 residual network;
[0051] The ResZNetC1 residual network performs feature extraction at multiple scales along the channel C direction of the third intermediate feature map to generate the first intermediate feature, and the ResZNetH1 residual network performs feature extraction at multiple scales along the height H direction of the third intermediate feature map to generate the second intermediate feature.
[0052] The ResZNetW1 residual network performs multi-scale feature extraction on the second intermediate feature along the width W direction of the second intermediate feature to generate the third intermediate feature;
[0053] The first fused feature is obtained by adding the first intermediate feature and the third intermediate feature together along the feature dimension.
[0054] The ResZNetC2 residual network extracts features from the third intermediate feature map along the channel C direction to generate the fourth intermediate feature, and the ResZNetH2 residual network extracts features from the third intermediate feature map along the height H direction to generate the fifth intermediate feature.
[0055] The ResZNetW2 residual network extracts features from the fourth intermediate feature along the width W direction to generate the sixth intermediate feature;
[0056] The fifth and sixth intermediate features are added together along the feature dimensions to obtain the second fused feature;
[0057] The ResZNetC3 residual network extracts features from the third intermediate feature map along the channel C direction of the third intermediate feature map to generate the seventh intermediate feature.
[0058] The ResZNetW3 residual network extracts features from the third intermediate feature map along the height W direction of the third intermediate feature map to generate the eighth intermediate feature.
[0059] The ResZNetH3 residual network extracts features from the seventh intermediate feature along the height H direction of the seventh intermediate feature to generate the ninth intermediate feature.
[0060] The third fused feature is obtained by adding the eighth and ninth intermediate features together along the feature dimensions.
[0061] The first, second, and third fusion features are concatenated using the CONCAT function and then used as the output of the ZBlock module.
[0062] Because convolutional operations are concentrated in local regions, it is difficult to obtain location-independent relational information even in deep networks. Multi-view feature fusion improves the information extraction capability of the network by fusing feature information from different perspectives. This is achieved through an edge attention mechanism module. This mechanism decomposes channel attention into two one-dimensional feature codes, and then performs spatial feature aggregation along these two directions respectively. This allows sensitive location information to be fully preserved, enhancing the extraction of edge information features. Multi-scale feature fusion models provide excellent multi-scale representation capabilities at a finer granular level in the channel dimension. Specifically, a set of convolutional kernels first extracts features from a set of input feature maps. Then, the output element of the previous set is added to the input element of the other set and fed into the next set of convolutional kernels. This process is repeated several times until all input feature maps have been processed. Finally, the feature maps from all sets are concatenated and sent to another set of 1×1 filters to fully fuse the information. For any possible path from input feature to output feature, as long as the feature passes through a 3×3 convolutional kernel, the equivalent receptive field increases. Due to the combinatorial effect, many equivalent feature scales are generated. The designed modules are as follows... Figure 2 This is based on the same method. The same approach is used for the height and width dimensions: the input is divided into multiple groups from different directions, and then, starting with the first group, the input of the group to be added is sent to the input of the next group, ultimately obtaining an expanded receptive field and thus acquiring more equivalent features. Therefore, the network can extract information not only from between channels but also from the other two dimensions, and additionally from between these three dimensions. This has a significant impact on utilizing edge information.
[0063] The Coordinate Attention (CA) module embeds location information into channel attention, forming a collaborative attention mechanism. It then decomposes channel attention into two one-dimensional feature encoding processes that collect features in two spatial directions, differing from standard channel attention mechanisms. Furthermore, it transforms the feature vectors into a single feature vector through two-dimensional global aggregation. This allows for the capture of long-range correlations along one spatial direction while preserving precise location information along the other. The resulting feature maps are then encoded into a pair of direction-aware and location-sensitive attention maps, which can be complementaryly applied to the input feature map to enhance the representation of the object of interest.
[0064] Preferably, the edge attention mechanism includes:
[0065]
[0066] Where, x c (i,j) represents the input pixel value. This indicates an input that has been transformed to the same channel along the width dimension. This represents the input transformed to the same channel along the height dimension, y c (i,j) represents the output pixel value;
[0067] In a preferred embodiment, the first multi-view multi-scale feature extraction model or the second multi-view multi-scale feature extraction model includes: two ZBlock modules and an edge attention mechanism module;
[0068] The ZBlock module includes three parallel layers: a first ZBlock-1-23 layer, a second ZBlock-2-13 layer, and a third ZBlock-3-12 layer.
[0069] The first ZBlock-1-23 layers include: ResZNetC1 residual network, ResZNetH1 residual network, and ResZNetW1 residual network;
[0070] The second ZBlock-2-13 layer includes: ResZNetC2 residual network, ResZNetH2 residual network, and ResZNetW2 residual network;
[0071] The three ZBlock-3-12 layers include: ResZNetC3 residual network, ResZNetH3 residual network, and ResZNetW3 residual network;
[0072] The third intermediate feature map is input into the first ZBlock module to extract features from multiple perspectives and scales to generate the first sub-intermediate feature map.
[0073] The first sub-intermediate feature map is input into the second ZBlock module to extract features from multiple perspectives and scales to generate the second sub-intermediate feature map.
[0074] The second sub-intermediate feature map is input into the edge attention mechanism module for weighted generation of the fourth intermediate feature map.
[0075] Preferably, the first sub-intermediate feature map and the second sub-intermediate feature map include:
[0076] The third intermediate feature map is input into the first ZBlock layer, the second ZBlock layer, and the third ZBlock layer respectively to extract features at multiple scales and from multiple perspectives, generating three first feature maps;
[0077] The three first feature maps are concatenated using the CONCAT function to generate the first sub-intermediate feature map. Similarly, the first sub-intermediate feature map is input into the second ZBlock module to obtain the second sub-intermediate feature map.
[0078] Please see Figure 5Preferably, the first ZBlock layer performs feature extraction on the third intermediate feature map, including:
[0079] The third intermediate feature map is input into the ResZNetC1 residual network and the ResZNetH1 residual network respectively. The ResZNetC1 residual network performs feature extraction at multiple scales along the channel C direction of the third intermediate feature map to generate the first intermediate feature. The ResZNetH1 residual network performs feature extraction at multiple scales along the height H direction of the third intermediate feature map to generate the second intermediate feature.
[0080] The second intermediate feature is input into the ResZNetW1 residual network to perform feature extraction at multiple scales along the width W direction of the second intermediate feature to generate the third intermediate feature;
[0081] The first fused feature is obtained by adding the first intermediate feature and the third intermediate feature along the feature dimension.
[0082] Preferably, the second ZBlock layer performs feature extraction on the third intermediate feature map, including:
[0083] The third intermediate feature map is input into the ResZNetC2 residual network and the ResZNetH2 residual network respectively. The ResZNetC2 residual network extracts features from the third intermediate feature map along the channel C direction to generate the fourth intermediate feature. The ResZNetH2 residual network extracts features from the third intermediate feature map along the height H direction to generate the fifth intermediate feature.
[0084] The fourth intermediate feature is input into the ResZNetW2 residual network to extract features from the fourth intermediate feature along the width W direction to generate the sixth intermediate feature;
[0085] The fifth and sixth intermediate features are added together along the feature dimension to obtain the second fused feature;
[0086] Preferably, the third ZBlock layer performs feature extraction on the third intermediate feature map, including:
[0087] The third intermediate feature map is input into the ResZNetC3 residual network and the ResZNetW3 residual network respectively. The ResZNetC3 residual network extracts features from the third intermediate feature map along the channel C direction to generate the seventh intermediate feature. The ResZNetW3 residual network extracts features from the third intermediate feature map along the height W direction to generate the eighth intermediate feature.
[0088] The seventh intermediate feature is input into the ResZNetH3 residual network to extract features from the seventh intermediate feature along the height H direction of the seventh intermediate feature to generate the ninth intermediate feature.
[0089] The third fused feature is obtained by adding the eighth and ninth intermediate features along the feature dimension.
[0090] The first, second, and third fusion features are concatenated using the CONCAT function and then used as the output of the ZBlock module.
[0091] The first multi-view multi-scale extraction module and the second multi-view multi-scale extraction model have the same network structure, but their internal parameters are different. The first multi-view multi-scale extraction model first enlarges the channel scale and then reduces it to extract multi-view features, while the second multi-view multi-scale extraction model first reduces the channel scale and then enlarges it to extract multi-view features.
[0092] Preferably, ResZNetHi, ResZNetCi, and ResZNetWi are all ResZNet residual networks, where i equals 1, 2, or 3.
[0093] S6: Add the fourth and fifth intermediate feature maps along the feature dimension to obtain the sixth intermediate feature map; input the sixth intermediate feature map into the attention mechanism model to weight the sixth intermediate feature map to obtain the seventh intermediate feature map; concatenate the first and seventh intermediate feature maps to obtain the eighth intermediate feature map;
[0094] Preferably, the seventh intermediate feature map includes:
[0095]
[0096] in, Let X represent a 3D convolution kernel, C′ represent the input, and u represent the number of output channels. c Indicates the output;
[0097]
[0098] Among them, F sq This indicates a squeeze operation, where H, W, and C represent height, width, and number of channels, respectively, and z... c This indicates the output after the extrusion operation;
[0099] s = F ex (z,W)=σ(g(z,W))=σ(W2ReLU(W1z))
[0100] Where s represents the final output through the SE attention mechanism, and F exThe expression represents the excitation operation. W1 and W2 represent two fully connected bottleneck structures, and z represents the output of the excitation operation as the input of the excitation operation.
[0101] S7: Input the eighth intermediate feature map into the decoder to obtain a high temporal and high spatial resolution image that integrates spatiotemporal information at time t1. Then, use the Adam algorithm to supervise the training of the model based on the error between the high temporal and high spatial resolution image that integrates spatiotemporal information and the reference image.
[0102] Preferably, the decoder includes: a first convolutional layer, a ReLU activation function, and a second convolutional layer connected in sequence; wherein the first convolutional layer is a 3×3 convolutional layer, and the second convolutional layer is a 1×1 convolutional layer.
[0103] Preferably, the loss function during supervised training of the model includes:
[0104] Loss=αL mse +βL ssim +(1-λ)L1+λL2
[0105] L1=||XY||1
[0106] L2=||XY||2
[0107]
[0108]
[0109] Where X represents the high temporal and high spatial resolution image generated at time t1 through fusion, and Y represents the reference image. i This represents the value of the i-th pixel in X. This represents the value of the i-th pixel in Y; Let X and Y represent the average values, respectively. Let η represent the variances of X and Y, respectively. XY Let C1 and C2 represent the covariance of X and Y, and C2 be two stable constants. λ is a weighting factor that changes as training progresses, with an initial value of 0.
[0110] In one embodiment, the present invention sets λ to increase by 0.02 per batch in the experiment. α is the control for L. mse The weighting factor for the loss was set to 0.8 based on experimental experience, and β was used to control L. ssimThe weighting factor for the loss function was set to 0.6 based on relevant experimental experience. The Adam algorithm was used to optimize the loss function, with the momentum term set to 0.5, other parameters kept at their default settings, batch size set to 4, initial learning rate set to 0.001, and a total of 70 training epochs.
[0111] To evaluate the performance of this invention, this embodiment selected a classic dataset for experiments and compared the results with six other classic experimental methods. The algorithms compared were: STARFM (based on weighted functions), FSDAF (based on hybrid algorithms), and DCSTFN, EDCSFN, StfNet, and AMNet (based on learning). In the spatiotemporal fusion comparison experiment, for subjective visual evaluation, the test set was input into the spatiotemporal fusion model to obtain the corresponding fused image. The fused image was then compared with the reference image to observe the differences in spatial and spectral details between the two images. Since subjective visual evaluation is prone to large errors due to inconsistencies in personal perception, environmental factors, and device display, it is often used as an auxiliary evaluation method for objective evaluation methods. In this embodiment, the objective evaluation methods mainly use the Spectral Angle Mapper (SAM), Relative Dimensionless Comprehensive Global Error (ERGAS), Similarity Index (CC), and Structural Similarity Index (SSIM) to evaluate quality.
[0112] Please see Figure 6 Compared to mainstream spatiotemporal fusion methods, the method proposed in this invention, with limited network depth, parameters, and computational resources, significantly improves the feature selection and representation capabilities of the network, reduces spectral distortion to a greater extent, and more fully preserves structural texture information, resulting in better image fusion quality. This improved fusion effect is of great significance for resource management, environmental protection, and economic development.
[0113] In summary, the multi-view, multi-scale feature extraction model constructed in this invention extracts and fuses feature information from different perspectives and scales, improving the network's feature extraction capability. It employs an edge attention mechanism to obtain the correlation of edge information and assigns different weights to edge details, thus enhancing the restoration of texture details when fusing edge features. Simultaneously, by decoding the extracted deep features and structural information through the image decoding module and using its own internal convolution to obtain its own related feature information, the image fusion quality is improved, spectral distortion is reduced to a greater extent, and structural and texture information is preserved more fully, resulting in better image fusion quality. This improved fusion effect is of great significance for land use and legend cover classification, natural disaster monitoring, urban planning and construction, agricultural management, environmental monitoring, and for achieving resource management, environmental protection, and economic development.
[0114] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-view and multi-scale based spatio-temporal fusion method for remote sensing images, characterized in that, Comprise: S1: obtaining and a high temporal, low spatial resolution image at a time instant, a low temporal, high spatial resolution image at a time instant, and a low temporal, high spatial resolution reference image at a time instant; S2: respectively for and Upsample the high temporal and low spatial resolution images at each time point to make them consistent with... The low temporal and high spatial resolution images at each time point are of the same size, resulting in two first intermediate images; S3: stitching the two first intermediate images with the low temporal, high spatial resolution image at the time t to obtain a second intermediate image; and encoding the low temporal, high spatial resolution image at the time t with a first encoder to generate a first intermediate feature map; the low temporal, high spatial resolution image at the time t to obtain a second intermediate image; and encoding the low temporal, high spatial resolution image at the time t with a first encoder to generate a first intermediate feature map; S4: encode the second intermediate image using the second encoder to generate the second intermediate feature map; add the first intermediate feature map and the second intermediate feature map in the feature dimension to obtain the third intermediate feature map; S5: input the third intermediate feature map into the first multi-view multi-scale feature extraction model and the second multi-view multi-scale feature extraction model respectively, and perform feature extraction on the third intermediate feature map in multiple scales and multiple views to obtain the fourth intermediate feature map and the fifth intermediate feature map; The first multi-view multi-scale feature extraction model and the second multi-view multi-scale feature extraction model both comprise two ZBlock modules and an edge attention mechanism module; the ZBlock module is composed of multiple ResZNet residual networks, which perform feature extraction on the third intermediate feature map in multiple views and multiple scales from the length, width and channel of the image respectively, and the edge attention mechanism module is used for weighting the extracted features; wherein the first ZBlock module in the first multi-view multi-scale feature extraction model performs feature extraction on the third intermediate feature map by changing the channel scale from small to large, and the second ZBlock module performs feature extraction by changing the channel scale from large to small; the first ZBlock module in the second multi-view multi-scale feature extraction model performs feature extraction on the third intermediate feature map by changing the channel scale from large to small, and the second ZBlock module performs feature extraction by changing the channel scale from small to large; S6: add the fourth intermediate feature map and the fifth intermediate feature map in the feature dimension to obtain the sixth intermediate feature map; input the sixth intermediate feature map into the attention mechanism model to weight the sixth intermediate feature map to obtain the seventh intermediate feature map; and splice the first intermediate feature map and the seventh intermediate feature map to obtain the eighth intermediate feature map; S7: input the eighth intermediate feature map into a decoder to obtain The high-temporal and high-spatial resolution image fused with the space-time information, and all the models are supervised and trained by using the Adam algorithm according to the error between the high-temporal and high-spatial resolution image fused with the space-time information and the reference image. 2.The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The first encoder comprises, sequentially connected, a first convolutional layer, a first ReLU activation function, a second convolutional layer, a second ReLU activation function, a third convolutional layer, an inner convolutional layer, and a third ReLU activation function; wherein the first convolutional layer, the second convolutional layer, and the third convolutional layer are convolutional layers. 3.The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The second encoder comprises a first convolutional layer, an inner convolutional layer and a second convolutional layer connected in sequence; wherein the first convolutional layer and the second convolutional layer are convolutional layers of .
4. The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The ZBlock module comprises three first ZBlock layers, second ZBlock layers and third ZBlock layers arranged side by side; The first ZBlock layer comprises: a ResZNetC1 residual network, a ResZNetH1 residual network and a ResZNetW1 residual network; The second ZBlock layer comprises: a ResZNetC2 residual network, a ResZNetH2 residual network and a ResZNetW2 residual network; The third ZBlock layer comprises: a ResZNetC3 residual network, a ResZNetH3 residual network and a ResZNetW3 residual network; The ResZNetC1 residual network performs feature extraction on the third intermediate feature map in multiple scales along the channel C direction of the third intermediate feature map to generate a first intermediate feature, and the ResZNetH1 residual network performs feature extraction on the third intermediate feature map in multiple scales along the height H direction of the third intermediate feature map to generate a second intermediate feature; The ResZNetW1 residual network performs feature extraction on the second intermediate feature in multiple scales along the width W direction of the second intermediate feature to generate a third intermediate feature; The first intermediate feature and the third intermediate feature are added in the feature dimension to obtain a first fused feature; The ResZNetC2 residual network extracts features along the channel C direction of the third intermediate feature map to generate a fourth intermediate feature, and the ResZNetH2 residual network extracts features along the height H direction of the third intermediate feature map to generate a fifth intermediate feature; The ResZNetW2 residual network extracts features along the width W direction of the fourth intermediate feature to generate a sixth intermediate feature; The fifth intermediate feature and the sixth intermediate feature are added in the feature dimension to obtain a second fusion feature; The ResZNetC3 residual network extracts features along the channel C direction of the third intermediate feature map to generate a seventh intermediate feature; The ResZNetW3 residual network extracts features along the height W direction of the third intermediate feature map to generate an eighth intermediate feature; The ResZNetH3 residual network extracts features along the height H direction of the seventh intermediate feature to generate a ninth intermediate feature; The eighth intermediate feature and the ninth intermediate feature are added in the feature dimension to obtain a third fusion feature; The first fusion feature, the second fusion feature and the third fusion feature are spliced by using a CONCAT function to obtain an output of the ZBlock module.
5. The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The decoder comprises a first convolutional layer, a ReLU activation function and a second convolutional layer connected in sequence; wherein the first convolutional layer is a 3 by 3 convolutional layer, and the second convolutional layer is a 3 by 3 convolutional layer. 6.The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The pair And Upsampling the high temporal resolution, low spatial resolution image at the moment includes: upsampling by interpolate() provided by PyTorch, and the specific upsampling multiple is .
7. The multi-view and multi-scale based spatio-temporal fusion method of remote sensing images according to claim 1, characterized in that, The loss function when the model is supervised training includes: where X represents the fused generated high temporal, high spatial resolution image at the moment, Y represents the reference image, represents the i-th pixel value in X, represents the i-th pixel value in Y; respectively represent the mean value of X and Y, respectively represent the variance of X and Y, represents the covariance of X and Y, are two stable constants; λ is a weight factor, which changes with the training and is initially set to 0.
Citation Information
Patent Citations
Remote sensing image space-time fusion method based on multi-scale mechanism and attention mechanism
CN111754404A
Remote sensing image fusion method based on large kernel attention mechanism for multi-scale feature enhancement
CN114936995A