Anti-time-varying sensitive remote sensing image space-time fusion method and system

By using a time-sensitive bidirectional convolutional-Transformer generative adversarial network, the problem of insufficient prediction accuracy in the spatiotemporal fusion of remote sensing images was solved, and the effective fusion of high temporal and high spatial resolution remote sensing images was achieved, thereby improving the robustness and accuracy of surface spatial monitoring.

CN122048670APending Publication Date: 2026-05-15GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610244465.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing spatiotemporal fusion methods for remote sensing images lack sufficient prediction accuracy in scenarios involving short-term abrupt changes, long-interval variations, and changes in land cover type. They are also susceptible to interference from prior data, making it difficult to achieve high temporal-spatial resolution remote sensing image fusion.

Method used

A time-sensitive bidirectional convolutional-transformer generative adversarial network is adopted. Through a time-sensitive bidirectional convolutional-transformer generator and a convolutional-transformer discriminator with multi-resolution input, it captures prior and arbitrary time-varying local-global features, dynamically calculates the correlation between spectral, spatial and time-varying information, and performs adversarial learning through a dual-guided three-attention fusion decoder and a discriminator with multi-resolution input to generate high-precision remote sensing images.

Benefits of technology

It significantly improves the robustness and accuracy of spatiotemporal fusion of remote sensing images, enabling better monitoring of changes in the Earth's surface, reducing the impact of prior data interference and resolution differences, and providing more reliable support for refined monitoring of the Earth's surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048670A_ABST
    Figure CN122048670A_ABST
Patent Text Reader

Abstract

The invention discloses an anti-time-varying sensitive remote sensing image space-time fusion method and system, and belongs to the technical field of remote sensing image processing, and the method comprises the steps: obtaining a priori moment fine resolution image, a priori moment coarse resolution image and a prediction moment coarse resolution image; extracting multi-scale prior detail semantic features from the prior moment fine resolution image through a resolution reduction encoder; multi-scale semantic features are extracted from the coarse resolution images at the prior moment and the prediction moment through a super-resolution encoder, and time-varying semantic features are obtained through subtraction operation; inputting the two types of features into a double-orientation three-attention fusion decoder, and sequentially generating a preliminary local-global cross fusion feature and a decision fusion feature; the reconstruction unit reconstructs the decision fusion feature into a prediction moment fine resolution image; and performing an adversarial learning optimization generation process on the predicted image and the real image. According to the method, the prediction robustness of time-space fusion on time-varying information is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, specifically to a time-sensitive remote sensing image spatiotemporal fusion method and system. Background Technology

[0002] Spatiotemporal fusion is an efficient method to provide high temporal and spatial resolution land cover observation products, offering policymakers timely and accurate data on land surface change monitoring. Single-source sensor data cannot simultaneously meet the demands for high spatial and temporal resolution. Spatiotemporal fusion of multi-source remote sensing imagery is an effective and convenient way to acquire high temporal and spatial resolution multispectral imagery.

[0003] Current spatiotemporal fusion models can be broadly categorized into three types based on the prior data strategies employed: First, generating a fine-resolution image for the predicted date using two pairs of coarse-to-fine resolution imagery for two prior dates and a coarse-resolution image for the predicted date. Second, generating a fine-resolution image for the predicted date using one pair of coarse-to-fine resolution imagery for one prior date and a coarse-resolution image for the predicted date. Third, generating a fine-resolution image for the predicted date using a pair of fine-resolution imagery for one prior date and a coarse-resolution image for the predicted date. While deep learning-based spatiotemporal fusion methods have developed rapidly, they still face the following technical challenges: Current deep learning-based spatiotemporal fusion models show superior prediction results for gradual phenological change scenarios, but their prediction accuracy still has significant room for improvement in scenarios with rapid abrupt changes, long-interval changes, and changes in land cover type. Although most fusion accuracy based on a two-pair prior data strategy is superior to the other two strategies, obtaining pairs of high-quality, cloud-free coarse-to-fine resolution remote sensing images for actual observations of changing scenarios is difficult, hindering practical applications. The spatiotemporal fusion method based on the third data strategy only utilizes coarse-resolution images from a single moment. The significant resolution difference between coarse and fine-resolution images results in very limited usable information, making it difficult to fully utilize prior spatiotemporal difference information. Furthermore, it suffers from sensor imaging differences compared to fine-resolution images. The second data strategy can provide spatiotemporal difference information between the prior and predicted moments, but the fusion result is easily influenced by prior information. In summary, spatiotemporal fusion technology still faces challenges such as severe prediction distortion in short-term abrupt changes, long-interval variations, and land cover type changes, as well as significant interference from prior data. Summary of the Invention

[0004] To address the aforementioned issues, this invention provides a time-sensitive remote sensing image spatiotemporal fusion method and system, aiming to improve the robustness and fusion capability of remote sensing image spatiotemporal fusion.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] On one hand, embodiments of the present invention provide a spatiotemporal fusion method for time-varying sensitive remote sensing images, the method comprising the following steps:

[0007] S100: Obtain the prior time-time fine-resolution image, the prior time-time coarse-resolution image, the prediction time coarse-resolution image, and the trained time-resistant bidirectional convolutional-transformer generative adversarial network; wherein, the time-resistant bidirectional convolutional-transformer generative adversarial network includes a time-resistant bidirectional convolutional-transformer generator and a multi-resolution input convolutional-transformer discriminator, and the time-resistant bidirectional convolutional-transformer generator includes a down-resolution encoder, a super-resolution encoder, a dual-guided three-attention fusion decoder, and a reconstruction unit;

[0008] S200, the prior time fine resolution image, the prior time coarse resolution image and the prediction time coarse resolution image are input into the trained time-resistant sensitive bidirectional convolutional-transformer generative adversarial network, and the prior time fine resolution image is extracted from the prior time coarse resolution image through the down-resolution encoder to obtain prior detail semantic features.

[0009] S300, the super-resolution encoder performs multi-scale feature extraction from low resolution to high resolution on the prior time coarse resolution image and the predicted time coarse resolution image to obtain time-varying semantic features;

[0010] S400, the prior detail semantic features and the time-varying semantic features are input into the dual-guided three-attention fusion decoder, and the preliminary local-global cross-fusion features are generated by dual-guided cross-convolution Transformer fusion. Then, the preliminary local-global cross-fusion features are weighed and integrated step by step by decision attention fusion to obtain the decision fusion features.

[0011] S500, the decision fusion features are reconstructed by the reconstruction unit to generate a high-resolution image at the prediction time; the high-resolution image at the prediction time and the corresponding real image are input into the multi-resolution input convolution-Transformer discriminator for adversarial learning to optimize the generation process of the high-resolution image at the prediction time.

[0012] On the other hand, embodiments of the present invention provide a time-sensitive remote sensing image spatiotemporal fusion system, comprising: at least one processor; at least one memory for storing at least one program; and when the at least one program is executed by the at least one processor, the at least one processor implements the above-described method.

[0013] On the other hand, embodiments of the present invention provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the above-described method.

[0014] The beneficial effects of this invention are as follows: This invention discloses a time-sensitive remote sensing image spatiotemporal fusion method and system. This invention captures prior and arbitrarily time-varying local-global features through a time-sensitive bidirectional encoder, dynamically calculates the correlation between spectral, spatial, and time-varying information using a dual-guided three-attention fusion decoder, and aggregates heterogeneous information. A discriminator with multi-resolution input is introduced to adversarially learn local-global structures and spectral information at different resolutions. This method overcomes the problems of severe prediction distortion and sensitivity to prior data interference in current spatiotemporal fusion methods under scenarios of short-term abrupt changes, long-interval variations, and changes in land cover type. It significantly improves the robustness and accuracy of spatiotemporal fusion, providing more reliable technical support for refined monitoring of land surface space. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating a time-sensitive remote sensing image spatiotemporal fusion method according to an embodiment of the present invention.

[0016] Figure 2 This is a model framework diagram of the time-sensitive bidirectional convolutional-Transformer generative adversarial network in an embodiment of the present invention;

[0017] Figure 3 This is a structural diagram of the time-sensitive bidirectional convolution-Transformer generator in an embodiment of the present invention;

[0018] Figure 4 This is a structural diagram of the down-resolution encoder and the super-resolution encoder in the embodiments of the present invention;

[0019] Figure 5 This is a structural diagram of each component unit of the model in this embodiment of the invention;

[0020] Figure 6 This is a structural diagram of the dual-guided cross-convolution Transformer in an embodiment of the present invention;

[0021] Figure 7 This is a structural diagram of the transposed attention of the cross-depth convolution head in an embodiment of the present invention;

[0022] Figure 8 This is a structural diagram of decision attention fusion in an embodiment of the present invention;

[0023] Figure 9This is a structural diagram of the multi-resolution input convolutional Transformer discriminator in an embodiment of the present invention;

[0024] Figure 10 This is a comparison chart of the prediction results of CIA test data in an embodiment of the present invention;

[0025] Figure 11 This is a comparison chart of the prediction results of LGC test data in an embodiment of the present invention. Detailed Implementation

[0026] The following will provide a clear and complete description of the concept, specific structure, and technical effects of the present invention in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0027] Spatiotemporal fusion models in related technologies can be broadly categorized into three types based on the prior data strategies employed: First, generating a fine-resolution image of the predicted date using two pairs of coarse-to-fine resolution images of prior dates and a coarse-resolution image of the predicted date; second, generating a fine-resolution image of the predicted date using one pair of coarse-to-fine resolution images of prior dates and a coarse-resolution image of the predicted date; and third, generating a fine-resolution image of the predicted date using a pair of fine-resolution images of prior dates and a coarse-resolution image of the predicted date. In recent years, the technologies used in spatiotemporal fusion have developed rapidly and can be broadly classified into weighted spatiotemporal fusion methods, unmixing-based spatiotemporal fusion methods, hybrid spatiotemporal fusion methods, and learning-based spatiotemporal fusion methods.

[0028] Current deep learning-based spatiotemporal fusion models have achieved promising prediction results, particularly in scenarios of gradual phenological changes. However, there is still significant room for improvement in prediction accuracy for scenarios with rapid abrupt changes, long-term changes, and changes in land cover type. While most fusion methods based on two pairs of prior data achieve higher accuracy than those based on other two strategies, obtaining pairs of high-quality, cloud-free coarse-to-fine resolution remote sensing images for changing scenarios is difficult in practical applications. Spatiotemporal fusion methods based on a third data strategy only utilize coarse-resolution images from a single moment. The significant resolution difference between coarse and fine-resolution images limits the available information, making it difficult to fully utilize prior spatiotemporal differences, and there are also sensor imaging differences between coarse and fine-resolution images. The second data strategy can provide spatiotemporal difference information between prior and predicted moments, but the fusion results are easily affected by prior moment information.

[0029] To address the above problems, this invention proposes a time-varying and sensitive spatiotemporal fusion method and system for refined monitoring of the Earth's surface. Ablation experiments and comparative experiments were conducted on the widely used public datasets CIA and LGC. The results demonstrate the superior fusion capability and time-varying robustness of the proposed model. The main contents are as follows:

[0030] (1) A time-sensitive bidirectional convolution-Transformer generative adversarial network is proposed, which includes a time-sensitive bidirectional convolution-Transformer generator and a convolution-Transformer discriminator with multi-resolution input, thereby improving the robustness of prediction of time-varying information and the spatiotemporal fusion capability.

[0031] (2) Time-sensitive bidirectional convolution-Transformer generator: A time-sensitive bidirectional encoder is designed to capture prior and arbitrary time-varying local-global features in both directions to improve the robustness of prediction of changing information and the ability to represent time-varying information.

[0032] (3) Anti-time-varying sensitivity bidirectional convolution-Transformer generator: A dual-guided three-attention fusion decoder was designed, and dual-guided cross-convolutional attention fusion and decision attention fusion were proposed. The correlation between spectral, spatial and time-varying information was dynamically calculated, heterogeneous information was aggregated, and adaptive step-by-step trade-offs and integration were performed to reduce the differences in heterogeneous imaging mechanisms and huge resolution differences.

[0033] (4) A convolutional-transformer discriminator with multi-resolution input is introduced to learn the local-global structure and spectral information of different resolutions and feed them back to the generator to generate more detailed images. A composite loss function is designed to form deep supervision to optimize the model's capabilities.

[0034] refer to Figure 1 ,like Figure 1 The image shown is an embodiment of the present invention providing a spatiotemporal fusion method for time-varying sensitive remote sensing images. The method includes the following steps:

[0035] S100: Obtain the prior time-time fine-resolution image, the prior time-time coarse-resolution image, the prediction time coarse-resolution image, and the trained time-resistant bidirectional convolutional-transformer generative adversarial network; wherein, the time-resistant bidirectional convolutional-transformer generative adversarial network includes a time-resistant bidirectional convolutional-transformer generator and a multi-resolution input convolutional-transformer discriminator, and the time-resistant bidirectional convolutional-transformer generator includes a down-resolution encoder, a super-resolution encoder, a dual-guided three-attention fusion decoder, and a reconstruction unit;

[0036] S200, the prior time fine resolution image, the prior time coarse resolution image and the prediction time coarse resolution image are input into the trained time-resistant sensitive bidirectional convolutional-transformer generative adversarial network, and the prior time fine resolution image is extracted from the prior time coarse resolution image through the down-resolution encoder to obtain prior detail semantic features.

[0037] S300, the super-resolution encoder performs multi-scale feature extraction from low resolution to high resolution on the prior time coarse resolution image and the predicted time coarse resolution image to obtain time-varying semantic features;

[0038] S400, the prior detail semantic features and the time-varying semantic features are input into the dual-guided three-attention fusion decoder, and the preliminary local-global cross-fusion features are generated by dual-guided cross-convolution Transformer fusion. Then, the preliminary local-global cross-fusion features are weighed and integrated step by step by decision attention fusion to obtain the decision fusion features.

[0039] S500, the decision fusion features are reconstructed by the reconstruction unit to generate a high-resolution image at the prediction time; the high-resolution image at the prediction time and the corresponding real image are input into the multi-resolution input convolution-Transformer discriminator for adversarial learning to optimize the generation process of the high-resolution image at the prediction time.

[0040] In the embodiments provided by this invention, a time-sensitive bidirectional encoder can adaptively capture prior and arbitrarily time-varying local-global features, thereby more effectively improving the robustness of predicting changing information. The dual-guided three-attention fusion decoder employs dual-guided cross-convolutional Transformer fusion and decision attention fusion to dynamically calculate the correlation between spectral, spatial, and time-varying information and aggregate heterogeneous information, reducing the impact of differences in heterogeneous imaging mechanisms and large resolution differences. A multi-resolution input convolutional-Transformer discriminator guides the generator to produce more refined images by adversarially learning local-global structures and spectral information at different resolutions. This series of processing steps not only improves the accuracy of spatiotemporal fusion of remote sensing images but also significantly enhances the system's robustness to time-varying information, providing more reliable technical support for refined monitoring of the Earth's surface.

[0041] In some embodiments, in S200, the step of extracting multi-scale features from high resolution to low resolution in the prior time-time fine-resolution image using the down-resolution encoder to obtain prior detail semantic features includes:

[0042] S210, low-level semantic features are extracted from the prior time-time fine-resolution image through multi-scale dilated convolution to obtain initial prior time-time fine features; wherein, the multi-scale dilated convolution uses dilated convolutions with dilation rates of 1, 2 and 3 in parallel to extract features from the prior time-time fine-resolution image, and the extracted multi-scale features are concatenated along the channel direction and then initially fused through 1×1 convolution to obtain the initial prior time-time fine features.

[0043] S220, the initial prior time-time fine features are input into a resolution-reducing feature extraction branch consisting of four convolutional Transformer modules. Each convolutional Transformer module consists of layer normalization, multi-head deep convolutional transpose attention, local enhancement forward pass network and residual connection. After local-global feature extraction of the input features, each convolutional Transformer module reduces the resolution through downsampling operation to obtain the prior detailed semantic features.

[0044] In this embodiment, features with different dilation rates are extracted in parallel through multi-scale dilated convolution, effectively capturing spatial information under different receptive fields and enhancing the model's ability to perceive multi-scale structures. The convolutional Transformer module combines the local feature extraction capability of convolution with the global dependency modeling capability of Transformer, while using depthwise convolution to reduce computational cost, making it suitable for high-resolution image processing. Downsampling operations progressively reduce the resolution, forming a multi-scale feature pyramid, providing rich hierarchical feature representations for subsequent fusion.

[0045] In some embodiments, S220, the multi-head depthwise convolutional transpose attention processes the input features, including:

[0046] S221: After layer normalization of the input features, pixel-level cross-channel context is aggregated through 1×1 convolution, and then spatial context is separated by 3×3 depth convolution to generate query matrix, key matrix and value matrix.

[0047] S222, After performing transformation operations on the query matrix and the key matrix respectively, perform a dot product operation, divide the dot product result by a learnable parameter to scale, and perform a Softmax operation on the scaled result to obtain a first attention map;

[0048] S223, after transforming the value matrix, multiply it with the first attention map, and after transforming the multiplication result, perform convolution processing to obtain the attention-corrected features;

[0049] S224, the attention-corrected features are added to the input features to obtain the local-global features of the multi-head depth convolution transposed attention output.

[0050] In this embodiment, the multi-head depthwise convolution transposed attention computes attention in the channel dimension rather than the spatial dimension through the transposed attention mechanism, significantly reducing computational complexity. The use of depthwise convolution further reduces the number of parameters, enabling the model to efficiently process high-resolution remote sensing images. The scaling mechanism for learnable parameters adaptively adjusts the attention distribution, avoiding gradient vanishing or exploding problems and improving training stability.

[0051] In some embodiments, in S300, the step of extracting multi-scale features from low resolution to high resolution in the prior time coarse-resolution image and the predicted time coarse-resolution image using the super-resolution encoder to obtain time-varying semantic features includes:

[0052] S310, low-level semantic features are extracted from the prior time coarse resolution image and the prediction time coarse resolution image respectively through multi-scale dilated convolution to obtain the initial prior time coarse features and the initial prediction time coarse features.

[0053] S320, the initial prior time coarse features and the initial prediction time coarse features are respectively input into the super-resolution feature extraction branch composed of 4 convolutional Transformer modules. Each convolutional Transformer module performs local-global feature extraction on the input features and then improves the resolution through upsampling operation to obtain the prior time multi-resolution semantic features and the prediction time multi-resolution semantic features.

[0054] S330, perform an element-wise subtraction operation on the prior time multi-resolution semantic features and the predicted time multi-resolution semantic features to obtain the time-varying semantic features.

[0055] In this embodiment, the super-resolution encoder gradually increases the resolution through upsampling operations, effectively uncovering key structural and spectral information hidden in the coarse-resolution data. Subtraction operations explicitly extract the change information between any prior time and the prediction time, providing direct time-varying features of arbitrary time intervals for subsequent fusion. This bidirectional encoding strategy enables the model to simultaneously obtain prior details from fine-resolution images and effective time-varying information from coarse-resolution images, achieving information complementarity.

[0056] In some embodiments, S400, the generation of preliminary local-global cross-fusion features through dual-guided cross-convolution Transformer fusion includes:

[0057] S411, take the prior detail semantic features as prior semantic features and the time-varying semantic features as time-varying information features, and perform layer normalization on the prior semantic features and the time-varying information features at the i-th resolution respectively to obtain prior normalized features and time-varying normalized features.

[0058] S412, in the prior semantic feature-guided cross-attention fusion, the time-varying normalized features are mapped to generate a query matrix through 1×1 convolution and 3×3 depth convolution, and the prior normalized features are mapped to generate a key matrix and a value matrix through 1×1 convolution and 3×3 depth convolution; in the time-varying information feature-guided cross-attention fusion, the prior normalized features are mapped to generate a query matrix through 1×1 convolution and 3×3 depth convolution, and the time-varying normalized features are mapped to generate a key matrix and a value matrix through 1×1 convolution and 3×3 depth convolution.

[0059] S413, after performing transformation operations on the generated query matrix and key matrix respectively, perform dot product operation, divide the dot product result by the learnable parameter to scale, and perform Softmax operation on the scaled result to obtain the second attention map;

[0060] S414, after transforming the generated value matrix, multiply it with the second attention map, transform the multiplication result, process it through convolution, and then add it to the prior semantic features and time-varying information features respectively to generate two directional preliminary cross-attention fusion features;

[0061] S415, after normalizing the preliminary cross-attention fusion features of the two directions, input them into the local enhancement forward pass network for local enhancement, and add them with the preliminary cross-attention fusion features of the two directions to obtain the enhanced local-global cross-fusion features of the two directions.

[0062] S416, the two guided enhanced local-global cross-fusion features are added together to generate the preliminary local-global cross-fusion feature.

[0063] In this embodiment, the dual-direction cross-attention mechanism achieves full interaction between prior details and time-varying information through two different information flows. In the prior semantic feature-guided branch, time-varying information serves as the query, and prior details serve as the key and value, which helps to dynamically extract and integrate time-varying information from prior details. In the time-varying information feature-guided branch, the roles are reversed, which helps to dynamically strengthen prior details from time-varying information. This bidirectional interactive design can dynamically calculate the correlation between spectral, spatial, and time-varying information, effectively aggregate heterogeneous information, and reduce the negative impact of differences in heterogeneous imaging mechanisms and huge resolution differences.

[0064] In some embodiments, S400, the step of progressively weighing and integrating the preliminary local-global cross-fusion features through decision attention fusion to obtain decision fusion features includes:

[0065] S421, the preliminary local-global cross-fusion feature at resolution i and the decision fusion feature at resolution i+1 are subjected to spectral attention enhancement and spatial attention enhancement, respectively, to obtain two enhanced features; wherein, the spectral attention enhancement compresses each spectral channel of the input feature through global average pooling and global max pooling, and then learns the nonlinear dependency between spectral channels through two convolutional layers to generate a spectral channel attention map, and multiplies the spectral channel attention map with the corresponding channel of the input feature to achieve spectral calibration; the spatial attention enhancement extracts spatial features along the channel dimension through global average pooling and global max pooling, concatenates the extracted spatial features along the channel and fuses them through a convolutional layer to generate a spatial attention map, and multiplies the spatial attention map with each channel of the input feature to achieve spatial region enhancement;

[0066] S422, the preliminary local-global cross-fusion feature at resolution i is concatenated with the decision fusion feature at resolution i+1, and then fused through a convolutional layer to obtain the intermediate fusion feature;

[0067] S423, add the two enhanced features to the intermediate fusion feature to obtain the decision fusion feature at the i-th resolution.

[0068] In this embodiment, decision attention fusion achieves a coarse-to-fine layered decision-making process through multi-resolution cascading. The spectral attention mechanism can adaptively assess the importance of different spectral bands, enhancing the spectral features crucial for land cover identification; the spatial attention mechanism can adaptively assess the importance of different spatial locations, enhancing key spatial information such as change regions and boundaries. This step-by-step trade-off and integration strategy can adaptively adjust the weighting of the fusion results at each stage, effectively reducing spurious change information that may be generated by different objects sharing the same spectrum or the same object having different spectra, and improving the fusion accuracy of dynamic change information such as details and spectra.

[0069] In some embodiments, S500, the step of performing adversarial learning on the convolutional-Transformer discriminator that inputs the predicted time-time fine-resolution image and the corresponding real image into the multi-resolution input to optimize the generation process of the predicted time-time fine-resolution image includes:

[0070] S510, the high-resolution image at the prediction time and its corresponding downsampled images by 2x and 4x are respectively input into the multi-resolution input convolution-Transformer discriminator along with the corresponding real image and its downsampled version;

[0071] S520 extracts primary local features through convolutional layers and the LeakyReLU activation function, and extracts local-global features through convolution-Transformer.

[0072] S530 reduces the feature size sequentially through at least three convolutional layers with a kernel size of 4×4 and a stride of 2, and then stitches and fuses them with image features of different resolutions.

[0073] The S540 uses convolutional layers, spectral normalization operations, and the Sigmoid function to output the discrimination results, thereby determining the relative authenticity of the input image.

[0074] In this embodiment, the discriminator design with multi-resolution inputs enables the model to simultaneously learn realistic features at different scales. High-resolution inputs focus on adversarial learning of fine-grained textures, edges, and small features; low-resolution inputs focus on adversarial learning of global information and relationships. The introduction of spectral normalization accelerates the model's convergence and improves training stability. The relative average least squares discriminator further enhances the stability and performance of the fusion model by judging the relative realism of the input images.

[0075] In some embodiments, the trained time-sensitive bidirectional convolutional-Transformer generative adversarial network is trained using a composite loss function, which includes a generator loss and a discriminator loss. The generator loss includes relative average least squares adversarial loss, pixel-level mean square error loss, multi-scale structural similarity loss, spectral angle loss, and feature loss. The discriminator loss is a relative average least squares loss. The composite loss function uses weighted summation of the various losses through weight coefficients to form deep supervision to optimize the model training process.

[0076] In this embodiment, the composite loss function constrains the generation process from multiple levels through multi-task learning. Pixel-level mean squared error loss ensures the closeness of the generated image to the real image in terms of pixel values; multi-scale structural similarity loss evaluates spatial structural consistency from three perspectives: brightness, contrast, and structure; spectral angle loss controls the degree of spectral distortion of the generated image; feature loss constrains high-level semantic consistency in the feature space of the pre-trained model VGG-19; and adversarial loss improves the realism of the generated image through a game-theoretic mechanism. This deep supervision mechanism can comprehensively optimize the model's capabilities and improve the radiometric, structural, and spectral fidelity of the fusion result.

[0077] The following are specific embodiments provided by the present invention:

[0078] Invention purpose and fusion model framework:

[0079] To address the challenges in current heterogeneous remote sensing spatiotemporal fusion, such as the significant resolution gap between fine and coarse resolution images, the substantial influence of prior data and time-varying factors, and the limited information represented by coarse resolution, this paper proposes a time-sensitive remote sensing image spatiotemporal fusion method and system for refined land spatial monitoring. The proposed time-sensitive bidirectional convolutional-transformer generative adversarial network (TRBCT-GAN) model framework is as follows: Figure 2 As shown, the TRBCT-GAN model consists of a time-sensitive bidirectional convolutional-Transformer generator and a multi-resolution input convolutional-Transformer discriminator, both adversarially learning time-varying information and prior details. First, when there are significant differences in time-varying information or abrupt scene changes between the test data and training samples, existing spatiotemporal fusion models often lead to distortions in the details and spectral accuracy of predictions for areas with drastic time changes, reducing their ability to capture real surface changes and accurately predicting changes in land cover. To address this challenge, this invention enhances the model's adaptive learning ability for changes in information over arbitrary time intervals, improving robustness and prediction accuracy under complex time-varying conditions, thereby reducing the performance degradation caused by differences in data distribution and abrupt scene changes. Second, information in coarse-resolution images is very limited; a single pixel often contains multiple land cover types, resulting in blurred surface details and boundary information, making it difficult to directly and effectively extract time-varying information from pure pixels. Moreover, the mixed pixel effect exacerbates the problem of different objects having the same spectral density and the same object having different spectral density, interfering with the model's ability to discriminate the true categories of land covers. In areas with drastic time-varying changes or sudden scene shifts, the aforementioned problems are further amplified, making it difficult to accurately predict rapid dynamic changes in land cover. To address this challenge, this invention performs super-resolution reconstruction of the time-varying information in coarse-resolution images, gradually improving spatial resolution and highlighting key structural and spectral information implicit in the coarse-resolution data. For example... Figure 2 As shown, the fine-resolution image at any time step is extracted from high to low resolution by a down-resolution encoder and a super-resolution encoder, respectively. Multi-scale features and extraction of coarse-resolution images at any time from low to high resolution Coarse-resolution image for prediction time The multi-scale features are obtained by subtraction operations to obtain the variation features at any time interval. Different sensors exhibit varying imaging characteristics due to their different imaging mechanisms. To reduce these differences and effectively aggregate complementary information from two sources, this invention designs a dual-guided, three-attention fusion decoder. Guided by both prior details and time-varying semantic features, it aggregates complementary information of spatial structure and semantic content through parallel attention pathways. To reduce the uncertainty and bias introduced by a single large-scale resolution upscaling, the fusion decoder utilizes small-scale ratios, multi-stage decision-making attention fusion for progressive upscaling, and reconstruction of intermediate resolution images. Furthermore, a multi-scale discriminator is used for hierarchical adversarial learning of local-global multi-resolution details and time-varying information to maintain the consistency of spatiotemporal evolution information.

[0080] Time-sensitive generator:

[0081] This invention proposes a time-resistant, sensitive bidirectional convolution-Transformer generator (TRBG), with the following structure: Figure 3 As shown, TRBG includes a down-resolution encoder, a super-resolution encoder, a dual-guided three-attention fusion decoder, and a reconstruction unit. The down-resolution encoder and super-resolution encoder extract multi-scale local-global prior features and time-varying features. The dual-guided three-attention fusion decoder consists of dual-guided cross-convolution-Transformer fusion and decision attention fusion, progressively aggregating complementary information of spatial structure and semantic content. The reconstruction unit is used to reconstruct the fused information to obtain an image at ideal resolution.

[0082] Time-sensitive bidirectional encoder:

[0083] When there are significant differences in time-varying information or abrupt scene changes between test data and training samples, existing spatiotemporal fusion models often lead to distortion of details and spectral accuracy in predictions of areas with drastic time variations, resulting in a decreased ability to capture real changes in land cover and an inability to accurately predict changes in land cover. To address this challenge, this invention obtains coarse-resolution images of the prediction time phase. Coarse-resolution images of any time phase The spatiotemporal differences enhance the adaptive learning ability of information changing at any time interval, thereby reducing the degradation of fusion performance caused by changes in prior data and sudden scene changes. Furthermore, coarse-resolution images often contain multiple land cover types within a single pixel, making it difficult to effectively extract time-varying information of different land cover types. Moreover, the mixed pixel effect exacerbates the problem of different land cover types sharing the same spectral density and the same land cover having different spectral densities, interfering with the model's ability to distinguish the true land cover categories. In areas with drastic time-varying or sudden scene changes, these problems are further amplified, making it difficult to accurately predict rapid dynamic changes in land cover. This invention performs super-resolution reconstruction of the time-varying information in coarse-resolution images, gradually improving spatial resolution and highlighting key structural and spectral information implicit in the coarse-resolution data. Considering the advantages of CNNs in Restormer and Uforme in extracting local features and Transformers in establishing long-range dependencies, and the fact that deep convolutional Transformers reduce computation and are suitable for high-resolution images, this invention designs a time-sensitive bidirectional encoder, such as... Figure 3 As shown, it consists of a down-resolution encoder and a super-resolution encoder of convolutional Transformer, which extract prior features from the high-to-low resolution direction and time-varying features from the low-to-high resolution direction, respectively.

[0084] like Figure 3 As shown, the down-resolution encoder extracts fine-resolution images at any given time from high to low resolution. Multi-resolution prior features The super-resolution encoder extracts coarse-resolution images at any time from low to high resolution. Coarse-resolution image for prediction time The multi-scale features are then used to obtain the variation features at any time interval through a subtraction operation. The structures of down-resolution encoders and super-resolution encoders are as follows: Figure 4 (a) and Figure 4 (b) The down-resolution encoder consists of multi-scale dilated convolution, convolutional Transformer and downsampling, while the super-resolution encoder consists of multi-scale dilated convolution, convolutional Transformer and upsampling.

[0085] refer to Figure 5 , Figure 5 The structures of each component are as follows: (a) multi-scale dilated convolution, (b) convolution-transformer, (c) multi-head depthwise convolutional transposed attention (MDTA), (d) locally enhanced forward pass network (LEFN), and (g) downsampling and upsampling.

[0086] Multi-scale dilated convolution is used to extract image, and Multi-scale local features of low-level semantics in images, structural composition such as Figure 5 (a) Extracted in parallel from convolutions with dilatancy ratios of 1, 2 and 3. , and The multi-scale features are then concatenated and preliminarily fused and adjusted using 1×1 convolution. Conv(k3d1) indicates that the convolution kernel is 3×3 and the dilation rate is 1. The process of extracting low-level semantic features by multi-scale dilated convolution is expressed as equations (1)-(3).

[0087] (1);

[0088] (2);

[0089] (3);

[0090] in, , and They represent extraction , and Low-level semantic features , and These represent dilated convolution operations with a kernel size of 3 and dilation rates of 1, 2, and 3, respectively. This represents a Leaky Rectified Linear Unit (LReLU) function. This represents a splicing operation along the channel direction. This indicates a convolution operation with a kernel of 1 followed by the LReLU activation function.

[0091] To effectively capture high-level, multi-scale prior details and time-varying information, a multi-resolution feature extraction module combining convolutional Transformer with downsampling or upsampling is introduced. For example... Figure 4 As shown in (a) and (b), the down-resolution encoder consists of four convolutional Transformers and three downsampling modules, extracting prior detail semantic features at four different resolutions from high resolution to low resolution. The super-resolution encoder consists of four convolutional Transformers and three upsampling modules, extracting... and Four semantic features at different resolutions are used to obtain time-varying semantic features at arbitrary intervals through a subtraction operation. The structures of convolutional Transformer, downsampling, and upsampling are as follows: Figure 5 (b) and Figure 5As shown in (e), the downsampling and upsampling operations are implemented by pixel de-reconstruction and pixel reconstruction, respectively, and the convolution (k3) is a 3×3 kernel convolution to adjust the number of channels. Extraction , and The implementation process of high-level multi-resolution semantic features is expressed as Equation (4)-(6), and the implementation of time-varying semantic features is expressed as Equation (7).

[0092] (4);

[0093] (5);

[0094] (6);

[0095] (7);

[0096] in, and It is the extracted first Resolution prior detail semantic features and time-varying semantic features and extract and The Resolution semantic features The function representing the convolution Transformer, and These represent downsampling and upsampling operations, respectively.

[0097] Because of the limitations of the local receptive field in CNNs, the global structure is ignored, and extracting only local features is insufficient to reconstruct high-quality images. Ordinary self-attention is suitable for capturing long-range dependencies but is computationally expensive. Therefore, a Multihead Depthwise Convolutional Transposed Attention (MDTA) module and a Local Enhancement Forward Network (LEFN) are introduced to collaboratively learn long-range dependencies within the image, enhance local features, and reduce computational cost. The structure of the Convolutional-Transformer is as follows... Figure 5 As shown in (b), the Convolutional Transformer consists of layer normalization, Multihead Depthwise Convolutional Transposed Attention (MDTA), a Local Enhancement Forward Network (LEFN), and residual connections. MDTA is similar in structure to the Multihead Depthwise Transposed Attention in the Restormer, as follows: Figure 5 As shown in (c). Specifically, as... Figure 5 (b) and Figure 5 (c) The input image Y is normalized to obtain LY; then a 1×1 convolutional layer is used to gather pixel-level cross-channel context, and a 3×3 depthwise convolutional layer is used to encode the channel separation spatial context and generate query (Q), key (K) and value (V); Q and K are transformed, dot product and softmax operation to obtain attention map SA; V is transformed and multiplied with SA, transformed and convolved to obtain SA-corrected features. , The local-global feature RY is obtained by adding it to the input Y, and the process is expressed as in equations (8)-(10).

[0098] (8);

[0099] (9);

[0100] (10);

[0101] Where Y, LY, and RY are the input features, layer-normalized features, and local-global features corrected by MDTA, respectively. It is a layer normalization operation. and It generates Q using 1×1 convolutions and 3×3 depthwise convolutions. and It generates K 1×1 convolutions and 3×3 depthwise convolutions. and It generates a 1×1 convolution and a 3×3 depthwise convolution for V. It is a learnable parameter used to adjust the amplitude. It is a transformation operation. It is a transformation operation. This represents a convolution operation with a kernel of 1.

[0102] like Figure 5 As shown in (b), RY is normalized by the layer and then passed to LEFN to enhance the local context. The enhanced feature ERY is then output via residual connections. The structure of LEFN is shown in the figure. Figure 5 (d) The implementation process is expressed as in equation (11).

[0103] (11);

[0104] in, It is the activation function of the Gaussian error linear unit (GELU). This indicates a depthwise convolution operation with a kernel of 3.

[0105] With the collaboration of MDTA and LEFN, the convolutional Transformer obtains local features and global correlations of regions, which helps to reconstruct high-quality remote sensing images.

[0106] Dual-direction three-attention fusion decoder:

[0107] To reduce the impact of differences in imaging mechanisms and large resolution differences, and to improve the perception and prediction capabilities of dynamically changing scenes, a dual-guided three-attention fusion decoder is proposed, such as... Figure 3 As shown, the dual-guided three-attention fusion decoder designs a dual-guided cross-convolution Transformer, namely, cross-attention fusion guided by prior semantic features and cross-attention fusion guided by time-varying semantic information, aggregating heterogeneous prior and time-varying information to dynamically calculate the correlation between spectral and spatial features; it also designs a decision attention fusion to perform high-order trade-offs and integration of the cross-fusion results at each level. The implementation of the dual-guided three-attention fusion decoder can be divided into the following two modules.

[0108] (1) Double-guided cross-convolution Transformer;

[0109] like Figure 3 As shown, to aggregate heterogeneous prior information and time-varying information, a dual-guided cross-convolutional Transformer (DGCT) was designed to fuse local-global prior semantic features at different resolutions. and time-varying semantic features Generate preliminary local-global cross-fusion features The structure of DGCT is as follows: Figure 5 Including prior semantic features Guided cross-attention fusion and time-varying semantic information The guided cross-attention fusion consists of layer normalization, cross-depthwise-convolution transposed attention (CDTA), LEFN, and residual connections.

[0110] refer to Figure 6 , Figure 6 The diagram shows the structure of a dual-guided cross-convolutional Transformer. Prior semantic features. Guided cross-attention fusion prior and Perform layer normalization to obtain normalized features , and Then, CDTA cross-fusion is used to generate correction features, and combined with... The initial cross-attention fusion features are obtained by summing them, and then passed to LEFN for further local enhancement. The implementation process is expressed as equations (12)-(15). The implementation structure of CDTA is as follows: Figure 7 The difference between CDTA and MDTA lies in the way Q, K, and V are obtained. In the prior semantic feature-guided cross-attention fusion, Q in CDTA is obtained by... K and V are obtained by mapping through 1×1 convolution and 3×3 depthwise convolution. It is obtained by mapping 1×1 convolution and 3×3 depth convolution, and the implementation process is expressed as in equation (13).

[0111] (12);

[0112] (13);

[0113] (14);

[0114] (15);

[0115] in, yes Normalized features are used to generate Q. and yes The normalized features are used to generate K and V. Is Preliminary cross-attention fusion features generated under guidance Is Enhanced local-global cross-fusion features generated under guidance.

[0116] Similarly, time-varying semantic information Guided cross-attention fusion prior and Perform layer normalization to obtain normalized features , and Then, CDTA cross-fusion is used to generate correction features, and combined with... The initial cross-fusion features are obtained by addition, and then passed to LEFN for further local enhancement. The implementation process is expressed as shown in equations (16)-(19). In CDTA, Q is determined by... K and V are obtained by mapping through 1×1 convolution and 3×3 depthwise convolution. It is obtained by mapping 1×1 convolution and 3×3 depth convolution, and the implementation process is expressed as in equation (17).

[0117] (16);

[0118] (17);

[0119] (18);

[0120] (19);

[0121] in, yes Normalized features are used to generate Q. and yes The normalized features are used to generate K and V. Is Preliminary cross-attention fusion features generated under guidance Is Enhanced local-global cross-fusion features generated under guidance.

[0122] Finally, and Guided cross-convolution attention fusion and Add to generate the first Preliminary local-global cross-fusion features of resolution , as in equation (20).

[0123] (20);

[0124] (2) Decision attention fusion;

[0125] To further reduce spurious changes that may arise from heterogeneous or homogeneous spectra and improve the fusion accuracy of dynamic information such as details and spectra, a multi-resolution cascaded decision attention fusion (DAF) module was designed, such as... Figure 3 As shown. A multi-resolution cascaded DAF will... Resolution cross-fusion results With the Resolution decision fusion features ( Through step-by-step fusion of spectral attention and spatial attention, spectral bands and spatial weights are adaptively evaluated and allocated to generate fused features with better time-varying information and prior details. For example... Figure 3 As shown, the three cascaded DAFs perform coarse-to-fine layer-by-layer decision attention fusion on the cross-fusion features of the four resolutions to generate the decision fusion features of the ideal resolution, as shown in Equation (21). Then, the ideal resolution image is generated by a reconstruction unit composed of three convolutions. And two intermediate resolution images and , as in equation (22).

[0126] (twenty one);

[0127] (twenty two);

[0128] in, It is a DAF module function. These are DAF module parameters. It is the first one generated by DAF. Resolution decision fusion features It is the final image generated by the spatiotemporal fusion model. and It is the generated reduced-resolution image.

[0129] refer to Figure 8 , Figure 8 The diagram shows the structure of Decision Attention Fusion (DAF). DAF consists of three parallel branches, with the upper branch used for adaptive evaluation of high spatial resolution fused features. The importance of spectral channels and spatial locations is determined, and important spectral and spatial information is adaptively enhanced; the lower branch is used to adaptively evaluate low spatial resolution decision fusion features. ( The importance of spectral channels and spatial locations is determined, and important spectral and spatial information is adaptively enhanced. An intermediate branch is used to fuse low-spatial-resolution and high-spatial-resolution fusion features. Finally, the results from all three are added together to obtain the decision fusion result at high spatial resolution. The implementation process is expressed as shown in equations (23)-(26).

[0130] (twenty three);

[0131] (twenty four);

[0132] (25);

[0133] (26);

[0134] in, It is a spectral attention module function. It is a spatial attention module function. Yes Features enhanced by spectral and spatial adaptive methods It is by Upsampling was obtained, Yes Features enhanced by spectral and spatial adaptive methods Yes and Features of fusion This is a decision fusion feature of DAF.

[0135] The upper and lower branches are respectively implemented by the spectral attention module (SpeAM) and the spatial attention module (SpaAM) to adaptively enhance spectral band information and spatial information of different importance levels, and adaptively adjust the weight allocation of the fusion results at each stage. The process is expressed as in equations (27)-(28). SpeAM uses global average pooling and global max pooling to compress the information of each spectral channel of the input, obtains the global information of the channel, and then adaptively learns the nonlinear dependency relationship between each spectral channel through two convolutions, generates weight coefficients, and performs addition and Sigmoid function to obtain the spectral channel attention map of [0,1]. The spectral channel attention map is multiplied with the corresponding channel of the input to achieve the calibration of the input. SpaAM performs global average pooling and global max pooling along the channel dimension to extract the most unique texture, edge and other features and overall background features of all channels, and splices them along the channels. Then, it is adaptively fused by convolution to learn context information and Sigmoid function mapping to generate the spatial attention map of [0,1]. The spatial attention map is multiplied with the features of each channel of the input to achieve the enhancement of the spatial region of the input.

[0136] (27);

[0137] (28);

[0138] in, It is input. and These are global average pooling and global max pooling operations, respectively. It is the Sigmoid function.

[0139] Convolutional Transformer discriminator with multi-resolution input:

[0140] To capture details and time-varying information of ground features at different sizes in a scene, a Multiresolution Inputs and Convolutional-Transformer Discriminator (MICTD) was designed to adversarially learn local and global spectral and detail information at different resolutions. This guides the generator output to possess high realism at different resolution levels, thereby improving the relative realism of the final generated image. Figure 9 As shown, the generated multi-resolution image , and Compared to ground truth (GT) images, images downgraded by 2x resolution 4x reduced resolution image The high resolution information is used as input to MICTD for adversarial learning of multi-resolution contextual information. The high resolution information focuses on adversarial learning of fine-grained textures, edges, and small features, while the low resolution information focuses on adversarial learning of global information and relationships.

[0141] like Figure 9 As shown, MICTD first uses a 3×3 kernel convolution and a convolutional Transformer to extract primary local-global features. Then, it uses three 4×4 kernel convolutions with a stride of 2 to successively reduce the feature size, and concatenates these features with images downgraded by a factor of 2 and 4. Deeper features are then extracted using a convolutional Transformer, and finally, the model outputs a realism judgment result using two 4×4 kernel convolutions with a stride of 1 and a sigmoid function. To accelerate model convergence and improve stability and performance, spectral normalization is introduced after each convolution operation. Furthermore, a relative average least squares discriminator is used to determine the relative realism of the input image to improve the stability and performance of the fusion model.

[0142] Composite loss function:

[0143] To optimize the stability and fusion capability of model training, a composite loss function was designed, including generator loss and discriminator loss. In addition to RaLS adversarial loss against the discriminator, the generator loss also includes pixel-level error loss, spectral consistency loss, structural consistency loss, and feature-level perceptual loss.

[0144] The adversarial loss between the generator and the discriminator is expressed as in expression (29).

[0145] (29);

[0146] The expressions for the RaLS functions of the generator and discriminator are as shown in (30) and (31).

[0147] (30);

[0148] (31);

[0149] Moreover, the discriminator output satisfies equations (32) and (33).

[0150] (32);

[0151] (33);

[0152] in, This represents the adversarial loss between the generator and the discriminator. Indicates discriminator loss. This represents the relative probability obtained by the discriminator. This represents the Sigmoid activation function. This represents the output of the untransformed discriminator. Represents the distribution of real images, Represents the distribution of generated images. and This indicates an operation that averages all reference and generated images in a batch of data.

[0153] To ensure that the generated image maintains pixel-level content information consistency with the real image, the content loss function... As shown in equation (34).

[0154] (34);

[0155] in, It's batch size. It is a real image. It generates an image.

[0156] To ensure that the generated image maintains spatial structural information consistency with the real image, multi-scale structural similarity is used to characterize spatial consistency, starting from brightness. Contrast and structure Three perspectives for evaluating spatial structure, spatial consistency loss function As shown in equation (35).

[0157] (35);

[0158] in, It is multi-scale structural similarity. , and It is the brightness at the Z-scale. Scale contrast and structure, , and It is a decision , and Parameters of component importance , and The expression is as shown in equation (36).

[0159] (36);

[0160] To ensure that the generated image maintains spectral information consistency with the real image, spectral angle loss is used to evaluate and control the degree of spectral distortion in the generated image; the spectral consistency loss function is employed. As shown in equation (37).

[0161] (37);

[0162] in, and It is the GT image pixel vector and its transpose. yes Image pixel vector and transpose, Is GT image and The angle between the pixel vectors of an image.

[0163] To ensure that the generated image maintains the same feature-level information as the real image, feature loss... On the pre-trained model VGG-19, the goal is to minimize the distance between two features before the activation function. The expression is as shown in equation (38).

[0164] (38);

[0165] in, and This represents the feature maps of real and generated images on the VGG model.

[0166] Therefore, the total loss function As in equation (39), where It is the weighting coefficient.

[0167] (39).

[0168] Experimental verification:

[0169] A. Experimental setup:

[0170] 1) Datasets: To evaluate the performance of the proposed method, we used two publicly available datasets: the widely used Coleambly Irrigation Area (CIA) dataset, which contains numerous phenological change scenarios, and the Lower Gwydir Catchment (LGC) dataset, which contains numerous abrupt changes in land cover types. The CIA remote sensing data size is 1024×1024×6, and the LGC remote sensing data size is 1024×1024×6.

[0171] 2) Implementation Details: The model is implemented using the PyTorch framework. The image patch size for training the network is 256×256×C, the batch size is 8, and the training epochs are 200. The number of input and output channels can be set according to the experimental data. The optimal hyperparameters in the loss function are set to... and The model objective was optimized using the Adam optimizer, with an initial learning rate set to 2 × 10⁻⁶. -4 Experiments based on deep learning models were implemented on GPUs, while experiments based on traditional models were implemented on CPUs.

[0172] 3) Evaluation metrics: To evaluate the fusion performance, the root mean square error (RMSE), spectral angle mapping (SAM), structural similarity (SSIM), ERGAS, and correlation coefficient (CC) were used to objectively evaluate the spatiotemporal fusion results.

[0173] B. Comparison with state-of-the-art (SOTA) methods on the CIA dataset:

[0174] In this section, we compare our proposed method with several state-of-the-art (SOTA) methods, including traditional models STARFM and FSDAF, and deep learning models GAN-STFM, MLFF-GAN, SwinSTFM, and ECPW-STFN.

[0175] As shown in Table 1, the method of this invention achieved the best RMSE, CC, SSIM, ERGAS, and SAM scores in all tested prediction scenarios, indicating that the fusion results of this method are superior in terms of radiometric, structural, and spectral fidelity. Specifically, in the prediction of the November 25, 2001 scenario using prior data from November 9, 2001, the method of this invention reduced RMSE to 0.0253, increased CC to 0.8692, increased SSIM to 0.8423, reduced ERGAS to 0.7835, and reduced SAM to 0.0705, significantly outperforming other comparative methods.

[0176] Table 1: Objective evaluation of the fusion results of the comparison method on CIA data

[0177]

[0178] For subjective visual evaluation, the CIA scene exhibits significant phenological variations. The prediction results for certain areas of the scene on November 25, 2001, are as follows: Figure 10As shown in the figure, it is evident that the fusion results of STARFM, FSDAF, GAN-STFM, MLFF-GAN, SwinSTFM, and ECPW-STFN exhibit significant spectral distortion and low accuracy in predicting changes in ground features within the scene, with their predictions heavily reliant on prior data. In contrast, the prediction results of the method presented in this invention are closest to the actual images, demonstrating better spectral and detail fidelity. In summary, the spatiotemporal fusion results of the method presented in this invention are superior in both visual and objective evaluations on the CIA dataset, which contains a large amount of phenological variation, demonstrating strong robustness.

[0179] C. Comparison with state-of-the-art (SOTA) methods on the LGC dataset:

[0180] In this section, we evaluate the performance of the proposed method on the LGC dataset, which contains a large number of abrupt land cover changes. Objective evaluation metrics are shown in Table 2. The proposed method achieves the best RMSE, CC, SSIM, ERGAS, and SAM metrics for all tested prediction dates. In the prediction for December 12, 2004, the proposed method reduced RMSE to 0.0294, improved CC to 0.7977, improved SSIM to 0.8111, reduced ERGAS to 1.1196, and reduced SAM to 0.1128. This verifies that the fusion results of the proposed method on abrupt land cover changes in different scenarios are relatively stable and have strong robustness.

[0181] Table 2: Quantitative evaluation indicators of the comparative methods on LGC data

[0182]

[0183] Regarding subjective visual evaluation, a massive flood occurred in December 2004, causing a drastic change in landform. The prediction results for a portion of the scene on December 12, 2004 are as follows: Figure 11 As shown in the figure, STARFM, GAN-STFM, MLFF-GAN, and ECPW-STFN clearly predict land cover type abrupt changes poorly, with ECPW-STFN exhibiting significant spectral distortion. FSDAF, SwinSTFM, and the method of this invention predict land cover type abrupt changes better, but considering both detail and spectral information, the spectral and detail information of the prediction results from the method of this invention is closer to the real images. In summary, the spatiotemporal fusion results of the method of this invention on the LGC dataset, which contains numerous abrupt changes in land cover, are superior in both visual and objective evaluation, demonstrating better spectral and detail fidelity.

[0184] In summary, this invention proposes a time-sensitive remote sensing image spatiotemporal fusion method. By constructing a time-sensitive bidirectional convolutional-Transformer generative adversarial network, it effectively solves the problems of severe prediction distortion in spatiotemporal fusion due to short-term abrupt changes, long-interval variations, and changes in land cover type, as well as the significant influence of prior data interference. Experimental results show that the proposed method outperforms state-of-the-art methods on two publicly available datasets, significantly improving the robustness and accuracy of spatiotemporal fusion.

[0185] This invention also provides a time-sensitive remote sensing image spatiotemporal fusion system, comprising: at least one processor; at least one memory for storing at least one program; and when the at least one program is executed by the at least one processor, the at least one processor implements the above-described method.

[0186] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0187] This invention also provides a computer program product, including a computer program or computer instructions, which are stored in a memory. A processor of a computer device reads the computer program or computer instructions from the memory and executes the computer program or computer instructions, causing the computer device to perform the above-described method.

[0188] It is understood that the content of the above method embodiments is applicable to the embodiments of this system, medium, and program product. The specific functions implemented by the embodiments of this system, medium, and program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0189] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0190] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

Claims

1. A spatiotemporal fusion method for time-varying sensitive remote sensing images, characterized in that, The method includes the following steps: S100: Obtain the prior time-time fine-resolution image, the prior time-time coarse-resolution image, the prediction time coarse-resolution image, and the trained time-resistant bidirectional convolutional-transformer generative adversarial network; wherein, the time-resistant bidirectional convolutional-transformer generative adversarial network includes a time-resistant bidirectional convolutional-transformer generator and a multi-resolution input convolutional-transformer discriminator, and the time-resistant bidirectional convolutional-transformer generator includes a down-resolution encoder, a super-resolution encoder, a dual-guided three-attention fusion decoder, and a reconstruction unit; S200, the prior time fine resolution image, the prior time coarse resolution image and the prediction time coarse resolution image are input into the trained time-resistant sensitive bidirectional convolutional-transformer generative adversarial network, and the prior time fine resolution image is extracted from the prior time coarse resolution image through the down-resolution encoder to obtain prior detail semantic features. S300, the super-resolution encoder performs multi-scale feature extraction from low resolution to high resolution on the prior time coarse resolution image and the predicted time coarse resolution image to obtain time-varying semantic features; S400, the prior detail semantic features and the time-varying semantic features are input into the dual-guided three-attention fusion decoder, and the preliminary local-global cross-fusion features are generated by dual-guided cross-convolution Transformer fusion. Then, the preliminary local-global cross-fusion features are weighed and integrated step by step by decision attention fusion to obtain the decision fusion features. S500, the decision fusion features are reconstructed by the reconstruction unit to generate a high-resolution image at the prediction time; the high-resolution image at the prediction time and the corresponding real image are input into the multi-resolution input convolution-Transformer discriminator for adversarial learning to optimize the generation process of the high-resolution image at the prediction time.

2. The method according to claim 1, characterized in that, In S200, the step of extracting multi-scale features from high resolution to low resolution in the prior time-time fine-resolution image using the down-resolution encoder to obtain prior detail semantic features includes: S210, low-level semantic features are extracted from the prior time-time fine-resolution image through multi-scale dilated convolution to obtain initial prior time-time fine features; wherein, the multi-scale dilated convolution uses dilated convolutions with dilation rates of 1, 2 and 3 in parallel to extract features from the prior time-time fine-resolution image, and the extracted multi-scale features are concatenated along the channel direction and then initially fused through 1×1 convolution to obtain the initial prior time-time fine features. S220, the initial prior time-time fine features are input into a resolution-reducing feature extraction branch consisting of four convolutional Transformer modules. Each convolutional Transformer module consists of layer normalization, multi-head deep convolutional transpose attention, local enhancement forward pass network and residual connection. After local-global feature extraction of the input features, each convolutional Transformer module reduces the resolution through downsampling operation to obtain the prior detailed semantic features.

3. The method according to claim 2, characterized in that, In S220, the multi-head depthwise convolutional transpose attention processes the input features, including: S221: After layer normalization of the input features, pixel-level cross-channel context is aggregated through 1×1 convolution, and then spatial context is separated by 3×3 depth convolution to generate query matrix, key matrix and value matrix. S222, After performing transformation operations on the query matrix and the key matrix respectively, perform a dot product operation, divide the dot product result by a learnable parameter to scale, and perform a Softmax operation on the scaled result to obtain a first attention map; S223, after transforming the value matrix, multiply it with the first attention map, and after transforming the multiplication result, perform convolution processing to obtain the attention-corrected features; S224, the attention-corrected features are added to the input features to obtain the local-global features of the multi-head depth convolution transposed attention output.

4. The method according to claim 1, characterized in that, In S300, the step of extracting multi-scale features from low resolution to high resolution in the prior time coarse-resolution image and the predicted time coarse-resolution image through the super-resolution encoder to obtain time-varying semantic features includes: S310, low-level semantic features are extracted from the prior time coarse resolution image and the prediction time coarse resolution image respectively through multi-scale dilated convolution to obtain the initial prior time coarse features and the initial prediction time coarse features. S320, the initial prior time coarse features and the initial prediction time coarse features are respectively input into the super-resolution feature extraction branch composed of 4 convolutional Transformer modules. Each convolutional Transformer module performs local-global feature extraction on the input features and then improves the resolution through upsampling operation to obtain the prior time multi-resolution semantic features and the prediction time multi-resolution semantic features. S330, perform an element-wise subtraction operation on the prior time multi-resolution semantic features and the predicted time multi-resolution semantic features to obtain the time-varying semantic features.

5. The method according to claim 1, characterized in that, In S400, the generation of preliminary local-global cross-fusion features through dual-guided cross-convolution Transformer fusion includes: S411, take the prior detail semantic features as prior semantic features and the time-varying semantic features as time-varying information features, and perform layer normalization on the prior semantic features and the time-varying information features at the i-th resolution respectively to obtain prior normalized features and time-varying normalized features. S412, in the prior semantic feature-guided cross-attention fusion, the time-varying normalized features are mapped to generate a query matrix through 1×1 convolution and 3×3 depth convolution, and the prior normalized features are mapped to generate a key matrix and a value matrix through 1×1 convolution and 3×3 depth convolution; in the time-varying information feature-guided cross-attention fusion, the prior normalized features are mapped to generate a query matrix through 1×1 convolution and 3×3 depth convolution, and the time-varying normalized features are mapped to generate a key matrix and a value matrix through 1×1 convolution and 3×3 depth convolution. S413, after performing transformation operations on the generated query matrix and key matrix respectively, perform dot product operation, divide the dot product result by the learnable parameter to scale, and perform Softmax operation on the scaled result to obtain the second attention map; S414, after transforming the generated value matrix, multiply it with the second attention map, transform the multiplication result, process it through convolution, and then add it to the prior semantic features and time-varying information features respectively to generate two directional preliminary cross-attention fusion features; S415, after normalizing the preliminary cross-attention fusion features of the two directions, input them into the local enhancement forward pass network for local enhancement, and add them with the preliminary cross-attention fusion features of the two directions to obtain the enhanced local-global cross-fusion features of the two directions. S416, the two guided enhanced local-global cross-fusion features are added together to generate the preliminary local-global cross-fusion feature.

6. The method according to claim 1, characterized in that, In S400, the step of progressively weighing and integrating the preliminary local-global cross-fusion features through decision attention fusion to obtain decision fusion features includes: S421, the preliminary local-global cross-fusion feature at resolution i and the decision fusion feature at resolution i+1 are subjected to spectral attention enhancement and spatial attention enhancement, respectively, to obtain two enhanced features; wherein, the spectral attention enhancement compresses each spectral channel of the input feature through global average pooling and global max pooling, and then learns the nonlinear dependency between spectral channels through two convolutional layers to generate a spectral channel attention map, and multiplies the spectral channel attention map with the corresponding channel of the input feature to achieve spectral calibration; the spatial attention enhancement extracts spatial features along the channel dimension through global average pooling and global max pooling, concatenates the extracted spatial features along the channel and fuses them through a convolutional layer to generate a spatial attention map, and multiplies the spatial attention map with each channel of the input feature to achieve spatial region enhancement; S422, the preliminary local-global cross-fusion feature at resolution i is concatenated with the decision fusion feature at resolution i+1, and then fused through a convolutional layer to obtain the intermediate fusion feature; S423, add the two enhanced features to the intermediate fusion feature to obtain the decision fusion feature at the i-th resolution.

7. The method according to claim 1, characterized in that, In S500, the step of inputting the predicted time-time fine-resolution image and the corresponding real image into the multi-resolution input convolution-Transformer discriminator for adversarial learning to optimize the generation process of the predicted time-time fine-resolution image includes: S510, the high-resolution image at the prediction time and its corresponding downsampled images by 2x and 4x are respectively input into the multi-resolution input convolution-Transformer discriminator along with the corresponding real image and its downsampled version; S520 extracts primary local features through convolutional layers and the LeakyReLU activation function, and extracts local-global features through convolution-Transformer. S530 reduces the feature size sequentially through at least three convolutional layers with a kernel size of 4×4 and a stride of 2, and then stitches and fuses them with image features of different resolutions. The S540 uses convolutional layers, spectral normalization operations, and the Sigmoid function to output the discrimination results, thereby determining the relative authenticity of the input image.

8. The method according to claim 1, characterized in that, The trained time-sensitive bidirectional convolutional-Transformer generative adversarial network is trained using a composite loss function, which includes a generator loss and a discriminator loss. The generator loss includes relative average least squares adversarial loss, pixel-level mean square error loss, multi-scale structural similarity loss, spectral angle loss, and feature loss. The discriminator loss is a relative average least squares loss. The composite loss function uses weighted summation of the various losses through weight coefficients to form deep supervision to optimize the model training process.

9. A time-sensitive remote sensing image spatiotemporal fusion system, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 8.