A method, system, terminal, and storage medium for multimodal remote sensing data representation and fusion.
By extracting temporal and spatial invariant features from multimodal remote sensing data using an autoencoder and then using geographic location information and bottleneck attention for feature fusion, the problem of low accuracy in multimodal remote sensing data fusion is solved, and high-precision feature-level fusion is achieved.
Patent Information
- Application Number
- CN202511164412.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-20
AI Technical Summary
Existing multimodal remote sensing data characterization and fusion methods suffer from low fusion accuracy, especially at the pixel and feature levels where it is difficult to achieve effective alignment and information complementarity of multi-source data. Furthermore, they lack adaptability and generalization capabilities across regions, time phases, and sensors.
An autoencoder is used to extract time-invariant features from multimodal image data. Positive and negative sample pairs are generated using geographic location information. The autoencoder is adjusted by a contrastive loss function constrained by cosine similarity. Feature fusion is performed by combining spatiotemporal invariant information integration. Bottleneck attention is introduced to evaluate modal importance.
It improves the fusion accuracy of multimodal remote sensing image features, enhances the alignment accuracy of cross-temporal and cross-modal features, and improves the robustness and adaptability of feature fusion.
Smart Images

Figure CN120726442B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing data processing technology, and in particular to a method, system, terminal, and storage medium for multimodal remote sensing data representation and fusion. Background Technology
[0002] Existing research on multimodal remote sensing data characterization and fusion can be mainly divided into pixel-level characterization and fusion methods, feature-level characterization and fusion methods, and decision-level characterization and fusion methods.
[0003] Pixel-level representation and fusion methods mainly include traditional numerical transformation methods and multi-scale decomposition methods. The essence of these methods is to fuse data using specific mathematical laws. Although they have made some progress, these methods have high requirements for image alignment, the fusion process is easily affected by the imaging environment, and there will be a certain degree of data redundancy, making it difficult to transfer to other tasks.
[0004] The method based on feature-level representation and fusion mainly consists of a feature extraction module, which fuses the extracted features at the feature level. However, the training cost of the neural network model required for feature extraction is too high, and the trained neural network model can only perform well under specific task objectives, and cannot adapt to scenarios of multi-task objectives and multi-scale feature extraction and fusion.
[0005] Decision-level representation and fusion methods are essentially fusions of decision results. The model takes multimodal data as input, but each modality's data performs a downstream task independently. The results from each individual modality are then aggregated to form the final downstream task result. Because this method makes limited use of correlation information across different modalities and its overall performance is not ideal, it is rarely used in multimodal remote sensing image fusion.
[0006] Existing multimodal remote sensing data representation and fusion processes have limitations. Data acquired from different sensors exhibit significant heterogeneity in image features such as spatial resolution, making effective alignment and information complementarity of multi-source data at the pixel and feature levels a technical challenge. Furthermore, inconsistencies in semantic information across different data sources and scales, as well as image differences arising from different imaging times at the same geographical location, all pose challenges to multimodal remote sensing data representation and fusion. Simultaneously, research on the adaptability and generalization capabilities for multidimensional data differences across regions, time phases, and sensors remains relatively scarce.
[0007] Existing technologies still suffer from low accuracy in the representation and fusion of multimodal remote sensing data, therefore, existing technologies need further improvement. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a method, system, terminal and storage medium for multimodal remote sensing data representation and fusion, in order to solve the problem of low fusion accuracy in existing multimodal remote sensing data representation and fusion methods.
[0009] The technical solution adopted by this invention to solve the technical problem is as follows:
[0010] In a first aspect, the present invention provides a method for multimodal remote sensing data characterization and fusion, comprising:
[0011] A multimodal image dataset is acquired, and the multimodal image dataset is preprocessed to obtain an initial image dataset;
[0012] The initial image dataset is input into the autoencoder corresponding to the modality to obtain the corresponding time-invariant features, and the autoencoder corresponding to the modality is pre-trained by feature reconstruction.
[0013] Using the geographic location information of the initial image data, the time-invariant features are classified to generate positive and negative sample pairs, and the trained autoencoder is adjusted by a contrastive loss function constrained by cosine similarity.
[0014] Based on the adjusted autoencoder, features with spatiotemporal invariant information are extracted from images of the corresponding modality, and feature fusion is performed by a spatiotemporal invariant information integration method that takes into account the importance of the modality.
[0015] Output the fused image features.
[0016] In one implementation, the step of acquiring a multimodal image dataset and preprocessing the multimodal image dataset to obtain an initial image dataset includes:
[0017] The multimodal image dataset is obtained by acquiring remote sensing images of synthetic aperture radar imaging, remote sensing images at a first resolution, and remote sensing images at a second resolution; wherein the second resolution is higher than the first resolution.
[0018] Using the remote sensing image at the second resolution as a reference, the resolution of all remote sensing images is adjusted to be consistent through bicubic convolution interpolation, and all remote sensing images are cropped into image blocks of a preset size.
[0019] The processed remote sensing images are used as the initial image dataset.
[0020] In one implementation, the step of inputting the initial image dataset into an autoencoder corresponding to each modality to obtain corresponding time-invariant features, and pre-training the autoencoder corresponding to the modality using feature reconstruction, includes:
[0021] The images in the initial image dataset are normalized, and random Gaussian noise is added to the normalized images to obtain the features with added noise.
[0022] The noise-added features are input into the corresponding modal autoencoder to generate features with time-invariant information in the corresponding modality.
[0023] Calculate the reconstruction loss of the features with time-invariant information, and update the corresponding autoencoder based on the reconstruction loss to obtain the trained autoencoder.
[0024] In one implementation, the step of calculating the reconstruction loss of the feature with time-invariant information and updating the corresponding autoencoder based on the reconstruction loss to obtain the trained autoencoder includes:
[0025] The mean squared error loss function is used as the feature reconstruction loss function of the autoencoder corresponding to each mode to calculate the reconstruction loss of the features with time-invariant information.
[0026] Based on the reconstruction loss, the corresponding autoencoder is updated using backpropagation of the error until the reconstruction error of the corresponding autoencoder is less than the first expected value, thus obtaining the trained autoencoder.
[0027] In one implementation, the step of using the geographic location information of the initial image data to classify the time-invariant features obtained, generating positive and negative sample pairs, and adjusting the trained autoencoder using a contrastive loss function constrained by cosine similarity, includes:
[0028] Using the geographic location information of the initial image data, the obtained features with time-invariant properties are classified, and the positive sample pairs are automatically generated under different observation conditions at the same geographic location, while the negative sample pairs are used with randomly selected images from other locations.
[0029] The features with time-invariant information, the positive sample pairs, and the negative sample pairs are used as vectors, and the cosine similarity is calculated for each vector to obtain the corresponding cosine similarity loss.
[0030] The parameters of the corresponding trained autoencoder are adjusted according to the cosine similarity loss to obtain the corresponding adjusted autoencoder; wherein, the adjusted autoencoder is used to extract features with spatiotemporally invariant information from images of a single modality.
[0031] In one implementation, adjusting the parameter output of the corresponding trained autoencoder based on the cosine similarity loss to obtain the corresponding adjusted autoencoder includes:
[0032] Based on the cosine similarity loss, the parameter output of the corresponding trained autoencoder is adjusted by backpropagation of error until the reconstruction error of the corresponding trained autoencoder is less than the second expected value, thus obtaining the adjusted autoencoder.
[0033] In one implementation, the step of extracting features with spatiotemporally invariant information from images of the corresponding modality based on the adjusted autoencoder, and performing feature fusion through a spatiotemporally invariant information integration method that takes into account modal importance, includes:
[0034] Images of all modalities at the same geographical location are input into the corresponding adjusted autoencoder to obtain the corresponding features with spatiotemporally invariant information.
[0035] By concatenating all features with spatiotemporally invariant information, a comprehensive feature is obtained.
[0036] Calculate channel attention scores in the channel space of the comprehensive feature and spatial attention scores in the pixel space of the comprehensive feature;
[0037] The integrated features are enhanced based on the channel attention score and the spatial attention score. The enhanced features are then integrated using a multilayer perceptron and the number of channels is reduced to obtain the fused image features.
[0038] Secondly, the present invention provides a multimodal remote sensing data characterization and fusion system, comprising:
[0039] The image data acquisition module is used to acquire a multimodal image dataset and preprocess the multimodal image dataset to obtain an initial image dataset.
[0040] An adaptive time-invariant information extraction module is used to input the initial image dataset into the autoencoder corresponding to the modality, obtain the corresponding features with time-invariant properties, and pre-train the autoencoder of the corresponding modality by feature reconstruction.
[0041] The contrast enhancement spatial invariant information extraction module is used to classify the time-invariant features obtained by utilizing the geographic location information of the initial image data, generate positive and negative sample pairs, and learn spatiotemporally and cross-source consistent feature representations through a contrast loss function constrained by cosine similarity.
[0042] The spatiotemporal invariant information integration module that takes modal importance into account is used to extract features with spatiotemporal invariant information from images of the corresponding modality based on the adjusted autoencoder, and to perform feature fusion through the spatiotemporal invariant information integration method that takes modal importance into account.
[0043] The fusion feature output module is used to output the fused image features.
[0044] Thirdly, the present invention provides a terminal, comprising: a processor and a memory, wherein the memory stores a multimodal remote sensing data characterization and fusion program, and the multimodal remote sensing data characterization and fusion program, when executed by the processor, is used to implement the operation of the multimodal remote sensing data characterization and fusion method as described in the first aspect.
[0045] Fourthly, the present invention also provides a computer-readable storage medium storing a multimodal remote sensing data characterization and fusion program, which, when executed by a processor, is used to implement the operation of the multimodal remote sensing data characterization and fusion method as described in the first aspect.
[0046] The present invention, by employing the above technical solution, has the following effects:
[0047] This invention employs an autoencoder to extract time-invariant features from images of different time phases within a single modality, reducing interference from temporal variations and enhancing cross-temporal feature consistency. Furthermore, it utilizes geographic index information to construct a contrastive learning task, enhancing spatially invariant feature representation and improving cross-modal feature alignment accuracy. It also introduces bottleneck attention to dynamically evaluate the importance of each modality, adjust fusion weights, strengthen the contribution of key modalities, and improve feature fusion robustness. This invention fully considers the spatiotemporal invariant features and importance of each modality, achieving adaptive multimodal remote sensing image feature-level fusion and improving the fusion accuracy of multimodal remote sensing image features. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0049] Figure 1 This is a flowchart of the multimodal remote sensing data representation and fusion method in this invention.
[0050] Figure 2 This is a flowchart of the multimodal remote sensing data representation and fusion algorithm in this invention.
[0051] Figure 3 This is a flowchart of model training and feature fusion in one implementation of the present invention.
[0052] Figure 4 This is a functional schematic diagram of the terminal in one implementation of the present invention.
[0053] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0055] Exemplary methods
[0056] Existing multimodal remote sensing data representation and fusion processes have limitations. Data acquired from different sensors exhibit significant heterogeneity in image features such as spatial resolution, making effective alignment and information complementarity of multi-source data at the pixel and feature levels a technical challenge. Furthermore, inconsistencies in semantic information across different data sources and scales, as well as image differences arising from different imaging times at the same geographical location, all pose challenges to multimodal remote sensing data representation and fusion. Simultaneously, research on the adaptability and generalization capabilities for multidimensional data differences across regions, time phases, and sensors remains relatively scarce.
[0057] To address the above-mentioned technical problems, this invention provides a method for multimodal remote sensing data representation and fusion, comprising: acquiring and preprocessing a multimodal image dataset to obtain an initial image dataset; inputting the initial image dataset into corresponding autoencoders to obtain corresponding time-invariant features, and pre-training the corresponding autoencoders by feature reconstruction; classifying the obtained features using the geographic location information of the initial image data to generate positive and negative sample pairs, and adjusting the trained autoencoders using a contrastive loss function constrained by cosine similarity; based on the adjusted autoencoders, extracting features with spatiotemporal invariant information from the images of the corresponding modal, and performing feature fusion by integrating spatiotemporal invariant information that takes into account modal importance; and outputting the fused image features. This invention improves the fusion accuracy of multimodal remote sensing image features.
[0058] like Figure 1 As shown, this embodiment of the invention provides a method for multimodal remote sensing data characterization and fusion, including the following steps:
[0059] Step S100: Obtain a multimodal image dataset and preprocess the multimodal image dataset to obtain an initial image dataset.
[0060] In this embodiment, a multimodal remote sensing data representation and fusion framework is provided. This framework is based on time-invariant features, spatially invariant features, and attention compression. The multimodal remote sensing data representation and fusion method is implemented based on this framework.
[0061] like Figure 2 As shown, this framework employs a multi-kernel convolutional-self-attention coupled autoencoder for end-to-end reconstruction pre-training of multi-temporal images. Then, while removing photosensitive noise, it integrates multi-scale latent features through self-attention, thereby extracting feature representations robust to temporal changes. Subsequently, positive and negative sample pairs are constructed based on geographic indexes, and a contrastive learning mechanism is used to enhance feature aggregation in spatially similar regions and feature separation in heterogeneous regions under unlabeled conditions, obtaining feature representations stable to spatial differences such as imaging angle, terrain undulation, and imaging scale. Finally, the feature representations with spatiotemporally invariant information are input into the bottleneck attention module, which, combined with a multilayer perceptron, completes cross-modal feature compression and global consistency fusion, outputting a low-dimensional, information-rich unified feature vector.
[0062] Based on the above framework, this embodiment needs to acquire a multimodal image dataset before training the autoencoder; wherein, the multimodal image dataset includes remote sensing images based on synthetic aperture radar imaging and remote sensing images acquired by optical remote sensing devices of different resolutions.
[0063] Specifically, in one implementation of this embodiment, step S100 includes the following steps:
[0064] Step S101: Acquire the remote sensing image of synthetic aperture radar imaging, the remote sensing image of first resolution, and the remote sensing image of second resolution to obtain the multimodal image dataset; wherein, the second resolution is higher than the first resolution;
[0065] Step S102: Using the remote sensing image at the second resolution as a reference, adjust the resolution of all remote sensing images to be consistent through bicubic convolution interpolation, and crop all remote sensing images into image blocks of a preset size.
[0066] Step S103: Use the processed remote sensing image as the initial image dataset.
[0067] As an example, the method for obtaining the multimodal image dataset in this embodiment is as follows:
[0068] The following remote sensing images are acquired within the area of District A of a certain city: Sentinel-1 remote sensing imagery (i.e., the remote sensing imagery acquired by synthetic aperture radar, hereinafter referred to as Sentinel-1 remote sensing imagery), Sentinel-2 remote sensing imagery (i.e., the remote sensing imagery with the first resolution, hereinafter referred to as Sentinel-2 remote sensing imagery), and Gaofen-1 remote sensing imagery (i.e., the remote sensing imagery with the second resolution, hereinafter referred to as Gaofen-1 remote sensing imagery). As an example, in this embodiment, the Sentinel-2 remote sensing imagery can be a remote sensing imagery with a 20-meter resolution band, and the Gaofen-1 remote sensing imagery can be a remote sensing imagery with a 2-meter resolution band; this is not a limitation.
[0069] In this embodiment, after acquiring the aforementioned multimodal image dataset, the preprocessing process for the multimodal image dataset is as follows:
[0070] Using the high-resolution Gaofen-1 remote sensing image as a benchmark, the resolution of Sentinel-1 and Sentinel-2 remote sensing images was adjusted to a consistent high resolution using bicubic convolution interpolation. Then, all remote sensing images in the adjusted multimodal image dataset were cropped into... Image patches of varying sizes. Then, the processed Sentinel-1 remote sensing imagery... Sentinel-2 remote sensing images And Gaofen-1 remote sensing images As the initial image dataset ( , , ).
[0071] like Figure 1 As shown, this embodiment of the invention provides a method for multimodal remote sensing data characterization and fusion, including the following steps:
[0072] Step S200: Input the initial image dataset into the autoencoder corresponding to the modality to obtain the corresponding time-invariant features, and pre-train the autoencoder corresponding to the modality by feature reconstruction.
[0073] In this embodiment, the initial image dataset ( , , The complete process of implementing the multimodal remote sensing data representation and fusion method is illustrated in the following diagram. Figure 2 As shown, it includes three modules:
[0074] 1) Adaptive Temporal-Invariant Feature Extraction Module (ATIFE): The ATIFE module consists of an autoencoder with a self-attention mechanism. This autoencoder is a noise reduction autoencoder, and it extracts time-invariant features from source domain remote sensing images at different time phases by feature reconstruction. The self-attention mechanism in this autoencoder is used to perform preliminary enhancement and integration of the extracted features.
[0075] 2) Contrastive-Enhanced Spatial-Invariant Feature Extraction Module (CESIFE): The CESIFE module is used to promote feature interaction between remote sensing image data of different modalities, enabling the model to extract spatial and modal invariant information from remote sensing image data.
[0076] 3) Modality-Aware Spatiotemporal-Invariant Feature Integration Module (MA-STIFI): The MA-STIFI module is used for adaptive modality importance consideration and to integrate and enhance the spatiotemporal invariant information contained in remote sensing image data of different modalities.
[0077] Based on the three modules mentioned above, this embodiment requires the initial image dataset ( , , The inputs are fed into the corresponding ATIFE modules for each modality, and the ATIFE modules for each modality are pre-trained using feature reconstruction to obtain single-modal time-invariant features. , , .
[0078] Specifically, in one implementation of this embodiment, step S200 includes the following steps:
[0079] Step S201: Normalize the images in the initial image dataset, and add random Gaussian noise to the normalized images to obtain the features with added noise.
[0080] Step S202: Input the noise-added features into the corresponding modal autoencoder to generate features with time-invariant information in the corresponding modality.
[0081] Step S203: Calculate the reconstruction loss of the feature with time-invariant information, and update the corresponding autoencoder according to the reconstruction loss to obtain the trained autoencoder.
[0082] In one implementation of this embodiment, step S203 includes the following steps:
[0083] Step S203a: Use the mean square error loss function as the feature reconstruction loss function of the autoencoder corresponding to each mode, and calculate the reconstruction loss of the features with time-invariant information.
[0084] Step S203b: Based on the reconstruction loss, update the corresponding autoencoder using error backpropagation until the reconstruction error of the corresponding autoencoder is less than the first expected value, and obtain the trained autoencoder.
[0085] In this embodiment, the specific training process of the ATIFE module is as follows:
[0086] For the initial image dataset ( , , After normalizing the image band by band to the range [0,1], random Gaussian noise is added to the normalized image with a mean of 0 and a standard deviation of 0.03, resulting in... , , Then, the normalized and Gaussian noise-added data... , , The inputs are fed into the corresponding autoencoders to generate features with time-invariant information in the corresponding modality. , , These time-invariant features provide effective suppression of seasonality, illumination, and imaging parameter differences caused by different observation dates. Moreover, these time-invariant features can enhance cross-modal discriminability while maintaining fine-grained spectral-spatial texture information, providing a robust and unified feature foundation for subsequent multimodal fusion and downstream classification / retrieval tasks.
[0087] In this embodiment, with Taking the input as an example, when training the ATIFE module corresponding to the modality, the Mean Square Error (MSE) loss function is used as the feature reconstruction loss calculation for the autoencoder. The complete process is as follows:
[0088] ;
[0089] ;
[0090] , , ;
[0091] ;
[0092] ;
[0093] ;
[0094] in, express It follows a mean of 0 and a variance of . It follows a normal distribution.
[0095] This means that the image with added noise will be processed through a feedforward neural network to obtain initial features. . and It is a weight matrix used to map the input dimension to the hidden dimension and then back to the output dimension. and This is a bias term.
[0096] , , These represent the query vector, key vector, and value vector, respectively, in the calculation of attention weights. , , These are the calculations of the weights of the corresponding vectors.
[0097] Represents the key vector The feature dimensions. This represents the final output feature. The attention weights are normalized using softmax (normalized exponential function), and the value vectors at each position are summed according to their weights to form the attention-enabled feature.
[0098] This represents the output feature of the decoder, which is obtained by using... As input, use transposed convolution Perform feature scale restoration, using the input original image as the label.
[0099] This represents the mean squared error loss. This represents the total number of pixels in the input image. Indicates in Decoder output at pixel level Indicates in The pixel values of the original image at the pixel level.
[0100] In this embodiment, the loss is calculated pixel by pixel using the MSE loss function, and then the parameters are iterated to achieve the training of the complete autoencoder. The specific iterative process is as follows:
[0101] An autoencoder with a self-attention mechanism is trained using the initial image dataset to obtain features with time-invariant information. The mean squared error loss function is used as the feature reconstruction loss function for the autoencoder corresponding to each modality. The feature reconstruction loss is calculated, and the autoencoder is updated by backpropagation of the error until the reconstruction error of the autoencoder is less than the expectation. The trained autoencoder is then saved.
[0102] In this embodiment, a multi-kernel convolution-self-attention coupled autoencoder is used to perform end-to-end reconstruction pre-training on multi-temporal images. This enables the trained autoencoder to extract time-invariant features from images of different temporal phases in a single modality, reducing temporal variation interference and enhancing cross-temporal feature consistency.
[0103] like Figure 1 As shown, this embodiment of the invention provides a method for multimodal remote sensing data characterization and fusion, including the following steps:
[0104] Step S300: Using the geographic location information of the initial image data, the obtained time-invariant features are classified to generate positive and negative sample pairs, and the trained autoencoder is adjusted by a contrastive loss function constrained by cosine similarity.
[0105] In this embodiment, after obtaining the trained autoencoder through the above training method, the CESIFE module is still needed to adjust the trained autoencoder.
[0106] Specifically, in one implementation of this embodiment, step S300 includes the following steps:
[0107] Step S301: Using the geographic location information of the initial image data, the obtained features with time-invariant properties are classified, and the positive sample pairs are automatically generated under different observation conditions at the same geographic location, and the negative sample pairs are used by randomly selected images from other locations.
[0108] Step S302: The features with time-invariant information, the positive sample pairs, and the negative sample pairs are used as vectors, and the cosine similarity is calculated for each vector to obtain the corresponding cosine similarity loss.
[0109] Step S303: Adjust the parameter output of the corresponding trained autoencoder according to the cosine similarity loss to obtain the corresponding adjusted autoencoder; wherein, the adjusted autoencoder is used to extract features with spatiotemporally invariant information from images of a single modality.
[0110] In one implementation of this embodiment, step S303 includes the following steps:
[0111] Step S303a: Based on the cosine similarity loss, adjust the parameter output of the corresponding trained autoencoder using error backpropagation until the reconstruction error of the corresponding trained autoencoder is less than the second expected value, and obtain the adjusted autoencoder.
[0112] In this embodiment, the process of adjusting the trained autoencoder mainly utilizes the initial image dataset ( , , The geographic location information of the image will be used to obtain features with time-invariant information. , , The system categorizes data and automatically generates positive sample pairs under different observation conditions within the same geographical location. It also uses randomly selected images from other locations as negative sample pairs. Finally, it learns spatiotemporal and cross-source consistent feature representations through a contrastive loss function constrained by cosine similarity, enabling the autoencoder to extract spatially and modally invariant information from remote sensing image data.
[0113] Specifically, for each incoming image Its corresponding geographical location is Subsequently, in all remaining images of the initial image dataset, [the image was found to be related to...]. Images from the same geographical location but different time periods or different data sources constitute a set of positive sample pairs. Furthermore, randomly select 10 pairs of... Images from different geographical locations, with no restrictions on data sources or time periods, are used to form a negative sample pair set. Subsequently, the features with time-invariant information obtained in the ATIFE module will be analyzed. Features of the positive sample pair set Features of the negative sample set As vectors, cosine similarity is calculated for each vector to obtain cosine similarity loss; then, the parameters of the autoencoder in the ATIFE module are adjusted to achieve contrast-enhanced spatial invariant information extraction. The complete process of the CESIFE module is as follows:
[0114] ;
[0115] ;
[0116] ;
[0117] + ;
[0118] in, , They represent the incoming images respectively. , The corresponding output characteristics of the ATIFE module, This represents the transpose of the corresponding eigenvector. , This represents the vector length of the corresponding eigenvector. Indicates the similarity between two features. Indicates image similarity loss. As a similarity safety interval, a gradient is generated when the similarity exceeds the safety interval. In this embodiment... Set it to 0.5.
[0119] In this embodiment, based on the aforementioned cosine similarity loss, the specific method for adjusting the trained autoencoder is as follows:
[0120] Positive and negative sample pairs are constructed using the geographic location information of the initial image. The output similarity of the autoencoder between positive sample pairs is maximized, and the output similarity of the autoencoder between negative sample pairs is minimized. The similarity loss is calculated, and the autoencoder is updated through error backpropagation until the reconstruction error of the autoencoder is less than expected. The adjusted autoencoder is then saved.
[0121] In this embodiment, geographic index information is used to construct a contrastive learning task, which enhances the spatial invariant feature representation and improves the cross-modal feature alignment accuracy.
[0122] like Figure 1 As shown, this embodiment of the invention provides a method for multimodal remote sensing data characterization and fusion, including the following steps:
[0123] Step S400: Based on the adjusted autoencoder, extract features with spatiotemporal invariant information from the image of the corresponding modality, and perform feature fusion by integrating spatiotemporal invariant information that takes into account the importance of the modality.
[0124] Step S500: Output the fused image features.
[0125] In this embodiment, after adjustment by the CESIFE module, the corresponding modality's ATIFE module can extract features with spatiotemporally invariant information from a single modality image. , , These features are input into the MA-STIFI module to achieve feature fusion that takes into account modal importance.
[0126] Specifically, in one implementation of this embodiment, step S400 includes the following steps:
[0127] Step S401: Input the images of all modalities at the same geographical location into the corresponding adjusted autoencoder to obtain the corresponding features with spatiotemporally invariant information.
[0128] Step S402: Concatenate all features with spatiotemporally invariant information to obtain comprehensive features;
[0129] Step S403: Calculate the channel attention score in the channel space of the integrated feature and calculate the spatial attention score in the pixel space of the integrated feature;
[0130] Step S404: Enhance the integrated features based on the channel attention score and the spatial attention score, integrate the enhanced features using a multilayer perceptron and reduce the number of channels to obtain the fused image features.
[0131] In this embodiment, during the feature fusion process that takes into account modal importance, images of all modalities at the same geographical location are considered. , , These are then input into the corresponding adjusted ATIFE modules to obtain features with spatiotemporally invariant information. , , These features are first directly spliced together to obtain the comprehensive features. and in channel space Calculate channel attention score and in pixel space Calculate spatial attention score Complete the channel attention score. Spatial attention score After the calculation, it will utilize and right Enhancement is then achieved by utilizing a multilayer perceptron. The enhanced features are integrated and the number of channels is reduced to achieve feature integration and compression.
[0132] During the training of the MA-STIFI module, a multilayer perceptron is also used. Then, a decoder is added, which also calculates the SmoothL1 loss based on the feature reconstruction target. The complete implementation process is as follows:
[0133] ;
[0134] ;
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] in, Indicates channel attention score, Represents spatial attention score, This represents the fusion features of multimodal remote sensing image data obtained through a multilayer perceptron after bottleneck attention enhancement. This indicates the differences in element-wise feature reconstruction. This describes the process for calculating the SmoothL1 loss. The threshold value is 1 in this embodiment.
[0140] In this embodiment, bottleneck attention is introduced to dynamically evaluate the importance of each modality, adjust the fusion weights, strengthen the contribution of key modalities, and improve the robustness of feature fusion.
[0141] like Figure 3 As shown, in the practical application scenario of this embodiment, the model training and feature fusion process mainly includes the following steps:
[0142] S11, Determine the initial image dataset;
[0143] S12, use the initial image dataset to train an autoencoder with a self-attention mechanism to obtain features with time-invariant information;
[0144] S13, calculate the feature reconstruction loss, update the autoencoder through error backpropagation until the reconstruction error of the autoencoder is less than the expectation, and save the trained autoencoder.
[0145] S14: Use the geographic location information of the initial image to construct positive sample pairs and negative sample pairs, and maximize the output similarity of the autoencoder between positive sample pairs and minimize the output similarity of the autoencoder between negative sample pairs;
[0146] S15, calculate the similarity loss, update the autoencoder through error backpropagation until the reconstruction error of the autoencoder is less than the expectation, and save the adjusted autoencoder.
[0147] S16, input the image data of each modality into the adjusted autoencoder to obtain spatiotemporally invariant features, and use bottleneck attention and multilayer perceptron to enhance and integrate multimodal features;
[0148] S17, calculate the feature reconstruction loss, update the autoencoder through error backpropagation until the reconstruction error of the autoencoder is less than the expectation, and save the trained feature integration model.
[0149] S18 uses a tuned autoencoder and feature integration model to fuse the input multimodal remote sensing images.
[0150] In this embodiment, by integrating multi-scale latent features through self-attention, feature representations that are robust to temporal changes can be extracted. Furthermore, by utilizing a contrastive learning mechanism, feature aggregation of spatially similar regions and feature separation of heterogeneous regions can be enhanced under unlabeled conditions. This can be combined with a multilayer perceptron to complete cross-modal feature compression and global consistency fusion, outputting a low-dimensional, information-rich unified feature vector. This achieves efficient noise reduction, spatiotemporal consistent feature learning, and compact representation of multi-temporal and multi-sensor remote sensing data.
[0151] This embodiment achieves the following technical effects through the above technical solution:
[0152] This embodiment employs an autoencoder to extract time-invariant features from images of different time phases within a single modality, reducing interference from temporal variations and enhancing cross-temporal feature consistency. Furthermore, it utilizes geographic index information to construct a contrastive learning task, enhancing spatially invariant feature representation and improving cross-modal feature alignment accuracy. It also introduces bottleneck attention to dynamically evaluate the importance of each modality, adjust fusion weights, strengthen the contribution of key modalities, and improve feature fusion robustness. This embodiment fully considers the spatiotemporal invariant features and importance of each modality, achieving adaptive multimodal remote sensing image feature-level fusion and improving the fusion accuracy of multimodal remote sensing image features.
[0153] Exemplary device
[0154] Based on the above embodiments, the present invention also provides a multimodal remote sensing data characterization and fusion system, comprising:
[0155] The image data acquisition module is used to acquire a multimodal image dataset and preprocess the multimodal image dataset to obtain an initial image dataset.
[0156] An adaptive time-invariant information extraction module is used to input the initial image dataset into the autoencoder corresponding to the modality, obtain the corresponding features with time-invariant properties, and pre-train the autoencoder of the corresponding modality by feature reconstruction.
[0157] The contrast enhancement spatial invariant information extraction module is used to classify the time-invariant features obtained by utilizing the geographic location information of the initial image data, generate positive and negative sample pairs, and learn spatiotemporally and cross-source consistent feature representations through a contrast loss function constrained by cosine similarity.
[0158] The spatiotemporal invariant information integration module that takes modal importance into account is used to extract features with spatiotemporal invariant information from images of the corresponding modality based on the adjusted autoencoder, and to perform feature fusion through the spatiotemporal invariant information integration method that takes modal importance into account.
[0159] The fusion feature output module is used to output the fused image features.
[0160] This embodiment achieves the following technical effects through the above technical solution:
[0161] This embodiment employs an autoencoder to extract time-invariant features from images of different time phases within a single modality, reducing interference from temporal variations and enhancing cross-temporal feature consistency. Furthermore, it utilizes geographic index information to construct a contrastive learning task, enhancing spatially invariant feature representation and improving cross-modal feature alignment accuracy. It also introduces bottleneck attention to dynamically evaluate the importance of each modality, adjust fusion weights, strengthen the contribution of key modalities, and improve feature fusion robustness. This embodiment fully considers the spatiotemporal invariant features and importance of each modality, achieving adaptive multimodal remote sensing image feature-level fusion and improving the fusion accuracy of multimodal remote sensing image features.
[0162] Based on the above embodiments, the present invention also provides a terminal, the principle block diagram of which can be as follows: Figure 4 As shown.
[0163] The terminal includes: a processor, a memory, an interface, a display screen, and a communication module connected via a system bus; wherein, the processor of the terminal provides computing and control capabilities; the memory of the terminal includes a computer-readable storage medium and internal memory; the computer-readable storage medium stores an operating system and computer programs; the internal memory provides an environment for the operation of the operating system and computer programs in the computer-readable storage medium; the interface is used to connect to external devices; the display screen is used to display relevant information; and the communication module is used to communicate with a cloud server or other devices.
[0164] When executed by a processor, this computer program is used to implement the operation of multimodal remote sensing data representation and fusion methods.
[0165] It will be understood by those skilled in the art that Figure 4The schematic diagram shown is merely a partial structural diagram related to the present invention and does not constitute a limitation on the terminal to which the present invention is applied. A specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0166] In one embodiment, a terminal is provided, comprising: a processor and a memory, the memory storing a multimodal remote sensing data representation and fusion program, which, when executed by the processor, is used to implement the operation of the multimodal remote sensing data representation and fusion method described above.
[0167] In one embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a multimodal remote sensing data characterization and fusion program, which, when executed by a processor, is used to implement the operation of the multimodal remote sensing data characterization and fusion method described above.
[0168] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, database, or other media used in the embodiments provided by this invention can include both non-volatile and volatile memory.
[0169] In summary, this invention provides a method, system, terminal, and storage medium for multimodal remote sensing data representation and fusion, comprising: acquiring and preprocessing a multimodal image dataset to obtain an initial image dataset; inputting the initial image dataset into corresponding autoencoders to obtain corresponding time-invariant features, and pre-training the corresponding modal autoencoders by feature reconstruction; classifying the obtained features using the geographic location information of the initial image data to generate positive and negative sample pairs, and adjusting the trained autoencoders using a contrastive loss function constrained by cosine similarity; based on the adjusted autoencoders, extracting features with spatiotemporal invariant information from the corresponding modal images, and performing feature fusion by integrating spatiotemporal invariant information that takes into account modal importance; and outputting the fused image features. This invention improves the fusion accuracy of multimodal remote sensing image features.
[0170] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for multimodal remote sensing data representation and fusion, characterized in that, include: A multimodal image dataset is acquired, and the multimodal image dataset is preprocessed to obtain an initial image dataset; The initial image dataset is input into the autoencoder corresponding to the modality to obtain the corresponding time-invariant features, and the autoencoder corresponding to the modality is pre-trained by feature reconstruction. Using the geographic location information of the initial image data, the time-invariant features are classified to generate positive and negative sample pairs, and the trained autoencoder is adjusted by a contrastive loss function constrained by cosine similarity. Based on the adjusted autoencoder, features with spatiotemporal invariant information are extracted from images of the corresponding modality, and feature fusion is performed by a spatiotemporal invariant information integration method that takes into account the importance of the modality. Output the fused image features; The step of inputting the initial image dataset into the autoencoder corresponding to each modality to obtain the corresponding time-invariant features, and pre-training the autoencoder corresponding to the modality by feature reconstruction, includes: The images in the initial image dataset are normalized, and random Gaussian noise is added to the normalized images to obtain the features with added noise. The noise-added features are input into the corresponding modal autoencoder to generate features with time-invariant information in the corresponding modality. Calculate the reconstruction loss of the features with time-invariant information, and update the corresponding autoencoder according to the reconstruction loss to obtain the trained autoencoder; The adjusted autoencoder extracts features with spatiotemporally invariant information from images of the corresponding modality, and performs feature fusion through a spatiotemporally invariant information integration method that takes into account modal importance, including: Images of all modalities at the same geographical location are input into the corresponding adjusted autoencoder to obtain the corresponding features with spatiotemporally invariant information. By concatenating all features with spatiotemporally invariant information, a comprehensive feature is obtained. Calculate channel attention scores in the channel space of the comprehensive feature and spatial attention scores in the pixel space of the comprehensive feature; The integrated features are enhanced based on the channel attention score and the spatial attention score. The enhanced features are then integrated using a multilayer perceptron and the number of channels is reduced to obtain the fused image features.
2. The multimodal remote sensing data representation and fusion method according to claim 1, characterized in that, The process of acquiring a multimodal image dataset and preprocessing the multimodal image dataset to obtain an initial image dataset includes: The multimodal image dataset is obtained by acquiring remote sensing images of synthetic aperture radar imaging, remote sensing images at a first resolution, and remote sensing images at a second resolution; wherein the second resolution is higher than the first resolution. Using the remote sensing image at the second resolution as a reference, the resolution of all remote sensing images is adjusted to be consistent through bicubic convolution interpolation, and all remote sensing images are cropped into image blocks of a preset size. The processed remote sensing images are used as the initial image dataset.
3. The multimodal remote sensing data representation and fusion method according to claim 1, characterized in that, The calculation of the reconstruction loss of the features with time-invariant information, and the updating of the corresponding autoencoder based on the reconstruction loss to obtain the trained autoencoder, includes: The mean squared error loss function is used as the feature reconstruction loss function of the autoencoder corresponding to each mode to calculate the reconstruction loss of the features with time-invariant information. Based on the reconstruction loss, the corresponding autoencoder is updated using backpropagation of the error until the reconstruction error of the corresponding autoencoder is less than the first expected value, thus obtaining the trained autoencoder.
4. The multimodal remote sensing data representation and fusion method according to claim 1, characterized in that, The step of using the geographic location information of the initial image data to classify the time-invariant features obtained, generating positive and negative sample pairs, and adjusting the trained autoencoder using a contrastive loss function constrained by cosine similarity includes: Using the geographic location information of the initial image data, the time-invariant features are classified, and positive sample pairs are automatically generated under different observation conditions at the same geographic location. Randomly selected images from other locations are used as negative sample pairs. The features with time-invariant information, the positive sample pairs, and the negative sample pairs are used as vectors, and the cosine similarity is calculated for each vector to obtain the corresponding cosine similarity loss. The parameters of the corresponding trained autoencoder are adjusted according to the cosine similarity loss to obtain the corresponding adjusted autoencoder; wherein, the adjusted autoencoder is used to extract features with spatiotemporally invariant information from images of a single modality.
5. The multimodal remote sensing data representation and fusion method according to claim 4, characterized in that, The step of adjusting the parameter output of the corresponding trained autoencoder based on the cosine similarity loss to obtain the corresponding adjusted autoencoder includes: Based on the cosine similarity loss, the parameter output of the corresponding trained autoencoder is adjusted using backpropagation of error until the reconstruction error of the corresponding trained autoencoder is less than the second expected value, thus obtaining the adjusted autoencoder.
6. A multimodal remote sensing data representation and fusion system, used to implement the multimodal remote sensing data representation and fusion method as described in any one of claims 1-5, characterized in that, include: The image data acquisition module is used to acquire a multimodal image dataset and preprocess the multimodal image dataset to obtain an initial image dataset. An adaptive time-invariant information extraction module is used to input the initial image dataset into the autoencoder corresponding to the modality, obtain the corresponding features with time-invariant properties, and pre-train the autoencoder of the corresponding modality by feature reconstruction. The contrast enhancement spatial invariant information extraction module is used to classify the time-invariant features obtained by utilizing the geographic location information of the initial image data, generate positive and negative sample pairs, and learn spatiotemporally and cross-source consistent feature representations through a contrast loss function constrained by cosine similarity. The spatiotemporal invariant information integration module that takes modal importance into account is used to extract features with spatiotemporal invariant information from images of the corresponding modality based on the adjusted autoencoder, and to perform feature fusion through the spatiotemporal invariant information integration method that takes modal importance into account. The fusion feature output module is used to output the fused image features.
7. A terminal, characterized in that, include: The processor and memory, wherein the memory stores a multimodal remote sensing data characterization and fusion program, which, when executed by the processor, is used to implement the operation of the multimodal remote sensing data characterization and fusion method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multimodal remote sensing data characterization and fusion program, which, when executed by a processor, is used to implement the operation of the multimodal remote sensing data characterization and fusion method as described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-modal remote sensing image ground object classification method based on modal information constraint
CN119399522A
Method, Device And Non-Transitory Computer-Readable Storage Medium For Processing A Sequence Of Top View Image Frames
US20220405945A1
Cited By
A shared-private decoupled optical-sar cross-modal unified representation method and system
CN122637114A