A missing modal remote sensing image fusion method based on a shared prototype space
By constructing a cross-modal shared prototype space and an adaptive selection fusion mechanism, the problem of unstable modal fusion in existing technologies is solved, and high-quality image reconstruction and fusion in complex scenes are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-17
AI Technical Summary
Existing missing modality fusion methods are unstable in pixel domain generation, have high uncertainty in single inference results, poor fusion reliability in complex scenes, and are prone to propagation of error information, affecting image quality and downstream task performance.
A cross-modal shared prototype space is constructed. Through multi-candidate representation reasoning and adaptive selection fusion mechanism, the interpretability reconstruction of missing modalities is completed using existing modal images. Regional adaptive fusion is performed in combination with candidate credibility evaluation.
It improves the interpretability of the missing modality reasoning process, reduces the risk of false textures and false edges, enhances the stability and accuracy of fusion results in complex scenarios, and suppresses the problem of error accumulation.
Smart Images

Figure CN122415346A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing and multimodal image fusion technology, specifically to a missing modality remote sensing image fusion method based on a shared prototype space, and more particularly to a method for completing missing modality representation inference and achieving high-quality fusion reconstruction using existing modality images under conditions of partial modality loss. Background Technology
[0002] The purpose of multimodal image fusion is to integrate complementary information acquired from different sensors, imaging mechanisms, or observation conditions to generate a result image with clearer structure, richer information, and more complete semantic expression. Most existing multimodal fusion methods assume that the images from each modality involved in the fusion can be acquired synchronously and completely. For example, in infrared and visible light fusion tasks, it is usually assumed that infrared and visible light images exist simultaneously; in medical multimodal image processing, it is usually assumed that structural and functional modal images can be stably acquired; and in remote sensing multi-source image processing, it is also usually assumed that multi-source observation data are completely available at the same or similar times.
[0003] However, modal missingness is a common problem in practical applications. On the one hand, the acquisition conditions for some modalities are more demanding, and they are easily affected by factors such as sensor failure, occlusion, environmental interference, communication limitations, or cost constraints, resulting in partial modal missingness or severe degradation. On the other hand, most existing methods for dealing with the problem of missing modalities adopt a technical approach of "first completing the missing modalities, and then fusing them," especially by directly generating pseudo-missing modal images through generative adversarial networks, diffusion models, or pixel regression networks, and then fusing them with existing modal images. While these methods can mitigate the impact of missing modalities to some extent, they still have the following shortcomings: First, directly generating missing modal images in the pixel domain can easily introduce pseudo-textures, pseudo-edges, or pseudo-responses, resulting in poor interpretability of the generated results. Second, the same existing modal input often corresponds to multiple possible missing modal response patterns in complex scenes, while existing methods usually only output a single prediction result, making it difficult to characterize this ambiguity. Third, if unreliable pseudo-missing modal images are directly used for fusion, the error information will be further propagated into the fusion result, affecting image quality and the performance of downstream tasks such as detection, segmentation, and recognition.
[0004] Therefore, a novel missing modality fusion method needs to be proposed, which no longer relies on pixel-level missing modality generation, but can complete missing modality reasoning within an interpretable shared representation space. Furthermore, the reliability of different reasoning results should be evaluated, and adaptive fusion should be implemented based on the reliability, thereby reducing the uncertainty caused by single reasoning and improving the stability, interpretability, and practicality of the fusion results in complex scenarios. Summary of the Invention
[0005] The purpose of this invention is to address the problems of unstable pixel-level generation, high uncertainty of single inference results, and poor fusion reliability in complex scenes in existing missing modality fusion methods. This invention proposes a missing modality remote sensing image fusion method based on a shared prototype space. This method constructs a cross-modality shared prototype space, performs multi-candidate representation inference for missing modalities within this space, and combines candidate credibility evaluation and an adaptive selection fusion mechanism to achieve stable reconstruction of image fusion results in missing modality scenarios.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] A missing modality remote sensing image fusion method based on a shared prototype space includes the following steps:
[0008] 11) Obtain a training sample set containing existing modal images and missing modal images: Construct training samples for missing modal scenarios based on complete bimodal image data to simulate the condition that some modalities are not available in real applications. During training, existing modal images are used as model inputs, and missing modal images and fused reference images are used together as supervision information.
[0009] 12) Construct a multi-candidate reasoning and adaptive fusion model based on a shared prototype space: This model consists of a multi-scale feature encoding module, a shared prototype space construction module, a missing modality multi-candidate reasoning module, a candidate credibility evaluation module, an adaptive selection fusion module, and a result reconstruction module connected in sequence.
[0010] 13) Supervised training of the model based on training samples: Input the existing modal images in the training samples into the model, and use the missing modal images and the fused reference images as supervision targets to complete the model parameter learning and optimization;
[0011] 14) Obtain image data to be processed containing only existing modal information: In the actual inference stage, only existing modal images that can be stably obtained are input, and the missing modal images are no longer relied upon;
[0012] 15) Input the image data to be processed into the trained model, and obtain the final fused image through multi-candidate missing modality inference, credibility evaluation and adaptive weighted fusion.
[0013] Obtaining a training sample set containing existing modality images and missing modality images includes the following steps:
[0014] 21) Collect or organize images containing existing modalities and missing modality images The paired sample data, wherein the existing mode For stable and acquireable modes, including visible light and SAR remote sensing images; the missing modes For modes that are prone to loss, including infrared and thermal imaging remote sensing images;
[0015] 22) Preprocess the paired sample data, including image cropping, size normalization, grayscale normalization, color normalization, noise reduction, and contrast enhancement, to ensure that the input image maintains consistency in spatial scale and pixel distribution.
[0016] 23) Construct a training scenario for missing modalities, and use the images of the missing modalities. As a training supervision reference, only existing modality images are retained. As input to the model, it is used to reproduce input conditions that are not available in real-world scenarios;
[0017] 24) Divide the preprocessed sample data into training set, validation set and test set according to the proportion, and use them for model training, hyperparameter tuning and performance evaluation, respectively;
[0018] 25) Repeat steps 21)-24) to expand and improve the training sample set and enhance the model's generalization ability in complex scenarios.
[0019] The construction of the multi-candidate reasoning and adaptive fusion model based on a shared prototype space includes the following steps:
[0020] 31) Construct a multi-scale feature encoding module for existing modal images. Multi-scale feature extraction is performed to obtain structural features, texture features, and global scene context features, outputting a compact and representative representation of existing modal features. ;
[0021] 32) Construct a shared prototype space construction module to establish a shared prototype space (SPS), which contains a basic shared prototype library. With semantic prototype library The basic shared prototype library is used to represent common structural information such as edges, contours, and texture skeletons shared among different modalities; the semantic prototype library is used to represent high-level semantic information such as target category, scene category, and thermal response mode.
[0022] 33) Representing existing modal features and missing modal feature representation The corresponding prototype response coefficients are obtained by projecting them onto the shared prototype space SPS. and Its expression is:
[0023]
[0024] in, This represents the prototype projection mapping function, used to map modal features to response distributions in the prototype space;
[0025] 34) Construct a missing modality multi-candidate inference module (MCGM) based on existing modality feature representations. and its prototype response coefficient K candidate missing modality representations are generated through multi-branch parallel inference. Its expression is:
[0026]
[0027] in, This represents the kth independent candidate inference branch, used to characterize the pluralism of missing modalities in complex scenarios;
[0028] 35) Construct a candidate credibility evaluation module (CEM). Based on the matching degree between the candidate missing modality representation and the shared prototype space, the consistency between candidates, the reconstruction error, and the scene context information, generate a pixel-level credibility map for each candidate. ;
[0029] 36) Construct an adaptive selection fusion module ASFM and a result reconstruction module RM. Based on the confidence map, dynamically select or adaptively weight and fuse multiple candidate missing modal representations, and reconstruct them in collaboration with existing modal features to output the final fusion result image.
[0030] The model training includes the following steps:
[0031] 41) The existing modality images in the training samples Input the multi-scale feature encoding module to obtain the feature representation of the existing modality. The prototype response coefficients are then projected onto the shared prototype space SPS to obtain the prototype response coefficients. ;
[0032] 42) Remove the missing modality images from the training samples. Similarly, mapping to the shared prototype space SPS yields the prototype response coefficients of the missing modes. Establish cross-modal sharing correspondence between existing modalities and missing modalities in the prototype space;
[0033] 43) Representing existing modal features and prototype response coefficient Input the missing modality multi-candidate inference module MCGM to generate multiple candidate missing modality representations covering different potential response patterns. ;
[0034] 44) Input multiple candidate missing mode representations into the candidate confidence evaluation module (CEM), calculate the confidence level of each candidate, and generate a confidence map corresponding to the image pixel position. ;
[0035] 45) Input multiple candidate missing mode representations and confidence maps into the adaptive selection fusion module ASFM, perform hard selection or soft weighted fusion on the candidates, and obtain the final latent representation of the missing mode. Its expression is:
[0036]
[0037] in, Let represent the k-th candidate fusion weight determined adaptively by credibility, and satisfy . ;
[0038] 46) The final missing modality latent representation Compared with existing modal feature representations Perform feature-level collaborative fusion to obtain a fused feature representation. The system inputs the results into the RM reconstruction module and outputs the final fused image. ;
[0039] 47) Construct the total loss function L for multi-constraint joint supervision:
[0040]
[0041] in, To rebuild the losses, To preserve structural losses, To share prototype consistency loss, For candidate diversity loss, For credibility constraint loss, , , , , The loss weighting coefficient is adjustable.
[0042] 48) Perform backpropagation based on the total loss function L, update the model parameters layer by layer, and determine whether the preset number of training rounds or convergence conditions have been reached. If so, complete the model training; otherwise, return to continue iterative training.
[0043] The candidate credibility assessment module CEM generates a credibility map. The basis includes:
[0044] 51) Matching residuals between candidate missing modal representations and the shared prototype space SPS Consistency measure among multiple candidate missing modal representations Reconstruction closed-loop error of candidate missing mode representation and existing modal scene context information ;
[0045] 52) The credibility diagram The ASFM module is used to guide region adaptive fusion by improving the representation of existing modal features in structurally clear regions. The fusion weights are used to improve the latent representation of the final missing modalities in the target salient region. The fusion weights reduce the impact of unreliable candidates in high-uncertainty regions, suppress artifacts and erroneous information propagation, and improve the stability and detail preservation of the fusion results. Beneficial effects
[0046] Compared with existing technologies, this invention proposes a missing modality remote sensing image fusion method based on a shared prototype space. When the missing modality is unavailable, it utilizes existing modality images to complete the missing modality representation inference and fusion result reconstruction. This method does not directly generate the missing modality deterministically in the pixel domain, but instead establishes a correspondence between existing and missing modalities at the cross-modal shared representation level by constructing a shared prototype space. This improves the interpretability of the missing modality inference process and reduces the risks introduced by pseudo-textures, pseudo-edges, and pseudo-responses. Furthermore, this invention generates multiple potential missing modality representations through a multi-candidate inference mechanism, which can better describe the diversity of missing modality responses in complex scenes. Combined with a candidate credibility evaluation mechanism, it performs region-adaptive judgment on the reliability of different candidate results, avoiding the error accumulation problem caused by the instability of a single inference result.
[0047] Furthermore, this invention employs an adaptive selection fusion mechanism to dynamically select or weightedly fuse multiple candidate missing modal representations based on candidate credibility, and collaboratively reconstructs the final fused image with existing modal features. This enhances the expressive power of the target region and complementary information region while preserving existing modal structural information and texture details, thereby improving the stability, accuracy, and practicality of the fusion result in complex scenarios. Attached Figure Description
[0048] Figure 1 This is a sequence diagram of the method of the present invention; Figure 2 This is a diagram illustrating the overall structure of the remote sensing image fusion model involved in this invention. Figure 3 This is a network structure diagram of the remote sensing image fusion model involved in this invention; Figure 4 This is a structural diagram of the multi-candidate reasoning and credibility evaluation involved in this invention; Figure 5 This is a schematic diagram of the fusion result of the present invention in a scenario with missing modalities. Detailed Implementation
[0049] To provide a better understanding of the structural features and effects achieved by the present invention, a detailed description is provided below, accompanied by preferred embodiments and accompanying drawings:
[0050] like Figure 1 As shown, the missing modality remote sensing image fusion method based on shared prototype space of the present invention includes the following steps:
[0051] The first step is to obtain a training sample set containing existing modality images and missing modality images: Training samples for missing modality scenarios are constructed based on complete bimodal image data to simulate the condition that some modalities are unavailable in real-world applications. During training, existing modality images are used as model input, and missing modality images and fused reference images are used together as supervision information. The specific steps are as follows:
[0052] (1) Collect or organize paired sample data containing existing modal images and missing modal images. The existing modal images are denoted as... The missing modality image is denoted as Existing modal images For modal images that can be stably acquired in practical applications, visible light remote sensing images or SAR remote sensing images are preferred; missing modal images For modal images that may not be stably acquired in practical applications due to factors such as environmental interference, changes in imaging conditions, sensor failure, obstruction, or communication limitations, infrared images or thermal imaging remote sensing images are preferred.
[0053] (2) Preprocess the paired sample data, including one or more of the following: image registration, image cropping, size normalization, grayscale normalization or color normalization, denoising, and contrast enhancement, to ensure the consistency of spatial location and pixel distribution of different modal images. Crop the existing modal images and the missing modal images to a size of Image patches are used to facilitate subsequent network training and feature extraction. Let the preprocessing function be... The preprocessed existing modality image and missing modality image are represented as follows:
[0054] in, This represents the preprocessed existing modal image. This represents the preprocessed image of the missing modalities.
[0055] (3) Construct a training scenario for missing modalities. During the training phase, only existing modal images are retained. As input to the model, the missing modality image As a training supervision reference, it simulates input conditions where missing modalities are unavailable in real-world applications. Simultaneously, a fusion reference image is constructed. The fusion reference image serves as the supervisory target for the final fusion result. It can be composed of standard fusion results under complete modal input conditions, or it can be generated by manual annotation or other high-quality fusion methods;
[0056] (4) Divide the preprocessed sample data into training set, validation set and test set. The training set is used for model parameter learning, the validation set is used for parameter adjustment and convergence monitoring during model training, and the test set is used for model performance evaluation;
[0057] (5) Repeat steps (1) to (4) to expand and improve the training sample set in order to improve the robustness and generalization ability of the model in complex scenes, different land cover categories and different imaging conditions.
[0058] The second step, as Figure 2 As shown, a missing modality remote sensing image fusion model based on a shared prototype space is constructed: a multi-scale feature encoding module is constructed to extract deep features from existing modality images and missing modality images; a shared prototype space module is constructed to establish cross-modality shared representations; a missing modality multi-candidate inference module is constructed to generate multiple candidate missing modality representations; a candidate credibility evaluation module is constructed to estimate the credibility of multiple candidate results; and an adaptive selection fusion module and a result reconstruction module are constructed to output the final fused result image. The specific steps are as follows:
[0059] (1) Construct a multi-scale feature encoding module to extract structural features, texture features and scene context features from existing modal images to obtain existing modal feature representations; at the same time, during the training phase, use feature extraction branches with the same structure to extract missing modal features;
[0060] (1-1) The preprocessed existing modal image The input is a multi-scale feature encoding module, which consists of 5 convolutional layers, 5 batch normalization layers, and 5 ReLU activation functions. The kernel size of each convolutional layer is [missing information]. The padding size is 1 for all layers, the stride of the first convolutional layer is 1, the stride of the second to fifth convolutional layers is 2, and the number of output channels is set to 32, 64, 128, 256 and 512 respectively.
[0061] (1-2) For an input size of The existing modal images are first processed by convolution, batch normalization, and ReLU activation function to obtain a size of The shallow feature map is used to extract edge information and local texture information in existing modal images;
[0062] (1-3) Then, the shallow feature map is sequentially input into the convolutional units from layer 2 to layer 5 for progressive downsampling and feature extraction, and the size of the output feature map is successively changed. , , and The second and third layers are mainly used to extract mid-scale structural texture features, while the fourth and fifth layers are mainly used to extract high-level semantic features and scene context features.
[0063] (1-4) After the above 5 layers of convolutional feature extraction, the existing modality feature representation is obtained. Its expression is:
[0064]
[0065] in, Represents the existing modal feature encoding function;
[0066] (1-5) During the training phase, the preprocessed missing modality images are... Input a missing modality encoding branch with the same structure as the existing modality feature extraction branch to obtain the missing modality feature representation. Its expression is:
[0067]
[0068] in, This represents the missing modality feature encoding function, with a final output size of... ;
[0069] (2) For example Figure 3 As shown, a shared prototype space building module is constructed to establish cross-modal shared representation relationships between existing modalities and missing modalities, enabling different modal features to be aligned and interact in a unified representation space;
[0070] (2-1) Construct a basic shared prototype library containing 32 basic shared prototypes to represent the common edge, contour, texture skeleton and structural layout information among different modalities;
[0071] (2-2) Construct a semantic prototype library containing 16 semantic prototypes to represent target category, scene category, hot response mode and other high-level semantic information;
[0072] (2-3) The basic shared prototype library and the semantic prototype library together constitute the shared prototype space SPS. The total number of prototypes in the shared prototype space is 48, and the dimension of each prototype vector is set to 512.
[0073] (2-4) Represent the existing modal features Input prototype projection mapping layer; the prototype projection mapping layer consists of a It consists of a convolutional layer, a global average pooling layer, and a fully connected layer. First, it utilizes... The convolutional layer performs channel integration on the input features, and then uses global average pooling to... The feature map is compressed into a global feature vector, and then mapped to a 48-dimensional prototype response vector through a fully connected layer to obtain the prototype response coefficients of the existing modes. Its expression is:
[0074]
[0075] Among them, it means Prototype projection mapping function;
[0076] (2-5) Representing missing modal features Input the same prototype projection mapping layer as in steps (2-4) to obtain the prototype response coefficients of the missing modes. Its expression is:
[0077]
[0078] (2-6) Using existing modal prototype response coefficients and missing modal prototype response coefficients Establish the correspondence between existing modalities and missing modalities in the shared prototype space, and provide constraints and guidance for subsequent multi-candidate reasoning of missing modalities;
[0079] (3) such as Figure 4 As shown, a missing modality multi-candidate inference module is constructed to represent existing modality features. and its prototype response coefficient Multiple candidate missing modality representations are generated to characterize the various potential response patterns that may exist for missing modalities in complex scenarios;
[0080] (3-1) Represent the existing modal features and prototype response coefficient To perform joint input, first, the 48-dimensional prototype response vector... Expanded into a prototype guiding vector through a fully connected mapping and copied to... At a spatial scale, a prototype guiding feature map is formed; then the prototype guiding feature map is combined with the existing modal feature representation. Perform splicing along the channel dimension;
[0081] (3-2) Construct a shared inference backbone network consisting of 2 convolutional layers, 2 batch normalization layers, and 2 ReLU activation functions, where the kernel size is [missing information]. All have a step size of 1, a fill size of 1, and 512 output channels.
[0082] (3-3) Construct four parallel candidate inference branches, with the number of branches K set to 4. Each candidate inference branch consists of two convolutional layers, two batch normalization layers, and two ReLU activation functions, where the kernel size is 1. All have a step size of 1, a fill size of 1, and 512 output channels.
[0083] (3-4) Input the common latent features output by the shared inference backbone network into the four candidate inference branches respectively to obtain four candidate missing modality representations. , , and Its expression is:
[0084]
[0085] in, This represents the k-th candidate inference branch. The output size of each candidate missing modality representation is... ;
[0086] (3-5) The four candidate missing modal representations correspond to different potential missing modal response patterns, which are used to reflect the multiple solutions of missing modal reasoning in complex remote sensing scenarios, thereby avoiding the problem of uncertainty accumulation caused by a single reasoning result in complex scenarios;
[0087] (4) Construct a candidate credibility evaluation module to perform region-level, pixel-level or channel-level credibility estimation on multiple candidate missing modal representations and generate corresponding credibility maps to guide the subsequent adaptive fusion process;
[0088] (4-1) Representation of the k-th candidate missing mode Compared with existing modal feature representations Channel dimensions are spliced together, and the shared prototype space matching information and candidate consistency information are combined to form the input features of the credibility evaluation module;
[0089] (4-2) Construct a credibility evaluation network consisting of 3 convolutional layers, 2 batch normalization layers, and 2 ReLU activation functions. The kernel size of the first two convolutional layers is 1. The stride is 1 for both layers, the padding size is 1 for both layers, and the number of output channels is 128 and 64 respectively; the third convolutional layer uses... Convolution with a stride of 1 and 1 output channel;
[0090] (4-3) After inputting the k-th candidate into the credibility evaluation network, the credibility graph corresponding to the k-th candidate is output. Its output size is ;
[0091] (4-4) Perform steps (4-1) to (4-3) on the four candidate missing modal representations respectively to obtain the four candidate confidence maps. , , and ;
[0092] (4-5) Perform softmax normalization on the four candidate confidence maps to obtain the fusion weight map corresponding to each candidate. , , and Its expression is:
[0093]
[0094] in, Let the fusion weight graph corresponding to the k-th candidate satisfy:
[0095]
[0096] (5) Construct an adaptive selection fusion module and a result reconstruction module, which are used to dynamically select or weightedly fuse multiple candidate missing modal representations based on the fusion weight map. The result reconstruction module is used to restore the fusion features to the original image resolution and output the final fusion result image.
[0097] (5-1) Based on the fusion weight graph The four candidate missing mode representations are weighted and fused pixel-wise to obtain the final latent representation of the missing mode. Its expression is:
[0098]
[0099] in, This represents the final latent representation of the missing modes, with an output size of [value missing]. ;
[0100] (5-2) The final missing modality latent representation Compared with existing modal feature representations Concatenate along the channel dimension, then input a convolutional kernel of size [size missing]. A fusion convolutional layer with a stride of 1 and a padding size of 1 is used to obtain the fused feature representation. Its output size is ;
[0101] (5-3) The result reconstruction module consists of 4 upsampling layers, 4 convolutional layers, 3 batch normalization layers, and 3 ReLU activation functions. Upsampling uses bilinear interpolation, and the upsampling scale factor is 1. The kernel size of each convolutional layer is [missing information]. The step size is 1, the fill size is 1, and the number of output channels is set to 256, 128, 64 and 32 respectively.
[0102] (5-4) Representing the fusion features The input result reconstruction module consists of 4 upsampling layers, 4 convolutional layers, 3 batch normalization layers, and 3 ReLU activation functions. Upsampling employs bilinear interpolation, and the upsampling scale factor is [missing value]. The kernel size of each convolutional layer is 1. The step size is 1 for all inputs, the padding size is 1 for all inputs, and the number of output channels is set to 256, 128, 64, and 32 respectively. The fusion feature representation, after the first upsampling and convolution, has an output size of After the second upsampling and convolution, the output size is After the third upsampling and convolution, the output size is... After the fourth upsampling and convolution, the output size is... This gradually restores the image's spatial resolution.
[0103] (5-5) Input the fourth upsampled output into the output convolutional layer, wherein the output convolutional layer preferably uses... Convolution with a stride of 1 is used to map the number of channels to the number of channels in the target fused image, resulting in the final fused image. Its expression is:
[0104]
[0105] in, This represents the reconstruction function. If the target fused image is a single-channel image, the final output size is... .
[0106] The third step is to train the model based on the training samples: The established network model is trained and its parameters are adjusted using the pre-defined training set and its corresponding supervision information until the preset number of iterations (epochs) is reached. Finally, the corresponding parameters and the trained network are retained. The specific steps are as follows:
[0107] (1) Input the existing modality images and missing modality images in the preprocessed training samples into the feature extraction module to obtain the existing modality feature representations respectively. and missing modal feature representation ;
[0108] (2) Represent the existing modal features and missing modal feature representation Input the shared prototype space construction module to obtain the existing modal prototype response coefficients. and missing modal prototype response coefficients Establish cross-modal shared representation relationships between existing modalities and missing modalities;
[0109] (3) Represent the existing modal features and prototype response coefficient The missing modality multi-candidate inference module is input to generate four candidate missing modality representations. , , and ;
[0110] (4) Input the four candidate missing mode representations into the candidate credibility evaluation module to generate the credibility graph corresponding to each candidate. , , and The corresponding fusion weight graph is then obtained through further normalization. , , and ;
[0111] (5) Input the four candidate missing mode representations and their corresponding fusion weight graphs into the adaptive selection and fusion module to dynamically select or weight and fuse multiple candidate missing mode representations to obtain the final missing mode latent representation. ;
[0112] (6) The final missing modality latent representation Compared with existing modal feature representations Perform collaborative fusion to obtain fusion feature representation. The system inputs the results into the reconstruction module and outputs the final fused image. ;
[0113] (7) Construct the total loss function using reconstruction loss, structure preservation loss, shared prototype consistency loss, candidate diversity loss, and credibility constraint loss. Its expression is:
[0114]
[0115] in, To rebuild the losses, To preserve structural losses, To share prototype consistency loss, For candidate diversity loss, For credibility constraint loss, , , , , The loss weighting coefficient is adjustable.
[0116] (8) Wherein, the reconstruction loss is used to constrain the pixel difference between the final fused image and the fused reference image, and its expression is:
[0117]
[0118] Structural retention loss The expression used to enhance the structural similarity between the fused image and the fused reference image is:
[0119]
[0120] Shared prototype consistency loss The expression used to constrain the response distribution of existing modes and missing modes to remain consistent in the shared prototype space is:
[0121]
[0122] Candidate diversity loss To prevent the output of multiple candidate branches from collapsing and to encourage different candidates to learn different potential missing modal response patterns, its expression is:
[0123]
[0124] Credibility constraint loss Used to constrain the matching relationship between candidate credibility and candidate quality;
[0125] (9) By calculating the gradient of the loss function, the error is backpropagated from the output layer back to each layer of the network. The gradient vector is determined by backpropagation of the loss value, and the parameters of the missing modality remote sensing image fusion model based on the shared prototype space are updated. It is determined whether the set number of training rounds has been reached. If so, the model training is completed; otherwise, training continues.
[0126] The fourth step involves acquiring image data containing only existing modal information. During model deployment or testing, only existing modal images are used as input; the actual missing modal images are no longer needed. The existing modal images are then preprocessed in the same way as during training before being input into the model for generating the final fused image.
[0127] The fifth step involves inputting the image data to be processed into the trained model, and obtaining the final fused image through multi-candidate missing modality inference, credibility evaluation, and adaptive weighted fusion.
[0128] like Figure 5 As shown, this is a comparison of the fusion results of the present invention in a missing modality scene. (a), (b), (c), and (d) are the existing modality input images in the missing modality scene, and (e), (f), (g), and (h) are the fusion result images obtained after processing the corresponding input images using the method described in this invention. Figure 5 As can be seen, even using only existing modal information, this invention can still generate fusion results with clear structure and rich details. Compared with the input image, the fusion result improves overall contrast, local texture preservation, target region saliency, and scene information representation, while effectively suppressing information loss and instability caused by the unavailability of missing modalities.
[0129] Specifically, the existing modal image to be processed is input into the multi-scale feature encoding module to obtain the existing modal feature representation; then the existing modal feature representation is mapped to the shared prototype space to obtain the prototype response coefficient; then the existing modal feature representation and the prototype response coefficient are input into the missing modal multi-candidate inference module to generate multiple candidate missing modal representations; then the candidate credibility evaluation module generates corresponding credibility maps and fusion weights for each candidate; then the adaptive selection fusion module dynamically selects or weights and fuses the multiple candidate missing modal representations according to the fusion weights to obtain the final missing modal latent representation; finally, the final missing modal latent representation is collaboratively fused with the existing modal feature representation, and the final fused result image is output through the result reconstruction module.
[0130] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for fusing missing modal remote sensing images based on a shared prototype space, characterized in that, Includes the following steps: 11) Obtain a training sample set containing existing modal images and missing modal images: Construct training samples for missing modal scenarios based on complete bimodal image data to simulate the condition that some modalities are not available in real applications. During training, existing modal images are used as model inputs, and missing modal images and fused reference images are used together as supervision information. 12) Construct a multi-candidate reasoning and adaptive fusion model based on a shared prototype space: This model consists of a multi-scale feature encoding module, a shared prototype space construction module, a missing modality multi-candidate reasoning module, a candidate credibility evaluation module, an adaptive selection fusion module, and a result reconstruction module connected in sequence. 13) Supervised training of the model based on training samples: Input the existing modal images in the training samples into the model, and use the missing modal images and the fused reference images as supervision targets to complete the model parameter learning and optimization; 14) Obtain image data to be processed containing only existing modal information: In the actual inference stage, only existing modal images that can be stably obtained are input, and the missing modal images are no longer relied upon; 15) Input the image data to be processed into the trained model, and obtain the final fused image through multi-candidate missing modality inference, credibility evaluation and adaptive weighted fusion.
2. The method according to claim 1, characterized in that, Obtaining a training sample set containing existing modality images and missing modality images includes the following steps: 21) Collect or organize images containing existing modalities and missing modality images The paired sample data, wherein the existing mode For stable and acquireable modes, including visible light and SAR remote sensing images; the missing modes For modes that are prone to loss, including infrared and thermal imaging remote sensing images; 22) Preprocess the paired sample data, including image cropping, size normalization, grayscale normalization, color normalization, noise reduction, and contrast enhancement, to ensure that the input image maintains consistency in spatial scale and pixel distribution. 23) Construct a training scenario for missing modalities, and use the images of the missing modalities. As a training supervision reference, only existing modality images are retained. As input to the model, it is used to reproduce input conditions that are not available in real-world scenarios; 24) Divide the preprocessed sample data into training set, validation set and test set according to the proportion, and use them for model training, hyperparameter tuning and performance evaluation, respectively; 25) Repeat steps 21)-24) to expand and improve the training sample set and enhance the model's generalization ability in complex scenarios.
3. The method according to claim 1, characterized in that, The construction of the multi-candidate reasoning and adaptive fusion model based on a shared prototype space includes the following steps: 31) Construct a multi-scale feature encoding module for existing modal images. Multi-scale feature extraction is performed to obtain structural features, texture features, and global scene context features, outputting a compact and representative representation of existing modal features. ; 32) Construct a shared prototype space construction module to establish a shared prototype space (SPS), which contains a basic shared prototype library. With semantic prototype library The basic shared prototype library is used to represent common structural information such as edges, contours, and texture skeletons shared among different modalities; the semantic prototype library is used to represent high-level semantic information such as target category, scene category, and thermal response mode. 33) Representing existing modal features and missing modal feature representation The corresponding prototype response coefficients are obtained by projecting them onto the shared prototype space SPS. and Its expression is: in, This represents the prototype projection mapping function, used to map modal features to response distributions in the prototype space; 34) Construct a missing modality multi-candidate inference module (MCGM) based on existing modality feature representations. and its prototype response coefficient K candidate missing modality representations are generated through multi-branch parallel inference. Its expression is: in, This represents the kth independent candidate inference branch, used to characterize the pluralism of missing modalities in complex scenarios; 35) Construct a candidate credibility evaluation module (CEM). Based on the matching degree between the candidate missing modality representation and the shared prototype space, the consistency between candidates, the reconstruction error, and the scene context information, generate a pixel-level credibility map for each candidate. ; 36) Construct an adaptive selection fusion module ASFM and a result reconstruction module RM. Based on the confidence map, dynamically select or adaptively weight and fuse multiple candidate missing modal representations, and reconstruct them in collaboration with existing modal features to output the final fusion result image.
4. The method according to claim 1, characterized in that, The model training includes the following steps: 41) The existing modal images in the training samples Input the multi-scale feature encoding module to obtain the feature representation of the existing modality. The prototype response coefficients are then projected onto the shared prototype space SPS to obtain the prototype response coefficients. ; 42) Remove the missing modality images from the training samples. Similarly, mapping to the shared prototype space SPS yields the prototype response coefficients of the missing modes. Establish cross-modal sharing correspondence between existing modalities and missing modalities in the prototype space; 43) Representing existing modal features and prototype response coefficient Input the missing modality multi-candidate inference module MCGM to generate multiple candidate missing modality representations covering different potential response patterns. ; 44) Input multiple candidate missing mode representations into the candidate confidence evaluation module (CEM), calculate the confidence level of each candidate, and generate a confidence map corresponding to the image pixel position. ; 45) Input multiple candidate missing mode representations and confidence maps into the adaptive selection fusion module ASFM, perform hard selection or soft weighted fusion on the candidates, and obtain the final latent representation of the missing mode. Its expression is: in, Let represent the k-th candidate fusion weight determined adaptively by credibility, and satisfy . ; 46) The final missing modality latent representation Compared with existing modal feature representations Perform feature-level collaborative fusion to obtain a fused feature representation. The system inputs the results into the RM reconstruction module and outputs the final fused image. ; 47) Construct the total loss function L for multi-constraint joint supervision: in, To rebuild the losses, To preserve structural losses, To share prototype consistency loss, For candidate diversity loss, For credibility constraint loss, , , , , The loss weighting coefficient is adjustable. 48) Perform backpropagation based on the total loss function L, update the model parameters layer by layer, and determine whether the preset number of training rounds or convergence conditions have been reached. If so, complete the model training; otherwise, return to continue iterative training.
5. The method according to claim 1, characterized in that, The candidate credibility assessment module CEM generates a credibility map. The basis includes: 51) Matching residuals between candidate missing modal representations and the shared prototype space SPS Consistency measure among multiple candidate missing modal representations Reconstruction closed-loop error of candidate missing mode representation and existing modal scene context information ; 52) The credibility diagram The ASFM module is used to guide region adaptive fusion by improving the representation of existing modal features in structurally clear regions. The fusion weights are used to improve the latent representation of the final missing modalities in the target salient region. The right to merge This method reduces the impact of unreliable candidates in high-uncertainty regions, suppresses artifacts and the propagation of erroneous information, and improves the stability and detail preservation of the fusion results.