Multi-modal image fusion method based on block-level self-supervised contrast learning

This multimodal image fusion method, which utilizes block-level self-supervised contrastive learning, addresses the problem of insufficient local feature discrimination in existing technologies. It achieves high-quality fusion of infrared and visible light images, improving image detail fidelity and structural consistency, and is applicable to fields such as intelligent monitoring and autonomous driving.

CN121788366APending Publication Date: 2026-04-03TIANJIN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods lack the ability to distinguish fine features when processing local small areas, resulting in smoothed texture details and blurred edge structures, which cannot meet the requirements of high-precision perception tasks, and lack a self-supervised internal mechanism to evaluate the quality of local fusion blocks.

Method used

A multimodal image fusion method based on block-level self-supervised contrastive learning is adopted. A self-supervised generative adversarial training framework is constructed through a generator, a dual-branch discriminator, and a block-level triplet selection module. Image-level content loss, adversarial loss, and local contrast loss are introduced to adaptively handle local feature heterogeneity.

Benefits of technology

It significantly improves the fidelity of texture details and the consistency of edge structure in fused images, enhances the robustness of the model in complex scenes, and is suitable for high-precision vision tasks such as intelligent monitoring and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121788366A_ABST
    Figure CN121788366A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image fusion method based on block-level self-supervised contrast learning, relates to the technical field of computer vision and image processing, and aims to solve the problem that the existing fusion technology neglects the complexity of image local space change and is difficult to process local heterogeneity in a self-adaptive manner. According to the method, a self-supervised block-level contrast learning framework is constructed, and a triple comprising an anchor point, a positive sample and a negative sample which is difficult to be subjected to spatial dislocation is mined and constructed by utilizing texture statistical characteristics and spatial structure constraints of an image; the invention aims to endow a fusion model with the capability of actively sensing the quality of a local structure, so that the fusion model can adaptively suppress artifacts and blurring caused by inconsistent space structures under the condition of no explicit supervision signal, the performance of a fusion image in the aspects of texture detail fidelity and edge structure consistency is remarkably improved, and the image quality is improved. Therefore, high-quality data support is provided for high-precision visual perception tasks such as intelligent monitoring and automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and image processing technology, and in particular to a method for fusing infrared and visible light images. Specifically, it is a method that uses a self-supervised contrastive learning strategy to capture local heterogeneity in images, so as to improve the fidelity of local details and structural consistency of the fused image. Background Technology

[0002] The core objective of infrared and visible light image fusion technology is to efficiently integrate the thermal radiation information captured by infrared sensors with the texture details acquired by visible light sensors. Infrared imaging relies on differences in object thermal radiation to achieve imaging, possessing strong anti-interference capabilities and all-weather operation. It can penetrate harsh environments such as darkness, smoke, and haze, accurately highlighting thermal targets. Visible light imaging, on the other hand, provides rich scene details and high spatial resolution, better aligning with human visual perception habits. Through the complementary fusion of these two modalities, the generated fused image possesses both high-contrast target features and retains clear background texture, significantly improving the information richness and environmental adaptability of the visual system. This technology demonstrates extremely high application value and broad market prospects in fields such as intelligent security monitoring, autonomous driving environmental perception, military reconnaissance, and multimodal target detection and recognition.

[0003] While deep learning has driven breakthroughs in image fusion, existing fusion methods often rely excessively on global constraint mechanisms when constructing loss functions or designing adversarial targets. These mechanisms include pixel intensity loss or gradient loss based on the entire image, or adversarial discrimination strategies based on the overall image distribution. However, such methods overlook the fundamental characteristics of infrared and visible light images, which stem from fundamentally different physical imaging mechanisms. This leads to significant differences in the feature distribution patterns and spatial mapping relationships of different local regions (patches) within an image—a characteristic defined as local heterogeneity. Existing techniques attempt to process the entire image region using a unified set of global weights or rules, preventing the model from adaptively extracting and fusing key feature information based on the characteristics of different local textures. This severely limits the adaptability of fusion models in complex scenes.

[0004] Due to inherent limitations imposed by global constraints, existing models often lack refined feature discrimination capabilities when processing small local regions. When two modalities exhibit nonlinear differences in local texture or structure, the model typically generates a "compromise" or "average" fusion result to minimize global error. While this approach can maintain the overall contour and low-frequency information of the image at a macroscopic level, it severely damages the high-frequency information, directly causing key texture details to be smoothed, edge structures to become blurred, and even leading to texture distortion and artifacts. This loss of detail makes it difficult for the fused image to meet the stringent requirements for image sharpness and structural fidelity in high-precision perception tasks such as object detection and semantic segmentation.

[0005] More critically, existing fusion frameworks generally employ a passive training strategy that fits a pre-set global optimization objective, lacking an active, self-supervised internal mechanism to evaluate the quality of local fusion blocks. Specifically, traditional models cannot effectively distinguish between two types of fusion blocks in the feature space: high-quality fusion blocks with clear structure and consistent semantics, and low-quality fusion blocks with chaotic structure and conflicting features. This methodological limitation prevents the model from implementing targeted optimization based on the difficulty of the samples when facing complex local feature conflicts. Ultimately, this results in insufficient robustness of the model when dealing with challenging local regions, making it difficult to ensure that the fused image exhibits consistently high-quality performance across all local regions.

[0006] Therefore, a multimodal image fusion method based on block-level self-supervised contrastive learning is proposed to solve the above problems. Summary of the Invention

[0007] In view of this, the technical problem to be solved by this invention is to propose a multimodal image fusion method based on block-level self-supervised contrastive learning, which can adaptively suppress artifacts and blurring caused by spatial structure inconsistencies, and significantly improve the performance of fused images in terms of texture detail fidelity and edge structure consistency, thereby providing high-quality data support for high-precision visual perception tasks such as intelligent monitoring and autonomous driving.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multimodal image fusion method based on block-level self-supervised contrastive learning. Dataset construction: Based on the MSRS dataset, it is divided into a training set (1083 pairs of images) and a test set (361 pairs of images). During the training phase, a random cropping strategy is used to select local regions of size 96×96 as input samples, and data augmentation operations such as random horizontal and vertical flipping are used. The network structure includes a generator, a dual-branch discriminator, and a block-level triplet selection module. The generator adopts a dual encoder-single decoder architecture. The dual-branch discriminator consists of an infrared discriminator and a visible light discriminator. The block-level triplet selection module includes a fine-grained block division mechanism and a self-supervised contrastive learning mechanism. Design loss function: Construct a multi-dimensional joint optimization objective, including image-level content loss, adversarial loss, and local contrast loss; Model training and validation: During the training phase, an end-to-end self-supervised generative adversarial training framework is constructed, with the generator and dual-branch discriminator being updated alternately, and block-level contrast loss is introduced as the core constraint; During the validation phase, the fusion result is generated from the input infrared image and visible light image, and six indicators, namely VIF, EN, AG, EI, SD, and SF, are used for quantitative evaluation.

[0009] Preferably, the dataset construction specifically includes: based on the publicly available MSRS dataset, which is one of the authoritative benchmarks in the field of infrared and visible light image fusion. It contains 1444 pairs of registered infrared and visible light images with a resolution of 640×480. This dataset covers rich road scenes, including various thermal radiation targets such as vehicles and pedestrians, as well as complex background textures, making it ideal for validating the model's ability to handle local details. To ensure the rigor of the experiment, we divided the dataset into training and testing sets: 1083 image pairs were used for model training and parameter optimization, and the remaining 361 image pairs were used for the final evaluation of model performance.

[0010] To enhance the model's generalization ability and prevent overfitting, this invention does not perform full-image training directly on the original resolution images during the training phase. Instead, a random cropping strategy is employed. Specifically, in each iteration, a local region of size 96×96 is randomly cropped from the original 640×480 image pair as the input sample. Furthermore, random horizontal and vertical flipping operations are used to expand the spatial diversity of the data, ensuring that the model can adapt to feature distributions in different directions and locations.

[0011] As a preferred option, the construction of the network structure specifically includes: The generator is responsible for receiving infrared and visible light images as input. These images are encoded into high-dimensional visual features by their respective encoders and then fused. The fused image is then output through the decoder.

[0012] Dual Discriminators: Employing a dual-branch discriminator architecture, consisting of an infrared discriminator and a visible light discriminator. They respectively receive the original modal image and the generated fused image as input, and through adversarial game, force the fused image to approximate the infrared thermal radiation characteristics and visible light texture characteristics in visual distribution, thus achieving adversarial learning at the visual distribution level.

[0013] The Patch-wise Triplet Selection Module, the core constraint module of this invention, includes a fine-grained block segmentation mechanism and a self-supervised contrastive learning mechanism. This module is responsible for mining local heterogeneity in the image, selecting anchor points and positive samples based on statistical features, and synthesizing difficult-to-negative samples using spatial misalignment. Furthermore, through contrastive loss, it forces the model to distinguish between structurally consistent and structurally disordered local features in the feature space, achieving self-supervised perception of local details.

[0014] As a preferred approach, the network structure specifically addresses the limitations of physical imaging mechanisms. Existing sensors cannot simultaneously capture the rich texture information of visible light images and the thermal radiation information of infrared images, resulting in a natural lack of a standard ground truth as a supervisory signal for image fusion tasks. To address this challenge, this invention constructs a multi-dimensional joint optimization objective, aiming to comprehensively constrain the fusion process from global distribution to local details. This loss function system mainly includes three aspects: image-level content loss (covering structural similarity, gradient, and intensity loss), used to ensure that the fused image is consistent with the source image in terms of basic visual features; adversarial loss, through the game between the generator and the discriminator, forcing the fused image to approximate the true modality in overall visual distribution; and local contrast loss, utilizing a self-supervised mechanism to mine image block-level structural information, specifically targeting the feature heterogeneity of local regions for refined constraints.

[0015] The model training and validation specifically include the following: During the training phase, this invention constructs an end-to-end self-supervised generative adversarial training framework. The model receives batches of input infrared and visible light images and dynamically mines the statistical characteristics within the images through a built-in block-level triplet selection module. This module constructs training triples containing anchor points, positive samples, and negative samples based on local texture richness and structural consistency. During optimization, the generator and the dual-branch discriminator are updated alternately, and a block-level contrastive loss is introduced as a core constraint. This loss function aims to maximize the feature similarity between anchor points and positive samples while minimizing the similarity with non-negative samples, thereby forcing the model to autonomously perceive the feature heterogeneity of local regions under unsupervised conditions and guiding the generator to focus on preserving high-frequency details and reconstructing local structures.

[0016] During the validation and testing phase, the model enters inference mode, generating a fusion result from the input infrared and visible light images. At this point, the generator utilizes its refined local feature discrimination capabilities learned during training to adaptively integrate complementary information from different regions, effectively avoiding blurring or conflicts caused by differences in local features, thus achieving high-fidelity preservation of the local structural consistency of the fused image. Finally, a comprehensive quantitative evaluation of the fused image quality is conducted using six mainstream objective evaluation metrics, including visual fidelity and edge strength.

[0017] Compared with existing technologies, the multimodal image fusion method based on block-level self-supervised contrastive learning provided by this invention has the following beneficial effects: The image fusion method proposed in this invention effectively overcomes the smoothing tendency of traditional global constraint methods when dealing with local feature heterogeneity by introducing a block-level triplet selection module, thus solving the technical problems of local texture blurring and high-frequency detail loss. Compared with existing technologies, this scheme, with its refined block-level contrastive learning mechanism, not only achieves high-fidelity preservation of micro-textures but also significantly enhances the model's spatial robustness in non-strict registration scenarios and can adaptively improve the saliency of infrared thermal targets through hard negative sample mining. Based on these advantages, this method can be widely extended to fields such as all-weather perception in intelligent transportation, complex evidence collection in security monitoring, and UAV power line inspection, providing high-definition and target recognition image support for multimodal vision tasks. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall process of the model of the present invention; Figure 2 This is a schematic diagram illustrating the image visualization comparison of the present invention; Figure 3 This is a schematic diagram of the network system structure for an embodiment of the present invention; Figure 4 This is a schematic diagram of the PTSM of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0021] For an example, please refer to... Figures 1 to 4 As shown: To address the issues of existing fusion techniques neglecting the complexity of local spatial variations in images and struggling to adaptively handle local heterogeneity, this application provides a multimodal image fusion method based on block-level self-supervised contrastive learning. The specific process is as follows: This invention constructs an infrared and visible light image fusion model based on a generative adversarial network (GAN). While leveraging the powerful global distribution fitting ability of GANs to ensure the overall visual naturalness of the images, this network structure innovatively utilizes a unique block-level triplet selection mechanism to address the limitation of traditional GANs in capturing local micro-differences. This approach specifically addresses the problem of local feature heterogeneity, thereby achieving an organic unity between macroscopic distribution alignment and microscopic structure fidelity. For example... Figure 3 As shown, the system mainly consists of three core modules: generator. Dual-channel discriminator , and block-level triple selection module .

[0022] (1) Generator ) Generator The architecture adopts a dual encoder-single decoder design, which is responsible for extracting features from the source image and reconstructing the fused image.

[0023] Feature extraction encoder: The generator front end is configured with two encoder branches with independent weights. , They are used to process infrared images. and visible light images The encoder is based on Block construction effectively extracts deep visual features and outputs infrared features respectively. and visible light characteristics .

[0024] Feature fusion and reconstruction: The extracted bimodal features are concatenated (Cat) along the channel dimension to form a joint feature representation. This joint feature is then fed into the backend image decoder. The decoder is also based on Block construction. The decoder maps high-dimensional features back to image space, outputting the final fused image. .

[0025] The entire generation process can be represented as:

[0026] In the formula, Represented as an infrared image, Indicates an image without visible light. Represented as an infrared image encoder, Represented as a visible light image encoder; (2) Dual Discriminators: To ensure that the fused image simultaneously approximates the true distribution of infrared and visible light in terms of visual attributes, this invention designs a dual-path discriminator architecture, including one infrared discriminator. A visible light discriminator Both have the same convolutional neural network structure, consisting of multiple convolutional layers, Leaky ReLU activation function, and fully connected linear layers. Infrared Discriminator Receive real infrared images As positive samples, fused images As an adversarial negative sample, it aims to determine whether the input image has the typical thermal radiation intensity distribution of an infrared image.

[0027] The visible light discriminator receives real visible light images. As positive samples, fused images As an adversarial negative sample, it aims to determine whether the input image has the typical texture detail distribution of a visible light image.

[0028] Through adversarial training, the discriminator forces the generator to output images that retain infrared intensity while possessing realistic visible light textures to ensure global feature consistency.

[0029] (3) Patch-wise Triplet Selection Module ): The block-level triplet selection module is the core constraint component of this invention, used to construct a contrastive loss, mine local heterogeneity of images during the training phase, and form a self-supervised signal. This module operates as a self-supervised sampler, dividing the input infrared, visible light, and generated fused images into non-overlapping, fine-grained image patches. The patching process is defined as follows:

[0030] In the formula, Indicates extracting the first Operators for non-overlapping image blocks N This represents the total number of image blocks, with the block size set to 32×32.

[0031] Anchor selection based on local texture statistics: To capture the most typical local features, the module first calculates the standard deviation response of each fusion block. This metric quantifies the pixel dispersion within an image patch, effectively reflecting the richness and contrast of local texture.

[0032]

[0033] In the formula, This indicates the location of the patch-level fused image. pixel values, H represents the average pixel value of the image, and H and W represent the height and width of the image, respectively. This formula measures the dispersion of the image's pixel values ​​and its overall contrast. Based on this model, all blocks are traversed, and [the following is a process of selecting...] The image patch with the highest value is used as the anchor sample. This represents the most challenging and detailed local area.

[0034] Positive Selection Based on Full Reference Quality Assessment: To determine the "ideal local structure," the module calculates the peak signal-to-noise ratio between the anchor block and other fused blocks in the same group. )Score. It can measure the structural fidelity and consistency of two image patches at the pixel level. This process can be represented as:

[0035] in This represents the PSNR score of the m-th patch block. express The patch block it belongs to. The module sorts and filters based on this metric. The highest value Each image patch is used as a positive sample, ensuring that the positive samples are highly consistent with the anchor points in terms of visual quality and structural features. A positive sample can be represented as:

[0036] The operation is used to select the highest score. 1 index to build a positive sample set .

[0037] Hard Negative Mining Based on Semantic Confusion: In order to construct negative samples with strong discriminative power.

[0038] First, a cross-spatial location fusion strategy is used to integrate spatial locations. infrared block and spatial position ( The visible light block input generator synthesizes a fused block with disordered spatial structure.

[0039]

[0040] in, Indicates position Infrared image blocks, Indicates position Visible light image patch, For generator, This is a merged block with a disordered spatial structure.

[0041] Semantic hard sample mining: To filter out hard negative samples that are more beneficial to the model, the module uses the encoder to extract features from the source image corresponding to the misaligned block and calculates the cosine similarity:

[0042] Spatial location infrared block and position The cosine similarity score between visible light patches is used. Based on this, the patch with the highest similarity (i.e., extremely similar semantic content but completely different spatial structure) is selected. Each misaligned block is used as a non-negative sample:

[0043] This mechanism forces the model to learn to distinguish between features that are "semantically similar but structurally incorrect," thereby significantly improving its sensitivity to local structures.

[0044] Triplet Feature Extraction: To measure the semantic distance between samples in the feature space, this module utilizes a pre-trained feature extractor. Feature mapping is performed on the constructed triplet samples. The feature extractor here is implemented based on the CLIP model image encoder, which has powerful semantic information extraction capabilities. To maintain the stability of feature extraction and utilize general semantic priors, the extractor keeps its parameters frozen. The feature extraction process is defined as follows:

[0045] These output features are used to calculate the contrastive loss, constructing a self-supervised local feature discrimination mechanism. This mechanism forces the model to learn to distinguish between "natural texture features with structural consistency" and "heterogeneous features with spatial distribution conflicts" in the feature space. Through this constraint, the model is endowed with a keen perception of local microstructures, thereby significantly improving the local detail clarity and texture fidelity of the fused image under unsupervised conditions.

[0046] 2. Loss function design.

[0047] To endow the fusion model with adaptive perception capabilities for local complex structures and ensure that the generated images appear natural and realistic from a macroscopic visual perspective, this invention constructs a multi-dimensional constraint system consisting of local structure identification, basic visual fidelity preservation, and global distribution fitting. The overall optimization objective function is... The definition is as follows:

[0048] in, , and These are the adjustment weights for each sub-loss term, which are set to 10, 10, and 1 respectively.

[0049] (1) Local structure discrimination loss based on triples ( ) This is the core constraint of the present invention in addressing local feature heterogeneity. The loss is based on a block-level triplet selection module (…). The sample relationships mined aim to reshape the distribution of local structures in the feature space. By maximizing the mutual information between anchor points and positive samples (structurally self-consistent blocks) while minimizing the mutual information between anchor points and hard-negative samples (spatially conflicting blocks), the model is forced to learn the ability to discriminate local microstructures. This is formally expressed as the InfoNCE loss:

[0050] in, This represents the cosine similarity measure between feature vectors. , , These represent the feature vectors of the anchor point, positive sample, and negative sample, respectively. This is a temperature hyperparameter, set to 0.1 here.

[0051] (2) Multi-scale visual content fidelity loss ) To prevent the loss of fundamental information from the source image during the pursuit of local structure optimization in the fused image, this invention introduces a hybrid constraint at the pixel and structural levels:

[0052] Significant intensity retention ( The pixel-level maximum value strategy is used to force the brightness distribution of the fused image to approximate the region with the most significant signal in the source image (infrared heat source or visible light high brightness area).

[0053] Texture gradient sharpening ( Gradient operator constraints ensure that the fused image inherits the sharpest edges and texture variations from the source image, avoiding smoothing of details.

[0054] Perceived structural consistency ( The structural similarity index (SSIM) is used to constrain brightness, contrast, and structural components at the sliding window scale to ensure that the fusion result conforms to human visual perception.

[0055] (3) Global visual distribution adversarial loss ( ) To further enhance the naturalness of the images, this invention utilizes a generative adversarial game mechanism, employing a dual-path discriminator to constrain the statistical distribution of the generated images. The generator's goal is to produce images capable of deceiving the infrared discriminator. and visible light discriminator The image possesses both infrared intensity characteristics and visible light texture characteristics in its macroscopic style:

[0056] in The sigmoid activation function is used, and this loss pushes the fused image to approximate the manifold of the real modality image at the probability distribution level.

[0057] 3. Model training and validation.

[0058] The training and verification process of this invention includes the following steps: (1) Training process During the training phase, this invention employs an end-to-end generative adversarial network framework for joint optimization. The model receives batches of input infrared and visible light images, extracts features through a generator, and fuses them to generate the final fused image. Subsequently, the generator engages in adversarial competition with a two-branch discriminator. The discriminator attempts to distinguish the fused image from the real modality image, while the generator strives to visually approximate the real modality. Based on this, image-level content loss, adversarial loss, and the core block-level triplet contrast loss are combined to comprehensively constrain the network, forcing the model to maintain global visual naturalness while self-supervisedly perceiving and correcting the heterogeneity of local features. The entire training process uses the AdamW optimizer for iterative parameter updates, with 100 training epochs and a batch size of 2 to ensure stable convergence and the acquisition of robust local detail preservation capabilities.

[0059] (2) Verification process During the validation phase, the model uses the trained generator for inference. Unlike the training phase, which relies on block-level triples to construct self-supervised constraints, the algorithm of this invention outputs the fused result in an end-to-end manner by inputting infrared and visible light images from the test set during the inference phase.

[0060] To comprehensively evaluate the quality of the generated images, this invention employs multiple objective evaluation metrics for quantitative analysis. These metrics can measure the visual fidelity and information richness of the fusion results from different perspectives: VIF (Visual Information Fidelity): This metric combines statistical patterns of natural scenes with models of the human visual system to quantitatively evaluate the degree of information sharing between the fused image and the source image. A higher VIF value indicates better visual fidelity in the fusion result and a greater consistency with human visual perception.

[0061] EN (Entropy): Used to measure the average amount of information contained in an image. The higher the EN value, the more effective information the fused image inherits from the source image, and the richer the detail in the image.

[0062] AG (Average Gradient): This metric characterizes the sharpness and texture variation of an image by calculating the rate of contrast change of minute details within the image. A higher AG value indicates sharper edges, higher sharpness, and that the algorithm more effectively preserves the gradient information of the source image.

[0063] EI (Edge Intensity): This is specifically used to evaluate the quality of edge information preservation in an image. A higher EI value indicates that the fused image performs better in maintaining the structural edges and contour features of the source image.

[0064] SD (Standard Deviation): This metric reflects the dispersion of pixel grayscale values ​​and overall contrast of an image. A higher SD value generally means that the image has stronger visual contrast and richer information hierarchy.

[0065] SF (Spatial Frequency): This metric is mainly used to measure the activity of spatial details in an image. The higher the SF value, the more high-frequency components are contained in the fused image, meaning it has richer texture details and finer spatial structure features.

[0066] (3) Experimental results To objectively evaluate the performance of the image fusion method proposed in this invention, we conducted a quantitative comparison with five existing mainstream and advanced fusion algorithms on a public test set, including LRRNET, DDFM, TGFuse, PromptF, and LUTFuse. The evaluation metrics covered visual information fidelity (VIF), entropy (EN), standard deviation (SD), spatial frequency (SF), average gradient (AG), and edge intensity (EI), comprehensively measuring the quality of the fused image.

[0067] Table 1. Comparison results with advanced algorithms:

[0068] As shown in Table 1, this invention achieved optimal results in five of the six key evaluation indicators (VIF, EN, SD, AG, and EI). Specifically, the visual information fidelity (VIF) of this method reached 1.0685, the entropy (EN) reached 6.7274, and the standard deviation (SD) reached 43.7134, all significantly better than the comparison algorithm. This indicates that the fused image generated by this invention not only best conforms to human visual perception in terms of overall visual perception, but also retains the complementary information and contrast in infrared and visible light images to the greatest extent.

[0069] Further analysis reveals that this method demonstrates a significant advantage in the mean gradient (AG) and edge intensity (EI) metrics, which reflect image sharpness and edge structure quality, reaching 3.8561 and 43.3514 respectively, substantially outperforming other GAN-based or traditional deep learning methods. This substantial improvement is directly attributed to the block-level difficult triplet selection introduced in this invention. The strategy involves constructing triples containing anchor points, positive samples, and non-negative samples during GAN adversarial training and introducing a self-supervised contrastive loss. This forces the generator to finely perceive the feature heterogeneity of local regions in the feature space. This mechanism allows the model to no longer focus solely on fitting the global distribution but to adaptively optimize local microstructures. Consequently, it achieves the ultimate preservation of high-frequency texture details and edge structures in the fusion result, effectively avoiding the local blurring phenomenon common in traditional fusion methods.

[0070] Table 2 Ablation experiments on key components;

[0071] Table 2 presents the ablation experiment results for the core component of the proposed method. To verify the effectiveness of this component and its role in feature alignment and representation learning, we compared the complete model with "without using..." "The baseline configuration and "will A quantitative comparison was made with the variant of "image-level learning".

[0072] First, compare "not using" The test results for the "complete model" show that the complete model achieved significant improvements across all metrics. Specifically, the increases in EN and SD indicate that the fused image contains richer source image information and higher contrast after the introduction of this component; while the improvement in VIF (from 1.0447 to 1.0685) verifies the component's advantage in maintaining image fidelity and consistency with human visual perception.

[0073] Secondly, to further explore the impact of different learning granularities on fusion performance, we compared "image-level" fusion performance. The experimental results show that although image-level constraints bring some improvement compared to the baseline model, they still lag behind the full model in key metrics reflecting texture details and edge strength. As can be seen from the table, the full model achieves 3.8561 for AG (average gradient) and 43.3514 for EI (edge ​​strength), significantly outperforming the image-level learning's 3.7721 and 41.3202. This result strongly demonstrates that simple global constraints are insufficient to accurately capture local misalignments and fine-grained features between infrared and visible light images. The patch-level strategy proposed in this paper effectively enhances the edge sharpness and texture detail preservation capabilities of the fused image through more refined feature mining.

[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0075] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multimodal image fusion method based on block-level self-supervised contrastive learning, characterized in that, Includes the following steps: S1, Dataset Construction: Based on the MSRS dataset, it is divided into training and testing sets. During the training phase, a random cropping strategy is used to select local regions of size 96×96 as input samples, and data augmentation operations such as random horizontal and vertical flipping are added. S2, Building the network structure: including a generator, a dual-branch discriminator and a block-level triplet selection module. The generator adopts a dual encoder-single decoder architecture. The dual-branch discriminator consists of an infrared discriminator and a visible light discriminator. The block-level triplet selection module includes a fine-grained block division mechanism and a self-supervised contrastive learning mechanism. S3, Design the loss function: Construct a multi-dimensional joint optimization objective, including image-level content loss, adversarial loss, and local contrast loss; the multi-dimensional joint optimization objective function... The definition is as follows: In the formula, , and These are the adjustment weights for each sub-loss term; S4, Model Training and Validation: During the training phase, an end-to-end self-supervised generative adversarial training framework is constructed, with the generator and dual-branch discriminator being updated alternately, and block-level contrast loss is introduced as the core constraint; During the validation phase, the fusion result of the input infrared image and visible light image is generated and quantitatively evaluated using six indicators, including visual information fidelity (VIF), entropy of information contained in the image (EN), average gradient (AG), edge strength (EI), standard deviation (SD), and spatial frequency (SF).

2. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The MSRS dataset contains 1,444 pairs of registered infrared and visible light images, covering a wide range of road scenes, as well as various thermal radiation targets and complex background textures.

3. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The generator's dual encoders have independent weights and are used to process infrared and visible light images respectively to extract deep visual features. The extracted dual-modal features are concatenated in the channel dimension to form a joint feature representation, which is then mapped back to the image space by the decoder to output a fused image.

4. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The infrared discriminator and the visible light discriminator of the dual-branch discriminator have the same convolutional neural network structure, both consisting of multiple convolutional layers, LeakyReLU activation function and fully connected linear layers; the infrared discriminator uses real infrared images as positive samples and fused images as adversarial negative samples, while the visible light discriminator uses real visible light images as positive samples and fused images as adversarial negative samples.

5. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The working process of the block-level triplet selection module described in S2 includes: S2.1, Fine-grained block division: Divide the input infrared, visible light and generated fused image into non-overlapping image blocks of size 32×32; S2.3, Anchor Point Selection: Calculate the standard deviation response of each fusion block and select the image block with the highest standard deviation response as the anchor point sample; calculate the standard deviation response of each fusion block. The formula is In the formula, This represents the fused image at the i-th patch level. Indicates its position pixel values, H represents the average pixel value of the image, and H and W represent the height and width of the image, respectively. This formula measures the dispersion of the image's pixel values ​​and its overall contrast. Based on this model, all blocks are traversed, and [the following is a process of selecting...] The image patch with the highest value is used as the anchor sample. This represents the most challenging and detailed local area at present; S2.3, Positive Sample Selection: Calculate the peak signal-to-noise ratio between the anchor block and other fused blocks in the same group, and select the image blocks with the highest peak signal-to-noise ratio as positive samples; In the formula, This represents the PSNR score of the m-th patch block. express The patch block it belongs to. The module sorts and filters based on this metric. The highest value Each image patch is used as a positive sample; the specific expression for a positive sample is: In the formula, Indicated as according to Sort, This represents the number of samples selected. This is represented as the set of positive samples used for subsequent comparative learning; S2.4, Difficult-to-Bear Sample Synthesis and Mining: Infrared and visible light blocks from different spatial locations are input into the generator to synthesize spatially disordered fusion blocks. The cosine similarity of the source image features corresponding to the disordered blocks is calculated, and the several disordered blocks with the highest similarity are selected as difficult-to-bear samples. The specific formula for the spatially disordered fusion block is as follows: In the formula, Indicates position Infrared image blocks, Indicates position Visible light image patch, For generator, This results in a fused block with a disordered spatial structure. The formula for calculating the cosine similarity of the source image features corresponding to the misaligned block is as follows: Select the most similar misaligned blocks as the hard-to-bear samples. P Neg The specific formula is as follows: In the formula, (m,n) represents the characteristic cosine similarity between the m-th infrared block and the n-th visible light block; S2.5, Triple Feature Extraction: The CLIP model image encoder with frozen parameters performs feature mapping on triples composed of anchor points, positive samples, and non-negative samples. The feature extraction process is defined by the following formula: 。 6. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The image-level content loss described in S3 includes significant intensity preservation loss, texture gradient sharpening loss, and perceptual structure consistency loss, wherein: S3.1, the saliency intensity preservation loss employs a pixel-level maximum value strategy, forcing the brightness distribution of the fused image to approximate the region with the most significant signal in the source image; the saliency intensity preservation loss The expression is: ; In the formula, Represented as a fused infrared image; Represented as the original visible light image, Represented as the raw infrared image: In the formula, S3.2, the texture gradient sharpening loss, constrained by the gradient operator, ensures that the fused image inherits the sharpest edges and texture variations from the source image; the texture gradient sharpening loss... The expression is: ; S3.3, the perceptual structural consistency loss uses a structural similarity index to constrain brightness, contrast, and structural components at a sliding window scale; the perceptual structural consistency loss The expression is: 。 7. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The adversarial loss is achieved through a game between the generator and the dual-branch discriminator. The generator aims to generate an image that can deceive the discriminator, making it possess both infrared intensity features and visible light texture features in terms of macroscopic style. The adversarial loss function is constructed using the sigmoid activation function.

8. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The local contrast loss is the InfoNCE loss, which maximizes the feature similarity between the anchor point and the positive sample while minimizing the similarity between the anchor point and the hard negative sample. The InfoNCE loss function is expressed as follows: In the formula, This represents the cosine similarity measure between feature vectors. , , These represent the feature vectors of the anchor point, positive samples, and non-negative samples, respectively. This is a temperature hyperparameter, set to 0.1 here.

9. The multimodal image fusion method based on block-level self-supervised contrastive learning according to claim 1, characterized in that, The model training uses the AdamW optimizer for parameter iterative updates, with 100 training epochs and a batch size of 2. Among the weight adjustments for each loss term, the image-level content loss has a weight of 10, the adversarial loss has a weight of 10, and the local contrast loss has a weight of 1.

Citation Information

Patent Citations

  • Bimodal image fusion method and device based on tuple disturbance and storage medium

    CN120783175A

  • Generative adversarial network-based MRI-PET mode conversion method and system

    CN121190599A

  • Target identification tracking method based on self-supervision mechanism

    CN121213951A