Infrared and visible image fusion method based on text semantic consistency guidance

By employing a dual-branch coding strategy guided by textual semantic consistency and decoupled from structure and intensity, the problem of semantic confusion in the fusion of infrared and visible light images is solved. This achieves synergistic enhancement of infrared target saliency and visible light structural details, thereby improving the semantic consistency and visual quality of the fused image.

CN121639491BActive Publication Date: 2026-05-08XIAMEN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN UNIV OF TECH
Filing Date
2026-02-04
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods struggle to accurately distinguish foreground targets from background clutter in complex scenes, leading to semantic confusion. Furthermore, they lack explicit modeling and utilization of high-level semantic information in images, affecting the semantic consistency and support capabilities for subsequent advanced vision tasks.

Method used

A text semantic consistency-guided approach is adopted, which constructs a unified text semantic prior through fine-grained text semantic generation and cross-modal semantic relationship modeling. Combined with a structure-strength decoupling dual-branch coding strategy, explicit and implicit semantic alignment of infrared and visible light visual features is achieved, and adaptive weighted fusion is performed.

Benefits of technology

While maintaining the saliency of infrared targets, it enhances visible light structural details, improves the semantic consistency and visual quality of fused images, and is suitable for multimodal image fusion tasks in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121639491B_ABST
    Figure CN121639491B_ABST
Patent Text Reader

Abstract

The application provides an infrared and visible light image fusion method based on text semantic consistency guidance, relates to the technical field of multi-modal image fusion, and comprises the following steps: fine-grained text semantics of infrared and visible light images are generated respectively and mapped to a unified embedding space; a cross-modal attention mechanism is used to bidirectionally compensate and enhance the text semantics, and a unified text semantic prior is constructed; a structure-intensity decoupled double-branch encoder is used to extract structure texture features of visible light and intensity saliency features of infrared respectively; under the guidance of the text semantic prior, explicit semantic consistency constraints and implicit semantic distribution consistency constraints are used to align the double-modal visual features in a shared semantic space; and finally, the text semantic prior is used as a global modulation signal to adaptively weight and fuse the aligned features and decode to generate a fusion image. The application effectively solves the problems of insufficient semantic modeling and poor fusion result consistency of existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multimodal image fusion technology, specifically to an infrared and visible light image fusion method guided by text semantic consistency. Background Technology

[0002] With the rapid development of computer vision and artificial intelligence technologies, multimodal image fusion technology is playing an increasingly important role in fields such as intelligent surveillance, autonomous driving, remote sensing, and military reconnaissance. Among them, infrared and visible light image fusion, by integrating the significant detection capability of infrared images for thermal targets with the rich texture and structural details of visible light images, can generate fused images that are more comprehensive and visually interpretable, thereby effectively improving scene perception and understanding capabilities in complex environments (such as nighttime, fog, and camouflage).

[0003] Early infrared and visible light image fusion methods were mainly based on traditional image processing techniques, such as multi-scale transformations (e.g., pyramid decomposition, wavelet transform), saliency analysis, and rule-based feature selection and weighted fusion. These methods rely on manually designed features and fusion rules. Although the principles are intuitive, their performance is heavily dependent on parameter tuning. They are poorly adaptable to complex and variable scenes and struggle to reconstruct visible light details and textures with high quality while preserving the saliency of infrared targets. This often leads to problems such as blurred edges, unbalanced contrast, or loss of semantic information in the fusion results.

[0004] In recent years, with the rise of deep learning, fusion methods based on Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), and Transformers have gradually become mainstream. These methods automatically learn the mapping from source images to fused images in an end-to-end manner, extracting deeper and more robust feature representations, thus achieving significant improvements in both objective metrics and subjective visual quality. However, most existing deep learning methods are essentially still a "vision-to-vision" mapping process, with their fusion decisions mainly relying on statistical regularities learned from pixels or low-level features, lacking explicit modeling and utilization of high-level semantic information of images. This leads to difficulties in accurately distinguishing foreground targets from background clutter in complex scenes, easily causing semantic confusion, such as misfusion of a high-temperature background as a target, or weakening low-temperature but structurally important objects. Thus, while improving visual effects, the fused image may compromise its semantic consistency and support capabilities for subsequent high-level visual tasks (such as object detection and recognition).

[0005] To introduce semantic guidance, some studies have attempted to utilize textual descriptions or cross-modal attention mechanisms. For example, image-annotated data or simple category labels can provide coarse semantic cues for the fusion process. However, these methods typically have significant limitations: First, the textual descriptions used are often coarse-grained (e.g., single object categories), failing to provide fine-grained semantic information about scene structure, target attributes, and interrelationships, thus limiting their guiding effect. Second, cross-modal interaction methods are relatively simple (e.g., direct splicing or shallow attention), failing to deeply model the profound semantic connections and complementarities between infrared and visible light modalities, making it difficult to establish stable and accurate semantic consistency constraints. Third, existing methods often mix and encode features from different modalities, failing to effectively decouple the structural texture information rich in the visible light modality from the intensity saliency information unique to the infrared modality. This feature coupling makes it difficult for the network to coordinate the two information flows during fusion, easily leading to semantic information imbalance—overemphasizing intensity may destroy structural integrity, while over-preserving structure may weaken target saliency.

[0006] In view of the above, this application is hereby submitted. Summary of the Invention

[0007] This invention provides an infrared and visible light image fusion method based on text semantic consistency guidance, which can at least partially improve the above-mentioned problems.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] An infrared and visible light image fusion method guided by text semantic consistency, comprising:

[0010] Visible light and infrared images are acquired, and fine-grained text semantic generation processing is performed on the visible light and infrared images to obtain a text description set. The text description set is then mapped to a unified text embedding space to generate multimodal text semantic features.

[0011] Cross-modal semantic relationship modeling is performed on the multimodal text semantic features, wherein the semantic interaction relationship between visible light text and infrared text is modeled to compensate for and enhance the semantics of visible light text and infrared text, and a unified text semantic prior is constructed.

[0012] Based on the structure-intensity decoupling dual-branch coding strategy, visual feature coding is performed on visible light images and infrared images respectively. Specifically, visible light visual features, mainly composed of structure and texture information, are extracted from visible light images, while infrared visual features, mainly composed of intensity distribution and thermal target response, are extracted from infrared images.

[0013] Guided by the unified text semantic prior, the visible light visual features and infrared visual features are subjected to bimodal semantic alignment processing. In this process, through explicit semantic consistency constraints and implicit structural consistency constraints, the visible light visual features and infrared visual features gradually converge in the unified semantic space.

[0014] Based on the visible light visual features after bimodal semantic alignment, the infrared visual features after bimodal semantic alignment, and the unified text semantic prior, semantic modulation fusion processing is performed to generate fused features. The fused features are then input into a preset decoder to obtain the final fused image.

[0015] In summary, this method generates fine-grained textual semantic descriptions for both infrared and visible light images, maps these textual semantics to a unified text embedding space, constructs cross-modal semantic relationships at the text level, compensates for and enhances the semantics of different modalities, forming a unified textual semantic prior. Based on this, a structure-intensity decoupled dual-branch coding strategy is employed to extract visible light structural texture features and infrared intensity saliency features respectively. Guided by the textual semantic prior, explicit semantic consistency constraints and implicit semantic distribution consistency constraints are used to align the bimodal visual features in a shared semantic space. Subsequently, the textual semantic prior is used as a global semantic modulation signal to adaptively weight and fuse the semantically aligned infrared and visible light features, and a decoding network is used to generate the final fused image. This invention can enhance visible light structural details while maintaining the saliency of infrared targets, effectively improving the semantic consistency and visual quality of the fused image, and is suitable for multimodal image fusion tasks in complex scenes.

[0016] Compared with existing technologies, this method can effectively align and fuse infrared and visible light visual features in a unified semantic space under the guidance of text semantics, enhance the semantic consistency and structural integrity of the fusion results, and take into account both the saliency of infrared targets and the expression of visible light details. It is suitable for multimodal image fusion tasks in complex scenes. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the infrared and visible light image fusion method based on text semantic consistency guidance provided in this embodiment of the invention.

[0018] Figure 2 This is a schematic diagram of the overall architecture of the infrared and visible light image fusion method based on text semantic consistency guidance provided in the embodiments of the present invention.

[0019] Figure 3 This is a schematic diagram of dual-modal feature alignment based on explicit and implicit semantic alignment of infrared and visible light features according to an embodiment of the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0021] refer to Figure 1 , Figure 2 As shown, the first embodiment of the present invention discloses an infrared and visible light image fusion method guided by text semantic consistency, which can be executed by an infrared and visible light image fusion device guided by text semantic consistency (hereinafter referred to as the fusion device), specifically, by one or more processors within the fusion device, to implement the following method:

[0022] S1, acquire visible light images and infrared images, perform fine-grained text semantic generation processing on the visible light images and infrared images to obtain a text description set, and map the text description set to a unified text embedding space to generate multimodal text semantic features;

[0023] Specifically, step S1 further includes: acquiring a visible light image and an infrared image as multimodal inputs to be fused, wherein the visible light image is used to provide detailed information, including scene structure, texture and edges, etc., and the infrared image is used to provide target thermal radiation intensity distribution and salient target information;

[0024] The visible light and infrared images are input into a large language model to obtain a fine-grained visible light text representation. and infrared light text description And based on visible light text representation and infrared light text description Constitutes a text description set;

[0025] The BLIP visual language encoder is used to map the text description set to a preset unified text embedding space to generate visible light text semantic features. and infrared light text semantic features And based on the semantic features of visible light text and infrared light text semantic features This yields multimodal text semantic features, among which, This is a mapping function.

[0026] In this embodiment, the visible light image and infrared image to be fused are first acquired as multimodal inputs. The visible light image primarily carries rich spatial details such as the scene's structure, texture, and edges, while the infrared image reflects the target's thermal radiation characteristics, clearly outlining the intensity distribution and salient areas of the thermal target. The two images are naturally complementary at the information level. To establish a bridge at the high-level semantic level, semantic transformation of the information from these two visual modalities is required.

[0027] The fine-grained text semantic generation process involves using a large language model to generate text descriptions containing scene structure information and target semantics based on the input image and unified prompt words. Then, a pre-trained visual language model is used to map the different text descriptions to a unified text embedding space. This conversion between the large language model and the visual language model helps alleviate the semantic differences in the expressive emphasis of texts from different modalities.

[0028] Specifically, the acquired visible light and infrared images are input into a large language model. Guided by unified prompts, this large language model can perform deep understanding and analysis of the image content. It doesn't simply identify object categories, but generates fine-grained natural language descriptions containing information such as scene composition, target attributes, and spatial relationships, thus obtaining corresponding visible light and infrared light textual descriptions, which together constitute a preliminary textual description set. This step elevates pixel-level visual information to language-level semantic information, laying the semantic foundation for subsequent cross-modal interaction. Its beneficial effect lies in its ability to extract high-level contextual and intent information beyond low-level visual features.

[0029] However, directly generated text representations remain in a discrete natural language space and cannot be directly computed with subsequent neural network features. Therefore, it is necessary to map these text representations into a continuous, computable vector space. Next, a pre-trained visual language model (e.g., BLIP or BLIP2) text encoder is used as the mapping function to map visible light and infrared light text representations to a predefined, unified text embedding space. This generates dense vector-form visible light and infrared light text semantic features, which together constitute multimodal text semantic features. Through this mapping operation, text information from different modalities and with different descriptive emphases is unified into the same semantic metric space, making semantic comparison, compensation, and fusion possible. This effectively alleviates the semantic discrepancies in the original descriptions of infrared and visible light images caused by different imaging principles, providing high-quality input for constructing globally consistent semantic priors.

[0030] In summary, a large language model, guided by a unified prompt word, generates a natural language description containing scene structure, target attributes, and semantic relationships based on two input modal images. Subsequently, a pre-trained visual language model, BLIP2, maps the text description to a unified text embedding space, yielding visible light and infrared text semantic features. This step transforms visual information into a high-level semantic representation that can be used for subsequent cross-modal alignment, thus providing prior constraints for semantic consistency modeling of visual features.

[0031] S2, perform cross-modal semantic relationship modeling on the multimodal text semantic features, wherein the semantic interaction relationship between visible light text and infrared text is modeled to compensate for and enhance the semantics of visible light text and infrared text, and a unified text semantic prior is constructed.

[0032] Specifically, step S2 further includes: based on the semantic features of visible light text and infrared light text semantic features The attention weights for the infrared and visible light modalities are calculated using the following formula: , ,in, Infrared semantic attention-weighted representation is aggregated for visible light text. It is an exponential function. , , , , , These are all learnable linear mapping matrices used for calculating attention weights. For embedded dimensions, The scaling factor. Aggregate visible light semantic attention-weighted representations for infrared text;

[0033] Based on the infrared semantic attention weighted representation of visible light text aggregation Infrared text aggregation with visible light semantic attention weighted representation The semantic consistency compensation signal is generated in both visible and infrared modes, and its calculation formula is as follows: , ,in, , All are learnable linear mapping coefficients used to generate a consistency compensation signal with the same dimension as the original embedding. This is a visible light compensation signal. For infrared compensation signals;

[0034] Compensation for visible light signals Infrared compensation signal visible light text semantic features and infrared light text semantic features By adding them together, we obtain the semantically enhanced visible light text semantics. and semantically enhanced infrared text semantics ;

[0035] Semantic enhancement of visible light text and semantically enhanced infrared text semantics The fusion process is performed to serve as a global semantic constraint for subsequent visual feature alignment and fusion, resulting in a unified text semantic prior. , It is a lightweight mapping function used to construct semantic priors for subsequent guiding visual features.

[0036] In this embodiment, the cross-modal semantic relationship modeling is achieved by constructing a bidirectional semantic compensation path to explicitly model the semantic association between two modalities of text, generating semantically enhanced text semantic features. This includes constructing bidirectional semantic compensation paths in the text embedding space for "visible light text aggregated with infrared semantics" and "infrared text aggregated with visible light semantics," respectively, to explicitly model the semantic association between the two modalities of text, ultimately obtaining the enhanced text semantics.

[0037] Specifically, firstly, by linearly mapping the semantic features of visible light text to those of infrared text, and combining this with an attention calculation mechanism, a weighted representation characterizing the semantic correlation between the two modalities of text is generated using QKV matrices of different modalities. Secondly, based on the aforementioned attention-weighted representation, semantic consistency compensation signals are generated for both the visible light and infrared directions. In one specific implementation, the attention-weighted representation is converted into a compensation vector with the same dimension as the original text semantic features through a linear mapping operation, used to characterize supplementary information from the other modality of text semantics. Finally, the semantic consistency compensation signal is fused with the corresponding original text semantic features to obtain semantically enhanced visible light and infrared text semantic features.

[0038] Among them, such as Figure 2 As shown in the text relation perception module, an attention mechanism is used to calculate the weighted representation of text semantic relations, and then a compensation signal is generated to finally obtain the enhanced visible light text semantic features and infrared text semantic features. Through the above steps, the enhanced visible light text semantics and infrared text semantics are fused to construct a unified text semantic prior, which is used for global semantic guidance in the subsequent visual feature alignment and fusion process, so as to alleviate the deviation in semantic expression emphasis of different modal texts.

[0039] In simple terms, this method calculates the weights between cross-modal text features through an attention mechanism and updates the original text semantic features with weights. That is, through the above steps, the semantic relationship between visible light text and infrared text can be explicitly modeled in the text embedding space, thereby obtaining semantically enhanced text semantic features, which can be used to construct a unified text semantic prior, thereby alleviating the semantic bias in the expression focus of different modal texts.

[0040] S3, based on the structure-intensity decoupling dual-branch coding strategy, performs visual feature coding processing on visible light images and infrared images respectively. Specifically, visible light visual features, mainly based on structure and texture information, are extracted from visible light images, while infrared visual features, mainly based on intensity distribution and thermal target response, are extracted from infrared images.

[0041] Specifically, step S3 further includes: the structure-intensity decoupled dual-branch coding strategy adopts a visible light branch and an infrared branch, wherein the visible light branch focuses on extracting scene structure, edge and texture information through multi-scale convolution and spatial attention mechanism, and the infrared branch focuses on extracting thermal target intensity distribution and significant region information;

[0042] The visible light image is input into the visible light branch to extract texture and structural features at different scales. The visible light branch consists of three convolutional layers with different scales to capture local texture and global structural information at different scales. A spatial attention module is then used to enhance key structural regions at different scales. , , The output features of three convolutional layers with different convolutional scales were obtained. , , ;

[0043] The output features of three convolutional layers with different convolutional scales , , After addition, the results are input into a preset spatial attention module to obtain visible light visual features. Among them, convolutional modules of different scales can adapt to textures and structures of different sizes in the image to extract the texture and structural features contained in the visible light image to the greatest extent, while spatial attention can filter different texture and structural features and select the relatively important features for subsequent modules to process.

[0044] The infrared image is input into the infrared branch to extract the intensity distribution and thermal response characteristics of the infrared image, thus obtaining infrared visual features. The infrared branch consists of three identical infrared feature extraction convolutional modules and channel attention modules, with each infrared feature extraction convolutional module containing a... The system employs a convolutional layer with a kernel size of [size missing] and a ReLU function (non-linear activation function). The infrared branch extracts the intensity distribution and salient region information of the thermal target through multi-layer convolution and a channel attention mechanism; channel attention then obtains the final infrared image features. Multi-layer convolution effectively extracts intensity distribution information and salient regions from the infrared image, while channel attention filters these distributions and regions, selecting the most important ones. Furthermore, the channel attention mechanism enhances salient channels related to thermal radiation. This dual-branch encoding method fully leverages the complementary characteristics of different modalities at both the structural and intensity levels.

[0045] In this embodiment, as Figure 2 As shown in the visual structure-intensity perception module, after inputting two modal images into the dual-branch structure, visual features of different modalities can be obtained. Specifically, the dual-branch encoder consists of two parallel, structure-specific branches: a visible light branch and an infrared branch. This heterogeneous design is the core embodiment of the "structure-intensity decoupling" concept. For visible light images, the rich scene details, object contours, and surface textures they contain are their core value. Therefore, the visible light branch is specifically designed to focus on extracting scene structure, edge, and texture information. In specific operation, the visible light image is input into this branch. This branch contains three convolutional layers of different scales, and this multi-scale design has a clear beneficial effect: small-scale ( Convolution helps integrate channel information and perceive subtle textures; medium-scale ( ) is a classic choice for extracting local features; while larger scales ( The convolutional kernels can capture broader contextual information, adapting to different sizes of structural and texture patterns in the image. Through these three parallel convolutional layers, output features corresponding to different receptive fields are obtained. Subsequently, these three multi-scale feature maps are added and fused to form an intermediate representation that integrates structural information from local to global perspectives. To further enhance the focus on key spatial regions, this fused feature is input into a pre-defined spatial attention module. This module can automatically learn and generate a spatial weight map, recalibrating the responses at different locations in the feature map. Its beneficial effect is that it can filter and enhance structural regions (such as clear edges and complex texture areas) that are crucial to subsequent fusion tasks, while suppressing redundant or noisy information, thus obtaining refined and task-oriented visible light visual features. The entire process realizes a progressive feature extraction from "multi-scale capture" to "spatial selective enhancement."

[0046] Meanwhile, for infrared images, their core value lies in the fact that pixel intensity directly reflects the thermal radiation characteristics of an object, highlighting the difference between thermal targets and the background. Therefore, the infrared branch is specifically designed to focus on extracting the intensity distribution and salient region information of thermal targets. In practice, the infrared image is input into this branch. This branch consists of three identical infrared feature extraction convolutional modules connected in series, each module containing a convolutional kernel of size [missing information]. The system employs convolutional layers and a ReLU nonlinear activation function. The advantage of this multi-layered, stacked convolutional structure lies in its ability to delve deeper into the intensity distribution patterns related to thermal radiation in infrared images through progressive nonlinear transformations, effectively highlighting regions of thermally responsive targets. After extracting rich channel features through the convolutional modules, a channel attention module is fed into the features to further differentiate the importance of the thermal features represented by different channels. This module evaluates the contribution of each feature channel and generates a channel weight vector, strengthening feature channels significantly related to thermal targets while suppressing secondary or redundant channels. By combining multi-layered convolutions with the channel attention mechanism, infrared visual features with enhanced significant thermal responses are ultimately obtained. The advantage of this design is its ability to adaptively focus on the feature channels most relevant to thermal targets, ensuring that the intensity saliency information of infrared modes is preserved and highlighted to the greatest extent during the encoding process.

[0047] S4. Under the guidance of the unified text semantic prior, the visible light visual features and infrared visual features are subjected to bimodal semantic alignment processing (explicit and implicit alignment are mainly achieved through loss functions). The visible light visual features and infrared visual features gradually converge in the unified semantic space through explicit semantic consistency constraints and implicit structural consistency constraints.

[0048] Specifically, step S4 further includes: the bimodal semantic alignment processing includes two mechanisms: explicit semantic alignment and implicit semantic alignment. The explicit semantic alignment is achieved by constraining the similarity between visual features and corresponding text semantic features in a shared semantic space, which is used to constrain the consistency between visual features and corresponding text semantic features. The implicit semantic alignment is achieved by constraining the consistency of the semantic distribution of visible light visual features and infrared visual features under text guidance in a shared semantic space, which is used to constrain the consistency of the semantic distribution of different modal visual features under text guidance.

[0049] Visible light visual features and infrared visual features Mapped to a predefined shared semantic space, the formula is: , , , These are all learnable linear transformation coefficients, used to give visual features a semantic representation consistent with text embedding. For visible light visual features mapped to a shared semantic space, Infrared visual features mapped to a shared semantic space;

[0050] In explicit semantic alignment, an explicit loss function is computed. To achieve similarity between constrained visual features and corresponding textual semantic features, specifically: based on visible light visual features mapped to a shared semantic space. Infrared visual features mapped to a shared semantic space visible light text semantic features and infrared light text semantic features The formula for calculating similarity is: , , The similarity between visible light visual features and text semantic features. The similarity between infrared visual features and text semantic features. This is the function for calculating cosine similarity.

[0051] Based on the similarity between visible light visual features and text semantic features Similarity between infrared visual features and text semantic features The formula for constructing explicit semantic consistency constraints is as follows: Used to constrain the consistency between visual features and corresponding textual semantic features.

[0052] In implicit semantic alignment, an implicit loss function is computed. To achieve semantic distribution consistency between constrained visible light visual features and infrared visual features under text guidance, specifically: calculate the visible light visual features mapped to the shared semantic space. Infrared visual features mapped to a shared semantic space visible light text semantic features and infrared light text semantic features The similarity between them yields the semantic distribution consistency, and its formula is: . The first term is used to constrain the similarity between the two modalities and their respective texts on the same scale, while the second term is used to constrain cross-modal consistency.

[0053] In this embodiment, the visible light visual features and infrared visual features obtained by the dual-branch visual feature encoding module are first mapped to a shared semantic space consistent with the text semantic features. Specifically, the visual features of the two modalities are mapped through a linear transformation to ensure that their semantic expression is consistent with the text embedding space, thereby providing a basic representation for subsequent semantic alignment.

[0054] Secondly, based on the visual features mapped to the shared semantic space, the similarity between visible light visual features and corresponding textual semantic features, as well as the similarity between infrared visual features and corresponding textual semantic features, are calculated to characterize the degree of consistency between visual features and textual semantics. In one specific implementation, the similarity can be calculated using cosine similarity. Furthermore, based on the above similarity results, explicit semantic consistency constraints are constructed to enhance the semantic correspondence between visual features and corresponding textual semantic features. By applying the explicit semantic consistency constraints to the visible light modality and the infrared modality respectively, the visual features of different modalities gradually converge towards their respective textual semantic features in the shared semantic space, thereby improving the semantic consistency between visual features and textual semantics.

[0055] Finally, guided by textual semantic features, the semantic distribution consistency between visible light visual features and infrared visual features is constrained to construct an implicit semantic alignment mechanism. Specifically, such as... Figure 3 As shown, explicit semantic alignment and implicit semantic alignment constrain bimodal visual features in the shared semantic space in different ways. In explicit semantic alignment, similarity constraints are applied to the visible light and infrared visual features mapped to the shared semantic space and their corresponding textual semantic features, ensuring that the visual features maintain semantic consistency with the corresponding textual semantics, thereby strengthening the one-to-one correspondence between visual features and textual semantics. In implicit semantic alignment, the consistency of semantic distribution between visible light and infrared visual features is further constrained. By simultaneously considering both intramodal visual-text relationships and cross-modal visual-text relationships, the overall semantic distribution of the two modalities' visual features in the shared semantic space tends to be consistent.

[0056] This step, by simultaneously constraining the similarity relationships between different modal visual features and their respective textual semantic features, as well as the similarity relationships between cross-modal visual features and heteromodal textual semantic features, ensures that the semantic distribution of visual features from both modalities remains consistent in the shared semantic space, thereby reducing cross-modal semantic bias and promoting the collaborative alignment of bimodal visual features.

[0057] S5. Based on the visible light visual features after bimodal semantic alignment, the infrared visual features after bimodal semantic alignment, and the unified text semantic prior, semantic modulation fusion processing is performed to generate fusion features. The fusion features are then input into a preset decoder to obtain the final fusion image.

[0058] Specifically, step S5 further includes: spatially replicating the unified text semantic prior. Extending to the same spatial dimension as the visual features after bimodal semantic alignment, the formula is: , For information of global semantic concern, This means extending the original textual semantic priors to the shared space dimension;

[0059] MLP linear mapping is performed on the visible light visual features and infrared visual features after bimodal semantic alignment to generate visible light adaptive weights. and infrared adaptive weights , It is a linear mapping function;

[0060] Adaptive weighting of visible light and infrared adaptive weights As a modulation factor, the visual features of the corresponding modality are subjected to semantically guided gating and residual compensation processing (i.e., the visible light modality is added for compensation) to form the features of the final fused image, and the fused features are jointly constructed by combining global semantic attention information. , This is element-wise multiplication;

[0061] The decoder combines infrared adaptive weights as modulation factors with global semantic information; infrared features form the dominant representation under the modulation of textual semantics, while visible light features serve as compensation terms to enhance global structural and textural information. The fused features are then input... The image is fed into a preset decoder to generate the final fused image. , It is a decoder that can decode the fused features into the final fused image.

[0062] In this embodiment, during the fusion stage, the unified text semantic prior is extended to the spatial dimension and used as a global semantic modulation signal to perform weighted fusion of semantically aligned infrared features and visible light features, so as to enhance the semantic consistency and structural integrity in the fusion result.

[0063] Specifically, firstly, the textual semantic prior is extended to a spatial dimension consistent with visual features, serving as a global semantic modulation signal to provide unified semantic guidance for the subsequent visual feature fusion process. This approach enables textual semantic information to participate in the modulation process of visual features at the spatial level, thereby enhancing the semantic consistency of the fusion result. Secondly, based on the semantically aligned infrared and visible light visual features, corresponding modality adaptive weights are generated. In one specific implementation, adaptive weights are generated for different modalities using a linear mapping method to characterize the contribution of each modality in the fusion process, thus achieving adaptive adjustment of different modal information. Thirdly, the adaptive weights of the infrared modality are used as modulation factors and combined with the global textual semantic modulation signal to construct a dominant fusion feature representation. Visible light modal features are then introduced as a compensation term to enhance the structural and textural information in the fusion result, forming the final fusion feature. Finally, the fusion feature is input into the decoding network to generate the final fused image.

[0064] Specifically, such as Figure 2 As shown, textual semantic priors participate in the fusion process as global semantic modulation signals, acting on the modality adaptive weight generation modules for both infrared and visible light visual features. During fusion, the modality adaptive weights generated based on infrared visual features modulate the infrared features under the guidance of the textual semantic priors, making the infrared modality the dominant representation in target saliency. Simultaneously, the modality adaptive weights generated based on visible light visual features modulate the visible light features to supplement the structural and texture information in the fused features. Subsequently, the modulated infrared and visible light features are weighted and fused, and then input into the decoding network to generate the final fused image, thereby enhancing visible light structural details while maintaining infrared target saliency. Through this method, the infrared modality plays a dominant role in target saliency representation, while the visible light modality supplements scene structure and detail information, thus improving the overall quality of the fused image.

[0065] Preferably, in this embodiment, an infrared intensity preservation constraint is introduced during the training process to reduce the loss of infrared energy distribution in the fused image, and the formula is as follows: ,in, For infrared energy distribution loss, The original infrared image, The value is the L1 norm. This formula assesses infrared energy loss by calculating the difference between the original infrared image and the fused image.

[0066] In this embodiment, to ensure that the final fused image fully incorporates visible light details without sacrificing the core detection capability of the infrared image for thermal targets—namely, its unique thermal radiation intensity distribution information—this invention introduces a special constraint mechanism during the overall training and optimization process of the model: the infrared intensity preservation constraint. This constraint does not act on the forward propagation network structure, but rather serves as an important component of the training objective function, guiding the optimization direction.

[0067] Specifically, in each iteration of the model training process, after the network generates the final fused image based on all the aforementioned steps (including textual semantic guidance, bi-branch encoding, semantic alignment, and fusion decoding), an additional loss term needs to be calculated to evaluate and constrain the fidelity of the fused image in the infrared energy distribution. The implementation involves calculating the difference between the fused image and the input original infrared image. Here, the L1 norm (i.e., the sum of absolute errors) is used as a measure of difference because it is less sensitive to outliers and can promote the generated image to more closely resemble the infrared source image in pixel intensity.

[0068] The core benefit of introducing this constraint lies in providing a crucial physical information conservation guide for the entire end-to-end fusion network training. In multimodal fusion tasks, when network models strive to integrate visible light textures and structures, they may unintentionally oversmooth or weaken key intensity regions representing thermal targets in infrared images, leading to a decrease in the saliency of thermal targets. By explicitly incorporating the loss into the total loss function (e.g., a weighted sum with other loss terms such as text alignment loss and reconstruction loss), the training process is forced to simultaneously optimize two objectives: achieving high-quality cross-modal semantic and texture fusion, and maximizing the preservation of the energy distribution characteristics of the input infrared image. This is equivalent to injecting a kind of memory into the model, making it aware that infrared intensity information is a valuable input that cannot be discarded.

[0069] In summary, this invention provides an innovative overall solution for infrared and visible light image fusion guided by textual semantic consistency, encompassing methods, systems, devices, and storage media. The core concept of this solution lies in overcoming the limitations of existing technologies that rely solely on visual features for fusion. It pioneers the introduction of fine-grained textual semantics as a global high-level guide, fundamentally improving the semantic consistency and information integrity of the fused images by constructing a dual-path alignment and fusion framework that facilitates "text-visual" collaboration.

[0070] Specifically, this invention first transforms input infrared and visible light images into fine-grained textual semantic features within a unified embedding space through the collaboration of a large language model and a visual language model (such as BLIP). It then innovatively designs a cross-modal bidirectional attention compensation mechanism to interactively enhance textual semantics, thereby generating a strengthened and unified textual semantic prior. This step integrates high-level scene understanding with target semantic injection, providing a priori knowledge foundation for bridging the semantic gap between different modalities. Next, a structure-intensity decoupled dual-branch encoding strategy is employed to extract the rich structural texture features of visible light images and the unique intensity saliency features of infrared images. This decoupling design respects the information complementarity between modalities, preserving a clean information flow for subsequent accurate fusion. In the feature processing stage, the key innovation of this invention lies in proposing an explicit and implicit dual semantic alignment mechanism. Guided by a unified textual semantic prior, explicit loss constrains the consistency between visual features and their own textual semantics, while implicit loss constrains the consistency of the semantic distribution of bimodal visual features, thus achieving deep collaboration and convergence of infrared and visible light visual features in a shared semantic space. Finally, the textual semantic prior is used as a global modulation signal to dynamically generate modality adaptive weights. These weights are then used to adaptively weight and fuse the aligned bimodal features, and the final fused image is generated through a decoding network. Furthermore, the infrared intensity preservation constraint introduced during training effectively ensures the preservation of thermal target saliency information in the fused image.

[0071] In summary, this invention achieves the following comprehensive benefits through the aforementioned interconnected technical chain: 1. Precise semantic guidance: Utilizing fine-grained text to guide the entire process from high-level semantics to low-level pixels significantly improves the semantic coherence and target-background differentiation of the fusion results; 2. Deeper feature alignment: Through dual visual-text alignment constraints, effective fusion of heterogeneous features in a unified semantic space is ensured, avoiding information conflicts and distortions; 3. Comprehensive information preservation: The combined effect of structure-intensity decoupling coding and infrared intensity constraints enables the fused image to perfectly maintain the salience and energy distribution of infrared targets while enhancing visible light details; 4. Robust scene adaptation: The entire method does not rely on manual rules and can adapt to various complex scenes (such as low illumination and complex backgrounds), generating fused images with high visual quality and strong task support capabilities, greatly promoting the practical application of multimodal image fusion technology in fields such as security monitoring, autonomous driving, and remote sensing analysis.

[0072] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for fusing infrared and visible light images based on text semantic consistency guidance, characterized in that, include: Visible light and infrared images are acquired, and fine-grained text semantic generation processing is performed on the visible light and infrared images to obtain a text description set. The text description set is then mapped to a unified text embedding space to generate multimodal text semantic features. Cross-modal semantic relationship modeling is performed on the multimodal text semantic features, specifically modeling the semantic interaction relationship between visible light text and infrared text to compensate for and enhance the semantics of visible light text and infrared text, thus constructing a unified text semantic prior. Based on the semantic features of visible light text and infrared light text semantic features The attention weights for the infrared and visible light modalities are calculated using the following formula: , ,in, Infrared semantic attention-weighted representation is aggregated for visible light text. It is an exponential function. , , , , , All are learnable linear mapping matrices. For embedded dimensions, The scaling factor. Aggregate visible light semantic attention-weighted representations for infrared text; Based on the infrared semantic attention weighted representation of visible light text aggregation Infrared text aggregation with visible light semantic attention weighted representation The semantic consistency compensation signal is generated in both visible and infrared modes, and its calculation formula is as follows: , ,in, , All are learnable linear mapping coefficients. This is a visible light compensation signal. For infrared compensation signals; Compensation for visible light signals Infrared compensation signal visible light text semantic features and infrared light text semantic features By adding them together, we obtain the semantically enhanced visible light text semantics. and semantically enhanced infrared text semantics ; Semantic enhancement of visible light text and semantically enhanced infrared text semantics By performing fusion processing, a unified textual semantic prior is obtained. , It is a lightweight mapping function; Based on the structure-intensity decoupling dual-branch coding strategy, visual feature coding is performed on visible light images and infrared images respectively. Specifically, visible light visual features, mainly composed of structure and texture information, are extracted from visible light images, while infrared visual features, mainly composed of intensity distribution and thermal target response, are extracted from infrared images. Guided by the unified text semantic prior, the visible light visual features and infrared visual features are subjected to bimodal semantic alignment processing. In this process, through explicit semantic consistency constraints and implicit structural consistency constraints, the visible light visual features and infrared visual features gradually converge in the unified semantic space. Based on the visible light visual features after bimodal semantic alignment, the infrared visual features after bimodal semantic alignment, and the unified text semantic prior, semantic modulation fusion processing is performed to generate fused features. The fused features are then input into a preset decoder to obtain the final fused image.

2. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 1, characterized in that, Visible light and infrared images are acquired, and fine-grained text semantic generation processing is performed on the visible light and infrared images to obtain a text description set. This text description set is then mapped to a unified text embedding space to generate multimodal text semantic features, specifically: Visible light images and infrared images are acquired as multimodal inputs to be fused. The visible light images are used to provide detailed information, including scene structure, texture, and edges. The infrared images are used to provide target thermal radiation intensity distribution and salient target information. The visible light and infrared images are input into a large language model to obtain a fine-grained visible light text representation. and infrared light text description And based on visible light text representation and infrared light text description Constitutes a text description set; The BLIP visual language encoder is used to map the text description set to a preset unified text embedding space to generate visible light text semantic features. and infrared light text semantic features And based on the semantic features of visible light text and infrared light text semantic features This yields multimodal text semantic features, among which, This is a mapping function.

3. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 2, characterized in that, Based on a structure-intensity decoupling dual-branch coding strategy, visual feature encoding is performed on visible light images and infrared images separately. Specifically, visible light visual features, primarily composed of structural and texture information, are extracted from the visible light images, while infrared visual features, primarily composed of intensity distribution and thermal target response, are extracted from the infrared images. The structure-intensity decoupled dual-branch coding strategy employs a visible light branch and an infrared branch. The visible light branch focuses on extracting scene structure, edge, and texture information through multi-scale convolution and spatial attention mechanisms, while the infrared branch focuses on extracting the intensity distribution of thermal targets and information on salient regions. The visible light image is input into the visible light branch to extract texture and structural features at different scales. The visible light branch consists of three convolutional layers with different scales. , , The output features of three convolutional layers with different convolutional scales were obtained. , , ; The output features of three convolutional layers with different convolutional scales , , After addition, the results are input into a preset spatial attention module to obtain visible light visual features. ; The infrared image is input into the infrared branch to extract the intensity distribution and thermal response characteristics of the infrared image, thus obtaining infrared visual features. The infrared branch consists of three identical infrared feature extraction convolutional modules and channel attention modules, with each infrared feature extraction convolutional module containing a... A convolutional layer with a kernel size of [size missing] and a ReLU function.

4. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 3, characterized in that, The bimodal semantic alignment processing includes two mechanisms: explicit semantic alignment and implicit semantic alignment. Explicit semantic alignment is achieved by constraining the similarity between visual features and corresponding text semantic features in a shared semantic space, thereby constraining the consistency between visual features and corresponding text semantic features. Implicit semantic alignment is achieved by constraining the consistency of the semantic distribution of visible light visual features and infrared visual features under text guidance in a shared semantic space, thereby constraining the consistency of the semantic distribution of different modal visual features under text guidance.

5. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 4, characterized in that, Guided by the unified textual semantic prior, bimodal semantic alignment processing is performed on visible light visual features and infrared visual features. Specifically, through explicit semantic consistency constraints and implicit structural consistency constraints, the visible light visual features and infrared visual features gradually converge in the unified semantic space. Visible light visual features and infrared visual features Mapped to a predefined shared semantic space, the formula is: , , , All are learnable linear transformation coefficients. For visible light visual features mapped to a shared semantic space, Infrared visual features mapped to a shared semantic space; In explicit semantic alignment, an explicit loss function is computed. To achieve the similarity between constrained visual features and corresponding text semantic features, specifically: Based on visible light visual features mapped to a shared semantic space Infrared visual features mapped to a shared semantic space visible light text semantic features and infrared light text semantic features The formula for calculating similarity is: , , The similarity between visible light visual features and text semantic features. The similarity between infrared visual features and text semantic features. This is the function for calculating cosine similarity. Based on the similarity between visible light visual features and text semantic features Similarity between infrared visual features and text semantic features The formula for constructing explicit semantic consistency constraints is as follows: .

6. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 5, characterized in that, Also includes: In implicit semantic alignment, an implicit loss function is computed. To achieve semantic consistency between constrained visible light visual features and infrared visual features under text guidance, the specific steps are as follows: Compute visible light visual features mapped to a shared semantic space Infrared visual features mapped to a shared semantic space visible light text semantic features and infrared light text semantic features The similarity between them yields the semantic distribution consistency, and its formula is: .

7. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 6, characterized in that, Based on the visible light visual features after bimodal semantic alignment, the infrared visual features after bimodal semantic alignment, and the unified text semantic prior, semantic modulation fusion processing is performed to generate fused features. These fused features are then input into a preset decoder to obtain the final fused image. Specifically: The unified text semantic prior is obtained through spatial replication. Extending to the same spatial dimension as the visual features after bimodal semantic alignment, the formula is: , For information of global semantic concern, This means extending the original textual semantic priors to the shared space dimension; MLP linear mapping is performed on the visible light visual features and infrared visual features after bimodal semantic alignment to generate visible light adaptive weights. and infrared adaptive weights , It is a linear mapping function; Adaptive weighting of visible light and infrared adaptive weights As a modulation factor, semantically guided gating and residual compensation are applied to the visual features of the corresponding modality, and fused features are constructed together with global semantic attention information. , This is element-wise multiplication; Input the fused features The image is fed into a preset decoder to generate the final fused image. , For decoders.

8. The infrared and visible light image fusion method based on text semantic consistency guidance according to claim 7, characterized in that, During training, an infrared intensity preservation constraint is introduced to reduce the loss of infrared energy distribution in the fused image. The formula is as follows: ,in, For infrared energy distribution loss, The original infrared image, It is an L1 norm.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on multi-semantic deep collaboration

    CN121214142A