A text-guided region enhanced infrared and visible image fusion method

CN122736885APending Publication Date: 2026-09-11ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610888077.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0005]针对现有技术中文本与图像局部区域缺乏显式关联、小目标区域响应较弱以及目标区域响应不稳定的技术问题,本发明提出一种基于文本引导的区域增强红外与可见光图像融合方法,记为R2T-Fusion(Region Response and Text-Residual Control Fusion),首先构建文本感知区域预测网络,以可见光图像、红外图像和文本特征为输入,生成与文本语义相关的区域响应图;随后采用基于Restormer的TransformerBlock的双分支特征提取骨干,并在图像内容驱动的基础融合得分上叠加文本残差得分,实现文本相关区域的差异化增强;进一步引入真值引导的渐进式掩码退火策略和区域面积归一化的目标—背景联合损失,以提高训练稳定性并增强模型对文本相关区域的响应能力

Benefits of technology

[0052]To address the lack of explicit correlation between natural language text and local image regions in existing text-guided image fusion, a text-aware region prediction network is proposed. By jointly modeling infrared images, visible light images, and text features, a region response map related to text semantics is generated, providing semantic prior support for subsequent text-guided fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736885A_ABST
    Figure CN122736885A_ABST
Patent Text Reader

Abstract

This invention proposes a text-guided region-enhanced infrared and visible light image fusion method. A multimodal feature extraction module extracts infrared image features, visible light image features, and textual description features respectively. A text-aware region prediction module models the joint features of the infrared and visible light images, sequentially performing channel-level text modulation and spatial gating weighting on the dual-modal joint features to output a continuous region response map. An enhanced joint feature is constructed through a feature fusion module. Guided by the continuous region response map and textual features, a basic fusion score is calculated based on the enhanced joint feature, and a text residual score based on the textual description features is superimposed to obtain the fused feature. A decoding and reconstruction module processes the fused feature to generate a fused image. During training, a target-background joint loss based on region area normalization and a progressive mask annealing strategy are introduced. This invention preserves the structure and detail information of the fused image while enabling more targeted enhancement of text-related regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image fusion technology, and more particularly to a method for fusion of infrared and visible light images. Background Technology

[0002] Infrared and Visible Image Fusion (IVIF) aims to combine the thermal radiation response of infrared imaging with the texture detail information of visible light imaging to generate a fused image that combines target saliency with background structure information. Infrared images have strong target perception capabilities under low light, complex weather, and occlusion conditions, while visible light images have advantages in texture detail, edge structure, and natural appearance representation. Therefore, IVIF has significant application value in scenarios such as intelligent surveillance, target detection, autonomous driving, and night vision. In recent years, with the development of deep learning, infrared and visible light image fusion methods have made significant progress in cross-modal feature modeling, detail preservation, and visual quality improvement.

[0003] Among these emerging cutting-edge methods, representative works such as ITFuse, SePT, and AITFuse have significantly improved fusion quality through feature interaction and long-distance dependency modeling. However, these methods are essentially still undifferentiated fusions of the entire image, making it difficult to achieve targeted fusion of specific target regions. To address this limitation, recent research has introduced semantic priors into the image fusion process, gradually shifting the fusion model from "globally undifferentiated fusion" to "semantically guided differentiated fusion," thereby enhancing its ability to perceive user intent. For example, a series of works such as HitFusion, SDSFusion, MDPPFuse, and CSDFusion have effectively integrated semantic features into the network through mechanisms such as Transformer interaction, high-level semantic constraints, or knowledge distillation. However, while the above methods improve the semantic perception ability in the fusion process, their semantic information mostly relies on pre-defined target categories or fixed labels, making it difficult to directly understand open-ended user intent in natural language text form.

[0004] To make fusion strategies more flexible, natural language text-guided image fusion methods are gradually becoming a new research direction. Visual-language pre-trained models, represented by CLIP, provide new technical support for image-text semantic alignment. TextFusion [CHENG C, XU T, WU XJ, et al. TextFusion: Unveiling the power of textual semantics for controllable image fusion[J]. Information Fusion, 2025,117:102790] and CLIPFusion [SUN D, WANG C, WANG T, et al. CLIPFusion: Infrared and visible image fusion network based on image-text large model and adaptive learning[J]. Displays, 2025, 89:103042.] and other works attempt to use text descriptions or visual-language models to regulate the fusion process, promoting IVIF from traditional global fusion to semantically controllable fusion. For example, the invention patent with publication number CN120765479A discloses an infrared and visible light image fusion method and system based on text-guided semantic perception. Through a text-guided semantic perception method, it uses independent encoders and cross-modal models to extract features and semantically align infrared and visible light images, solving the problem of poor image fusion effect in the prior art. However, existing methods still face two major bottlenecks in complex scenarios: First, there is a lack of accurate mapping mechanism between text semantics and local image regions, making it impossible to establish an explicit association between natural language and local image regions, which makes it difficult to accurately locate abstract text instructions to specific targets; Second, the feature responses of small targets are easily submerged by large areas of background information, resulting in insufficient model discrimination of different text instructions; Third, the transition between ground truth and predicted mask is not smooth during training, leading to unstable responses in text-related regions. Summary of the Invention

[0005] To address the technical problems of existing technologies, such as the lack of explicit correlation between text and local image regions, weak response of small target regions, and unstable response of target regions, this invention proposes a text-guided region enhancement method for infrared and visible light image fusion, denoted as R2T-Fusion (Region Response and Text-Residual Control Fusion). First, a text-aware region prediction network is constructed, using visible light images, infrared images, and text features as input to generate region response maps related to text semantics. Then, a dual-branch feature extraction backbone based on TransformerBlock and Restormer is used, and text residual scores are superimposed on the image content-driven basic fusion score to achieve differentiated enhancement of text-related regions. Furthermore, a ground-value-guided progressive mask annealing strategy and a target-background joint loss with region area normalization are introduced to improve training stability and enhance the model's response capability to text-related regions. Experimental results show that this method can more effectively enhance text-related regions while preserving the structural and detailed information of the fused image.

[0006] To achieve the above objectives, the technical solution of the present invention is implemented as follows:

[0007] A text-guided method for region enhancement and fusion of infrared and visible light images, comprising the following steps:

[0008] S1: Obtain a dataset containing paired infrared images, visible light images, and text descriptions;

[0009] S2: Construct an R2T-Fusion model. A multimodal feature extraction module extracts infrared image features, visible light image features, and text description features respectively. A text-aware region prediction module models the dual-modal joint features of the infrared and visible light images. Based on the text description features, the dual-modal joint features are sequentially subjected to channel-level text modulation and spatial gating weighting, outputting a continuous region response map corresponding to the text description semantics. A feature fusion module enhances the infrared and visible light image features and constructs enhanced joint features. Guided by the continuous region response map and text features, a basic fusion score is calculated based on the enhanced joint features, and the text residual score based on the text description features is superimposed to obtain the fused features. A decoding and reconstruction module processes the fused features to generate a fused image.

[0010] S3: The R2T-Fusion model is trained using the dataset and based on the target-background joint loss normalized by the region area. During training, the fusion mask of the region ground truth mask and the continuous region response map is calculated using the progressive mask annealing strategy. The fusion mask is used as the region response map input to the feature fusion module.

[0011] S4: Obtain the infrared image, visible light image, and required text description to be fused, and use the trained R2T-Fusion model to achieve the fusion of infrared and visible light images.

[0012] Furthermore, infrared image features, visible light image features, and text description features are extracted separately using a multimodal feature extraction module, including:

[0013] Feature extraction from infrared and visible light images: Two independent convolutional modules are used to extract features from the visible light image. and infrared images Perform shallow feature mapping to obtain shallow features of the visible light image. and shallow features of infrared images Two Restormer modules are used to process shallow features of visible light images. and shallow features of infrared images Perform feature encoding to obtain infrared image features and visible light image features ;

[0014] Text description feature extraction: A pre-trained CLIP text encoder is used to extract features from the input text description. The text description features are encoded and normalized to obtain the text description features. .

[0015] Furthermore, based on the text description features, the dual-modal joint features are sequentially subjected to channel-level text modulation and spatial gating weighting, outputting a continuous region response map corresponding to the semantics of the text description, including:

[0016] a1. Regarding infrared images and visible light images Dual-modal joint features spliced ​​along the channel dimension Proceed to the first step in sequence The initial features are obtained by using the first ReLU activation function. ;

[0017] a2. Textual description features Perform the first linear mapping to generate channel-level FiLM scaling factors. and FiLM bias coefficient ; Using FiLM scaling factor and FiLM bias coefficient For initial features Channel-level text modulation features are obtained by performing FiLM text modulation. , Indicates the strength coefficient;

[0018] a3. Textual description features The second linear mapping, the second ReLU activation function, and the third linear mapping are performed sequentially to generate a channel-dimensional text modulation vector. , channel dimension text modulation vector Text modulation features After element-wise multiplication, spatial gated convolution is performed. Generate a text space gating graph using the first Sigmoid function. ; Gated graphs of text space After adding 1 to the whole, the text modulation features Perform element-wise multiplication, then pass through the second... The third ReLU activation function, the third Generate unnormalized regional response score maps The continuous regional response map is generated by normalizing the regional response score map Z using the second Sigmoid function. .

[0019] Furthermore, the infrared image features and visible light image features are enhanced through a feature fusion module to construct enhanced joint features. Guided by continuous region response maps and text features, a basic fusion score is calculated based on the enhanced joint features, and the text residual score based on text description features is superimposed to obtain the fused features, including:

[0020] b1. Utilize the feature enhancement module to enhance infrared image features and visible light image features Infrared image enhancement features are obtained by performing feature enhancement operations based on text feature mapping. and visible light image enhancement features ;

[0021] b2. Using a joint feature construction module based on infrared image enhancement features and visible light image enhancement features Constructing enhanced joint features , This indicates a channel splicing operation;

[0022] b3. Enhance joint features using a dual-branch weight prediction module. The basic fusion score and text residual score are obtained by performing continuous region response map injection for basic fusion and text residual control respectively. Then, the fusion weight is calculated under the guidance of the continuous region response map, and the fusion weight is used to analyze the infrared image features. and visible light image features Weighted fusion is performed to obtain fusion features. .

[0023] Furthermore, the feature enhancement operation based on text feature mapping in step b1 includes:

[0024] b11. Using the fourth and fifth linear mappings to describe text features Mapping these maps to the channel spaces of the visible light branch and the infrared branch respectively, yielding visible light branch mapping results and infrared branch mapping results. These results are then input into the first... Function and Second Adding 1 to each function result in a text channel control vector used for visible light image features. Text channel control vectors of infrared image features ;

[0025] b12, Set the text channel control vector Features of infrared images Perform element-wise multiplication and pass through the fourth Infrared attention response map generated by the third Sigmoid function. ;Text channel control vector Features of visible light images Perform element-wise multiplication, and then pass through the fifth... Generate visible light attention response map using the fourth Sigmoid function. Using infrared attention response maps Features of infrared images Infrared image enhancement features are obtained by element-wise multiplication. Using visible light attention response maps Features of visible light images Element-wise multiplication yields visible light image enhancement features .

[0026] Furthermore, in step b3, the basic fusion of continuous region response map injection is performed, including: through the sixth Enhanced joint features Perform feature reduction to obtain the reduced features. ; Continuous region response map By adjusting the coefficient After adjustment and dimensionality reduction features Perform channel splicing, and go through the seventh... The operation implements mask injection, and the mask injection features are obtained. Injecting features into the mask To carry out the first After depthwise convolution, it is processed by the fourth ReLU activation function and the eighth And utilize learnable scaling parameters Scaling is performed to obtain the base fusion score. ;

[0027] In step b3, text residual control is performed, including: through the ninth... Enhanced joint features Perform feature reduction to obtain the reduced features. Textual descriptive features are derived through the sixth linear mapping. Mapped to FiLM scaling factor and FiLM bias coefficient Using scaling FiLM coefficients and FiLM bias coefficient Dimensionality reduction features Perform text FiLM modulation to obtain dimensionality-reduced features after text modulation. , Indicator intensity coefficient; dimensionality reduction features after text modulation. To proceed to the second After depthwise convolution, it is processed by the fifth ReLU activation function and the tenth ReLU activation function. Obtain text residual scores .

[0028] Furthermore, in step b3, the fusion weights are calculated under the guidance of the continuous region response map, and the fusion weights are used to analyze the infrared image features. and visible light image features Weighted fusion includes:

[0029] Based on the continuous region response map Construction region guiding coefficient This is used to adjust the influence of text residual scores in different regions. Indicates the base offset parameter;

[0030] Using regional guiding coefficients Text residual score Element-wise multiplication and using learnable scaling factors After scaling and using weighting coefficients Weighted base fusion score Add them together to get the fusion score. The fusion score is obtained by using the fifth Sigmoid function. Mapped to a fusion weight graph ;

[0031] Using fusion weight graph Infrared image features and visible light image features Perform pixel-by-pixel weighted fusion to obtain fused features. .

[0032] Furthermore, the fused image is generated by processing the fused features through the decoding and reconstruction module, including: using the seventh linear mapping to convert the text description features... Mapped to FiLM scaling factor and FiLM bias coefficient Using FiLM scaling factor and FiLM bias coefficient Fusion features Text modulation is performed to obtain the fused features after text modulation. , Indicates the strength coefficient, passed through the eleventh... The normalization operation modulates the fused features of the text. Convert to final fused image .

[0033] Furthermore, the progressive masked annealing strategy is as follows:

[0034] ;

[0035] in, Indicates the fusion mask, For the region truth mask, This is the annealing factor, which is set to increase sequentially with each training phase. This indicates post-processing operations, including changing the response value. Crop to Obtain the cropped mask image within the range ; for mask image Perform a power transform to obtain the power-transformed mask image. From the mask image Responses with confidence levels below a preset threshold are filtered out, while responses with confidence levels greater than or equal to the preset threshold are retained.

[0036] The joint loss for area normalization includes:

[0037] ;

[0038] in, This represents the joint loss normalized to area. Indicates the region image loss. Indicates regional monitoring losses, Indicates the semantic alignment loss of the region. Indicates cross-text difference loss. These are the weighting coefficients for each loss term.

[0039] Furthermore, the region image loss is expressed as:

[0040] ;

[0041] ;

[0042] ;

[0043] in, Loss in the target area For background area loss, For reference map of the target area, For single-channel infrared images Representation after copying to three channels. , representing the gradient reference of the target region, and These represent the weighting coefficients for the visible light gradient and the infrared gradient, respectively. This represents the image edge difference operator. Represents the target region weight map. Represents the background region weight map. Indicates the effective area of ​​the target region Indicates the effective area of ​​the background region. To prevent tiny constants with a denominator of zero, Represents the square of the L2 norm. Represents the L1 norm;

[0044] The regional monitoring loss adopts the joint monitoring loss of BCE loss and Dice loss;

[0045] The region semantic alignment loss is:

[0046]

[0047] in, and These represent CLIP image encoder and text encoder, respectively.

[0048] The cross-text difference loss is:

[0049]

[0050] in, This indicates the number of text elements corresponding to the same image. and This represents the predicted region response maps generated for the same image under different text inputs. This represents the minimum difference interval.

[0051] The beneficial effects of this invention are as follows:

[0052] To address the lack of explicit correlation between natural language text and local image regions in existing text-guided image fusion, a text-aware region prediction network is proposed. By jointly modeling infrared images, visible light images, and text features, a region response map related to text semantics is generated, providing semantic prior support for subsequent text-guided fusion.

[0053] To address the issue of weak response in target regions under text conditions, a fusion mechanism based on text residual control is designed. Text-related residual scores are superimposed on the image-driven basic fusion scores, and text semantic injection is retained during the decoding stage, thereby enhancing the model's ability to differentiate fusion of text-related regions.

[0054] To address the issue of unstable responses in text-related regions due to the uneven transition between ground truth and predicted mask during training, a ground truth-guided progressive mask annealing training strategy is proposed. A target-background joint loss with normalized region area is constructed to ensure stable training while simultaneously enhancing the target region and preserving the background structure. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is a schematic diagram of the network structure corresponding to a text-guided region-enhanced infrared and visible light image fusion method of the present invention.

[0057] Figure 2 This is a schematic diagram of the text-aware region prediction and progressive mask annealing strategy of the present invention.

[0058] Figure 3 This is a structural diagram of the fusion weight prediction module for text residual control in the feature fusion module of the present invention.

[0059] Figure 4 This is a comparison of the fusion results of the present invention and different methods in typical nighttime scenes of the IVT dataset.

[0060] Figure 5 This is a comparison of the fusion results of the present invention and different methods in a generalized scenario. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] This invention proposes a text-guided region-enhanced infrared and visible light image fusion method, the network structure of which is shown in the figure below. Figure 1 As shown, the model consists of four core modules: a multimodal feature extraction module, a text-aware region prediction module, a text residual-controlled feature fusion module, and a decoding and reconstruction module. First, the model uses a dual-branch visual encoder to extract visible light images separately. With infrared images Features , The pre-trained CLIP text encoder is used to obtain text features. Secondly, the text-aware region prediction module is used to predict based on... , and Output region response score map The region response map is obtained by mapping with the Sigmoid function. Based on this, during the training phase, further... With region truth mask The mixture is then fused to obtain a mask that is actually used for predicting the fusion weights. Next, Inject fusion weight prediction process to generate adaptive fusion weights. and utilize The fusion ratio of visible light features and infrared features at different spatial locations is dynamically adjusted to complete feature fusion; finally, the final fused image is obtained through a decoder. It should be noted that the region truth mask... It is used only during the training phase for region supervision and progressive masked annealing, and is not used as input for the inference phase. During the inference phase, the model only receives visible light images, infrared images, and text descriptions, and generates response maps through the text-aware region prediction module. And it is used as a regional prior in the prediction of fusion weights.

[0063] A text-guided method for region enhancement and fusion of infrared and visible light images, such as Figure 1 As shown, the steps are as follows:

[0064] S1: Obtain a dataset containing paired infrared images, visible light images, and text descriptions.

[0065] In this embodiment, the IVT dataset [CHENG C, XU T, WU XJ, et al. TextFusion: Unveiling the power of textual semantics for controllable imagefusion[J]. Information Fusion, 2025, 117:102790] is used as the main training and testing data source. The IVT dataset is constructed for text-guided infrared-visible image fusion tasks and includes paired infrared images, visible images, text descriptions, and their corresponding region of interest (ROI) annotations. According to the training process of the model according to the present invention, the data is uniformly organized into four types of files: ir, vis, text, and association, and it is ensured that the image number, text number, and region annotation file correspond one-to-one. The organized data contains a total of 2000 pairs of infrared-visible image samples, each pair of images is accompanied by 5 text descriptions and their corresponding ROI annotations. During the training phase, each pair of images participates in optimization under different texts to enhance the model's responsiveness to text semantic regions.

[0066] S2: Construct the R2T-Fusion model, such as Figure 1 As shown, the multimodal feature extraction module extracts infrared image features, visible light image features, and text description features respectively; the text-aware region prediction module performs dual-modal joint feature modeling on the infrared and visible light images, and performs channel-level text modulation and spatial gating weighting on the dual-modal joint features based on the text description features, outputting a continuous region response map corresponding to the text description semantics; the feature fusion module obtains the fused features by superimposing the basic fusion scores of the infrared and visible light image features with the text residual scores based on the text description features under the guidance of the continuous region response map; and the decoding and reconstruction module processes the fused features to generate a fused image.

[0067] In this embodiment of the application, infrared image features, visible light image features, and text description features are extracted respectively through a multimodal feature extraction module, including:

[0068] Feature extraction from infrared and visible light images:

[0069] First, two independent convolutional modules are used to process the visible light image separately. and infrared images Perform shallow feature mapping:

[0070] ;

[0071] ;

[0072] in, This represents a two-dimensional convolution operation. These are shallow features of visible light images. These are shallow features of infrared images.

[0073] Furthermore, a Restormer-style dual-branch feature extraction module is adopted, utilizing two Restormer modules to encode features for the two modalities respectively. For any modality , its first The layer feature update process is represented as follows:

[0074] ;

[0075] ;

[0076] go through After layer stacking, two modalities of deep features are obtained:

[0077] ;

[0078] in, Presentation layer normalization operation, This represents the attention mechanism module (self-attention of Transformer). Indicates a feedforward network. Indicates the first Layer output feature map, Indicates the first Layer attention mechanism module and the first -1 layer output residual connection output.

[0079] Text description feature extraction: A pre-trained CLIP text encoder is used to extract features from the input text description. The text description features are encoded and normalized to obtain the text description features. :

[0080] ;

[0081] in, This represents the text description feature extraction process implemented by the pre-trained CLIP text encoder. This indicates a normalization operation.

[0082] In this embodiment of the application, to establish an explicit association between the text and the target region, thereby endowing the model with local control over differences, a text-aware region prediction module is designed, such as... Figure 2 As shown in (a), this module uses visible light images Infrared images and text features As input, output a region response map related to the current text. .

[0083] In this embodiment, a text-aware region prediction module acquires dual-modal joint features of infrared and visible light images, and performs channel-level text modulation and spatial gating weighting on the dual-modal joint features based on text description features, outputting a continuous region response map corresponding to the text description semantics, including:

[0084] First, the infrared and visible light images are stitched together along the channel dimension to form a dual-modal joint feature. :

[0085] ;

[0086] in, This indicates a channel splicing operation.

[0087] Furthermore, dual-modal joint features Convolution Using ReLU activation function to extract initial features :

[0088] ;

[0089] in, This represents the ReLU activation function.

[0090] Furthermore, Feature-wise Linear Modulation (FiLM) is introduced to perform a first linear mapping on the text description features to generate channel-level modulation parameters. Text modulation features are obtained by using channel-level modulation parameters to perform text modulation on regional features. ,in, This represents the mapping function (linear mapping operation of the linear layer) used to generate FiLM modulation parameters from text features. and These represent the FiLM scaling factor and FiLM bias factor generated from the text description features, respectively. This represents element-wise multiplication; This is the intensity coefficient, used to control the degree to which text features influence the initial features of the image.

[0091] It should be noted that by introducing a residual FiLM modulation mechanism, text features are transformed into channel-level affine modulation parameters and injected into region features in the form of residuals. This mechanism can enhance the region response related to the current text semantics while maintaining the stability of the original features, thereby improving the response variability of the text-aware region prediction module to different text inputs.

[0092] Furthermore, semantic text space mapping (linear mapping, ReLU activation function, linear mapping) is performed on the text description features to generate channel-dimensional text modulation vectors. ; Modulate the channel-dimensional text vector Text modulation features After element-wise multiplication, a spatially gated convolution and a sigmoid function are used to generate a text spatially gated graph. , Represents spatial gated convolution. This is the Sigmoid function.

[0093] Furthermore, the text space gating graph After adding 1 to the whole, the text modulation features Perform element-wise multiplication to preserve the original text modulation features. Basic values, accessed only through text space gating graphs Text modulation features Incremental enhancement is performed to selectively enhance different spatial locations based on text features, i.e., spatial gating weighting, through convolutional mapping (…). ReLU Generate unnormalized regional response score maps The region response score map Z is normalized using the Sigmoid function to generate a continuous region response map. The range of values ​​is .

[0094] In this embodiment of the application, when obtaining visible light image features Infrared image features Textual description features and the actual mask used Subsequently, this invention constructs a fusion weight prediction module for text residual control. This module does not simply concatenate text features with image features, but rather superimposes text residual scores on the image content-driven basic fusion scores, enabling text semantics to directly adjust the fusion ratio of visible light and infrared features.

[0095] In this embodiment, the infrared image features and visible light image features are enhanced by a feature fusion module to construct enhanced joint features. Guided by continuous region response maps and text features, a basic fusion score is calculated based on the enhanced joint features, and the text residual score based on text description features is superimposed to obtain the fused features, such as... Figure 3 As shown, it includes:

[0096] First, the feature enhancement module is used to enhance the features of the infrared image. and visible light image features Infrared image enhancement features are obtained by performing feature enhancement operations based on text feature mapping. and visible light image enhancement features .

[0097] Specifically, feature enhancement operations based on text feature mapping are performed, such as... Figure 3 As shown, it includes:

[0098] First, describe the text features By mapping the channel spaces of the visible light branch and the infrared branch respectively, text channel control vectors for visible light image features are obtained. Text channel control vectors of infrared image features This is used to independently control two modal features:

[0099] ;

[0100] ;

[0101] in, This represents the hyperbolic tangent activation function. and These represent linear mappings from textual description features to the visible light branch and infrared branch channel spaces, respectively.

[0102] Secondly, the text channel control vector Features of infrared images Element-wise multiplication is performed, and an infrared attention response map is generated using a convolutional layer and a sigmoid function. Text channel control vector Features of visible light images Element-wise multiplication is performed, and a visible light attention response map is generated through convolutional layers and a sigmoid function. Using infrared attention response maps Features of infrared images Infrared image enhancement features are obtained by element-wise multiplication. Using visible light attention response maps Features of visible light images Element-wise multiplication yields visible light image enhancement features .

[0103] Furthermore, a joint feature construction module is used to enhance features based on infrared images. and visible light image enhancement features Constructing enhanced joint features .

[0104] Furthermore, a dual-branch weight prediction module is used to enhance joint features. The basic fusion score and text residual score are obtained by performing continuous region response map injection for basic fusion and text residual control respectively. Then, the fusion weight is calculated under the guidance of the continuous region response map, and the fusion weight is used to analyze the infrared image features. and visible light image features Weighted fusion is performed to obtain fusion features. .

[0105] Specifically, the basic fusion of continuous region response map injection is performed, such as... Figure 3 As shown, this includes: enhancing joint features through convolution pairs. Perform feature reduction to obtain the reduced features. ; Continuous region response map (Continuous region response map during training) With region truth mask Fusion mask By adjusting the coefficient After adjustment and dimensionality reduction features The data is stitched together and then subjected to convolutional operations to achieve mask injection, thereby giving the target region more attention in the basic fusion process and obtaining mask injection features. During reasoning: During training: Injecting features into the mask conduct After depthwise convolution, the system undergoes ReLU activation and convolution operations, and learnable scaling parameters are utilized. Scaling is performed to obtain the base fusion score. .

[0106] Specifically, perform text residual control, such as Figure 3 As shown, this includes: enhancing joint features through convolution pairs. Perform feature reduction to obtain the reduced features. Using the modulation parameter mapping function of FiLM modulation, text description features are mapped. Mapped to FiLM scaling factor and FiLM bias coefficient ,Right now , This represents the mapping function that generates FiLM modulation parameters from textual description features. Here, a linear mapping is used, utilizing the FiLM scaling factor. and FiLM bias coefficient Dimensionality reduction features Perform text FiLM modulation to obtain dimensionality-reduced features after text modulation. , Indicator intensity coefficient; dimensionality reduction features after text modulation. conduct The text residual score is obtained after depthwise convolution followed by ReLU activation and convolution operation. .

[0107] Specifically, fusion weights are calculated under the guidance of continuous region response maps, and these fusion weights are used to analyze infrared image features. and visible light image features Perform weighted fusion, such as Figure 3 As shown, it includes:

[0108] First, based on the continuous region response map Construction region guiding coefficient :

[0109] During reasoning

[0110] During training ;

[0111] in, Used to adjust the influence of text residual scores in different regions. The fusion mask calculated for the progressive mask annealing strategy, when a certain position When it is large, The impact of the text residual score is also relatively large; when the value at a certain position is large, the impact is stronger. When smaller, Smaller in scale, but still retaining some textual influence, preventing the textual guidance function from being overshadowed. Complete truncation; in this invention, Set it to 0.4.

[0112] Secondly, utilizing regional guidance coefficients Text residual score Element-wise multiplication and using learnable scaling factors After scaling and using weighting coefficients Weighted base fusion score Add them together to get the fusion score. The fusion score is obtained by using the Sigmoid function. Mapped to a fusion weight graph Its value range is It is used to control the fusion ratio of visible light and infrared features in different regions.

[0113] Finally, using the fusion weight graph Infrared image features and visible light image features Perform pixel-by-pixel weighted fusion to obtain fused features. .

[0114] In this embodiment of the application, the fused image is generated by processing the fusion features through a decoding and reconstruction module, including:

[0115] Using the modulation parameter mapping function of FiLM modulation, text description features are... Mapped to FiLM scaling factor and FiLM bias coefficient ,Right now , This represents the text modulation mapping function in the decoding and reconstruction module. Here, a linear mapping is used, utilizing the FiLM scaling factor. and FiLM bias coefficient Fusion features Text modulation is performed to obtain the fused features after text modulation. , The intensity coefficients are represented by the fused features of the modulated text obtained through convolutional mapping and normalization operations. Convert to final fused image .

[0116] S3: The R2T-Fusion model is trained using the dataset and based on the target-background joint loss normalized by region area. During training, a progressive mask annealing strategy is used to calculate the fusion mask of the ground truth mask of the region and the continuous region response map. The fusion mask is used as the region response map input to the feature fusion module.

[0117] This invention employs a progressive masked annealing strategy for training, such as... Figure 2 As shown in (b), in the early stages of training, more reliance is placed on region ground truth masks. As training progresses, the system gradually transitions to predicting regional response maps. The progressive masked annealing strategy is as follows: Let the first... The annealing coefficient corresponding to each training round is: The mask corresponding to the response map actually used in the feature fusion module is defined as follows:

[0118] ;

[0119] in, This represents the mask that actually participates in the fusion weight prediction after annealing, i.e., the fusion mask. This indicates a post-processing operation, first setting the response value... Crop to Within the range, obtain the cropped mask image. ;Then the mask image Perform a power transform to obtain the power-transformed mask image. After the power transformation, the numerical difference between the target and the background is widened, thus suppressing the weak response region of the background and improving the mask image. Responses with confidence levels below a preset threshold are filtered out, while responses with confidence levels greater than or equal to the preset threshold are retained to obtain the mask image. This is to reduce the interference of background noise on subsequent fusion weight prediction. In the implementation of this invention, Set by stage:

[0120] .

[0121] In this embodiment of the application, to enhance the practical use of the mask To improve the fusion quality of high-response regions while maintaining the stability of the background structure in low-response regions, this invention employs a joint loss with region area normalization to optimize the model. The target region weight map and the background region weight map are defined as follows:

[0122] , ;

[0123] The effective areas of the corresponding target region and background region are defined as follows:

[0124] ;

[0125] in, and These represent the effective areas of the target area and the background area, respectively. To prevent tiny constants with a denominator of zero.

[0126] The total loss of this invention is defined as:

[0127] ;

[0128] in, Indicates the region image loss. Indicates regional monitoring losses, Indicates the semantic alignment loss of the region. Indicates cross-text difference loss. These are the weighting coefficients for each loss term.

[0129] Let the fused image be Visible light image is The representation of a single-channel infrared image after being copied into a three-channel image is denoted as... The target region is defined as follows (reference image):

[0130] ;

[0131] in, Used to preserve strong information in the target area, including significant infrared response and visible light brightness details.

[0132] Considering that the target region not only needs brightness enhancement but also needs to maintain edge and region texture structure, this invention introduces gradient constraints in the target region. Let the gradient reference for the target region be:

[0133] ;

[0134] in, and These represent the weighting coefficients for the visible light gradient and the infrared gradient, respectively. This represents the image edge difference operator, used to capture the edges and local texture features of an image.

[0135] In this embodiment of the application, based on the above definition, the region image loss is expressed as:

[0136] ;

[0137] The target area loss is defined as follows:

[0138] ;

[0139] Background region loss is defined as:

[0140] ;

[0141] In the formula, express Norm, This represents the squared error term. Target area loss. Simultaneously constraining the fusion result to approximate the target brightness reference. It preserves the edge structure of the target region, making the text-related areas more prominent and clearer; background region loss. The main constraint is to ensure that the fusion result approximates the brightness and gradient distribution of the visible light image in non-target areas, in order to preserve the naturalness of background texture, road edges, and scene structure. Because... and Passing through and Normalization ensures that even small target regions are not diluted by large background areas in the total loss. Therefore, the region image loss... This is the key constraint for achieving a balance between text-related region enhancement and background structure stability in this invention.

[0142] In this embodiment, in addition to region image loss, region supervision loss is further employed. Constrain the region response map output by the text-aware region prediction module. Region supervision is performed using a joint form of binary cross-entropy loss (BCE) and Dice loss.

[0143]

[0144] The BCE loss is used for pixel-level binary classification supervision and can be expressed as:

[0145] ;

[0146] Dice loss is used to constrain the overall overlap between the predicted region and the ground truth region, and can be expressed as:

[0147] ;

[0148] in, Represents the predicted regional response map The response value of the i-th pixel, with a value range of . ; Indicates the region truth label for the corresponding pixel; Indicates the total number of pixels; For smoothing terms, BCE focuses on constraining whether the foreground / background classification of each pixel is correct, while Dice focuses more on the consistency of the predicted region with the ground truth region in terms of overall shape and coverage.

[0149] In this embodiment, the region semantic alignment loss This is used to constrain the semantic similarity between high-response regions in the fused image and text features, thereby improving the consistency of the fusion result's response to the text description. Cross-textual difference loss. This is used to enhance the discriminative power of the same image response under different texts, and to prevent the model from generating approximately the same fusion result under different text conditions.

[0150] ;

[0151] ;

[0152] in, and These represent the CLIP image encoder and text encoder, respectively. This indicates the number of text elements corresponding to the same image. and This represents the predicted region response maps generated for the same image under different text inputs. This represents the minimum difference interval.

[0153] S4: Obtain the infrared image, visible light image, and required text description to be fused, and use the trained R2T-Fusion model to achieve the fusion of infrared and visible light images.

[0154] experiment:

[0155] Part 1: Experimental Platform and Dataset

[0156] The IVT dataset was used as the primary training and testing data source. To further verify the model's generalization ability in unknown scenarios, cross-scene testing was also conducted on the RoadScene dataset. The RoadScene dataset mainly contains infrared and visible light image pairs in complex road traffic environments and is often used to evaluate the adaptability of infrared and visible light image fusion methods in road environments. It should be noted that, due to the lack of fine-grained text and region annotations on this dataset, this invention only uses it for objective generalization evaluation of the global fusion quality. During the inference phase, a unified road scene prompt word was used as the text prior input, and cross-scene inference fusion was directly performed. The model of this invention is implemented based on the PyTorch framework and runs on a single NVIDIA RTX 4090 GPU. The number of training epochs was set to 30, and the batch size was set to 1. The optimizer used was AdamW, and the base learning rate was set to... The learning rate is updated using a cosine annealing strategy.

[0157] Part Two: Evaluation Indicators

[0158] To comprehensively evaluate the fusion performance of the proposed method, this invention provides a quantitative evaluation from two levels: full-image fusion quality and text-guided region enhancement quality. The full-image evaluation employs six objective indicators: entropy (EN), standard deviation (SD), spatial frequency (SF), average gradient (AG), visual information fidelity (VIF), and correlation coefficient (CC). EN and SD measure the information content and overall contrast of the fused image, respectively; SF and AG reflect the richness of texture details and edge sharpness, respectively; and VIF and CC evaluate the ability of the fusion result to preserve the visual information and structural relevance of the source image. Considering that full-image indicators cannot directly reflect the effect of text semantics on specific regions, this invention further introduces region-level evaluation indicators: based on the ground truth annotation of the region corresponding to each text. The fusion result is divided into target region and non-target background region, and EN is calculated within the target region. T SD T SF T and AG TT represents the target region: the text region / region of interest, used to calculate the metrics for that region, not the whole image metrics, used to evaluate the information content, contrast, texture detail, and edge response of the specified text region; simultaneously, the Structural Similarity Index (SSIM) is calculated within the background region. B (Structural Similarity Index) and L1 norm B B represents the background and is used to evaluate the structure preservation ability and pixel perturbation degree of non-target areas.

[0159] Part Three: Experimental Results and Analysis

[0160] To verify the full-image fusion quality of the proposed method, this invention compared it with several representative infrared-visible image fusion methods on the IVT dataset. The quantitative results are shown in Table 1. The proposed method achieved the best results in the SF and AG metrics (26.93 and 10.05 respectively), indicating its significant advantages in high-frequency detail preservation and edge sharpness enhancement. In the VIF metric, the proposed method achieved 0.74, second only to TextFusion's 0.81, demonstrating strong visual information fidelity. Simultaneously, the proposed method tied for the best in the CC metric with MaeFuse, indicating its ability to maintain a good overall correlation between the fused result and the source image. The proposed method demonstrates superior performance in detail preservation, edge sharpness, and source image correlation, exhibiting competitive overall fusion performance.

[0161] Table 1 Quantitative comparison results on the IVT dataset

[0162]

[0163] To further analyze the visual differences reflected in the quantitative results in Table 1, Figure 4 Visual comparison results for typical nighttime scenes from the IVT dataset are presented. For the method of this invention, the input text focuses on the pedestrian target area; therefore, the green box is used to demonstrate the local fusion effect of the text-related area. It can be seen that while some comparison methods can retain some scene brightness information, the target area still suffers from problems such as blurred outlines, unclear edges, or strong background interference. In contrast, the method of this invention can present the pedestrian target area more clearly, while better preserving road edges and surrounding structural information, making the human figure outline in the local area more complete, the details clearer, and the overall visual performance more natural.

[0164] To more comprehensively verify the generalization ability of the proposed method, in addition to comparative analysis on the main experimental dataset, this invention further conducted cross-scene generalization experiments on the RoadScene dataset. The experimental results are shown in Table 2 and... Figure 5As shown, the RoadScene dataset differs from the data used in the training phase in terms of scene distribution, target type, and visible and infrared imaging characteristics, thus enabling a more effective evaluation of the adaptability of different methods in unknown scenarios.

[0165] Table 2 Quantitative comparison results on the RoadScene dataset

[0166]

[0167] Comparative analysis reveals that the method of this invention achieves optimal results in SF, AG, and VIF metrics (23.03, 9.44, and 0.74 respectively), indicating that it effectively preserves high-frequency textures, edge details, and visual information even in unknown road scenarios. Simultaneously, the SD metric reaches 57.25, approaching the optimal DIDFuse method, demonstrating strong overall contrast retention. Although the method of this invention does not achieve the best results in EN and CC metrics, it exhibits significant advantages in key metrics that better reflect detail expression and visual fidelity. Figure 5 As can be seen, compared with methods such as DAF Tuse and Swin Fusion, the method of this invention can more clearly present the outline, windows, and body edges of the target vehicle, while maintaining the natural brightness and texture distribution of background areas such as roads, houses, and the sky, without obvious over-enhancement or pseudo-texture phenomena. The combined quantitative and qualitative results show that the proposed method has good detail enhancement capabilities and generalized performance in full-image fusion quality under cross-scene conditions.

[0168] Part Four: Quantitative Evaluation of Enhanced Text-Guided Areas

[0169] To further verify the text-guided region enhancement capability of the proposed method, this invention conducts region-level quantitative evaluation in addition to the overall image objective indicators. As shown in Table 3, the Text-IF and LDFusion methods achieve higher values ​​on some target region enhancement indicators, indicating their strong performance in terms of information content and gradient response in text-related regions. However, both methods fall short of SSIM... B Lower and L1 B The higher value indicates that the enhancement process significantly perturbs non-target regions. In contrast, although the method of this invention is not optimal in terms of statistical indicators for some target regions, it achieves higher SD values. T SF T and AG T It still outperforms the TextFusion method, demonstrating strong target region enhancement capabilities; at the same time, the method of this invention achieves the highest SSIM. B and the lowest L1 BThis indicates that while enhancing text-related regions, it can better maintain the stability of the background structure. Overall, the method of this invention achieves a more reasonable balance between target region enhancement and background structure preservation.

[0170] Table 3. Quantitative comparison results of text-guided region enhancement on the IVT dataset.

[0171]

[0172] Part 5. Analysis of Ablation Experiment Results:

[0173] To verify the effectiveness of the key modules and training strategies in the proposed method, this invention uses the control variable method to construct four ablation configurations: removing the text-aware region prediction module (w / o TRN), removing the progressive mask annealing strategy (w / o PMA), removing the text residual control branch (w / o TRM), and removing the cross-text difference loss (w / o CTD). The experimental results are shown in Table 4.

[0174] Table 4 Ablation Experiment Results of Each Component Module and Training Strategy

[0175]

[0176] Table 4 shows that after removing TRN, the model lacks explicit text-related region priors, and SF and VIF decrease significantly, indicating that the text-aware region prediction module plays an important role in enhancing local details and preserving visual information. After removing PMA, all indicators are lower than the complete model, indicating that progressive mask annealing can alleviate the interference of unstable prediction regions on the learning of fusion weights in the early stage of training. After removing TRM, the model has difficulty effectively adjusting the fusion score under different text conditions, and SF and VIF decrease significantly, indicating that the text residual control branch is the key to improving the fusion difference under text conditions. After removing CTD, EN, VIF, and CC all decrease, indicating that cross-text difference constraints help avoid the degradation of region responses to similar results under different text inputs. In summary, the modules can effectively improve the fusion performance of infrared and visible light images through synergistic effects, verifying the effectiveness of the method design of this invention.

[0177] In summary, this invention constructs a text-aware region prediction network, using visible light images, infrared images, and text features as inputs to generate region response maps related to text semantics. Building upon this, a fusion mechanism based on a Restormer-style dual-branch backbone and text residual control is introduced. The final fusion weights are obtained by jointly calculating the image content-driven basic fusion score and the text residual score, and the fused image is generated in the decoding and reconstruction module. Simultaneously, by combining a ground-value-guided progressive mask annealing training strategy and a target-background joint loss with region area normalization, the interference of unstable region prediction in the early stages of training on the fusion network is effectively mitigated, and the target region enhancement capability and background structure preservation capability are improved.

[0178] Experimental results show that the proposed method exhibits good performance in enhancing text-related regions, preserving detail and texture, improving edge sharpness, and maintaining visual information fidelity. Compared with traditional methods that only perform global fusion, the method of this invention not only maintains the structural integrity and detail of the fused image, but also enhances text-related regions more specifically, demonstrating better text-guided region enhancement capabilities. Overall, the proposed method achieves a reasonable balance between target enhancement and background structure preservation, and has a certain degree of cross-scene generalization ability.

[0179] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A text-guided method for region-enhanced infrared and visible light image fusion, characterized in that, The steps are as follows: S1: Obtain a dataset containing paired infrared images, visible light images, and text descriptions; S2: Construct an R2T-Fusion model. A multimodal feature extraction module extracts infrared image features, visible light image features, and text description features respectively. A text-aware region prediction module models the dual-modal joint features of the infrared and visible light images. Based on the text description features, the dual-modal joint features are sequentially subjected to channel-level text modulation and spatial gating weighting, outputting a continuous region response map corresponding to the text description semantics. A feature fusion module enhances the infrared and visible light image features and constructs enhanced joint features. Guided by the continuous region response map and text features, a basic fusion score is calculated based on the enhanced joint features, and the text residual score based on the text description features is superimposed to obtain the fused features. The fusion features are processed by the decoding and reconstruction module to generate a fused image; S3: The R2T-Fusion model is trained using the dataset and based on the target-background joint loss normalized by the region area. During training, the fusion mask of the region ground truth mask and the continuous region response map is calculated using the progressive mask annealing strategy. The fusion mask is used as the region response map input to the feature fusion module. S4: Obtain the infrared image, visible light image, and required text description to be fused, and use the trained R2T-Fusion model to achieve the fusion of infrared and visible light images.

2. The text-guided region-enhanced infrared and visible light image fusion method according to claim 1, characterized in that, The multimodal feature extraction module extracts infrared image features, visible light image features, and text description features, including: Feature extraction from infrared and visible light images: Two independent convolutional modules are used to extract features from the visible light image. and infrared images Perform shallow feature mapping to obtain shallow features of the visible light image. and shallow features of infrared images Two Restormer modules are used to process shallow features of visible light images. and shallow features of infrared images Perform feature encoding to obtain infrared image features and visible light image features ; Text description feature extraction: A pre-trained CLIP text encoder is used to extract features from the input text description. The text description features are encoded and normalized to obtain the text description features. .

3. The text-guided region-enhanced infrared and visible light image fusion method according to claim 2, characterized in that, Based on textual description features, the dual-modal joint features are sequentially subjected to channel-level text modulation and spatial gating weighting, outputting a continuous region response map corresponding to the semantic description of the text, including: a1. Regarding infrared images and visible light images Dual-modal joint features spliced ​​along the channel dimension Proceed to the first step in sequence The initial features are obtained by using the first ReLU activation function. ; a2. Textual description features Perform the first linear mapping to generate channel-level FiLM scaling factors. and FiLM bias coefficient ; Using FiLM scaling factor and FiLM bias coefficient For initial features Channel-level text modulation features are obtained by performing FiLM text modulation. , Indicates the strength coefficient. This indicates element-wise multiplication; a3. Textual description features The second linear mapping, the second ReLU activation function, and the third linear mapping are performed sequentially to generate a channel-dimensional text modulation vector. , channel dimension text modulation vector Text modulation features After element-wise multiplication, spatial gated convolution is performed. Generate a text space gating graph using the first Sigmoid function. ; Gated graphs of text space After adding 1 to the whole, the text modulation features Perform element-wise multiplication, then pass through the second... The third ReLU activation function, the third Generate unnormalized regional response score maps The continuous regional response map is generated by normalizing the regional response score map Z using the second Sigmoid function. .

4. The text-guided region-enhanced infrared and visible light image fusion method according to claim 3, characterized in that, The feature fusion module enhances infrared and visible light image features and constructs enhanced joint features. Guided by continuous region response maps and text features, a basic fusion score is calculated based on the enhanced joint features, and the text residual score based on text description features is superimposed to obtain the fused features, including: b1. Utilize the feature enhancement module to enhance infrared image features and visible light image features Infrared image enhancement features are obtained by performing feature enhancement operations based on text feature mapping. and visible light image enhancement features ; b2. Using a joint feature construction module based on infrared image enhancement features and visible light image enhancement features Constructing enhanced joint features , Indicates a channel splicing operation; b3. Enhance joint features using a dual-branch weight prediction module. Basic fusion scores and text residual scores are obtained by performing basic fusion based on continuous region response map injection and text residual control, respectively. Fusion weights are then calculated under the guidance of continuous region response maps, and the fusion weights are used to analyze infrared image features. and visible light image features Weighted fusion is performed to obtain fusion features. .

5. The text-guided region-enhanced infrared and visible light image fusion method according to claim 4, characterized in that, Step b1 includes the following feature enhancement operations based on text feature mapping: b11. Using the fourth and fifth linear mappings to describe text features Mapping these maps to the channel spaces of the visible light branch and the infrared branch respectively, yielding visible light branch mapping results and infrared branch mapping results. These results are then input into the first... Function and Second Adding 1 to each function result in a text channel control vector used for visible light image features. Text channel control vectors of infrared image features ; b12, Set the text channel control vector Features of infrared images Perform element-wise multiplication and pass through the fourth Infrared attention response map generated by the third Sigmoid function. ;Text channel control vector Features of visible light images Perform element-wise multiplication, and then pass through the fifth... Generate visible light attention response map using the fourth Sigmoid function. Using infrared attention response maps Features of infrared images Element-wise multiplication yields infrared image enhancement features Using visible light attention response maps Features of visible light images Element-wise multiplication yields visible light image enhancement features .

6. The text-guided region-enhanced infrared and visible light image fusion method according to claim 4 or 5, characterized in that, In step b3, basic fusion based on continuous region response map injection is performed, including: through the sixth Enhanced joint features Perform feature reduction to obtain the reduced features. ; Continuous region response map By adjusting the coefficient After adjustment and dimensionality reduction features Perform channel splicing, and go through the seventh... The operation implements mask injection, and the mask injection features are obtained. Injecting features into the mask To carry out the first After depthwise convolution, it is processed by the fourth ReLU activation function and the eighth And utilize learnable scaling parameters Scaling is performed to obtain the base fusion score. ; In step b3, text residual control is performed, including: through the ninth... Enhanced joint features Perform feature reduction to obtain the reduced features. Textual descriptive features are derived through the sixth linear mapping. Mapped to FiLM scaling factor and FiLM bias coefficient Using scaling FiLM coefficients and FiLM bias coefficient Dimensionality reduction features Perform text FiLM modulation to obtain dimensionality-reduced features after text modulation. , Indicator intensity coefficient; dimensionality reduction features after text modulation. Proceed to the second After depthwise convolution, it is processed by the fifth ReLU activation function and the tenth ReLU activation function. Obtain text residual scores .

7. The text-guided region-enhanced infrared and visible light image fusion method according to claim 6, characterized in that, In step b3, the fusion weights are calculated under the guidance of the continuous region response map, and the fusion weights are used to analyze the infrared image features. and visible light image features Weighted fusion includes: Based on the continuous region response map Construction region guiding coefficient This is used to adjust the influence of text residual scores in different regions. Indicates the base offset parameter; Using regional guiding coefficients Text residual score Element-wise multiplication and using learnable scaling factors After scaling and using weighting coefficients Weighted base fusion score Add them together to get the fusion score. The fusion score is obtained by using the fifth Sigmoid function. Mapped to a fusion weight graph ; Using fusion weight graph Infrared image features and visible light image features Perform pixel-by-pixel weighted fusion to obtain fused features. .

8. The text-guided region-enhanced infrared and visible light image fusion method according to any one of claims 1-5 or 7, characterized in that, The fused image is generated by processing the fused features through the decoding and reconstruction module, including: using the seventh linear mapping to convert text description features... Mapped to FiLM scaling factor and FiLM bias coefficient Using FiLM scaling factor and FiLM bias coefficient Fusion features Text modulation is performed to obtain the fused features after text modulation. , Indicates the strength coefficient, passed through the eleventh... The normalization operation modulates the fused features of the text. Convert to final fused image .

9. The text-guided region-enhanced infrared and visible light image fusion method according to claim 8, characterized in that, The progressive masked annealing strategy is as follows: ; in, Indicates the fusion mask, For the region truth mask, This is the annealing coefficient, which is set to increase sequentially with each training phase. This indicates post-processing operations, including: converting the response value... Crop to Obtain the cropped mask image within the range ; for mask image Perform a power transform to obtain the power-transformed mask image. From the mask image Responses with confidence levels below a preset threshold are filtered out, while responses with confidence levels greater than or equal to the preset threshold are retained. The joint loss due to area normalization includes: ; in, This represents the joint loss normalized to area. Indicates the region image loss. Indicates regional monitoring losses, Indicates the semantic alignment loss of the region. Indicates cross-text difference loss. These are the weighting coefficients for each loss term.

10. The text-guided region-enhanced infrared and visible light image fusion method according to claim 9, characterized in that, The region image loss is expressed as: ; ; ; in, Loss in the target area For background area loss, For reference map of the target area, For single-channel infrared images Representation after copying to three channels. , representing the gradient reference of the target region, and These represent the weighting coefficients for the visible light gradient and the infrared gradient, respectively. This represents the image edge difference operator. Represents the target region weight map. Represents the background region weight map. Indicates the effective area of ​​the target region Indicates the effective area of ​​the background region. To prevent tiny constants with a denominator of zero, Represents the square of the L2 norm. Represents the L1 norm; The regional monitoring loss adopts the joint monitoring loss of BCE loss and Dice loss; The region semantic alignment loss is: ; in, and These represent the CLIP image encoder and text encoder, respectively. The cross-text difference loss is: ; in, This indicates the number of text elements corresponding to the same image. and This represents the predicted region response maps generated for the same image under different text inputs. This represents the minimum difference interval.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method and system based on text-guided semantic perception

    CN120765479A