An infrared and visible light image fusion method for night scene

CN122312400BActive Publication Date: 2026-09-29ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610770210.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-09-29
Estimated Expiration
2046-06-01

AI Technical Summary

Technical Problem

[0004]然而,当上述方法应用于红外与可见光图像融合任务时,尽管在一定程度上提升了融合效果,但仍存在显著局限性

Benefits of technology

[0013]本发明在融合前增强可见光图像的亮度,在融合过程中同时建模跨模态差异、互补信息与全局依赖关系,并结合能量感知保真度损失函数让融合模型能够准确评估红外与可见光特征的重要性。具体而言,首先通过现有光照增强模块以缓解低照度退化影响;随后引入 跨域差异感知注意力融合模块,利用差分放大及三重注意力机制实现多尺度的跨模态特征增强。接着,通过跨域Transformer模块建模红外与可见光之间的全局交互,红外分支采用自注意力保持显著目标一致性,可见光分支通过交叉注意力获取红外引导以强化长距离纹理。最后,提出能量感知保真度损失,利用显著性掩模和能量变化掩模引导模型关注关键区域,迫使融合模型能够准确评估红外与可见光特征的重要性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122312400B_ABST
    Figure CN122312400B_ABST
Patent Text Reader

Abstract

The application discloses a night scene-oriented infrared and visible light image fusion method, which enhances the brightness of the visible light image before fusion, simultaneously models the cross-modal difference, complementary information and global dependency relationship in the fusion process, and combines the energy-aware fidelity loss function to enable the fusion model to accurately evaluate the importance of infrared and visible light features. By adaptively enhancing the brightness of the night visible light image before fusion, the interference of low-illumination conditions on feature extraction and fusion process is relieved; a cross-modal difference-aware attention mechanism is introduced at the feature level to model and amplify the complementary difference between infrared and visible light; and through the energy-aware fidelity loss function, the fusion result is adaptively guided to preferentially inherit the modal features with more significant information, thereby improving the comprehensive performance of the fusion image in target saliency, structural integrity and visual consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and deep learning technology, and in particular to an infrared and visible light image fusion method for nighttime scenes that incorporates illumination enhancement, cross-modal difference perception attention mechanism, and energy perception fidelity constraint. Background Technology

[0002] With the rapid development of multimodal sensing technology, visible light and infrared images are playing an increasingly important role in fields such as surveillance, target recognition, and intelligent transportation. Visible light images capture the rich texture and structural details of a scene by collecting light reflections, but their quality is highly dependent on external lighting conditions and is easily affected by factors such as low light, obstructions, and weather changes. In nighttime scenes, problems such as increased noise, blurring, and underexposure can occur. In contrast, infrared images, based on thermal radiation imaging, can maintain good target saliency in low-light or no-light environments, but lack high-frequency texture information for object identification. Therefore, effectively fusing the structural details of visible light images with the salient target areas of infrared images to obtain images that combine rich detail and target saliency has become the core objective of infrared-visible light fusion technology, and has been widely applied in fields such as military reconnaissance, target detection, traffic monitoring, and environmental monitoring. For example, in traffic monitoring, fused images can provide clear vehicle and pedestrian information at night; in ecological monitoring, fused images can assist in vegetation analysis and wildlife tracking. Therefore, research on nighttime infrared and visible light image fusion has significant theoretical and practical value.

[0003] Compared to daytime scenes, nighttime image fusion faces more complex imaging degradation. Firstly, nighttime visible light images suffer from severe visibility limitations and spectral contamination, manifesting as underexposure in localized areas, enhanced shadows, and overall low brightness. Therefore, it is necessary to enhance the brightness of visible light images during nighttime image fusion to further extract scene information in dimly lit environments, resulting in a more comprehensive fusion outcome. Consequently, many studies introduce low-light enhancement or brightness correction modules before fusion to restore the brightness distribution of nighttime visible light images and enhance key details. Meanwhile, recent image fusion methods can be broadly categorized into two types: traditional methods and deep learning-based methods. Traditional methods typically rely on multi-scale decomposition, sparse representation, or transform domain processing, using hand-designed fusion rules for feature selection. However, on the one hand, hand-designed rules struggle to cover diverse imaging patterns in complex scenes; on the other hand, the transformation process becomes increasingly complex, leading to reduced computational efficiency. With the development of deep learning technology, data-driven fusion methods have achieved significant improvements due to their end-to-end optimization capabilities.

[0004] However, when the aforementioned methods are applied to the task of fusing infrared and visible light images, although they improve the fusion effect to some extent, significant limitations still exist. First, single attention mechanisms typically focus only on a specific attribute of the source image, making it difficult to simultaneously characterize the complex relationships between global semantics and local details during multimodal information interaction, thus limiting the full fusion of cross-modal features. Second, existing methods generally lack modeling of the cross-modal spatial dependencies between infrared and visible light features, especially tending to ignore long-range structures and texture details in visible light images. Furthermore, during multi-stage feature interaction, low-light degradation information continuously interferes with the fusion process, making it difficult for the model to assess the importance of source features from both infrared and visible light images. Summary of the Invention

[0005] To address the aforementioned problems, this invention proposes an infrared and visible light image fusion method for nighttime scenes, comprising the following steps:

[0006] S1: Obtain infrared and nighttime visible light images: Construct a multimodal image dataset with infrared and nighttime visible light images of the same size, and use the infrared and nighttime visible light images as input data; build a fusion network structure including a brightness enhancement module, an encoder, and a decoder; the encoder adopts a densely connected multi-layer convolutional structure to extract multi-scale feature information, and enhances the representation ability of image details and structural information through cross-layer feature fusion.

[0007] S2: The nighttime visible light image is subjected to brightness enhancement processing to obtain an enhanced visible light image; specifically, a brightness enhancement module is introduced into the input nighttime visible light image, and the channel-level brightness is adjusted according to the brightness distribution of the visible light image to enhance the recognizability of dark areas and alleviate the impact of low-light degradation on feature extraction; the brightness enhancement module is based on illumination decomposition, decouples the brightness component and reflection component of the nighttime visible light image, and enhances the brightness component channel by channel through adaptive weight mapping to maintain color consistency while improving overall brightness.

[0008] Furthermore, the cross-domain difference perception attention fusion module is used to model and enhance the difference information between infrared image features and enhanced visible light image features, so as to improve the ability of fused features to express different spatial structures.

[0009] S3: The infrared image and the enhanced visible light image are respectively input into a dual-branch encoder for feature extraction to obtain infrared features and visible light features; specifically, the infrared image and the enhanced visible light image are respectively encoded to obtain infrared features and visible light features; the two modal features are differentially modeled through a cross-domain difference-aware attention fusion module, and the features are weighted by combining channel attention, spatial attention and corner attention to highlight infrared salient target information and enhance visible light texture and structural details; the cross-domain difference-aware attention fusion module includes a differential enhancement unit, power average pooling and a triple attention module;

[0010] S4: Input the infrared features and visible light features into the cross-feature Transformer module, amplify the difference information between the two through the differential enhancement unit, and perform multi-dimensional weighting on the amplified difference features through the triple attention mechanism to obtain the fused features;

[0011] S5: Input the fused features into the decoder for image reconstruction to generate a fused image. Specifically, input the fused features into the decoder for reconstruction to generate a fused image, and constrain the training process using an energy-aware fidelity loss function to make the fusion result adaptively inherit more significant source modality features, thereby obtaining a fused image with complete structure, clear details, and conformity to human visual perception.

[0012] The beneficial effects of this invention are:

[0013] This invention enhances the brightness of visible light images before fusion, and simultaneously models cross-modal differences, complementary information, and global dependencies during the fusion process. It also incorporates an energy-aware fidelity loss function to enable the fusion model to accurately assess the importance of infrared and visible light features. Specifically, it first mitigates the degradation effects of low-light conditions using an existing illumination enhancement module. Then, it introduces a cross-domain difference-aware attention fusion module, utilizing differential amplification and a triple attention mechanism to achieve multi-scale cross-modal feature enhancement. Next, it models the global interaction between infrared and visible light using a cross-domain Transformer module. The infrared branch employs self-attention to maintain salient target consistency, while the visible light branch uses cross-attention to acquire infrared guidance to enhance long-range textures. Finally, it proposes an energy-aware fidelity loss function, using saliency masks and energy change masks to guide the model to focus on key regions, forcing the fusion model to accurately assess the importance of infrared and visible light features. Attached Figure Description

[0014] Figure 1 This is the overall flowchart of an energy-sensing cross-modal attention network for nighttime infrared-visible image fusion.

[0015] Figure 2This is a diagram of an energy-sensing cross-modal attention network architecture for nighttime infrared-visible image fusion.

[0016] Figure 3 This is the architecture diagram of the cross-domain difference perception attention fusion module.

[0017] Figure 4 This presents the fusion experiment results of an energy-aware cross-modal attention network for nighttime infrared-visible light image fusion on the LLVIP dataset. The first column represents the infrared image, the second column the visible light image, and the tenth column the fused image generated using this invention.

[0018] Figure 5 This presents the fusion experiment results of an energy-aware cross-modal attention network for nighttime infrared-visible light image fusion on the MSRS dataset. The first column represents the infrared image, the second column the visible light image, and the tenth column the fused image generated using this invention.

[0019] Figure 6 This presents the fusion experiment results of an energy-aware cross-modal attention network for nighttime infrared-visible light image fusion on the M3FD dataset. The first column represents the infrared image, the second column the visible light image, and the tenth column the fused image generated using this invention.

[0020] Figure 7 This presents the fusion experiment results of an energy-aware cross-modal attention network for nighttime infrared-visible light image fusion on the KAIST dataset. The first column represents the infrared image, the second column the visible light image, and the tenth column the fused image generated using this invention. Detailed Implementation

[0021] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.

[0022] like Figure 1 As shown: This invention includes the following steps:

[0023] S1: Obtain infrared and nighttime visible light images: Construct a multimodal image dataset with infrared and nighttime visible light images of the same size, and use the infrared and nighttime visible light images as input data; build a fusion network structure including a brightness enhancement module, an encoder, and a decoder; the encoder adopts a densely connected multi-layer convolutional structure to extract multi-scale feature information, and enhances the representation ability of image details and structural information through cross-layer feature fusion.

[0024] S2: The nighttime visible light image is subjected to brightness enhancement processing to obtain an enhanced visible light image; specifically, a brightness enhancement module is introduced into the input nighttime visible light image, and the channel-level brightness is adjusted according to the brightness distribution of the visible light image to enhance the recognizability of dark areas and alleviate the impact of low-light degradation on feature extraction; the brightness enhancement module is based on illumination decomposition, decouples the brightness component and reflection component of the nighttime visible light image, and enhances the brightness component channel by channel through adaptive weight mapping to maintain color consistency while improving overall brightness.

[0025] Furthermore, the cross-domain difference perception attention fusion module is used to model and enhance the difference information between infrared image features and enhanced visible light image features, so as to improve the ability of fused features to express different spatial structures.

[0026] S3: The infrared image and the enhanced visible light image are input into a dual-branch encoder for feature extraction to obtain infrared features and visible light features. Specifically, the infrared image and the enhanced visible light image are encoded to obtain infrared features and visible light features. The two modal features are differentially modeled through a cross-domain difference-aware attention fusion module, and the features are weighted by combining channel attention, spatial attention, and corner attention to highlight the infrared salient target information and enhance the visible light texture and structural details. The cross-domain difference-aware attention fusion module includes a differential enhancement unit, power-mean pooling, and a triple attention module. The enhanced visible light image features and infrared image features are first input into the differential enhancement unit for modal difference amplification processing, and the feature update method is as follows:

[0027]

[0028]

[0029] Where i represents the infrared image features obtained from the i-th layer of the encoder. Features of visible light images . Indicates shared components, and Indicates complementary components;

[0030] Constructing bidirectional difference features:

[0031]

[0032]

[0033] in, , This represents the difference information between infrared and visible light modes in the i-th layer feature space; this difference enhancement strategy ensures that the model fully utilizes the modal differences during the fusion process to improve detail preservation and structural integrity.

[0034] To enhance the discriminative power of the differential features, power-average pooling (PAP) is introduced to adaptively modulate the amplitude of the differential information. Specifically, PAP is defined as:

[0035]

[0036] in, Representing the difference information, p is the power exponent, used to adjust the nonlinear response strength of the pooling result. Unlike traditional global average pooling, PAP can highlight significant local differences while statistically analyzing global information. The PAP module enhances its sensitivity to local differences through the power exponent operation, effectively improving the preservation of edge and detail information in the features. Specifically, when p > 1, PAP assigns higher weights to larger numerical differences, highlighting significant edge or detail features in the image; conversely, when p < 1, it helps to smooth local differences. Thanks to this, PAP can more effectively highlight key edges, thermal targets, and texture differences in infrared and visible light images, providing high-quality difference feature support for subsequent fusion.

[0037] After obtaining the differential features, nonlinear activation and channel enhancement techniques are introduced to further enhance them, and the updated features are as follows:

[0038]

[0039]

[0040] in This represents the PReLU(PAP(·)) operation. This represents element-wise multiplication. This difference amplification strategy enables the network to pay more attention to key regions of difference between modalities in subsequent attention modeling.

[0041] The triple attention module consists of channel attention (CA), spatial attention (SA), and corner attention (CoA);

[0042] The channel attention (CA) method measures the importance of each channel at a global semantic level by performing global average pooling and applying sigmoid activation on each channel, thereby obtaining the channel weights for infrared and visible light features:

[0043]

[0044]

[0045] in, This represents the Sigmoid function, where H and W represent the height and width of the feature extracted from the image, respectively. Based on the obtained weights, the weighted channel enhancement feature is:

[0046]

[0047]

[0048] in, This indicates element-wise addition;

[0049] Spatial attention (SA) further models the importance of pixels from a spatial dimension. Its principle is that spatial attention (SA) utilizes pixel positions... Information generates spatial weights, constructs a one-dimensional linear position index sequence, and rearranges it into a two-dimensional spatial position mapping:

[0050]

[0051] Each spatial location The corresponding index value is:

[0052]

[0053] This location mapping provides a unique spatial identifier for each pixel in the feature map, characterizing its relative position within the overall structure. To obtain a smooth and normalizable spatial weight distribution, a sigmoid function is applied to the location index mapping. The location weights for infrared features and enhanced visible light features are calculated as follows:

[0054]

[0055]

[0056] Based on this spatial weight, the enhanced features generated by the spatial attention mechanism are:

[0057]

[0058]

[0059] The Corner Attention (CoA) method starts from the geometric features of the image structure and extracts four corner points ( The feature vectors of the region ∈{(0,0),(0,W),(H,0),(H,W)},k=0,…,3) are averaged and then activated by Sigmoid to generate corner weights, thereby enhancing the expressive power of image edges and corner regions.

[0060]

[0061]

[0062] The enhanced features generated by the corner attention mechanism are:

[0063]

[0064]

[0065] After obtaining the attention features in three dimensions—channel, space, and corner—the cross-domain difference-aware attention fusion module sums them element-wise with the differentially enhanced modal features to finally obtain the first... Layer fusion output characteristics:

[0066]

[0067]

[0068] in, , These represent the infrared features and enhanced visible light features of the (i+1)th layer, respectively.

[0069] S4: The infrared and visible light features are input into the cross-feature Transformer module. The difference information between the two is amplified by the differential enhancement unit, and the amplified difference features are weighted in multiple dimensions by the triple attention mechanism to obtain the fused features. The cross-feature Transformer module processes the different dependencies between the infrared and visible light features respectively.

[0070] For infrared features, the module employs a self-attention mechanism, denoted as... To enhance the long-range dependencies within infrared features, highlight consistency and saliency, thereby maintaining the integrity of the key target's structure;

[0071] In visible light features, a cross-modal cross-attention mechanism is introduced, denoted as... Using infrared features as the query and visible light features as the key and value, this method guides the reconstruction of visible light texture features under infrared saliency constraints. This strategy encourages the network to focus on visible light structures and details highly correlated with the infrared target region, while effectively suppressing background noise and redundant texture interference. The process is as follows: Figure 1 As shown, the infrared features output by the cross-feature Transformer module and visible light characteristics Its mathematical form is expressed as:

[0072]

[0073]

[0074] in, and These are infrared features and enhanced visible light features from the last layer of the preceding convolutional module, respectively.

[0075] The aforementioned differentiated attention design enables the cross-feature Transformer module to adaptively establish effective feature associations between different modalities. While avoiding the weakening of infrared discrimination information, it guides visible light features to align with the infrared salient region, thereby obtaining a fused feature representation with higher structural integrity and more complete texture expression. Considering the sensitivity of the Transformer architecture to input size, this invention introduces a dynamic size adaptation strategy in the cross-feature Transformer module. During the training phase, the input image is uniformly adjusted to an integer multiple of 16×16 to match the 256-dimensional embedding space and improve computational efficiency. During the testing phase, the input is expanded to a suitable size using zero-padding, and then cropped back to the original resolution after fusion. This strategy ensures model stability and training efficiency while allowing it to flexibly adapt to the input requirements of different resolutions and application scenarios.

[0076] S5: Input the fused features into the decoder for image reconstruction to generate a fused image. Specifically, input the fused features into the decoder for reconstruction to generate a fused image, and constrain the training process using an energy-aware fidelity loss function to make the fusion result adaptively inherit more significant source modality features, thereby obtaining a fused image with complete structure, clear details, and conformity to human visual perception.

[0077] The energy-sensing fidelity loss function is as follows: During the multi-stage feature interaction process, low-light degradation information continuously interferes with the fusion process, making it difficult for the model to assess the importance of source features from infrared and visible light images. To address this challenge, an image-level saliency mask is used. and feature-level energy change mask Integrating into the fidelity loss, the energy-aware fidelity loss-guided fusion model can accurately assess the importance of infrared and visible light features, thereby achieving more discriminative and robust feature learning. Specifically, to enhance the retention of saliency information in the fusion result, this invention introduces an image-level saliency mask to guide the model to adaptively weight salient regions. This mask selectively emphasizes the modalities with more prominent information by comparing the infrared image and the enhanced visible light image, and its definition is as follows:

[0078]

[0079] in, Representing modes At pixel position The saliency intensity at a location is used to measure the importance of that location within the spatial domain.

[0080]

[0081] in, Therefore A local neighborhood window centered on the center, This represents the number of pixels within the window. This definition characterizes the contrast of a pixel relative to its local context, effectively reflecting the salience intensity of target edges and structural regions.

[0082] Considering that high-level features focus more on expressing the semantic and content information of an image, this invention further characterizes the differences in content information in different modal source images from the perspective of feature energy changes. Specifically, it calculates the average response of features in the decoder stage and encoder stage in the channel dimension, and uses the difference as a measure of energy change.

[0083]

[0084] Where k represents different modes of infrared or enhanced visible light. and These represent the decoder features and encoder features of the c-th channel, respectively. and These are the number of feature channels for the decoder and encoder, respectively, set to 3. It reflects the energy changes of the source image features and is used to measure changes in content information.

[0085] Feature-level energy change masks can be constructed based on energy changes in different modes:

[0086]

[0087] Comprehensive image-level saliency mask and feature-level energy change mask The energy-sensing fidelity loss ultimately defined in this invention is:

[0088]

[0089] in This represents the fusion result. The first term in the formula is used to constrain the consistency between the fusion result and the infrared image in regions where infrared information is dominant. Represents a mask of characteristic-level energy changes With image-level saliency mask The weighting coefficients obtained through fusion tend to approach 1 when the infrared mode is more dominant in terms of significance or characteristic energy variation at a certain spatial location, thus affecting the... and The differences between them impose stronger constraints, guiding the fusion result to preferentially retain stable target structures and salient region information in the infrared image. The latter term is used to supervise the fusion result in regions where the visible light enhancement mode is more reliable. When the importance of the infrared mode is low, The value of decreases accordingly. Increase the size to make the fusion result closer to an enhanced visible light image. This effectively preserves the rich texture details and background structure information. The two losses form complementary constraints in the spatial dimension, enabling dynamic selection and fusion of information from different modalities. This loss function can adaptively allocate supervision weights for different modalities based on saliency and energy changes, thereby suppressing the interference of degraded information while more effectively preserving the key content and structural information in the source image.

[0090] Figure 2 This is an architecture diagram of the method of the present invention. (a) Overall architecture of the fusion network: The network input includes a visible light image. With infrared images Visible light images First, an image enhancement module is used to generate an enhanced image. Subsequently, infrared images and The features extracted by the encoders are then fed into the dual-branch encoder. Finally, the fused image is obtained through the decoder. . To mitigate energy-sensing fidelity loss, the model adaptively retains modal features with higher information content and more significant structure in degraded scenarios by constraining the changes in the energy of source image features during the fusion process. (b) Encoder: The encoder consists of two branches. The cross-domain difference-aware attention fusion module is responsible for realizing cross-domain information interaction between visible light and infrared modalities, effectively fusing complementary features of the two modalities. Subsequently, the features are fed into the cross-feature Transformer module, which uses self-attention and cross-attention mechanisms to establish global dependencies between features. (c) Fusion module: The fusion module receives features from the dual-branch encoder and uses channel attention and spatial attention mechanisms to adaptively weight the features. (d) Cross-feature Transformer module: The visible light and infrared features are compressed by channels and then enter two consecutive Transformer modules. For the last Transformer module, the visible light branch uses a cross-attention mechanism, with features from the visible light branch serving as values ​​(V) and keys (K), and features from the infrared branch serving as queries (Q). The infrared branch uses a self-attention mechanism, with Q, K, and V all coming from features within the infrared branch itself.

[0091] Figure 3 This is an architecture diagram of the cross-domain difference-aware attention fusion module of the method of the present invention. (a) Cross-domain difference-aware attention fusion module: First, the input infrared features are processed... and enhanced visible light characteristics Perform a difference operation to obtain the difference features. and Subsequently, the two feature paths are input into a parallel dual-branch structure. Within each branch, the features sequentially pass through a triple attention module consisting of channel attention, spatial attention, and corner attention, achieving multi-dimensional interactive fusion of cross-modal features. (b) Channel Attention Module: This module aggregates spatial information through global average pooling and generates channel weights using the Sigmoid function. These weights are then multiplied element-wise with the original features and added to the differential features to obtain the enhanced features. (c) Spatial attention module: in Based on this, the module uses the Sigmoid function to generate weights for each spatial location and then compares the weight map with the input features. Multiply each element point by point, then add the result to the differential features to output the feature. (d) Focus Attention Module: This module first extracts the key corner location information from the feature map and encodes it as corner response features; then, after processing by global average pooling and the Sigmoid function, it is combined with... Multiply by the features, then add to the difference features to obtain the final output features. .

[0092] Figure 4 This image shows the experimental results of infrared image colorization based on topological semantic structure loss on the LLVIP dataset. From the area in the first row of boxes, it can be observed that SwinFusion, DAFTFuse, and FISCNet all exhibit significant spectral contamination and noise in text details, resulting in blurred text and unclear edges on the wall. In contrast, the method of this invention can accurately recover text outlines under low-light conditions, maintaining high contrast and clarity. In the fourth row of boxes, it can be seen that only LENFusion, RDFUSE, and Ours can better preserve the stripe and edge details of zebra stripes. This advantage is attributed to the cross-domain difference-aware attention fusion module proposed in this invention. It enhances the complementarity of infrared salient features and visible light detail features through a triple attention mechanism of channels, space, and corners. Simultaneously, it combines a cross-feature Transformer module to capture cross-modal global dependencies, making the fusion result more natural and consistent in terms of brightness, texture, and structure.

[0093] Figure 5 This image shows the experimental results of the infrared image colorization method based on topological semantic structure loss on the MSRS dataset. From the comparison in the second row of boxes, it can be observed that CS2Fusion and RDMFSuse exhibit varying degrees of gradient vanishing, resulting in blurred boundaries between buildings and the night sky, making them difficult to distinguish effectively. This is mainly because these methods fail to adequately model the energy balance between infrared salient features and visible light texture features, leading to the loss of detail information in dark areas or low-contrast scenes. In contrast, the method of this invention maintains clear structural edges in the transition area between buildings and the background within the red box in the second row, demonstrating superior detail fidelity and visual hierarchy. In the fourth row of boxes, it is evident that only CS2Fusion and Ours can fully preserve the texture structure of the brick surface while highlighting salient targets. This advantage is attributed to the energy-perceived fidelity loss proposed in this invention. This loss effectively maintains the energy distribution balance between salient and background areas by constraining energy changes between feature layers, thereby simultaneously improving structural consistency, detail integrity, and brightness balance during the fusion process.

[0094] Figure 6This image shows the experimental results of infrared image colorization using the topological semantic structure loss-based method on the M3FD dataset. Observation of the bounding boxes reveals that RDFuse is significantly too dark in the fused image. As seen in the second row of boxes, SwinFusion, DATFuse, and HAIAFusion exhibit significant blurring and texture loss in the building window area. In the third row of boxes, only HAIAFusion and Ours retain the true color and edge contours of the white vehicle. This performance is attributed to the introduction of a cross-feature Transformer, which enables the model to capture cross-modal global dependencies, achieving color and structure consistency under infrared saliency guidance, while the energy-aware fidelity loss effectively suppresses the problem of over-compression in dark areas.

[0095] Figure 7 The images show the experimental results of an infrared image colorization method based on topological semantic structure loss on the KAIST dataset. Observation of the bounding boxes reveals that the fused image from HAIAFusion is almost indistinguishable from the original visible light image, indicating that it fails to fully utilize infrared features. From the bounding boxes in the second row of images, both LENFusion and RDMFuse exhibit blurred images of cars parked on the roadside. In the bounding boxes of the fourth row, it is clearly observed that the fused images generated by DATFuse, CS2Fusion, HAIAFusion, and RDMFuse do not display the features of trees and bushes in the infrared image. In contrast, the method of this invention, through the difference enhancement mechanism of the cross-domain difference-aware attention fusion module, effectively balances infrared salient features and visible light structural features, resulting in fusion results that exhibit higher clarity and depth in both salient target preservation and background texture rendering.

[0096] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.

Claims

1. A method for fusing infrared and visible light images for nighttime scenes, characterized in that, Includes the following steps: S1: Construct a multimodal image dataset with infrared images and nighttime visible light images of the same size, and use the infrared images and nighttime visible light images as input data; build a fusion network structure that includes a brightness enhancement module, an encoder, and a decoder; S2: A brightness enhancement module is introduced into the input nighttime visible light image. The brightness of the image is adjusted at the channel level according to the brightness distribution of the visible light image to enhance the recognizability of dark areas and alleviate the impact of low-light degradation on feature extraction. S3: Perform feature encoding on the infrared image and the enhanced visible light image respectively to obtain infrared features and visible light features; The two modal features are differentially modeled by a cross-domain difference perception attention fusion module, and the features are weighted by combining channel attention, spatial attention and corner attention to highlight infrared salient target information and enhance visible light texture and structural details. S4: Input the infrared features and visible light features into the cross-feature Transformer module, amplify the difference information between the two through the differential enhancement unit, and perform multi-dimensional weighting on the amplified difference features through the triple attention mechanism to obtain the fused features; S5: Input the fusion features into the decoder to reconstruct the image and generate a fused image; The cross-domain difference-aware attention fusion module in step S3 includes a differential enhancement unit, power-mean pooling, and a triple attention module. The enhanced visible light image features and infrared image features are first input to the differential enhancement unit for modal difference amplification processing, and the feature update method is as follows: Where i represents the infrared image features obtained from the i-th layer of the encoder. Features of visible light images ; Indicates shared components, and Indicates complementary components; Constructing bidirectional difference features: in, , This represents the difference information between infrared and visible light modes in the i-th layer feature space; Power-mean pooling (PAP) is introduced to adaptively modulate the differential information. Specifically, PAP is defined as: in, This represents the difference information, where p is the power exponent, used to adjust the nonlinear response strength of the pooling result; After obtaining the differential features, nonlinear activation and channel enhancement techniques are introduced to further enhance them, and the updated features are as follows: in This represents the PReLU(PAP(·)) operation. This indicates element-wise multiplication.

2. The infrared and visible light image fusion method for nighttime scenes according to claim 1, characterized in that: In step S1, the encoder employs a densely connected multi-layer convolutional structure to extract multi-scale feature information and enhances the representation of image details and structural information through cross-layer feature fusion.

3. The infrared and visible light image fusion method for nighttime scenes according to claim 1, characterized in that: In step S2, the brightness enhancement module decouples the brightness component and reflection component of the nighttime visible light image based on illumination decomposition, and enhances the brightness component channel by channel through adaptive weight mapping to maintain color consistency while improving overall brightness.

4. The infrared and visible light image fusion method for nighttime scenes according to claim 3, characterized in that: A cross-domain difference perception attention fusion module is used to model and enhance the difference information between infrared image features and enhanced visible light image features, so as to improve the ability of fused features to express different spatial structures.

5. The infrared and visible light image fusion method for nighttime scenes according to claim 1, characterized in that: The triple attention module consists of channel attention (CA), spatial attention (SA), and corner attention (CoA); The channel attention (CA) method measures the importance of each channel at a global semantic level by performing global average pooling and applying sigmoid activation on each channel, thereby obtaining the channel weights for infrared and visible light features: in, This represents the Sigmoid function, where H and W represent the height and width of the feature extracted from the image, respectively. Based on the obtained weights, the weighted channel enhancement feature is: in, This indicates element-wise addition; The spatial attention (SA) utilizes pixel location Information generates spatial weights, constructs a one-dimensional linear position index sequence, and rearranges it into a two-dimensional spatial position mapping: Each spatial location The corresponding index value is: Applying a Sigmoid function to the location index map, the location weights for infrared features and enhanced visible light features are calculated as follows: Based on this spatial weight, the enhanced features generated by the spatial attention mechanism are: The Corner Attention (CoA) method starts from the geometric features of the image structure and extracts four corner points ( The feature vectors of the region ∈{(0,0),(0,W),(H,0),(H,W)},k=0,…,3) are averaged and then activated by Sigmoid to generate corner weights, thereby enhancing the expressive power of image edges and corner regions. The enhanced features generated by the corner attention mechanism are: After obtaining the attention features in three dimensions—channel, space, and corner—the cross-domain difference-aware attention fusion module sums them element-wise with the differentially enhanced modal features to finally obtain the first... Layer fusion output characteristics: in, , These represent the infrared features and enhanced visible light features of the (i+1)th layer, respectively.

6. The infrared and visible light image fusion method for nighttime scenes according to claim 1, characterized in that: The cross-feature Transformer module handles the different dependencies between infrared features and visible light features respectively; For infrared features, the module employs a self-attention mechanism, denoted as... ; In visible light features, a cross-modal cross-attention mechanism is introduced, denoted as... Using infrared features as the query and visible light features as the key and value, the visible light texture features are reconstructed under the constraint of infrared saliency. Infrared features output by the cross-feature Transformer module and visible light characteristics Its mathematical form is expressed as: in, and These are infrared features and enhanced visible light features from the last layer of the preceding convolutional module, respectively.

7. The infrared and visible light image fusion method for nighttime scenes according to claim 1, characterized in that: Step S5 specifically involves: inputting the fused features into the decoder for reconstruction to generate a fused image, and constraining the training process through an energy-aware fidelity loss function to make the fusion result adaptively inherit more significant source modality features, thereby obtaining a fused image that is structurally complete, detailed, and conforms to human visual perception.

8. The infrared and visible light image fusion method for nighttime scenes according to claim 7, characterized in that: The energy-sensing fidelity loss function is: [Image-level saliency mask] and feature-level energy change mask An energy-aware fidelity loss is constructed by integrating it into the fidelity loss. This mask selectively emphasizes the modalities with more prominent information by making a saliency comparison between the infrared image and the enhanced visible light image. It is defined as follows: in, Representing modes At pixel position The saliency intensity at a location is used to measure the importance of that location within the spatial domain. in, Therefore A local neighborhood window centered on the center, This represents the number of pixels within the window.

9. The infrared and visible light image fusion method for nighttime scenes according to claim 8, characterized in that: To characterize the differences in content information in source images of different modalities from the perspective of feature energy change, specifically: calculate the average response of features in the decoder stage and encoder stage in the channel dimension, and use the difference as a measure of energy change: Where k represents different modes of infrared or enhanced visible light. and These represent the decoder features and encoder features of the c-th channel, respectively; and These are the number of feature channels for the decoder and encoder, respectively, and are set to 3. It reflects the energy changes of the source image features and is used to measure changes in content information; Feature-level energy change masks can be constructed based on energy changes in different modes: The final defined energy-sensing fidelity loss is: in This indicates the fusion result.

Citation Information

Patent Citations

  • Multi-modal image target detection method based on image fusion

    CN110322423A

  • KR20250135520A