An infrared and visible light image fusion method and system

CN122798641APending Publication Date: 2026-09-22WUHAN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611275816.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

第一,现有的图像融合方法缺乏有效的人类视觉先验信息引导机制,传统的融合方法往往以无差别的方式处理整幅图像,无法根据用户关注的重点区域进行针对性的融合优化,导致显著目标的融合效果不佳;

Benefits of technology

1、提出了一种涂鸦引导红外和可见光图像融合框架,利用视觉语言大模型将稀疏涂鸦提示转换为语义掩码,避免了语言驱动指导中固有的歧义,通过低成本交互支持任意曲线、折线和不规则形状,并产生鲁棒的零样本分割,对未见过的场景具有很强的泛化能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798641A_ABST
    Figure CN122798641A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of deep learning image fusion, and particularly relates to an infrared and visible light image fusion method and system, comprising the following steps: S1: importing a to-be-tested image; S2: using an image editing tool to scribble a region of interest in the infrared or visible light image, and inputting the scribble and an image obtained by splicing the infrared and visible light image channels into a visual language large model to construct a semantic mask; S3: introducing a correlation decoupling loss, decoupling and enhancing the high-frequency and low-frequency feature information between the infrared and visible light image features, and constructing three layers of cross-modal features with the first two layers of encoding features of a pre-trained convolutional neural network; the correlation of cross-modal low-frequency shared information and high-frequency specific information is enhanced and decoupled, the fusion capability of the cross-modal features is improved, and a fused image with a prominent target, clear structure and rich texture information is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of deep learning image fusion, and in particular to a method and system for fusion of infrared and visible light images. Background Technology

[0002] Due to limitations in physical imaging mechanisms, a single sensor often struggles to fully perceive complete scene information in complex or extreme environments. To address this issue, image fusion technology integrates complementary information from multiple sensors to construct a more comprehensive and robust scene representation. Among various multimodal fusion tasks, Infrared and Visible Image Fusion (IVIF) has garnered significant attention due to its important applications in autonomous driving and video surveillance. Infrared images, by capturing the thermal radiation of objects, can effectively emphasize salient targets under low-light or adverse weather conditions. However, they lack rich color information and fine-grained structural details. In contrast, visible light images can provide high-resolution structural textures and rich visual details, but their imaging quality largely depends on lighting conditions. Therefore, IVIF aims to fully combine salient target information from infrared images with structural texture information from visible images to produce a fusion output that preserves prominent thermal targets and a natural visual appearance, enhancing scene understanding for downstream tasks such as object detection, semantic segmentation, and visual tracking.

[0003] However, the significant differences in intensity distribution, contrast, and texture structure between infrared and visible light images make the efficient aggregation of cross-modal complementary features a challenging task. Existing technologies mainly suffer from the following shortcomings: First, existing image fusion methods lack effective mechanisms to guide the use of prior information from human vision. Traditional fusion methods often process the entire image indiscriminately, failing to perform targeted fusion optimization based on the key areas of interest to the user, resulting in poor fusion performance for prominent targets. Second, existing methods fail to adequately consider the correlation differences between infrared and visible light images at different frequency components during the shallow feature extraction stage. The low-frequency components of infrared and visible light images exhibit a certain statistical correlation (including background content and large-scale environmental features), while their high-frequency components are mode-specific due to modal differences. Existing methods fail to effectively perform differentiated feature enhancement and decoupling processing for these two frequency components. Third, in the process of cross-modal feature interaction, existing methods either adopt a global self-attention mechanism, which leads to excessive computational cost and is prone to noise interference, or adopt local convolution operations, which cannot model long-distance cross-modal dependencies, making it difficult to achieve a good balance between global semantic relevance and preservation of local fine structural details. Fourth, during the multi-scale fusion process, due to unpredictable object boundaries and severe feature degradation during network propagation, low-resolution object features are difficult to align with their high-resolution counterparts, resulting in insufficient semantic region responses in the fusion results. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides an infrared and visible light image fusion method and system that utilizes graffiti prompts to introduce prior information about human vision and enhances and decouples cross-modal low-frequency shared information and high-frequency specific information based on the correlation differences of different frequency features. Simultaneously, this invention constructs a semantically guided global and local feature interaction mechanism and uses semantic information to modulate cross-scale features, thereby improving the fusion capability of cross-modal features and obtaining a fused image with prominent targets, clear structure, and rich texture information.

[0005] An infrared and visible light image fusion method of the present invention includes the following steps: S1: Import the image to be tested; S2: Use image editing tools to doodle on the region of interest in the infrared or visible light image, and input the doodle with the image after stitching the infrared and visible light image channels into the visual language big model to construct a semantic mask; S3: Using the correlation decoupling shallow feature enhancement module, the correlation decoupling loss is introduced to decouple the high-frequency and low-frequency feature information between infrared and visible light image features and enhance the feature expression. It also constructs a three-layer cross-modal feature with the first two layers of the pre-trained convolutional neural network encoding features. S4: By using a masked proxy cross-attention layer, a good balance is achieved between modeling global cross-modal semantic relevance and preserving local fine structural details by introducing semantic mask-guided proxy cross-attention and adaptive local convolutional interaction; S5: Through the multi-scale feature calibration module, deep fusion features and semantic masks generated by graffiti are used as prior guidance, and shallow fusion features are calibrated through spatial adaptive normalization mechanism, thereby enhancing the cross-scale consistency of key semantic regions while preserving local details. S6: Construct a fused image of infrared and visible light images with clear details and rich semantics through a feature decoder.

[0006] Preferably, the specific process of S2 includes: The specific process of S2 includes: S21: Use red graffiti to draw on areas of interest in infrared images through image editing; S22: By extracting the skeleton and sampling discrete points, the graffiti prompts are sampled into discrete points with a certain shape, which serve as prompt information for the visual segmentation model. S23: Replace the R channel of the visible image with an infrared image of the same scene, while retaining the visible structural and texture information in the G and B channels to construct a cross-modal image. Input the reconstructed cross-modal image and sampling point cues into the large visual segmentation model to generate semantic masks for the subsequent fusion network. Because visible light images contain too many colors, they can pose a certain obstacle to graffiti extraction. Therefore, image editing is used to select red graffiti as the area of ​​interest in infrared images for graffiti extraction. By extracting the skeleton and sampling discrete points, the graffiti prompts are sampled as discrete points with a certain shape, which serve as prompt information for the visual segmentation model. Different sensors capture different types of information, and single-modal images may not provide enough cues to generate masks that cover infrared target responses and visible structural details. To construct a unified cross-modal cue input, this invention uses a cross-modal channel shuffling strategy. The R channel in the RGB image is located in the long wavelength range of visible light and is closer to infrared imaging in terms of spectral characteristics. Therefore, the R channel of the visible image is replaced with an infrared image of the same scene, while the G and B channels retain visible structural and texture information. In this way, the reconstructed image retains the original RGB input format, while incorporating significant infrared target features and preserving scene details from the visible image. Finally, the reconstructed cross-modal image and sampling point cues are input into a large visual segmentation model to generate semantic masks for subsequent fusion networks.

[0007] Preferably, the specific process of S3 includes: S31: First, the infrared and visible light images and semantic mask are input into the fusion network. For the visible light image with RGB channels, it is converted into YcrCb channels and only the Y channel is taken as the input of the visible light image in the fusion network. S32: Infrared and visible light images of the same scene, their low-frequency components should have statistical correlation, including background content and large-scale environmental features. Secondly, due to modal differences, their high-frequency components are modal-specific and generally exhibit relatively weak cross-modal correlation. The goal is to maintain the consistency of relevant information by enhancing modal shared features, while integrating relatively independent complementary high-frequency cues in the two modalities. This invention proposes a related decoupling shallow feature enhancement module to replace the first convolutional layer of ResNet, aiming to enhance the consistent low-frequency structural response across modalities while decoupling high-frequency cues for specific modalities; Specifically, the source image The input is fed into basic convolutional mapping and normalization operations to obtain the corresponding shallow features. ; Subsequently, in order to decouple and enhance modality-specific high-frequency information and share low-frequency representation, the related decoupling shallow feature enhancement module processes shallow features through two dedicated branches (i.e., high-frequency enhancement branch and low-frequency enhancement branch); S33: Shallow features After high-frequency branching, differential edge convolution and local structure modeling are used as embedded learnable network operations to extract high-frequency features. ; S34: For low-frequency enhancement branches, shallow features First, shallow features are processed by low-pass filtering to eliminate high-frequency noise and preserve robust shared low-frequency components, resulting in low-frequency features. It is then fed into the proposed dual-scale enhanced residual block for further enhancement, as shown below: ; Specifically, low-frequency characteristics The input is first fed into two parallel convolutional branches: a 3×3 branch to capture local structural details, and a 5×5 branch to aggregate broader neighborhood contextual information through a larger receptive field. The normalized 3×3 and 5×5 convolutional responses are then summed to produce intermediate multi-scale features. Subsequently, It is passed to the dual-path feature interaction unit, one path applies a fully connected projection, and then... One path activates the other fully connected projection without nonlinear activation, and the two projected features are fused through element-wise multiplication to achieve high-order nonlinear feature interaction. S35: To enhance feature representation by strengthening the correlation of low-frequency features while weakening the correlation of high-frequency features, this invention constructs a correlation-constrained loss adapted to cross-modal image fusion. Specifically, it defines the similarity relationship between two features using cosine similarity, as shown below: ; in This represents the pooling operation; in order to enhance the low-frequency correlation across modalities while constraining the high-frequency features of cross-modal images and the correlation between high-frequency and low-frequency features within the same modality, the correlation constraint loss consists of three parts, specifically; ; in , and It's about adjusting parameters. The aim is to maximize the correlation between low-frequency features across modalities. The aim is to minimize the correlation between low-frequency features across modalities and the correlation between high-frequency features across modalities. The aim is to minimize the correlation between low-frequency and high-frequency features within each modality; therefore, the three losses are specifically constructed as follows: ; in and This represents the low-frequency characteristics of visible light and infrared images, while and This indicates its corresponding high-frequency characteristics; furthermore, in and Absolute value constraints are introduced to maintain the relative independence of cross-modal high-frequency features and intra-modal low-frequency features; Specifically, positive and negative correlations represent the same and opposite trends of change between two features, respectively. Both of these cases indicate strong feature coupling. Therefore, minimizing absolute correlation suppresses strong positive and negative correlations, pushes the correlation to zero, reduces linear correlation, and enhances feature decoupling.

[0008] Preferably, in step S34, the processing procedure for the dual-scale enhanced residual block is as follows: Low-frequency features are first fed into two parallel convolutional branches: a 3×3 convolutional branch to capture local structural details and a 5×5 convolutional branch to aggregate broader neighborhood contextual information through a larger receptive field. Adding the normalized 3×3 convolutional responses and 5×5 convolutional responses produces intermediate multi-scale features; Intermediate multi-scale features are passed to the dual-path feature interaction unit. One path is activated after applying a fully connected projection, while the other path performs another fully connected projection without nonlinear activation. The two projected features are fused through element-wise multiplication to achieve high-order nonlinear feature interaction. Preferably, the specific process of S4 includes: S41: Due to the physical differences in the imaging mechanisms of infrared and visible light sensors, there are inherent domain differences between the source image features. Direct fusion performed in the original image domain often blurs fine-grained details and produces obvious artifacts in the fusion result. To mitigate these drawbacks, sufficient cross-modal feature interaction is usually implemented beforehand. This invention proposes a mask proxy cross-attention layer. Specifically, the enhanced features are decomposed by the relevant decoupled shallow feature enhancement module and encoded by the first two layers of ResNet to construct three pairs of infrared and visible light image features at different scales. ,in Then, infrared and visible light features of the same scale are input into the mask proxy cross-attention layer to fully interact with global semantics and local details; S42: For the global information interaction branch of the masked proxy cross-attention layer, this invention constructs mask-guided proxy features, embedding semantic constraints into cross-modal interactions, thereby prioritizing foreground region interactions rather than interactions from irrelevant non-critical regions. Specifically, the semantic mask generated from the graffiti prompt is first mapped by a convolutional layer to obtain the mask feature Fmask. Subsequently, the mask feature is used to calibrate the projection query features in the global proxy token construction. Semantic-guided feature refinement and enhancement are achieved by emphasizing meaningful foreground regions and reducing background interference, which can be described as follows: ; Where parameters Set to 0.5. Indicates the modal features used for querying; S43: Subsequently, by constructing a small number of proxy tokens As a global proxy representation, to avoid the high computational cost and noise interference caused by direct, dense pairwise correlation calculations on the original spatial features, under the guidance of mask constraints, the proxy token acts as a semantic mediator in cross-modal attention computation, transforming pixel-level dense interactions into a two-stage process of global agent aggregation and spatial location broadcasting. This mitigates the modality gap caused by low-level response differences, and can be expressed as follows: ; in and It is the learned relative positional deviation. Indicates the value used for attention calculation The function, in order to fully interact with the global semantic information of infrared and visible light images, uses infrared and visible light image features as bidirectional queries to perform global interaction and obtain updated global interactive features. ; S43: Although agent cross-attention excels at modeling long-range cross-modal dependencies, its global aggregation nature tends to discard distinctive local details unique to each modality. Therefore, the agent cross-attention layer also introduces a CNN-based local detail interaction branch to preserve high-frequency structural information. Specifically, it first connects multi-scale features along the channel dimension. ,in Indicates from the The features of a layer, the connection features are represented as Subsequently, Two inputs based on The weight generation branch is used to obtain mode-specific local weights for infrared and visible light features, respectively. and Meanwhile, through convolution, normalization, and Activate from Extracting local spatial information and further introducing learnable parameters. The contributions from the two modes of adaptive equilibrium, the local feature interaction process can be described as follows: ; in This represents element-wise multiplication, and finally, integrated local features are used. Corresponding features to the global Gating mechanisms to generate enhanced interactive features This design enables effective complementary feature learning between global semantics and local details, and finally updates the features after the interaction. The input is fed into the channel fusion module, which fuses the infrared and visible light image features updated at the same scale to obtain the corresponding fused features. .

[0009] Preferably, the specific process of S5 includes: S51: For infrared and visible image fusion tasks involving high-resolution input and multi-scale objects, single-scale fusion features are insufficient to balance infrared salient objects and visible image details due to feature differences between infrared and visible images. Multi-scale cross-modal fusion enables high-level features to encode infrared targets and global context, while low-level features can preserve visible textures and fine structures in low-level features, improving the model's generalization ability to complex real-world scenes. However, due to unpredictable object boundaries and severe feature degradation during network propagation, it is difficult to align low-resolution object features with their high-resolution counterparts. To address this issue, this invention proposes a multi-scale feature calibration module that uses mask information and deep features as semantic guidance to further modulate shallow fusion features. S52: Specifically, firstly, with Under guidance Taking calibration as an example, the graffiti is used to generate a semantic mask. And upsampling processing, generating with Soft masks with the same spatial size Subsequently, on Perform upsampling and The connectivity features are processed through convolution and normalization, followed by two independent 3 × 3 convolutional layers to predict the scaling factor. and shift factor , The function is used to... Limiting to the range [-1, 1] to achieve stable modulation, then shallow features are... Calibration, the process is as follows: ; in This represents element-wise multiplication. Finally, the calibrated features and the original features are adaptively combined using a soft mask. ; S53: After receiving the updated calibration In the same way, using masks and right Further modulation and calibration are performed to obtain the final fusion characteristics.

[0010] Preferably, an infrared and visible light image fusion system includes a semantic mask acquisition module, a correlation decoupling shallow feature enhancement module, a mask proxy cross-attention layer, a fusion module, a multi-scale feature calibration module, and a mask decoder; Semantic mask acquisition module: used to construct corresponding semantic masks from imported infrared and visible light image pairs and constructed graffiti prompts; Correlation decoupling shallow feature enhancement module: used before integrating the Rsenet network, it performs shallow encoding on infrared and visible light images, decouples high-frequency and low-frequency feature information, and constructs three layers of multi-scale encoded features together with the first two layers of encoding of the integrated network through correlation decoupling loss constraint; Masked proxy cross-attention layer: Introduces semantically guided proxy cross-attention and local adaptive convolution to perform global semantic interaction and local detail interaction within the same layer on feature spectral pairs of infrared and visible light at three different scales; Fusion module: used to fuse infrared and visible light images of the same size after full interaction to obtain the feature spectrum of the final fused image at different scales; Multi-scale feature calibration module: Modulates deep fusion features using semantic masks and shallow fusion features, and enhances and enriches the semantic information of the fusion features through explicit semantic injection; Mask decoder: Used to input the modulated fusion features into the mask decoder for decoding and outputting the fused image; Preferably, in the semantic mask acquisition module, the R channel of the visible image is replaced with an infrared image of the same scene, while the G and B channels retain the visible structural and texture information to construct a cross-modal image; the reconstructed cross-modal image and sampling point cues are input into the visual segmentation large model to generate a semantic mask; Preferably, the related decoupling shallow feature enhancement module processes shallow features through two dedicated branches: a high-frequency enhancement branch and a low-frequency enhancement branch. The high-frequency enhancement branch extracts high-frequency features using differential edge convolution and local structure modeling; The low-frequency enhancement branch is processed by low-pass filtering and then input into the dual-scale enhancement residual block for enhancement. Preferably, it further includes a computer storage medium storing a computer program executable by a processor, characterized in that the computer program performs an infrared and visible light image fusion method.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. A framework for fusing infrared and visible light images guided by graffiti is proposed. It uses a large visual language model to convert sparse graffiti prompts into semantic masks, avoiding the ambiguity inherent in language-driven guidance. It supports arbitrary curves, polylines and irregular shapes through low-cost interaction and produces robust zero-sample segmentation. It has a strong generalization ability for unseen scenes. 2. In order to improve shallow feature representation and mitigate modal differences by decoupling and constraining the correlation of low-frequency features and minimizing the correlation of high-frequency features, this invention proposes a decoupled shallow feature enhancement module with novel correlation decoupling loss; 3. In order to better facilitate the interaction between cross-modal features and provide features containing cross-modal feature information for subsequent fusion, this invention introduces mask-guided proxy cross-attention and local convolution interaction during the interaction process, utilizes global context relevance to achieve high-fidelity feature interaction, and preserves fine local texture details through adaptive local feature modulation. 4. In order to enhance the feature response in salient object regions and suppress unreliable responses caused by feature degradation, thereby mitigating feature distortion during deep multi-scale fusion and feature propagation, this invention constructs a multi-scale feature calibration module that utilizes semantic masks derived from graffiti hints to enhance the feature response in salient object regions. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the process flow of the present invention; Figure 2 This is a schematic diagram of the system structure of the present invention; Figure 3 This is a schematic diagram of the semantic mask generation process of the present invention. Detailed Implementation

[0013] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be thorough and complete.

[0014] Example 1

[0015] like Figure 1 As shown, an infrared and visible light image fusion method of the present invention includes the following steps: S1: Import the image to be tested; S2: Use image editing tools to doodle on the region of interest in the infrared or visible light image, and input the doodle with the image after stitching the infrared and visible light image channels into the visual language big model to construct a semantic mask; This step generates a semantic mask containing prior information about the target region through graffiti interaction and a large visual model, which serves as semantic guidance for the subsequent fusion network. S21: Graffiti Marking: In this embodiment, since visible light images contain rich color distributions, directly drawing annotations on them can easily introduce color and texture interference, which is not conducive to the accurate extraction of prompt information. Therefore, this embodiment chooses to draw annotations in infrared images. Users use image editing tools (such as Photoshop or GIMP) to draw freely on the target area of ​​interest (such as pedestrians, vehicles, and other areas with significant thermal radiation) in the infrared image with a red pen, forming a graffiti mark covering the target subject. This graffiti mark serves as prior information for human vision and is used to indicate the foreground area that the fusion network should focus on. S22: Graffiti Point Sampling: To adapt the graffiti markers to the input format of the visual segmentation large model, this embodiment performs point sampling processing on the graffiti markers drawn in S21. First, skeletonization is performed on the graffiti area to obtain a single-pixel wide skeleton structure representing the geometric center line of the graffiti. Then, equally spaced discrete point sampling is performed along the skeleton line to generate a set of spatially uniform two-dimensional coordinate points. This set of points serves as point prompts for the visual segmentation large model, which not only preserves the shape features of the graffiti but also avoids redundant calculations caused by dense pixel input. S23: Cross-modal input suggestion construction: Since single-modal images cannot simultaneously provide complete clues to both infrared target response and visible light structural details, if the original infrared or visible light images are directly input into a large visual segmentation model, the generated masks often cannot take into account both significant thermal targets and fine texture boundaries. Therefore, this embodiment proposes a cross-modal channel shuffling strategy to construct a unified cross-modal input image. Specifically, considering that the R channel in the RGB color space is in the long wavelength range of visible light and its spectral characteristics are closest to those of infrared imaging, the R channel of the original visible light image is replaced with an infrared image of the same scene, while keeping the G and B channels unchanged. This preserves the structural and texture information of the original visible light image. The reconstructed image obtained by this operation retains the standard RGB input format and incorporates significant infrared target thermal radiation features and visible light scene details, providing sufficient cross-modal cues for the segmentation model. Finally, the graffiti discrete point set obtained in S22 and the reconstructed cross-modal image are input into a large visual segmentation model (e.g., Segment Anything Model, SAM). Guided by point cues, the model performs segmentation inference on the cross-modal image and outputs a semantic mask corresponding to the graffiti region. This semantic mask, as a soft label, represents the spatial distribution probability of salient foreground targets in the image and will serve as a prior semantic constraint for the fusion network in subsequent steps S3 to S5. It is used to guide the network to focus on salient foreground regions during feature extraction, cross-modal interaction, and multi-scale calibration, thereby enhancing target response and suppressing background interference. If the color distribution of the visible light image is relatively simple or the target area is in stark contrast to the background, the graffiti annotation can be directly applied to the visible light image, omitting the graffiti step in the infrared image, in order to simplify the user interaction process. In this case, the graffiti prompt is directly applied to the visible light image. After skeleton extraction and discrete point sampling, it is input into the visual segmentation big model along with the cross-modal reconstructed image to generate a semantic mask. S3: Using the correlation decoupling shallow feature enhancement module, the correlation decoupling loss is introduced to decouple the high-frequency and low-frequency feature information between infrared and visible light image features and enhance the feature expression. It also constructs a three-layer cross-modal feature with the first two layers of the pre-trained convolutional neural network encoding features. The specific process of S3 includes: S31: First, the infrared and visible light images and semantic mask are input into the fusion network. For the visible light image with RGB channels, it is converted into YcrCb channels and only the Y channel is taken as the input of the visible light image in the fusion network. S32: Infrared and visible light images of the same scene, their low-frequency components should have statistical correlation, including background content and large-scale environmental features. Secondly, due to modal differences, their high-frequency components are modal-specific and generally exhibit relatively weak cross-modal correlation. The goal is to maintain the consistency of relevant information by enhancing modal shared features, while integrating relatively independent complementary high-frequency cues in the two modalities. This invention proposes a related decoupling shallow feature enhancement module to replace the first convolutional layer of ResNet, aiming to enhance the consistent low-frequency structural response across modalities while decoupling high-frequency cues for specific modalities; Specifically, the source image The input is fed into basic convolutional mapping and normalization operations to obtain the corresponding shallow features. ; Subsequently, in order to decouple and enhance modality-specific high-frequency information and share low-frequency representation, the related decoupling shallow feature enhancement module processes shallow features through two dedicated branches (i.e., high-frequency enhancement branch and low-frequency enhancement branch); S33: Shallow features After high-frequency branching, differential edge convolution and local structure modeling are used as embedded learnable network operations to extract high-frequency features. ; S34: For low-frequency enhancement branches, shallow features First, shallow features are processed by low-pass filtering to eliminate high-frequency noise and preserve robust shared low-frequency components, resulting in low-frequency features. It is then fed into the proposed dual-scale enhanced residual block for further enhancement, as shown below: ; Specifically, low-frequency characteristics The input is first fed into two parallel convolutional branches: a 3×3 branch to capture local structural details, and a 5×5 branch to aggregate broader neighborhood contextual information through a larger receptive field. The normalized 3×3 and 5×5 convolutional responses are then summed to produce intermediate multi-scale features. Subsequently, It is passed to the dual-path feature interaction unit, one path applies a fully connected projection, and then... One path activates the other fully connected projection without nonlinear activation, and the two projected features are fused through element-wise multiplication to achieve high-order nonlinear feature interaction. S35: To enhance feature representation by strengthening the correlation of low-frequency features while weakening the correlation of high-frequency features, this invention constructs a correlation-constrained loss adapted to cross-modal image fusion. Specifically, it defines the similarity relationship between two features using cosine similarity, as shown below: ; in This represents the pooling operation; in order to enhance the low-frequency correlation across modalities while constraining the high-frequency features of cross-modal images and the correlation between high-frequency and low-frequency features within the same modality, the correlation constraint loss consists of three parts, specifically; ; in , and It's about adjusting parameters. The aim is to maximize the correlation between low-frequency features across modalities. The aim is to minimize the correlation between low-frequency features across modalities and the correlation between high-frequency features across modalities. The aim is to minimize the correlation between low-frequency and high-frequency features within each modality; therefore, the three losses are specifically constructed as follows: ; in and This represents the low-frequency characteristics of visible light and infrared images, while and This indicates its corresponding high-frequency characteristics; furthermore, in and Absolute value constraints are introduced to maintain the relative independence of cross-modal high-frequency features and intra-modal low-frequency features; Specifically, positive correlation and negative correlation represent the same and opposite trends of change between two features, respectively. Both of these situations indicate that the features are strongly coupled. Therefore, minimizing the absolute correlation suppresses strong positive and negative correlations. It pushes the correlation to zero, reduces linear correlation, and enhances feature decoupling. S4: By using a masked proxy cross-attention layer, a good balance is achieved between modeling global cross-modal semantic relevance and preserving local fine structural details by introducing semantic mask-guided proxy cross-attention and adaptive local convolutional interaction; The specific process of S4 includes: S41: Due to the physical differences in the imaging mechanisms of infrared and visible light sensors, there are inherent domain differences between the source image features. Direct fusion performed in the original image domain often blurs fine-grained details and produces obvious artifacts in the fusion result. To mitigate these drawbacks, sufficient cross-modal feature interaction is usually implemented beforehand. This invention proposes a mask proxy cross-attention layer. Specifically, the enhanced features are decomposed by the relevant decoupled shallow feature enhancement module and encoded by the first two layers of ResNet to construct three pairs of infrared and visible light image features at different scales. ,in Then, infrared and visible light features of the same scale are input into the mask proxy cross-attention layer to fully interact with global semantics and local details; S42: For the global information interaction branch of the masked proxy cross-attention layer, this invention constructs mask-guided proxy features, embedding semantic constraints into cross-modal interactions, thereby prioritizing foreground region interactions rather than interactions from irrelevant non-critical regions. Specifically, the semantic mask generated from the graffiti prompt is first mapped by a convolutional layer to obtain the mask feature Fmask. Subsequently, the mask feature is used to calibrate the projection query features in the global proxy token construction. Semantic-guided feature refinement and enhancement are achieved by emphasizing meaningful foreground regions and reducing background interference, which can be described as follows: ; Where parameters Set to 0.5. Indicates the modal features used for querying; S43: Subsequently, by constructing a small number of proxy tokens As a global proxy representation, to avoid the high computational cost and noise interference caused by direct, dense pairwise correlation calculations on the original spatial features, under the guidance of mask constraints, the proxy token acts as a semantic mediator in cross-modal attention computation, transforming pixel-level dense interactions into a two-stage process of global agent aggregation and spatial location broadcasting. This mitigates the modality gap caused by low-level response differences, and can be expressed as follows: ; in and It is the learned relative positional deviation. Indicates the value used for attention calculation The function, in order to fully interact with the global semantic information of infrared and visible light images, uses infrared and visible light image features as bidirectional queries to perform global interaction and obtain updated global interactive features. ; S43: Although agent cross-attention excels at modeling long-range cross-modal dependencies, its global aggregation nature tends to discard distinctive local details unique to each modality. Therefore, the agent cross-attention layer also introduces a CNN-based local detail interaction branch to preserve high-frequency structural information. Specifically, it first connects multi-scale features along the channel dimension. ,in Indicates from the The features of a layer, the connection features are represented as Subsequently, Two inputs based on The weight generation branch is used to obtain mode-specific local weights for infrared and visible light features, respectively. and Meanwhile, through convolution, normalization, and Activate from Extracting local spatial information and further introducing learnable parameters. The contributions from the two modes of adaptive equilibrium, the local feature interaction process can be described as follows: ; in This represents element-wise multiplication, and finally, integrated local features are used. Corresponding features to the global Gating mechanisms to generate enhanced interactive features This design enables effective complementary feature learning between global semantics and local details, and finally updates the features after the interaction. The input is fed into the channel fusion module, which fuses the infrared and visible light image features updated at the same scale to obtain the corresponding fused features. ; S5: Through the multi-scale feature calibration module, deep fusion features and semantic masks generated by graffiti are used as prior guidance, and shallow fusion features are calibrated through spatial adaptive normalization mechanism, thereby enhancing the cross-scale consistency of key semantic regions while preserving local details. The specific process of S5 includes: S51: For infrared and visible image fusion tasks involving high-resolution input and multi-scale objects, single-scale fusion features are insufficient to balance infrared salient objects and visible image details due to feature differences between infrared and visible images. Multi-scale cross-modal fusion enables high-level features to encode infrared targets and global context, while low-level features can preserve visible textures and fine structures in low-level features, improving the model's generalization ability to complex real-world scenes. However, due to unpredictable object boundaries and severe feature degradation during network propagation, it is difficult to align low-resolution object features with their high-resolution counterparts. To address this issue, this invention proposes a multi-scale feature calibration module that uses mask information and deep features as semantic guidance to further modulate shallow fusion features. S52: Specifically, firstly, with Under guidance Taking calibration as an example, the graffiti is used to generate a semantic mask. And upsampling processing, generating with Soft masks with the same spatial size Subsequently, on Perform upsampling and The connectivity features are processed through convolution and normalization, followed by two independent 3 × 3 convolutional layers to predict the scaling factor. and shift factor , The function is used to... Limiting to the range [-1, 1] to achieve stable modulation, then shallow features are... Calibration, the process is as follows: ; in This represents element-wise multiplication. Finally, the calibrated features and the original features are adaptively combined using a soft mask. ; S53: After receiving the updated calibration In the same way, using masks and right Further modulation and calibration are performed to obtain the final fusion characteristics; S6: Construct a fused image of infrared and visible light images with clear details and rich semantics through a feature decoder; In this step, the final fused features obtained after modulation by the multi-scale feature calibration module are input into the feature decoder. Through step-by-step upsampling and feature recombination, a fused image with clear details and rich semantics is reconstructed. Specifically, S6 includes the following sub-steps: S61: Fusion Feature Input: The final fused feature map obtained by the multi-scale feature calibration module in S5 is input to the feature decoder. This fused feature map has fused the significant target thermal radiation response of the infrared image and the rich structural texture information of the visible light image. After semantic mask-guided calibration, the response of the foreground region is enhanced and the background interference is suppressed. S62: Stepwise upsampling and feature decoding: The feature decoder employs a progressive upsampling structure symmetrical to the encoder. In this embodiment, the decoder consists of multiple upsampling convolutional blocks, each of which includes: Bilinear interpolation upsampling layer is used to gradually restore the spatial size of the feature map to the original image resolution; Convolutional layers (preferably 3×3 convolutions) are used to smooth and reorganize the upsampled features to eliminate the checkerboard effect; Batch normalization layers and ReLU activation functions are used to accelerate convergence and introduce nonlinear expressive power; Skip connection is used to concatenate the high-resolution spatial detail features preserved in the corresponding layer of the encoder with the upsampled deep semantic features, thereby restoring spatial resolution while preserving fine-grained texture information. By upsampling step by step, the spatial size of the feature map is gradually restored from the lowest resolution to the original resolution of the input image (e.g., from 16×16 to 256×256 or 512×512). S63: Output Channel Mapping and Fusion Image Generation: In the final layer of the decoder, a 1×1 convolutional layer maps the multi-channel feature maps to a single-channel or three-channel output: If grayscale fusion output is used, the number of output channels is 1, and a grayscale fusion image is directly generated; If color fusion output is used, the number of output channels is 3, generating an RGB color fusion image; In this embodiment, since the visible light image has been converted into YCrCb channels and only the Y channel is used in the fusion network, the output of the fusion network is a single-channel luminance component. Then, the fused Y channel is merged with the Cr channel (chromatic red) and Cb channel (chromatic blue) retained in the original visible light image and converted back to the RGB color space to generate a color fusion image with a natural color appearance. This processing method ensures that the fusion image retains the significant thermal radiation information of the infrared target and inherits the natural color and texture details of the visible light image. S64: Output the blended image: After the above processing, a fused infrared and visible light image with clear details and rich semantics is output. This image is superior to a single source image in terms of salient target response, texture preservation, contrast and visual naturalness, and can be directly used for downstream vision tasks. As an alternative: In another embodiment, the feature decoder may use a transposed convolution-based upsampling method instead of bilinear interpolation upsampling to achieve a learnable upsampling process and further enhance the detail reconstruction capability. In this case, each upsampling convolution block in the decoder contains a deconvolution layer, a batch normalization layer and a ReLU activation function, and is combined with skip connections to fuse the spatial detail features of the corresponding layer of the encoder. In another embodiment, if the visible light image itself contains rich color information and the user has high requirements for color fidelity, the YCrCb channel conversion step can be omitted, and the RGB three channels can be directly input into the fusion network. The decoder directly outputs the three-channel RGB fused image to avoid color difference loss caused by color space conversion.

[0016] Example 2

[0017] Based on Example 1, in the infrared and visible light image fusion method of the present invention, the processing procedure of the dual-scale enhancement residual block in S34 is as follows: Low-frequency features are first fed into two parallel convolutional branches: a 3×3 convolutional branch to capture local structural details and a 5×5 convolutional branch to aggregate broader neighborhood contextual information through a larger receptive field. Adding the normalized 3×3 convolutional responses and 5×5 convolutional responses produces intermediate multi-scale features; Intermediate multi-scale features are passed to the dual-path feature interaction unit. One path is activated after applying a fully connected projection, while the other path performs another fully connected projection without nonlinear activation. The two projected features are fused through element-wise multiplication to achieve high-order nonlinear feature interaction.

[0018] Example 3

[0019] This invention has been extensively tested on three publicly available datasets: M3FD, RoadScene, and TNO; The M3FD dataset contains a total of 4200 pairs of infrared and visible light images, of which 3900 were used for training in this invention and the remaining 300 were used for official testing. To further verify the effectiveness of this invention, we performed cross-dataset evaluation on the RoadScene and TNO datasets using a model trained on M3FD. Furthermore, this invention employed six metrics, including spatial frequency (SF), entropy (EN), standard deviation (SD), visual fidelity (VIF), average gradient (AG), and edge information preservation. Comprehensive measurement and fusion results; These metrics evaluate fused images in terms of information richness, structural integrity, and perceptual fidelity. Generally, higher metric values ​​indicate better fusion performance and image quality. This invention is implemented on an NVIDIA RTX A6000 GPU based on PyTorch. We use a pre-trained ResNet as the backbone network, and the model is optimized using the Adam optimizer with a learning rate of [missing information]. The momentum parameter is set to [value] during the training phase. and With a batch size of 8, we applied data augmentation, including random cropping of 128×128 blocks and random rotation, to enhance the model's robustness and prevent overfitting. During the inference phase, the network processes the source images at their original resolution, preserving full spatial detail. During training and testing, the RGB visible images are first converted to the YCrCb color space, and then the visible and infrared images are converted to the YCrCb color space. Channels are used as network inputs. Furthermore, we compare CDN-SAA (the infrared and visible light image fusion method proposed in this application) with 15 performance-priority model methods to verify its effectiveness and generalization, including LRR (a novel representation learning-guided fusion network for infrared and visible light images); CDDFuse (correlation-driven bi-branch eigenvalue decomposition for multimodal image fusion); MaeFuse (an infrared and visible light image fusion method that utilizes pre-trained masked autoencoders to transfer full-level features through guided training); Text-IF (degradation-aware and interactive image fusion guided by semantic text); PSFusion (re-examining the necessity of image fusion in high-level vision tasks: a practical infrared and visible light image fusion network based on progressive semantic injection and scene fidelity); and BDLFusion (a multimodal image fusion and...). Its extended tasks include: Two-layer dynamic learning); SeAFusion (Image fusion in high-level visual task closed loop: Semantically perceptive real-time infrared and visible light image fusion network); EMMA (Equally variable multimodal image fusion); CAF (Elegance and precision: A compact, automatic, and flexible framework for multimodal image fusion and its applications); EVAFusion (Introducing human evaluation into infrared and visible light image fusion); SHIP (Exploring collaborative high-order interactions in infrared and visible light image fusion); SAGE (Cherishing every piece of general segmentation model information: A semantic prior utilization method for multimodal image fusion and its extended tasks); DCEvo (Discriminative cross-dimensional evolutionary learning for infrared and visible light image fusion); SIBA (Source image is the best attention for infrared and visible light image fusion); ISFL (Multimodal image fusion based on intervention-stabilized feature learning); In the following experiments, we used 3900 pairs of infrared and visible light image pairs from the M3FD training set as the training and validation sets for the model, and then tested the model's performance on the M3FD, RoadScene, and TNO test sets, respectively. Table 1 shows the results of comparing this invention with other state-of-the-art (SOTA) methods:

[0020] Note: The values ​​in the table are all measured results for each indicator. The larger the value of each indicator, the better the fusion effect. The English terms in Table 1 have the following meanings: M3FD: Multi-scene multimodal fusion dataset; RoadScene: Road scene dataset; TNO: Scientific research organization dataset; EN: Information entropy; SD: Standard deviation; SF: Spatial frequency; AG: Average gradient; Qabf: Edge information retention; VIF: Visual information fidelity; CDN-SAA: An infrared and visible light image fusion method proposed in this application; LRR: A novel representation learning-guided fusion network for infrared and visible light images; CDDFuse: A correlation-driven bi-branch feature decomposition for multimodal image fusion; MaeFuse: An infrared and visible light image fusion method that uses a pre-trained mask autoencoder to transmit full-level features through guided training; Text-IF: Degradation-aware and interactive image fusion guided by semantic text; PSFusion: A practical infrared and visible light image fusion network based on progressive semantic injection and scene fidelity; BDLFusion: A bi-layer dynamic learning method for multimodal image fusion and its extended tasks; Se AFusion: Semantically Aware Real-Time Infrared and Visible Image Fusion Network; EMMA: Equivariant Multimodal Image Fusion; CAF: A Compact, Automatic, and Flexible Framework for Multimodal Image Fusion and Its Applications; EVAFusion: Introducing Human Evaluation into Infrared and Visible Image Fusion; SHIP: Exploring Collaborative Higher-Order Interactions in Infrared and Visible Image Fusion; SAGE: A Semantic Prior Utilization Method for Multimodal Image Fusion and Its Extended Tasks; DCEvo: Discriminative Cross-Dimensional Evolutionary Learning for Infrared and Visible Image Fusion; SIBA: Optimal Attention for Infrared and Visible Image Fusion Based on Source Image; ISFL: Multimodal Image Fusion Based on Intervention-Stable Feature Learning. This invention achieves best or better experimental results on spatial frequency (SF), entropy (EN), standard deviation (SD), visual fidelity (VIF), and average gradient (AG) metrics on the Multi-Scene Multimodal Fusion Dataset (M3FD), RoadScene Dataset, and Test Set of the Scientific Research Organization Dataset (TNO). This indicates that the method proposed in this application can generate fused images with richer information, stronger contrast, clearer gradient changes, and more complete preservation of visual information. Specifically, higher information entropy (EN) and visual information fidelity (VIF) indicate that the fusion result retains more effective information, higher standard deviation (SD) indicates better contrast and target saliency, and higher spatial frequency (SF) and average gradient (AG) reflect richer texture details and edge gradient information in the fused image. These advantages are due to the explicit guidance of the graffiti-guided semantic mask in this invention to key regions. This mechanism helps the network pay more attention to foreground targets and salient structures. Therefore, this invention obtains a fused image with richer information and a clearer background in M3FD scenes with clear thermal targets and complex visible textures. However, this invention also has limitations in edge information preservation. The indicator yielded a relatively low value. More attention is paid to the edge consistency between the fused image and the source image. For example, the SIBA source image is the best attention for the fusion of infrared and visible light images, and it gets a higher value on this metric because their fusion result more directly preserves the original edge response of the source image. In contrast, the present invention performs adaptive fusion of foreground targets and significant structures through semantic mask-guided feature interaction and multi-scale feature calibration.

[0021] Example 4

[0022] An infrared and visible light image fusion system includes a semantic mask acquisition module, a correlation decoupling shallow feature enhancement module, a mask proxy cross-attention layer, a fusion module, a multi-scale feature calibration module, and a mask decoder; Semantic mask acquisition module: used to construct corresponding semantic masks from imported infrared and visible light image pairs and constructed graffiti prompts; Specifically, the semantic mask acquisition module first acquires the red graffiti marks drawn by the user on the infrared image using an image editing tool. Then, it converts the graffiti into discrete point set cues through skeleton extraction and discrete point sampling. Simultaneously, the module executes a cross-modal channel shuffling strategy, replacing the R channel of the visible light image with an infrared image of the same scene, while leaving the G and B channels unchanged, thus constructing a reconstructed cross-modal image. Finally, the module inputs the discrete point set and the reconstructed cross-modal image into a large visual segmentation model (such as the Segment Anything Model, SAM) to generate a semantic mask corresponding to the graffiti area. This semantic mask serves as a soft label, representing the spatial distribution probability of salient foreground objects in the image, providing prior semantic constraints for the subsequent fusion network. Correlation decoupling shallow feature enhancement module: used before integrating the Rsenet network, it performs shallow encoding on infrared and visible light images, decouples high-frequency and low-frequency feature information, and constructs three layers of multi-scale encoded features together with the first two layers of encoding of the integrated network through correlation decoupling loss constraint; Specifically, this module replaces the first convolutional layer of ResNet. After obtaining shallow features from the input source image through basic convolutional mapping and normalization operations, it processes the shallow features through two dedicated branches: a high-frequency enhancement branch and a low-frequency enhancement branch. The high-frequency enhancement branch uses differential edge convolution and local structure modeling as embedded learnable network operations to extract high-frequency features. The low-frequency enhancement branch first uses low-pass filtering to eliminate high-frequency noise and retain shared low-frequency components. Then, the obtained low-frequency features are input into a dual-scale enhancement residual block for enhancement. The dual-scale enhancement residual block contains two parallel 3×3 convolutional branches and a 5×5 convolutional branch, as well as a dual-path feature interaction unit. This module achieves full enhancement of low-frequency features through high-order nonlinear feature interaction. It also constructs a correlation constraint loss, defines feature similarity with cosine similarity, and achieves decoupling and enhancement of cross-modal features by maximizing the correlation of cross-modal low-frequency features, minimizing the correlation of cross-modal high-frequency features, and minimizing the correlation of low and high-frequency features within the same modality. Masked proxy cross-attention layer: Introduces semantically guided proxy cross-attention and local adaptive convolution to perform global semantic interaction and local detail interaction within the same layer on feature spectral pairs of infrared and visible light at three different scales; This module combines the enhanced features output from the related decoupled shallow feature enhancement module with the encoded features from the first two layers of ResNet to construct three pairs of infrared and visible light image features at different scales. Then, the infrared and visible light features at the same scale are input into the mask proxy cross-attention layer. This layer contains two parallel branches: the global information interaction branch constructs mask-guided proxy features as semantic mediators, and transforms pixel-level dense interactions into a two-stage process of global agent aggregation and spatial location broadcasting through a small number of proxy tokens. Global interactions are performed using infrared and visible light features as bidirectional queries. The local detail interaction branch connects multi-scale features along the channel dimension based on CNN, obtains modality-specific local weights through the weight generation branch, and extracts local spatial information through convolution, normalization, and activation. Learnable parameters are introduced to adaptively balance the contributions of the two modalities. The two branches are fused through a gating mechanism to achieve complementary feature learning of global semantics and local details. Fusion module: used to fuse infrared and visible light images of the same size after full interaction to obtain the feature spectrum of the final fused image at different scales; For the updated infrared and visible light image features output by the mask proxy cross-attention layer at the same scale, the fusion module merges them through channel fusion to generate the fused features corresponding to that scale, thereby obtaining fused image feature spectra at multiple different scales.

[0023] Multi-scale feature calibration module: Modulates deep fusion features using semantic masks and shallow fusion features, and enhances and enriches the semantic information of the fusion features through explicit semantic injection; This module first downsamples and upsamples the semantic mask generated by the graffiti to generate a soft mask with the same size as the shallow fusion feature space. Then, it upsamples the deep fusion features and connects them with the shallow fusion features. After convolution and normalization, the connected features are predicted by two independent 3×3 convolutional layers. The scaling factor and shift factor are limited to the range [-1,1] using the Tanh function. Subsequently, the shallow fusion features are calibrated using the scaling factor and shift factor. The calibrated features and the original features are adaptively combined using the soft mask to obtain the updated calibrated features. Finally, in the same way, the mask and the updated calibrated deep features are used to further modulate and calibrate another shallow fusion feature to obtain the final fusion feature. Mask decoder: Used to input the modulated fusion features into the mask decoder for decoding and outputting the fused image; Specifically, the mask decoder adopts a progressive upsampling structure symmetrical to the encoder, consisting of multiple upsampling convolutional blocks. Each upsampling convolutional block contains a bilinear interpolation upsampling layer, a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation function. It also uses skip connections to concatenate the high-resolution spatial detail features of the corresponding layer of the encoder with the upsampled deep semantic features. The last layer of the decoder maps the multi-channel feature map to a single-channel Y channel output through a 1×1 convolutional layer. Then, it is merged with the Cr and Cb channels retained in the original visible light image and converted back to the RGB color space to generate a color fusion image with a natural color appearance. In the semantic mask acquisition module, the R channel of the visible image is replaced with an infrared image of the same scene, while the G and B channels retain the visible structural and texture information to construct a cross-modal image; the reconstructed cross-modal image and sampling point cues are input into the visual segmentation large model to generate a semantic mask; The related decoupling shallow feature enhancement module processes shallow features through two dedicated branches: a high-frequency enhancement branch and a low-frequency enhancement branch. The high-frequency enhancement branch extracts high-frequency features using differential edge convolution and local structure modeling; The low-frequency enhancement branch is processed by low-pass filtering and then input into the dual-scale enhancement residual block for enhancement; this invention introduces prior information of human vision through graffiti guidance and uses a large visual language model to convert sparse graffiti prompts into semantic masks, thus avoiding the ambiguity inherent in language-driven guidance. By using a related decoupling shallow feature enhancement module to enhance and decouple the correlation of features at different frequencies, the representation of shallow features is improved and modal differences are reduced. A good balance is achieved between global semantic modeling and local detail preservation by using a mask proxy cross-attention layer; The multi-scale feature calibration module enhances the feature response in salient object regions by utilizing semantic masks, thus suppressing unreliable responses caused by feature degradation. The above modules work together to effectively improve the fusion capability of cross-modal features, and obtain fused images with prominent targets, clear structures and rich texture information.

[0024] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for fusing infrared and visible light images, characterized in that, Includes the following steps: S1: Import the image to be tested; S2: Use image editing tools to doodle on the region of interest in the infrared or visible light image, and input the doodle with the image after stitching the infrared and visible light image channels into the visual language big model to construct a semantic mask; S3: Introducing correlation decoupling loss, the high-frequency and low-frequency feature information between infrared and visible light image features is decoupled from each other and the feature expression is enhanced. It is also used to construct a three-layer cross-modal feature with the first two layers of the pre-trained convolutional neural network encoding features. S4: By introducing semantic mask-guided proxy cross-attention and adaptive local convolutional interactions, global cross-modal semantic relevance is modeled while preserving local fine-grained structural details; S5: Use deep fusion features and graffiti-generated semantic masks as prior guidance, and calibrate shallow fusion features through a spatial adaptive normalization mechanism to enhance cross-scale consistency of key semantic regions while preserving local details; S6: Construct a fused image of infrared and visible light images with clear details and rich semantics.

2. The infrared and visible light image fusion method as described in claim 1, characterized in that, The specific process of S2 includes: S21: Use red graffiti to draw on areas of interest in infrared images through image editing; S22: By extracting the skeleton and sampling discrete points, the graffiti prompts are sampled into discrete points with a certain shape, which serve as prompt information for the visual segmentation model. S23: Replace the R channel of the visible image with an infrared image of the same scene, while retaining the visible structural and texture information in the G and B channels to construct a cross-modal image. Input the reconstructed cross-modal image and sampling point cues into the large visual segmentation model to generate semantic masks for the subsequent fusion network.

3. The infrared and visible light image fusion method as described in claim 1, characterized in that, The specific process of S3 includes: S31: Input infrared and visible light images and semantic masks into the fusion network. For visible light images with RGB channels, convert them to YCrCb channels and only take the Y channel as the input of the visible light image in the fusion network. S32: Input the source image into basic convolution mapping and normalization operations to obtain the corresponding shallow features; the related decoupling shallow feature enhancement module processes shallow features through two dedicated branches: high-frequency enhancement branch and low-frequency enhancement branch; S33: Shallow features are processed through high-frequency branches, and differential edge convolution and local structure modeling are used as embedded learnable network operations to extract high-frequency features; S34: For the low-frequency enhancement branch, the shallow features are first processed by low-pass filtering to eliminate high-frequency noise and retain shared low-frequency components; the obtained low-frequency features are then input into the dual-scale enhancement residual block for enhancement. S35: Construct a correlation constraint loss, which defines the similarity relationship between two features using cosine similarity. The correlation constraint loss consists of three parts: maximizing the correlation between low-frequency features across modalities, minimizing the correlation between high-frequency features across modalities, and minimizing the correlation between low-frequency features and high-frequency features within each modality.

4. The infrared and visible light image fusion method as described in claim 3, characterized in that, In S34, the processing procedure for the dual-scale enhanced residual block is as follows: Low-frequency features are first fed into two parallel convolutional branches: a 3×3 convolutional branch to capture local structural details and a 5×5 convolutional branch to aggregate broader neighborhood contextual information through a larger receptive field. Adding the normalized 3×3 convolutional responses and 5×5 convolutional responses produces intermediate multi-scale features; Intermediate multi-scale features are passed to the dual-path feature interaction unit. One path is activated after applying a fully connected projection, while the other path performs another fully connected projection without nonlinear activation. The two projected features are fused through element-wise multiplication to achieve high-order nonlinear feature interaction.

5. The infrared and visible light image fusion method as described in claim 1, characterized in that, The specific process of S4 includes: S41: The enhanced features are decomposed by the related decoupling shallow feature enhancement module and encoded by the first two layers of ResNet to construct three pairs of infrared and visible light image features at different scales. The infrared and visible light features at the same scale are input into the mask proxy cross attention layer to fully interact with global semantics and local details. S42: For the global information interaction branch, the semantic mask generated from the graffiti prompt is first mapped by the convolutional layer to obtain mask features. The mask features are then used to calibrate the projection query features in the global proxy token construction, and to perform semantically guided feature refinement and enhancement. S43: Construct a small number of proxy tokens as global proxy representations. Under the guidance of mask constraints, the proxy tokens act as semantic mediators in cross-modal attention computation, transforming pixel-level dense interactions into a two-stage process of global agent aggregation and spatial location broadcasting. Global interactions are performed using infrared and visible light image features as bidirectional queries to obtain updated global interaction features. S44: The proxy cross-attention layer introduces a CNN-based local detail interaction branch, connecting multi-scale features along the channel dimension. The connected features are input into two weight generation branches to obtain modality-specific local weights for infrared and visible light features, respectively. Simultaneously, local spatial information is extracted through convolution, normalization, and activation. Learnable parameters are introduced to adaptively balance the contributions of the two modalities. An enhanced interactive feature is generated using a gating mechanism that integrates local and global corresponding features. The updated features after interaction are input into the channel fusion module, which fuses the infrared and visible light image features updated at the same scale to obtain the corresponding fused features.

6. The infrared and visible light image fusion method as described in claim 1, characterized in that, The specific process of S5 includes: S51: The semantic mask generated from the graffiti is downsampled and upsampled to generate a soft mask with the same size as the shallow fusion feature space. S52: Upsample the deep fusion features and connect them with the shallow fusion features. The connected features are processed by convolution and normalization, and then the scaling factor and shift factor are predicted through two independent 3×3 convolutional layers. S53: The shallow fusion features are calibrated using scaling and shift factors, and the calibrated features and the original features are adaptively combined using a soft mask to obtain the updated calibrated features; S54: In the same manner, the deep feature is further modulated and calibrated using a mask and updated calibration to obtain the final fused feature.

7. An infrared and visible light image fusion system, characterized in that, It includes a semantic mask acquisition module, a related decoupling shallow feature enhancement module, a mask proxy cross-attention layer, a fusion module, a multi-scale feature calibration module, and a mask decoder; Semantic mask acquisition module: used to construct corresponding semantic masks from imported infrared and visible light image pairs and constructed graffiti prompts; Correlation decoupling shallow feature enhancement module: used before integrating the Rsenet network, it performs shallow encoding on infrared and visible light images, decouples high-frequency and low-frequency feature information, and constructs three layers of multi-scale encoded features together with the first two layers of encoding of the integrated network through correlation decoupling loss constraint; Masked proxy cross-attention layer: Introduces semantically guided proxy cross-attention and local adaptive convolution to perform global semantic interaction and local detail interaction within the same layer on feature spectral pairs of infrared and visible light at three different scales; Fusion module: used to fuse infrared and visible light images of the same size after full interaction to obtain the feature spectrum of the final fused image at different scales; Multi-scale feature calibration module: Modulates deep fusion features using semantic masks and shallow fusion features, and enhances and enriches the semantic information of the fusion features through explicit semantic injection; Mask decoder: Used to input the modulated fusion features into the mask decoder for decoding and outputting the fused image.

8. The infrared and visible light image fusion system as described in claim 7, characterized in that, In the semantic mask acquisition module, the R channel of the visible image is replaced with an infrared image of the same scene, while the G and B channels retain the visible structural and texture information to construct a cross-modal image; The reconstructed cross-modal images and sampling point cues are input into a large visual segmentation model to generate a semantic mask.

9. The infrared and visible light image fusion system as described in claim 7, characterized in that, The related decoupling shallow feature enhancement module processes shallow features through two dedicated branches: a high-frequency enhancement branch and a low-frequency enhancement branch. The high-frequency enhancement branch extracts high-frequency features using differential edge convolution and local structure modeling; The low-frequency enhancement branch is processed by low-pass filtering and then input into the dual-scale enhancement residual block for enhancement.

10. The infrared and visible light image fusion system as described in claim 7, characterized in that, It also includes a computer storage medium storing a computer program executable by a processor, the computer program performing the infrared and visible light image fusion method as described in any one of claims 1-6.