Infrared and visible light image fusion method for prompting representation analysis

By using a cue-based parsing method, the content and style information of infrared and visible light images are separated to generate high-quality fused images. This solves the problem of poor image fusion effects in existing technologies and is applicable to fields such as night surveillance, autonomous driving, public safety, and military reconnaissance.

CN120953443AActive Publication Date: 2025-11-14DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511475475.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2025-11-14
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion models mix information from different modalities during the fusion process, resulting in poor fused image quality that cannot simultaneously meet the needs of human observation and intelligent vision systems.

Method used

A cue representation parsing method is adopted, which extracts cue vectors through a shared encoder, uses cosine similarity to calculate the parsing of content semantics and style semantics, and combines reconstruction loss and cosine similarity loss to separate the content and style information of infrared and visible light images. In the image reconstruction stage, a cross-attention mechanism is used to generate a high-quality fused image.

Benefits of technology

The generated fused images can satisfy both human visual perception and be efficiently utilized by intelligent vision systems, improving image quality and robustness in complex environments. They are suitable for fields such as nighttime surveillance, autonomous driving, public safety, and military reconnaissance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953443A_ABST
    Figure CN120953443A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning, and discloses an infrared and visible light image fusion method for prompt representation analysis, and the method comprises the steps: inputting paired infrared and visible light images into a shared encoder in a prompt extraction stage, and preliminarily extracting a prompt vector; in the prompt analysis stage, mutual analysis of content semantic prompt and style semantic prompt is restrained through two cosine similarity calculations; in the image reconstruction stage, the reconstruction loss and cosine similarity loss in the prompt analysis stage are utilized to restrain the two image reconstructors to prompt and reconstruct infrared and visible light images respectively by utilizing contents and respective styles. Through the above stages, pure content prompts are obtained, so that the pure content prompts can be utilized to guide the image fusion process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning technology and relates to a method for fusing infrared and visible light images for cue representation parsing. Background Technology

[0002] Prompt learning is a class of efficient transfer learning methods that has emerged in recent years with the rise of large-scale pre-trained models (PMs). Its core idea is to introduce manually designed or learnable prompts at the input stage to guide the large model to exhibit better adaptability and generalization ability in downstream tasks. Unlike traditional model fine-tuning methods, prompt learning does not require updating a large number of model parameters. Instead, it adjusts the input format or trains a small number of prompt parameters, enabling the pre-trained model to better perform specific tasks, thus significantly reducing computational and storage overhead.

[0003] The concept of cue learning first gained attention in the field of Natural Language Processing (NLP). Pre-trained models, such as BERT and GPT, which represent large-scale language models, have learned rich semantic and syntactic knowledge from massive amounts of text data. However, how to efficiently transfer this knowledge to downstream tasks has become a key issue.

[0004] As research has deepened, the concept of cue learning has been introduced into the field of computer vision and developed into Visual Prompt Learning (VPT). Within this framework, researchers embed learnable cue vectors into network structures such as the Visual Transformer, enabling models to efficiently adapt to different data distributions in tasks such as image classification, object detection, and semantic segmentation. Practice has shown that this method can achieve results close to or even exceeding full fine-tuning with only a few parameter updates, making it particularly suitable for scenarios with limited training samples or highly variable tasks. Simultaneously, cue learning has also demonstrated significant value in cross-modal tasks. By constructing a unified cue space, it can effectively promote feature alignment and information interaction between different modalities, thus achieving significant progress in tasks such as image-text retrieval, visual question answering, and image generation. In conclusion, cue learning, as an emerging technology with high parameter efficiency and strong transferability, is becoming an important development direction in computer vision and multimodal tasks, providing new technical paths for image processing, feature fusion, and generative modeling.

[0005] Multimodal image fusion technology is an important research direction in image processing and computer vision. Its goal is to effectively integrate multi-source image information with different sources and perceptual mechanisms to generate a more comprehensive fused image in terms of visual effects and semantic expression. Compared with single-modal images, the fused image can simultaneously retain the advantageous features of multiple sensors, enhancing both information expression capabilities and robustness in complex environments. Taking infrared and visible light images as examples, the former reflects the thermal radiation characteristics of objects and has strong target salience in nighttime, foggy, or low-light environments, but lacks fine texture and color information; the latter provides rich structural details and scene content, but its performance degrades under insufficient lighting or severe interference. By fusing these two types of images, the clarity of infrared targets and the richness of visible light details can be presented simultaneously in the same image, thereby improving the perception and recognition capabilities of human observation and machine vision systems.

[0006] Research on multimodal image fusion has evolved from traditional methods to deep learning approaches. Early methods were largely based on transform domain and filtering theory, with typical examples including methods based on Laplacian pyramids, wavelet transforms, and sparse representations. These methods typically decompose images into frequency or structural information at different scales and then reconstruct them using certain fusion rules. Although these methods are simple to implement and can preserve the features of the source image to some extent, the lack of modeling for high-level semantic information often leads to deficiencies in detail preservation and overall consistency in the fused image. With the development of computer vision, researchers have gradually introduced deep learning into image fusion tasks. Convolutional Neural Networks (CNNs), with their powerful feature extraction capabilities, became the mainstream framework for early deep learning fusion methods. Through end-to-end training, CNN models can automatically learn fusion rules and achieve good results in preserving details and highlighting targets. However, convolutional structures have limitations in capturing long-range dependencies and cross-modal relationships, which means that information loss or distortion may still occur in fused images in complex scenes.

[0007] In recent years, the emergence of new deep generative technologies such as Generative Adversarial Networks (GANs), Transformers, and Diffusion Models has brought new development opportunities to multimodal image fusion. GANs, through adversarial training mechanisms, enable generators to produce fused results that closely resemble the distribution of real images, demonstrating advantages in improving image clarity and realism. However, due to the instability of adversarial training and mode collapse, GANs have certain limitations in practical applications. Transformer-based fusion methods, utilizing their self-attention mechanism, can better capture the global dependencies between cross-modal features, thus achieving more reasonable feature interactions and fusion in complex scenes. Diffusion models, as a representative of generative models that have emerged in recent years, have achieved significant results in image generation and editing tasks thanks to their progressive denoising generation mechanism and more stable training process. Applying diffusion models to multimodal image fusion can not only improve the realism and diversity of fused images but also maintain the spatial consistency of different modalities during the generation process.

[0008] In "CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion," Z. Zhao et al. used Transformer and CNN to extract low-frequency global information and high-frequency local information respectively, and then fused and concatenated them into a fusion feature during the fusion process. In "Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion," X. Yi et al. used semantic text to guide image fusion to achieve degradation-aware and interactive image fusion tasks. However, their method simply concatenated the features of different modalities after attention calculation and then used text guidance to modulate the fusion features. All these methods directly fuse or concatenate the features of different modalities before proceeding to the next step, ignoring the mutual interference between different modalities, thus leading to suboptimal results. Summary of the Invention

[0009] For the task of infrared and visible light image fusion, the goal of this invention is to obtain high-quality fused images that are convenient for human observation and can be used in fields such as night surveillance, autonomous driving, public safety, and military reconnaissance. Current infrared and visible light image fusion models extract features from the infrared and visible light images separately during the fusion process, perform fusion operations at the feature level, and then reconstruct the fused features into a fused image. Because information from different modalities of visible light and infrared is mixed during the fusion process, the final fused image is of poor quality. To address this problem, this invention proposes an infrared and visible light image fusion method based on cue representation parsing. In the cue extraction stage, paired infrared and visible light images are input into a shared encoder to initially extract cue vectors. In the cue parsing stage, two cosine similarity calculations are used to constrain the mutual parsing of content semantic cues and style semantic cues. In the image reconstruction stage, reconstruction loss and the cosine similarity loss from the cue parsing stage are used to constrain two image reconstructors to reconstruct the infrared and visible light images using content and their respective style cues. After these stages, a clean content cue is obtained, which can be used to guide the image fusion process.

[0010] The technical solution of the present invention:

[0011] A method for fusing infrared and visible light images with suggested resolution, comprising the following steps:

[0012] (1) Prompt extraction stage:

[0013] Given paired infrared and visible light datasets ,in , Representing the infrared image dataset and the visible light image dataset respectively, where m represents the number of samples in the dataset; using a shared encoder The shared encoder extracts cue vectors and consists of a series of downsampling modules; paired infrared and visible light images are input into the shared encoder. Initially, the cue vector is extracted, and corresponding infrared cue vector and visible light cue vector are obtained;

[0014] (2) The prompt indicates the parsing stage:

[0015] The content semantic vectors and style semantic vectors in the obtained infrared cue vectors and visible light cue vectors are analyzed together: the infrared cue vectors and visible light cue vectors are respectively divided into content cue vectors and style cue vectors. Two cosine similarities are used to constrain the relationship between the two sets of content cue vectors and the two sets of style cue vectors, ensuring that the two sets of content cue vectors are close to each other and the two sets of style cue vectors are perpendicular to each other. This can be expressed by the formula:

[0016] (1)

[0017] (2)

[0018] (3)

[0019] (4)

[0020] in, Represents the content cue vector for visible light. Indicates the content cue vector of infrared. The style cue vector representing visible light. Indicates the style cue vector for infrared. This indicates the calculation of the norm of the cue vector. This indicates the calculation of cosine similarity. This represents the calculation of the cosine similarity between the content cue vectors of visible light and infrared light. This represents the calculation of the cosine similarity between the style cue vectors for visible light and infrared light. This represents the loss caused by constraint hint vectors moving closer together. This represents the loss where the constraint hint vectors are perpendicular.

[0021] (3) Image reconstruction stage:

[0022] After cosine similarity constraints, the content semantic vectors and style semantic vectors of the infrared cue vectors and visible light cue vectors are parsed. To ensure the completeness of the semantics contained in the infrared and visible light cue vectors, the two sets of content cue vectors and the two sets of style cue vectors are concatenated to obtain two sets of cue vectors containing complete semantics. The two complete cue vectors are then fed into the visible light feature decoder. and infrared feature decoder Image reconstruction; visible light feature decoder and infrared feature decoder Each module consists of a series of upsampling modules; the reconstruction loss is used to supervise the reconstruction of the image, expressed by the formula:

[0023] (5)

[0024] (6)

[0025] in, This represents the loss for constrained visible light image reconstruction. This represents the loss of the constrained infrared reconstructed image. This indicates the number of pixels contained in the image. Represents the pixels in a visible light image. Represents the pixels of the visible light image of the reconstructed image. Represents the pixels in an infrared image. Represents the pixels of the infrared image used to reconstruct the image;

[0026] (4) Image fusion stage:

[0027] First, the paired infrared and visible light images are input into the infrared image encoder IE2 and the visible light image encoder IE1, respectively, to extract infrared features. and visible light characteristics Then, the similarity vectors of corresponding features in the same regions of the paired infrared and visible light images are calculated; after normalizing the similarity vectors, the weights of the semantic content contained in different regions are obtained, denoted as... ;

[0028] Using a projection matrix, infrared features and visible light characteristics Convert to visible light query matrix Infrared query matrix Visible light bond matrix Infrared key matrix The conversion formula is as follows:

[0029] (7)

[0030] (8)

[0031] (9)

[0032] (10)

[0033] in, The projection matrix of the visible light query matrix. The projection matrix of the visible light bond matrix. The projection matrix of the infrared query matrix. The projection matrix representing the infrared bond matrix;

[0034] The hint vector obtained during the parsing phase is used to represent the content hint vector. Weights of content semantics obtained in the image fusion stage The value matrix is ​​obtained. This can be expressed as a formula:

[0035] (11)

[0036] Through the cross-attention mechanism, the value matrix Injected into the cross-attention weights, a fused feature without domain-specific style interference is obtained, expressed by the formula:

[0037] (12)

[0038] (13)

[0039] in, This represents the cross-attention feature of visible light. This represents the cross-attention feature of infrared radiation. This represents the Softmax activation function;

[0040] Then and spliced ​​into a feature , sent to the fusion decoder The fused image is obtained after training. The loss in the training process to obtain the fused image is the SSIM loss.

[0041] The beneficial effects of this invention are as follows: The infrared and visible light image fusion method based on cue representation parsing of this invention mitigates the impact of modal differences during image fusion by mutually parsing content semantic cues and style semantic cues, resulting in high-quality multimodal fused images. The goal of this invention is to obtain a fused image that both conforms to human visual perception and can be efficiently utilized by intelligent vision systems. This technology can be widely applied in various fields such as nighttime surveillance, autonomous driving, public safety, and military reconnaissance, providing key technical support for multimodal information fusion and intelligent perception systems. Attached Figure Description

[0042] Figure 1 It is a parsing architecture for content and style hints.

[0043] Figure 2 It is an infrared and visible light image fusion architecture guided by content semantic prompts. Detailed Implementation

[0044] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0045] Figure 1 The content and style cues represent the parsing architecture, first pairing the infrared image with the visible light image ( and Input to shared encoder Initially, cue vectors are extracted. Then, the content cue vector and style cue vector are parsed using formulas (1) and (2). Finally, the parsed content cue vector and style cue vector are concatenated and input into their respective visible light feature decoders. and infrared feature decoder To reconstruct infrared and visible light images ( and After the first phase of training, the shared encoder It has the ability to parse content semantic cues and style semantic cues, and has also obtained the content cue vectors needed for subsequent processing. .

[0046] Figure 2 It is a content semantic prompt-guided infrared and visible light image fusion architecture that inputs paired infrared and visible light images into infrared image encoder IE2 and visible light image encoder IE1, respectively, to extract infrared features. and visible light characteristics Then, the similarity vectors of corresponding features in the same regions of the paired infrared and visible light images are calculated; after normalizing the similarity vectors, the weights of the semantic content contained in different regions are obtained, denoted as... Using content hint vectors Weight of content semantics The value matrix can be obtained through formula (11). Using a projection matrix, infrared features are... and visible light characteristics Convert to visible light query matrix Infrared query matrix Visible light bond matrix Infrared key matrix AND-value matrix Cross-attention calculation is performed together, as shown in formulas (12) and (13), to obtain fused features without domain-specific style interference. and Then, these features are concatenated into a single feature and fed into the fusion decoder. Obtain the fused image The loss in the training process of obtaining the fused image is the SSIM loss.

Claims

1. A method for fusing infrared and visible light images with cue representation analysis, characterized in that, The steps are as follows: (1) Prompt extraction stage: Given paired infrared and visible light datasets ,in , Representing the infrared image dataset and the visible light image dataset respectively, where m represents the number of samples in the dataset; using a shared encoder The shared encoder extracts cue vectors and consists of a series of downsampling modules; paired infrared and visible light images are input into the shared encoder. Initially, the cue vector is extracted, and corresponding infrared cue vector and visible light cue vector are obtained; (2) Hints indicate the parsing phase: The content semantic vector and style semantic vector in the obtained infrared cue vector and visible light cue vector are mutually analyzed: the infrared cue vector and visible light cue vector are respectively divided into content cue vector and style cue vector, and two cosine similarities are used to constrain the relationship between the two sets of content cue vectors and the two sets of style cue vectors, so that the two sets of content cue vectors are close to each other and the two sets of style cue vectors are perpendicular to each other. (3) Image reconstruction stage: The two sets of content cue vectors and the two sets of style cue vectors are concatenated to obtain two sets of cue vectors containing complete semantics. The two sets of complete cue vectors are then fed into the visible light feature decoder. and infrared feature decoder Reconstructing the image; Visible light feature decoder and infrared feature decoder Each module consists of a series of upsampling modules; the reconstruction loss is used to supervise the reconstruction of the image. (4) Image fusion stage: First, the paired infrared and visible light images are input into the infrared image encoder IE2 and the visible light image encoder IE1, respectively, to extract infrared features. and visible light characteristics Then, the similarity vector of the corresponding features of the same region in the paired infrared image and the visible light image is calculated; After normalizing the similarity vector, the weights of the semantic content contained in different regions are obtained, denoted as... ; Using a projection matrix, infrared features and visible light characteristics Convert to visible light query matrix Infrared query matrix Visible light bond matrix Infrared key matrix The conversion formula is as follows: (7) (8) (9) (10) in, The projection matrix of the visible light query matrix. The projection matrix of the visible light bond matrix. The projection matrix of the infrared query matrix. The projection matrix representing the infrared bond matrix; The hint vector obtained during the parsing phase is used to represent the content hints. Weights of content semantics obtained in the image fusion stage The value matrix is ​​obtained. This can be expressed as a formula: (11) Through the cross-attention mechanism, the value matrix Injected into the cross-attention weights, a fused feature without domain-specific style interference is obtained, expressed by the formula: (12) (13) in, This represents the cross-attention feature of visible light. This represents the cross-attention feature of infrared radiation. This represents the Softmax activation function; Then and spliced ​​into a feature , sent to the fusion decoder The fused image is obtained after training. The loss in the training process to obtain the fused image is the SSIM loss.

2. The infrared and visible light image fusion method for prompt representation analysis according to claim 1, characterized in that, Step (2) can be expressed by the formula: (1) (2) (3) (4) in, Represents the content cue vector for visible light. Indicates the content cue vector of infrared. The style cue vector representing visible light. Indicates the style cue vector for infrared. This indicates the calculation of the norm of the cue vector. This indicates the calculation of cosine similarity. This represents the calculation of the cosine similarity between the content cue vectors of visible light and infrared light. This represents the calculation of the cosine similarity between the style cue vectors for visible light and infrared light. This represents the loss caused by constraint hint vectors moving closer together. This represents the loss when the constraint hint vectors are perpendicular to each other.

3. The infrared and visible light image fusion method for prompt representation analysis according to claim 2, characterized in that, Step (3) can be expressed by the formula: (5) (6) in, This represents the loss for constrained visible light image reconstruction. This represents the loss of the constrained infrared reconstructed image. This indicates the number of pixels contained in the image. Represents the pixels in a visible light image. Represents the pixels of the visible light image of the reconstructed image. Represents the pixels in an infrared image. Represents the pixels of the infrared image from which the image is reconstructed.

Citation Information

Patent Citations

  • Visible light and infrared image fusion target tracking method based on visual prompt learning

    CN119888569A

  • Visible light-thermal infrared image semantic segmentation method and system driven by plug-and-play prompt

    CN120411499A