Semantic prior guidance infrared and visible light image fusion method, system and device and storage medium
By using pre-trained visual models to extract semantic priors and apply them in fusion networks, the problem of infrared and visible image fusion methods in the prior art relying on manual labels and lack of adaptability is solved, and more efficient multimodal information integration and fusion effect is achieved.
Patent Information
- Application Number
- CN202510237254.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-06-17
AI Technical Summary
In the prior art, infrared and visible image fusion methods rely on manual labels, lack adaptability, and are difficult to adapt to the fusion needs of different objects.
Using a semantic prior boot method based on pretrained vision models, global and local semantic priors are extracted through CLIP and DINO models, and these priors, adaptively modulated modulation of modular specific features of infrared and visible light inputs in semantic adaptive fusion networks, and finally generating a fusion image through the visual feature decoder.
Reduce dependence on manual labels, enhance adaptive integration of multimodal information, improve semantic consistency and detail retention of fusion images, and has broad application prospects.
Smart Images

Figure CN120164067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a semantic prior-guided infrared and visible light image fusion method based on a pre-trained vision model. The present invention also relates to the semantic prior-guided infrared and visible light image fusion system, device and storage medium. Background Art
[0002] In the field of infrared and visible light image processing, fusion technology aims to generate more accurate and reliable results by integrating the advantages of the two image modalities. Infrared images can capture the thermal radiation characteristics of a scene, while visible light images provide rich texture and detail information. However, how to effectively combine the characteristics of the two modalities is the main challenge faced by current technologies.
[0003] Most traditional image fusion methods are based on manually designed features and lack a deep understanding of the global and local semantics of images, resulting in unsatisfactory detail retention and semantic consistency in the fusion results. In recent years, deep learning models, especially pre-trained vision models such as CLIP and DINO, have demonstrated excellent capabilities in understanding image semantics and visual features. Applying these models to infrared and visible light image fusion can significantly improve the fusion accuracy and semantic consistency.
[0004] Due to theoretical and technical limitations, neither visible light images nor infrared images can comprehensively describe the imaging scene. Visible light images capture rich texture details but are sensitive to lighting changes, while infrared images provide thermal information that is insensitive to lighting but lack sufficient texture. Image fusion aims to adaptively integrate the complementary features of these modalities, merge semantic information to enhance human perception and support advanced vision tasks. By combining the advantages of visible light and infrared images, the fusion process can generate more information-rich images, thereby improving the understanding of the scene by humans and machines. This method has been widely applied in fields such as object tracking, object detection, and security monitoring. However, achieving true adaptive fusion that captures high-level semantic information remains challenging.
[0005] In recent years, the focus of image fusion research has been on enhancing high-level semantic representations. Existing semantic-driven methods improve semantic information through training on high-level tasks, but they cannot adaptively fuse different semantic targets, limiting flexibility and detail richness.
[0006] Related patent literature: CN111861960A discloses an infrared and visible light image fusion method, which is based on variational method and local gradient similarity, and can generate a fusion image that retains more pixel gradient and intensity information. Taking the registered infrared and visible light images as the background and aiming to obtain a more informative fusion image, the image fusion problem is studied. First, the fusion gradient of the source image is calculated by using the structure tensor, and the local gradient similarity is used to make the fusion gradient direction more accurate; secondly, according to the comparison of pixel intensities, the source image is reconstructed into a saliency map and a non-saliency map, and a weight map for discriminating and retaining the effective details of the non-saliency map is calculated; furthermore, based on the gradient features and pixel intensity information of the source image, an image fusion model is established; finally, the variational method is used to solve the optimization model to obtain the fusion image.
[0007] The above technologies do not give specific guidance on how the present invention solves the limitations of the existing fusion methods that rely on manual labels and the problem of difficult adaptive fusion for different objects. Summary of the Invention
[0008] The purpose of the present invention is to overcome the deficiencies of the above technologies and provide a semantic prior-guided infrared and visible light image fusion method to solve the limitations of the existing fusion methods that rely on manual labels and the problem of difficult adaptive fusion for different objects.
[0009] For this reason, another purpose of the present invention is to provide a semantic prior-guided infrared and visible light image fusion system, device and storage medium for realizing the above fusion method.
[0010] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0011] A semantic prior-guided infrared and visible light image fusion method (based on a pre-trained visual model), the technical solution thereof is that it includes the following steps:
[0012] S100: Obtain a public dataset, including MFD, MSRS, LLVIP, FMB, and divide it into a training set and a test set;
[0013] S200: Establish a semantic prior-guided infrared and visible light image fusion model based on a pre-trained visual model, that is, a deep learning network model, and use the public dataset to pre-train the deep learning network model to generate an initial recognition model, wherein the deep learning network model is composed of three parts: a semantic prior generation module, a semantic adaptive fusion network, and a visual feature decoder;
[0014] S300: Extract global and local semantic priors using the pre-trained models CLIP and DINO, then use these priors in the semantic adaptive fusion network, and then guide the fusion process through the semantic adaptive fusion module and the deep semantic adaptive fusion module, adaptively modulate the modality-specific features from the infrared and visible light inputs, and finally the visual decoder reconstructs the fused features into the final output image;
[0015] S400: Conduct experiments using the public dataset, obtain experimental results through the divided dataset and test set, and quantitatively evaluate the fusion performance using multiple (which can be 6) statistical metrics.
[0016] In the above technical solution, the preferred technical solution may be that in the public dataset of step S100, the MSRS and FMB datasets are used to evaluate the performance of the deep learning model in the semantic segmentation task, and the LLVIP dataset is also used to test the ability of the deep learning model in the object detection task; the number of image pairs for testing in these data such as MFD, MSRS, LLVIP, and FMB are 300, 361, 300, and 280 respectively; the MSRS dataset includes 1083 pairs of images for training the model. In step S200, the deep learning network model processes the infrared and visible light images through semantic priors, extracts global and local semantic priors using the pre-trained models CLIP and DINO, and then uses these priors in the semantic adaptive fusion network, where the fusion process is guided by the semantic adaptive fusion module and the deep semantic adaptive fusion module, and the modality-specific features from the infrared and visible light inputs are adaptively adjusted. This adaptive fusion is used to enhance the semantic consistency and detail retention in the fused image. Finally, the visual feature decoder reconstructs the fused features into the final output image. In step S200, the semantic prior generation module is responsible for generating semantic priors from the paired visible light and infrared images. Specifically, using the pre-trained visual models, CLIP and DINO, with fixed parameters, visual features are independently extracted from the visible light and infrared inputs. The process is as follows:
[0017]
[0018] where and θ c and θ d respectively represent the image encoders from CLIP and DINO, and the visible light and infrared images are represented by I vi and I ir respectively. It is proposed to combine the semantic features of CLIP and DINO through the correlation pool. For clarity, F c represents the CLIP semantic features extracted from the visible light and infrared images F c vi and Fd Represent DINO semantic features similar to extraction
[0019] First, calculate F d The cosine similarity of each patch of to generate an affinity graph, denoted as ξ. Then use this affinity graph on F c to perform correlation pooling, thereby generating a new set of semantic features F cd , which encapsulates global and local semantic details. The process of performing correlation pooling is as follows:
[0020]
[0021] In the formula, is the value of the spatial position (i, j) and ξ ij is the affinity weight at the same position . This weight reflects the semantic similarity between the patches indexed by (i, j). ε is a small constant to prevent division by zero. As shown in the semantic prior generation component of Figure 2 , the CLIP semantic features extracted from the source image are subjected to correlation pooling operations. This process generates semantic features with enhanced correlation, denoted as
[0022] In the above technical solution, a preferred technical solution may further be that in step S200, in the component of the semantic adaptive fusion network, the paired visible light and infrared images are fed into the image encoder, and a Transformer-based block is used as the basic feature extractor. The feature extraction process can be described as follows:
[0023]
[0024] where I = 2, 3, 4, tii and Tiir respectively represent the feature extractors of visible light and infrared images. For i = 1, 2, 3, the CSAF module is used for direct fusion However, when i = 4, fusion is performed through the proposed PSAF module, which uses cross-attention before the CSAF module and self-attention after it.
[0025] In the above technical solution, a preferred technical solution may further be that the semantic prior-guided infrared and visible light image fusion method further includes: (recognizing that depth features encapsulate rich contextual semantic information, the depth semantic adaptive fusion module enhances the interaction between the two modalities by combining cross-attention through exchanging the queries Q of the two modalities. Next), in the (general) semantic adaptive fusion module, the input feature map is subjected to non-parametric instance normalization, and then the semantic prior is mapped into the feature space to generate spatially adaptive modulation parameters. Subsequently, the output of the depth semantic adaptive fusion module is processed by a feature decoder composed of Transformer-based blocks, and then the obtained features are concatenated with the feature map output by the general semantic adaptive fusion module of the previous layer, and then convolution is performed. This series of operations is repeated multiple times to obtain the integrated visual feature F fu That is:
[0026] F fu ={Conv(C(T f (F i '),F i-1 ))} r ,
[0027] where T f represents the feature decoder, C(.) represents concatenation along the channel dimension, and {.} r represents multi-layer repetition. Note that the upsampling of the features processed by T f needs to correspond to the downsampling performed by the feature extractor
[0028] Using a visual feature decoder, finally, the fused visual feature is passed through the Restomer decoder to obtain the fused image
[0029] The semantic prior-guided infrared and visible light image fusion method further includes a calculation method for the loss function:
[0030] Using the source visible light image I vi , infrared image I ir and the final fused image I f as the input of the fusion loss function L f , and this loss function consists of the following parts: structural similarity (SSIM) loss, (maximum) gradient loss, intensity loss, and color consistency loss. Each component of the loss function is as follows:
[0031] Calculation method of structural similarity loss: Using the structural similarity loss to evaluate the similarity between the fused image and the original source image to ensure that the fused image retains the structural integrity of the input image, which is expressed as:
[0032]
[0033] Calculation method of color consistency loss (color difference): To judge the degree of color loss and ensure that the fused image and the source visible image maintain the consistency of color information, so as to maintain the naturalness of the image and the accuracy of color representation in the fused image, especially when dealing with scenes with rich colors, it is expressed as:
[0034]
[0035] In the formula, T cbcr
[39] represents the transfer function from RGB to CbCr;
[0036] Calculation method of Intensity Loss: This loss measures the pixel-level difference between the fused image and the source image. Therefore, the intensity loss is defined as:
[0037]
[0038] Calculation method of Gradient Loss: The purpose of gradient loss is to ensure that the fused image contains rich texture information by retaining the maximum edges of the two source images, which is expressed as:
[0039]
[0040] Calculation method of total loss, i.e., overall loss: The overall loss function is the weighted sum of the above loss terms, and its formula is:
[0041]
[0042] Where α, α ζ and α z are hyperparameters that control the trade-off of the fusion loss function.
[0043] A semantic prior-guided infrared and visible light image fusion system. The fusion system uses the semantic prior-guided infrared and visible light image fusion method. The technical solution is that the fusion system includes:
[0044] A semantic prior generation module, which is responsible for generating semantic priors from paired visible light and infrared images. Specifically, it uses pre-trained vision models CLIP and DINO, with fixed parameters, to independently extract visual features from visible light and infrared inputs;
[0045] Semantic Adaptive Fusion Network Module. In this module, paired visible light and infrared images are fed into an image encoder. A Transformer-based block is used as the basic feature extractor. The extracted features undergo general semantic adaptive fusion and deep semantic adaptive fusion to fully utilize the complementary properties in the multi-modal setting, and finally refined output features are obtained.
[0046] Visual Feature Decoder Module. In this module, the fused visual features pass through a Restomer decoder to obtain the fused image, which is the final output of this method.
[0047] Object Detection Module. The main function of the object detection module is to automatically identify and locate target objects from the input image, and at the same time generate the bounding boxes and class labels of each target. Through the feature extraction network, this module can analyze the multi-scale features of the image and effectively capture global and local features using deep learning methods (such as convolutional neural networks or Transformer architectures). On this basis, the object detection module predicts the target category through the classification branch and precisely locates the target position through the regression branch, ensuring that the output results have high precision and robustness and can be widely applied in fields such as autonomous driving, video surveillance, and scene analysis.
[0048] An electronic device (i.e., a semantic prior-guided infrared and visible light image fusion device), the technical solution of which is that it includes: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, the steps of the semantic prior-guided infrared and visible light image fusion method are implemented.
[0049] A computer-readable storage medium, the technical solution of which is that a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the semantic prior-guided infrared and visible light image fusion method are implemented.
[0050] The present invention provides a semantic prior-guided infrared and visible light image fusion method based on a pre-trained vision model, which combines the ability of CLIP to integrate visual features with human knowledge in natural language and the ability of DINO to cluster semantically similar features, creating global and local semantic priors for a comprehensive and label-free understanding of the source images. These priors guide the fusion process through a Semantic Adaptive fusion Network, achieving adaptive and semantic-aware fusion that highlights modality-specific features. Finally, a visual feature decoder synthesizes the fused image, capturing key semantic details from each source. By leveraging robust, label-free semantic priors, the present invention enables a deeper understanding of infrared and visible light source images, thereby achieving adaptive fusion of fundamental features across modalities. Experimental results show that the present invention achieves adaptive fusion of multimodal information through its semantic adaptive fusion network, retaining key details of various semantic objects under comprehensive semantic understanding. Extensive experiments demonstrate that the present invention reduces the dependence on manual labels and enhances the adaptive integration of multimodal information, providing good performance in terms of visual quality and visual tasks, and having broad application prospects and popularization value.
[0051] Compared with the existing technologies, the features and beneficial effects of the present invention are as follows:
[0052] The semantic prior-guided infrared and visible light image fusion method based on a pre-trained vision model provided by the present invention integrates the semantic capabilities of pvm into the field of image fusion. This method uses CLIP and DINO encoders to extract and synthesize visual semantic features from source images, forming robust semantic priors that capture global and local context information. The method of the present invention includes a designed semantic adaptive fusion network that combines general and deep semantic adaptive fusion (C / PSAF) modules. CSAF uses spatial modulation to achieve adaptive multimodal fusion, while PSAF uses cross-modal attention to achieve deeper semantic integration. These modules work together to dynamically adjust modality-specific information for different semantic objects, achieving precise and context-aware fusion. This method reduces the dependence on manual labels, a common limitation of existing fusion methods, and enhances the adaptive integration of multimodal information.
[0053] In summary, the present invention provides a semantic prior-guided infrared and visible light image fusion method, system, device, and storage medium, which solves the limitations of existing fusion methods that rely on manual labels and the problem of difficulty in adapting to the adaptive fusion of different objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flowchart of the semantic prior-guided infrared and visible light image fusion method based on a pre-trained vision model of the present invention.
[0055] Figure 2 This is the processing flowchart for building and training the deep network model of the semantic prior-guided infrared and visible light image fusion method based on the pre-trained vision model of the present invention.
[0056] Figure 3 This is the flowchart of the model processing steps of the semantic prior-guided infrared and visible light image fusion method based on the pre-trained vision model of the present invention.
[0057] Figure 4 This is the model architecture diagram of the semantic prior-guided infrared and visible light image fusion method based on the pre-trained vision model of the present invention.
[0058] Figure 5 This is the schematic diagram of the functional modules of the semantic prior-guided infrared and visible light image fusion method based on the pre-trained vision model of the present invention. Detailed implementation manners
[0059] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on these embodiments, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of the present invention.
[0060] Embodiment 1: As Figure 1 shown, the semantic prior-guided infrared and visible light image fusion method (based on the pre-trained vision model) of the present invention includes the following steps:
[0061] S100: Obtain a public dataset, including MFD, MSRS, LLVIP, FMB, and divide it into a training set and a test set. The division ratio of the training set to the test set is 7:3.
[0062] S200: Establish a semantic prior-guided infrared and visible light image fusion model based on the pre-trained vision model, that is, a deep learning network model, and use the public dataset to pre-train the deep learning network model to generate an initial recognition model. Among them, the deep learning network model consists of three parts: a semantic prior generation module, a semantic adaptive fusion network, and a visual feature decoder.
[0063] S300: Use the pre-trained models CLIP and DINO to extract global and local semantic priors, then use these priors in the semantic adaptive fusion network, and then guide the fusion process through the semantic adaptive fusion module and the deep semantic adaptive fusion module, adaptively modulate the modality-specific features from the infrared and visible light inputs, and finally the visual decoder reconstructs the fused features into the final output image.
[0064] S400: Conduct experiments using the public dataset, obtain experimental results through the divided dataset and test set, and quantitatively evaluate the fusion performance using multiple (six) statistical metrics.
[0065] The semantic prior-guided infrared and visible image fusion method based on a pre-trained vision model provided in this embodiment integrates the semantic capabilities of the pvm into the field of image fusion. This method uses CLIP and DINO encoders to extract and synthesize visual semantic features from source images to form a robust semantic prior that captures global and local context information. This method includes a designed semantic adaptive fusion network that combines general and deep semantic adaptive fusion (C / PSAF) modules. CSAF uses spatial modulation to achieve adaptive multimodal fusion, while PSAF uses cross-modal attention to achieve deeper semantic integration. These modules work together to dynamically adjust modality-specific information for different semantic objects to achieve precise and context-aware fusion. This method reduces the dependence on manual labels, a common limitation of existing fusion methods, and enhances the adaptive integration of multimodal information. It is instructive in fields such as object detection.
[0066] In the public dataset of step S100, the MSRS and FMB datasets are used to evaluate the performance of the deep learning model in semantic segmentation tasks, and the LLVIP dataset is also used to test the ability of the deep learning model in object detection tasks; the number of image pairs used for testing in MFD, MSRS, LLVIP, and FMB are 300, 361, 300, and 280 respectively; the MSRS dataset includes 1083 pairs of images for training the model.
[0067] In steps S200 and S300, design an algorithm for semantic prior-guided infrared and visible image fusion based on a pre-trained vision model, build a deep neural network model, and use the public dataset to pre-train the deep learning model to generate an initial recognition model, where the deep learning model consists of three parts: a semantic prior generation module, a semantic adaptive fusion network, and a visual feature decoder. In step S200, the deep learning network model processes infrared and visible images through semantic prior generation, extracts global and local semantic priors using the pre-trained models CLIP and DINO, and then uses these priors in the semantic adaptive fusion network, where the fusion process is guided by the semantic adaptive fusion module and the deep semantic adaptive fusion module to adaptively adjust modality-specific features from infrared and visible inputs. This adaptive fusion is used to enhance semantic consistency and detail retention in the fused image. Finally, the visual feature decoder reconstructs the fused features into the final output image.
[0068] Specifically, in this embodiment, an algorithm for semantic prior-guided infrared and visible light image fusion applicable to pre-trained vision models should be designed according to the scenario requirements. This method integrates the semantic capabilities of pvm into the field of image fusion. Referring to the above discussion, this method reduces the dependence on manual labels, which is a common limitation of existing fusion methods, and enhances the adaptive integration of multimodal information.
[0069] Furthermore, the architecture designed in this embodiment for collaborative dense scene text detection and recognition regards the text detection task as a set prediction problem, and iteratively updates the text queries in the text detector to achieve information refinement in the detection stage and loss gradient backflow in the recognition stage, optimizing the collaborative architecture of dense scene text detection and recognition end-to-end. In recent years, the focus of image fusion research has been on enhancing high-level semantic representations. Existing semantic-driven methods improve semantic information through training on high-level tasks, but they cannot adaptively fuse different semantic targets, limiting flexibility and detail richness. This method adaptively fuses and dynamically adjusts according to the characteristics of each target, selectively emphasizing relevant details.
[0070] To achieve this level of adaptive fusion in this embodiment, a large-scale pre-trained vision model (pvm) with powerful semantic understanding capabilities is utilized. The emergence of pvm has revolutionized the application of computer vision in various fields. Specifically, CLIP aligns visual features with human knowledge in natural language, enhancing global semantic understanding and improving tasks such as object detection and semantic segmentation. At the same time, DINO is good at clustering semantically similar features, allowing for fine-grained local feature extraction, which is crucial for unsupervised object discovery. Recognizing these complementary advantages, it is recommended to use CLIP for macroscopic-level understanding and DINO for capturing fine textures, creating global and local semantic priors to guide the fusion process. By using these priors, the model adaptively extracts semantically meaningful features from different image modalities.
[0071] In this embodiment, the deep learning network model uses a pre-trained vision model in semantic prior generation. Thanks to large-scale datasets and carefully designed pre-training strategies, the pre-trained vision model (pvm) has shown good potential in fields such as classification, vision-language tasks, and generation tasks. Among them, CLIP is trained by contrastive learning on a wide range of vision-language pairs, and is good at extracting text and global image features to achieve robust semantic alignment. Text-IF utilizes the text-driven semantic function of CLIP to achieve text-guided image fusion by coupling text and image features. Similarly, DINO is trained through an unsupervised teacher-student framework without a labeled dataset, can capture fine-grained textures, and exhibits strong local spatial semantics. Its ability to retain detailed features makes it very suitable for tasks that require complex spatial understanding. Inspired by these advancements, the goal of this embodiment is to integrate the unique semantic extraction advantages of CLIP and DINO into image fusion. By integrating the global semantic understanding of CLIP and the local detail capture of DINO, the model of this embodiment is guided to learn an adaptive fusion strategy from a high-level semantic perspective.
[0072] In this embodiment, the deep learning network model is divided into three parts: semantic prior generation, semantic adaptive fusion network, and visual feature decoder. First, the deep learning model processes infrared and visible light images through semantic prior generation, and uses the pre-trained models CLIP and DINO to extract global and local semantic priors. Then these priors are used in the semantic adaptive fusion network, where the General Semantic Adaptive Fusion (CSAF) module and the Progressive Semantic Adaptive Fusion (PSAF) module guide the fusion process, adaptively modulating modality-specific features from infrared and visible light inputs. This adaptive fusion enhances semantic consistency and detail retention in the fused image. Finally, the visual feature decoder reconstructs the fused features into the final output image.
[0073] In step S200, the semantic prior generation module is responsible for generating semantic priors from paired visible light and infrared images. Specifically, using pre-trained vision models, CLIP and DINO, with fixed parameters, visual features are independently extracted from visible light and infrared inputs. The process is as follows:
[0074]
[0075] where and θ c and θ d respectively represent the image encoders from CLIP and DINO, and the visible light and infrared images are represented by I vi and I ir respectively. It is proposed to combine the semantic features of CLIP and DINO through correlation pooling. For clarity, F cDenote the CLIP semantic features \(F\) extracted from visible and infrared images c vi and \(F\) d denote the DINO semantic features extracted similarly
[0076] First, calculate the cosine similarity of each patch of \(F\) d to generate an affinity graph, denoted as \(\xi\). Then use this affinity graph to perform correlation pooling on \(F\) c to produce a new set of semantic features \(F\) cd which encapsulates global and local semantic details. The process of performing correlation pooling is as follows:
[0077]
[0078] where, is the value at spatial location \((i, j)\) \(\xi\) ij is the affinity weight at the same location . This weight reflects the semantic similarity between the patches indexed by \((i, j)\). \(\epsilon\) is a small constant to prevent division by zero. As shown in the semantic prior generation component of Figure 2 , the CLIP semantic features extracted from the source image are subjected to correlation pooling operations. This process produces semantic features with enhanced correlation, denoted as
[0079] In step S200, in this component of the semantic adaptive fusion network, the paired visible and infrared images are fed into an image encoder, and a Transformer-based block is used as the basic feature extractor. The feature extraction process can be described as follows:
[0080]
[0081] where \(I = 2, 3, 4\), \(t_{ii}\) and \(T_{ii}^r\) respectively represent the feature extractors of the visible and infrared images. For \(i = 1, 2, 3\), direct fusion is performed using the CSAF module However, when \(i = 4\), fusion is performed through the proposed PSAF module, which uses cross-attention before the CSAF module and self-attention after it
[0082] Recognizing that deep features encapsulate rich contextual semantic information, the deep semantic adaptive fusion module enhances the interaction between the two modalities by combining cross-attention through swapping the queries Q of the two modalities. Next, in the general semantic adaptive fusion module, the input feature map undergoes non-parametric instance normalization, and then the semantic prior is mapped into the feature space to generate spatial adaptive modulation parameters. Subsequently, the output of the deep semantic adaptive fusion module is processed by a feature decoder consisting of Transformer-based blocks, and then the resulting features are concatenated with the feature map output by the general semantic adaptive fusion module of the previous layer, and then convolution is performed. This series of operations is repeated multiple times to obtain the fused visual feature F fu That is:
[0083] F fu ={Conv(C(T f (F i '),F i-1 ))} r ,
[0084] where T f represents the feature decoder, C(.) represents concatenation along the channel dimension, and {.} r represents multi-layer repetition. Note that the upsampling of the features processed by T f needs to correspond to the downsampling performed by the feature extractor
[0085] Using a visual feature decoder, finally, the fused visual feature is passed through the Restomer decoder to obtain the fused image
[0086] In this embodiment, step S400 is a method for fusing infrared and visible light images guided by the semantic prior of a pre-trained visual model in a real-world scenario through an experimentally obtained model. The processing steps of the deep network model are as Figure 3 、 Figure 4 shown
[0087] The datasets used in the experiment. To verify the generalization, the fusion performance of the deep learning model on widely used public datasets of visible light and infrared image fusion was comprehensively evaluated, including M3FD, MSRS, LLVIP, and FMB. In addition, the performance of the deep learning model on the semantic segmentation task was evaluated using the MSRS and FMB datasets. The ability in the object detection task was also tested on the LLVIP dataset. The number of image pairs used for testing in these datasets is 300, 361, 300, and 280 respectively. In addition, the MSRS dataset includes 1,083 pairs of images for training the model
[0088] The proposed deep learning model was trained for 120 epochs under the PyTorch framework. A four-level encoder with the same structure was used for two images, where each level contains two Transformer blocks with the number of attention heads increasing as (1 / 2 / 4 / 8). The decoder consists of two Restormer
[40] blocks for reconstructing the final fused image. The hyperparameters controlling the trade-off of various loss terms were empirically set as αs = 1, ac = 12, αi = 4, and αg = 10. The initial learning rate was set to 0.0001, the batch size was 4, and an optimizer was utilized. The source images were cropped. The frozen ViT-B / 16 model was used for CLIP and DINO, and pre-training was performed after Open-CLIP and DINO respectively. Six statistical evaluation metrics were used to quantitatively evaluate the fusion performance, including entropy value (EN), spatial frequency (SF), average gradient (AG), standard deviation (SD), visual information fidelity (VIF), and NAB / F. The higher the values of these metrics, the better the fusion performance.
[0089] It should be noted that the above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0090] For the specific architecture of the semantic prior-guided infrared and visible light image fusion model of the pre-trained visual model of the present invention, please refer to Figure 5 . The semantic prior-guided infrared and visible light image fusion architecture of this pre-trained visual model mainly includes semantic prior generation, semantic adaptive fusion network, and visual feature decoder.
[0091] In the optimization stage, in some embodiments, the semantic prior-guided infrared and visible light image fusion method based on the pre-trained visual model further includes a calculation method for the loss function:
[0092] Using the source visible light image I vi , infrared image I ir and the final fused image I f as the input of the fusion loss function L f . This loss function consists of the following parts: structural similarity (SSIM) loss, maximum gradient loss, intensity loss, and color consistency loss. Each component of the loss function is as follows:
[0093] Calculation method of structural similarity loss: The structural similarity loss is used to evaluate the similarity between the fused image and the original source image, ensuring that the fused image retains the structural integrity of the input image, which is expressed as:
[0094]
[0095] Calculation method of color consistency loss (color difference): The degree of color loss is judged to ensure that the fused image and the source visible image maintain the consistency of color information, so as to maintain the naturalness of the image and the accuracy of color representation in the fused image, especially when dealing with scenes with rich colors, which is expressed as:
[0096]
[0097] In the formula, T cbcr
[39] represents the transfer function from RGB to CbCr;
[0098] Calculation method of intensity loss: This loss measures the pixel-level difference between the fused image and the source image. Therefore, the intensity loss is defined as:
[0099]
[0100] Calculation method of gradient loss: The purpose of gradient loss is to ensure that the fused image contains rich texture information by retaining the maximum edges of the two source images, which is expressed as:
[0101]
[0102] Calculation method of total loss, i.e., overall loss: The overall loss function is the weighted sum of the above loss terms, and its formula is:
[0103]
[0104] where α, α ζ and α z are hyperparameters that control the trade-off of the fusion loss function.
[0105] Example 2: As Figure 5 shown, a semantic prior-guided infrared and visible light image fusion system, the fusion system uses the semantic prior-guided infrared and visible light image fusion method, and the fusion system includes:
[0106] A semantic prior generation module 11, which is responsible for generating semantic priors from paired visible light and infrared images. Specifically, pre-trained visual models CLIP and DINO are used, with fixed parameters, to independently extract visual features from visible light and infrared inputs;
[0107] Semantic Adaptive Fusion Network Module 12. In this module, paired visible light and infrared images are fed into an image encoder. A Transformer-based block is used as the basic feature extractor. The extracted features are subjected to general semantic adaptive fusion and deep semantic adaptive fusion to make full use of the complementary attributes in the multi-modal setting, and finally refined output features are obtained.
[0108] Visual Feature Decoder Module 13. In this module, the fused visual features are passed through a Restomer decoder to obtain the fused image, which is the final output of this method.
[0109] Object Detection Module 14. The main function of the object detection module is to automatically identify and locate target objects from the input image, and at the same time generate the bounding boxes and class labels of each target. Through the feature extraction network, this module can analyze the multi-scale features of the image and effectively capture global and local features using deep learning methods (such as convolutional neural networks or Transformer architectures). On this basis, the object detection module predicts the target category through the classification branch and accurately locates the target position through the regression branch, ensuring that the output results have high accuracy and robustness and can be widely applied in fields such as autonomous driving, video surveillance, and scene analysis.
[0110] Embodiment 3: An electronic device (i.e., a semantic prior-guided infrared and visible light image fusion device), which includes: a processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, the steps of the semantic prior-guided infrared and visible light image fusion method are implemented.
[0111] Embodiment 4: A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the semantic prior-guided infrared and visible light image fusion method are implemented.
[0112] Compared with the prior art, the semantic prior-guided infrared and visible light image fusion method based on a pre-trained vision model provided by the present invention integrates the semantic capabilities of the pvm into the field of image fusion. This method uses CLIP and DINO encoders to extract and synthesize visual semantic features from source images, forming a robust semantic prior that captures global and local context information. The method of the present invention includes a designed semantic adaptive fusion network that combines general and deep semantic adaptive fusion (C / PSAF) modules. CSAF achieves adaptive multimodal fusion using spatial modulation, while PSAF achieves deeper semantic integration using cross-modal attention. These modules work together to dynamically adjust modality-specific information for different semantic objects, achieving precise and context-aware fusion. This method reduces the dependence on manual labels, a common limitation of existing fusion methods, and enhances the adaptive integration of multimodal information. It is meaningful in the field of object detection.
[0113] In summary, the above embodiments of the present invention provide a semantic prior-guided infrared and visible light image fusion method, system, device, and storage medium, which solve the limitations of the existing fusion methods that rely on manual labels and the problem of difficult adaptive fusion for different objects.
Claims
1. A semantic prior-guided infrared and visible light image fusion method, characterized in that: It includes the following steps: S100: Obtain public data sets, including MFD, MSRS, LLVIP, and FMB, and divide them into training sets and test sets; S200: Establishing a semantic prior-guided infrared and visible light image fusion model based on a pre-trained visual model, i.e., a deep learning network model, and pre-training the deep learning network model using the public data set to generate an initial recognition model, wherein the deep learning network model consists of three parts: a semantic prior generation module, a semantic adaptive fusion network, and a visual feature decoder; S300: extract global and local semantic priors using pre-trained models CLIP and DINO, then use these priors in the semantic adaptive fusion network, then guide the fusion process through the semantic adaptive fusion module and the deep semantic adaptive fusion module, adaptively modulate the modality-specific features from infrared and visible light inputs, and finally the visual decoder reconstructs the fused features into the final output image; S400: Conduct experiments using the public dataset, obtain experimental results through the divided dataset and test set, and quantitatively evaluate the fusion performance using multiple statistical indicators.
2. The semantic prior-guided infrared and visible light image fusion method according to claim 1, characterized in that: In the public data set of step S100, the MSRS and FMB data sets are used to evaluate the performance of the deep learning model in the semantic segmentation task, and the ability of the deep learning model in the target detection task is also tested on the LLVIP data set; the number of image pairs used for testing in the MFD, MSRS, LLVIP and FMB data sets are 300, 361, 300 and 280 respectively; the MSRS data set includes 1083 pairs of images for training the model.
3. The semantic prior-guided infrared and visible light image fusion method according to claim 1, characterized in that: In step S200, the deep learning network model processes the infrared and visible light images through semantic prior generation, extracts global and local semantic priors using pre-trained models CLIP and DINO, and then uses these priors in a semantic adaptive fusion network, wherein the fusion process is guided by a semantic adaptive fusion module and a deep semantic adaptive fusion module, and modality-specific features from infrared and visible light inputs are adaptively adjusted. This adaptive fusion is used to enhance semantic consistency and detail retention in the fused image. Finally, a visual feature decoder reconstructs the fused features into a final output image.
4. The semantic prior guided infrared and visible light image fusion method according to claim 1, characterized in that: In step S200, the semantic prior generation module is responsible for generating semantic priors from paired visible light and infrared images. Specifically, the pre-trained visual models, CLIP and DINO, with fixed parameters are used to independently extract visual features from visible light and infrared inputs. The process is as follows: where and θ c and θ d I represents the image encoders from CLIP and DINO respectively, and the visible light and infrared images are represented by I vi and I ir It is proposed to combine the semantic features of CLIP and DINO through association pooling to make it clear. c represents the CLIP semantic feature F extracted from visible light and infrared images c vi and F d Represents similar extracted DINO semantic features and 5. The semantic prior guided infrared and visible light image fusion method according to claim 1, characterized in that: In step S200, in the semantic adaptive fusion network component, the paired visible light and infrared images are fed into the image encoder, and the Transformer-based block is used as the basic feature extractor. The feature extraction process can be described as follows: Where I = 2, 3, 4, tii and Tiir represent the feature extractors of visible light and infrared images respectively. For i = 1, 2, 3, the CSAF module is used to directly fuse and When i = 4, fusion is performed via the proposed PSAF module, which uses cross-attention before the CSAF module and self-attention after it.
6. The semantic prior guided infrared and visible light image fusion method according to claim 1, characterized in that It also includes: In the semantic adaptive fusion module, the input feature map is subjected to non-parametric instance normalization, and then the semantic prior is mapped into the feature space to generate spatial adaptive modulation parameters. Subsequently, the output of the deep semantic adaptive fusion module is processed by a feature decoder composed of Transformer-based blocks, and then the obtained features are connected with the feature map output by the general semantic adaptive fusion module of the previous layer, and then convolution is performed. This series of operations is repeated many times, so that the fused visual feature F fu The integration of F fu ={Conv(C(T f (F i '),F i-1 ))} r , Where T f represents the feature decoder, C(.) represents the connection along the channel dimension, {.} r Indicates multiple layers of repetition, T f The upsampling of processed features needs to correspond to the downsampling performed by the feature extractor, Use the visual feature decoder and finally pass the fused visual features through the Restomer decoder to obtain the fused image.
7. The semantic prior guided infrared and visible light image fusion method according to claim 1, characterized in that It also includes the calculation method of the loss function: Using the source visible light image I vi 、Infrared image I ir And the final fused image I f As the fusion loss function L f The loss function consists of the following components: structural similarity (SSIM) loss, gradient loss, intensity loss, and color consistency loss. Each component of the loss function is as follows: Calculation method of structural similarity loss: Structural similarity loss is used to evaluate the similarity between the fused image and the original source image to ensure that the fused image retains the structural integrity of the input image, which is expressed as: Calculation method of color consistency loss: The degree of color loss is determined to ensure that the fused image maintains the consistency of color information with the source visible image, so as to maintain the naturalness of the image and the accuracy of the color representation in the fused image, especially when dealing with colorful scenes, which is expressed as: Where T cbcr [39] represents the transfer function from RGB to CbCr; Calculation method of intensity loss: This loss measures the pixel-level difference between the fused image and the source image, so the intensity loss is defined as: Calculation method of gradient loss: The purpose of gradient loss is to ensure that the fused image contains rich texture information by retaining the maximum edge of the two source images, which is expressed as: The calculation method of total loss is the overall loss: the overall loss function is the weighted sum of the above loss terms, and its formula is: Among them, α ζ and α z is a hyperparameter that controls the trade-off of the fusion loss function.
8. A semantic prior-guided infrared and visible light image fusion system, the fusion system using the semantic prior-guided infrared and visible light image fusion method according to any one of claims 1 to 7, characterized in that The fusion system comprises: Semantic prior generation module (11), which is responsible for generating semantic priors from paired visible and infrared images. Specifically, it uses pre-trained visual models CLIP and DINO with fixed parameters to independently extract visual features from visible and infrared inputs; Semantic adaptive fusion network module (12), in which the paired visible and infrared images are fed into the image encoder, and the transformer-based block is used as the basic feature extractor. The extracted features are subjected to general semantic adaptive fusion and deep semantic adaptive fusion, taking advantage of the complementary properties in the multimodal setting, and finally the refined output features are obtained; A visual feature decoder module (13), in which the fused visual features are passed through a Restomer decoder to obtain a fused image as the final output of the method; The target detection module (14) is used to automatically identify and locate the target object from the input image, and generate the bounding box and category label of each target. Through the feature extraction network, the module can analyze the multi-scale features of the image and effectively capture the global and local features using the deep learning method. On this basis, the target detection module predicts the target category through the classification branch and accurately locates the target position through the regression branch, ensuring that the output result has high accuracy and robustness.
9. An electronic device (i.e., a semantic prior-guided infrared and visible light image fusion device), characterized in that It comprises: a processor, a memory and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the semantic prior guided infrared and visible light image fusion method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored on a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the semantic prior-guided infrared and visible light image fusion method as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Infrared and visible light image fusion method
CN111861960A
Cited By
Infrared and visible light image fusion method and system based on text-guided semantic perception
CN120765479A
Extreme illumination-oriented visible light and infrared image fusion method, system and device based on text-guided spatial frequency domain interaction and medium
CN121214128A
Infrared weak and small target detection method based on double-granularity semantic prompt
CN121459031A
A method, apparatus, electronic device, and storage medium for infrared and visible light image fusion based on semantic prior transfer.
CN122573724A