Infrared and visible image fusion method based on semantic guidance and scene graph reasoning
Patent Information
- Application Number
- CN202610763430.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-08-28
AI Technical Summary
然而,该方法没有解决红外与可见光图像间的几何未对准问题,导致融合结果中出现重影或结构模糊;同时,特征融合机制仅依赖于空间光谱注意力机制,缺乏对场景结构的深入理解,导致在复杂场景下融合效果受限;此外,训练策略采用两阶段训练,缺乏自监督学习机制,无法充分利用未标注数据提升模型性能
(1)本发明设计了一种基于自监督微调的场景图构建模块,通过关系对比学习与跨模态一致性约束生成结构化语义先验,为融合过程提供像素级语义重要性权重图,增强模型对场景内容的深层理解与语义一致性表达能力。
Smart Images

Figure CN122656871A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image fusion and multimodal image processing, and is a method for fusion of infrared and visible light images based on semantic guidance and scene graph reasoning. Background Technology
[0002] Image fusion and multimodal image processing is an important research direction in computer vision, aiming to obtain fused images that combine thermal target saliency with texture detail preservation. Image fusion and multimodal image processing techniques are commonly used in scenarios such as autonomous driving, nighttime surveillance, remote sensing, military reconnaissance, and medical auxiliary diagnosis, playing a significant role in improving efficiency and accuracy in these fields. However, existing visible light and infrared image fusion methods suffer from problems such as intermodal geometric misalignment, insufficient semantic guidance, and computational efficiency limitations of complex models. Therefore, there is an urgent need to design a visible light and infrared image fusion method that can simultaneously address the issues of geometric misalignment in bimodal images, insufficient high-level semantic guidance, and computational efficiency limitations of complex models, thereby improving the structural consistency, target saliency, texture detail preservation, and semantic consistency of the fusion results.
[0003] Chinese patent publication number "CN121481864A" entitled "A Semantically Guided Mamba Infrared and Visible Image Fusion Method and System" describes a method that first constructs an image reconstruction network using a first-stage loss function; second, it constructs an image fusion network including a trained reconstruction encoder, a fusion module, and a semantically guided dual-branch decoder; then, it freezes the reconstruction encoder parameters and trains the image fusion network using a second-stage loss function; finally, it uses the trained image fusion network for image fusion. However, this method fails to address the geometric misalignment problem between infrared and visible light images, resulting in ghosting or structural blurring in the fusion result. Furthermore, the feature fusion mechanism relies solely on spatial spectral attention, lacking a deep understanding of scene structure, thus limiting fusion performance in complex scenes. Additionally, the two-stage training strategy lacks a self-supervised learning mechanism, failing to fully utilize unlabeled data to improve model performance. Therefore, designing an infrared and visible light image fusion method that effectively solves the geometric misalignment problem, deeply understands scene structure, and utilizes a self-supervised learning mechanism to improve fusion quality is a problem that this invention urgently needs to solve. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides an infrared and visible light image fusion method based on semantic guidance and scene graph reasoning, which improves alignment accuracy in areas with large initial registration errors or complex scenes, and ensures the robustness of semantic priors to modal differences.
[0005] To achieve the above objectives, the present invention specifically adopts the following technical solution: The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning includes the following steps: S1, Prepare the dataset: Prepare three datasets for visible light and infrared image fusion. Dataset 1 and Dataset 2 are used for training and fine-tuning, and Dataset 3 is used for testing. S2, Construct a visible light and infrared image fusion network model: The fusion network model consists of three parts: a scene graph construction module based on self-supervised fine-tuning, a semantically guided deformable alignment module, and a dynamic calculation and fusion module; S3, Training the network model: Train the visible light and infrared image fusion network model by inputting the dataset prepared in step S1 into the visible light and infrared image fusion network model constructed in step S2 and training it. S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output image fusion result and the input image and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can then be considered to have been pre-trained and saved. Select test images from dataset 3 and input them into the solidified model to obtain the image fusion result. Use the optimal evaluation metric for image fusion effect to measure the accuracy and performance of the model. S5, fine-tuning the model: the model is trained and fine-tuned using the visible light and infrared image fusion dataset 2 to optimize the model parameters, further improve the performance of the fusion network, and obtain target prediction results with smaller registration errors and more accurate positions; S6, save the model. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters to determine the final fused network model.
[0006] Furthermore, in S1, dataset one is the LLVIP dataset; dataset two is the RoadScene dataset; and dataset three is the TNO dataset.
[0007] Furthermore, in S2, the visible light and infrared image fusion model comprises three parts: a scene graph construction module based on self-supervised fine-tuning, a semantically guided deformable alignment module, and a dynamic computation and fusion module. The scene graph construction module based on self-supervised fine-tuning can effectively bridge the semantic gap between infrared and visible light images, providing high-quality semantic priors for subsequent alignment and fusion. The semantically guided deformable alignment module can use high-level semantic information to gradually guide and refine the feature-level alignment process, solving the geometric misalignment problem between infrared and visible light images. The dynamic computation and fusion module can utilize Visual Mamba to efficiently process high-resolution images, effectively meeting the dual requirements of global perception and computational efficiency for the fusion task.
[0008] Furthermore, in S2, the scene graph construction module based on self-supervised fine-tuning consists of a BLIP-2 model, a relation prediction Transformer decoder, and a self-supervised fine-tuning mechanism; the BLIP-2 model is used to obtain the bounding boxes and preliminary category labels of all salient instances in the image; the relation prediction Transformer decoder is used to predict the attributes of objects and the semantic relationships between them and other objects; the self-supervised fine-tuning mechanism is used to optimize the relation prediction network and improve its robustness to cross-modal differences. Furthermore, in S2, the semantically guided deformable alignment module consists of an initial offset field prediction module, a semantic evaluation and guidance module, and a parameter optimization module. The initial offset field prediction module is used to solve the geometric misalignment problem between infrared and visible light images. The semantic evaluation and guidance module is used to evaluate the semantic information of the current alignment features and generate a new semantic weight map to guide the next round of offset field prediction, gradually improving the accuracy of feature-level alignment. The parameter optimization module is used to ensure that the alignment features after each iteration are more accurate and output the optimized alignment features, providing a precise geometric basis for the subsequent fusion stage.
[0009] Furthermore, in S2, the dynamic computation and fusion module consists of a Visual Mamba backbone and a lightweight fusion network encoder-decoder; the Visual Mamba backbone is used to extract initial features from infrared and visible light images; the lightweight fusion network encoder-decoder is used to adaptively fuse features, combining channel attention, spatial attention, and semantic importance graph guidance to achieve fine-grained alignment and fusion of features.
[0010] Further, in S2, the loss function is a joint loss, including reconstruction loss, feature content loss, relation consistency loss, semantic weighted alignment loss, and scene graph fine-tuning loss. The reconstruction loss is used to ensure the pixel-level fidelity between the fused image and the reference image. The feature content loss is used to ensure the consistency of the high-level feature representation between the fused image and the source image. The relation consistency loss is used to maintain the consistency between the relationships between objects in the fused image and the source image. The semantic weighted alignment loss is used to optimize the geometric misalignment problem in the feature alignment process. The scene graph fine-tuning loss is used to improve the semantic expressive power of the scene graph inference module.
[0011] Furthermore, in step S4, the performance of the algorithm's prediction results is evaluated using evaluation metrics during the training of the network model.
[0012] The beneficial effects of this invention are as follows: (1) The present invention designs a scene graph construction module based on self-supervised fine-tuning. It generates structured semantic priors through relational comparison learning and cross-modal consistency constraints, providing pixel-level semantic importance weight graphs for the fusion process, thereby enhancing the model's deep understanding of scene content and semantic consistency expression capabilities.
[0013] (2) This invention constructs a semantically weighted lightweight fusion module, which deeply integrates the efficient global modeling capability of the Visual Mamba backbone with the semantic importance weight graph and channel / space attention mechanism to realize a semantically guided adaptive feature fusion strategy, which significantly improves the complementarity and fusion quality of multimodal information in complex scenarios.
[0014] (3) The present invention designs a joint loss function that includes reconstruction loss, feature content loss, relation consistency loss, semantic weighted alignment loss and scene graph fine-tuning loss. Through multi-objective collaborative optimization, it fully guarantees the comprehensive performance of the fused image in terms of pixel fidelity, high-level semantic consistency and geometric alignment accuracy.
[0015] (4) The present invention was validated on multiple public datasets such as LLVIP, TNO and RoadScene. The experimental results show that the present method is significantly better than the existing mainstream methods in seven core indicators: edge preservation ability (Qabf), structural similarity (SSIM), visual information fidelity (VIF), average gradient (AG), spatial frequency (SF), mutual information (MI) and peak signal-to-noise ratio (PSNR). The subjective visual quality is consistent with the objective evaluation results, proving that it has excellent generalization ability and practical value under different lighting conditions, complex backgrounds and multi-scale scenes. Attached Figure Description
[0016] Figure 1 This is a flowchart of an infrared and visible light image fusion method based on semantic guidance and scene graph reasoning; Figure 2 This is a schematic diagram of the visible light and infrared image fusion model constructed in this invention; Figure 3 This is a structural diagram of the scene graph construction module based on self-supervised fine-tuning in this invention; Figure 4 This is a schematic diagram illustrating the self-supervised fine-tuning mechanism of the present invention. Figure 5 This is a structural diagram of the semantically guided deformable feature alignment module of the present invention; Figure 6 This is a structural diagram of the dynamic calculation and fusion module of the present invention; Figure 7 This is a structural diagram of the lightweight fusion network encoder-decoder of the present invention; Figure 8This is a qualitative comparison diagram of the visible light and infrared image fusion method of the present invention and existing methods; Figure 9 This diagram illustrates a comparison of evaluation metrics between the visible light and infrared image fusion method of the present invention and existing methods. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example
[0018] like Figure 1 As shown, Implementation Example 1 of the Invention provides a flowchart of an infrared and visible light image fusion method based on semantic guidance and scene graph reasoning. The method specifically includes the following steps: S1, Prepare the dataset: Prepare the LLVIP dataset 1 for network training, to train the entire fusion network; prepare the RoadScene dataset 2 for fine-tuning; prepare the TNO dataset 3 for testing. S2, Construct a visible light and infrared image fusion network model: The fusion network model consists of three parts: a scene graph construction module based on self-supervised fine-tuning, a semantically guided deformable alignment module, and a dynamic calculation and fusion module; The scene graph construction module based on self-supervised fine-tuning consists of a BLIP-2 model, a relation prediction Transformer decoder, and a self-supervised fine-tuning mechanism. The BLIP-2 model is used to obtain bounding boxes and preliminary category labels for all salient instances in the image. The relation prediction Transformer decoder is used to predict object attributes and semantic relationships with other objects. The self-supervised fine-tuning mechanism is used to optimize the relation prediction network and improve robustness to cross-modal differences. The semantically guided deformable alignment module consists of an initial offset field prediction module, a semantic evaluation and guidance module, and a parameter optimization module. The initial offset field prediction module is used to address the geometric misalignment problem between infrared and visible light images. The semantic evaluation and guidance module is used to evaluate the semantic information of the current alignment features and generate a new semantic weight map to guide the next round of offset field prediction, gradually improving the accuracy of feature-level alignment. The parameter optimization module is used to ensure that the alignment features after each iteration are more accurate, outputting optimized alignment features to provide a precise geometric basis for subsequent fusion stages. The dynamic computation and fusion module consists of a VisualMamba backbone and a lightweight fusion network encoder-decoder. The Mamba backbone is used to extract initial features from infrared and visible light images; a lightweight fusion network encoder-decoder is used to adaptively fuse features, combining channel attention, spatial attention, and semantic importance graph guidance to achieve fine-grained alignment and fusion of features.
[0019] S3, Training the network model: Train the visible light and infrared image fusion network model by inputting the dataset prepared in step S1 into the visible light and infrared image fusion network model constructed in step S2 and training it.
[0020] S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output image fusion result and the input image, and minimizes the loss. Set a training loss threshold, and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters are then considered pre-trained and saved. Select test images from dataset three and input them into the fixed model to obtain the image fusion result. Use the optimal evaluation metric for image fusion effect to measure the model's accuracy and performance. During training, select a joint loss function, including reconstruction loss, feature content loss, etc. The joint loss function, which combines relation consistency loss, semantic weighted alignment loss, and scene graph fine-tuning loss, can balance the performance of fused images across multiple dimensions, including pixel fidelity, high-level semantic consistency, geometric alignment accuracy, and scene structure understanding. This effectively guides the fusion network to focus on the fusion quality of semantically important regions. Simultaneously, it provides stable gradients for the training process, ensuring efficient model convergence and significantly improving training stability and reliability. Appropriate evaluation metrics include edge preservation ability (Qabf), structural similarity (SSIM), visual information fidelity (VIF), average gradient (AG), spatial frequency (SF), mutual information (MI), and peak signal-to-noise ratio (PSNR).
[0021] S5, Fine-tuning the model: The model is trained and fine-tuned using dataset two to optimize model parameters and further improve the performance of the fusion network; S6, Save the model: After the fine-tuning training in step S5 is completed, solidify the fine-tuned network parameters and determine the final fusion network model; if performing a visible light and infrared image fusion task, the visible light and infrared images can be directly input into the trained end-to-end network model to obtain the final fusion result. Example
[0022] like Figure 1 As shown, the infrared and visible light image fusion method based on semantic guidance and scene graph reasoning specifically includes the following steps: S1. Prepare the datasets. Prepare Dataset 1 for training the fusion network. Dataset 1 is the LLVIP dataset, which contains 3096 pairs of strictly registered infrared and visible light images, suitable for evaluating the ability to preserve thermal targets and details under low light conditions. Prepare Dataset 2 for fine-tuning. The dataset contains 221 pairs of urban road images with rich traffic elements. Prepare Dataset 3 for testing. The dataset contains 30 pairs of images, including complex scenes from multiple sensors, which can test the robustness of the fusion method under modal differences and geometric deformations.
[0023] S2, Construct a visible light and infrared image fusion network model: The fusion network model consists of three parts: a scene graph construction module based on self-supervised fine-tuning, a semantically guided deformable alignment module, and a dynamic calculation and fusion module; The scene graph construction module based on self-supervised fine-tuning consists of a BLIP-2 model, a relation prediction Transformer decoder, and a self-supervised fine-tuning mechanism; The BLIP-2 model is used to obtain bounding boxes and preliminary class labels for all salient instances in the image. For each detected object region, the visual features extracted from the BLIP-2 visual encoder are input into a relation prediction Transformer decoder, which simultaneously predicts the object's attributes and semantic relationships with other objects. A pair of dual-stream encoders based on the BLIP-2 model with shared weights extract depth features from the infrared and visible light images respectively, mapping pixel information to a unified deep semantic feature space. The relation prediction Transformer decoder is used to predict the attributes of objects and their semantic relationships with other objects. Features extracted from the BLIP-2 model are fed into the relation prediction decoder to identify key objects such as pedestrians and vehicles in the image, and predict the attributes of the objects and the semantic and spatial relationships between them, such as "beside" and "driving", forming preliminary "subject-verb-object" relation triples. The resulting scene graph will serve as a strong semantic prior to identify key semantic regions such as pedestrians and vehicles in the image that need to be aligned and preserved. A self-supervised fine-tuning mechanism is used to optimize the relationship prediction network and improve robustness to cross-modal differences; the BLIP-2 model is used to generate an initial scene map from aligned features using zero-sample initialization. The model includes object nodes, attributes, and relation edges. Then, the backbone parameters of the BLIP-2 model are fixed, and only its relation prediction decoder is subjected to self-supervised fine-tuning for the fusion task. The self-supervised signal originates entirely from the input image pairs themselves, requiring no external annotation. Specifically, it includes: relation contrast learning loss, which constructs positive and negative sample pairs by performing data augmentation operations simulating modal changes on regions corresponding to the same relation triplet; and cross-modal relation consistency loss, which constrains the consistency of semantic representation by matching relation features of objects identified as the same pair in infrared and visible light images. The total loss function of the fine-tuning process is defined as... By optimizing the loss function, a scene graph generator robust to cross-modal differences is obtained, providing accurate semantic priors for feature alignment and image fusion. This further enables the generation of discrete scene graph information... Mapping back to pixel space to guide alignment, this invention utilizes the attention map of the Cross-Attention layer of the BLIP-2 model for salient object nodes detected in the scene graph. Extract the cross-attention map of the last layer of the decoder and its corresponding attributes. The final semantic importance weight graph Calculate as a weighted sum of attention maps for all objects, and normalize to... Interval: ; in, For object The confidence score.
[0024] The semantically guided deformable alignment module consists of an initial offset field prediction module, a semantic evaluation and guidance module, and a parameter optimization module; The initial offset field prediction module is used to address the geometric misalignment between infrared and visible light images; it consists of three 3×3 convolutional layers; infrared features. With visible light characteristics After concatenation along the channel dimensions, input the data into the initial offset field prediction network to predict the initial offset field. Initial alignment features are generated using deformable convolution D. ; The parameter optimization module ensures that the alignment features are more accurate after each iteration, outputting optimized alignment features to provide a precise geometric basis for subsequent fusion stages. The initial offset field prediction module outputs a loop that runs for up to K iterations, where each iteration adjusts the alignment features from the previous round. Reference infrared features and the semantic weight graph generated from the scene graph Concatenate along the channel dimension and input offset prediction network Predict the current residual migration field Cumulative offset field and current alignment features The update formula is defined as: ;
[0025] in, Represents spatial transformation, utilizing accumulated offset field Original visible light characteristics Perform resampling; The dynamic computation and fusion module consists of a Visual Mamba backbone and a lightweight fusion network encoder-decoder. The Visual Mamba backbone is used to extract initial features from infrared and visible light images. The lightweight fusion network encoder-decoder is used for adaptive feature fusion, combining channel attention, spatial attention, and semantic importance graph guidance to achieve fine-grained feature alignment and fusion. Input features and After 3×3 convolution, ReLU activation, and depthwise separable convolution, basic encoded features with compact expressive power are obtained. These basic encoded features are then subjected to global pooling and convolutional mapping to generate a fusion weight map W. Global pooling extracts the overall response information of the features, characterizing the importance of different channels, while convolutional mapping extracts local spatial distribution information and combines it with semantic response to generate unified guiding weights. This weight map simultaneously includes channel attention, spatial attention, and semantic importance expressions, adaptively representing the contribution of different positions, channels, and semantic regions to the fusion result. Then, the two input features are processed... and Perform weighted combination; where, features Element-wise multiplication with the weighted graph W, features With supplementary weight map Element-wise multiplication is performed, and then the two results are summed pixel by pixel to obtain the fused feature. Its expression is: ; in, This indicates element-wise multiplication.
[0026] S3, Training the network model: Train the visible light and infrared image fusion network model by inputting the dataset prepared in step S1 into the visible light and infrared image fusion network model constructed in step S2 and training it.
[0027] S4. Select a suitable loss function and determine the optimal evaluation metric for this method: Select a suitable loss function that minimizes the difference between the output image fusion result and the input image and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can then be considered to have been pre-trained and saved. Select test images from dataset 3 and input them into the solidified model to obtain the image fusion result. Use the optimal evaluation metric for image fusion effect to measure the accuracy and performance of the model. The loss function calculated by the network output and label in S4 uses a joint loss function, including reconstruction loss, feature content loss, relation consistency loss, semantic weighted alignment loss, and scene graph fine-tuning loss. Reconstruction loss It is used to ensure pixel-level fidelity between the fused image and the reference image, effectively preserving the key structural and texture information of the source image. This can be expressed using the following formula: ; Where H and W represent the height and width of the image, respectively. Indicates the location of the fused image Pixel value at that location, This represents the pixel value of the reference image at the corresponding location; Feature content loss This is used to ensure the consistency between the fused image and the reference image in the high-level semantic feature space. Multi-layer features are extracted and differences are calculated through a pre-trained VGG network. This can be expressed using the following formula: ; Where L represents the set of selected VGG network feature layers, Indicates the first Feature maps extracted from layers, These represent the number of channels, height, and width of the feature map for that layer, respectively. Relationship consistency loss This is used to constrain the distribution of relationships between objects in the fused image to remain consistent with the source images (infrared and visible light images), thereby enhancing the semantic rationality of the fusion result. This can be expressed using the following formula: ; in, Denotes KL divergence, These represent the probability distribution vectors of object relationships in the infrared image, visible light image, and fused image extracted by the scene graph generator, respectively. Semantic weighted alignment loss It is used to optimize the geometric misalignment problem in the feature alignment process and guides the accurate registration of key regions through a semantic importance weight map. This can be expressed using the following formula: ; in, This represents the total number of pixels in the feature map. The semantic weight at the location, and Let represent the feature vectors of the aligned feature and the reference feature at position k, respectively; Scene graph fine-tuning loss It is used to improve the robustness of scene graph generators to cross-modal differences, including relation contrast learning loss and cross-modal relation consistency loss. This can be expressed using the following formula: ; ; ; in, For the set of positive sample relation pairs, For cosine similarity, For temperature parameters, for Positive sample relation embedding, and These are the embedding vectors for the infrared and visible light mode relationships, respectively. This is the balance coefficient; Appropriate evaluation metrics include edge preservation ability (Qabf), structural similarity (SSIM), visual information fidelity (VIF), average gradient (AG), spatial frequency (SF), mutual information (MI), and peak signal-to-noise ratio (PSNR).
[0028] The appropriate evaluation metrics selected in S4 are edge preservation ability (Qabf), structural similarity (SSIM), visual information fidelity (VIF), average gradient (AG), spatial frequency (SF), mutual information (MI), and peak signal-to-noise ratio (PSNR). Edge Preservation Failure (Qabf) quantifies the degree to which the fused image retains the edge structure of the source image; a higher value indicates more complete preservation of edge details. Qabf can be expressed by the following formula: ; in, Image size, The local edge preservation of the source images A (infrared), B (visible light) and the fused image F. These are weighting coefficients based on gradient magnitude; Structural similarity (SSIM) measures the degree of similarity between the fused image and the reference image in terms of brightness, contrast, and structural information. Its value ranges from [value range missing]. The closer the value is to 1, the higher the structural fidelity. SSIM can be expressed by the following formula: ; in, These are the mean, standard deviation, and covariance, respectively. It is the stability constant; Visual Information Fidelity (VIF) is a statistical model of natural scenes used to evaluate the amount of visual information retained in a fused image; a higher value indicates higher perceptual quality. VIF can be expressed using the following formula: ; Where j is the sub-band index, For variance, subscript These correspond to the fused image, the reference image, and the noise, respectively. The average gradient (AG) reflects the sharpness of image details and the richness of texture; a higher value indicates stronger edge and detail representation. AG can be expressed by the following formula: ; in, These are the Sobel gradients in the x and y directions, respectively. Spatial frequency (SF) comprehensively characterizes the spatial detail activity of an image in the horizontal and vertical directions; a higher value indicates richer texture information. SF can be expressed by the following formula: ; ; ; Where RF and CF are the row frequency and column frequency, respectively, and the boundaries are filled with mirror images. Mutual information (MI) measures the amount of information shared between the fused image and its source images. A higher MI value indicates that more key information from the source images is preserved. MI can be expressed by the following formula: ; in, For joint probability distribution, It represents a marginal probability distribution; Peak Signal-to-Noise Ratio (PSNR) assesses the overall fidelity of an image based on mean square error, and is measured in dB. A higher PSNR indicates less distortion. PSNR can be expressed by the following formula: ; ; in, The maximum pixel value. Mean square error; During network training, the feature extraction backbone uses Visual Mamba to load ImageNet pre-trained weights, the semantically guided alignment network is initialized with Xavier, and the BLIP-2 scene graph module only has its relationship prediction network fine-tuned. Training uses the Adam optimizer with a batch size of 16, and the learning rate is set in stages: the alignment module training stage... (100 rounds), the scene graph and fusion network training phase is as follows: The final end-to-end fine-tuning stage is reduced to The entire process uses cosine annealing scheduling; the input image is uniformly 256×256 and is enhanced by random flipping and ±10° rotation; the dynamic early termination threshold τ is set to 0.7 through grid search.
[0029] S5, Fine-tuning the model: The model was trained and fine-tuned again using the RoadScene dataset 2, with the learning rate set to 0.005 and 500 iterations, while other parameters remained unchanged, to further improve the performance of the fusion network; S6, Save the model: After training is completed in step S4, solidify the fine-tuned network parameters. In step S5, fine-tune the model and determine the final fusion network model. If performing a visible light and infrared image fusion task, the visible light and infrared images can be directly input into the trained end-to-end network model to obtain the final prediction result. The implementation of convolution, activation functions, etc., are algorithms well known to those skilled in the art, and the specific processes and methods can be found in relevant textbooks or technical documents.
[0030] This invention constructs an end-to-end infrared and visible light image fusion method based on semantically guided iterative alignment and scene graph reasoning. This method effectively solves the geometric misalignment problem between infrared and visible light images and fully utilizes the structured semantic priors provided by the scene graph to achieve adaptive feature fusion, directly generating high-quality fused images. This method avoids the tedious process of manually designing feature alignment strategies and fusion rules, simplifying and increasing the efficiency of the fusion process, and significantly improving the overall performance of the fused image in terms of detail preservation, semantic consistency, and visual quality. A qualitative comparison of the image fusion results of existing technologies and the method proposed in this invention is as follows: Figure 8 As shown, under the same conditions, the feasibility and superiority of the proposed method were further verified by calculating the correlation index between the image fusion result and the labeled image obtained by existing methods.
[0031] A comparative diagram of evaluation indicators between existing technologies and the method proposed in this invention is shown below. Figure 9 As shown in the figure, the method proposed in this invention significantly outperforms existing mainstream methods in seven core fusion evaluation indicators: Qabf, SSIM, VIF, AG, SF, MI, and PSNR. It comprehensively verifies its advantages in the collaborative preservation of infrared thermal radiation targets and visible light texture details, the enhancement of image edge sharpness and structural integrity, and the improvement of multimodal information fusion quality. It effectively suppresses ghosting and structural blurring, and achieves dual optimization of geometric alignment accuracy and semantic fusion quality.
[0032] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for fusing infrared and visible light images based on semantic guidance and scene graph reasoning, characterized in that: The method specifically includes the following steps: S1, Prepare the dataset: Prepare three datasets for visible light and infrared image fusion. Dataset 1 and Dataset 2 are used for network training and fine-tuning, and Dataset 3 is used for testing. S2, Construct a visible light and infrared image fusion network model: The fusion network model consists of three parts: a scene graph construction module based on self-supervised fine-tuning, a semantically guided deformable alignment module, and a dynamic calculation and fusion module; S3, Training the network model: Train the visible light and infrared image fusion network model by inputting the dataset prepared in step S1 into the visible light and infrared image fusion network model constructed in step S2 and training it. S4. Select a suitable loss function and determine the optimal evaluation metric: Select a suitable loss function that minimizes the difference between the output image fusion result and the input image and minimizes the loss. Set a training loss threshold and iteratively optimize the model until the number of training iterations reaches the set threshold or the value of the loss function reaches the set threshold range. The model parameters can then be considered to have been pre-trained and saved. Select the test images from dataset 3 and input them into the solidified model to obtain the image fusion result. Use the optimal evaluation metric for image fusion effect to measure the accuracy and performance of the model. S5, fine-tuning the model: the model is trained and fine-tuned using the visible light and infrared image fusion dataset 2 to optimize model parameters, improve the performance of the fusion network, and obtain target prediction results with smaller registration errors and more accurate positions; S6, save the model. After the fine-tuning training in S5 is completed, solidify the fine-tuned network parameters to determine the final fused network model.
2. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 1, characterized in that: In S1, dataset one is the LLVIP dataset; dataset two is the RoadScene dataset; and dataset three is the TNO dataset.
3. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 1, characterized in that: In S2; The scene graph construction module based on self-supervised fine-tuning can effectively bridge the semantic gap between infrared and visible light images, providing high-quality semantic priors for subsequent alignment and fusion. The semantically guided deformable alignment module can use high-level semantic information to gradually guide and refine the feature-level alignment process, solving the geometric misalignment problem between infrared and visible light images. The dynamic computing and fusion module can utilize Visual Mamba to efficiently process high-resolution images, which better meets the dual requirements of fusion tasks for global perception and computational efficiency.
4. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 3, characterized in that: In S2, the scene graph construction module based on self-supervised fine-tuning consists of a BLIP-2 model, a relation prediction Transformer decoder, and a self-supervised fine-tuning mechanism; The BLIP-2 model is used to obtain bounding boxes and preliminary class labels for all salient instances in an image; The relation prediction Transformer decoder is used to predict the attributes of an object and the semantic relationships between it and other objects. The self-supervised fine-tuning mechanism is used to optimize the relationship prediction network and improve its robustness to cross-modal differences.
5. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 3, characterized in that: In S2, the semantically guided deformable alignment module consists of an initial offset field prediction module, a semantic evaluation and guidance module, and a parameter optimization module. The initial offset field prediction module is used to solve the geometric misalignment problem between infrared and visible light images; The semantic evaluation and guidance module is used to evaluate the semantic information of the current aligned features and generate a new semantic weight map to guide the next round of offset field prediction, thereby gradually improving the accuracy of feature-level alignment. The parameter optimization module is used to ensure that the alignment features are more accurate after each iteration, and outputs the optimized alignment features to provide a precise geometric basis for the subsequent fusion stage.
6. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 3, characterized in that: In S2, the dynamic computing and fusion module consists of a Visual Mamba backbone and a lightweight fusion network encoder-decoder; The Visual Mamba backbone is used to extract initial features from infrared and visible light images; The lightweight fusion network encoder-decoder is used to adaptively fuse features, combining channel attention, spatial attention, and semantic importance graph guidance to achieve fine-grained alignment and fusion of features.
7. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 1, characterized in that: In S4, the loss function is a joint loss, including reconstruction loss, feature content loss, relation consistency loss, semantic weighted alignment loss, and scene graph fine-tuning loss; The reconstruction loss is used to ensure pixel-level fidelity between the fused image and the reference image; The feature content loss is used to ensure the consistency of the high-level feature representation between the fused image and the source image; The relationship consistency loss is used to maintain the consistency between the relationships between objects in the fused image and the source image; The semantically weighted alignment loss is used to optimize the geometric misalignment problem in the feature alignment process; The scene graph fine-tuning loss is used to improve the semantic expressive power of the scene graph reasoning module.
8. The infrared and visible light image fusion method based on semantic guidance and scene graph reasoning according to claim 7, characterized in that: In step S4, the training of the network model also includes evaluating the performance of the algorithm's prediction results using evaluation metrics.