Visible light-thermal infrared saliency target detection method based on correlation modeling

By extracting and fusing the features of visible and thermal infrared images, and using correlation modeling and homography matrix distortion technology, the problem of difficult processing of unaligned image pairs is solved, achieving efficient significance target detection.

CN120198647AActive Publication Date: 2025-06-24ANHUI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510374304.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-24
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

Existing visible-hot infrared saliency object detection methods are difficult to process unaligned image pairs, resulting in performance degradation.

Method used

Multi-layer features of visible and thermal infrared images are extracted through a pre-trained feature extractor, combined with semantic feature fusion module and spatial-level correlation module, gradually model the correlation between modes and within modes, and estimate that the homography matrix distorts the thermal infrared images so that they align the target area in the visible light image.

Benefits of technology

The performance improvement of visible-hot infrared saliency target detection without alignment is achieved, allowing more accurate detection of saliency targets in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198647A_ABST
    Figure CN120198647A_ABST
Patent Text Reader

Abstract

The invention discloses a visible light-thermal infrared saliency target detection method based on correlation modeling, and the method comprises the steps: modeling the correlation of two modes from the level of space, region and image step by step, so as to predict an accurate saliency detection result in a non-aligned visible light-thermal infrared image pair. On the spatial level, an adapter containing semantic information of the salient objects is embedded to guide correlation modeling to pay attention to a salient region, and the salient objects in the two modalities are aligned. At a region level, a partial correlation between the two modalities is modeled for the aligned salient regions. On the image level, the region level correlation is expanded to the whole image of the visible light mode to predict the accurate salient object region in the visible light image. According to the method, the best effect is achieved on non-aligned, weak-aligned and aligned saliency target detection data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and image detection, and particularly relates to a visible light-thermal infrared salient object detection method based on correlation modeling. Background Art

[0002] Salient object detection aims to identify and segment the most attractive regions in a visual scene. It can help eliminate redundant information and has been applied to various visual tasks, such as image compression, video analysis, and visual tracking. Although great progress has been made in visible light images, existing methods still have difficulty in distinguishing salient regions in complex scenes, such as cluttered backgrounds, low light, and similar foregrounds and backgrounds.

[0003] Thermal infrared sensors can capture the overall shape of an object and provide complementary information for visible light images. Therefore, some recent studies have introduced corresponding thermal infrared images for visible light images for information complementarity to improve the detection performance in complex scenes. Early traditional visible light-thermal infrared salient object detection methods mainly used handcrafted features, such as texture, color, and brightness information, to highlight the common salient object regions in the two modalities through comparison and differential feature fusion. With the improvement of computing power, deep learning-based methods have become the mainstream. They use visible light-thermal infrared salient object detection datasets to train network models and are supervised and optimized through corresponding loss functions. The trained models can extract high-level features from visible light and thermal infrared images and fuse them to predict more accurate saliency maps in complex scenes. Deep learning-based methods can be divided into early fusion, mid-level fusion, and late fusion according to different fusion methods: early fusion is to fuse the information of the two modalities before the visible light and thermal infrared image data are fed into the network model, mid-level fusion is carried out at the feature level, and late fusion is carried out at the prediction result level. Among them, mid-level fusion has better results and is also the most widely used fusion method.

[0004] However, existing visible light-thermal infrared salient object detection methods are almost designed based on the alignment of images in the two modalities. Moreover, the originally captured visible light-thermal infrared image pairs are unaligned in space and scale. Manually aligning them will consume a large amount of human costs and is not conducive to practical application and deployment. In addition, due to the difficulty of directly using the multimodal correspondence in unaligned image pairs, directly applying existing manually aligned methods to unaligned data will lead to a significant drop in performance.

[0005] In existing work, DCNet attempts to establish multimodal correspondences in weakly aligned image pairs through affine transformation and dynamic convolution. Although the local receptive field of the convolution operation is effective in weakly aligned image pairs with small spatial offsets, it is difficult to handle large deviations in space and scale between non-aligned image pairs. To improve it, SACNet uses a pair of asymmetric windows to cover the corresponding information in non-aligned image pairs. However, the fixed windows cannot flexibly achieve accurate coverage of the corresponding information in different scenarios, resulting in the introduction of irrelevant noise in the relevant modeling. In the invention, we explicitly align the common information of the two modalities and gradually model the inter-modal and intra-modal correlations.

[0006] Therefore, how to directly model the correlation of the corresponding information between the two modalities in non-aligned image pairs to utilize the complementary information between the two modalities is the core of improving the performance of visible light-thermal infrared salient object detection without alignment. Summary of the Invention

[0007] Objective of the Invention: The objective of the present invention is to solve the deficiencies existing in the prior art and provide a visible light-thermal infrared salient object detection method based on correlation modeling.

[0008] Technical Solution: A visible light-thermal infrared salient object detection method based on correlation modeling of the present invention includes the following steps:

[0009] Step S1: Input a pair of non-aligned visible light image and thermal infrared image directly captured by a camera. The same target captured in this image pair has different spatial positions and scales. Then, use a pre-trained feature extractor to extract features from the visible light image and the thermal infrared image respectively to obtain the features of the corresponding modalities, that is, the multi-layer features of the visible light image and the multi-layer features of the thermal infrared image

[0010] Here, i represents the i-th layer of features. Four layers of features are extracted for each modality's image. Among them, f r 4 and f t 4 are the high-level features of the corresponding modalities, containing rich target semantic information;

[0011] Step S2: Send the obtained two high-level features f r 4 and f t 4 into the semantic feature fusion module to integrate and obtain the high-level semantic feature f s of the common target region in the two modalities. Use the high-level semantic feature f s to guide the model to focus on the target region;

[0012] Step S3: Feed the obtained high-level semantic feature f s together with the original unaligned visible light image and thermal infrared image into the semantic-guided spatial-level correlation module. The spatial-level correlation module models the spatial correlation of the target regions in the two-modal images, and on this basis, estimates the homography matrix H between the two modalities. Use the homography matrix H to perform a warping operation on the thermal infrared image to align it with the corresponding target region in the visible light image. The warped thermal infrared image is denoted as I′ t ;

[0013] Step S4: Feed the thermal infrared image I′ t , high-level semantic feature f s , multi-level features of the visible light image and multi-level features of the thermal infrared image together into the region-level inter-modal correlation module; the region-level inter-modal correlation module first maps the corresponding regions in the visible light modality and the warped thermal infrared modality, and then models the inter-modal correlation of the features of the corresponding regions in the two modalities. Denote the obtained region-level inter-modal correlation features as

[0014] Step S5: Feed the region-level inter-modal correlation features and multi-level features of the visible light image into the image-level intra-modal correlation module. The image-level intra-modal correlation module extends the region correlation to the entire image range through self-attention to model the correlation of the significant regions in the entire visible light image. Denote the obtained intra-modal correlation features as

[0015] Step S6: Feed the intra-modal correlation features into the feature decoder, and fuse the correlation features of adjacent layers from top to bottom to obtain the predicted saliency map S, which is supervised by its labeled ground truth G.

[0016] To extract discriminative features from the input images, the pre-trained feature extractor in Step S1 is based on the Swin-B model of Transformer. The Swin-B model extracts four layers of features with different resolutions, which are 96×96, 48×48, 24×24, and 12×12 respectively.

[0017] To make full use of the semantic information in the high-level features to locate the target regions, the specific method for obtaining the high-level semantic feature f s in Step S2 through the semantic feature fusion module is as follows:

[0018] First, the two high-level features f r 4 and f t4 Level and splice to obtain multi-modal high-level features

[0019] Next, the multi-modal high-level features are fed into self-attention to enhance the features of the target regions jointly attended to by the two modal features by modeling the dependencies in the features;

[0020] Then, a convolutional block is used to refine and smooth the enhanced multi-modal high-level features to obtain the high-level semantic feature f containing the target location information s ; here, the semantic fusion module includes splicing, cross-attention, and convolutional operations.

[0021] To explicitly align the corresponding regions in the two modalities, the detailed process of step S3 is as follows:

[0022] Step S3.1: Use the pre-trained multi-modal homography estimator IHN to regard the visible light image and the thermal infrared image as the source image and the target image respectively, and extract the features of the source image and the target image through a siamese encoder. Then, calculate the correlation between the source image and the target image in the feature space and estimate the homography matrix H. The estimation formulas for the source image and the target image are:

[0023]

[0024] where I rgb and I t represent the unaligned visible light image and thermal infrared image respectively, Φ(.) represents the feature encoder, Ψ(,) represents the correlation calculation method, represents the homography estimator;

[0025] Step S3.2: Feed the features of each layer extracted by the siamese encoder into the semantics-guided adapter to make IHN adapt to the visible light-thermal infrared data. The calculation process is as follows:

[0026]

[0027] where is the extracted feature after adaptation of the l-th layer and will be propagated to the next layer of features. S-Adapter l is the adapter of the l-th layer. S-Adapter l uses the high-level semantic feature f s to guide the current layer of features F l to focus on the salient regions. The calculation process is as follows:

[0028]

[0029] where W dn and Wup is a mapping operation represents the fusion unit, φ represents the ReLU activation function, CAP(.) represents channel average pooling, σ represents the Sigmoid activation function, and ⊙ represents the element-wise multiplication operation;

[0030] Step S3.3: Adjust the formula for obtaining the homography matrix H by the multi-modal homography estimator IHN in step 3.1 to:

[0031]

[0032] where Φ Adapter (,) represents the feature encoder fine-tuned by the adapter;

[0033] Step S3.4: Use the homography matrix H obtained in step S3.3 to perform a warping operation on the thermal infrared image, and finally obtain the thermal infrared image I′ aligned with the visible light, t , and the calculation process is as follows:

[0034] I′ t = Warp(I t ,H)

[0035] where Warp(,) is the warping operation.

[0036] To model the correlation of corresponding regions in the two modalities, the specific method of step S4 is:

[0037] First, use the warped thermal infrared image I′ t to map out the corresponding region in the visible light image, and this region is used to interact with the unwarped visible light image. Here, the mapping means that IHN takes a pair of source images and target images and outputs the estimated homologous matrix H, and the matrix H can map the points in the target image to the corresponding points in the source image;

[0038] Then, introduce the high-level semantic feature f s to promote the correlation modeling to focus on the target region and obtain the inter-modal correlation feature The calculation process is as follows:

[0039]

[0040] where is the RGB feature of the i-th layer (there are four layers of features), represents the mapping operation, represents the correlation operation, and Q, K, and V represent query, key, and value respectively.

[0041] To expand the regional correlation between the two modalities modeled, the specific method of step S5 is:

[0042] First, use the residual connection to add the cross-modal correlation features to the entire visible light modal features;

[0043] Then, propagate the cross-modal correlation to the entire visible light modal through the correlation operation to obtain the intra-modal correlation features at the image level The calculation process is as follows:

[0044]

[0045] Among them, is the RGB feature of the i-th layer, is the cross-modal correlation feature at the regional level.

[0046] Furthermore, the specific method of step S6 is:

[0047] Use the labeled ground truth G through the binary cross-entropy loss function L bce and the dice coefficient loss function L dice to jointly supervise the predicted saliency map S. The formula for the total loss function L in the training stage is:

[0048] L = L bce + L dice

[0049]

[0050] Among them, N is the total number of pixels, G j represents the ground truth label of the j-th pixel, P j represents the predicted probability of the j-th pixel, |P| represents the number of pixels in the significant target area in the prediction map, |G| represents the number of pixels in the significant target detection area in the ground truth label map, and |P∩G| represents the number of pixels in the intersection area between P and G.

[0051] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0052] (1) Based on a progressive multi-modal fusion strategy, the present invention proposes a progressive correlation modeling strategy from the spatial level to the regional level and then to the image level and achieves the best performance compared with the existing technical solutions.

[0053] (2) The spatial-level correlation modeling module of the present invention can fine-tune the pre-trained multi-modal homography estimator to align the corresponding regions in the two modalities.

[0054] (3) The regional-level correlation modeling module of the present invention can establish the correlation between the corresponding region features in the two modalities.

[0055] (4) The image-level correlation modeling module of the present invention can extend the modeled regional correlation in the visible light modality to the entire image.

[0056] (5) The semantic feature fusion module of the present invention can integrate the semantic information in the high-level features to help the model locate the target region position. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a schematic diagram of the overall process of the present invention;

[0058] Figure 2 is a schematic diagram of the semantic feature fusion module of the present invention;

[0059] Figure 3 is a schematic diagram of the semantic-guided spatial-level correlation module of the present invention;

[0060] Figure 4 is a schematic diagram of the regional-level inter-modal correlation module of the present invention;

[0061] Figure 5 is a schematic diagram of the image-level intra-modal correlation module of the present invention;

[0062] Figure 6 is a schematic diagram of the visual comparison between the present invention and the existing methods. DETAILED DESCRIPTION OF THE INVENTION

[0063] The technical solution of the present invention will be described in detail below, but the protection scope of the present invention is not limited to the described embodiments.

[0064] As Figure 1 shown, the visible light-thermal infrared salient object detection method based on correlation modeling of the present invention includes the following steps:

[0065] Step S1: Input a pair of unaligned visible light image and thermal infrared image directly captured by a camera, and then use a pre-trained feature extractor to extract features from the visible light image and the thermal infrared image respectively to obtain the features of the corresponding modality, that is, the multi-layer features of the visible light image and the multi-layer features of the thermal infrared image

[0066] Here, i represents the i-th layer of features, and four layers of features are extracted for each modality image, where f r 4 and f t 4 are the high-level features of the corresponding modality, containing rich target semantic information;

[0067] Step S2: Combine the two obtained high-level features f r 4 and f t4 It is sent to the semantic feature fusion module to integrally obtain the high-level semantic features f of the common target regions in the two modalities. s Using the high-level semantic features f s to guide the model to focus on the target regions;

[0068] Step S3: The obtained high-level semantic features f s , the original unaligned visible light image, and the thermal infrared image are sent to the semantic-guided spatial-level correlation module together. The spatial-level correlation module models the spatial correlation of the target regions in the two-modal images and, on this basis, estimates the homography matrix H between the two modalities. The homography matrix H is used to perform a warping operation on the thermal infrared image to align it with the corresponding target region in the visible light image. The warped thermal infrared image is denoted as I'. t ;

[0069] Step S4: The thermal infrared image I' t , the high-level semantic features f s , the multi-level features of the visible light image and the multi-level features of the thermal infrared image are sent to the region-level inter-modal correlation module together. The region-level inter-modal correlation module first maps the corresponding regions in the visible light modality and the warped thermal infrared modality, and then models the inter-modal correlation of the features of the corresponding regions in the two modalities. The obtained region-level inter-modal correlation features are denoted as

[0070] Step S5: The region-level inter-modal correlation features and the multi-level features of the visible light image are sent to the image-level intra-modal correlation module. The image-level intra-modal correlation module extends the region correlation to the entire image range through self-attention to model the correlation of the salient regions in the entire visible light image. The obtained image-level intra-modal correlation features are denoted as

[0071] Step S6: The intra-modal correlation features are sent to the feature decoder to fuse the correlation features of adjacent layers from top to bottom to obtain the predicted saliency map S, which is supervised by its labeled ground truth G.

[0072] The pre-trained feature extractor in step S1 of this embodiment is the Swin-B model based on Transformer. The Swin-B model extracts four layers of features with different resolutions, which are 96×96, 48×48, 24×24, and 12×12 respectively.

[0073] The specific method for obtaining the high-level semantic features f s in step S2 of this embodiment through the semantic feature fusion module is as follows:

[0074] First, level and splice two high-level features f r 4 and f t 4 to obtain a multi-modal high-level feature

[0075] Next, send the multi-modal high-level feature into self-attention to enhance the feature of the target area jointly attended by the two modal features by modeling the dependencies in the features;

[0076] Then, use a convolutional block to refine and smooth the enhanced multi-modal high-level feature to obtain a high-level semantic feature f s .

[0077] The detailed process of step S3 in this embodiment is as follows:

[0078] Step S3.1: Use the pre-trained multi-modal homography estimator IHN to regard the visible light image and the thermal infrared image as the source image and the target image respectively, extract the features of the source image and the target image through the siamese encoder, then calculate the correlation between the source image and the target image in the feature space and estimate the homography matrix H. The estimation formulas for the source image and the target image are:

[0079]

[0080] where I rgb and I t represent the unaligned visible light image and the thermal infrared image respectively, Φ(.) represents the feature encoder, Ψ(,) represents the correlation calculation method, represents the homography estimator;

[0081] Step S3.2: Send the features of each layer extracted by the siamese encoder into the semantic-guided adapter to make IHN adapt to the visible light-thermal infrared data. The calculation process is as follows:

[0082]

[0083] where is the extracted feature adapted at the l-th layer and will be propagated to the next layer of features. S-Adapter l is the adapter at the l-th layer. S-Adapter l uses the high-level semantic feature f s to guide the current layer of features F l to focus on the salient region. The calculation process is as follows:

[0084]

[0085] Where W dn and W up are mapping operations, represents the fusion unit, φ represents the ReLU activation function, CAP(.) represents channel average pooling, σ represents the Sigmoid activation function, and ⊙ represents the element-wise multiplication operation;

[0086] Step S3.3: Adjust the formula for obtaining the homography matrix H by the multi-modal homography estimator IHN in Step 3.1 to:

[0087]

[0088] Where Φ Adapter (,) represents the feature encoder fine-tuned by the adapter;

[0089] Step S3.4: Use the homography matrix H obtained in Step S3.3 to perform a warping operation on the thermal infrared image, and finally obtain the thermal infrared image I′ aligned with the visible light, t , and the calculation process is as follows:

[0090] I′ t = Warp(I t ,H)

[0091] Where Warp(,) is the warping operation.

[0092] The specific method of Step S4 in this embodiment is:

[0093] First, use the warped thermal infrared image I′ t to map the corresponding region in the visible light image, and this region is used to interact with the unwarped visible light image. Here, the mapping means that IHN takes a pair of source images and target images and outputs the estimated homologous matrix H;

[0094] Then, introduce the high-level semantic feature f s to promote the correlation modeling to focus on the target region, and obtain the region-level inter-modal correlation feature The calculation process is as follows:

[0095]

[0096]

[0097] Where is the RGB feature of the i-th layer, represents the mapping operation, represents the correlation operation, and Q, K, and V represent query, key, and value respectively.

[0098] The specific method of Step S5 in this embodiment is:

[0099] First, use the residual connection to add the inter-modal correlation features to the entire visible light modal features;

[0100] Then, propagate the inter-modal correlation to the entire visible light modal through the correlation operation to obtain the image-level intra-modal correlation features The calculation process is as follows:

[0101]

[0102] Among them, is the RGB feature of the i-th layer, is the region-level inter-modal correlation feature.

[0103] The specific method of step S6 in this embodiment is:

[0104] Use the labeled ground truth G through the binary cross-entropy loss function L bce and the dice coefficient loss function L dice to jointly supervise the predicted saliency map S. The formula for the total loss function L in the training stage is:

[0105] L = L bce + L dice

[0106]

[0107]

[0108] Among them, N is the total number of pixels, G j represents the ground truth label of the j-th pixel, P j represents the predicted probability of the j-th pixel, |P| represents the number of pixels in the significant target area in the prediction map, |G| represents the number of pixels in the significant target detection area in the ground truth label map, and |P∩G| represents the number of pixels in the intersection area between P and G.

[0109] Example 1:

[0110] For the non - aligned model, in this embodiment, a total of 10,000 pairs of images, namely UVT20K - Train in the existing non - aligned visible - thermal infrared salient object detection dataset UVT20K, are used as the training set, and five non - aligned datasets, namely UVT20K - Test, UVT2000, un - VT5000 - Test, un - VT1000, and un - VT821, are used as the test sets to verify the effectiveness of the network. For the aligned model, in this embodiment, a total of 2,500 pairs of images, namely VT5000 - Train in the existing aligned visible - thermal infrared salient object detection dataset VT5000, are used as the training set, and three existing publicly available aligned datasets, namely VT5000 - Test, VT1000, and VT821, are used as the test sets to verify the effectiveness of the network.

[0111] As Figure 1 shown, the visible - thermal infrared image pairs are respectively fed into the feature extractor to extract the corresponding visible - light modality features and thermal - infrared modality features The highest - level features f r 4 and f t 4 are fed into the semantic feature fusion module to integrate and obtain the semantic feature f s to guide the model to focus on the target area. Meanwhile, the input visible - thermal infrared image pairs and the semantic feature f s are fed into the semantic - guided spatial - level correlation module to align the corresponding regions in the thermal - infrared image and the visible - light image, obtaining the aligned thermal - infrared image I′ t . Then, the semantic feature, the thermal - infrared image, the visible - light modality features, and the thermal - infrared modality features are fed into the region - level inter - modality correlation module to model the correlation between the corresponding region features of the two modalities, obtaining the inter - modality correlation features Then, the inter - modality correlation features and the visible - light modality features are fed into the image - level inter - modality correlation module to expand the region - level correlation to the entire image level, obtaining the intra - modality correlation features of the visible - light modality Finally, the intra - modality correlation features of the visible - light modality are fed into the feature decoder to integrate multi - layer features, obtaining the finally predicted saliency map S.

[0112] The network in this embodiment is based on the PyTorch framework, trained for 80 epochs with a batch size of 4 on two GeForce RTX 3090 GPUs, setting the learning rate to 1e - 5 and the weight decay to 1e - 4, and the resolution of the input image pairs is 384 * 384.

[0113] This embodiment uses three quantitative metrics widely used in visible-light to thermal-infrared salient object detection: E-measure (Em), F-measure (Fm), and S-measure (Sm) to quantitatively evaluate the performance of the network. The technical solution of the present invention is compared with 13 other existing technical solutions based on three quantitative metrics, including MIDD, CSRNet, CGFNet, SwinNet, OSRNet, TNet, DCNet, HRTransNet, MCFNet, LSNet, CAVER, LAFB, SACNet.

[0114] Quantitative comparison:

[0115] The specific comparison experiment results between this embodiment and 13 existing technical solutions are shown in Table 1, where the bold font represents the optimal performance. It can be seen from Table 1 that compared with the sub-optimal method CAVER, the performance of the present invention has been improved by 6.2%, 2.8%, and 7.8% on average for the three metrics (i.e., Em, Fm, and Sm) on two non-aligned datasets.

[0116] Table 1 Performance comparison experiment results

[0117]

[0118]

[0119] Visual comparison:

[0120] The visual comparison results of the salient prediction maps generated by this embodiment and 13 existing technical solutions are as Figure 6 shown. This embodiment selects 4 challenging samples, including: complex shapes of salient objects, cluttered images, similar foreground and background of salient objects, low light, low contrast of thermal-infrared images, etc. For example, Figure 6 on the left side of the first row in, in the case where the shape of the salient object is complex and the quality of the thermal-infrared image is poor, existing advanced methods will have difficulty in locating the target area or accurately segmenting the target, while our method can completely and accurately segment the salient object, which proves that our method can fully exploit useful multi-modal complementary information through progressive correlation modeling. Another example is, Figure 6 on the right side of the second row, in the case of low light, the visible-light image is difficult to provide effective target information. Our method can use the target contour information in the thermal-infrared image to achieve accurate prediction, while the compared methods will be seriously interfered by the background and predict incorrect salient regions, which indicates that our method can flexibly use the useful information in the dominant modality to achieve more accurate prediction. Therefore, the present invention can more accurately achieve non-aligned visible-light to thermal-infrared salient object detection.

Claims

1. A visible light-thermal infrared salient target detection method based on correlation modeling, characterized in that: The following steps are involved: Step S1: Input a pair of unaligned visible light images and thermal infrared images directly captured by a camera, and then use a pre-trained feature extractor to extract features from the visible light image and the thermal infrared image respectively to obtain features of the corresponding modality, i.e., multi-layer features of the visible light image. and multi-layer features of thermal infrared images Here, i represents the i-th layer of features, and four layers of features are extracted from each modality image, among which and It is the high-level feature of the corresponding modality, containing rich target semantic information; Step S2: The two high-level features and The high-level semantic features f of the common target area in the two modalities are obtained by integrating them. s , using high-level semantic features f s Guide the model to focus on the target area; Step S3: The obtained high-level semantic features f s The original unaligned visible light image and thermal infrared image are sent to the semantically guided spatial correlation module. The spatial correlation module models the spatial correlation of the target area in the two modal images and estimates the homography matrix H between the two modalities on this basis. The thermal infrared image is distorted using the homography matrix H to align it with the corresponding target area in the visible light image. The distorted thermal infrared image is recorded as I t ′; Step S4: thermal infrared image I t ′, high-level semantic features f s , Multi-layer features of visible light images and multi-layer features of thermal infrared images The two modalities are sent to the regional inter-modal correlation module together; the regional inter-modal correlation module first maps the corresponding areas in the visible light mode and the distorted thermal infrared mode, and then models the inter-modal correlation of the corresponding regional features of the two modes. The obtained regional inter-modal correlation feature is recorded as Step S5: regional level inter-modal correlation features and multi-layer features of visible light images The image-level intra-modal correlation module is used to extend the regional correlation to the entire image range through self-attention to model the correlation of the salient regions in the entire visible light image. The obtained image-level intra-modal correlation feature is recorded as Step S6: The intra-modal correlation features It is sent to the feature decoder to fuse the correlation features of adjacent layers from top to bottom to obtain the predicted saliency map S, which is supervised by the labeled true value G.

2. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1 is characterized in that: The pre-trained feature extractor in step S1 is based on the Swin-B model of Transformer. The Swin-B model extracts four layers of features with different resolutions, namely 96×96, 48×48, 24×24 and 12×12.

3. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1 is characterized in that: The step S2 obtains the high-level semantic feature f by using the semantic feature fusion module s The specific method is: First, two high-level features and Flatten and concatenate to obtain multimodal high-level features Next, the multimodal high-level features Feed self-attention to enhance the target area features that are jointly focused on by the two modal features by modeling the dependencies in the features; Then, a convolutional block is used to refine and smooth the enhanced multimodal high-level features to obtain the high-level semantic features f containing the target location information. s .

4. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1 is characterized in that: The detailed process of step S3 is as follows: Step S3.1: Use the pre-trained multimodal homography estimator IHN to treat the visible light image and the thermal infrared image as the source image and the target image respectively, and extract the features of the source image and the target image through the twin encoder. Then, calculate the correlation between the source image and the target image in the feature space and estimate the homography matrix H. The estimation formula of the source image and the target image is: Among them I rgb and I t denote the unaligned visible light image and thermal infrared image respectively, Φ(.) denotes the feature encoder, Ψ(,) denotes the correlation calculation method, represents the homography estimator; Step S3.2: Feed the features of each layer extracted by the twin encoder into the semantically guided adapter to adapt the IHN to the visible light-thermal infrared data. The calculation process is as follows: in It is the extracted features after adaptation of the lth layer and will be propagated to the next layer of features, S-Adapter l It is the adapter of layer l, S-Adapter l Using high-level semantic features f s To guide the current layer feature F l Focusing on the salient area, the calculation process is as follows: Where W dn and W up is a mapping operation, represents the fusion unit, φ represents the ReLU activation function, CAP(.) represents the channel average pooling, σ represents the Sigmoid activation function, and ⊙ represents the element-wise multiplication operation; Step S3.3: Adjust the formula for obtaining the homography matrix H by the multimodal homography estimator IHN in step 3.1 to: where Φ Adapter (,) represents the feature encoder after fine-tuning by the adapter; Step S3.4: Use the homography matrix H obtained in step S3.3 to distort the thermal infrared image, and finally obtain a thermal infrared image I aligned with the visible light. t ′, the calculation process is as follows: I t ′=Warp(I t ,H) Where Warp(,) is the warping operation.

5. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1 is characterized in that: The specific method of step S4 is: First, using the distorted thermal infrared image I t ′ to map out the corresponding area in the visible light image, which is used to interact with the undistorted visible light image. Here, mapping means that IHN takes a pair of source and target images and outputs an estimated homology matrix H; Then, refer to the high-level semantic feature f s To promote correlation modeling to focus on the target area and obtain regional-level inter-modal correlation features The calculation process is as follows: in is the RGB feature of the i-th layer, Represents a mapping operation, Table correlation operation, Q, K, V represent query, key, and value respectively.

6. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1 is characterized in that: The specific method of step S5 is: First, the inter-modality correlation features are added to the entire visible light modality features using residual connections; Then, the correlation between the modalities is propagated to the entire visible light modality through the correlation operation to obtain the image-level intra-modal correlation features. The calculation process is as follows: in, is the RGB feature of the i-th layer, It is the regional level inter-modal correlation feature.

7. The visible light-thermal infrared salient target detection method based on correlation modeling according to claim 1, characterized in that: The specific method of step S6 is: Use the labeled truth value G through the binary cross entropy loss function L bce and the dice coefficient loss function L dice The formula of the total loss function L in the training phase for the jointly supervised predicted saliency map S is: L=L bce +L dice Where N is the total number of pixels, G j Represented as the true value label of the j-th pixel, P j represents the predicted probability of the jth pixel, |P| represents the number of pixels in the salient target area in the predicted image, |G| represents the number of pixels in the salient target detection area in the true value label image, and |P∩G| represents the number of pixels in the intersection area between P and G.

Citation Information

Patent Citations

  • Visible light thermal infrared visual tracking method for weak registration data

    CN116205959A

  • Visible light-thermal infrared salient target detection method based on bidirectional alternating fusion strategy

    CN118038228A

  • RGB-D salient object detection method

    GB202403824D0