Salient target detection method and system based on deep interactive fusion of three-modal features
By building a parallelized three-stream encoder network, combined with the deep separation convolution and edge guidance module, the deep interactive fusion of three-modal features of RGB, depth and thermal imaging is achieved, solving the problem of insufficient accuracy and robustness of significant object detection in complex scenarios, and achieving higher quality target segmentation.
Patent Information
- Application Number
- CN202510452999.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
The existing single-modal significance object detection methods are insufficient in complex scenarios with occlusion, similar background colors or low contrast areas, and it is difficult to effectively utilize the unique information of the three modes of RGB, depth and thermal imaging.
A parallelized three-stream encoder network architecture is built, and the depth of the pyramid structure can be separated convolution and multi-scale expansion convolution can be combined with the adaptive weight allocation mechanism and edge guidance module to realize the deep interactive fusion of multimodal features, improving the selectivity and boundary clarity of significant regional features.
The accuracy and robustness of significant target detection are significantly improved in complex scenarios, achieving higher quality target segmentation effects, especially in occlusion and background chaos areas.
Smart Images

Figure CN120374946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to a saliency object detection method and system for deep interaction and fusion of three-modal features. Background Art
[0002] Saliency object detection is an important computer vision task that can automatically identify the most attractive and prominent objects in a scene. Saliency object detection plays an important role in image segmentation, tracking, cropping, redirection, activity prediction, etc. Although significant progress has been made in the research of SOD in recent years, most of the research focuses on single-modal data, especially RGB images. These methods rely on various cues such as color, texture, and edge information to distinguish salient objects from the background. However, the information provided by single-modal data is limited, which often challenges the accurate detection of salient objects, especially in complex scenes with occlusion, similar background colors, or low-contrast regions.
[0003] With the development of sensor technology, depth and thermal imaging photography have been successfully applied to a wide range of scenarios, especially RGB-D and RGB-T saliency object detection tasks. Different from RGB visible light images, depth images and thermal images can provide spatial structure cues attached with depth and thermal imaging information attached with temperature respectively, which can be regarded as complementary information to RGB images. In the past few years, many methods for RGB-D saliency object detection and RGB-T saliency object detection have been designed and achieved advanced performance.
[0004] There are relatively few studies on saliency object detection using these three modalities of RGB images, depth images, and thermal images simultaneously. By fusing the three modalities of RGB, depth (D), and thermal infrared (T) (RGB-D-T), the model can make full use of the unique information provided by each modality, thereby achieving more robust and accurate saliency object detection. Each modality has its own characteristics and noise patterns, and it is not easy to find an optimal fusion method. Existing multi-modal methods often rely on simple fusion techniques such as addition, multiplication, and concatenation to combine different modalities, which are ineffective for challenging scenarios such as low illuminance and chaotic backgrounds. The information contained in visible light, depth images, and thermal images is different. Visible light contains rich texture details and diverse color features. Depth images mainly provide three-dimensional geometric structures of the scene and spatial position information of objects. Thermal images focus on the thermal radiation characteristics and temperature distribution of objects. Therefore, in view of the characteristics of the three, designing corresponding feature enhancement and feature extraction schemes to realize a high-quality three-modal saliency object detection network is an urgent problem for those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a saliency object detection method and system with deep interactive fusion of three-modal features, which can improve the detection accuracy and robustness in complex scenarios such as occlusion, similar background colors, or low-contrast regions, so as to solve the technical problems proposed in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions: A saliency object detection method with deep interactive fusion of three-modal features, at least including the following steps:
[0007] S1: Construct a parallel three-stream encoder network architecture to achieve hierarchical feature extraction of multi-modal inputs of visible light, depth, and thermal imaging;
[0008] S2: The dynamic feature enhancement module designed for the visible light modality realizes selective enhancement of salient region features by integrating depthwise separable convolutions of the pyramid structure;
[0009] S3: Construct a three-modal information interaction and fusion module, which improves the cross-modal feature fusion effect while maintaining parameter efficiency through the synergistic effect of multi-scale dilated convolutions and bidirectional attention gating;
[0010] S4: Construct an edge guidance module, integrate the edge guidance module to build a multi-scale edge guidance progressive fusion decoder, and realize boundary-clear object segmentation by guiding the foreground and background feature correlation through edge features, thereby realizing saliency object detection with deep interactive fusion of three-modal features.
[0011] Furthermore, the S1 at least includes the following steps:
[0012] Perform geometric transformation, photometric transformation, and Gaussian model operation on visible light images, depth images, and thermal images to expand sample diversity and improve model generalization;
[0013] Perform data augmentation on the original RGB visible light image, Depth depth image, and Thermal thermal imaging. The data augmentation includes but is not limited to rotation, translation transformation, and motion blur;
[0014] Crop all the enhanced images on the template frame into template images with a size of 352×352×3 as the input data of the three-stream encoder network;
[0015] Use VGG16 as the feature extraction backbone network of the three-stream encoder; extract the 1 / 2, 1 / 4, 1 / 8, 1 / 16 downsampled feature maps fea i ;
[0016] Denote them as fea 1 / 2 、fea 1 / 4 、fea 1 / 8 、fea1 / 16 。
[0017] Furthermore, the S2 at least includes the following steps:
[0018] Perform depthwise separable convolution operations on the feature maps of each scale obtained in S1, fully extract features while reducing the number of parameters. The depthwise separable convolution operations independently perform spatial convolution on each channel and use 1×1 convolution to adjust the number of channels for channel fusion. Specifically, refer to the following formula:
[0019] DepthwiseConv(fea i ) = γ(fea i , K depth )
[0020] PointwiseConv(fea i ) = fea i *K point
[0021] where K depth ∈R C×1×3×3 ; γ represents dilated convolution; fea i represents the feature maps of each branch;
[0022] Adaptive dynamically adjust the weights of each branch through the weight generation network. Global average pooling captures global information. Use the fully connected layer to generate branch weights and Softmax normalization. Finally, weighted fusion of multi-scale features. Specifically, refer to the following formula:
[0023] α = Softmax(W2·ReLU(W1·GlobalPool(fea i )))(i = 1,2,3,4)
[0024]
[0025] where W1 represents the first layer weight matrix for feature compression and non-linear activation; W2 is the second layer weight matrix to generate the weight coefficients of N branches; α represents the branch weight vector.
[0026] Furthermore, the three-modal information interaction and fusion module realizes deep cross-modal feature interaction and complementary information mining through the adaptive weight allocation mechanism. The three-modal feature deep interaction and fusion module contains four different-scale feature interaction branches. The application of the three-modal feature deep interaction and fusion module at least includes the following steps:
[0027] First, capture multi-modal features through convolutional kernels with dilation rates d ∈ {1, 3, 5, 7} to form a receptive field radius r = 2d + 1;
[0028] Branch 1 with a dilation rate of 1, a receptive field size of 3×3, captures local texture features; Branch 2 with a dilation rate of 3, a receptive field size of 7×7, captures part-level features; Branch 3 with a dilation rate of 5, a receptive field size of 11×11, captures object-level structures; Branch 4 with a dilation rate of 7, a receptive field size of 15×15, captures global context features;
[0029] Establish dynamic feature interaction among RGB-D-T modalities, improve the focusing ability of important regions and channels, use the attention maps of other modalities to adjust the features of the current modality, apply the channel attention mechanism to the feature maps F of different scales, generate weights through average pooling and max pooling, and pass through an MLP to obtain the feature M with dynamically enhanced channel attention c , see the following formula:
[0030] M c = Sigmoid(MLP(AvgPool(F)) + Sigmoid(MLP(MaxPool(F)))
[0031] where AvgPool(F) is the average pooling operation on F, MaxPool(F) is the max pooling operation on F, and Sigmoid is the corresponding Sigmoid activation function;
[0032] Parallelly apply the spatial attention mechanism to the feature map F, fuse max pooling and average pooling, and obtain the feature map M with dynamically enhanced spatial attention s :
[0033] M s = Sigmoid(COnv 7×7 ([MaxPool(F); AvgPool(F)]))
[0034] Establish dynamic feature interaction among RGB-D-T modalities. The specific method for cross-modal feature interaction fusion is as follows:
[0035]
[0036] where l represents the hierarchical index (l = 1, 2, 3, 4); ⊙ represents element-wise multiplication; represents the feature maps of depth (d) and thermal imaging (t) at the l-th layer; m represents traversing the depth (d) and thermal imaging (t) modalities; represents the spatial attention weights of the depth (d) and thermal imaging (t) modalities; This is the feature map of the RGB modality at the l-th layer;
[0037] Cascade the features of each scale, apply channel attention, and fuse with the original input to achieve feature aggregation from local to global, obtaining tri-modal depth fusion features The core implementation of hierarchical feature residual aggregation is shown in the following formula:
[0038]
[0039] Among them, X rgb represents the original RGB feature map; CA global (F fusion ) represents applying the global channel attention mechanism to the fused feature map F fusion .
[0040] Furthermore, the edge guidance module amplifies the high-frequency edge signals through residual learning, guides the association between the foreground and the background, realizes boundary-clear segmentation, and obtains the edge-enhanced feature EnFea i , and the specific implementation method is:
[0041]
[0042] EdgeWeight i = Sigmoid(BN(Conv 1×1 (E i )))
[0043]
[0044] Among them, i = 1, 2, 3, 4; E i represents the edge residual feature map; ⊙ represents element-wise multiplication;
[0045] Integrate the edge guidance module to construct a multi-scale edge-guided progressive fusion decoder, and use the progressive feature fusion upsampling technique to gradually restore the feature map. The progressive cascade fusion method is:
[0046]
[0047] Among them, Decoder i-1 represents the previous-level decoder, which accepts the cascaded features for processing and gradually restores the spatial resolution of the feature map; represents and EnFea i are cascaded in the channel dimension.
[0048] The saliency object detection system for tri-modal feature depth interaction fusion at least includes a data acquisition module, a feature extraction module, a feature interaction module, a multi-scale edge-guided progressive fusion decoder, and a saliency object detection prediction module
[0049] The data acquisition module is used to acquire and preprocess the original data;
[0050] The feature extraction module is used to segment features at different scales through the backbone network;
[0051] The feature interaction module is used to enhance the cross-modal feature fusion effect;
[0052] The multi-scale edge-guided progressive fusion decoder is used to gradually upsample to obtain an accurate saliency object detection map;
[0053] The saliency object detection prediction module converts the probability map into a binary map after predicting the pixel category of the input image through the model, and obtains the visualization result.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] 1. By constructing a parallel three-stream encoder network architecture, the present invention realizes hierarchical feature extraction of multi-modal inputs of visible light, depth, and thermal imaging. By introducing a three-modal information interaction and fusion module at the encoder level and designing an adaptive weight allocation mechanism, in-depth interaction and complementary information mining of cross-modal features are realized in the spatial and channel dimensions. The three-modal information interaction and fusion module includes an inter-modal attention mechanism and a feature enhancement unit, which can dynamically adjust the contribution degree of different modal features;
[0056] 2. The present invention designs a dynamic feature enhancement module for the visible light modality. By integrating depthwise separable convolutions with a pyramid structure, selective enhancement of features in significant regions is realized. Further, a multi-scale edge-guided progressive fusion decoder is constructed, and an edge perception module is designed. By guiding the feature correlation between the foreground and background regions through edge features, object segmentation with clear boundaries is realized;
[0057] 3. The present invention simultaneously introduces a progressive feature fusion upsampling technique to gradually refine and restore the feature map. Furthermore, by fusing the complementarity of multi-modal features and combining dynamic feature enhancement and edge guidance optimization, the present invention significantly improves the detail retention and boundary integrity of saliency object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0059] Figure 1 It is a flowchart of the method of the present invention;
[0060] Figure 2It is the flow chart of the system of the present invention;
[0061] Figure 3 It is the visualization result diagram of the present invention. Specific embodiments
[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0063] Embodiment 1:
[0064] Please refer to Figure 1 and Figure 3 , a saliency object detection method for tri-modal feature deep interactive fusion, at least including the following steps:
[0065] S1: Construct a parallel three-stream encoder network architecture to achieve hierarchical feature extraction of multi-modal inputs of visible light, depth, and thermal imaging;
[0066] S2: The dynamic feature enhancement module designed for the visible light modality realizes selective enhancement of salient region features by integrating depthwise separable convolutions of the pyramid structure;
[0067] S3: Construct a tri-modal information interaction and fusion module, and through the synergistic effect of multi-scale dilated convolution and bidirectional attention gating, improve the cross-modal feature fusion effect while maintaining parameter efficiency;
[0068] S4: Construct an edge guidance module, integrate the edge guidance module to construct a multi-scale edge-guided progressive fusion decoder, and realize clear boundary object segmentation by guiding the foreground and background feature association through edge features, so as to realize saliency object detection of tri-modal feature deep interactive fusion.
[0069] Furthermore, S1 at least includes the following steps:
[0070] Perform geometric transformation, photometric transformation, and Gaussian model operation on visible light images, depth images, and thermal images to expand sample diversity and improve model generalization;
[0071] Perform data augmentation processing on the original RGB visible light image, Depth depth image, and Thermal thermal imaging. The data augmentation includes but is not limited to rotation, translation transformation, and motion blur;
[0072] Crop all the enhanced images into template images with a size of 352×352×3 on the template frame as the input data of the three-stream encoder network;
[0073] Use VGG16 as the feature extraction backbone network of the three-stream encoder; extract the 1 / 2, 1 / 4, 1 / 8, and 1 / 16 downsampled feature maps fea of the VGG16 network i ;
[0074] which are respectively denoted as fea 1 / 2 , fea 1 / 4 , fea 1 / 8 , fea 1 / 16 .
[0075] Furthermore, S2 at least includes the following steps:
[0076] Perform depthwise separable convolution operations on the feature maps of each scale obtained in S1, fully extract features while reducing the number of parameters. The depthwise separable convolution operations perform spatial convolution independently in each channel and use 1×1 convolution to adjust the number of channels for channel fusion. Specifically, refer to the following formula:
[0077] DepthwiseConv(fea i ) = γ(fea i , K depth )
[0078] PointwiseConv(fea i ) = fea i * K point
[0079] where K depth ∈ R C×1×3×3 ; γ represents dilated convolution; fea i represents the feature maps of each branch;
[0080] Adaptive dynamically adjust the weights of each branch through the weight generation network, use global average pooling to capture global information, use a fully connected layer to generate branch weights, perform Softmax normalization, and finally weighted fuse multi-scale features. Specifically, refer to the following formula:
[0081] α = Softmax(W2·ReLU(W1·GlobalPool(fea i )))(i = 1,2,3,4)
[0082]
[0083] where W1 represents the first-layer weight matrix for feature compression and non-linear activation; W2 is the second-layer weight matrix to generate the weight coefficients of N branches; α represents the branch weight vector.
[0084] Furthermore, the three-modal information interaction and fusion module realizes deep cross-modal feature interaction and complementary information mining through an adaptive weight allocation mechanism. The three-modal feature deep interaction and fusion module contains four feature interaction branches of different scales. The application of the three-modal feature deep interaction and fusion module includes at least the following steps:
[0085] First, capture multi-modal features through convolutional kernels with dilation rates d ∈ {1, 3, 5, 7} to form a receptive field radius r = 2d + 1;
[0086] Branch 1 with a dilation rate of 1 has a receptive field size of 3×3 and captures local texture features; Branch 2 with a dilation rate of 3 has a receptive field size of 7×7 and captures part-level features; Branch 3 with a dilation rate of 5 has a receptive field size of 11×11 and captures object-level structures; Branch 4 with a dilation rate of 7 has a receptive field size of 15×15 and captures global context features;
[0087] Establish dynamic feature interaction among the RGB-D-T modalities, improve the focusing ability of important regions and channels, use the attention maps of other modalities to adjust the features of the current modality, apply the channel attention mechanism to feature maps F of different scales, generate weights through average pooling and max pooling, and pass through an MLP to obtain the features M with dynamically enhanced channel attention c , as shown in the following formula:
[0088] M c = Sigmoid(MLP(AvgPool(F)) + Sigmoid(MLP(MaxPool(F)))
[0089] where AvgPool(F) is the average pooling operation on F, MaxPool(F) is the max pooling operation on F, and Sigmoid is the corresponding Sigmoid activation function;
[0090] Parallelly apply the spatial attention mechanism to the feature map F, fuse max pooling and average pooling, and obtain the feature map M with dynamically enhanced spatial attention s :
[0091] M s = Sigmoid(Conv 7×7 ([MaxPool(F); AvgPool(F)]))
[0092] Establish dynamic feature interaction among the RGB-D-T modalities. The specific method for cross-modal feature interaction and fusion is:
[0093]
[0094] where l represents the hierarchical index (l = 1, 2, 3, 4); ⊙ represents element-wise multiplication; Denote the feature maps of depth (d) and thermal imaging (t) at the l-th layer; m represents traversing the depth (d) and thermal imaging (t) modalities; Denote the spatial attention weights of the depth (d) and thermal imaging (t) modalities; This is the feature map of the RGB modality at the l-th layer;
[0095] After concatenating the features of each scale, channel attention is applied and fused with the original input to achieve feature aggregation from local to global, obtaining the three-modal depth fusion features The core implementation of hierarchical feature residual aggregation is shown in the following formula:
[0096]
[0097] Among them, X rgb Denote the original RGB feature map; CA global (F fusion ) denotes applying the global channel attention mechanism to the fused feature map F fusion .
[0098] Furthermore, the edge guidance module amplifies the high-frequency edge signals through residual learning, guides the association between the foreground and the background, realizes clear boundary segmentation, and obtains the edge-enhanced feature EnFea i , and the specific implementation method is:
[0099]
[0100] EdgeWeight i = Sigmoid(BN(Conv 1×1 (E i )))
[0101]
[0102] Among them, i = 1, 2, 3, 4; E i Denote the edge residual feature map; ⊙ denotes element-wise multiplication;
[0103] Integrate the edge guidance module to construct a multi-scale edge guidance progressive fusion decoder, and use the progressive feature fusion upsampling technology to gradually restore the feature map layer by layer. The progressive cascade fusion method is:
[0104]
[0105] Among them, Decoder i-1 Denote the previous-level decoder, which accepts the concatenated features for processing and gradually restores the spatial resolution of the feature map; Denote and EnFea iPerform cascading in the channel dimension.
[0106] Specifically, when implementing, verify the saliency object detection method of the deep cross-modal feature interaction and fusion of the present invention. For F-measure (F m ), MAE, S-measure (S m ), E-measure (E m ), conduct comparative verification on the open-source dataset VDT2048. Table 1 shows the comparison results of each evaluation index of the method of the present invention and other 13 advanced methods under different attributes.
[0107] It can be analyzed from Table 1 that for the saliency object detection method of the deep cross-modal feature interaction and fusion described in the present invention, on the open dataset VDT2048 for tri-modal saliency object detection, the MAE is 0.0028, the F-measure is 0.9070, the E-measure is 0.9859, and the S-measure is 0.9281, which is superior to other 13 advanced bi-modal and tri-modal saliency object detection methods and has strong competitiveness.
[0108] Table 1 Comparison results of each advanced model on the VDT2048 validation set
[0109]
[0110]
[0111] Embodiment 2:
[0112] Refer to Figure 2 , the saliency object detection system of the deep cross-modal feature interaction and fusion includes at least a data acquisition module, a feature extraction module, a feature interaction module, a multi-scale edge-guided progressive fusion decoder, and a saliency object detection prediction module
[0113] The data acquisition module is used to acquire and preprocess the original data;
[0114] The feature extraction module is used to segment features of different scales through the backbone network;
[0115] The feature interaction module is used to improve the cross-modal feature fusion effect;
[0116] The multi-scale edge-guided progressive fusion decoder is used to gradually upsample to obtain an accurate saliency object detection map;
[0117] The saliency object detection prediction module converts the probability map into a binary map after predicting the pixel category of the input image through the model, and obtains a visualization result.
[0118] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in all respects, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
Claims
1. A saliency object detection method based on deep interaction and fusion of three-modal features, characterized in that: At least include the following steps: S1: Construct a parallel three-stream encoder network architecture to achieve hierarchical feature extraction of multi-modal inputs of visible light, depth, and thermal imaging; S2: The dynamic feature enhancement module designed for the visible light modality realizes selective enhancement of salient region features by integrating depthwise separable convolutions with a pyramid structure; S3: Construct a three-modal information interaction and fusion module, which improves the cross-modal feature fusion effect while maintaining parameter efficiency through the synergistic effect of multi-scale dilated convolutions and bidirectional attention gating; S4: Construct an edge guidance module, integrate the edge guidance module to construct a multi-scale edge-guided progressive fusion decoder, and achieve salient object detection with clear boundaries by guiding the association of foreground and background features through edge features, thereby realizing deep interaction and fusion of three-modal features for salient object detection.
2. The saliency object detection method based on the deep interaction and fusion of three-modal features according to claim 1, wherein: The S1 at least includes the following steps: Perform geometric transformation, photometric transformation, and Gaussian model operation on visible light images, depth images, and thermal images to expand sample diversity and improve model generalization; Perform data augmentation on the original RGB visible light image, Depth depth image, and Thermal thermal imaging. The data augmentation includes but is not limited to rotation, translation transformation, and motion blur; Crop all the enhanced images on the template frame into template images with a size of 352×352×3 as the input data of the three-stream encoder network; Use VGG16 as the feature extraction backbone network of the three-stream encoder; extract the 1 / 2, 1 / 4, 1 / 8, and 1 / 16 downsampled feature maps fea of the VGG16 network i ; Denoted as fea respectively 1 / 2 、fea 1 / 4 、fea 1 / 8 、fea 1 / 16 。 3. The saliency object detection method based on deep interaction and fusion of three-modal features according to claim 1, wherein: The S2 at least includes the following steps: Perform depthwise separable convolution operations on the feature maps of each scale obtained in S1 to fully extract features while reducing the number of parameters. The depthwise separable convolution operations independently perform spatial convolution on each channel and use 1×1 convolution to adjust the number of channels for channel fusion. See the following formula for details: DepthwiseConv(fea i ) = γ(fea i , K depth ) PointwiseConv(fea i ) = fea i * K point Among them, K depth ∈R C×1×3×3 ; γ represents dilated convolution; fea i represents the feature maps of each branch; Adaptive dynamically adjust the weights of each branch through a weight generation network, capture global information through global average pooling, generate branch weights using a fully connected layer, perform Softmax normalization, and finally weighted fuse multi-scale features. See the following formula for details: α = Softmax(W2·ReLU(W1·GlobalPool(fea i )))(i = 1, 2, 3, 4) Among them, W1 represents the first-layer weight matrix for feature compression and non-linear activation; W2 is the second-layer weight matrix for generating weight coefficients of N branches; α represents the branch weight vector.
4. The saliency object detection method based on deep interaction and fusion of three-modal features according to claim 3, characterized in that: The three-modal information interaction and fusion module realizes deep cross-modal feature interaction and complementary information mining through an adaptive weight allocation mechanism. The three-modal feature deep interaction and fusion module contains four different-scale feature interaction branches. The application of the three-modal feature deep interaction and fusion module at least includes the following steps: First, capture multi-modal features through convolutional kernels with dilation rates d∈{1, 3, 5, 7} to form a receptive field radius r = 2d + 1; Branch 1 with a dilation rate of 1 has a receptive field size of 3×3 and captures local texture features; Branch 2 with a dilation rate of 3 has a receptive field size of 7×7 and captures part-level features; Branch 3 with a dilation rate of 5 has a receptive field size of 11×11 and captures object-level structures; Branch 4 with a dilation rate of 7 has a receptive field size of 15×15 and captures global context features; Establish dynamic feature interaction between RGB-D-T modalities, improve the focusing ability of important regions and channels, use the attention maps of other modalities to adjust the features of the current modality, apply the channel attention mechanism to feature maps F of different scales, generate weights through average pooling and max pooling and passing through an MLP, and obtain the features M with dynamically enhanced channel attention c , see the following formula: M c = Sigmoid(MLP(AvgPool(F)) + Sigmoid(MLP(MaxPool(F))) Among them, AvgPool(F) is the average pooling operation on F, MaxPool(F) is the maximum pooling operation on F, and Sigmoid is the corresponding Sigmoid activation function; Apply the spatial attention mechanism to the feature map F in parallel, fuse the max pooling and average pooling to obtain the feature map M with dynamically enhanced spatial attention s : M s = Sigmoid(Conv 7×7 ([MaxPool(F); AvgPool(F)])) Establish dynamic feature interaction between RGB-D-T modalities. The specific method of cross-modal feature interaction and fusion is as follows: where \(l\) represents the hierarchical index (\(l = 1, 2, 3, 4\)); \(\odot\) represents element-wise multiplication; represents the feature maps of depth (\(d\)) and thermal imaging (\(t\)) at the \(l\)-th layer; \(m\) represents traversing the depth (\(d\)) and thermal imaging (\(t\)) modalities; represents the spatial attention weights of the depth (\(d\)) and thermal imaging (\(t\)) modalities; This is the feature map of the RGB modality at the \(l\)-th layer; Cascade the features at each scale, apply channel attention, and fuse them with the original input to achieve feature aggregation from local to global, obtaining the three-modal depth fusion features The core implementation of hierarchical feature residual aggregation is shown in the following formula: Among them, X rgb represents the original RGB feature map; CA global (F fusion ) represents applying the global channel attention mechanism to the fused feature map F fusion .
5. The saliency object detection method based on the deep interaction and fusion of three-modal features according to claim 4, wherein: The edge guidance module amplifies high-frequency edge signals through residual learning, guides the correlation between the foreground and the background, realizes boundary-clear segmentation, and obtains edge-enhanced features EnFea i , and the specific implementation method is as follows: EdgeWeight i = Sigmoid(BN(Conv 1×1 (E i ))) where i = 1, 2, 3, 4; E i represents the edge residual feature map; ⊙ represents element-wise multiplication; Integrate the edge guidance module to construct a multi-scale edge guidance progressive fusion decoder, and use the progressive feature fusion upsampling technology to restore the feature map layer by layer. The progressive cascade fusion method is as follows: Among them, Fecoder i-1 represents the decoder at the previous level, which accepts the cascaded features for processing and gradually restores the spatial resolution of the feature map; represents and EnFea i are cascaded in the channel dimension.
6. A saliency object detection system for deep interactive fusion of three-modal features, which is used for the method for saliency object detection by deep interactive fusion of three-modal features according to any one of the above claims 1-5, and is characterized in that: It includes at least a data acquisition module, a feature extraction module, a feature interaction module, a multi-scale edge guidance progressive fusion decoder, and a saliency object detection prediction module The data acquisition module is used to acquire and preprocess the original data; The feature extraction module is used to segment features of different scales through the backbone network; The feature interaction module is used to improve the cross-modal feature fusion effect; The multi-scale edge guidance progressive fusion decoder is used to gradually upsample to obtain an accurate saliency object detection map; The saliency object detection prediction module converts the probability map into a binary map after predicting the pixel category of the input image through the model, and obtains the visualization result.
Citation Information
Cited By
Image recognition method and device, computer equipment and storage medium
CN121366397A
Multi-modal detection device and method for space target
CN121878866A
A multi-modal detection apparatus and method for space objects
CN121878866B