An RGB-D Salient Object Detection Method Based on a Multimodal Difference Fusion Network

By introducing a multimodal differential fusion network in RGB-D significance object detection, the difference between RGB and Depth modes is used to optimize the significance inference process, and the fusion is carried out through the third-stream differential supervision mechanism, the accuracy problem of existing methods in dealing with challenging scenarios is solved, and more efficient RGB-D significance object detection is achieved.

CN114693952BActive Publication Date: 2025-07-01ANHUI UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210308520.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-07-01
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

Existing RGB-D significance object detection methods are difficult to accurately and effectively detect significant targets when dealing with challenging scenarios, and most methods use an indiscriminate fusion method to integrate RGB and Depth features, ignoring the differences between modalities.

Method used

A RGB-D significance object detection method based on multimodal differential fusion network is designed. By utilizing the differential analysis between the RGB mode and the Depth mode, the significance inference process of RGB stream and the Depth stream is optimized, and the differential fusion of cross-modal through the three-stream differential supervision mechanism is carried out.

Benefits of technology

Through modal differential analysis and differential fusion, significance target detection is more accurate and effective in challenging scenarios, improving the performance of RGB-D significance target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114693952B_ABST
    Figure CN114693952B_ABST
Patent Text Reader

Abstract

The present invention provides an RGB-D salient object detection method based on a multi-modal difference fusion network, belonging to the field of image saliency detection technology. The method uses a Swin Transformer to extract RGB and Depth features containing global context information for inferring salient objects in the scene. The present invention mainly explores the differences between the RGB and Depth modalities to analyze the connections and differences of saliency in these two modalities, and designs a difference fusion network to fuse cross-modal features for capturing complete salient objects. The present invention includes the following steps: (1) using a Swin Transformer to extract cross-modal features; (2) using a bidirectional fusion method to fuse RGB and Depth features to generate a Fusion stream; (3) using a three-stream difference supervision mechanism to obtain the differences between modalities; (4) using this difference to fuse cross-modal features; (5) using a target cascade aggregation decoder to perform saliency inference and decoding on the fused cross-modal features to generate a predicted saliency map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and particularly to an RGB-D salient object detection method based on a multi-modal difference fusion network. Background Art

[0002] With the development and progress of information technology, and the explosive growth of multimedia data volume (such as pictures, texts, audios, videos, etc.) in daily life, it has promoted the vigorous development of image processing technology. As a very important technology in the field of image processing, the salient object detection technology mainly analyzes the most attention-grabbing objects or regions in an image and automatically separates the salient objects from the background. As one of the most basic density prediction tasks, it is widely used in many other downstream tasks, such as image retrieval, semantic segmentation, visual tracking, content-based image editing, and robot navigation. In addition, salient object detection is also widely used in the analysis and acquisition processes of many social media, such as the technology applications of emphasizing portraits and blurring backgrounds in mobile phone photography technology.

[0003] Most of the early salient object detection methods were for RGB images and could achieve satisfactory results. Generally, real RGB scenes mostly contain some challenging scenes, such as low contrast, multiple objects, transparent objects, complex backgrounds, etc. Facing these challenging scenes, it is difficult for RGB-based salient object detection to accurately and effectively detect salient objects and segment them completely. Facing this problem, depth images (Depth maps) are used in the field of salient detection. By using the spatial information and 3D layout information in the Depth map to provide supplementary clues, thus helping the salient object detection method to effectively process these challenging scenes, this technology is called RGB-D salient object detection.

[0004] With the popularization of depth acquisition devices (such as Microsoft Kinect, Huawei Mate 30, iPhone XR, etc.), depth information can be obtained at a relatively low cost. This phenomenon has also accelerated the vigorous development of RGB-D salient detection. Currently, most RGB-D salient object detection methods improve the performance of salient detection by integrating RGB features and Depth features to obtain gain information. However, most of these methods use an undifferentiated fusion method to integrate RGB features and Depth features, which regards RGB information and Depth information as having the same status. However, the human visual mechanism acts on RGB scenes, so obviously the roles played by RGB and the Depth map are different.

[0005] To address the above-mentioned problems, the present invention designs an RGB-D saliency object detection method based on a multi-modal difference fusion network, which uses the difference analysis between the RGB modality and the Depth modality to give the saliency object of the scene. The saliency inference processes of the RGB stream and the Depth stream are respectively optimized by using the differences between these modalities. Finally, by fusing the differences between the RGB and Depth modalities, the final saliency result is obtained. Specifically, the present invention designs a three-stream difference supervision mechanism, which respectively performs saliency and edge inference through the RGB stream, the Depth stream, and the fusion stream, and implements cross-modal difference fusion by integrating these inference results. Summary of the Invention

[0006] To address the above problems, the present invention provides an RGB-D saliency object detection method based on a multi-modal difference fusion network. The specific technical solutions adopted are as follows:

[0007] 1. Obtain and organize the RGB-D dataset for training and testing.

[0008] 1.1) Inductively organize the obtained RGB-D dataset (DUT-RGB dataset, NJU2K dataset, NLPR dataset, LFSD dataset, RGBD135 dataset), and divide a single sample into an RGB image P RGB , a depth image P depth , a manually annotated saliency object segmentation image S GT and a manually annotated saliency object edge segmentation image E GT .

[0009] 1.2) Divide the collected RGB-D dataset into a training set and a test set. The training set is a 2985-sample set composed of 800 samples from the DUT dataset, 1400 samples from the NJU2K dataset, and 650 samples from the NLPR dataset. The remaining samples of the above five datasets are used as the test set.

[0010] 2. The present invention uses the SwinTransformer network in deep learning as the backbone network of the present invention to extract RGB and Depth features.

[0011] 2.1 Respectively construct two SwinTransformer-based encoders to extract RGB features and Depth features. Among them, the Swin Transformer encoder is composed of four basic Swin Transformer blocks, and its definition is as follows:

[0012] S = MLP(LN(W m (LN(F f)) + F f )) + W m (LN(F f )) + F f Formula (1)

[0013] ST = MLP(LN(W s (LN(S)) + S)) + W s (LN(S)) + S Formula (2)

[0014] where MLP represents a multi - layer perceptron, LN represents layer normalization, W m represents the multi - head self - attention mechanism, and W s represents the self - attention mechanism based on the shifted window.

[0015] 2.2 Based on Step 2.1, the outputs of the RGB and Depth encoders can be obtained, denoted as RGB features and Depth features

[0016] 3. Based on the RGB and Depth features generated in Step 2, the present invention designs a cross - modal bi - directional fusion module (Bi - directional Fusion Module, BFM) to preliminarily fuse cross - modal features and prepare for the three - stream difference supervision mechanism in the next stage.

[0017] 3.1 First, a 3×3 convolution operation is used to enhance the receptive field information, and then two cross - modal features are obtained by cross - multiplication, which are used to enhance the RGB and Depth features respectively, defined as follows:[[]]

[0018]

[0019] where α ∈ {r, d}, i ∈ {1, 2, 3, 4} represents the layer where the feature is located in the encoder, and Sigmoid represents the sigmoid activation function. Thus, the enhanced RGB feature and Depth feature can be generated.

[0020] 3.2 The enhanced RGB feature and Depth feature generated in Step 3.1 are fused through a concatenation operation, and the operation is as follows:[[]]

[0021]

[0022] where cat represents the concatenation operation, and BCov represents the convolution operation and batch normalization (Batch Normal).

[0023] 4. The three-stream difference supervision mechanism proposed by the present invention is used to achieve the differential fusion between multiple modalities. Specifically, it can be expressed as three branches, which are respectively represented as the RGB branch, the Depth branch, and the Fusion branch.

[0024] 4.1 RGB features generated based on the Swin Transformer in step 2 Construct the RGB branch in the three-stream difference supervision mechanism, and use the cascaded aggregation decoder proposed by the present invention to predict the saliency map. Before the RGB features are input into the CAD, the present invention uses the ASPP technique to strengthen the receptive field of the RGB features and enhance the global information of the RGB features. And use the significant object segmentation map S GT for supervised learning. The operations of the RGB branch are described as follows:

[0025]

[0026] Among them, CAD represents the cascaded aggregation decoder, A represents the ASPP technique, represents the saliency map predicted by the RGB branch.

[0027] 4.2 Depth features generated based on the Swin Transformer in step 2 Construct the Depth branch in the three-stream difference supervision mechanism, and use the cascaded aggregation decoder proposed by the present invention to predict the saliency map. Before the Depth features are input into the cascaded aggregation decoder, the present invention uses the ASPP technique to strengthen the receptive field of the Depth features and enhance the global information of the Depth features. And use the significant object segmentation map S GT for supervised learning. The operations of the Depth branch are described as follows:

[0028]

[0029] Among them, CAD represents the cascaded aggregation decoder, A represents the ASPP technique, represents the saliency map predicted by the RGB branch.

[0030] 4.3 Cross-modal fusion features generated based on step 3 Use the four obtained fusion features to construct the Fusion branch, and use the significant object edge segmentation image for supervised learning. Use the cascaded aggregation decoder to integrate the four-scale features and predict the significant object edge map. The Fusion branch is defined as follows:

[0031]

[0032] 5. RGB saliency prediction maps formed based on the three-stream difference supervision mechanism described in step 4 and Depth saliency prediction maps and the predicted salient object segmentation maps The present invention designs a difference supervision module that utilizes and to fuse RGB features and Depth features.

[0033] 5.1 Use an interactive method to respectively constrain RGB features and Depth features. Specifically, use to constrain Depth features, use to constrain RGB features, and then use to constrain the fused features.

[0034] The process is as follows:

[0035]

[0036] 5.2 Based on the three-stream enhanced features (RGB enhanced features, Depth enhanced features, and Fusion enhanced features) obtained in step 5.1, use the channel attention mechanism to enhance the correlation degree of the channel dimension. Finally, use the concatenation operation to obtain the final difference fusion feature, which is defined as follows:

[0037]

[0038]

[0039] where CA represents the channel attention mechanism, and F i represents the difference fusion feature.

[0040] 6. Based on steps 4 and 5, the present invention designs a cascaded aggregation decoder structure for saliency inference. And embed this cascaded aggregation decoder structure into the three-stream difference supervision mechanism and the final saliency result prediction.

[0041] 6.1 The cascaded aggregation decoder uses a top-down method to gradually aggregate multi-scale features and generates an attention mask map through the spatial attention mechanism to enhance the next-level features, which is defined as follows:

[0042] F3 = UP(F4) + F3 × SA(F4) Equation (11)

[0043] where UP represents the upsampling operation, and SA represents the spatial attention mechanism.

[0044] 6.2 Repeat the operation in step 6.1 above to obtain the second-layer features, the first-layer features of the cascaded aggregation decoder. Finally, use the sigmoid activation function for the cascaded aggregation decoder to process the underlying features to obtain the final prediction S. pre 。

[0045] 7) The saliency map S predicted by the present invention pre and the manually annotated saliency object segmentation map S GT are used to calculate the loss function, and the parameter weights of the model proposed by the present invention are gradually updated through the Adam optimizer and the backpropagation algorithm. Finally, the structure and parameter weights of the RGB-D saliency object detection algorithm are determined.

[0046] 8) On the basis of determining the structure and parameter weights of the model in steps 2-6, the RGB-D image pairs in the test set involved in step 1 are tested to generate saliency maps, and evaluation metrics such as MAE, S-measure, F-measure, and E-measure are used for evaluation.

[0047] The RGB and Depth multi-modal saliency object detection based on the Swin Transformer network in the present invention mainly starts from the perspective of the differences between multi-modal data and proposes a novel RGB-D saliency object detection method based on a multi-modal difference fusion network. This method predicts the understanding and reasoning of saliency for different modalities from the RGB branch, the Depth branch, and the Fusion branch respectively, and integrates the differences of multi-modal through the proposed multi-modal difference fusion module. Compared with the previous RGB-D saliency object detection methods, the present invention has the following advantages:

[0048] (1) The present invention uses Swin Transformer as the encoder to extract RGB and Depth features, and the multi-modal features based on Swin Transformer can extract global context dependencies. (2) The present invention designs a three-stream difference supervision mechanism to respectively perceive the differences in the saliency expression of the RGB modality and the Depth modality. (3) The present invention designs a multi-modal difference fusion module to fuse the differences between the RGB and Depth modalities to achieve the effect of mutual enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 shows the overall structural schematic diagram of the present invention

[0050] Figure 2 shows the schematic diagram of the bidirectional fusion module proposed by the present invention

[0051] Figure 3Indicates the multi-modal difference fusion module proposed by the present invention

[0052] Figure 4 Indicates the cascaded aggregation decoder proposed by the present invention

[0053] Figure 5 Indicates the result comparison diagram of the present invention and other RGB-D saliency object detection methods Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. In addition, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art in this research direction without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.

[0055] Refer to the attached Figure 1 , an RGB-D saliency object detection method based on a multi-modal difference fusion network mainly includes the following steps:

[0056] 1. Obtain and organize the RGB-D data sets for training and testing.

[0057] 1.1) Inductively organize the obtained RGB-D data sets (DUT-RGB data set, NJU2K data set, NLPR data set, LFSD data set, RGBD135 data set), and divide a single sample into an RGB image P RGB , a depth image P depth , a manually annotated saliency object segmentation image S GT and a manually annotated saliency object edge segmentation image E GT .

[0058] 1.2) Divide the collected RGB-D data sets into a training set and a test set. The training set is a 2985-sample set composed of 800 samples in the DUT data set, 1400 samples in the NJU2K data set, and 650 samples in the NLPR data set. The remaining samples of the above five data sets are used as the test set.

[0059] 2. The present invention uses the SwinTransformer network in deep learning as the backbone network of the present invention to extract RGB and Depth features.

[0060] 2.1 Construct two SwinTransformer-based encoders to extract RGB features and Depth features respectively. Among them, the Swin Transformer encoder is composed of four basic Swin Transformer blocks, and its definition is as follows:

[0061] S = MLP(LN(W m (LN(F f )) + F f )) + W m (LN(F f )) + F f Formula (1)

[0062] ST = MLP(LN(W s (LN(S)) + S)) + W s (LN(S)) + S Formula (2)

[0063] Wherein, MLP represents a multi - layer perceptron, LN represents layer normalization, W m represents the multi - head self - attention mechanism, W s represents the self - attention mechanism based on the transformed window.

[0064] 2.2 Based on Step 2.1, the outputs of the RGB and Depth encoders can be obtained, denoted as RGB features and Depth features

[0065] 3. Based on the RGB and Depth features generated in Step 2, the present invention designs a cross - modal bi - directional fusion module (Bi - directional Fusion Module, BFM) to preliminarily fuse cross - modal features and prepare for the three - stream difference supervision mechanism in the next stage.

[0066] 3.1 First, a 3×3 convolution operation is used to enhance the receptive field information, and then two cross - modal features are obtained by cross - multiplication, which are used to enhance the RGB and Depth features respectively, and are defined as follows:[[]]

[0067]

[0068] Wherein, α ∈ {r, d}, i ∈ {1, 2, 3, 4} represents the layer where the feature is located in the encoder, and Sigmoid represents the sigmoid activation function. Thus, the enhanced RGB feature and Depth feature can be generated.

[0069] 3.2 The enhanced RGB feature and Depth feature generated in Step 3.1 are fused through a concatenation operation, and the operation is as follows:[[]]

[0070]

[0071] Among them, "cat" represents the concatenation operation, and "BCov" represents the convolution operation and batch normalization (Batch Normal).

[0072] 4. The three-stream differential supervision mechanism proposed by the present invention is used to achieve differential fusion between multiple modalities. Specifically, it can be represented as three branches, namely the RGB branch, the Depth branch, and the Fusion branch.

[0073] 4.1 RGB features generated based on the Swin Transformer in step 2 Construct the RGB branch in the three-stream differential supervision mechanism, and use the cascade aggregation decoder proposed by the present invention to predict the saliency map. Before the RGB features are input into the CAD, the present invention uses the ASPP technology to enhance the receptive field of the RGB features and enhance the global information of the RGB features. And use the significant object segmentation map S GT for supervised learning. The operations of the RGB branch are described as follows:

[0074]

[0075] Among them, "CAD" represents the cascade aggregation decoder, "A" represents the ASPP technology, represents the saliency map predicted by the RGB branch.

[0076] 4.2 Depth features generated based on the Swin Transformer in step 2 Construct the Depth branch in the three-stream differential supervision mechanism, and use the cascade aggregation decoder proposed by the present invention to predict the saliency map. Before the Depth features are input into the cascade aggregation decoder, the present invention uses the ASPP technology to enhance the receptive field of the Depth features and enhance the global information of the Depth features. And use the significant object segmentation map S GT for supervised learning. The operations of the Depth branch are described as follows:

[0077]

[0078] Among them, "CAD" represents the cascade aggregation decoder, "A" represents the ASPP technology, represents the saliency map predicted by the RGB branch.

[0079] 4.3 Cross-modal fusion features generated based on step 3 Use the four obtained fusion features to construct the Fusion branch, and use the significant object edge segmentation image for supervised learning. Use the cascade aggregation decoder to integrate the four-scale features and predict the significant object edge map. The Fusion branch is defined as follows:

[0080]

[0081] 5. RGB saliency prediction map formed based on the three-stream difference supervision mechanism described in step 4 and Depth saliency prediction map and the predicted salient object segmentation map The present invention designs a difference supervision module, which uses and to fuse RGB features and Depth features.

[0082] 5.1 Use an interactive method to respectively constrain RGB features and Depth features. Specifically, use to constrain Depth features, use to constrain RGB features, and then use to constrain the fused features.

[0083] The process is as follows:

[0084]

[0085] 5.2 Based on the three-stream enhanced features (RGB enhanced feature, Depth enhanced feature, and Fusion enhanced feature) obtained in step 5.1, use the channel attention mechanism to enhance the correlation degree of the channel dimension. Finally, use the concatenation operation to obtain the final difference fusion feature, which is defined as follows:

[0086]

[0087]

[0088] where CA represents the channel attention mechanism, and F i represents the difference fusion feature.

[0089] 6. Based on steps 4 and 5, the present invention designs a cascaded aggregation decoder structure for saliency inference. And embed this cascaded aggregation decoder structure into the three-stream difference supervision mechanism and the final saliency result prediction.

[0090] 6.1 The cascaded aggregation decoder adopts a top-down manner to gradually aggregate multi-scale features, and generates an attention mask map through the spatial attention mechanism to enhance the next-level features, which is defined as follows:

[0091] F3 = UP(F4) + F3 × SA(F4) Equation (11)

[0092] where UP represents the upsampling operation, and SA represents the spatial attention mechanism.

[0093] 6.2 Repeat the operation in step 6.1 to obtain the second-layer features, first-layer features of the cascaded aggregation decoder. Finally, use the sigmoid activation function for the cascaded aggregation decoder to process the bottom-layer features to obtain the final prediction S. pre 。

[0094] 7) The saliency map S predicted by the present invention pre and the manually annotated salient object segmentation map S GT are used to calculate the loss function, and the parameter weights of the model proposed by the present invention are gradually updated through the Adam optimizer and the backpropagation algorithm. Finally, the structure and parameter weights of the RGB-D saliency object detection algorithm are determined.

[0095] 8) Based on the structure and parameter weights of the model determined in steps 2-6, test the RGB-D image pairs in the test set involved in step 1 to generate saliency maps, and evaluate them using evaluation metrics such as MAE, S-measure, F-measure, and E-measure.

[0096] The above is the preferred embodiment of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A RGB-D salient object detection method based on a multi-modal difference fusion network, characterized in that, The method includes the following steps: 1) Use the Swin Transformer network in deep learning as the RGB and Depth encoders to extract hierarchical visual features of RGB and Depth images. Among them, the Swin Transformer encoder consists of four basic Swin Transformer blocks, which are defined as follows: S = MLP(LN(W m (LN(F f )) + F f )) + W m (LN(F f )) + F f Formula (1) ST = MLP(LN(W s (LN(S)) + S)) + W s (LN(S)) + S, Formula (2) Among them, MLP represents a multi-layer perceptron, LN represents layer normalization, and w m represents the multi-head self-attention mechanism, and w s represents the self-attention mechanism based on the shifted window; the outputs of the RGB and Depth encoders are denoted as RGB features and Depth features 2) The cross-modal bidirectional fusion module is used to preliminarily fuse cross-modal features to prepare for the three-stream difference supervision mechanism in the next stage; 3) Construct a three-stream difference supervision mechanism to achieve difference fusion between multiple modalities, which is represented by three branches, namely the RGB branch, the Depth branch, and the Fusion branch: 3.1) Construct the RGB branch in the three-stream difference supervision mechanism, and use a cascaded aggregation decoder to predict the saliency map; before the RGB features are input into the CAD, use the ASPP technique to enhance the receptive field of the RGB features, enhance the global information of the RGB features, and use the significant object segmentation map S GT Perform supervised learning, and the operations of the RGB branch are described as follows: Among them, CAD represents the cascaded aggregation decoder, and A represents the ASPP technology, represents the saliency map predicted by the RGB branch; 3.2) Use a cascaded aggregation decoder to predict the saliency map. Before the Depth feature is input into the cascaded aggregation decoder, use the ASPP technique to enhance the receptive field of the Depth feature, enhance the global information of the Depth feature, and use the salient object segmentation map S GT for supervised learning. The operations of the Depth branch are described as follows: Among them, CAD represents the cascaded aggregation decoder, and A represents the ASPP technology, represents the saliency map predicted by the RGB branch; 3.3) Based on the cross-modal fusion features generated in step 2.2 Using the four obtained fusion features, construct the Fusion branch, and use the salient object edge segmentation image for supervised learning. Integrate the four-scale features with a cascaded aggregation decoder to predict the salient object edge map. The Fusion branch is defined as follows: 4) Explore the three-stream difference supervision mechanism to generate RGB saliency prediction maps and Depth saliency prediction maps and predicted salient object segmentation maps and design a difference supervision module that utilizes and to fuse RGB features and Depth features; 5) Aggregate the second-layer features and the first-layer features of the obtained cascaded aggregation decoder, and use the sigmoid activation function for the cascaded aggregation decoder to process the underlying features to obtain the final prediction S pre , and the predicted saliency map S pre is used to calculate the loss function with the manually annotated salient object segmentation map S GT . Then, the parameters and weights of the model are gradually updated through the Adam optimizer and the backpropagation algorithm, and finally, the structure and parameter weights of the RGB-D saliency object detection algorithm are determined.

2. The RGB-D salient object detection method based on a multi-modal difference fusion network according to claim 1, wherein The above step 2) is specifically as follows: 2.1) First, use a 3×3 convolution operation to enhance receptive field information, and then use the cross-multiplication method to obtain two cross-modal features, which are used to enhance RGB and Depth features respectively, and are defined as follows: where α ∈ {r, d}, i ∈ {1, 2, 3, 4} represents the level at which the feature is located in the encoder, and Sigmoid represents the sigmoid activation function. Thus, the enhanced RGB features and Depth features can be generated; 2.2) Fuse the enhanced RGB features generated in step 2.1 and the Depth features through a concatenation operation, which is described as follows: Among them, cat represents the concatenation operation, and BCov represents the convolution operation and batch normalization.

3. The RGB-D salient object detection method based on a multi-modal difference fusion network according to claim 1, wherein The above step 4) is specifically as follows: 4.1) Enhance the RGB feature and the Depth feature respectively using an interactive method. Use to enhance the Depth feature, use to enhance the RGB feature, and then use to enhance the fused feature. The process is as follows: 4.2) Based on the three-stream enhanced features (RGB enhanced feature, Depth enhanced feature, and Fusion enhanced feature) obtained in step 2.1, use the channel attention mechanism to enhance the correlation degree of the channel dimension. Finally, use the concatenation operation to obtain the final difference fusion feature, which is defined as follows: Among them, CA represents the channel attention mechanism, and F i represents the differential fusion feature; 4.3) Design a cascaded aggregation decoder structure for the inference of salient objects, and embed the cascaded aggregation decoder structure into the three-stream difference supervision mechanism and the final salient result prediction; 4.4) Use the cascaded aggregation decoder to gradually aggregate multi-scale features in a top-down manner, and generate an attention mask map through the spatial attention mechanism to enhance the next-level features, which is defined as follows: F3 = UP(F4)+F3×SA(F4) Equation (11) Among them, UP represents the upsampling operation, and SA represents the spatial attention mechanism.

Citation Information

Patent Citations

  • RGB-D image saliency target detection method based on cross-modal feature fusion

    CN113076957A

  • RGB-D image-based CLANet steel rail surface defect detection system and method

    CN114170174A