An RGB-D Image Salient Object Detection Method Based on Dual Backbone Networks
Through the RGB-D significance object detection method of dual backbone network, the problem of poor RGB-D image detection effect is solved through the implicit and explicit feature fusion module, and more accurate significance object detection is achieved.
Patent Information
- Application Number
- CN202211241392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-10-11
AI Technical Summary
The existing RGB-D significance object detection method fails to fully utilize the information difference between the two modes due to the simple tandem RGB image and the Depth image, resulting in poor detection effect and poor generalization.
The dual backbone network is used to detect the significance target of RGB-D image. Through the implicit and explicit multimodal feature fusion network module, the layer-by-layer complementary fusion of RGB images and Depth image features is realized. The dual attention cross-modal fusion and spatial-related feature fusion module are used to improve the detection accuracy.
It achieves significant target detection results with clear edges and accurate positioning, and improves detection effect and generalization capabilities.
Smart Images

Figure CN115908250B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of saliency object detection in computer vision using RGB images and Depth images, and particularly relates to an end-to-end object detection algorithm using deep learning technology. Background Art
[0002] Generally, image saliency object detection only processes RGB color images for detection, while RGB-D saliency object detection combines two image modalities, RGB color images and Depth, to further improve the performance of saliency object detection methods. However, traditional RGB-D saliency object detection extracts features manually, and the detection performance is limited by the feature extraction method, resulting in poor generalization and unsatisfactory detection effects of such methods. The development of deep learning has brought a significant improvement in object detection performance.
[0003] Most existing RGB-D saliency object detection models based on deep learning algorithms simply concatenate RGB images and Depth images as the input of the network model, then use an encoder-decoder structure for feature extraction and processing, and finally the decoder directly outputs a thresholded template (mask) image as the detection result. However, the imaging principles of RGB images and Depth images are two completely different ways, and the information they contain and the ways of presenting information are also very different. It is not accurate enough to simply treat the two equally. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and disclose an RGB-D image saliency object detection method based on a dual backbone network.
[0005] Technical Solution:
[0006] An RGB-D image saliency object detection method based on a dual backbone network, characterized in that, based on an RGB-D saliency object detection model of a dual backbone network, through the design of an implicit-explicit multimodal feature fusion network module, it realizes the full fusion of the feature information of RGB images and Depth images, and makes accurate decisions for saliency object detection.
[0007] The described RGB-D image saliency object detection method based on a dual backbone network is characterized in that the specific implementation method is as follows: In the encoding stage of the dual backbone network, the feature information between each network layer is fused through implicit dual-attention cross-modal fusion to achieve layer-by-layer complementary fusion of features in each modality; secondly, at the end of the encoder, through an explicitly re-fusion module for spatially related features, the fusion effect of multi-modal features is further enhanced; in the decoder stage, the features fused twice are decoded in sequence to obtain a saliency object detection result with clear edges and accurate positioning.
[0008] The core of the present invention is to design a multi-modal feature learning network to achieve full fusion of RGB image and Depth image data in two modalities, thereby improving the saliency object detection effect based on RGB-D images. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 is the structural diagram of the dual backbone network model proposed by the present invention
[0010] Figure 2 is the qualitative result of saliency object detection of the embodiment of the present invention on the experimental dataset
[0011] Figure 3 is the quantitative result of saliency object detection of the embodiment of the present invention on seven publicly available datasets DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the present invention with reference to examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0013] The present invention proposes an RGB-D saliency object detection model based on a dual backbone network. Through the design of an implicit-explicit multi-modal feature fusion network module, the model realizes full fusion of the feature information of RGB images and Depth images, and makes accurate decisions for saliency object detection.
[0014] Specifically, in the encoding stage of the dual backbone network, the feature information between each network layer is fused through implicit dual-attention cross-modal fusion to achieve layer-by-layer complementary fusion of features in each modality; secondly, at the end of the encoder, through an explicitly re-fusion module for spatially related features, the fusion effect of multi-modal features is further enhanced. In the decoder stage, the features fused twice are decoded in sequence to obtain a saliency object detection result with clear edges and accurate positioning.
[0015] Specifically, it includes the following steps:
[0016] S1. Obtain the RGB image R and the Depth image D with the same shooting scene in pairs, and scale both pairs of images to a size of 224X224;
[0017] S2. Send these two images into the two input ends of the model encoder respectively for feature extraction and fusion;
[0018] S3. Input the fused features into the decoder end, decode and obtain the output result of the model;
[0019] Among them, in the above-mentioned saliency target calculation process S2, it specifically includes the following steps: S2.1 - S2.3
[0020] S2.1. Multi-modal feature extraction of the dual backbone network
[0021] The two branches of this dual backbone network have similar network structures, and both use VGG-16 as the backbone: that is, it consists of five convolutional layers and one dilated convolutional layer, and is connected by Max-Pooling (maximum pooling layer) between layers. After each maximum pooling layer in each layer, the corresponding features of R and D can be obtained in the present invention, denoted as
[0022] S2.2. Cross-modal implicit feature fusion module (CSM) based on dual attention
[0023] Considering the features extracted by the shallow network and are mainly low-level features and contain more noise, which is not very helpful for saliency target detection. Therefore, in the present invention, starting from the output of the second layer network to the fifth layer network, cross-modal feature fusion of the corresponding R and D features between layers is performed. The cross-modal feature fusion module proposed in the present invention mainly includes spatial attention operation and channel attention operation. The following are introduced separately:
[0024] Since the spatial channel attention mechanism was proposed, its strengthening of key content and suppression of noise data have been widely confirmed. Therefore, in the present invention, the idea of spatial channel attention is borrowed, and the specific data flow and its calculation formula are as follows. First, perform spatial attention weighting operation on the features obtained by each layer of the backbone network:
[0025]
[0026]
[0027] In the formula refers to the feature maps of R and D in the nth layer of the backbone network, is the feature obtained after calculation by the spatial attention mechanism, It refers to the spatial attention maps of R and D calculated for the nth layer. The symbol “+” represents the corresponding addition of each element of the matrix. Among them, The calculation formula is as follows:
[0028]
[0029]
[0030] Among them, Sigmoid is the activation function, and Conv refers to the Convolution operation. After obtaining the spatially attention-weighted R and D features, in the channel attention stage, the corresponding feature calculation formula is as follows:
[0031]
[0032]
[0033] Among them, GAP refers to global average pooling, refers to the per-pixel multiplication operation, and Conv refers to the Convolution operation. Finally, the features and are concatenated respectively to obtain the features after multi-modal fusion for the nth layer:
[0034]
[0035] In the formula, Con refers to the concatenation operation (Concatenation).
[0036] S2.3, Spatial Correlated Feature Explicit Re-Fusion (SCFF)
[0037] The trained encoder convolutional layer can extract complex feature information from the picture. After five layers of convolution and pooling, an informative feature map will be obtained. Continuing to use convolution at this time will cause information loss and affect the performance of salient object detection. Using Atrous Spatial Pyramid Pooling (ASPP) dilated convolution for feature extraction can better preserve feature information. Therefore, in the last layer, that is, the present invention uses the dilated convolution combination of ASPP for further feature extraction. Finally, the present invention uses Spatial Correlated Feature Fusion (SCFF) to cross-fuse the features of the two modalities to obtain the multi-modal fusion feature F co The specific calculation formula is as follows:
[0038]
[0039]
[0040] Among them is the features respectively obtained after passing through two ASPPs. Conv refers to the Convolution operation
[0041] S3, the layer-by-layer decoding of features
[0042] The features after fusing the second to fifth layers are concatenated and then layer-by-layer input into the decoder Decoder for decoding. Its calculation formula is as follows
[0043] Y 1 = Conv(Up(F co )),
[0044] Y n = Conv(Con(Up(Y n-1 ), CSM 7-n )) n ∈ {2, 3, 4, 5},
[0045] where Up() refers to upsampling
[0046] The final output result of the Decoder is Y 5 a binary image, where white represents the salient object and the black area represents the background. This result is used as the salient object detection result of the RGB-D image
[0047] A dual backbone network for salient object detection in RGB-D images provided by an embodiment of the present invention includes the following steps
[0048] Step 1: Read the paired RGB image X and Depth image D in the dataset and scale the two images to a size of 224X224
[0049] Step 2: Send X and D into the encoders of the dual backbone network respectively to perform layer-by-layer extraction of two-modal features
[0050] Step 3: Send the layer-by-layer learned RGB image features and Depth image features into the cross-modal feature fusion module (CSM) based on dual attention and the spatial correlation feature explicit re-fusion module (SCFF) respectively to perform implicit-explicit information fusion of multi-modal features
[0051] Step 4: Send the implicitly-explicitly fused features, concatenated in sequence, into the decoder Decoder to decode and obtain the final output salient object detection result, the Mask image Y
[0052] Figure 2 is the qualitative result of salient object detection on the dataset used in the experiment
[0053] InFigure 2 Among them, by comparing the significant targets predicted by the present invention (the last column) with the significant targets manually marked (the third column), it can be seen that the prediction results of the present invention are very close to the true significant targets, further demonstrating the effectiveness of the method of the present invention.
[0054] Figure 3 They are the quantitative results of significant target detection on seven public data sets.
[0055] Figure 3 Among them, each row represents the detection quantization results of the method of the present invention on different data sets. From the quantization results of the four evaluation indexes (MAE, F-measure, E-measure, S-measure), it can be seen that the method of the present invention has achieved good significant target detection results. Among them, the smaller the value of MAE, the better (the minimum value is 0), and the larger the values of F-measure, E-measure and S-measure, the better, and their maximum values are all 1.
Claims
1. A saliency object detection method for RGB-D images based on a dual backbone network, characterized in that, Its RGB-D saliency object detection model based on a dual backbone network realizes the full fusion of the feature information of RGB images and Depth images through the design of an implicit-explicit multimodal feature fusion network module, and makes accurate decisions for saliency object detection; Specifically, it includes the following steps: S1, obtain pairs of RGB images with the same shooting scene and Depth Image , and scale the paired images to the same size; S2. Send this pair of images into the two input ends of the model encoder respectively for feature extraction and fusion; S3. Input the fused features into the decoder end to decode and obtain the output result of the model; Among them, in S2, it specifically includes the following steps: S2.1 - S2.3: S2.
1. Multimodal feature extraction of the dual backbone network; The two branches of the dual backbone network have similar network structures, both using VGG-16 as the backbone: that is, it consists of five convolutional layers and one dilated convolutional layer, connected by Max-Pooling between layers, and obtaining and the corresponding features, denoted as { 、 、 、 、 、 、 、 、 、 }; S2.
2. Cross-modal implicit feature fusion module based on dual attention; The cross-modal implicit feature fusion module includes spatial attention operation and channel attention operation; From the output of the second - layer network to the fifth - layer network, perform cross - modal feature fusion on the and features between corresponding layers; S2.
3. Explicit re-fusion of spatially correlated features; In the last layer, namely { 、 }, the dilated convolution combination of ASPP is further used for feature extraction; finally, Spatial Correlated Feature Fusion (SCFF) is used to cross-fuse the features of the two modalities to obtain the features after multi-modal fusion , and its specific calculation formula is as follows: , , where refers to series operation; Among them , are the features respectively obtained after passing through two ASPPs for the set { 、 }, refers to the Convolution convolution operation; S3. Layer-by-layer decoding of features; After splicing the features fused in the second to fifth layers, input them into the decoder Decoder layer by layer for decoding. Its calculation formula is as follows: , , Among them, refers to upsampling; The final output result of the decoder is a binarized image, where white represents the salient object and the black area represents the background; this result is used as the salient object detection result of the RGB-D image.
2. A method for RGB-D image saliency object detection based on a dual backbone network according to claim 1, characterized in that Specific implementation method: In the encoding stage of the dual backbone network, the feature information between network layers is fused layer by layer with complementary features of each modality through implicit dual-attention cross-modal fusion; Secondly, at the end of the encoder, through the designed explicit re-fusion module of spatially correlated features, the fusion effect of multimodal features is further enhanced; In the decoder stage, the features fused twice are decoded in sequence, and then a saliency object detection result with clear edges and accurate positioning is obtained.
3. A method for RGB-D image saliency object detection based on a dual backbone network according to claim 1, characterized in that In S1, both paired pictures are scaled to a size of 224X224.
4. A method for RGB-D image saliency object detection based on a dual backbone network according to claim 1, characterized in that Spatial channel attention operation mechanism: The data stream and its calculation formula are as follows: First, perform spatial attention weighting operation on the features obtained from each layer of the backbone network: , ; In the formula , denote and the feature maps of the n-th layer in the backbone network, , are the features obtained after calculation by the spatial attention mechanism, , denote the calculated spatial attention maps of the n-th layer and , and "+" denotes the element-wise addition of the matrices; among them, , are calculated as follows: , ; Among them, is an activation function, refers to the Convolution convolution operation; After obtaining the spatially attention-weighted and features, in the channel attention stage, the corresponding feature calculation formula is as follows: , ; Among them, refers to global average pooling, refers to per-pixel multiplication operation, refers to Convolution convolution operation; finally, the features and are concatenated respectively to obtain the features after multi-modal fusion of the nth layer: , In the formula means series operation.
Citation Information
Patent Citations
Cascade multi-mode fusion video target tracking method based on attention model
CN108171141A
Salient target detection method based on RGB-T multi-source image data
CN114898106A