A method for saliency detection of RGBD images based on depth estimation
By introducing depth estimation and feature fusion technology in RGBD significance detection, the problem of incomplete object detection under low-quality depth maps is solved, and higher detection accuracy and better significance object recognition effect are achieved.
Patent Information
- Application Number
- CN202210944652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-08-08
AI Technical Summary
When the existing RGBD significance detection method processes low-quality depth maps, the target detection is incomplete and the prediction accuracy is low, making it difficult to accurately detect significant targets in complex backgrounds and low contrast scenarios.
Using a depth estimation-based method, RGB features are extracted for depth estimation, an estimated depth map is generated, and the original depth map and estimated depth map are fused as the depth mode input of the cross-modal fusion module, and feature fusion is combined with attention mechanism and multiple fusion methods (addition and multiplication) to achieve RGBD significance detection.
Improved the accuracy and integrity of significance detection, especially performance in low-quality depth maps and complex backgrounds, and enhanced the ability to identify significant targets.
Smart Images

Figure CN115272268B_ABST
Abstract
Description
Technical Field
[0001] The technical solution of the present invention relates to the field of image data processing, and more specifically to an RGBD saliency detection method based on depth estimation. Background Art
[0002] Saliency detection aims to locate the most attractive and prominent objects in a scene, which is a fundamental research in computer vision applications and plays a key role in a series of computer vision applications, such as image understanding, video segmentation, and semantic segmentation. Although RGB saliency detection methods have made great progress by introducing convolutional neural networks, RGB images lack sufficient spatial information, and the single-modal input makes saliency detection still face many challenges, such as small target tasks, complex backgrounds, and low-contrast scenes. The RGBD saliency detection technology that combines RGB images and corresponding depth maps can overcome the above problems to a certain extent, so it has received the attention and research of researchers and the industry.
[0003] Most traditional RGBD saliency detection methods are based on low-level handcrafted features (such as color, edges, textures, etc.). These methods are typical heuristic methods. The representability of low-level features is limited, which restricts the generalization ability and results in low prediction accuracy. In recent years, deep neural networks have begun to be applied to the field of RGBD saliency detection. Deep learning methods refer to using convolutional neural networks to extract high-level semantic features of images to calculate the saliency value of images. However, low-quality depth maps will introduce a large amount of noise, and there are still problems in the existing technology such as incomplete detection of salient objects and inaccurate algorithm detection when the quality of depth maps is low. Summary of the Invention
[0004] The present invention provides an RGBD saliency detection method based on depth estimation. The present invention uses the extracted RGB features to perform depth estimation through a depth estimation module to obtain an estimated depth map, extracts depth features from the original depth map and the estimated depth map, and uses a depth fusion module (DFM) for fusion as the depth modality input of a cross-modal fusion module (CFM). Then, the RGB features and the fused depth features are fused using the cross-modal fusion module (CFM) to complete the RGBD saliency detection based on depth estimation.
[0005] The technical solution of the present invention is as follows:
[0006] An RGBD image saliency detection method based on depth estimation, characterized by including the following:
[0007] Obtain RGBD image data, including an RGB image and a depth map;
[0008] Construct an encoder-decoder network: The encoder includes a two-stream backbone network for RGB and depth, a depth estimation module, a depth fusion module, and a cross-modal fusion module. The two-stream backbone networks are both constructed based on VGG-16, denoted as the VGG-16rgb backbone network and the VGG-16depth backbone network respectively, which are used to extract the appearance features of RGB images and the spatial features of depth maps; the RGB features at each scale extracted by the VGG-16rgb backbone network are subjected to depth estimation by the depth estimation module to obtain the estimated depth map D e ; the number of scales of the RGB features extracted by the VGG-16rgb backbone network is n, where i = 1 to n is an integer;
[0009] The original depth map D o and the estimated depth map D e are each separately input into the same VGG-16depth backbone network to obtain the original depth features at each scale and the estimated depth features at each scale
[0010] The depth fusion module adaptively fuses the depth features at the corresponding scales of the two depth streams of the original depth map D o and the estimated depth map D e through channel adaptive weighting to obtain the fused depth features
[0011] The RGB features at the corresponding scale and the fused depth features are adaptively fused using the cross-modal fusion module to obtain the cross-modal fusion features at the corresponding scale
[0012] The decoder is composed of n saliency detection decoders connected in series. Each saliency detection decoder Decoderi performs upsampling at the corresponding scale to obtain saliency maps S at different scales i and selects the saliency map S at the largest scale 1 as the final saliency prediction result map S;
[0013] Calculate the loss of the constructed encoder-decoder network for RGBD saliency detection based on depth estimation.
[0014] The depth estimation module includes a convolutional unit Conv6 and multiple depth estimation decoders DED connected in series in sequence. The number of DEDs is n + 1. The convolutional unit Conv6 includes a max pooling layer and three convolutional layers. The input of the max pooling layer is VGG16 rgbThe last-scale RGB features, the output of the max pooling layer is sequentially connected to three cascaded convolutional layers; the output of the last convolutional layer is connected to multiple sequentially cascaded depth estimation decoders DED, from VGG16 rgb RGB features of different scales of the backbone network Integrated into the depth estimation decoder DED of the corresponding scale through skip connections, and the output of the last multiple depth estimation decoders DED is the estimated depth map D e ;
[0015] The depth estimation decoder DED includes two cascaded ConvBR units, and the output features of the previous depth estimation decoder DED and VGG16 rgb RGB features of the corresponding scale of the backbone network After the cascading operation, it is input into two cascaded ConvBR units, and the output is the input features of the next depth estimation decoder DED The ConvBR unit is an ordered operation composed of convolution, batch normalization, and ReLU function; there is no cascading operation in the first depth estimation decoder DED connected to the convolution unit Conv6;
[0016] The depth fusion module includes a cascading operation, channel attention, and an element-wise multiplication operation. The input of the depth fusion module is the depth feature f of a certain scale extracted from the original depth map od and the depth feature f extracted from the estimated depth map of the corresponding scale ed , the input of the channel attention is the output of the cascading operation, and the output of the channel attention is multiplied element-wise with the output of the cascading operation to obtain the output of the depth fusion module, that is, the fused depth feature f d ;
[0017] The cross-modal fusion module includes two paths, each path includes a spatial attention map and a ConvBR3 unit, and the RGB feature f of a certain scale r is input into a spatial attention map SA r , and the output of the spatial attention map SA r is multiplied element-wise with the fused depth feature f of the corresponding scale d , and the result after the element-wise multiplication operation is added element-wise with the fused depth feature f of the corresponding scale d to output the depth enhancement feature f d' ;
[0018] The fused depth feature f of a certain scale d is input into a spatial attention map SA d , and the output of the spatial attention map SA d is multiplied element-wise with the RGB feature f of the corresponding scaler The result after performing element-wise multiplication operation is then combined with the RGB feature f at the corresponding scale r to output the RGB enhanced feature f through element-wise addition operation r' ;
[0019] The depth enhanced feature f d' and the RGB enhanced feature f r' are respectively processed by a ConvBR3 unit, and then the two outputs are respectively subjected to element-wise addition operation and element-wise multiplication operation. The results of the two operations are then cascaded and processed by a third ConvBR3 unit to obtain the RGBD fusion feature f rd , which is the output of the cross-modal fusion module CFM; the ConvBR3 unit consists of sequential operations of 3×3 convolution, batch normalization, and ReLU function;
[0020] The number of depth fusion modules and the number of cross-modal fusion modules are both the same as the number n of scales of the RGB features output by the VGG16 rgb backbone network.
[0021] The RGBD image data comes from the NJU2K dataset, STERE dataset, NLPR dataset, SIP dataset, or DES dataset.
[0022] A method for RGBD image saliency detection based on depth estimation, and the specific steps of the detection method are as follows:
[0023] In the first step, the RGB feature f of the RGB image I is extracted r :
[0024] The RGB image I is fed into the VGG16 rgb backbone network to extract the RGB feature f r , as shown in the following formula (1):
[0025] f r = VGG16 rgb (I) (1),
[0026] In formula (1), VGG16 rgb (·) is the VGG16 deep network,
[0027] The VGG16 deep network includes a convolutional layer, a pooling layer, a non-linear activation function Relu layer, and a residual connection;
[0028] In the second step, the estimated depth map D is obtained e :
[0029] In VGG16 rgbA Conv6 unit is added after the last block of the backbone network, which consists of a max pooling layer and three convolutional layers to expand the receptive field of the RGB stream; the present invention designs a depth estimation decoding process implemented in the style of a U-net network, and a total of six depth estimation decoders DED are set, from VGG16 rgb RGB features of different scales of the backbone network (i = 1, 2, 3, 4, 5) are integrated into the depth estimation decoder DED of the corresponding scale through skip connections, and RGB features of different scales are integrated in a bottom-up manner from the U-bottom (i = 1, 2, 3, 4, 5) and the output features of the depth estimation decoder DED of different scales (i = 1, 2, 3, 4, 5), the intermediate estimated depth maps of different levels obtained by the six depth estimation decoders Using the output features of the depth estimation decoder DED at the largest scale The corresponding intermediate estimated depth map As the estimated depth map D e , the specific formula is as follows:
[0030]
[0031] DED(f) = ConvBR(ConvBR(f)) (3),
[0032]
[0033]
[0034]
[0035] In formula (3), ConvBR() represents a sequential operation composed of convolution, batch normalization and ReLU function, f is the input feature, and in formula (5), [·, ·] represents a concatenation operation;
[0036] The third step is to calculate the loss of depth estimation:
[0037] To ensure that the intermediate estimated depth maps of different levels Are close to the original depth map D o , the SmoothL1 loss function is used during training, and the formula is as follows:
[0038]
[0039]
[0040] In formula (7), W and H represent the width and height of the depth map, and Δ(x, y) represents the error of each pixel (x, y) between the intermediate estimated depth map and the original depth map; when calculating the loss between the intermediate estimated depth map and the original depth map, it is necessary to perform a resize operation on the original depth map at an appropriate scale to make its size consistent with the intermediate estimated depth maps of different scales.
[0041] Step 4: Extract depth features f o and f e from the original depth map D od and the estimated depth map D ed respectively:
[0042] Feed the original depth map D o and the estimated depth map D e into the VGG16 depth backbone network respectively to extract depth features f od and f ed of different scales. The formula is as follows:
[0043] f od = VGG16 depth (D o ) (9),
[0044] f ed = VGG16 depth (D e ) (10),
[0045] In formulas (9) and (10), VGG16 depth (·) is the VGG16 depth network;
[0046] Step 5: Fuse the depth features f od and f ed :
[0047] Input the depth features f od and f ed of different scales obtained in the above Step 4 into a depth fusion module DEM, and fuse them under the guidance of channel attention. Concatenate the corresponding-scale f od and f ed to obtain the concatenated depth features, and extract the channel attention of the concatenated depth features. Multiply it with the concatenated depth features through a residual connection to obtain the fused depth feature f d . The formula is as follows:
[0048] CA(f) = FC σ (FC φ (GMP s (f))) (11),
[0049]
[0050] In formula (11), FC σ (·) represents a fully connected layer with a Sigmoid activation function, FC φ (·) represents a fully connected layer with a ReLU activation function, GMP s (·) represents a global maximum pooling operation in space, f is the input feature, in formula (12), represents element-wise multiplication, [·,·] represents a concatenation operation;
[0051] Step 6, fuse the RGB feature f r and the fused depth feature f d :
[0052] The RGB feature f r obtained in the first step above and the fused depth feature f d obtained in the fifth step are cross-enhanced under the guidance of spatial attention using the cross-modal fusion module CFM. Then, the RGB feature f r and the fused depth feature f d The features of these two modalities are fused using element-wise addition and element-wise multiplication respectively, and the features obtained by the two fusion methods are fused to obtain the RGBD fused feature f rd , and the specific operations are as follows:
[0053] Extract the spatial attention maps SA r of the RGB feature f d obtained in the first step above and the fused depth feature f d and SA r , and the formula is as follows:
[0054] SA r = Sigmoid(Conv(f r )) (13),
[0055] SA d = Sigmoid(Conv(f d )) (14),
[0056] In formulas (13) and (14), Sigmoid(·) refers to the Sigmoid activation function, and Conv(·) is a convolution operation,
[0057] To retain the original information of each modality, the enhanced features are combined with their original features using residual connections to obtain the cross-enhanced features f r' and f d' , and the formula is as follows:
[0058]
[0059]
[0060] In formulas (15) and (16), f r' represents the RGB enhanced feature, and f d' represents the depth enhanced feature. represents element-wise multiplication.
[0061] After obtaining the enhanced features of the two modalities, element-wise addition and element-wise multiplication are used for fusion respectively, and the cross-enhanced features of the two different fusion methods are concatenated together. The convolution operation is used to adaptively weight these two features to obtain the RGBD fusion feature f rd , and the formula is as follows:
[0062]
[0063]
[0064] f rd = Relu(BN(Conv([f add , f mul ))) (19),
[0065] In formula (17), f add represents the addition fusion feature, represents element-wise addition, and ConvBR3() represents a sequential operation composed of a 3×3 convolution, batch normalization, and ReLU function. In formula (18), f mul represents the multiplication fusion feature, represents element-wise multiplication. In formula (19), Conv(·) is the convolution operation, BN(·) is the batch normalization operation, Relu(·) is the non-linear activation function, and [·, ·] represents the concatenation operation;
[0066] Step 7: Obtain the final saliency prediction result map S:
[0067] In order to obtain a more refined saliency map, the deep features of the 5th layer of the encoder network (i.e., the RGBD fusion feature ) are integrated into the shallow layer in a bottom-up manner and combined with the shallow layer features to obtain more accurate semantics, thereby obtaining the final saliency prediction result map S; specifically, a saliency detection decoder with the same number of scales as the RGB features output by the VGG16 rgb backbone network is set. In this embodiment, there are five saliency detection decoders. The input of the saliency detection decoder is the corresponding scale of the RGBD fusion feature f rd , and the five saliency detection decoders are connected in series and output saliency maps S of different scales respectively.i , select the saliency map S with the largest scale 1 as the final saliency prediction result map S;
[0068] Eighth step, calculate the loss of saliency detection:
[0069] To measure the difference between the saliency maps S of different scales obtained in the above seventh step i and the ground truth G, the binary cross-entropy loss function is used to calculate the saliency loss L during training s , as shown in the following formula (20):
[0070] L s (S i , G) = G log S i + (1 - G) log(1 - S i ) (20),
[0071] During the calculation of the ground truth G, a resize process is required to ensure that it is consistent with the size of the saliency map at the corresponding scale.
[0072] Ninth step, calculate the total loss L total :
[0073] The total loss L total combines the depth estimation loss L in the third step d and the saliency loss L in the eighth step s , as shown in the following formula (21):
[0074]
[0075] The network is trained by continuously reducing the size of L total , and the total loss L is optimized using the stochastic gradient descent method total ;
[0076] Thus, the RGBD saliency detection based on depth estimation is completed.
[0077] The present invention also protects an RGBD image saliency detection system, which is characterized by including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor runs the computer program, the process of the above-mentioned RGBD image saliency detection method based on depth estimation is realized.
[0078] Compared with the prior art, the beneficial effects of the present invention are:
[0079] The prominent substantive features of the present invention are:
[0080] The method of the present invention proposes a depth estimation module, which uses RGB features for depth estimation to generate an estimated depth map as a supplement to the original depth map. After the two are fused, they are used as the input of the depth modality. The fused depth features provide more spatial information for the network to help locate significant targets. The features of the two modalities use a cross-modal fusion module to complement and select the advantageous parts of the two modalities, and more effective features can be obtained. Specifically, after strengthening the features using spatial attention, two methods, addition and multiplication, are used to fuse the features of the two modalities. Addition utilizes feature complementarity, and multiplication emphasizes feature commonality more. Then, the fused features obtained by the two methods are adaptively fused, combined with the calculation of the total loss, to achieve end-to-end training and significantly improve the detection accuracy.
[0081] The significant progress of the present invention is as follows:
[0082] (1) The present invention does not discard the low-quality depth map. Feature fusion is performed during the encoding stage, and an estimated depth map is generated as a supplement to the depth modality, thereby obtaining more effective depth features, extracting more spatial information, making the depth features extracted by the encoder (the first to the sixth steps belong to the encoder part. The seventh step, the bottom-up process, belongs to the decoder part.) not easily affected by the low-quality depth map, the extracted features contain less noise, and the prediction accuracy is higher.
[0083] (2) The method adopted by the present invention is a depth-estimation-based method. The features extracted from the original depth map and the estimated depth map are fused as the input of the depth modality, and then the attention mechanism and multiple fusion methods (addition and multiplication) are used to combine the complementarity and commonality of the features of the two modalities. The extracted features are more refined, the prediction accuracy is higher, and the effect of image saliency detection is improved.
[0084] (3) The present invention uses a depth fusion module to adaptively select depth features for fusion between the estimated depth map and the original depth map under the guidance of channel attention, thereby obtaining more effective depth features. A cross-modal fusion module based on spatial attention is used to enhance the features of the two modalities and generate more complete saliency features, so as to achieve accurate detection of RGBD significant targets. Brief Description of the Drawings
[0085] The present invention will be further described below in conjunction with the drawings and embodiments.
[0086] Figure 1 It is a flowchart of the RGBD saliency detection method based on depth estimation of the present invention.
[0087] Figure 2 It is a saliency prediction result graph S of the RGB image I with a butterfly as the significant target in the embodiment of the present invention.
[0088] Figure 3It is a schematic diagram of the network structure of the RGBD saliency detection method based on depth estimation of the present invention.
[0089] Figure 4 It is a schematic diagram of the structure of the depth estimation module (DEM).
[0090] Figure 5 It is a schematic diagram of the structure of the depth fusion module (DFM).
[0091] Figure 6 It is a schematic diagram of the structure of the cross-modal fusion module (CFM).
[0092] Figure 7 It is a comparison chart of the results of the present invention and other RGBD saliency detection methods. Detailed implementation manners
[0093] Figure 1 The illustrated embodiments show that the process of the RGBD saliency detection method based on depth estimation of the present invention is as follows:
[0094] Prepare publicly available datasets commonly used for RGBD saliency object detection, such as NJU2K dataset, STERE dataset, NLPR dataset, SIP dataset, DES dataset. The pictures in each dataset are labeled with salient objects.
[0095] Extract the RGB feature f of the RGB image I r → Obtain the estimated depth map D e → Calculate the loss of depth estimation → From the original depth map D o and the estimated depth map D e Extract the depth features f od and f ed → Fuse the depth features f od and f ed → Fuse the RGB feature f r and the fused depth feature f d → Obtain the final saliency prediction result map S → Calculate the loss of saliency detection → Calculate the total loss L total → Complete the RGBD saliency detection based on depth estimation.
[0096] The two modalities in the present invention are the RGB modality and the depth modality. The features extracted from both the original depth map and the estimated depth map are fused and used as the input of the depth modality of the network, which can process low-quality depth images. In this application, low-quality depth images refer to those depth images that are affected by factors such as the environment and sensors, and the generated depth maps cannot highlight salient objects but instead contain a lot of noise, such as Figure 7 (f)'s Depth map.
[0097] The network structure of the present invention is asFigure 3 As shown, the inputs of the RGB image and the depth map are respectively set to 256×256×3 and 256×256×1. The encoder is a two-stream backbone network for RGB and depth. Both two-stream backbone networks are constructed based on VGG-16, denoted as the VGG-16rgb backbone network and the VGG-16depth backbone network respectively, which are used to extract the appearance features of the RGB image and the spatial features of the depth map. The RGB features at each scale extracted by the VGG-16rgb backbone network (i = 1, 2, 3, 4, 5) will undergo depth estimation through the Depth Estimation Module (DEM) to obtain the estimated depth map D e .
[0098] Next, the original depth map D o and the estimated depth map D e are each separately input into the same VGG-16depth backbone network to obtain the original depth features at each scale (i = 1, 2, 3, 4, 5) and the estimated depth features at each scale (i = 1, 2, 3, 4, 5).
[0099] Secondly, the Depth Fusion Module (DFM) adaptively fuses the depth features at each scale of the two depth streams of the above-mentioned original depth map D o and the estimated depth map D e through channel adaptive weighting to obtain the fused depth features (i = 1, 2, 3, 4, 5). This fused depth feature contains more spatial information about the salient objects.
[0100] Then, the RGB features at the corresponding scale and the fused depth features are adaptively fused using the Cross-modal Fusion Module (CFM) based on the spatial attention mechanism to obtain the cross-modal fusion features at different scales (i = 1, 2, 3, 4, 5), fully exploiting the complementarity between the RGB modality and the depth modality.
[0101] Finally, the saliency detection decoder Decoderi (i = 1, 2, 3, 4, 5) consists of two or three convolutional layers and one deconvolutional layer, where the deconvolutional layer is used to gradually restore the resolution. The bottom-up saliency detection decoder performs upsampling at five scales to obtain the saliency maps S i (i = 1, 2, 3, 4, 5), and the saliency map S at the largest scale 1 is selected as the final saliency prediction result map S, Figure 3 where G represents the ground truth map.
[0102] Figure 4 It is a schematic structural diagram of the depth estimation module DEM, including a convolutional unit Conv6 and multiple depth estimation decoders DED connected in series in sequence. The number of DEDs is the number of scales of the RGB features output by the backbone network + 1. The convolutional unit Conv6 includes a max pooling layer and three 3×3 convolutional layers. The input of the max pooling layer is the RGB feature of the last scale of VGG16 rgb The output of the backbone network is the RGB feature of the last scale of VGG16. The output of the max pooling layer is connected to three consecutive 3×3 convolutional layers in sequence to expand the receptive field of the RGB stream. The output of the last 3×3 convolutional layer is connected to multiple depth estimation decoders DED connected in series in sequence, which come from different scales of RGB features of the VGG16 backbone network rgb (i = 1, 2, 3, 4, 5) are integrated into the corresponding scale of the depth estimation decoder DED through skip connections. The output of the last multiple depth estimation decoders DED is the estimated depth map D rgb The RGB features of different scales of the backbone network (i = 1, 2, 3, 4, 5) are integrated into the corresponding scale of the depth estimation decoder DED through skip connections. The output of the last multiple depth estimation decoders DED is the estimated depth map D e .
[0103] The depth estimation decoder DED includes two consecutive ConvBR units. The output feature of the previous depth estimation decoder DED and the RGB feature of the corresponding scale of the VGG16 rgb backbone network are cascaded and then input into two consecutive ConvBR units, and the output is the input feature of the next depth estimation decoder DED The ConvBR unit is an ordered operation composed of convolution, batch normalization, and ReLU function. There is no cascading operation in the first depth estimation decoder DED connected to the convolutional unit Conv6
[0104] The depth estimation module DEM is a depth estimation decoding process in the style of a U-net network. In this embodiment, a total of six depth estimation decoders DED are set, which come from different scales of RGB features of the VGG16 rgb backbone network (i = 1, 2, 3, 4, 5) are integrated into the corresponding scale of the depth estimation decoder DED through skip connections, and the RGB features are integrated in a bottom-up manner from the bottom of the U (i = 1, 2, 3, 4, 5) and the output features of the depth estimation decoders DED of different scales (i = 1, 2, 3, 4, 5), where the intermediate estimated depth maps at different levels obtained by the six depth estimation decoders The output feature of the depth estimation module DEM is the output feature the corresponding intermediate estimated depth image Denoted as the estimated depth map D e .
[0105] Figure 5 is a schematic structural diagram of the Depth Fusion Module (DFM), including a cascading operation, channel attention, and an element-wise multiplication operation. The input of the Depth Fusion Module (DFM) is the depth feature f at a certain scale extracted from the original depth map od and the depth feature f extracted from the estimated depth map at the corresponding scale ed . The input of the channel attention is the output of the cascading operation. The output of the channel attention is multiplied element-wise with the output of the cascading operation to obtain the output of the Depth Fusion Module (DFM), that is, the fused depth feature f d . f od and f ed are cascaded to obtain the cascaded depth feature, and the channel attention of the cascaded depth feature is extracted and multiplied with the cascaded depth feature through a residual connection to obtain the fused depth feature f d , realizing the adaptive fusion of depth images under the guidance of channel attention
[0106] Figure 6 is a schematic structural diagram of the Cross-modal Fusion Module CFM. The Cross-modal Fusion Module CFM includes two paths, and each path includes a spatial attention map and a ConvBR3 unit. The RGB feature f at a certain scale r is input into a spatial attention map SA r . The output of the spatial attention map SA r is multiplied element-wise with the corresponding scale of the fused depth feature f d , and the result after the element-wise multiplication operation is then added element-wise with the corresponding scale of the fused depth feature f d to output and obtain the depth enhancement feature f d' ;
[0107] The fused depth feature f at a certain scale d is input into a spatial attention map SA d . The output of the spatial attention map SA d is multiplied element-wise with the corresponding scale of the RGB feature f r , and the result after the element-wise multiplication operation is then added element-wise with the corresponding scale of the RGB feature f r to output and obtain the RGB enhancement feature f r' ;
[0108] The depth enhancement feature f d' and the RGB enhancement feature f r'The two outputs respectively processed by a ConvBR3 unit are respectively subjected to element-wise addition operation processing and element-wise multiplication operation processing. The results of the two processes are then cascaded and processed by a third ConvBR3 unit to obtain the RGBD fusion feature f rd , which is the output of the cross-modal fusion module CFM. The ConvBR3 unit is a sequential operation composed of a 3×3 convolution, batch normalization, and ReLU function
[0109] The number of deep fusion modules (DFMs) and the number of cross-modal fusion modules CFM are both the same as the number of scales of the RGB features output by the VGG16 rgb backbone network
[0110] Embodiment 1
[0111] In this embodiment, the salient object is a butterfly. The method for RGBD saliency detection based on depth estimation described in this embodiment specifically includes the following steps
[0112] The first step is to extract the RGB feature f of the RGB image I r :
[0113] The RGB image I is fed into the VGG16 rgb backbone network to extract the RGB feature f r , as shown in the following formula (1)
[0114] f r = VGG16 rgb (I) (1)
[0115] In formula (1), VGG16 rgb (·) is the VGG16 deep network, which includes a convolutional layer, a pooling layer, a non-linear activation function Relu layer, and a residual connection
[0116] The second step is to obtain the estimated depth map D e :
[0117] A Conv6 unit is added after the last block of the VGG16 rgb backbone network. It consists of a max pooling layer and three convolutional layers to expand the receptive field of the RGB stream. The present invention designs a depth estimation decoding process implemented in the style of a U-net network, and a total of six depth estimation decoders DED are set. The RGB features of different scales from the VGG16 rgb backbone network (i = 1, 2, 3, 4, 5) are integrated into the corresponding scale depth estimation decoder DED through skip connections, and the RGB features of different scales are integrated in a bottom-up manner from the bottom of the U (i = 1, 2, 3, 4, 5) and the output features of the depth estimation decoder DED at different scales (i = 1, 2, 3, 4, 5), the intermediate estimated depth maps at different levels obtained by six depth estimation decoders Using the output features of the depth estimation decoder DED at the largest scale The corresponding intermediate estimated depth map As the estimated depth map D e , the specific formula is as follows:
[0118]
[0119] DED(f) = ConvBR(ConvBR(f)) (3),
[0120]
[0121]
[0122]
[0123] In formula (3), ConvBR() represents a sequential operation composed of convolution, batch normalization, and ReLU function, and f is the input feature. In formula (5), [·,·] represents the concatenation operation;
[0124] The third step is to calculate the loss of depth estimation:
[0125] To ensure that the intermediate estimated depth maps at different levels are close to the original depth map D o , the SmoothL1 loss function is used during training, and the formula is as follows:
[0126]
[0127]
[0128] In formula (7), W and H represent the width and height of the depth map, and Δ(x,y) represents the error of each pixel (x,y) between the intermediate estimated depth map and the original depth map;
[0129] The fourth step is to extract the depth features f o and f e from the original depth map D od and the estimated depth map D ed respectively:
[0130] Send the original depth map D o and the estimated depth map D e into VGG16 depthThe backbone network extracts depth features f at different scales respectively od and f ed , and the formula is as follows:
[0131] f od = VGG16 depth (D o ) (9),
[0132] f ed = VGG16 depth (D e ) (10),
[0133] In formulas (9) and (10), VGG16 depth (·) is the VGG16 deep network; both the RGB and depth modalities use VGG16 as the backbone network, but do not share parameters, and are two independent backbone networks.
[0134] Step 5, fuse the depth features f od and f ed :
[0135] Input the depth features f od and f ed at different scales obtained in the above Step 4 into a depth fusion module DEM, and fuse them under the guidance of channel attention. Concatenate the corresponding scale f od and f ed to obtain the concatenated depth features, and extract the channel attention of the concatenated depth features, and multiply it with the concatenated depth features through residual connection to obtain the fused depth feature f d , and the formula is as follows:
[0136] CA(f) = FC σ (FC φ (GMP s (f))) (11),
[0137]
[0138] In formula (11), FC σ (·) represents a fully connected layer with a Sigmoid activation function, FC φ (·) represents a fully connected layer with a ReLU activation function, GMP s (·) represents a spatial global maximum pooling operation, f is the input feature. In formula (12), represents element-wise multiplication, [·,·] represents a concatenation operation;
[0139] Step 6, fuse the RGB feature f r and the fused depth feature f d :
[0140] The RGB feature f obtained in the above first step r and the fused depth feature f obtained in the fifth step d are cross-enhanced under the guidance of spatial attention by using the cross-modal fusion module CFM. Then, the RGB feature f r and the fused depth feature f d of these two modalities are fused using element-wise addition and element-wise multiplication respectively, and the features obtained by the two fusion methods are fused to obtain the RGBD fused feature f rd , and the specific operations are as follows:
[0141] Extract the spatial attention maps SA r of the RGB feature f obtained in the above first step d and the fused depth feature f obtained in the fifth step d and SA r , and the formulas are as follows:
[0142] SA r = Sigmoid(Conv(f r )) (13),
[0143] SA d = Sigmoid(Conv(f d )) (14),
[0144] In formulas (13) and (14), Sigmoid(·) refers to the Sigmoid activation function, and Conv(·) is the convolution operation.
[0145] To retain the original information of each modality, the enhanced features are combined with their original features using residual connections to obtain the cross-enhanced features f r' and f d' of the two modalities, and the formulas are as follows:
[0146]
[0147]
[0148] In formulas (15) and (16), f r' represents the RGB enhanced feature, f d' represents the depth enhanced feature, represents element-wise multiplication,
[0149] After obtaining the enhanced features of the two modalities, they are fused using element-wise addition and element-wise multiplication respectively, and the cross-enhanced features of the two different fusion methods are concatenated, and the two features are adaptively weighted using the convolution operation to obtain the RGBD fused feature f rd, the formula is as follows:
[0150]
[0151]
[0152] f rd = Relu(BN(Conv([f add , f mul ))) (19),
[0153] In formula (17), f add represents the addition fusion feature, represents element-wise addition, ConvBR3() represents a sequential operation composed of a 3×3 convolution, batch normalization, and ReLU function. In formula (18), f mul represents the multiplication fusion feature, represents element-wise multiplication. In formula (19), Conv(·) is the convolution operation, BN(·) is the batch normalization operation, Relu(·) is the non-linear activation function, and [·, ·] represents the concatenation operation;
[0154] Step 7, obtain the final saliency prediction result map S:
[0155] To obtain a more refined saliency map, the deep features of the 5th layer of the encoder network (i.e., the RGBD fusion feature ) are integrated into the shallow layer in a bottom-up manner, combined with the shallow layer features to obtain more accurate semantics, and the final saliency prediction result map S is obtained; specifically, there is a saliency detection decoder with the same number of scales as the RGB features output by the VGG16 rgb backbone network. In this embodiment, there are five saliency detection decoders. The input of the saliency detection decoder is the corresponding scale of the RGBD fusion feature f rd , and the five saliency detection decoders are connected in series in sequence, and saliency maps S i of different scales are output respectively. Select the saliency map S 1 with the largest scale as the final saliency prediction result map S;
[0156] Step 8, calculate the loss of saliency detection:
[0157] To measure the difference between the saliency maps S i of different scales obtained in the above Step 7 and the ground truth G, the binary cross-entropy loss function is used to calculate the saliency loss L s during training, as shown in the following formula (20):
[0158] L s (S i , G) = G log Si +(1 - G)log(1 - S i ) (20),
[0159] During the calculation of the true value G, a resize process is required to ensure that it is consistent with the size of the saliency map at the corresponding scale.
[0160] Step 9: Calculate the total loss L total :
[0161] Total loss L total Combines the depth estimation loss L in Step 3 d and the saliency loss L in Step 8 s , as shown in the following formula (21):
[0162]
[0163] By continuously reducing the size of L total , the network is trained, and the total loss L is optimized using the stochastic gradient descent method total ;
[0164] Thus, the RGBD saliency detection based on depth estimation is completed.
[0165] Figure 2 This is the final saliency prediction result map S of the RGB image I of this embodiment, in which there is a significant object, a butterfly.
[0166] In the above embodiment, the VGG16 deep network, the true value, and the stochastic gradient descent method are all well-known in the technical field.
[0167] To more prominently show the effect of the saliency detection of the present invention, the detection results of the detection method of the present application are compared with those of other advanced models below, and the comparison results are as Figure 7 shown. The present invention is in Figure 7The visual comparison with eight advanced algorithms (shown in columns 5 - 12) is presented. None of these methods introduce the idea of depth estimation. The visual comparison includes ordinary life scenes (a) and various challenging scenes, including small objects (b), multiple objects (c), complex backgrounds (d), low - contrast scenes (e), and low - quality depth maps (f). First, (a) shows an ordinary life scene where the chair and basin in the foreground are prominent in the original RGB image, and the depth map also provides useful depth information. However, the existing models compared in this application still cannot completely label the chair as a salient object, while the present invention can completely detect this salient object. (b) shows an example of small objects. Larger non - salient objects in the scene are likely to cause confusion. The present invention is the only one that correctly labels the sculpture on the table, while the comparison methods all detect the table as a salient object without exception. When faced with a small - object scene, it is easy to mis - detect the larger object in the scene as a salient object, but the present invention does not make such a mistake. (c) shows an example of multiple objects. Even though the object on the far right in the depth map introduces misleading information, the present invention correctly and completely predicts three salient objects. In (d), the background is complex and the salient object is very small, but the present invention still produces reliable results, while other methods cannot identify the salient object from the complex background. In (e), the contrast between the salient object and the background is very low. The existing models cannot detect and segment the person on the left. In contrast, the present invention produces satisfactory results. Finally, (f) shows the case of a low - quality depth map. Besides the depth map not providing useful information, the salient object in the RGB image is also small and located in a complex scene. The present invention can more effectively utilize modality complementarity and depth estimation to eliminate the side effects of the depth map.
[0168] Figure 7 The saliency detection image in (a) is from the DES dataset, the saliency detection images in (b) and (c) are from the NLPR dataset, the saliency detection image in (d) is from the NJU2K dataset, the saliency detection image in (e) is from the SIP dataset, and the saliency detection image in (f) is from the STERE dataset. The pictures in each dataset are labeled with salient object tags.
[0169] Matters not described in the present invention are applicable to the prior art.
Claims
1. A method for saliency detection of RGBD images based on depth estimation, characterized in that, It includes the following: Obtain RGBD image data, including RGB images and depth maps; Construct an encoder-decoder network: The encoder includes a two-stream backbone network for RGB and depth, a depth estimation module, a depth fusion module, and a cross-modal fusion module. The two-stream backbone networks are both constructed based on VGG-16, denoted as the VGG-16rgb backbone network and the VGG-16depth backbone network respectively, which are used to extract the appearance features of RGB images and the spatial features of depth maps. The RGB features at each scale extracted by the VGG-16rgb backbone network are estimated by the depth estimation module to obtain the estimated depth map D e ; The number of scales of the RGB features extracted by the VGG-16rgb backbone network is n, where i = 1 to n is an integer; Input the original depth map D o and the estimated depth map D e separately into the same VGG-16 depth backbone network to obtain the original depth features at each scale and the estimated depth features at each scale The original depth map D is processed using a deep fusion module o and the estimated depth map D e The depth features of the corresponding scales of these two depth streams are adaptively fused through channel adaptive weighting, and the obtained fused depth features RGB features of the corresponding scale and fused depth features Use the cross-modal fusion module for adaptive fusion to obtain cross-modal fusion features of the corresponding scale The decoder is composed of n saliency detection decoders connected in series. Each saliency detection decoder Decoderi performs upsampling at the corresponding scale to obtain saliency maps S at different scales. i The saliency map S at the largest scale is selected. 1 as the final saliency prediction result map S; Calculate the loss of the constructed encoder-decoder network for RGBD saliency detection based on depth estimation.
2. The RGBD image saliency detection method based on depth estimation according to claim 1, wherein The depth estimation module includes a convolutional unit Conv6 and multiple depth estimation decoders DED connected in series in sequence. The number of DEDs is n + 1. The convolutional unit Conv6 includes a max pooling layer and three convolutional layers. The input of the max pooling layer is the RGB feature of the last scale of WGG16 rgb The output of the max pooling layer is connected to three convolutional layers connected in series in sequence. The output of the last convolutional layer is connected to multiple depth estimation decoders DED connected in series in sequence. The RGB features of different scales from the VGG16 rgb backbone network are integrated into the corresponding scale of the depth estimation decoder DED through skip connections. The output of the last multiple depth estimation decoders DED is the estimated depth map D e ; The depth estimation decoder DED includes two consecutive ConvBR units, and the output features of the previous depth estimation decoder DED and the corresponding scale RGB features of the VGG16 rgb backbone network are input into the two consecutive ConvBR units after concatenation operation, and the input features of the next depth estimation decoder DED are output The ConvBR unit is an ordered operation composed of convolution, batch normalization, and ReLU function; there is no concatenation operation in the first depth estimation decoder DED connected to the convolution unit Conv6; The depth fusion module includes a cascading operation, channel attention, and an element-wise multiplication operation. The input to the depth fusion module is the depth feature f at a certain scale extracted from the original depth map od and the depth feature f extracted from the estimated depth map at the corresponding scale ed . The input to the channel attention is the output of the cascading operation. The output of the channel attention is multiplied element-wise with the output of the cascading operation to obtain the output of the depth fusion module, that is, the fused depth feature f d ; The cross-modal fusion module includes two paths, each path includes a spatial attention map and a ConvBR3 unit, and the RGB feature f at a certain scale r is input into a spatial attention map SA r In the spatial attention map SA r the output of which is subjected to an element-wise multiplication operation with the fused depth feature f at the corresponding scale d and then the result is subjected to an element-wise addition operation with the fused depth feature f at the corresponding scale d to output the depth enhanced feature f d' ; The fused depth feature f at a certain scale d is input into a spatial attention map SA d where the output of the spatial attention map SA d is subjected to an element-wise multiplication operation with the RGB feature f at the corresponding scale r and then the result is subjected to an element-wise addition operation with the RGB feature f at the corresponding scale r to output the RGB enhanced feature f r' ; Depth-enhanced feature f d' and RGB-enhanced feature f r' After being processed by a ConvBR3 unit respectively, the two outputs are respectively subjected to element-wise addition operation and element-wise multiplication operation. The results of the two operations are then cascaded and processed by a third ConvBR3 unit to obtain the RGBD fusion feature f rd , which is the output of the cross-modal fusion module CFM; the ConvBR3 unit is a sequential operation composed of 3×3 convolution, batch normalization, and ReLU function; The number of deep fusion modules and the number of cross-modal fusion modules are both the same as the number of scales n of the RGB features output by the VGG16 rgb backbone network.
3. The method for saliency detection of RGBD images based on depth estimation according to claim 1, characterized in that, The RGBD image data comes from the NJU2K dataset, STERE dataset, NLPR dataset, SIP dataset, or DES dataset.
4. A saliency detection method for RGBD images based on depth estimation, characterized in that, The steps of this detection method are as follows: Step 1, extract the RGB feature f of the RGB image I r : Send the RGB image I into VGG16 rgb backbone network to extract RGB feature f r , as shown in the following formula (1): f r = VGG16 rgb (I) (1), In formula (1), VGG16 rgb (·) is the VGG16 deep network; Step 2: Obtain the estimated depth map D e : In VGG16 rgb A Conv6 unit is added after the last block of the backbone network, which consists of a max pooling layer and three convolutional layers; A depth estimation decoding process is designed in the style of a U-net network, and a total of six depth estimation decoders DED are set up, coming from VGG16 rgb RGB features of different scales from the backbone network (i = 1, 2, 3, 4, 5) are integrated into the depth estimation decoder DED of the corresponding scale through skip connections, and RGB features of different scales are integrated in a bottom-up manner from the U-bottom (i = 1, 2, 3, 4, 5) and the output features of the depth estimation decoder DED of different scales (i = 1, 2, 3, 4, 5), the intermediate estimated depth maps of different levels obtained by the six depth estimation decoders Using the output features of the depth estimation decoder DED at the largest scale The corresponding intermediate estimated depth map As the estimated depth map D e , the specific formula is as follows: DED(f) = ConvBR(ConvBR(f)) (3), In formula (3), ConvBR() represents a sequential operation composed of convolution, batch normalization, and the ReLU function, and f is the input feature. In formula (5), [·,·] represents the concatenation operation; Third, calculate the loss of depth estimation: During training, the SmoothL1 loss function is used to calculate the intermediate estimated depth maps at different levels and the original depth map D o for the depth estimation loss L d , and the formula is as follows: In formula (7), W and H represent the width and height of the depth map, and Δ(x,y) represents the error of each pixel (x,y) between the intermediate estimated depth map and the original depth map; Step 4: Extract depth features f o and f e from the original depth map D od and the estimated depth map D ed respectively: Send the original depth map D o and the estimated depth map D e into the VGG16 depth backbone network respectively to extract depth features f od and f ed at different scales. The formula is as follows: f od = VGG16 depth (D o ) (9), f ed = VGG16 depth (D e ) (10), In Formulas (9) and (10), VGG16 depth (·) is the VGG16 deep network; Step 5, fuse the depth feature f od and f ed : Input the depth features f of different scales obtained in the above fourth step od and f ed into a depth fusion module DEM, and fuse them under the guidance of channel attention. Fuse the f of the corresponding scale od and f ed by cascading to obtain the cascaded depth features, and extract the channel attention of the cascaded depth features, and multiply it with the cascaded depth features through residual connection to obtain the fused depth feature f d , and the formula is as follows: CA(f) = FC σ (FC φ (GMP s (f))) (11), In formula (11), FC σ (·) represents a fully connected layer with a Sigmoid activation function, FC φ (·) represents a fully connected layer with a ReLU activation function, GMP s (·) represents a global maximum pooling operation in the spatial domain, f is the input feature, in formula (12), represents element-wise multiplication, represents a concatenation operation; Step 6, fuse the RGB feature f r and the fused depth feature f d : The RGB feature f obtained in the first step above r and the fused depth feature f obtained in the fifth step d are cross-enhanced under the guidance of spatial attention. Then, the RGB feature f r and the fused depth feature f d The features of these two modalities are fused using element-wise addition and element-wise multiplication respectively, and the features obtained by the two fusion methods are fused to obtain the RGBD fusion feature f rd , and the specific operations are as follows: Extract the spatial attention maps SA of the RGB feature f obtained in the first step above r and the fused depth feature f obtained in the fifth step d respectively, and the formula is as follows: d and SA r , as follows: SA r = Sigmoid(Conv(f r )) (13), SA d = Sigmoid(Conv(f d )) (14), In formulas (13) and (14), Sigmoid(·) refers to the Sigmoid activation function, and Conv(·) is the convolution operation, The enhanced features are combined with their original features using residual connections to obtain cross-enhanced features f r' and f d' , as shown in the following formula: In formulas (15) and (16), f r' represents the RGB enhancement feature, and f d' represents the depth enhancement feature, represents element-wise multiplication, After obtaining the enhanced features of the two modalities, element-wise addition and element-wise multiplication are used for fusion respectively, and the cross-enhanced features of the two different fusion methods are concatenated together. Then, a convolutional operation is used to adaptively weight these two features to obtain the RGBD fusion feature f rd , and the formula is as follows: f rd = Relu(BN(Conv([f add , f mul ))) (19), In formula (17), f add represents the additive fusion feature, represents element-wise addition, ConvBR3() represents a sequential operation composed of a 3×3 convolution, batch normalization, and the ReLU function. In formula (18), f mul represents the multiplicative fusion feature, represents element-wise multiplication. In formula (19), Conv(·) is the convolution operation, BN(·) is the batch normalization operation, Relu(·) is the non-linear activation function, and [·,·] represents the concatenation operation; Seventh, obtain the final saliency prediction result map S: Integrate the RGBD fusion features Integrate them into the shallower layer in a bottom-up manner, combine them with the features of the shallower layer to obtain more accurate semantics, and obtain the final saliency prediction result map S; Specifically, five saliency detection decoders are set. The input of the saliency detection decoder is the RGBD fusion feature f of the corresponding scale rd , and the five saliency detection decoders are connected in series in turn, and saliency maps S of different scales are output respectively i , and the saliency map S of the largest scale is selected 1 as the final saliency prediction result map S; Eighth, calculate the loss of saliency detection: To measure the difference between the different-scale saliency maps S obtained in the above seventh step i and the ground truth G, the binary cross-entropy loss function is used during training to calculate the saliency loss L s , as shown in the following formula (20): L s (S i ,G) = G log S i + (1 - G) log(1 - S i ) (20), Step 9, calculate the total loss L total : Total loss L total Combined with the depth estimation loss L in the third step d and the saliency loss L in the eighth step s , as shown in the following formula (21): By continuously reducing the size of L total to train the network and using the stochastic gradient descent method to optimize the total loss L total ; Thus, the RGBD saliency detection based on depth estimation is completed.
5. An RGBD image saliency detection system, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor runs this computer program, it realizes the process of the RGBD image saliency detection method based on depth estimation described in any one of claims 1-4.
Citation Information
Patent Citations
Cross-modal interaction RGB-D image salient region detection method
CN114445618A
KR20220029335A