RGB-D image saliency target detection method based on progressive hierarchical fusion
By adopting a progressive hierarchical fusion method in RGB-D image significance object detection, combining multi-scale enhancement module, multi-modal fusion module and edge enhancement module, the problem of information loss in the prior art is solved, and a significant object detection effect with higher accuracy and rich details is achieved.
Patent Information
- Application Number
- CN202510173288.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing RGB-D image significance object detection methods have the problem of information loss in multimodal feature fusion and hierarchical feature processing, especially the progressive upsampling method at the decoder stage leads to the loss of useful information.
Using a progressive hierarchical fusion method, the four-layer features of RGB and depth images are extracted through Swin Transformer, and feature enhancement and fusion are used in the backbone network. The decoder stage uses a multi-layer decoding structure and a two-layer fusion module for progressive hierarchical fusion, and low-level depth features are extracted in the edge enhancement module to increase edge details of the significant graph.
It effectively reduces the loss of useful information, improves the accuracy and detailed information of significance target detection, and shows competitive performance on multiple disclosed benchmark data sets.
Smart Images

Figure CN120107620A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a method for detecting salient objects in RGB-D images based on progressive hierarchical fusion. Background Art
[0002] Human visual perception is highly selective and attentive, and can quickly locate important objects in a scene. The salient object detection task is inspired by this mechanism and detects the most visually attractive objects or regions in the input data. With the continuous development of technology, salient object detection is widely used in engineering fields such as medicine, autonomous driving, and intelligent interaction, and also plays an important role in computer tasks such as image classification, image retrieval, data enhancement, and image quality assessment. At the same time, with the widespread application of depth sensors, many scholars have introduced depth information into the salient object detection task, using the spatial information provided by depth information to predict salient objects, which has promoted the development of salient object detection tasks in RGB-D images.
[0003] The salient object detection method of RGB-D images makes full use of the characteristics of RGB and depth information. By fusing the color information of RGB images and the spatial information of depth images, the feature information of the two modalities complement each other, and the salient map can be accurately predicted. In recent years, deep learning algorithms have been continuously improved and innovated. According to the modal fusion method, it can be divided into three types: early fusion, late fusion, and mid-term fusion. Early fusion refers to the direct fusion of the original information of RGB and depth images at the data input stage. This fusion method is easily affected by low-quality depth maps, while late fusion is the fusion of salient object detection results. The design is relatively simple and cannot make full use of multimodal information. Recently, most scholars have adopted mid-term fusion to fuse the multi-layer features of RGB and depth images extracted by the backbone network, which can learn cross-modal information and increase the robustness of the network. In 2024, Tang et al. proposed a deep mask guiding network in the article "DMGNet: Depth mask guiding network for RGB-D salient object detection". The salient objects of the depth image are pre-segmented to create a mask to enhance the feature extraction ability of RGB. In addition, the feature fusion pyramid model is used to enhance the cross-modal fusion results. In 2024, Jiang et al. proposed a global-aware interaction network in "Global-aware Interaction Network for RGB-D salient object detection" to effectively integrate RGB and depth images. The network uses an attention mechanism to enhance multimodal features and uses deep features to guide RGB features to reduce the redundancy generated during the fusion process. In 2024, Sun et al. proposed a bifurcated multimodal fusion network in "BMFNet: Bifurcated multi-modal fusion network for RGB-D salient object detection". The designed multimodal feature interaction module uses an attention mechanism to achieve RGB and deep feature complementarity, and introduces a multimodal feature fusion module to aggregate internal modalities of different groups. The multi-scale feature learning module learns contextual information of different scales.Wang et al. proposed a feature cross-dimensional hybrid network in the article "FCDHNet: A feature cross-dimensional hybrid network for RGB-D salient object detection" in 2025, which uses self-attention and dilated convolution to enhance the convolution extraction ability of high-level features, and uses the SLIC algorithm (Simple Linear Iterative Clustering) to generate superpixel mapping to extract channel information. Finally, a hybrid RGB feature enhancement module is used to improve multimodal feature fusion and aggregate useful information of RGB and depth images. CN118334365A discloses an HDFNet model for salient object detection in RGB-D images. The model uses an asymmetric fusion module to fuse high-level modal information in the encoding stage. Although the dynamic expansion pyramid model is used to optimize the decoder features, the decoder still uses a progressive upsampling method, which is easy to lose useful information when integrating hierarchical features. CN119027771A discloses a multimodal variational fusion method for RGB-D salient target detection, which calculates shared latent features and private latent features based on extracted high-level features, and generates attention fusion based on the two features, but does not process too much fused hierarchical features on the decoder side, and lacks detailed information of low-level features. CN119152176A discloses an RGB-D salient target detection method in a multi-target scene, which estimates the original depth map using an RGB-based estimated depth map model to reduce the impact of poor-quality depth maps, and learns the correlation between hierarchical features before multimodal feature fusion, and then uses an attention mechanism to fuse modal information, but the decoder side only performs simple decoding of RGB, depth and fusion features, and lacks processing of public hierarchical features.
[0004] However, the above-mentioned prior art has the following disadvantages and shortcomings: (1) The existing methods focus on the fusion of multimodal features in the encoder, and the processing of hierarchical features is relatively simple; (2) The decoder stage mostly adopts a progressive upsampling method, which simply aggregates the fused multimodal features, which easily leads to the loss of useful information. Summary of the invention
[0005] The purpose of the present invention is to provide a salient object detection method for RGB-D images based on progressive hierarchical fusion, so as to fully exploit the cross-modal and cross-level information of RGB and depth images and improve the detection accuracy.
[0006] To achieve the purpose of the present invention, the technical solution provided by the present invention is specifically as follows:
[0007] A method for detecting salient objects in RGB-D images based on progressive hierarchical fusion includes the following steps:
[0008] Step 1: Dataset preprocessing;
[0009] Step 2: Use Swin Transformer as the backbone network to extract the features of the four-layer RGB image and depth image respectively. The RGB image encoder provides the RGB feature information required by the network, and the depth image encoder provides the depth feature information required by the network.
[0010] Step 3: Use the Multi-Scale Enhancement Module (MSEM) in the last layer of the backbone network to perform multi-scale enhancement on the high-level RGB and depth semantic information;
[0011] Step 4: The first three layers of features extracted by the backbone network and the fourth layer of features enhanced by the multi-scale enhancement module are respectively input into the Multi-Modal Fusion Module (MMFM) to fuse the RGB and depth features of each layer;
[0012] Step 5: In the decoder stage, a multi-layer decoding structure (MLDS) is used to decode the RGB-D features in a progressive hierarchical fusion manner, and a dual-layer fusion module (DLFM) is used to fuse the hierarchical features.
[0013] Step 6: After the decoder generates the saliency map, an edge enhancement module (EEM) is used to increase the edge details of the saliency map.
[0014] Furthermore, the step 1 specifically includes the following:
[0015] Step 1.1: Obtain a data set, which is divided into a test set and a training set. The test set uses NJU2K, NLPR, SIP, LFSD and STERE;
[0016] NJU2K is a large dataset, which contains 1985 images, of which the stereo images are provided by the Internet and 3D movies, and the photos are taken by Fuji W3 camera; NLPR contains 1000 images, captured by standard Microsoft Kinect, including indoor and outdoor scenes; SIP is taken by Huawei Mate10, and the depth map is automatically estimated, with a total of 1000 images containing one or more prominent people; LFSD is a 100 light field dataset collected using Lytro light field camera, including 60 indoor scenes and 40 outdoor scenes; STERE consists of 1000 stereo images and is the first stereo image dataset in the field of saliency detection; the training set uses 1485 images from NJU2K and 700 images from NLPR, a total of 2185 images;
[0017] The training set uses 1485 images from NJU2K and 700 images from NLPR, totaling 2185 images;
[0018] Step 1.2: In order to improve the generalization ability of the model, the acquired data set is preprocessed, and random flipping, random cropping, random rotation and standard preprocessing operations are used to enhance the image data, and the image size is uniformly adjusted to 384×384.
[0019] Furthermore, the step 2 specifically includes the following steps: Swin Transformer is initialized using pre-trained network parameters to improve network efficiency; the RGB and depth images pre-trained in the first step are input into the network model to obtain four layers of features, where the RGB feature is recorded as The deep feature is recorded as i represents the level of the backbone network.
[0020] Furthermore, the step 3 specifically includes the following:
[0021] Step 3.1: For the high-level features extracted in step 2 and Feed it into the multi-scale enhancement module to obtain rich contextual information;
[0022] for First, in order to reduce the amount of calculation and memory consumption, reduce the input features The number of channels is obtained, and the output f MSEM1 , the specific formula is as follows:
[0023]
[0024] Where: Conv 1×1 (·) indicates the use of 1×1 convolution operation;
[0025] Step 3.2: Next, dilated convolutions with different dilation rates are used to expand the receptive field of the input features, and the Convolutional Block Attention Module (CBAM) is used to enhance the modal features. The specific formula is as follows:
[0026]
[0027] Where: CBAM(·) represents the convolutional attention operation, represents a dilated convolution operation with a dilation rate of 6. represents a dilated convolution operation with a dilation rate of 4. represents a dilated convolution operation with a dilation rate of 2, + represents an addition operation, and formula (2) is used to calculate the original input feature Feature enhancement is performed. In order to increase the correlation between different receptive fields, the result of formula (2) is MSEM2 Added as input to formula (3), the result of formula (3) is f MSEM3 Add as input to formula (4);
[0028] Step 3.3: In order to capture long-range dependencies, in addition to using global average pooling to obtain global features, strip pooling is also used. Strip pooling deploys two rectangular convolution kernels in the horizontal and vertical directions in the spatial dimension, respectively. While capturing long-range dependencies, it also pays attention to local details. The specific formula is as follows:
[0029]
[0030] Where: Strip(·) is the strip pooling operation, GAP(·) is the global average pooling operation, and f MSEM5 express The result of the strip pooling operation, f MSEM6 express The result of the average pooling operation of the entire play;
[0031] Step 3.4: Concatenate all the output results of the previous three steps to obtain enhanced high-level features The formula is as follows:
[0032]
[0033] Among them: Concat(·) is a cascade operation, It is the result of multi-scale enhancement of RGB high-level features, deep high-level features The same principle applies to the multi-scale enhancement module, and the enhanced depth features are finally obtained.
[0034] Furthermore, the step 4 specifically includes the following:
[0035] Step 4.1: Feed the input features into the Feature Enhancement Module (FEM) to enable RGB and depth features to learn cross-modal complementary features;
[0036] For the RGB input features, the channel attention operation is first used to obtain the spatial attention map of the deep feature. Then the original RGB feature is multiplied by the deep spatial attention map to supplement the spatial information missing from the RGB feature. Finally, the channel attention operation is used to calculate the feature correlation in the channel dimension, thereby assigning attention weights to different channels and multiplying them with the original RGB feature to obtain the enhanced RGB feature. The specific formula is as follows:
[0037]
[0038] Among them: SA(·) is the spatial attention operation, CA(·) is the channel attention operation, and × represents the element-wise multiplication operation; at the same time, when inputting the depth feature, the principle is the same, using the RGB detail information to supplement the depth feature, and finally obtaining the enhanced depth feature
[0039] Step 4.2: Enhanced RGB features and deep features Feed to the Global Feature Guidance Module (GFGM);
[0040] First, the two features are added together to obtain the fusion feature, and the global average pooling is used to capture the global spatial information. Then, the global spatial weight is calculated using the Sigmoid activation function, and the weight is used to compare with the original RGB feature. and deep features Multiply to reduce background noise. The specific formula is as follows:
[0041]
[0042] Where: Sig(·) represents the Sigmoid activation function, Represents RGB features After the results of the global feature guidance module, Representing deep features The result of the module guided by global features;
[0043] Step 4.3: and The feature is input into the Feature Fusion Module (FFM), and cross addition and multiplication operations are used. The cascade method is used to retain all the information of the feature map, and finally the fused feature is obtained. The specific formula is as follows:
[0044]
[0045] Furthermore, the step 5 specifically includes the following:
[0046] Step 5.1: Input the four-layer features obtained by the encoder into the decoder. Compared with the cascade fusion features, a two-layer fusion module is designed to fuse the adjacent layer features. For the high-level feature f rgbd_h and low-level features f rgbd_l ;
[0047] First, the high-level features f rgbd_h Upsample it and make it consistent with the low-level feature f rgbd_l The scale is the same, and then the channel attention mechanism is used to enhance the prominent area of the feature to avoid the interference of useless information. The two layers of features are fused by element-level addition to obtain f rgbd_ca Finally, the feature is used to supplement the original feature, and the hierarchical features are fused by element-level addition to obtain f DLFM ;
[0048] f rgbd_ca =CA(UP2(f rgbd_h ))+CA(f rgbd_l ) (11)
[0049] f DLFM =f rgbd_ca ×UP2(f rgbd_h )+f rgbd_ca ×f rgbd_l (12)
[0050] Where: UP2(·) represents the upsampling operation;
[0051] Step 5.2: Based on the double-layer fusion module of step 5.1, high-level fusion features First and Input into the double-layer fusion module to obtain the fused hierarchical features And use the residual to fuse the original high-level features and Add to reduce information loss, and then add the result of the addition Input into the second two-layer fusion module to obtain the second fused level feature Finally and As the original information Add and input into the third two-layer fusion module to get the first predicted saliency map The specific formula is as follows:
[0052]
[0053] Where: DLFM(·) means the use of double-layer fusion module;
[0054] Step 5.3: Preliminary comparison of the In order to increase the number of hierarchical interactions and refine the modal features, the three output results in the previous step are further fused using a double-layer fusion module. And the original input features Add as one input of the first two-layer fusion module in this step, and the other input is and the original input features The result of adding these two parts is sent to the double-layer fusion module to obtain the first fused hierarchical feature of this step Similarly, and The original input result is fed into the second double-layer fusion module of this step to obtain the second predicted saliency map of this step
[0055]
[0056] Step 5.4: Finally, the two output results of step 5.3 and their original input are fed into the two-layer fusion module to obtain the final predicted saliency map All original input and fused hierarchical feature information are preserved. The specific formula is as follows:
[0057]
[0058] Furthermore, the step 6 specifically includes the following:
[0059] Step 6.1: The low-level features contain rich details and contour information, so the low-level features of the deep encoder are and As the input of the edge enhancement module;
[0060] First, Upsample to make it consistent with the low-level features have the same scale and will and after upsampling The convolution blocks are cascaded together to organize the hierarchical features. Finally, the channel attention operation is used to capture the cross-channel information, and then the clear boundary features f are generated through element-wise multiplication and residual connection. edge , the specific formula is as follows:
[0061]
[0062] f edge =CA(f c )×f c +f c (20)
[0063] Where: Conv 3×3 (·) is a convolutional layer with a convolution kernel of size 3×3;
[0064] Step 6.2: Compare the three predicted saliency maps obtained in step 5 with the boundary features f in step 6.1 edge Cascading improves the edge quality of the final predicted saliency map. The specific formula is as follows:
[0065]
[0066] Where: UP4(·) is a 4-fold upsampling operation, S 1 , S 2 and S are the final predicted saliency maps, and GT is used to supervise them, and S is selected as the final predicted saliency map.
[0067] Furthermore, during the training process, the input image size is 384×384×3, the batch size is 4, the number of epochs is 200, the initial learning rate is 0.00005, and the Adam optimizer is used to optimize the network. The model finally uses the binary cross entropy loss function. The specific formula is as follows:
[0068] BCE(p,y)=-[y·log(p)+(1-y)·log(1-p)] (24)
[0069] Where: p represents the predicted value, y is the true value of the sample, and log(·) represents the logarithmic calculation.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] The present invention provides a new fusion decoding method for features at different levels, and proposes a network based on progressive hierarchical fusion. In addition, the present invention designs four modules in the network design (multi-scale enhancement module, multi-modal fusion module, edge enhancement module and multi-layer decoding structure) to fuse and enhance multi-modal features and multi-level features at different stages, thereby achieving accurate prediction of saliency maps. The method of the present invention has achieved competitive performance on multiple public benchmark datasets. The details are as follows:
[0072] (1) Compared with CN119131415A, the method of the present invention adopts a progressive hierarchical fusion method to fuse high-level features and low-level features. The decoder stage is more detailed about the interaction between features of different layers. The final saliency map contains more detailed information, reducing the loss of useful information.
[0073] (2) Compared with CN117953236A, the method of the present invention adopts an attention mechanism to enhance multimodal features, so that RGB features and depth features use complementary information, reduce information redundancy, and also adopts global features to reduce the impact of background noise.
[0074] (3) Compared with CN118587449A, the method of the present invention adopts a convolutional attention mechanism module to enhance the representation ability of features of different receptive fields, uses strip pooling to focus on local details, and extracts low-level deep features to generate sharp edge information to improve detection accuracy.
[0075] (4) The present invention proposes an end-to-end progressive hierarchical fusion network, which interacts and corrects multimodal information through an attention mechanism, fuses hierarchical features in a progressive hierarchical interaction manner, and effectively reduces the loss of useful information for salient object detection in RGB-D images.
[0076] (5) The present invention designs a multi-scale enhancement module to capture the contextual information of high-level features. The multi-scale enhancement module makes full use of cross-modal complementary information. The edge enhancement module produces sharp edge details. The multi-layer decoding structure uses a progressive hierarchical fusion method to fully integrate high-level features and low-level features to generate a complete and clear saliency map. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 A schematic diagram of a method flow chart provided in an embodiment of the present application;
[0078] Figure 2 A schematic diagram of the overall architecture of a progressive hierarchical fusion network provided in an embodiment of the present application;
[0079] Figure 3 A schematic diagram of a visualization example of salient object detection using RGB-D images in this application. DETAILED DESCRIPTION
[0080] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0081] Figure 1 A schematic diagram of a method flow chart provided in an embodiment of the present application;
[0082] Figure 2The schematic diagram of the overall architecture of the progressive hierarchical fusion network provided in the embodiment of the present application follows the encoder-decoder architecture. The framework mainly consists of four parts: 1) Multi-scale enhancement module. It increases the feature receptive field through hole convolution to capture long-distance dependencies. 2) Multimodal fusion module. The attention mechanism is used to achieve information complementarity of RGB and depth features, and global average pooling is used to reduce background noise. 3) Edge enhancement module. The edge contours of the low-level depth information are extracted to provide edge details for the final prediction image. 4) Multi-layer decoding structure. The fusion features of the encoder are fused in a progressive hierarchical manner to reduce the loss of useful information.
[0083] Figure 3 This is a schematic diagram of a visualization example of salient object detection using RGB-D images of the present application. The first column is the RGB image, the second column is the depth image, the third column is the true value map of RGB-D salient object detection, and the last column is the predicted salient map of the present invention. From the results, it can be seen that the present invention achieves good results in many challenging scenes. Under conditions such as poor depth maps, complex scenes, low contrast, and fine-grained objects, the present invention can clearly and accurately predict salient objects.
[0084] The present application provides a method for detecting salient objects in RGB-D images based on progressive hierarchical fusion, comprising the following steps:
[0085] Step 1: Dataset preprocessing;
[0086] Step 2: Use Swin Transformer as the backbone network to extract the features of the four-layer RGB image and depth image respectively. The RGB image encoder provides the RGB feature information required by the network, and the depth image encoder provides the depth feature information required by the network.
[0087] Step 3: In the last layer of the backbone network, a multi-scale enhancement module MSEM is used to perform multi-scale enhancement on the high-level RGB and depth semantic information;
[0088] Step 4: The first three layers of features extracted by the backbone network and the fourth layer of features enhanced by the multi-scale enhancement module are respectively input into the multimodal fusion module MMFM to fuse the RGB and depth features of each layer;
[0089] Step 5: In the decoder stage, a multi-layer decoding structure MLDS is used to decode the RGB-D features in a progressive hierarchical fusion manner, and a double-layer fusion module DLFM is used to fuse the hierarchical features;
[0090] Step 6: After the decoder generates the saliency map, the edge enhancement module EEM is used to increase the edge details of the saliency map.
[0091] It should be noted that the network designed by the present invention adopts Swin Transformer as the backbone network, adopts an encoder-decoder architecture, and two encoders are used in the encoder stage to extract the hierarchical features of RGB and depth images respectively. The overall network consists of a multi-scale enhancement module, a multi-modal fusion module, an edge enhancement module and a multi-layer decoding structure. Specifically, the multi-scale enhancement module is used to perform multi-scale enhancement on the high-level features of RGB and depth, so that the features capture contextual information, and also pay attention to local details while capturing global features; the first three layers of the backbone network and the enhanced fourth layer of multi-modal information are respectively input into the multi-modal fusion module, and the attention mechanism is used to make the information of RGB and depth features complement each other and reduce the generation of redundant information; in the decoder stage, the RGB-D fusion features generated by the encoder are fused in a progressive hierarchical manner, and the fusion features are continuously input into the double-layer fusion module to integrate high-level and low-level features, so that the predicted saliency map has rich detail information; finally, the edge enhancement module extracts edge information from the low-level depth encoder to add details to the final predicted saliency map.
[0092] As the network layer deepens, the scale of the input image continues to decrease and the channels continue to increase. Although rich semantic information is captured, information will be lost due to the reduction in resolution. In order to solve this problem, the present invention designs a multi-scale enhancement module, which uses dilated convolutions with different dilation rates to avoid the loss of internal data structure information of the image caused by pooling, and uses global average pooling and strip pooling to capture global and local information.
[0093] Specifically, high-level features are extracted from the Swin Transformer backbone network and Feed it into the multi-scale enhancement module to obtain rich context information; First, in order to reduce the amount of calculation and memory consumption, reduce the input features The number of channels is obtained, and the output f MSEM1 , the specific formula is as follows:
[0094]
[0095] Where: Conv 1×1 (·) indicates the use of 1×1 convolution operation;
[0096] Next, we use dilated convolutions with different dilation rates to expand the receptive field of the input features, and use the convolutional attention module (CBAM) to enhance the modal features. The specific formula is as follows:
[0097]
[0098] Where: CBAM(·) represents the convolutional attention operation, represents a dilated convolution operation with a dilation rate of 6. represents a dilated convolution operation with a dilation rate of 4. represents a dilated convolution operation with a dilation rate of 2, + represents an addition operation, and formula (2) is used to calculate the original input feature Feature enhancement is performed. In order to increase the correlation between different receptive fields, the present invention converts the result of formula (2) into MSEM2 Added as input to formula (3), the result of formula (3) is f MSEM3 Add as input to formula (4);
[0099] In order to capture long-range dependencies, in addition to using global average pooling to obtain global features, the present invention also uses strip pooling. Strip pooling deploys two rectangular convolution kernels in the horizontal and vertical directions in the spatial dimension, respectively. While capturing long-range dependencies, it also pays attention to local details. The specific formula is as follows:
[0100]
[0101] Where: Strip(·) is the strip pooling operation, GAP(·) is the global average pooling operation, and f MSEM5 express The result of the strip pooling operation, f MSEM6 express The result of the average pooling operation of the entire play;
[0102] Finally, all the above output results are cascaded together to obtain enhanced high-level features The formula is as follows:
[0103]
[0104] Among them: Concat(·) is a cascade operation, It is the result of multi-scale enhancement of RGB high-level features, deep high-level features The same principle applies to the multi-scale enhancement module, and the enhanced depth features are finally obtained.
[0105] RGB images are composed of the superposition of three channels of red, green and blue, and contain rich color and texture information. Depth images are single-channel grayscale images. Each pixel represents its distance from the camera and provides spatial information. Reasonable fusion of multimodal information can increase the robustness of the network. However, two problems often arise in the fusion process. Complex RGB images have a lot of interference information, which is prone to redundancy during fusion, and low-quality depth maps will mislead the saliency map. Therefore, in response to these two problems, the present invention proposes a multimodal fusion module, which consists of a feature enhancement module, a global feature guidance module and a feature fusion module. The feature enhancement module adopts a complementary attention mechanism to enable RGB and depth features to learn from each other in the process of enhancing modal features. The global feature guidance module adopts global average pooling to reduce background noise. Finally, the feature fusion module effectively fuses multimodal information.
[0106] Specifically, the backbone network-level feature input features are fed into the feature enhancement module to enable RGB and deep features to learn cross-modal complementary features. For the RGB input features, the channel attention operation is first used to obtain the spatial attention map of the deep features, and then the original RGB features are multiplied by the deep spatial attention map to supplement the spatial information missing from the RGB features. Finally, the channel attention operation is used to calculate the feature correlation in the channel dimension, thereby assigning attention weights to different channels and multiplying them with the original RGB features to obtain the enhanced RGB features. The specific formula is as follows:
[0107]
[0108] Among them: SA(·) is the spatial attention operation, CA(·) is the channel attention operation, and × represents the element-wise multiplication operation; at the same time, when inputting the depth feature, the principle is the same, using the RGB detail information to supplement the depth feature, and finally obtaining the enhanced depth feature
[0109] Next, the enhanced RGB features and deep features The global feature guidance module first adds the two features to obtain the fused feature, uses global average pooling to capture the global spatial information, and then uses the Sigmoid activation function to calculate the global spatial weight, which is used to compare the original RGB feature and deep features Multiply to reduce background noise. The specific formula is as follows:
[0110]
[0111] Where: Sig(·) represents the Sigmoid activation function, Represents RGB features After the results of the global feature guidance module, Representing deep features The result of the module guided by global features;
[0112] Finally and Input into the feature fusion module, use cross addition and multiplication operations, and use the cascade method to retain all the information of the feature map, and finally get the fused feature The specific formula is as follows:
[0113]
[0114] High-level features are rich in semantic information and can help people quickly identify the main features of objects, while low-level features contain detailed information such as contours, edges, and colors. Indiscriminate fusion of features at different levels can easily lead to the loss of useful information. Most existing methods use a simple way to cascade hierarchical features, ignoring the impact of different hierarchical features on saliency maps. Therefore, the present invention designs a multi-layer decoding structure. Compared with the simple fusion of hierarchical features, the multi-layer decoder uses a double-layer fusion module to integrate hierarchical features, and uses residual connections to reduce information loss, and fuses high-level features and low-level features layer by layer in a progressive hierarchical fusion manner.
[0115] Specifically, the four-layer features obtained by the encoder are input into the decoder. Compared with the cascade fusion feature, the present invention designs a double-layer fusion module to fuse the adjacent layer features. rgbd_h and low-level features f rgbd_l First, the high-level features f rgbd_h Upsample it and make it consistent with the low-level feature f rgbd_l The scale is the same, and then the channel attention mechanism is used to enhance the prominent area of the feature to avoid the interference of useless information. The two layers of features are fused by element-level addition to obtain f rgbd_ca Finally, the feature is used to supplement the original feature, and the hierarchical features are fused by element-level addition to obtain f DLFM ;
[0116] f rgbd_ca =CA(UP2(f rgbd_h ))+CA(f rgbd_l ) (11)
[0117] f DLFM =f rgbd_ca ×UP2(f rgbd_h )+f rgbd_ca ×f rgbd_l (12)
[0118] Where: UP2(·) represents the upsampling operation;
[0119] Based on the double-layer fusion module, high-level fusion features First and Input into the double-layer fusion module to obtain the fused hierarchical features And use the residual to fuse the original high-level features and Add to reduce information loss, and then add the result of the addition Input into the second two-layer fusion module to obtain the second fused level feature Finally and As the original information Add and input into the third two-layer fusion module to get the first predicted saliency map The specific formula is as follows:
[0120]
[0121]
[0122] Where: DLFM(·) means the use of double-layer fusion module;
[0123] The above formula is preliminary In order to increase the number of hierarchical interactions and refine the modal features, the present invention continues to use a double-layer fusion module for the three output results in the previous step. And the original input features Add as one input of the first two-layer fusion module in this step, and the other input is and the original input features The result of adding these two parts is sent to the double-layer fusion module to obtain the first fused hierarchical feature of this step Similarly, and The original input result is fed into the second double-layer fusion module of this step to obtain the second predicted saliency map of this step
[0124]
[0125] Finally, the two output results of formula (16) and formula (17) and their original input are fed into the double-layer fusion module to obtain the final predicted saliency map All original input and fused hierarchical feature information can be preserved. The specific formula is as follows:
[0126]
[0127] Compared with high-level features, low-level features have detailed information, and their edges and contours can directly reflect the image content. Compared with complex RGB images, depth images can extract salient contours of depth images. Therefore, the present invention designs an edge enhancement module, which extracts low-level depth features, pays attention to the edge contours of the salient map, and increases the edge details of the final predicted image.
[0128] Specifically, low-level features contain rich details and contour information, so the low-level features of the deep encoder and As the input of the edge enhancement module, we first Upsample to make it consistent with the low-level features have the same scale and will and after upsampling The convolution blocks are cascaded together to organize the hierarchical features. Finally, the channel attention operation is used to capture the cross-channel information, and then the clear boundary features f are generated through element-wise multiplication and residual connection. edge , the specific formula is as follows:
[0129]
[0130] Where: Conv 3×3 (·) is a convolutional layer with a convolution kernel of size 3×3;
[0131] Finally, the three predicted saliency maps obtained by the multi-layer decoding structure are respectively compared with the boundary features f edge Cascading improves the edge quality of the final predicted saliency map. The specific formula is as follows:
[0132]
[0133]
[0134]
[0135] Where: UP4(·) is a 4-fold upsampling operation, S 1 , S 2 and S are the final predicted saliency maps, and GT is used to supervise them, and S is selected as the final predicted saliency map.
[0136] The above describes in detail the optional implementation modes of the embodiments of the present invention in combination with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above implementation modes. Within the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical scheme of the embodiments of the present invention, and these simple modifications all belong to the protection scope of the embodiments of the present invention.
Claims
1. A method for detecting salient objects in RGB-D images based on progressive hierarchical fusion, characterized in that: The steps include: Step 1: Dataset preprocessing; Step 2: Use Swin Transformer as the backbone network to extract the features of the four-layer RGB image and depth image respectively. The RGB image encoder provides the RGB feature information required by the network, and the depth image encoder provides the depth feature information required by the network. Step 3: In the last layer of the backbone network, a multi-scale enhancement module MSEM is used to perform multi-scale enhancement on the high-level RGB and depth semantic information; Step 4: The first three layers of features extracted by the backbone network and the fourth layer of features enhanced by the multi-scale enhancement module are respectively input into the multimodal fusion module MMFM to fuse the RGB and depth features of each layer; Step 5: In the decoder stage, a multi-layer decoding structure MLDS is used to decode the RGB-D features in a progressive hierarchical fusion manner, and a double-layer fusion module DLFM is used to fuse the hierarchical features; Step 6: After the decoder generates the saliency map, the edge enhancement module EEM is used to increase the edge details of the saliency map.
2. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 1, characterized in that: The step 1 specifically includes the following: Step 1.1: Obtain a dataset, which is divided into a test set and a training set. The test set uses NJU2K, NLPR, SIP, LFSD, and STERE; the training set uses 1485 images from NJU2K and 700 images from NLPR, for a total of 2185 images; Step 1.2: Preprocess the acquired data set, perform data augmentation on the images using random flipping, random cropping, random rotation, and standard preprocessing operations, and resize the images to 384×384.
3. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 2, characterized in that: The step 2 specifically includes the following steps: Swin Transformer is initialized using pre-trained network parameters; the RGB and depth images pre-trained in the first step are input into the network model to obtain four layers of features, where the RGB feature is recorded as The deep feature is recorded as i represents the level of the backbone network.
4. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 3, characterized in that: The step 3 specifically includes the following: Step 3.1: For the high-level features extracted in step 2 and Feed it into the multi-scale enhancement module to obtain rich contextual information; for Reduce input features The number of channels is obtained, and the output f MSEM1 , the specific formula is as follows: Where: Conv 1×1 (·) indicates the use of 1×1 convolution operation; Step 3.2: Next, dilated convolutions with different dilation rates are used to expand the receptive field of the input features, and the convolutional attention module CBAM is used to enhance the modal features. The specific formula is as follows: Where: CBAM(·) represents the convolutional attention operation, represents a dilated convolution operation with a dilation rate of 6. represents a dilated convolution operation with a dilation rate of 4. represents a dilated convolution operation with a dilation rate of 2, + represents an addition operation, and formula (2) is used to calculate the original input feature The feature enhancement is performed and the result f of formula (2) is MSEM2 Added as input to formula (3), the result of formula (3) is f MSEM3 Add as input to formula (4); Step 3.3: In addition to using global average pooling to obtain global features, strip pooling is also used. Strip pooling deploys two rectangular convolution kernels in the horizontal and vertical directions in the spatial dimension, respectively. While capturing long-range dependencies, it also pays attention to local details. The specific formula is as follows: Where: Strip(·) is the strip pooling operation, GAP(·) is the global average pooling operation, and f MSEM5 express The result of the strip pooling operation, f MSEM6 express The result of the average pooling operation of the entire play; Step 3.4: Concatenate all the output results of the previous three steps to obtain enhanced high-level features The formula is as follows: Among them: Concat(·) is a cascade operation, It is the result of multi-scale enhancement of RGB high-level features, deep high-level features The same principle applies to the multi-scale enhancement module, and the enhanced depth features are finally obtained.
5. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 4, characterized in that: The step 4 specifically includes the following: Step 4.1: Feed the input features into the feature enhancement module FEM to enable RGB and depth features to learn cross-modal complementary features; For the RGB input features, the channel attention operation is first used to obtain the spatial attention map of the deep feature. Then the original RGB feature is multiplied by the deep spatial attention map to supplement the spatial information missing from the RGB feature. Finally, the channel attention operation is used to calculate the feature correlation in the channel dimension, thereby assigning attention weights to different channels and multiplying them with the original RGB feature to obtain the enhanced RGB feature. The specific formula is as follows: Among them: SA(·) is the spatial attention operation, CA(·) is the channel attention operation, and × represents the element-wise multiplication operation; at the same time, when inputting the depth feature, the principle is the same, using the RGB detail information to supplement the depth feature, and finally obtaining the enhanced depth feature Step 4.2: Enhanced RGB features and deep features Feed to the global feature guidance module GFGM; First, the two features are added together to obtain the fusion feature, and the global average pooling is used to capture the global spatial information. Then, the global spatial weight is calculated using the Sigmoid activation function, and the weight is used to compare with the original RGB feature. and deep features Multiply to reduce background noise. The specific formula is as follows: Where: Sig(·) represents the Sigmoid activation function, Represents RGB features After the results of the global feature guidance module, Representing deep features The result of the module guided by global features; Step 4.3: and The input is sent to the feature fusion module FFM, and cross addition and multiplication operations are used. The cascade method is used to retain all the information of the feature map, and finally the fused feature is obtained. The specific formula is as follows:
6. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 5, characterized in that: The step 5 specifically includes the following: Step 5.1: Input the four-layer features obtained by the encoder into the decoder. Compared with the cascade fusion features, a two-layer fusion module is designed to fuse the adjacent layer features. For the high-level feature f rgbd_h and low-level features f rgbd_l ; First, the high-level features f rgbd_h Upsample it and make it consistent with the low-level feature f rgbd_l The scale is the same, and then the channel attention mechanism is used to enhance the prominent area of the feature, and the two layers of features are fused by element-level addition to obtain f rgbd_ca Finally, the feature is used to supplement the original feature, and the hierarchical features are fused by element-level addition to obtain f DLFM ; f rgbd_ca =CA ( UP 2( f rgbd _h ))+ CA ( f rgbd_l )(11) f DLFM =f rgbd_ca ×UP2(f rgbd_h )+f rgbd_ca ×f rgbd_l (12) Where: UP2(·) represents the upsampling operation; Step 5.2: Based on the double-layer fusion module of step 5.1, high-level fusion features First and Input into the double-layer fusion module to obtain the fused hierarchical features And use the residual to fuse the original high-level features and Add to reduce information loss, and then add the result of the addition Input into the second two-layer fusion module to obtain the second fused level feature Finally and As the original information Add and input into the third two-layer fusion module to get the first predicted saliency map The specific formula is as follows: Where: DLFM(·) means the use of double-layer fusion module; Step 5.3: Preliminary comparison of the After fusion, the three output results in the previous step continue to use the double-layer fusion module. And the original input features Add as one input of the first two-layer fusion module in this step, and the other input is and the original input features The result of adding these two parts is sent to the double-layer fusion module to obtain the first fused hierarchical feature of this step Similarly, and The original input result is fed into the second double-layer fusion module of this step to obtain the second predicted saliency map of this step Step 5.4: Finally, the two output results of step 5.3 and their original input are fed into the two-layer fusion module to obtain the final predicted saliency map All original input and fused hierarchical feature information are preserved. The specific formula is as follows:
7. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 6, characterized in that: The step 6 specifically includes the following: Step 6.1: The low-level features contain rich details and contour information, so the low-level features of the deep encoder are and As the input of edge enhancement module; First, Upsample to make it consistent with the low-level features have the same scale and will and after upsampling The convolution blocks are cascaded together to organize the hierarchical features. Finally, the channel attention operation is used to capture the cross-channel information, and then the clear boundary features f are generated through element-wise multiplication and residual connection. edge , the specific formula is as follows: f edge =CA(f c )×f c +f c (20) Where: Conv 3×3 (·) is a convolutional layer with a convolution kernel of size 3×3; Step 6.2: Compare the three predicted saliency maps obtained in step 5 with the boundary features f in step 6.1 edge Cascading improves the edge quality of the final predicted saliency map. The specific formula is as follows: Among them: UP4(·) is a 4-fold upsampling operation, S1, S2 and S are the final predicted saliency maps, and GT is used to supervise them, and S is selected as the final predicted saliency map.
8. The method for detecting salient objects in RGB-D images based on progressive hierarchical fusion according to claim 7, characterized in that: During the training process, the input image size is 384×384×3, the batch size is 4, the number of epochs is 200, the initial learning rate is 0.00005, and the Adam optimizer is used to optimize the network. The model finally uses the binary cross entropy loss function. The specific formula is as follows: BCE(p,y)=-[y·log(p)+(1-y)·log(1-p)](24) Where: p represents the predicted value, y is the true value of the sample, and log(·) represents the logarithmic calculation.
Citation Information
Patent Citations
RGB-D saliency target detection method based on cross-modal fusion network
CN117953236A
Novel RGB-D image saliency target detection method
CN118334365A
RGB-D saliency detection method based on progressive weighted decoding
CN118587449A
Multi-modal variation fusion method based on RGB-D salient target detection
CN119027771A
RGB-D collaborative saliency target detection method based on prompt learning
CN119131415A
Cited By
RGB-D saliency detection method based on progressive multi-scale feature fusion
CN119810423A
RGB-D salient target detection method based on potential perception and hierarchical fusion
CN120953577A