An Interactive Salience Mining Method for RGB-D Salient Object Detection
Through the interactive significance mining method, the multi-level cross-modal features of RGB and depth images are integrated using the dual-stream encoder network and cross-modal interaction module, the problem of difficulty in integrating multi-modal information and weak combination of context information in RGB-D significant object detection is solved, and more accurate and precise significant object detection effect is achieved.
Patent Information
- Application Number
- CN202310536823.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-05-12
AI Technical Summary
The existing RGB-D significant object detection technology has problems such as difficulty in integrating multimodal information, weak context information combination, and insufficient significant target boundaries.
The interactive significance mining method is used to extract multi-level cross-modal features of RGB and depth images through a dual-stream encoder network, and feature interaction and aggregation are achieved using a cross-modal interaction module. The significance mining module integrates multi-level features in a gradual manner, separates significant perceptual features and complex background features, and further extracts context information through the scene exploration module.
It effectively solves the problems of difficulty in integrating multimodal information and weak integration of context information, and the effect of significant target detection is significantly improved, and the significant figure is more accurate and detailed.
Smart Images

Figure CN116524208B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and image processing, and particularly to an interactive saliency mining method for RGB-D salient object detection. Background Art
[0002] The human visual system can quickly focus on the target or area of interest from a complex scene through unique perception units and neural structures, which is called the visual attention mechanism. Generally speaking, human visual attention includes two pathways: bottom-up and top-down. The bottom-up attention mechanism is data-driven and is the non-autonomous perception of external things by the human brain; while the top-down attention mechanism is consciousness-driven and is the process of the human brain actively issuing instructions to capture the target of interest. Inspired by this, people hope that computers also have this ability, so the salient object detection (SOD) task came into being, aiming to let the computer automatically detect the most attention-grabbing area or target from a given scene, and has been widely applied in a large number of computer vision tasks including image segmentation, video compression, and object detection. After more than a decade of development, the salient object detection task has spawned many branches, including salient object detection for RGB images, salient object detection for high-resolution RGB images, salient object detection for RGB-D images, collaborative salient object for image groups, salient object detection for RGB-T images, salient object detection for light field images, salient object detection for panoramic images, salient object detection for remote sensing images, and salient object detection for video sequences. What the present invention relates to is this branch of salient object detection for RGB-D images.
[0003] Although significant object detection based on RGB images has made great progress, in some low-contrast and complex background scenarios, satisfactory detection results cannot be achieved solely by relying on RGB images. In fact, the human eye can not only perceive appearance information such as color, shape, and texture in a scene, but also capture the depth information of the scene through the binocular vision system to form a sense of stereoscopy. Depth images are an intuitive manifestation of depth information, and their numerical values reflect the distance between objects in the scene and the depth sensor. By introducing depth images into the significant object detection task, on the one hand, it can simulate the human stereoscopic vision perception ability and enhance the computer's distance recognition ability between different objects; on the other hand, some excellent characteristics of depth images (such as internal consistency of objects, shape priors, etc.) also provide new ideas for solving the detection of some difficult scenarios. In recent years, the rise of depth cameras (such as Kinect, RealSense) has made the acquisition of depth images more and more convenient, which also lays a data foundation for the rise of the RGB-D significant object detection task. Since the initial exploration of the impact of depth information on visual saliency in 2012, significant object detection based on RGB-D images has received increasing attention.
[0004] Inspired by the human visual attention mechanism, the significant object detection task aims to locate the most attention-grabbing objects or regions in a given scene. In recent years, with the development and popularization of depth cameras, depth images have been successfully applied to various computer vision tasks, which also provides new ideas for significant object detection technology. By introducing depth images, not only can the computer more comprehensively simulate the human visual system, but also the supplementary information such as structure and position provided by depth images can provide new solutions for the detection of difficult scenarios such as low contrast and complex backgrounds.
[0005] However, there are still problems in the field of significant object detection, such as difficult integration of multi-modal information, weak combination of context information, and unclear boundaries of significant objects.
[0006] Therefore, the present invention proposes an interactive saliency mining method for RGB-D significant object detection to solve the above problems. Summary of the Invention
[0007] (1) Technical problems to be solved
[0008] Aiming at the deficiencies of the prior art, the present invention provides an interactive saliency mining method for RGB-D significant object detection, which solves the problems of difficult integration of multi-modal information, weak combination of context information, and unclear boundaries of significant objects in the current field of significant object detection.
[0009] (2) Technical solutions
[0010] In the encoding stage of the present invention, multi-level cross-modal fusion features are obtained. The progressive saliency mining module is used to decode the multi-level features, gradually filtering out complex background interference to make the saliency map more accurate, and gradually refining the saliency region to make the saliency map more refined, achieving an optimal decoding process. The present invention first uses a two-stream encoder network to extract multi-level cross-modal features of RGB and depth images; then proposes a cross-modal interaction module to achieve the interaction and aggregation of these different features through a series of matrix operations; secondly, attempts to separate the saliency perception information from the complex environment and establish a fusion method between adjacent features under the guidance of a relatively rough saliency map; finally, further extracts context information from the saliency perception information and background information.
[0011] To achieve the above object, the present invention specifically adopts the following technical solutions:
[0012] An interactive saliency mining method for RGB-D salient object detection, comprising the following steps:
[0013] S1. Analyze the advantages and existing problems of the existing saliency detection algorithms based on RGB-D images, and build a saliency object detection network and each module;
[0014] S2. Collect and organize the existing RGB-D image datasets available for saliency object detection, extract multi-level RGB image features and Depth image features respectively, and cross-modally fuse the RBG features and Depth features to form multi-level cross-modal fusion features with multi-resolution;
[0015] S3. The saliency mining module integrates the multi-level fusion features in a progressive manner, separates the saliency perception features and complex background features, and gradually outputs the rough saliency maps of each layer and the final saliency map;
[0016] S4. The scene exploration module further extracts the context information of the saliency perception features and complex background features;
[0017] S5. Conduct multiple experiments, optimize different modules, network structures and parameters, and record the experimental results of each time;
[0018] S6. Complete the evaluation of the quantitative indicators by calculating four evaluation indicators: mean absolute error, mean F-measure, E-measure and S-measure, evaluate whether the constructed network is effective, and verify whether each proposed module is effective;
[0019] Among them, the mean absolute error is used to measure the mean of the absolute error between the saliency map and the ground truth map pixel by pixel, and the calculation method is as follows:
[0020] The mean F-measure is used to calculate the harmonic mean of average precision and recall, and the calculation method is as follows:
[0021] And β 2 takes a value of 0.3;
[0022] The E-measure is used to measure the statistical information at the image level and its local pixel matching information, and the calculation method is as follows:
[0023] The S-measure is used to compare the structural similarity information, where s o is the object structural similarity, s r is the regional structural similarity, α is the balance parameter, taking a value of 0.5, and the calculation method is as follows: S α = α * s o +(1 - α) * s r .
[0024] Furthermore, in the step S1, the constructed saliency object detection network is used to detect the most attention-grabbing region in the image.
[0025] Furthermore, in the step S2, the method for extracting multi-level RGB image features and Depth image features is to use two PVT backbone networks with shared weights pre-trained on ImageNet to extract RGB features and Depth features respectively, forming and i = 1, … 4, where i represents the layer number, representing the output of each layer of PVT, and the four-layer features obtained have different resolutions and numbers of channels;
[0026] The method for fusing the RBG features and Depth features is implemented using a cross-modal interaction module. The cross-modal interaction module introduces matrix operations to establish the correlation between RGB features and depth features; first, the RGB feature F RGB is reshaped and transposed, the Depth feature F Depth is reshaped, and then the two tensor matrices are multiplied to obtain the correlation matrix S and the difference perception information matrix S D ; after reshaping F RGB and F Depth and multiplying them with S D through matrix multiplication and then splicing, the final fused feature F fuse is obtained.
[0027] Furthermore, in the step S3, the saliency mining module progressively fuses multi-level features, and in the process of progressive fusion, the adjacent higher-level features E i and the rough saliency map Si to mine the saliency information T under the guidance of S and separate it from the complex background information T B ; Gradually output the rough saliency map and the final saliency map from the fourth layer to the first layer.
[0028] Furthermore, the progressive fusion method refers to: using the saliency feature U of the i-th layer i (i = 2, 3, 4) and the rough saliency map S i to guide the fusion feature E of the i-1-th layer i-1 , and generate U i-1 and Si -1 through the saliency mining module. By analogy, the saliency mining module of the first layer generates the final saliency map S 1 , which is used as the prediction result of this network;
[0029] The saliency information T S and the complex background information T B are realized through a series of upsampling, sigmoid function, addition and multiplication:
[0030] P i = Sigmoid(Upsample(S i ))
[0031] T = E i-1 + Upsample(U i )
[0032] T S = T × P i ,
[0033] T B = T × (1 - P i )
[0034] After highlighting the saliency perception information and extracting the background perception features, use the scene exploration module to further extract the context information of the saliency features and background features to improve the segmentation accuracy; perform a weighted subtraction operation on the saliency features and background features to obtain the fused saliency feature U:
[0035] U = α × SEB(T s ) - β × SEB(T B ),
[0036] where SEB represents the operation of the scene exploration modality. After U passes through a simple convolution, the saliency map output by the saliency mining module of this layer can be obtained.
[0037] Furthermore, the scenario exploration module consists of four stages. First, four parallel 1×1 convolutions are used to reduce the channels of the input features, and both ordinary convolutions and dilated convolutions are used in the four stages to further extract features.
[0038] (III) Beneficial effects
[0039] Compared with the prior art, the present invention provides an interactive saliency mining method for RGB-D salient object detection, having the following beneficial effects:
[0040] In the present invention, the complementary and integration of cross-modal information of RGB images and depth images are realized by constructing a dual-branch encoder and a cross-modal interaction module, and the progressive fusion of multi-level cross-modal features and the further extraction of context information are completed by designing a saliency mining module and a scenario exploration module; the present invention solves the problems of difficult integration of multi-modal information, insufficient utilization of multi-level information with different resolutions, and weak combination of context information in the field of salient object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a flowchart of the interactive saliency mining algorithm for RGB-D salient object detection according to the present invention;
[0042] Figure 2 is an overall network structure diagram of the interactive saliency mining algorithm for RGB-D salient object detection according to the present invention;
[0043] Figure 3 is a structure diagram of the cross-modal interaction module (CMIM) designed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment
[0046] As Figures 1-3 shown, an interactive saliency mining method for RGB-D salient object detection proposed in an embodiment of the present invention includes the following steps:
[0047] S1. Collect and organize existing RGB-D image datasets available for salient object detection. A total of 4 datasets available for RGB-D salient object detection are collected, namely the STERE dataset, the NJU2K dataset, the LFSD dataset, and the NLPR dataset. 700 images are selected from the NLPR dataset and 1485 images are selected from the NJU2K dataset to train the method in this paper. The remaining images in the NJU2K and NLPR datasets and the images in the entire STERE and LFSD datasets are used for testing.
[0048] S2. Extract multi-level RGB image features and depth image features respectively; and cross-modal fuse the RBG features and Depth features to form multi-resolution multi-level cross-modal fusion features.
[0049] S3. The saliency mining module integrates multi-level fusion features in a progressive manner, separates saliency-aware features and complex background features, and further extracts the context information of saliency features and background features with the scene exploration module, gradually outputting rough saliency maps of each layer and the final saliency map.
[0050] S4. Complete the evaluation of quantitative metrics by calculating four evaluation metrics: mean absolute error (MAE), mean F-measure, E-measure, and S-measure, to evaluate whether the constructed network is effective and verify whether each proposed module is effective.
[0051] Among them, the mean absolute error (MAE) is used to measure the mean of the absolute error between the saliency map and the ground truth map pixel by pixel, and the calculation method is as follows:
[0052] mean F-measure is used to calculate the harmonic mean of average precision and recall rate, and the calculation method is as follows:
[0053] And β 2 Generally takes a value of 0.3;
[0054] E-measure is used to measure the statistical information at the image level and its local pixel matching information, and the calculation method is as follows:
[0055] S-measure is used to compare structural similarity information, where s o is the object structural similarity, s r is the regional structural similarity, α is the balance parameter, taking a value of 0.5, and the calculation method is as follows: S α =α*s o +(1 - α)*s r .
[0056] Further, in step S1, the method for extracting RGB image features and Depth image features is to use two PVT backbone networks with shared weights pre-trained on ImageNet to extract RGB features and depth features respectively. The four-layer features obtained have different resolutions and numbers of channels, and are represented by and (where i = 1, …, 4, and i represents the layer number), representing the output of each layer of PVT.
[0057] Furthermore, the method for fusing RBG features and Depth features is implemented using a cross-modal interaction module to obtain multi-level cross-modal fusion features E i (i = 1, 2, 3, 4), where i represents the i-th layer. The cross-modal interaction module introduces matrix operations to establish the correlation between RGB features and depth features. First, reshape and transpose the RGB feature F RGB , reshape the Depth feature F Depth , and then perform matrix multiplication on the two tensor matrices to obtain the correlation matrix S and the difference perception information matrix S D . After reshaping F RGB and F Depth , perform matrix multiplication with S D and then concatenate to obtain the final fusion feature F fuse :
[0058] Q RGB = Reshape&Transpose(F RGB )
[0059] Q Depth = Reshape(F Depth )
[0060]
[0061] S D = 1 - Softmax(S)
[0062]
[0063] Among them, V RGB , V Depth (V RGB , V Depth ∈ R N×C ) are obtained by reshaping F RGB and F Depth .
[0064] In Figure 1 , the fusion feature F i at different levels is represented by E Fuse .
[0065] Further, the saliency mining module designed in S3 progressively fuses multi-level features, and utilizes the adjacent higher-level feature E i and the rough saliency map S i to guide the mining of saliency information T S , and separates it from the complex background information T B . The rough saliency map and the final saliency map are gradually output from the fourth layer to the first layer.
[0066] The progressive fusion method means that the saliency feature U i (i = 2, 3, 4) of the i-th layer and the rough saliency map S i are used to guide the fusion feature E i-1 of the (i - 1)-th layer. After passing through the saliency mining module, U i-1 and Si -1 are generated. By analogy, the saliency mining module of the first layer generates the final saliency map S 1 , which is used as the prediction result of this network.
[0067] The above saliency information T S and the complex background information T B are realized through a series of upsampling, sigmoid function, addition, and multiplication:
[0068] P i = Sigmoid(Upsample(S i ))
[0069] T = E i-1 + Upsample(U i )
[0070] T S = T × P i ,
[0071] T B = T × (1 - P i )
[0072] After highlighting the saliency perception information and extracting the background perception features, the scene exploration module is used to further extract the context information of the saliency features and background features to improve the segmentation accuracy.
[0073] Furthermore, the scene exploration module consists of four stages, which are specifically defined as follows:
[0074] G i = Conv 1×1 (G)
[0075]
[0076]
[0077]
[0078]
[0079] First, four parallel 1×1 convolutions are used to reduce the channels of the input feature G, which can solve the computational burden to a certain extent and obtain four features (G i , {i = 1, 2, 3, 4}). For the first stage, in this embodiment, a 1×1 convolution and a normal 3×3 convolution (dilation rate = 1) are combined to process G 1 , and the output feature O can be obtained from the first stage 1 . For the second stage, in this embodiment, element-wise addition is used to fuse G 2 and O 1 together, and then two 3×3 convolutions with dilation rates of 1 and 3 are used in sequence, and the output feature O of the second stage can be obtained therefrom 2 . For the third stage, in this embodiment, element-wise addition is still used to fuse G 3 and O 2 together, and then a 5×5 convolution and a 3×3 convolution with a dilation rate of 4 are introduced to process them in sequence, and the output feature O of the third stage can be obtained 3 . And so on, O 4 is obtained
[0080] Under the action of such a scene exploration module, the context information of saliency-aware information and background-aware features can be better explored. In addition, after using SEB to expand the context information of TS and TB, two learnable parameters α and β are introduced to multiply with TS and TB respectively, and the refined feature U can be obtained through subtraction operation:
[0081] U = α × SEB(T s ) - β × SEB(T B ),
[0082] SEB represents the operation of the scene exploration modality. After U passes through a simple convolution, the saliency map output by the saliency mining module on this layer can be obtained
[0083] SMM is used as an incremental refinement strategy to realize the fusion of adjacent features, and the saliencies S3, S2, and S1 are obtained in sequence. The present invention regards S1 as the final refined saliency map
[0084] The interactive saliency mining algorithm for RGB-D salient object detection in the present invention uses cross-modal interaction modalities to establish the correlation and difference of cross-modal features to complete complementarity; through the saliency mining module, the saliency information is mined and the complex background information is suppressed by using the adjacent higher-level saliency features and the guidance of the rough saliency map in a progressive fusion manner, and the receptive field is expanded and the segmentation accuracy is greatly improved by using the scene exploration module, realizing the optimal decoding process.
[0085] In another aspect of the present invention, the interactive saliency mining algorithm for RGB-D salient object detection in this embodiment selects 700 images and 1485 images from the NLPR and NJU2K data sets respectively to form a training set, and uses the remaining images of the NJU2K and NLPR data sets and the images on the entire STERE and LFSD data sets as test sets for testing.
[0086] In the training and testing phases, the input RGB-D images are resized to 352×352×3. The model uses the Adam optimizer with an initial learning rate of 10-4, and the learning rate decays by 10 times every 50 epochs. The batch size is set to 8. All experiments are implemented on an NVIDIA GeForce RTX 3080Ti.
[0087] The proposed method is compared with 15 RGB-D saliency object detection methods, namely AFNet[1], MMCI[2], TANet[3], CPFP[4], DMRA[5], D3Net[6], BiANet[7], UCNet[8], JL-DCF[9], BBSNet
[10] , BTS-Net
[11] , BPGNet
[12] , DCMF
[13] , SG-MWMG
[14] and SPSN
[15] . The results are shown in Table 1.
[0088] Table 1 Experimental results
[0089]
[0090] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An interactive saliency mining method for RGB-D salient object detection, characterized in that: It includes the following steps: S1. Analyze the advantages and existing problems of existing saliency object detection algorithms based on RGB-D images, and build a saliency object detection network and each module; S2. Collect and organize existing RGB-D image datasets available for saliency object detection, extract multi-level RGB image features and Depth image features respectively, and cross-modal fuse RBG features and Depth features to form multi-resolution multi-level cross-modal fusion features; S3. The saliency mining module integrates multi-level fusion features in a progressive manner, separates saliency perception features and complex background features, and gradually outputs rough saliency maps of each layer and the final saliency map; S4. The scene exploration module further extracts the context information of saliency perception features and complex background features; S5. Conduct multiple experiments, optimize different modules, network structures and parameters, and record the experimental results each time; S6. Complete the evaluation of quantitative indicators by calculating four evaluation indicators: mean absolute error, mean F-measure, E-measure and S-measure, evaluate whether the constructed network is effective, and verify whether each proposed module is effective; Among them, the mean absolute error is used to measure the mean of the absolute error between the saliency map and the ground truth map pixel by pixel, and the calculation method is as follows: The mean F-measure is used to calculate the harmonic mean of average precision and recall rate, and the calculation method is as follows: and β 2 takes a value of 0.3; The E-measure is used to measure the statistical information of the image hierarchy and its local pixel matching information, and the calculation method is as follows: The S-measure is used to compare structurally similar information, where s o is the structural similarity of an object, and s r is the structural similarity of a region. α is a balancing parameter with a value of 0.5, and the calculation method is as follows: S α = α * s o + (1 - α) * s r ; In the step S2, the method for extracting multi-level RGB image features and Depth image features is to use two PVT backbone networks with shared weights pre-trained on ImageNet to extract RGB features and Depth features respectively, forming and i = 1, …, 4, where i represents the layer number, representing the output of each layer of PVT. The four layers of features obtained have different resolutions and numbers of channels; The method for fusing the RBG feature and the Depth feature is implemented by a cross-modal interaction module, and the cross-modal interaction module introduces matrix operations to establish the correlation between the RGB feature and the depth feature; first, the RGB feature F RGB is reshaped and transposed, and the Depth feature F Depth is reshaped, and then the two tensor matrices are multiplied to obtain the correlation matrix S and the difference perception information matrix S D ; F RGB and F Depth after being reshaped by the matrix are multiplied by S D and then concatenated to obtain the final fused feature F fuse .
2. An interactive saliency mining method for RGB-D salient object detection according to claim 1, characterized in that: In step S1, the constructed saliency object detection network is used to detect the region in the image that attracts the most attention.
3. An interactive saliency mining method for RGB-D salient object detection according to claim 1, characterized in that: In the step S3, the saliency mining module progressively fuses multi-level features, and utilizes the adjacent higher-level feature E i and the rough saliency map S i to guide the mining of saliency information T S , and separates it from the complex background information T B ; gradually outputs the rough saliency map and the final saliency map from the fourth layer to the first layer.
4. An interactive saliency mining method for RGB-D salient object detection according to claim 3, characterized in that: The progressive fusion method refers to: using the saliency feature U of the i-th layer i (i = 2, 3, 4) and the rough saliency map S i to guide the fusion feature E of the (i - 1)-th layer i-1 , generating U i-1 and Si -1 through the saliency mining module, and so on. The saliency mining module of the first layer generates the final saliency map S 1 , which is used as the prediction result of this network; The significance information T S and the complex background information T B are implemented through a series of upsampling, sigmoid functions, addition, and multiplication: P i = Sigmoid(Upsample(S i )) T = E i-1 + Upsample(U i ) T S = T × P i , T B = T × (1 - P i ) After highlighting saliency perception information and extracting background perception features, use the scene exploration module to further extract the context information of saliency features and background features to improve the segmentation accuracy; perform a weighted subtraction operation on the saliency features and background features to obtain the fused saliency feature U: U = α × SEB(T s ) - β × SEB(T B ), where SEB represents the operation of the scene exploration modality. After U passes through a simple convolution, the saliency map output by the saliency mining module on this layer can be obtained.
5. An interactive saliency mining method for RGB-D salient object detection according to claim 4, characterized in that: The scene exploration module consists of four stages. First, use four parallel 1×1 convolutions to reduce the channels of the input features. Each of the four stages further extracts features using ordinary convolutions and dilated convolutions.
Citation Information
Patent Citations
RGB-D saliency target detection method based on interactive attention guidance and trapezoidal pyramid fusion
CN114283315A
Salient target detection algorithm based on efficient multi-scale context exploration network
CN114612684A
Cited By
RGB-D salient target detection method based on potential perception and hierarchical fusion
CN120953577A