RGB-D salient target detection method based on semantic features and biological inspiration
By constructing an encoder, a multi-stage fusion module, and a cortical decoder, and combining the DSSE module and loss function optimization, the problems of high computational complexity and insufficient feature fusion in existing RGB-D salient object detection methods are solved, achieving efficient and robust RGB-D salient object detection.
Patent Information
- Application Number
- CN202511416032.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing Transformer-based RGB-D salient target detection methods suffer from high computational complexity, insufficient feature fusion, and inadequate multi-scale processing capabilities, leading to a decrease in detection efficiency and accuracy.
We employ a semantic feature-based and bio-inspired RGB-D salient object detection method. By constructing an encoder, a multi-stage fusion module, and a cortical decoder, we achieve efficient feature fusion and hierarchical processing of RGB and depth images. We utilize the DSSE module for sparse semantic enhancement and combine it with a loss function to optimize model parameters.
It effectively reduces feature redundancy, improves feature extraction efficiency and model generalization ability, optimizes the performance of multimodal salient object detection, and enhances the robustness and accuracy of detection.
Smart Images

Figure CN120894541A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, and particularly relates to an RGB-D salient object detection method based on semantic features and biological inspiration. BACKGROUND
[0002] RGB-D salient object detection is a key research direction in the field of computer vision, and its core task is to realize the accurate positioning and segmentation of the most salient object in the image by fusing the color and texture information of the RGB image and the three-dimensional geometric features of the depth image. This technology shows important application value in image understanding, target recognition, intelligent robot navigation, medical image analysis, video monitoring and other fields.
[0003] In the prior art, models based on the Transformer architecture have attracted much attention due to their excellent global context modeling ability and long-distance dependency capturing ability. These methods significantly improve the recognition ability of salient regions by fusing hierarchical feature extraction mechanisms and self-attention mechanisms. For example, some methods use pure Transformer structures for saliency detection, which enhances the recognition of salient regions; some methods strengthen local structure modeling through token conversion; and some methods combine the encoder-decoder architecture and the multi-head attention mechanism to achieve more accurate saliency prediction.
[0004] However, there are still some problems to be solved in the current Transformer-based methods. First, the quadratic computational complexity of the Transformer limits its efficiency in practical applications. Many methods try to improve efficiency by pruning redundant tokens, but these methods often lead to a significant decrease in accuracy and face difficulties when applied to local visual Transformers, which limits the detection accuracy and generalization ability of the model. Second, cross-modal feature fusion methods usually rely on simple concatenation or weighted fusion, which fails to fully consider the independence of RGB and depth modalities, resulting in feature redundancy and failing to fully utilize the complementary information between the two. Finally, the multi-scale feature processing capability is insufficient, and existing methods mostly use basic feature aggregation techniques, which are difficult to effectively capture features of different scales, limiting the adaptability of the model to complex scenes. SUMMARY
[0005] The present application aims to provide an RGB-D salient object detection method based on semantic features and biological inspiration to solve at least one technical problem in the prior art.
[0006] The first aspect of the present application provides an RGB-D salient object detection method based on semantic features and biological inspiration, comprising the following steps:
[0007] inputting a to-be-detected RGB-D image into a salient object detection model, and outputting a salient object detection result from the salient object detection model; wherein a training method of the salient object detection model comprises the following steps:
[0008] constructing a neural network comprising an encoder, a multi-stage fusion module, and a cortical decoder;
[0009] extracting features of the RGB-D image sample through the encoder to obtain low-level, middle-level, and high-level semantic features of the RGB image and the depth image;
[0010] fusing the low-level semantic features of the RGB image and the depth image through the multi-stage fusion module to obtain low-level fusion features; fusing the middle-level semantic features of the RGB image and the depth image to obtain middle-level fusion features; and fusing the high-level semantic features of the RGB image and the depth image to obtain high-level fusion features;
[0011] decoding the low-level, middle-level, and high-level fusion features through the cortical decoder to obtain a saliency prediction map;
[0012] updating parameters of the neural network according to a loss function to obtain the salient object detection model.
[0013] The to-be-detected RGB-D image refers to an RGB-D image that needs to be subjected to salient object detection, and the RGB-D image sample refers to an RGB-D image used for training the salient object detection model. The RGB-D image contains RGB information and depth information, wherein the RGB information constitutes an RGB image, and the depth information constitutes a depth image. The salient object detection result is usually represented in the form of a saliency prediction map.
[0014] Further, the above-mentioned encoder comprises a DSSE module, in which the input features are divided into a plurality of patches, adjacent patches are fused to obtain local clustering tokens, the local clustering tokens and original image tokens are subjected to multi-head attention interaction to generate preliminary semantic tokens, the preliminary semantic tokens are subjected to spatial pooling to capture global semantic information to generate global center tokens, and the global center tokens and the preliminary semantic tokens are fused through a heterogeneous attention mechanism to generate final semantic tokens.
[0015] Further, the above-mentioned fusing of the low-level semantic features of the RGB image and the depth image comprises:
[0016] performing element-wise multiplication and splicing on the low-level semantic features of the RGB image and the depth image, and then performing 1x1 convolution on the two results to obtain a feature map ;
[0017] the feature map The low-level fusion feature is obtained by sequentially passing through a combination of 3x3 convolution and 3 3x3 dilated convolution, splicing, passing through 1x1 convolution, and splicing.
[0018] Further, the above-mentioned fusion of the middle-level semantic features of the RGB image and the depth image comprises:
[0019] The middle-level semantic features of the RGB image and the depth image are divided into two branches, one branch is subjected to element-level multiplication, and then the significant features are enhanced through the selective region attention mechanism to obtain a feature map and a feature map ; the other branch is subjected to splicing operation, and then fused with the feature map and the feature map through 1x1 convolution to obtain a middle-level fusion feature.
[0020] Further, the above-mentioned fusion of the high-level semantic features of the RGB image and the depth image comprises:
[0021] The high-level semantic features of the RGB image and the depth image are subjected to element-level multiplication, and then spliced with the splicing result of the high-level semantic features of the RGB image and the depth image, and subjected to dimension reduction optimization and channel fusion, and the spatial features are enhanced using the spatial attention mechanism to obtain a high-level fusion feature.
[0022] Further, the above-mentioned decoding of the low-level, middle-level and high-level fusion features comprises:
[0023] The middle-level and high-level fusion features are subjected to dilated convolution, and the combined features are input into three 3x3 convolution layers for parallel processing, and the three parallel processed features are subjected to a splicing operation to fuse into a single feature .
[0024] The feature is subjected to channel fusion to obtain a feature , and then subjected to dimension optimization, and the optimized feature is spliced with the low-level fusion feature, and the splicing result is subjected to channel fusion and dimension optimization again, and the processed result is spliced with to obtain a saliency prediction map.
[0025] Further, the above-mentioned loss function is:
[0026]
[0027] wherein, represents the total loss of the salient object detection model; , , , respectively represent the weight coefficients of the four loss terms. 、 、 、 respectively represent the features fused into a single feature in the cortical decoder losses generated in the stage of fusing the results of the splicing into a single feature, the stage of performing channel fusion on the splicing results again, the stage of performing dimension optimization on the splicing results again, and the stage of generating a saliency prediction map.
[0028] The second aspect of the present application provides an electronic device, comprising:
[0029] at least one processor;
[0030] and a memory connected in communication with the at least one processor;
[0031] wherein the memory stores instructions that, when executed by the at least one processor, implement the above-mentioned RGB-D salient object detection method based on semantic features and biological inspiration.
[0032] The third aspect of the present application provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the above-mentioned RGB-D salient object detection method based on semantic features and biological inspiration.
[0033] The technical scheme of the embodiment of the present application has the following beneficial effects:
[0034] (1) Through the two-stage sparse semantic enhancement mechanism, feature redundancy is effectively reduced, feature extraction efficiency and model generalization ability are improved, and computational cost is reduced;
[0035] (2) The multi-stage fusion module realizes efficient fusion of RGB and depth features, integrates information from different modalities at different scales, and optimizes the performance of multi-modal salient object detection;
[0036] (3) The cortical decoder can simulate the hierarchical visual processing mechanism of the brain, realize efficient layered feature refinement and computational optimization, and improve the robustness and generalization ability of salient object detection. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 is a method flowchart of the embodiment of the present application.
[0038] Figure 2 is a structural schematic diagram of the two-stage sparse semantic enhancement module in the embodiment of the present application.
[0039] Figure 3 is a structural schematic diagram of the multi-stage fusion module in the embodiment of the present application. DETAILED DESCRIPTION
[0040] With reference to the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described, obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0041] According to the embodiments of the present application, the salient object detection model is first trained, and then the trained salient object detection model is used for RGB-D salient object detection. For convenience of representation, the salient object detection model of the embodiments of the present application is marked as OSFBiF-Net model.
[0042] As shown in Figure 1 , the OSFBiF-Net model includes an encoder, a multi-stage fusion module and a cortical decoder. Among them, the encoder is based on the encoder of Swin Transformer, and a double-stage sparse semantic enhancement (DSSE) module is introduced. Specifically, the encoder includes four stages of encoding modules, and the third stage of encoding module is the DSSE module.
[0043] In the encoder, feature extraction is performed on the RGB-D image sample to obtain high-level semantic features of the RGB image and high-level semantic features of the depth image. Specifically, the encoder performs multi-stage feature extraction on the RGB image and the depth image of the RGB-D image sample respectively to obtain semantic feature representations at different levels. For example, four-stage feature extraction can obtain semantic features 、 、 and of the RGB image 、 、 and of the depth image; wherein may be referred to as low-level semantic features of the RGB image, 、 may be referred to as intermediate-level semantic features of the RGB image, may be referred to as high-level semantic features of the RGB image; may be referred to as low-level semantic features of the depth image, 、 may be referred to as intermediate-level semantic features of the depth image, may be referred to as high-level semantic features of the depth image.
[0044] For RGB images, the RGB image is segmented into multiple patches, and each patch is converted into a token to form an original image token (IT). The IT is processed by the encoding module of the first stage of the encoder to generate semantic features . The semantic features are generated by processing the encoding module of the second stage of the encoder In the third stage of the encoder, the features are semantically enhanced by a DSSE module. In the DSSE module, the segmented into multiple patches, and adjacent patches are fused to obtain local clustering tokens (LCTs); the LCTs and the ITs are interacted by multi-head attention to generate preliminary semantic tokens (ISTs), the ISTs are spatially pooled to capture global semantic information to generate global center tokens (GCTs), and the GCTs and the ISTs are fused by a heterogeneous attention mechanism to generate semantic features The encoding module of the fourth stage of the encoder combines the and the IT to restore the spatial details of the image to generate advanced semantic features .
[0045] For depth images, the depth image is segmented into multiple patches, and each patch is converted into a token to form an original image token (IT). The IT is processed by the encoding module of the first stage of the encoder to generate semantic features . The semantic features are generated by processing the encoding module of the second stage of the encoder In the third stage of the encoder, the features are semantically enhanced by a two-stage sparse semantic enhancement (DSSE) module. In the DSSE module, the segmented into multiple patches, and adjacent patches are fused to obtain local clustering tokens (LCTs); the LCTs and the ITs are interacted by multi-head attention to generate preliminary semantic tokens (ISTs), the ISTs are spatially pooled to capture global semantic information to generate global center tokens (GCTs), and the GCTs and the ISTs are fused by a heterogeneous attention mechanism to generate semantic features The encoding module of the fourth stage of the encoder combines the and the IT to restore the spatial details of the image to generate advanced semantic features .
[0046] The specific structure of the DSSE module used in the embodiments of the application is as follows Figure 2where "IT" denotes the tokenized representation of the input raw image (RGB image tokens or depth image tokens), "LCT" denotes the local cluster token, "IST" denotes the initial semantic token, "GCT" denotes the global center token, and "FST" denotes the final semantic token (i.e., semantic feature or semantic feature ).
[0047] In the DSSE module, the second layer feature map of the input RGB image or depth image is segmented into multiple patches, and adjacent patches are fused to obtain a local cluster token (LCT). Then, the LCT interacts with the original image token sequence IT using a multi-head attention mechanism to generate a preliminary semantic token (IST), as shown in the following formula:
[0048]
[0049] where, denotes a feedforward network, denotes the preliminary semantic token after being processed by the feedforward network, denotes the initial semantic token.
[0050] The preliminary semantic token (IST) is then spatially pooled to generate a global center token (GCT) for capturing global semantic information. The GCT is fused with the IST using a heterogeneous attention mechanism to generate a final semantic token (FST), as shown in the following formula:
[0051]
[0052] ;
[0053] where, denotes the initial semantic token, denotes the global center token, denotes a multi-head attention mechanism, denotes the key used in the multi-head attention mechanism, denotes the value used in the multi-head attention mechanism, denotes the feature after being processed by the multi-head attention mechanism, denotes a feedforward network, denotes the final semantic token.
[0054] The spatial details of the image are then restored using the self-attention layer of the Swin Transformer encoder in combination with IT and FST to obtain the high-level semantic feature of the processed RGB image or the high-level semantic feature of the processed depth image .
[0055] AsFigure 3 As shown, semantic features at different levels ( , , , , , , , The input is fed into a multi-stage fusion module, where semantic features at different levels are processed in parallel across three stages.
[0056] Phase 1: Input consists of low-level semantic features and ,Will and Perform element-wise multiplication and concatenation separately, then pass the two results together through a 1x1 convolution and concatenate them again to obtain the feature map. ;
[0057] The results are sequentially processed through a combination of 3×3 convolutions and three 3×3 dilated convolutions, then concatenated and output as the final low-level fusion feature through a 1×1 convolution. The specific formula is as follows:
[0058]
[0059]
[0060] in, This represents a 1×1 convolution operation. This represents element-wise addition. This represents element-wise multiplication. This represents the low-level feature map after initial fusion. This represents a 3×3 convolution operation. Indicates low-level fusion features, This represents the feature map after dilated convolution, where i equals 3, 5, or 7.
[0061] Second stage: Input is intermediate semantic features , and , The process is divided into two branches. One branch performs element-wise multiplication, and then feeds the result into a selective region attention (SRA) module to enhance salient features, thus obtaining a feature map. and feature map The other branch first performs a concatenation operation, then a 1×1 convolution, and finally merges it with the feature map. and feature map The fusion process yields the final intermediate fusion characteristics. and The specific formula is as follows:
[0062]
[0063]
[0064] wherein, represents the output of the selective region attention module, represents the output of the spatial attention branch, represents the output of the channel attention branch, represents the output feature map after SRA processing, represents a convolution operation, represents an intermediate semantic feature map from the RGB path or , represents an intermediate semantic feature map from the depth path or , represents an element-wise addition operation, represents an element-wise multiplication operation, represents a selective region attention mechanism.
[0065] The third stage: the input is a high-level semantic feature and , first element-level multiplication, and then splicing with the splicing result of and ; a dimension reduction optimization (DO) module is used, which first compresses the fused high-dimensional feature to a low-dimensional feature by using 3x3 convolution + batch normalization (BN) + Gaussian error linear unit activation function (GELU), and removes redundant information; then 1x1 convolution is used to refine the channel again, significantly reducing the parameter quantity and calculation quantity, while retaining key semantics, providing lighter and more efficient input features for subsequent channel fusion (CF) and spatial attention (SA); then, the channel fusion (CF) module is used, the spatial attention (SA) module is used to enhance the spatial features, and finally the high-level fusion feature is outputted, and the specific formula is as follows:
[0066]
[0067] wherein, represents a high-level fusion feature, represents an element-wise addition operation, represents an element-wise multiplication operation, represents a spatial attention mechanism, represents a channel fusion operation, represents a dimension reduction optimization operation.
[0068] In the cortical decoder, firstly, the intermediate-level fusion features are fused into a single feature ; The two-stage decoding is performed on the input intermediate-level fusion features and high-level fusion features , , respectively, different dilated convolutions are used to extract multi-scale features, and the output is merged by element-wise multiplication, the merged features are input into three 3x3 convolution layers for parallel processing to further extract features; finally, the three parallel processed extracted features are spliced once to fuse into a single feature , and the specific formula is as follows:
[0069]
[0070] wherein, represents the feature after fusing multi-layer high-level visual features, represents the high-level visual feature after preliminary fusion and processing, represents the dilated convolution operation on the up-sampled feature , represents the 3x3 convolution operation on the up-sampled feature map , represents the feature of the intermediate-level fusion feature after dilated convolution, represents the feature splicing operation, represents the element-wise multiplication operation.
[0071] Secondly, the spliced result is subjected to a channel fusion stage: the fusion feature is subjected to channel fusion to obtain a feature , and then the spliced result is subjected to a dimension optimization stage: the optimized feature is subjected to dimension optimization, and the low-level fusion feature is spliced, and the spliced result is subjected to channel fusion and dimension optimization again, and finally a saliency prediction map generation stage: the processed result is spliced with to obtain the saliency prediction map finally output by the model, and the specific formula is as follows:
[0072] ;
[0073] wherein, represents the saliency prediction map finally output, represents the channel fusion operation, represents the dimension optimization operation, represents the splicing the up-sampled feature map, representing low-level fusion features, representing feature stitching operations.
[0074] In training the OSFBiF-Net model, the saliency prediction map and the boundary prediction map are compared with the corresponding ground truth, the loss function is calculated, and the model parameters are optimized by using the loss signals for back propagation.
[0075] Specifically, the binary cross-entropy (CE) loss is calculated at multiple scales, and a corresponding weight coefficient is applied to the loss of each scale to comprehensively consider the accuracy of saliency and boundary prediction at different scales. The specific expression of the multi-scale CE loss function is as follows:
[0076]
[0077] wherein, represents the cross-entropy loss, represents the i-th element of the true label, represents the probability value of the i-th element in the probability distribution predicted by the model, represents the total number of classes, and i represents the class index.
[0078] Moreover, the intersection over union (IOU) loss is calculated at multiple scales, and a corresponding weight coefficient is applied to the loss of each scale to comprehensively consider the accuracy of saliency and boundary prediction at different scales, as shown in the following formula:
[0079]
[0080] wherein, represents the intersection over union loss, represents the calculated intersection over union value, A represents the predicted saliency object region, and B represents the true saliency object region.
[0081] The total loss function is the weighted sum of the four prediction losses, and the binary cross-entropy (CE) loss and the intersection over union (IOU) loss are calculated for each prediction loss. The loss of the i-th prediction stage is given by the following formula:
[0082] wherein,
[0083] represents the total loss of the i-th prediction stage, represents the binary cross-entropy loss between the overall prediction map and the true value, represents the intersection over union loss between the overall prediction map and the true value. wherein, intermediate stage prediction map intersection over union loss between the ground truth .
[0084] The final loss function is the weighted sum of the four prediction losses, and the formula is:
[0085]
[0086] wherein, denotes the total loss of the model; , , , denote the weight coefficients of the four loss terms respectively; , , , denote the losses generated in the stage of fusing features into a single feature in the cortical decoder, the stage of channel fusion of the spliced results again, the stage of dimension optimization of the spliced results again and the stage of generating a saliency prediction map.
[0087] Finally, the RGB-D salient object detection is performed by using the OSFBiF-Net model trained through the above steps. The RGB-D image to be detected is input into the OSFBiF-Net model, and the saliency prediction map and the boundary prediction map are generated by the OSFBiF-Net model.
[0088] The following describes experiments using the OSFBiF-Net model, and the technical effects of the embodiments of the present application are illustrated in combination with specific experimental data.
[0089] In the experiment, four commonly used image description evaluation indexes are used to evaluate the performance of the model, which are structural similarity measurement (SSIM), maximum F value (Fmax), maximum enhanced alignment measurement (EAIM) and mean absolute error (MAE). . Details of the evaluation indexes are as follows:
[0090] (1) used to evaluate the structural similarity between the prediction map and the real map, combining the structural similarity at the region level and the object level, and further combining the comparison of brightness, contrast and structure. The calculation formula of SSIM is as follows:
[0091]
[0092] wherein, and denote the prediction map and the real map respectively, and respectively represent the mean of and and respectively represent the variance of and represent the covariance of and and respectively represent constants used in the stabilization formula.
[0093] (2) Calculate precision and recall at different thresholds and report the highest score at the optimal threshold to evaluate the model's accuracy in object detection tasks, calculated as follows:
[0094]
[0095] where, precision; recall; is a parameter used to balance precision and recall, usually set to 1.
[0096] (3) Combining local pixel values and image-level mean, while considering global statistical information and local pixel matching, to evaluate the model's robustness in complex scenes, calculated as follows:
[0097]
[0098] where, predicted map, true map, similarity measure, is a parameter used to adjust the steepness of the curve.
[0099] (4) Calculate the pixel-level mean absolute error between the predicted saliency map and the true value, providing an accurate evaluation of the model's error range, especially valuable for fine evaluation, calculated as follows:
[0100]
[0101] where, total number of pixels, and respectively represent the value of the predicted map and the true map at the i-th pixel.
[0102] To verify the effectiveness of the two-stage sparse semantic enhancement module, experiments were conducted on the LFSD, NLPR and ReDweb-S datasets, and the experimental results are shown in Tables 1-3. w / o DSSE represents a model variant without using the DSSE module; Swin-DESS(S) represents a model variant that uses a smaller scale Swin Transformer as the backbone network and integrates the DSSE module; Swin-DESS(B) represents a model variant that uses a larger scale version of the Swin Transformer as the backbone network and integrates the DSSE module. The backbone network refers to the core part of the model, which is used to extract features from the input image.
[0103] Table 1 Values of various performance indicators of the DSSE module on different model variants and the LFSD dataset
[0104]
[0105] Table 2 Values of various performance indicators of the DSSE module on different model variants and the NLPR dataset
[0106]
[0107] Table 3 Values of various performance indicators of the DSSE module on different model variants and the ReDweb-S dataset
[0108]
[0109] It can be seen that compared with w / o DSSE, the flops indicator of Swin-DESS(S) is reduced by 33% while not sacrificing accuracy; it is proved that integrating the DESS module into the third stage of the Swin Transformer architecture significantly reduces the computational cost of the backbone architecture.
[0110] To verify the effectiveness of the multi-stage fusion module, experiments were conducted on the NLPR and DUT-Depth datasets, and the experimental results are shown in Tables 4 and 5. BaseLine represents a model variant that uses a simple additive fusion method; +Stage1 represents a model variant that adds the first stage of the multi-stage fusion module to the BaseLine; +Stage1+Stage2 represents a model variant that further adds the second stage of the multi-stage fusion module to +Stage1; +Stage1+Stage2+Stage3 represents a complete model variant that adds all three stages of the multi-stage fusion module.
[0111] Table 4 Values of various performance indicators of the multi-stage fusion module on the NLPR dataset
[0112]
[0113] Table 5. Values of various performance indicators of the multi-stage fusion module on the DUT-Depth dataset
[0114]
[0115] It can be seen that the multi-stage fusion module significantly improves the evaluation indicators , and while reducing the error indicators on both the NLPR and DUT-Depth datasets. The experimental results show that the introduction of the multi-stage fusion module helps to enhance feature representation and reduce noise, thereby optimizing the overall performance of the multi-stage fusion module.
[0116] To verify the effectiveness of the cortical decoder, experiments were conducted on the NLPR and DUT-Depth datasets, and the experimental results are shown in Tables 6-7. Baseline represents a model variant using a simple hierarchical decoding and summation method, w / oHighLevelVisual represents a variant that removes high-level fusion features from the base model, w / o lowLevelVisual represents a variant that excludes low-level fusion features from the base model, and HighLevelVisual+LowLevelVisual represents a model variant using a complete cortical decoder.
[0117] Table 6. Values of various performance indicators of the cortical decoder on the NLPR dataset
[0118]
[0119] Table 7. Values of various performance indicators of the cortical decoder on the DUT-Depth dataset
[0120]
[0121] It can be seen that the cortical decoder performs well on both the NLPR and DUT-Depth datasets. Compared with variants that do not use the cortical decoder, the cortical decoder significantly improves the evaluation indicators , , and reduces . Compared with the w / oHighLevelVisual and w / o lowLevelVisual variants, the complete model achieves better optimization of parameter and computational complexity, demonstrating the effectiveness of the cortical decoder.
[0122] To prove the effectiveness and generalization of the embodiments of the present application, the OSFBiF-Net model of the embodiments of the present application and the existing RGB-D salient object detection models are evaluated and compared on five commonly used data sets, and the experimental results are shown in Tables 8-12.
[0123] Table 8 Performance index results of different models on the RGBD135 data set
[0124]
[0125] Table 9 Performance index results of different models on the SIP data set
[0126]
[0127] Table 10 Performance index results of different models on the LSFD data set
[0128]
[0129] Table 11 Performance index results of different models on the ReDWeb-S data set
[0130]
[0131] Table 12 Performance index results of different models on the DUTLF-Depth data set
[0132]
[0133] It can be seen that the OSFBiF-Net exhibits excellent performance, and the performance on the RGBD135, SIP, LFSD, ReDWeb-S and DUTLF-Depth data sets is significantly better than that of other models. Compared with the existing models, the OSFBiF-Net has made obvious improvement in multiple evaluation indexes, especially in processing complex scenes and multi-modal data, which is particularly prominent. The possible reason is that the OSFBiF-Net better captures the features of the salient object by effectively fusing multi-modal features and bidirectional information flow, thereby improving the overall detection accuracy.
[0134] The above merely describes preferred embodiments of the present application, and is not intended to limit the present application in any form or in essence. It should be noted that those skilled in the art can make some improvements and supplements without departing from the method of the present application, and these improvements and supplements should also be considered as the protection scope of the present application. For those skilled in the art, some slight changes, modifications and equivalent changes made by using the disclosed technical content without departing from the spirit and scope of the present application are equivalent embodiments of the present application; meanwhile, any equivalent changes, modifications and evolution made according to the essential technology of the present application to the above embodiments are still within the scope of the technical solutions of the present application.
Claims
1. A semantic feature- and bio-inspired RGB-D salient target detection method, characterized in that, Includes the following steps: The RGB-D image to be detected is input into the salient object detection model, which outputs the salient object detection result. The training method of the salient object detection model includes the following steps: Construct a neural network that includes an encoder, a multi-stage fusion module, and a cortical decoder; By extracting features from RGB-D image samples using an encoder, low-level, mid-level, and high-level semantic features of RGB and depth images are obtained. The multi-stage fusion module fuses the low-level semantic features of RGB images and depth images to obtain low-level fusion features; it fuses the mid-level semantic features of RGB images and depth images to obtain mid-level fusion features; and it fuses the high-level semantic features of RGB images and depth images to obtain high-level fusion features. The saliency prediction map is obtained by decoding the low-level, mid-level and high-level fused features through the cortical decoder. The parameters of the neural network are updated based on the loss function to obtain a salient object detection model.
2. The method according to claim 1, characterized in that, The encoder includes a DSSE module, in which the input features are segmented into multiple patches, and adjacent patches are fused to obtain local clustering tokens. Multi-head attention interaction is performed between local clustering tokens and original image tokens to generate preliminary semantic tokens; Spatial pooling is performed on the initial semantic tokens to capture global semantic information and generate a global central token; The global central token and the initial semantic token are fused together using a heterogeneous attention mechanism to generate the final semantic token.
3. The method according to claim 1, characterized in that, The fusion of low-level semantic features of RGB images and depth images includes: The low-level semantic features of the RGB image and the depth image are multiplied and concatenated element-wise. The two results are then concatenated again through a 1×1 convolution to obtain the feature map. ; Feature map The results are sequentially combined using a combination of 3×3 convolutions and three 3×3 dilated convolutions, then concatenated and processed by a 1×1 convolution to obtain low-level fusion features.
4. The method according to claim 1, characterized in that, The fusion of intermediate semantic features from RGB and depth images includes: The intermediate semantic features of the RGB and depth images are divided into two branches. One branch performs element-wise multiplication, and then a selective region attention mechanism is used to enhance salient features to obtain the feature map. and feature map The other branch first performs a concatenation operation, then a 1×1 convolution, and finally merges it with the feature map. and feature map Fusion yields intermediate fusion characteristics.
5. The method according to claim 1, characterized in that, The fusion of high-level semantic features from RGB and depth images includes: The high-level semantic features of the RGB and depth images are first multiplied element-wise, and then concatenated with the concatenation result of the high-level semantic features of the RGB and depth images. After dimensionality reduction optimization and channel fusion, spatial features are enhanced using a spatial attention mechanism to obtain high-level fused features.
6. The method according to claim 1, characterized in that, The decoding of low-level, mid-level, and high-level fused features includes: Dilated convolutions are applied to the intermediate and high-level fusion features, which are then merged through element-wise multiplication. The merged features are then fed into three 3×3 convolutional layers for parallel processing. Finally, the three parallel-processed features are concatenated into a single feature. ; Features Channel fusion is performed to obtain features. Then, dimensionality optimization is performed, and the optimized features are concatenated with the low-level fusion features. The concatenated result is then subjected to channel fusion and dimensionality optimization again. The processed result is compared with... The data is then stitched together to obtain a significance prediction map.
7. The method according to claim 1, characterized in that, The loss function is: in, This represents the total loss of the salient object detection model; , , , These represent the weighting coefficients of the four loss terms; , , , These respectively represent the features fused into a single feature in the cortical decoder. The losses incurred in the stages of channel fusion, dimension optimization, and saliency prediction map generation.
8. An electronic device, characterized in that, include: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that, when executed by at least one processor, implement the semantic feature- and bio-inspired RGB-D salient target detection method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The device stores instructions that, when executed by a processor, implement the RGB-D salient object detection method based on semantic features and bio-inspired principles as described in any of claims 1-7.
Citation Information
Patent Citations
RGB-D salient target detection method based on multi-scale adaptive fusion
CN115690516A
RGB-D saliency target detection method based on depth quality weighting
CN116310396A
RGB-D saliency target detection method based on multi-level feature and context information fusion
CN116778180A
RGB-D salient object detection method
GB202403824D0
Cited By
Image recognition model training method, image recognition method and storage medium
CN121190910A
Image recognition model training method, image recognition method, and storage medium
CN121190910B
RGB-D video salient target detection method based on depth-guided adaptive query
CN121640350A