RGB-D salient target detection method based on decoupling contrast learning

By employing wavelet convolution and Swin Transformer collaborative modeling and cross-modal interactive parallel Transformer fusion mechanism, combined with pixel-level structure-aware contrastive learning, the problem of intermodal feature differences and boundary ambiguity in RGB-D salient object detection is solved, achieving high-precision and robust salient object detection, applicable to scenarios such as robot navigation, 3D vision analysis, and autonomous driving.

CN120997483APending Publication Date: 2025-11-21NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511119607.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing RGB-D salient object detection methods suffer from significant differences in feature distribution between modalities, lack of long-distance cross-modal dependency modeling capabilities, insufficient supervision signals to drive discriminative embedding space learning, and weak boundary modeling capabilities, leading to problems such as semantic conflicts, feature confusion, boundary blurring, and insufficient recovery of structural details.

Method used

A collaborative modeling mechanism combining wavelet convolution and Swin Transformer is adopted, and a cross-modal interactive parallel Transformer fusion mechanism is designed. A pixel-level structure-aware contrastive learning strategy is introduced. Through the cross-modal interactive parallel Transformer module and the pixel-level contrastive learning mechanism, efficient fusion of RGB and depth maps and clear boundary detection of significant target regions are achieved.

Benefits of technology

It improves the semantic consistency and structural integrity of salient regions, enhances cross-modal understanding capabilities, and achieves high-precision, robust salient target detection, making it suitable for intelligent perception scenarios such as robot navigation, 3D vision analysis, and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997483A_ABST
    Figure CN120997483A_ABST
Patent Text Reader

Abstract

The invention discloses an RGB-D salient target detection method based on decoupling contrast learning, and designs a saliency detection framework integrating expression enhancement, modal collaborative perception and structural discrimination learning by combining a structural heterogeneity problem in multi-modal modeling and utilizing the frequency domain structural advantage of a deep mode and the long-distance modeling capability of Transform. By introducing wavelet convolution and Transform joint modeling, a cross-modal interaction parallel fusion mechanism and a pixel-level structure perception contrast learning strategy, high-precision, multi-scale and boundary clear detection of a salient target area in a complex scene is realized. The method can effectively solve the problems of large information difference between modes of the RGB and the depth map, difficulty in structure alignment, fuzzy boundary prediction, weak feature expression ability and the like, significantly improves semantic consistency and structural integrity of the salient region, and has good cross-modal generalization ability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of target detection, and particularly relates to an RGB-D salient target detection method based on decoupled contrast learning. BACKGROUND

[0002] With the development of deep learning and multi-modal fusion technology, the salient object detection (SOD) task has been widely applied in image segmentation, target tracking, automatic driving and other fields. Unlike the traditional detection method based only on RGB images, the RGB-D salient object detection method can more accurately perceive the spatial boundary and stereoscopic features of the target by introducing a depth map to enhance the spatial structure information.

[0003] However, the existing RGB-D salient object detection method still faces the following technical problems:

[0004] 1. Significant difference in feature distribution between modalities, limited fusion effect: RGB images mainly encode appearance and texture information, while depth maps provide geometric structure and spatial layout, and there is a significant modal difference between the two. If they are fused without structure alignment, it is easy to cause semantic conflict and feature confusion, affecting the final detection performance.

[0005] 2. Lack of long-distance cross-modal dependency modeling capability: Most methods focus on local perception and are difficult to effectively align the semantic structure between RGB and Depth in a global range, resulting in limited cross-modal context interaction and limiting the consistent modeling of fine-grained targets in complex scenes.

[0006] 3. Insufficient supervised signal to drive discriminative embedding space learning: Current mainstream methods mostly use BCE or IoU pixel-level loss, which can provide local supervision but is difficult to constrain the aggregation and separation of the embedding space from the structure level, resulting in insufficient compactness of the salient region and blurred boundaries.

[0007] 4. Weak boundary modeling capability, insufficient structure detail recovery: Conventional decoders lack a special modeling mechanism for edge regions and are difficult to preserve fine-grained boundary information of the target, especially in weak contrast or occlusion conditions, the saliency map presents edge blur and contour loss phenomenon.

[0008] The existing RGB-D SOD technology mainly includes the following categories:

[0009] 1. Convolutional Neural Network (CNN) based methods: such as D3Net, CPFP, and JLDCF, these methods mainly rely on multi-layer convolutional feature extractors and attention mechanisms to fuse RGB and depth maps, emphasizing multi-scale context modeling and edge information reinforcement. Although good performance is achieved on standard datasets, the handling of modal differences is relatively rough, lacking explicit structural alignment mechanisms, and stability is insufficient in depth map quality fluctuations or complex scenarios.

[0010] 2. Transformer-based methods: such as DCMNet, TPCL, and D2Net, these methods introduce Transformer structures to capture long-range dependencies between RGB and Depth modalities, usually using parallel or interactive encoding structures to enhance modal fusion capabilities. Although they have advantages in structural modeling, they are highly dependent on computational resources, and still struggle to effectively address modal inconsistencies and boundary ambiguities in real-world diverse scenarios.

[0011] 3. Methods based on modal guidance and collaborative mechanisms: such as S2MA, CoNet, and HDFNet, these methods explicitly design modal guidance strategies, such as cross-modal attention, mutual information enhancement, or feature reweighting mechanisms, to achieve collaborative modeling between RGB and Depth. Although they can improve fusion accuracy, most are limited to local-level modeling, lacking deep constraints on overall semantic consistency, and have limited adaptability to complex structures or low-quality depth maps.

[0012] Although existing methods have made some progress in RGB-D SOD salient object detection, they generally struggle to address the following issues simultaneously: (1) Most existing methods use early concatenation or attention fusion strategies, lacking collaborative modeling of geometric boundaries and texture information, which can lead to incorrect or ambiguous salient region detection in modal inconsistency scenarios; (2) Although Transformer-based methods have advantages in modeling global dependencies, they lack explicit constraints on the semantic correspondence between RGB and depth modalities, especially in complex structures or regions with dramatic scale changes; (3) Existing supervision mechanisms still rely on pixel-level losses such as BCE or IoU, which cannot effectively learn structure-aware feature embedding spaces, leading to issues such as boundary ambiguity and background leakage in complex scenarios; (4) Most methods only use depth maps as auxiliary guidance information, failing to fully exploit their potential value in boundary structures, depth levels, and other aspects, resulting in insufficient geometric perception capabilities of the network. SUMMARY

[0013] In order to overcome the shortcomings of the prior art, the present application provides an RGB-D salient object detection method based on decoupled contrast learning, which combines the structural heterogeneity problem in multi-modal modeling, utilizes the frequency domain structural advantage of the depth mode and the long-distance modeling capability of the Transformer, and designs a saliency detection framework integrating expression enhancement, modal collaborative perception and structural discriminative learning. By introducing wavelet convolution and Transformer joint modeling, cross-modal interaction parallel fusion mechanism and pixel-level structure perception contrast learning strategy, high-precision, multi-scale and clear boundary detection of the salient target region in a complex scene is realized. The present application can effectively solve the problems of large information difference between RGB and depth map modalities, difficult structural alignment, fuzzy boundary prediction and weak feature expression capability, significantly improve the semantic consistency and structural integrity of the salient region, and has good cross-modal generalization ability and robustness. The method is suitable for various intelligent perception scenes with depth information, such as robot navigation, three-dimensional vision analysis, augmented reality and automatic driving, and provides stable and efficient technical support for multi-modal understanding and high-precision target extraction tasks.

[0014] The technical solution adopted by the present application to solve its technical problems is as follows:

[0015] Step 1: cross-modal feature extraction and preprocessing;

[0016] The RGB image and the depth image are respectively subjected to feature extraction;

[0017] For the RGB image: a pre-trained Swin Transformer backbone network is adopted to extract multi-scale semantic features;

[0018] For the depth image: first, a wavelet convolution module is used to extract frequency domain multi-scale structural information, and then a depth modal Swin Transformer backbone network is used to perform global modeling and strengthen the depth structure expression capability;

[0019] Step 2: cross-modal interaction parallel Transformer fusion mechanism;

[0020] A cross-modal interaction parallel Transformer module is designed, which is based on the Swin Transformer window division and position encoding mechanism, respectively performs self-attention and cross-attention calculation within and between modalities to ensure information complementarity and consistency;

[0021] Step 3: pixel-level structure perception contrast learning strategy;

[0022] The pixel-level contrast learning mechanism utilizes the cross-modal consistency of target and background pixels to improve the structural discrimination ability and boundary expression quality.

[0023] Step 4: saliency map generation and boundary structure recovery;

[0024] The saliency map is generated using the fusion features, and the edge detection module is combined to refine the boundaries of the salient regions.

[0025] Step 5: joint optimization training;

[0026] The model is trained using a comprehensive loss function to ensure the accuracy, structural integrity, and boundary clarity of the saliency detection.

[0027] Preferably, the step 1 is specifically:

[0028] Step 1-1: input RGB image and depth map where H represents the height of the input image, and W represents the width of the input image;

[0029] Step 1-2: RGB modal feature extraction, obtaining RGB feature F Rgb :

[0030]

[0031] where, is the RGB modal Swin Transformer backbone network, is the size after downsampling, s is the spatial scaling ratio, and C is the number of feature channels;

[0032] Step 1-3: depth modal wavelet convolution preprocessing:

[0033]

[0034] where F wavelet is the wavelet convolution output feature, and the wavelet convolution extracts the multi-scale frequency domain structure of the depth map, and C' is the number of output channels;

[0035] Step 1-4: depth modal global feature modeling:

[0036]

[0037] where F D is the depth modal feature, is the depth modal Swin Transformer backbone network.

[0038] Preferably, the step 2 is specifically:

[0039] Step 2-1: Obtain query Q, key K, and value V from RGB and depth features respectively through linear mapping:

[0040]

[0041] Among them, Q Rgb K Rgb V Rgb Q represents the query, key, and value for the RGB modality, respectively. D K D V D These represent the query, key, and value of the deep modality, respectively. Let m be the linear transformation matrix, and m be the attention dimension.

[0042] Step 2-2: Calculate intramodal self-attention within a fixed window:

[0043]

[0044] Among them, A Rgb A is the self-attention calculated in the RGB mode. D The self-attention is calculated in the deep modality, Softmax(·) is the Softmax operation, and B self The offset is used to encode the position within the window;

[0045] Steps 2-3: Calculate cross-modal attention to achieve information exchange:

[0046]

[0047] Among them, B cross For cross-modal attention position bias;

[0048] Steps 2-4: Output fused features:

[0049]

[0050] in, This indicates the characteristics of RGB modal fusion. F represents the deep modality fusion feature. fus This represents the fusion feature of RGB and depth modes.

[0051] Preferably, step 3 specifically comprises:

[0052] Step 3-1: From the fusion feature F fus Extract the feature vector of each pixel:

[0053] p = 1, ..., N, N = H′ × W′

[0054] Among them, fp represents the feature vector of each pixel, p represents the pth pixel;

[0055] Step 3-2: According to the salient region label M ∈ {0, 1} H′×W′ , construct the positive sample and negative sample pair;

[0056] Step 3-3: Define the pixel-level contrastive loss:

[0057]

[0058] wherein, is the pixel-level contrastive loss, is the cosine similarity, τ > 0 is a temperature parameter, is the positive sample feature, is the negative sample feature.

[0059] Preferably, the step 4 is specifically:

[0060] Step 4-1: Salient map prediction:

[0061]

[0062] wherein, S is the obtained saliency map, is the decoder network;

[0063] Step 4-2: Edge detection;

[0064] E = ε (I Rgb ,I D ) ∈ [0, 1] H′×W′

[0065] wherein, E is the output edge probability map, ε (·, ·) is based on the joint extraction of RGB and depth images to extract the edge probability;

[0066] Step 4-3: Combine the edge probability to refine the saliency map;

[0067] S refined = S ⊙ (1 + αE)

[0068] wherein, S refined is the refined saliency map, ⊙ is the element-wise product, and α > 0 is the edge weight.

[0069] Preferably, the step 5 is specifically:

[0070] The total loss is defined as:

[0071]

[0072] wherein, represents the pixel-level saliency segmentation loss, pixel-level structure-aware contrastive loss, a salient object edge supervision loss, λ1, λ2, λ3 are weight hyperparameters of each loss term;

[0073] Each loss is defined as:

[0074] Pixel-level cross-entropy salient segmentation loss:

[0075]

[0076] where N represents the total number of all pixels in the image, y i is the true label of the i-th pixel, is the saliency probability value of the i-th pixel predicted by the model;

[0077] Pixel-level contrastive loss:

[0078]

[0079] where, is the pixel-level contrastive loss, is the cosine similarity, τ>0 is the temperature parameter, is the positive sample feature, is the negative sample feature;

[0080] Edge loss:

[0081]

[0082] where E i represents whether the i-th pixel is on the real edge, is the edge probability value of the i-th pixel predicted by the model, ||·||1 represents the L1 norm.

[0083] An electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above-mentioned RGB-D salient object detection method.

[0084] A computer readable storage medium, which stores a computer program, the computer program is executed by a processor to realize the above-mentioned RGB-D salient object detection method.

[0085] A chip, comprising: a processor, used to call and run a computer program from a memory, so that the device installed with the chip executes the above-mentioned RGB-D salient object detection method.

[0086] A computer program product, the computer program product comprising a computer storage medium storing a computer program comprising instructions executable by at least one processor, when the instructions are executed by the at least one processor, the above-mentioned RGB-D salient object detection method is realized.

[0087] The beneficial effects of the present application are as follows:

[0088] 1. Because the wavelet convolution and the Swin transformer collaborative modeling mechanism are adopted, the present application can solve the problems of weak structure modeling capability of depth map and fuzzy edge expression, and realize strong structure preservation and scale adaptive depth feature modeling effect. By jointly introducing wavelet convolution and Swin Transformer structure in depth modal encoding, the multi-scale edge perception ability of wavelet in frequency domain and the global modeling ability of Transformer in spatial domain are utilized, the modeling capability of depth features to key structures such as boundaries and geometric shapes is enhanced, and the expression quality and stability of depth map in complex scenes are significantly improved.

[0089] 2. Because the cross-modal interaction parallel Transformer fusion mechanism is proposed, the present application can solve the problems of strong structural heterogeneity between RGB and depth map modal and difficult information collaboration, and realize the unified fusion effect of inter-modal semantic alignment and complementary enhancement. By designing a double-branch parallel interaction Transformer structure, semantic feature extraction is performed on RGB and depth modal respectively, and significant area alignment and complementary information interaction are guided in the middle layer, the present application realizes deep collaboration while maintaining modal independence, effectively promotes the semantic consistency of salient regions and cross-modal understanding ability.

[0090] 3. Because the structure perception pixel-level decoupling contrast learning strategy is introduced, the present application can solve the problem of insufficient boundary discriminability between salient regions and background, and realize the salient object extraction effect with clear boundaries and stable structure. By designing a contrast loss mechanism guided by saliency mask, structure-sensitive positive and negative sample pairs are constructed at the pixel level, the structure contrast relationship between different modalities and different regions is constrained, the discriminability of the model to target edges and structure contours is improved, and the boundary clarity and structure integrity of salient regions are significantly enhanced.

[0091] 4. Because the end-to-end unified optimization of the significance detection framework is constructed, the application can solve the problems of complex structure and fragmented training process of traditional multi-modal detection methods, and realize the overall system effect of stable training and efficient reasoning. The wavelet structure modeling, modal interaction fusion and structure discrimination learning are integrated in an end-to-end trainable framework, which has the characteristics of simple structure, stable performance and strong adaptability, can be efficiently deployed in various RGB-D intelligent perception scenes, and has good practical application value and generalization.

[0092] In summary, the application breaks through the key technical bottlenecks of strong inter-modal structural heterogeneity, difficult feature collaboration and blurred boundary prediction in RGB-D salient object detection. By introducing the wavelet convolution and Swin Transformer collaborative modeling mechanism, cross-modal interaction parallel fusion structure, and structure-aware pixel-level decoupling contrast learning strategy, the structural expression ability of the salient region, the modal collaborative consistency and the boundary discrimination accuracy are greatly improved. The method has high precision, high robustness and good deployment efficiency, and provides reliable technical support for multi-modal perception and high-precision visual understanding tasks in complex environments such as robot navigation, autonomous driving and augmented reality. BRIEF DESCRIPTION OF DRAWINGS

[0093] Figure 1 is a flowchart of the method of the application;

[0094] Figure 2 is an effect diagram of an embodiment of the application. DETAILED DESCRIPTION

[0095] The application will be further described below in combination with the drawings and embodiments.

[0096] The application proposes an RGB-D salient object detection method based on decoupled contrast learning, which realizes the unified optimization of inter-modal consistency modeling, salient region discrimination enhancement and boundary structure recovery by innovative technologies such as wavelet convolution and Transformer collaborative modeling of deep structural features, cross-modal interaction parallel Transformer fusion mechanism, and pixel-level structure-aware contrast learning strategy, effectively improving the robustness and generalization ability of RGB-D salient detection in complex scenes.

[0097] As shown in Figure 1 , the method of the application consists of five steps, which are as follows:

[0098] Step 1: Cross-modal feature extraction and preprocessing

[0099] In order to solve the problem of insufficient extraction of deep modal structural features, first, the features of RGB images and depth images are extracted respectively. The RGB modal adopts a pre-trained Swin Transformer backbone network Multi-scale semantic features are extracted. The depth modality is first processed by a wavelet convolution module to extract the multi-scale structure information in the frequency domain, and then a depth modality SwinTransformer backbone network is used for global modeling to enhance the depth structure expression capability. The specific process is as follows:

[0100] 1. Input RGB image and depth map

[0101] 2. RGB modality feature extraction to obtain RGB feature F Rgb :

[0102]

[0103] wherein, is the RGB modality Swin Transformer backbone network, is the size after downsampling, s is the spatial scaling ratio, and C is the number of feature channels.

[0104] 3. Depth modality wavelet convolution preprocessing:

[0105]

[0106] wherein, F wavelet is the wavelet convolution output feature, the wavelet convolution extracts the multi-scale structure of the depth map in the frequency domain,

[0107] C' is the number of output channels.

[0108] 4. Depth modality global feature modeling:

[0109]

[0110] wherein, F D is the depth modality feature, is the depth modality Swin Transformer backbone network.

[0111] This step makes full use of the multi-scale structure extraction capability of wavelet convolution in the frequency domain and combines the global modeling advantage of Swin Transformer in the spatial domain to ensure that the geometric and structural features of the depth modality are completely preserved, effectively reducing the interference caused by information redundancy and noise in the process of modality interaction, and improving the accuracy and robustness of cross-modality feature fusion.

[0112] Step two: cross-modality interaction parallel Transformer fusion mechanism;

[0113] To achieve deep fusion and information collaboration between RGB and depth modalities, a cross-modal interactive parallel Transformer module is designed. This module, based on the Swing Transformer window partitioning and positional encoding mechanism, performs qualitative self-attention and cross-attention calculations within and between modalities to ensure information complementarity and consistency. The specific implementation is as follows:

[0114] 1. Obtain the query (Q), key (K), and value (V) by linear mapping for RGB and depth features respectively:

[0115]

[0116] Among them, Q Rgb K Rgb V Rgb Q represents the query, key, and value for the RGB modality, respectively. D K D V D These represent the query, key, and value for deep modality, respectively. Let m be the linear transformation matrix, and m be the attention dimension.

[0117] 2. Calculate intramodal self-attention within a fixed window:

[0118]

[0119] Among them, A Rgb A is the self-attention calculated in the RGB mode. D The self-attention is calculated in the deep modality, Softmax(·) is the Softmax operation, and B self The offset is used to encode the position within the window.

[0120] 3. Calculate cross-modal attention to achieve information exchange:

[0121]

[0122] Among them, B cross For cross-modal attention position bias.

[0123] 4. Output fusion features:

[0124]

[0125] in, This indicates the characteristics of RGB modal fusion. F represents the deep modal fusion feature. fus This represents the fusion feature of RGB and depth modes.

[0126] To solve the problem of insufficient structural modeling ability and significant information redundancy of the depth map modality, a collaborative coding structure that fuses wavelet convolution and Swin Transformer is constructed to extract fine local structures in the frequency domain and model global context dependencies in the spatial domain. Through this structure enhancement, the depth features can accurately capture multi-scale geometric relationships and boundary contours, providing stronger structural expression and discriminability. This design ensures effective alignment and collaboration of cross-modal features, promoting the unified expression of semantic and structural information.

[0127] Step three: pixel-level structure-aware contrastive learning strategy

[0128] To address the structural heterogeneity and feature differences between modalities, a pixel-level contrastive learning mechanism is designed to improve the structural discriminability and boundary expression quality by leveraging the cross-modal consistency of salient objects and background pixels. The specific steps are as follows:

[0129] 1. Extract each pixel feature vector from the fusion feature F fus :

[0130] p = 1, …, N, N = H' x W'

[0131] where f p represents the per-pixel feature vector, and p represents the pixel number.

[0132] 2. Construct positive and negative sample pairs based on the salient region label M e {0, 1} H′×W′ .

[0133] 3. Define the pixel-level contrastive loss:

[0134]

[0135] where L is the pixel-level contrastive loss, cos is the cosine similarity, τ > 0 is the temperature parameter, f is the positive sample feature, and f is the negative sample feature.

[0136] To address the problem of blurred boundaries and difficulty in distinguishing objects from backgrounds in salient object detection, a structure-aware pixel-level decoupled contrastive learning strategy is introduced to improve the expression quality of salient regions from both class discrimination and structure preservation dimensions. This mechanism not only achieves cross-modal semantic consistency alignment but also significantly enhances the perception of weak saliency, small targets, and fine-grained edges. Ultimately, this module effectively improves the model's saliency discrimination ability in complex environments, providing stable support for fine structure segmentation and accurate target extraction.

[0137] Step four: saliency map generation and boundary structure recovery

[0138] The saliency map is generated by using fusion features, and the edge detection module is used to refine the boundary of the salient region. The specific implementation steps are as follows:

[0139] 1. Saliency map prediction:

[0140]

[0141] where S is the obtained saliency map, is the decoder network.

[0142] 2. Edge detection:

[0143] E = ε(I Rgb ,I D ) ∈ [0, 1] H′×W′

[0144] where E is the output edge probability map, and ε(·,·) is based on the joint extraction of RGB and depth images to extract edge probability.

[0145] 3. Refine the saliency map combined with the edge probability:

[0146] S refined = S ⊙ (1 + αE)

[0147] where S refined is the refined saliency map, ⊙ is the element-wise product, and α > 0 is the edge weight.

[0148] This step aims to solve the problem of traditional multi-modal saliency detection method structure fragmentation, training complexity, and deployment difficulty, and constructs an end-to-end unified optimization saliency detection framework. The framework integrates key modules such as deep structural feature modeling, cross-modal interaction fusion, and structure perception contrast learning into the same network system, forming an efficient collaborative training closed loop.

[0149] Step five: joint optimization training;

[0150] The model is trained using a comprehensive loss function to ensure the accuracy, structural integrity, and boundary clarity of saliency detection. The total loss is defined as:

[0151]

[0152] where, represents the pixel-level saliency segmentation loss, represents the pixel-level structure perception contrast loss, represents the salient target edge supervision loss, and λ1, λ2, λ3 are the weight hyperparameters of each loss term.

[0153] Each loss is defined as:

[0154] 1. Pixel-level cross-entropy saliency segmentation loss:

[0155]

[0156] where N represents the total number of all pixels in the image, y i is the true label of the i-th pixel, is the saliency probability value of the i-th pixel predicted by the model.

[0157] 2. Pixel-level contrastive loss:

[0158]

[0159] where, is the pixel-level contrastive loss, is the cosine similarity, and τ>0 is the temperature parameter, is the positive sample feature, is the negative sample feature.

[0160] 3. Edge loss:

[0161]

[0162] where E i represents whether the i-th pixel is on the real edge (obtained by the edge extraction operator through the true label), is the edge probability value of the i-th pixel predicted by the model, and ||·||1 represents the L1 norm.

[0163] This step aims to achieve the unified optimization of accurate prediction of salient regions, structural integrity preservation and boundary detail recovery, and constructs a joint optimization objective composed of three types of loss functions. Through this joint training strategy, the model achieves an efficient trade-off between accuracy, discriminability and boundary clarity, providing a stable and reliable learning foundation for RGB-D salient object detection in complex scenes.

[0164] Innovations of the present application:

[0165] 1. Because the wavelet convolution and Swin transformer collaborative modeling mechanism are adopted, the present application can solve the problems of weak structure modeling capability of depth map and fuzzy edge expression, and achieve strong structure preservation and scale adaptive deep feature modeling effect. By jointly introducing wavelet convolution and Swin Transformer structure in depth modal encoding, the multi-scale edge perception ability of wavelet in frequency domain and the global modeling ability of Transformer in spatial domain are utilized, the modeling capability of depth features to key structures such as boundaries and geometric shapes is enhanced, and the expression quality and stability of depth map in complex scenes are significantly improved.

[0166] 2. By proposing a cross-modal interactive parallel Transformer fusion mechanism, this invention addresses the challenges of strong structural heterogeneity and difficulty in information collaboration between RGB and depth map modalities, achieving a unified fusion effect of semantic alignment and complementary enhancement between modalities. Through the design of a dual-branch parallel interactive Transformer structure, semantic features are extracted from both RGB and depth modalities, and they are guided to perform salient region alignment and complementary information interaction in the intermediate layer. This invention achieves deep collaboration while maintaining modal independence, effectively promoting semantic consistency of salient regions and cross-modal understanding capabilities.

[0167] 3. By introducing a structure-aware pixel-level decoupled contrastive learning strategy, this invention can solve the problem of insufficient boundary discrimination between salient regions and the background, achieving salient target extraction with clear boundaries and stable structure. Through the design of a contrastive loss mechanism guided by saliency masks, structure-sensitive positive and negative sample pairs are constructed at the pixel level. By constraining the structural contrast relationships between different modalities and different regions, the model's ability to discriminate target edges and structural contours is improved, significantly enhancing the boundary clarity and structural integrity of salient regions.

[0168] 4. Because it constructs an end-to-end unified and optimized saliency detection framework, this invention can solve the problems of complex structure and fragmented training process in traditional multimodal detection methods, achieving stable training and efficient inference as a whole system. By integrating wavelet structure modeling, modal interaction fusion, and structure discrimination learning into a unified end-to-end trainable framework, it features simple structure, stable performance, and strong adaptability. It can be efficiently deployed in various RGB-D intelligent sensing scenarios, demonstrating good practical application value and scalability.

[0169] In summary, this invention overcomes key technical bottlenecks in RGB-D salient object detection, such as strong intermodal structural heterogeneity, difficulty in feature collaboration, and fuzzy boundary prediction. By introducing a wavelet convolution and Swing Transformer collaborative modeling mechanism, a cross-modal interactive parallel fusion structure, and a structure-aware pixel-level decoupled contrastive learning strategy, it significantly improves the structural representation ability of salient regions, modal collaboration consistency, and boundary discrimination accuracy. This method possesses high accuracy, high robustness, and good deployment efficiency, providing reliable technical support for multimodal perception and high-precision visual understanding tasks in complex environments such as robot navigation, autonomous driving, and augmented reality.

[0170] Experimental results:

[0171] Table 1: Performance comparison of different methods on benchmark datasets DUT-RGBD, NJU2K, and NLPR

[0172]

[0173] From the test results, it can be seen that

[0174] As Figure 2 , the application achieves the optimal performance on three mainstream RGB-D salient object detection datasets (DUT-RGBD, NJU2K, NLPR). In terms of precision indicators, the application leads in F-measure (F β ), E-measure (E ξ ) and structural similarity measure (S α ), and is significantly better than the existing representative methods TPCL, CATNet, CIRNet and DCMF, showing stronger salient region discrimination ability and structure preservation ability; in terms of error indicators, the average absolute error of the application is the lowest on all datasets, further indicating that the model prediction is more fine and the boundary is clearer. In addition, while ensuring performance leading, the parameter amount and the calculation amount of the application are far lower than CATNet, and are better than TPCL within a reasonable complexity range, indicating that the application realizes a good balance between performance and efficiency in model design. In summary, the experimental results verify the effectiveness of the key strategies such as frequency domain modeling, cross-modal interaction and structure contrast learning proposed by the application, which can realize high-quality RGB-D salient object detection while maintaining computational efficiency, and has good generalization ability and engineering practical value.

Claims

1. An RGB-D salient object detection method based on decoupled contrastive learning, characterized in that, Comprising the following steps: Step 1: Cross-modal feature extraction and preprocessing; Feature extraction is performed on RGB images and depth images respectively; For RGB images: adopt pre-trained Swin Transformer backbone network Extract multi-scale semantic features; For depth image: first through wavelet convolution module Extract frequency domain multi-scale structure information, and then use depth modal SwinTransformer backbone network Global modeling is performed to strengthen the depth structure expression capability; Step 2: Cross-modal interaction parallel Transformer fusion mechanism; A cross-modal interaction parallel Transformer module is designed, which is based on Swin Transformer window division and position encoding mechanism, and performs self-attention and cross-attention calculation within and between modalities respectively to ensure information complementarity and consistency; Step 3: Pixel-level structure perception contrast learning strategy; Pixel-level contrast learning mechanism is used to improve the structure discrimination ability and boundary expression quality by utilizing the cross-modal consistency of target and background pixels; Step 4: Saliency map generation and boundary structure recovery; A saliency map is generated using the fused features, and an edge detection module is used to refine the boundaries of the salient regions; Step 5: Joint optimization training; A comprehensive loss function is used to train the model to ensure the accuracy, structural integrity and boundary clarity of salient detection.

2. The RGB-D salient object detection method based on decoupled contrastive learning according to claim 1, characterized in that, Said step 1 is specifically: Step 1-1: inputting an RGB image and a depth map where H denotes the height of the input image and W denotes the width of the input image. Step 1-2: RGB modality feature extraction, get RGB feature F Rgb : wherein, is the RGB modality Swin Transformer backbone network, is the size after downsampling, s is the spatial scaling ratio, and C is the number of feature channels. Step 1-3: Depth modality wavelet convolution preprocessing: where F wavelet is the wavelet convolution output feature, and the wavelet convolution extracts the multi-scale frequency domain structure of the depth map, and C' is the output channel number; Step 1-4: Depth modality global feature modeling: wherein F D is a deep modal feature, is a deep modal Swin Transformer backbone network.

3. The RGB-D salient object detection method based on decoupled contrastive learning according to claim 2, characterized in that, Said step 2 is specifically: Step 2-1: Linear mapping is performed on RGB and depth features to obtain query Q, key K and value V respectively: where Q Rgb , K Rgb , V Rgb represent the query, key and value of RGB modality respectively, Q D , K D , V D represent the query, key and value of depth modality respectively, is a linear transformation matrix, and m is the attention dimension. Step 2-2: Calculate the intra-modal self-attention in the fixed window: wherein A Rgb is the self-attention calculated under the RGB modality, A D is the self-attention calculated under the depth modality, Softmax(·) is the Softmax operation, B self is the position encoding bias within the window; Step 2-3: Calculate the cross-modal cross-attention to realize information interaction: wherein B cross is a cross-modal cross-attention position bias; Step 2-4: Output fused features: wherein, denotes the RGB modality fused feature, denotes the depth modality fused feature, F fus denotes the RGB and depth modality fused feature.

4. The RGB-D salient object detection method based on decoupled contrastive learning according to claim 3, characterized in that, Said step 3 is specifically: Step 3-1: Extract each pixel feature vector from the fused feature F fus Step 3-2: Compute the pixel-wise similarity score where f p represents the feature vector of each pixel, p represents the pth pixel; Step 3-2: Construct positive and negative sample pairs based on the salient region label M e {0, 1} H′×W′ , construct positive and negative sample pairs; Step 3-3: Define pixel-level contrast loss: wherein, is a pixel-level contrastive loss, is a cosine similarity, and τ > 0 is a temperature parameter, is a positive sample feature, is a negative sample feature.

5. The RGB-D salient object detection method based on decoupled contrastive learning according to claim 4, characterized in that, Said step 4 is specifically: Step 4-1: Saliency map prediction: wherein S is the resulting saliency map, is a decoder network; Step 4-2: Edge detection; E = ε(I Rgb , D ) ∈ [0, 1] H′×W′ Where E is the output edge probability map, and ε(·,·) is based on RGB and depth images to extract edge probability; Step 4-3: Refine the saliency map combined with the edge probability; S refined = S ⊙ (1 + aE) where S refined is the refined saliency map, is the element-wise product, and a > 0 is the edge weight.

6. The RGB-D salient object detection method based on decoupled contrastive learning according to claim 5, characterized in that, Said step 5 is specifically: The total loss is defined as: wherein, represents a pixel-level saliency segmentation loss, represents a pixel-level structure-aware contrastive loss, represents a salient object edge supervision loss, λ1, λ2, λ3 are weight hyperparameters of each loss term. The loss is defined as: Pixel-level cross-entropy saliency segmentation loss: where N represents the total number of all pixels in the image, y i is the true label of the i-th pixel, is the saliency probability value of the i-th pixel predicted by the model. Pixel-level contrast loss: wherein, is a pixel-level contrastive loss, is a cosine similarity, τ > 0 is a temperature parameter, is a positive sample feature, is a negative sample feature; Edge loss: where E i denotes whether the i-th pixel is on a real edge or not, is the edge probability value of the i-th pixel predicted by the model, and ||·||1 denotes the L1 norm.

7. An electronic device, comprising: Comprise: Processor and memory; The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to make the electronic device execute the method as claimed in any one of claims 1 to 6.

8. A computer-readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to realize the method as claimed in any one of claims 1 to 6.

9. A chip, characterized by Comprise: The processor is used to call and run the computer program from the memory, so that the device installed with the chip executes the method as claimed in any one of claims 1 to 6.

10. A computer program product, characterised in that, The computer program product comprises a computer storage medium storing a computer program, and the computer program comprises instructions executable by at least one processor, which realize the method as claimed in any one of claims 1 to 6 when executed by the at least one processor.

Citation Information

Cited By

  • Target tracking method based on multi-modal matching, product, equipment and storage medium

    CN122134760A