RGB-D salient target detection method based on texture enhancement guidance
By constructing a texture enhancement module and a dual-path adaptive interaction module, and utilizing high-frequency texture priors to constrain noise suppression of deep features, the problem of decreased detection accuracy caused by modal differences in RGB-D salient object detection is solved, achieving higher detection accuracy and model robustness.
Patent Information
- Application Number
- CN202510808852.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
AI Technical Summary
Existing RGB-D salient object detection methods ignore the essential modal differences between depth data and RGB images, resulting in decreased detection accuracy. In addition, noise and errors are amplified during the feature extraction process, affecting the accuracy of detection results and the robustness of the model.
An RGB-D salient object detection method based on texture enhancement guidance is adopted. By constructing a texture enhancement module, a dual-path adaptive interaction module and a dynamic decoding module, the noise suppression of deep features is constrained by high-frequency texture priors, cross-modal semantic associations are established, and the calibration of multi-level features is achieved through deformable cross-scale transformation technology.
In multi-target and low-quality depth input scenarios, the boundary integrity and noise suppression capabilities are significantly improved, thereby enhancing the accuracy of detection results and the robustness of the model.
Smart Images

Figure CN120635585A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of salient object detection methods, and specifically relates to an RGB-D salient object detection method based on texture enhancement guidance. Background Art
[0002] Salient object detection aims to locate and segment the most visually appealing objects in an image. This technology has been widely used in image retrieval, scene classification, object tracking, person re-identification, and other fields, often serving as a preprocessing step for algorithmic models. With the development and widespread adoption of depth sensors, cross-modal learning methods that fuse RGB images with depth information have gradually gained attention and become a hot topic in the field of computer vision. This multimodal data fusion strategy significantly improves object perception and understanding in complex scenes by exploiting the complementarity between color information and spatial geometric features.
[0003] In recent years, a variety of cross-modal fusion methods have emerged in the field of RGB-D salient object detection to improve detection performance. The mainstream methods can be divided into two categories: one that directly concatenates the depth map as a supplementary channel with the RGB image at the initial encoding stage, and the other that combines bimodal features through an intermediate-layer feature fusion strategy. However, existing methods ignore the essential modal differences between depth data and RGB images. Specifically, as a single-channel geometric representation, the depth map's pixel values reflect the spatial distance distribution of the scene, but lack the rich high-frequency texture details of the RGB image. This modality difference leads to two key issues: first, sensor noise and depth estimation errors are amplified layer by layer during the feature extraction process, forming cumulative interference that propagates across layers; second, directly adopting channel concatenation or element-by-element addition fusion strategies makes it difficult to effectively establish cross-modal semantic correspondences, resulting in the gradual degradation of high-frequency detail information during the decoding stage, ultimately affecting the accuracy of detection results and the robustness of the model. Summary of the Invention
[0004] The purpose of this invention is to provide an RGB-D salient object detection method based on texture enhancement guidance, which solves the problem that existing detection methods ignore the essential modal differences between depth data and RGB images, resulting in reduced detection accuracy.
[0005] The technical solution adopted by the present invention is an RGB-D salient object detection method based on texture enhancement guidance, which is specifically implemented in the following steps: Step 1: Build the dataset and encoder; Step 2: construct a texture enhancement module; Step 3: construct a dual-path adaptive interaction module; Step 4: Build a dynamic decoding module.
[0006] The present invention is also characterized in that: Step 1 is implemented as follows: Step 1.1, collect salient object dataset; Step 1.2, classify and organize the data set collected in step 1.1; Step 1.3: Divide all dataset files into training set, validation set, and test set in proportion; Step 1.4: Use PVTv2 as the backbone for feature extraction, load pre-trained weights, and build the encoder. In step 1.5, the resolution of the RGB image and depth image is adjusted to 384×384 and input into the encoder respectively.
[0007] Step 2 is implemented as follows: Step 2.1, build a high-frequency feature extraction module; Step 2.2, build the deep confidence module; Step 2.3: Build a cross-modal gating fusion module.
[0008] Step 2.1 is implemented as follows: Step 2.1.1: Construct RGB multi-scale residual high-frequency feature extraction branch, and the input feature is , then enter two parallel branches. Branch 1 uses a 3×3 convolution layer to capture local high-frequency texture features. The convolution kernel size is 3×3, the padding is 1, and the number of output channels is 1 / 2 of the target channel number. Branch 2 uses a 5×5 convolution layer to capture large-scale high-frequency structural features. The convolution kernel size is 5×5, the padding is 2, and the number of output channels is 1 / 2 of the target channel number. The outputs of the two branches are spliced along the channel dimension, and after batch normalization and ReLU activation function, they are fused with the input residual connection. The formula is shown in formula (1) (2):
[0009] Where, The convolution kernel is The convolutional layer, Indicates that features are merged along the channel, express activation function, represents batch normalization, It means element-by-element addition; Step 2.1.2: Construct a dynamic frequency-channel attention joint mechanism to combine the features obtained in step 2.1.1 As input, the high-frequency components are separated by residual calculation, and the spatial dimension is compressed by global average pooling to generate channel-level weights. The formula is shown in Equation (3) and (4):
[0010] Where, represents 3×3 local average pooling, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, Represents the Sigmoid activation function; Step 2.1.3, dynamic feature fusion, input the high-frequency information obtained in step 2.1.2 and the channel machine weights normalized by softmax Perform cross-dimensional splicing, and finally generate the final weight matrix, and weighted optimize the input features. The formula is shown in formula (5):
[0011] Where, Indicates that features are merged along the channel, represents a convolution layer with a convolution kernel of 1×1. represents the Sigmoid activation function, express .
[0012] Step 2.2 is implemented as follows: Step 2.2.1, build a depth confidence predictor and input depth features , through 3×3 convolutional layers, The activation function is combined with 1×1 convolution to extract important features. Finally, the Sigmoid function generates a deep feature confidence map to suppress interference in the noise area. The formula is shown in formula (6):
[0013] Where, represents a convolution layer with a convolution kernel of 1×1. express Activation function, output weight matrix ; Step 2.2.2, use the data generated in step 2.2.1 Deep features Perform confidence weighting processing, and the formula is shown in formula (7):
[0014] Where, express .
[0015] Step 2.3 is implemented as follows: Step 2.3.1, build a cross-modal attention module to take the weighted deep features generated in step 2.2.2 High-frequency features of RGB images The channel dimension is spliced, and the spliced fusion features are compressed in the spatial dimension through global average pooling, and then the feature interaction and dimension adjustment are realized through the fully connected layer. Finally, the Sigmoid activation function is used to decouple the output into two dynamically adjusted fusion coefficients, which are used to guide the adaptive fusion of the features of the depth mode and the RGB mode respectively. The formula is shown in Equation (8) and (9):
[0016] Where, Indicates that features are merged along the channel, represents global average pooling, represents the fully connected layer, activation , represents the feature splitting operation, , are the channel attention weights of depth and RGB features respectively; Step 2.3.2, construct the residual enhancement fusion, using the attention weights generated in step 2.3.1 Weighted fusion of two modal features, the weighted depth features , RGB high-frequency features Compared with the original depth feature Add together and output the enhanced depth representation, which is expressed as shown in formula (10):
[0017] Where, express .
[0018] Step 3 is implemented as follows: Step 3.1, construct the residual enhancement feature transformation module; For RGB features and deep features Multi-scale nonlinear transformations are performed separately. First, a deep convolution with a kernel size of 3×3 is used to extract local spatial features. The number of output channels is kept at C. Then, a 1×1 convolution is used to expand the number of channels to 2C. The activation function uses GELU to enhance the nonlinear expression ability. Finally, a second 1×1 convolution is used to restore the number of channels to C. The residual is connected with the original input to retain the original feature information while enhancing the modal specificity. The formula is shown in Equations (11) and (12):
[0019] Where, represents a depth convolution layer with a convolution kernel of 3×3. represents a convolution layer with a convolution kernel of 1×1. express activation function, It means element-by-element addition; Step 3.2, construct a dynamic gated multimodal fusion module; Step 3.3, build a dual attention collaborative optimization module.
[0020] Step 3.2 is implemented as follows: Step 3.2.1, the enhanced RGB features and deep features Splicing along the channel dimension to generate joint features ; Step 3.2.2: Construct dynamic weights, compress the spatial dimension to 1×1 through global average pooling, retain channel-level statistical information, and then learn the inter-modal weight distribution through the fully connected layer to apply Function, generate normalized weight distribution, and finally dynamically superimpose the two modal features according to the weight. The formula is shown in formula (13) (14):
[0021] Where, represents global average pooling, represents the fully connected layer, express activation function, Represents element-wise multiplication.
[0022] Step 3.3: Follow the steps below to implement Step 3.3.1, fusion features Perform global average pooling to generate a channel description vector, then generate a channel-level attention map through a fully connected layer and a Sigmoid function, and apply the attention map to the original features. The formula is shown in Equation (15):
[0023] Where, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, activation , Represents element-wise multiplication; Step 3.3.2, for the features generated in step 3.3.1 Do processing, extract the maximum and mean of the fusion features along the channel dimension, generate a dual-channel spatial description map, use 7×7 convolution to capture long-range spatial dependencies, and use The function generates spatial weights and applies the spatial attention weights to the channel-optimized features. The formula is shown in Equation (16):
[0024] Where, and Respectively represent the maximum pooling and average pooling operations along the channel latitude, represents a convolution layer with a convolution kernel of 7×7. Indicates that features are merged along the channel, activation , Represents element-wise multiplication.
[0025] Step 4 is implemented as follows: Step 4.1, build dynamic token interaction module; Step 4.1 is implemented as follows: Step 4.1.1: Use dynamic token global context modeling and input high-level feature maps , flattening it into a spatial sequence , semantic tokens are generated by linear projection, spatial features are converted into learnable semantic token sequences, and global context information is captured. The formula is shown in Equation (17):
[0026] Step 4.1.2, the token sequence Input the Transformer encoder and optimize cross-region semantic associations through the multi-head self-attention mechanism. The formula is shown in Equations (18), (19), and (20):
[0027] Where, Represents a linear transformation matrix, generating queries, keys, and values. , Indicates the value used to scale the similarity to avoid gradient explosion. express activation function, Represents element-wise multiplication; In step 4.1.3, the similarity between the original features and the optimized tokens is calculated, and the spatial attention weights are generated and fused. The formula is shown in Equations (21) and (22):
[0028] Where, It is used to scale the similarity value to avoid gradient explosion, T represents the transposition operation, express activation function, represents element-wise multiplication, It means element-by-element addition; Step 4.2: construct a deformable cross-scale transformation and residual fusion module; Step 4.2 is as follows: Apply separable convolution and Sigmoid activation to generate a mask, and then weight the low- and middle-level features and the mask and concatenate them to obtain And through average pooling and channel attention optimization, finally with convolution The features are connected with residuals, and the formula is shown in (23) (24) (25):
[0029] Where, ( ) represents a depth-wise separable convolution layer with a convolution kernel of 3×3. ( )express activation function, represents element-wise multiplication, ( ) indicates that features are merged along the channel, ( ) represents a convolution layer with a convolution kernel of 3×3. ( ) represents average pooling, ( ) represents the channel attention mechanism.
[0030] The beneficial effects of the present invention are: This invention is based on a texture enhancement-guided RGB-D salient target detection method. Through the collaboration of a texture-guided depth enhancement module, a dual-path adaptive interaction module, and a dynamic decoding module, an innovative solution for RGB-D salient target detection is constructed. This method addresses the key issues of heterogeneous modal feature degradation, cross-layer propagation of depth noise, and multi-scale semantic mismatch in the prior art. It uses high-frequency texture prior constraints to achieve noise suppression of depth features, establishes cross-modal semantic associations through a dynamic interaction mechanism of channel-space collaboration, and uses deformable cross-scale transformation technology to achieve progressive calibration of multi-level features. In input scenarios such as multiple targets and low-quality depth, this method exhibits significant advantages in boundary integrity and noise suppression. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1This is the overall model framework diagram of the RGB-D salient object detection method based on texture enhancement guidance in the present invention; Figure 2 This is a network architecture diagram of the texture enhancement module in step 2 of the RGB-D salient object detection method based on texture enhancement guidance of the present invention; Figure 3 This is a network architecture diagram of the dual-path adaptive interaction module in step 3 of the texture enhancement-guided RGB-D salient object detection method of the present invention; Figure 4 This is a network architecture diagram of the dynamic decoding module in step 4 of the texture enhancement-guided RGB-D salient target detection method of the present invention. DETAILED DESCRIPTION
[0032] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] The present invention is based on the RGB-D salient object detection method guided by texture enhancement, which is specifically implemented by the following steps: Step 1: Build the dataset and encoder; Step 2: Build a texture enhancement module, such as Figure 2 As shown; Step 3: Construct a dual-path adaptive interaction module, such as Figure 3 As shown; Step 4: Build a dynamic decoding module, such as Figure 4 shown.
[0034] Example 1 The present invention is based on the RGB-D salient object detection method guided by texture enhancement, wherein step 1 is specifically implemented according to the following steps: Step 1.1, collect salient object dataset; Step 1.2, classify and organize the data set collected in step 1.1; Step 1.3: Divide all dataset files into training set, validation set, and test set in proportion; Step 1.4, use PVTv2 as the backbone for feature extraction, load pre-trained weights, and build an encoder. The constructed network structure is shown in the following figure: Figure 1 As shown; In step 1.5, the resolution of the RGB image and depth image is adjusted to 384×384 and input into the encoder respectively.
[0035] Example 2 The present invention is based on the RGB-D salient object detection method guided by texture enhancement, wherein step 2 is specifically implemented according to the following steps: Step 2.1, build a high-frequency feature extraction module, such as Figure 2As shown in the RGB branch; Step 2.1.1: Construct RGB multi-scale residual high-frequency feature extraction branch, and the input feature is , then enter two parallel branches. Branch 1 uses a 3×3 convolution layer to capture local high-frequency texture features. The convolution kernel size is 3×3, the padding is 1, and the number of output channels is 1 / 2 of the target channel number. Branch 2 uses a 5×5 convolution layer to capture large-scale high-frequency structural features. The convolution kernel size is 5×5, the padding is 2, and the number of output channels is 1 / 2 of the target channel number. The outputs of the two branches are spliced along the channel dimension, and after batch normalization and ReLU activation function, they are fused with the input residual connection. The formula is shown in formula (1) (2):
[0036] Where, The convolution kernel is The convolutional layer, Indicates that features are merged along the channel, express activation function, represents batch normalization, It means element-by-element addition; Step 2.1.2: Construct a dynamic frequency-channel attention joint mechanism to combine the features obtained in step 2.1.1 As input, the high-frequency components are separated by residual calculation, and the spatial dimension is compressed by global average pooling to generate channel-level weights. The formula is shown in Equation (3) and (4):
[0037] Where, represents 3×3 local average pooling, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, Represents the Sigmoid activation function; Step 2.1.3, dynamic feature fusion, input the high-frequency information obtained in step 2.1.2 and the channel machine weights normalized by softmax Perform cross-dimensional splicing, and finally generate the final weight matrix, and weighted optimize the input features. The formula is shown in formula (5):
[0038] Where, Indicates that features are merged along the channel, represents a convolution layer with a convolution kernel of 1×1. represents the Sigmoid activation function, express .
[0039] Step 2.2, build the deep confidence module; Step 2.3, construct a cross-modal gating fusion module, such as Figure 2 As shown in the attention weight part.
[0040] Example 3 The present invention is based on the texture enhancement guided RGB-D salient object detection method, wherein step 2.2 is specifically implemented as follows: Step 2.2.1, build a depth confidence predictor and input depth features , through 3×3 convolutional layers, The activation function is combined with 1×1 convolution to extract important features. Finally, the Sigmoid function generates a deep feature confidence map to suppress interference in the noise area. The formula is shown in formula (6):
[0041] Where, represents a convolution layer with a convolution kernel of 1×1. express Activation function, output weight matrix ; Step 2.2.2, use the data generated in step 2.2.1 Deep features Perform confidence weighting processing, and the formula is shown in formula (7):
[0042] Where, express .
[0043] Step 2.3 is specifically implemented as follows: Step 2.3.1, build a cross-modal attention module to take the weighted deep features generated in step 2.2.2 High-frequency features of RGB images The channel dimension is spliced, and the spliced fusion features are compressed in the spatial dimension through global average pooling, and then the feature interaction and dimension adjustment are realized through the fully connected layer. Finally, the Sigmoid activation function is used to decouple the output into two dynamically adjusted fusion coefficients, which are used to guide the adaptive fusion of the features of the depth mode and the RGB mode respectively. The formula is shown in Equation (8) and (9):
[0044] Where, Indicates that features are merged along the channel, represents global average pooling, represents the fully connected layer, activation , represents the feature splitting operation, , are the channel attention weights of depth and RGB features respectively; Step 2.3.2, construct the residual enhancement fusion, using the attention weights generated in step 2.3.1 Weighted fusion of two modal features, the weighted depth features , RGB high-frequency features Compared with the original depth feature Add together and output the enhanced depth representation, which is expressed as shown in formula (10):
[0045] Where, express .
[0046] Example 4 The present invention is based on the RGB-D salient object detection method guided by texture enhancement, wherein step 3 is specifically implemented as follows: Step 3.1, construct the residual enhancement feature transformation module, such as Figure 3 As shown in the feature enhancement module; For RGB features and deep features Multi-scale nonlinear transformations are performed separately. First, a deep convolution with a kernel size of 3×3 is used to extract local spatial features. The number of output channels is kept at C. Then, a 1×1 convolution is used to expand the number of channels to 2C. The activation function uses GELU to enhance the nonlinear expression ability. Finally, a second 1×1 convolution is used to restore the number of channels to C. The residual is connected with the original input to retain the original feature information while enhancing the modal specificity. The formula is shown in Equations (11) and (12):
[0047] Where, represents a depth convolution layer with a convolution kernel of 3×3. represents a convolution layer with a convolution kernel of 1×1. express activation function, It means element-by-element addition; Step 3.2, construct a dynamic gated multimodal fusion module, such as Figure 3 As shown in the RGB and depth weight distribution section; Step 3.2 is implemented as follows: Step 3.2.1, the enhanced RGB features and deep features Splicing along the channel dimension to generate joint features ; Step 3.2.2: Construct dynamic weights, compress the spatial dimension to 1×1 through global average pooling, retain channel-level statistical information, and then learn the inter-modal weight distribution through the fully connected layer to apply Function, generate normalized weight distribution, and finally dynamically superimpose the two modal features according to the weight. The formula is shown in formula (13) (14):
[0048] Where, represents global average pooling, represents the fully connected layer, express activation function, Represents element-wise multiplication; Step 3.3, construct the dual attention collaborative optimization module, such as Figure 3 As shown in the channel attention and spatial attention modules; Step 3.3: Follow the steps below to implement Step 3.3.1, fusion features Perform global average pooling to generate a channel description vector, then generate a channel-level attention map through a fully connected layer and a Sigmoid function, and apply the attention map to the original features. The formula is shown in Equation (15):
[0049] Where, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, activation , Represents element-wise multiplication; Step 3.3.2, for the features generated in step 3.3.1 Do processing, extract the maximum and mean of the fusion features along the channel dimension, generate a dual-channel spatial description map, use 7×7 convolution to capture long-range spatial dependencies, and use The function generates spatial weights and applies the spatial attention weights to the channel-optimized features. The formula is shown in Equation (16):
[0050] Where, and Respectively represent the maximum pooling and average pooling operations along the channel latitude, represents a convolution layer with a convolution kernel of 7×7. Indicates that features are merged along the channel, activation , Represents element-wise multiplication.
[0051] Example 5 The present invention is based on the RGB-D salient object detection method guided by texture enhancement, wherein step 4 is specifically implemented according to the following steps: Step 4.1, build dynamic token interaction module; Step 4.1 is implemented as follows: Step 4.1.1: Use dynamic token global context modeling and input high-level feature maps , flattening it into a spatial sequence , semantic tokens are generated by linear projection, spatial features are converted into learnable semantic token sequences, and global context information is captured. The formula is shown in Equation (17):
[0052] Step 4.1.2, the token sequence Input the Transformer encoder and optimize cross-region semantic associations through the multi-head self-attention mechanism. The formula is shown in Equations (18), (19), and (20):
[0053] Where, Represents a linear transformation matrix, generating queries, keys, and values. , Indicates the value used to scale the similarity to avoid gradient explosion. express activation function, Represents element-wise multiplication; In step 4.1.3, the similarity between the original features and the optimized tokens is calculated, and the spatial attention weights are generated and fused. The formula is shown in Equations (21) and (22):
[0054] Where, It is used to scale the similarity value to avoid gradient explosion, T represents the transposition operation, express activation function, represents element-wise multiplication, It means element-by-element addition; Step 4.2: construct a deformable cross-scale transformation and residual fusion module; Step 4.2 is as follows: Apply separable convolution and Sigmoid activation to generate a mask, and then weight the low- and middle-level features and the mask and concatenate them to obtain And through average pooling and channel attention optimization, finally with convolution The features are connected with residuals, and the formula is shown in (23) (24) (25):
[0055] Where, ( ) represents a depth-wise separable convolution layer with a convolution kernel of 3×3. ( )express activation function, represents element-wise multiplication, ( ) indicates that features are merged along the channel, ( ) represents a convolution layer with a convolution kernel of 3×3. ( ) represents average pooling, ( ) represents the channel attention mechanism.
[0056] Example 6 The experimental results of our texture-enhancement-guided RGB-D salient object detection method are shown in Tables 1 and 2 below. These results demonstrate the significant accuracy advantage of the network model trained using our proposed method. On five salient object detection datasets, NJU2K, NLPR, DUT-RGBD, STERE, and SIP, the model's overall performance surpasses existing methods. The best results are in bold, while data marked with a "-" indicates unpublished data.
[0057] Table 1
[0058] Table 2
[0059] This invention addresses the key issues in the existing technology, such as the degradation of heterogeneous modal features, cross-layer propagation of depth noise, and multi-scale semantic mismatch. It innovatively utilizes high-frequency texture prior constraints to achieve noise suppression of depth features, establishes cross-modal semantic associations through a dynamic interaction mechanism of channel-space collaboration, and adopts deformable cross-scale transformation technology to achieve progressive calibration of multi-level features. In input scenarios such as multi-target and low-quality depth, this method exhibits significant advantages in boundary integrity and noise suppression.
Claims
1. An RGB-D salient object detection method based on texture enhancement guidance, characterized by: Please follow the steps below to implement: Step 1: Build the dataset and encoder; Step 2: construct a texture enhancement module; Step 3: construct a dual-path adaptive interaction module; Step 4: Build a dynamic decoding module.
2. The texture enhancement-guided RGB-D salient object detection method according to claim 1, characterized in that: The step 1 is specifically implemented according to the following steps: Step 1.1, collect salient object dataset; Step 1.2, classify and organize the data set collected in step 1.1; Step 1.3: Divide all dataset files into training set, validation set, and test set in proportion; Step 1.4: Use PVTv2 as the backbone for feature extraction, load pre-trained weights, and build the encoder. In step 1.5, the resolution of the RGB image and depth image is adjusted to 384×384 and input into the encoder respectively.
3. The texture enhancement-guided RGB-D salient object detection method according to claim 1, characterized in that: The step 2 is specifically implemented according to the following steps: Step 2.1, build a high-frequency feature extraction module; Step 2.2, build the deep confidence module; Step 2.3: Build a cross-modal gating fusion module.
4. The texture enhancement-guided RGB-D salient object detection method according to claim 1, characterized in that: The step 2.1 is specifically implemented according to the following steps: Step 2.1.1, construct RGB multi-scale residual high-frequency feature extraction branch, the input feature is , then enter two parallel branches. Branch 1 uses a 3×3 convolution layer to capture local high-frequency texture features. The convolution kernel size is 3×3, the padding is 1, and the number of output channels is 1 / 2 of the target channel number. Branch 2 uses a 5×5 convolution layer to capture large-scale high-frequency structural features. The convolution kernel size is 5×5, the padding is 2, and the number of output channels is 1 / 2 of the target channel number. The outputs of the two branches are spliced along the channel dimension, and after batch normalization and ReLU activation function, they are fused with the input residual connection. The formula is shown in formula (1) (2): Where, The convolution kernel is The convolutional layer, Indicates that features are merged along the channel, express activation function, represents batch normalization, It means element-by-element addition; Step 2.1.2: Construct a dynamic frequency-channel attention joint mechanism to combine the features obtained in step 2.1.1 As input, the high-frequency components are separated by residual calculation, and the spatial dimension is compressed by global average pooling to generate channel-level weights. The formula is shown in Equation (3) and (4): Where, represents 3×3 local average pooling, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, Represents the Sigmoid activation function; Step 2.1.3, dynamic feature fusion, input the high-frequency information obtained in step 2.1.2 and the channel machine weights normalized by softmax Perform cross-dimensional splicing, and finally generate the final weight matrix, and weighted optimize the input features. The formula is shown in formula (5): Where, Indicates that features are merged along the channel, represents a convolution layer with a convolution kernel of 1×1. represents the Sigmoid activation function, express .
5. The texture enhancement-guided RGB-D salient object detection method according to claim 3, characterized in that: The step 2.2 is specifically implemented as follows: Step 2.2.1, build a depth confidence predictor and input depth features , through 3×3 convolutional layers, The activation function is combined with 1×1 convolution to extract important features. Finally, the Sigmoid function generates a deep feature confidence map to suppress interference in the noise area. The formula is shown in formula (6): Where, represents a convolution layer with a convolution kernel of 1×1. express Activation function, output weight matrix ; Step 2.2.2, use the data generated in step 2.2.1 Deep features Perform confidence weighting processing, and the formula is shown in formula (7): Where, express .
6. The texture enhancement-guided RGB-D salient object detection method according to claim 3, characterized in that: The step 2.3 is specifically implemented as follows: Step 2.3.1, build a cross-modal attention module to take the weighted deep features generated in step 2.2.2 High-frequency features of RGB images The channel dimension is spliced, and the spliced fusion features are compressed in the spatial dimension through global average pooling, and then the feature interaction and dimension adjustment are realized through the fully connected layer. Finally, the Sigmoid activation function is used to decouple the output into two dynamically adjusted fusion coefficients, which are used to guide the adaptive fusion of the features of the depth mode and the RGB mode respectively. The formula is shown in Equation (8) and (9): Where, Indicates that features are merged along the channel, represents global average pooling, represents the fully connected layer, activation , represents the feature splitting operation, , are the channel attention weights of depth and RGB features respectively; Step 2.3.2, construct the residual enhancement fusion, using the attention weights generated in step 2.3.1 Weighted fusion of two modal features, the weighted depth features , RGB high-frequency features Compared with the original depth feature Add together and output the enhanced depth representation, which is expressed as shown in formula (10): Where, express .
7. The texture enhancement-guided RGB-D salient object detection method according to claim 1, characterized in that: The step 3 is specifically implemented as follows: Step 3.1, construct the residual enhancement feature transformation module; For RGB features and deep features Multi-scale nonlinear transformations are performed separately. First, a deep convolution with a kernel size of 3×3 is used to extract local spatial features. The number of output channels is kept at C. Then, a 1×1 convolution is used to expand the number of channels to 2C. The activation function uses GELU to enhance the nonlinear expression ability. Finally, a second 1×1 convolution is used to restore the number of channels to C. The residual is connected with the original input to retain the original feature information while enhancing the modal specificity. The formula is shown in Equations (11) and (12): Where, represents a depth convolution layer with a convolution kernel of 3×3. represents a convolution layer with a convolution kernel of 1×1. express activation function, It means element-by-element addition; Step 3.2, construct a dynamic gated multimodal fusion module; Step 3.3, build a dual attention collaborative optimization module.
8. The texture enhancement-guided RGB-D salient object detection method according to claim 7, characterized in that: The step 3.2 is specifically implemented as follows: Step 3.2.1, the enhanced RGB features and deep features Splicing along the channel dimension to generate joint features ; Step 3.2.2: Construct dynamic weights, compress the spatial dimension to 1×1 through global average pooling, retain channel-level statistical information, and then learn the inter-modal weight distribution through the fully connected layer to apply Function, generate normalized weight distribution, and finally dynamically superimpose the two modal features according to the weight. The formula is shown in formula (13) (14): Where, represents global average pooling, represents the fully connected layer, express activation function, Represents element-wise multiplication.
9. The texture enhancement-guided RGB-D salient object detection method according to claim 7, characterized in that: The step 3.3 is specifically implemented as follows: Step 3.3.1, fusion features Perform global average pooling to generate a channel description vector, then generate a channel-level attention map through a fully connected layer and a Sigmoid function, and apply the attention map to the original features. The formula is shown in Equation (15): Where, represents global average pooling, represents a convolution layer with a convolution kernel of 1×1. express activation function, activation , Represents element-wise multiplication; Step 3.3.2, for the features generated in step 3.3.1 Do processing, extract the maximum and mean of the fusion features along the channel dimension, generate a dual-channel spatial description map, use 7×7 convolution to capture long-range spatial dependencies, and use The function generates spatial weights and applies the spatial attention weights to the channel-optimized features. The formula is shown in Equation (16): Where, and Respectively represent the maximum pooling and average pooling operations along the channel latitude, represents a convolution layer with a convolution kernel of 7×7. Indicates that features are merged along the channel, activation , Represents element-wise multiplication.
10. The texture enhancement-guided RGB-D salient object detection method according to claim 1, characterized in that: The step 4 is specifically implemented according to the following steps: Step 4.1, build dynamic token interaction module; Step 4.1 is implemented as follows: Step 4.1.1: Use dynamic token global context modeling and input high-level feature maps , flattening it into a spatial sequence , semantic tokens are generated by linear projection, spatial features are converted into learnable semantic token sequences, and global context information is captured. The formula is shown in Equation (17): Step 4.1.2, the token sequence Input the Transformer encoder and optimize cross-region semantic associations through the multi-head self-attention mechanism. The formula is shown in Equations (18), (19), and (20): Where, Represents a linear transformation matrix, generating queries, keys, and values. , Indicates the value used to scale the similarity to avoid gradient explosion. express activation function, Represents element-wise multiplication; In step 4.1.3, the similarity between the original features and the optimized tokens is calculated, and the spatial attention weights are generated and fused. The formula is shown in Equations (21) and (22): Where, It is used to scale the similarity value to avoid gradient explosion, T represents the transposition operation, express activation function, represents element-wise multiplication, It means element-by-element addition; Step 4.2: construct a deformable cross-scale transformation and residual fusion module; Step 4.2 is as follows: Apply separable convolution and Sigmoid activation to generate a mask, and then weight the low- and middle-level features and the mask and concatenate them to obtain And through average pooling and channel attention optimization, finally with convolution The features are connected with residuals, and the formula is shown in (23) (24) (25): Where, ( ) represents a depth-wise separable convolution layer with a convolution kernel of 3×3. ( )express activation function, represents element-wise multiplication, ( ) indicates that features are merged along the channel, ( ) represents a convolution layer with a convolution kernel of 3×3. ( ) represents average pooling, ( ) represents the channel attention mechanism.