Cross-modal adaptive feature integrated RGB-D saliency target detection method
Through the cross-modal adaptive feature integration method and combined with the polarization attention mechanism, the problem of insufficient feature fusion in RGB-D significance target detection is solved, and the detection performance and robustness are improved.
Patent Information
- Application Number
- CN202510757956.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing RGB-D significance object detection method, how to fully integrate the complementary properties of RGB and depth images to improve detection performance.
The cross-modal adaptive feature integration method is adopted, and the cross-modal feature integration module TFM and the adaptive feature fusion module AFI are constructed, combined with the polarized attention mechanism PSA, the feature fusion of RGB and depth images is achieved, and the feature expression and detection performance are enhanced.
Improve the accuracy and efficiency of RGB-D significance object detection, and enhance the robustness of complex backgrounds and lighting changes.
Smart Images

Figure CN120279283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision and image processing technologies, and particularly relates to an RGB-D saliency object detection method for cross-modal adaptive feature integration. Background Art
[0002] Saliency object detection is a fundamental but challenging task in computer vision, mainly used for automatically detecting the most distinctive regions in images, and has been widely applied in multiple fields such as video segmentation, image retrieval, video detection, visual tracking, etc. In recent years, with the development of deep learning and neural networks, significant progress has been made in RGB saliency object detection, and many effective methods have been proposed. However, due to problems such as the similarity of image textures and complex backgrounds, the detection effect achieved only through RGB images is still not satisfactory. In fact, the human eye can not only perceive information such as colors and textures in a scene, but can also capture depth information in the scene through the binocular visual system. This depth information is not affected by complex backgrounds and light intensities, and depth images are a form of expression of depth information. With the rise of depth cameras, it enables us to obtain corresponding depth images at a lower cost. Introducing it to participate in the saliency object detection task can simulate the visual perception ability of the human eye and enhance the computer's recognition of objects. Thus, it has attracted the interest of researchers in developing RGB-D SOD methods. A large number of studies have shown that the task of saliency object detection using RGB-D images can achieve better detection effects in many challenging scenarios. However, how to fully integrate the complementary attributes between the two modalities and make full use of their respective advantages to fuse RGB features and depth features remains an issue that needs attention.
[0003] As a pixel-level prediction task, most deep learning-based RGB-D saliency object detection methods adopt an encoder-decoder network architecture. Among them, the encoder part mostly uses a pre-trained convolutional neural network model to map the input image to the hidden layer to extract multi-level feature representations. Then, the decoder is used to restore the image resolution and generate the saliency detection result. Currently, according to the different positions of the RGB image and the depth image fused in the network, the fusion methods can be divided into three types: early fusion, mid-level fusion, and late fusion. Specifically, early fusion is performed at the input stage. Generally, the RGB image and the depth image are simply concatenated at the channel level and directly formed into a four-channel image as the input, and then the subsequent encoder and decoder operations are completed. However, this fusion method obviously does not consider the inherent differences between the RGB image and the depth image, which may lead to inaccurate detection results. Late fusion means using two independent network models to detect the saliency maps of the RGB image and the depth image respectively, and then simply multiplying or adding the two saliency maps as the final result. This method loses a lot of deep information in the image and cannot explore the internal relationship between the RGB image and the depth image, often resulting in problems such as blurred edges or insufficient prediction in the result image. Mid-level fusion mainly explores how to fuse multi-level features to extract complementary information, so as to achieve a balance between accuracy and efficiency. However, few RGB-D saliency detection models explicitly utilize modality-specific features. Therefore, there is an urgent need for a method that can not only utilize the information specific to the two modalities but also adaptively integrate the features of the two modalities to explore cross-modal integrated features and thus improve the performance of saliency object detection. Summary of the Invention
[0004] In view of the above problems, the purpose of the present invention is to provide an RGB-D saliency object detection method with cross-modal adaptive feature integration, which effectively fuses cross-modal features, explores shared information, and utilizes modality-specific features. Effectively aggregates the cross-modal complementarity of RGB and depth images, enhances feature information, and improves the performance of the RGB-D saliency object detection task.
[0005] An RGB-D saliency object detection method with cross-modal adaptive feature integration provided by the present invention specifically includes the following steps: Step S1: Data preparation, Obtain the RGB-D dataset for the task for training and testing. The RGB-D dataset includes the NJU2K dataset, NLPR dataset, SIP dataset, and STERE dataset. Among them, a part of the NJU2K dataset and a part of the NLPR dataset are used as the training set, and the remaining part of the NJU2K dataset, NLPR dataset, SIP dataset, and STERE dataset are jointly used as the test set; Step S2: Construct a network model. First, construct a backbone network through feature extraction, then construct a cross-modal feature integration module TFM, then construct a modality-specific decoder network, and finally construct an adaptive feature integration module AFI and calculate the loss function; Step S3: Evaluation metrics Four evaluation metrics are used to evaluate the effectiveness of the network model. The four evaluation metrics are: F-measure, Mean Absolute Error, S-measure, and E-measure.
[0006] As a preference of the present invention, the following steps are further included in step S1: Step S11: The RGB-D dataset includes RGB images, depth images, and label ground-truth maps manually annotated; Step S12: First, adjust the resolutions of the RGB images and depth images to 352×352, and then use them as the input of the backbone network.
[0007] As a preference of the present invention, the following steps are further included in step 2: Step S21: Construction of the feature extraction backbone network Use the Res2Net-50 network, a deep convolutional neural network model improved based on the residual network architecture pre-trained on the ImageNet dataset, and add a polarization attention mechanism PSA as the backbone network to extract multi-scale features of RGB images and depth images respectively, and obtain five groups of multi-level features of RGB images and depth images at different levels and , where ∈{1, 2, 3, 4, 5}, represents the number of encoder layers, and represent the RGB image and the depth map respectively. Then the number of channels of the feature in the th layer of the encoder is , .
[0008] As a preference of the present invention, the following steps are further included in step 2: Step S22: Cross-modal feature integration module TFM In the integrated encoder subnet, it consists of five cross-modal feature integration modules TFM, and integrates the features from the modality-specific network and the features output by the previous cross-modal feature integration module TFM to generate , where ∈{1, 2, 3, 4, 5}, where represents the integrated feature, represents the number of encoder layers. The cross-modal feature integration module TFM is used to learn the integrated features of RGB images and depth images. 1×1 convolution is used to reduce the number of channels for the input RGB features and depth features, respectively. The single-modal features extracted by the modality-specific encoder are and Cross-modal fusion is performed layer by layer and passed to the next layer to obtain multi-level information.
[0009] As the preferred embodiment of the present invention, step 2 also includes the following steps: Step S221: extract the single-mode features from the modality-specific encoder and Use 1×1 convolution to reduce the number of channels to 1 / 2 of the input. Use the cross-refinement enhancement strategy and use the attention mechanism to automatically select important features of the cross-refinement image. Specifically, use channel attention and spatial attention for RGB features and depth features respectively to highlight the feature response of the salient area. The specific process includes: ; ; in and They represent the input features of the RGB branch and the depth branch after 1×1 convolution, and Represent the global average pooling and global maximum pooling operations respectively, and represents the convolutional layer, represents element-wise multiplication, represents the ReLU activation function, Indicates the concatenation operation in the channel dimension. and Represents the RGB features and depth features after attention selection respectively.
[0010] As the preferred embodiment of the present invention, step 2 also includes the following steps: Step S222: By designing sub-modules DA and RA, the modal features are cross-enhanced and refined from a cross-modal perspective. Sub-module DA first uses 1×1 convolution to The number of channels of the feature is converted to , and the two feature matrices are and , similarly, Map to , Represents the number of channels, height and width of the input feature map respectively. The number of channels is set to 1 / 6 of the original image, and use the mapped result to calculate the enhanced features. and Multiply, express The transposed matrix of Function gets the attention weight matrix , the attention weight matrix The size is , corresponding to the attention weights at each position in the input feature map, and then the attention weight matrix After projection Matrix multiplication to obtain enhanced eigenvectors Finally, the enhanced feature vector It is transformed back into the same shape as the original input feature map, that is, C×H×W. The specific process includes: ; ; Submodule RA first uses 1×1 convolution to The number of channels of the feature is converted to , and the two feature matrices are and , then Map to , by calculating the input and Multiply, then apply Function gets the attention weight matrix , the attention weight matrix After projection The matrix is multiplied to obtain the enhanced feature vector, which is named , which means using RGB features to enhance the depth features. Finally, the enhanced feature vector It is transformed back into the same shape as the original input feature map, that is, C×H×W. In order to retain the information of the original modality image, the enhanced features are further combined with the original features by using residual connections. The residual connection representation process of RGB modality and depth modality is as follows: ; and They represent the results of RGB stream features and depth stream features after residual connection enhancement. First, and After 3×3 convolution, the smooth representation of the features is obtained and fused by element-by-element multiplication. Finally, the result is combined with the output of the previous layer cross-modal feature integration module TFM. After fusion, the second 3×3 convolution enhances the representation ability of the network, and the output result is Output of a cross-modal feature integration module TFM , the specific process is as follows: .
[0011] As a preference of the present invention, the following steps are further included in step 2: Step S23: Construction of a modality-specific decoder network Construct a modality-specific decoder using a U-Net network structure, jump-connect the multi-level features generated by the modality-specific encoder to the corresponding decoder levels to enhance the features of the corresponding levels, and input the input of the modality-specific decoder and as well as the features after the jump connection combination through a DFR module, where the DFR module represents the combination of feature connection operations and feature enhancement operations, extracts different-scale information through multi-scale convolution and dilated convolution and fuses them to achieve an enhanced representation of the feature map, and finally input the features output by the modality-specific decoder into the modality integration network.
[0012] As a preference of the present invention, the following steps are further included in step 2: Step S24: Adaptive Feature Integration Module AFI In the integrated decoder subnet, use the Adaptive Feature Integration Module AFI to adaptively fuse the features of the modality-specific subnet and the integrated features of the previous layer, and utilize the jump connection structure of the U-Net network structure and represent the features after the jump connection fusion of the modality-specific encoder, represent the features after the jump connection fusion of the corresponding layer of the cross-modal feature integration module TFM, where ( = 1, 2, 3, 4, 5), and make , and pass through a 3×3 convolutional layer to reduce the number of channels and obtain smooth features. In the designed attention weight layer GSA, where GSA represents a module that fuses global average pooling, a 1×1 convolutional layer, and a sigmoid function to generate weights for the RGB and depth branches and enhance their features, and use the features generated by the previous layer of the modality integration network to enhance and fuse the features of the RGB modality and the depth modality. The specific process is as follows: ; ; Among them, and represent intermediate variables, represents element-wise multiplication, represents connecting two features together, and after obtaining the fused result After that, a smoothed representation of the features is obtained through a 3×3 convolutional layer, and then combined with to obtain enhanced features with rich cross-modal information through an addition operation.
[0013] As a preference of the present invention, the following steps are further included in step 2: Step S25: Calculation of the loss function, The loss function consists of three parts, namely the structural similarity loss function SSIMLoss, the binary cross-entropy loss function BCELoss, and the intersection over union loss function IOULoss, which are respectively given weights of 0.8, 1, and 1 and combined, denoted as SL; and respectively represent the saliency result maps of the RGB branch and the depth branch, and GT respectively represent the result generated by the modality integration network and the label of the saliency image; the overall loss function is expressed as: .
[0014] 1. The present invention introduces a polarization attention mechanism (PSA) to enhance feature expression and adopts a mid-term fusion method to achieve multi-level cross-modal interaction. A two-stream interaction network including a cross-modal feature integration module (TFM) and an adaptive feature fusion module (AFI) is constructed. The cross-modal feature integration module TFM is applied in the encoder. Through the DA, RA sub-modules and the cross-refinement enhancement strategy, it effectively reduces the redundancy of single-modal features, aggregates cross-modal complementary information, and enhances the feature expression ability. The adaptive feature fusion module AFI is applied in the decoder. It uses the U-Net skip connection structure to obtain the features of the cross-modal feature integration module TFM in the encoder, adaptively calculates the weights of different modalities through the designed attention weight layer GSA and fuses them, and fuses features of different scales into a full-scale feature map, improving the model's perception ability of multi-scales. It fully demonstrates the performance in the RGB-D saliency object detection task and provides a more effective solution for the research and practical applications in this field. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By referring to the following description in conjunction with the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more obvious and easier to understand. In the drawings: Figure 1 is the overall network structure diagram of the present invention; Figure 2 is the structural diagram of the cross-modal feature integration module TFM of the present invention; Figure 3 is the structural diagram of the DA sub-module of the present invention; Figure 4This is the structural diagram of the Adaptive Feature Integration module AFI in the present invention. Detailed implementation manner
[0016] Refer to Figures 1-4 In this embodiment, a cross-modal adaptive feature integration RGB-D saliency object detection method is provided, which specifically includes the following steps: Step S1: Data preparation, Obtain the RGB-D data set of the task for training and testing. The RGB-D data set includes the NJU2K data set, the NLPR data set, the SIP data set, and the STERE data set (all of the above data sets are publicly available data sets). Among them, a part of the NJU2K data set and a part of the NLPR data set are used as the training set, and the remaining part of the NJU2K data set, the NLPR data set, the SIP data set, and the STERE data set are jointly used as the test set; Step S11: The RGB-D data set includes RGB images, depth images, and manually annotated label ground truth maps; among them, for each sample in the data set, it includes an RGB image, the corresponding depth image, and the corresponding manually annotated ground truth map; Step S12: First, adjust the resolution of the RGB image and the depth image to 352×352, and then use them as the input of the backbone network. Among them, when preparing the data, it is necessary to ensure that the formats of these data are unified, and the RGB image, the depth image, and the ground truth map are in one-to-one correspondence, which is convenient for subsequent data reading and processing.
[0017] Step S2: Build a network model. First, build a backbone network through feature extraction, then build a cross-modal feature integration module TFM, then build a modality-specific decoder network, and finally build an adaptive feature integration module AFI and calculate the loss function; Step S21: Build the backbone network for feature extraction, Refer to Figure 1 As shown, use the Res2Net-50 network, a deep convolutional neural network model improved based on the Residual Network (ResNet) architecture pre-trained on the ImageNet data set in the article "ImageNet: A Large-Scale Hierarchical Image Database" published by Deng et al. in CVPR2009, and add a Polarized Self-Attention (PSA) mechanism as the backbone network to extract multi-scale features of the RGB image and the depth image respectively, and obtain five groups of multi-level features of the RGB image and the depth image at different levels and where ∈ {1, 2, 3, 4, 5}, represents the number of encoder layers, and represent the RGB image and the depth image respectively. Then, the number of channels of the features in the -th layer of the encoder is , .
[0018] In a specific implementation, the RGB image is input into the modality-specific encoder for processing the RGB image. The Res2Net-50 network in this modality-specific encoder extracts features according to its pre-trained parameters and structure. At the same time, the polarization attention mechanism PSA enhances the extracted features. For the depth image , it is also input into the corresponding modality-specific encoder for the same operation; the number of channels of each layer of features is fixed, .
[0019] Step S22: Cross-modal feature integration module TFM (Cross modal feature integration module), As shown in Figure 2 , in the integrated encoder subnet, it consists of five cross-modal feature integration modules TFM, which integrate the features from the modality-specific network and the features output by the previous layer of the cross-modal feature integration module TFM to generate , where ∈ {1, 2, 3, 4, 5}, and where represents the integrated feature, represents the number of encoder layers. The cross-modal feature integration module TFM is used to learn the integrated features of the RGB image and the depth image. The 1×1 convolution is used to reduce the number of channels for the input RGB features and depth features respectively. The cross-refinement enhancement strategy is adopted, and the channel attention and spatial attention are used to obtain the attention weight matrix to enhance the feature response of the salient regions, and the residual connection is used to retain the original modality image information. Finally, the enhanced features are fused and passed to the next layer. Among them, the single-modal features and extracted by the modality-specific encoder are fused cross-modally layer by layer and passed to the next layer to obtain multi-level information.
[0020] Step S221: Respectively, for the single-modal features and extracted by the modality-specific encoderReduce the number of channels to 1 / 2 of the input using 1×1 convolutions, and use a cross-refinement enhancement strategy. Utilize the attention mechanism to automatically select important features for cross-refinement of the image. Specifically, channel attention and spatial attention are respectively used for RGB features and depth features to highlight the feature responses of the salient regions. The specific process includes: ; ; Among them and respectively represent the input features of the RGB branch and the depth branch after 1×1 convolutions, and respectively represent global average pooling and global max pooling operations, and represent convolutional layers, represents element-wise multiplication, represents the ReLU activation function, represents concatenation operation in the channel dimension, and represent the RGB features and depth features respectively after attention selection (i.e., and ).
[0021] Refer to Figure 3 as shown. Step S222: Cross-enhance and refine the modal features from a cross-modal perspective by designing sub-module DA and sub-module RA. Taking sub-module DA as an example, since the depth image provides spatial and position information for the RGB image, the features of the depth image are used to generate a weight to multiply with the RGB image to enhance the features of the RGB image. Specifically, sub-module DA first uses 1×1 convolution to convert the number of channels of the features to and to obtain two feature matrices respectively. Similarly, map to . respectively represent the number of channels, height, and width of the input feature map. The purpose is to use convolution operations to generate different feature matrices for facilitating the subsequent generation of the attention weights of the depth features to weight the features of the RGB image. The number of channels is set to 1 / 6 to improve the running speed of the network model. Use the mapped results to calculate the enhanced features. By calculating the multiplication of the input and . represents the transposed matrix of , and then apply the , the attention weight matrix The size is , corresponding to the attention weights at each position in the input feature map, and then the attention weight matrix After projection Matrix multiplication to obtain enhanced eigenvectors Finally, the enhanced feature vector It is transformed back into the same shape as the original input feature map, that is, C×H×W. The specific process includes: ; ; Another submodule RA and submodule DA are complementary in meaning, that is, using RGB image features to generate spatial weights to provide rich texture and color information for deep image features. Submodule RA first uses 1×1 convolution to The number of channels of the feature is converted to , and the two feature matrices are and , then Map to , by calculating the input and Multiply, then apply Function gets the attention weight matrix , the attention weight matrix After projection The matrix is multiplied to obtain the enhanced feature vector, which is named , which means using RGB features to enhance the depth features. Finally, the enhanced feature vector It is transformed back into the same shape as the original input feature map, that is, C×H×W. In addition, in order to retain the information of the original modality image, the enhanced features are further combined with the original features by using residual connections. The residual connection representation process of RGB modality and depth modality is as follows: ; and They represent the results of RGB stream features and depth stream features after residual connection enhancement. After obtaining the enhanced features, the most important thing is to fuse the enhanced features. Specifically, first and After 3×3 convolution, the smooth representation of the features is obtained and fused by element-by-element multiplication. Finally, the result is combined with the output of the previous layer cross-modal feature integration module TFM. After fusion, the second 3×3 convolution enhances the representation ability of the network, and the output result is Output of a cross-modal feature integration module TFM , the specific process is as follows: .
[0022] Step S23: Construction of a modality-specific decoder network Use the U-Net network structure to construct a modality-specific decoder, jump-connect the multi-level features generated by the modality-specific encoder to the corresponding decoder levels to enhance the features of the corresponding levels, and input the input of the modality-specific decoder and as well as the features after the jump-connection combination pass through the DFR module, where the DFR module represents the combination of feature connection operations and feature enhancement operations, extracts different-scale information through multi-scale convolution and dilated convolution and fuses it to achieve an enhanced representation of the feature map, which helps to improve the feature detection ability of the network model for different scales and complex patterns. Finally, the features output by the modality-specific decoder are input into the modality integration network.
[0023] Refer to Figure 4 shown in Step S24: Adaptive Feature Integration Module AFI (Adaptive feature fusion module) In the integrated decoder subnet, use the adaptive feature integration module AFI to adaptively fuse the features of the modality-specific subnet and the integrated features of the previous layer, and utilize the jump-connection structure of the U-Net network structure and represent the features after the jump-connection fusion of the modality-specific encoder, represent the features after the jump-connection fusion of the corresponding layer of the cross-modal feature integration module TFM, where ([[]] = 1, 2, 3, 4, 5), and pass , and through a 3×3 convolutional layer to reduce the number of channels and obtain smooth features. In the designed attention weight layer GSA, where GSA represents a module that fuses global average pooling, a 1×1 convolutional layer, and a sigmoid function to generate weights for the RGB and depth branches and enhance their features, and use the features generated by the previous layer of the modality integration network to enhance and fuse the features of the RGB modality and the depth modality. The specific process is as follows: ; ; Among them, and represent intermediate variables, represents element-wise multiplication, represents connecting two features together, and after obtaining the fused result After that, a smoothed representation of the features is obtained through a 3×3 convolutional layer, and then combined with using an addition operation to obtain enhanced features with rich cross-modal information, so as to improve the model's perception ability of multi-scales.
[0024] Step S25: Loss function calculation, The loss function consists of three parts, namely the structural similarity loss function SSIMLoss, the binary cross-entropy loss function BCELoss, and the intersection over union loss function IOULoss. Weights of 0.8, 1, and 1 are assigned to the above three loss functions respectively and combined, denoted as SL; and represent the saliency result maps of the RGB branch and the depth branch respectively, and GT represent the result generated by the modality integration network and the label of the saliency image respectively; the overall loss function is expressed as: .
[0025] In specific calculations, the corresponding loss values are calculated according to the calculation formulas of SSIMLoss, BCELoss, and IOULoss respectively, and then weighted summation is performed according to the weights to obtain the final loss function.
[0026] In the model training stage, the Adam algorithm is used as the optimizer to optimize the network; the initial learning rate is set to 1e-4, and the learning rate drops to one-tenth of the original every 60 training epochs of the model; During the training process, data augmentation is performed by means of random flipping, cropping, and rotation; specifically, for the input RGB images and depth images, horizontal flipping and vertical flipping are randomly performed, randomly cropped into a fixed size (such as 352×352), and randomly rotated by a certain angle, which can increase the robustness of the model; The resolution of the input images is uniformly adjusted to 352×352, and the model is trained for 200 epochs; in each epoch, the training set data is batched and input into the model for forward propagation and backward propagation, and the parameters of the model are updated according to the value of the loss function.
[0027] Step S3: Model testing and evaluation metrics, After determining the structure and parameter weights of the model, the RGB-D image pairs on the test set are tested; First, adjust the resolution of the RGB images and depth images in the test set to 352×352, and then input them into the trained network to obtain saliency predictions. In the network, the images are processed successively through the feature extraction network, the cross-modal feature integration module TFM, the modality-specific decoder network, the adaptive feature integration module AFI, etc., and finally a saliency map with the same size as the input image is output. The output of the branch where the modality integration network is located is used as the final predicted result map. Four commonly used evaluation metrics are used to evaluate the effectiveness of the network, namely F-measure, Mean Absolute Error, S-measure, and E-measure. The detailed definitions of these four metrics are as follows: F-measure ( ) is a widely used comprehensive evaluation metric that takes into account both the precision and recall scores, and its definition is as follows: ; where Precision and Recall represent the precision score and recall score respectively, and is set to 0.3 to emphasize precision; Mean Absolute Error ( ) represents the average error between the saliency map and the ground truth , and its definition is as follows: ; where and represent the width and height of the saliency map respectively, and the variables and represent the coordinate variables in the image coordinate system respectively, and are used to traverse each pixel point in the image; S-measure ( ) evaluates the spatial structure similarity between the saliency map and the saliency label map , takes into account the saliency of regions and boundaries, and is also known as the structured measurement, and its definition is as follows: ; where ∈[0, 1] is the balance parameter, and the default setting is 0.5; E-measure ( ) is an evaluation metric based on the enhanced alignment map between the saliency map and the saliency label map, and its definition is as follows: ; where and respectively represent the width and height of the saliency map, represents the enhanced alignment matrix.
[0028] The above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for RGB-D salient object detection with cross-modal adaptive feature integration, characterized in that It includes the following steps: Step S1: Data preparation, Obtain the RGB-D dataset for the task for training and testing. The RGB-D dataset includes the NJU2K dataset, NLPR dataset, SIP dataset, and STERE dataset. Among them, a part of the NJU2K dataset and a part of the NLPR dataset are used as the training set, and the remaining part of the NJU2K dataset, NLPR dataset, SIP dataset, and STERE dataset are jointly used as the test set; Step S2: Build a network model. First, build a backbone network through feature extraction, then build a cross-modal feature integration module TFM, then build a modality-specific decoder network, and finally build an adaptive feature fusion module AFI and calculate the loss function; Step S3: Evaluation metrics, Use four evaluation metrics to evaluate the effectiveness of the network model. The four evaluation metrics are: F-measure, Mean Absolute Error, S-measure, and E-measure.
2. The RGB-D saliency object detection method for cross-modal adaptive feature integration according to claim 1, wherein In step S1, it also includes the following steps: Step S11: The RGB-D dataset includes RGB images, depth images, and manually annotated ground truth maps; Step S12: First, adjust the resolution of the RGB images and depth images to 352×352, and then use them as the input of the backbone network.
3. A cross-modal adaptive feature integration RGB-D salient object detection method according to claim 1, characterized in that, In step 2, it also includes the following steps: Step S21: Build the feature extraction backbone network, Using the Res2Net-50 network, a deep convolutional neural network model improved based on the residual network architecture pre-trained on the ImageNet dataset, and adding the polarization attention mechanism PSA as the backbone network to extract multi-scale features of RGB images and depth images respectively, obtaining five groups of multi-level features of RGB images and depth images at different levels and , where ∈﹛1,2,3,4,5﹜, represents the number of encoder layers, and represent RGB images and depth maps respectively, then the number of channels of the features of the th layer in the encoder is , .
4. A method for RGB-D salient object detection based on cross-modal adaptive feature integration according to claim 3, wherein In step 2, it also includes the following steps: Step S22: Cross-modal feature integration module TFM, In the integrated encoder subnet, it consists of five cross-modal feature integration modules TFM, which integrate the features from the modality-specific networks and the features output by the previous cross-modal feature integration module TFM to generate , where ∈{1, 2, 3, 4, 5}, where represents the integrated features, denotes the encoder layer number. The cross-modal feature integration module TFM is used to learn the integrated features of RGB images and depth images. The 1×1 convolution is used to reduce the number of channels for the input RGB features and depth features respectively. Among them, the single-modal features extracted by the modality-specific encoder and are fused cross-modally layer by layer and passed to the next layer to obtain multi-level information.
5. A cross-modal adaptive feature integration-based RGB-D salient object detection method according to claim 4, characterized in that In step 2, it also includes the following steps: Step S221: Respectively, for the features of a single modality extracted by the modality-specific encoder and Use a 1×1 convolution to reduce the number of channels to 1 / 2 of the input. Use the cross-refinement enhancement strategy, and utilize the attention mechanism to automatically select the important features of the cross-refinement image. Specifically, for RGB features and depth features, channel attention and spatial attention are respectively used to highlight the feature responses of significant regions. The specific process includes: ; ; Among them and respectively represent the input features of the RGB branch and the depth branch after 1×1 convolution, and respectively represent global average pooling and global max pooling operations, and represent convolutional layers, represents element-wise multiplication, represents the ReLU activation function, represents concatenation operation in the channel dimension, and represent the RGB features and depth features respectively after attention selection.
6. The RGB-D salient object detection method based on cross-modal adaptive feature integration according to claim 5, characterized in that, In step 2, it also includes the following steps: Step S222: Cross - enhance and refine the modal features from a cross - modal perspective by designing sub - module DA and sub - module RA. Sub - module DA first uses a 1×1 convolution to convert the number of channels of the features to and to obtain two feature matrices, respectively. Similarly, map to . represent the number of channels, height, and width of the input feature map respectively. The number of channels is set to 1 / 6 of and . Calculate the enhanced features by multiplying the input represents 's transposed matrix, and then apply the function to obtain the attention weight matrix . The size of the attention weight matrix is , corresponding to the attention weights at each position in the input feature map. Immediately afterwards, multiply the attention weight matrix with the projected matrix to obtain the enhanced feature vector . Finally, reshape the enhanced feature vector back to the same shape as the original input feature map, i.e., C×H×W. The specific process includes: ; ; The sub-module RA first uses a 1×1 convolution to convert the number of channels of the features to and , and then maps to . By calculating the product of the input and , and then applying the function to obtain the attention weight matrix . Multiply the attention weight matrix with the projected matrix to obtain an enhanced feature vector, and name the vector , indicating that the depth features are enhanced with RGB features. Finally, reshape the enhanced feature vector back to the same shape as the original input feature map, i.e., C×H×W. To retain the information of the original modality image, the enhanced features are further combined with the original features by using residual connections. The process of residual connections for the RGB modality and the depth modality is as follows: ; and respectively represent the results after the RGB stream feature and the depth stream feature are enhanced by residual connection. First, and are convolved through 3×3 convolution to obtain a smooth representation of the feature and fused by element-wise multiplication. Finally, the result is integrated with the output of the previous cross-modal feature integration module TFM. After fusion, the representational ability of the network is enhanced through a second 3×3 convolution, and the output result is the output of the -th cross-modal feature integration module TFM , and the specific process is as follows: 。 7. A cross-modal adaptive feature integration-based RGB-D salient object detection method according to claim 6, wherein In step 2, it also includes the following steps: Step S23: Build the modality-specific decoder network, Construct a modality-specific decoder using the U-Net network structure, skip-connect the multi-level features generated by the modality-specific encoder to the corresponding decoder levels to enhance the features at the corresponding levels, and input the input of the modality-specific decoder and and the features after the skip-connection combination pass through the DFR module, where the DFR module represents the combination of feature connection operations and feature enhancement operations, extracts different-scale information through multi-scale convolution and dilated convolution and fuses them to achieve an enhanced representation of the feature map. Finally, the features output by the modality-specific decoder are input into the modality integration network.
8. A cross-modal adaptive feature integration-based RGB-D salient object detection method according to claim 7, characterized in that In step 2, it also includes the following steps: Step S24: Adaptive feature fusion module AFI In the integrated decoder subnet, the Adaptive Feature Integration (AFI) module is used to adaptively fuse the features of the modality-specific subnet and the integrated features of the previous layer. By leveraging the skip connection structure of the U-Net architecture, and represent the features after fusion through skip connections in the modality-specific encoder, represents the features after fusion through skip connections in the corresponding layer of the cross-modal feature integration module (TFM), where ([[]]END]] = 1, 2, 3, 4, 5). Then, , and undergo a 3×3 convolutional layer to reduce the number of channels and obtain smoothed features. In the designed Global Spatial Attention (GSA) layer, GSA represents a module that fuses global average pooling, a 1×1 convolutional layer, and a sigmoid function to generate weights for the RGB and depth branches and enhance their features. The features of the RGB modality and the depth modality are enhanced and fused with the features generated by the previous layer's modality integration network. The specific process is as follows: ; ; Among them, and represent intermediate variables, represents element-wise multiplication, represents concatenating two features, and after obtaining the fused result it passes through a 3×3 convolutional layer to obtain a smooth representation of the features, and then is connected with using an addition operation to obtain enhanced features with rich cross-modal information.
9. A method for RGB-D salient object detection based on cross-modal adaptive feature integration according to claim 8, characterized in that, In step 2, it also includes the following steps: Step S25: Calculate the loss function, The loss function consists of three parts, namely the structural similarity loss function SSIMLoss, the binary cross-entropy loss function BCELoss, and the intersection over union loss function IOULoss. They are combined with weights of 0.8, 1, and 1 respectively, and denoted as SL. and respectively represent the saliency result maps of the RGB branch and the depth branch. and GT respectively represent the result generated by the modality integration network and the label of the saliency image; the overall loss function is expressed as: 。
Citation Information
Patent Citations
RGB-D saliency target detection method based on depth quality weighting
CN116310396A
Cited By
Cross-modal unsupervised PCB welding spot defect detection method
CN121190442A
A cross-modal unsupervised PCB solder joint defect detection method
CN121190442B