RGB-D-oriented fusion model and target detection method thereof

By introducing edge maps and designing multiple fusion modules, the problem that complex boundaries are easily overlooked in RGB-D object detection is solved, and more accurate and semantic information-rich detection results are achieved.

CN119963807AActive Publication Date: 2025-05-09NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510018172.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-09
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing RGB-D object detection algorithm is easily ignored under complex or detailed boundaries, resulting in significant target boundaries blurring, thereby reducing detection accuracy.

Method used

The generated edge map is used as additional input features, combined with the multi-scale perception fusion module and the global fusion module, efficient fusion of RGB features and depth features is carried out, and edge enhancement fusion module is introduced during the decoding process to enhance edge details and global semantic information.

Benefits of technology

It effectively alleviates the problem of detection performance degradation due to blurred target boundaries, improves the accuracy of significance detection and the richness of semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963807A_ABST
    Figure CN119963807A_ABST
Patent Text Reader

Abstract

The invention provides an RGB-D-oriented fusion model and a target detection method thereof.The method comprises the steps that firstly, an edge image of a truth value image is generated through a Canny edge detection algorithm, the edge image serves as an additional input feature of a backbone network to enrich feature information and assist in deep learning of the model, then, multi-level feature extraction is conducted on an RGB image through MobileNetV2, and a target detection result is obtained; the depth image and the edge image are subjected to multi-layer feature extraction through an encoder based on a residual network ResNet-152, edge feature information is integrated into RGB image features extracted by a MobileNetV2 encoder, a multi-scale perception fusion module and a global fusion module are utilized, and in combination with complementary semantic information of the RGB features and the depth features, the depth image and the edge image are obtained; and finally, decoding a plurality of obtained fusion features, inputting the fusion features into an edge reinforcement fusion module in a decoding process layer by layer, monitoring fusion of RGB flow and depth flow detection results through designed network loss, and outputting a final saliency detection result. Therefore, the detection effect under the conditions of complex background and fuzzy salient target boundary is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer image processing technology, and specifically, is a fusion model for RGB-D and a target detection method thereof. Background Art

[0002] The continuous progress in the field of computer vision has made salient object detection an important focus of current research. Most of the current conventional salient object detection algorithms are oriented to data collected by ordinary optical cameras and are not very robust. In the case of complex backgrounds and blurred boundaries of salient objects, complex or delicate boundaries are easily ignored, and the generated saliency map usually has blurred edges. The depth data collected by the depth camera can effectively alleviate the above problems, extract effective information from the depth image information to fill in the unclear information in the RGB image, and improve the detection accuracy. The basic goal of salient object detection is to automatically identify and segment the areas or targets that the human eye is most interested in in a given scene. It has been widely used in many computer vision tasks, such as image retrieval, visual tracking, action recognition, and automatic driving.

[0003] As time goes by, many salient object detection algorithms based on RGB-D have been proposed. For example, traditional image algorithms use different types of manual features such as color cues, boundary cues, texture cues, global contrast and spatial priors to extract useful feature information from RGB images and depth images, and use this information to predict salient maps. For example, salient prediction networks and boundary preservation networks that can focus on the areas and boundaries of salient objects at the same time. With the rapid development of deep learning, compared with the limitations of manually designed features in information expression, the feature information extracted by convolutional neural networks has rich semantic information and powerful representation capabilities. Most RGB-D methods focus on exploring the fusion of RGB modal features and deep modal features. Fusion strategies are divided into three categories: early fusion, multi-scale fusion and late fusion. Early fusion directly concatenates RGB features with depth features to add an input channel into the network; while late fusion The fusion uses two independent network streams to extract high-level features of RGB and depth data respectively. These features are connected and then used to generate the final saliency prediction. Multi-scale fusion is proposed to effectively explore the correlation between RGB images and depth maps. For example, a multi-scale, multi-path fusion network integrating RGB images and depth maps with a cross-modal interaction module (MMCI) is used to explore the complementarity between low-level and high-level representations. The BiANet network uses a multi-scale bilateral attention module (MBAM) to capture better global information from multiple layers. The BBS-Net network uses a forked backbone strategy (BBS) and develops a deep enhancement module (DEM) to explore the information part of the depth map from the spatial and channel views. However, the above salient object detection algorithms do not solve the problem that the salient object boundary is blurred, resulting in a decrease in accuracy when complex or delicate boundaries are easily ignored. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the fusion of RGB image and depth image features is difficult in the existing target detection algorithm, the unreasonable design of the fusion mechanism may lead to the decline of detection performance, and the phenomenon that the detection effect is poor due to the blurred target boundary in the significant target detection. Specifically, a fusion model for RGB-D and its target detection method are used. The algorithm further enriches the expression of edge information by introducing the generated edge map, and effectively alleviates the problem of decreased detection performance due to the blurred target boundary. At the same time, three fusion modules are designed to efficiently fuse the complementary semantic information between RGB features and depth features at different feature levels, thereby enhancing the collaborative expression ability of cross-modal information and ensuring the full integration of RGB and depth modal features.

[0005] The specific technical solution adopted by the present invention is as follows:

[0006] A fusion model for RGB-D and a target detection method thereof, comprising the following steps:

[0007] (1) The Canny edge detection algorithm is used to generate the edge map of the true value image, which is used as an additional input feature of the backbone network to enrich the feature information and assist the deep learning of the model.

[0008] (2) MobileNetV2 is used to extract multi-level features from RGB images, and an encoder based on the residual network ResNet-152 is used to extract multi-level features from depth images, thereby analyzing and capturing the key features of RGB images and depth images at different levels. At the same time, the edge map is extracted through the ResNet-152 encoder, and the obtained edge feature information is integrated into the RGB image features extracted by the MobileNetV2 encoder, thereby enhancing the feature representation capability of the RGB image.

[0009] (3) Using the multi-scale perception fusion module (MSP-Fusion) and the global fusion module (Global-Fusion), the complementary semantic information of RGB features and depth features is combined at the low-level and high-level feature levels to complete the hierarchical fusion of cross-modal features.

[0010] (4) After decoding, the multiple fusion features obtained in step 3 are input layer by layer into the edge-refinement fusion module in the decoding process. This module effectively enhances the detail information of the edge of the salient object and the overall global semantic information by combining the RGB image features with the low-level and high-level RGB-D fusion features. Through the careful fusion of features at different levels, the edge-refinement module can not only improve the detail perception ability of the edge area, but also optimize the global semantic expression, thereby producing more accurate and semantically rich saliency detection results.

[0011] (5) Design image fusion loss and network overall loss to achieve the fusion of RGB stream and depth stream detection results and supervised learning of the network, and output the final saliency detection results.

[0012] As a further improvement of the present invention, the step 1) first performs coloring on the truth map. Since the truth map is a structure obtained by pixel classification, it is necessary to perform pixel coloring on it for further processing. Then, the edge of the colored true label map is extracted using the classic Canny edge detection algorithm. The Canny algorithm can accurately extract the target edge and generate a high-quality salient target edge map through multi-level gradient detection and non-maximum suppression. In this process, the key edge information of the salient target is effectively extracted, which not only enhances the boundary clarity of the feature expression, but also significantly enriches the feature information of the model. The introduction of the edge map provides the model with an accurate description of the boundary of the salient target, which can serve as additional guidance information in multimodal fusion.

[0013] As a further improvement of the present invention, in step 2), MobileNetV2 downsamples the RGB image layer by layer to extract multi-level feature information. The feature extraction of the depth image and the edge image is completed by the ResNet-152 encoder, and the residual connection structure of the network alleviates the gradient vanishing problem while extracting deep features. The edge map features are embedded in the RGB features of the corresponding scale after downsampling at each layer, and feature fusion is achieved through addition operations, which improves the diversity of features, supports model learning, and strengthens the ability to extract edge information.

[0014] As a further improvement of the present invention, the RGB image in step 3) carries rich scene and detail information, while the depth image provides the distance information from each pixel to the camera and contains more spatial structure information. Since the low-level feature map contains rich detail information, we introduced the MSPF module in the fusion process. By adding the RGB and depth features, the module generates preliminary fusion features. Next, the number of channels is adjusted by 1×1 convolution, and the horizontal and vertical features of the image are extracted by 1×3 and 3×1 convolution respectively, so as to capture spatial information in different directions. Next, the extracted features are deeply fused by 1×1 and 3×3 convolution operations to enhance the expressiveness of the features. Then, the fused feature map is combined with the original RGB and depth features to form a richer feature representation.

[0015] On this basis, the concatenated feature maps are processed by deep separable convolution (DW), batch normalization (BN) and Swish activation function to further extract deep features. Subsequently, the feature maps are reduced in dimension and nonlinearly mapped by pointwise convolution (PW), BN and Swish activation function to optimize their representation capabilities. In order to capture a wider range of spatial feature associations, the feature maps are further processed by the spatial attention mechanism (SA) to enhance the feature map's perception of global information and promote the effective fusion of RGB and deep features. In order to enhance the feature expression capability, the module designs three independent branches: the main branch extracts features through deep DW-BN, and the shortcut branch is processed through BN and PW-BN. Finally, the feature maps of the three branches are fused by element-by-element addition, and the GELU activation function is applied to perform nonlinear transformation to generate the final comprehensive feature map. The calculation formula is:

[0016]

[0017] This fusion process effectively combines the detail information of the RGB image and the spatial information of the depth image to generate a richer and more comprehensive feature map.

[0018] When processing high-level feature maps, they contain more semantic information and have a lower data volume, so the GF module is used to efficiently fuse these features and make better use of semantic information. First, 1×1 convolution operations are performed on RGB and depth features respectively, and then 3×3 convolution and Sigmoid activation functions are applied to extract deeper semantic information. Then, cross-fusion is performed by element-by-element multiplication to fuse RGB features with depth features. The results are added and concatenated to form an initial fused feature map. The initial fused feature map is concatenated with the previously extracted deep semantic features to further enrich the feature representation. Subsequently, feature extraction is performed through DW-BN and Swish activation functions, and the features are further optimized through PW-BN and Swish. On this basis, SA is used to enhance the spatial correlation between RGB and depth features to promote more effective fusion. Finally, the feature map processed by SA is added to the original RGB feature, and nonlinear mapping is performed through the GELU activation function to obtain the final fused feature map, which provides rich representation for the ERF module. Its calculation process can be expressed as:

[0019]

[0020] Through this fusion strategy, GF can effectively integrate the semantic information of RGB and depth images, and at the same time, enhance the ability to capture global features with the help of the attention mechanism, generating a feature expression that integrates semantic information with rich global information.

[0021] As a further improvement of the present invention, in step 4), during the upsampling decoding of the RGB image and the depth image, the ERF module is interspersed layer by layer. This module is intended to better combine the RGB image with the low-level and high-level RGB-D fusion features, thereby improving the accuracy of saliency detection. The low-level detail information extracted by MSPF is fused in the ERF modules of the first two layers to enhance the detail perception ability; the high-level abstract information extracted by GF is fused in the ERF modules of the last two layers to provide more globally perceived semantic information.

[0022] The fusion information extracted by MSPF or GF is mainly used to fill the edge information lost in the RGB feature to enhance the expression of edge details. First, the fusion information is added element by element to the RGB feature, and then the fused features are processed by the channel attention mechanism (CA) to further extract more edge information. In this process, the RGB feature and the extracted edge information are multiplied and added element by element again to further supplement the details and edge information that may be lost in the RGB feature.

[0023] Subsequently, 1×1 convolution is used to adjust the channel dimension of the feature map to prepare for subsequent splicing and fusion operations, thereby obtaining a preliminary fused feature representation. On this basis, in order to further enhance the semantic information of the deep features, the spatial attention mechanism (SA) is used to process the deep feature map to strengthen the global spatial correlation of the deep features. Subsequently, two convolution operations are applied to extract the multi-scale information of the deep features to capture richer semantic features.

[0024] Finally, the RGB feature map and the depth feature map are multiplied element by element and spliced ​​in the channel dimension to further fuse the feature information of the two. After the DW-BN-Swish operation, the final fused feature map is obtained. In order to obtain the final segmentation result map, the number of channels of the fused feature map is adjusted to the number of categories, and the final segmentation prediction is performed in combination with high-level abstract information. The calculation process can be expressed as:

[0025]

[0026] As a further improvement of the present invention, in step 5), in order to extract the residual features of the RGB feature map from the depth feature map for semantic segmentation, the depth stream needs to have the ability to independently implement semantic segmentation. Therefore, during the training process, it is necessary to calculate the true value map y and the depth stream output The loss between. The loss function is calculated using cross entropy, and the specific formula is as follows:

[0027]

[0028] Final classification results It also needs to be supervised by the true value graph y, and the cross entropy is still used for calculation. The formula is as follows:

[0029]

[0030] At the same time, in the edge enhancement fusion module, in order to better combine the RGB image features with the low-level and high-level RGB-D fusion features, the fusion module also needs to be supervised by the true value map y. In , it needs to be compared with the true label map y and the error is calculated through the cross entropy loss function. The specific formula is as follows:

[0031]

[0032] Therefore, the final loss function is:

[0033] L total =L d +L rgb +L n

[0034] This total loss function will combine the loss of the depth stream, the loss of the RGB stream, and the loss of the edge enhancement fusion module to optimize the model training process to improve the final semantic segmentation accuracy.

[0035] The beneficial effects of the present invention are as follows: by adopting the Canny algorithm to generate an edge map based on true value annotation, the problem of fuzzy boundaries of salient targets is effectively solved, especially when complex or delicate boundaries are easily ignored, the phenomenon of fuzzy edges of saliency maps is significantly improved. At the same time, in view of the significant difference in information between depth images and RGB images, the present invention proposes a strategy for cross-modal fusion at different feature levels, which fully combines the detail information of RGB images with the spatial information of depth images, thereby significantly improving the network's comprehensive perception ability of multimodal information. In addition, the edge enhancement fusion module introduced in the upsampling stage can make full use of the fusion information of different feature levels in the early stage to achieve high-precision restoration of the details of the original image. Finally, the present invention designs a joint optimization loss function, which can not only provide strong supervision for the fusion module, but also globally constrain the training of the overall network. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is the overall network structure diagram of the present invention.

[0037] Figure 2 This is the network structure diagram of the multi-scale perception fusion module (MSPF) in the present invention.

[0038] Figure 3 This is the network structure diagram of the global fusion module (GF) in the present invention.

[0039] Figure 4 This is the network structure diagram of the edge enhanced fusion module (ERF). DETAILED DESCRIPTION

[0040] The overall network structure of the model is as follows Figure 1 As shown in the figure, a fusion model and target detection method for RGB-D are presented. Figure 1 The specific steps are as follows:

[0041] In step 1, the true value map is first colored. Since the true value map is a structure obtained by pixel classification, it needs to be colored for further processing. Then, the classic Canny edge detection algorithm is used to extract the edge of the colored true value map. The Canny algorithm can accurately extract the edge of the target and generate a high-quality salient target edge map through multi-level gradient detection and non-maximum suppression. In this process, the key edge information of the salient target is effectively extracted, which not only strengthens the boundary clarity of the feature expression, but also significantly enriches the feature information of the model. The introduction of the edge map provides the model with an accurate description of the boundary of the salient target, which can serve as additional guidance information in multimodal fusion.

[0042] In step 2, MobileNetV2 is used to perform multi-level feature extraction on the RGB image, and the encoder based on the residual network ResNet-152 is used to perform multi-level feature extraction on the depth image, so as to analyze and capture the key features of the RGB image and the depth image at different levels. At the same time, the edge map is extracted by the ResNet-152 encoder, and the obtained edge feature information is integrated into the RGB image features extracted by the MobileNetV2 encoder, thereby enhancing the feature representation ability of the RGB image, such as Figure 1 as shown.

[0043] In step 3, the MSPF module and GF module are used in the last four convolutional layers in feature extraction. The MSPF module is used to process low-level feature maps, which contain rich detail information and are an important source of target edge and texture features. In order to extract this information more effectively, the module adds the RGB and depth feature information to obtain preliminary fusion features. Then, the fusion feature is concatenated with the original RGB and depth features, followed by depth separable convolution (DW), batch normalization (BN) and Swish activation function processing, and point-by-point convolution (PW), BN and S are performed on the processed feature maps. The wish activation function is used to reduce the dimension of the feature map. In order to capture a wider range of spatial feature associations, the feature map is processed by the spatial attention mechanism (SA). To improve the feature expression ability, the module is designed with three branches: the main branch extracts features through DW and BN, and the shortcut branch is processed through BN and PW-BN. Finally, the feature maps of the three branches are fused by element-by-element summation, and the final feature map is obtained through the GELU activation function. While maintaining the low-level feature details, this process effectively combines the detail information of the RGB image with the spatial features of the depth image, generating a richer feature map, such as Figure 2 as shown.

[0044] The GF module is mainly used to process high-level feature maps. High-level feature maps usually contain rich semantic information and have low feature dimensions, which are suitable for cross-modal feature fusion operations. First, the RGB and depth features are extracted separately and cross-fused by multiplication, that is, the RGB features are multiplied element by element with the depth features, and the obtained feature maps are combined by addition and splicing operations to obtain a preliminary comprehensive feature representation. Then, the features are extracted by DW-BN and Swish activation functions, and then the features are further optimized by PW-BN and Swish to effectively extract the global spatial information in the high-level features, thereby improving the model's expression ability and segmentation accuracy for high-level semantic features. On this basis, SA is used to enhance the spatial correlation between RGB and depth features to promote more effective fusion. Finally, the feature map processed by SA is added to the original RGB feature, and nonlinear mapping is performed through the GELU activation function to obtain the final fused feature map, which provides rich representation for the ERF module, such as Figure 3 as shown.

[0045] In step 4, during the upsampling decoding of RGB images and depth images, ERF modules are interspersed layer by layer. This module aims to better combine RGB images with low-level and high-level RGB-D fusion features to improve the accuracy of saliency detection. The low-level detail information extracted by MSPF is fused in the ERF modules of the first two layers to enhance the detail perception ability; the high-level abstract information extracted by GF is fused in the ERF modules of the last two layers to provide more globally perceived semantic information.

[0046] After the extracted fusion features are spliced ​​with the RGB feature information, the channel attention mechanism (CA) is used to further enhance the feature expression ability. The feature representation is further enriched by performing element-by-element multiplication operations on the RGB features and the fusion features and adding them together. Next, a 1×1 convolution is used to adjust the number of channels in preparation for the subsequent splicing operation, and finally the integrated feature map is generated. At the same time, in order to further enhance the semantic information of the deep features, SA is applied to process the deep features, and the deep feature map is extracted through two convolution operations. Finally, the RGB feature map and the deep feature map are element-by-element multiplied and spliced ​​in the channel dimension, and the final fusion feature map is obtained through the DW operation. Finally, combined with the high-level abstract information, the number of channels is adjusted to the number of classifications to obtain the final segmentation result map.

[0047] In step 5, in order to extract the residual features of the RGB feature map from the depth feature map for semantic segmentation, the depth stream needs to have the ability to independently implement semantic segmentation. Therefore, during the training process, it is necessary to calculate the true value map y and the depth stream output The loss between. The loss function is calculated using cross entropy, and the specific formula is as follows:

[0048]

[0049] Final classification results It also needs to be supervised by the true value graph y, and the cross entropy is still used for calculation. The formula is as follows:

[0050]

[0051] At the same time, in the edge enhancement fusion module, in order to better combine the RGB image features with the low-level and high-level RGB-D fusion features, the fusion module also needs to be supervised by the true value map y. In , it needs to be compared with the true label map y, and the error is calculated through the cross entropy loss function. The specific formula is as follows:

[0052]

[0053] Therefore, the final loss function is:

[0054] L total =L d +L rgb +L n

[0055] This total loss function will combine the loss of the depth stream, the loss of the RGB stream, and the loss of the edge enhancement fusion module to optimize the model training process to improve the final semantic segmentation accuracy.

Claims

1. A fusion model for RGB-D and a target detection method thereof, characterized in that: The following steps are involved: Step 1: Use the Canny edge detection algorithm to generate the edge map of the true value image, and use it as an additional input feature of the backbone network to enrich the feature information and assist the deep learning of the model. Step 2: Use MobileNetV2 to perform multi-level feature extraction on RGB images, and use the encoder based on the residual network ResNet-152 to perform multi-level feature extraction on the depth image, so as to analyze and capture the key features of RGB images and depth images at different levels. At the same time, the edge map is extracted through the ResNet-152 encoder, and the obtained edge feature information is integrated into the RGB image features extracted by the MobileNetV2 encoder, thereby enhancing the feature representation ability of the RGB image. Step 3: Using the multi-scale perception fusion module (MSP-Fusion) and the global fusion module (Global-Fusion), at the low-level and high-level feature levels, the complementary semantic information of RGB features and depth features is combined to complete the hierarchical fusion of cross-modal features. Step 4: After the multiple fusion features obtained in step 3 are decoded, they are input layer by layer into the edge-refinement fusion module in the decoding process. This module effectively enhances the detail information of the edges of salient objects and the overall global semantic information by combining RGB image features with low-level and high-level RGB-D fusion features. Step 5: Design the image fusion loss and the overall network loss to achieve the fusion of RGB stream and depth stream detection results and supervised learning of the network, and output the final saliency detection results.

2. The RGB-D fusion model and target detection method thereof according to claim 1, characterized in that: In the step 1, the true value map is firstly colored. Then, the Canny algorithm is used to perform edge detection on the edge map after coloring, so as to generate an edge map of the salient target. This processing can effectively extract the edge information of the salient target, enrich the feature information, and promote the auxiliary learning of the model in salient target detection.

3. The RGB-D fusion model and target detection method thereof according to claim 1, characterized in that: In the step 2, MobileNetV2 downsamples the RGB image layer by layer to extract multi-level feature information. The feature extraction of the depth image and the edge image is completed by the ResNet-152 encoder. The residual connection structure of the network alleviates the gradient vanishing problem while extracting deep features. The edge map features are embedded in the RGB features of the corresponding scale after downsampling at each layer, and feature fusion is achieved through addition operations, which improves the diversity of features, supports model learning, and strengthens the ability to extract edge information.

4. The RGB-D fusion model and target detection method thereof according to claim 1, characterized in that: In the step 3, since the low-level feature map contains rich detail information, we introduced the MSPF module in the fusion process. By adding the RGB and depth features, the module generates preliminary fusion features. Then, the fusion feature is spliced ​​with the original RGB and depth features, followed by depth separable convolution (DW), batch normalization (BN) and Swish activation function processing, and the processed feature map is subjected to pointwise convolution (PW), BN and Swish activation function to reduce the dimension of the feature map. In order to capture a wider range of spatial feature associations, the feature map is processed by the spatial attention mechanism (SA). In order to enhance the feature expression ability, the module designs three independent branches: the main branch extracts features through DW and BN, and the shortcut branch is processed through BN and PW-BN. Finally, the feature maps of the three branches are fused through element-by-element summation operation, and the final feature representation is obtained through the GELU activation function. This process effectively combines the detail information of the RGB image with the spatial features of the depth image to generate a richer feature map. When processing high-level feature maps, they contain more semantic information and have a lower data volume, so the GF module is used to efficiently fuse these features and make better use of semantic information. First, the RGB and depth features are extracted separately and cross-fused by multiplication, that is, the RGB features are multiplied element by element with the depth features, and the obtained feature maps are combined by addition and splicing operations to obtain a preliminary comprehensive feature representation. Then, the features are extracted through DW-BN and Swish activation functions, and then the features are further optimized through PW-BN and Swish. On this basis, SA is used to enhance the spatial correlation between RGB and depth features to promote more effective fusion. Finally, the feature map processed by SA is added to the original RGB feature, and nonlinear mapping is performed through the GELU activation function to obtain the final fused feature map, which provides rich representation for the ERF module.

5. The RGB-D fusion model and target detection method thereof according to claim 1, characterized in that: In step 4, during the upsampling decoding of the RGB image and the depth image, the ERF module is interspersed layer by layer. This module is designed to better combine the RGB image with the low-level and high-level RGB-D fusion features, thereby improving the accuracy of saliency detection. The low-level detail information extracted by MSPF is fused in the ERF modules of the first two layers to enhance the detail perception ability; the high-level abstract information extracted by GF is fused in the ERF modules of the last two layers to provide more globally perceived semantic information. After the extracted fusion features are spliced ​​with the RGB feature information, the channel attention mechanism (CA) is used to further enhance the feature expression ability. The feature representation is further enriched by performing element-by-element multiplication operations on the RGB features and the fusion features and adding them together. Next, a 1×1 convolution is used to adjust the number of channels in preparation for the subsequent splicing operation, and finally an integrated feature map is generated. At the same time, in order to further enhance the semantic information of the deep features, SA is applied to process the deep features, and the deep feature map is extracted through two convolution operations. Finally, the RGB feature map and the deep feature map are element-by-element multiplied and spliced ​​in the channel dimension, and the final fused feature map is obtained through the DW operation. Finally, combined with the high-level abstract information, the number of channels is adjusted to the number of classifications to obtain the final segmentation result map.

6. The RGB-D fusion model and target detection method thereof according to claim 1, characterized in that: In step 5, in order to extract the residual features of the RGB feature map from the depth feature map for semantic segmentation, the depth stream needs to have the ability to independently implement semantic segmentation. Therefore, during the training process, it is necessary to calculate the true value map y and the depth stream output The loss between. The loss function is calculated using cross entropy, and the specific formula is as follows: Final classification results It also needs to be supervised by the true value graph y, and the cross entropy is still used for calculation. The formula is as follows: At the same time, in the edge enhancement fusion module, in order to better combine the RGB image features with the low-level and high-level RGB-D fusion features, the fusion module also needs to be supervised by the true value map y. In , it needs to be compared with the true label map y and the error is calculated through the cross entropy loss function. The specific formula is as follows: Therefore, the final loss function is: L total =L d +L rgb +L n This total loss function will combine the loss of the depth stream, the loss of the RGB stream, and the loss of the edge enhancement fusion module to optimize the model training process to improve the final semantic segmentation accuracy.

Citation Information

Patent Citations

  • RGB-D saliency target detection method based on boundary deformable convolution guidance

    CN115830420A

  • RGB-D saliency target detection method

    CN116206133A

  • RGB-D semantic segmentation method based on feature enhancement and global feature fusion

    CN118365877A

  • RGB-D underwater saliency target detection method based on semantic guidance fusion

    CN118570623A

  • RGB-D saliency detection method based on progressive weighted decoding

    CN118587449A

Cited By

  • Rock debris image segmentation method based on multi-scale feature enhancement and edge perception gating

    CN120355926A

  • Rock chip image segmentation method based on multi-scale feature enhancement and edge-aware gating

    CN120355926B

  • RGB-D salient target detection method based on potential perception and hierarchical fusion

    CN120953577A