DCDU-Mama-based transparent object depth completion method
Through the transparent object depth completion method based on DCDU-Mamba, the problem of difficulty in obtaining transparent object depth information in the prior art is solved, and efficient capture of transparent object edge areas and accurate recovery of depth information are achieved.
Patent Information
- Application Number
- CN202510274976.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The prior art is difficult to accurately obtain the depth information of transparent objects, resulting in the failure of grabbing transparent objects.
Using the transparent object depth completion method based on DCDU-Mamba, the DCDU-Mamba network framework is constructed, including an encoder, jump connection module and decoder, and the RGB image and depth image of transparent objects are obtained, feature extraction and fusion are performed to generate transparent object depth completion image.
Effectively capture the complex details and edge features of transparent objects, significantly improving the network's ability to capture transparent objects' edge areas and restore the resolution of transparent object images.
Smart Images

Figure CN120147394A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and specifically to a method for depth completion of transparent objects based on DCDU-Mamba. Background Art
[0002] In contemporary medical and industrial production, the widespread use of transparent objects has become an undeniable trend. Robots play an important role in modern industry, and in fields such as manufacturing, healthcare, and home services, it is inevitable to grasp highly reflective or transparent objects. However, current robot grasping methods rely heavily on the depth information collected by RGB-D cameras. For non-transparent objects, this information can be obtained by depth cameras. But for transparent objects such as chemical test tubes, beakers, and glass bottles, due to their refractive and highly reflective properties, they do not conform to the geometric optical path assumptions in classical vision algorithms, which makes it complex to capture the edges and details related to transparent objects. In addition, the lack of texture features makes it challenging to determine the actual shape and position of transparent objects. Therefore, it is difficult for depth cameras to accurately measure the depth of transparent objects and obtain their coordinates in three-dimensional space, resulting in the failure of grasping transparent objects.
[0003] Regarding the problem of depth completion of transparent objects, although neural networks based on convolution or Transformer have made some progress in depth completion of transparent objects, there are still significant limitations. Convolution-based models have difficulty capturing the global correlations of long-range light propagation with local receptive fields, while although Transformer can model global dependencies, it is prone to losing high-frequency details and has high computational costs. Its linear complexity is not suitable for processing high-resolution images, which is particularly important for the task of accurately extracting edge features of transparent objects. The depth maps predicted by these methods still have inaccurate depth values and blurred contours. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for depth completion of transparent objects based on DCDU-Mamba, including the following steps:
[0005] 1) Construct a DCDU-Mamba network framework, including an encoder, a skip connection module, and a decoder;
[0006] 2) Obtain the RGB image and depth image of the transparent object;
[0007] 3) The encoder inputs the RGB image and depth image of the transparent object into two independent VMamba feature extraction networks respectively for feature extraction to obtain RGB features and depth features;
[0008] 4) The skip connection module fuses the RGB features and depth features through a dual-head cross-attention mechanism to obtain a shallow feature map;
[0009] The decoder is used to fuse the RGB features and depth features to obtain a deep feature map;
[0010] 5) Concatenate the deep feature map and the shallow feature map, and perform upsampling through the DOSA network in the decoder to generate a depth completion image of the transparent object;
[0011] 6) Restore the size of the depth completion image of the transparent object through the linear embedding layer to align the depth completion image of the transparent object with the size of the original depth image.
[0012] Furthermore, in step 2), the RGB image and the depth image of the transparent object are acquired by a depth camera, denoted as I c and D s .
[0013] Furthermore, in step 3), the two independent VMamba feature extraction networks are mirror images of each other.
[0014] Furthermore, in step 3), the VMamba feature extraction network includes a Patch Embedding module and a four-stage structure;
[0015] The four-stage structure sequentially includes a VSS module, a VSS and Patch Merging integration module I, Patch Merging II, and Patch Merging III;
[0016] The Patch Embedding module divides the RGB image or the depth image into non-overlapping image patches, and projects these image patches onto dimension C through a linear embedding layer; the size of each image patch is s×s×3; H and W are the height and width of the RGB image;
[0017] The VSS and Patch Merging integration module I, Patch Merging II, and Patch Merging III each include two VSS modules and one Patch Merging module;
[0018] The VSS module is used to extract the local feature information of the image patches to generate RGB features or depth features;
[0019] The Patch Merging module is used to perform a downsampling operation to reduce the size of the image patches and increase the number of channels.
[0020] Furthermore, in the four-stage structure of the VMamba feature extraction network, the outputs of each stage are as follows:
[0021]
[0022] Among them, F c is the feature map of the RGB image, and F s is the feature map of the sparse depth image. i ∈ {1, 2, 3, 4} respectively represents each module in the four-stage structure of VMamba.
[0023] Furthermore, in the four-stage structure, the size of the image block in each stage is where C is the dimension.
[0024] Furthermore, in step 4), the steps of fusing the RGB feature and the depth feature include:
[0025] 4.1) Given the RGB feature F c and the depth feature F s , multiply them by the weight matrices W q , W k and W v respectively to generate the query matrix (Query, Q), the key matrix (Key, K), and the value matrix (Value, V);
[0026] 4.2) Use the depth feature F s as the query source and the RGB feature F c as the key-value source, that is:
[0027]
[0028] 4.3) Construct a fusion block from the cross RGB feature to the depth feature as shown in the following formula:
[0029]
[0030] 4.4) Use the RGB feature F c as the query source and the depth feature F s as the key-value source, that is:
[0031]
[0032] 4.5) Construct a fusion block from the cross depth feature to the RGB feature as shown in the following formula:
[0033]
[0034] Among them, W q c , W k c and W k c are learnable matrices for RGB feature mapping, and W q s 、Wq s and W v s are learnable matrices for depth feature mapping, c represents the RGB feature map, s represents the depth feature map, and i represents the i-th stage among the four stages in the VMamba feature extraction network. and respectively represent cross-attention mechanism operations in different directions, and d is a scaling factor;
[0035] 4.6) Through concatenation operations and convolution operations using a 3×3 convolution kernel, fuse the cross-RGB feature to depth feature fusion block and the cross-RGB feature to depth feature fusion block together to obtain the fused shallow feature map That is:
[0036]
[0037] Furthermore, in step 5), the steps of generating the depth completion image of the transparent object include:
[0038] 5.1) Adopt the channel halving operation to divide the channels of the fused features into two equal and independent branches, denoted as the first branch and the second branch respectively;
[0039] 5.2) The first branch, through convolution operations, mixes the fusion result from the lower decoder and the shallow feature map F of the current stage f i to obtain the fused feature h;
[0040] Among them, the input H i , output h i of the i-th layer of the first branch are respectively as follows:
[0041] H i =Concat(h 0 ,h 1 ,…,h i-1 )(8)
[0042] h i =ReLU(BN(Conv(Conv(H i ,1×1,k×4),3×3,k)))(9)
[0043] Among them, k represents the number of output channels of the 1×1 convolution operation;
[0044] 5.3) The second branch, through convolution operations, mixes the fusion result from the lower decoder and the shallow feature map of the current stage to obtain the fused feature y cat , that is:
[0045] y i = Conv3×3(x i-1 , Z), i = 1, 2, ..., L (10)
[0046] y cat = Concat(y 1 , y 2 , ..., y L ) (11)
[0047] Wherein, x 0 is the input, L is the number of layers, and Z is the number of output channels of the i-th convolutional layer; y i is the fused feature corresponding to the i-th convolutional layer;
[0048] 5.4) Merge the fused feature y cat and the fused feature h through a channel concatenation operation, and shuffle the merged feature map in the channel dimension to generate a feature fusion image;
[0049] 5.5) Upsample the feature fusion image to twice the original size through a dynamic sampling layer, thereby generating a transparent object depth completion image.
[0050] Furthermore, the first branch adopts an improved DenseNet network structure, including multiple dense connection blocks and a size alignment module;
[0051] Each dense connection block contains multiple cascaded convolutional layers;
[0052] The second branch adopts an improved OSA network, including multiple convolutional layers.
[0053] Furthermore, the steps of upsampling the feature fusion image to the current scale through the dynamic sampling layer in step 5.5) include:
[0054] 5.5.1) Initialize to generate a normalized two-dimensional grid. Duplicate the second dimension of the grid g times so that each group of features in the two-dimensional grid shares the same sampling set and adapts to the grouping mechanism to obtain the original sampling grid ε, that is:
[0055]
[0056] 5.5.2) Given the upsampling scale factor s and the feature map λ of size C×H×W; map the input to an offset λ 1 through a linear layer, with a size of 2gs 2 ×H×W; g is the number of offsets;
[0057] 5.5.3) Adjust the offset λ 1 to obtain the offset λ 2 , that is:
[0058] λ 2 = 0.25λ 1 = 0.25ear(λ) (13)Lin
[0059] 5.5.4) After passing the offset λ through the pixel_shuffling function 2 reshape it to obtain an offset grid λ of size 2g × sH × s, i.e.: 3 That is:
[0060] λ 3 = Pixel_shuffling(λ 2 ) (14)
[0061] 5.5.5) Construct a sampling set η, i.e.:
[0062] η = λ 2 + ε (15)
[0063] 5.5.6) Using the bilinear interpolation method in the grid_sample function, resample the feature map to an output λ' of size C × sH × sW, as follows:
[0064] λ' = Grid_sample(η, λ) (16).
[0065] The technical effect of the present invention is beyond doubt. By integrating VMamba and U-Net, the present invention introduces the U-Mamba framework. This framework enhances the multi-scale characteristics, can effectively capture the context information related to transparent objects, and at the same time retains local details. In addition, through the selective state space mechanism adopted by VMamba, U-Mamba dynamically focuses on the key regions related to transparent objects in the image, which helps to comprehensively understand and estimate transparent objects.
[0066] The present invention integrates VMamba and U-Net, introducing a dual-head U-Mamba structure, enhancing the utilization of multi-scale features, enabling the model to effectively capture the context information related to transparent objects, and at the same time retaining local details. In addition, through the selective state space mechanism of VMamba, the dual-head U-Mamba structure can dynamically focus on the key regions related to transparent objects in the image, such as edge refraction and background distortion, while ignoring irrelevant noise.
[0067] The present invention uses a dual-head cross-attention mechanism to fuse RGB features and depth features, improving the complex details and edge features of transparent objects, further integrating and optimizing the effective regions of multi-modalities in the RGB feature map and the depth feature map, and significantly enhancing the network's ability to capture the edge regions of transparent objects.
[0068] The present invention uses a DOSA network in the decoder, combines low-level and multi-scale fusion features, enhances multi-level feature representation, and can effectively restore the resolution of transparent object images.
[0069] The present invention evaluates the proposed algorithm on the TransCG dataset and compares the test results with those of other advanced methods, verifying the effectiveness and superiority of the proposed method. Brief Description of the Drawings
[0070] Figure 1 is the overall network framework based on DCDU-Mamba;
[0071] Figure 2 is the feature extraction network based on VMamba;
[0072] Figure 3 is the feature fusion network based on the dual-head cross-attention mechanism;
[0073] Figure 4 is the upsampling network based on DOSA.
[0074] Figure 5 is the dynamic upsampling module based on DySample.
[0075] Figure 6 is the depth prediction effect of three depth completion methods in real scenes. Detailed Embodiments
[0076] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject matter scope of the present invention is limited to the following embodiments. Without departing from the above-mentioned technical idea of the present invention, various substitutions and changes made according to the common general knowledge and customary means in the art shall be included within the protection scope of the present invention.
[0077] Embodiment 1:
[0078] A method for depth completion of transparent objects based on DCDU-Mamba, comprising the following steps:
[0079] 1) Construct a DCDU-Mamba network framework, including an encoder, a skip connection module, and a decoder;
[0080] 2) Obtain the RGB image and depth image of the transparent object;
[0081] 3) The encoder inputs the RGB image and depth image of the transparent object into two independent VMamba feature extraction networks respectively for feature extraction to obtain RGB features and depth features;
[0082] 4) The skip connection module fuses the RGB features and depth features through a dual-head cross-attention mechanism to obtain a shallow feature map;
[0083] The decoder is used to fuse the RGB features and depth features to obtain a deep feature map;
[0084] 5) The deep feature map and the shallow feature map are concatenated, and upsampling is achieved through the DOSA network in the decoder to generate a transparent object depth completion image;
[0085] Specifically, the deep feature map, that is, the deep high-level fusion feature map, is Figure 1 the feature map at the bottom layer of the decoder; the shallow feature map is the feature map fused through the dual-head cross-attention mechanism in step 4).
[0086] The concatenation process here is Figure 1 the process of the dotted box in the decoder part. The deep high-level fusion feature map with a "+" corresponds to the fusion result of the lower-layer decoder;
[0087] The shallow feature map corresponds to the fusion feature at the current stage;
[0088] 6) The size of the transparent object depth completion image is restored through a linear embedding layer to align the transparent object depth completion image with the size of the original depth image.
[0089] In step 2), the RGB image and depth image of the transparent object are acquired by a depth camera, denoted as I c and D s .
[0090] In step 3), the two independent VMamba feature extraction networks are mirror images of each other.
[0091] In step 3), the VMamba feature extraction network includes a Patch Embedding module and a four-stage structure;
[0092] The four-stage structure successively includes a VSS module, an integrated module I of VSS and Patch Merging, Patch Merging II, and Patch Merging III;
[0093] The Patch Embedding module divides the RGB image or depth image into non-overlapping image patches, and projects these image patches onto dimension C through a linear embedding layer; the size of each image patch is s×s×3; H and W are the height and width of the RGB image;
[0094] The VSS and Patch Merging integration modules I, Patch Merging II, and Patch Merging III each include two VSS modules and one Patch Merging module;
[0095] The VSS module is used to extract the local feature information of the image patch, thereby generating RGB features or depth features;
[0096] The Patch Merging module is used to perform downsampling operations, reduce the size of the image patch, and increase the number of channels.
[0097] In the four-stage structure of the VMamba feature extraction network, the outputs of each stage are as follows:
[0098]
[0099] Among them, F c is the feature map of the RGB image, F s is the feature map of the sparse depth image, and i ∈ {1, 2, 3, 4} respectively represent the various modules in the four-stage structure of VMamba.
[0100] In the four-stage structure, the size of the image patch in each stage is C is the dimension.
[0101] In step 4), the steps for fusing the RGB features and depth features include:
[0102] 4.1) Given the RGB feature F c and the depth feature F s , multiply them by the weight matrices W q , W k and W v respectively to generate the query matrix (Query, Q), the key matrix (Key, K), and the value matrix (Value, V);
[0103] 4.2) Use the depth feature F s as the query source and the RGB feature F c as the key-value source, that is:
[0104]
[0105] 4.3) Construct a fusion block from the cross RGB feature to the depth feature, as shown in the following formula:
[0106]
[0107] In the formula, d is the scaling factor; the parameter
[0108] 4.4) Use the RGB feature F c as the query source and the depth feature F s as the key-value source, i.e.:
[0109]
[0110] 4.5) Construct a fusion block from cross-depth features to RGB features as shown in the following formula:
[0111]
[0112] where W q c 、W k c and W k c are learnable matrices for RGB feature mapping, W q s 、W q s and W v s are learnable matrices for depth feature mapping, c represents the RGB feature map, s represents the depth feature map, and i represents the i-th stage among the four stages of the VMamba feature extraction network. and represent cross-attention mechanism operations in different directions respectively, and d is a scaling factor;
[0113] 4.6) Through the splicing operation and the convolution operation using a 3×3 convolution kernel, fuse the fusion block from cross-RGB features to depth features and the fusion block from cross-RGB features to depth features together to obtain the fused shallow feature map i.e.:
[0114]
[0115] In step 5), the steps for generating the depth completion image of the transparent object include:
[0116] 5.1) Adopt the channel halving operation to divide the channels of the fused features into two equal and independent branches, denoted as the first branch and the second branch respectively;
[0117] 5.2) The first branch mixes the fusion result from the lower decoder and the shallow feature map of the current stage through the convolution operation to obtain the fused feature h;
[0118] where the input H i and the output h i of the i-th layer of the first branch are shown as follows respectively:
[0119] H i= Concat(h 0 , h 1 ,..., h i-1 )(8)
[0120] h i = ReLU(BN(Conv(Conv(H i , 1×1, k×4), 3×3, k)))(9)
[0121] Among them, k represents the number of output channels of the 1×1 convolution operation;
[0122] 5.3) The second branch performs a convolution operation to mix the fusion result from the lower decoder and the shallow feature map at the current stage to obtain the fusion feature y cat , that is:
[0123] y i = Conv3×3(x i-1 , Z), i = 1, 2,..., L(10)
[0124] y cat = Concat(y 1 , y 2 ,..., y L )(11)
[0125] In the formula, x 0 is the input, L is the number of layers, Z is the number of output channels of the i-th convolutional layer; y i is the fusion feature corresponding to the i-th convolutional layer;
[0126] 5.4) The fusion feature y cat and the fusion feature h are merged through a channel concatenation operation, and the merged feature map is shuffled in the channel dimension to generate a feature fusion image;
[0127] 5.5) The feature fusion image is upsampled to twice the original size through a dynamic sampling layer, thereby generating a transparent object depth completion image.
[0128] The first branch adopts an improved DenseNet network structure, including multiple dense connection blocks and a size alignment module;
[0129] Each dense connection block contains multiple cascaded convolutional layers;
[0130] The second branch adopts an improved OSA network, including multiple convolutional layers.
[0131] The steps of upsampling the feature fusion image to the current scale through the dynamic sampling layer in step 5.5) include:
[0132] 5.5.1) Initialize and generate a normalized two-dimensional grid. Duplicate the second dimension of the grid g times so that the features in each group of the two-dimensional grid share the same sampling set, adapting to the grouping mechanism, to obtain the original sampling grid ε, i.e.:
[0133]
[0134] 5.5.2) Given the upsampling scale factor s and the feature map λ of size C×H×W; map the input through a linear layer to obtain the offset λ 1 , with a size of 2gs 2 ×H×W; g is the number of offsets;
[0135] 5.5.3) Adjust the offset λ 1 to obtain the offset λ 2 , i.e.:
[0136] λ 2 = 0.25λ 1 = 0.25ear(λ) (13)Lin
[0137] 5.5.4) Reshape the offset λ 2 through the pixel_shuffling function to obtain the offset grid λ 3 of size 2g×sH×s, i.e.:
[0138] λ 3 = Pixel_shuffling(λ 2 ) (14)
[0139] 5.5.5) Construct the sampling set η, i.e.:
[0140] η = λ 2 + ε (15)
[0141] 5.5.6) Use the bilinear interpolation method in the grid_sample function to resample the feature map to the output λ' of size C×sH×sW, as follows:
[0142] λ' = Grid_sample(η,λ) (16).
[0143] Example 2:
[0144] A depth completion method for transparent objects based on DCDU-Mamba, comprising the following steps:
[0145] 1) Construct a DCDU-Mamba network framework, including an encoder, a skip connection module, and a decoder;
[0146] 2) Obtain the RGB image and depth image of the transparent object;
[0147] 3) The encoder inputs the RGB image and depth image of the transparent object into two independent VMamba feature extraction networks respectively for feature extraction, obtaining RGB features and depth features;
[0148] 4) The skip connection module fuses the RGB features and depth features through a dual-head cross-attention mechanism to obtain a shallow feature map;
[0149] The decoder is used to fuse the RGB features and depth features to obtain a deep feature map;
[0150] 5) Concatenate the deep feature map and the shallow feature map, and perform upsampling through the DOSA network in the decoder to generate a depth completion image of the transparent object;
[0151] 6) Restore the size of the depth completion image of the transparent object through a linear embedding layer to align the depth completion image of the transparent object with the size of the original depth image.
[0152] Embodiment 3:
[0153] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as that of Embodiment 2. Further, in step 2), the RGB image and depth image of the transparent object are obtained by a depth camera, and are denoted as I c and D s .
[0154] Embodiment 4:
[0155] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Embodiments 2-3. Further, in step 3), the two independent VMamba feature extraction networks are mirror images of each other.
[0156] Embodiment 5:
[0157] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Embodiments 2-4. Further, in step 3), the VMamba feature extraction network includes a Patch Embedding module and a four-stage structure;
[0158] The four-stage structure sequentially includes a VSS module, a VSS and Patch Merging integration module I, Patch Merging II, and Patch Merging III;
[0159] The Patch Embedding module divides the RGB image or depth image into non-overlapping A number of image patches are projected onto dimension C through a linear embedding layer; the size of each image patch is s×s×3;
[0160] The VSS and Patch Merging integration modules I, Patch Merging II, and Patch Merging III each include two VSS modules and one Patch Merging module;
[0161] The VSS module is used to extract the local feature information of the image patches, thereby generating RGB features or depth features;
[0162] The Patch Merging module is used to perform downsampling operations, reduce the size of the image patches, and increase the number of channels.
[0163] Example 6:
[0164] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Examples 2-5. Further, in the four-stage structure of the VMamba feature extraction network, the outputs of each stage are as follows:
[0165]
[0166] Among them, F c is the feature map of the RGB image, F s is the feature map of the sparse depth image, and i∈{1,2,3,4} respectively represent the various modules in the four-stage structure of VMamba.
[0167] Example 7:
[0168] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Examples 2-6. Further, in the four-stage structure, the size of the image patches in each stage is C is the dimension.
[0169] Further, in step 4), the steps for fusing the RGB features and the depth features include:
[0170] 4.1) Given the RGB feature F c and the depth feature F s , multiply them by the weight matrices W q , W k and W v respectively to generate the query matrix (Query, Q), the key matrix (Key, K), and the value matrix (Value, V);
[0171] 4.2) Use the depth feature F s as the query source, and the RGB feature F cAs the key-value source, i.e.:
[0172]
[0173] 4.3) Construct a fusion block from cross-RGB features to depth features, as shown in the following formula:
[0174]
[0175] 4.4) Use the RGB feature F c as the query source, and the depth feature F s as the key-value source, i.e.:
[0176]
[0177] 4.5) Construct a fusion block from cross-depth features to RGB features, as shown in the following formula:
[0178]
[0179] Among them, W q c 、W k c and W k c are learnable matrices for RGB feature mapping, W q s 、W q s and W v s are learnable matrices for depth feature mapping, c represents the RGB feature map, s represents the depth feature map, and i represents the i-th stage among the four stages of the VMamba feature extraction network. and respectively represent cross-attention mechanism operations in different directions, and d is the scaling factor;
[0180] 4.6) Through concatenation operations and convolution operations using a 3×3 convolution kernel, fuse the fusion block from cross-RGB features to depth features and the fusion block from cross-RGB features to depth features together to obtain the fused shallow feature map i.e.:
[0181]
[0182] Example 8:
[0183] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Examples 2-7. Further, in step 5), the steps for generating the depth completion image of the transparent object include:
[0184] 5.1) Adopt the channel halving operation to divide the channels of the fused feature into two equal and independent branches, denoted as the first branch and the second branch respectively;
[0185] 5.2) The first branch mixes the fused result from the lower-layer decoder and the shallow feature map of the current stage through convolution operation to obtain the fused feature h;
[0186] Among them, the input H i and output h i of the i-th layer of the first branch are respectively as follows:
[0187] H i = Concat(h 0 , h 1 ,..., h i-1 )(8)
[0188] h i = ReLU(BN(Conv(Conv(H i , 1×1, k×4), 3×3, k)))(9)
[0189] Among them, k represents the number of output channels of the 1×1 convolution operation;
[0190] 5.3) The second branch mixes the fused result from the lower-layer decoder and the shallow feature map of the current stage through convolution operation to obtain the fused feature y cat , that is:
[0191] y i = Conv3×3(x i-1 , Z), i = 1, 2,..., L(10)
[0192] y cat = Concat(y 1 , y 2 ,..., y L )(11)
[0193] In the formula, x 0 is the input, L is the number of layers, and Z is the number of output channels of the i-th convolutional layer;
[0194] 5.4) Merge the fused feature y cat and the fused feature h through channel concatenation operation, and shuffle the merged feature map in the channel dimension to generate a feature fusion image;
[0195] 5.5) Upsample the feature fusion image to twice the original size through the dynamic sampling layer to generate a transparent object depth completion image.
[0196] Example 9:
[0197] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Examples 2-8. Further, the first branch adopts an improved DenseNet network structure, including multiple densely connected blocks and a size alignment module;
[0198] Each densely connected block contains multiple cascaded convolutional layers;
[0199] The second branch adopts an improved OSA network, including multiple convolutional layers.
[0200] Example 10:
[0201] A method for depth completion of transparent objects based on DCDU-Mamba, the technical content is the same as any one of Examples 2-9. Further, step 5.5) The steps of upsampling the feature fusion image to the current scale through the dynamic sampling layer include:
[0202] 5.5.1) Initialize to generate a normalized two-dimensional grid. Duplicate the second dimension of the grid g times so that the features in each group of the two-dimensional grid share the same sampling set and adapt to the grouping mechanism to obtain the original sampling grid ε, that is:
[0203]
[0204] 5.5.2) Given the upsampling scale factor s and the feature map λ of size C×H×W; map the input to an offset λ through a linear layer 1 , with a size of 2gs 2 ×H×W; g is the number of offsets;
[0205] 5.5.3) Adjust the offset λ 1 to obtain the offset λ 2 , that is:
[0206] λ 2 = 0.25λ 1 = 0.25ear(λ) (13)Lin
[0207] 5.5.4) Reshape the offset λ 2 through the pixel_shuffling function to obtain an offset grid λ 3 of size 2g×sH×s, that is:
[0208] λ 3 = Pixel_shuffling(λ 2 ) (14)
[0209] 5.5.5) Construct the sampling set η, that is:
[0210] η = λ 2 + ε (15)
[0211] 5.5.6) Use the bilinear interpolation method in the grid_sample function to resample the feature map to the output λ' of size C × sH × sW as follows:
[0212] λ' = Grid_sample(η, λ) (16).
[0213] Example 11:
[0214] A method for depth completion of transparent objects based on DCDU-Mamba is as follows:
[0215] Due to the reflection and refraction effects of light, it is difficult for a depth camera to accurately obtain the depth information of transparent objects; at the same time, transparent objects usually do not have rich surface textures, resulting in unclear image contours based on visual feature matching. To address these problems, the present invention proposes a depth completion algorithm for transparent objects based on DCDU-Mamba. The overall network framework of DCDU-Mamba is as Figure 1 shown. The network framework mainly consists of three parts: an encoder, skip connections, and a decoder, including the following steps:
[0216] Step 101, the encoder inputs the RGB image and the depth image into two independent VMamba networks respectively for feature extraction.
[0217] Step 102, the skip connections fuse the multi-scale features of the RGB image and the depth image through a dual-head cross-attention mechanism to capture multi-level features from coarse to fine.
[0218] Step 103, splice the deep high-level fusion feature map with the shallow feature map obtained at this stage, and perform upsampling through the DOSA network in the decoder.
[0219] Step 104, restore the feature size through a linear embedding layer to ensure a consistent size alignment relationship with the original depth map, thereby forming a depth map that fuses color, texture, and depth information.
[0220] As Figure 2 shown, it is a flowchart of Step 101, including the following steps:
[0221] Step 201, the input of the feature extraction network based on VMamba includes the RGB image I c and the sparse depth map D s , both of which are obtained by a depth camera, where and To balance the difference in the number of channels between the RGB image and the depth image, the depth image is extended to a size of H×W×3.
[0222] Step 202: Input the RGB image and the extended depth image into the mirrored VMamba encoder simultaneously, and extract RGB features and depth features separately at multiple scales. In one of the VMamba encoders, the Patch Embedding module divides the input RGB image or depth image into non-overlapping blocks (patches), where s×s×3 represents the size of the patches, represents the number of patches. And these patches are projected onto dimension C through a linear embedding layer.
[0223] Step 203: Use multiple VSS modules and Patch Merging modules to construct a four-stage hierarchical structure. In the first stage, only two VSS modules are used, while in the latter three stages, two VSS modules and a Patch Merging module are adopted. The output channel numbers corresponding to each stage are [C, 2C, 4C, 8C] in sequence. The VSS module is responsible for extracting local feature information and concentrating on transparent objects. The Patch Merging module performs downsampling operations, halving the size and doubling the number of channels to capture more features or information. The output of each stage can be expressed as:
[0224]
[0225] where F c is the feature map of the RGB image, F s is the feature map of the sparse depth image, i∈{1,2,3,4} respectively represents each stage of the VMamba, and the size of each stage is
[0226] As Figure 3 shown, it is a flowchart of Step 102, including the following steps:
[0227] Step 301: Given the RGB feature F c and the depth feature F s , multiply them by the weight matrices W q , W k and W v respectively to generate the query matrix (Query, Q), the key matrix (Key, K) and the value matrix (Value, V).
[0228] Step 302: Use the depth feature F s as the query source and the RGB feature F c as the key-value source, as shown in the following formula:
[0229]
[0230] Step 303: Construct a fusion block from cross RGB features to depth features as shown in the following formula:
[0231]
[0232] Step 304: Use the RGB feature F c as the query source and the depth feature F s as the key-value source, as shown in the following formula:
[0233]
[0234] Step 305: Construct a fusion block from cross depth features to RGB features as shown in the following formula:
[0235]
[0236] where, W q c 、W k c and W k c are learnable matrices for RGB feature mapping, W q s 、W q s and W v s are learnable matrices for depth feature mapping, c represents the RGB feature map, s represents the depth feature map, and i represents the i-th stage among the four stages in the VMamba feature extraction network. and respectively represent cross-attention mechanism operations in different directions, and d is a scaling factor.
[0237] Step 306: To promote the direct transmission of RGB information and depth information and solve the problem of gradient disappearance, by introducing skip connections, the interaction between the input information and the cross-attention mechanism fusion information is enhanced. Through concatenation operations and convolution operations using a 3×3 convolution kernel, the two multimodal information is further fused together, focusing on the depth discontinuity regions at the edges of transparent objects, and the finally fused multimodal feature is obtained as shown in the following formula:
[0238]
[0239] As Figure 4 shown, it is a flowchart of Step 103, including the following steps:
[0240] Step 401, mix the fusion result from the lower decoder, the original depth map of the same size, and the fusion features of the current stage through 3×3 or 1×1 convolution operations.
[0241] Step 402, adopt the channel halving operation to divide the channels of the fusion features into two equal and independent branches.
[0242] Step 403, the first branch adopts an improved DenseNet network structure, the core of which consists of multiple dense connection blocks (Dense Blocks). At the same time, to keep the size of the feature maps consistent with that of the other branch, a size alignment module with 3×3 pointwise convolution operations is designed at the end of the branch. Each dense connection block contains 7 cascaded convolutional layers, and the input of each convolutional layer is the concatenation (i.e., dense connection) of the outputs and the input of all previous layers, realizing feature reuse and explicitly retaining shallow details such as the edges and textures of transparent objects and deep semantic information such as the contours of transparent objects. Therefore, the input Hi of the i-th layer i is:
[0243] Hi i = Concat(hi 0 , hi 1 ,..., hi i-1 )(8)
[0244] where hi i is the output of the i-th layer, and Concat(·) is the concatenation in the channel dimension.
[0245] For each layer, first apply 1×1 convolution and 3×3 convolution operations in sequence, and finally perform batch normalization (BN) processing and use the ReLU activation function to obtain the output hi i of the i-th layer as:
[0246] hi i = ReLU(BN(Conv(Conv(Hi i , 1×1, k×4), 3×3, k)))(9)
[0247] where k represents the number of output channels of the 1×1 convolution operation and is a constant, set to 20.
[0248] Step 405, the second path adopts an improved OSA network (One-Shot Aggregation Network), the core of which improves the feature reuse efficiency through layer-by-layer connection and feature concatenation.
[0249] Each layer of the network consists of a 3×3 convolutional layer, and the outputs of all layers are combined through concatenation operations. Therefore, the output yi of the i-th layer i is:
[0250] y i = Conv3×3(x i-1 , Z), i = 1, 2, ..., L(10)
[0251] where x 0 is the input, L is the number of layers, and Z is the number of output channels of the i-th convolutional layer.
[0252] The output y of the final concatenation operation cat is:[[]]
[0253] y cat = Concat(y 1 , y 2 , ..., y L )(11)
[0254] Step 406, merge the output feature maps of the two branches in Steps 404 and 405 through a channel concatenation (Concat) operation, and shuffle the merged feature maps in the channel dimension to improve the network generalization performance.
[0255] Step 407, to better adapt to features of different scales, upsample the features to the current scale through a DySample (dynamic sampling) layer, which can effectively restore the resolution of the image.
[0256] As Figure 5 shown, it is a flowchart of Step 407, including the following steps:[[]]
[0257] Step 501, initialize to generate a normalized two-dimensional grid. Duplicate the second dimension of the grid g times so that each group of features shares the same sampling set to adapt to the grouping mechanism, and obtain the original sampling grid ε as:[[]]
[0258]
[0259] Step 502, given an upsampling scale factor s and a feature map λ of size C × H × W. Map the input through a linear layer to an offset λ 1 , whose size is 2gs 2 × H × W, where g divides the feature map into g groups along the channel dimension and generates g groups of offsets. Multiply the offset λ 1 by 0.25 to satisfy the theoretical marginal condition between overlap and non-overlap, then the output offset λ 2 is:[[]]
[0260] λ 2 = 0.25λ 1 = 0.25ear(λ)(13)Lin
[0261] Step 503, offset λ 2 is reshaped into the final output offset grid λ through the pixel_shuffling function 3 , with a size of 2g×sH×s. Then the finally output offset grid λ 3 is:
[0262] λ 3 = Pixel_shuffling(λ 2 )(14)
[0263] Step 504, the sampling set η is the sum of the offset grid λ 2 and the original sampling grid ε, that is:
[0264] η = λ 2 + ε (15)
[0265] Step 505, using the bilinear interpolation method in the grid_sample function, the feature map is resampled to the output λ' with a size of C×sH×sW, as follows:
[0266] λ' = Grid_sample(η,λ) (16).
[0267] (1) Dataset
[0268] To evaluate the performance of DCDU-Mamba compared with other methods, experiments were conducted on the TransCG dataset. TransCG is the first large-scale, real transparent object dataset, which consists of 57,715 RGB-D images taken from different angles of 130 scenes by a camera connected to a robotic arm.
[0269] (2) Loss function
[0270] Depth completion of transparent objects is regarded as a dense regression task, whose goal is to predict a complete depth map from RGB image and sparse depth map inputs. A joint loss function framework is proposed for the optical properties of refraction and reflection of transparent objects and the depth camera noise problem, dividing the loss into accuracy loss and smoothness loss, and optimizing the geometric accuracy and spatial continuity of the predicted transparent objects respectively.
[0271] The accuracy loss aims to minimize the difference between the predicted depth and the true value, making the model converge towards a high-precision depth map, and is defined as follows:
[0272]
[0273] where L c is L 1 and L2 The combined loss penalizes depth inaccuracy; α is the 1 balance coefficient between the L 2 loss and the L pre loss. D gt represents the predicted depth map, D
[0274] The smoothness loss L s is composed of the cosine distance between the pre predicted depth map D
[0275]
[0276] where and are the gradient vectors along the height axis and width axis of the predicted depth map respectively, and are the gradient vectors along the height axis and width axis of the true depth map respectively. The smoothness loss L s is also calculated on the mask of the transparent region.
[0277] Combining the accuracy loss and the smoothness loss, the total loss function is defined as the weighted sum of the accuracy loss and the smoothness loss:
[0278]
[0279] where β is the weight coefficient that balances the accuracy loss and the smoothness loss.
[0280] (3) Evaluation metrics
[0281] To comprehensively evaluate the performance of the DCDU-Mamba model, the following four key evaluation metrics are used: Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Mean Absolute Relative Error (REL), and threshold δ (where δ is set to 1.05, 1.10, and 1.25). Accordingly, these metrics are calculated on the mask of the transparent region.
[0282] The Mean Absolute Error MAE is the average of the absolute errors between the predicted depth map and the actual depth map, and its calculation formula is:
[0283]
[0284] The root mean square error RMSE is the square root of the ratio of the sum of the squares of the differences between the predicted depth map and the actual depth map to the number of valid depth image pixels. Its calculation formula is:
[0285]
[0286] The average relative error REL is the average of the ratios of the absolute errors to the actual depth map, and its calculation formula is:
[0287]
[0288] where represents the predicted depth map, represents the true depth map, V represents the set of pixels of the depth map under the mask region, and |V| represents the size of the set V.
[0289] The threshold δ is defined such that the predicted depth satisfies the following conditions:
[0290]
[0291] where and represent the pixel values of the predicted depth map and the corresponding true depth map respectively, and δ is set to 1.05, 1.10, and 1.25 respectively.
[0292] (4) Parameter settings
[0293] To update the network parameters and accelerate the convergence speed of the model, the Adam optimizer is used to train the DCDU-Mamba model. The basic model is constructed using the DFNet network framework. The initial learning rate is set to 0.001 and is reduced to 0.2 times the original after every 10 training epochs. The weighting coefficients α and β in the loss function are set to 0.01 and 0.001 respectively. For all experiments, the number of training epochs and the batch size are set to 40 and 8 respectively, and the resolutions of all input RGB images and depth images are uniformly 320×240 for training and testing. To prevent overfitting, data augmentation methods such as random flipping, rotation, adding noise, and color transformation in the HLS color space are used during training. For the feature extraction network based on VMamba, its encoder weights are initialized using the VMamba-T weights pre-trained on ImageNet-1k.
[0294] (5) Comparative experiments and analysis
[0295] The comparison results with other advanced methods on the TransCG dataset are shown in Table 1, including the ClearGrasp algorithm based on three Deeplabv3+ models, the LIDF-Refine algorithm based on local implicit neural representation, the DFNet algorithm based on dense block network, the FDCT algorithm based on improved DFNet, and the TODE-Trans algorithm based on Swin Transformer feature extraction.
[0296] The results in Table 1 confirm that the proposed method not only significantly outperforms the basic model of DFNet but also outperforms many state-of-the-art techniques. In particular, the DCDU-Mamba model reached 91.25 at the threshold δ 1.05 which is 0.82 higher than the best-performing TODE-Trans model, representing an improvement of 0.82%. It is worth noting that, as shown in Table 1, the DCDU-Mamba model has similar error metrics such as RMSE, REL, and MAE compared to TODE-Trans.
[0297] Comparison results with other advanced methods on the TransCG dataset
[0298]
[0299] Apply different depth completion methods to estimate the depth of transparent objects. From left to right, use DFNet, TODE-Trans, and the DCDU-Mamba of the present invention for depth estimation, as Figure 6 shown. The depth map after depth completion can better reflect the contour of the original object to prove the quality of the depth completion algorithm. The DFNet method has lower accuracy compared to the other two methods. The depth map predicted by the TODE-Trans method has some incorrect pixel points. The depth map predicted by the DCDU-Mamba of the present invention is smoother and can better reflect the contour of the transparent object.
Claims
1. A transparent object depth completion method based on DCDU-Mamba, characterized in that: The following steps are involved: 1) Construct the DCDU-Mamba network framework, including encoder, skip connection module and decoder. 2) Obtain RGB image and depth image of transparent object; 3) The encoder inputs the RGB image and depth image of the transparent object into two independent VMamba feature extraction networks for feature extraction to obtain RGB features and depth features; 4) The skip connection module fuses RGB features and depth features through a double-headed cross attention mechanism to obtain a shallow feature map; Use the decoder to fuse RGB features and depth features to obtain a deep feature map; 5) The deep feature map is concatenated with the shallow feature map, and up-sampled through the DOSA network in the decoder to generate a depth completion image of the transparent object; 6) The size of the transparent object depth completion image is restored through a linear embedding layer so that the transparent object depth completion image is aligned with the original depth image size.
2. A method for depth completion of transparent objects based on DCDU-Mamba according to claim 1, characterized in that: In step 2), the RGB image and depth image of the transparent object are obtained by the depth camera, which are respectively denoted as I and c and D s .
3. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 1, characterized in that: In step 3), two independent VMamba feature extraction networks mirror each other.
4. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 1, characterized in that: In step 3), the VMamba feature extraction network includes a Patch Embedding module and a four-stage structure; The four-stage structure includes VSS module, VSS and Patch Merging integrated module I, Patch Merging II, and Patch Merging III in sequence; The Patch Embedding module divides the RGB image or depth image into non-overlapping image blocks are projected onto dimension C through a linear embedding layer; the size of each image block is s×s×3; H and W are the height and width of the RGB image; The VSS and Patch Merging integrated modules I, Patch Merging II and Patch Merging III each include two VSS modules and one Patch Merging module; The VSS module is used to extract local feature information of the image block, thereby generating RGB features or depth features; The Patch Merging module is used to perform a downsampling operation, reduce the size of the image block, and increase the number of channels.
5. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 4, characterized in that: In the four-stage structure of the VMamba feature extraction network, the output of each stage is as follows: Among them, F c is the feature map of the RGB image, F s is the feature map of the sparse depth map, and i∈{1,2,3,4} represents each module in the four-stage structure of VMamba.
6. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 5, characterized in that: In the four-stage structure, the image block size of each stage is 7. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 1, characterized in that: In step 4), the step of fusing RGB features and depth features includes: 4.1) Given RGB feature F c and deep features F s , multiplied by the weight matrix W q , W k and W v , generate query matrix (Query, Q), key matrix (Key, K) and value matrix (Value, V); 4.2) The depth feature F s As the query source, the RGB feature F c As a key-value source, that is: In the formula, is the weight; 4.3) Construct a fusion block from RGB features to depth features, as shown in the following formula: Where d is the scaling factor; 4.4) Transform the RGB feature F c As the query source, the deep feature F s As a key-value source, that is: 4.5) Construct a fusion block from deep features to RGB features, as shown in the following formula: Among them, W q c , and is the learnable matrix of RGB feature map, and is a learnable matrix of deep feature maps, c represents the RGB feature map, s represents the deep feature map, and i represents the i-th stage of the four stages in the VMamba feature extraction network. and They represent the cross attention mechanism operations in different directions, and d is the scaling factor; 4.6) Through the concatenation operation and the convolution operation using a 3×3 convolution kernel, the fusion blocks from RGB features to deep features and the fusion blocks from RGB features to deep features are fused together to obtain the fused shallow feature map. Right now:
8. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 1, characterized in that: In step 5), the step of generating a transparent object depth completion image includes: 5.1) Using the channel half-split operation, the channel of the fused feature is divided into two equal and independent branches, which are respectively recorded as the first branch and the second branch; 5.2) The first branch uses convolution operation to combine the fusion results from the lower decoder and the shallow feature map of the current stage Mix and obtain the fusion feature h; Among them, the input H of the i-th layer of the first branch i 、output h i They are as follows: H i =Concat(h0,h1,…,h i-1 )(8) h i =ReLU(BN(Conv(Conv(H i ,1×1,k×4),3×3,k)))(9) Where k represents the number of output channels of the 1×1 convolution operation; 5.3) The second branch uses convolution operation to combine the fusion results from the lower decoder and the shallow feature map of the current stage Mix and get the fusion feature y cat ,Right now: y i =Conv3×3(x i-1 ,Z),i=1,2,…,L(10) and cat =Concat(y1,y2,…,y L )(11) Where x0 is the input, L is the number of layers, Z is the number of output channels of the i-th convolutional layer; y i is the fusion feature corresponding to the i-th convolutional layer; 5.4) Fusion feature y through channel concatenation operation cat Merge it with the fusion feature h, and shuffle the merged feature map in the channel dimension to generate a feature fusion image; 5.5) The feature fusion image is upsampled to twice its original size through a dynamic sampling layer to generate a transparent object depth completion image.
9. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 8, characterized in that: The first branch adopts an improved DenseNet network structure, including multiple densely connected blocks and a size alignment module; Each densely connected block contains multiple cascaded convolutional layers; The second branch adopts an improved OSA network including multiple convolutional layers.
10. The method for depth completion of transparent objects based on DCDU-Mamba according to claim 8, characterized in that: Step 5.5) The step of upsampling the feature fusion image to the current scale through the dynamic sampling layer includes: 5.5.1) Initialize and generate a normalized two-dimensional grid, and replicate the second dimension of the grid g times so that the features of each group of the two-dimensional grid share the same sampling set and adapt to the grouping mechanism to obtain the original sampling grid ε, that is: 5.5.2) Given an upsampling factor s and a feature map λ of size C×H×W; the input is mapped to an offset λ1 of size 2gs through a linear layer 2 ×H×W; g is the amount of offset; 5.5.3) Adjust the offset λ1 to obtain the offset λ2, that is: λ2=0.25λ1=0.25ear(λ) (13)Lin 5.5.4) After the pixel_shuffling function, the offset λ2 is reshaped to obtain an offset grid λ3 of size 2g×sH×s, that is: λ3=Pixel_shuffling(λ2) (14) 5.5.5) Construct the sampling set η, that is: η=λ2+ε (15) 5.5.6) Using the bilinear interpolation method in the grid_sample function, the feature map is resampled to an output λ' of size C×sH×sW, as shown below: λ′=Grid_sample(η,λ) (16).
Citation Information
Patent Citations
Transparent object depth completion method based on double cross attention network
CN118134983A
Depth completion method based on distance perception mask converter and sparse attention mask mechanism
CN118154654A
A multi-scale cascaded hourglass depth map completion method guided by RGB images
JP7610899B1
Method and device for depth image completion
US20230245282A1