Multi-scale cascade hourglass depth map complementing method guided by RGB image
The multi-scale cascaded hourglass depth map completion method addresses the limitations of existing methods by performing early fusion of RGB and depth images and using attention modules for improved accuracy and resolution in depth map completion.
Patent Information
- Application Number
- JP2024164969
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-09
- Filing Date
- 2024-09-24
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2044-09-24
AI Technical Summary
Existing depth map completion methods fail to fully utilize RGB image information and suffer from inaccuracies and loss of global and local information, leading to incomplete and inaccurate depth maps.
A multi-scale cascaded hourglass depth map completion method guided by RGB images, which performs early fusion of sparse depth maps and RGB images, utilizes a multi-scale configuration and attention modules for encoding and decoding, and includes an optimization augmentation module for edge detail enhancement.
The method improves the accuracy, spatial resolution, and robustness of completed depth maps by fully utilizing RGB image information and optimizing the calculation efficiency, resulting in more complete and accurate depth maps.
Smart Images

Figure 2025079313000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to the technical field of computer vision, and in particular to a multi-scale cascaded hourglass depth map completion method guided by RGB images. [Background technology]
[0002] The task of depth map completion refers to completing the incomplete depth map obtained, and aims to obtain as complete and accurate scene depth information as possible by estimating and filling in the missing depth information using deep learning models. Depth maps play an important role in various tasks, such as scene understanding, autonomous driving, robot navigation, simultaneous localization and mapping (SLAM), intelligent agriculture, and augmented reality. Therefore, obtaining accurate pixel-level scene depth is a long-term goal for future research.
[0003] In recent years, convolutional neural networks (CNNs) have promoted the rapid development of depth map completion tasks due to their excellent ability in multi-scale feature extraction. Depth map completion generally relies on CNNs, and these networks can capture the features and structures of depth maps by learning from a huge amount of depth map data. In practical applications, when depth information is missing, these networks can predict the depth value of the missing area based on the existing depth information and other image features. With the development of new mechanisms, the depth map completion task has made further progress. Among them, the attention mechanism is a key technology that imitates the attention allocation process in human information processing and enables the model to pay attention to different parts of the input data, thereby improving the accuracy of learning and prediction. Self-attention mechanisms have been introduced and widely applied to the depth map completion task, bringing about significant performance improvements. The application of multi-scale networks can effectively capture features of different levels of detail and fully utilize low-level and high-level features. Although the depth map completion task has made significant progress, the completeness of the completion information is still largely lacking. Therefore, in order to solve the problem of sparse depth map completion, when a sparse depth map sD and a guide RGB image aligned to sD are given, a multi-scale cascaded hourglass network guided by an RGB image is proposed to solve the problem. Although some prior patents have proposed sparse depth map completion means, there are still some problems. For example, in a depth map completion method based on a weakly aligned RGB-D image (Patent Document 1), a neural network is used to distinguish flat regions from deep structure regions, and surface normals and Gaussian weights are used to optimize the smoothness and structural accuracy of depth values. However, in such a method, the color and texture information in the RGB image and the effective information of the sparse depth map are not fully utilized, and the completed dense depth map is insufficient in accuracy and effectiveness. In addition, in a depth map completion method based on deformable convolution (Patent Document 2), the RGB image is used as a guide, and the ENet structure is improved by increasing deformable convolution and additional teaching information, but global and local information is lost during the process of cropping image data.Because large-scale depth maps generally have lower resolution and more global information, while small-scale depth maps have higher resolution and more local detail information, this method cannot provide accurate scene information and object edge information. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] China Patent Application Publication No. 116012430 [Patent Document 2] China Patent Publication No. 113538278 Summary of the Invention [Problem to be solved by the invention]
[0005] In view of the shortcomings of the prior art, the present invention provides a multi-scale cascaded hourglass depth map completion method guided by an RGB image, which performs early fusion processing on the input sparse depth map and RGB image to maintain the integrity of depth information and reduce information loss, and then uses a multi-scale configuration and attention module to perform encoding and decoding operations to improve the integrity and edge resolution of the completed depth map. [Means for solving the problem]
[0006] The technical means of the present invention are as follows. According to one embodiment of the present invention, a multi-scale cascaded hourglass depth map completion method guided by an RGB image includes the steps of: Step S1: obtaining a sparse depth map to be complemented and a corresponding RGB color guide image, performing channel-dimensional preprocessing on the sparse depth map and the RGB color guide image, and connecting them in series to obtain an RGB_D image to be processed; a step S2 of inputting the target RGB_D image into a trained depth map completion network model, the depth map completion network model including an early fusion encoder, a sparse depth map teaching module, a multi-scale hourglass completion module and an optimization augmentation module, the early fusion encoder is used to generate a feature map with a scale decreasing at each layer based on the RGB_D image, the sparse depth map teaching module is used to process the RGB_D image with three different layers of sub-networks to obtain down-sampled sparse image teaching feature maps at three different scales, the multi-scale hourglass completion module is used to complement the sparse depth map based on the output of the early fusion encoder and the output of the sparse depth map teaching module to obtain a completed dense depth map, and the optimization augmentation module is used to perform edge detail enhancement processing on the completed dense depth map; The method includes a step S3 of taking the output of the optimization enhancement module as an interpolated dense depth map.
[0007] Furthermore, the step of serially connecting the sparse depth map and the RGB color image through channel dimension preprocessing includes: performing a convolution process on the RGB color guide image to obtain a 48-channel feature map; performing a convolution process on the sparse depth map to obtain a 16-channel feature map; The method includes a step of serially connecting the processed 48-channel RGB image feature map and the 16-channel sparse depth map feature map in the channel dimension to obtain an RGB_D image to be processed.
[0008] Furthermore, the early fusion encoder further comprises: One pre-processing layer, which includes one 3x3 convolutional layer and one ReLU activation function, performs initialization convolution on the serially connected RGB_D image; Each of them consists of one 3x3 convolutional layer and two ReLU activation functions, and contains five sequence containers that are used to make the output feature maps have multiple scales, such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the scale of the initial RGB_D image.
[0009] Furthermore, the multi-scale hourglass completion module: The present invention includes four cascaded hourglass encoding / decoding modules, each of which includes an hourglass encoder and an hourglass decoder, the hourglass encoder extracting depth information and different-scale outputs from different-scale early fusion features output from the early fusion encoder and different-scale sparse depth features output from the sparse depth map instruction module, and the hourglass decoder being used to output a dense depth map after completion; The hourglass encoder includes three convolutional attention sequence containers, each of which includes a 3x3 convolutional layer, a ReLU activation function layer, and a dual attention module; The hourglass decoder includes three transposed convolutional attention sequence containers, each of which includes a 3x3 transposed convolutional layer, a convolutional layer, two ReLU activation function layers, and a dual attention module.
[0010] Furthermore, the dual attention module: It includes one convolution module, one spatial attention module and one channel attention module; The convolution module includes two 3x3 convolution layers, one ReLU activation function layer, and two batch normalization layers; The channel attention module is used to: First, two transformations are performed on the input feature map to generate two auxiliary feature maps. Next, a 1x1 convolution kernel is used to perform channel compression on the input feature map to reduce the number of channels to in_planes / ratio, where in_planes is the number of input channels and ratio is a parameter that controls the channel compression ratio. Next, a nonlinear transformation is performed using the ReLU activation function, and then the number of channels is restored to in_planes using a 1x1 convolution kernel again. Finally, the results of average pooling and max pooling are added together, and the output of the channel attention is scaled by a sigmoid layer to obtain the final channel attention weights, which weight the channels of the input feature map. The spatial attention module is used to: First, the input feature maps are first sent to average pooling and max pooling operations to calculate the average and max feature maps, respectively; Then, these two feature maps are serially connected to form a feature map containing two channels; This feature map is then passed through a 7x7 convolutional layer to calculate the spatial attention weights, Finally, we scale the output of the spatial attention using a sigmoid layer to obtain the final spatial attention weight graph.
[0011] Furthermore, the optimization reinforcement module is a recurrent four-layer U-net network module, which is composed of four encoders, a decoder, and an output layer; The encoder includes an input layer, a downsampling module, and a skip connection. The input layer is initially used to accept input data; The downsampling module is composed of a 3x3 convolution layer, a pooling layer and an activation function, which is used to gradually reduce the size and number of channels of the feature map and extract high-level semantic features; The skip connections are used to preserve the feature maps of each layer of the encoder for feature fusion at the decoder; The encoder includes an upsampling module, a post-upsampling processing module, a skip connection, and an output layer; The upsampling module is composed of a deconvolution layer and an activation function, and is used to gradually increase the size and number of channels of the feature map to restore it to the original input size; The skip connection is used to connect the feature map of the corresponding layer in the encoder and the feature map of the decoder to realize feature fusion; The post-upsampling processing module is composed of a convolution layer, a batch normalization layer and an activation function, and is used to further process the upsampling features; The output layer is used to combine the low-level and high-level features to generate the final completion result.
[0012] Further, the step of training the depth map completion network model includes: Obtaining a depth map completion network training dataset from the Kitti dataset, which is a publicly available dataset used in computer vision and autonomous driving research, including GroundTruth with high labeling information coverage and paired sparse depth maps and color RGB images; training a network using the RGB image and the sparse depth map as input data for the network; Calculating a loss between the outputted depth map of the network and the complete depth map of GroundTruth in the training dataset, and performing backpropagation using the loss to update the weights of the network; The method includes a step of minimizing the loss value using gradient descent for the loss value of the completed depth map obtained by the depth map completion network and the actual complete depth map to obtain an optimal model. Effect of the Invention
[0013] Compared with the prior art, the present invention has the following advantages: 1. The present invention uses a method of preprocessing and early fusing RGB images and sparse depth maps, and fully utilizes the color, texture and edge information of RGB images to fill in missing areas in the depth map, which helps to improve the accuracy, spatial resolution and robustness of the depth map, and improves the completion effect of the depth map. In addition, a sparse depth map teaching module is proposed to use sparse convolution instead of traditional convolution, which optimizes the calculation and reduces parameters, and improves the efficiency and speed of the model. 2. The present invention proposes a mechanism of a multi-scale cascade hourglass network model, which extracts multi-scale features to refine the depth, and the hourglass network can capture feature information of different scales through multi-layer feature extraction from top to bottom and bottom to top. In addition, lower layer features can be transmitted to upper layers via a path from bottom to top, and upper layer features can be transmitted to lower layers via a path from top to bottom. Such feature multiplexing can improve the parameter efficiency and calculation efficiency of the network, and can obtain more depth information corresponding to different configurations. 3. Considering the feature of many missing depths in the depth map, the present invention designs a dual attention module that combines spatial attention mechanism and channel attention mechanism, which helps the model to understand and process the input data more accurately, and improves the model's performance in the depth map completion task. [Brief description of the drawings]
[0014] In order to more clearly explain the technical means in the embodiments of the present invention or the prior art, the accompanying drawings necessary for the description of the embodiments or the prior art will be briefly introduced below. It goes without saying that the drawings below are some embodiments of the present invention, and those skilled in the art can further derive other drawings based on these drawings without creative effort.
[0015] [Figure 1]1 is a flow chart of a multi-scale cascaded hourglass depth map completion method guided by an RGB image according to the present invention. [Diagram 2] FIG. 2 is a schematic diagram of a depth map completion network model according to an embodiment of the present invention. [Diagram 3] FIG. 2 is a block diagram of an early fusion encoder according to an embodiment of the present invention. [Figure 4] FIG. 2 is a block diagram of a sparse depth map teaching module in an embodiment of the present invention; [Diagram 5] 1 is a block diagram of a multi-scale hourglass completion module according to an embodiment of the present invention; [Figure 6] FIG. 1 is a diagram for explaining a customer device information DB stored in a secondary storage device included in the voice processing device according to the first embodiment; [Figure 7] FIG. 2 is a configuration diagram of an optimization reinforcement module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] In order to make clearer the purpose, technical means and advantages of the embodiments of the present invention, the technical means of the embodiments of the present invention will be described clearly and completely below with reference to the drawings in the embodiments of the present invention, but it goes without saying that the described embodiments are not all the embodiments but only some of the embodiments of the present invention. Any other embodiments that a person skilled in the art can obtain based on the embodiments of the present invention without any creative effort shall be included in the protection scope of the present invention.
[0017] It should be noted that the terms "first" and "second" in the present specification, claims, and drawings are used to distinguish between similar objects, and are not used to describe a particular order or priority. It is understood that the embodiments of the present invention described herein can be performed according to orders other than those illustrated herein, and therefore the data thus employed may be interchanged where appropriate. Moreover, the terms "comprise" and "have," as well as any variations thereof, refer to a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units need not be limited to those steps or units that are explicitly recited, but may further include other steps or units that are not explicitly recited or that are inherent in the process, method, product, or apparatus.
[0018] As shown in FIG. 1, the present invention provides a multi-scale cascaded hourglass depth map completion method guided by RGB image, which includes the following steps S1, S2 and S3.
[0019] S1: Obtain a sparse depth map to be complemented and a corresponding RGB color guide image, perform channel-dimensional preprocessing on the sparse depth map and the RGB color guide image, and connect them in series to obtain an RGB_D image to be processed.
[0020] Specifically, considering that RGB images can provide rich context information, the depth map completion model can help better understand the composition and semantics of a scene. By combining the RGB image and the depth map, the model can more accurately estimate the missing depth information. The convolution operation is performed on the RGB image and the sparse depth map, and the channel dimensions of the two resulting feature maps are serially connected to form a new input RGB_D image, which can be expressed as follows: RGBD=C(RGB,Depth) Here, C denotes the series connection in the channel dimension.
[0021] S2: The RGB_D image is input into the trained depth map completion network model (see FIG. 2), and the RGB_D image first enters the early fusion encoder (see FIG. 3), which generates a feature map whose scale decreases with each layer based on the RGB_D image. In addition, in order to effectively utilize the sparsity of the sparse depth map data, a sparse depth map teaching module (see FIG. 4) is designed, which uses a mask to avoid unnecessary calculations for depth value missing positions, thereby improving calculation efficiency and model performance. The sparse depth map of the above dataset is input into this teaching module, and three different-layer sub-networks of this module obtain three different scales of down-sampled sparse image teaching feature maps. Then, the output of the early fusion encoder and the output of the sparse depth map teaching module are input into the multi-scale hourglass completion module (see FIG. 5) to complete the sparse depth map and obtain a completed dense depth map.
[0022] Specifically, this depth map completion network model mainly includes an early fusion encoder, a sparse depth map teaching module, a multi-scale hourglass completion module and an optimization augmentation module (see FIG. 7).
[0023] Furthermore, the structure of the early fusion encoder includes one preprocessing layer and five convolutional sequence containers. The preprocessing layer includes one 3x3 convolutional layer, which performs initialization convolution on the serially connected RGB_D feature map. The convolutional sequence container includes one 3x3 convolutional layer and one ReLU activation function. The output feature map by the encoder configured with these five convolutional sequence containers has multiple scales, such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the scale of the initial RGB_D feature map.
[0024] The working flow of the early fusion encoder is as follows.
[0025] First, the RGB_D feature map obtained in step S1 is input to the initialization sequence container init as input, and an initialization operation is performed on input to obtain x 0 get.
[0026] Next, x 0 Input the following to the first downsampled convolution sequence container to generate a downsampled feature map x 1 get.
[0027] x 1 We feed this to the second downsampled convolution sequence container to generate a downsampled feature map x with a scale of 1 / 4. 2 get.
[0028] x 2 We feed this to the third downsampled convolution sequence container to generate a downsampled feature map x 3 get.
[0029] x 3 We feed the fourth downsampled convolution sequence container with the input, x, to generate a downsampled feature map with a scale of 1 / 16. 4 get.
[0030] x 4 The final downsampled convolution sequence container is fed with the input, resulting in a downsampled feature map x with a scale of 1 / 32. 5 To summarize the above, they can be expressed as follows: x i =enc_i(x i-1 ) Here, the values corresponding to i are 1, 2, 3, 4, and 5. enc_1 is the first downsampled convolution sequence container, enc_2 is the second downsampled convolution sequence container, enc_3 is the third downsampled convolution sequence container, enc_4 is the fourth downsampled convolution sequence container, and enc_5 is the fifth downsampled convolution sequence container.
[0031] Furthermore, we use a sparse convolution module, and obtain downsampled sparse image guidance feature maps at three different scales through three different hierarchical submodules of this module.
[0032] Specifically, after conducting related experiments on the sparse depth map teaching module, it was found that the number of sparse convolution layers also affected the experimental results, so different layers of sparse convolution operations were performed on the sparse depth map. The principle is that a value greater than 0 is taken for the sparse depth map as a mask, and the mask and sparse depth map are input into a sparse convolution network, and the sparse depth map is convolutionally processed by this network. This sparse depth map teaching module contains a total of three sparse convolution networks to process the sparse depth map. Three layers of sparse convolution, two layers of sparse convolution, and one layer of sparse convolution are respectively performed, and different scales of image sizes are obtained, which are 1 / 8, 1 / 4, and 1 / 2 of the initial image scale, respectively.
[0033] The working principle of the sparse depth map teaching module is as follows.
[0034] a. The sparse depth map sparsedeph is input to the first submodule of the sparse depth map training module, and through three sparse convolutional layers, a feature map x with a scale 1 / 8 of the original sparse depth map is generated. 1 get. x 1 =sparseconv(sparseconv(sparseconv(sparsedepth)))
[0035] b. The sparse depth map is input to the second submodule of the sparse depth map training module, and passed through three sparse convolutional layers to generate a feature map x 2 get. x 2 =sparseconv(sparseconv(sparsedepth))
[0036] c. The sparse depth map is input to the third submodule of the sparse depth map training module and passed through three sparse convolutional layers to obtain a feature map x with a scale of 1 / 2 that of the first sparse depth map. 3 get. x 3 =sparseconv(sparsedepth)
[0037] Furthermore, the structure of the multi-scale hourglass interpolation module includes four hourglass encoding / decoding modules cascaded together, where each hourglass encoding / decoding module includes an hourglass encoder and an hourglass decoder. The hourglass encoder extracts more depth information and different scale outputs from the fusion features of different scales and the sparse depth features of different scales. The hourglass decoder is designed to output the dense depth map after interpolation.
[0038] Here, one hourglass encoder includes three convolutional attention sequence containers, each of which includes one 3x3 convolutional layer, one ReLU activation function layer, and one dual attention module (see Fig. 6). The convolutional attention sequence container first serially connects the different scale depth maps from the sparse depth map teaching module to the output of the previous hourglass decoder, which is upsampled, and then serves as the input of this hourglass encoder. Then, this encoder performs layer-by-layer matrix addition with the three inputs from the previous hourglass decoder (the input of the first hourglass encoder is only the input of the sparse depth map teaching module), and then obtains the output results of different scales through these three sequence containers.
[0039] Preferably, in the network model of this embodiment, the input of the first hourglass encoder is the 1 / 8 scale sparse depth map after being processed by the sparse depth map teaching module, the input of the second hourglass encoder is the feature map connected in series to the 1 / 4 scale sparse depth map after being processed by the sparse depth map teaching module after the result output by the first hourglass decoder is upsampled, and the previous three elements (total of four elements) in the result array output by the previous hourglass encoder. The input of the third hourglass encoder is the feature map connected in series to the 1 / 2 scale sparse depth map after being processed by the sparse depth map teaching module after the result output by the second hourglass decoder is upsampled, and the previous three elements (total of four elements) in the result array output by the previous hourglass encoder. The input of the fourth hourglass encoder is the feature map connected in series to the original sparse depth map after being upsampled.
[0040] Preferably, the working flow of this hourglass encoder is as follows:
[0041] First, the output of the sparse depth map teaching module is taken as input x, and then it is passed through the first convolutional attention sequence container enc1 to obtain x. 0 Then, it is judged whether there is an input from the previous hourglass decoder (only the first hourglass encoder has no input from the previous stage), and if there is, it performs upsampling processing by the function to obtain x 0 This is then used as the input for the second convolutional attention container by performing matrix addition with . x 0 =enc_1(x) x 0 =x 0 +F.interpolate(x_b3)
[0042] b, followed by x 0 Take x as input and pass it through the second convolutional attention sequence container enc2 to obtain x 1Then, it is judged whether there is an input from the previous hourglass decoder (only the first hourglass encoder has no input from the previous stage), and if there is, it performs upsampling processing by the function to obtain x 1 Then, we perform matrix addition with this as the input for the third convolutional attention container. x 1 =enc_2(x 0 ) x 1 =x 1 +F.interpolate(x_b2)
[0043] c, followed by x 1 Take x as input and pass it through the third convolutional attention sequence container enc3 to obtain x 2 Then, it is judged whether there is an input from the previous hourglass decoder (only the first hourglass encoder has no input from the previous stage), and if there is, it performs upsampling processing by the function to obtain x 2 Then, perform matrix addition with x 0 , x 1 , x 2 Provide feedback. x 2 =enc_3(x 1 ) x 2 =x 2 +F.interpolate(x_b1) where F.interpolate is the upsampling function and + is matrix addition.
[0044] Here, one hourglass decoder includes three transposed convolutional attention sequence containers, each of which includes one 3x3 transposed convolutional layer, one convolutional layer, two ReLU activation function layers and one dual attention module. The output of the hourglass encoder and the output of the early fusion encoder are added pointwise and then input to the decoder, and the decoder performs depth map completion processing (the completion results of the previous three hourglass decoders are all input to the next hourglass encoder, and the output of the last hourglass decoder is the completed dense depth map).
[0045] Preferably, in the network model of this embodiment, the first input order for inputting the five outputs of the early fusion encoder to the hourglass decoder described below is such that 1 / 32 and 1 / 16 are input to the first hourglass decoder, 1 / 16 and 1 / 8 are input to the second hourglass decoder, 1 / 8 and 1 / 4 are input to the third hourglass decoder, and 1 / 4 and 1 / 2 are input to the final hourglass decoder.
[0046] The operation flow of this hourglass decoder is as follows.
[0047] a. First, the output of the hourglass encoder and the output of the early fusion encoder are input to the hourglass decoder, and the following operations are performed on the pre_dx and pre_cx arrays, respectively. x 2 =pre_dx[2]+pre_cx[2] x 1 =pre_dx[1]+pre_cx[1] x 0 =pre_dx[0]+pre_cx[0] where x 2 is the input of the first convolutional transpose sequence container, and x 1 is the input of the second convolutional transpose sequence container, and x 0 is the input of the third convolution transpose sequence container.
[0048] b. x as shown below 2 Enter into the first convolution transpose sequence container and get x 3 Get x 3 and x 1 Put both into the second convolution transpose sequence container to get x 4 get. x 3 =dec3(x 2 ) x 4 =dec2(x 1 +x 3 ) Here, dec3 is the first convolution-transposed sequence container, and dec2 is the second convolution-transposed sequence container.
[0049] c. Finally, x as shown below. 4 and x 0 are input together into the predicted sequence container. output=dec1(x 4 +x 0 ) Here, output is the final prediction and dec1 is the third convolutional transposed sequence container.
[0050] In order to learn to more depth, a dual attention module is designed in this invention, which includes one preprocessing layer, a spatial attention module and a channel attention module.
[0051] The preprocessing layer includes two 3x3 convolutional layers, one ReLU activation function layer, and two batch normalization BatchNorm layers. The input of this module first passes through one 3x3 convolutional layer, followed by batch normalization and ReLU activation function. The output then passes through another 3x3 convolutional layer and is batch normalized again.
[0052] In this basic module, Channel Attention and Spatial Attention mechanisms are further introduced. Here, the Channel Attention mechanism mainly pays attention to the relationship and importance between different channels in the feature map. By learning the channel weights, the network can self-adaptively adjust the importance of each channel. The Spatial Attention mechanism mainly pays attention to the relationship and importance of different spatial locations in the feature map. By learning the spatial weights, the network can self-adaptively adjust the importance of each spatial location and generate a weight graph to highlight the spatial information of a specific region. The processing of these two modules can enhance the feature modeling ability of the network, and the network can better understand and utilize the key information in the input data, improving the performance and generalization ability of the model. Furthermore, the flow of this module is as follows:
[0053] First, the feature map depthen after being processed by the 3x3 convolution layer and the ReLU activation layer in the sequence container in which this dual attention module exists is taken as the input of this module. In this module, a series of convolution normalization and activation function operations are first performed to obtain the channel attention input, which is denoted as cattinput and can be expressed as the following equation: cattinput=bn(conv(relu(bn(conv(depthen))))) Here, conv stands for 3x3 convolution, relu stands for activation function, and bn stands for batch normalization.
[0054] b. Next, the obtained cattinput is input to the channel attention module ca, and a matrix dot product is performed with the original input to obtain the spatial attention input sattinput, as shown in the following equation. sattinput=ca(cattinput) o cattinput where ca represents the channel attention module and o represents the pixel-level matrix dot product operation.
[0055] c. Finally, the obtained spatial attention input sattinput is input to the spatial attention module sa, and a matrix dot product operation with the original input is performed to obtain the next sequence container input nextinput, as shown in the following equation. nextinput=sa(sattinput) o sattinput where sa represents the spatial attention module and o represents the pixel-level matrix dot product operation.
[0056] Furthermore, the optimization module is based on the U-net network and is a recurrent four-layer U-net network module, which is composed of four encoders, a decoder and an output layer, and the encoder part of the optimization reinforcement module includes an input layer, a downsampling module and a skip connection. The input layer is used to initially accept input data, and the downsampling module is composed of a 3x3 convolution layer, a pooling layer and an activation function, which is used to gradually reduce the size and number of channels of the feature map and extract high-level semantic features. The skip connection preserves the feature map of each layer of the encoder so that the decoder can perform feature fusion. These features are transmitted to the decoder part.
[0057] The encoder part includes an upsampling module and a skip connection. The upsampling module is composed of a deconvolution layer and an activation function, and is used to gradually increase the size and number of channels of the feature map and restore it to the original input size. The skip connection connects the feature map of the corresponding layer in the encoder with the feature map of the decoder to realize feature fusion. The post-upsampling processing module is composed of a convolution layer, a batch normalization layer, and an activation function, and is used to further process the features after upsampling.
[0058] The output layer of the final layer is used to generate the final completion result, which aims to combine the low-level and high-level features to obtain a more accurate completion result.
[0059] According to the above method, a depth map completion network model is constructed. In order to obtain an optimal completion model, the model needs to be trained. The present invention obtains a depth map completion network training dataset from the Kitti dataset, which is publicly available for research in computer vision and autonomous driving, and trains a depth map completion network model. Specifically, the present invention includes the following steps:
[0060] 1. Obtaining a depth map completion network training dataset from the Kitti dataset, which is a publicly available dataset used in computer vision and autonomous driving research, including GroundTruth with high labeling information coverage and paired sparse depth maps and color RGB images; 2. Training a network using the RGB image and the sparse depth map as input data for the network; 3. Calculate a loss between the outputted depth map of the network and the complete depth map of GroundTruth in the training dataset, and perform backpropagation according to the loss to update the weights of the network; 4. The method includes a step of minimizing the loss value of the completed depth map obtained by the depth map completion network and the actual complete depth map using gradient descent to obtain an optimal model.
[0061] In the present invention, the training parameters of the network model were set as follows: All experiments of the present invention were completed in a Python 3.7 (Ubuntu 18.04) environment, the video card used in the experiment was A800-SXM, and the PyTorch deep learning mechanism was used to train the network. In addition, the Adam optimizer was used for training, the parameters were updated by the Adam optimizer, and finally the weights of the network training were recorded and saved. The size of the input data set was 30,000 pairs, and the test images were 1,000 pairs. In the parameter settings, the learning rate was 0.001, the batchsize was 14, and the training period epochs was 50.
[0062] During the training process, to obtain the optimal model, the learning rate was adjusted using a linear decay policy, with the base learning rate being a maximum of 0.001, and the learning rate decaying to half the initial rate every five training rounds.
[0063] Furthermore, the L1 loss function is used as the loss function for this model training, which has a stable gradient for any input value, does not cause the problem of gradient explosion, and has a robust solution.
[0064] In this experiment, the loss between the output image and the target image is calculated using the following L1 loss function:
number
number
number
number
number
[0065] S3: Take the output of the optimization augmentation module as an interpolated dense depth map.
[0066] The present invention performs an early fusion process on the input sparse depth map and RGB image to maintain the integrity of the depth information and reduce information loss, and then uses a multi-scale configuration and attention module to perform encoding / decoding operations to improve the integrity and edge resolution of the completed depth map.
[0067] Finally, the following should be noted: The above embodiments are merely for describing the technical means of the present invention, and are not intended to limit the same, and the present invention has been described in detail with reference to the above embodiments. However, the technical means described in the above embodiments can be modified or equivalently replaced in part or all of the technical features, and it is obvious to those skilled in the art that the essence of the corresponding technical means does not depart from the scope of the technical means of the embodiments of the present invention through such modifications or replacements.
[0068] (Additional Note) (Appendix 1) Step S1: obtaining a sparse depth map to be complemented and a corresponding RGB color guide image, performing channel-dimensional preprocessing on the sparse depth map and the RGB color guide image, and connecting them in series to obtain an RGB_D image to be processed; a step S2 of inputting the target RGB_D image into a trained depth map completion network model, the depth map completion network model including an early fusion encoder, a sparse depth map teaching module, a multi-scale hourglass completion module and an optimization augmentation module, the early fusion encoder is used to generate a feature map with a scale decreasing at each layer based on the RGB_D image, the sparse depth map teaching module is used to process the RGB_D image with three different layers of sub-networks to obtain down-sampled sparse image teaching feature maps at three different scales, the multi-scale hourglass completion module is used to complement the sparse depth map based on the output of the early fusion encoder and the output of the sparse depth map teaching module to obtain a completed dense depth map, and the optimization augmentation module is used to perform edge detail enhancement processing on the completed dense depth map; and step S3 of obtaining an output of the optimization augmentation module as an interpolated dense depth map.
[0069] (Appendix 2) The step of serially connecting the sparse depth map and the RGB color image through channel dimension preprocessing includes: performing a convolution process on the RGB color guide image to obtain a 48-channel feature map; performing a convolution process on the sparse depth map to obtain a 16-channel feature map; and a step of serially connecting the processed 48-channel RGB image feature map and the 16-channel sparse depth map feature map in the channel dimension to obtain an RGB_D image to be processed.
[0070] (Appendix 3) The early fusion encoder includes: One pre-processing layer, which includes one 3x3 convolutional layer and one ReLU activation function, performs initialization convolution on the serially connected RGB_D image; The RGB image-guided multi-scale cascaded hourglass depth map completion method described in Appendix 1, characterized in that it includes five sequence containers, each of which comprises one 3x3 convolutional layer and two ReLU activation functions, and is used to make the output feature map have multiple scales, such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the scale of the initial RGB_D image.
[0071] (Appendix 4) The multi-scale hourglass completion module includes: The present invention includes four cascaded hourglass encoding / decoding modules, each of which includes an hourglass encoder and an hourglass decoder, the hourglass encoder extracting depth information and different-scale outputs from different-scale early fusion features output from the early fusion encoder and different-scale sparse depth features output from the sparse depth map instruction module, and the hourglass decoder being used to output a dense depth map after completion; The hourglass encoder includes three convolutional attention sequence containers, each of which includes a 3x3 convolutional layer, a ReLU activation function layer, and a dual attention module; The multi-scale cascaded hourglass depth map completion method guided by RGB images described in Appendix 1, characterized in that the hourglass decoder includes three transposed convolutional attention sequence containers, each of which includes a 3x3 transposed convolutional layer, a convolutional layer, two ReLU activation function layers and a dual attention module.
[0072] (Appendix 5) The dual attention module: It includes one convolution module, one spatial attention module and one channel attention module; The convolution module includes two 3x3 convolution layers, one ReLU activation function layer, and two batch normalization layers; The channel attention module is used to: First, two transformations are performed on the input feature map to generate two auxiliary feature maps. Next, a 1x1 convolution kernel is used to perform channel compression on the input feature map to reduce the number of channels to in_planes / ratio, where in_planes is the number of input channels and ratio is a parameter that controls the channel compression ratio. Next, a nonlinear transformation is performed using the ReLU activation function, and then the number of channels is restored to in_planes using a 1x1 convolution kernel again. Finally, the results of average pooling and max pooling are summed and the output of the channel attention is scaled by a sigmoid layer to obtain the final channel attention weights, which are used to weight the channels of the input feature maps; The spatial attention module is used to: First, the input feature maps are first sent to average pooling and max pooling operations to calculate the average and max feature maps, respectively; Then, these two feature maps are serially connected to form a feature map containing two channels; This feature map is then passed through a 7x7 convolutional layer to calculate the spatial attention weights, Finally, the output of the spatial attention is scaled using a sigmoid layer to obtain a final spatial attention weight graph.
[0073] (Appendix 6) The optimization reinforcement module is a recurrent four-layer U-net network module, which is composed of four encoders, a decoder, and an output layer; The encoder includes an input layer, a downsampling module, and a skip connection. The input layer is initially used to accept input data; The downsampling module is composed of a 3x3 convolution layer, a pooling layer and an activation function, which is used to gradually reduce the size and number of channels of the feature map and extract high-level semantic features; The skip connections are used to preserve the feature maps of each layer of the encoder for feature fusion at the decoder; The encoder includes an upsampling module, a post-upsampling processing module, a skip connection, and an output layer; The upsampling module is composed of a deconvolution layer and an activation function, and is used to gradually increase the size and number of channels of the feature map to restore it to the original input size; The skip connection is used to connect the feature map of the corresponding layer in the encoder and the feature map of the decoder to realize feature fusion; The post-upsampling processing module is composed of a convolution layer, a batch normalization layer and an activation function, and is used to further process the upsampling features; The RGB image-guided multi-scale cascaded hourglass depth map completion method of claim 1, wherein the output layer is used to combine low-level and high-level features to generate a final completion result.
[0074] (Appendix 7) The step of training the depth map completion network model includes: Obtaining a depth map completion network training dataset from the Kitti dataset, which is a publicly available dataset used in computer vision and autonomous driving research, including GroundTruth with high labeling information coverage and paired sparse depth maps and color RGB images; training a network using the RGB image and the sparse depth map as input data for the network; Calculating a loss between the outputted depth map of the network and the complete depth map of GroundTruth in the training dataset, and performing backpropagation using the loss to update the weights of the network; and minimizing a loss value using a gradient descent method for the loss values of the completed depth map obtained by the depth map completion network and the actual complete depth map to obtain an optimal model.
Claims
1. Step S1: obtaining a sparse depth map to be complemented and a corresponding RGB color guide image, performing channel-dimensional preprocessing on the sparse depth map and the RGB color guide image, and connecting them in series to obtain an RGB_D image to be processed; S2, inputting the target RGB_D image into a trained depth map completion network model, the depth map completion network model including an early fusion encoder, a sparse depth map teaching module, a multi-scale hourglass completion module and an optimization augmentation module, the early fusion encoder is used to generate a feature map with a scale decreasing at each layer based on the RGB_D image, the sparse depth map teaching module is used to process the RGB_D image with three different layers of sub-networks to obtain down-sampled sparse image teaching feature maps at three different scales, the multi-scale hourglass completion module is used to complete the sparse depth map based on the output of the early fusion encoder and the output of the sparse depth map teaching module to obtain a completed dense depth map, and the optimization augmentation module is used to perform edge detail enhancement processing on the completed dense depth map; and step S3 of taking the output of the optimization augmentation module as an interpolated dense depth map.
2. The step of serially connecting the sparse depth map and the RGB color image through channel-dimensional pre-processing includes: performing a convolution process on the RGB color guide image to obtain a 48-channel feature map; performing a convolution process on the sparse depth map to obtain a 16-channel feature map; The method of claim 1, further comprising: serially connecting the processed 48-channel RGB image feature map and the 16-channel sparse depth map feature map in the channel dimension to obtain a target RGB_D image.
3. The early fusion encoder comprises: A pre-processing layer including a 3x3 convolution layer and a ReLU activation function, which performs initialization convolution on the serially connected RGB_D image; The method of claim 1, further comprising: five sequence containers, each of which comprises one 3x3 convolutional layer and two ReLU activation functions, and each of which is used to make the output feature map have multiple scales, such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the scale of the original RGB_D image.
4. The multi-scale hourglass completion module includes: The present invention includes four cascaded hourglass encoding / decoding modules, each of which includes an hourglass encoder and an hourglass decoder, the hourglass encoder extracting depth information and different-scale outputs from different-scale early fusion features output from the early fusion encoder and different-scale sparse depth features output from the sparse depth map instruction module, and the hourglass decoder being used to output a dense depth map after completion; The hourglass encoder includes three convolutional attention sequence containers, each of which includes a 3x3 convolutional layer, a ReLU activation function layer, and a dual attention module; The method of claim 1, wherein the hourglass decoder includes three transposed convolutional attention sequence containers, each of which includes a 3x3 transposed convolutional layer, a convolutional layer, two ReLU activation function layers, and a dual attention module.
5. The dual attention module: It includes one convolution module, one spatial attention module and one channel attention module; The convolution module includes two 3x3 convolution layers, one ReLU activation function layer, and two batch normalization BatchNorm layers; The channel attention module is used to: First, two transformations are performed on the input feature map to generate two auxiliary feature maps. Next, a 1x1 convolution kernel is used to perform channel compression on the input feature map to reduce the number of channels to in_planes / ratio, where in_planes is the number of channels entered and ratio is a parameter that controls the channel compression ratio. Next, a nonlinear transformation is performed using the ReLU activation function, and then the number of channels is restored to in_planes using a 1x1 convolution kernel again. Finally, the results of the average pooling and the max pooling are summed, and the output of the channel attention is scaled by a sigmoid layer to obtain the final channel attention weights, which are used to weight the channels of the input feature maps; The spatial attention module is used to: First, the input feature maps are first sent to average pooling and max pooling operations to calculate the average and max feature maps, respectively; Then, these two feature maps are serially connected to form a feature map containing two channels; This feature map is then passed through a 7x7 convolutional layer to compute the spatial attention weights. Finally, a sigmoid layer is used to scale the output of the spatial attention to obtain a final spatial attention weight graph.
6. The optimization reinforcement module is a recurrent four-layer U-net network module, which is composed of four encoders, a decoder, and an output layer; The encoder includes an input layer, a downsampling module, and a skip connection. The input layer is initially used to accept input data; The downsampling module is composed of a 3x3 convolution layer, a pooling layer and an activation function, which is used to gradually reduce the size and number of channels of the feature map and extract high-level semantic features; The skip connections are used to preserve the feature maps of each layer of the encoder for feature fusion at the decoder; The encoder includes an upsampling module, a post-upsampling processing module, a skip connection, and an output layer; The upsampling module is composed of a deconvolution layer and an activation function, and is used to gradually increase the size and number of channels of the feature map to restore it to the original input size; The skip connection is used to connect the feature map of the corresponding layer in the encoder and the feature map of the decoder to realize feature fusion; The post-upsampling processing module is composed of a convolution layer, a batch normalization layer and an activation function, and is used to further process the upsampling features; The method of claim 1 , wherein the output layer is used to combine low-level and high-level features to generate a final completion result.
7. The step of training the depth map completion network model includes: Obtaining a depth map completion network training dataset from the Kitti dataset, which is used in publicly available computer vision and autonomous driving research, including GroundTruth with high labeling information coverage and paired sparse depth maps and color RGB images; training a network using the RGB image and the sparse depth map as input data for the network; Calculating a loss between the completed depth map output from the network and the complete depth map of the GroundTruth in the training dataset, and performing backpropagation using the loss to update the weights of the network; The method of claim 1, further comprising: minimizing the loss value of the interpolated depth map obtained by the depth map completion network and the actual complete depth map using gradient descent to obtain an optimal model.
Citation Information
Patent Citations
Point cloud image depth completion method based on cascade feature interaction
CN115511759A
RGB-D semantic segmentation method and system based on adaptive context sensing network
CN116580192A
Depth completion method based on sparse representation
CN116862965A
Image-enhanced depth sensing using machine learning
JP2021517685A
Depth map completion method based on deformable convolution
CN113538278A