A method for estimating depth information from images

CN116468769BActive Publication Date: 2026-09-01BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310217308.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2026-09-01
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

基于图像的深度估计方法可以分为单目深度估计和多目深度估计,多目深度估计通常需要两个摄像头拍摄的同一个场景的两张图像,通过一致的相机参数-基线和焦距,基于立体视觉技术对两幅图像进行匹配从而获取深度信息,但是当场景中的纹理较少或没有时,很难在图像中捕捉到足够的特征来进行匹配,所以局限性较大

Benefits of technology

[0055]本发明可以在网络训练完成的前提下预测出更精确的深度信息,对比已有深度学习技术,可以充分利用输入图像的局部相关性和远程关系依赖提升低纹理区域的预测效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468769B_ABST
    Figure CN116468769B_ABST
Patent Text Reader

Abstract

This invention provides an image-based depth information estimation method, comprising: inputting an unlabeled image sequence of the same scene into a deep neural network to extract image features; sequentially using channel attention and spatial attention mechanisms to adaptively optimize the image features; using bilinear interpolation for upsampling to restore image resolution; using the restored feature image as the target image to predict depth information, and reconstructing the target image based on the predicted depth information and adjacent frames; calculating the photometric error and smoothing error between the target image and the reconstructed image at multiple scales to obtain a loss function; performing unsupervised model training, updating the model parameters according to the loss function to obtain the trained model; and using the trained model to predict depth information of the input scene image. This invention can fully utilize the local correlation and long-range dependency of the input image to improve the prediction effect in low-texture regions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent depth estimation technology and relates to a method for predicting corresponding depth information based on an image. Background Technology

[0002] Ordinary cameras, when capturing images, can only record the color information of a scene. When projected into a 2D image from 3D space, the distance between the scene and the camera—that is, depth information—is lost. Acquiring scene depth information is a crucial branch of computer vision and a vital component in applications such as 3D reconstruction, autonomous driving, and robot localization. More specifically, for each pixel in a given RGB image, we need to estimate a metric depth value. Traditional methods for acquiring depth information rely on hardware devices. The most common device is LiDAR, which estimates depth information by measuring the time it takes for laser light to reflect off an object's surface. However, LiDAR devices are expensive, and acquiring high-precision, dense depth information requires significant manpower, making widespread application in real-world scenarios difficult. Another common hardware device is a depth camera. Depth cameras obtain scene depth information based on Time-of-Flight (TOF) technology. They continuously send light pulses to a target and then use sensors to receive the light returning from the object. By detecting the round-trip time of these emitted and received light pulses, the distance to the target is determined. The sensor calculates the distance to the photographed object by measuring the time difference or phase difference between light emission and reflection, thus generating depth information. In addition, when combined with traditional camera shooting, the three-dimensional outline of the object can be presented as a topographic map with different colors representing different distances. However, due to the short range of its ranging sensor and the high requirements of the scene environment, its application range in outdoor environments is limited.

[0003] Compared to traditional hardware-based measurement methods, image-based depth estimation methods only require image capture and have lower hardware requirements, thus offering greater application value in real-world scenarios. Image-based depth estimation methods can be divided into monocular depth estimation and multi-view depth estimation. Multi-view depth estimation typically requires two images of the same scene captured by two cameras. By using consistent camera parameters—baseline and focal length—the two images are matched using stereo vision techniques to obtain depth information. However, when there is little or no texture in the scene, it is difficult to capture enough features in the images for matching, thus limiting its effectiveness. Monocular depth estimation, on the other hand, uses only one camera to acquire images or video sequences, requiring no additional complex equipment or specialized techniques. In most cases, depth estimation can be achieved with just one camera, thus possessing broad application value and significant research importance.

[0004] Therefore, how to provide a depth information estimation method based on monocular images is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention proposes an image-based depth information estimation method to solve the technical problems in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] This invention discloses an image-based depth information estimation method, including a model construction step and a depth information prediction step:

[0008] The model construction steps include:

[0009] S1: Input an unlabeled image sequence of the same scene into a deep neural network to extract image features, including local and global features.

[0010] S2: Adaptive feature optimization is performed on the image features by sequentially using channel attention mechanism and spatial attention mechanism.

[0011] S3: Upsample the optimized image features using bilinear interpolation to restore the image resolution.

[0012] S4: Perform depth information prediction on the feature image recovered in S3 as the target image, calculate the relative pose change of the current feature image; reconstruct the target image using the depth information and the relative pose change, i.e., reconstruct the image.

[0013] S5: Calculate the photometric error and smoothing error of the target image and the reconstructed image at multiple scales, and further obtain the loss function;

[0014] S6: Repeat S2-S6 to train the unsupervised model, update the model parameters according to the loss function, and obtain the trained model;

[0015] The step of predicting depth information includes:

[0016] S7: Use the trained model to predict depth information from the input scene image.

[0017] Preferably, S1 includes:

[0018] S11: Input an unlabeled image sequence of the same scene into a deep network and divide the images into patches of the same size;

[0019] S12: Use CNN to extract local features and Transformer to extract global features;

[0020] S13: The local features and the global features are concatenated and then output through convolution.

[0021] Preferably, S2 includes:

[0022] S21: Calculate the dependencies between different channels using the channel attention mechanism for the image features and obtain the corresponding attention weights, then output the channel attention map;

[0023] S22: Utilize spatial attention mechanism to enhance attention to key regions of the channel attention map and extract key information to obtain optimized image features.

[0024] Preferably, the specific execution steps of S2 include:

[0025] The image features are spatially compressed using a max pooling layer and an average pooling layer to obtain two tensors.

[0026] The tensor is fed into the multilayer perceptron to output intermediate features;

[0027] The intermediate features are summed and then processed using a sigmoid function to obtain a channel attention map.

[0028] The channel attention map is passed sequentially through a max pooling layer and an average pooling layer to obtain a tensor 2;

[0029] Spatial attention is calculated by passing the tensor 2 through convolutional layers and sigmoid to obtain optimized image features.

[0030] Preferably, S3 includes:

[0031] The optimized image features are linearly interpolated sequentially in the x and y directions, and scale recovery is performed through upsampling.

[0032] Preferably, S4 includes:

[0033] S41: Deep Network Receives Target View I t As input, predict depth maps d corresponding to n scales, where n≥4;

[0034] S42: The pose network will display the target view I t and adjacent frame source view I t-1 ,I t+1 As input, and output the relative pose change T t→t' t'∈{t-1,t+1};

[0035] S43: Based on the assumption that the shooting scene is static and the changes in view are only caused by the movement of the camera, the target image can be reconstructed using the source view of adjacent frames, depth map and pose changes.

[0036] Preferably, the specific execution steps of S41 include:

[0037] Deep networks are used to predict depth maps; deep networks consist of encoders and decoders.

[0038] The encoder is used to extract features from the input image. It consists of multiple encoder blocks. After each encoder block, the size of the image is reduced to half of the input size.

[0039] The decoder is used to scale the extracted features and output depth maps of different sizes to construct multi-scale features. The decoder block uses upsampling to restore the size. The output of each decoder block is twice the input. The input of the decoder consists of two parts: the first part comes from the output of the decoder in the previous stage, and the second part comes from the output of the encoder block.

[0040] Preferably, the specific execution steps of S43 include:

[0041] I t'→t =I t' [proj(reproj(I t ,d,T t→t' ),K)]

[0042] T t→t' =Θ pose (I t ,I t' ),t∈{t-1,t+1}

[0043] Among them, I t'→t To reconstruct the image, K is a known intrinsic parameter of the camera, [] is the sampling operator, reproj returns the 3D point cloud of camera t', and proj output projects the point cloud onto I. t' 2D coordinates, T t→t' For relative pose changes, Θ pose For posture networks.

[0044] Preferably, S5 includes:

[0045] Structural similarity (SSIM) is used to calculate the similarity between the reconstructed image and the target image;

[0046] The photometric error is obtained by superimposing the L1 norm on the aforementioned similarity. p (I t ,I t'→t );

[0047] The smoothing error l is obtained by weighting the depth information using the image gradient. smooth (d);

[0048] The operation is repeated at n scales to obtain photometric error and smoothing error, and their weighted sum is calculated to obtain the loss function, where n≥4:

[0049]

[0050] Where u is the mask. For minimum photometric loss, β is the sum of photometric loss and smoothing loss. smoot Weighting coefficients between h;

[0051] u = [min(l p (I t ,I t'→t )) <min(l p (I t ,I t' ))]t'∈{t-1,t+1}

[0052]

[0053] Preferably, the automatic masking method is used to ignore pixels in the image sequence that do not change between adjacent frames, and the mask is set to binary.

[0054] As can be seen from the above technical solution, compared with the prior art, the beneficial effects of the present invention include:

[0055] This invention can predict more accurate depth information after the network training is completed. Compared with existing deep learning techniques, it can make full use of the local correlation and long-range relation dependence of the input image to improve the prediction effect of low-texture regions. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0057] Figure 1 A flowchart illustrating an image-based depth information estimation method provided in one embodiment of the present invention;

[0058] Figure 2 This is a schematic diagram of training image sequence data provided in one embodiment of the present invention;

[0059] Figure 3 This is a schematic diagram of an extracted feature image sequence provided in one embodiment of the present invention;

[0060] Figure 4 This is a schematic diagram illustrating the comparison of depth information prediction provided in one embodiment of the present invention;

[0061] Figure 5This is a network architecture diagram of an image-based depth information estimation method provided in one embodiment of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] like Figure 1 As shown, this embodiment of the invention provides an image-based depth information estimation method, including a model construction step and a depth information prediction step:

[0064] The steps for building a model include:

[0065] S1: Input an unlabeled image sequence of the same scene into a deep neural network to extract image features, including local and global features.

[0066] In this step, image sequences are used to build a dataset. The training dataset consists of multiple images taken by the same camera, used for unsupervised training of the model. The input images are three adjacent frames of the same scene. See the appendix for a specific example. Figure 2 The image shown is a schematic diagram of the training image sequence data, where I t I t-1 I t+1 It is the image sequence input to S1.

[0067] In one embodiment, the image feature extraction steps are as follows:

[0068] The image is input into a deep network model, which divides the image into patches of the same size.

[0069] The input image is encoded using CNN and Transformer to obtain the extracted local correlations and long-range dependencies;

[0070] The aforementioned visual features are concatenated and stitched together, and then output as features through convolution.

[0071] In practice: For image features, the MpViT model pre-trained on ImageNet is used to divide the image into several patches, and the local feature extraction capabilities of CNN and the global feature extraction capabilities of Transformer are used to achieve effective feature extraction.

[0072] Among them, the CNN module is used to extract local features L∈R H×W×CThe Transformer module is used to extract global features G∈R. H×W×C .

[0073] Finally, the local feature L and the global feature G are concatenated and then output through convolution. As shown below:

[0074] X = Concat([L,G])

[0075] X′=H(X)

[0076] Here, X represents the features concatenated from local and global features, and H(·) represents the learning function that maps the concatenated features X to the final features X′. The H(·) function is implemented using a 1×1 convolution.

[0077] like Figure 3 As shown, this is the feature image after feature extraction from the input image sequence. Besides qualitative evaluation metrics and visualized depth images, intermediate feature maps are also a visual indicator of a depth estimation model's information extraction capabilities. In this embodiment, the intermediate feature map is overlaid on the input image for observation. Figure 4 As shown, from bottom to top, the images are: input image, feature map of this embodiment, feature map of other methods, predicted depth map of this embodiment, and predicted depth map of other methods. It can be seen that the method of this invention is able to extract more object details, thus obtaining a clearer depth map.

[0078] S2: Adaptive feature optimization of image features is performed by sequentially using channel attention mechanism and spatial attention mechanism.

[0079] In one embodiment, the purpose of the channel attention module is to calculate the dependencies between different color channels and obtain the corresponding attention weights. The input features are spatially compressed through a max pooling layer (MaxPool) and an average pooling layer (AvgPool) to obtain two tensors, denoted as Tensor 1. These are then fed into a multilayer perceptron (MLP). Finally, the channel attention map is passed through a max pooling layer and an average pooling layer sequentially to obtain a tensor, denoted as Tensor 2. Tensor 2 is summed and passed through a sigmoid function to obtain the channel attention map Att. C The spatial attention module aims to enhance attention to important regions to extract key information. The input features are sequentially processed through max pooling and average pooling layers, and finally through convolutional layers and a sigmoid function to obtain the spatial attention map Att. S The final characteristic can be described as Y = Att S (Att C (X)), where X and Y represent the input features and output features, respectively.

[0080] S3: Upsample the optimized image features using bilinear interpolation to restore the image resolution.

[0081] In one embodiment, the input feature map is upsampled using bilinear interpolation to restore resolution, which can be achieved using the torch.nn.functional.grid_sample function.

[0082] S4: Use the feature image recovered in S3 as the target image to predict depth information and calculate the relative pose change of the current feature image; use the depth information and relative pose change to reconstruct the target image, i.e., reconstruct the image, and train the network in an unsupervised manner.

[0083] In one embodiment, a deep network is first used to predict the depth information of the target image. A deep network Θdepth is designed based on an autoencoder architecture, consisting of an encoder and a decoder. The encoder extracts features from the input image and consists of five encoder modules. The decoder is responsible for scale recovery of the extracted features and outputting depth maps of different sizes, constructing multi-scale features. After each encoder module, the size of the feature map is reduced to half of the input. The encoder contains five encoder blocks, each with an output size of (H / 2, W / 2), (H / 4, W / 4), (H / 8, W / 8), (H / 16, W / 16), and (H / 32, W / 32). The decoder blocks use upsampling to recover the size, and the output of each decoder block is twice the input. The decoder input consists of two parts: the first part comes from the output of the previous decoder, and the second part corresponds to the encoder output. The details of the decoder output feature map are enhanced by fusing feature maps of different scales.

[0084] The decoder predicts the corresponding depth map d, while the pose network Θ... pose Target View I t and nearby source view I' t Given t'∈{t-1,t+1} as input, output the relative pose change T. t→t' =Θ pose (I t ,I t' ), t∈{t-1,t+1}. Based on the assumption that the shooting scene is static and the changes in the view are only caused by camera movement, the source view I can be used. t' The target view I is reconstructed from the pixels t'∈{t-1,t+1}. t , target view I t and adjacent frame source view I' t Let t'∈{t-1,t+1} be the input. The construction here can be summarized as the following formula:

[0085] I t'→t =I t' [proj(reproj(I t ,d,T t→t' ),K)]

[0086] Where d represents the predicted depth information, and T t→t' Represents pose change, K is the known intrinsic parameters of the camera, [] is the sampling operator, reproj returns the 3D point cloud of camera t', and proj output projects the point cloud onto I. t' The 2D coordinates are used to obtain the reconstructed image I. t'→t .

[0087] S5: Calculate the photometric error and smoothing error of the target image and the reconstructed image at multiple scales, and further obtain the loss function.

[0088] In one embodiment, given an input target image I t and reconstructed image I t'→t The structural similarity index measure (SSIM) is used to calculate the similarity between the reconstructed image and the target image, and then the L1 norm is added to obtain the photometric error. Where α is the weighting parameter, which was set to 0.85 in the experiment.

[0089] Since depth discontinuities often occur at image gradients, an L1 penalty on the disparity gradient is used to encourage local disparity smoothing, resulting in a smoothing error:

[0090]

[0091] in, and These represent the depth gradients in the x and y directions, respectively. d t Is with I t The corresponding depth value.

[0092] To prevent getting stuck in local minima during training, the photometric error and smoothing error are calculated as a weighted sum in the form of multi-scale errors.

[0093] In one embodiment, pixels that do not change between adjacent frames in a sequence are ignored using automatic masking techniques. The mask u is set to binary:

[0094] u = [min(l p (I t ,I t'→t )) <min(l p (I t,I t' ))]t'∈{t-1,t+1}

[0095] The final error is obtained by multiplying it by the photometric loss.

[0096] This embodiment uses a multi-scale loss function, which consists of two parts: photometric error and smoothing error. To minimize the loss function and achieve the goal of predicting high-precision depth information, the loss function is designed as follows:

[0097]

[0098] Where u is the mask. To minimize luminous loss,

[0099]

[0100] β represents the photometric loss and smoothing loss. smooth The weighting coefficients between them.

[0101] In this embodiment, the network model will produce outputs at four scales during training. The output sizes are 1 / 8, 1 / 4, 1 / 2, and 1 of the input protrusion size, respectively. Therefore, our final loss is the average of the losses at the four scales.

[0102] S6: Repeat S2-S5 to train the unsupervised model, update the model parameters according to the loss function, and obtain the trained model.

[0103] This embodiment employs unsupervised learning, eliminating the need for real depth values ​​as supervisory signals during training. Each input image consists of three adjacent frames. During iterations S2 and S5, a portion of the data is processed based on the available GPU space on the training machine. The model parameters are updated using backpropagation based on the loss function in each loop. Each iteration consists of feeding all data into the network, and training continues until a specified number of iterations is reached. In this embodiment, the number of images processed per training iteration is set to 16, and training lasts for 22 epochs. A dynamic learning rate is used to prevent learning instability; the learning rate is initialized to 1×10⁻⁶. -4 The loss is reduced to half its original value over the next 18 epochs, with the weight β for smoothing loss set to 0.001. After each training epoch, the model is validated on a validation set. Finally, after the chain is completed, the model that performs best on the validation set is selected as the final training result.

[0104] The steps for predicting depth information include:

[0105] S7: Use the trained model to predict depth information from the input scene image.

[0106] In one embodiment, testing the model simulates its performance in real-world use. Images outside the training set are selected as input to the model, and the model's output is the depth information for each pixel of that image.

[0107] To demonstrate the superior predictive performance of this invention, the following explanation is provided in conjunction with specific image prediction results:

[0108] Figure 4 The top image shows the input image sequence, and the bottom image shows the depth information predicted from the input images, displayed as a depth map after color visualization. In the image, the darker the color, the farther the distance; the more yellow the color, the closer the distance. It can be seen that this invention performs better in the details and edge features of objects, such as the lampposts of streetlights and the canopies of trees, and can clearly predict their outlines.

[0109] The image-based depth information estimation method provided by the present invention has been described in detail above. Specific examples have been used in this embodiment to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of ​​the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

[0110] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined in these embodiments may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A depth information estimation method based on images, characterized in that, This includes the steps of building the model and predicting depth information: The model construction steps include: S1: Input an unlabeled image sequence of the same scene into a deep neural network to extract image features, including local and global features, including: S11: Input an unlabeled image sequence of the same scene into a deep network and divide the images into patches of the same size; S12: Use CNN to extract local features and Transformer to extract global features; S13: The local features and the global features are concatenated and then output through convolution; S2: Adaptive feature optimization of the image features is performed sequentially using channel attention mechanism and spatial attention mechanism; S3: Upsample the optimized image features using bilinear interpolation to restore the image resolution; S4: Input the feature image restored in S3 as the target image into the deep network to predict depth information; input the feature image restored in S3 as the target image into the pose network to calculate the relative pose change of the current feature image; reconstruct the target image using the depth information and the relative pose change, i.e., reconstruct the image. S5: Calculate the photometric error and smoothing error of the target image and the reconstructed image at multiple scales, and further obtain the loss function; S6: Repeat S2-S6 to train the unsupervised model, update the parameters of the deep network and the pose network according to the loss function, and obtain the trained model; The step of predicting depth information includes: S7: Use the trained model to predict depth information from the input scene image.

2. The image-based depth information estimation method according to claim 1, characterized in that, S2 includes: S21: Calculate the dependencies between different channels using the channel attention mechanism for the image features and obtain the corresponding attention weights, then output the channel attention map; S22: Utilize spatial attention mechanism to enhance attention to key regions of the channel attention map and extract key information to obtain optimized image features.

3. The image-based depth information estimation method according to claim 1, characterized in that, The specific execution steps of S2 include: The image features are spatially compressed using a max pooling layer and an average pooling layer to obtain two tensors. The tensor is fed into the multilayer perceptron to output intermediate features; The intermediate features are summed and then processed using a sigmoid function to obtain a channel attention map. The channel attention map is passed sequentially through a max pooling layer and an average pooling layer to obtain a tensor 2; Spatial attention is calculated by passing the tensor 2 through convolutional layers and sigmoid to obtain optimized image features.

4. The image-based depth information estimation method according to claim 1, characterized in that, S3 includes: The optimized image features are linearly interpolated sequentially in the x and y directions, and scale recovery is performed through upsampling.

5. The image-based depth information estimation method according to claim 1, characterized in that, S4 includes: S41: Deep network receives target view As input, predict depth maps d corresponding to n scales, where n≥4; S42: The pose network will display the target view. and adjacent frame source view As input, output relative pose change ; S43: Reconstruct the target image using adjacent frame source views, depth maps, and pose changes.

6. The image-based depth information estimation method according to claim 5, characterized in that, The specific execution steps of S41 include: Using deep networks to view the target To perform depth information prediction and obtain a depth map, a deep network includes an encoder and a decoder. The encoder is used to extract features from the input image. It consists of multiple encoder blocks. After each encoder block, the size of the image is reduced to half of the input size. The decoder is used to scale the extracted features and output depth maps of different sizes to construct multi-scale features. The decoder block uses upsampling to restore the size. The output of each decoder block is twice the input. The input of the decoder consists of two parts: the first part comes from the output of the decoder in the previous stage, and the second part comes from the output of the encoder block.

7. The image-based depth information estimation method according to claim 5, characterized in that, The specific execution steps of S43 include: ; ; in, To reconstruct the image, K is a known intrinsic parameter of the camera, and [] represents the sampling operator. Return to camera 3D point cloud, The output projects the point cloud onto 2D coordinates For relative pose changes, For posture networks.

8. The image-based depth information estimation method according to claim 5, characterized in that, S5 includes: Structural similarity (SSIM) is used to calculate the similarity between the reconstructed image and the target image; The photometric error is obtained by superimposing the L1 norm on the aforementioned similarity. ; The smoothing error is obtained by weighting depth information using image gradients. ; The operation is repeated at n scales to obtain photometric error and smoothing error, and their weighted sum is calculated to obtain the loss function: ; in, For the mask, To minimize luminous loss, It is the loss of photometric properties and the loss of smoothness. The weighting coefficients between them; ; 。 9. The image-based depth information estimation method according to claim 8, characterized in that, The automatic masking method ignores pixels in the image sequence that do not change between adjacent frames, and sets the mask to binary.

Citation Information

Patent Citations

  • Monocular depth estimation system and method for enhancing feature fusion in three-dimensional scene reconstruction

    CN115294282A

  • Road crack real-time detection method fusing CNN and Tannformer

    CN115690042A