A method for depth estimation of indoor scene monocular images based on deep learning
Through the multi-stage encoding feature extraction and parallel partition prediction deep information, the problem of deep loss of fine-grained information and difficulty in considering global local relationships in convolutional neural networks is solved, and high-precision monocular depth estimation is achieved.
Patent Information
- Application Number
- CN202210251724.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-03-15
AI Technical Summary
In monocular depth estimation encoding networks using convolutional neural networks, fine-grained image information is easily lost at deep levels, and it is difficult to effectively consider global and local relationships in complex scenarios, resulting in low accuracy of depth estimation.
A multi-stage encoding feature extraction method is adopted, and parallel partitions are designed to predict global to local depth information separately in the decoding network. Through specific loss functions, model training is optimized to ensure that the model can effectively consider the global and local relationships in the scenario.
Effectively utilizing the fine-grained information of the input image improves the accuracy of depth estimation. Especially in complex scenarios, depth information can be predicted more accurately without a large amount of labeled data to achieve higher results.
Smart Images

Figure CN114638870B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for estimating the depth of an indoor scene monocular image based on deep learning, and belongs to the technical field of three-dimensional scene perception. Background Art
[0002] Depth estimation from 2D RGB images has a wide range of applications, such as 3D reconstruction, scene understanding, autonomous driving, robotics, etc. With the emergence of large-scale datasets and the improvement of hardware computing power, recent research on image depth estimation has focused on 2D to 3D reconstruction using deep learning and convolutional neural networks. Depth estimation from a single RGB image is an ill-posed problem because a single image can correspond to an infinite number of 3D scenes. In addition, problems such as lack of scene coverage, semi-transparent or reflective materials may lead to ambiguous situations where geometry cannot be inferred from appearance.
[0003] The method of monocular depth estimation based on deep learning began with the dual-scale network proposed by Eigen et al. Then some researchers proposed many effective methods based on deep learning using convolutional neural networks. The paper "Laina et al., Deeper Depth Prediction with Fully Convolutional Residual Networks" uses a fully convolutional residual network based on ResNet-50 and replaces the fully connected layer with a series of upsampling blocks. The paper "Alhashim et al., High Quality Monoocular Depth Estimation via Transfer Learning" introduces skip connections in a simple encoder-decoder network architecture and trains the model using transfer learning. The paper "Lee et al., FromBig to Small: Multi-Scale Local Planar Guidance for Monocular DepthEstimation" proposes to replace the standard upsampling layer with a local plane guidance layer to guide the features to full resolution in the decoder. The paper "Fu et al., Deep ordinal regression network for monocular depthestimation" found that the performance can be improved if the depth regression task is converted into a classification task. The paper "Bhat et al., AdaBins: Depth Estimation using Adaptive Bins" designs the AdaBins module, which divides the depth value range into 256 intervals. The center value of each interval is the depth value of the pixel that falls within the interval. The final depth of a pixel is the value of the linear combination interval of the center depth. The paper "Ranftl et al., Vision Transformers for DensePrediction" applies Vision Transformer to monocular depth estimation and obtains a high-precision depth estimation model through training with a large dataset.
[0004] Although there has been great progress in indoor monocular image depth estimation based on deep learning, there are still some problems: 1) In most of the codec structures used in deep learning neural networks, the encoder will suffer from insufficient feature extraction and spatial information loss due to operations such as explicit downsampling during the feature extraction stage, which makes the network easily lose the fine-grained information of the image; 2) The actual scene structure faced by monocular depth estimation in indoor scenes is usually more complex. If the global and local relationships in the scene are not effectively considered, the accuracy of depth estimation will be very low; 3) Although the emergence of VisionTransformer can greatly improve the problem of image granularity loss, its model parameters are large and requires a large amount of labeled data to drive training. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method for depth estimation of indoor scene monocular images based on deep learning. In order to solve the problem that fine-grained information of images is easily lost in the deep layers of the monocular depth estimation encoding network using convolutional neural networks, the features of multi-stage encoding are comprehensively used. In the decoding network, in view of the difficulty of traditional networks in effectively considering the global and local relationship in the scene in complex scenes, a parallel partition is designed to predict the global to local depth information respectively, and then adjust the fused decoder, and a corresponding loss function is designed to solve the above problem.
[0006] The technical solution of the present invention is: a method for indoor scene monocular image depth estimation based on deep learning, the specific steps are:
[0007] Step 1: Introduce the neural network EfficientNet-b7 pre-trained for image classification on ImageNet and construct an encoder.
[0008] Step 2: At different stages of the encoder, SENet-based residual connections and convolution and resampling operations are introduced to obtain predictions at different stages.
[0009] Step 3: Based on the depth interval division method, a loss function focusing on the global to local aspects of the image is constructed and applied to predictions at different stages.
[0010] Step 4: Use the Transformer structure based on the self-attention mechanism to fuse the depth information predicted at different stages and output the scene depth prediction result.
[0011] The specific step of Step 1 is as follows: download the EfficientNet-b7 network pre-trained on ImageNet from the Internet, and obtain the feature vectors encoded in the 3rd, 5th, 6th, 8th, and 12th blocks. The resolutions of these feature vectors are respectively
[0012] The Step 2 is specifically as follows:
[0013] Step 2.1: Input the feature vector encoded by the 3rd block into 4 residual blocks based on SENet, the feature vector encoded by the 5th block into 3 residual blocks based on SENet, the feature vector encoded by the 6th block into 2 residual blocks based on SENet, and the feature vector encoded by the 8th block into 1 residual block based on SENet.
[0014] Step 2.2: Add a channel attention layer after the last residual block in each stage and add a residual connection from the encoder to this layer.
[0015] Step 2.3: The features of each stage are gradually passed through double upsampling and convolution layers to obtain five stages of features with the same number of channels (30) and the same resolution (half of the input resolution).
[0016] Step 2.4: Add and fuse the features of the 1st, 2nd and 5th stages pixel by pixel, add and fuse the features of the 2nd, 3rd and 5th stages pixel by pixel, add and fuse the features of the 1st, 3rd and 4th stages pixel by pixel, and then pass through the convolution layer to get four predictions, which are marked as prediction 1 to prediction 4 from shallow to deep of the neural network.
[0017] This fusion selection can predict the local and global from shallow to deep, and then the features of the first stage are used as references for the next two predictions, and the features of the fifth stage are used as references for the above two predictions, so as to better complete the global and local predictions; when fusing, the rich spatial information contained in the shallow features can also improve the accuracy of the fusion. Compared with the traditional method of serial fusion step by step and then output, it not only improves efficiency, but also improves accuracy.
[0018] The Step 3 is specifically as follows:
[0019] Step 3.1: Get the maximum depth d_max and minimum depth d_min from the real depth map.
[0020] Step 3.2: Divide the depth interval [d_min, d_max] into 10 small intervals evenly. The calculation formula for the length of a small interval is as follows:
[0021]
[0022] Among these 10 intervals, the depth value range of the i-th interval is calculated as follows:
[0023] [d_min+(i-1)×len, d_min+i×len]
[0024] Step 3.3: Make a histogram for the real depth map to find the interval with the largest proportion of scene depth among the 10 intervals. This interval contains most of the global information. Correspondingly, the interval with a smaller proportion contains more local information.
[0025] Step 3.4: Arrange the 10 depth intervals in descending order according to their proportions, and calculate the mean square error of prediction 1 in Step 2.4 in the 5th to 10th intervals, the mean square error of prediction 2 in the 4th to 8th intervals, the mean square error of prediction 3 in the 2nd to 4th intervals, and the mean square error of prediction 4 in the 1st and 2nd intervals.
[0026] Step 3.5: Combine the four errors as a loss term that constrains predictions 1 to 4 to focus on the local to global aspects during model training. The calculation formula is as follows:
[0027]
[0028] Among them, λ1=0.5, λ2=λ3=0.6, λ4=1, n i is the total number of pixels in the true depth map after the interval mask, and They are the real depth map and the pixel p in the prediction i. i The depth value of .
[0029] Through the depth interval partitioning method used in Step 3, a specific loss function is obtained as a supplementary constraint for partial prediction. Applying it to predictions at different stages can gradually separate global and local predictions into different stages during model training, so that the predictions at each stage of the model can focus on the depth interval that can be predicted more accurately at that stage.
[0030] The Step 4 is specifically as follows:
[0031] Step 4.1: Concatenate the prediction results of the four stages into a four-channel tensor
[0032] Step 4.2: Transform the four-channel tensor Perform a convolution operation with a convolution kernel of 16×16, a step size of 16, and an output channel of 4, that is:
[0033]
[0034] Step 4.3: Flatten the two-dimensional tensor obtained after convolution into one dimension, that is:
[0035]
[0036] Step 4.4: Input the one-dimensional tensor into the Transformer Encoder and restore its output one-dimensional tensor to a two-dimensional tensor as the weight matrix
[0037] Step 4.5: Convert the four-channel tensor Perform a convolution operation with a convolution kernel of 3×3, a stride of 1, and an output channel of 128, and the shape is Tensor
[0038] Step 4.6: Weight Matrix With tensor After performing pixel-by-pixel dot product operations, the final prediction result is output through a series of convolutional layers.
[0039] Compared with the traditional channel splicing and convolution method, the fusion method adopted in Step 4 can greatly improve the accuracy of model prediction without introducing more parameters.
[0040] The beneficial effects of the present invention are:
[0041] (1) In order to make full use of the fine-grained information of the input image, the present invention extracts feature vectors from multiple stages of the encoder to solve the problem that fine-grained information is easily lost in the deep part of the encoding network in the traditional method;
[0042] (2) The present invention designs a decoder that predicts global to local depth information between parallel partitions and then adjusts the fusion, and designs a corresponding loss function to optimize the problem that traditional networks are difficult to effectively consider the global and local relationships in complex scenes.
[0043] (3) The present invention uses a convolutional neural network as its basis and can achieve more accurate results without the need for a large data set to drive training.
[0044] (4) The present invention does not require a very large data set to drive training, and can effectively improve the accuracy of monocular depth estimation.
[0045] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, those skilled in the art can be guided by the practice of the present invention based on the following investigation and study. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flow chart of the steps of the present invention;
[0047] Figure 2 This is a schematic diagram comparing the depth map predicted by the monocular depth estimation network used in the present invention and the current most advanced networks AdaBins and DPT-Hybrid in certain scenarios. In the figure:
[0048] (a) is the input RGB image;
[0049] (b) is the real depth map;
[0050] (c) is the depth map predicted by AdaBins;
[0051] (d) is the depth map predicted by DPT-Hybrid;
[0052] (e) is the depth map predicted by the present invention;
[0053] Figure 3 This is an example diagram of the present invention generating a three-dimensional point cloud from a single RGB image by predicting depth. DETAILED DESCRIPTION
[0054] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0055] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0056] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0057] like Figure 1 As shown in FIG. 1 , a method for estimating the depth of a monocular image of an indoor scene based on deep learning is described. The specific steps are as follows:
[0058] In the feature extraction stage of the encoder, the feature vectors encoded by the 3rd, 5th, 6th, 8th, and 12th blocks in the EfficientNet-b7 encoder are extracted, and the shapes are Where H and W are the height and width of the input image respectively.
[0059] Then the feature vector encoded by the third block is input into 4 residual blocks based on SENet, the feature vector encoded by the fifth block is input into 3 residual blocks based on SENet, the feature vector encoded by the sixth block is input into 2 residual blocks based on SENet, and the feature vector encoded by the eighth block is input into 1 residual block based on SENet.
[0060] The next step is to add a channel attention layer after the last residual block of each stage, and add a residual connection from the encoder to this layer to construct a large residual block. The features of each stage are then gradually passed through double upsampling and convolution layers to obtain five stages with the same number of channels (30) and the same resolution (half of the input resolution).
[0061] In the next step, the features of stages 1, 2, and 5 are added and fused pixel by pixel, the features of stages 2, 3, and 5 are added and fused pixel by pixel, the features of stages 1, 3, and 4 are added and fused pixel by pixel, and the features of stages 1, 4, and 5 are added and fused pixel by pixel, and then four predictions are obtained through the convolution layer, which are marked as prediction 1 to prediction 4 from shallow to deep of the neural network.
[0062] Next, a loss function focusing on the local to global depth is designed for predictions 1 to 4.
[0063] First, obtain the maximum depth d_max and the minimum depth d_min from the real depth map, and then divide the depth interval [d_min, d_max] into 10 small intervals evenly. The calculation formula for the length of a small interval is as follows:
[0064]
[0065] Among these 10 intervals, the depth value range of the i-th interval is calculated as follows:
[0066] [d_min+(i-1)×len, d_min+i×len]
[0067] Then, a histogram is made for the true depth map to find the interval with the largest proportion of scene depth among the 10 intervals. This interval contains most of the global information. Correspondingly, the interval with a smaller proportion contains more local information.
[0068] The next step is to sort the 10 depth intervals in descending order according to their proportions, and calculate the mean square error of prediction 1 in the 5th to 10th intervals, the mean square error of prediction 2 in the 4th to 8th intervals, the mean square error of prediction 3 in the 2nd to 4th intervals, and the mean square error of prediction 4 in the 1st and 2nd intervals.
[0069] The next step is to combine the four errors as a loss term to constrain predictions 1 to 4 to focus on the local to global during model training. The calculation formula is as follows:
[0070]
[0071] Among them, λ1=0.5, λ2=λ3=0.6, λ4=1, n i is the total number of pixels in the true depth map after the interval mask, and They are the real depth map and the pixel p in the prediction i. i The depth value of .
[0072] After obtaining predictions 1 to 4, the four predictions need to be fused.
[0073] First, concatenate the prediction results of the four stages into a four-channel tensor Then take this four-channel tensor Perform a convolution operation with a convolution kernel of 16×16, a step size of 16, and an output channel of 4, that is:
[0074]
[0075] Then the two-dimensional tensor obtained after convolution is flattened into one dimension, that is:
[0076]
[0077] The next step is to input the one-dimensional tensor into the Transformer Encoder and restore its output one-dimensional tensor to a two-dimensional tensor as the weight matrix
[0078] The next step is to convert the four-channel tensor Perform a convolution operation with a convolution kernel of 3×3, a stride of 1, and an output channel of 128, and the shape is Tensor
[0079] Finally, the weight matrix With tensor After performing pixel-by-pixel dot product operations, the final prediction result is output through a series of convolutional layers.
[0080] The present invention uses data from the NYUDepth v2 and SUN RGB-D datasets to experiment with the proposed indoor scene monocular image depth estimation method based on deep learning. The NYUDepth v2 dataset is obtained by collecting indoor scenes using the Microsoft KinectRGBD camera, while the SUN RGB-D dataset is collected using devices including Intel Realsense, AsusXtion, Kinect v1, and Kinect v2. Both datasets are indoor scene datasets, and the SUN RGB-D dataset contains more complex scenes.
[0081] Table 1 shows the average relative error, root mean square error, logarithmic mean error and accuracy under the threshold of the parallel decoder used in the present invention and the traditional simple serial decoder after training on the NYUDepthv2 dataset, as well as the number of parameters contained in the model. It can be seen from the data in Table 1 that the method of the present invention has achieved better results than the traditional method, has a certain improvement in the accuracy of depth map estimation, and has reduced the number of model parameters by 29.4%.
[0082]
[0083] Table 1
[0084] Table 2 shows the average relative error, root mean square error, logarithmic mean error and accuracy under the threshold in the NYUDepth v2 dataset after using the loss term designed by the present invention for predicting i and not using it. From the data in Table 2, it can be seen that the loss term designed by the present invention for predicting i can effectively reduce the model prediction error.
[0085]
[0086] Table 2
[0087] Table 3 shows the average relative error, root mean square error, logarithmic mean error and accuracy under the threshold of the Transformer fusion method used in the present invention and another fusion method directly calculated by softmax after training on the NYUDepth v2 dataset, as well as the number of parameters contained in the model, where the calculation method by softmax is as follows:
[0088]
[0089] Where blocki is Figure 1 Specifically, the prediction i is first convolved, then mapped to between 0 and 1 using the sigmoid function and multiplied by the prediction i, and finally all the prediction i are summed up pixel by pixel to get the output. The encoder in this set of experiments uses EfficientNet-B3 pre-trained on ImageNet, which has fewer parameters and faster training speed. From the data in Table 3, it can be seen that the use of Transformer has advantages in both accuracy and error compared with simple computational fusion, and the number of additional parameters is not large.
[0090]
[0091] Table 3
[0092] Table 4 shows the average relative error, root mean square error, logarithmic mean error and accuracy under threshold in the loss term designed by the present invention for prediction i after training on the NYU Depth v2 dataset when the number of intervals is set to 1, 4, and 10 respectively. The encoder used in this group of experiments is also EfficientNet-B3 pre-trained on ImageNet. It can be seen from the data in Table 4 that the division of 10 intervals can achieve better results.
[0093]
[0094] Table 4
[0095] Figure 2 The figure shows a comparison between the depth map predicted by the monocular depth estimation network used in the present invention and the current most advanced networks AdaBins and DPT-Hybrid in certain scenarios, in which: (a) input RGB image; (b) true depth map; (c) depth map predicted by AdaBins; (d) depth map predicted by DPT-Hybrid; (e) depth map predicted by the present invention. It can be seen from the figure that the present invention can more accurately predict the depth information of indoor monocular RGB images, and the edge contours of objects are clearer than AdaBins.
[0096] Figure 3 It shows a schematic diagram of a three-dimensional point cloud generated by using the depth map predicted by the present invention. It can be seen from the figure that the present invention can effectively restore three-dimensional depth information from a two-dimensional image, and has a good guiding role in tasks such as three-dimensional reconstruction and scene understanding.
[0097] Table 5 shows the average relative error, root mean square error, logarithmic mean error and accuracy under threshold of the method of the present invention and the current most advanced methods AdaBins and DPT-Hybrid under the NYUDepthv2 dataset. It can be seen from the data in Table 5 that the method of the present invention has achieved good results in many indicators and has a certain improvement in the accuracy of depth map estimation. Although the encoders of these two most advanced methods and the method of the present invention are pre-trained on ImageNet, DPT-Hybrid requires a large amount of additional training data during model training and fine-tuning on NYU Depth v2 to obtain better results. Specifically, DPT-Hybrid needs to first train 60 epochs on a dataset containing 1.4 million images and then fine-tune on NYU Depth v2, but AdaBins and the model of the present invention only need to be trained on a subset of 50,000 images of NYUDepth v2 for 25 epochs and 20 epochs, respectively.
[0098]
[0099]
[0100] Table 5
[0101] Table 6 shows the average relative error, root mean square error, logarithmic mean error and accuracy under threshold of the method of the present invention and the current most advanced methods AdaBins and DPT-Hybrid after pre-training on the NYUDepthv2 dataset and testing on the test set of the SUN RGB-D dataset. Since no effective method was found to process the inverse depth in the real depth map during testing, only the missing values were masked in the experiment, and since DPT-Hybrid can only be input at a resolution of 480×640, the image resolution was uniformly adjusted to 480×640 in the experiment. It can be seen from the data in Table 6 that the generalization ability of the model of the present invention is roughly ranked second among the three.
[0102]
[0103] Table 6
[0104] Table 7 summarizes the model parameters of the method of the present invention, the current most advanced methods AdaBins and DPT-Hybrid, and the time required for a prediction. It can be seen from the data in Table 7 that the model of the present invention has fewer parameters than the other two models. As for the time required for a prediction, the model of the present invention is slightly slower than AdaBins, ranking second. The inference speed experiment was completed on a machine equipped with an NVIDIA GeForce GTX 1660ti GPU, and the input image resolution was 480×640. Since the output of DPT-Hybrid is twice the output resolution of the present invention and AdaBins, the output of the present invention and AdaBins in the experiment also include the time for a double resampling. In addition, the time in the table is obtained by averaging the time after 5,000 inferences.
[0105]
[0106] Table 7
[0107] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A method for indoor scene monocular image depth estimation based on deep learning, characterized by: Step 1: Introduce the neural network EfficientNet-b7 pre-trained for image classification on ImageNet and construct an encoder; Step 2: At different stages of the encoder, introduce residual connections based on SENet as well as convolution and resampling operations to obtain predictions at different stages; Step 3: Based on the depth interval division method, a loss function focusing on the global to local aspects of the image is constructed and applied to predictions at different stages; Step 4: Use the Transformer structure based on the self-attention mechanism to fuse the depth information predicted at different stages and output the scene depth prediction result; The specific step 1 is as follows: download the EfficientNet-b7 network pre-trained on ImageNet from the Internet, and obtain the feature vectors encoded in the 3rd, 5th, 6th, 8th, and 12th blocks. The resolutions of these feature vectors are respectively The Step 2 is specifically as follows: Step 2.1: Input the feature vector encoded by the third block into four residual blocks based on SENet, the feature vector encoded by the fifth block into three residual blocks based on SENet, the feature vector encoded by the sixth block into two residual blocks based on SENet, and the feature vector encoded by the eighth block into one residual block based on SENet; Step 2.2: Add a channel attention layer after the last residual block in each stage and add a residual connection from the encoder to this layer; Step 2.3: The features of each stage are gradually passed through double upsampling and convolution layers to obtain the features of five stages with the same number of channels of 30 and the same resolution of half the input resolution; Step 2.4: Add and fuse the features of the 1st, 2nd and 5th stages pixel by pixel, add and fuse the features of the 2nd, 3rd and 5th stages pixel by pixel, add and fuse the features of the 1st, 3rd and 4th stages pixel by pixel, and add and fuse the features of the 1st, 4th and 5th stages pixel by pixel, and then pass through the convolution layer to get four predictions, which are marked as prediction 1 to prediction 4 from shallow to deep in the neural network; The Step 3 is specifically as follows: Step 3.1: Get the maximum depth d_max and minimum depth d_min from the real depth map; Step 3.2: Divide the depth interval [d_min, d_max] into 10 small intervals evenly. The calculation formula for the length of a small interval is as follows: Among these 10 intervals, the depth value range of the i-th interval is calculated as follows: [d_min+(i-1)×len, d_min+i×len] Step 3.3: Make a histogram of the real depth map to find the interval with the largest proportion of scene depth among the 10 intervals; Step 3.4: Arrange the 10 depth intervals in descending order according to their proportions, and calculate the mean square error of prediction 1 in Step 2.4 in the 5th to 10th intervals, the mean square error of prediction 2 in the 4th to 8th intervals, the mean square error of prediction 3 in the 2nd to 4th intervals, and the mean square error of prediction 4 in the 1st and 2nd intervals; Step 3.5: Combine the four errors as a loss term that constrains predictions 1 to 4 to focus on the local to global aspects during model training. The calculation formula is as follows: Among them, λ1=0.5, λ2=λ3=0.6, λ4=1, n i is the total number of pixels in the true depth map after the interval mask, and They are the real depth map and the pixel p in the prediction i. i The depth value of .
2. The method for indoor scene monocular image depth estimation based on deep learning according to claim 1, characterized in that: The Step 4 is specifically as follows: Step 4.1: Concatenate the prediction results of the four stages into a four-channel tensor Step 4.2: Transform the four-channel tensor Perform a convolution operation with a convolution kernel of 16×16, a step size of 16, and an output channel of 4, that is: Step 4.3: Flatten the two-dimensional tensor obtained after convolution into one dimension, that is: Step 4.4: Input the one-dimensional tensor into TransformerEncoder and restore its output one-dimensional tensor to a two-dimensional tensor as the weight matrix Step 4.5: Convert the four-channel tensor Perform a convolution operation with a convolution kernel of 3×3, a stride of 1, and an output channel of 128, and the shape is Tensor Step 4.6: Weight Matrix With tensor After performing pixel-by-pixel dot product operations, the final prediction result is output through a series of convolutional layers.
Citation Information
Patent Citations
Depth map prediction method based on visual angle fusion
CN110443842A
Monocular image depth estimation method based on multi-scale residual pyramid attention network model
CN112001960A