Monocular Depth Estimation Method for Visual Augmented Reality of Wearable Smart Devices
Through image activity segmentation and depth similarity loss of recursive networks, the depth estimation error of monocular depth estimation in texture complex and low-texture areas is solved, achieving higher depth estimation accuracy and consistency.
Patent Information
- Application Number
- CN202211171962.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-09-26
AI Technical Summary
The existing monocular depth estimation method has deteriorated depth estimation performance in texture complex areas and low-texture areas, and the photometric constraint method fails in fuzzy scenes, resulting in depth estimation error and error constraint relationships.
Image activity measure (IAM) is used to segment image features, combined with the depth similarity loss of the recursive network, and the similarity measurement is performed through the cosine similarity of the style matrix, improving the network's feature extraction ability in complex areas, and constraining the similarity of the network output through the depth consistency loss.
It improves the depth estimation performance of the network in complex textured areas, makes up for the insufficient constraints of the photometric loss in low-textured areas, and improves the accuracy and consistency of depth estimation.
Smart Images

Figure CN115496787B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to image processing technology, and specifically to a monocular depth estimation method for visual augmented reality of wearable intelligent equipment. Background Art
[0002] Monocular depth estimation is widely used in augmented reality, intelligent equipment, robot navigation, and autonomous driving. Compared with traditional Structure From Motion (SFM) and Simultaneous Localization and Mapping (SLAM) algorithms, monocular depth estimation can obtain the relative depth of a scene without relying on ground truth, thereby determining the relative positional relationship of each object in the image. The acquisition of ground truth generally uses expensive lidar or rendering by a computer simulation engine. However, lidar is not conducive to scenes with frequently changing backgrounds, and the simulation engine has the disadvantage of poor generalization ability for real-world scenes. Self-supervised monocular depth estimation unifies these two tasks into a single framework, uses monocular video as input, and uses a self-supervised constraint mechanism for view synthesis to implement a depth information estimation algorithm that is easy to deploy and set up. With the continuous development of computer computing power, the information mining ability of deep learning algorithms driven by big data has been continuously enhanced, making it possible to obtain depth information from a single image. Depth estimation from a single image is to establish a mapping relationship, which is an ill-posed problem in itself, that is, it is impossible to obtain the absolute depth relationship between objects like a depth sensor, and only the relative depth of each object in the field of view can be obtained. In practical applications, obtaining the relative depth between objects is sufficient to calculate the relative positional relationship of each object in the scene, thus meeting the requirements of the 3D reconstruction task.
[0003] Depth estimation algorithms based on single images are divided into two types: supervised algorithms and self-supervised algorithms. Supervised algorithms require the assistance of ground truth, resulting in difficult deployment of supervised algorithms. Self-supervised algorithms establish a loss relationship based on their own photometric loss, do not require the assistance of ground truth, and there is a large amount of data without ground truth in the entity three-dimensional space. Therefore, self-supervised algorithms are more in line with the actual situation of nature, which makes self-supervised algorithms gradually become the main research direction in the field of depth estimation. With the development and use of multi-task frameworks and auxiliary judgments, the indicators and performance of self-supervised algorithms have surpassed some older supervised algorithms and are gradually approaching the latest supervised algorithms. However, self-supervised depth estimation algorithms still have some defects. Facing the pixels in the textureless area, a small photometric loss does not necessarily mean that the algorithm has good depth estimation ability, which will lead to errors in the overall network when estimating depth.
[0004] Monocular depth estimation is a long-standing problem in the field of computer vision. In 2016, Laina proposed FCRN, which uses a method similar to transposed convolution to replace the fully connected layer, reducing not only the parameters but also adapting to input images of all sizes. Recently, the main research direction of monocular supervised depth estimation has shifted from improving accuracy to simplifying the model and reducing network parameters. Self-supervised depth estimation has been widely studied. Godard and Zhou were the first to adopt a self-supervised monocular depth estimation method that trains a depth network and an independent pose network (Godard C, Aodha O M, Brostow GJ. Unsupervised Monocular Depth Estimation with Left-Right Consistency[C]. Computer Vision&Pattern Recognition, 2017.) (Zhou J, Wang Y, Qin K, et al. Unsupervised High-Resolution Depth Learning From Videos With Dual Networks[C]. 2019 IEEE / CVF International Conference on Computer Vision(ICCV), 2019.). The latest methods have achieved good results in directions such as modeling the scene flow of moving objects and performing separate warping projections on moving objects. With the construction of auxiliary judgment and multi-tasks, the modeling of moving objects becomes more accurate, the occlusion problem is continuously optimized, and the performance of the self-supervised monocular depth estimation network is also continuously improved.
[0005] Defects of existing monocular depth estimation methods:
[0006] First, monocular depth estimation uses the warping reprojection of adjacent frames to establish self-supervised constraint relationships. This leads to the situation that when the depth information is estimated well, due to errors in the warping reprojection, incorrect constraint relationships will still occur, which is particularly obvious in areas with complex textures. In areas with complex textures such as the edges of trees and the overlapping parts of the human body, the depth estimation performance will significantly decline.
[0007] Second, in multiple scenarios, existing photometric constraint methods cannot provide effective constraint relationships. Due to the strong dependence of interpolation reprojection on the original image, when the position relationship of the original image is relatively ambiguous, the performance of photometric constraints will significantly decline. Summary of the Invention
[0008] To solve the above-mentioned shortcomings of the prior art, the present invention provides a monocular depth estimation method for visual augmented reality of wearable intelligent equipment.
[0009] First, the Image Activity Measure (IAM) is used to separately process texture complex regions, enhancing the network's feature extraction ability for complex regions.
[0010] Second, a recurrent neural network is used for secondary constraint, a depth consistency loss is proposed, and the cosine similarity of the style matrix is used for similarity measurement.
[0011] The technical solution of the present invention is: a monocular depth estimation method for wearable intelligent equipment visual augmented reality. First, the image activity measure is used to segment image features. Based on the multi-directional distribution of the image contour, the input picture is segmented into high-order features and low-order features. The high-order features and low-order features are feature-fused and sent to a common decoder to enhance the network's perception intensity for different features;
[0012] Secondly, based on the depth similarity loss of the recurrent network, for the depth map estimated by the network, it is reconstructed by the recurrent network and then recursively sent to the depth estimation network with the same structure. The photometric loss is used to constrain the similarity of the input images of the first-order network and the second-order network, and the depth similarity loss is used to constrain the similarity of the output depth maps of the first-order network and the second-order network to make up for the defect that the photometric loss has poor constraint when facing low-texture regions; the depth consistency loss uses the cosine similarity of the style matrix for similarity measurement.
[0013] Further, the method of the image activity measure is the Image Variance Method IAM.
[0014] Further, the image gradient method calculates the variance of the image block along four directions, namely the horizontal direction, the vertical direction, the lower left diagonal direction, and the lower right diagonal direction.
[0015] Further, the calculation formula of the image variance method is as follows, where α ∈ [0,1] is a weight factor, representing the percentage of the sum of two variances. Here, α is set to 0.5; where V1 is the sum of the variances of the central pixel from the lower left diagonal and the lower right diagonal, V2 is the sum of the variances of the central pixel from the horizontal and vertical directions, M and N are the length and width of the image block, and b i,j is the pixel at position [i,j] in the image block.
[0016]
[0017]
[0018]
[0019] Further, the depth similarity loss based on the recurrent network is specifically defined as follows:
[0020] L dc = α1||D t , D′ t || + α2||I′ t , I″ t || + α3||I t , I″ t ||
[0021] Where α1, α2, and α3 are consistency weights, || || is the similarity loss calculation, the input image I t is the input image, I′ t is the reconstructed image, the depth map D′ t estimated from the reconstructed image I′ t , and the pose parameters estimated by the second-order pose network are used to recover the reconstructed image I″ t , the output D t of the depth estimation network, and the depth consistency loss L dc .
[0022] Further, the total loss L of the network is jointly calculated by the photometric loss L pe , the smoothness loss L s , and the depth consistency loss L dc :
[0023] L = L pe + L s + L dc .
[0024] Beneficial effects:
[0025] First, the present invention uses the Image Activity Measure (IAM) to segment image features. Based on the multi-directional distribution of the image contour, the input image is segmented into high-order features and low-order features. The high-order features and low-order features are feature-fused and fed into a common decoder to enhance the network's perception intensity for different features, and optimize the depth estimation performance in the texture review area.
[0026] Second, the present invention proposes a depth similarity loss based on the recurrent network. For the depth map estimated by the network, it is reconstructed by the recurrent network and then recursively fed into the depth estimation network with the same structure. The similarity between the input images of the first-order network and the second-order network is constrained by the photometric loss, and the similarity between the output depth maps of the first-order network and the second-order network is constrained by the depth similarity loss, so as to make up for the defect that the photometric loss has poor fuzzy photometric constraint in the low-texture area. Description of the Drawings
[0027] Figure 1 It is the IAM feature fusion diagram;
[0028] Figure 2 It is a schematic diagram of the processing of the depth consistency loss based on the recursive network;
[0029] Figure 3 It is a comparison diagram of the KITTI dataset results;
[0030] Figure 4 It is a comparison diagram of the Cityscapes dataset results. Specific implementation manner
[0031] The technical solution of the present invention will be further explained in conjunction with the accompanying drawings.
[0032] This study proposes a recursive depth estimation network algorithm based on image activity. First, an Image activity measure (IAM) is used to segment image features. Based on the multi-directional distribution of the image contour, the input picture is segmented into high-order features and low-order features. The high-order features and low-order features are feature-fused and sent to a common decoder to enhance the network's perception intensity of different features. Secondly, a depth similarity loss based on the recursive network is proposed. For the depth map estimated by the network, it is reconstructed by the recursive network and then recursively sent to the depth estimation network with the same structure. The photometric loss is used to constrain the similarity of the input images of the first-order network and the second-order network, and the depth similarity loss is used to constrain the similarity of the output depth maps of the first-order network and the second-order network, so as to make up for the defect that the photometric loss has poor constraint in the face of low-texture regions. The depth consistency loss uses the cosine similarity of the style matrix for similarity measurement.
[0033] I. Image activity
[0034] Image activity is an index used to measure the image content in the field of image compression coding, and to a certain extent, the activity of the image can also distinguish the texture regions in the image. The active regions of the image generally refer to the regions with strong edges and strong textures in the image. Generally, in the pixel regions with stronger textures or clearer edges, the image activity is generally higher. According to different measurement formulas, effective segmentation of the high- and low-texture regions of different pictures can be made.
[0035] Commonly used image activity measurement methods include the local variance method IAM1, the edge operator method IAM2, the wavelet transform method IAM3, and the image gradient method IAM4.
[0036] 1. Local Variance Method IAM1: The local variance method is the simplest way to measure the activity of a region. It assumes that the pixels in the measurement area are stationary. For an image block from M to N, it measures using the formula (1). Since the size of the region directly affects the IAM of that region, the local variance method is not a good method for measuring image activity.
[0037]
[0038] 2. Edge Operator Method IAM2: A more commonly used method is to use feature extraction methods such as edge extraction to measure IAM. In edge detection, when detecting the boundaries where the pixel values in the image change rapidly or discontinuously, some edge extraction operators are often used, including the Sobel operator, Prewitt operator, Roberts operator, Canny operator, and Marr-Hildreth operator. These operators usually return a binary image identical to the original image, with the positions where the edges are found being 1 and other positions being 0. After the edges are detected, the edge operator method uses the normalized magnitude of the edges to calculate IAM.
[0039]
[0040] 3. Wavelet Transform Method IAM3: By performing a two-dimensional wavelet transform on the calculated image, the two-dimensional image is decomposed into four sub-bands, namely LL, LH, HL, and HH according to frequency characteristics.
[0041] IAM is calculated by performing wavelet transforms on LH, HL, and HH among them.
[0042]
[0043] 4. Image Gradient Method IAM4: As Figure 1 shown, the image gradient method calculates IAM by applying methods such as logarithmic or square root functions to the horizontal or vertical gradients of the image.
[0044]
[0045] To correctly distinguish the textureless regions and strong texture regions in the image through the activity of different regions of the image, IAM should be calculated from multiple angles of the edges. Therefore, in this study, the variances of the image blocks are calculated along four directions, namely the horizontal direction, vertical direction, lower left diagonal direction, and lower right diagonal direction. Due to the smoothness of the middle pixels, the sum of the variances v2 in the horizontal and vertical directions is much larger than the sum of the variances v1 in the lower left diagonal and lower right diagonal directions. The formula is as follows, where α ∈ [0,1] is a weight factor representing the percentage of the sum of the two variances. Here, α is set to 0.5.
[0046]
[0047]
[0048]
[0049] The input image is segmented into regions through the IAM formula to obtain a high-order feature map rich in texture and a low-order feature map with weak texture. The two images are sent into the same encoder for feature extraction. The extracted high-order features and low-order features are fused and sent into a deep decoder to recover depth information. By separately processing the high-order features and low-order features in the image, the texture effectiveness of the high-order features is ensured and the problem of weakening of the low-order features during the encoder process is avoided.
[0050] II. Depth Consistency Loss Based on Recurrent Network
[0051] In daily life, based on a large amount of life experience, humans can infer the three-dimensional structure of self-motion and the surrounding scene in a very short time. This is because extensive daily movement and careful observation of the surrounding scene enable our brains to have a rich and structural understanding of the world. This principle is summarized as Structure from motion (SFM). Based on SFM, Zhou established the initial self-supervised monocular depth estimation framework SFMLearner, using weak supervision of view synthesis as the constraint method, and using the depth map D estimated by DeptNet t and the self-motion parameters estimated by PoseNet to recover the reconstructed image I′ of the input image I t t . By comparing the similarity between the reconstructed image I′ of the input image I t t , the entire depth estimation network is constrained to achieve a self-supervised monocular depth estimation framework.
[0052] However, the early network frameworks were too large and complex. To address this issue, Godard proposed the Monodepth2 network framework. In terms of network structure, Monodepth2 uses a standard, fully convolutional U-Net as the depth prediction network and an independent pose network to predict inter-frame motion. In terms of algorithms, Monodepth2 proposed full-resolution multi-scale, upsampling the intermediate depth images to the input resolution to reduce texture replication artifacts. In terms of the loss function, Monodepth2 adopts per-pixel minimum projection and dynamic culling masks. The former emphasizes that good reprojections should not be averaged, and the latter uses masks to cull dynamic objects in the image to reduce the reprojection error phenomenon of dynamic objects. Monodepth2 established the basic framework for subsequent monocular depth estimation and achieved relatively high depth estimation accuracy using a simple and practical network framework. In terms of the loss function, it uses the L1 norm, photometric loss L pe and smoothness loss L s to jointly calculate the reconstruction loss.
[0053] L pe (I t , I′ t ) = λ1SSIM(I t , I′ t ) + λ2L1(8)
[0054]
[0055] Traditional photometric loss overly relies on the reconstructed image I′ t of the input image I t . When there are weak-texture or textureless regions in the scene, the output D t of the depth estimation network will exhibit a certain degree of hole effect, resulting in the reconstructed image I′ t no longer being representative. At this time, calculating the photometric loss will degrade the performance of the network. To address this issue, this study proposed a depth consistency loss based on a recurrent network to strengthen the constraint.
[0056] As Figure 2 shown, the reconstructed image I′ t recovered by the traditional network is fed into a second-order depth estimation network with the same structure. The reconstructed image I′ t and the adjacent frame I t-1 of the input image are simultaneously fed into a second-order pose estimation network with the same structure. The depth map D′ t estimated from the reconstructed image I′ t and the pose parameters estimated by the second-order pose network are used to recover the reconstructed image I″ t . When the network converges completely, the input image I tThe reconstructed graph I′ t should be the same. Derived from I t and I′ t the estimated first-order depth D t and the second-order depth D′ t should be the same, and I″ t recovered by the second-order network should be the same as the input image I t 's reconstructed graph I′ t should be the same. The depth consistency loss L dc based on the recursive network is defined as follows.
[0057] L dc = α1||D t , D′ t || + α2||I′ t , I″ t || + α3||I t , I″ t || (10)
[0058] where α1, α2, and α3 are consistency weights, and || || is the similarity loss calculation.
[0059] Finally, the total loss of the depth estimation network is calculated jointly by the photometric loss L pe , the smoothness loss L s and the depth consistency loss L dc .
[0060] L = L pe + L s + L dc (11)
[0061] III. Dataset Evaluation
[0062] The algorithm in this paper is trained and tested on the KITTI dataset. Using the dataset segmentation method proposed by Eigen [1] etc., the same transformation is applied to all input images. The center point of the camera is set as the center point of the image, and the focal length of the camera is set as the average of all focal lengths in KITTI. The test set includes 697 pairs of images with a resolution of 1242×375, covering a total of 29 scenes. 39806 images are used for training, and the remaining 4424 images are used for validation. The network framework in this paper is implemented using Pytorch. During training, the images are resized to 640×192. The network is optimized by the Adam optimizer, and the optimizer parameters are set as β1 = 0.9, β2 = 0.999. The weights of the loss function are λ1 = 0.90, λ2 = 0.05, λ3 = 0.05, the initial learning rate is set to 0.0001, and a total of 20 batches are trained. After every 15 batches of training, the learning rate is reduced to one-tenth of the original.
[0063] Table 1 shows a qualitative comparison of our method with several related works. The models are trained on different types of data - monocular video (M), stereo pairs (S), and binocular video (MS), while all of them are tested using a single image as input. The best results are marked in bold and the second-best are underlined. In the monocular training mode, by comparing all objective metrics, our model has better accuracy performance than other algorithms. Compared with the second-best Monodepth2, the Abs Rel index improves from 0.115 to 0.111. For the more penalizing metric Sq Rel, our model improves from 0.837 to 0.816 compared to the second-best DualNet. Additionally, our model has a good result of 0.884 for the most important accuracy metric in object depth estimation. For the stereo training mode, the MonoDisp model achieves the best results in terms of the Sq Rel and RMSE metrics, and the RMSE even reaches the best among all the compared algorithms. The reason why MonoDisp achieves better results in these metrics is that the model is specifically designed for stereo training. They use the bilateral cyclic relationship of the left-right disparity, which enables their model to achieve higher indexing accuracy during stereo training. MonoResM incorporates traditional stereo matching methods in the deep learning algorithm and replaces the ground truth value with the semi-global matching method. This is why MonoResM achieves the best results in the RMSE log and the last two accuracy metrics. By inserting proxy supervised ground truth, the depth estimation metrics are enhanced during the training process. Although our model is made through the monocular training mode, it still exceeds the results of MonoDisp and MonoResM in terms of Abs Rel, and our model improves to 0.874 for the most important accuracy metric. In the binocular training, the proposed model achieves the best results in terms of Abs Rel, RMSE log, and accuracy, and we achieve the second-best results in other metrics. The results show that the algorithm in this paper has significantly better effects than the results of other unsupervised monocular depth estimation algorithms.
[0064] The most commonly used quantitative metrics for evaluating monocular depth estimation currently are the absolute relative difference (AbsRel), relative error (SqRel), root mean square error (RMSE), and root mean square error of logarithm RMSE(log). The following are the calculation formulas for each quantitative metric.
[0065]
[0066]
[0067]
[0068]
[0069] The accuracy of monocular depth estimation is the percentage that meets the following conditions
[0070]
[0071] Table 1 Test results on the KITTI dataset
[0072] Among them, M represents monocular training, S represents stereo pair training, and MS represents monocular plus stereo pair training.
[0073]
[0074]
[0075] [1]Eigen D,Fergus R.Predicting Depth,Surface Normals and SemanticLabels with a Common Multi-Scale Convolutional Architecture[J].IEEE,2014.
[0076] [2]Casser V,Pirk S,Mahjourian R,et al.Depth Prediction Without theSensors:Leveraging Structure for Unsupervised Learning from Monocular Videos[J],2018.
[0077] [3]Godard C,Aodha O M,Brostow G J.Unsupervised Monocular DepthEstimation with Left-Right Consistency[C].Computer Vision&PatternRecognition,2017.
[0078] [4] Gordon A, Li H, Jonschkowski R, et al. Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown Cameras[C]. 2019 IEEE / CVF International Conference on Computer Vision(ICCV), 2019.
[0079] [5] Zhou J, Wang Y, Qin K, et al. Unsupervised High-Resolution Depth Learning From Videos With Dual Networks[C]. 2019 IEEE / CVF International Conference on Computer Vision(ICCV), 2019.
[0080] [6] Klingner M, Termhlen J A, Mikolajczyk J, et al. Self-Supervised Monocular Depth Estimation: Solving the Dynamic Object Problem by Semantic Guidance[J], 2020.
[0081] [7] Godard C, Aodha O M, Firman M, et al. Digging Into Self-Supervised Monocular Depth Estimation[C]. 2019 IEEE / CVF International Conference on Computer Vision(ICCV), 2020.
[0082] [8] Garg R, Bg V K, Carneiro G, et al. Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue[C]. European Conference on Computer Vision, 2016.
[0083] [9] Mehta I, Sakurikar P, Narayanan PJ. Structured Adversarial Training for Unsupervised Monocular Depth Estimation[C]. 2018 International Conference on 3D Vision(3DV), 2018.
[0084]
[10] Poggi M, Tosi F, Mattoccia S. Learning monocular depth estimation with unsupervised trinocular assumptions[C]. 2018 International Conference on 3D Vision(3DV), 2018:324 - 333.
[0085]
[11] Li R, Wang S, Long Z, et al. UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning[J], 2017.
[0086]
[12] Luo C, Yang Z, Peng W, et al. Every Pixel Counts++: Joint Learning of Geometry and Motion with 3D Holistic Understanding[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
[0087] Figure 3 The depth estimation effect diagrams on the KITTI dataset are shown. In this paper, the depth map results of five algorithms, namely Struct2 [2] , Mono [9] , SGD [6] , Mono2 [7] and the algorithm in this paper, are compared. The obvious parts are highlighted with rectangular frames. Compared with other methods, when the proposed model in this paper processes small objects at far distances such as people, vehicles, billboards, and tree trunks, the depth maps obtained by the algorithm in this paper can outline the overall contour more clearly, with obvious edges and rich details. For objects such as roadside guardrails at close range, the depth maps obtained by the proposed model in this paper are more continuous, smooth, sharp and clear.
[0088] To evaluate the generalization of the model in this paper, the model was directly used to test the Cityscapes dataset, and the results were compared with other unsupervised monocular depth estimation methods. The results are as Figure 4 shown, where the obvious areas are highlighted with rectangular boxes. Compared with other methods, the depth maps obtained by the model in this paper are more discriminative and recognizable when dealing with challenging areas such as multiple people overlapping and complex textures. This shows that the model has good generalization.
Claims
1. A monocular depth estimation method for vision-augmented reality of wearable intelligent devices, characterized in that, First, an image activity metric is used to segment image features. Based on the multi-directional distribution of the image contour, the input image is segmented into high-order features and low-order features. The high-order features and low-order features are fused and fed into a common decoder to enhance the network's perception intensity for different features; Secondly, based on the depth similarity loss of the recursive network, for the depth map estimated by the network, it is reconstructed by the recursive network and then recursively fed into the depth estimation network with the same structure. The photometric loss is used to constrain the similarity of the input images of the first-order network and the second-order network, and the depth similarity loss is used to constrain the similarity of the output depth maps of the first-order network and the second-order network to make up for the defect that the photometric loss has poor constraint when facing low-texture regions; the depth consistency loss uses the cosine similarity of the style matrix for similarity measurement; The method of the image activity metric is the Image Variance Method (IAM); The image gradient method calculates the variance of the image patch along four directions, namely the horizontal direction, the vertical direction, the lower left diagonal direction, and the lower right diagonal direction. The calculation formula of the image gradient method is as follows, where α ∈ [0, 1] is a weight factor representing the percentage of the sum of the two variances, and α is set to 0.5; Where V1 is the sum of the variances of the central pixel from the lower left diagonal and the lower right diagonal; V2 is the sum of the variances of the central pixel from the horizontal and vertical directions; M and N are the length and width of the image block, and b i,j is the pixel at position [i, j] in the image block, 2. The monocular depth estimation method for visual augmented reality of wearable intelligent devices according to claim 1, characterized in that, The depth similarity loss based on the recursive network is specifically defined as follows: L dc = α1||D t , D′ t || + α2||I′ t , I″ t || + α3||I t , I″ t || Among them, α1, α2, α3 are consistency weights, || || is the calculation of similarity loss, and the input image is I t is the input image, and I′ t is the reconstructed image, and the reconstructed image I′ t is the estimated depth map D′ t , and the pose parameters estimated by the second-order pose network are used to restore the reconstructed image I″ t , the output D of the depth estimation network t , and the depth consistency loss L dc .
3. The monocular depth estimation method for visual augmented reality of wearable intelligent devices according to claim 1, characterized in that The total loss L of the depth estimation network is calculated jointly by the photometric loss L pe , the smoothness loss L s and the depth consistency loss L dc : L = L pe + L s + L dc 。