Depth estimation using neural networks
Patent Information
- Application Number
- CN202180004351.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-19
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-04-19
AI Technical Summary
然而,根据单个图像估计深度图可能需要密集计算能力,这在一些情况下不很好地适合于移动应用
Smart Images

Figure CN115500083B_ABST
Abstract
Description
Technical Field
[0001] This description generally involves depth estimation using neural networks. Background Technology
[0002] Depth estimation is a computer vision task designed to estimate depth (parallax) from image data (e.g., receiving an RGB image and outputting a depth image). In some traditional methods, multiple cameras and / or physical markers in a scene are used to reconstruct a depth map from multiple views of the same scene / object. However, estimating a depth map from a single image can require computationally intensive power, which is not well-suited for mobile applications in some cases. Summary of the Invention
[0003] According to one aspect, a method for depth estimation includes: receiving image data from a sensor system; generating a first depth map based on the image data by a neural network, wherein the first depth map has a first scale; obtaining a depth estimate associated with the image data; and using the depth estimate to transform the first depth map into a second depth map, wherein the second depth map has a second scale.
[0004] According to some aspects, the method may include one or more (or any combination thereof) of the following features: The method includes generating surface normals from image data by a neural network, wherein a first depth map is transformed into a second depth map using the surface normals and a depth estimate. The method may include generating visual feature points based on image data, the visual feature points being associated with the depth estimate. The method may include obtaining the depth estimate from a depth sensor. The depth estimate may be obtained during an augmented reality (AR) session that can be performed by a mobile computing device. The method may include estimating affine parameters based on an offset between the first depth map and the depth estimate, wherein the affine parameters include scale and shift, and transforming the first depth map into a second depth map based on the affine parameters. The method may include: predicting a first surface normal from image data by a neural network; predicting a second surface normal based on the second depth map; calculating a self-consistency loss based on the first and second surface normals; and updating the neural network based on the self-consistency loss. The method may include using the second depth map to estimate at least one planar region in the image data, wherein the at least one planar region is configured to be used as a surface to which a virtual object is attached.
[0005] According to one aspect, a depth estimation system includes: a sensor system configured to acquire image data; a neural network configured to generate a first depth map based on the image data, wherein the first depth map has a first scale; a depth estimation generator configured to acquire a depth estimate associated with the image data; and a depth map transformer configured to estimate affine parameters based on the depth estimate and the first depth map and use the affine parameters to transform the first depth map into a second depth map, wherein the second depth map has a second scale.
[0006] According to some aspects, the depth estimation system may include one or more (or any combination thereof) of the above / below features. The neural network is configured to execute on a mobile computing device. The depth estimation system may include a visual inertial motion tracker configured to generate visual feature points associated with the depth estimate. The depth estimation system may include a depth sensor configured to obtain the depth estimate. The depth estimation generator is configured to obtain the depth estimate during an augmented reality (AR) session, wherein the depth estimation generator is also configured to obtain pose data, gravity orientation, and identification of one or more planar regions in the image data during the AR session. Affine parameters may include scale and shift for each depth estimate in a first depth map. The depth map transformer may include a Random Sample Consensus (RANSAC)-based solver that minimizes the objective function to estimate the scale and shift. The depth estimation system may include a convolutional neural network trainer configured to use a neural network to predict a first surface normal based on image data, predict a second surface normal based on a second depth map, calculate a self-consistency loss based on the first and second surface normals, calculate a loss based on the first surface normal and the ground reality normal, and update the neural network based on the self-consistency loss and the loss. The depth map transformer may include a plane generator configured to use the second depth map to estimate at least one planar region in the image data, wherein the at least one planar region is configured to be used as a surface to which a virtual object is attached, wherein the plane generator includes: a graph converter configured to convert the second depth map into a point cloud; and a plane detector configured to use the point cloud to detect at least one planar region according to a plane fitting algorithm.
[0007] According to one aspect, a non-transitory computer-readable medium stores executable instructions that, when executed by at least one processor, cause the at least one processor to receive image data from a sensor system, generate a first depth map based on the image data by a neural network having a first scale, obtain a depth estimate associated with the image data, transform the first depth map into a second depth map using the depth estimate, wherein the second depth map has a second scale, and use the second depth map to estimate at least one planar region in the image data, the at least one planar region being configured to be used as a surface to attach virtual objects during an augmented reality (AR) session.
[0008] The non-transitory computer-readable medium may include any one of the above / below features (or any combination thereof). The executable instructions include, when executed by at least one processor, instructions that cause the at least one processor to estimate affine parameters based on an offset between a first depth map and a depth estimate, wherein the affine parameters include scale and shift, and that the first depth map is transformed into a second depth map based on the affine parameters. The depth estimate may be obtained from at least one of a visual-inertial motion tracker, a depth sensor, a dual-pixel depth estimator, a moving stereo depth estimator, a sparse active depth estimator, or a pre-computed sparse map. The executable instructions include, when executed by at least one processor, instructions that cause the at least one processor to generate surface normals from image data using a neural network, wherein the first depth map is transformed into a second depth map using the surface normals and the depth estimate. Attached Figure Description
[0009] Figure 1A The diagram illustrates a depth estimation system based on one aspect.
[0010] Figure 1B The diagram illustrates a depth estimation generator that obtains depth estimates based on one aspect.
[0011] Figure 1C The illustration shows an example of visual feature points in one aspect of image data.
[0012] Figure 1D The diagram illustrates a depth map transformer based on one aspect.
[0013] Figure 1E The illustration shows an example operation of a parameter estimation solver based on one aspect of a depth map transformer.
[0014] Figure 1F The diagram shows an accelerometer that measures the direction of gravity from one perspective.
[0015] Figure 1G The diagram illustrates a plane generator configured to use visual feature points to detect one or more planar regions.
[0016] Figure 1H The illustration shows an example of information captured during an augmented reality (AR) session.
[0017] Figure 1I The diagram illustrates a neural network trainer based on one aspect.
[0018] Figure 2 The diagram illustrates a neural network based on one aspect.
[0019] Figure 3 The diagram illustrates a plane generator configured to detect one or more planar regions in image data from a depth map.
[0020] Figure 4 The illustration shows an AR system with a depth estimation system based on one aspect.
[0021] Figure 5 The diagram depicts a flowchart of an example operation of a depth estimation system based on one aspect.
[0022] Figure 6 The diagram illustrates a flowchart of an example operation of adjusting a neural network according to one aspect.
[0023] Figure 7 The diagram illustrates a flowchart of an example operation of a depth estimation system based on another aspect.
[0024] Figure 8 The illustration shows an example computing device based on a depth estimation system. Detailed Implementation
[0025] An embodiment provides a depth estimation system comprising: a sensor system that acquires image data; and a neural network configured to generate a depth map based on image frames of the image data (e.g., using a single image frame to generate the depth map). In some examples, the depth map generated by the neural network may be associated with a first scale (e.g., a non-metric map). The depth map generated by the neural network may be an affine-invariant depth map, which is a depth map adapted for scaling / shifting but not associated with a metric scale (or imperial numeral system). The depth estimation system includes: a depth estimation generator that acquires depth estimates from one or more sources (e.g., depth estimates having depth values according to a second scale (e.g., a metric scale); and a depth map transformer configured to use the depth estimates to transform the depth map generated by the neural network into a depth map with a second scale (e.g., a metric scale). The first and second scales may be different scales capable of being based on two different measurement systems with different standards. In some examples, a metric depth map can refer to an image in which each pixel represents a metric depth value according to the metric scale (e.g., meters) of the corresponding pixel in the image. A metric depth estimate obtained by a depth estimation generator can be considered a sparse depth estimate (e.g., a depth estimate for some pixels rather than all pixels in the image data). In some examples, the metric depth estimate is associated with a subset of pixels in the image data. The depth map converter uses the sparse depth estimate to provide a second scale (e.g., a metric scale) for the depth map generated by the neural network. In some examples, embodiments provide a system capable of providing a metric scale for all pixels when a metric depth estimate may only exist for a sparse subset, and dense metric depth maps offer technical advantages over sparse metric depth maps for downstream applications (e.g., 3D reconstruction, planar lookup, etc.).
[0026] This depth estimation system can provide a solution for scale / shift ambiguity (or commonly referred to as affine ambiguity) in monocular deep neural networks. For example, this depth estimation system can use a sparse source of metric depth to address affine ambiguity in monocular machine learning (ML) depth models. Affine ambiguity can be difficult for some applications that require (or benefit from) real-world scales (e.g., metric scales). For example, mobile augmented reality (AR) applications may involve placing virtual objects in a camera view with real-world dimensions. To render the object at a real-world scale, it may be necessary to estimate the depth of the surface on which the virtual object is placed in metric units. According to the embodiments discussed herein, the metric depth map generated by this depth estimation system can be used to estimate planar regions in image data, where the planar regions are used as the surfaces to which the virtual objects are attached.
[0027] In some traditional AR applications, surfaces are estimated in a 3D point cloud. However, these methods may not allow users to quickly (e.g., immediately) place virtual objects in a scene. Instead, the user scans a planar surface with sufficient texture to detect a sufficient number of 3D points and then performs plane detection, which can result in the AR session failing to detect many planes and / or taking a relatively long time to detect them. However, by using a metric depth map generated by a depth estimation system, it is possible to reduce the latency for detecting planar regions. For example, the depth estimation system can reduce placement latency by using a neural network to predict the scale of the depth of the object / planar surface to be placed (e.g., estimating depth from a single image or a small number of images, thus requiring less user movement). Furthermore, this depth estimation system can predict depth based on a low-textured surface such as a white table. Additionally, note that the metric depth map generated by this depth estimation system can be used in a wide variety of applications, including robotics (beyond AR applications).
[0028] In some examples, the depth map transformer uses one or more additional signals to help provide a second scale (e.g., a metric scale) for the depth map generated by the neural network. In some examples, the neural network predicts surface normals, and the depth map transformer uses the predicted surface normals along with sparse depth estimates to provide a second scale (e.g., a metric scale) for the depth map generated by the neural network.
[0029] The accuracy of depth prediction can be improved by predicting both depth and surface normals. To encourage consistency between predicted depth and surface normals, a self-consistency loss (e.g., unsupervised self-consistency loss) is used during the training or tuning of the neural network. For example, a neural network can predict a first surface normal based on an RGB image, while a depth map transformer can predict a second surface normal based on a metric depth map. The self-consistency loss is calculated based on the difference between the first and second surface normals and is added to the supervised loss. The supervised loss is calculated based on the difference between the first surface normal and the ground truth normal. The self-consistency loss encourages the network to minimize any deviation between the first and second surface normals.
[0030] In some examples, the depth map transformer can receive a gravity direction and a planar region. The gravity direction is obtained from an accelerometer. The planar region can be estimated by a planar generator using visual feature points (e.g., SLAM points) during an AR session. The depth map transformer can use the gravity direction and the planar region (along with sparse depth estimation) to provide a second scale (e.g., a metric scale) for the depth map generated by the neural network.
[0031] The depth map transformer may include a parameter estimator solver configured to perform a parameter estimation algorithm to estimate affine parameters (e.g., shift, scale) based on the offset between the sparse depth estimate and the depth map generated by the neural network. In some examples, the parameter estimator solver is a Random Sample Consensus (RANSAC)-based solver that solves an objective function to estimate the scale and shift. In some examples, the parameter estimator solver is configured to solve a least-squares parameter estimation problem within a RANSAC loop to estimate the affine parameters for the depth map to transform it to a second scale (e.g., a metric scale).
[0032] In some examples, the neural network is considered a monocular deep neural network because it predicts a depth map based on a single image frame. In some examples, the neural network includes a U-net architecture configured to predict pixel-wise depth from a red-green-blue (RGB) image. In some examples, the neural network includes features that enable it to run on mobile computing devices such as smartphones, tablets, etc. For example, the neural network can use depthwise separable convolutions. Depthwise separable convolutions involve factoring a standard convolution into depthwise convolutions and 1x1 convolutions called pointwise convolutions. This factorization has the effect of reducing computation and model size. In some examples, the neural network can use a Blurpool encoder, which can be a combination of anti-aliasing and subsampling operations that make the network more robust and stable against corruptions such as rotation, scaling, blurring, and noise variants. In some examples, the neural network can include bilinear upsampling, which reduces the parameters that transpose the convolution and thus reduces the network size. Refer to the figures for further illustration of these and other features.
[0033] Figures 1A to 1G The illustration illustrates a depth estimation system 100. The depth estimation system 100 generates a depth map 138 based on a depth estimate 108 (obtained from one or more sources) and a depth map 120 generated by a neural network 118. The depth map 120 generated by the neural network 118 has a first scale. In some examples, the first scale is a non-metric scale. The depth map 138 has a second scale. The first and second scales are based on two different measurement systems with different standards. In some examples, the second scale is a metric scale. The depth estimation system 100 is configured to convert the depth map 120 with the first scale into a depth map 138 with the second scale. The depth map 138 with the second scale can be used to control augmented reality, robotics, natural user interface technologies, games, or other applications.
[0034] The depth estimation system 100 includes a sensor system 102 that acquires image data 104. The sensor system 102 includes one or more cameras 107. In some examples, the sensor system 102 includes a single camera 107. In some examples, the sensor system 102 includes two or more cameras 107. The sensor system 102 may include an inertial motion unit (IMU). The IMU can detect motion, movement, and / or acceleration of the computing device. The IMU may include various different types of sensors, such as, for example, an accelerometer (e.g., Figure 1C The sensor system 102 may include other types of sensors, such as light sensors, audio sensors, distance and / or proximity sensors, contact sensors such as capacitive sensors, timers and / or other sensors and / or different combinations of sensors.
[0035] The depth estimation system 100 includes one or more processors 140, which may be formed in a substrate configured to execute one or more machine-executable instructions or software, firmware, or combinations thereof. The processors 140 may be semiconductor-based—that is, the processors may include semiconductor materials capable of performing digital logic. The depth estimation system 100 may also include one or more memory devices 142. The memory devices 142 may include any type of storage device storing information in a format capable of being read and / or executed by the processors 140. The memory devices 142 may store applications and modules that, when executed by the processors 140, perform any of the operations discussed herein. In some examples, applications and modules may be stored in external storage devices and loaded into the memory devices 142.
[0036] Neural network 118 is configured to generate depth map 120 based on image data 104 captured by sensor system 102. In some examples, neural network 118 receives image frame 104a of image data 104 and generates depth map 120 based on image frame 104a. Image frame 104a is a red-green-blue (RGB) image. In some examples, neural network 118 uses a single image frame 104a to generate depth map 120. In some examples, neural network 118 uses two or more image frames 104a to generate depth map 120. Depth map 120 generated by neural network 118 may be an affine invariant depth map, which is a depth map adapted for scaling / shifting but not associated with a first scale (e.g., a metric scale). Depth map 120 may refer to an image in which each pixel represents a depth value according to a non-metric scale (e.g., 0 to 1) of the corresponding pixel in the image. The non-metric scale may be a scale not based on the metric, International System of Units (SI), or Imperial measurement systems. Although embodiments are described with reference to metric (or metric values) and non-metric (or non-metric) scales, the first and second scales can be based on any two different measurement systems with different standards. Depth map 120 can be used to describe an image containing information related to the distance from the camera viewpoint to the surface of an object in the scene. The depth value is inversely correlated with the distance from the camera viewpoint to the surface of an object in the scene.
[0037] The neural network 118 can be any type of deep neural network configured to generate a depth map 120 using one or more image frames 104a (or a single image frame 104a). In some examples, the neural network 118 is a convolutional neural network. In some examples, the neural network 118 is considered a monocular deep neural network because the neural network 118 predicts the depth map 120 based on a single image frame 104a. The neural network 118 is configured to predict pixel-wise depth based on image frames 104a. In some examples, the neural network 118 includes a U-net architecture, for example, an encoder-decoder with skip connections and learnable parameters.
[0038] In some examples, the neural network 118 has a size capable of running on mobile computing devices (e.g., smartphones, tablets, etc.). In some examples, the size of the neural network 118 is less than 150 Mb. In some examples, the size of the neural network 118 is less than 100 Mb. In some examples, the size of the neural network 118 is approximately 70 Mb or less. In some examples, the neural network 118 uses depthwise separable convolutions, which are the form of factorized convolutions that factorize a standard convolution into depthwise convolutions and 1x1 convolutions called pointwise convolutions. This factorization can have the effect of reducing computation and model size. In some examples, the neural network 118 can use a Blurpool encoder, which can be a combination of anti-aliasing and subsampling operations that make the network more robust and stable to damage such as rotation, scaling, blurring, and noise variants. In some examples, the neural network 118 can include bilinear upsampling, which can reduce the parameters that transpose the convolution and thus reduce the size of the network.
[0039] In some examples, the neural network 118 also predicts surface normals 122a that describe the surface orientation of image frame 104a (e.g., all visible surfaces in the scene). In some examples, surface normals 122a include per-pixel normals or per-pixel surface orientations. In some examples, surface normals 122a include surface normal vectors. The surface normal 122a of a pixel in an image can be defined as a three-dimensional vector corresponding to the orientation of a 3D surface represented by that pixel in the real world. The orientation of the 3D surface is represented by a direction vector perpendicular to the real-world 3D surface. In some examples, the neural network 118 is also configured to detect planar regions 124 within image frame 104a. Planar regions 124 may include vertical planes and / or horizontal planes.
[0040] Depth estimation system 100 includes a depth estimation generator 106 that obtains a depth estimate 108 (e.g., a metric depth estimate) associated with image data 104. Depth estimate 108 may include depth values in a metric scale for a subset of pixels in image data 104. For example, a metric scale can refer to any type of measurement system such as the metric system and / or the imperial system. The depth estimate 108 obtained by depth estimation generator 106 can be considered a sparse depth estimate (e.g., a depth estimate for a subset of pixels rather than all pixels in the image data). For example, if image frame 104a is 10x10, then image frame 104a includes one hundred pixels. However, depth estimate 108 may include depth estimates in a metric scale for a subset of these pixels. In contrast, a dense depth map (e.g., depth map 120) provides depth values (e.g., non-metric depth values) for a large number of pixels in an image or for all pixels in an image.
[0041] Depth estimation generator 106 can be any type of component configured to generate (or obtain) depth estimation 108 based on image data 104. In some examples, depth estimation generator 106 also obtains pose data 110 and identifies planar regions 114 within image data 104. Pose data 110 can identify the pose (e.g., position and orientation) of a device (e.g., a smartphone with depth estimation system 100) performing depth estimation system 100. In some examples, pose data 110 includes the device's five degrees of freedom (DoF) position. In some examples, pose data 110 includes the device's six DoF position. In some examples, depth estimation generator 106 includes a plane generator 123 configured to detect planar regions 114 within image data 104 using any type of plane detection algorithm (or plane fitting algorithm). Planar regions 114 can be planar surfaces of objects (e.g., tables, walls, etc.) within image data 104.
[0042] refer to Figure 1B The depth estimation generator 106 may include a visual-inertial motion tracker 160, a depth sensor 164, a moving stereo depth estimator 168, a sparse active depth estimator 170, and / or a pre-computed sparse map 172. Each component of the depth estimation generator 106 may represent a separate source for obtaining the depth estimate 108. For example, each component may generate the depth estimate 108 independently, where the depth estimation generator 106 may include one or more components. In some examples, the depth estimation generator 106 may include a single source, such as one of the visual-inertial motion tracker 160, the depth sensor 164, the dual-pixel depth estimator 166, the moving stereo depth estimator 168, the sparse active depth estimator 170, or the pre-computed sparse map 172. In some examples, if the depth estimation generator 106 includes multiple sources (e.g., multiple components), the depth estimation generator 106 may select one of the sources for use when generating the depth map 138. In some examples, if the depth estimation generator 106 includes multiple sources (e.g., multiple components), the depth estimation generator 106 can select multiple sources to use when generating the depth map 138.
[0043] The visual inertial motion tracker 160 is configured to generate visual feature points 162 representing image data 104. The visual feature points 162 are associated with depth estimation 108. For example, each visual feature point 162 may include a depth value in a metric scale. Figure 1CThe illustration shows a scene 125 captured by camera 107, where scene 125 depicts visual feature points 162 generated by visual inertial motion tracker 160 using image data 104. Visual feature points 162 may include depth values in metric scales, where the depth value is inversely correlated with the distance from the camera viewpoint to the surface of an object in scene 125.
[0044] Visual feature points 162 are multiple points in 3D space representing the user's environment (e.g., points of interest). In some examples, each visual feature point 162 includes an approximation of a fixed location and orientation in 3D space and can be updated over time. For example, a user can move her mobile phone camera around scene 125 during AR session 174, where a visual inertial motion tracker 160 can generate visual feature points 162 representing scene 125. In some examples, visual feature points 162 include Simultaneous Localization and Mapping (SLAM) points. In some examples, visual feature points 162 are referred to as a point cloud. In some examples, visual feature points 162 are referred to as feature points. In some examples, visual feature points 162 are referred to as 3D feature points. In some examples, visual feature points 162 are in the range of 200-400 per image frame 104a.
[0045] Return to reference Figure 1B In some examples, the visual inertial motion tracker 160 is configured to execute a SLAM algorithm, which is a tracking algorithm capable of estimating the movement of a device (e.g., a smartphone) in space using camera 107. In some examples, the SLAM algorithm is also configured to detect planar regions 114. In some examples, the SLAM algorithm iteratively calculates the device's position and orientation (e.g., pose data 110) by analyzing keypoints (e.g., visual feature points 162) and descriptors in each image and tracking these descriptors frame by frame, which enables 3D reconstruction of the environment.
[0046] Depth sensor 164 is configured to generate depth estimate 108 based on image data 104. In some examples, depth sensor 164 includes a light detection and ranging (LiDAR) sensor. A dual-pixel depth estimator 166 uses a machine learning model to estimate the depth of the camera's dual-pixel autofocus system. Dual-pixel operates by splitting each pixel in half so that each half-pixel observes a different half of the primary lens aperture. By reading out each of these half-pixel images individually, two slightly different views of the scene are obtained, and these different views are used by dual-pixel depth estimator 166 to generate depth estimate 108. A moving stereo depth estimator 168 can use multiple images in a stereo matching algorithm to generate depth estimate 108. In some examples, a single camera can be moved around scene 125 to capture multiple images, which are used for stereo matching to estimate metric depth. A sparse active depth estimator 170 can include a sparse time-of-flight estimator or a sparse phase detection autofocus (PDAF) estimator. In some examples, a pre-computed sparse map 172 is a sparse map used by a visual positioning service.
[0047] Return to reference Figure 1A The depth estimation system 100 includes a depth map transformer 126 configured to use depth estimation 108 to transform a depth map 120 generated by a neural network 118 into a depth map 138. The depth map 138 may refer to an image in which each pixel represents a depth value according to a metric scale (e.g., meters) of the corresponding pixel in image data 104. The depth map transformer 126 is configured to use depth estimation 108 to provide a metric scale for the depth map 120 generated by the neural network 118.
[0048] Depth map transformer 126 is configured to estimate affine parameters 132 based on depth map 120 generated by neural network 118 and depth estimate 108. Affine parameters 132 include scale 134 and shift 136 of depth map 120. Scale 134 includes a scale value indicating a resizing amount of depth map 120. Shift 136 includes a shift value indicating a shift amount of pixels in depth map 120. Note that scale 134 (or scale value) refers to a resizing amount that is completely different from the aforementioned "first scale" and "second scale," and is completely different from the aforementioned "first scale" and "second scale" that reference a different measurement system (e.g., the first scale may be a non-metric scale while the second scale may be a metric scale). Depth map transformer 126 is configured to use affine parameters 132 to transform depth map 120 into depth map 138. In some examples, scale 134 and shift 136 comprise two numbers (e.g., s = scale, t = shift) that, when multiplied and summed with the values in each pixel of depth map 120, produce depth map 138 (e.g., D138(x,y) = s * D120(x,y) + t), where D120(x,y) is the value at pixel position (x,y) in depth map 120). Affine parameter 132 can be estimated from the sparse set of depth estimate 108 and then applied to each pixel in depth map 120 using the above equation. Since depth map 120 has effective depth for all pixels, depth map 138 will also have a metric scale for all pixels.
[0049] Depth map transformer 126 is configured to perform a parameter estimation algorithm to solve an optimization problem (e.g., an objective function) that minimizes the objective of aligning depth estimate 108 with depth map 120. In other words, depth map transformer 126 is configured to minimize the objective function that aligns depth estimate 108 with depth map 120 to estimate affine parameters 132. For example, as indicated above, depth estimate 108 obtained by depth estimate generator 106 can be considered a sparse depth estimate (e.g., a depth estimate of some pixels rather than all pixels in image data 104). For example, if image frame 104a is 10x10, then image frame 104a comprises one hundred pixels. Depth estimate 108 may include depth estimates in metric scales for a subset of pixels in image frame 104a (e.g., a number less than one hundred in the example of a 10x10 image). However, depth map 120 includes a depth value for each pixel in the image, where the depth value is a non-metric unit such as a number between zero and one. For each pixel having a metric depth estimate 108 (e.g., a metric depth value), the depth map transformer 126 can obtain a corresponding depth value (e.g., a non-metric depth value) in the depth map 120 and use the metric and non-metric depth values to estimate scale 134 and shift 136. This can include minimizing the error when scale 134 is multiplied by the non-metric depth value plus shift 136 minus the metric depth value is zero. In some examples, the depth map transformer 126 is configured to solve a least-squares parameter estimation problem within a Random Sample Consensus (RANSAC) cycle to estimate affine parameters 132.
[0050] refer to Figure 1D The depth map transformer 126 may include a data projector 176 configured to project a depth estimate 108 onto a depth map 120. If the depth estimate 108 includes visual feature points 162, the data projector 176 projects the visual feature points 162 onto the depth map 120. The depth map transformer 126 may include a parameter estimation solver 178 configured to solve an optimization problem to estimate affine parameters 132 (e.g., scale 134, shift 136), wherein the optimization problem minimizes the objective of aligning the depth estimate 108 with the depth map 120. In some examples, the parameter estimation solver 178 includes a RANSAC-based parameter estimation algorithm. In some examples, the parameter estimation solver 178 is configured to solve a least-squares parameter estimation problem within a RANSAC loop to estimate the affine parameters 132.
[0051] Figure 1EThe illustration shows an example operation of the parameter estimation solver 178. In operation 101, the parameter estimation solver 178 determines a scale 134 and a shift 136 based on the depth offset between the depth estimate 108 and the depth map 120. The parameter estimation solver 178 uses any two points from the depth estimate 108 and the depth map 120 to compute the scale 134 (e.g., the scale for inverse depth) and the shift 136 (e.g., the shift for inverse depth) based on the following equation:
[0052] Equation (1):
[0053] Equation (2): c = l i -kd i
[0054] Parameter k indicates a scale of 134, and parameter c indicates a shift of 136. Parameter l i It is the inverse depth (e.g., a metric depth value) of the i-th estimate (which corresponds to the i-th depth prediction). Parameter d i It is the inverse depth predicted from the i-th depth (e.g., a non-metric depth value). Parameter l j It is the inverse depth (e.g., a metric depth value) of the j-th estimate (which corresponds to the j-th depth prediction). Parameter d j It is the inverse depth predicted from the j-th depth (e.g., a non-metric depth value). For example, l i and l j This can represent the metric depth value of two points (e.g., two pixels) in depth estimation 108, and d i and d j It can represent the non-metric depth value of two corresponding points (e.g., two pixels) in the depth map 120.
[0055] In operation 103, parameter estimation solver 178 performs an evaluation method based on the following equations to identify which other points (e.g., pixels) are inliers of the above solutions (e.g., equations (1) and (2)):
[0056] Equation (3): e = (d i -l i ) 2 , where e < t, and t is the interior point threshold (e.g., RANSAC interior point threshold). For example, for a specific point (e.g., a pixel) with both non-metric and metric depth values, the parameter estimation solver 178 obtains the non-metric depth value (d). i ) and metric depth values (l i If the squared difference is less than the inlier threshold, the point is identified as an inlier.
[0057] In operation 105, parameter estimation solver 178 is configured to perform a least-squares solver for scale 134(k) and shift 136(c) to refine the estimate based on the desired estimate from the evaluation method.
[0058] Return to reference Figure 1A The depth map transformer 126 can use one or more other signals to help provide a metric scale for the depth map 120 generated by the neural network 118. In some examples, the neural network 118 can predict surface normals 122a, and the depth map transformer 126 can use the predicted surface normals 122a and the depth estimate 108 to determine the metric scale for the depth map 120 generated by the neural network 118. For example, the depth map transformer 126 can predict surface normals 122b based on the depth map 138 and use the offset between the surface normals 122b predicted from the depth map 138 and the surface normals 122a predicted from the neural network 118 to help determine the affine parameter 132. For example, the depth map transformer 126 can minimize an objective function that can penalize the offset between the depth map 120 and the depth estimate 108, as well as the offset between the surface normals 122a predicted from the neural network 118 and the surface normals 122b predicted from the depth map 138.
[0059] In some examples, the depth map transformer 126 receives a gravity direction 112 and / or a planar region 114. The depth map transformer 126 is configured to use the gravity direction 112 and the planar region 114 (and depth estimation 108) to provide a metric scale for the depth map 120 generated by the neural network 118. Figure 1F As shown, the direction of gravity 112 can be obtained from the accelerometer 121. A planar region 114 can be detected from the image data 104. In some examples, such as... Figure 1G As shown, planar region 114 can be estimated by planar generator 123 using visual feature points 162 (e.g., SLAM points). For example, planar generator 123 can perform a planar detection algorithm (or planar fitting algorithm) to detect planar region 114 in image data 104. Using gravity direction 112 and planar region 114, depth map transformer 126 can minimize an objective function that can penalize surface normal 122b in the horizontal surface region to match gravity direction 112 (or the opposite direction of gravity direction 112 depending on the coordinate system).
[0060] like Figure 1H As shown, depth estimation 108, pose data 110, gravity direction 112, and planar region 114 can be obtained during an AR session 174 that can be performed by a client AR application 173. This can be achieved as follows: Figure 4As further discussed, an AR session 174 is initiated when a user has created or joined a multi-user AR collaborative environment. A client AR application 173 can be installed on (and executed by) a mobile computing device. In some examples, the client AR application 173 is a software development kit (SDK) that combines the operation of one or more AR applications. In some examples, in conjunction with other components of the depth estimation system 100 (e.g., depth estimation generator 106, sensor system 102, etc.), the client AR application 173 is configured to detect and track the device's position relative to physical space to obtain pose data 110, detect the size and position of different types of surfaces (e.g., horizontal, vertical, angled surfaces) to obtain planar regions 114, obtain the direction of gravity 112 from the accelerometer 121, and generate depth estimates 108 (e.g., visual feature points 162). During the AR session 174, users can add virtual objects to the scene 125, and multiple users can then join the AR environment to simultaneously view and interact with these virtual objects from different locations in a shared physical space.
[0061] like Figure 1I As shown, the depth estimation system 100 may include a convolutional neural network (CNN) trainer 155 configured to train or update a neural network 118. In some examples, the accuracy of the depth map 138 can be improved by predicting both depth and surface normals 122a. Surface normals can be considered as higher-order structural priors because all pixels belonging to the same 3D plane will have the same normal but not necessarily the same depth. Therefore, by training the neural network 118 to predict surface normals 122a in the same way, the neural network 118 is trained to reason / infer higher-order knowledge about planes in scene 125. This can result in a smoother depth for planar regions in scene 125 where virtual objects are typically placed.
[0062] To encourage consistency between the predicted depth and the surface normal 122a, a self-consistency loss 182 (e.g., unsupervised self-consistency loss) is used during the training of the neural network 118. For example, the neural network 118 predicts a depth map 120 and a surface normal 122a based on image frame 104a, and a depth map transformer 126 predicts a surface normal 122b based on a depth map 138. The self-consistency loss 182 is calculated based on the difference between the surface normal 122a and the surface normal 122b. A loss 180 (e.g., supervised loss) is calculated based on the difference between the surface normal 122a and the ground truth normal 122c. A total loss 184 is calculated based on the loss 180 and the self-consistency loss 182 (e.g., loss 180 is added to the self-consistency loss 182). The self-consistency loss 182 encourages the neural network 118 to minimize any deviation between the surface normal 122a and the surface normal 122b.
[0063] Figure 2 The diagram illustrates an example of a neural network 218. Neural network 218 can be... Figures 1A to 1I Examples of neural network 118 may be provided, and any details discussed with reference to those figures may be included. In some examples, neural network 218 is a convolutional neural network. Neural network 218 receives image frame 204a and generates depth map 220. Depth map 220 may be... Figures 1A to 1I Examples of depth maps 120 may be provided, and any details discussed may be included with reference to those maps. Additionally, in some examples, the neural network 218 is configured to predict surface normals (e.g., Figures 1A to 1I Surface normal 122a) and planar region 124 (e.g., Figures 1A to 1I (planar region 124). In some examples, neural network 218 includes a U-net architecture configured to predict pixel-wise depth based on a red-green-blue (RGB) image, wherein the U-net architecture is an encoder-decoder with skip connections having learnable parameters.
[0064] The neural network 118 may include a plurality of downsampler units such as downsampler unit 248-1, downsampler unit 248-2, downsampler unit 248-3, downsampler unit 248-4, and downsampler unit 248-5, and a plurality of upsampler units such as upsampler unit 249-1, upsampler unit 249-2, upsampler unit 249-3, upsampler unit 249-4, and upsampler unit 249-5. Each downsampler unit (e.g., 248-1, 248-2, 248-3, 248-4, 248-5) includes a depthwise separable convolution 252, a rectified linear activation function (ReLU) 254, and a max pooling operation 256. Each upsampling unit (e.g., 249-1, 249-2, 249-3, 249-4, 249-5) includes a depthwise separable convolution 252, a rectified linear activation function (ReLU) 254, and a bilinear upsampling operation 258. Finally, the output of the upsampling unit (e.g., 249-5) is fed to the depthwise separable convolution 252, followed by the rectified linear activation function (ReLU).
[0065] The depthwise separable convolution 252 includes factorization convolutions that factorize the standard convolution into depthwise convolutions and 1x1 convolutions called pointwise convolutions. This decomposition has the effect of reducing computation and model size. Additionally, the use of bilinear upsampling operation 258 reduces the parameters that transpose the convolution and thus reduces the network size. In some examples, the neural network 218 may use a Blurpool encoder, which can be a combination of anti-aliasing and subsampling operations that make the neural network 218 more robust and stable to damage such as rotation, scaling, blurring, and noise variations.
[0066] Figure 3 The illustration shows an example of a plane generator 390 that uses a metric depth map 338 to detect or identify one or more planar regions 395 (e.g., metric planar regions). For example, the location and size of the planar region 395 can be identified based on information from the metric scale. In some examples, the plane generator 390 is included... Figures 1A to 1I The depth estimation system 100 may include any details discussed with reference to those graphs. Metric planar regions may be planar surfaces of objects within an image having metric scales. In some examples, the planar generator 390 may receive a metric depth map 338 and pose data 310 and detect one or more planar regions 395 from the metric depth map 338.
[0067] As indicated above, affine ambiguity can pose difficulties for some applications that require (or benefit from) real-world scales. For example, mobile AR applications may involve placing virtual objects within a camera view that has real-world dimensions. However, in order to render objects at real-world scales, it may be necessary to estimate the depth of the surface on which the virtual object is placed in metric units. According to the embodiments discussed herein, metric depth map 338 (e.g., by...) Figures 1A to 1I A depth estimation system 100 (generated by the system) can be used to estimate at least one planar region 395 in image data, wherein the at least one planar region 395 is configured to be used as a surface to which a virtual object is attached. By using a metric depth map 338, the latency for detecting the planar region 395 can be reduced. For example, a depth estimation system (e.g., Figures 1A to 1I The depth estimation system 100 can reduce placement latency by using a convolutional neural network to predict the scale of the depth of the placed object / plane surface (e.g., estimating depth from a single image or a small number of images, thus requiring less user movement). Furthermore, the depth estimation system can predict depth based on low-texture surfaces such as a white table.
[0068] The plane generator 390 may include a mapping converter 392 configured to convert a metric depth map 338 into a point cloud 394. The plane generator 390 may include a plane detector 396 that performs a plane fitting algorithm configured to detect one or more planar regions 395 using the point cloud 394. The plane generator 390 includes a verification model 398 configured to process the planar regions 395, which may reject one or more planar regions 395 based on visibility and other constraints.
[0069] Figure 4 The illustration is based on one aspect of the AR system 450. (Reference) Figure 4The AR system 450 includes a first computing device 411-1 and a second computing device 411-2, wherein users of the first computing device 411-1 and the second computing device 411-2 are able to view and interact with one or more virtual objects 430 included in a shared AR environment 401. Although Figure 4 The illustration shows two computing devices, but the embodiment includes any number of computing devices (e.g., more than two) capable of joining the shared AR environment 401. The first computing device 411-1 and the second computing device 411-2 are configured to communicate with an AR collaborative service 415 executable by a server computer 461 via one or more application programming interfaces (APIs).
[0070] AR Collaborative Service 415 is configured to create multi-user or collaborative AR experiences that users can share. AR Collaborative Service 415 communicates via network 451 with multiple computing devices, including a first computing device 411-1 and a second computing device 411-2, where users of the first computing device 411-1 and users of the second computing device 411-2 can share the same AR environment 401. AR Collaborative Service 415 can allow users to create 3D maps for creating multi-player or collaborative AR experiences that users can share with other users. Users can add virtual objects 430 to a scene 425, and multiple users can then simultaneously view and interact with these virtual objects 430 from different locations in a shared physical space.
[0071] The first computing device 411-1 and / or the second computing device 411-2 can be any type of mobile computing system such as a smartphone, tablet, laptop, wearable device, etc. Wearable devices can include head-mounted display (HMD) devices such as optical head-mounted displays (OHMD), transparent head-up displays (HUD), augmented reality (AR) devices, or other devices such as goggles or headsets with sensors, displays, and computing capabilities. In some examples, wearable devices include smart glasses. Smart glasses are optical head-mounted display devices designed in the shape of eyeglasses. For example, smart glasses are glasses that add information next to what the wearer sees through eyeglasses.
[0072] AR environment 401 may involve a physical space within the user's view and a virtual space in which one or more virtual objects 430 are positioned. Figure 4The virtual object 430 shown is illustrated as a box but can include any type of virtual object added by the user. Providing (or rendering) the AR environment 401 can then involve altering the user's view of the physical space by displaying the virtual objects 430, such that they appear to the user to be presented in the physical space within the user's view, or overlaid on or within the physical space within the user's view. The display of the virtual objects 430 is therefore based on a mapping between virtual and physical space. Overlaying the virtual objects 430 can be achieved, for example, by superimposing the virtual objects 430 onto the user's light field in the physical space, by reproducing the user's view of the physical space on one or more display screens, and / or in other ways—e.g., by using a heads-up display, a mobile device display screen, etc.
[0073] The first computing device 411-1 and / or the second computing device 411-2 include a depth estimation system 400. The depth estimation system 400 is... Figures 1A to 1I Examples of depth estimation system 100 may be provided, and may include any details discussed with reference to those diagrams. Depth estimation system 400 uses image data captured by first computing device 411-1 to generate a metric depth map, and this metric depth map is used to detect one or more planar regions 495 according to any of the techniques discussed above. In some examples, planar regions 495 may be visually illustrated to a user, allowing the user to view planar regions 495 and attach virtual objects 430 to them. For example, a user of first computing device 411-1 may use planar regions 495 to attach virtual objects 430. When second computing device 411-2 enters the same physical space, AR collaborative service 415 may render AR environment 401 onto the screen of second computing device 411-2, where a user may view and interact with virtual objects 430 added by the user of first computing device 411-1. The second computing device 411-2 may include a depth estimation system 400 configured to generate a metric depth map and use the metric depth map to detect one or more planar regions 495, wherein a user of the second computing device 411-2 may add one or more other virtual objects 430 to the detected planar regions 495, wherein a user of the first computing device 411-1 will be able to view and interact with them.
[0074] Figure 5 The diagram illustrates a flowchart 500 depicting an example operation of a depth estimation system. (Although reference...) Figures 1A to 1I The depth estimation system 100 describes the operation, but Figure 5 The operation can be applied to any system described in this article. Although Figure 5Flowchart 500 illustrates operations in a sequential order; however, it should be understood that this is merely an example and may include additional or alternative operations. Furthermore, operations may be performed in a different order than shown, or in parallel or overlapping manner. Figure 5 The operation and related operations.
[0075] Operation 502 includes receiving image data 104 from sensor system 102. Operation 504 includes generating a depth map 120 (e.g., a first depth map) based on the image data 104 by neural network 118, wherein the depth map 120 has a first scale. Operation 506 includes obtaining a depth estimate 108 associated with the image data 104. Operation 508 includes using the depth estimate 108 to transform the depth map 120 into a depth map 138 (e.g., a second depth map), wherein the depth map 138 has a second scale. The first scale and the second scale are different scales that can be based on two different measurement systems with different standards. In some examples, the first scale is a non-metric scale. In some examples, the second scale is a metric scale. Additionally, the depth estimate 108 has a depth value corresponding to the second scale.
[0076] Figure 6 The diagram illustrates a flowchart 600 depicting an example operation of a depth estimation system. (Although reference...) Figures 1A to 1I The depth estimation system 100 describes the operation, but Figure 6 The operation can be applied to any system described in this article. Although Figure 6 Flowchart 600 illustrates operations in a sequential order; however, it should be understood that this is merely an example and may include additional or alternative operations. Furthermore, operations may be performed in a different order than shown, or in parallel or overlapping manner. Figure 6 The operation and related operations.
[0077] Operation 602 includes predicting a depth map 120 (e.g., a first depth map) and a first surface normal 122a based on image frame 104a by neural network 118, wherein the depth map 120 has a first scale (e.g., a non-metric scale). Operation 604 includes obtaining a depth estimate 108 associated with image data 104. In some examples, the depth estimate 108 has depth values according to a second scale (e.g., a metric scale). Operation 606 includes using the depth estimate 108 to transform the depth map 120 into a depth map 138 (e.g., a second depth map), wherein the depth map 138 has a second scale (e.g., a metric scale). Operation 608 includes estimating a second surface normal 122b based on the depth map 138. Additionally, note that the first and second scales are different scales that can be based on two different measurement systems with different standards.
[0078] Operation 610 includes calculating a self-consistency loss 182 based on the difference between the first surface normal 122a and the second surface normal 122b. In some examples, the self-consistency loss 182 is an unsupervised loss. In some examples, flowchart 600 includes calculating a loss 180 (e.g., a supervised loss) based on the difference between the first surface normal 122a and the ground reality normal 122c. Operation 612 includes updating the neural network 118 based on the self-consistency loss 182. In some examples, the neural network 118 is updated based on both the self-consistency loss 182 and the loss 180.
[0079] Figure 7 The illustration depicts a flowchart 700 illustrating an example operation of a depth estimation system. (Although reference...) Figures 1A to 1I Depth estimation system 100 and Figure 4 The AR system 450 describes the operation, but Figure 7 The operation can be applied to any system described in this article. Although Figure 7 Flowchart 700 illustrates operations in a sequential order; however, it should be understood that this is merely an example and may include additional or alternative operations. Furthermore, operations may be performed in a different order than shown, or in parallel or overlapping manner. Figure 7 The operation and related operations.
[0080] Operation 702 includes receiving image data 104 from sensor system 102. Operation 704 includes generating a depth map 120 (e.g., a first depth map) based on image data 104 by neural network 118, wherein depth map 120 has a first scale (e.g., a non-metric scale). Operation 706 includes obtaining a depth estimate 108 associated with image data 104. In some examples, depth estimate 108 has depth values according to a second scale (e.g., a metric scale). Operation 708 includes using depth estimate 108 to transform depth map 120 into depth map 138 (e.g., a second depth map), wherein depth map 138 has a second scale (e.g., a metric scale). Operation 710 includes using depth map 138 to estimate at least one planar region 495 in image data 104, wherein at least one planar region 495 is configured to be used as a surface to attach virtual object 430 during augmented reality (AR) session 174.
[0081] Example 1. A method for depth estimation, the method comprising: receiving image data from a sensor system; generating a first depth map based on the image data by a neural network, wherein the first depth map has a first scale; obtaining a depth estimate associated with the image data; and using the depth estimate to transform the first depth map into a second depth map, wherein the second depth map has a second scale.
[0082] Example 2. The method according to Example 1 further includes: generating surface normals by the neural network based on the image data.
[0083] Example 3. The method according to any one of Examples 1 to 2, wherein the surface normal and the depth estimate are used to transform the first depth map into the second depth map.
[0084] Example 4. The method according to any one of Examples 1 to 3 further includes: generating visual feature points based on the image data, the visual feature points being associated with the depth estimation.
[0085] Example 5. The method according to any one of Examples 1 to 4 further includes: obtaining the depth estimate from a depth sensor.
[0086] Example 6. The method according to any one of Examples 1 to 5, wherein the depth estimate is obtained during an augmented reality (AR) session that can be performed by a mobile computing device.
[0087] Example 7. The method according to any one of Examples 1 to 6, further comprising: estimating affine parameters based on an offset between the first depth map and the depth estimate, the affine parameters including scale and shift, wherein the first depth map is transformed into the second depth map based on the affine parameters.
[0088] Example 8. The method according to any one of Examples 1 to 7 further includes: predicting a first surface normal by the neural network based on the image data; and predicting a second surface normal based on the second depth map.
[0089] Example 9. The method according to any one of Examples 1 to 8 further includes: calculating a self-consistency loss based on the first surface normal and the second surface normal.
[0090] Example 10. The method according to any one of Examples 1 to 9 further includes: updating the neural network based on the self-consistency loss.
[0091] Example 11. The method according to any one of Examples 1 to 10, further comprising: using the second depth map to estimate at least one planar region in the image data, the at least one planar region being configured as a surface to which a virtual object is to be attached.
[0092] Example 12. A depth estimation system comprising: a sensor system configured to acquire image data; a neural network configured to generate a first depth map based on the image data, the first depth map having a first scale; a depth estimation generator configured to acquire a depth estimate associated with the image data; and a depth map transformer configured to estimate affine parameters based on the depth estimate and the first depth map and to use the affine parameters to transform the first depth map into a second depth map having a second scale.
[0093] Example 13. The depth estimation system according to Example 12, wherein the neural network is configured to execute on a mobile computing device.
[0094] Example 14. The depth estimation system according to any one of Examples 12 to 13 further includes: a visual inertial motion tracker configured to generate visual feature points associated with the depth estimation.
[0095] Example 15. The depth estimation system according to any one of Examples 12 to 14 further includes: a depth sensor configured to obtain the depth estimate.
[0096] Example 16. A depth estimation system according to any one of Examples 12 to 15, wherein the depth estimation generator is configured to obtain the depth estimate during an augmented reality (AR) session, and the depth estimation generator is configured to also obtain pose data, gravity direction, and identification of one or more planar regions in the image data during the AR session.
[0097] Example 17. A depth estimation system according to any one of Examples 12 to 16, wherein the affine parameters include a scale and shift for each depth estimate in the first depth map.
[0098] Example 18. A depth estimation system according to any one of Examples 12 to 17, wherein the depth map transformer includes a Random Sample Consensus (RANSAC)-based solver that minimizes an objective function to estimate the scale and the shift.
[0099] Example 19. A depth estimation system according to any one of Examples 12 to 18, further comprising: a neural network trainer configured to use the neural network to predict a first surface normal based on the image data; predict a second surface normal based on the second depth map; calculate a self-consistency loss based on the first surface normal and the second surface normal; calculate a loss based on the first surface normal and the ground reality normal; and / or update the neural network based on the self-consistency loss and the loss.
[0100] Example 20. A depth estimation system according to any one of Examples 12 to 19, further comprising: a plane generator configured to use the second depth map to estimate at least one planar region in the image data, the at least one planar region being configured to be used as a surface to which a virtual object is attached, the plane generator including a graph converter configured to convert the second depth map into a point cloud; and a plane detector configured to use the point cloud to detect the at least one planar region according to a plane fitting algorithm.
[0101] Example 21. A non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause the at least one processor to: receive image data from a sensor system; generate a first depth map having a first scale by a neural network based on the image data; obtain a depth estimate associated with the image data; transform the first depth map into a second depth map having a second scale using the depth estimate; and use the second depth map to estimate at least one planar region in the image data, the at least one planar region being configured to be used as a surface to attach virtual objects during an augmented reality (AR) session.
[0102] Example 22. A non-transitory computer-readable medium according to Example 21, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to: estimate affine parameters based on an offset between the first depth map and the depth estimate, the affine parameters including scale and shift, wherein the first depth map is transformed into the second depth map based on the affine parameters.
[0103] Example 23. A non-transitory computer-readable medium according to any one of Examples 21 to 22, wherein the depth estimate is obtained from at least one of a visual inertial motion tracker, a depth sensor, a dual-pixel depth estimator, a moving stereo depth estimator, a sparse active depth estimator, and / or a pre-computed sparse graph.
[0104] Example 24. A non-transitory computer-readable medium according to any one of Examples 21 to 23, wherein the executable instructions include instructions that, when executed by the at least one processor, cause the at least one processor to: generate surface normals from the neural network based on the image data, wherein the surface normals and the depth estimation are used to transform the first depth map into the second depth map.
[0105] Figure 8 Examples of example computer devices 800 and 850 that can be used with the techniques described herein are shown. The computing device 800 includes a processor 802, a memory 804, a storage device 806, a high-speed interface 808 connected to the memory 804 and a high-speed expansion port 810, and a low-speed interface 812 connected to a low-speed bus 814 and the storage device 806. Each of components 802, 804, 806, 808, 810, and 812 is interconnected using various buses and can be mounted on a general-purpose motherboard or otherwise. The processor 802 can process instructions for execution within the computing device 800, including instructions stored in the memory 804 or storage device 806 to display graphical information for a GUI on an external input / output device—such as a display 816 coupled to the high-speed interface 808. In other embodiments, multiple processors and / or multiple buses, as well as multiple memories and memory types, may be suitably used. In addition, multiple computing devices 800 can be connected, each providing some of the necessary operations (e.g., as a server library, a set of blade servers, or a multiprocessor system).
[0106] The memory 804 stores information within the computing device 800. In one embodiment, the memory 804 is one or more volatile storage units. In another embodiment, the memory 804 is one or more non-volatile storage units. The memory 804 may also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0107] Storage device 806 provides large-capacity storage for computing device 800. In one embodiment, storage device 806 may be or include computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices, or arrays of devices, including devices in storage area networks or other configurations. A computer program product may be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 804, storage device 806, or memory on processor 802.
[0108] High-speed controller 808 manages bandwidth-intensive operations of computing device 800, while low-speed controller 812 manages lower bandwidth-intensive operations. This functional allocation is merely exemplary. In one embodiment, high-speed controller 808 is coupled to memory 804, display 816 (e.g., via a graphics processor or accelerator), and is coupled to high-speed expansion port 810, which can accept various expansion cards (not shown). In another embodiment, low-speed controller 812 is coupled to storage device 806 and low-speed expansion port 814. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, Wireless Ethernet), may be coupled to one or more input / output devices, such as keyboards, pointing devices, scanners, or networking devices such as switches or routers, for example, via a network adapter.
[0109] As shown, the computing device 800 can be implemented in a variety of different forms. For example, it can be implemented as a standard server 820, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 824. Furthermore, it can be implemented in a personal computer such as a laptop computer 822. Alternatively, components from the computing device 800 can be combined with other components in a mobile device (not shown), such as device 850. Each such device can contain one or more of the computing devices 800, 850, and the entire system can consist of multiple computing devices 800, 850 communicating with each other.
[0110] Among other components, computing device 850 includes processor 852, memory 864, input / output devices such as display 854, communication interface 866, and transceiver 868. Storage devices, such as microdrives or other devices, may also be provided to device 850 to provide additional storage. Each of components 850, 852, 864, 854, 866, and 868 is interconnected using various buses, and multiple components may be mounted on a general-purpose motherboard or otherwise, as appropriate.
[0111] Processor 852 can execute instructions within computing device 850, including instructions stored in memory 864. The processor can be implemented as a chipset comprising individual or multiple analog and digital processors. For example, the processor can provide coordination for other components of device 850, such as control of the user interface, applications running on device 850, and wireless communication of device 850.
[0112] Processor 852 can communicate with the user via control interface 858 and display interface 856 coupled to display 854. For example, display 854 can be a TFT LCD (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technology. Display interface 856 may include appropriate circuitry for driving display 854 to present graphics and other information to the user. Control interface 858 can receive commands from the user and translate them for submission to processor 852. Furthermore, an external interface 862 can be provided to communicate with processor 852 to enable near-field communication between device 850 and other devices. For example, external interface 862 may provide wired communication in some embodiments or wireless communication in others, and multiple interfaces may be used.
[0113] Memory 864 stores information within computing device 850. Memory 864 can be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 874 may also be provided and connected to device 850 via an extended interface 872, which may include a SIMM (Single In-line Memory Module) card interface. This extended memory 874 can provide additional storage space for device 850, or it can store applications or other information for device 850. Specifically, extended memory 874 may include instructions for performing or supplementing the above processes, and may also include security information. Therefore, for example, extended memory 874 can be provided as a security module for device 850 and can be programmed with instructions that allow secure use of device 850. Furthermore, secure applications and additional information can be provided via a SIMM card, such as by placing identification information on the SIMM card in an intrusive manner.
[0114] For example, the memory may include flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 864, extended memory 874, or memory on processor 852, which, for example, can be received via transceiver 868 or external interface 862.
[0115] Device 850 can communicate wirelessly via a communication interface 866, which may include digital signal processing circuitry if necessary. The communication interface 866 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, etc. For example, such communication can occur via a radio frequency transceiver 868. Furthermore, short-range communication, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown), is possible. Additionally, a GPS (Global Positioning System) receiver module 870 can provide device 850 with additional navigation and location-related wireless data, which can be appropriately used by applications running on device 850.
[0116] Device 850 can also communicate audibly using an audio codec 860 that can receive spoken information from a user and convert it into usable digital information. The audio codec 860 can similarly generate audible sounds for the user, such as through a speaker, for example, in the handset of device 850. Such sounds can include sounds from voice telephone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications operating on device 850.
[0117] As shown in the figure, the computing device 850 can be implemented in many different forms. For example, it can be implemented as a cellular phone 880. It can also be implemented as a smartphone 882, a personal digital assistant, or part of another similar mobile device.
[0118] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and coupled to transfer data and instructions to the storage system, at least one input device, and at least one output device. Additionally, the term "module" can include both software and / or hardware.
[0119] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) to display information to the user and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, speech, or tactile input.
[0121] The systems and techniques described herein can be implemented in computing systems that include back-end components (e.g., as data servers), middleware components (e.g., application servers), or front-end components (e.g., client computers having a graphical user interface or web browser through which users can interact with implementations of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. Components of the system can be interconnected via any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.
[0122] A computing system may include clients and servers. Clients and servers are typically geographically separated and interact via communication networks. The relationship between clients and servers is created by computer programs running on their respective computers that have a client-server relationship with each other.
[0123] In some implementations... Figure 8 The computing device depicted may include sensors that interface with virtual reality (VR headset 890). For example, including... Figure 8One or more sensors on the computing device 850 or other computing device depicted can provide input to the VR headset 890, or generally, to the VR space. Sensors may include, but are not limited to, touchscreens, accelerometers, gyroscopes, pressure sensors, biometric sensors, temperature sensors, humidity sensors, and ambient light sensors. The computing device 850 can use the sensors to determine the absolute position of the computing device in the VR space and / or detect rotation, which can then be used as input to the VR space. For example, the computing device 850 can be incorporated into the VR space as a virtual object such as a controller, laser pointer, keyboard, weapon, etc. User placement of the computing device / virtual object, when incorporated into the VR space, allows the user to position the computing device to view the virtual object in a certain way within the VR space. For example, if the virtual object represents a laser pointer, the user can manipulate the computing device as if it were an actual laser pointer. The user can move the computing device left and right, up and down, in circles, etc., and use the device in a manner similar to using a laser pointer.
[0124] In some implementations, one or more input devices included in or connected to the computing device 850 can be used as input to the VR space. Input devices may include, but are not limited to, touchscreens, keyboards, one or more buttons, trackpads, touchpads, pointing devices, mice, trackballs, joysticks, cameras, microphones, headphones or earbuds with input capabilities, game controllers, or other connectable input devices. When the computing device is incorporated into the VR space, a user interacting with the input devices included on the computing device 850 may cause specific actions to occur within the VR space.
[0125] In some implementations, the touchscreen of computing device 850 can be rendered as a touchpad in VR space. Users can interact with the touchscreen of computing device 850. For example, in VR headset 890, the interaction is rendered as movement on a touchpad rendered in VR space. The rendered movement can control objects in VR space.
[0126] In some implementations, one or more output devices included on the computing device 850 may provide output and / or feedback to a user of the VR headset 890 in the VR space. The output and feedback may be visual, tactile, or audio. The output and / or feedback may include, but is not limited to, vibration, turning one or more lights or strobe lights on and off or flashing and / or blinking, issuing alarms, playing ringtones, playing songs, and playing audio files. Output devices may include, but are not limited to, vibration motors, vibration coils, piezoelectric devices, electrostatic devices, light-emitting diodes (LEDs), strobe lights, and speakers.
[0127] In some implementations, computing device 850 may appear as another object in a computer-generated 3D environment. User interactions with computing device 850 (e.g., rotating, shaking, touching a touchscreen, swiping a finger on the touchscreen) can be interpreted as interactions with an object in VR space. In the example of a laser pointer in VR space, computing device 850 appears as a virtual laser pointer in a computer-generated 3D environment. When the user manipulates computing device 850, the user in VR space perceives the laser pointer as moving. The user receives feedback from interactions with computing device 850 within VR space on computing device 850 or VR headset 890.
[0128] In some implementations, in addition to computing devices, one or more input devices (e.g., a mouse, a keyboard) may be rendered in a computer-generated 3D environment. The rendered input devices (e.g., a rendered mouse, a rendered keyboard) can be used during rendering in VR space to control objects in VR space.
[0129] Computing device 800 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 850 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the invention described and / or claimed herein.
[0130] Several embodiments have been described. However, it will be understood that various modifications may be made without departing from the spirit and scope of this specification.
[0131] Furthermore, the logical flow depicted in the accompanying drawings does not require the specific or sequential order shown to achieve the desired result. Additionally, other steps can be provided or removed from the described flow, and other components can be added to or removed from the described system. Therefore, other embodiments are within the scope of the appended claims.
Claims
1. A method for depth estimation, the method comprising: Receive image data from the sensor system; A first depth map and a first surface normal are generated by a neural network based on the image data, wherein the first depth map has a first scale; Obtain depth estimation data associated with the image data; The depth estimation data is used to transform the first depth map into a second depth map, the second depth map having a second scale; Estimate the second surface normal from the second depth map; The loss is calculated based on the first surface normal and the second surface normal; as well as The neural network is updated based on the loss.
2. The method according to claim 1, wherein, The first depth map is transformed into the second depth map using the first surface normal and the depth estimation data.
3. The method according to claim 1, further comprising: Visual feature points are generated based on the image data, and the depth estimation data includes the visual feature points.
4. The method according to claim 1, further comprising: The depth estimation data is obtained from the depth sensor.
5. The method according to claim 1, wherein, The depth estimation data was obtained during an augmented reality session.
6. The method of claim 1, further comprising: At least one affine parameter is estimated based on the offset between the first depth map and the depth estimation data, the at least one affine parameter including at least one of scale or shift, wherein the first depth map is transformed into the second depth map based on the at least one affine parameter.
7. The method according to any one of claims 1 to 6, further comprising: The second depth map is used to estimate at least one planar region in the image data, the at least one planar region being configured as a surface to which a virtual object is to be attached.
8. A depth estimation system, comprising: At least one processor; as well as A non-transitory computer-readable medium storing executable instructions that cause the at least one processor to perform operations, the operations including: Receive image data from the sensor system; A first depth map and a first surface normal are generated by a neural network based on the image data, wherein the first depth map has a first scale; Obtain depth estimation data associated with the image data; The depth estimation data is used to transform the first depth map into a second depth map, the second depth map having a second scale; Estimate the second surface normal from the second depth map; The loss is calculated based on the first surface normal and the second surface normal; and The neural network is updated based on the loss.
9. The depth estimation system according to claim 8, wherein, The neural network is configured to execute on a mobile computing device.
10. The depth estimation system according to claim 8, wherein, The operation further includes: Visual feature points are generated based on the image data, and the depth estimation data includes the visual feature points.
11. The depth estimation system according to claim 8, wherein, The operation further includes: The depth estimation data is obtained from the depth sensor.
12. The depth estimation system according to claim 8, wherein, The depth estimation data is acquired during an augmented reality session, and the operation further includes acquiring pose data, gravity direction, and identification of one or more planar regions in the image data during the augmented reality session.
13. The depth estimation system according to claim 8, wherein, The operation further includes: At least one affine parameter is estimated based on the offset between the first depth map and the depth estimation data, the at least one affine parameter including at least one of scale or shift.
14. The depth estimation system according to claim 13, wherein, The operation further includes: Execute the objective function to estimate at least one of the scale or the shift.
15. The depth estimation system according to claim 8, wherein, The operation further includes: The second depth map is used to estimate at least one planar region in the image data, the at least one planar region being configured to be used as a surface to which a virtual object is to be attached; Convert the second depth map into a point cloud; and The point cloud is used to detect the at least one planar region according to a planar fitting algorithm.
16. A non-transitory computer-readable medium storing executable instructions, which, when executed by at least one processor, cause the at least one processor to perform operations, the operations including: Receive image data from the sensor system; A first depth map and a first surface normal are generated by a neural network based on the image data, wherein the first depth map has a first scale; Obtain depth estimation data associated with the image data; The depth estimation data is used to transform the first depth map into a second depth map, the second depth map having a second scale; Estimate the second surface normal from the second depth map; The loss is calculated based on the first surface normal and the second surface normal; as well as The neural network is updated based on the loss.
17. The non-transitory computer-readable medium according to claim 16, wherein, The operation further includes: At least one affine parameter is estimated based on the offset between the first depth map and the depth estimation data, the at least one affine parameter including at least one of scale or shift, wherein the first depth map is transformed into the second depth map based on the at least one affine parameter.
18. The non-transitory computer-readable medium according to claim 16, wherein, The depth estimation data is obtained from at least one of the following: a visual inertial motion tracker, a depth sensor, a dual-pixel depth estimator, a moving stereo depth estimator, a sparse active depth estimator, or a pre-computed sparse map.
19. The non-transitory computer-readable medium according to claim 16, wherein, The first depth map is transformed into the second depth map using the surface normal and the depth estimation data.