Boundary Estimation Based on a Monocular Video of a Posed Person

Through the deep network method and sensor data, depth maps and wall segmentation maps are generated, and the three-dimensional boundaries of the indoor environment are then estimated, which solves the problems of inaccurate and robustness of boundary estimation in the prior art, and realizes accurate and robust estimation of environmental boundaries in augmented reality systems.

CN113748445BActive Publication Date: 2025-05-27MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080030596.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-15
Filing Date
2020-04-23
Publication Date
2025-05-27
Estimated Expiration
2040-04-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately estimate the three-dimensional boundaries of the indoor environment in an augmented reality system, especially when angles and edges are obscured.

Method used

By using a deep network method, combining data captured by image sensors and posture sensors, depth maps and wall segmentation maps are generated, thereby generating point clouds and clustering to estimate the boundaries of the environment.

Benefits of technology

Accurate estimation of room boundaries is achieved, robust, capable of handling angle and edge occlusion without relying on prior room shapes or high-quality internal point clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113748445B_ABST
    Figure CN113748445B_ABST
Patent Text Reader

Abstract

Techniques are disclosed for estimating boundaries of a room environment that is at least partially surrounded by a set of adjacent walls. A set of images and a set of poses are obtained. A depth map is generated based on the set of images and the set of poses. A set of wall segmentation maps is generated based on the set of images, where each of the set of wall segmentation maps indicates a target region of a corresponding image that includes the set of adjacent walls. A point cloud is generated based on the depth map and the set of wall segmentation maps, the point cloud including a plurality of points sampled along a portion of the depth map that is aligned with the target region. The boundaries of the environment along the set of adjacent walls are estimated based on the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 838,265, filed on April 24, 2019, entitled "SYSTEMS AND METHODS FOR DEEPINDOOR BOUNDARY ESTIMATION FROM POSED MONOCULAR VIDEO" and U.S. Provisional Patent Application No. 62 / 848,492, filed on May 15, 2019, entitled "SYSTEMS AND METHODS FOR DEEP INDOOR BOUNDARYESTIMATION FROM POSED MONOCULAR VIDEO", the entire contents of which are incorporated herein by reference. Background Art

[0003] Modern computing and display technologies have facilitated the development of systems for so - called "virtual reality" or "augmented reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that they appear or may be perceived as real. Virtual reality or "VR" scenarios typically involve the presentation of digital or virtual image information and are opaque to other actual real - world visual inputs; augmented reality or "AR" scenarios typically involve the presentation of digital or virtual image information as an enhancement to the visualization of the actual world surrounding the user.

[0004] Despite the progress made in these display technologies, there remains a need in the art for improved methods, systems, and devices related to augmented reality systems, particularly display systems. Summary of the Invention

[0005] The present disclosure relates to computing systems, methods, and configurations, and more particularly to computing systems, methods, and configurations in which understanding the three - dimensional (3D) geometric aspects of an environment is important, such as in applications that may involve computing systems for augmented reality (AR), navigation, and general scene understanding.

[0006] A system of one or more computers can be configured to perform particular operations or actions by virtue of software, firmware, hardware, or a combination thereof installed on the system, which in operation causes or induces the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of instructions that, when executed by a data processing apparatus, cause the apparatus to perform the actions. One general aspect includes a method that includes: obtaining a set of images and a set of poses corresponding to the set of images, the set of images having been captured from an environment that is at least partially surrounded by a set of adjacent walls. The method further includes generating a depth map of the environment based on the set of images and the set of poses. The method further includes generating a set of wall segmentation maps based on the set of images, each wall segmentation map in the set of wall segmentation maps indicating a target region of a corresponding image in the set of images that includes the set of adjacent walls. The method further includes generating a point cloud based on the depth map and the set of wall segmentation maps, the point cloud including a plurality of points sampled along a portion of the depth map that is aligned with the target region. The method further includes estimating a boundary of the environment along the set of adjacent walls based on the point cloud. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0007] Implementations may include one or more of the following features. The method further includes: capturing a set of images and a set of poses using one or more sensors, wherein obtaining a set of images and a set of poses includes receiving a set of images and a set of poses from one or more sensors. In the method, the one or more sensors include an image sensor configured to capture the set of images and a pose sensor configured to capture the set of poses. The method further includes: identifying a set of clusters of the point cloud, each cluster in the set of clusters including a subset of the plurality of points, and wherein each cluster in the set of clusters is determined to correspond to a different wall in the set of adjacent walls. The plurality of points includes 2D points and the method of estimating a boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting a line to the plurality of points, resulting in a plurality of lines. The method may further include forming a closed loop by extending the plurality of lines until an intersection point is reached. The plurality of points includes 3D points and the method of estimating a boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting a plane to the plurality of points, resulting in a plurality of planes. The method may further include forming a closed loop by extending the plurality of planes until an intersecting line is reached. In the method, the set of images includes RGB images. In the method, the set of poses includes the camera orientation of the image sensor that captured the set of images. In the method, the plurality of points includes 3D points. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.

[0008] One general aspect includes a system that includes: one or more processors; a computer-readable medium including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: obtaining a set of images and a set of poses corresponding to the set of images, the set of images being captured in an environment that is at least partially surrounded by a set of adjacent walls. The operations further include generating a depth map of the environment based on the set of images and the set of poses. The operations further include generating a set of wall segmentation maps based on the set of images, each wall segmentation map in the set of wall segmentation maps indicating a target region of the corresponding image in the set of images that includes the set of adjacent walls. The operations further include generating a point cloud based on the depth map and the set of wall segmentation maps, the point cloud including a plurality of points sampled along a portion of the depth map that is aligned with the target region. The operations further include estimating a boundary of the environment along the set of adjacent walls based on the point cloud. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0009] Implementations may include one or more of the following features. In the system, the operations further include: capturing a set of images and a set of poses using one or more sensors, where obtaining a set of images and a set of poses includes receiving a set of images and a set of poses from the one or more sensors. In the system, the one or more sensors include an image sensor configured to capture the set of images and a pose sensor configured to capture the set of poses. In the system, the operations further include: identifying a set of clusters of the point cloud, each cluster in the set of clusters including a subset of the plurality of points, and wherein each cluster in the set of clusters is determined to correspond to a different wall in the set of adjacent walls. The plurality of points include 2D points and the system for estimating a boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting a line to the plurality of points, resulting in a plurality of lines. The system may further include forming a closed loop by extending the plurality of lines until an intersection point is reached. The plurality of points include 3D points and the system for estimating a boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting a plane to the plurality of points, resulting in a plurality of planes. The system may further include forming a closed loop by extending the plurality of planes until an intersecting line is reached. In the system, the set of images includes RGB images. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.

[0010] Compared with traditional techniques, many benefits are obtained through the present invention. For example, embodiments of the present invention are capable of precisely inferring room boundaries using currently available deep network methods without enumerating a possible set of room types and are robust to corner and edge occlusions. For example, the described embodiments do not rely on a list of possible prior room shapes. Additionally, the described embodiments do not rely on the availability of high-quality internal point clouds at the model input. Furthermore, the results using the described embodiments have established important benchmarks for boundary estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 Shows an example implementation of the present invention for boundary estimation of a room environment using a head-mounted display.

[0012] Figure 2 Shows an example system for estimating room boundaries using pose images.

[0013] Figure 3A and Figure 3B Shows an example of wall segmentation.

[0014] Figures 4A - 4D Shows an example step for generating a boundary.

[0015] Figure 5 Shows example training data for training a clustering identifier.

[0016] Figure 6 Shows a method for estimating the boundary of an environment.

[0017] Figure 7 Shows an example system or device in which one or more of the described techniques can be implemented.

[0018] Figure 8 Shows a simplified computer system. DETAILED DESCRIPTION

[0019] Understanding the internal three-dimensional (3D) layout can be important for understanding the long-range geometry of a space through numerous applications in augmented reality (AR), navigation, and general scene understanding. Such layouts can be presented in various ways, including cuboid parameters, monocular corner coordinates and their connectivity, and more semantically rich complete floor plans. The various methods differ in the amount of information they utilize at the input and their assumptions about the room geometry. For example, some methods utilize a clean 3D point cloud at the input, while other methods utilize monocular perspective or panoramic images. The lack of consistency among this set of related problems reveals a general disagreement about what the standard setup should be for indoor scene layout prediction.

[0020] In terms of sensor data, timestamped red, green, and blue (RGB) camera and pose data can be obtained from many modern devices such as smartphones, AR and virtual reality (VR) head-mounted displays (HMDs), etc. For a complete video sequence corresponding to the interior, the problem to be solved goes beyond the corner and edge estimation prevalent in monocular layout estimation and becomes an estimation of the full boundary layout of the interior space. Such metric information about the spatial extent and the shape of the space can be regarded as the first step for various downstream 3D applications. Embodiments of the present invention are capable of using current depth methods to accurately infer this boundary without enumerating a possible set of room types and are robust against corner and edge occlusions. In some cases, the horizontal boundary (i.e., the location of the exterior wall) can be predicted because it contains the vast majority of the structure within the room layout, and the walls and ceilings are typically well approximated by a single plane.

[0021] In some embodiments, the disclosed pipeline begins with a deep depth estimation of the RGB frames of a video sequence. One of the most restrictive bottlenecks in general 3D reconstruction applications of deep learning is the accuracy of the deep depth estimation model. In cluttered indoor scenes such as in the NYUv2 dataset, given a monocular input, such networks may still struggle to achieve an RMS error better than 0.5 - 0.6 meters. Through topic configuration, the performance bottleneck can be bypassed by incorporating temporal information into the depth estimation module by switching to a modern multi-view stereo method.

[0022] Using such embodiments, depth segmentation can be trained to isolate depth predictions corresponding to wall points. These predictions are projected into a 3D point cloud and then clustered by a novel depth network that is tuned to detect points belonging to the same plane instance. Once the point clusters are assigned, a method is employed to transform the clusters into a complete set of planes forming the complete boundary layout. By directly clustering the wall points, the embodiments provided herein perform well even when corners are occluded.

[0023] The embodiments disclosed herein relate to an unsupervised pipeline for generating a complete indoor boundary (i.e., an external boundary map) from a monocular sequence of posed RGB images. In some embodiments of the present invention, various robust depth methods can be employed for depth estimation and wall segmentation to generate an external boundary point cloud, and then depth unsupervised clustering is used to fit the wall plane to obtain the final boundary map of the room. The embodiments of the present invention achieve excellent performance on the popular ScanNet dataset and are applicable to various complex room shapes as well as multi-room scenarios.

[0024] Figure 1 An example implementation of the present invention for boundary estimation of a room environment 100 using an HMD (such as an AR / VR HMD) is shown. In the example shown, a user wearing a wearable device 102 navigates in a room along a trajectory 104, allowing the image capture device of the wearable device 102 to capture a series of images I 1 , T 2 ,..., T N at a series of timestamps. Each image may include a portion of one or more of a set of walls 108 forming the room environment 100. The wearable device 102 may further capture a series of poses P 1 , I 2 ,..., I N at a series of timestamps T 1 , T 2 ,..., T N such that each image can be associated with a pose. 1 , P 2 ,..., P N , such that each image can be associated with a pose.

[0025] Each pose may include a position and / or an orientation. The position may be a 3D value (e.g., X, Y, and Z coordinates) and may correspond to the position from which the corresponding image is captured. The orientation may be a 3D value (e.g., pitch angle, yaw angle, roll angle) and may correspond to the orientation from which the corresponding image is captured. Any sensor (referred to as a pose sensor) that captures data indicating the movement of the wearable device 102 can be used to determine the pose. Based on a series of images I 1 , I 2 ,..., I N and the corresponding poses P 1 , P 2 ,..., P N , an estimated boundary 106 of the room environment 100 can be generated.

[0026] Figure 2An example system 200 for estimating room boundaries using pose images according to some embodiments of the present invention is shown. In some implementations, the system 200 can be incorporated into a wearable device. The system 200 can include a sensor 202, which includes an image sensor 206 and a pose sensor 208. The image sensor 206 can capture a series of images 210 of a set of walls. Each image 210 can be an RGB image, a grayscale image, and other possibilities. The pose sensor 208 can capture pose data 212, which can include a series of poses corresponding to the images 210. Each pose can include a position and / or orientation corresponding to the image sensor 206, such that the position and / or orientation from which each image 210 is captured can be determined.

[0027] The images 210 and the pose data 212 can be provided (e.g., sent, transmitted, etc. via a wired or wireless connection) to a processing module 204 for processing of the data. The processing module 204 can include a depth map generator 214, a wall segmentation generator 216, a wall point cloud generator 222, a clustering recognizer 226, and a boundary estimator 230. Each of the components in the processing module 204 can correspond to a hardware component, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC) and other possible integrated circuit devices. In some cases, one or more of the components in the processing module 204 can be software implemented and can be executed by a central processing unit (CPU), a graphics processing unit (GPU), or a dedicated processing unit (such as a neural network accelerator). For example, two or more of the components of the processing module 204 can be executed using the same set of neural network accelerators or the same CPU or GPU.

[0028] The depth map generator 214 can receive the images 210 and the pose data 212 and can generate a depth map 218 based on the images 210 and the pose data 212. To generate a depth map from a posed monocular sequence, multiple observations of the same real-world scene from different poses can be utilized to generate a per-frame disparity map, which can then be inverted to generate a per-frame dense depth map. In some embodiments, the depth map generator 214 can be implemented as a machine learning model, such as a neural network. When projected into the reference frame I i , the input to the network is the input RGB image I j and a 3D cost volume V i constructed by computing the per-pixel absolute intensity difference between I i and an adjacent frame I j . To project the intensity p at the position u=(u, v) i in I T i ​u Pixels of i, pose T of the reference frame i and pose T of the adjacent frame j and the assumed depth dn can be used as:

[0029]

[0030] where is the position projected in I j and π represents the pinhole projection using the camera intrinsic parameters. By changing d min between d max and d n the value of the 3D cost volume at position u and sampled depth n in I i can be calculated as:

[0031]

[0032] To generate a cost volume from multiple adjacent frames, pairwise cost volumes can be averaged.

[0033] The wall segmentation generator 216 can receive the image 210 and can generate a wall segmentation map 220 based on the image 210. In some embodiments, the wall segmentation generator 216 can be implemented as a machine learning model, such as a neural network. One goal of a semantic segmentation network may be to classify the position of walls in a scene as they are the only points belonging to internal boundaries. However, it has been found that the floor / wall / ceiling annotations of the ScanNet dataset are rather incomplete and incorrect; thus different datasets are collected and used for floor / wall / ceiling segmentation. In some cases, the segmentation network architecture can include a standard feature pyramid network based on a ResNet-101 backbone. The predictions can be output at a stride of 4, and the upsampling layer can be implemented as a pixel shuffle operation to improve network efficiency, especially for higher stride upsampling layers (up to a stride of 8).

[0034] The wall point cloud generator 222 can receive the depth map 218 and the wall segmentation map 220 and can generate a point cloud 224 based on the depth map 218 and the wall segmentation map 220. A combined point cloud can be generated from several depth images for use by the clustering network. Sets of depth images with known pose trajectories can be fused in an implicit surface representation, and the point cloud can be extracted by a derivative of the marching cubes method. One benefit of using an implicit surface representation rather than simply back-projecting each depth pixel is that it removes redundant points and averages the noise over multiple observations, resulting in a set of smoother and clearer vertices. To remove interior wall vertices, the concept of the alpha shape can be used to create a subset of the point cloud representing its concave hull. Then, any points not within the concave hull radius r can be discarded, and the point cloud can be subsampled to N vertices.

[0035] The clustering recognizer 226 can receive the point cloud 224 and can generate a clustered point cloud 228 based on the point cloud 224. A completely unsupervised technique for clustering unordered point clouds based on plane sections can be employed, without explicitly computing the normals or plane parameters during inference. The clustering recognizer 226 can be implemented as a machine learning model, such as a neural network, and in some embodiments, the PointNet architecture along with PointNet global features can be used to output the clustering probability for each input point.

[0036] To generate unique clustering assignments for individual wall instances, a clustering technique that is robust to 3D position noise, occlusion, and variable point density is needed. Additionally, it is desirable for the clustering to be able to distinguish between the following planar walls: parallel and thus having the same point normals, but different positions in 3D space. A pairwise loss function can be formulated that allows a penalty to be incurred when two points belonging to clearly different classes are assigned the same label. Over-segmentation is not penalized because clusters belonging to the same true plane can easily be merged during post-processing. The input to the network is a set of N points P, with 3D coordinates P x =(x, y, z), point normals P n =(nx, ny, nz) and a predicted clustering probability vector Pp of length k for k clusters. The clustering loss L cluster is given by:

[0037]

[0038] where:

[0039]

[0040] and:

[0041]

[0042] To label the noise points that do not belong to the valid walls, when the maximum number of clusters is set to k, a probability vector of length k + 1 can be predicted such that the (k + 1)-th label is reserved for the noise points. To prevent the trivial solution where all points are assigned to the (k + 1)-th cluster, a regularization loss L reg is calculated as follows:

[0043]

[0044] where P plane i is the sum of the probability vectors of the first k classes, excluding the (k + 1)-th noise class. The total loss to be minimized is the sum of L cluster and α·L reg where α is a hyperparameter used to balance the losses.

[0045] The boundary estimator 230 can receive the clustered point cloud 228 and can generate an estimated boundary 232 based on the clustered point cloud 228. Given a point cloud with cluster labels generated by the cluster identifier 226, a closed layout can be generated as described below. To keep the system design fairly modular, no assumptions about the modality of the input points are required, and thus the design may be robust to spurious points, outliers in the labels, and missing walls. Assuming all walls are parallel to the Z-axis, all points can be projected onto the X-Y plane to generate a top-down view of the point cloud. To establish connections between clusters, the problem is formulated as a traveling salesman problem to find the shortest closed path along the medians of all clusters. In some embodiments, a 2-opt algorithm can be used to compute the solution.

[0046] If the number of true walls in the scene is less than the maximum of k hypothesized walls, there may be over-segmentation by the cluster identifier 226. Thus, 2D line parameters can be estimated (e.g., using RANSAC), and walls with a relative normal deviation of less than 30° and an inter-cluster point-to-line error of less than e merge can be assigned the same label. After the merging step, according to the Manhattan assumption, the lines are snapped to the nearest orthogonal axes and extended to intersect. The intersection of two wall segments connected by a 2-opt edge is defined as a corner. For cases of major occlusion where an entire wall segment is not represented in the point cloud, the two connected segments may be parallel. To generate a corner for such pairs, the endpoints of one of the segments can be extended in the orthogonal direction to force an intersection.

[0047] Figure 3A and Figure 3BAn example of wall segmentation performed by a wall segmentation generator 216 according to some embodiments of the present invention is shown. In Figure 3A , an image 302 captured by an image sensor is shown. When the image 302 is provided as an input to the wall segmentation generator 216, a wall segmentation map 304 as shown in Figure 3B is generated. The wall segmentation map 304 may include regions determined to correspond to walls and regions not determined to correspond to walls. The former may each be designated as a target region 306 such that the target regions 306 of the corresponding image 302 are similarly determined to correspond to (e.g., include) walls.

[0048] Figures 4A - 4D An example of steps for generating a boundary according to some embodiments of the present invention is shown. In Figure 4A , a raw output is generated by a clustering recognizer 226 with possible over-segmentation. In Figure 4B , a more compact clustering is obtained by the clustering recognizer 226 by merging duplicate clusters. In Figure 4C , estimated line parameters and inter-cluster connectivity are produced. In some cases, the center of each cluster is determined by identifying the median or average point. As described above, the centers of the clusters are then connected by finding the closed shortest path along all cluster centers.

[0049] In Figure 4D , a line or plane is fitted to each cluster using, for example, a curve fitting algorithm. For example, if the points in each cluster are 2D points, a line is fitted to the 2D points for each cluster. If the points in each cluster are 3D points, a plane is fitted to the 3D points for each cluster. In some embodiments, orthogonality and intersection may be forced on connected parallel lines to generate a closed boundary. For example, since lines 402-1 and 402-3 are parallel and would otherwise not connect, line 402-2 is formed to connect the lines and close the boundary. Similarly, since lines 402-3 and 402-5 are parallel and would otherwise not connect, line 402-4 is formed to connect the lines and close the boundary.

[0050] Figure 5An example training data 500 for training a clustering recognizer 226 according to some embodiments of the present invention is shown. In some embodiments, the training data 500 may be a complete synthetic data set with synthetic normals. To generate the training data 500, room boundaries may be drawn on a 2D domain having a room shape randomly sampled from a rectangle, an L shape, a T shape, or a U shape, where the length of each side is uniformly sampled in the range [1m, 5m], and the orientation is uniformly sampled in the range [0, 2π]. Then the line drawing may be vertically projected to obtain a 3D model with a height randomly sampled in the range [1.5m, 2.5m]. Then a point cloud input may be generated by uniformly sampling from the 3D faces of the model. The point normals are calculated according to the 3D faces of the model.

[0051] To better simulate data generated from imperfect sensors or depth estimation algorithms, points may be placed in a plurality of cylinders (e.g., 5 cylinders), each cylinder having a center defined by randomly sampled points, a radius randomly sampled from [0.5m, 1.5m], and an infinite length. If the remaining number of points will be less than 10% of the original number of points, the deletion process is stopped. Finally, Gaussian distributed noise with σ = 0 and μ = 0.015 is added to each remaining point.

[0052] Figure 6 A method 600 for estimating the boundaries of an environment (e.g., a room environment 100) at least partially surrounded by a set of adjacent walls (e.g., wall 108) according to some embodiments of the present invention is shown. One or more steps of method 600 may be omitted during the execution of method 600, and one or more steps of method 600 need not be executed in the order shown. One or more steps of method 600 may be performed or facilitated by one or more processors included in a system or device (e.g., a wearable device such as an AR / VR device).

[0053] At step 602, a set of images (e.g., image 210) and a set of poses (e.g., pose data 212) corresponding to the set of images are obtained. The set of images can be images of an environment and each image can include one or more of the set of walls. Each pose in the set of poses can be captured simultaneously with an image in the set of images such that the position and / or orientation of each image in the set of images can be determined. In some embodiments, the set of images and the set of poses can be captured by a set of sensors (e.g., sensors 202) of a system or device (e.g., system 200). In some embodiments, the set of images can be captured by an image sensor (e.g., image sensor 206) and the set of poses can be captured by a pose sensor (e.g., pose sensor 208). The image sensor can be a camera or some other image capture device. The pose sensor can be an inertial measurement unit (IMU), an accelerometer, a gyroscope, a tilt sensor, or any combination thereof. In some embodiments, the set of poses can be determined from the set of images themselves such that the pose sensor and the image sensor can be the same sensor.

[0054] At step 604, a depth map of the environment (e.g., depth map 218) can be generated based on the set of images and the set of poses. The depth map can be a cumulative depth map that combines multiple (or all) of the images in the set of images, or in some embodiments, a depth map can be generated for each image in the set of images. In some embodiments, the depth map can be generated by a depth map generator (e.g., depth map generator 214), which can be a machine learning model (e.g., a neural network) trained to output a depth map when provided with the set of images and the set of poses as input.

[0055] At step 606, a set of wall segmentation maps (e.g., wall segmentation map 220) can be generated based on the set of images. Each wall segmentation map in the set of wall segmentation maps (e.g., wall segmentation map 304) can indicate a target region (e.g., target region 306) including a set of adjacent walls of the corresponding image in the set of images. In some embodiments, the set of wall segmentation maps can be generated by a wall segmentation generator (e.g., wall segmentation generator 216), which can be a machine learning model (e.g., a neural network) trained to output a wall segmentation map when provided with an image as input.

[0056] At step 608, a point cloud (e.g., point cloud 224) is generated based on the depth map and the set of wall segmentation maps. The point cloud can include a plurality of points sampled along a portion of the depth map aligned with the target region. In some embodiments, the point cloud can be generated by a wall point cloud generator (e.g., wall point cloud generator 222), which can be a machine learning model (e.g., a neural network) trained to output a point cloud when provided with a depth map and a set of wall segmentation maps as input.

[0057] At step 610, a set of clusters (e.g., the clusters of cluster point cloud 408) is identified for the point cloud. Each cluster in the set of clusters can include a subset of the plurality of points. Each cluster in the clusters is intended to correspond to a different wall in a set of adjacent walls. Clusters determined to correspond to the same wall in the set of adjacent walls can be combined into a single cluster. In some embodiments, the set of clusters can be identified by a cluster identifier (e.g., cluster identifier 226), which can be a machine learning model (e.g., a neural network) trained to output a set of clusters and / or a cluster point cloud (e.g., cluster point cloud 228) when provided with a point cloud as input.

[0058] At step 612, the boundaries of the environment along the set of adjacent walls are estimated based on the point cloud (e.g., estimated boundaries 106 and 232). In some embodiments, step 612 includes one or both of steps 614 and 616.

[0059] At step 614, if the plurality of points includes 2D points, for each cluster in the set of clusters, a line is fitted to the plurality of points, resulting in multiple lines. If the plurality of points includes 3D points, for each cluster in the set of clusters, a plane is fitted to the plurality of points, resulting in multiple planes. The line fitting or plane fitting can be accomplished using, for example, a curve fitting method.

[0060] At step 616, if the plurality of points includes 2D points, a closed loop is formed by extending the multiple lines until they reach an intersection point. If the plurality of points includes 3D points, a closed loop is formed by extending the multiple planes until they reach an intersecting line.

[0061] Figure 7 An example system or device is shown that can implement one or more of the described techniques. Specifically, Figure 7FIG. shows a schematic diagram of a wearable system 700 according to some embodiments of the present invention. The wearable system 700 may include a wearable device 701 and at least one remote device 703 (e.g., separate hardware but communicatively coupled) that is remote from the wearable device 701. Although the wearable device 701 is worn by a user (typically as a head-mounted device), the remote device 703 may be held by the user (e.g., as a handheld controller) or mounted in various configurations, such as fixedly attached to a frame, fixedly attached to a helmet or hat worn by the user, embedded in headphones, or otherwise detachably attached to the user (e.g., in a backpack configuration, in a belt-coupled configuration, etc.).

[0062] The wearable device 701 may include a left eyepiece 702A and a left lens assembly 705A arranged in a side-by-side configuration and constituting a left optical stack. The left lens assembly 705A may include an accommodation lens on the user side of the left optical stack and a compensation lens on the world side of the left optical stack. Similarly, the wearable device 701 may include a right eyepiece 702B and a right lens assembly 705B arranged in a side-by-side configuration and constituting a right optical stack. The right lens assembly 705B may include an accommodation lens on the user side of the right optical stack and a compensation lens on the world side of the right optical stack.

[0063] In some embodiments, the wearable device 701 includes one or more sensors, including but not limited to: a left front-facing world camera 706A directly attached to or near the left eyepiece 702A, a right front-facing world camera 706B directly attached to or near the right eyepiece 702B, a left-side world camera 706C directly attached to or near the left eyepiece 702A, a right-side world camera 706D directly attached to or near the right eyepiece 702B, a left-eye tracking camera 726A directly facing the left eye, a right-eye tracking camera 726B directly facing the right eye, and a depth sensor 728 attached between the eyepieces 702. The wearable device 701 may include one or more image projection devices, such as a left projector 714A optically linked to the left eyepiece 702A and a right projector 714B optically linked to the right eyepiece 702B.

[0064] The wearable system 700 may include a processing module 750 for collecting, processing, and / or controlling data within the system. The components of the processing module 750 may be distributed between the wearable device 701 and the remote device 703. For example, the processing module 750 may include a local processing module 752 on the wearable portion of the wearable system 700 and a remote processing module 756 that is physically separated from and communicatively linked to the local processing module 752. Each of the local processing module 752 and the remote processing module 756 may include one or more processing units (e.g., a central processing unit (CPU), a graphics processing unit (GPU), etc.) and one or more storage devices, such as non-volatile memory (e.g., flash memory).

[0065] The processing module 750 may collect data captured by various sensors of the wearable system 700, such as a camera 706, an eye-tracking camera 726, a depth sensor 728, a remote sensor 730, an ambient light sensor, a microphone, an IMU, an accelerometer, a compass, a global navigation satellite system (GNSS) unit, a radio device, and / or a gyroscope. For example, the processing module 750 may receive an image 720 from the camera 706. Specifically, the processing module 750 may receive a left-front image 720A from the world camera 706A facing left-front, a right-front image 720B from the world camera 706B facing right-front, a left-side image 720C from the world camera 706C facing left-side, and a right-side image 720D from the world camera 706D facing right-side. In some embodiments, the image 720 may include a single image, a pair of images, a video including an image stream, a video including a pair of image streams, etc. The image 720 may be periodically generated and sent to the processing module 750 while the wearable system 700 is powered on, or may be generated in response to an instruction sent by the processing module 750 to one or more cameras.

[0066] The camera 706 may be configured at various positions and orientations along the outer surface of the wearable device 701 to capture images around the user. In some cases, the cameras 706A, 706B may be positioned to capture images that substantially overlap with the FOVs of the user's left and right eyes, respectively. Thus, the placement of the camera 706 may be close to the user's eyes but not so close as to block the user's FOV. Alternatively or additionally, the cameras 706A, 706B may be positioned to align with the incoupling positions of the virtual image lights 722A, 722B, respectively. The cameras 706C, 706D may be positioned to capture images of the sides of the user, e.g., within or outside the user's peripheral vision. The images 720C, 720D captured using the cameras 706C, 706D do not have to overlap with the images 720A, 720B captured using the cameras 706A, 706B.

[0067] In some embodiments, the processing module 750 may receive ambient light information from an ambient light sensor. The ambient light information may indicate a range of luminance values or spatially resolved luminance values. The depth sensor 728 may capture a depth image 732 in the forward-facing direction of the wearable device 701. Each value of the depth image 732 may correspond to the distance between the depth sensor 728 and the nearest detected object in a particular direction. As another example, the processing module 750 may receive eye tracking data 734 from the eye tracking camera 726, which may include images of the left and right eyes. As another example, the processing module 750 may receive projection image luminance values from one or both of the projectors 714. The remote sensor 730 located within the remote device 703 may include any one of the above sensors having similar functions.

[0068] The projector 714 and the eyepiece 702, along with other components in the optical stack, are used to deliver virtual content to the user of the wearable system 700. For example, the eyepieces 702A, 702B may include transparent or translucent waveguides configured to respectively guide and outcouple the light generated by the projectors 714A, 714B. Specifically, the processing module 750 may cause the left projector 714A to output left virtual image light 722A onto the left eyepiece 702A, and may cause the right projector 714B to output right virtual image light 722B onto the right eyepiece 702B. In some embodiments, the projector 714 may include a microelectromechanical systems (MEMS) spatial light modulator (SLM) scanning device. In some embodiments, each of the eyepieces 702A, 702B may include multiple waveguides corresponding to different colors. In some embodiments, the lens assemblies 705A, 705B may be coupled to and / or integrated with the eyepieces 702A, 702B. For example, the lens assemblies 705A, 705B may be incorporated into a multi-layer eyepiece and may form one or more layers that make up one of the eyepieces 702A, 702B.

[0069] Figure 8 A simplified computer system 800 is shown in accordance with embodiments described herein. As Figure 8 shown, the computer system 800 may be incorporated into the devices described herein. Figure 8 A schematic diagram of one embodiment of the computer system 800 is provided, which may execute some or all of the steps of the methods provided by the various embodiments. It should be noted that Figure 8 only a general description of the various components is intended, and any or all of them may be used depending on the circumstances. Thus, Figure 8 it is shown broadly how the individual system elements may be implemented in a relatively separated or relatively more integrated manner.

[0070] The computer system 800 is shown as including hardware components that can be electrically coupled via a bus 805 or can communicate in other ways as appropriate. The hardware components can include: one or more processors 810, including but not limited to one or more general-purpose processors and / or one or more special-purpose processors, such as digital signal processing chips, graphics acceleration processors, etc.; one or more input devices 815, which can include but are not limited to a mouse, keyboard, camera, etc.; and one or more output devices 820, which can include but are not limited to a display device, printer, etc.

[0071] The computer system 800 can further include one or more non-transitory storage devices 825 and / or communicate with one or more non-transitory storage devices 825, which can include but are not limited to local and / or network-accessible storage devices, and / or can include but are not limited to disk drives, drive arrays, optical storage devices, solid-state storage devices, such as random access memory (“RAM”) and / or read-only memory (“ROM”) that can be programmable, flash-updatable, etc. Such storage devices can be configured to implement any suitable data storage, including but not limited to various file systems, database structures, etc.

[0072] The computer system 800 may also include a communication subsystem 819, which can include but is not limited to a modem, network card (wireless or wired), infrared communication device, wireless communication device, and / or chipset, such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication facilities, etc., and / or similar devices. The communication subsystem 819 can include one or more input and / or output communication interfaces to allow data exchange with a network (such as, for example, the network described below, as an example), other computer systems, a television, and / or any other device described herein. Depending on the required functionality and / or other implementation issues, a portable electronic device or similar device can transmit images and / or other information via the communication subsystem 819. In other embodiments, a portable electronic device (e.g., a first electronic device) can be incorporated into the computer system 800, such as an electronic device of the input device 815. In some embodiments, as described above, the computer system 800 will further include a working memory 835, which can include a RAM or ROM device.

[0073] The computer system 800 may also include software elements shown as currently residing within the working memory 835, including an operating system 840, device drivers, executable libraries, and / or other code, such as one or more application programs 845, which may include computer programs provided by various embodiments, and / or may be designed to implement methods and / or configure systems provided by other embodiments, as described herein. By way of example only, one or more of the processes described with respect to the methods above may be implemented as code and / or instructions executable by a computer and / or a processor within a computer; in one aspect, then, such code and / or instructions may be used to configure and / or adapt a general purpose computer or other device to perform one or more operations in accordance with the described methods.

[0074] The set of instructions and / or code may be stored on a non-transitory computer-readable storage medium, such as the storage device 825 described above. In some cases, the storage medium may be incorporated within the computer system, such as the computer system 800. In other embodiments, the storage medium may be separate from the computer system (e.g., a removable medium, such as an optical disk), and / or provided in an installation package such that the storage medium can be used to program, configure, and / or adapt a general purpose computer with the instructions / code stored thereon. The instructions may take the form of executable code executable by the computer system 800, and / or may take the form of source code and / or installable code, which, when compiled and / or installed on the computer system 800, for example using any of a variety of generally available compilers, installers, compression / decompression utilities, etc., then take the form of executable code.

[0075] It will be apparent to those skilled in the art that substantial variations may be made in accordance with specific requirements. For example, customized hardware may also be used, and / or particular elements may be implemented in hardware, software including portable software such as applets, or both. Additionally, connections to other computing devices such as network input / output devices may be employed.

[0076] As described above, on the one hand, some embodiments may employ a computer system such as computer system 800 to execute methods according to various embodiments of the present technology. According to a set of embodiments, in response to a processor 810 executing one or more sequences of one or more instructions, some or all of the processes of such methods are executed by computer system 800, and such instructions may be incorporated into an operating system 840 and / or other code contained in working memory 835, such as application 845. Such instructions may be read into working memory 835 from another computer-readable medium, such as one or more storage devices 825. Merely by way of example, the execution of the sequence of instructions contained in working memory 835 may cause processor 810 to execute one or more processes of the methods described herein. Additionally or alternatively, portions of the methods described herein may be executed by dedicated hardware.

[0077] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any medium that participates in providing data that causes a machine to operate in a particular fashion. In embodiments implemented using computer system 800, various computer-readable media may be involved in providing instructions / code to one or more processors 810 for execution and / or may be used to store and / or carry such instructions / code. In many implementations, the computer-readable medium is a physical and / or tangible storage medium. Such media may take the form of non-volatile media or volatile media. Non-volatile media includes, for example, optical disks and / or magnetic disks, such as one or more storage devices 825. Volatile media includes, but is not limited to, dynamic memory, such as working memory 835.

[0078] Common forms of physical and / or tangible computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, or any other magnetic medium, CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.

[0079] In carrying one or more sequences of one or more instructions to one or more processors 810 for execution, various forms of computer-readable media may be involved. Merely by way of example, the instructions may initially be carried on a magnetic disk and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions as a signal through a transmission medium to be received and / or executed by computer system 800.

[0080] The communication subsystem 819 and / or its components will generally receive signals, and then the bus 805 can carry the signals and / or the data, instructions, etc. carried by the signals to the working memory 835, from which the processor 810 fetches and executes the instructions. The instructions received by the working memory 835 can optionally be stored on the non-transitory storage device 825 before or after being executed by the processor 810.

[0081] The methods, systems, and devices discussed above are examples. Various configurations may omit, substitute, or add various procedures or components as appropriate. For example, in alternative configurations, the methods may be performed in a different order than described, and / or various stages may be added, omitted, and / or combined. Additionally, features described with respect to certain configurations may be combined in various other configurations. Different aspects and elements of the configurations may be combined in a similar manner. Moreover, technology is constantly evolving, and thus, many of the elements are examples and do not limit the scope of the present disclosure or the claims.

[0082] Specific details are given in the description to provide a thorough understanding of the exemplary configurations including the implementation. However, the configurations may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail to avoid obscuring the configurations. This description only provides example configurations and does not limit the scope, applicability, or configurations of the claims. Instead, the foregoing description of the configurations will provide those skilled in the art with an enabling description for implementing the described techniques. Various changes may be made to the function and arrangement of the elements without departing from the spirit or scope of the present disclosure.

[0083] In addition, a configuration may be described as a process depicted as a schematic flowchart or block diagram. Although each may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. Additionally, the order of the operations may be rearranged. The process may have additional steps not included in the figures. Moreover, examples of the methods may be implemented by hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the necessary tasks may be stored in a non-transitory computer-readable medium (such as a storage medium). The processor may execute the described tasks.

[0084] Several example configurations have been described, and various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the present disclosure. For example, the elements described above may be components of a larger system, where other rules may take precedence over or otherwise modify the application of the techniques. Additionally, many steps may be taken before, during, or after considering the elements described above. Accordingly, the foregoing description does not limit the scope of the claims.

[0085] As used herein and in the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an" and "the" include plural references. Thus, for example, reference to "a user" includes a plurality of such users, reference to "a processor" includes reference to one or more processors and equivalents thereof known to those skilled in the art, and the like.

[0086] Further, as used in this specification and the following claims, the words "comprises", "comprising", "includes", "including", "has", "having", "with" and "contains" are intended to specify the presence of the stated feature, integer, component or step, but they do not preclude the presence or addition of one or more other features, integers, components, steps, acts or groups.

[0087] It should also be understood that the examples and embodiments described herein are for illustrative purposes only and will suggest various modifications or alterations to those skilled in the art, and will be included within the spirit and scope of this application and the scope of the appended claims.

Claims

1. A method, comprising: obtaining a set of images and a set of poses corresponding to the set of images, the set of images having been captured from an environment that is at least partially surrounded by a set of adjacent walls; generating a depth map of the environment based on the set of images and the set of poses; generating a set of wall segmentation maps by providing the set of images to a machine learning model that has been trained to output wall segmentation maps based on an image input, each wall segmentation map in the set of wall segmentation maps indicating one or more regions of a corresponding image in the set of images that include the set of adjacent walls; generating a point cloud including a plurality of points by sampling the depth map along portions thereof aligned with the one or more regions; identifying a set of clusters of the point cloud, each cluster in the set of clusters including a subset of the plurality of points, and wherein each cluster in the set of clusters is determined to correspond to a different wall in the set of adjacent walls; and for each cluster in the set of clusters, estimating a boundary of the environment along the set of adjacent walls by fitting a line or a plane to the plurality of points.

2. The method according to claim 1, further comprising: capturing the set of images and the set of poses using one or more sensors, wherein obtaining the set of images and the set of poses includes: receiving the set of images and the set of poses from the one or more sensors.

3. The method according to claim 2, wherein, the one or more sensors include an image sensor configured to capture the set of images and a pose sensor configured to capture the set of poses.

4. The method according to claim 1, wherein, the plurality of points include 2D points, and wherein estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting the line to the plurality of points to obtain a plurality of lines; and forming a closed loop by extending the plurality of lines until an intersection point is reached.

5. The method according to claim 1, wherein, the plurality of points include 3D points, and wherein estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: for each cluster in the set of clusters, fitting the plane to the plurality of points to obtain a plurality of planes; and forming a closed loop by extending the plurality of planes until an intersecting line is reached.

6. The method according to claim 1, wherein, the set of images includes RGB images.

7. The method according to claim 1, wherein, the set of poses includes the camera orientation of an image sensor that captured the set of images.

8. The method according to claim 1, wherein, the plurality of points include 3D points.

9. A system, comprising: one or more processors; and a computer-readable medium including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including the following: Obtain a set of images and a set of poses corresponding to the set of images, where the set of images has been captured from an environment that is at least partially surrounded by a set of adjacent walls; Generate a depth map of the environment based on the set of images and the set of poses; Generate a set of wall segmentation maps by providing the set of images to a machine learning model that has been trained to output wall segmentation maps based on image inputs, where each wall segmentation map in the set of wall segmentation maps indicates one or more regions of the corresponding image in the set of images that include the set of adjacent walls; Generate a point cloud including a plurality of points by sampling the depth map along portions thereof aligned with the one or more regions; Identify a set of clusters of the point cloud, where each cluster in the set of clusters includes a subset of the plurality of points, and where each cluster in the set of clusters is determined to correspond to a different wall in the set of adjacent walls; And For each cluster in the set of clusters, estimate the boundary of the environment along the set of adjacent walls by fitting a line or a plane to the plurality of points.

10. The system according to claim 9, Wherein, The operations further include: Capturing the set of images and the set of poses using one or more sensors, where obtaining the set of images and the set of poses includes: receiving the set of images and the set of poses from the one or more sensors.

11. The system according to claim 10, Wherein, The one or more sensors include an image sensor configured to capture the set of images and a pose sensor configured to capture the set of poses.

12. The system according to claim 9, Wherein, The plurality of points include 2D points, and where estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: For each cluster in the set of clusters, fitting the line to the plurality of points to obtain a plurality of lines; and Forming a closed loop by extending the plurality of lines until an intersection point is reached.

13. The system according to claim 9, Wherein, The plurality of points include 3D points, and where estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: For each cluster in the set of clusters, fitting the plane to the plurality of points to obtain a plurality of planes; and Forming a closed loop by extending the plurality of planes until an intersecting line is reached.

14. The system according to claim 9, Wherein, The set of images includes RGB images.

15. A non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations including the following: Obtain a set of images and a set of poses corresponding to the set of images, where the set of images has been captured from an environment that is at least partially surrounded by a set of adjacent walls; Generate a depth map of the environment based on the set of images and the set of poses; Generating a set of wall segmentation maps by providing the set of images to a machine learning model that has been trained to output wall segmentation maps based on image inputs, wherein each wall segmentation map in the set of wall segmentation maps indicates one or more regions of the corresponding image in the set of images that include the set of adjacent walls; Generating a point cloud including a plurality of points by sampling the depth map along portions of the depth map that are aligned with the one or more regions; Identifying a set of clusters in the point cloud, wherein each cluster in the set of clusters includes a subset of the plurality of points, and wherein each cluster in the set of clusters is determined to correspond to a different wall in the set of adjacent walls; And For each cluster in the set of clusters, estimating a boundary of the environment along the set of adjacent walls by fitting a line or a plane to the plurality of points.

16. The non-transitory computer-readable medium according to claim 15, Wherein, The operations further include: Capturing the set of images and the set of poses using one or more sensors, wherein obtaining the set of images and the set of poses includes: receiving the set of images and the set of poses from the one or more sensors.

17. The non-transitory computer-readable medium according to claim 16, Wherein, The one or more sensors include an image sensor configured to capture the set of images and a pose sensor configured to capture the set of poses.

18. The non-transitory computer-readable medium according to claim 15, Wherein, The plurality of points include 2D points, and wherein estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: For each cluster in the set of clusters, fitting the line to the plurality of points to obtain a plurality of lines; and Forming a closed loop by extending the plurality of lines until an intersection point is reached.

19. The non-transitory computer-readable medium according to claim 15, Wherein, The plurality of points include 3D points, and wherein estimating the boundary of the environment along the set of adjacent walls based on the point cloud includes: For each cluster in the set of clusters, fitting the plane to the plurality of points to obtain a plurality of planes; and Forming a closed loop by extending the plurality of planes until an intersecting line is reached.

20. The non-transitory computer-readable medium according to claim 15, Wherein, The set of images includes RGB images.

Citation Information

Patent Citations

  • Estimating dimensions for an enclosed space using a multi-directional camera

    CN109564690A

  • Methods and systems for creating virtual and augmented reality

    WO2015192117A1