Method and system for detecting and combining structural features in 3D reconstruction

By introducing shape-aware technology into 3D reconstruction technology, including shape detection and shape-aware volume fusion, the problem of inaccurate shape-aware and camera posture alignment in traditional technologies is solved, and a more accurate and visually comfortable 3D mesh reconstruction is achieved.

CN114863059BActive Publication Date: 2025-05-30MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210505506.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-09-25
Filing Date
2016-09-23
Publication Date
2025-05-30
Estimated Expiration
2036-09-23

AI Technical Summary

Technical Problem

Traditional 3D reconstruction technology has inaccuracies in shape perception and camera pose alignment, resulting in noise and artifacts in the generated 3D mesh, affecting the visual effect and user experience.

Method used

Using shape-aware technology, including shape detection, shape-aware pose estimation and shape-aware volume fusion algorithm, aligning to form a more accurate and robust 3D mesh by detecting shapes in the scene and updating camera poses.

Benefits of technology

The generated 3D mesh is achieved with clearer and more realistic shapes and edges, improving user experience, and improving the alignment accuracy of depth maps by detecting shapes, reducing noise and artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863059B_ABST
    Figure CN114863059B_ABST
Patent Text Reader

Abstract

A method for forming a reconstructed 3D mesh includes: receiving a set of captured depth maps associated with a scene, performing an initial camera pose alignment associated with the set of captured depth maps, and overlaying the set of captured depth maps in a reference frame. The method further includes detecting one or more shapes in the overlaid set of captured depth maps and updating the initial camera pose alignment to provide a shape-aware camera pose alignment. The method further includes performing shape-aware volume fusion and forming a reconstructed 3D mesh associated with the scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201680054286.9, "Method and System for Detecting and Combining Structural Features in 3D Reconstruction" (filed on September 23, 2016).

[0002] Cross - reference to related applications

[0003] This application claims the priority of U.S. Provisional Patent Application No. 62,232,833, entitled "Method and System for Detecting and Combining Structural Features in 3D Reconstruction", filed on September 25, 2015, the disclosure of which is hereby incorporated by reference in its entirety for all purposes. Summary of the Invention

[0004] The present invention generally relates to the field of computerized three - dimensional (3D) image reconstruction, and more particularly, to methods and systems for detecting and combining structural features in 3D reconstruction.

[0005] As described herein, embodiments of the present invention are directed to solving problems not fully addressed by conventional techniques, as well as providing additional features that will become apparent by reference to the following detailed description in conjunction with the accompanying drawings.

[0006] Some embodiments disclosed herein relate to methods and systems for providing shape - aware 3D reconstruction. Some implementations incorporate improved shape - aware techniques, such as shape detection, shape - aware pose estimation, shape - aware volume fusion algorithms, etc.

[0007] According to an embodiment of the present invention, a method for forming a reconstructed 3D mesh is provided. The method includes receiving a set of captured depth maps associated with a scene, performing an initial camera pose alignment associated with the set of captured depth maps, and overlaying the set of captured depth maps in a reference frame. The method further includes detecting one or more shapes in the overlaid set of captured depth maps and updating the initial camera pose alignment to provide a shape - aware camera pose alignment. The method also includes performing shape - aware volume fusion and forming a reconstructed 3D mesh associated with the scene.

[0008] According to another embodiment of the present invention, a method for detecting shapes present in a scene is provided. The method includes determining a vertical direction associated with a point cloud including a plurality of captured depth maps and forming a virtual plane orthogonal to the vertical direction. The method further includes projecting points of the point cloud onto the virtual plane and calculating projection statistics of the points of the point cloud. The method further includes detecting one or more lines based on the calculated projection statistics, the one or more lines being associated with vertical walls, and detecting shapes present in the scene based on the projection statistics and the one or more detected lines.

[0009] According to a particular embodiment of the present invention, a method for performing shape-aware camera pose alignment is provided. The method includes receiving a set of captured depth maps. Each depth map in the set of captured depth maps is associated with a physical camera pose. The method further includes receiving one or more detected shapes. Each shape in the one or more detected shapes is characterized by dimensions and a position / orientation. The method further includes creating a 3D mesh for each shape in the one or more detected shapes and creating one or more virtual cameras associated with each 3D mesh in a local reference frame. Additionally, the method includes rendering one or more depth maps. Each of the one or more rendered depth maps is associated with each virtual camera associated with each 3D mesh. Furthermore, the method includes jointly solving for the physical camera pose and the position / orientation of each shape in the one or more detected shapes by optimizing the alignment between the one or more rendered depth maps and the set of captured depth maps.

[0010] In an embodiment, the shape-aware 3D reconstruction method includes one or more of the following steps: performing pose estimation on a set of captured depth maps; performing shape detection for an aligned pose after pose estimation; performing shape-aware pose estimation based on the detected shape; and performing shape-aware volume fusion based on the aligned pose and shape to generate one or more 3D meshes.

[0011] Compared with the prior art, many benefits are achieved by the present invention. For example, embodiments of the present invention provide clear and sharp shapes and edges in the 3D mesh, and thus look more realistic than 3D meshes generated without using shape-aware 3D reconstruction. Accordingly, the 3D meshes provided by embodiments of the present invention are more comfortable for viewers. Another advantage is that, due to the presence of detected shapes during the 3D reconstruction process, more accurate and robust alignment of the captured depth maps can be achieved. In addition, an end-to-end 3D reconstruction framework is provided that applies prior knowledge of artificial scenes and at the same time maintains flexibility in terms of scene heterogeneity. These and other embodiments of the present invention, as well as many of its advantages and features, are described in more detail in conjunction with the following text and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present disclosure is described in detail with reference to the following drawings. The drawings are provided for illustrative purposes only and depict exemplary embodiments of the present disclosure. These drawings are provided to facilitate the reader's understanding of the present disclosure and should not be considered as limiting the breadth, scope, or applicability of the present disclosure. It should be noted that these drawings are not necessarily drawn to scale for clarity and ease of illustration.

[0013] Figure 1 is a simplified flowchart showing a method for creating a 3D mesh of a scene using multiple frames of captured depth maps.

[0014] Figure 2 It is a simplified flowchart showing a method for generating a 3D mesh of a scene using multiple frames of a captured depth map according to an embodiment of the present invention.

[0015] Figure 3 It is a simplified flowchart showing a method for detecting shapes present in a point cloud according to an embodiment of the present invention.

[0016] Figure 4 It is a simplified flowchart showing a method for performing shape-aware camera pose alignment according to an embodiment of the present invention.

[0017] Figure 5 It is a simplified flowchart showing a method for performing shape-aware volume fusion according to an embodiment of the present invention.

[0018] Figure 6A It is a sketch of a 3D mesh of a wall end according to an embodiment of the present invention.

[0019] Figure 6B It is a sketch of a 3D mesh of a door frame according to an embodiment of the present invention.

[0020] Figure 7A It is a simplified schematic diagram showing a rendered depth map associated with an internal view of a door frame and an associated virtual camera according to an embodiment of the present invention.

[0021] Figure 7B It is a simplified schematic diagram showing a rendered depth map associated with an external view of a door frame and an associated virtual camera according to an embodiment of the present invention.

[0022] Figure 7C It is a simplified schematic diagram showing a rendered depth map associated with two wall corners and an associated virtual camera according to an embodiment of the present invention.

[0023] Figure 8A It is a simplified schematic diagram showing a rendered depth map of an internal view of a door frame rendered from a virtual camera according to an embodiment of the present invention.

[0024] Figure 8B It is a simplified schematic diagram showing a rendered depth map of an external view of a door frame rendered from a virtual camera according to an embodiment of the present invention.

[0025] Figure 9A It is a simplified point cloud diagram showing a captured depth map.

[0026] Figure 9B It is a simplified point cloud diagram showing a captured depth map and a rendered depth map using the shape-aware method provided by an embodiment of the present invention.

[0027] Figure 10A It is shown using with respect to Figure 1An image of a first reconstructed 3D mesh reconstructed by the described method.

[0028] Figure 10B is an image showing a second reconstructed 3D mesh reconstructed using the method described with respect to Figure 2 An image of a second reconstructed 3D mesh reconstructed by the described method.

[0029] Figure 11A is an image showing a third reconstructed 3D mesh reconstructed using the method described with respect to Figure 1 An image of a third reconstructed 3D mesh reconstructed by the described method.

[0030] Figure 11B is an image showing a fourth reconstructed 3D mesh reconstructed using the method described with respect to Figure 2 An image of a fourth reconstructed 3D mesh reconstructed by the described method.

[0031] Figure 12 is a simplified schematic diagram showing a system for reconstructing a 3D mesh using captured depth maps according to an embodiment of the present invention.

[0032] Figure 13 is a block diagram of a computer system or information processing device that may include embodiments, be incorporated into embodiments, or be used to practice any innovations, embodiments, and / or examples found within the present disclosure. Detailed Description

[0033] Embodiments of the present invention relate to methods and systems for computerized three-dimensional (3D) scene reconstruction, and more particularly, to methods and systems for detecting and combining structural features in 3D reconstruction.

[0034] The following description is presented to enable a person of ordinary skill in the art to make and use the present invention. The description of specific devices, techniques, and applications is provided only as an example. Various modifications to the examples described herein will be apparent to a person of ordinary skill in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the present invention. Thus, embodiments of the present invention are not intended to be limited to the examples described and shown herein, but rather to the scope consistent with the claims.

[0035] As used herein, the word "exemplary" means "serving as an example or illustration". Any aspect or design described herein as "exemplary" is not necessarily to be construed as preferred or superior to other aspects or designs.

[0036] Aspects of the present technology will now be described in detail, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements.

[0037] It should be understood that the specific order or hierarchy of steps in the processes disclosed herein are examples of exemplary methods. Based on design preferences, it is understood that the specific order or hierarchy of steps in a process may be rearranged while remaining within the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not meant to be limited to the specific order or hierarchy presented.

[0038] The embodiments disclosed herein relate to methods and systems for providing shape-aware 3D reconstruction. As described herein, some embodiments of the present invention incorporate improved shape-aware techniques, such as shape detection, shape-aware pose estimation, shape-aware volume fusion algorithms, etc. According to an embodiment of the present invention, a shape-aware 3D reconstruction method may include one or more of the following steps: performing pose estimation on a set of depth images; performing shape detection in an aligned pose after pose estimation; performing shape-aware pose estimation based on the detected shape; and performing shape-aware volume fusion based on the aligned pose and shape to generate a 3D mesh.

[0039] 3D reconstruction is one of the most popular topics in 3D computer vision. It takes images (e.g., color / grayscale images, depth images, etc.) as input and generates a 3D mesh (e.g., automatically) representing the observed scene. 3D reconstruction has many applications in virtual reality, drawing, robotics, gaming, movie production, etc.

[0040] As an example, a 3D reconstruction algorithm may receive an input image (e.g., a color / grayscale image, a color / grayscale image + depth image, or just depth), and process the input image as appropriate to form a captured depth map. For example, a multi-view stereo algorithm from a color image may be used to generate a passive depth map, and an active sensing technique (such as a structured light depth sensor) may be used to obtain an active depth map. Although the foregoing examples are shown, embodiments of the present invention may be configured to process any type of depth map. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0041] Figure 1 is a simplified flowchart showing a method for creating a 3D mesh of a scene using multiple frames of a captured depth map. Refer to Figure 1, which shows a method of creating a 3D model (e.g., a 3D triangular mesh representing the 3D surface associated with the scene) of a scene based on multiple frames of captured depth maps. Method 100 includes receiving a set of captured depth maps (110). A captured depth map is a depth image where each pixel has an associated depth value that represents the depth from the pixel to the camera that acquired the depth image. Compared to a color image having three or more channels per pixel (e.g., an RGB image having red, green, and blue components), a depth map can have a single channel per pixel (i.e., the pixel distance from the camera). The process of receiving the set of captured depth maps can include processing an input image (e.g., an RGB image) to generate one or more captured depth maps, also referred to as frames of captured depth maps. In other embodiments, the captured depth maps are obtained using a time-of-flight camera, lidar (LIDAR), a stereo camera, etc., and are thus received by the system.

[0042] The set of captured depth maps includes depth maps from different camera angles and / or positions. As an example, a depth map stream can be provided by a moving depth camera. As the moving depth camera pans and / or moves, depth maps are generated as a stream of depth images. As another example, a static depth camera can be used to collect multiple depth maps of part or all of a scene from different angles and / or different positions or a combination thereof.

[0043] The method also includes aligning the camera poses associated with a set of captured depth maps in a reference frame (112) and overlaying the set of captured depth maps in the reference frame (112). In an embodiment, a pose estimation process is utilized to align depth points from all cameras and create a locally and globally consistent point cloud in 3D world coordinates. Depth points from the same location in world coordinates should be aligned as close to each other as possible. However, due to inaccuracies in the depth maps, the pose estimation is usually not perfect, especially for structural features such as corners of walls, ends of walls, door frames in indoor scenes, etc., which cause artifacts on these structural features when they are present in the generated mesh. Additionally, these inaccuracies can be exacerbated when the mesh boundaries are considered as occluders (i.e., objects that block background objects), as the artifacts can be more apparent to the user.

[0044] To align the camera pose indicating the position and orientation of the camera associated with each depth image, the depth maps are overlaid and the differences in the positions of adjacent and / or overlaid pixels are reduced or minimized. Once the positions of the pixels in the reference frame have been adjusted, the camera pose is adjusted and / or updated to align the camera pose with the adjusted pixel positions. Thus, the camera pose is aligned (114) in the reference frame. In other words, a rendered depth map can be created by projecting the depth points of all depth maps into a reference frame (e.g., a 3D world coordinate system) based on the estimated camera pose.

[0045] The method further includes performing volume fusion (116) to form a reconstructed 3D mesh (118). The volume fusion process can include fusing multiple captured depth maps into a volumetric representation as a discrete form of the signed distance function of the observed scene. 3D mesh generation can include extracting a polygonal mesh from the volumetric representation in 3D space using the marching cubes algorithm or other suitable methods.

[0046] To reduce the artifacts discussed above, embodiments of the present invention provide methods and systems for performing shape-aware 3D reconstruction that incorporate improved shape-aware techniques such as shape detection, shape-aware pose estimation, shape-aware volume fusion algorithms, etc.

[0047] For indoor structures, since they are man-made, these structures typically have regular shapes compared to organic outdoor structures. Additionally, inexpensive depth cameras can produce captured depth maps that contain a relatively high level of noise, which results in errors in the depth values associated with each pixel. These depth errors can lead to inaccuracies in the camera pose estimation process. These errors can propagate through the system, resulting in errors including noise and inaccuracies in the reconstructed 3D mesh. As an example, waves or bends in a corner (such as a wavy appearance of a wall that should be flat) are visually unsatisfactory to the user. Thus, using embodiments of the present invention, the reconstructed 3D mesh is characterized by increased accuracy, reduced noise, etc., resulting in a 3D mesh that is visually satisfactory to the user.

[0048] It should be understood that Figure 1 the specific steps shown provide a specific method for creating a 3D mesh of a scene using multiple frames of captured depth maps according to embodiments of the present invention. Other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, Figure 1 each of the steps shown may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, depending on the specific application, additional steps may be added or removed. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0049] Figure 2 is a simplified flowchart showing a method of generating a 3D mesh of a scene using multiple frames of a captured depth map according to an embodiment of the present invention. In Figure 2 The method shown can be considered a process of generating a reconstructed 3D mesh from a captured depth map by using a shape-aware 3D reconstruction method and system.

[0050] Referring Figure 2 , method 200 includes receiving a set of captured depth maps (210). As discussed with respect to Figure 1 , the set of captured depth maps can be received as depth maps, a processed form of depth maps, or generated from other images to provide a set of captured depth maps. The method also includes performing an initial camera pose estimation (212) and overlaying the set of captured depth maps in a reference frame (214). In the initial camera pose estimation, the depth maps are overlaid and the differences in the positions of adjacent and / or overlaid pixels are reduced or minimized. Once the positions of the pixels in the reference frame have been adjusted, the camera pose is adjusted and / or updated to align the camera pose with the adjusted pixel positions and provide an initial camera pose estimation.

[0051] During this initial refinement of the set of captured depth maps, it is possible that the initial estimate of the camera pose includes some inaccuracies. As a result, particularly in regions of structural features, the overlaid depth maps may exhibit some misalignment. Accordingly, embodiments of the present invention apply shape detection to the aligned camera pose to detect structural shapes that may have strong characteristics using the point distribution of the point cloud, as described more fully below. As Figure 2 shown, the method detects shapes (218) in an overlaid set of captured depth maps.

[0052] Figure 3 is a simplified flowchart showing a method of detecting shapes present in a point cloud according to an embodiment of the present invention. The point cloud can be formed by overlaying a set of captured depth maps in a reference frame. Additional description related to forming the point cloud based on captured depth maps, rendered depth maps, or a combination thereof is provided with respect to FIG. 9. As Figure 3 shown, the method is useful for detecting structures such as door frames, windows, wall corners, wall ends, walls, furniture, and other man-made structures present in the point cloud.

[0053] Although the camera pose can be determined, the relationship between the camera pose and the vertical reference frame may be unknown. In some embodiments, the z-axis of the reference frame can be aligned with the direction of gravity. Accordingly, method 300 includes determining a vertical direction associated with the point cloud using point normals (310). In particular for indoor scenes, the presence of walls and other structural features can be used to determine the vertical direction associated with the point cloud, also referred to as the vertical direction of the point cloud. For example, for a given pixel in the point cloud, the pixels in the vicinity of the given pixel are analyzed to determine the normal vector of the given pixel. This normal vector is referred to as the point normal. As an example, for a pixel representing a part of a wall, the adjacent pixels will typically lie in a plane. Accordingly, the normal vector of this plane can be used to define the normal vector of the pixel of interest.

[0054] Given the normal vectors of some or all of the pixels in the point cloud, the direction orthogonal to the normal vectors will define the vertical direction. In other words, the normal vectors will typically lie in parallel, horizontal planes, the vertical direction of which is orthogonal to these parallel, horizontal planes.

[0055] In some embodiments, determining the vertical direction includes estimating the vertical direction and then refining the estimated vertical direction, although these steps can be combined into a single process that provides the desired vertical direction vector. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0056] The method further includes forming a virtual plane orthogonal to the vertical direction (312) and projecting the points in the point cloud onto the virtual plane orthogonal to the vertical direction and computing their projection statistic (314). Given a vertical direction aligned with gravity, a plane orthogonal to the vertical direction can be defined, which will represent a horizontal surface, such as the floor of a room. In addition to the term virtual plane, this plane orthogonal to the vertical direction can be referred to as the projection plane. An example of the computed projection statistic is the point distribution that can be collected for each two-dimensional position on the virtual plane.

[0057] By projecting the points in the point cloud onto the virtual plane orthogonal to the vertical direction, all the points in the point cloud can be represented as a two-dimensional data set. This two-dimensional data set will represent the position of the points in the x-y space of the point, the height range of the points projected onto the x-y positions, and the density of the points associated with the x-y positions.

[0058] For a given position in the projection plane, which can be referred to as the x-y space, the density of the points projected onto the given position represents the number of points present in the point cloud at the height above the given position. As an example, consider a wall with a door in it. The density of the points at the position below the wall will be high and remains high until it reaches the door frame. The projection onto the projection plane will result in a line extending along the bottom of the wall. The point density at the position below the door frame is low (only the points associated with the top of the door frame and the wall above the door frame). Once on the other side of the door frame, the density will increase again.

[0059] After projecting the point cloud onto the projection plane, the density of the points in the projection plane will effectively provide a floor plan of the scene. Each pixel in the projection plane can have a gray value that indicates the number of points associated with the specific pixel onto which the projection occurs. Given the point distribution, the method also includes detecting lines from the projection statistics as vertical walls (316). The projection statistics can be considered as elements of the projected image.

[0060] Thus, embodiments of the present invention utilize one or more projection statistics, including a predetermined number of points projected onto a specific x / y position on a 2D virtual plane. Another projection statistic is the distribution of the point normals of the points projected onto a specific x / y position. Additionally, another projection statistic is the height range of the points projected onto a specific x / y position. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0061] Based on the projection statistics and one or more detected lines, the method includes detecting one or more shapes (e.g., wall corners, door frames, doors, etc.) (318). The one or more shapes can be different shapes (wall corners and door frames) or multiple examples of a shape (two wall corners in different parts of a room). The inventors have determined that most regular shapes are associated with walls. For example, a wall corner is the connection of two orthogonal walls, a wall end is the end of a wall, and a door frame is an opening in a wall. These structural features are identified and detected by analyzing the point distribution.

[0062] The method also includes determining the dimensions and positions of the one or more detected shapes (320). In addition to the density of the points projected onto each two-dimensional position, the point height distribution at each two-dimensional position above the available projection plane can be used to determine the vertical extent or extension of the detected shape. As an example, if a two-dimensional position has multiple points, all with heights greater than 7 feet, then that two-dimensional position may be located below a door frame that opens at the top of the door frame and is then solid above the door frame. A histogram can be created for each two-dimensional position, where the points projected onto the two-dimensional positions set along the histogram are functions of their heights above the projection plane.

[0063] In some embodiments, determining the dimensions and positions of one or more detected shapes is determining the initial dimensions and positions of each shape, which will be parameterized depending on the type of shape. For example, two-dimensional position, orientation, and vertical extent are determined for a corner. For a door frame, the thickness and width can be determined. For a door, the height and width can be determined.

[0064] It should be understood that Figure 3 the specific steps shown in provide a particular method for detecting shapes present in a point cloud according to embodiments of the present invention. Other sequences of steps may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, Figure 3 each of the steps shown in may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, depending on the particular application, additional steps may be added or removed. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0065] Referring again to Figure 2 , after a shape in the point cloud (i.e., a set of captured depth images that are covered) has been detected, the method includes performing shape-aware camera pose estimation, also referred to as shape-aware camera pose alignment (218). Thus, embodiments of the present invention perform a second camera pose alignment process that is informed by the presence of the shape detected in the point cloud, thereby providing a camera pose associated with each depth image in the set of depth images, which is optimized with the detected shape as a constraint. In addition to aligning the camera pose based on the overlap between the covered captured depth maps, embodiments also align the camera pose based on the overlap between the covered captured depth maps and the detected shape. By aligning the depth maps with the detected shape, as a result of using the detected shape as an additional constraint, the reconstructed 3D mesh has higher accuracy. By using the detected shape as a constraint, errors that can propagate through the system can be reduced or eliminated, thereby improving the accuracy of the 3D mesh.

[0066] Figure 4 is a simplified flowchart showing a method of forming shape-aware camera pose alignment according to an embodiment of the present invention. Regarding the method 400 discussed in may be a method of performing shape-aware camera pose alignment as discussed for the process 218 in. As discussed below, the detected shape is used for optimizing the camera pose estimation. Figure 4 The method 400 discussed with respect to Figure 2 may be a method of performing shape-aware camera pose alignment as discussed for the process 218 in. As discussed below, the detected shape is used for optimizing the camera pose estimation.

[0067] Method 400 includes receiving a set of captured depth maps (410). Each depth map in the captured depth maps is associated with a physical camera pose. The method also includes receiving one or more detected shapes (412). Each shape in the one or more detected shapes is characterized by dimensions and position / orientation. The method includes creating a 3D mesh (414) for each of the one or more detected shapes. Examples of the created shape meshes can be found in Figure 6A and Figure 6B . As shown in Figure 6A , a 3D mesh of a wall end is shown. In Figure 6B , a 3D mesh of a door frame is shown. These shapes can be detected using the methods discussed in Figure 3 . As shown in Figure 6B , the door frame mesh consists of multiple adjacent triangular regions. Although the door frame can have different heights, widths, opening widths, etc., the angles between the sides and the top of the door frame and other features will generally be regular and predictable. The 3D mesh associated with the door frame or other structural features will be separate from the mesh obtained by the process 118 in Figure 1 . As described herein, shape-aware volume fusion utilizes the meshes associated with the structural features in forming the shape-aware reconstructed 3D mesh.

[0068] The method also includes creating one or more virtual cameras (416) for each 3D mesh in a local reference frame. One or more virtual cameras are created in the local reference frame with reference to the detected shapes. For a given detected shape, the virtual cameras will be positioned in the reference frame of the detected shape. If the position and / or orientation of the detected shape is adjusted, the virtual cameras will be adjusted to maintain a constant position in the reference frame. If the dimensions of the detected shape change, such as a decrease in the door frame thickness, the virtual cameras on the opposite sides of the door frame will be closer to each other in combination with the decrease in the door frame thickness. Thus, each triangle in the 3D mesh of the shape can be viewed by at least one virtual camera. For example, for a corner of a wall, one virtual camera is sufficient to cover all the triangles, while for a wall end or a door frame, usually at least two virtual cameras are required to cover all the triangles. It should be understood that these virtual cameras are special because they have the detected shapes associated with the virtual cameras.

[0069] Referring to Figure 6B , a 3D mesh associated with a door frame is shown. After detecting the door frame as discussed in Figure 2 , a 3D mesh as shown in Figure 6B is created. To create virtual cameras for the 3D mesh, a rendered depth map associated with the door frame is formed as shown in Figure 7A . Based on the rendered depth map, virtual camera 710 can be created at a predetermined position and orientation.

[0070] Figure 7A is a simplified schematic diagram showing a rendered depth map associated with an internal view of a door frame and an associated virtual camera according to an embodiment of the present invention. Figure 7B is a simplified schematic diagram showing a rendered depth map associated with an external view of a door frame and an associated virtual camera according to an embodiment of the present invention. The rendered depth map is a subset of a point cloud. The point cloud is formed by combining depth maps (i.e., frames of depth maps). The point cloud can be formed by combining captured depth maps, rendered depth maps, or a combination of captured and rendered depth maps. Refer to Figure 7A and Figure 7B , the rendered depth map includes a set of depth points associated with a structure (i.e., the door frame).

[0071] Viewed from the inside of the door frame, the rendered depth map 705 can be considered to represent the distance from the pixels constituting the door frame to the virtual camera 710 for the part of the depth map including the door frame. Viewed from the outside of the door frame, the rendered depth map 715 can be considered to represent the distance from the pixels constituting the door frame to the virtual camera 720 for the part of the depth map including the door frame. The part 717 of the rendered depth map 715 represents the door that opens once swung outwards from the door frame.

[0072] As Figure 7A shown, the virtual camera can be placed at a position centered on the door frame and at a predetermined distance, for example 2 meters, from the door frame. Thus, for each different shape, different camera positions and orientations can be utilized.

[0073] Figure 7C is a simplified schematic diagram showing a rendered depth map associated with two wall corners and an associated virtual camera according to an embodiment of the present invention. In the illustrated embodiment, the two walls intersect at an angle of 90°. As Figure 7C shown, the virtual camera 730 is centered at the corner where two adjacent walls intersect.

[0074] The method further includes synthesizing depth maps (418) from each virtual camera of each 3D mesh for each detected shape. In other words, for each detected shape, depth maps from each virtual camera are synthesized based on the shape's 3D mesh. Thus, the embodiment provides depth maps associated with each virtual camera.

[0075] Figure 8A is a simplified schematic diagram showing a rendered depth map of an internal view of a door frame rendered from a virtual camera according to an embodiment of the present invention. Figure 8B is a simplified schematic diagram showing a rendered depth map of an external view of a door frame rendered from a virtual camera according to an embodiment of the present invention. In these depth maps, grayscale can be used to represent depth values. As Figure 8BAs shown, the door opens on the left side of the depth map. Accordingly, the open door obscures a portion of the left side of the door frame. It should be understood that the door frame and the door can be regarded as two different shapes. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0076] Figure 8A The depth map shown in Figure 7A is associated with the virtual camera 710 shown in Figure 8B The depth map shown in Figure 7B is associated with the virtual camera 720 shown in

[0077] The method further includes performing a joint optimization (420) of the camera pose and / or the dimensions and positions of each detected shape. The position of each detected shape is related to the pose of the rendered depth map. The dimensions are similar. These camera pose alignments utilize the rendered depth map from process 414 and the captured depth map (e.g., passive or active) as part of the joint optimization. The joint optimization can be accomplished using ICP-based alignment or other techniques, which can also be referred to as pose estimation / refinement. It is noted that the pose of the rendered depth map is optionally optimized as part of this process.

[0078] Further referring to Figure 4 and the description provided by process 416, the process of shape-aware camera pose alignment can include the following steps:

[0079] Step 1: Find the nearest point pairs between each frame-frame pair.

[0080] Step 2: Find the nearest point pairs between each frame-shape pair.

[0081] Step 3: Jointly optimize R, T for each frame and F, G, and D for each shape using the following objective function.

[0082] Step 4: Iterate starting from Step 1 until the optimization converges.

[0083] Objective function:

[0084] In the objective function, the first term involves the alignment between the captured depth maps. The second term involves the alignment between the captured depth map and the rendered depth map (i.e., the detected shape). The third and fourth terms involve ensuring the smoothness of the pose trajectory.

[0085] In the above equation,

[0086] i provides an index for each frame

[0087] j provides an index for each other frame

[0088] m provides an index for each nearest point pair

[0089] p i (·) and qj ( (·) represents the depth point p from frame i and its corresponding nearest depth point q from frame j

[0090] p i (·) and h k (·) represents the depth point p from frame i and its corresponding nearest depth point h from shape k

[0091] R i and T i relate to the rotation and translation of frame i (i.e., the camera pose)

[0092] F k and G k relate to the rotation and translation of shape k (i.e., the camera pose)

[0093] D k specifies the dimensions of shape k

[0094] w represents the weight of each term

[0095] After the joint optimization of the camera pose has been performed, the original depth image is aligned with the rendered depth map and thus also with one or more detected shapes. Consequently, the point cloud for 3D mesh reconstruction will become more accurate and consistent, especially in regions close to prominent shapes and structures. In Figure 9A and Figure 9B a comparison of the point cloud alignment with and without detected shapes is shown. Figure 9A is a simplified point cloud diagram showing the captured depth map. Figure 9B is a simplified point cloud diagram showing the captured depth map and the rendered depth map using the shape-aware method provided by an embodiment of the present invention. It can be observed that, as shown in the image Figure 9B the points are better aligned with the shape-aware camera pose estimation.

[0096] It should be understood that Figure 4 the specific steps shown in provide a specific method for forming a shape-aware camera pose alignment according to an embodiment of the present invention. Other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally Figure 4 each of the steps shown in may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, depending on the specific application, additional steps may be added or removed. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0097] Returning again toFigure 2 , method 200 includes performing shape-aware volumetric fusion (220) and forming a reconstructed 3D mesh (222) using shape-aware volumetric fusion techniques. Regarding Figure 5 Additional descriptions related to embodiments of shape-aware volumetric fusion are provided.

[0098] It should be understood that Figure 2 The specific steps shown in provide a specific method for generating a 3D mesh of a scene using depth maps captured from multiple frames according to an embodiment of the present invention. Other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, Figure 2 Each of the steps shown in may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, depending on the specific application, additional steps may be added or removed. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0099] Figure 5 is a simplified flowchart showing a method of performing shape-aware volumetric fusion according to an embodiment of the present invention. When applying this technique, the detected shape is utilized, resulting in a sharper and clearer shape mesh than other methods.

[0100] Method 500 includes recreating a shape mesh (510) for each detected shape with an optimized shape size. The method also includes rendering depth maps for each virtual camera from each shape mesh (512) and performing joint volumetric fusion (514) using the captured depth maps and the rendered depth maps.

[0101] Joint volumetric fusion (514) is developed on top of the classic work of volumetric fusion, first introduced in "Volumetric Methods for Building Complex Models from Range Images". More specifically, a 3D volume is first created, which can be uniformly subdivided into a 3D mesh of voxels and mapped to the 3D physical space of the capture area. Each voxel of this volume representation will store a value that specifies the relative distance from the actual surface. These values are positive in front of the actual surface and negative behind it, so this volume representation implicitly describes the 3D surface: the location where the value changes sign. Volumetric fusion can convert a set of captured depth maps into this volume representation. The distance value, truncated signed distance function (TSDF), in each voxel is calculated as follows:

[0102]

[0103] where

[0104] v is the position of the voxel

[0105] tsdf(v) is the relative distance value of the voxel

[0106] proj i (v) is the projection of v on the captured depth map i

[0107] is the weight of the voxel v projected onto the captured depth map i

[0108] D i (·) is the captured depth map i

[0109] T i is the position of the camera i

[0110] If (1) the voxel v is outside the frustum of the camera i or (2) |D i (proj i (v)) - ||v - T i ||| is greater than the predetermined truncation distance M, then will always be set to zero. For other cases, can be set to 1 or the confidence value of the corresponding point in the captured depth map.

[0111] For shape-aware volume fusion performed according to an embodiment of the present invention, the truncated signed distance function is calculated from both the captured depth map and the rendered depth map (i.e., the detected shape).

[0112]

[0113] Where

[0114] D i (·) is the rendered depth map s

[0115] G i is the position of the virtual camera s

[0116] If (1) the voxel v is outside the frustum of the virtual camera s or (2) |E s (proj s (v)) - ||v - G s ||| is greater than the predetermined truncation distance M, then will also be set to zero. When it is not zero, will be set to a value (i.e., 20) greater than that of the captured depth map (i.e., 1), such that the points from the rendered depth map will dominate. For points closer and closer to the boundary of the detected shape, some embodiments also gradually reduce the value (i.e., from 20 to 1). Reducing the weight around the boundary can create a smooth transition of the detected shape (sharper) to the original mesh generated using the captured depth map.

[0117] After shape-aware volumetric fusion, the main structures (e.g., door frames, wall corners, wall ends, etc.) in the final mesh will be sharper and clearer.

[0118] It should be understood that Figure 5 the specific steps shown in provide a specific method for performing shape-aware volumetric fusion according to an embodiment of the present invention. Other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the steps outlined above in a different order. Additionally, Figure 5 each of the steps shown in may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, depending on the specific application, additional steps may be added or removed. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0119] Figure 10A is an image showing a first reconstructed 3D mesh reconstructed using the method described with respect to Figure 1 Figure 10B is an image showing a second reconstructed 3D mesh reconstructed using the method described with respect to Figure 2 Figure 10A and Figure 10B respectively provide a comparison of 3D meshes reconstructed without and with shape-aware 3D reconstruction techniques.

[0120] In the image shown in representing a door in a wall, the reconstructed 3D mesh includes ripples along the left edge of the door jamb and along the right edge of the door jamb. In the image shown in Figure 10A Figure 10B the same door shown in is shown, and shape-aware 3D mesh reconstruction produces a clearer and more accurate output. Considering the left edge of the door jamb, the wall appears to recede from the viewer. This bow that inaccurately represents the physical scene is likely caused by an error in estimating the camera pose. Figure 10A

[0121] As shown in Figure 10B the transition from the door frame shape to the rest of the mesh is smoother and is clearly defined by straight vertical door jambs. Thus, for indoor scenes, embodiments of the present invention provide visually pleasing and accurate 3D mesh reconstruction.

[0122] Figure 11A is an image showing a third reconstructed 3D mesh reconstructed using the method described with respect to Figure 1 Figure 11B is an image showing a fourth reconstructed 3D mesh reconstructed using the method described with respect to Figure 2 Figure 11A and Figure 11BProvide a comparison of 3D meshes reconstructed without and with shape-aware 3D reconstruction techniques, respectively.

[0123] In the image shown in Figure 11A representing the booth and table in the niche, the reconstructed 3D mesh includes ripples in the wall ends that form the left side of the niche and ripples in the wall ends that form the right side of the niche. Additionally, the wall above the niche exhibits ripples and irregularities on the left side of the wall above the bench. In Figure 11B the image shown, Figure 11A showing the same niche, bench, and table as shown in Figure 11A shape-aware 3D mesh reconstruction produces a clearer and more accurate output. In particular, the wall forming the right edge of the niche appears to extend into the Figure 11A next niche. However, in Figure 11B the right side of the left wall is flat with a clean wall end, clearly separating adjacent niches and accurately representing the physical scene.

[0124] Figure 12 is a simplified schematic diagram of a system for reconstructing a 3D mesh using depth images according to an embodiment of the present invention. The system includes a depth camera 1220 that can be used to collect a series of captured depth maps. In this example, a depth camera (position 1) is employed to capture a first depth map of scene 1210, and when the camera (1222) is in position 2, a second depth map of scene 1210 is captured.

[0125] The set of captured depth maps is sent to a computer system 1230 that can be integrated with or separate from the depth camera. The computer system is operable to execute the computational methods described herein and generate a reconstructed 3D mesh of scene 1210 for display to a user via a display 1232. The reconstructed 3D mesh can be sent to other systems via an I / O interface 1240 for display, storage, etc. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0126] Figure 13 is a block diagram of a computer system or information processing device that may incorporate embodiments, be incorporated into embodiments, or be used to practice any innovation, embodiment, and / or example found within the present disclosure.

[0127] Figure 13 is a block diagram of computer system 1300. Figure 13 This is merely illustrative. In some embodiments, the computer system includes a single computer device, where the subsystems may be components of the computer device. In other embodiments, the computer system may include multiple computer devices with internal components, each computer device being a subsystem. Computer system 1300 and any of its components or subsystems may include hardware and / or software elements configured to execute the methods described herein.

[0128] The computer system 1300 may include familiar computer components such as one or more data processors or central processing units (CPUs) 1305, one or more graphics processors or graphics processing units (GPUs) 1310, a memory subsystem 1315, a storage subsystem 1320, one or more input / output (I / O) interfaces 1325, a communication interface 1330, and the like. The computer system 1300 may include a system bus 1335 that interconnects the above components and provides functionality, such as connections for communication between devices.

[0129] One or more data processors or central processing units (CPUs) 1305 may execute logic or program code or provide application-specific functionality. Some examples of the CPU 1305 may include one or more microprocessors (e.g., single-core and multi-core) or microcontrollers, one or more field programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs). As a user here, the processor includes a multi-core processor located on the same integrated chip, or multiple processing units on a single circuit board or networked.

[0130] One or more graphics processors or graphics processing units (GPUs) 1310 may execute logic or program code associated with graphics or provide graphics-specific functionality. The GPU 1310 may include any conventional graphics processing unit, such as those provided by conventional video cards. In various embodiments, the GPU 1310 may include one or more vector or parallel processing units. These GPUs may be user programmable and include hardware elements for encoding / decoding specific types of data (e.g., video data) or for accelerating 2D or 3D drawing operations, texturing operations, shading operations, etc. One or more graphics processors or graphics processing units (GPUs) 1310 may include any number of registers, logic units, arithmetic units, caches, memory interfaces, and the like.

[0131] The memory subsystem 1315 may store information, for example, using machine-readable articles, information storage devices, or computer-readable storage media. Some examples may include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, and other semiconductor memories. The memory subsystem 1315 may include data and program code 1340.

[0132] The storage subsystem 1320 can also store information using machine-readable articles, information storage devices, or computer-readable storage media. The storage subsystem 1320 can use the storage medium 1345 to store information. Some examples of the storage medium 1345 used by the storage subsystem 1320 can include floppy disks, hard disks, optical storage media such as CD-ROMs, DVDs, and barcodes, removable storage devices, network storage devices, and the like. In some embodiments, the storage subsystem 1320 can be used to store all or part of the data and program code 1340.

[0133] One or more input / output (I / O) interfaces 1325 can perform I / O operations. One or more input devices 1350 and / or one or more output devices 1355 can be communicatively coupled to one or more I / O interfaces 1325. One or more input devices 1350 can receive information from one or more sources of the computer system 1300. Some examples of one or more input devices 1350 can include computer mice, trackballs, trackpads, joysticks, wireless remote controls, graphics tablets, voice command systems, eye-tracking systems, external storage systems, monitors appropriately configured as touchscreens, communication interfaces appropriately configured as transceivers, and the like. In various embodiments, one or more input devices 1350 can allow a user of the computer system 1300 to interact with one or more non-graphical or graphical user interfaces to input comments, select objects, icons, text, user interface widgets, or other user interface elements that appear on the monitor / display device via commands, button clicks, and the like.

[0134] One or more output devices 1355 can output information to one or more destinations of the computer system 1300. Some examples of one or more output devices 1355 can include printers, fax machines, feedback devices for mice or joysticks, external storage systems, monitors or other display devices, communication interfaces appropriately configured as transceivers, and the like. One or more output devices 1355 can allow a user of the computer system 1300 to view objects, icons, text, user interface widgets, or other user interface elements. A display device or monitor can be used with the computer system 1300 and can include hardware and / or software elements configured to display information.

[0135] The communication interface 1330 can perform communication operations, including sending and receiving data. Some examples of the communication interface 1330 can include network communication interfaces (e.g., Ethernet, Wi-Fi, etc.). For example, the communication interface 1330 can be coupled to a communication network / external bus 1360, such as a computer network, a USB hub, etc. A computer system can include multiple identical components or subsystems, for example, connected together via the communication interface 1330 or via an internal interface. In some embodiments, a computer system, subsystem, or device can communicate via a network. In this case, one computer can be considered a client and another computer is a server, where each can be part of the same computer system. The client and the server can each include multiple systems, subsystems, or components.

[0136] The computer system 1300 can also include one or more applications (e.g., software components or functions) to be executed by the processor to perform, carry out, or otherwise implement the techniques disclosed herein. These applications can be implemented as data and program code 1340. Additionally, computer programs, executable computer code, human-readable source code, shader code, rendering engines, etc., as well as data such as image files, models including geometric descriptions of objects, ordered geometric descriptions of objects, process descriptions of models, scene descriptor files, etc., can be stored in the memory subsystem 1315 and / or the storage subsystem 1320.

[0137] Such a program can also be encoded and transmitted using a carrier signal suitable for transmission via wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, a computer-readable medium according to an embodiment of the present invention can be created using a data signal encoded with such a program. The computer-readable medium encoded with the program code can be packaged with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or within a single computer product (e.g., a hard disk drive, a CD, or an entire computer system) and can exist on or within different computer products in a system or network. The computer system can include a monitor, a printer, or other suitable display for providing any of the results mentioned herein to a user.

[0138] Any method described herein may be performed wholly or in part by a computer system including one or more processors that may be configured to perform these steps. Accordingly, embodiments may relate to a computer system configured to perform the steps of any method described herein, potentially employing different components to perform the corresponding steps or groups of corresponding steps. Although presented as numbered steps, the method steps herein may be performed simultaneously or in a different order. Additionally, portions of these steps may be used in conjunction with portions of other steps from other methods. Further, all or part of these steps may be optional. Additionally, any step of any method may be performed with a module, circuit, or other means for performing these steps.

[0139] While the various embodiments of the invention have been described above, it should be understood that they have been presented by way of example only, and not as limitations. Similarly, the various figures may depict an exemplary architecture or other configuration for the present disclosure, which is done to assist in understanding the features and functions that may be included in the present disclosure. The present disclosure is not limited to the exemplary architecture or configuration shown, but may be implemented using a variety of alternative architectures and configurations. Additionally, while the present disclosure has been described above in terms of various exemplary embodiments and implementations, it should be understood that the various features and functions described in one or more of the individual embodiments are not limited to their applicability to the particular embodiments in which they are described. Rather, they may be applied individually or in some combination to one or more other embodiments of the present disclosure, whether or not those embodiments are described, and whether or not those features are presented as part of the embodiments described. Accordingly, the breadth and scope of the present disclosure should not be limited by any of the above exemplary embodiments.

[0140] In this document, the term "module" as used herein refers to software, firmware, hardware, and any combination of these elements for performing the related functions described herein. Additionally, for purposes of discussion, the various modules are described as discrete modules; however, it will be apparent to one of ordinary skill in the art that two or more modules may be combined to form a single module that performs the related functions in accordance with embodiments of the invention.

[0141] It should be understood that, for clarity, the embodiments of the invention have been described above with reference to different functional units and processors. However, it will be apparent that any suitable distribution of functionality between different functional units, processors, or domains may be used without departing from the invention. For example, functions illustrated as being performed by separate processors or controllers may be performed by the same processor or controller. Accordingly, the reference to a particular functional unit is only considered as a reference to a suitable means for providing the described functionality, and not as indicating a strict logical or physical structure or organization.

[0142] Unless otherwise expressly stated, the terms and phrases used herein and their variants shall be interpreted as open-ended rather than limiting. By way of example of the foregoing: the term "including" shall be understood to mean "including but not limited to", etc.; the term "example" is used to provide exemplary instances of the items under discussion, rather than an exhaustive or limiting list thereof; and adjectives such as "conventional", "traditional", "normal", "standard", "known" and terms of similar import shall not be construed as limiting the items described to a given time period or to items available at a given time. Rather, these terms should be understood to encompass conventional, traditional, normal or standard techniques now known or available at any time in the future. Similarly, a group of items associated with the conjunction "and" shall not be understood to require that each of these items be present in the grouping, but rather shall be understood as "and / or", unless otherwise expressly stated. Similarly, a group of items associated with the conjunction "or" shall not be understood to require mutual exclusivity within the group, but rather shall be understood as "and / or", unless otherwise expressly stated. Further, although the items, elements or components of the present disclosure may be described or claimed in the singular, the plural is contemplated within their scope, unless expressly stated to be limited to the singular. In some instances, the presence of expansive words and phrases such as "one or more", "at least", "but not limited to" or other similar phrases shall not be construed to imply an intention or requirement for a narrower case in instances where such expansive phrases may be absent.

[0143] It should also be understood that the examples and embodiments described herein are for illustrative purposes only and that various modifications or changes thereto will be suggested to those skilled in the art and will be included within the spirit and scope of this application and within the scope of the appended claims.

Claims

1. A method for updating a camera pose, the method comprises: receiving, at one or more processors, a set of captured depth maps, each captured depth map in the set of captured depth maps being associated with a physical camera pose, the set of captured depth maps including a scene; using the one or more processors to detect a first shape and a second shape present in the scene, the first shape being characterized by a first size and / or a first position / orientation, and the second shape being characterized by a second size and / or a second position / orientation; using the one or more processors to create a first 3D mesh for the first shape; using the one or more processors to create a first virtual camera associated with the first 3D mesh in a first local reference frame; using the one or more processors to render a first depth map associated with the first virtual camera; using the one or more processors to create a second 3D mesh for the second shape; using the one or more processors to create a second virtual camera associated with the second 3D mesh in a second local reference frame; using the one or more processors to render a second depth map associated with the second virtual camera; using the one or more processors to identify a subset of the captured depth maps, each captured depth map in the subset including at least a first portion of the first shape or a second portion of the second shape; and using the one or more processors, the first shape, and the second shape to jointly solve for the physical camera pose, the first size and first position / orientation, and the second size and second position / orientation by optimizing the alignment between the first depth map, the second depth map, and the subset of the captured depth maps, thereby updating the physical camera pose associated with the subset of the captured depth maps to provide an updated physical camera pose.

2. The method according to claim 1, wherein detecting the first shape and the second shape comprises: using the one or more processors and determining, for a plurality of pixels in a point cloud, a plurality of horizontal planes, the plurality of horizontal planes being defined for each pixel in the plurality of pixels by a point normal from an adjacent pixel to each pixel in the plurality of pixels; using one or more processors to calculate a vector perpendicular to the plurality of horizontal planes, the vector defining a vertical direction associated with the point cloud; using the one or more processors to form a virtual plane orthogonal to the vertical direction; using the one or more processors to project points in the point cloud onto the virtual plane to generate a two-dimensional data set representing a plurality of points in the point cloud associated with a predetermined position in the virtual plane; using the one or more processors to calculate projection statistics of the points in the point cloud; using the one or more processors to detect a plurality of lines associated with vertical walls based on the calculated projection statistics; and using the one or more processors to detect the first shape and the second shape based on the projection statistics and the detected plurality of lines.

3. The method according to claim 1, wherein, the physical camera poses associated with the subset of the captured depth maps include a first physical camera pose and a second physical camera pose, the first physical camera pose is associated with a first captured depth map in the subset of the captured depth maps, the second physical camera pose is associated with a second captured depth map in the subset of the captured depth maps, and wherein the first captured depth map corresponds to a portion of the first depth map, and the second captured depth map corresponds to a portion of the second depth map.

4. The method according to claim 1, wherein, optimizing the alignment between the first depth map, the second depth map and the subset of the captured depth maps includes: optimizing the alignment between the first depth map and the second depth map; optimizing the alignment between the first depth map and the subset of the captured depth maps; and optimizing the alignment between the second depth map and the subset of the captured depth maps.

5. The method according to claim 1, wherein, the first 3D mesh for the first shape includes a plurality of triangles, and wherein at least a portion of the plurality of triangles is within a first field of view of the first virtual camera.

6. The method according to claim 5, further comprising: creating, using the one or more processors, a third virtual camera associated with the first 3D mesh, wherein a second portion of the plurality of triangles is within a second field of view of the third virtual camera, and wherein each triangle of the plurality of triangles is within at least one of the first field of view and the second field of view.

7. The method according to claim 1, wherein, the first local reference frame includes a first reference frame of the first shape.

8. The method according to claim 1, wherein, the second local reference frame includes a second reference frame of the second shape.

9. The method according to claim 1, wherein, the set of captured depth maps is obtained from different positions relative to the scene.

10. The method according to claim 1, wherein, the set of captured depth maps is obtained from a single position relative to the scene at different times.

11. A system for updating camera poses, the system comprising: a depth camera; and one or more processors communicatively coupled to the depth camera, wherein the one or more processors are configured to perform operations including the following: obtaining, via the depth camera, a set of captured depth maps, each captured depth map in the set of captured depth maps being associated with a physical camera pose, the set of captured depth maps including a scene; detecting a first shape and a second shape present in the scene, the first shape being characterized by a first size and / or a first position / orientation, and the second shape being characterized by a second size and / or a second position / orientation; creating a first 3D mesh for the first shape; creating a first virtual camera associated with the first 3D mesh in a first local reference frame; rendering a first depth map associated with the first virtual camera; Create a second 3D mesh for the second shape; Create a second virtual camera associated with the second 3D mesh in a second local reference frame; Render a second depth map associated with the second virtual camera; Identify a subset of the captured depth maps, each of the captured depth maps in the subset including at least a first portion of the first shape or a second portion of the second shape; and Jointly solve for the physical camera pose, the first dimensions and first position / orientation, and the second dimensions and second position / orientation using the first shape and the second shape by optimizing the alignment between the first depth map, the second depth map, and the subset of the captured depth maps, thereby updating the physical camera pose associated with the subset of the captured depth maps to provide an updated physical camera pose.

12. The system according to claim 11, wherein, the one or more processors are configured to detect the first shape and the second shape by performing additional operations including: Determine a plurality of horizontal planes for a plurality of pixels in a point cloud, the plurality of horizontal planes being defined for each pixel in the plurality of pixels by the point normals from adjacent pixels to each pixel in the plurality of pixels; Calculate a vector perpendicular to the plurality of horizontal planes, the vector defining a vertical direction associated with the point cloud; Using the one or more processors, form a virtual plane orthogonal to the vertical direction; Project the points in the point cloud onto the virtual plane to generate a two-dimensional data set representing a plurality of points in the point cloud associated with a predetermined position in the virtual plane; Calculate projection statistics of the points in the point cloud; Detect a plurality of lines based on the calculated projection statistics, the plurality of lines being associated with vertical walls; and and Detect the first shape and the second shape based on the projection statistics and the plurality of detected lines.

13. The system according to claim 11, wherein, the physical camera pose associated with the subset of the captured depth maps includes a first physical camera pose and a second physical camera pose, the first physical camera pose being associated with a first captured depth map in the subset of the captured depth maps, the second physical camera pose being associated with a second captured depth map in the subset of the captured depth maps, and wherein the first captured depth map corresponds to a portion of the first depth map, and the second captured depth map corresponds to a portion of the second depth map.

14. The system according to claim 11, wherein, the one or more processors are configured to optimize the alignment between the first depth map, the second depth map, and the subset of the captured depth maps by performing additional operations including: Optimize the alignment between the first depth map and the second depth map; Optimize the alignment between the first depth map and the subset of the captured depth maps; and and Optimize the alignment between the second depth map and the subset of the captured depth maps.

15. The system according to claim 11, wherein, The first 3D mesh for the first shape includes a plurality of triangles, and at least a portion of the plurality of triangles is within a first field of view of the first virtual camera.

16. The system according to claim 15, wherein, the one or more processors are configured to perform additional operations including: create a third virtual camera associated with the first 3D mesh, wherein a second portion of the plurality of triangles is within a second field of view of the third virtual camera, and wherein each triangle of the plurality of triangles is within at least one of the first field of view and the second field of view.

17. A system for forming a reconstructed 3D mesh, the system comprising: a depth camera; and one or more processors communicatively coupled to the depth camera, wherein the one or more processors are configured to perform operations including: acquire a set of captured depth maps associated with a scene via the depth camera; perform an initial camera pose alignment associated with each captured depth map in the set of captured depth maps; overlay the set of captured depth maps in a reference frame; detect one or more shapes in the overlaid set of captured depth maps, thereby providing one or more detected shapes; use the one or more detected shapes to update the initial camera pose alignment based on an overlap between the overlaid set of captured depth maps and the one or more detected shapes to provide a shape-aware camera pose alignment associated with each captured depth map in the set of captured depth maps; perform shape-aware volume fusion using the one or more detected shapes; and form the reconstructed 3D mesh associated with the scene.

18. The system according to claim 17, wherein, the reference frame includes a reference frame of one of the one or more detected shapes.

19. The system according to claim 17, wherein, the one or more detected shapes include at least one of a corner or a door frame.

20. The system according to claim 17, wherein, the one or more processors are configured to provide the shape-aware camera pose alignment associated with each captured depth map in the set of captured depth maps by performing additional operations including: create a 3D mesh for each of the one or more detected shapes; wherein the overlaid set of captured depth maps is associated with a physical camera pose, and each of the one or more detected shapes is characterized by a size and a position / orientation; create one or more virtual cameras associated with each 3D mesh in a local reference frame; render one or more depth maps, each of the rendered one or more depth maps being associated with each virtual camera associated with each 3D mesh; and jointly solve for the physical camera pose and the position / orientation of each of the one or more detected shapes by optimizing an alignment between the rendered one or more depth maps and the set of captured depth maps.

Citation Information

Patent Citations

  • Moving object segmentation using depth images

    CN102663722A

  • Method and device for capturing markerless motion and reconstructing scene

    CN102842148A