Simple environment solver using planar extraction
By identifying surface planes and inferring corner points, a simple mesh model is constructed, which solves the resource waste problem of dense mesh models in existing XR systems and achieves efficient and accurate environment reconstruction and virtual object rendering.
Patent Information
- Application Number
- CN202080060983.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-03
- Filing Date
- 2020-06-25
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2040-06-25
AI Technical Summary
Existing XR systems, when reconstructing physical environments, use dense mesh models that contain unnecessary details, leading to a waste of computational resources and storage space, and are unable to accurately handle area calculations of decorative effects or rendering of virtual objects.
By identifying surface planes and inferring corner points, a simple mesh model of the environment is constructed, reducing computational and storage requirements and enabling rapid reconstruction of a 3D representation of the environment.
It improves the efficiency and accuracy of environment reconstruction, reduces the consumption of computing resources and network bandwidth, and supports multi-user experience and accurate rendering of virtual objects.
Smart Images

Figure CN114341943B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to cross-reality systems that use 3D world reconstruction to render scenes. Background Technology
[0002] Computers can control human user interfaces to create X-Reality (XR or Cross-Reality) environments, in which the computer generates parts or all of the XR environment that is perceived by the user. These XR environments can be Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR) environments, where parts or all of the XR environment can be generated by the computer using data describing the environment. This data can describe, for example, virtual objects that can be rendered in a way that the user senses or perceives as part of the physical world and can interact with. Because the data is rendered and presented through user interface devices (e.g., head-mounted displays), the user can experience these virtual objects. The data can be displayed to the user, or can control audio played to the user, or can control a tactile (or haptic) interface, allowing the user to experience the tactile sensation of sensing or perceiving the virtual objects.
[0003] XR systems are useful for many applications, spanning scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. Compared to VR, AR and MR involve one or more virtual objects that are related to real-world objects. The experience of interacting with real-world objects greatly enhances the user experience of XR systems and opens doors to a variety of applications that present realistic and easily understood information about how to change the physical world.
[0004] XR systems can represent the physical world surrounding the user as a "mesh." A mesh can be represented by multiple interconnected triangles. Each triangle has edges connecting points on a surface of objects within the physical world, such that each triangle represents a portion of a surface. Information about a portion of the surface (such as color, texture, or other attributes) can be stored associatively within the triangle. In operation, an XR system can process image information to detect points and surfaces, thereby creating or updating the mesh. Summary of the Invention
[0005] This application relates to methods and apparatus for rapidly generating environments containing computer-generated objects. The techniques described herein can be used together, individually, or in any suitable combination.
[0006] Some embodiments relate to a portable electronic system. The portable electronic system includes a sensor configured to capture information about a physical world, and a processor configured to execute computer-executable instructions to compute a three-dimensional 3D representation of a portion of the physical world, at least in part based on the captured information about the physical world. The computer-executable instructions include instructions for: extracting a plurality of planar segments from the information captured by the sensor; identifying a plurality of surface planes, at least in part based on the plurality of planar segments; and inferring a plurality of corner points of the portion of the physical world, at least in part based on the plurality of surface planes.
[0007] In some embodiments, the computer-executable instructions further include instructions for using the corner points to construct a mesh model of the portion of the physical world.
[0008] In some embodiments, the plurality of surface planes are identified at least in part based on input from a user wearing at least a portion of the portable electronic system.
[0009] In some embodiments, the portable electronic system includes a transceiver configured to communicate with a device providing remote storage via a computer network.
[0010] In some embodiments, the processor implements a service configured to provide an application with a 3D representation of the portion of the physical world.
[0011] In some embodiments, the service stores the corner point in local storage or transmits the corner point to cloud storage as the three-dimensional 3D representation of the part of the physical world.
[0012] In some embodiments, identifying the plurality of surface planes includes: determining whether a principal plane segment normal exists in a group of plane segment normals of the plurality of plane segments; setting the principal plane segment normal as a surface plane normal when the determination indicates a principal plane segment normal in the group; and calculating the surface plane normal based on at least a portion of the plane segment normals in the group when the determination indicates no principal plane segment normal in the group.
[0013] In some embodiments, calculating the surface plane normal includes calculating a weighted average of at least a portion of the plane segment normals in the group.
[0014] In some embodiments, inferring the plurality of corner points of the portion of the physical world includes: extending a first surface plane and a second surface plane of the plurality of surface planes to infinity; and obtaining a boundary line intersecting the first surface plane and the second surface plane.
[0015] In some embodiments, inferring a plurality of corner points of the portion of the physical world further includes inferring one of the plurality of corner points by intersecting the boundary line with a third surface plane.
[0016] Some embodiments relate to a non-transitory computer-readable medium encoded with a plurality of computer-executable instructions, which, when executed by at least one processor, perform a method for providing a three-dimensional 3D representation of a portion of a physical world, in which the portion of the physical world is represented by a plurality of corner points. The method includes: capturing information about a portion of the physical world within a user's field of view (FOV); extracting a plurality of planar segments from the captured information; identifying a plurality of surface planes from the plurality of planar segments; and calculating a plurality of corner points representing the portion of the physical world based on the intersection of the surface planes among the identified plurality of surface planes.
[0017] In some embodiments, the method includes calculating whether a first plurality of corner points form a closure.
[0018] In some embodiments, calculating whether a closure is formed includes: determining whether the boundary lines joining the first plurality of corner points can be connected to define the surfaces that join and define the closure volume.
[0019] In some embodiments, the portion of the physical world is a first portion of the physical world, the user is a first user, and the plurality of corner points are a first plurality of corner points; and the method further includes: receiving a second plurality of corner points of a second portion of the physical world from a second user; and providing the 3D representation of the physical world based at least in part on the first plurality of corner points and the second plurality of corner points.
[0020] In some embodiments, the user is a first user, and the method further includes: transmitting, via a computer network, corner points calculated based on captured information about the portion of the physical world within the first user's field of vision; receiving the transmitted corner points at an XR device used by a second user; and rendering information about the portion of the physical world to the second user using the XR device based on the received corner points.
[0021] In some embodiments, the method includes: calculating metadata for the corner points, the metadata indicating the positional relationship between the corner points.
[0022] In some embodiments, the method includes: storing the corner point including corresponding metadata such that the corner point can be retrieved by multiple users, including the user.
[0023] Some embodiments relate to a method of operating an interreal system to reconstruct an environment. The interreal system includes a processor configured to process image information in communication with sensors worn by a user, the sensors generating depth information for various regions within the sensor's field of view. The method includes: extracting a plurality of planar segments from the depth information; displaying the extracted planar segments to the user; receiving user input indicating a plurality of surface planes, each surface plane representing a surface defining the environment; and calculating a plurality of corner points of the environment based at least in part on the plurality of surface planes.
[0024] In some embodiments, the method includes determining whether the plurality of corner points form a closed loop.
[0025] In some embodiments, the method includes storing the corner point when it is determined that the closure has been formed. Attached Figure Description
[0026] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various figures is represented by the same numbers. For clarity, not every component will be labeled in every drawing. In the drawings:
[0027] Figure 1 This is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments.
[0028] Figure 2 This is a sketch illustrating an exemplary simplified AR scene that includes exemplary world reconstruction use cases including visual occlusion, physical-based interaction, and environmental reasoning according to some embodiments.
[0029] Figure 3 This is a schematic diagram illustrating a data flow in an AR system according to some embodiments, the AR system being configured to provide an experience of AR content interacting with the physical world.
[0030] Figure 4 This is a schematic diagram illustrating an example of an AR display system according to some embodiments.
[0031] Figure 5A This is a schematic diagram illustrating an AR display system according to some embodiments that renders AR content as the user moves through the physical world environment while wearing the device.
[0032] Figure 5B This is a schematic diagram illustrating viewing optical components and accessories according to some embodiments.
[0033] Figure 6This is a schematic diagram illustrating an AR system using a world reconstruction system according to some embodiments.
[0034] Figure 7A This is a schematic diagram illustrating a 3D space discrete as voxels according to some embodiments.
[0035] Figure 7B This is a schematic diagram illustrating the reconstruction range relative to a single viewpoint according to some embodiments.
[0036] Figure 7C This is a schematic diagram illustrating the perceived range relative to the reconstruction range at a single location, according to some embodiments.
[0037] Figures 8A to 8F This is a schematic diagram illustrating, according to some embodiments, the reconstruction of a surface into a voxel model by an image sensor viewing the surface in the physical world from multiple locations and viewpoints.
[0038] Figure 9 This is a schematic diagram illustrating a scene represented by bricks including voxels, surfaces in the scene, and a depth sensor capturing a depth image of the surfaces, according to some embodiments.
[0039] Figure 10A This is a schematic diagram showing a 3D space represented by eight bricks.
[0040] Figure 10B It is shown Figure 10A A schematic diagram of the voxel grid in the brick.
[0041] Figure 11 This is a schematic diagram illustrating a planar extraction system according to some embodiments.
[0042] Figure 12 This illustrates details regarding planar extraction according to some embodiments. Figure 11 A schematic diagram of the various parts of the planar extraction system.
[0043] Figure 13 This is a schematic diagram illustrating a scene represented by bricks including voxels, and exemplary planar data in the scene, according to some embodiments.
[0044] Figure 14 This illustrates some embodiments. Figure 11 A schematic diagram of a planar data repository.
[0045] Figure 15 This illustrates, according to some embodiments, when a planar query is sent to... Figure 11 A schematic diagram of planar geometry extraction when storing planar data in a planar data repository.
[0046] Figure 16AThis illustrates the generation according to some embodiments. Figure 15 A schematic diagram of the planar coverage points.
[0047] Figure 16B This is a schematic diagram illustrating various exemplary planar geometric representations that can be extracted from an exemplary rasterized planar mask according to some embodiments.
[0048] Figure 17 A mesh for a scene is shown according to some embodiments.
[0049] Figure 18A The diagram illustrates a representation of an outer rectangular plane according to some embodiments. Figure 17 The scene.
[0050] Figure 18B The diagram illustrates a representation of an inner rectangular plane according to some embodiments. Figure 17 The scene.
[0051] Figure 18C The diagram illustrates a polygonal plane according to some embodiments. Figure 17 The scene.
[0052] Figure 19 The diagram illustrates the use of some embodiments to... Figure 17 The mesh shown is planarized to create a denoised mesh. Figure 17 The scene.
[0053] Figure 20 This is a flowchart illustrating a method for operating an AR system to generate a 3D reconstruction of an environment according to some embodiments.
[0054] Figure 21 It is shown that, according to some embodiments, it is at least partially based on... Figure 20 The flowchart shows a method for identifying surface planes using the obtained planar segments.
[0055] Figure 22 This illustrates an inference based on some embodiments. Figure 20 The flowchart shows the method for determining corner points.
[0056] Figure 23 This illustrates a configuration for execution according to some embodiments. Figure 20 A simplified schematic diagram of the AR system using this method.
[0057] Figure 24 This is a simplified schematic diagram illustrating the extraction of planar fragments of an environment according to some embodiments.
[0058] Figure 25 This illustrates a method based on some embodiments. Figure 24A simplified schematic diagram of the surface plane identified by extracting planar fragments from the environment.
[0059] Figure 26 This illustrates the intersecting according to some embodiments. Figure 25 A simplified diagram of the boundary line obtained from the two wall planes in the diagram.
[0060] Figure 27 This illustrates how, according to some embodiments, by making Figure 26 The boundary lines in the middle are respectively with Figure 25 A simplified diagram showing the corner point deduced from the intersection of the floor plane and the ceiling plane.
[0061] Figure 28 This illustrates, according to some embodiments, at least in part, based on corner points. Figure 24 A schematic diagram of the 3D reconstruction of the environment. Detailed Implementation
[0062] This paper describes methods and apparatus for reconstructing a three-dimensional (3D) world for creating and using environments (e.g., indoor environments) in X-Reality (XR or Cross-Reality) systems. Typically, a 3D representation of an environment is constructed by scanning the entire environment, including, for example, walls, floors, and ceilings, where the XR system is held and / or worn by a user. The XR system generates a dense mesh to represent the environment. The inventors have recognized and acknowledged that dense meshes can include unnecessary detail for the specific tasks performed by the XR system. For example, the system might construct a dense mesh model containing many triangles to represent small imperfections on a wall and any decorations on the wall, but this model might be used by an application that renders virtual objects covering the surface of the wall or identifies the location of the wall or calculates the area of the wall—tasks that may not be affected by small imperfections on the wall, or that may be impossible to complete if the area of the wall cannot be accurately calculated due to the decorations on the wall's surface. Examples of such applications might include home-contracting apps where data representing the room structure may suffice, and games like "Dr. Grordbort's Invaders" that require data representing areas of the walls to allow portholes to be opened for evil robots, and which might give error messages about insufficient wall space due to decorations covering the wall surfaces.
[0063] The inventors have recognized and appreciated techniques for quickly and accurately representing other parts of a room or environment as a set of corner points. Corner points can be obtained by identifying surface planes, representing surfaces of the environment, such as any walls, floors, and / or ceilings. Surface planes can be calculated based on information collected by sensors on a wearable device that can be used to scan a portion of the environment. The sensors can provide depth and / or image information. The XR system can obtain plane segments from the depth and / or image information. Each plane segment can indicate the orientation of the plane, represented by, for example, a plane normal. The XR system can then identify surface planes of the environment from a group of one or more plane segments. In some embodiments, the surface planes can be specifically selected by a user operating the XR system. In some embodiments, the XR system can automatically identify surface planes.
[0064] A 3D representation of an environment can be quickly and accurately reconstructed using corner points. For example, a simple mesh representation of the environment can be generated from corner points, replacing or supplementing a mesh calculated in a conventional manner. In some embodiments, for multi-user experiences, corner points can be transmitted between multiple users in an XR experience involving the environment. Corner points of the environment can be transmitted faster than a dense mesh of the environment. Furthermore, building a 3D representation of the environment based on corner points consumes less computational power, storage space, and network bandwidth compared to scanning a dense mesh of the entire environment.
[0065] The techniques described herein can be used with or alone in a variety of devices and scenarios, including wearable or portable devices with limited computing resources to provide cross-reality scenarios. In some embodiments, these techniques may be implemented by a service that forms part of an XR system. Applications performing simple tasks that require only sufficient information to reconstruct the environment can interact with the service to obtain a set of corner points, with or without associated metadata about the points and / or surfaces that define those points, to present information about the environment. For example, the application may render virtual objects relative to these surfaces. For instance, the application may render virtual images or other objects hanging on a wall. As another example, the application may render a virtual color overlay on a wall to change its perceived color, or it may display labels on the surfaces of the environment, such as labels indicating areas of the surface, the amount of paint needed to cover the surface, or other information about the environment that may be calculated.
[0066] AR System Overview
[0067] Figure 1-2 Such a scenario is illustrated. For illustrative purposes, the AR system is used as an example of an XR system. Figure 3-8 illustrates an exemplary AR system that includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.
[0068] refer to Figure 1 The text describes an outdoor AR scene 4 where the AR user sees a park-like setting 6 in the physical world, characterized by people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the AR user also perceives that they "see" a robot statue 10 standing on the physical world concrete platform 8, and a flying, anthropomorphic cartoon avatar 2 that looks like a bumblebee, even though these elements (e.g., avatar 2 and robot statue 10) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, producing an AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements within other virtual or physical world image elements is challenging.
[0069] Such AR scenarios can be realized through a system that includes a world reconstruction component, which can construct and update a representation of the physical world surface around the user. This representation can be used for occlusion rendering in physically based interactions, placing virtual objects, and for path planning and navigation of virtual characters, or for other operations in which information about the physical world is used. Figure 2 Another example of AR scene 200 is depicted, illustrating exemplary world reconstruction use cases according to some embodiments, including visual occlusion 202, physical-based interaction 204, and environmental reasoning 206.
[0070] Exemplary scenario 200 is a living room with walls, a bookshelf on one side of the wall, a floor lamp in a corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, the user of AR technology also perceives virtual objects such as images on the wall behind the sofa, birds flying through the door, a deer peeking out from the bookshelf, and a decorative item in the form of a windmill placed on the coffee table. For the image on the wall, AR technology needs not only information about the surface of the wall but also information about objects and surfaces in the room that are occluding the image (e.g., the shape of the lamp) to correctly render the virtual object. For the flying bird, AR technology needs information about all objects and surfaces around the room to render the bird with realistic physical effects, such as avoiding objects and surfaces or bouncing upon collision. For the deer, AR technology needs information about surfaces (e.g., the floor or the coffee table) to calculate the deer's placement. For the windmill, the system can identify it as an object detached from the table and infer that it is movable, while the corner of the bookshelf or the corner of the wall can be inferred to be fixed. Such distinctions can be used to infer which parts of the scene are used or updated in each of various operations.
[0071] A scene can be presented to a user through a system comprising multiple components, including a user interface that can stimulate one or more user senses (including visual, auditory, and / or tactile). Additionally, the system may include one or more sensors capable of measuring parameters of the physical portions of the scene, including the user's position and / or movement within those portions. Furthermore, the system may include one or more computing devices with associated computer hardware (e.g., memory). These components may be integrated into a single device or distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.
[0072] Figure 3 An AR system 302, configured to provide an experience of interacting with AR content and the physical world 306 according to some embodiments, is depicted. The AR system 302 may include a display 308. In the illustrated embodiment, the display 308 may be worn by a user as part of a headset, allowing the user to wear the display over their eyes like goggles or glasses. At least a portion of the display may be transparent, allowing the user to observe see-through reality 310. See-through reality 310 may correspond to the portion of the physical world 306 within the current viewpoint of the AR system 302, where the current viewpoint of the AR system 302 may correspond to the user's viewpoint, provided the user is wearing a headset incorporating the display and sensors of the AR system to acquire information about the physical world.
[0073] AR content can also be displayed on display 308, overlaid on perspective reality 310. To provide accurate interaction between AR content and perspective reality 310 on display 308, AR system 302 may include sensor 322 configured to capture information about the physical world 306.
[0074] Sensor 322 may include one or more depth sensors that output depth maps 312. Each depth map 312 may have multiple pixels, each pixel representing a distance to a surface in the physical world 306 in a specific direction relative to the depth sensor. Raw depth data may be derived from the depth sensors to create the depth maps. Such depth maps can be updated as quickly as the depth sensors can form new images, reaching hundreds or thousands of times per second. However, this data may be noisy and incomplete, and may have holes displayed as black pixels on the shown depth map.
[0075] The system may include other sensors, such as image sensors. Image sensors can acquire information that can be processed in other ways to represent the physical world. For example, images can be processed in world reconstruction component 316 to create a mesh, which represents the connected parts of objects in the physical world. Metadata about these objects, including, for example, color and surface texture, can similarly be acquired by sensors and stored as part of the world reconstruction.
[0076] The system can also acquire information about the user's head pose relative to the physical world. In some embodiments, sensor 310 may include an inertial measurement unit (IMU) that can be used to calculate and / or determine head pose 314. Head pose 314 for depth mapping may instruct the sensor to capture the current viewpoint of the depth map in, for example, six degrees of freedom (6DoF), but head pose 314 may be used for other purposes, such as relating image information to specific parts of the physical world or relating the position of a display worn on the user's head to the physical world. In some embodiments, head pose information may be derived in ways other than those of an IMU (e.g., analyzing objects in an image).
[0077] The world reconstruction component 316 can receive depth maps 312 and head pose 314 from sensors, as well as any other data, and integrate this data into a reconstruction 318, which can at least appear as a single, combined reconstruction. The reconstruction 318 can be more complete and less noisy than the sensor data. The world reconstruction component 316 can update the reconstruction 318 using spatial and temporal averaging of sensor data from multiple viewpoints that varies over time.
[0078] Reconstruction 318 may include a representation of the physical world having one or more data formats, such as voxels, meshes, planes, etc. Different formats may represent alternative representations of the same part of the physical world, or may represent different parts of the physical world. In the example shown, on the left side of reconstruction 318, a portion of the physical world is rendered as a global surface; on the right side of reconstruction 318, a portion of the physical world is rendered as a mesh.
[0079] Reconstruction 318 can be used for AR functions, such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation can change as the user moves or objects in the physical world change. Aspects of Reconstruction 318 can be used, for example, by component 320 that generates a global surface representation that varies in world coordinates, which can be used by other components.
[0080] AR content can be generated based on this information, for example, by AR application 304. AR application 304 can be a game program that performs one or more functions, for example, based on information about the physical world (e.g., visual occlusion, physics-based interaction, and environmental reasoning). It can perform these functions by querying data in different formats from reconstruction 318 generated by world reconstruction component 316. In some embodiments, component 320 can be configured to update its output when the representation in a region of interest in the physical world changes. This region of interest can be set, for example, to a portion of the physical world near the user of the system, such as a portion within the user's field of vision, or to be projected (predicted / determined) to enter the user's field of vision.
[0081] AR application 304 can use this information to generate and update AR content. The virtual portion of the AR content can be combined with perspective reality 310 and displayed on display 308 to create a realistic user experience.
[0082] In some embodiments, an AR experience can be provided to the user through a wearable display system. Figure 4 An example of a wearable display system 80 (hereinafter referred to as "System 80") is shown. System 80 includes a head-mounted display device 62 (hereinafter referred to as "Display Device 62") and various mechanical and electronic modules and systems to support the functionality of Display Device 62. Display Device 62 may be coupled to a frame 64, which may be worn by a display system user or viewer 60 (hereinafter referred to as "User 60") and configured to position Display Device 62 in front of User 60's eyes. According to various embodiments, Display Device 62 may be a sequential display. Display Device 62 may be monocular or binocular. In some embodiments, Display Device 62 may be... Figure 3 Example of display 308 in the image.
[0083] In some embodiments, speaker 66 is coupled to frame 64 and positioned near the ear canal of user 60. In some embodiments, another speaker (not shown) is positioned near another ear canal of user 60 to provide stereo / shapeable sound control. Display device 62 is operatively coupled to local data processing module 70, for example via a wired lead or wireless connection 68. Local data processing module 70 can be mounted in various configurations, such as being fixedly attached to frame 64, fixedly attached to a helmet or hat worn by user 60, embedded in headphones, or otherwise detachably connected to user 60 (e.g., in a backpack configuration or a belt-coupled configuration).
[0084] The local data processing module 70 may include a processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which can be used to assist in data processing, caching, and storage. Data includes a) data captured from sensors (which may be operatively coupled to frame 64, for example) or otherwise attached to user 60 (e.g., image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, radios, and / or gyroscopes), and / or b) data acquired using remote processing module 72 and / or remote data repository 74, which may be used to transmit data to display device 62 after such processing or retrieval. The local data processing module 70 may be operatively coupled to remote processing module 72 and remote data repository 74 via communication links 76, 78 (e.g., via wired or wireless communication links), such that these remote modules 72, 74 are operatively coupled to each other and can be used as resources for the local processing and data module 70. In some embodiments, Figure 3 The world reconstruction component 316 can be implemented at least partially in the local data processing module 70. For example, the local data processing module 70 can be configured to execute computer-executable instructions to generate a physical world representation based at least partially on at least a portion of the data.
[0085] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 70 may include a single processor (e.g., a single-core or multi-core ARM processor), which limits the computational budget of module 70 but enables a smaller device. In some embodiments, the world reconstruction component 316 may use a computational budget smaller than that of a single ARM core to generate a physical world representation in real time in a non-predefined space, allowing access to the remaining computational budget of a single ARM core for other purposes, such as mesh extraction.
[0086] In some embodiments, the remote data repository 74 may include a digital data storage facility that can be accessed via the Internet or other network configurations in a “cloud” resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, thereby allowing fully autonomous use from the remote module. World reconstruction may, for example, be stored wholly or partially in this repository 74.
[0087] In some embodiments, the local data processing module 70 is operatively coupled to the battery 82. In some embodiments, the battery 82 is a removable power source, for example, above a counter battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during the non-operational period of the system 80 and a removable battery, such that the user 60 can operate the system 80 for a longer period of time without having to connect to a power source to charge the lithium-ion battery or have to shut down the system 80 to replace the battery.
[0088] Figure 5A The illustration depicts a user 30 wearing an AR display system that renders AR content as the user moves within a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 places the AR display system at location 34, and the AR display system records environmental information about the passable world relative to location 34 (e.g., digital representations of real objects in the physical world, which can be stored and updated as real objects in the physical world change), such as gestures related to mapping features or directional audio input. Location 34 is aggregated to data input 36 and processed at least by the passable world module 38, which can be done, for example, by... Figure 4 The processing is implemented on the remote processing module 72. In some embodiments, the traversable world module 38 may include a world reconstruction component 316.
[0089] The walkable world module 38 determines the location and manner in which AR content 40 can be placed in the physical world, as determined from data input 36. By presenting a representation of the physical world and the AR content via a user interface, the AR content is "placed" in the physical world, where the AR content is rendered as if interacting with objects in the physical world, and objects in the physical world are rendered as if the AR content is obstructing the user's view of these objects when appropriate. In some embodiments, the shape and position of AR content 40 can be determined by appropriately selecting portions of a fixed element 42 (e.g., a table) from a reconstruction (e.g., reconstruction 318). As an example, the fixed element could be a table, and the virtual content could be placed so that it appears to be on that table. In some embodiments, the AR content can be placed within a structure in a field of view 44, which could be the current field of view or an estimated future field of view. In some embodiments, the AR content can be placed relative to a mapped mesh model 46 of the physical world.
[0090] As depicted, fixed element 42 serves as a proxy for any fixed element in the physical world. It can be stored in the walkable world module 38, allowing user 30 to perceive content on fixed element 42 without the system needing to map it to fixed element 42 every time user 30 sees it. Fixed element 42 can therefore be a mapped mesh model from a previous modeling session, or determined by an individual user, but both are stored in the walkable world module 38 for future reference by multiple users. Thus, the walkable world module 38 can recognize environment 32 from previously mapped environments and display AR content without requiring user 30's device to first map environment 32, saving computation time and cycles and avoiding any delays in rendering AR content.
[0091] A physical world-mapping mesh model 46 can be created by the AR display system, and appropriate surfaces and measurements for interacting with and displaying AR content 40 can be mapped and stored in the accessible world module 38 for future retrieval by user 30 or other users without remapping or modeling. In some embodiments, data input 36 is input such as geolocation, user ID, and current activity to indicate to the accessible world module 38 which of one or more fixed elements 42 is available, which AR content 40 was last placed on a fixed element 42, and whether that same content is displayed (the AR content is "persistent" regardless of whether the user views a particular accessible world model).
[0092] Even in embodiments where objects are considered fixed, the walkable world module 38 can be updated periodically to account for the possibility of changes in the physical world. The model of fixed objects can be updated at a very low frequency. Other objects in the physical world may be moving or not considered fixed. To render a realistic AR scene, the AR system can update the positions of these non-fixed objects at a much higher frequency than it would be for updating fixed objects. To accurately track all objects in the physical world, the AR system can extract information from multiple sensors, including one or more image sensors.
[0093] Figure 5BThis is a schematic diagram of the viewing optical assembly 48 and its associated components. In some embodiments, two eye-tracking cameras 50, pointed at the user's eye 49, detect measurements of the user's eye 49, such as eye shape, eyelid occlusion, pupil direction, and flickering on the user's eye 49. In some embodiments, one of the sensors may be a depth sensor 51, such as a time-of-flight sensor, that emits signals into the world and detects reflections of these signals from nearby objects to determine the distance to a given object. For example, a depth sensor can quickly determine whether an object has entered the user's field of vision due to the movement of these objects or due to a change in the user's posture. However, alternatively or additionally, information about the position of an object in the user's field of vision may be collected using other sensors. For example, depth information may be obtained from a stereoscopic image sensor or an all-light sensor.
[0094] In some embodiments, the world camera 52 records a view larger than its periphery to map the environment 32 and detect input that may affect the AR content. In some embodiments, the world camera 52 and / or camera 53 may be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. Camera 53 may also capture images of the physical world within the user's field of view at specific times. Pixels of the frame-based image sensor may also be repeatedly sampled, even if their values remain unchanged. Each of the world camera 52, camera 53, and depth sensor 51 has a corresponding field of view 54, 55, and 56 to collect data from and record the physical world scene, for example... Figure 5A The physical world environment depicted in the text 32.
[0095] The inertial measurement unit 57 can determine the motion and orientation of the viewing optical assembly 48. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 51 is operatively coupled to the eye-tracking camera 50 as confirmation of the measured accommodation relative to the actual distance the user's eyes 49 are looking at.
[0096] It should be understood that viewing optical component 48 may include Figure 5B The viewing optics 48 may include some of the components shown, and may include components that replace or are not included in the examples shown. In some embodiments, for example, the viewing optics 48 may include two world cameras 52 instead of four world cameras. Alternatively or additionally, cameras 52 and 53 are not required to capture visible light images of their entire field of view. The viewing optics 48 may include other types of components. In some embodiments, the viewing optics 48 may include one or more dynamic vision sensors (DVS) whose pixels are asynchronously responsive to relative changes in light intensity exceeding a threshold.
[0097] In some embodiments, the viewing optics 48 may not include a depth sensor 51 based on time-of-flight information. In some embodiments, for example, the viewing optics 48 may include one or more all-light cameras whose pixels can capture light intensity and the angle of incident light, from which depth information can be determined. For example, the all-light camera may include an image sensor covered with a transmission diffraction mask (TDM). Alternatively or additionally, the all-light camera may include an image sensor comprising angle-sensitive pixels and / or phase-detection autofocus pixels (PDAF) and / or microlens arrays (MLA). Such a sensor can serve as a source of depth information, replacing or supplementing the depth sensor 51.
[0098] It should also be understood that Figure 5B The configuration of the components is shown as an example. The viewing optics 48 may include components with any suitable configuration, which can be set to provide the user with the maximum field of view useful for a particular set of components. For example, if the viewing optics 48 has a world camera 52, the world camera may be placed in the central area of the viewing optics rather than on one side.
[0099] Information from sensors in the viewing optics component 48 can be coupled to one or more processors in the system. The processors can generate data that can be rendered to allow a user to perceive interaction with objects in the physical world. This rendering can be implemented in any suitable manner, including generating image data depicting physical and virtual objects. In other embodiments, physical and virtual content can be depicted in a scene by modulating the opacity of a display device through which a user views the physical world. The opacity can be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data may include only virtual content that can be modified so that, when viewed through a user interface, the virtual content is perceived by the user as realistically interacting with the physical world (e.g., clipping content to resolve occlusion). Regardless of how the content is presented to the user, a model of the physical world is required so that the characteristics of virtual objects that can be affected by physical objects, including the shape, position, motion, and visibility of the virtual objects, can be correctly calculated. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 318.
[0100] The model can be created based on data collected from sensors on a user's wearable device. However, in some embodiments, the model can be created based on data collected from multiple users, which can be aggregated on computing devices remotely from all users (and may be in the "cloud").
[0101] This can be at least partially achieved through the world reconstruction system (e.g., Figure 6 A more detailed description Figure 3The world reconstruction component 316 is used to create the model. The world reconstruction component 316 may include a sensing module 160 that generates, updates, and stores a representation of a portion of the physical world. In some embodiments, the sensing module 160 may represent a portion of the physical world within the reconstruction range of a sensor as a plurality of voxels. Each voxel may correspond to a 3D cube of a predetermined volume in the physical world and includes surface information indicating the presence of a surface within the volume represented by the voxel. Values may be assigned to voxels indicating whether their corresponding volume has been determined to contain a surface of a physical object, is determined to be empty, or has not yet been measured by a sensor and therefore its value is unknown. It should be understood that it is not necessary to explicitly store the values of voxels indicating that they are determined to be empty or unknown, as the values of voxels can be stored in computer memory in any suitable manner, including not storing information about voxels determined to be empty or unknown.
[0102] Figure 7A An example of a 3D space 100 discrete as voxels 102 is depicted. In some embodiments, the perception module 160 can identify objects of interest and set the volume of the voxels to capture features of the objects of interest and avoid redundant information. For example, the perception module 160 can be configured to identify large objects and surfaces, such as walls, ceilings, floors, and large furniture. Therefore, the volume of the voxels can be set to a relatively large size, such as 4 cm. 3 A cube.
[0103] The reconstruction of the physical world, including voxels, can be called a volumetric model. As sensors move through the physical world, information is created over time to build the volumetric model. This occurs when a user of a wearable device that includes sensors moves around. Figure 8A -F depicts an example of reconstructing the physical world as a volumetric model. In the example shown, the physical world includes... Figure 8A A portion 180 of the surface shown. Figure 8A In the first position, the sensor 182 may have a field of view 184, within which a portion 180 of the surface is visible.
[0104] Sensor 182 can be of any suitable type, such as a depth sensor. However, depth data can be obtained from an image sensor or otherwise. Sensing module 160 can receive data from sensor 182 and then... Figure 8B The values of multiple voxels 186 are set to represent portions 180 of the surface visible in the field of view 184 by the sensor 182.
[0105] exist Figure 8C In this configuration, sensor 182 can be moved to a second position and has a field of view 188. For example... Figure 8DAs shown, another set of voxels becomes visible, and the values of these voxels can be set to indicate the location where the surface has entered the field of view 188 of sensor 182. The values of these voxels can be added to the volume model used for the surface.
[0106] exist Figure 8E In this embodiment, sensor 182 can be further moved to a third position and has a field of view 190. In the example shown, an additional portion of the surface becomes visible in the field of view 190. Figure 8F As shown, another set of voxels can become visible, and the values of these voxels can be set to indicate the location of the portion of the surface that has entered the field of view 190 of sensor 182. The values of these voxels can be added to the volumetric model used for the surface. Figure 6 As shown, this information can be stored as volume information 162a as part of the persistent world. Information about the surface, such as color or texture, can also be stored. Such information can be stored, for example, as volume metadata 162b.
[0107] In addition to generating information for the representation of the persistent world, the perception module 160 can also identify and output indications of changes in the area surrounding the user of the AR system. Such indications of change can trigger updates to volumetric data stored as part of the persistent world, or trigger other functions, such as triggering the triggering component 304 to update the AR content.
[0108] In some embodiments, the sensing module 160 may identify changes based on a signed distance function (SDF) model. The sensing module 160 may be configured to receive sensor data such as depth map 160a and head pose 160b, and then fuse the sensor data into an SDF model 160c. Depth map 160a may directly provide SDF information, and the image may be processed to obtain SDF information. SDF information represents the distance to the sensor used to capture this information. Since those sensors may be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and therefore from the user's perspective. Head pose 160b enables the SDF information to be correlated with voxels in the physical world.
[0109] Back Figure 6 In some embodiments, the sensing module 160 may generate, update, and store a representation of a portion of the physical world within the sensing range. The sensing range may be determined at least in part based on the sensor's reconstructed range, which may be determined at least in part based on the limits of the sensor's observation range. As a specific example, an active depth sensor operating with active IR pulses can reliably operate over a distance range, thereby creating the sensor's observation range, which can range from a few centimeters or tens of centimeters to several meters.
[0110] Figure 7B The reconstruction range relative to the sensor 104 with viewpoint 106 is depicted. A reconstruction of the 3D space within viewpoint 106 can be constructed based on data captured by the sensor 104. In the example shown, the observation range of the sensor 104 is 40 cm to 5 m. In some embodiments, the reconstruction range of the sensor can be determined to be smaller than the observation range of the sensor, because sensor output near its observation limit can be noisier, incomplete, and inaccurate. For example, in the 40 cm to 5 m example shown, the corresponding reconstruction range could be set from 1 to 3 m, and data collected by the sensor indicating surfaces outside this range could be omitted.
[0111] In some embodiments, the sensing range may be greater than the sensor's reconstructed range. If component 164, which uses data about the physical world, requires data about regions within the sensing range that are outside the portion of the physical world currently within the reconstructed range, this information can be provided from the persistent world 162. Accordingly, information about the physical world can be easily accessed through querying. In some embodiments, an API may be provided to respond to such queries, thereby providing information about the user's current sensing range. Such a technique reduces the time required to access existing reconstructions and provides an improved user experience.
[0112] In some embodiments, the sensing range can be a 3D space corresponding to a bounding box centered around the user's location. As the user moves, portions of the physical world within the sensing range that can be queried via component 164 can move with the user. Figure 7C A bounding box 110 centered at position 112 is depicted. It should be understood that the size of the bounding box 110 is set to reasonably expand the observation range surrounding the sensor, as the user cannot move at unreasonable speeds. In the example shown, the observation limit of the sensor worn by the user is 5m. The bounding box 110 is set to 20m. 3 A cube.
[0113] return Figure 6 The world reconstruction component 316 may include additional modules that can interact with the perception module 160. In some embodiments, the persistent world module 162 may receive a representation of the physical world based on data acquired by the perception module 160. The persistent world module 162 may also include representations of the physical world in various formats. For example, it may store volume metadata 162b such as voxels, as well as meshes 162c and planes 162d. In some embodiments, other information, such as depth maps, may be stored.
[0114] In some embodiments, the perception module 160 may include modules that generate representations for the physical world in various formats, including, for example, mesh 160d, planes, and semantics 160e. These modules may generate representations based on data within the perception range of one or more sensors at the time of representation generation, as well as data captured in previous time and information in the persistent world 162. In some embodiments, these components may operate on depth information captured using a depth sensor. However, the AR system may include a vision sensor and may generate such representations by analyzing monocular or binocular visual information.
[0115] In some embodiments, these modules can operate on regions of the physical world. When the sensing module 160 detects changes in other sub-regions of the physical world, those modules can be triggered to update the sub-regions of the physical world. For example, such changes can be detected by detecting new surfaces or other criteria in the SDF model 160c (e.g., changing the values of a sufficient number of voxels representing that sub-region).
[0116] The world reconstruction component 316 may include a component 164 capable of receiving a representation of the physical world from the perception module 160. Information about the physical world may be pulled by these components based on, for example, a usage request from an application. In some embodiments, information may be pushed to the usage component, for example, via indication of changes in a pre-identified area or changes in the representation of the physical world within the perception range. Component 164 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.
[0117] In response to a query from component 164, perception module 160 may send a representation of the physical world in one or more formats. For example, when component 164 indicates that the use is for visual occlusion or physical-based interaction, perception module 160 may send a representation of a surface. When component 164 indicates that the use is for environmental reasoning, perception module 160 may send a mesh, plane, and semantics of the physical world.
[0118] In some embodiments, the perception module 160 may include a component that formats information to provide component 164. An example of such a component may be a ray projection component 160f. Using a component (e.g., component 164), information about the physical world can be queried from a specific viewpoint, for example. The ray projection component 160f can be selected from one or more representations of physical world data within the field of view from that viewpoint.
[0119] As should be understood from the preceding description, the perception module 160, or another component of the AR system, can process data to create a 3D representation of a portion of the physical world. The amount of data to be processed can be reduced by at least partially selecting portions of the 3D reconstructed volume based on camera frustum and / or depth images, extracting and preserving planar data, capturing, preserving, and updating 3D reconstructed data in blocks that allow for local updates while maintaining neighbor consistency, providing occlusion data (wherein the occlusion data is derived from a combination of one or more depth data sources) to applications generating such scenes, and / or performing multi-level mesh simplification.
[0120] World reconstruction systems can integrate sensor data over time from multiple viewpoints in the physical world. As devices including sensors move, the sensor poses (e.g., position and orientation) can be tracked. Since the sensor frame poses and how they relate to other poses are known, each of these multiple viewpoints in the physical world can be fused into a single, combined reconstruction. By using spatial and temporal averaging (i.e., averaging the data from multiple viewpoints over time), the reconstruction can be more complete and less noisy than the original sensor data.
[0121] Reconstruction can contain data of varying levels of complexity, including, for example, raw data (e.g., real-time depth data), fused volumetric data (e.g., voxels), and computational data (e.g., meshes).
[0122] In some embodiments, AR and MR systems represent 3D scenes with a regular voxel grid, where each voxel may contain a signed distance field (SDF) value. The SDF value describes whether the voxel is located inside or outside a surface in the scene to be reconstructed, and the distance from the voxel to that surface. Calculating 3D reconstruction data to represent the volume required for the scene requires significant memory and processing power. These requirements increase cubically for scenes representing large spaces, as the number of variables required for 3D reconstruction increases with the number of depth images processed.
[0123] This document describes an efficient method for reducing processing. According to some embodiments, a scene may be represented by one or more bricks. Each brick may include multiple voxels. The bricks processed to generate a 3D reconstruction of the scene can be selected by choosing a set of bricks representing the scene based on a frustum derived from the field of view (FOV) of an image sensor and / or a depth image (or “depth map”) of the scene created using a depth sensor.
[0124] A depth image may have one or more pixels, each representing a distance to a surface in the scene. These distances may be related to the position relative to the image sensor, allowing the data output from the image sensor to be processed selectively. Image data may be processed for those bricks representing the portion of the 3D scene that contains the surface visible from the image sensor's viewpoint (or "viewpoint"). Processing of some or all of the remaining bricks may be omitted. In this way, the selected bricks may be those that could contain new information, obtained by picking bricks whose output from the image sensor is unlikely to provide useful information about them. The data output from the image sensor is unlikely to provide useful information about bricks that are closer to or farther from the image sensor than the surface indicated by the depth map, because these bricks are empty spaces or behind the surface and therefore not depicted in the image from the image sensor.
[0125] Figure 9 A cross-sectional view of scene 400 along a plane parallel to the y and z coordinates is shown. The XR system can represent scene 400 using a voxel grid 504. A conventional XR system can update each voxel of the voxel grid based on each new depth image captured by sensor 406, which can be an image sensor or a depth sensor, so that the 3D reconstruction generated from the voxel grid can reflect changes in the scene. Updating in this way consumes significant computational resources and also introduces artifacts at the output of the XR system due to time delays caused by computationally intensive tasks.
[0126] This describes a technique for providing accurate 3D reconstruction data with low computational resource utilization, for example, by selecting portions of a voxel grid 504 based at least in part on a camera frustum 404 of an image sensor 406 and / or a depth image captured by the image sensor.
[0127] In the example shown, image sensor 406 captures a depth image (not shown) of surface 402 including scene 400. This depth image can be stored in computer memory in any convenient manner that captures the distance between a reference point and the surface in scene 400. In some embodiments, the depth image can be represented as values in a plane parallel to the x-axis and y-axis, such as... Figure 9As shown, the reference point is the origin of the coordinate system. Positions in the XY plane can correspond to directions relative to the reference point, and the values at those pixel positions can indicate the distance from the reference point to the nearest surface in the direction indicated by the coordinates in the plane. Such a depth image can comprise a grid of pixels (not shown) in a plane parallel to the x and y axes. Each pixel can indicate the distance from image sensor 406 to surface 402 in a specific direction. In some embodiments, the depth sensor may fail to measure the distance to the surface in a specific direction. This can occur, for example, if the surface is beyond the range of image sensor 406. In some embodiments, the depth sensor can be an active depth sensor that measures distance based on reflected energy, but the surface may not reflect enough energy for accurate measurement. Therefore, in some embodiments, the depth image may have “holes” where pixels are not assigned values.
[0128] In some embodiments, the reference point of the depth image can vary. Such a configuration allows the depth image to represent a surface across the entire 3D scene, rather than being limited to a portion with a predetermined and finite angular range relative to a specific reference point. In such embodiments, the depth image may indicate the distance to the surface as the image sensor 406 moves through six degrees of freedom (6DOF). In these embodiments, the depth image may include a set of pixels for each of a plurality of reference points. In these embodiments, a portion of the depth image may be selected based on a “camera pose,” which refers to the direction and / or orientation in which the image sensor 406 is pointed when capturing image data.
[0129] Image sensor 406 may have a field of view (FOV), which may be represented by camera frustum 404. In some embodiments, by assuming a maximum depth 410 and / or a minimum depth 412 that image sensor 406 can provide, the depicted infinite camera frustum can be reduced to a finite 3D trapezoidal prism 408. 3D trapezoidal prism 408 may be a convex polyhedron defined at six planes.
[0130] In some embodiments, one or more voxels 504 may be grouped into bricks 502. Figure 10A A portion 500 of scene 400 is shown, which includes eight bricks 502. Figure 10B It shows 8 3 Example brick 502 of voxel 504. (Reference) Figure 9 Scene 400 may include one or more bricks, in Figure 4 The view shown illustrates sixteen of them. Each brick can be identified by brick identifiers such as
[0000] -
[0015] .
[0131] Geometric shape extraction system
[0132] In some embodiments, the geometry extraction system can extract geometry while scanning a scene with a camera and / or sensors, allowing for fast and efficient extraction that can adapt to dynamic environmental changes. In some embodiments, the geometry extraction system can retain the extracted geometry in local and / or remote memory. The retained geometry may have a unique identifier, allowing the retained geometry to be shared, for example, across different timestamps and / or different queries from different applications. In some embodiments, the geometry extraction system can support different representations of the geometry based on individual queries. (See below...) Figure 11-19 In the description, a plane is used as an exemplary geometric shape. It should be understood that the geometry extraction system can detect other geometric shapes, in place of a plane or other than a plane, for subsequent processing, including, for example, cylinders, cubes, lines, corners, or semantics such as glass surfaces or holes. In some embodiments, the principles described herein with respect to geometry extraction can be applied to object extraction, etc.
[0133] Figure 11 A planar extraction system 1300 according to some embodiments is illustrated. The planar extraction system 1300 may include a depth fusion 1304 that can receive multiple depth maps 1302. The multiple depth maps 1302 may be created by one or more users wearing depth sensors and / or downloaded from local / remote memory. The multiple depth maps 1302 may represent multiple views of the same surface. Differences may exist between the multiple depth maps, which can be reconciled by the depth fusion 1304.
[0134] In some embodiments, deep fusion 1304 can generate SDF 1306. Mesh tiles 1308 can be generated, for example, in the corresponding tiles (e.g., Figure 13 The planar extraction 1310 extracts planes from SDF 1306 by applying the marching cube algorithm to the bricks
[0000] to
[0015] in the grid. Plane extraction 1310 can detect planes in the grid bricks 1308 and extract planes at least partially based on the grid bricks 1308. Plane extraction 1310 can also extract facets of each brick at least partially based on the corresponding grid bricks. The facet mesh may include vertices in the mesh but not the edges connecting adjacent vertices, making storing facets consume less storage space than storing the mesh. The planar data repository 1312 can hold the extracted planes and facets.
[0135] In some embodiments, an XR application or other component 164 of an XR system may request and obtain planes from a plane data repository 1312 via a plane query 1314, which may be sent via an application programming interface (API). For example, an application may send information about its location to a plane extraction system 1300 and request all nearby planes (e.g., within a five-meter radius). The plane extraction system 1300 may then search its plane data repository 1312 and send the selected planes to the application. The plane query 1314 may include information such as where the application needs the planes, what kind of planes the application needs, and / or how the planes should look (e.g., horizontal, vertical, or angled, which can be determined by examining the primitive normals of the planes in the plane data repository).
[0136] Figure 12 A portion 1400 of a plane extraction system 1300 according to some embodiments is shown, illustrating details regarding plane extraction 1310. Plane extraction 1310 may include dividing each grid tile 1308 into sub-tiles 1402. Plane detection 1404 may be performed on each sub-tile 1402. For example, plane detection 1404 may: compare the original normals of each grid triangle in the sub-tile; merge those grid triangles with an original normal difference less than a predetermined threshold into a single grid triangle; and identify grid triangles with an area greater than a predetermined area value as planes.
[0137] Figure 13 This is a schematic diagram illustrating a scene 1500 represented by bricks
[0000] to
[0015] including voxels according to some embodiments, and exemplary planar data including a brick plane 1502, a global plane 1504, and a facet 1506 in the scene. Figure 13 The image shows a brick divided into four sub-bricks 1508
[0011] . It should be understood that the grid bricks can be divided into any suitable number of sub-bricks. The granularity of the plane detected by the plane detector 1404 can be determined by the size of the sub-bricks, and the size of the bricks can be determined by the granularity of the local / remote storage of the 3D reconstruction data.
[0138] Back Figure 12 Plane detection 1404 can determine the brick plane (e.g., brick plane 1502) of each grid brick based at least in part on the detection plane for each sub-brick in the grid brick. Plane detection 1404 can also determine a global plane that extends beyond one brick (e.g., global plane 1504).
[0139] In some embodiments, plane extraction 1310 may include plane update 1406, which may update existing brick planes and / or global planes stored in the plane data repository 1312 based at least in part on the planes detected by plane detection 1404. Plane update 1406 may include adding additional brick planes, removing some existing brick planes, and / or replacing some existing brick planes with brick planes detected by plane detection 1404 that correspond to the same bricks, so that real-time changes in the scenario are preserved in the plane data repository 1312. Plane update 1406 may also include aggregating brick planes detected by plane detection 1404 into existing global planes, for example, when a brick plane is detected to be adjacent to an existing global plane.
[0140] In some embodiments, plane extraction 1310 may further include plane merging and splitting 1408. For example, plane merging can combine multiple global planes into a large global plane when a brick plane is added and connects two global planes. Plane splitting can split a global plane into multiple global planes, for example, when a brick plane in the middle of a global plane is removed.
[0141] Figure 14 A data structure in a planar data repository 1312 according to some embodiments is shown. A global plane 1614, indexed by plane ID 1612, can be at the highest level of this data structure. Each global plane 1614 may include multiple brick planes and facets of bricks adjacent to the corresponding global plane, such that a brick plane can be reserved for each brick, and the global plane can be accurately rendered when the edges of the global plane are not qualified as brick planes for the corresponding bricks. In some embodiments, facets of bricks adjacent to the global plane, rather than facets of all bricks in the scene, are reserved, as this is sufficient to accurately render the global plane. For example, as... Figure 13 As shown, global plane 1504 extends across bricks
[0008] to
[0010] and
[0006] . Brick
[0006] has brick plane 1502, which is not part of global plane 1504. Using data structures in plane data repository 1312, when a plane query requests global plane 1504, face elements of bricks
[0006] and
[0012] are examined to determine whether global plane 1504 extends into bricks
[0006] and
[0012] . In the example shown, face element 1506 indicates that global plane 1504 extends into brick
[0006] .
[0142] Back Figure 14Global plane 1614 can be bidirectionally associated with the corresponding brick plane 1610. A brick can be identified by brick ID 1602. Bricks can be divided into planar bricks 1604 that include at least one plane and non-planar bricks 1606 that do not include a plane. The face elements of both planar and non-planar bricks can be retained, depending on whether the brick is adjacent to the global plane and not on whether the brick includes a plane. It should be understood that while the XR system is viewing the scene, planes can be continuously retained in the plane data repository 1312 regardless of whether a plane query 1314 exists.
[0143] Figure 15 A planar geometry extraction 1702, illustrating the extraction of a plane for use by an application when the application sends a planar query 1314 to a planar data repository 1312, is illustrated according to some embodiments. The planar geometry extraction 1702 may be implemented as an API. The planar query 1314 may indicate a requested planar geometry representation, such as an outer rectangular plane, an inner rectangular plane, or a polygonal plane. Based on the planar query 1314, a planar search 1704 may search for and obtain planar data in the planar data repository 1312.
[0144] In some embodiments, rasterization from planar cover point 1706 can generate planar cover points. Figure 16A An example is shown. There are four bricks
[0000] -
[0003] , each with a brick plane 1802. Plane overlay points 1806 (or "rasterized points") are generated by projecting the boundary points of the brick planes onto a global plane 1804.
[0145] Back Figure 15 The rasterization from planar overlay point 1706 can also generate a rasterized planar mask from that planar overlay point. Based on the planar geometry representation requested by planar query 1314, the inner rectangle planar representation, outer rectangle planar representation, and polygon planar representation can be extracted respectively via inner rectangle extraction 1708, outer rectangle extraction 1710, and polygon extraction 1712. In some embodiments, the application can receive the requested planar geometry representation within milliseconds of sending the planar query.
[0146] Figure 16BAn exemplary rasterized planar mask 1814 is shown. Various planar geometric representations can be generated from the rasterized planar mask. In the example shown, a polygon 1812 is generated by connecting some planar overlay points of the rasterized planar mask such that none of the planar overlay points in the mask are outside the polygon. An outer rectangle 1808 is generated such that the outer rectangle 1808 is the smallest rectangle surrounding the rasterized planar mask 1814. The inner rectangle 1810 is generated by the following operations: assigning “1” to bricks with two planar coverage points and assigning “0” to bricks without two planar coverage points to form a raster grid, identifying groups of bricks marked with “1” and aligned on lines parallel to the edges of the bricks (e.g., bricks
[0001] ,
[0005] ,
[0009] and
[0013] as one group, and bricks
[0013] -
[0015] as another group), and generating an inner rectangle for each identified group such that the inner rectangle is the smallest rectangle surrounding the corresponding group.
[0147] Figure 17 A mesh for scene 1900 is shown according to some embodiments. Figure 18A -C illustrates a scenario 1900 represented by an outer rectangular plane, an inner rectangular plane, and a polygonal plane, respectively, according to some embodiments.
[0148] Figure 19 The diagram shows a less noisy 3D representation of scene 1900, which is based on extracted planar data (e.g. Figure 18A The plane shown in -C) will Figure 17 The mesh shown is obtained by planarizing it.
[0149] Environmental Reconstruction
[0150] In some embodiments, a portion of the environment, such as an indoor environment, may be represented by a set of corner points, with or without metadata about those points or surfaces they define. This can be achieved through the sensing module 160. Figure 6 This can be implemented using any other suitable component to simply represent a portion of the environment as described herein. This information describing the representation can be stored as part of a persistent world model 162, rather than as a dense mesh or other representation as described above, or as a supplement to such a dense mesh or other representation. Alternatively or additionally, those corner points can be converted to a format that mimics a representation calculated using other techniques as described above. As a concrete example, a representation of the environment as a set of corner points can be converted to a simple mesh and stored as mesh 162c. Alternatively or additionally, each identified wall can be stored as plane 162d.
[0151] When component 164 needs information about a portion of the physical world that stores such a representation, it can obtain the representation by invoking the perception module 160 or by accessing the persistent world model 162 in any other suitable manner. In some embodiments, the representation can be provided whenever information about that portion of the physical world is requested. In some embodiments, the component 164 requesting information can specify whether a simplified representation is appropriate or requested. For example, an application can invoke the persistent world model 162 using parameters indicating whether a simplified representation of the physical world is appropriate or requested.
[0152] Corner points can be determined based on surface planes representing surfaces in the environment (e.g., walls, floors, and / or ceilings). These surface planes can be derived from data collected using sensors on wearable or other components within an XR system. The data can represent plane segments that correspond to portions of walls or other surfaces detected by processing image, distance, or other sensor data as described above.
[0153] In some embodiments, an environment, such as an indoor environment, can be reconstructed by identifying surface planes of the environment based on planar segments. In some embodiments, a planar segment can be represented by three or more points defining a flat surface. A planar segment normal can be calculated for the planar segment. As a specific example, four points can be used to define a quadrilateral in space that fits within or surrounds the planar segment. The planar segment normal can indicate the orientation of the planar segment. The planar segment normal can be associated with a value indicating the region of the planar segment.
[0154] In some embodiments, planar fragments can be obtained from, for example, stored data from a persistent world 162. Alternatively or additionally, functionality provided by a software development kit (SDK) provided with a wearable device can process sensor data to generate planar fragments. In some embodiments, planar fragments can be obtained by scanning a portion of the environment using an AR device, which can be used as an “initial scan.” In some embodiments, the length of the initial scan can be predetermined, for example, a few minutes or seconds, to obtain enough planar fragments to reconstruct the environment. In some embodiments, the length of the initial scan can be dynamic, ending when enough planar fragments are obtained to reconstruct the environment.
[0155] In some embodiments, the planar fragment may originate from one or more brick planes. In some embodiments, the shape of the planar fragment may be random, depending on the depth and / or image information captured by the AR device within a predetermined time period. In some embodiments, the shape of the planar fragment may be the same as or different from another planar fragment.
[0156] In some embodiments, the AR system can be configured to acquire multiple planar segments sufficient to identify all surface planes of the environment. In some embodiments, the number of planar segments may correspond to the number of surface planes in the environment. For example, for a room with a ceiling, a floor, and four walls, when the distance between the ceiling and the floor is known, four planar segments, each corresponding to one of the four walls, may be sufficient to reconstruct the room. In this exemplary scenario, the AR system can set each planar segment as a surface plane.
[0157] In some embodiments, the initial scan may last for several minutes or seconds. In some embodiments, the number of planar fragments obtained may exceed the number of planar surfaces in the environment. A surface plane filter may be configured to process the obtained planar fragments of surface planes for the environment by, for example, filtering out unnecessary planar fragments and combining multiple planar fragments into one. In some embodiments, planar fragments may be processed to form groups of planar fragments that may represent different portions of the same surface. Surface planes that can be modeled as an infinite range of surfaces can be derived from each group and used to determine the intersections of surfaces defining the environment, thereby identifying corner points. In some embodiments, a surface plane may be represented by three or more points that define the group of planar fragments from which the surface plane is derived. Surface plane normals can be calculated based on the points representing the surface planes. Surface plane normals may indicate the orientation of the planes from which the surface plane extends.
[0158] In some embodiments, corner points of the environment can be inferred by considering the first and second adjacent surface planes of the environment as extending to infinity and identifying the boundary lines where these planes intersect. The boundary lines can then intersect with a third surface plane, and in some embodiments, with a fourth surface plane, to identify the endpoints of the boundary lines. The first and second surface planes can be selected as adjacent based on one or more criteria, such as whether the planes intersect without penetrating other surface planes, whether the planes intersect at a distance from the user of the wearable device that corresponds to an estimated maximum size of the room, or the proximity of the centroids of the planar segments used to identify each surface plane. The first and second surface planes can be selected from those planes having a generally vertical orientation, such that they represent walls defining portions of the environment. Criteria, such as a normal within a certain threshold angle of 0 degrees, can be used to select vertical planes. The third and fourth surface planes can be selected from planes having a horizontal orientation, such as a normal within a certain threshold angle of 90 degrees. In some embodiments, the threshold angle may differ for horizontal and vertical planes. For example, the threshold angle may be larger for horizontal planes to identify planes representing, for example, a sloping ceiling. In some embodiments, the SDK equipped with a wearable device can provide a surface plane based on a user request, for example, indicating the type of surface plane requested (including, for example, a ceiling plane, a floor plane, and a wall plane).
[0159] Regardless of the method chosen for processing the surface planes, combinations of adjacent vertical planes can be processed to identify additional boundary lines, which can then be further intersected with one or more horizontal planes to identify corner points. The AR system can iterate this process until a complete set of corner points in the environment is inferred / determined. When the identified corner points define an enclosed space, the complete set can be determined. If the information obtained during the initial scan is insufficient to form a closure, additional scans can be performed.
[0160] Figure 20 An exemplary method 2000 for operating an AR system to generate a 3D reconstruction of an environment, according to some embodiments, is shown. Combined with... Figures 20 to 2 The method described in 9 can be executed in one or more processors of an XR system. Method 2000 can begin by extracting a planar fragment of the environment (Action 2002).
[0161] Extracting planar fragments Figure 24 The diagram illustrates an exemplary environment 2500. In this example, environment 2500 is room 2502 in an art museum. Room 2502 may be defined by a ceiling 2504, a floor 2506, and walls 2508. Walls 2508 may be used to display one or more works of art, such as a painting 2510 with a flat surface and an artwork 2512 with a curved surface. It should be understood that method 2000 is not limited to reconstructing a room such as room 2502, but can be applied to reconstructing any kind of environment, including but not limited to rooms with multiple ceilings, multiple floors, or vaulted ceilings.
[0162] Users 2516A-2516C within the environment 2500 can wear corresponding AR devices 2514A-2514C. AR devices 2514A-2514C can have corresponding fields of view 2518A-2518C. Within their respective fields of view, during an initial scan, the AR device can acquire a planar segment 2520. In some embodiments, the AR device can extract the planar segment from depth and / or image and / or other information captured by sensors in the AR device. The AR device can represent a planar segment having information defining its location, orientation, and the area it encompasses. For example, a planar segment can be represented by three or more points that define an area having a similar depth relative to the user or otherwise identified as a flat surface. In some embodiments, each point can be represented by its position in a coordinate system. In some embodiments, the coordinate system can be fixed, for example, using a reference in the environment as the origin, which makes it easier to share among users. In some embodiments, the coordinate system can be dynamic, for example, using the user's position as the origin, which simplifies data and reduces the computation required to display the environment to the user.
[0163] For each planar segment, a planar segment normal 2522 can be calculated. In some embodiments, the planar segment normal can be in vector format. The planar segment normal can indicate the orientation of the corresponding planar segment. The planar segment normal can also have a value representing the size of the area covered by the planar segment.
[0164] Back Figure 20 As shown, method 2000 may include identifying (action 2004) one or more surface planes, at least in part, based on the obtained planar fragments. In some embodiments, the surface planes may be specifically identified by a user (ad hoc). Figure 24 In the example shown, the planar segment extracted within FOV 2518B can be displayed on AR device 2514B and is visible to user 2516B. The user can, for example, point to the planar segment displayed on AR device 2514B to indicate to AR device 2514B that the planar segment normal 2524 should be set to the plane normal of the dominant surface of the wall.
[0165] In some embodiments, surface planes can be identified automatically or semi-automatically by a method, requiring little or no user input. Figure 21 An exemplary method 2100 for action 2004 is described according to some embodiments. Method 2100 may begin by dividing the obtained planar segments into different groups (action 2102) based on one or more criteria.
[0166] In some embodiments, planar segments can be grouped using user input. Figure 24 In the example shown, user 2516A can instruct AR device 2514A that one of the four planar segments displayed within FOV 2518A should be set to the first group. The AR device can be configured to interpret this user input as an instruction that the remaining planar segments within FOV 2518A should be set to a second group, different from the first group, allowing the user to point to a planar segment and then quickly move forward to scan different spaces in the environment. In some cases, partial user input can accelerate the execution of method 2100.
[0167] In the illustrated embodiments, at least one such criterion is based on the corresponding plane segment normal. In some embodiments, plane segments having plane segment normals within an error range can be grouped together. In some embodiments, the error range can be predetermined, for example, less than 30 degrees, less than 20 degrees, or less than 10 degrees. In other embodiments, the boundaries of the groups can be determined dynamically, for example, by using a clustering algorithm. As another example, plane segments can be grouped based on the distance between them. The distance between two plane segments can be measured, for example, at the centroid of the first plane segment, in a direction parallel to the normal of that plane segment. The distance from the centroid to the intersection point with the plane containing the second plane segment can be measured.
[0168] Alternatively or additionally, one or more criteria can be used to select or exclude planar segments from a group. For example, only planar segments larger than a threshold size can be processed. Alternatively or additionally, planar segments that are too large, such as those exceeding the possible dimensions of a room, can be excluded. Classification can be based on other criteria, such as the orientation of the planar segments, allowing only horizontal or vertical wall segments, or wall segments with orientations that may otherwise represent walls, floors, or ceilings, to be processed. Other characteristics, such as the color or texture of the wall segments, can be used alternatively or additionally. Similarly, distance from the user or other reference point can be used to assign planar segments to groups. For example, distance can be measured in a direction parallel to the normal of the planar segment, or using other criteria that group segments on the same surface together.
[0169] Regardless of how the planar fragments are selected and grouped, the groups can then be processed to derive surface planes from the planar fragments in each group. In action 2104, it can be determined whether a dominant planar fragment normal exists among the planar fragment normals for each group. In some embodiments, the dominant planar fragment normal can be determined by comparing the sizes of the planar fragments associated with the planar fragment normal, which can be indicated by a value representing the region of the planar fragment associated with the planar fragment normal. If a planar fragment normal is associated with a planar fragment that is significantly larger than the other planar fragments, the planar fragment normal associated with the significantly larger planar fragment can be selected as the dominant planar fragment normal.
[0170] The statistical distribution of the sizes of plane segments can be used to determine whether a principal plane normal exists. For example, if the largest plane segment exceeds the average size of other plane segments in a set by more than one standard deviation, it can be selected as the principal plane segment normal. Regardless of how the principal plane segment normal is determined, in action 2106, the principal plane segment normal can be set as the surface plane normal, which can then be used to define the surface plane representing the surface.
[0171] If no primary plane segment normal is detected as a result of the processing at action 2104, then at action 2108, the surface plane normal can be derived from the plane segment normals of the plane segments in the group. Figure 21 This illustrates, for example, that a weighted average of the normals of selected planar segments in the group can be set as the surface normal. The weights can be proportional to the size of the planar segments associated with the planar segment normals, such that the results may be affected by measuring surface regions with the same or similar orientations.
[0172] The selected plane segment normals can include all plane segment normals in the group, or a subset of all plane segment normals in the group selected by, for example, removing plane segment normals associated with plane segments that are less than a value threshold (e.g., multiple brick planes) and / or have orientations outside the threshold. This selection can be performed at any suitable time and can be performed, for example, as part of action 2108, or alternatively or additionally as part of grouping plane segments in action 2102.
[0173] In action 2110, method 2000 can derive a surface plane based on surface plane normals. When a primary plane segment normal is detected, the plane segment corresponding to the primary plane segment can be saved as a surface plane. When no primary plane segment normal is detected, the surface plane can be derived via action 2108. Regardless of how the surface plane normals are determined, a surface plane can be represented by three or more points, which can be considered as representing an infinite plane perpendicular to the surface plane normals (e.g., surface plane 2616 as shown in the figure). In some embodiments, the surface plane can include metadata, including, for example, where the group of plane segments is centered or how much surface area is detected in the plane segments within that group. In some embodiments, the surface plane can be a fragment of the surface of the environment (e.g., surface plane 2614 as shown in the figure), which can have the same format as the plane segments.
[0174] refer to Figure 25 It depicts a simplified schematic diagram 2600 of an environment 2500 according to some embodiments. Schematic diagram 2600 may include surface planes 2602-2620, which may be based on... Figure 24 The planar segment 2520 extracted from the image is used for identification. Figure 25 In the illustration, surface planes are shown as truncated so that they can be shown; however, it should be understood that surface planes may not have an associated area, as they can be considered infinite. As an example, surface plane 2610 can be identified at least in part based on a fragment of the plane captured by AR device 2514B worn by user 2516B. Figure 24In the example shown, a painting 2510 with a flat surface and an artwork 2512 with a curved surface are within the field of view 2518B. Planar segments extracted from the artwork can be removed before calculating the surface plane because the plane normal for a curved plane may be outside the error range. Planar segments extracted from the painting can also be removed before calculating the surface plane because the value of the plane normal for the painting may be significantly smaller than the value of a planar segment extracted from the wall behind it. As another example, surface planes 2614 and 2616 can be identified at least in part based on planar segments extracted by an AR device 2514A worn by user 2516A. The extracted planar segments can be divided into two groups for calculating surface planes 2614 and 2616, respectively.
[0175] Back Figure 20 Method 2000 may include inferring one or more corner points of the environment (Action 2006) based at least in part on surface planes. Figure 22 A method 2300 for action 2006 is described according to some embodiments. Method 2300 may begin by extending (action 2302) two wall planes (e.g., surface planes 2602 and 2606) to infinity. In some embodiments, the wall planes may be selected from those planes having a generally vertical orientation, such that they represent walls defining portions of an environment. In some embodiments, method 2300 may include tracing a path from one intersecting surface plane to another, e.g., in a direction defined by the wall planes. If the path forms a loop, method 2300 may determine that a closure is formed. If the path encounters an identified surface plane that intersects one plane but not the second plane, indicating a path interruption, method 2300 may determine that no closure is formed. In some embodiments, center-to-center distances between the wall planes (e.g., ...) may be calculated. Figure 26 (d) Method 2300 can begin with two wall planes having the minimum center-to-center distance.
[0176] Method 2300 may include obtaining (action 2304) the boundary line intersecting both wall planes. This line can be obtained by manipulating a computational representation of the geometry described herein. For example, as Figure 26 As shown, boundary line 2702 intersects two surface planes 2602 and 2606. In action 2306, one or more first corner points can be inferred by making the boundary line intersect the corresponding floor plane. For example, as Figure 27 As shown, the first corner point 2802 is deduced by intersecting the boundary line 2702 with the floor plane 2622. In action 2308, one or more second corner points can be deduced by intersecting the boundary line with the corresponding ceiling plane. Figure 27As shown, the second corner point 2804 is deduced by intersecting boundary line 2702 with ceiling plane 2620. Although the room 2500 shown has one ceiling and one floor, it should be understood that some rooms may have multiple ceilings or floors. For example, the boundary line at the transition between two floor planes may result in two first corner points. Similarly, the boundary line at the transition between two ceiling planes may result in two second corner points.
[0177] Reference Back Figure 20 Method 2000 may include determining whether the corner points inferred in action 2008 define an enclosed space, such that the set of points can be confirmed as an accurate representation of a room or interior environment. In some embodiments, the determination indicates closure if the boundary lines connecting the corner points can be connected to define surfaces that engage and define a closed volume. In some embodiments, the determination indicates closure if the wall plane defined by the corner points forms a loop.
[0178] In some embodiments, when the calculation indicates a closure has been formed, corner points can be saved in action 2010—for example, they can be stored in local memory and / or communicated via a computer network to be stored in remote memory. When the calculation result indicates that a closure has not yet been formed, some actions in actions 2002-2006 can be repeated multiple times until a closure is formed or a data error is detected. For example, in each iteration, additional data may be collected or accessed to use more planar fragments in action 2002, thereby increasing the chance of forming a closure. In some embodiments, an error may mean that the inferred corner point is not in a space that can be corrected to a simple representation. In some embodiments, in response to an error, the system can present a user interface providing the user with intervention options, such as indicating the presence of a surface plane at a location where the sensor did not detect a planar fragment, or indicating the absence of a surface plane at a location where the system incorrectly grouped some extracted planar fragments.
[0179] In the example shown, in action 2012, the AR system determines whether new planar fragments are needed. When it is determined that new planar fragments are needed, method 2000 may perform action 2002 to capture additional planar fragments. When it is determined that new planar fragments are not needed and instead, for example, the captured planar fragments should be regrouped, method 2000 may perform action 2004 to obtain additional surface planes and / or replace existing surface planes.
[0180] Figure 23A simplified schematic diagram illustrating a system 2400 configured to perform method 2000 according to some embodiments is depicted. System 2400 may include a cloud storage 2402 and AR devices 2514A, 2514B, and 2514C configured to communicate with the cloud storage 2402 and other AR devices. The first AR device 2514A may include a first local memory storing a first corner point inferred from a planar segment extracted by the first AR device. The second AR device 2514B may include a second local memory storing a second corner point inferred from a planar segment extracted by the second AR device. The third AR device 2514C may include a third local memory storing a third corner point inferred from a planar segment extracted by the third AR device.
[0181] In some embodiments, when an application running on the first AR device requires a 3D representation of environment 2500, the first AR device can retrieve second and third corner points from a cloud storage to construct a 3D representation of room 2500, for example... Figure 28 The mesh model 2900 shown is an example. Mesh models can have a variety of uses. For instance, mesh model 2900 can be used by a home contracting application running on an AR device to calculate the amount of paint needed to cover room 2500. Compared to traditional models obtained by scanning the entire room, mesh model 2900 can achieve this by partially scanning the room, while providing accurate and fast results.
[0182] In some embodiments, during a multi-user experience, AR devices can scan the environment together. AR devices can communicate with each other at identified corner points, thereby identifying all corner points of the environment in less time than a single user.
[0183] in conclusion
[0184] Therefore, several aspects of some embodiments have been described, and it should be understood that various changes, modifications and improvements will readily occur to those skilled in the art.
[0185] As an example, the embodiments are described in conjunction with an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein can be applied in MR environments or more generally in other XR and VR environments.
[0186] As another example, embodiments are described in conjunction with devices such as wearable devices. It should be understood that some or all of the technologies described herein can be implemented via networks (e.g., the cloud), discrete applications and / or devices, or any suitable combination of networks and discrete applications.
[0187] Such changes, modifications, and improvements are intended to be part of this disclosure and are intended to fall within the spirit and scope of this disclosure. Furthermore, while advantages of this disclosure have been indicated, it should be understood that not every embodiment of this disclosure will include every described advantage. Some embodiments may not implement any features described herein and in certain circumstances as advantageous. Therefore, the foregoing description and figures are by way of example only.
[0188] The embodiments described above in this disclosure can be implemented in any of a variety of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided in a single computer or distributed across multiple computers. Such a processor can be implemented as an integrated circuit having one or more processors within an integrated circuit component, including commercially available integrated circuit components known in the art as such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor can be implemented in custom circuitry such as ASICs or in semi-custom circuitry generated by configuring programmable logic devices. As another alternative, the processor can be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a particular example, some commercially available microprocessors have multiple cores, such that one or a subset of these cores can constitute a processor. However, a processor can be implemented using circuitry of any suitable format.
[0189] Furthermore, it should be understood that a computer can be embodied in any of a variety of forms, such as a rack-mounted computer, desktop computer, laptop computer, or tablet computer. Additionally, a computer can be embedded in a device that is not typically considered a computer but has appropriate processing capabilities, including a personal digital assistant (PDA), a smartphone, or any other suitable portable or fixed electronic device.
[0190] In addition, the computer may have one or more input and output devices. These devices can be used, in particular, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or displays for visual presentation output, and speakers or other sound-generating devices for auditory presentation output. Examples of input devices that can be used for a user interface include keyboards and pointing devices such as mice, touchpads, and digitizing tablets. As another example, the computer may receive input information via speech recognition or other audible formats. In the illustrated embodiments, the input / output devices are shown to be physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated into the same unit as the processor or other components of the computing device. For example, the keyboard may be implemented as a soft keyboard on a touchscreen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated via a wireless connection.
[0191] Such computers can be interconnected through one or more networks of any suitable form, including local area networks (LANs) or wide area networks (WANs) such as corporate networks or the Internet. Such networks can be based on any suitable technology and can operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0192] Furthermore, the various methods or processes outlined in this paper can be encoded as software that can be executed on one or more processors employing any of a variety of operating systems or platforms. Additionally, such software can be written using any of a variety of suitable programming languages and / or programming or scripting tools, and can also be compiled into executable machine language code or intermediate code that executes on a framework or virtual machine.
[0193] In this respect, this disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy disks, compressed optical discs (CDs), optical discs, digital video discs (DVDs), magnetic tape, flash memory, circuit configurations in field-programmable gate arrays or other semiconductor devices, or other tangible computer storage media) encoding one or more programs that, when executed on one or more computers or other processors, perform methods implementing the various embodiments of this disclosure discussed above. It is apparent from the foregoing examples that a computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. One or more such computer-readable storage media may be removable, such that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of this disclosure as described above. As used herein, the term "computer-readable storage medium" includes only computer-readable media that can be considered as an edifice (i.e., an article of manufacture) or a machine. In some embodiments, this disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagating signal.
[0194] In this document, the terms "program" or "software" are used in a general sense to refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement the various aspects of this disclosure as described above. Furthermore, it should be understood that, according to one aspect of this embodiment, one or more computer programs that perform the methods of this disclosure when executed do not need to reside on a single computer or processor, but can be distributed in a modular manner across multiple different computers or processors to implement the various aspects of this disclosure.
[0195] Computer-executable instructions can take many forms, such as program modules that are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. Generally, in various embodiments, the functionality of program modules can be combined or distributed as needed.
[0196] Furthermore, data structures can be stored in any suitable form on a computer-readable medium. For simplicity, a data structure can be shown as having fields related by position within the data structure. Such relationships can also be implemented by allocating storage devices for the position fields in a computer-readable medium that conveys the relationships between the fields. However, any suitable mechanism can be used to establish relationships between the information in the fields of a data structure, including other mechanisms that establish relationships between data elements using pointers, labels, or other methods.
[0197] The various aspects of this disclosure can be used individually, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and are therefore not limited to the details and arrangements of the components set forth in the foregoing description or shown in the accompanying drawings. For example, an aspect described in one embodiment can be combined in any way with aspects described in other embodiments.
[0198] Furthermore, this disclosure can be embodied as a method, examples of which have been provided. Actions performed as part of the method can be ordered in any suitable manner. Thus, embodiments can be constructed in which actions are performed in an order different from the order shown, which may include performing some actions simultaneously, even those shown as sequential actions in the illustrative embodiments.
[0199] The use of ordinal terms (e.g., "first", "second", "third", etc.) in claims to modify claim elements does not imply that one claim element has any priority, precedence, or order of action or method of execution relative to another claim. Rather, it is merely used as a label to distinguish one claim element with a certain name from another element with the same name (but using ordinal terms) to differentiate these claim elements.
[0200] Furthermore, the wording and terminology used herein are for descriptive purposes and should not be considered limiting. The use of “including” or “having,” “comprising,” “involving,” and variations thereof in this document is intended to cover the items listed thereafter and their equivalents, as well as additional items.
Claims
1. A portable electronic system, comprising: Sensors are configured to capture information about the physical world; A processor is configured to execute computer-executable instructions to compute a three-dimensional 3D representation of a portion of the physical world, at least in part, based on captured information about the physical world, wherein the computer-executable instructions include instructions for: Extract multiple planar segments from the information captured by the sensor; Multiple surface planes are identified, at least in part, based on the multiple planar segments; and Based at least in part on the plurality of surface planes, a plurality of corner points of the portion of the physical world are inferred such that the portion of the physical world is represented by the plurality of corner points.
2. The portable electronic system according to claim 1, wherein, The plurality of planar segments are associated with the orientation of the plane represented by each planar segment.
3. The portable electronic system according to claim 1, wherein, The computer-executable instructions also include instructions for using the corner points to construct a mesh model of the portion of the physical world.
4. The portable electronic system according to claim 1, wherein, The plurality of surface planes are identified, at least in part, based on input from at least a portion of the user wearing the portable electronic system.
5. The portable electronic system according to claim 1, comprising: A transceiver is configured to communicate with a device that provides remote storage over a computer network.
6. The portable electronic system according to claim 1, wherein, The processor implements a service configured to provide the application with a 3D representation of the portion of the physical world.
7. The portable electronic system according to claim 6, wherein, The service stores the corner point in local storage or transmits the corner point to cloud storage as the three-dimensional 3D representation of the part of the physical world.
8. The portable electronic system according to any one of claims 1 to 7, wherein, Identifying the plurality of surface planes includes: Determine whether a dominant plane segment normal exists within the group of plane segment normals of the plurality of plane segments; When the determination indicates the normal of the principal plane segment in the group, the principal plane segment normal is set as the surface plane normal; and When the determination indicates that there is no major plane segment normal in the group, the surface plane normal is calculated based on at least a portion of the plane segment normals in the group.
9. The portable electronic system according to claim 8, wherein, Calculating the surface plane normals includes: calculating a weighted average of at least a portion of the plane segment normals in the group.
10. The portable electronic system according to any one of claims 1 to 7, wherein, The plurality of corner points inferring the portion of the physical world include: Extend the first and second surface planes of the plurality of surface planes to infinity; and Obtain the boundary line that intersects with the first surface plane and the second surface plane.
11. The portable electronic system according to claim 10, wherein, The multiple corner points inferred from the aforementioned portion of the physical world also include: One of the plurality of corner points is inferred by making the boundary line intersect with the third surface plane.
12. At least one non-transitory computer-readable medium encoded with a plurality of computer-executable instructions, which, when executed by at least one processor, perform a method for providing a three-dimensional 3D representation of a portion of a physical world, wherein the portion of the physical world is represented by a plurality of corner points, the method comprising: Capture information about the portion of the physical world within the user's field of view (FOV); Extract multiple planar fragments from the captured information; Identify multiple surface planes from the plurality of planar segments; as well as Based on the intersection of the surface planes among the identified multiple surface planes, multiple corner points representing the portion of the physical world are calculated such that the portion of the physical world is represented by the multiple corner points.
13. The at least one non-transitory computer-readable medium according to claim 12, wherein, The plurality of planar segments are associated with the orientation of the plane represented by each planar segment.
14. The at least one non-transitory computer-readable medium according to claim 12, wherein, The method includes: Calculate whether the first set of multiple corner points form a closed loop.
15. The at least one non-transitory computer-readable medium according to claim 14, wherein, Calculating whether a closure is formed includes: determining whether the boundary lines joining the first plurality of corner points can be connected to define the surfaces that join and define the closed volume.
16. At least one non-transitory computer-readable medium according to any one of claims 12 to 15, wherein: The portion of the physical world is the first part of the physical world. The user mentioned is the first user. The plurality of corner points are the first plurality of corner points; as well as The method further includes: Receive the second plurality of corner points of the second part of the physical world from the second user; as well as The 3D representation of the physical world is provided at least in part based on the first plurality of corner points and the second plurality of corner points.
17. At least one non-transitory computer-readable medium according to any one of claims 12 to 15, wherein: The user mentioned is the first user, and The method further includes: The corner point, calculated based on the captured information about the portion of the physical world within the first user's field of vision, is transmitted via a computer network. The corner point transmitted is received at the XR device used by the second user; and Based on the received multiple corner points, the XR device renders information about the portion of the physical world to the second user.
18. At least one non-transitory computer-readable medium according to any one of claims 12 to 15, wherein, The method includes: Calculate metadata for the corner points, the metadata indicating the positional relationship between the corner points.
19. The at least one non-transitory computer-readable medium according to claim 18, wherein, The method includes: The corner point, including the corresponding metadata, is stored so that it can be retrieved by multiple users, including the user.
20. A method for operating an interreal system to reconstruct the physical world, the interreal system including a processor configured to process image information in communication with sensors worn by a user, the sensors generating depth information for various regions within the sensor's field of view, the method comprising: Extract multiple planar segments from the depth information; The extracted planar fragment is displayed to the user; Receive user input indicating multiple surface planes, each surface plane representing a surface that defines the physical world; as well as Based at least in part on the plurality of surface planes, a plurality of corner points of the physical world are calculated such that the portion of the physical world is represented by the plurality of corner points.
21. The method according to claim 20, wherein, The plurality of planar segments are associated with the orientation of the plane represented by each planar segment.
22. The method of claim 20, further comprising: Determine whether the plurality of corner points form a closed loop.
23. The method of claim 22, further comprising: When it is determined that the closure has been formed, the corner point is stored.
24. The method according to any one of claims 20 and 22 to 23, wherein: The portion of the physical world is the first part of the physical world. The user mentioned is the first user. The plurality of corner points are the first plurality of corner points; as well as The method further includes: Receive the second plurality of corner points of the second part of the physical world from the second user; as well as A 3D representation of the physical world is provided, at least in part, based on the first plurality of corner points and the second plurality of corner points.
25. The method according to any one of claims 20 and 22 to 23, wherein, The method includes: Calculate metadata for the corner points, the metadata indicating the positional relationship between the corner points.
Citation Information
Patent Citations
Structural modeling using depth sensors
US20150070387A1