Caching and Updating of Dense 3D Reconstruction Data

By selectively updating and simplifying the representation of 3D reconstruction data on mobile devices, combining the correlation between high information and the physical world, the real-time update problem of AR/MR/VR systems under limited resource conditions is solved, and fast and accurate 3D reconstruction data updates and efficient resource utilization are achieved.

CN114051629BActive Publication Date: 2025-07-04MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080044530.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-26
Filing Date
2020-05-21
Publication Date
2025-07-04
Estimated Expiration
2040-05-21

AI Technical Summary

Technical Problem

Existing AR/MR/VR systems require a large amount of computing resources when updating intensive 3D reconstruction data in real time, especially on mobile devices with limited computing power and storage space, resulting in problems of processing delays and excessive resource consumption.

Method used

By selectively updating and simplifying 3D reconstruction data representations using sensor data on mobile devices, combining the association of high-information with the physical world, block representation and multi-layer caching mechanisms are adopted to reduce computing burdens and improve data processing efficiency.

Benefits of technology

It realizes the rapid and accurate update of 3D reconstruction data under limited resource conditions, reduces computing and storage requirements, and improves the real-time rendering capability and user experience of AR/MR/VR systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114051629B_ABST
    Figure CN114051629B_ABST
Patent Text Reader

Abstract

A method is provided for effectively updating and managing the output of real-time or offline 3D reconstruction and scanning in a mobile device with limited resources and a connection to the Internet. The method shares and updates the same 3D reconstruction data in a single-user application or a multi-user application, providing fresh, accurate, and comprehensive 3D reconstruction data for various mobile XR applications. The method includes a block-based 3D data representation that allows local updates while maintaining neighborhood consistency, and a multi-level caching mechanism for efficient retrieval, prefetching, and storage of 3D data for XR applications. Height information for indoor environments that can be represented as building floors can be associated with a sparse representation and / or a dense representation of the physical world to improve the accuracy of positioning results and / or render virtual content more realistically.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to a cross-reality system that uses 3D world reconstruction to render scenes. Background Art

[0002] A computer may control a human user interface to create an X Reality (XR or cross-reality) environment, in which the computer generates some or all of the XR environment that is perceived by the user. These XR environments may be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments, where some or all of the XR environment may be generated by the computer partially using data that describes the environment. The data may describe, for example, virtual objects, which may be rendered in a manner that the user senses or perceives as part of the physical world and may interact with the virtual objects. Since the data is rendered and presented through a user interface device (e.g., a head-mounted display device), the user may experience these virtual objects. The data may be shown to the user, or may control the audio played to the user, or may control a tactile (or haptic) interface, so that the user can experience the touch sensation of the virtual objects that the user senses or perceives.

[0003] XR systems can be useful for many applications, spanning fields such as scientific visualization, medical training, engineering design and prototyping, teleoperation and telepresence, and personal entertainment. Compared with VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment of using XR systems and also opens the door to various applications that present information about how to change the reality of the physical world in an easy-to-understand manner. Summary of the Invention

[0004] Aspects of the present application relate to methods and apparatuses for caching and updating 3D reconstruction data. The inventors have recognized and understood techniques for caching and updating dense 3D reconstruction data in real time on devices with limited computing resources, such as mobile devices. These techniques may be used together, used separately, or used in any suitable combination.

[0005] Some embodiments relate to a portable electronic system. The portable electronic system includes: at least one sensor configured to capture three-dimensional (3D) information about an object in the physical world; a local memory storing computer-executable instructions; a transceiver configured to communicate with a remote memory via a computer network; and a processor communicatively coupled to the at least one sensor, the local memory, and the transceiver, wherein the processor executes the computer-executable instructions to provide a 3D representation of a portion of the physical world based at least in part on the 3D information about the object in the physical world. The 3D representation of the portion of the physical world includes a plurality of blocks, each block including an indication of an object in a region of a portion of the physical world at a corresponding point in time and a value indicating a floor. The computer-executable instructions include instructions for processing to identify a subset of the plurality of blocks corresponding to information captured by an incoming sensor.

[0006] In some embodiments, a value indicating the floor of the block is determined based at least in part on the 3D information about the object in the physical world.

[0007] In some embodiments, the at least one sensor includes an image sensor.

[0008] In some embodiments, at least a portion of the at least one sensor is configured to provide at least a portion of the information captured by the incoming sensor on which determining the position of the portable electronic system is at least partially based.

[0009] In some embodiments, the position of the portable electronic system has different degrees of granularity.

[0010] In some embodiments, the position of the portable electronic system includes a value indicating the floor of a corresponding block.

[0011] In some embodiments, the portable electronic system includes one or more of a magnetometer, an altimeter, and a GPS sensor configured to provide at least a portion of the information captured by the incoming sensor.

[0012] In some embodiments, a value indicating the floor of the block is determined based at least in part on data captured by one or more of a magnetometer, an altimeter, and a GPS sensor.

[0013] In some embodiments, the position of the portable electronic system is determined based at least in part on data captured by one or more of a magnetometer, an altimeter, and a GPS sensor.

[0014] In some embodiments, the position of the portable electronic system has different degrees of granularity.

[0015] Some embodiments relate to a method of operating a portable electronic system. The method includes: capturing sensor information about the physical world with at least one sensor; and processing the sensor information with at least one processor to: generate a representation of a portion of the physical world proximate to the portable electronic device; and associate height information with the portion of the physical world.

[0016] In some embodiments, the height information includes an identifier of a floor of a building.

[0017] In some embodiments, associating height information with the portion of the physical world includes: processing the output of a vision sensor to detect an indicator of the floor.

[0018] In some embodiments, associating height information with the portion of the physical world includes: tracking a change in floor.

[0019] In some embodiments, tracking a change in floor includes: processing a visual image to detect a traversal of a staircase or an escalator.

[0020] In some embodiments, the representation of the portion of the physical world is a sparse representation.

[0021] In some embodiments, the method includes: sending a location request based on (a) the representation of the portion of the physical world and (b) including the height information.

[0022] In some embodiments, the representation of the portion of the physical world is a first part of a dense representation.

[0023] In some embodiments, the dense representation includes a plurality of parts, the plurality of parts including the first part, each part including associated height information. The method further includes: performing occlusion processing based on the plurality of parts, the occlusion processing including: sorting surface information in a corresponding part based on the height information associated with each part of the plurality of parts.

[0024] Some embodiments relate to a non-transitory computer-readable medium having instructions stored thereon that, when executed on a processor, perform actions including: processing sensor information about the physical world to: generate a representation of a portion of the physical world proximate to the portable electronic device; and associate height information with the portion of the physical world.

[0025] The foregoing summary is provided by way of example and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in the various drawings is represented by the same numeral. For clarity, not every component will be labeled in every drawing. In the drawings:

[0027] Figure 1 is a schematic diagram showing an example of a simplified augmented reality (AR) scene according to some embodiments.

[0028] Figure 2 is a sketch of an exemplary simplified AR scene showing an exemplary world reconstruction use case including visual occlusion, physics-based interaction, and environmental reasoning according to some embodiments.

[0029] Figure 3 is a schematic diagram showing the data flow in an AR system according to some embodiments, the AR system being configured to provide an experience of interacting AR content with the physical world.

[0030] Figure 4 is a schematic diagram showing an example of an AR display system according to some embodiments.

[0031] Figure 5A is a schematic diagram showing an AR display system according to some embodiments that renders AR content as the user moves through a physical world environment while wearing it.

[0032] Figure 5B is a schematic diagram showing a viewing optical component and accompanying components according to some embodiments.

[0033] Figure 6 is a schematic diagram showing an AR system using a world reconstruction system according to some embodiments.

[0034] Figure 7A is a schematic diagram showing a 3D space discretized into voxels according to some embodiments.

[0035] Figure 7B is a schematic diagram showing the reconstruction range relative to a single viewpoint according to some embodiments.

[0036] Figure 7C is a schematic diagram showing the perception range of the reconstruction range at a single location according to some embodiments.

[0037] Figures 8A to 8F is a schematic diagram showing the reconstruction of a surface into a voxel model by an image sensor viewing the surface in the physical world from multiple positions and viewpoints according to some embodiments.

[0038] Figure 9Is a schematic diagram showing a scene represented by bricks including voxels, a surface in the scene, and a depth sensor that captures the surface in a depth image according to some embodiments.

[0039] Figure 10A Is a schematic diagram showing a 3D space represented by eight bricks.

[0040] Figure 10B Is a schematic diagram showing Figure 10A The voxel grid in the bricks of.

[0041] Figure 11 Is a schematic diagram showing a volume representation hierarchy according to some embodiments.

[0042] Figure 12 Is a flowchart showing a method of operating a computing system to generate a 3D reconstruction of a scene according to some embodiments.

[0043] Figure 13 Is a flowchart showing a method of selecting a portion of a plurality of bricks for a Figure 12 Depth sensor in according to some embodiments.

[0044] Figure 14 Is a flowchart showing a method of performing a Figure 13 Camera frustum acceptance test in according to some embodiments.

[0045] Figure 15 Is a flowchart showing a method of selecting a portion of a first plurality of bricks for a Figure 12 Depth image in according to some embodiments.

[0046] Figure 16 Is a flowchart showing a method of performing a Figure 15 First depth image acceptance test in according to some embodiments.

[0047] Figure 17 Is a flowchart showing a method of performing a Figure 15 Second depth image acceptance test in according to some embodiments.

[0048] Figure 18 Shows a table used in a method of classifying all pixels in a rectangle relative to a Figure 17 Minimum brick value (bmin) and maximum brick value (bmax) in according to some embodiments.

[0049] Figures 19A to 19F Is a schematic diagram showing a method of selecting bricks for a camera frustum according to some embodiments.

[0050] Figures 20A to 20BIs a schematic diagram showing the selection of bricks for a depth image including a surface according to some embodiments.

[0051] Figure 21 Is a schematic diagram showing a plane extraction system according to some embodiments.

[0052] Figure 22 Is a schematic diagram showing, according to some embodiments, Figure 21 Of the parts of a plane extraction system with details on plane extraction.

[0053] Figure 23 Is a schematic diagram showing a scene represented by bricks including voxels and exemplary plane data in the scene according to some embodiments.

[0054] Figure 24 Is a schematic diagram showing, according to some embodiments, Figure 21 Of a plane data repository.

[0055] Figure 25 Is a schematic diagram showing plane geometry extraction when a plane query is sent to Figure 21 Of a plane data repository according to some embodiments.

[0056] Figure 26A Is a schematic diagram showing the generation of Figure 25 Of plane coverage points according to some embodiments.

[0057] Figure 26B Is a schematic diagram showing various exemplary plane geometry representations that can be extracted from an exemplary rasterized plane mask according to some embodiments.

[0058] Figure 27 Shows a grid for a scene according to some embodiments.

[0059] Figure 28A Shows a Figure 27 Scene represented by an external rectangular plane according to some embodiments.

[0060] Figure 28B Shows a Figure 27 Scene of A represented by an internal rectangular plane according to some embodiments.

[0061] Figure 28C Shows a Figure 27 Scene represented by a polygonal plane according to some embodiments.

[0062] Figure 29 Shows a Figure 27 Scene of a denoised grid by planarizing the grid shown in Figure 27 According to some embodiments.

[0063] Figure 30 is a flowchart showing a method of generating a model of an environment represented by a grid according to some embodiments.

[0064] Figure 31 is a schematic diagram showing a 2D representation of a portion of the physical world through four blocks according to some embodiments.

[0065] Figures 32A - 32D is a schematic diagram showing the grid evolution of an exemplary grid block during multi-level simplification according to some embodiments.

[0066] Figure 33A and 33B show representations of the same environment without simplification and simplified by triangle reduction, respectively.

[0067] Figure 34A and 34B show close-up representations of the same environment without simplification and simplified by triangle reduction, respectively.

[0068] Figure 35A and 35B show representations of the same environment without planarization and with planarization, respectively.

[0069] Figure 36A and 36B show representations of the same environment without simplification and simplified by removing disconnected components, respectively.

[0070] Figure 37 is a schematic diagram showing an electronic system enabling an interactive X reality environment for multiple users according to some embodiments.

[0071] Figure 38 is a schematic diagram showing the Figure 37 interaction of components of the electronic system in according to some embodiments.

[0072] Figure 39 is a flowchart showing a method of operating the Figure 37 electronic system in according to some embodiments.

[0073] Figure 40 is a flowchart showing a method of capturing 3D information about an object in the physical world and representing the physical world as Figure 39 3D reconstructed blocks in according to some embodiments.

[0074] Figure 41 is a flowchart showing a method of selecting a version of a block representing a subset of blocks in Figure 39 according to some embodiments.

[0075] Figure 42 is a flowchart showing a method of operating an electronic system in accordance with some embodiments. Figure 37 in the

[0076] Figure 43A is a simplified schematic diagram showing an update detected in a portion of the physical world represented by grid blocks in accordance with some embodiments.

[0077] Figure 43B is a simplified schematic diagram showing a grid block in accordance with some embodiments.

[0078] Figure 43C is a simplified schematic diagram showing a crack at the edge of two adjacent grid blocks in accordance with some embodiments.

[0079] Figure 43D is a simplified schematic diagram showing masking (paper) of a crack in accordance with some embodiments by implementing a grid side edge that overlaps with an adjacent grid block. Figure 43C in the

[0080] Figure 44 is a schematic diagram showing a 2D representation of a portion of the physical world through four blocks in accordance with some embodiments.

[0081] Figure 45 is a schematic diagram showing a 3D representation of a portion of the physical world through eight blocks in accordance with some embodiments.

[0082] Figure 46 is a schematic diagram showing a 3D representation of a portion of the physical world obtained by updating the 3D representation in accordance with some embodiments. Figure 45 in the

[0083] Figure 47 is a schematic diagram showing an example of an augmented world viewable by a first and a second user wearing an AR display system in accordance with some embodiments.

[0084] Figure 48 is a schematic diagram showing an example of an augmented world obtained by updating the augmented world with a new version of a block in accordance with some embodiments. Figure 47 in the

[0085] Figure 49 is a schematic diagram showing an occlusion rendering system in accordance with some embodiments.

[0086] Figure 50 is a schematic diagram showing a depth image with holes.

[0087] Figure 51 is a flowchart showing a method of performing occlusion rendering in an augmented reality environment in accordance with some embodiments.

[0088] Figure 52 is a flowchart showing details of generating surface information based on depth information captured by a depth sensor worn by a user in Figure 51 .

[0089] Figure 53 is a flowchart showing details of filtering depth information to generate a depth map in Figure 52 .

[0090] Figure 54A is a sketch of imaging from a first point of view with a depth camera to identify regions of voxels occupied by a surface and empty voxels.

[0091] Figure 54B is a sketch of imaging from multiple points of view with a depth camera to identify regions of voxels occupied by a surface and empty voxels and indicate regions of "holes" for which there is no available volume information because the voxels in the region of the "hole" are not imaged by the depth camera. DETAILED DESCRIPTION

[0092] Methods and apparatus for creating and using three-dimensional (3D) world reconstructions in augmented reality (AR), mixed reality (MR), or virtual reality (VR) systems are described herein. To provide a realistic AR / MR / VR experience to a user, an AR / MR / VR system must understand the user's physical environment in order to correctly associate the position of virtual objects with real objects. A world reconstruction can be built based on image and depth information about those physical environments, which is collected by sensors that are part of the AR / MR / VR system. Then, this world reconstruction can be used by any one of a number of components of such a system. For example, this world reconstruction can be used by components that perform visual occlusion processing, compute physics-based interactions, or perform environmental reasoning.

[0093] Occlusion processing identifies portions of a virtual object that should not be rendered and / or displayed to the user because there is an object in the physical world that would prevent the user from viewing the position where the user would perceive the virtual object. Physics-based interactions are computed to determine the position or manner in which a virtual object is displayed to the user. For example, a virtual object can be rendered to appear to be resting on a physical object, moving through empty space, or colliding with the surface of a physical object. The world reconstruction provides a model from which information about objects in the physical world can be obtained for such computations.

[0094] Environmental reasoning can also use world reconstruction in the process of generating information that can be used to calculate how to render virtual objects. For example, environmental reasoning can involve identifying a transparent surface by recognizing whether it is a windowpane or a glass tabletop. Based on such identification, an area containing a physical object can be classified as not occluding a virtual object, but can be classified as interacting with the virtual object. Environmental reasoning can also generate information used in other ways, such as identifying fixed objects that can be tracked relative to the user's field of view to calculate the movement of the user's field of view.

[0095] However, there are significant challenges in providing such a system. Computing world reconstruction can require a large amount of processing. In addition, an AR / MR / VR system must correctly know how to position virtual objects relative to the user's head, body, etc. As the user's position changes relative to the physical environment, the relevant part of the physical world can also change, which requires further processing. Moreover, it is often necessary to update the 3D reconstruction data as an object moves in the physical world (e.g., a cup moves on a table). The update of the data representing the environment that the user is experiencing must be performed quickly without using a large amount of computing resources of the computer generating the AR / MR / VR environment, because other functions cannot be performed while performing world reconstruction. In addition, processing the reconstruction data by components that "consume" the data exacerbates the demand for computer resources.

[0096] Known AR / MR / VR systems require high computing power (e.g., GPU) only within a predefined reconstruction volume (e.g., a predefined voxel grid) to run real-time world reconstruction. The inventors have recognized and understood techniques for operating an AR / MR / VR system to provide accurate 3D reconstruction data in real time with little use of computing resources (e.g., computing power (e.g., a single ARM core), memory (e.g., less than 1GB), and network bandwidth (e.g., less than 100Mbps)). These techniques involve reducing the processing required to generate and maintain world reconstruction, and providing and consuming data with low computational overhead.

[0097] These techniques can include, for example, reducing the amount of data processed when updating world reconstruction by identifying portions of sensor data that are readily available for creating or updating world reconstruction. For example, the sensor data can be selected based on whether it represents a part of the physical world that is likely to be close to the surface of an object represented in the world reconstruction.

[0098] In some embodiments, computing resources can be reduced by simplifying the data representing world reconstruction. A simpler representation can reduce the resources used for processing, storing, and / or managing the data and its use.

[0099] In some embodiments, the physical world can be represented in chunks that can be stored and retrieved separately, but combined in a way that provides a realistic representation of the physical world. These chunks can be managed in memory to limit computing resources, and in some embodiments, can be shared between AR / MR / VR systems that can operate in the same physical space, so that each AR / MR / VR system performs less processing to build the world reconstruction.

[0100] In some embodiments, by associating altitude information with a representation of a portion of the physical world, the computational burden can be reduced and / or the accuracy of processing information representing the physical world can be increased. For a dense representation of the physical world, these representations can be mesh chunks as described herein, or connected meshes representing larger portions of the physical world, such as the floor of a building.

[0101] In some embodiments, the use of computing resources can be reduced by selecting from different representations of the physical world when accessing information about the physical world. World reconstruction can include, for example, information about the physical world captured from different sensors and / or stored in different formats. The simplest data consumed or provided can be provided to components that use world reconstruction to render virtual objects. If the simpler data is not available, data obtained using other sensors can be accessed, which can generate a higher computational load. As an example, world reconstruction can include depth maps collected using depth sensors and more cumbersome representations of the 3D world, such as meshes that can be stored as computed from image information. Information about the physical world can be provided to components that perform occlusion processing based on the depth map, if available. In cases where there are holes in the depth map, information to fill these holes can be extracted from the mesh. In some embodiments, the depth map can be "live", representing the physical world captured by the depth sensor at the time the data is accessed.

[0102] The techniques described herein can be used with or alone in a variety of types of devices and for a variety of types of scenarios, including wearable or portable devices with limited computing resources that provide augmented reality scenarios.

[0103] Overview of the AR system

[0104] Figures 1 - 2 Such a scenario is shown. For illustrative purposes, an AR system is used as an example of an XR system. Figure 3 -8 shows an exemplary AR system that includes one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.

[0105] Reference Figure 1, depicts an AR scene 4 in which a user of AR technology sees a physical-world park-like setting 6 featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the users of AR technology also perceive that they "see" a robotic statue 10 standing on the physical-world concrete platform 8 and an anthropomorphic, cartoon-like avatar character 2 that appears to be flying like a bumblebee, even though these elements (e.g., avatar character 2 and robotic statue 10) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to produce an AR technology that promotes a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or physical-world image elements.

[0106] Such an AR scene can be implemented by a system including a world reconstruction component that can construct and update a representation of the physical-world surfaces around the user. This representation can be used for occlusion rendering, placing virtual objects in physics-based interactions, and for virtual character path planning and navigation, or for other operations that use information about the physical world. Figure 2 Depicts another example of an AR scene 200, which shows an exemplary world reconstruction use case according to some embodiments, including visual occlusion 202, physics-based interaction 204, and environmental reasoning 206.

[0107] The exemplary scene 200 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, the users of AR technology also perceive virtual objects such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking out from the bookshelf, and an ornament in the form of a windmill placed on the coffee table. For the image on the wall, the AR technology requires not only information about the surface of the wall but also information about the objects and surfaces (e.g., the shape of the lamp) in the room that are occluding the image to correctly render the virtual object. For the flying bird, the AR technology requires information about all the objects and surfaces around the room to render the bird with realistic physics effects, to avoid objects and surfaces or bounce when the bird collides. For the deer, the AR technology requires information about the surface (e.g., the floor or the coffee table) to calculate the placement of the deer. For the windmill, the system can identify that it is an object separate from the table and can reason that it is movable, while the corner of the bookshelf or the corner of the wall can be reasoned to be fixed. Such distinctions can be used to reason about which parts of the scene are used or updated in each of various operations.

[0108] A scene can be presented to a user through a system including multiple components, the multiple components including a user interface that can stimulate one or more user senses (including vision, sound, and / or touch). Additionally, the system can include one or more sensors that can measure parameters of a physical portion of the scene (including the position and / or movement of the user within the physical portion of the scene). Further, the system can include one or more computing devices having associated computer hardware (such as memory). These components can be integrated into a single device or distributed among multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.

[0109] Figure 3 Depicted is an AR system 302 configured to provide an experience of interacting AR content with the physical world 306 according to some embodiments. The AR system 302 can include a display 308. In the illustrated embodiment, the display 308 can be worn by the user as part of a head-mounted headset such that the user can wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display can be transparent such that the user can observe see-through reality 310. The see-through reality 310 can correspond to the portion of the physical world 306 within the current viewpoint of the AR system 302, and in the case where the user is wearing a head-mounted headset incorporating the display and sensors of the AR system to obtain information about the physical world, the current viewpoint of the AR system 302 can correspond to the user's viewpoint.

[0110] AR content can also be presented on the display 308, overlaying the see-through reality 310. To provide accurate interaction between the AR content and the see-through reality 310 on the display 308, the AR system 302 can include sensors 322 configured to capture information about the physical world 306.

[0111] The sensors 322 can include one or more depth sensors that output depth maps 312. Each depth map 312 can have multiple pixels, and each pixel can represent the distance to a surface in the physical world 306 in a specific direction relative to the depth sensor. Raw depth data can be received from the depth sensors to create the depth maps. Such depth maps can be updated as fast as the depth sensors can form new images, up to hundreds or thousands of times per second. However, the data can be noisy and incomplete, and have holes shown as black pixels on the illustrated depth maps.

[0112] The system may include other sensors, such as image sensors. The image sensors may acquire information that can be otherwise processed to represent the physical world. In some embodiments, the system may include other sensors, such as, for example, one or more of the following sensors: infrared camera sensors, visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyro sensors, accelerometers, magnetometers, altimeters, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and haptic interfaces.

[0113] Sensor data can be used by the portable system to enable the system to track its position relative to the physical world. Additionally, this information can enable the device to be positioned relative to a coordinate system that persists across system sessions and / or is shared with other systems. For example, the ability to determine position relative to a shared coordinate system is described in co-pending U.S. application No. 62 / 982,694, the entire content of which is incorporated herein by reference. In that application, the position is determined based on a sparse representation of the physical world (e.g., a canonical map), and each of a plurality of portable user systems can be positioned against the canonical map by providing a localization request that includes sparse information about a portion of the current environment of the portable system.

[0114] Dense data about the physical world can also be acquired. For example, sensor data can be processed in the world reconstruction component 316 to create a mesh representing connected portions of objects in the physical world. Metadata about these objects, including, for example, color and surface texture, can be similarly acquired with the sensors and stored as part of the world reconstruction. Metadata about the position of the device, including the system, can be determined or inferred based on the sensor data. For example, localization relative to a canonical map can be used to determine the position. Alternatively or additionally, magnetometers, altimeters, GPS sensors, etc. can be used to determine or infer the position of the device. The position of the device can have different degrees of granularity. For example, the accuracy of the determined device position can vary from coarse (e.g., accurate to a sphere with a 10-meter diameter) to fine (e.g., accurate to a sphere with a 3-meter diameter) to ultra-fine (e.g., accurate to a sphere with a 1-meter diameter) to super-fine (e.g., accurate to a sphere with a 0.5-meter diameter), etc.

[0115] In some embodiments, sensor data can be processed to determine which floor of a multi-story building the device is located on. As the user of the device navigates around one or more floors of a multi-story building, the sensor data can be processed in the world reconstruction component 316 to create a mesh representing connected portions of objects in the physical world, which may be on different floors. For example, as the user of the device navigates around one or more floors of a multi-story building, the sensor data can be processed in the world reconstruction component 316 to create a continuous mesh including objects from the respective floors of the user's navigation or other objects at different heights otherwise. Thus, in some embodiments, height and floor information can be used interchangeably, with relative differences in height indicating differences in floors and vice versa. Alternatively or additionally, the mesh can have multiple layers, where each layer represents one or more floors in the building or otherwise represents objects at different heights.

[0116] In some embodiments, height can be measured directly, for example using an altimeter. Relative changes in height can enable the device to detect that it is moving between floors, such that the mesh data used or generated by the device is associated with the appropriate floor at any given time.

[0117] Alternatively or additionally, height can be inferred from other measurements. For example, in some embodiments of an XR system, an environment such as an office building can have markings identifying the floors of the building, which can be recognized by processing the sensor output on a portable system. For example, the markings can be in the form of machine-readable QR codes, or can be the color of a wall or other indicators that can be detected by the sensors. As another example, the portable system can be configured to read signs or other human-readable indicators of location, such as elevator buttons. As yet another example, a canonical map can be created for each floor of the building, such that positioning relative to the map that uniquely represents a particular floor can indicate height. In embodiments where multiple floors of a building have a similar layout, other indicators can be used in place of or in addition to the positioning results. As yet another example, when the device tracks its position relative to the physical world, it can detect actions indicating movement between building floors. Such detected actions can include traversing a set of stairs or an escalator.

[0118] Merging height information into a dense representation and / or a sparse representation of the environment can enable more robust handling of sparse information and / or dense information. For example, in a hotel or office building with many floors that look similar, traditional localization methods may result in inaccurate results and / or require multiple iterations to resolve the location between similar floors. Merging height information into a sparse representation and / or a dense representation of the physical world can improve accuracy and / or reduce the processing time required to determine a localization result, because the localization service can limit the localization process by only attempting to localize relative to a map that represents a portion of the physical world having the same height as the current height of the portable system. In embodiments where the localization service performs localization, the localization request may include height information.

[0119] As another example, associating height information with grid information can enable vertical ordering of dense information, such that processing can more accurately determine a surface between a user and an object in a direction the user may be looking. For example, processing of a user on a building floor can determine that a virtual object below that user's floor is not visible because it is occluded by the floor. In another scenario, processing can determine that the virtual object is visible because the user is looking down in a stairwell.

[0120] The system can also obtain information about the user's head pose relative to the physical world. In some embodiments, sensor 310 may include an inertial measurement unit that can be used to compute and / or determine head pose 314. Head pose 314 for a depth map may indicate the current viewing point at which the sensor captures the depth map in, for example, six degrees of freedom (6DoF), but head pose 314 can be used for other purposes, such as correlating image information with a particular portion of the physical world or correlating the position of a display worn on the user's head with the physical world. In some embodiments, head pose information can be derived in other ways different from an IMU (e.g., analyzing objects in an image).

[0121] The world reconstruction component 316 can receive the depth map 312 and head pose 314 from the sensor and any other data, and integrate the data into a reconstruction 318, which can at least appear as a single combined reconstruction. Reconstruction 318 can be more complete and less noisy than the sensor data. The world reconstruction component 316 can update reconstruction 318 using spatial and temporal averaging of sensor data from multiple viewpoints over time.

[0122] The reconstruction 318 may include a representation of the physical world having one or more data formats including, for example, voxels, meshes, planes, etc. The different formats may represent alternative representations of the same part of the physical world or may represent different parts of the physical world. In the example shown, on the left side of the reconstruction 318, a part of the physical world is presented as a global surface; on the right side of the reconstruction 318, a part of the physical world is presented as a mesh.

[0123] The reconstruction 318 can be used for AR functions such as generating a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation can change as the user moves or as objects in the physical world change. Aspects of the reconstruction 318 can be used, for example, by a component 320 that generates a global surface representation that varies in world coordinates, and this varying global surface representation can be used by other components.

[0124] AR content can be generated based on this information, for example, by an AR application 304. The AR application 304 can be, for example, a game program that performs one or more functions based on information about the physical world (e.g., visual occlusion, physics-based interactions, and environmental reasoning). It can perform these functions by querying data in different formats of the reconstruction 318 generated by the world reconstruction component 316. In some embodiments, the component 320 can be configured to output an update when the representation in the region of interest of the physical world changes. The region of interest can be set, for example, to approximate a part of the physical world near the user of the system, such as the part within the user's field of view, or projected (predicted / determined) to enter the user's field of view.

[0125] The AR application 304 can use this information to generate and update AR content. The virtual part of the AR content can be presented on the display 308 in combination with the perspective reality 310, thereby creating a realistic user experience.

[0126] In some embodiments, an AR experience can be provided to the user through a wearable display system. Figure 4 An example of a wearable display system 80 (hereinafter referred to as "system 80") is shown. The system 80 includes a head-mounted display device 62 (hereinafter referred to as "display device 62") and various mechanical and electronic modules and systems to support the functions of the display device 62. The display device 62 can be coupled to a frame 64, and the frame 64 can be worn by a user or viewer 60 of the display system (hereinafter referred to as "user 60") and is configured to position the display device 62 in front of the eyes of the user 60. According to various embodiments, the display device 62 can be a sequential display. The display device 62 can be monocular or binocular. In some embodiments, the display device 62 can be Figure 3 an example of the display 308 in

[0127] In some embodiments, speaker 66 is coupled to frame 64 and is positioned near the ear canal of user 60. In some embodiments, another speaker (not shown) is positioned near the other ear canal of user 60 to provide stereo / shapeable sound control. Display device 62 is operatively coupled to local data processing module 70, for example, via a wired lead or wireless connection 68, and local data processing module 70 can be mounted in a variety of configurations, such as fixedly attached to frame 64, fixedly attached to a helmet or hat worn by user 60, embedded in the earphones, or otherwise removably connected to user 60 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0128] Local data processing module 70 can include a processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which can be used to assist in the processing, caching, and storage of data. The data includes a) data captured from sensors (which can be operatively coupled to frame 64, for example) or otherwise attached to user 60 (e.g., an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio device, and / or a gyroscope), and / or b) data obtained using remote processing module 72 and / or remote data repository 74, which may be used to be passed to display device 62 after such processing or retrieval. Local data processing module 70 can be operatively coupled to remote processing module 72 and remote data repository 74, respectively, via communication links 76, 78 (e.g., via a wired or wireless communication link), such that these remote modules 72, 74 are operatively coupled to each other and can be used as resources for local processing and data module 70. In some embodiments, Figure 3 the world reconstruction component 316 in can be at least partially implemented in local data processing module 70. For example, local data processing module 70 can be configured to execute computer-executable instructions to generate a physical world representation based at least in part on at least a portion of the data.

[0129] In some embodiments, local data processing module 70 can include one or more processors (e.g., a graphics processing unit (GPU)) configured to analyze and process data and / or image information. In some embodiments, local data processing module 70 can include a single processor (e.g., a single-core or multi-core ARM processor), which will limit the computational budget of module 70 but enable a smaller device. In some embodiments, world reconstruction component 316 can use a computational budget less than that of a single ARM core to generate a physical world representation in real time on a non-predefined space, such that the remaining computational budget of a single ARM core can be accessed for other purposes, such as mesh extraction.

[0130] In some embodiments, the remote data repository 74 may include a digital data storage facility that can be used via the Internet or other network configurations in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, allowing for fully autonomous use from remote modules. World reconstruction, for example, may be stored in whole or in part in the repository 74.

[0131] In some embodiments, the local data processing module 70 is operatively coupled to a battery 82. In some embodiments, the battery 82 is a removable power source, such as above a counter battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during non-operating times of the system 80 and a removable battery, such that the user 60 can operate the system 80 for longer periods of time without having to connect to a power source to charge the lithium-ion battery or having to turn off the system 80 to replace the battery.

[0132] Figure 5A A user 30 wearing an AR display system that renders AR content is shown as the user 30 moves in a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 places the AR display system at location 34, and the AR display system records environmental information of the passable world relative to location 34 (e.g., digital representations of real objects in the physical world that can be stored and updated as the real objects in the physical world change), such as poses related to mapped features or directional audio inputs. Location 34 is aggregated into data input 36 and processed at least by the passable world module 38, which may be achieved, for example, by processing on Figure 3 the remote processing module 72. In some embodiments, the passable world module 38 may include a world reconstruction component 316.

[0133] The traversable world module 38 determines the location and manner in which AR content 40 can be placed in the physical world as determined from data input 36. By presenting a representation of the physical world and the AR content via the user interface, the AR content is "placed" in the physical world, where the AR content is rendered as if it is interacting with objects in the physical world, and the objects in the physical world are rendered as if the AR content is occluding the user's view of those objects when appropriate. In some embodiments, the shape and location of the AR content 40 can be determined by appropriately selecting portions of a stationary element 42 (such as a table) from the reconstruction (e.g., reconstruction 318) to place the AR content. As an example, the stationary element can be a table, and the virtual content can be placed such that it appears to be on the table. In some embodiments, the AR content can be placed within a structure in the field of view 44, which can be the current field of view or an estimated future field of view. In some embodiments, the AR content can be placed relative to a mapped grid model 46 of the physical world.

[0134] As depicted, the stationary element 42 serves as a proxy for any stationary element within the physical world, which can be stored in the traversable world module 38 such that the user 30 can perceive the content on the stationary element 42 without the system having to map to the stationary element 42 every time the user 30 sees it. The stationary element 42 can thus be a mapped grid model from a previous modeling session or determined according to an individual user, but both are stored on the traversable world module 38 for future reference by multiple users. Thus, the traversable world module 38 can identify the environment 32 from a previously mapped environment and display the AR content without the user 30's device first mapping the environment 32, thereby saving computational processes and cycles and avoiding latency of any rendered AR content.

[0135] The mapped grid model 46 of the physical world can be created by the AR display system, and the appropriate surfaces and metrics for interacting with and displaying the AR content 40 can be mapped and stored in the traversable world module 38 for future retrieval by the user 30 or other users without remapping or modeling. In some embodiments, the data input 36 is an input such as a geographical location, user identification, and current activity to indicate to the traversable world module 38 which stationary element 42 among one or more stationary elements is available, which AR content 40 was last placed on the stationary element 42, and whether to display that same content (the AR content is "persistent" content regardless of whether the user views a particular traversable world model).

[0136] Figure 5BA schematic diagram showing the viewing optical assembly 48 and accompanying components is presented. In some embodiments, facing the user's eyes 49, two eye tracking cameras 50 detect metrics of the user's eyes 49, such as eye shape, eyelid occlusion, pupil direction, and glint on the user's eyes 49. In some embodiments, a depth sensor 51 (e.g., a time-of-flight sensor) emits a relay signal into the world to determine the distance to a given object. In some embodiments, a world camera 52 records a greater-than-peripheral view to map the environment 32 and detect inputs that may affect the AR content. Camera 53 may further capture a physical world image within the user's field of view at a specific timestamp. Each of the world camera 52, camera 53, and depth sensor 51 has a corresponding field of view 54, 55, and 56 to collect data from a physical world scene such as the physical world environment 32 as shown in Figure 3 and record that physical world scene.

[0137] An inertial measurement unit 57 may determine the motion and orientation of the viewing optical assembly 48. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 51 is operatively coupled to the eye tracking camera 50 as a confirmation of the measured accommodation relative to the actual distance the user's eyes 49 are looking at.

[0138] Information from these sensors in the viewing optical assembly 48 may be coupled to one or more processors in the system. The processor may generate data that can be rendered to enable the user to perceive interacting with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data depicting physical and virtual objects. In other embodiments, physical and virtual content may be depicted in a scene by modulating the opacity of a display device through which the user views the physical world. The opacity may be controlled to create the appearance of virtual objects and also prevent the user from seeing objects in the physical world that are occluded by the virtual objects. Regardless of how the content is presented to the user, a model of the physical world is required so that the characteristics of virtual objects that can be affected by physical objects, including the shape, position, motion, and visibility of the virtual objects, can be correctly calculated. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 318.

[0139] The model may be created based on data collected from sensors on the user's wearable device. However, in some embodiments, the model may be created based on data collected from multiple users, which may be aggregated in a computing device remote from all users (and may be "in the cloud").

[0140] It may be at least partially through a world reconstruction system (e.g., Figure 6 more detailedly depicted inFigure 3 using the world reconstruction component 316) to create a model. The world reconstruction component 316 may include a perception module 160, and the perception module 160 may generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 160 may represent a portion of the physical world within the reconstruction range of the sensor as a plurality of voxels. Each voxel may correspond to a 3D cube of a predetermined volume in the physical world and include surface information that indicates whether there is a surface in the volume represented by the voxel. A value may be assigned to the voxel that indicates whether its corresponding volume has been determined to include the surface of a physical object, has been determined to be empty, or has not been measured by the sensor and thus its value is unknown. It should be understood that it is not necessary to explicitly store the values of the voxels determined to be empty or unknown, as the values of the voxels may be stored in computer memory in any suitable manner, including not storing information about the voxels determined to be empty or unknown.

[0141] Figure 7A FIG. depicts an example of a 3D space 100 discretized into voxels 102. In some embodiments, the perception module 160 may determine an object of interest and set the volume of the voxels in order to capture the characteristics of the object of interest and avoid redundant information. For example, the perception module 160 may be configured to identify larger objects and surfaces, such as walls, ceilings, floors, and large furniture. Thus, the volume of the voxels may be set to a relatively large size, such as 4 cm 3 cube.

[0142] The reconstruction of the physical world including voxels may be referred to as a volume model. As the sensor moves in the physical world, information for creating a volume model may be created over time. This movement may occur when the user of the wearable device including the sensor moves around. Figure 8A -F depicts an example of reconstructing the physical world as a volume model. In the example shown, the physical world includes Figure 8A a portion 180 of the surface shown in Figure 8A . In

[0143] The sensor 182 may be of any suitable type, such as a depth sensor. However, depth data may be obtained from an image sensor or otherwise. The perception module 160 may receive data from the sensor 182 and then set the values of a plurality of voxels 186, as shown in Figure 8B , to represent the portion 180 of the surface visible in the field of view 184 by the sensor 182.

[0144] In Figure 8C , the sensor 182 may move to a second position and have a field of view 188. As shown inFigure 8D As shown, another set of voxels becomes visible, and the values of these voxels can be set to indicate the position of the portion of the surface that has entered the field of view 188 of the sensor 182. The values of these voxels can be added to the volume model for the surface.

[0145] In Figure 8E the sensor 182 can be further moved to a third position and has a field of view 190. In the example shown, an additional portion of the surface becomes visible in the field of view 190. As Figure 8F shown, another set of voxels can become visible, and the values of these voxels can be set to indicate the position of the portion of the surface that has entered the field of view 190 of the sensor 182. The values of these voxels can be added to the volume model for the surface. As Figure 6 shown, this information can be stored as part of the persistent world as volume information 162a. Information about the surface, such as color or texture, can also be stored. Such information can be stored as, for example, volume metadata 162b.

[0146] In addition to generating information for the persistent world representation, the perception module 160 can also identify and output an indication of a change in the area around the user of the AR system. This indication of change can trigger an update to the volume data stored as part of the persistent world, or trigger other functions, such as triggering the trigger component 304 that generates AR content to update the AR content.

[0147] In some embodiments, the perception module 160 can identify changes based on a signed distance function (SDF) model. The perception module 160 can be configured to receive sensor data such as a depth map 160a and a head pose 160b, and then fuse the sensor data into the SDF model 160c. The depth map 160a can directly provide SDF information, and the image can be processed to obtain SDF information. The SDF information represents the distance from the sensor used to capture the information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit and thus from the perspective of the user. The head pose 160b can enable the SDF information to be related to voxels in the physical world.

[0148] Returning to Figure 6 , in some embodiments, the perception module 160 can generate, update, and store a representation of the portion of the physical world within the perception range. The perception range can be determined at least in part based on the reconstruction range of the sensor, which can be determined at least in part based on the limits of the observation range of the sensor. As a specific example, an active depth sensor operating with active IR pulses can operate reliably within a certain distance range, creating an observation range of the sensor that can range from a few centimeters or tens of centimeters to several meters.

[0149] Figure 7B Depicts the reconstruction range relative to sensor 104 with viewpoint 106. A reconstruction of the 3D space within viewpoint 106 can be constructed based on data captured by sensor 104. In the example shown, the viewing range of sensor 104 is from 40 cm to 5 m. In some embodiments, the reconstruction range of the sensor can be determined to be less than the viewing range of the sensor because sensor outputs near its viewing limit can be noisier, more incomplete, and less accurate. For example, in the 40 cm to 5 m example shown, the corresponding reconstruction range can be set to from 1 to 3 m, and data collected by the sensor indicating surfaces outside of that range may not be used.

[0150] In some embodiments, the sensing range can be greater than the reconstruction range of the sensor. If component 164 that uses data about the physical world needs data about regions within the sensing range that are outside of the portion of the physical world that is currently within the reconstruction range, that information can be provided from the persistent world 162. Accordingly, information about the physical world can be easily accessed through queries. In some embodiments, an API can be provided to respond to such queries to provide information about the user's current sensing range. Such techniques can reduce the time required to access an existing reconstruction and provide an improved user experience.

[0151] In some embodiments, the sensing range can be a 3D space corresponding to a bounding box centered around the user's location. As the user moves, the portion of the physical world that can be queried by component 164 within the sensing range can move with the user. Figure 7C Depicts a bounding box 110 centered at location 112. It should be understood that the size of bounding box 110 can be set to surround the viewing range of the sensor with a reasonable expansion since the user cannot move at an unreasonable speed. In the example shown, the viewing limit of the sensor worn by the user is 5 m. The bounding box 110 is set to a 3 cube of 20 m.

[0152] Return Figure 6 , the world reconstruction component 316 can include additional modules that can interact with the sensing module 160. In some embodiments, the persistent world module 162 can receive a representation of the physical world based on data acquired by the sensing module 160. The persistent world module 162 can also include representations of the physical world in various formats. For example, volumetric metadata 162b such as voxels, as well as meshes 162c and planes 162d can be stored. In some embodiments, other information such as depth maps can be saved.

[0153] In some embodiments, the perception module 160 may include modules that generate representations for the physical world in various formats, including, for example, a mesh 160d, planar, and semantic 160e. These modules may generate the representations based on data within the perception range of one or more sensors at the time of generating the representation, as well as data captured at previous times and information in the persistent world 162. In some embodiments, these components may operate on depth information captured by a depth sensor. However, the AR system may include visual sensors and may generate such representations by analyzing monocular or binocular visual information.

[0154] In some embodiments, as described below, these modules may operate on regions of the physical world, such as regions represented by blocks or tiles. When the perception module 160 detects a change in the physical world in other sub-regions, those modules may be triggered to update the blocks or tiles of the physical world or that sub-region. For example, such a change may be detected by detecting a new surface in the SDF model 160c or other criteria (such as changing the values of a sufficient number of voxels representing the sub-region).

[0155] The world reconstruction component 316 may include a component 164 that can receive a representation of the physical world from the perception module 160. Information about the physical world may be pulled by these components according to, for example, a usage request from an application. In some embodiments, the information may be pushed to the usage component, for example, via an indication of a change in a pre-identified region or a change in the representation of the physical world within the perception range. The component 164 may include, for example, game programs and other components that perform processing for visual occlusion, physics-based interaction, and environmental reasoning.

[0156] In response to a query from the component 164, the perception module 160 may send a representation of the physical world in one or more formats. For example, when the component 164 indicates that the usage is for visual occlusion or physics-based interaction, the perception module 160 may send a representation of the surface. When the component 164 indicates that the usage is for environmental reasoning, the perception module 160 may send a mesh, planar, and semantic of the physical world.

[0157] In some embodiments, the perception module 160 may include components that format information to provide to the component 164. An example of such a component may be a ray casting component 160f. The usage component (e.g., component 164) may, for example, query information about the physical world from a specific perspective. The ray casting component 160f may select from one or more representations of the physical world data within the field of view from that viewpoint.

[0158] Viewpoint - dependent voxel selection for fast volume reconstruction

[0159] It should be understood from the foregoing description that the perception module 160 or another component of the AR system may process data to create a 3D representation of a portion of the physical world. The portion of the 3D reconstruction volume may be selected by at least partially based on the camera frustum and / or depth image, plane data may be extracted and retained, 3D reconstruction data in blocks that allow local updates while maintaining neighbor consistency may be captured, retained and updated, occlusion data (where the occlusion data is derived from a combination of one or more depth data sources) may be provided to an application that generates such a scene, and / or multi-level mesh simplification may be performed to reduce the data to be processed.

[0160] The world reconstruction system may integrate sensor data over time from multiple viewpoints of the physical world. When the device including the sensor moves, the pose (e.g., position and orientation) of the sensor may be tracked. Since the frame pose of the sensor is known and how it relates to other poses, each of these multiple viewpoints of the physical world may be fused in a single combined reconstruction. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time), the reconstruction may be more complete and less noisy than the original sensor data.

[0161] The reconstruction may include data at different levels of complexity, including for example raw data (e.g., real-time depth data), fused volume data (e.g., voxels), and computed data (e.g., meshes).

[0162] In some embodiments, AR and MR systems represent a 3D scene with a regular voxel grid, where each voxel may contain a signed distance field (SDF) value. The SDF value describes whether the voxel is inside or outside the surface of the scene to be reconstructed and the distance from the voxel to that surface. Computing the 3D reconstruction data for the volume required to represent the scene requires a large amount of memory and processing power. Since the number of variables required for 3D reconstruction grows cubically with the number of depth images processed, these requirements also increase for scenes representing larger spaces.

[0163] Described herein are efficient ways to reduce processing. According to some embodiments, the scene may be represented by one or more bricks. Each brick may include a plurality of voxels. The bricks to be processed to generate a 3D reconstruction of the scene may be selected by picking a set of bricks representing the scene based on a frustum derived from the field of view (FOV) of an image sensor and / or a depth image (or "depth map") of the scene created using a depth sensor.

[0164] A depth image can have one or more pixels, each pixel representing the distance to a surface in a scene. These distances can be related to the position relative to the image sensor, such that the data output from the image sensor can be selectively processed. The image data can be processed for those voxels representing portions of the 3D scene that include surfaces visible from the perspective (or "viewpoint") of the image sensor. Processing of some or all of the remaining voxels can be omitted. By this method, the selected voxels can be voxels that are likely to contain new information, which can be obtained by picking voxels for which the output of the image sensor is unlikely to provide useful information about them. The data output from the image sensor is unlikely to provide useful information about voxels that are closer to or farther from the image sensor than the surfaces indicated by the depth map, because these voxels are empty space or behind the surfaces and thus not depicted in the image from the image sensor.

[0165] In some embodiments, one or more criteria can be applied to effectively select a set of voxels for processing. An initial set of voxels can be limited to those within the frustum of the image sensor. Then a large number of voxels outside the frustum can be picked. Then more computationally resource-intensive processing can be performed on a subset of the voxels accepted for processing after the picking to update the 3D reconstruction. Thus, with processing using a reduced number of voxels, the 3D representation of the scene to be updated can be computed more efficiently.

[0166] Greater reduction in processing can be achieved by picking voxels based on the depth image. According to some embodiments, picking and / or acceptance of voxels can be performed by projecting the contour of each voxel in the initial set onto the depth image. Such picking can be based on whether the voxel corresponds to a portion of the scene indicated by the depth image that is near the surface. Voxels that can be simply identified as being entirely in front of the surface or entirely behind the surface can be picked. In some embodiments, such determination can be made efficiently. For example, the bounding box around the projection of the voxel into the depth map can be used to determine the maximum voxel value and the minimum voxel value in the z - coordinate direction, which can be substantially perpendicular to the 2D plane of the depth image. By comparing these maximum and minimum voxel values with the distances represented by the pixels in the depth map, voxels can be picked and / or accepted for further processing. Such processing can result in the selection of voxels for initial processing that intersect and / or are in front of the surfaces reflected in the depth image. In some embodiments, such processing can distinguish between voxels in front of solid surfaces and voxels in front of perforated surfaces (i.e., voxels representing regions for which the depth sensor cannot reliably measure the distance to the surface).

[0167] In some embodiments, the pick / accept criteria may result in classifying some or all of the bricks accepted for further processing such that the processing algorithms for computing the volume reconstruction can be customized for the characteristics of the bricks. In some embodiments, different processing may be selected based on whether the brick is classified as intersecting a surface, in front of a solid surface, or in front of a perforated surface.

[0168] Figure 9 A cross-sectional view of a scene 400 along a plane parallel to the y and z coordinates is shown. The XR system may represent the scene 400 by a voxel grid 504. Conventional XR systems may update each voxel of the voxel grid based on each new depth image captured by a sensor 406, which may be an image sensor or a depth sensor, such that the 3D reconstruction generated from the voxel grid may reflect changes in the scene. Updating in this way can consume a large amount of computing resources and also cause artifacts at the output of the XR system due to, for example, time delays caused by heavy computations.

[0169] Described herein are techniques for providing accurate 3D reconstruction data with low computational resource utilization, such as by selecting portions of the voxel grid 504 based at least in part on a frustum 404 of a camera of the image sensor 406 and / or depth images captured by the image sensor.

[0170] In the example shown, the image sensor 406 captures a depth image (not shown) including a surface 402 of the scene 400. The depth image may be stored in a computer memory in any convenient manner that captures the distance between a certain reference point and the surface in the scene 400. In some embodiments, the depth image may be represented as values in a plane parallel to the x and y axes, as Figure 9 shown, where the reference point is the origin of the coordinate system. Positions in the X-Y plane may correspond to directions relative to the reference point, and the values at those pixel positions may indicate the distance from the reference point to the nearest surface in the direction indicated by the coordinates in the plane. Such a depth image may include a grid of pixels (not shown) in a plane parallel to the x and y axes. Each pixel may indicate the distance from the image sensor 406 to the surface 402 in a particular direction. In some embodiments, the depth sensor may not be able to measure the distance to the surface in a particular direction. For example, this may occur if the surface is outside the range of the image sensor 406. In some embodiments, the depth sensor may be an active depth sensor that measures distance based on reflected energy, but the surface may not reflect enough energy for an accurate measurement. Thus, in some embodiments, the depth image may have "holes" where pixels are not assigned values.

[0171] In some embodiments, the reference points of the depth image can vary. Such a configuration can allow the depth image to represent surfaces in an entire 3D scene, rather than being limited to a portion having a predetermined and limited angular range relative to a specific reference point. In such embodiments, the depth image can indicate the distance to a surface as the image sensor 406 moves through six degrees of freedom (6DOF). In these embodiments, the depth image can include a set of pixels for each of a plurality of reference points. In these embodiments, a portion of the depth image can be selected based on the "camera pose", which represents the direction and / or orientation in which the image sensor 406 is pointing when the image data is captured.

[0172] The image sensor 406 can have a field of view (FOV), which can be represented by the camera frustum 404. In some embodiments, the depicted infinite camera frustum can be reduced to a finite 3D trapezoidal prism 408 by assuming a maximum depth 410 that the image sensor 406 can provide and / or a minimum depth 412 that the image sensor 406 can provide. The 3D trapezoidal prism 408 can be a convex polyhedron bounded by six planes.

[0173] In some embodiments, one or more voxels 504 can be grouped into bricks 502. Figure 10A A portion 500 of the scene 400 is shown, which includes eight bricks 502. Figure 10B Shows including 8 3 exemplary bricks 502 of individual voxels 504. Refer Figure 9 , the scene 400 can include one or more bricks, sixteen of which are shown in the view Figure 4 shown. Each brick can be identified by a passable brick identifier such as

[0000] -

[0015] .

[0174] Figure 11 Depicts a volume representation hierarchy that can be implemented in some embodiments. In some embodiments, such a volume representation hierarchy can reduce the latency of data transmission. In some embodiments, the voxel grid of the physical world can be mapped to conform to the structure of the storage architecture of a processor (such as the processor on which component 304 is executed) for computing AR content. One or more voxels can be grouped into "bricks". One or more bricks can be grouped into "tiles". The size of a tile can correspond to the storage page of the storage medium local to the processor. Tiles can be moved between local memory and remote memory, for example, via a wireless connection, based on usage or expected usage according to a storage management algorithm.

[0175] In some embodiments, the upload and / or download between the perception module 160 and the persistent world module 162 can be performed on multiple tiles in a single operation. One or more tiles can be grouped into a "RAM tile set". The size of the RAM tile set can correspond to the area within the reconstruction range of the sensors worn by the user. One or more RAM tile sets can be grouped into a "global tile set". The size of the global tile set can correspond to the perception range of the world reconstruction system (e.g., the perception range of the perception module 160).

[0176] Figure 12 FIG. 600 is a flow chart showing a method 600 of operating a computing system to generate a 3D reconstruction of a scene. Method 600 may begin by representing a scene (e.g., scene 400) with one or more bricks (e.g., brick 502), each brick including one or more voxels (e.g., voxel 504). Each brick may represent a portion of the scene. The bricks may be identifiable relative to a persistent coordinate system such that the same brick represents the same volume in the scene even as the pose of the image sensor (e.g., image sensor 406) changes.

[0177] In act 604, method 600 may capture a depth image (e.g., a depth image including surface 402) from a depth sensor (e.g., depth sensor 406). The depth sensor may be an active depth sensor that sends, for example, IR radiation for reflection and measures the time of flight. Each such measurement represents the distance from the depth sensor to the surface in a particular direction. This depth information may represent the same volume as the volume represented by the bricks.

[0178] In act 606, method 600 may pick a portion of one or more bricks for a camera frustum (e.g., a finite 3D trapezoidal prism 408 derived from camera frustum 404) to produce a first one or more bricks, which is a reduced set of bricks from the one or more bricks. Such picking may eliminate bricks that represent portions of the scene that are outside the field of view of the image sensor when acquiring the image data being processed. The image data is thus less likely to contain information useful for creating or updating the bricks.

[0179] In act 608, method 600 may pick a portion of the first one or more bricks for the depth image to produce a second one or more bricks, the second one or more bricks being a reduced set of bricks from the first one or more bricks. In act 610, method 600 may generate a 3D reconstruction of the scene based on the second one or more bricks.

[0180] Return Figure 9, given the surface 402 captured by the depth image and the corresponding camera pose, the voxels between the image sensor 406 and the surface 402 can be empty. The farther a voxel is from the image sensor 406 and behind the surface 402, the less likely it is to determine whether the voxel represents the interior of an object or empty space. The degree of certainty can be represented by a weight function that weights the voxel updates based on the distance to the surface 402. When the weight function of a voxel located behind the surface 402 (farther from the image sensor 402) is higher than a threshold, the voxel may not be updated or may have a zero update (e.g., an update with zero change). Moreover, all voxels that do not fall within the camera frustum 404 may not be updated or investigated for this depth image.

[0181] Method 600 can not only improve the processing speed of volumetric depth image fusion, but also consume less memory storage, which enables Method 600 to run on wearable hardware. For example, a small reconstruction volume of 5m * 5m * 3m with a voxel size of 1cm 3 and 8 bytes per voxel (4 bytes for the distance value and 4 bytes for the weight value) would already require approximately 600MB. Method 600 can classify bricks according to the distance of the bricks to the surface relative to a truncated threshold. For example, Method 600 can identify empty bricks (e.g., selected bricks, or bricks that are farther from the surface than the truncated threshold) so as not to allocate memory space for non-empty bricks. Method 600 can also identify bricks that are far from the surface by the truncated threshold so as to store these bricks with a constant distance value of a negative truncated threshold and a weight of 1. Method 600 can also identify bricks with a distance to the surface between zero and the truncated threshold so as to store these bricks with a constant SDF value of a positive truncated threshold but with a varying weight. Storing a constant distance or weight value for bricks with a single value can be entropy-based compression for a zero-entropy field.

[0182] Method 600 can allow bricks to be marked as "not containing any part of the surface" during voxel updates, which can significantly accelerate the processing of the bricks. The processing can include, for example, converting the image of the part of the scene represented by the bricks into a mesh.

[0183] Figure 13Illustrates an exemplary method 606 for selecting one or more portions of bricks for a camera frustum 404 of an image sensor 406 according to some embodiments. Method 606 may begin by finding an axis-aligned bounding box (AABB) with cubic shape that encloses the camera frustum 404. The AABB may contain one or more bricks in the scene. Method 606 may include dividing (act 704) the AABB into one or more sub-AABBs and performing (act 706) a camera frustum acceptance test. If method 606 determines at act 708 that a sub-AABB reaches the size of a brick, method 606 may produce (act 710) the first one or more bricks. If method 606 determines at act 708 that a sub-AABB is larger than the size of a brick, method 606 may repeat acts 704 - 708 until the sub-AABB reaches the size of a brick.

[0184] For example, given a 3D trapezoidal prism 408 corresponding to the camera frustum 404, an AABB with side lengths that are powers of two and that encloses the 3D trapezoidal prism 408 can be found in constant time. The AABB can be divided into eight sub-AABBs. Each of the eight sub-AABBs can be tested for intersection with the camera frustum 404. When it is determined that a sub-AABB does not intersect the camera frustum 404, the brick corresponding to that sub-AABB can be selected. The selected brick can be rejected from further processing. When it is determined that a sub-AABB intersects the camera frustum 404, that sub-AABB can be further divided into eight sub-AABBs of that sub-AABB. Then, each of the eight sub-AABBs of that sub-AABB can be tested for intersection with the camera frustum 404. The iteration of dividing and testing continues until the sub-AABB corresponds to a single brick. To determine whether the camera frustum 404 intersects the AABB, a two-step test can be performed. First, it can be tested whether at least one corner point of the AABB lies within each plane that bounds the camera frustum 404. Second, it can be tested whether each corner point of the camera frustum 404 lies within the AABB so that some cases of AABBs that do not intersect the camera frustum 404 but are misclassified as partially inside (e.g., only one corner point on the frustum edge) can be captured.

[0185] An ideal byproduct of this two-step test is that for each brick that intersects the camera frustum 404, it can be known whether it is completely within the camera frustum 404 or only partially within the camera frustum 404. For bricks that are completely within the camera frustum 404, when updating individual voxels later, the test of whether it is inside the camera frustum 404 can be skipped for each voxel.

[0186] Figure 14Illustrates an exemplary method 706 for performing a camera frustum acceptance test according to some embodiments. Method 706 may begin by testing (act 802) each of one or more sub-AABBs relative to each plane defining the camera frustum 404. At act 804, method 706 may determine whether the tested sub-AABB is completely outside the camera frustum 404. At act 806, if it is determined that the tested sub-AABB is completely outside the camera frustum 404, method 706 may pick all the bricks included in the tested sub-AABB. At act 808, if it is determined that the tested sub-AABB is not completely outside the camera frustum 404, method 706 may determine whether the tested sub-AABB is completely inside the camera frustum 404.

[0187] At act 810, if it is determined that the tested sub-AABB is completely inside the camera frustum 404, method 706 may add all the bricks included in the tested sub-AABB to the first one or more bricks. At act 708, if it is determined that the tested sub-AABB is not completely inside the camera frustum 404, which may indicate that the tested sub-AABB intersects the camera frustum 404, method 706 may determine whether the tested sub-AABB reaches the size of a brick.

[0188] At act 814, if it is determined that the tested sub-AABB is equal to the size of a brick, method 706 may further determine whether each corner point of the camera frustum 404 is located inside the tested sub-AABB. If it is determined that each corner point of the camera frustum 404 is located inside the brick of the tested sub-AABB, method 706 may pick (act 806) the brick of the tested sub-AABB. If it is determined that not every corner point of the camera frustum is located inside the brick of the tested sub-AABB, method 706 may add (act 810) the brick of the tested sub-AABB to the first one or more bricks.

[0189] Figure 15 Illustrates an exemplary method 608 for picking a portion of the first one or more bricks for a depth image according to some embodiments. Method 608 may begin by performing (act 902) a first depth image acceptance test on each of the first one or more bricks. At act 904, method 808 may determine whether the first depth image acceptance test accepts the tested brick. If it is determined that the tested brick is accepted by the first depth image acceptance test, which may indicate that the tested brick intersects a surface in the scene, method 608 may apply (act 906) an incremental change to the selected voxels and add (act 914) the tested brick to the second one or more bricks.

[0190] At operation 908, if it is determined that the first depth image acceptance test does not accept the tested brick, method 608 may perform (operation 908) a second depth image acceptance test for the tested brick. At operation 910, it is determined whether the second depth image acceptance test accepts the tested brick. If it is determined that the second depth image acceptance test accepts the tested brick, which may indicate that the tested brick is in front of a solid or perforated background in the scene, method 608 may apply (operation 912) a constant increment to all voxels or selected voxels and then add (operation 914) the tested brick to the second one or more bricks. If it is determined that the second depth image acceptance test also does not accept the tested brick, method 608 may pick (operation 916) the tested brick.

[0191] Figure 16 Exemplary method 902 for performing a first depth image acceptance test in accordance with some embodiments is shown. For each brick to be tested, method 902 may begin by determining (operation 1002) a minimum brick value (bmin) and a maximum brick value (bmax) along a direction parallel to the z coordinate. The bmin value and the bmax value may be filled to illustrate an integration threshold beyond which depth values indicate continuous updating of voxels in the brick. At operation 1004, method 902 may calculate the 2D pixel positions of the corners of the tested brick by projecting the corners of the brick onto the depth image. At operation 1006, method 902 may calculate a rectangle by establishing a convex hull of the 2D pixel positions of the corners of the brick. At operation 1008, method 902 may test each pixel in the rectangle against the bmin value and the bmax value. At operation 1010, method 902 may determine whether all pixels in the rectangle have depth values between the bmin value and the bmax value. If it is determined that all pixels in the rectangle have depth values between the bmin value and the bmax value, method 902 may accept (operation 1012) the brick. If it is determined that not all pixels in the rectangle have depth values between the bmin value and the bmax value, method 902 may perform (operation 908) a second depth image acceptance test for the brick.

[0192] Figure 17 Exemplary method 908 for performing a second depth image acceptance test in accordance with some embodiments is shown. Method 908 may begin at operation 1102 by classifying all pixels in the rectangle relative to the bmin value and the bmax value. At operation 1104, method 908 may, for example, by using Figure 18Use the table shown in to determine whether the tested brick is in front of a solid or perforated background. If it is determined that the tested brick is in front of a solid or perforated background, method 908 may accept (action 1106) the brick. If it is determined that the tested brick is not in front of a solid or perforated background, then at action 916, method 908 may pick the brick.

[0193] Figure 19A -F depicts an example of picking bricks representing scene 190 for camera frustum 192. In Figure 19A , scene 190 is represented by a single AABB 194a, which in the example shown includes 16×16 bricks. In Figure 19B , the single AABB 194a is divided into four sub-AABBs 194b, each sub-AABB 194b including 8×8 bricks. After performing the camera frustum acceptance test (e.g., method 706), one of the four sub-AABBs 194b fails the camera frustum acceptance test, so the 8×8 bricks in the failed sub-AABB 194b are picked and illustrated as white bricks. In Figure 19C , each of the three sub-AABBs 194 that pass the camera frustum acceptance test is further divided into four sub-AABBs 194c, each sub-AABB 194c including 4×4 bricks. After performing the camera frustum acceptance test (e.g., method 706), eight of the sixteen sub-AABBs 194c fail the camera frustum acceptance test, so the bricks in the failed sub-AABB 194c are picked and illustrated as white bricks. Similarly, in Figure 19D , sub-AABB 194d includes 2×2 bricks. In Figure 19E , sub-AABB 194e includes a single brick, and thus the sub-AABB 194e and the corresponding brick that pass the camera frustum test are produced as the first plurality of bricks and are illustrated as gray bricks 196f in Figure 19F . In the example shown, if picking is not performed, the world reconstruction component will compute all 256 bricks. With brick picking for the camera frustum, the world reconstruction component only needs to compute the first plurality of bricks, i.e., 34 bricks, and thus can render the result faster.

[0194] Figure 20A depicts an example of further picking 34 bricks 196f for a depth image including surface 220 by performing, for example, method 608. Figure 20BDepicts the selection results for the depth image, showing that 12 out of 34 bricks 222a in 196f passed the first depth image acceptance test (e.g., method 904), 9 out of 34 bricks 222b in 196f passed the second depth image acceptance test (e.g., method 910), and finally, after selection for the depth image including surface 220, 13 out of 34 bricks 222c were selected. As a result, in the example shown, through brick selection for the depth image, the number of bricks calculated by the world reconstruction component was further reduced to 21 bricks. It should also be understood that due to the results of the first and second depth image acceptance tests, the calculation speed of the world reconstruction component can be accelerated not only by reducing the number of bricks but also by classifying the bricks. For example, as discussed with respect to Figure 15 A constant increment can be applied to the 9 bricks 222b that did not pass the first depth image acceptance test but passed the second depth image acceptance test. Applying the constant increment in batches can further improve the calculation speed compared to applying a variable increment to each voxel. The geometry (e.g., planes) in the scene can be obtained in the XR system to support applications such as the wall for placing the virtual screen and / or the floor for navigating the virtual robot. A common representation of the geometry of the scene is a mesh, which may include groups of connected triangles with vertices and edges. By convention, the geometry in the scene is obtained by generating a mesh for the scene and searching for the geometry in the mesh, which takes some time (e.g., a few seconds) to process and does not indicate the relationship between the geometries requested by different queries. For example, the first query may be for the table plane. In response to the first query, the system can find the table plane and place the watch on the table plane. Then the second query may be for the watch. In response to the second query, the system can find all possible table planes and check each table plane to see if the watch exists until the watch is found, because the response to the first query does not indicate whether the table plane is the one for the watch.

[0195] A geometry extraction system is described herein. In some embodiments, the geometry extraction system can extract geometry when scanning a scene with a camera and / or sensor, which allows for fast and efficient extraction that can adapt to dynamic environmental changes. In some embodiments, the geometry extraction system can retain the extracted geometry in local and / or remote memory. The retained geometry can have a unique identifier such that, for example, different queries at different timestamps and / or from different applications can share the retained geometry. In some embodiments, the geometry extraction system can support different representations of the geometry according to individual queries. In the following Figures 21 - 29In the description of FIG, a plane is used as an exemplary geometric figure. It should be understood that the geometry extraction system can detect other geometric figures instead of or in addition to the plane for subsequent processing, including, for example, cylinders, cubes, lines, corners, or semantics such as glass surfaces or holes. In some embodiments, the principles described herein for geometry extraction can be applied to object extraction, etc.

[0196] Figure 21 A plane extraction system 1300 is shown according to some embodiments. The plane extraction system 1300 may include a depth fusion 1304, which may receive a plurality of depth maps 1302. The plurality of depth maps 1302 may be created by one or more users wearing a depth sensor, and / or downloaded from a local / remote memory. The plurality of depth maps 1302 may represent multiple views of the same surface. There may be differences between the plurality of depth maps, which may be reconciled by the depth fusion 1304.

[0197] In some embodiments, deep fusion 1304 may generate SDF 1306 based at least in part on method 600. Grid bricks 1308 may be generated by, for example, placing blocks on corresponding bricks (e.g., Figure 23

[0000] to

[0015] in the mesh bricks 1308). Plane extraction 1310 can detect plane planes in the mesh bricks 1308 and extract planes based at least in part on the mesh bricks 1308. Plane extraction 1310 can also extract facets for each brick based at least in part on the corresponding mesh bricks. The facet mesh may include vertices in the mesh, but not edges connecting adjacent vertices, so that storing facets consumes less storage space than a mesh. A plane data repository 1312 can retain the extracted planes and facets.

[0198] In some embodiments, the XR application can request and obtain a plane from the plane data repository 1312 via a plane query 1314, which can be sent by an application program interface (API). For example, an application can send information about its location to the plane extraction system 1300 and request all planes near it (e.g., within a five-meter radius). The plane extraction system 1300 can then search its plane data repository 1312 and send the selected plane to the application. The plane query 1314 may include information such as where the application needs a plane, what kind of plane the application needs, and / or how the plane should look (e.g., horizontal, vertical, or angled, which can be determined by examining the primitive normal of the plane in the plane data repository).

[0199] Figure 22Shows a portion 1400 of a plane extraction system 1300 according to some embodiments, which shows details regarding plane extraction 1310. Plane extraction 1310 may include dividing each grid brick 1308 into sub-bricks 1402. Plane detection 1404 may be performed on each sub-brick 1402. For example, plane detection 1404 may: compare the original normals of each grid triangle in the sub-brick; merge those grid triangles with an original normal difference less than a predetermined threshold into one grid triangle; and identify grid triangles with an area greater than a predetermined area value as a plane.

[0200] Figure 23 Is a schematic diagram showing a scene 1500 represented by bricks

[0000] to

[0015] including voxels and exemplary plane data including brick planes 1502, global planes 1504, and surface elements 1506 in the scene. Figure 23 Shows brick

[0011] divided into four sub-bricks 1508. It should be understood that a grid brick can be divided into any suitable number of sub-bricks. The granularity of the plane detected by plane detection 1404 can be determined by the size of the sub-brick, and the size of the brick can be determined by the granularity of the local / remote memory storing the volumetric 3D reconstruction data.

[0201] Returning to Figure 22 , plane detection 1404 can determine the brick plane (e.g., brick plane 1502) of each grid brick at least partially based on the detected planes for each sub-brick in the grid brick. Plane detection 1404 can also determine global planes (e.g., global plane 1504) that extend across more than one brick.

[0202] In some embodiments, plane extraction 1310 may include plane update 1406, which can update the existing brick planes and / or global planes stored in the plane data repository 1312 at least partially based on the planes detected by plane detection 1404. Plane update 1406 may include adding additional brick planes, removing some existing brick planes, and / or replacing some existing brick planes with brick planes detected by plane detection 1404 and corresponding to the same brick, so that real-time changes in the scene are retained in the plane data repository 1312. Plane update 1406 may also include aggregating the brick planes detected by plane detection 1404 into existing global planes, for example, when a brick plane is detected adjacent to an existing global plane.

[0203] In some embodiments, plane extraction 1310 may further include plane merging and splitting 1408. For example, when a brick plane is added and connects two global planes, plane merging may merge multiple global planes into one large global plane. Plane splitting may split one global plane into multiple global planes, for example, when the brick plane in the middle of the global plane is removed.

[0204] Figure 24 A data structure in the plane data repository 1312 according to some embodiments is shown. The global plane 1614 indexed by the plane ID 1612 may be at the highest level of the data structure. Each global plane 1614 may include multiple brick planes and the face elements of the bricks adjacent to the corresponding global plane, such that a brick plane can be reserved for each brick, and the global plane can be accurately presented when the edge of the global plane is not qualified as the brick plane for the corresponding brick. In some embodiments, the face elements of the bricks adjacent to the global plane rather than the face elements of all the bricks in the scene are reserved, because this is sufficient to accurately present the global plane. For example, as Figure 23 shown, the global plane 1504 extends across bricks

[0008] to

[0010] and

[0006] . The brick

[0006] has a brick plane 1502, and the brick plane 1502 is not part of the global plane 1504. Using the data structure in the plane data repository 1312, when a plane query requests the global plane 1504, the face elements of the bricks

[0006] and

[0012] are checked to determine whether the global plane 1504 extends into the bricks

[0006] and

[0012] . In the example shown, the face element 1506 indicates that the global plane 1504 extends into the brick

[0006] .

[0205] Returning to Figure 24 , the global plane 1614 may be bidirectionally associated with the corresponding brick plane 1610. Bricks may be identified by brick IDs 1602. Bricks may be divided into plane bricks 1604 that include at least one plane and non-plane bricks 1606 that do not include a plane. The face elements of both plane bricks and non-plane bricks may be reserved, depending on whether the brick is adjacent to a global plane rather than on whether the brick includes a plane. It should be understood that the planes may be continuously reserved in the plane data repository 1312 regardless of whether there is a plane query 1314 when the XR system is viewing the scene.

[0206] Figure 25Illustrated is a planar geometry extraction 1702 according to some embodiments where a plane can be extracted for use by an application when the application sends a planar query 1314 to a planar data repository 1312. The planar geometry extraction 1702 can be implemented as an API. The planar query 1314 can indicate a requested planar geometry representation, such as an outer rectangular plane, an inner rectangular plane, or a polygon plane. According to the planar query 1314, a planar search 1704 can search for and obtain planar data in the planar data repository 1312.

[0207] In some embodiments, rasterization from planar coverage points 1706 can generate planar coverage points. An example is shown in Figure 26A . There are four bricks

[0000] -

[0003] , each brick having a brick plane 1802. By projecting the boundary points of the brick planes onto a global plane 1804, planar coverage points 1806 (or "rasterized points") are generated.

[0208] Returning to Figure 25 , rasterization from planar coverage points 1706 can also generate a rasterized planar mask from the planar coverage points. According to the planar geometry representation requested by the planar query 1314, an inner rectangular plane representation, an outer rectangular plane representation, and a polygon plane representation can be extracted by inner rectangle extraction 1708, outer rectangle extraction 1710, and polygon extraction 1712, respectively. In some embodiments, the application can receive the requested planar geometry representation within a few milliseconds after sending the planar query.

[0209] Figure 26B An exemplary rasterized planar mask 1814 is shown in. According to the rasterized planar mask, various planar geometry representations can be generated. In the example shown, a polygon 1812 is generated by connecting some of the planar coverage points of the rasterized planar mask such that none of the planar coverage points in the mask are outside the polygon. An outer rectangle 1808 is generated such that the outer rectangle 1808 is the smallest rectangle surrounding the rasterized planar mask 1814. The inner rectangle 1810 is generated by: assigning "1" to bricks having two planar coverage points and "0" to bricks not having two planar coverage points to form a rasterized grid, determining groups of bricks marked as "1" and aligned on lines parallel to the edges of the bricks (e.g., bricks

[0001] ,

[0005] ,

[0009] , and

[00013] as a group, bricks

[0013] -

[0015] as a group), and generating an inner rectangle for each determined group such that the inner rectangle is the smallest rectangle surrounding the corresponding group.

[0210] Figure 27 Illustrated is a grid for a scene 1900 according to some embodiments. Figure 28A-C shows a scene 1900 represented by an outer rectangular plane, an inner rectangular plane, and a polygonal plane, respectively, according to some embodiments.

[0211] Figure 29 A less noisy 3D representation of the scene 1900 is shown, which is obtained by planarizing the mesh shown in Figure 28A -C based on the extracted plane data (e.g., the planes shown in Figure 27 -C).

[0212] Multi - level block grid simplification

[0213] In some embodiments, before the representation of the XR environment is stored or used for rendering functions (such as occlusion handling or computing physical interactions between objects in the XR environment), processing may be employed to reduce the complexity of the representation. For example, the mesh component 160d may simplify the mesh or a part of the mesh before storing the mesh or a part of the mesh as the mesh 162c in the persistent world 162.

[0214] Such processing may require performing operations hierarchically on the representation of the XR environment. These levels may include simplification operations before and after region-based operations. Like the simplification operations, the region-based operations may reduce the complexity of the representation of the XR environment. By hierarchically performing the operations in this way, the total processing for generating a simplified representation of the XR environment can be reduced while maintaining the quality of the representation of the XR environment. As a result, the simplified high-quality representation can be updated frequently, so that the XR environment can be updated frequently, thereby improving the performance of the XR system, such as by presenting a more realistic environment to the user.

[0215] The XR environment may represent the physical world, and the data representing the XR environment may be captured by one or more sensors. However, the techniques described herein can be applied to the XR environment regardless of the data source representing the environment. In some embodiments, the XR environment may be represented by a mesh including one or more points and polygons (e.g., triangles) defined by subsets of the points. A first simplification operation before the region-based operation may reduce the complexity of the representation of the environment. For example, the mesh may be simplified by reducing the number of such polygons in the mesh. As a specific example, the first simplification operation may employ a triangle reduction algorithm, which may reduce the number of triangles used to represent the XR environment.

[0216] Region-based operations can be shape detection operations that can detect one or more shapes. A common shape detection operation is a planarization operation, where the detected shape is a plane. The detected plane can represent an object or a part of an object. The detection of the plane can simplify the process of rendering an XR environment. For example, a moving object rendered in an XR environment can move in an easily computable manner when it collides with a plane. Thus, identifying a plane can simplify the subsequent rendering of the moving object compared to performing calculations with multiple polygons representing the same part of the environment. Instead of or in addition to a plane, other shapes can be detected and used in subsequent processing, including cylinders, cubes, lines, corners, or semantics such as glass surfaces or holes. Such operations can group polygons that represent the surfaces of the detected shapes.

[0217] A second simplification operation after the region-based operation can further simplify the representation of the environment, for example, by further reducing the number of polygons in the representation. The second simplification operation can focus on reducing the number of polygons within each region detected by the region-based operation.

[0218] Such processing can enable a meshing service that processes sensor data collected in a physical environment and provides a mesh for an application that generates content. In some embodiments, the processing can provide a simplified representation of virtual objects in a virtual environment.

[0219] In XR systems (such as virtual reality (VR), augmented reality (AR), and mixed reality (MR) systems), three-dimensional (3D) mesh data is typically used for a variety of purposes, including, for example, occluding virtual content in a graphics / game engine based on physical objects in the environment, or calculating the rigid body collision effects of virtual objects in the physics engine of a game engine. In some embodiments, the requirements for the mesh can be different for different uses of the mesh, and a simplified mesh can be suitable for many such uses, where some simplification techniques are more suitable for certain uses than others.

[0220] Accordingly, the processes described herein can be implemented with any of a variety of simplification techniques and / or simplification techniques that can be configured based on the intended use of the simplified mesh. The processes described herein can be used to improve the utility of a meshing service that provides simplified meshes to multiple client applications, which can use the mesh in different ways. Each client application may require a mesh with a different level of simplification. In some embodiments, an application accessing the meshing service can specify a target simplification or the mesh to be provided to it. The mesh simplification methods described herein can be used for multiple client applications, including, for example, those that perform virtual content occlusion, physical simulation, or environmental geometry visualization. The mesh processing described herein can have low latency and can be flexible because it can optimize / bias operations for different uses (e.g., flat surfaces, varying triangle counts).

[0221] The mesh simplification methods described herein can provide real-time performance (e.g., low latency to support fly (real-time) environment changes), local update capabilities (e.g., renew the portions of the mesh that have changed since the last update), and flattened surfaces (e.g., flattened flat surfaces to support robust physical simulation).

[0222] In some embodiments, the representation of the XR environment can be segmented into multiple blocks, some or all of which can be processed in parallel. In some embodiments, the resulting blocks can then be recombined. In some embodiments, the blocks can be defined with a "skirt" that overlaps with adjacent blocks. The skirt enables the blocks to be recombined at the docking of the reassembled blocks with fewer and / or less noticeable discontinuities.

[0223] Accordingly, in some embodiments, the mesh simplification method can include mesh block segmentation, pre-simplification, mesh planarization, and post-simplification. To accelerate the process, the global mesh can first be segmented into blocks of component meshes such that the mesh blocks can be addressed (e.g., processed) in parallel. Then, the mesh blocks can be extended with a skirt at the boundaries between adjacent blocks. For mesh blocks with skirts, individual mesh blocks can be simplified while the global mesh can be visually seamless although topologically disconnected.

[0224] In some embodiments, the mesh simplification method may be suitable for use by an application that represents interactions of objects in an XR environment using a simplified mesh, e.g., by making the simplification process plane-aware. To simplify the mesh, a three-step simplification process may be implemented. The mesh may first be moderately pre-simplified using a relatively high target triangle count. Then, planar regions may be detected via a region growing algorithm. The mesh may be planarized by projecting the corresponding triangles onto the detected plane. In some embodiments, the mesh may be normalized by adjusting the plane (or original) normal to be substantially perpendicular and parallel to the detected plane. Thereafter, a post-simplification process may be run on the planarized mesh. The post-simplification process may focus more on the detected planar regions, e.g., simplifying the mesh of each detected planar region to reach a desired level of complexity (e.g., metric complexity), e.g., indicated by a target value of one or more metrics.

[0225] Figure 30 Method 3000 for generating a model of an environment represented by a mesh according to some embodiments is shown. In some embodiments, method 3000 may be performed on a meshing service on an XR platform. Method 3000 may start with an input mesh representing the environment at action 3002. In some embodiments, the input mesh may have a high resolution, which may be indicated by the number of triangles. The input mesh may be generated by a reconstruction system (e.g., a volumetric 3D reconstruction system), and the input mesh may include 3D reconstruction data.

[0226] In some embodiments, the reconstruction system may generate a volumetric 3D representation of the environment, which may create a data hierarchy of 3D information of the environment captured by one or more sensors. For example, the sensor may be a depth camera, which may capture 3D information of the environment, e.g., a depth image stream with the corresponding pose of the depth camera (i.e., camera pose). The 3D information of the environment may be processed into a voxel grid. Each voxel may contain one or more signed distance functions (SDFs) that describe whether the voxel is inside or outside the geometry of an object in the environment. The voxels may be grouped into "bricks". Each brick may include, for example, a plurality of voxels of a cubic volume, e.g., 8 3 voxels. The bricks may be further grouped into "tiles". Each tile may include a plurality of bricks.

[0227] The size of the tiles may be selected to facilitate memory operations in a computing device. For example, the size may be selected based on the amount of information about the environment maintained in the active memory of the device that processes such data. For example, the system may transfer tiles between the active memory, which is typically local to the device, and other memories with greater latency (e.g., non-volatile memory or remote memory in the cloud). One or more complete or partial tiles may contain information representing "blocks" in the mesh or other representations of the environment.

[0228] In some embodiments, the volumetric 3D reconstruction system may generate the input mesh 3002 as a globally meshed that is topologically connected. In some embodiments, the volumetric 3D reconstruction system may generate the input mesh 3002 as a globally meshed that is visually seamless although topologically disconnected. For example, a topologically disconnected globally meshed may consist of multiple mesh blocks, each mesh block being generated by a block.

[0229] The reconstruction system may be configured to capture substantial details of the environment, which enables the system to distinguish differences between adjacent parts of the representation that have relatively small feature differences. Adjacent regions with different properties may be identified as different surfaces, resulting in the system identifying a large number of surfaces in the environment. However, such a system may capture details that are unnecessary but still processed for many applications. For example, for a client application requesting a mesh from a meshing service, when two triangles that form a rectangle would be a sufficient representation of a wall, the reconstruction system may unnecessarily display bumps on the wall with many triangles. In some embodiments, when requesting a mesh from a meshing service, the application may specify a target simplification level for the requested mesh. The target simplification level may be expressed as a degree of compression, the number of triangles per unit area, or in any other suitable way.

[0230] Method 3000 may effectively generate a model of the environment sufficient for a client application based on the input mesh. In action 3004, the input mesh may be segmented into one or more first mesh blocks, each first mesh block corresponding to a block in the data hierarchy of the volumetric 3D representation of the environment.

[0231] Each first mesh block may represent a part of the environment, and a complexity metric (e.g., mesh resolution) may be a first value. In some embodiments, the complexity metric of a mesh block indicates the number of triangles in that mesh block. In some embodiments, processing of the mesh blocks may be performed sequentially and / or in parallel. However, the simplification processing as described herein may be applied to the entire mesh or any suitable portion (e.g., one or more mesh blocks).

[0232] Action 3006 represents a sub-process performed on each of the multiple mesh blocks. The sub-processing may be performed independently on the multiple mesh blocks such that the processing may be easily performed in parallel on some or all of the mesh blocks. The sub-process may be performed on all mesh blocks or a subset of the mesh blocks selected for further processing. The subset of mesh blocks may be selected at least in part based on the field of view of the device on which the application requesting the simplified mesh is executing.

[0233] In operation 3006, some of the first mesh blocks may be selected based on, for example, the objects described in the first mesh block or the location of the first mesh block. For each selected first mesh block, a multi-level simplification may be performed. In some embodiments, the multi-level simplification on the selected first mesh blocks may be performed in parallel, and as a result, the simplification on the selected first mesh blocks may be completed at approximately the same time point, although this may depend on the complexity metric of each mesh block of the selected first mesh blocks.

[0234] The multi-level simplification may include a pre-simplification operation, a region-based operation (e.g., a planarization operation), and a post-simplification operation. In some embodiments, the multi-level simplification may be performed based on an input value from a client application. The input value may indicate the mesh complexity (e.g., mesh resolution) required by the client application. For each selected first mesh block, the input value from the client application may be the same or different.

[0235] In operation 3012, a pre-simplification operation may be performed on the selected first mesh blocks to generate second mesh blocks. The pre-simplification operation may reduce the complexity of the block. For a mesh block, the pre-simplification may reduce the number of polygons in the mesh block. In some embodiments, the amount of pre-simplification at operation 3012 may be configurable. For example, a target value may be provided by, for example, a client application as an input to the processing at operation 3012. The target value may be a single value or multiple values of one or more specified or predetermined metrics. One or more metrics may include, for example, an absolute triangle count, a percentage of the initial triangle count, and / or a quadratic error metric, which may measure the average squared distance between the simplified mesh and the original mesh (e.g., the input mesh 3002).

[0236] The target value may be provided in any suitable manner. For example, an instance of method 3000 may be pre-configured with the target value. In some embodiments, the target value may be provided by an application that requests a mesh from a meshing service that executes method 3000 via an API. For example, the target value of operation 3012 may be the ultimate goal requested by a rendering function (e.g., the requesting application). In some embodiments, the target value provided as an input may be adjusted or overridden to ensure that sufficient data is retained in the mesh for subsequent processing. For example, the processing in operation 3014 may require a minimum number of triangles, and if the target value is below the minimum number of triangles, the target value provided by the application may be replaced by the minimum value.

[0237] In such embodiments, the pre-simplified mesh may have values of one or more metrics such that the pre-simplified mesh may be processed faster than the input mesh with original block partitioning during a region-based operation while still containing all or most of the regions of the input mesh with original block partitioning.

[0238] Because the values of one or more metrics are not controlled, the simplified grid may be too coarse, unevenly distributed, and / or lose many regions of the original block-divided input grid required for the following region-based operations.

[0239] The complexity metric of the second grid block generated in action 3012 can be a second value, which can be less than the first value of the metric complexity. In some embodiments, a triangulation reduction algorithm can be used to perform the pre-simplification operation of action 3012.

[0240] In action 3014, a shape detection operation can be performed on the second grid block to generate a third grid block. Take the planarization operation as an example. The complexity metric of the third grid block can be a third value. In some embodiments, the third value of the metric complexity can be the same as the second value of the metric complexity. In some embodiments, the third value of the metric complexity can be less than the second value of the metric complexity. The planarization operation can include: for example, using a region growing algorithm to detect planar regions in the second grid block, projecting the grids of the detected planar regions onto the corresponding planes, adjusting the plane normals of the detected planar regions to be substantially perpendicular to the corresponding planes, and simplifying the projected grids on each corresponding plane based on, for example, a target triangle count. In some embodiments, the plane normals of the detected planar regions can be adjusted before projecting the grids of the detected planar regions onto the corresponding planes.

[0241] In action 3016, a post-simplification operation can be performed on the third grid block to generate a fourth grid block. In some embodiments, the processing at action 3014 can desirably be performed on the grid at a resolution higher than the resolution required in the simplified grid to be output from method 3000. In some embodiments, the processing at action 3016 can simplify the entire grid block to achieve a desired complexity level (e.g., metric complexity), which can be indicated by a target value of one or more metrics, and the target value can be the same as or different from the target provided to action 3012. In some embodiments, the post-simplification operation at action 3016 can focus on reducing the number of polygons in each plane detected by the planarization operation at action 3014.

[0242] The complexity measure of the fourth grid block may be a fourth value, which may be less than the third value that measures complexity. In some embodiments, the percentage reduction between the third value that measures complexity and the fourth value that measures complexity may be greater than the percentage reduction between the first value that measures complexity and the second value that measures complexity. In some embodiments, the percentage reduction between the third value that measures complexity and the fourth value that measures complexity may be at least twice as large as the percentage reduction between the first value that measures complexity and the second value that measures complexity. In some embodiments, a triangular reduction algorithm may be used to perform the post-simplification operation at action 3016. In some embodiments, the same simplification algorithm as the pre-simplification operation at action 3012 may be used to perform the post-simplification operation at action 3016.

[0243] In action 3008, the simplified selected block may be combined with other selected grid blocks similarly processed in action 3006, and / or may be combined with unselected blocks into a new grid of the environment. In action 3010, the new grid of the environment may be provided to the client application. In some embodiments, the new grid of the environment may be referred to as a simplified grid.

[0244] In some embodiments, action 3008 may be skipped. The simplified grid block may be sent directly to the client application, where the grid block may be visually seamless although topologically disconnected.

[0245] Figure 31 An example of dividing a grid representation 3100 of an environment into grid blocks according to some embodiments is shown. The grid representation 3100 may be divided into four grid blocks: grid blocks A - D. In some embodiments, the grid blocks may correspond to regions of the physical world of the environment that have the same volume. In some embodiments, the grid blocks may correspond to regions of the physical world of the environment that have different volumes. For example, when the physical world is an office, the office may be divided into multiple regions, each region being one cubic foot. One block may include a 3D representation of one region of the office.

[0246] Although the grid representation 3100 of the environment is shown in two dimensions (2D), it should be understood that the environment may be three-dimensional and correspondingly represented by a 3D grid representation. Although the grid representation 3100 of the environment is illustrated as a combination of four grid blocks, it should be understood that the environment may be represented by any suitable number (e.g., two, three, five, six, or more) of grid blocks.

[0247] The representation 3100 can be divided into four parts: for example, parts 3102, 3104, 3106, and 3108 shown by the solid line 3110. In some embodiments, parts 3102, 3104, 3106, and 3108 can be respectively designated as grid blocks A - D.

[0248] When a grid block is updated, it can continue to dock with adjacent blocks that are not updated. As a result, discontinuities can occur at the boundaries between grid blocks. If the regions represented by adjacent blocks have discontinuities, the fused grid can be interpreted in subsequent processing as indicating the presence of a crack between adjacent blocks. In some embodiments, such a crack in the representation of the physical world space can be interpreted as a space with infinite depth. Thus, the space can be an artifact of the representation of the physical world rather than an actual feature. Any application that uses such a fused grid to generate a representation of an object in the physical world may not correctly generate the output. For example, an application that renders a virtual character on a surface in the physical world may render the character as if it has fallen into the crack, which does not create the desired appearance of the object.

[0249] To reduce the appearance of such cracks, in some embodiments, a portion of an adjacent block can represent the same region of the physical world. For example, the docking region between adjacent blocks can be represented by a portion of each adjacent block, which can enable easy independent updates and / or rendering considering level of detail (LOD) (e.g., reducing the complexity of 3D reconstruction of a portion of the physical world as that portion moves out of the user's field of view). Even if one block is updated while its adjacent block is not updated, the fused grid can represent by combining data representing the docking region from both blocks. As a specific example, when fusing an updated block with an adjacent block, the physics engine can determine the overlapping region of the adjacent blocks based on, for example, which of the adjacent blocks is observable in its overlapping region. The data structure based on the block can use a side edge, a zipper, or any other suitable method to represent the docking region between adjacent blocks such that when the block is updated, it will continue to dock with the non - updated adjacent block. The appearance of this method can be to "paper over" the crack between adjacent blocks. Thus, a block can be updated independently of adjacent blocks.

[0250] In Figure 31In the example shown, regions at the boundaries of portions 3102, 3104, 3106, and 3108 may be designated as side edges, as shown by dashed line 3112. In some embodiments, each of the grid blocks A - D may include one of portions 3102, 3104, 3106, and 3108 and a corresponding side edge. For example, grid block B may include portion 3104 and side edge 3114, and side edge 3114 overlaps with boundary portions of adjacent grid blocks A, C, and D of grid block B such that cracks between the grid blocks may be masked when the blocks are joined into a single grid. Grid blocks A, C, and D may also include corresponding side edges. Thus, before returning a single connected 3D grid representation to an application, the processor may mask any cracks between the grid blocks.

[0251] In some embodiments, the block grid including the side edges may be sent directly to the application without combining it into a topologically connected global grid. The application may have a global grid composed of the block grids, which is visually seamless although topologically disconnected.

[0252] Figures 32A - 32D Shows the mesh evolution of an exemplary grid block 3201 during multi - level simplification. Grid block 3201 may include vertices 3206, edges 3208, and faces 3210. Each face may have a normal, which may be represented by multiple coordinates (e.g., shown as x, y, z in FIG. 32).

[0253] Pre - simplification operations may be performed on grid block 3201 to generate grid block 3202. An edge collapse transformation may be used. In the example shown, grid block 3202 reduces the number of faces of grid block 3201 from ten to eight. The resulting faces of grid block 3202 may each have a corresponding set of normals (e.g., x1, y1, z1; x2, y2, z2;...; x8, y8, z8).

[0254] A planarization operation may be performed on the mesh block 3202 to generate the mesh block 3203. The planarization operation may include detecting a plane region in the mesh 3202 based on, for example, a plane (or original) normal of a face. The values ​​of the plane normals x1, y1, z1 of the first face 3212 and the plane normals x2, y2, z2 of the second face 3214 may be compared. The comparison result of the plane normals of the first face and the second face may indicate an angle between the plane normals (e.g., an angle between x1 and x2). When the comparison result is within a threshold, it may be determined that the first plane and the second plane are on the same plane region. In the example shown, planes 3212, 3214, 3216, and 3218 may be determined to be on a first plane region corresponding to plane 3228; planes 3220, 3222, 3224, and 3226 may be determined to be on a second same plane region corresponding to plane 3230.

[0255] The planarization operation may further include: projecting a triangle formed by the edges of planes 3212, 3214, 3216, and 3218 onto plane 3228 as indicated by dashed line 3232; and projecting a triangle formed by the edges of planes 3220, 3222, 3224, and 3226 onto plane 3230 as indicated by dashed line 3234. The planarization operation may further include: adjusting the plane normals of planes 3212, 3214, 3216, and 3218 to be the same as the plane normal (x_a, y_a, z_a) of plane 3228; and adjusting the plane normals of planes 3220, 3222, 3224, and 3226 to be the same as the plane normal (x_b, y_b, z_b) of plane 3230.

[0256] A post-simplification operation may be performed on mesh tile 3203 to generate mesh tile 3204. In the example shown, mesh tile 3204 reduces the number of faces of mesh tile 3203 from eight to four.

[0257] Figure 33A and 33B 36A and 36B illustrate the effect of simplification, showing side by side the same portion of the physical world with and without simplification applied. These figures provide a graphical illustration of how simplification can provide usable information to operate an AR system while providing less data to process.

[0258] Figure 33A and 33B Representations of the same environment without simplification and after simplification by triangle reduction are shown. Figure 30 An example of the processing performed at the pre-simplification box 3012 and the post-simplification box 3016 in.

[0259] Figure 34A and 34BClose-up representations of the same environment are shown, respectively, without and with simplification by triangle reduction.

[0260] Figure 35A and 35B representations of the same environment are shown, respectively, without and with planarization. Such processing is an example of processing that can be performed at the planarization block 3014 in Figure 30 .

[0261] Figure 36A and 36B representations of the same environment are shown, respectively, without and with simplification by removing disconnected components. Such processing is an example of an alternative embodiment of a region-based operation that can be performed at the block 3014 in Figure 30 .

[0262] Caching and updating of dense 3D reconstruction data

[0263] In some embodiments, 3D reconstruction data can be captured, retained, and updated in the form of blocks, which can allow for local updates while maintaining adjacent consistency. The block-based 3D reconstruction data representation can be used in conjunction with a multi-level caching mechanism that efficiently retrieves, pre-fetches, and stores 3D data for AR and MR applications (including single-device and multi-device applications). For example, the volume information 162a and / or the mesh 162c ( Figure 6 ) can be stored in blocks. The component 164 can use this block-based representation to receive information about the physical world. Similarly, the sensing component 160 can store and retrieve such information in blocks.

[0264] These techniques expand the functionality of portable devices with limited computational resources to present AR and MR content in a highly realistic manner. Such techniques can be used, for example, to efficiently update and manage the output of real-time or offline reconstructions and scans in mobile devices with limited resources and connected to the Internet (continuously or discontinuously). These techniques can provide up-to-date, accurate, and comprehensive 3D reconstruction data for various mobile AR and MR applications in single-device or multi-device applications that share and update the same 3D reconstruction data. The 3D reconstruction data can be in any suitable format, including meshes, point clouds, voxels, etc.

[0265] Some AR and MR systems have attempted to simplify the rendering of MR and AR scenes by limiting the amount of 3D reconstruction data processed at any given time. Sensors used to capture 3D information may have a maximum reconstruction range, which can limit the bounding volume around the field of view of the sensor. To reduce the amount of 3D reconstruction data, some reconstruction systems keep only the regions near the field of view of the sensor in active working memory and store other data in auxiliary storage devices. For example, the regions near the field of view of the sensor are stored in the CPU memory, while other data are retained in a local cache (such as a disk) or are retained over a network to a remote storage device (such as in the cloud).

[0266] The computational cost of generating the information stored in the CPU memory, although limited, may still be relatively high. Some AR and MR systems continuously recompute the global representation of the environment seen by the reconstruction system in order to select the information to store in the CPU memory, which can be very expensive for interactive applications. Other AR and MR systems that use certain methods to compute only local updates to the connected representation can be equally expensive, especially for simplified meshes, as it requires decomposing the existing mesh, computing another mesh with the same boundaries, and then reconnecting the mesh parts.

[0267] In some embodiments, the 3D reconstruction data can be segmented into blocks. The 3D reconstruction data can be sent between storage media based on the blocks. For example, blocks can be paged out of active memory and retained in a local or remote cache. The system can implement a paging algorithm, where the active memory associated with a wearable device (e.g., a head-mounted display device) stores blocks that represent a portion of the 3D reconstruction of the physical world in the field of view of the user of the wearable device. The wearable device can capture data about the portion of the physical world commensurate with the field of view of the user of the wearable device. As the physical world changes in the user's field of view, the blocks representing that region of the physical world can be in active memory and can be easily updated from the active memory. As the user's field of view changes, blocks representing the region of the physical world that has moved out of the user's field of view can be moved to the cache, such that blocks representing the region of the physical world that has entered the user's field of view can be loaded into active memory.

[0268] In some embodiments, a coordinate system can be created for a portion of the physical world to be 3D reconstructed. Each block in the 3D representation of that portion of the physical world can correspond to a different region of the physical world that can be identified using the coordinate system.

[0269] In some embodiments, when a block is updated, the updated block may continue to dock with adjacent blocks that may not have been updated yet. If the regions represented by the adjacent blocks do not overlap, there may be cracks in the merged grid of the adjacent blocks. In some embodiments, such cracks in the representation of the physical world space may be interpreted as spaces with infinite depth. Thus, the spaces may be artifacts of the representation of the physical world rather than actual features. Any application that uses such merged grids to generate a representation of an object in the physical world may not correctly generate the output. For example, an application that renders a virtual character on a surface in the physical world may render the character as if it has fallen into a crack, which does not create the desired appearance of the object. Therefore, in some embodiments, a portion of the adjacent blocks may represent the same region of the physical world, e.g., the docking between adjacent blocks may represent the same region of the physical world, which can enable easy independent updates and / or rendering considering level of detail (LOD) (e.g., reducing the complexity of 3D reconstruction of a portion of the physical world as that portion moves out of the user's field of view). For example, when a block is updated, its adjacent blocks may not be updated. When merging the updated block with the adjacent blocks, the physics engine may determine the overlapping regions of the adjacent blocks based on, for example, which of the adjacent blocks is observable in their overlapping regions. Based on the data structure of the blocks, a side edge, zipper, or any other suitable method may be employed to represent the docking region between adjacent blocks such that when a block is updated, it will continue to dock with the non-updated adjacent blocks. The appearance of such a method may be "covering" over the crack between the adjacent blocks. Thus, the blocks that have changed can be updated independently of the adjacent blocks.

[0270] In some embodiments, these techniques may be used in an AR and / or MR "platform" that receives and processes data from sensors worn by one or more users. The sensor data may be used to create and update 3D reconstruction data representing a portion of the physical world encountered by the user. While the sensors are capturing and updating data, the reconstruction service may continuously reconstruct a 3D representation of the physical world. One or more techniques may be used to determine the blocks affected by changes in the physical world, and those blocks may be updated. The 3D reconstruction data may then be provided to an application that uses the 3D reconstruction data to render a scene to depict virtual reality objects placed in or interacting with objects in the physical world. The data may be provided to the application via an application programming interface (API). The API may be a push or pull interface that pushes the data to the application when relevant portions change or in response to a request from the application for the latest information.

[0271] In an example of a pull interface, when an application requests 3D reconstruction data of the physical world, the reconstruction service can determine the appropriate version of each block to be provided to the application, enabling the reconstruction service to start from the latest block. The reconstruction service can, for example, search for previously reserved blocks. A single-device system can enable a single device to contribute 3D reconstruction data of the physical world. In a single-device system, if the requested region of the physical world is within the active region (e.g., the region within the current field of view of the device) or extends beyond the active region, the reserved blocks can be directly used as the latest blocks because these reserved blocks will not be updated as they are retained when the region moves out of the device's field of view. On the other hand, a multi-device system can enable multiple devices to provide 3D reconstruction data of the physical world, such as using cloud persistence or peer-to-peer local caching of blocks. Each device can update the regions within its active region that can be reserved. The multi-device system can create a coordinate system such that blocks generated by different devices can be identified using the coordinate system. Thus, if those updates are made after any version made by the first device, the blocks requested for the data generated by the application for the first device can be based on updates from other devices. Blocks constructed using data from the first device and other devices can be merged by using the coordinate system.

[0272] The selected blocks can be used to provide 3D reconstruction data of the physical world in any suitable format, but a mesh is used here as an example of a suitable representation. The mesh can be created by processing image data to identify points of interest in the environment (e.g., the edges of objects). These points can be connected to form the mesh. A group of points (usually three points) in the mesh associated with the same object or a part thereof defines the surface of the object or the part. The information stored with the group of points describes the surface in the environment. This information can then be used in various ways to render and / or display information about the environment. The selected blocks can be used to provide 3D reconstruction data in any suitable manner. In some embodiments, the latest blocks can be provided. In some embodiments, the latest blocks can be used to determine whether the blocks need to be updated.

[0273] For example, in some embodiments, in a multi-device system, when the blocks requested by the application have been identified, the reconstruction service can check the blocks reserved by other devices to determine whether there are any significant updates (e.g., via a geometric change threshold or a timestamp), re-run meshing on the changed blocks, and then retain these updated mesh blocks.

[0274] In some embodiments, when a block set requested by an application has been identified, if the application has requested a connected mesh, the block set can be processed as a global mesh, which can be topologically connected or visually seamless although topologically disconnected using any suitable technique (e.g., edges and zippers).

[0275] In some embodiments, when a block change occurs, an application (e.g., a graphics / game engine) can update its internal blocks (e.g., blocks stored in active memory and / or local cache). The reconstruction service can know which blocks the application has and can thus calculate which other (e.g., adjacent) blocks need to be updated in the engine to maintain correct overlap with the edges / zippers when the blocks in the field of view are updated.

[0276] In some embodiments, an AR and / or MR platform can be implemented to support, for example, the execution of AR and / or MR applications on a mobile device. An application that executes or generates data for presentation on a user interface can request 3D reconstruction data representing the physical world. The 3D reconstruction data can be provided from active memory on the device and can be updated with 3D reconstruction data representing the physical world in the device's field of view as the user changes their field of view. The 3D reconstruction data in active memory can represent the active area of the mobile device. In some embodiments, 3D reconstruction data outside the active area of the device can be stored in other memories, such as in a local cache on the device or coupled to the device via a low-latency connection. In some embodiments, 3D reconstruction data outside the active area of the device can also be stored in a remote cache, such as in the cloud, which can be accessed by the device via a higher-latency connection. When the user changes their field of view, the platform can access / load 3D reconstruction data from the cache to add to active memory to represent the area moving into the user's field of view. The platform can move other data representing the area moving out of the user's field of view into the cache.

[0277] Prediction of the movement of the device (which may cause areas to move into the device's field of view while other areas move out of the device's field of view) can be used to initiate the transfer of 3D reconstruction data between active memory and the cache. The prediction of the movement can be used, for example, to select 3D reconstruction data for transfer into and / or out of active memory. In some embodiments, the predicted movement can be used to transfer 3D reconstruction data into and / or out of the local cache by retrieving or transferring 3D reconstruction data from / to the remote cache. Exchanging 3D reconstruction data between the local cache and the remote cache based on the user's predicted movement can ensure that the 3D reconstruction data is available with low latency to move into active memory.

[0278] In embodiments where regions of the physical world are represented by blocks, initiating the transmission of blocks may require a prior request for blocks representing regions predicted to enter the user's field of view. For example, if the platform determines, based on sensor data or other data, that the user is walking at a particular speed in a particular direction, it can identify regions that may enter the user's field of view and transmit blocks representing those regions to a local cache on the mobile device. If the mobile device is a wearable device, such as a pair of glasses, predicting movement may require receiving sensor data indicating the position, orientation, and / or rotation of the user's head.

[0279] Figure 37 FIG. 3700 shows a system 3700 for enabling an interactive X reality environment for multiple users. The system 3700 may include a computing network 3705 that includes one or more computer servers 3710 connected by one or more high-bandwidth interfaces 3715. The servers in the computing network need not be co-located. Each of the one or more servers 3710 may include one or more processors for executing program instructions. The servers also include a memory for storing the program instructions and data used and / or generated by processes executed by the servers under the guidance of the program instructions. The system 3700 may include one or more devices 3720 that include, for example, an AR display system 80 (e.g., Figure 3 the viewing optical assembly 48 in FIG. B).

[0280] The computing network 3705 transfers data between the servers 3710 and between the servers and the devices 3720 via one or more data network connections 3730. Examples of such data networks include, but are not limited to, any and all types of public and private data networks, mobile and wired, including, for example, the interconnection of many such networks, commonly referred to as the Internet. The figure is not intended to imply any particular media, topology, or protocol.

[0281] In some embodiments, the devices may be configured to communicate directly with the computing network 3705 or any of the servers 3710. In some embodiments, the devices 3720 may communicate with a remote server 3710 and, optionally, communicate locally with other devices and the AR display system via a local gateway 3740 for processing and / or transferring data between the network 3705 and the one or more devices 3720.

[0282] As shown in the figure, the gateway 3740 is implemented as a separate hardware component, which includes a processor for executing software instructions and a memory for storing software instructions and data. The gateway has its own wired and / or wireless connection to the data network for communicating with the server 3710 including the computing network 3705. In some embodiments, the gateway 3740 can be integrated with the device 3720 worn or carried by the user. For example, the gateway 3740 can be implemented as a downloadable software application installed and running on the processor included in the device 3720. In one embodiment, the gateway 3740 provides access for one or more users to the computing network 3705 via the data network 3730. In some embodiments, the gateway 3740 may include communication links 76 and 78.

[0283] Each of the servers 3710 includes, for example, a working memory and a storage device for storing data and software programs, a microprocessor for executing program instructions, a graphics processor, and other special processors for rendering and generating graphics, images, videos, audio, and multimedia files. The computing network 3705 may also include a device for storing data accessed, used, or created by the servers 3710. In some embodiments, the computing network 3705 may include a remote processing module 72 and a remote data repository 74.

[0284] The software programs running on the servers as well as the optional device 3720 and gateway 3740 are used to generate a digital world (also referred to herein as a virtual world) through which the user interacts with the device 3720. The digital world is represented by data and processes that describe and / or define virtual, non-existent entities, environments, and conditions that can be presented to and interacted with by the user via the device 3720. For example, a certain type of object, entity, or item that appears physically present when instantiated in a scene being viewed or experienced by the user may include descriptions of its appearance, its behavior, how it allows the user to interact with it, and other characteristics. The data for creating the environment of the virtual world (including virtual objects) may include, for example, atmospheric data, terrain data, weather data, temperature data, location data, and other data for defining and / or describing the virtual environment. Additionally, the data defining the various conditions that govern the operation of the virtual world may include, for example, physical laws, time, spatial relationships, and other data for defining and / or creating the conditions governing the operation of the virtual world (including virtual objects).

[0285] Unless the context otherwise indicates, this document will generally refer to entities, objects, conditions, characteristics, behaviors, or other features in the digital world as objects (e.g., digital objects, virtual objects, rendered physical objects, etc.). An object can be any type of living or non-living object, including but not limited to buildings, plants, vehicles, people, animals, organisms, machines, data, videos, text, pictures, and other users. Objects can also be defined in the digital world to store information about items, behaviors, or conditions that actually exist in the physical world. Data that describes or defines an entity, object, or item or stores its current state is generally referred to as object data in this document. This data is processed by the server 3710, or by the gateway 3740 or the device 3720 depending on the implementation, to instantiate an instance of the object and render the object in a manner that enables a user to experience it through the device 3720.

[0286] Programmers who develop and / or curate the digital world can create or define objects and the conditions for instantiating the objects. However, the digital world can allow others to create or modify the objects. Once an object is instantiated, one or more users who experience the digital world can be allowed to change, control, or manipulate the state of the object.

[0287] For example, in one embodiment, the development, generation, and management of the digital world are generally provided by one or more system management programmers. In some embodiments, this can include the development, design, and / or execution of storylines, themes, and events in the digital world, as well as the distribution of the narrative through various forms of events and media (e.g., movies, digital, web, mobile, augmented reality, and live entertainment). The system management programmers can also handle the technical management, moderation, and curation of the digital world and the associated user community, as well as other tasks typically performed by network administrators.

[0288] Users interact with one or more digital worlds using some type of local computing device, typically referred to as the device 3720. Examples of such devices include but are not limited to smartphones, tablet devices, head-up displays (HUDs), gaming consoles, or any other device that can deliver data to a user and provide an interface or display, as well as combinations of these devices. In some embodiments, the device 3720 can include or communicate with local peripheral devices or input / output components (e.g., keyboards, mice, joysticks, game controllers, haptic interface devices, motion capture controllers, audio devices, voice devices, projector systems, 3D displays, and holographic 3D contact lenses).

[0289] Figure 38 is a schematic diagram showing an electronic system 3800 according to some embodiments. In some embodiments, the system 3800 can be Figure 37Part of system 3700. System 3800 may include a first device 3810 (e.g., a first portable device of a first user) and a second device 3820 (e.g., a second portable device of a second user). Devices 3810 and 3820 may be, for example, Figure 37 Devices 3720 and AR display system 80. Devices 3810 and 3820 may communicate with cloud cache 3802 via networks 3804a and 3804b, respectively. In some embodiments, cloud cache 3802 may be implemented in the memory of one or more servers 3710 of Figure 37 . Networks 3804a and 3804b may be Figure 37 Examples of data network 3730 and / or local gateway 3740.

[0290] Devices 3810 and 3820 may be separate AR systems (e.g., device 3720). In some embodiments, devices 3810 and 3820 may include AR display systems worn by their respective users. In some embodiments, one of devices 3810 and 3820 may be an AR display system worn by a user; the other may be a smartphone held by the user. Although two devices 3810 and 3820 are shown in this example, it should be understood that system 3800 may include one or more devices, and the one or more devices may run the same type of AR system or different types of AR systems.

[0291] Devices 3810 and 3820 may be portable computing devices. The first device 3810 may include, for example, a processor 3812, a local cache 3814, and one or more AR applications 3816. The processor 3812 may include a computing portion 3812a configured to execute computer-executable instructions based at least in part on data collected by one or more sensors (e.g., Figure 3 Depth sensor 51, world camera 52, and / or inertial measurement unit 57 of B) to provide a 3D representation (e.g., 3D reconstruction data) of a portion of the physical world.

[0292] The computing portion 3812a may represent the physical world as one or more blocks. Each block may represent objects in different regions of the physical world. Each region may have a corresponding volume. In some embodiments, the blocks may represent regions having the same volume. In some embodiments, the blocks may represent regions having different volumes. For example, when the physical world is an office, the office may be divided into cubes, each cube being one cubic foot. A block may include a 3D representation (e.g., 3D reconstruction data) of one cube of the office. In some embodiments, the office may be divided into regions having various volumes, and each volume may include a similar amount of 3D information (e.g., 3D reconstruction data) such that the data size of the 3D representation of each region may be similar. The representation may be formatted to facilitate further processing such as occlusion handling to determine whether a virtual object is occluded by a physical object or physical process, and to determine how a virtual object should move or deform when interacting with physical objects in the physical world. The blocks may be formatted, for example, as grid blocks where features (e.g., corners) of objects in the physical world become points in the grid blocks, or are used as points to create the grid blocks. Connections between points in the grid may indicate groups of points on the same surface of a physical object.

[0293] Each block may have one or more versions, each version containing data representing its corresponding region based on data at a point in time (e.g., volumetric 3D reconstruction data such as voxels, and / or a mesh that may represent surfaces in the region represented by the corresponding block). When additional data becomes available, the computing portion 3812a may create a new version of the block, such as data indicating that an object in the physical world has changed or additional data from which a more accurate representation of the physical world may be created. The additional data may come from sensors on the device (e.g., devices 3810 and / or 3820). In some embodiments, the additional data may come from remote sensors and may be obtained, for example, via a network connection.

[0294] The processor 3812 may also include an active memory 3812b that may be configured to store blocks in the field of view of the device. In some embodiments, the active memory 3812b may store blocks outside the field of view of the device. In some embodiments, the active memory 3812b may store blocks adjacent to blocks in the field of view of the device. In some embodiments, the active memory 3812b may store blocks predicted to be in the field of view of the device. In some embodiments, if a block is within the field of view of the device at that time, the processor 3812 maintains the block in the active memory 3812b. The field of view may be determined by the imaging area of one or more sensors. In some embodiments, the field of view may be determined by the amount of the physical world presented to a user of the device or perceivable by an average user without using an AR system. Thus, the field of view may depend on the position of the user in the physical world and the orientation of the wearable component of the device.

[0295] If a block becomes outside the field of view of the device 3810 as the user moves, the processor 3812 may consider the block to be inactive. Inactive blocks may be paged out of the active memory to a cache. The cache may be a local cache or a remote cache. In Figure 38 embodiments, the block is first paged out to the local cache 3814 via the local gateway 3818b. In some embodiments, the local cache 3814 may be the only available cache.

[0296] In some embodiments, there may be a remote cache accessible via a network. In the illustrated embodiment, the cloud cache (e.g., remote cache) 3802 accessed via the network 3804a is an example of a remote cache. The processor 3812 may manage when to move blocks between the local cache 3814 and the cloud cache 3802. For example, when the local cache 3814 is full, the processor 3812 may page out a block to the cloud cache 3802 via the network 3804a. Since blocks in the local cache 3814 are accessible to render a scene with lower latency than blocks in the cloud cache, the processor 3812 may use an algorithm designed to keep the blocks most likely to become active in the local cache 3814 to select the blocks to page out of the local cache 3814. Such an algorithm may be based on the time of access. In some embodiments, the algorithm may be based on a prediction of the movement of the device that will change the field of view of the device.

[0297] Applications that render scenes (e.g., computer games) can obtain information representing portions of the physical world that affect the scene to be rendered. Application 3816 can obtain active chunks from active memory 3812b via local gateway 3818a. In some embodiments, local gateway 3818a can be implemented as an application programming interface (API) such that processor 3812 implements a "service" for application 3816. In embodiments where the data of the physical world is represented as a mesh, the service can be a "meshing service". The API can be implemented as a push or pull interface, or can have attributes of both. In a pull interface, for example, application 3816 can indicate the portions of the physical world for which it needs data, and the service can provide data for those portions. For example, in a push system, the service can provide data about portions of the physical world when such data changes or becomes available.

[0298] The portions of the physical world for which data is provided can be limited to the relevant portions indicated by application 3816, e.g., data within the field of view of the device or data representing portions of the physical world within a threshold distance of the device's field of view. In a pull / push system, application 3816 can request data for a portion of the physical world, and the service can provide data about the requested portion plus any adjacent portions in which the data has changed. To limit the information to that which has changed, in addition to maintaining chunks that describe the physical world, the service can also keep track of which versions of the chunks have been provided to each application 3816. The operation of determining which portion of the representation of the physical world will be updated and where the update occurs can be partitioned between application 3816 and the service in any suitable way. Similarly, the location where the updated data is incorporated into the representation of the physical world can be partitioned in any suitable way.

[0299] In some embodiments, when the sensors are capturing and updating data, the reconstruction service can continuously reconstruct a 3D representation of the physical world. The data can then be provided to application 3816 that uses the 3D reconstruction data to render a scene depicting the physical world and virtual reality objects located in or interacting with objects in the physical world. The data can be provided to application 3816 via an API, which can be implemented as a push interface that pushes the data to application 3816 when the relevant portion changes, or a pull interface that responds to a request from application 3816 for the latest information, or both.

[0300] For example, the application 3816 can operate on a grid representation of a portion of the physical world that constitutes a 45-degree view at a distance of 10 meters from an origin defined by the current position of the device and the direction the device is facing. When the area changes or data indicating physical changes within the area becomes available, the grid can be calculated to represent the area. The grid can be calculated within the application 3816 based on data provided by the service, or it can be calculated within the service and provided to the application 3816. In both cases, the service can store information in the physical world, thus simplifying the calculation of the grid. As described herein, blocks with zippers, side edges, or other techniques implemented to facilitate "masking" over cracks between adjacent blocks can be used so that only the changing portions of the representation of the physical world are processed. The changing portions of the representation of the physical world can then replace the corresponding portions in the previous representation of the physical world.

[0301] Efficiently accessing the representation of the portion of the physical world used to generate the grid that will be used by the application 3816 to render a scene to the user can reduce computer resources, thus making the XR system easier to implement on portable devices or other devices with limited computing resources, and can produce a more realistic user experience because the XR scene better matches the physical world. Accordingly, instead of or in addition to using blocks with side edges, zippers, or other techniques to facilitate masking over cracks between blocks, as described elsewhere herein, the algorithms used to page blocks into and out of active memory and / or the local cache can be selected to reduce the access time to the blocks required to calculate the grid at any given time.

[0302] In Figure 38 an exemplary embodiment, the gateway 3818a is a pull interface. When the AR application 3816 requests information about a region of the physical world, but the block representing the region is not in the active memory 3812b, the processor 3812 can search for the block retained in the local cache 3814. If the processor 3812 cannot find the block in both the active memory 3812b and the local cache 3814, the processor 3812 can search for the block retained in the cloud cache 3802. Since accessing the active memory 3812b has a lower latency than accessing the data in the local cache 3814, and accessing the data in the local cache 3814 has a lower latency than accessing the data in the cloud cache 3802, the overall speed of generating the grid can be increased by a service implementing a paging algorithm that loads the block into the active memory before the block is requested, or moves the block from the cloud cache 3802 to the local cache 3814 before the block is requested.

[0303] Similar to the first device 3810, the second device 3820 may include a processor 3822 having a computing portion 3822a and an active memory 3822b, a local cache 3824, and one or more AR applications 3826. The AR applications 3826 may communicate with the processor 3822 via a local gateway 3828a. The local cache 3824 may communicate with the processor 3822 via a local gateway 3828b.

[0304] Thus, the cloud cache 3802 may retain chunks sent from both devices 3810 and 3820. The first device 3810 may access in the cloud cache 3802 chunks captured and sent from the second device 3820; similarly, the second device 3820 may access in the cloud cache 3802 chunks captured and sent from the first device 3810.

[0305] Devices 3801 and 3802 are provided as examples of portable AR devices. Any suitable device, such as a smartphone, may be similarly used and implemented.

[0306] Figure 39 is a flowchart showing a method 3900 of an operating system (e.g., system 3700) according to some embodiments. At action 3902, the device may capture 3D information of a physical world including objects in the physical world and represent the physical world as chunks including 3D reconstruction data. In some embodiments, the 3D reconstruction data may be captured by a single system and used only for rendering information on that system. In some embodiments, the 3D reconstruction data may be captured by multiple systems and may be used for rendering information on any one of the multiple systems or any other system. In these embodiments, the 3D reconstruction data from multiple systems may be combined and accessed by multiple systems or any other system.

[0307] For example, several users each wearing an AR system may set their devices to enhanced mode while exploring a warehouse. The sensors of each device may be capturing 3D information of the warehouse (e.g., 3D reconstruction data including depth maps, images, etc.) in the sensor's field of view, the field of view including objects in the warehouse (e.g., tables, windows, doors, floors, ceilings, walls). Each device may divide the warehouse into regions having corresponding volumes and represent the respective regions as chunks. These chunks may have versions. Each version of a chunk may have values representing the objects in the region of the physical world at a certain point in time.

[0308] When an application needs information about the physical world, the version of the block representing that part of the physical world can be selected and used to generate the information. Although this selection process can be performed by any suitable processor or distributed across any suitable processors, according to some embodiments, the process can be performed locally on the device on which the application requesting the data is executing.

[0309] Thus, at action 3904, a processor (e.g., processor 3812 or 3822) can respond to a request for 3D reconstruction data from an application (e.g., AR application 3816 or 3826). In some embodiments, regardless of whether the application requests 3D reconstruction data, the device can continue to capture 3D information including 3D reconstruction data about the physical world and represent the physical world as blocks of 3D reconstruction data. The 3D reconstruction data can be used to create a new version of the block.

[0310] If the application requests 3D reconstruction data, the process can proceed to action 3906, where the processor can identify a subset of blocks corresponding to the part of the physical world required to deliver the 3D reconstruction data, according to the request. The identification of the blocks can be based on data collected by sensors (e.g., depth sensor 51, world camera 52, inertial measurement unit 57, global positioning system, etc.), for example. A multi-device system can create a common coordinate system, such that the common coordinate system can be used to create blocks generated by different devices associated with the corresponding part of the physical world, without regard to which device provided the 3D reconstruction data to reconstruct the part of the physical world represented by the block. As an example of how a common coordinate system can be created, data from devices in a roughly the same vicinity can be routed to the same server or one or more servers for processing. There, the data from each device can initially be represented in a device-specific coordinate system. Once enough data has been collected from each device to identify features in the common part of the physical world, those features can be associated, thus providing a transformation from one device-specific coordinate system to other device-specific coordinate systems. One of these device-specific coordinate systems can be designated as the common coordinate system, and the other coordinate systems can be transformed thereto, and this coordinate system can be used to transform data from the device-specific coordinate systems to the coordinate system designated as the common coordinate system. Regardless of the specific mechanism for creating the common coordinate system, once the common coordinate system is created, the 3D reconstruction data requested by an application that generated data for a first device can be based on updates from other devices (if those updates were made after any version made by the first device). Blocks from the first device and other devices can be merged by using, for example, the common coordinate system.

[0311] The specific processing in operation 3906 can depend on the nature of the request. In some embodiments, if an application requesting 3D reconstruction data maintains its own information about the blocks and requests a specific block, then the request for 3D reconstruction data at operation 3904 can include a reference to a specific subset of blocks, and identifying the subset of blocks at operation 3906 can include determining the subset of blocks corresponding to the specific subset of blocks. In some embodiments, the request for 3D reconstruction data at operation 3904 can include a reference to the field of view of the device on which the application is executing, and identifying the subset of blocks at operation 3906 can include determining the subset of blocks corresponding to the reference field of view of the device.

[0312] Regardless of the manner in which the blocks are identified / determined, at operation 3908, the processor can select a version of the blocks of the subset of blocks. The selection can be based on one or more criteria. The criteria can be, for example, based on the most recent version of the blocks from an available source. In the illustrated embodiment, the version of the blocks can be stored in an active memory, a local cache, or a remote cache. Operation 3908 can include, for example, selecting the version in the active memory (if available), or if not available, selecting the version in the local cache (if available), or selecting the version from the remote cache (if available). If no version of the blocks is available, the selection may need to generate the blocks, for example, based on data (such as 3D reconstruction data) collected using sensors (such as depth sensor 51, world camera 52, and / or inertial measurement unit 57). Such an algorithm for block selection can be used in a system for managing, for example, background processes, the versions of the blocks stored at each possible location. This is described below in conjunction with Figure 41 an exemplary management process.

[0313] At operation 3910, the processor can provide information based on the selected version of the blocks to the application. The processing at operation 3910 may simply involve providing the blocks to the application, which can be appropriate when the application directly uses the blocks. In the case where the application receives a mesh, the processing at operation 3910 may involve generating a mesh from the blocks and / or the subset of blocks and providing the mesh or any appropriate portion of the mesh to the application.

[0314] Figure 40 is a flowchart showing details of blocks that capture 3D information about an object in the physical world and represent the physical world as 3D reconstruction data according to some embodiments. In some embodiments, Figure 40 is a flowchart showing details of operation 3902 of Figure 39 At operation 4002, one or more sensors (such as depth sensor 51, world camera 52, inertial measurement unit 57, etc.) of a system (such as system 3700) capture 3D information about an object in the physical world, where the physical world includes the object in the physical world.

[0315] At operation 4004, a processor of the system (e.g., processor 3812 or 3822) may create a version of a block that includes 3D reconstruction data of the physical world based on 3D information captured by one or more sensors. In some embodiments, each block may be formatted as one or more portions of a mesh. In some embodiments, other representations of the physical world may be used.

[0316] A block may have versions such that whenever information about a region of the physical world is captured by any device, a new version of the block may be stored. Each version of the block may have 3D reconstruction data that includes values representing objects in a region of the physical world at a certain point in time. In some embodiments, such processing may be performed locally on the device, resulting in a new version of the block being stored in active memory. In some embodiments, in a multi-device system, similar processing may be performed in a server (e.g., Figure 37 server 3710), and the server may manage the versions of the blocks such that the most recent version available in its remote cache is provided whenever requested by any device.

[0317] Because these blocks represent the physical world, and most of it will remain unchanged, a new version of a block may not necessarily be created when new 3D reconstruction data representing a corresponding region of the physical world becomes available. Instead, managing the versions of the blocks may require processing the 3D reconstruction data representing the physical world to determine whether there have been sufficient changes since the last version of the blocks representing those regions of the physical world to warrant a change. In some embodiments, sufficient changes may be indicated by the size of a block metric becoming higher than a threshold since the last version has been stored.

[0318] In some embodiments, when a block is requested, other criteria may be applied to determine which version of the block to provide as the current version, such as the version with the minimum value of a metric indicating the integrity or accuracy of the data in the block. Similar processing may be performed on each device, resulting in block versions stored in the local cache on the device.

[0319] One or more techniques can be used to manage the versions of the blocks available for services on each device. For example, if an acceptable version of a block that has already been computed exists, rather than creating a new version of the block based on sensor data, the processor can access the previously stored block. Such access can be performed efficiently by managing the storage of the versions of the blocks. At action 4006, the processor of the device can page out the version of the block of the 3D reconstruction data of the physical world from the active memory (e.g., active memory 3812b or 3822b). Paging can include the processor accessing sensor data to continuously update the block in the active memory / local cache / cloud cache, for example, according to the field of view of the device. When the field of view of the device changes, the block corresponding to the new field of view can be transferred (e.g., paged) from the local cache and / or cloud cache to the active memory, and the block corresponding to the region just outside the new field of view (e.g., the block adjacent to the block in the new field of view) can be transferred (e.g., paged) from the active memory and / or cloud cache to the local cache. For example, at action 4008, the version of the block paged out by the processor can be retained in the local memory (e.g., local cache 3814 or 3824) and / or remote memory (e.g., cloud cache 3802). In some embodiments, when, for example, each new version of a block is created on the device, the version can be sent to the remote memory so that other users can access it.

[0320] Figure 41 is a flowchart showing an exemplary process for performing an action of selecting a version of a block that represents a subset of blocks according to some embodiments. In some embodiments, Figure 41 is a flowchart showing Figure 39 the details of action 3908 of. To select the version of each block in the subset of blocks, at action 4102, the processor (e.g., processor 3812 or 3822) can query whether the latest version is stored in the active memory (e.g., active memory 3812b or 3822b). In some embodiments, it can be determined whether the version is the latest by comparing the value attached to the version (e.g., geometric change size, timestamp, etc.) with the data collected by sensors (e.g., depth sensor 51, world camera 52, and / or inertial measurement unit 57). In some embodiments, a comparison can be made between the current sensor data and the version of the block stored in the active memory. Based on the degree of difference, which can represent the change in the physical world or, for example, the quality of the version in the active memory, the version in the active memory can be considered the latest.

[0321] If the latest version is stored in the active memory, the process proceeds to operation 4104, where the latest version is selected. If the latest version is not stored in the active memory, the process proceeds to operation 4106, where the processor can query whether the latest version is stored in the local memory (e.g., local caches 3814 or 3824). The query can be performed using the criteria described above in connection with operation 4102 or any other suitable criteria. If the latest version is stored in the local memory, then in operation 4108, the latest version is selected.

[0322] If the latest version is not stored in the local memory, then in operation 4110, the processor can query whether the latest version is stored in the remote memory (e.g., cloud cache 3802). The query can also be performed using the criteria described above in connection with operation 4102 or any other suitable criteria. If the latest version is stored in the remote memory, then in operation 4112, the latest version is selected.

[0323] If the latest version is not stored in the remote memory, the process can proceed to operation 4114, where the processor of the device can generate a new version of the block based on 3D information (e.g., 3D reconstruction data) captured by the sensors. In some embodiments, in operation 4116, the processor can identify adjacent blocks of the block with the new version and update the identified adjacent blocks according to the new version of the block.

[0324] Figure 42 FIG. 10 is a flow diagram of a method 4200 of an operating system according to some embodiments. In method 4200, instead of pulling blocks into the active memory and / or local caches when the device needs those blocks, paging can be managed based on the projection of the device's field of view based on device movement.

[0325] Similar to operation 3902, in operation 4202, sensors on the device can capture 3D information about the physical world including objects in the physical world and represent the physical world as blocks including 3D reconstruction data.

[0326] At action 4204, a processor (e.g., processor 3812 or 3822) may calculate a region of the physical world based at least in part on the output of sensors, which a portable pointable component (e.g., depth sensor 51, world camera 52, and / or inertial measurement unit 57) will be pointed at a region at some future time. In some embodiments, the processor may perform the calculation based on motion data from an inertial sensor or the result of an analysis of a captured image. In a simple calculation, for example, to obtain a quick result, the processor may perform the calculation based on the translation and rotation of the user's head. When applying a more comprehensive algorithm, the processor may perform the calculation based on objects in the scene. For example, the algorithm may consider that a user walking towards a wall or a table is unlikely to walk through the wall or the table.

[0327] At action 4206, the processor may select a block based on the calculated region. At action 4208, the processor may use the selected block to update an active memory (e.g., active memory 3812b or 3822b). In some embodiments, the processor may select a block based on Figure 41 the flowchart of. At action 4210, the processor may select a block from the active memory to provide to an application (e.g., application 3816 or 3826) via an API, for example, based on changes to each block since the block's version was last provided to the application.

[0328] In some embodiments, at action 4206, the processor may request the selected block from a remote memory (e.g., cloud cache 3802) and update the information stored in a local cache (e.g., 3814 or 3824) such that the local cache stores the selected block. Action 4206 may be similar to Figure 39 action 3908 described in.

[0329] The block-based processing as described above may be based on blocks that allow parts of a 3D representation to be processed separately and then combined with other blocks. According to some embodiments, the blocks may be formatted such that when a block changes, the changed representation largely or completely preserves the value of the block at the interface with other blocks. Such processing enables a changed version of a block to be used with versions of adjacent blocks that have not changed, without creating unacceptable artifacts in a scene rendered based on the changed and unchanged blocks. Figures 43A - 48 Such a block is shown.

[0330] A 3D representation of the physical world can be provided through volumetric 3D reconstruction, which can create a hierarchy of 3D reconstruction data of the physical world captured by sensors. For example, the sensor can be a depth camera, which can capture 3D information of the physical world, such as a stream of depth images with the respective poses of the depth camera (i.e., camera poses). The 3D information of the physical world can be processed into a voxel grid. Each voxel can contain one or more signed distance functions (SDFs), which describe whether the voxel is inside or outside the geometry of an object in the physical world. The voxels can be grouped into "bricks". Each brick can include, for example, a plurality of voxels of a cubic volume, such as 8 3 individual voxels. The bricks can be further grouped into "tiles". Each tile can include a plurality of bricks.

[0331] In some embodiments, the voxel grid can be mapped to conform to the memory structure. The tiles can correspond to the storage pages of the storage medium. The size of the tiles can be variable, for example, depending on the size of the storage pages of the storage medium used. Thus, the 3D reconstruction data can be sent between storage media (e.g., the active memory and / or local memory of the device, and / or remote memory in the cloud) based on the tiles. In some embodiments, one or more tiles can be processed to generate blocks. The blocks can be updated, for example, when at least one voxel in one or more tiles changes.

[0332] The blocks may not necessarily be limited to corresponding to the tiles. In some embodiments, the blocks can be generated based on one or more bricks, one or more voxels, or one or more SDF samples, etc. The blocks can be any suitable partition of the physical world. The blocks do not necessarily have to be limited to the format of a grid. The blocks can be in any suitable format of the 3D reconstruction data.

[0333] Figure 43A -D shows an exemplary physical world 4300 represented by a grid block 4302. Each grid block 4302 can be extracted from voxels 4304 corresponding to a predetermined volume of the grid block. In the example shown, each block can be the output of a cubic region of voxels (e.g., 1m 3 ) in a low-level reconstruction model. Each grid block 4302 can contain a part of the world grid and can be processed independently. Since certain blocks change when things move in exploring new areas or environments, scalability can be achieved through fast local updates. In the example shown, except for the grid block 4306, which has a new object 4308 placed in front of an existing surface 4310, the grid blocks do not change. In this case, the AR system only needs to update the grid block 4306, which can save a large amount of computing power compared to arbitrarily updating the entire grid of the world.

[0334] Figure 43B is a simplified schematic diagram showing a grid block according to some embodiments. In the example shown, the grid block may have a fully connected grid internally, which means that vertices are shared by multiple triangles.

[0335] On the other hand, a separate grid block can be an unconnected independent grid. Figure 43C is a simplified schematic diagram showing a crack according to some embodiments, which may exist at the edge of two adjacent grid blocks. Figure 44 D is a simplified schematic diagram showing masking Figure 43C the crack in by implementing a grid side edge that overlaps with an adjacent grid block according to some embodiments.

[0336] Figure 44 is a schematic diagram showing a representation 4400 of dividing a part of the physical world in 2D according to some embodiments. The 2D representation 4400 can be obtained by connecting four sets of blocks (blocks A - D). The representation 4400 can be divided into four blocks: for example, blocks 4402, 4404, 4406, and 4408 shown by solid line 4410. In some embodiments, blocks 4402, 4404, 4406, and 4408 can be designated as blocks A - D respectively. An application may need 3D reconstruction data in a grid format for further processing, such as occlusion testing, and generating physical effects in a physics engine. In some embodiments, the set of blocks can be in the format of a grid, which can be generated by a device (e.g., devices 3810, 3820), a network (e.g., a cloud including cloud cache 3802), or a discrete application (e.g., applications 3816, 3826).

[0337] In some embodiments, the regions at the boundaries of blocks 4402, 4404, 4406, and 4408 can be side edges shown by dashed line 4412, for example. In some embodiments, each of the blocks A - D can include a block and a corresponding side edge. For example, block B can include block 4404 and a side edge 4414 that overlaps with the boundary portions of adjacent blocks A, C, and D of block B, so that cracks between the blocks can be masked when the blocks are connected into a global grid. Blocks A, C, and D can also include corresponding side edges. Therefore, before returning the blocks including 3D reconstruction data to the application, the processor can mask any cracks between the blocks.

[0338] In some embodiments, the global grid can be a topologically connected global grid. For example, adjacent blocks in the set of blocks can share grid vertices at block boundaries such as line 4410. In some embodiments, the global grid can be visually seamless using any suitable technique (e.g., side edges and zippers) although it is topologically disconnected.

[0339] Although a method using side edges is shown, other methods can be used to enable the changed block to be combined with the adjacent unchanged blocks, such as zippers. Although in the illustrated example, a part of the physical world is represented by four 2D blocks, it should be appreciated that a part of the physical world can be represented by any suitable number (e.g., two, three, five, six, or more) of 2D and / or 3D blocks. Each block can correspond to a space in the physical world. In some embodiments, the blocks in the 2D and / or 3D representation of a part of the physical world can correspond to spaces of the same size (e.g., area / volume) in the physical world. In some embodiments, the blocks in the 2D and / or 3D representation of a part of the physical world can correspond to spaces of different sizes in the physical world.

[0340] Figure 45 FIG. is a schematic diagram showing a 3D representation 4500 of a part of the physical world according to some embodiments. Similar to the 2D representation 4400, the 3D representation 4500 can be obtained by connecting eight blocks (blocks A-H). In some embodiments, blocks A-H can be exclusive of each other, e.g., having no overlapping regions. In some embodiments, blocks A-H can have regions that overlap with adjacent blocks (e.g., side edges 4516). In some embodiments, each of the blocks A-H can have a version. Each version of a block can have values representing objects in the region of the physical world at a certain point in time. In the illustrated example, the 3D representation 4500 includes versions of blocks A-H: version 4502 of block A, version 4504 of block B, version 4514 of block C, version 4512 of block D, version 4534 of block E, version 4506 of block F, version 4508 of block G, and version 4510 of block H. Version 4502 of block A can include value 4518; version 4504 of block B can include value 4522; version 4514 of block C can include value 4528; version 4512 of block D can include value 4532; version 4534 of block E can include value 4520; version 4506 can include value 4524; version 4508 can include value 4526; version 4510 can include value 4530.

[0341] Figure 46FIG. is a schematic diagram showing a 3D representation 4600 of a part of the physical world obtained by updating a 3D representation 4500 according to some embodiments. Compared with the 3D representation 4500, the 3D representation 4600 may have a new version 4610 of a block H including information 4630. The information 4630 may be different from the information 4530. For example, a first device may retain a version 4510 of the block H in a remote memory. The version 4510 of the block H may include information 4530 corresponding to a table with an empty surface. After the first device leaves the area (e.g., the field of view of the first device no longer includes the block H), a second device may place a virtual and / or physical box on the surface of the table and then retain a version 4610 of the block H in the remote memory. The version 4610 of the block H may include information 4630 corresponding to a table with a virtual and / or physical box. If the first device returns, the first device is capable of selecting the version 4610 of the block H for viewing from the available versions of the block H (including the versions 4610 and 4510 of the block H).

[0342] Figure 47 FIG. is a schematic diagram showing an enhanced world 4700 viewable by a first device 4702 (e.g., device 3810) and a second device 4712 (e.g., device 3820). The first and second devices may include AR display systems 4704 and 4714 (e.g., AR display system 80) operating in an enhanced mode. The enhanced world 4700 may be obtained by connecting four blocks (blocks A-D). In the example shown, the enhanced world 4700 includes versions of blocks A-D: a version 4702A of block A, a version 4702B of block B, a version 4702C of block C, and a version 4702D of block D. The first device 4702 may be looking in a first direction 4706 and have a first field of view (FOV) 4708. In the example shown, the first FOV includes the version 4702B of block B and the version 4702D of block D. A processor (e.g., 3812) of the second device 4704 may include computer-executable instructions for identifying blocks B and D corresponding to the first FOV and selecting the version 4702B of block B and the version 4702D of block D. The second device 4714 may be looking in a second direction 4716 and have a second FOV 4718. In the example shown, the second FOV includes the version 4702C of block C and the version 4702D of block D. A processor (e.g., 3822) of the second device 4714 may include computer-executable instructions for identifying blocks C and D corresponding to the second FOV and selecting the version 4702C of block C and the version 4702D of block D.

[0343] Figure 48FIG. is a schematic diagram showing an enhanced world 4800 obtained by updating the enhanced world 4700 with new versions of blocks. Compared with the enhanced world 4700, the enhanced world 4800 may include a version 4802C of block C different from version 4702C, and a version 4802D of block D different from version 4702D. The first device 4702 can look in the third direction 4806 and has a third FOV 4808. In the example shown, the third FOV includes the version 4802C of block C and the version 4802D of block D. The processor of the first device 4702 may include computer-executable instructions for determining which of the versions 4702C, 4802C, 4702D, and 4802D to provide to the application based on, for example, changes in the FOV and / or information collected by the sensors of the first device 4702. In some embodiments, the processor of the first device 4702 may include computer-executable instructions for generating the version 4802C of block C and the version 4802D of block D when the corresponding latest versions of blocks C and D are not available in the local memory (e.g., local caches 3814, 3824) or remote memory (e.g., cloud cache 3802). In some embodiments, the first device 4702 is capable of estimating a change in its FOV (e.g., from the first FOV 4708 to the third FOV 4808), selecting block C based on the estimate, and storing the version of block C in a memory closer to the processor (e.g., moving the version of block C from the remote memory to the local cache, or from the local cache to the active memory).

[0344] Method for occlusion rendering using ray casting and real - time depth

[0345] The realism of presenting AR and MR scenes to the user can be enhanced by providing occlusion data to the applications that generate the AR and MR scenes, where the occlusion data is derived from a combination of one or more depth data sources. The occlusion data may represent the surfaces of physical objects in the scene and may be formatted in any suitable way, such as by depth data that indicates the distance from the viewpoint of the scene to be rendered to the surface. For example, the occlusion data can be received from the perception module 160 ( Figure 6 ) using the component 164.

[0346] However, in some embodiments, a data source can be one or more depth cameras that directly sense and capture the positions between the depth cameras and real objects in the physical world. Data from the depth cameras can be provided directly to the usage component 164, or can be provided indirectly, for example, through the sensing module 160. One or more depth cameras can provide an immediate view of the physical world at a frame rate high enough to capture changes in the physical world but low enough not to increase the processing burden. In some embodiments, the frame rate can be 5 frames per second, 10 frames per second, 12 frames per second, 15 frames per second, 20 frames per second, 24 frames per second, 30 frames per second, and so on. In some embodiments, the frame rate can be less than 5 frames per second. In some embodiments, the frame rate can be greater than 30 frames per second. Thus, in some embodiments, the frame rate can be in the range of, for example, 1 - 5 frames per second, 5 - 10 frames per second, 10 - 15 frames per second, 15 - 20 frames per second, or 20 - 30 frames per second, etc.

[0347] A second data source can be a stereo vision camera that can capture a visual representation of the physical world. The depth data from the depth cameras and / or the image data from the vision cameras can be processed to extract points representing real objects in the physical world. Images from vision cameras such as stereo cameras can be processed to compute a three-dimensional (3D) reconstruction of the physical world. In some embodiments, the depth data can be generated, for example, using deep learning techniques based on the images from the vision cameras. Some or all of the 3D reconstructions can be computed and stored in memory before the occlusion data. In some embodiments, the 3D reconstructions can be maintained in computer memory by a process independent of any process that generates the depth information for occlusion processing, and the occlusion processing can access the stored 3D reconstructions as needed. In some embodiments, the 3D reconstructions can be maintained in memory, and portions thereof can be updated in response to indications, for example, that there are changes in the physical world corresponding to portions of the 3D reconstructions as calculated based on the depth information. In some embodiments, the second data source can be implemented by ray casting into the 3D reconstruction of the physical world to obtain low-level 3D reconstruction data (e.g., ray cast point cloud). Through ray casting, data from the second data source can be selected to fill any holes in the occlusion data, enabling the integration of data from two (or more) sources.

[0348] According to some embodiments, the depth data and / or the image data and / or the low-level data of the 3D reconstruction can be oriented relative to the user of the AR or MR system. For example, such orientation can be achieved by using data from sensors worn by the user. The sensors can be worn, for example, on a head-mounted display device / unit.

[0349] In a system where occlusion data can be generated based on multiple depth data sources, the system can include a filter that identifies which portions of a 3D region are represented by data from each of the multiple depth data sources. The filter can apply one or more criteria to identify portions of the region for which data from a second data source will be collected. These criteria can be an indication of the reliability of the depth data. Since depth data is collected, another criterion can be a detected change in a portion of the region.

[0350] Selecting between multiple depth data sources to provide data for different portions of the representation of a 3D region can reduce processing time. For example, when the processing required to derive occlusion data from data collected by a first data source is less than that required by a second data source, the selection may favor data from the first data source, but data from the second source can be used when data from the first data source is unavailable or unacceptable. As a specific example, the first data source can be a depth camera and the second data source can be a stereo vision camera. Data from the stereo camera can be formatted for a 3D reconstruction of the physical world. In some embodiments, the 3D reconstruction can be computed before occlusion data is needed. Alternatively or additionally, the 3D reconstruction can be recomputed when occlusion data is needed. In some embodiments, criteria can be applied to determine whether the 3D reconstruction should be recomputed.

[0351] In some embodiments, the occlusion data is computed by a service that provides the occlusion data to an application executing on a computing device that will render an XR scene by applying a programming interface (API). The service can execute on the same computing device as the application or can execute on a remote computer. The service can include one or more of the components discussed herein, such as a filter for data from the first data source and / or an engine for selectively obtaining data from the second data source based on the filtered data from the first data source. The service can also include components that combine the filtered data from the first data source with the selected data from the second data source to generate the occlusion data.

[0352] Occlusion data can be formatted in any suitable way that represents surfaces in the physical world. For example, the occlusion data can be formatted as a depth buffer of a surface, thereby storing data that identifies the position of the surface in the physical world. The occlusion data can then be used in any suitable way. In some embodiments, the occlusion data can be provided to one or more applications that want virtual objects to be occluded by real objects. In some embodiments, the occlusion data can be formatted as a depth filter created by the system for an application that requests occlusion data from an occlusion service to render virtual objects at one or more locations. The depth filter can identify locations where the application should not render image information for a virtual object because the virtual object at those locations will be occluded by a surface in the physical world. It should be understood that "occlusion data" can have a suitable format to provide information about surfaces in the physical world and is not necessarily used for occlusion processing. In some embodiments, the occlusion data can be used in any application that performs processing based on the representation of surfaces in a scene of the physical world.

[0353] Compared to conventional AR and MR systems in which an application uses mesh data to perform occlusion processing, the methods described herein provide occlusion data with less latency and / or using lower computational resources. Mesh data can be obtained from geometric data extracted by a processing image sensor using multiple time- or cost-intensive steps, including marching cubes algorithms, mesh simplification, and applying triangle count limits. The computation of mesh data can take hundreds of milliseconds to several seconds, and when the environment changes dynamically and the application renders the scene using an outdated mesh, the latency of having an up-to-date mesh can result in visible artifacts. These artifacts, for example, manifest as: when virtual content is supposed to be rendered behind a real object, the virtual content appears to be superimposed on top of the real object, which disrupts the user's immersive perception / sensation of such applications and provides the user with incorrect cues for 3D depth perception.

[0354] For an application that uses a mesh for occlusion processing to have an up-to-date mesh, the application must continuously query the mesh (resulting in a large amount of continuous processing) or utilize a mechanism to determine if there are changes and then query the new mesh (which reduces overall processing but still has a high latency between changes in the physical world and when the mesh reflecting those changes reaches the application).

[0355] By performing occlusion directly using low-level data of 3D reconstruction data (such as point clouds) and real-time depth data instead of a mesh, the latency between when a change occurs in the environment and when it is reflected in the occlusion data can be reduced, thereby maintaining closer synchronization with the physical world and thus achieving a higher perceived visual quality.

[0356] In some embodiments, a real-time depth map of a physical environment can be obtained from a depth sensor (e.g., a depth camera). Each pixel in the depth map can correspond to a discrete distance measurement captured from a 3D point in the environment. In some embodiments, these depth cameras can provide a depth map including a set of points at a real-time rate. However, the depth map may have holes, which can be caused by the depth camera being unable to acquire sensor data representing a region or acquiring incorrect or unreliable data for the represented region. In some embodiments, if the depth sensor uses infrared (IR) light, these holes can be generated, for example, due to materials or structures in the physical environment not reflecting IR light well or not reflecting IR light at all. In some embodiments, these holes can be produced, for example, by very thin structures or surfaces at grazing incident angles that do not reflect light towards the depth sensor. The depth sensor may also encounter motion blur when moving rapidly, which can also result in lost data. Additionally, the "holes" in the depth map represent regions of the depth map that are not suitable for use in occlusion processing for any other reason. Any suitable processing can be used to detect such holes, such as processing the depth map to detect insufficient connectivity between points or regions in the depth map. As another example, a process that calculates a quality metric for regions of the depth map and treats regions with a low quality metric as holes can be used to detect holes. One such metric can be the inter-image variation of pixels representing the same location in the physical world in the depth map. Pixels having such variation exceeding a threshold can be classified as holes. In some embodiments, holes can be identified by pixels that meet a predetermined statistical criterion for a cluster of pixels where the quality metric is below the threshold.

[0357] In some embodiments, the depth map can first be "filtered" to identify holes. Then, the rays from the perspective from which the scene will be rendered to the holes can be determined. These rays can be "projected" into a 3D representation of the physical world created using multiple sensors rather than a single depth sensor to identify the data representing the regions of the holes. The 3D representation of the physical world can be, for example, a 3D reconstruction created from data from stereo vision cameras. The data from the 3D reconstruction identified by this ray projection can be added to the depth map, thereby filling the holes.

[0358] When the holes are identified, the 3D reconstruction can be calculated based on the image sensor data. Alternatively, some or all of the 3D reconstruction can be pre-calculated and stored in a memory. For example, the 3D reconstruction can be maintained in a computer memory by a process independent of any process for generating depth information for occlusion processing, and the process can access the stored 3D reconstruction as needed. As a further alternative, the 3D reconstruction can be maintained in the memory, but portions of it can be updated in response to an indication calculated based on the depth information that there is a change in the physical world corresponding to a portion of the 3D reconstruction.

[0359] In an XR system, light rays can have the same pose as the user's eye gaze. In the exemplary system described below, a depth map can be obtained similarly in the same way as the user's eye gaze, because the depth sensor can be worn by the user and can be mounted on the user's head near the eyes. The user can similarly wear a visual camera for forming 3D reconstruction data such that the images and the data derived from these images can be related to a coordinate system in which light rays defined relative to the depth map can be projected into the 3D reconstruction calculated from the visual images. Inertial measurement units and / or other sensors similarly worn by the user and / or associated with the sensors can provide data to perform coordinate transformation to add the data to the 3D representation, independent of the pose of the visual camera, and to relate the light rays defined by the depth map to the 3D reconstruction.

[0360] In some embodiments, the user's focus or associated virtual content placement information can guide the light ray projection to make the light ray projection adaptive in the image space by projecting denser light rays at depth discontinuities to obtain high-quality occlusions at object boundaries and sparse light rays in the center of the object to reduce processing requirements. The light ray projection can additionally provide local 3D surface information (such as normals and positions), which can be used to improve the temporal warping process with depth information and mitigate the loss of visible pixels that need to be rendered or ray-traced in a typical rendering engine. Temporal warping is a technique in XR that modifies the rendered image before sending it to the display to correct for the calculated head movement that occurs between rendering and display. In some embodiments, temporal warping can be used to synchronize the data from the depth map with the 3D representation of the physical world, and the 3D representation of the physical world can be used to generate data to fill the holes in the depth map. The data from the two data sources can be temporally warped to represent the pose calculated at the time of display. In some embodiments, the data from the 3D representation can be temporally warped to represent the pose calculated when the data was captured with the depth map.

[0361] In some embodiments, advanced features such as temporal warping can utilize the 3D local surface information from the light ray projection. When rendering a content frame in the absence of physical world occlusions or with eroded depth imaging, temporal warping can fill all the lost visible pixels that were previously occluded. Therefore, it may not be necessary for the rendering engine to fill the pixels, thus enabling a more loosely decoupled rendering application (or more independent temporal warping).

[0362] The processing described above can be performed on data acquired by a number of suitable sensors and presented on a number of suitable interfaces in many suitable forms of hardware processors. Examples of suitable systems including sensors, processing, and user interfaces are presented below. In the illustrated embodiments, a "service" can be implemented as part of an XR system with computer-executable instructions. Execution of these instructions can control one or more processors to access sensor data, then generate depth information and provide it to an application executing on the XR system. These instructions can be executed on the same processor or the same device that executes the application presenting the XR scene to the user, or can be executed on a remote device accessed by the user device via a computer network.

[0363] Figure 49 An occlusion rendering system 4900 according to some embodiments is shown. The occlusion rendering system 4900 can include a reconstruction filter 4902. The reconstruction filter 4902 can receive depth information 4904. In some embodiments, the depth information 4904 can be a sequence of depth images captured by a depth camera. In some embodiments, the depth information 4904 can be derived from a sequence of images captured by a vision camera, for example, using structure from motion based on a single camera and / or using stereo computation based on two cameras. Figure 50 A depth image 5000 according to some embodiments is shown. In some embodiments, surface information can be generated from the depth information. The surface information can indicate the distance to physical objects in the field of view (FOV) of a head-mounted display device including a depth camera and / or a vision camera. The surface information can be updated in real time as the scene and FOV change.

[0364] A second source of depth information is shown as a 3D reconstruction 4908. The 3D reconstruction 4908 can include a 3D representation of the physical world. The 3D representation of the physical world can be created and / or maintained in a computer memory. In some embodiments, the 3D representation can be generated from images captured by a vision camera, for example, using structure from motion based on a single camera and / or using stereo computation based on two cameras. In some embodiments, the 3D representation can be generated from depth images captured by a depth camera. For example, the 3D reconstruction 4908 can use the depth information 4904 in combination with the pose of the depth camera relative to the world origin to create and / or update. The representation can be built and modified over time, for example, when a user wearing the camera looks around in the physical world. In some embodiments, the depth information 4904 can also be used to generate a 3D representation of the physical world. The 3D reconstruction 4908 can be a volume reconstruction including 3D voxels. In some embodiments, each 3D voxel can represent a cubic space (e.g., 0.5 meters by 0.5 meters by 0.5 meters), and each 3D voxel can include data related to and / or describing a surface in the real world within the cubic space.

[0365] The 3D reconstruction 4908 of the world can be stored in any suitable manner. In some embodiments, the 3D reconstruction 4908 can be stored as a “cloud” of points representing the features of objects in the physical world. In some embodiments, the 3D reconstruction 408 can be stored as a mesh, where groups of points define the vertices of triangles representing surfaces. In some embodiments, other techniques such as room layout detection systems and / or object detection can be used to generate the 3D reconstruction 4908. In some embodiments, multiple techniques can be used together to generate the 3D reconstruction 4908. For example, object detection can be used for known physical objects in the physical world, 3D modeling can be used for unknown physical objects in the physical world, and a room layout detection system can also be used to identify boundaries in the physical world, such as walls and floors.

[0366] The reconstruction filter 4902 can include computer-executable instructions for generating a depth map based on the depth information 4904. The depth map can include one or more pixels. Each pixel can indicate the distance to a point on a surface in the physical world. In some embodiments, the reconstruction filter 4902 can synthesize the depth information 4904 and data from the ray casting engine 4906. In some embodiments, the reconstruction filter 4902 can reduce or remove noise from the depth information 4904 based at least in part on data synthesized from the ray casting engine 4902 and / or data from the depth information 4904 and the ray casting engine 4906. In some embodiments, the reconstruction filter 4902 can upsample the depth information 4904 using deep learning techniques.

[0367] The reconstruction filter 4902 can identify regions of the depth map based on a quality metric. For example, when the quality metric of a pixel is above a threshold, the pixel can be determined to be incorrect or noisy. A region of the depth map that contains incorrect or noisy pixels can be referred to as a hole (e.g., hole 5002).

[0368] Figure 54A and 54B Alternative examples of how holes can be generated in a depth map are provided in embodiments where the depth map is composed of multiple depth images. Figure 54A is a schematic diagram of imaging with a depth camera from a first perspective to identify regions of voxels occupied by a surface and empty voxels. Figure 54B is a schematic diagram of imaging with a depth camera from multiple perspectives to identify regions of voxels occupied and empty. Figure 54BShows a plurality of voxels determined to be occupied or empty by fusing data from multiple camera images. However, the voxels in region 5420 have not been imaged yet. Region 5420 may have been imaged using a depth camera at position 5422, but the camera has not yet moved to the position. Thus, region 5420 is the space being observed, and no volume information is available for the space being observed. The AR system can guide the user wearing it to scan the space being observed.

[0369] Return Figure 49 , the reconstruction filter 4902 can notify the ray casting engine 4906 about the positions of the holes. The ray casting engine 4906 can generate a view of the physical world given the user's pose and can remove the holes from the depth map. The data can represent a portion of the user's current view of the physical world at the current time or at the time when the data will be time warped using time warping. The ray casting engine 4906 can generate one or more 2D images, for example, one image for each eye. In some embodiments, the reconstruction filter 4902 can remove regions of the depth map that are spaced apart from the position of the virtual object by more than a threshold distance, as these regions may be irrelevant to the occlusion testing of the virtual object.

[0370] The ray casting engine 4906 can be implemented by any suitable technique. The ray casting engine 4906 generates a view of the physical world given the user's pose. In some embodiments, the ray casting engine 4906 can implement a ray casting algorithm on the 3D reconstruction 4908 to extract data therefrom. The ray casting algorithm can take the user's pose as input. The ray casting engine 4906 can project rays from a virtual camera into the 3D reconstruction 4908 of the physical world to obtain surface information missing from the depth map (such as holes). The ray casting engine 4906 can project dense rays at the boundaries of physical objects in the physical world to obtain high-quality occlusion at the object boundaries and project sparse rays in the central regions of the physical objects to reduce processing requirements. Then, the ray casting engine 4906 can provide the ray casting point cloud to the reconstruction filter 4902. The ray casting engine 4906 is shown as an example. In some embodiments, the ray casting engine 4906 can be a meshing engine. The meshing engine can implement a meshing algorithm on the 3D reconstruction 4908 to extract data therefrom, such as including triangles and the connectivity of these triangles. The meshing algorithm can take the user's pose as input.

[0371] The reconstruction filter 4902 can synthesize the depth information 4904 and the data from the ray casting engine 4906 to compensate for the holes in the depth map from the depth information 4904 with the ray cast point cloud data from the ray casting engine 4906. In some embodiments, the resolution of the depth map can be improved. This method can be used to generate a high-resolution depth image from a sparse or low-resolution depth image.

[0372] The reconstruction filter 4902 can provide the updated depth map to the occlusion service 4910. The occlusion service 4910 can calculate occlusion data based on the updated depth map and information about the positions of virtual objects in the scene. The occlusion data can be a depth buffer of the surfaces in the physical world. The depth buffer can store the depth of pixels. In some embodiments, the occlusion service 4910 can be an interface to the application 4912. In some embodiments, the occlusion service 4910 can interface with the graphics system. In these embodiments, the graphics system can expose the depth buffer to the application 4912, where the depth buffer is pre-filled with the occlusion data.

[0373] The occlusion service 4910 can provide the occlusion data to one or more applications 4912. In some embodiments, the occlusion data can correspond to the user's pose. In some embodiments, the occlusion data can be a per-pixel representation. In some embodiments, the occlusion data can be a mesh representation. The application 4912 can be configured to execute computer-executable instructions based on the occlusion data to render virtual objects in the scene. In some embodiments, the occlusion rendering can be performed by a separate graphics system rather than the application 4912. The separate graphics system can use time warping techniques.

[0374] In some embodiments, the reconstruction filter 4902, the ray casting engine 4906, and the occlusion service 4910 can be remote services, such as the remote processing module 72; or the 3D reconstruction 4908 can be stored in a remote memory, such as the remote data repository 74; and the application 4912 can be on the AR display system 80.

[0375] Figure 51Method 5100 for occlusion rendering in an augmented reality (AR) environment according to some embodiments is shown. At action 5102, depth information can be captured from, for example, a depth sensor (e.g., depth sensor 51) on a head-mounted display device. The depth information can indicate the distance between the head-mounted display device and a physical object. At action 5104, surface information can be generated from the depth information. The surface information can indicate the distance to a physical object in the field of view (FOV) of the head-mounted display device and / or the user of the head-mounted display device. The surface information can be updated in real time as the scene and FOV change. At action 5106, the portion of the virtual object to be rendered can be calculated based on the surface information and information about the position of the virtual object in the scene.

[0376] Figure 52 Details of action 5104 according to some embodiments are shown. At action 5202, the depth information can be filtered to generate a depth map. The depth map can include one or more pixels. Each pixel can indicate the distance to a point on the physical object. At action 5204, low-level data of a 3D reconstruction of the physical object can be selectively obtained, for example, from 3D reconstruction 4908. At action 5206, the surface information can be generated based on the depth map and the selectively obtained low-level data of the 3D reconstruction of the physical object.

[0377] Figure 53 Details of action 5202 according to some embodiments are shown. At action 5302, a quality metric for a region of the depth map can be determined. The quality metric can indicate whether the region of the depth map is incorrect or noisy. At action 5304, holes in the depth map can be identified based on the quality metric, for example, by comparing with a threshold. At action 5306, the identified holes can be removed from the depth map.

[0378] Conclusion

[0379] Accordingly, several aspects of some embodiments have been described. It should be understood that various changes, modifications, and improvements will occur readily to those skilled in the art.

[0380] As an example, embodiments are described in the context of an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein can be applied in a mixed reality (MR) environment or more generally in other XR environments and virtual reality (VR) environments.

[0381] As another example, embodiments are described in the context of a device such as a wearable device. It should be understood that some or all of the techniques described herein can be implemented via a network (e.g., the cloud), discrete applications, and / or devices, or any suitable combination of a network and discrete applications.

[0382] Such changes, modifications, and improvements are intended to be part of this disclosure and are intended to fall within the spirit and scope of this disclosure. Additionally, although the advantages of this disclosure are indicated, it should be understood that not every embodiment of this disclosure will include every described advantage. Some embodiments may not implement any of the features described herein as advantageous and in some cases. Accordingly, the foregoing description and drawings are provided by way of example only.

[0383] The above-described embodiments of this disclosure can be implemented in any of a variety of ways. For example, an embodiment can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. Such a processor can be implemented as an integrated circuit having one or more processors in the integrated circuit components, including commercially available integrated circuit components known in the art by names such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor can be implemented in a custom circuit such as an ASIC or in a semi-custom circuit generated by configuring a programmable logic device. As another alternative, the processor can be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores such that one or a subset of these cores can constitute the processor. However, the processor can be implemented using any suitable format of circuitry.

[0384] Moreover, it should be understood that a computer can be embodied in any of a variety of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer can be embedded in a device that is not generally considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smart phone, or any other suitable portable or stationary electronic device.

[0385] In addition, a computer may have one or more input and output devices. These devices can be used in particular to present a user interface. Examples of output devices that can be used to provide a user interface include printers or displays for visual presentation of output, and speakers or other sound generating devices for auditory presentation of output. Examples of input devices that can be used for a user interface include keyboards and pointing devices such as mice, touchpads, and digitizing tablets. As another example, a computer can receive input information via speech recognition or other audible formats. In the illustrated embodiment, the input / output devices are shown physically separate from the computing device. However, in some embodiments, the input and / or output devices can be physically integrated into the same unit as the processor or other elements of the computing device. For example, a keyboard can be implemented as a soft keyboard on a touchscreen. In some embodiments, the input / output devices can be completely disconnected from the computing device and functionally integrated via a wireless connection.

[0386] Such computers can be interconnected via one or more networks in any suitable form, including as a local area network or a wide area network such as a corporate network or the Internet. Such networks can be based on any suitable technology and can operate according to any suitable protocol, and can include wireless networks, wired networks, or fiber optic networks.

[0387] In addition, the various methods or processes outlined herein can be encoded as software executable on one or more processors employing any of a variety of operating systems or platforms. Additionally, such software can be written using any of a variety of suitable programming languages and / or programming or scripting tools, and can also be compiled into executable machine language code or intermediate code for execution on a framework or virtual machine.

[0388] In this regard, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy disks, compact discs (CDs), optical discs, digital video discs (DVDs), magnetic tapes, flash memories, circuit configurations in field programmable gate arrays or other semiconductor devices, or other tangible computer storage media) encoding one or more programs that, when executed on one or more computers or other processors, perform methods implementing various embodiments of the present disclosure discussed above. From the foregoing examples, it is apparent that a computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transitory form. One or more such computer-readable storage media can be removable, such that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" includes only computer-readable media that can be considered to be a manufacture (i.e., a manufactured article) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.

[0389] As used herein in a general sense, the terms "program" or "software" refer to any type of computer code or set of computer-executable instructions that can be used to program a computer or other processor to implement various aspects of the present disclosure as described above. Additionally, it should be understood that, according to one aspect of this embodiment, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion among multiple different computers or processors to implement various aspects of the present disclosure.

[0390] Computer-executable instructions can take many forms, such as program modules executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Generally, in various embodiments, the functions of program modules can be combined or distributed as needed.

[0391] In addition, data structures can be stored in a computer-readable medium in any suitable form. For simplicity of illustration, a data structure may be shown as having fields related by their positions in the data structure. Such relationships can equally be implemented by assigning storage devices with position fields in a computer-readable medium that conveys relationships between fields. However, any suitable mechanism can be used to establish relationships among the information in the fields of a data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.

[0392] Aspects of the present disclosure can be used individually, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and are thus not limited to the details and arrangements of the components set forth in the previous description or shown in the drawings. For example, aspects described in one embodiment can be combined with aspects described in other embodiments in any manner.

[0393] In addition, the present disclosure can be embodied as a method, examples of which have been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments can be constructed in which the acts are performed in an order different from that shown, which can include performing some acts simultaneously even though they are shown as sequential acts in illustrative embodiments.

[0394] The use of ordinal terms (e.g., "first," "second," "third," etc.) in the claims to modify a claim element itself does not denote any priority, precedence, or order of one claim element with respect to another or the temporal order of acts of a method of performing, but is merely used as a label to distinguish one claim element having a certain name from another element having the same name (but for the ordinal term) to distinguish these claim elements.

[0395] In addition, the language and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of "including" or "having," "comprising," "involving," and variations thereof herein is intended to cover the items listed thereafter and equivalents thereof as well as additional items.

Claims

1. A portable electronic system, comprising: At least one sensor configured to capture three-dimensional (3D) information about an object in the physical world; A local memory storing computer-executable instructions; A transceiver configured to communicate with a remote memory via a computer network; And A processor communicatively coupled to the at least one sensor, the local memory, and the transceiver, wherein the processor executes the computer-executable instructions to provide a 3D representation of a portion of the physical world based at least in part on the 3D information about the object in the physical world, wherein: The 3D representation of the portion of the physical world includes a plurality of blocks, each block including a value representing an object in a region of the portion of the physical world at a corresponding time point and a value indicating a floor; and The computer-executable instructions include instructions for processing as follows: Identifying a subset of the plurality of blocks corresponding to information captured by an incoming sensor, generating a new version of a block in the subset of the plurality of blocks based on a comparison between the information captured by the incoming sensor and the subset of the plurality of blocks, and Rendering a virtual object based on the new version of the block and adjacent blocks of the block in the subset.

2. The portable electronic system according to claim 1, wherein, Determining, at least in part based on the 3D information about an object in the physical world, a value indicating the floor of the block.

3. The portable electronic system according to claim 2, wherein, The at least one sensor includes an image sensor.

4. The portable electronic system according to claim 1, wherein, At least a portion of the at least one sensor is configured to: provide at least a portion of the information captured by the incoming sensor on which determining the position of the portable electronic system is at least partially based.

5. The portable electronic system according to claim 4, wherein, The position of the portable electronic system has different degrees of granularity.

6. The portable electronic system according to claim 4, wherein, The position of the portable electronic system includes a value indicating the floor of a corresponding block.

7. The portable electronic system according to claim 1, further comprising: One or more of a magnetometer, an altimeter, and a GPS sensor, configured to provide at least a portion of the information captured by the incoming sensor.

8. The portable electronic system according to claim 7, wherein, Determining, at least in part based on data captured by one or more of a magnetometer, an altimeter, and a GPS sensor, a value indicating the floor of the block.

9. The portable electronic system according to claim 7, wherein, Determining, at least in part based on data captured by one or more of a magnetometer, an altimeter, and a GPS sensor, the position of the portable electronic system.

10. The portable electronic system according to claim 9, wherein, The position of the portable electronic system has different degrees of granularity.

11. A method of operating a portable electronic system, the method comprising: Capturing sensor information about the physical world with at least one sensor; And Processing the sensor information with at least one processor to: Generate a representation of a portion of the physical world proximate to the portable electronic device, the representation of the portion of the physical world including a plurality of blocks, each block including a value representing an object in a region of the portion of the physical world at a corresponding time point; Associating height information with the portion of the physical world; Generating a new version of a block in the plurality of blocks based on a comparison between the information captured by the incoming sensor and the plurality of blocks, and Render a virtual object based on the new version of the block and adjacent blocks of the block among the multiple blocks.

12. The method according to claim 11, wherein, The height information includes identifiers of floors of a building.

13. The method according to claim 12, wherein, Associating the height information with the portion of the physical world includes: processing the output of a visual sensor to detect an indicator of the floor.

14. The method according to claim 12, wherein, Associating the height information with the portion of the physical world includes: tracking changes in the floor.

15. The method according to claim 12, wherein, Tracking changes in the floor includes: processing visual images to detect crossings of stairs or escalators.

16. The method according to claim 11, wherein, The representation of the portion of the physical world is a sparse representation.

17. The method according to claim 16, further comprising: Sending a positioning request based on (a) the representation of the portion of the physical world and (b) including the height information.

18. The method according to claim 11, wherein, The representation of the portion of the physical world is a first part of a dense representation.

19. The method according to claim 18, wherein: The dense representation includes a plurality of parts, the plurality of parts including the first part, each part including associated height information; and The method further comprises: performing occlusion processing based on the plurality of parts, the occlusion processing including: sorting surface information in a corresponding part based on the height information associated with each part of the plurality of parts.

20. A non-transitory computer-readable medium storing instructions that, when executed on a processor, perform actions including the following: Processing sensor information about the physical world to: Generate a representation of a portion of the physical world proximate to a portable electronic device, the representation of the portion of the physical world including a plurality of blocks, each block including values representing objects in a region of the portion of the physical world at a corresponding point in time; Associate height information with the portion of the physical world; Generate a new version of a block among the plurality of blocks based on a comparison between information captured by an incoming sensor and the plurality of blocks, and Render a virtual object based on the new version of the block and adjacent blocks of the block among the multiple blocks.

Citation Information

Patent Citations

  • Systems and methods for augmented reality preparation, processing, and application

    US20170352192A1

  • Crowd sourced mapping with robust structural features

    US20190025062A1