Fast 3D reconstruction with depth information
By capturing information through a depth sensor, identifying valid and invalid pixels, and combining voxel representation and confidence level to update the 3D environment of the XR system, the problem of high computational resource consumption in XR systems is solved, and a fast-updating and realistically interactive 3D environment representation is achieved.
Patent Information
- Application Number
- CN202080055363.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-07
- Filing Date
- 2020-08-06
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-08-06
AI Technical Summary
Existing XR systems consume excessive computational resources when rapidly creating and updating 3D environment representations, making it difficult to provide a realistic interactive experience with limited computing resources, especially when the user environment changes and the occlusion and movement of virtual and physical objects need to be updated quickly.
By using a depth sensor to capture information about the physical world, calculating depth images and identifying valid and invalid pixels, updating the 3D representation using confidence levels and weights, and combining voxel representation and truncated signed distance functions, the environment map is dynamically updated, reducing the computational burden.
It enables rapid updating of the 3D environment representation under limited computing resources, improves the realism of interaction between virtual and physical objects, reduces the consumption of computing resources, and improves the response speed and accuracy of the XR system.
Smart Images

Figure CN114245908B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] This application relates generally to cross reality systems that render a scene using three-dimensional (3D) reconstruction. BACKGROUND
[0002] Computers can control human user interfaces to create X Reality (XR or cross reality) environments in which some or all of the XR environment perceived by a user is generated by a computer. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which some or all of the XR environment can be generated in part by a computer using data that describes the environment. This data can describe, for example, virtual objects that can be rendered in a way that a user senses or perceives as being part of the physical world and can interact with the virtual objects. Because the data is rendered and presented through a user interface device (e.g., a head-mounted display device), the user can experience these virtual objects. The data can be displayed for the user to see, or can control audio that is played for the user to hear, or can control a tactile (or haptic) interface, enabling the user to experience a tactile sensation of feeling the virtual objects.
[0003] XR systems can be useful for many applications, spanning fields such as scientific visualization, medical training, engineering design and prototyping, tele-manipulation and tele-presence, and personal entertainment. In contrast to VR, AR and MR include one or more virtual objects that are related to real objects of the physical world. The experience of virtual objects interacting with real objects greatly enhances the enjoyment of users using XR systems, and opens the door to a variety of applications that present information about how to change the real world in a realistic and easily understood way.
[0004] An XR system can represent the physical surfaces of the world around a user of the system as a “mesh.” The mesh can be represented by a plurality of interconnected triangles. Each triangle has edges connecting points on surfaces of objects within the physical world, such that each triangle represents a portion of a surface. Information about the portion of the surface (e.g., color, texture, or other attributes) can be stored in association within the triangle. In operation, the XR system can process image information to detect points and surfaces, creating or updating the mesh. SUMMARY
[0005] Aspects of the present application relate to methods and apparatus for fast 3D reconstruction with depth information. The technology described herein can be used together, separately, or in any suitable combination.
[0006] Some embodiments relate to a portable electronic system. The portable electronic system includes a depth sensor configured to capture information about a physical world, and at least one processor configured to execute computer executable instructions to compute a three-dimensional (3D) representation of a portion of the physical world based at least in part on the captured information about the physical world. The computer executable instructions include instructions to: compute, from the captured information, a depth image comprising a plurality of pixels, each pixel indicating a distance to a surface in the physical world; determine, based at least in part on the captured information, valid pixels and invalid pixels of the plurality of pixels of the depth image; update the 3D representation of the portion of the physical world with the valid pixels; and update the 3D representation of the portion of the physical world with the invalid pixels.
[0007] In some embodiments, computing the depth image includes computing a confidence level about the distance indicated by the plurality of pixels, and determining the valid pixels and the invalid pixels includes, for each pixel of the plurality of pixels, determining whether the corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
[0008] In some embodiments, updating the 3D representation of the portion of the physical world with the valid pixels includes modifying a geometry of the 3D representation of the portion of the physical world with the distances indicated by the valid pixels.
[0009] In some embodiments, updating the 3D representation of the portion of the physical world with the valid pixels includes adding an object to an object map.
[0010] In some embodiments, updating the 3D representation of the portion of the physical world with the invalid pixels includes removing an object from the object map.
[0011] In some embodiments, updating the 3D representation of the portion of the physical world with the invalid pixels includes removing one or more reconstructed surfaces from the 3D representation of the portion of the physical world based at least in part on the distances indicated by the invalid pixels.
[0012] In some embodiments, removing one or more reconstructed surfaces from the 3D representation of the portion of the physical world when the distance indicated by the corresponding invalid pixel is outside an operating range of the sensor.
[0013] In some embodiments, the sensor includes a light source configured to emit light modulated at a frequency, a pixel array including a plurality of pixel circuits and configured to detect light reflected by an object at the frequency, and a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicating an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicating a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array. The depth image is computed based at least in part on the phase image.
[0014] In some embodiments, determining the valid pixels and the invalid pixels includes, for each pixel of the plurality of pixels of the depth image, determining whether a corresponding amplitude in the amplitude image is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding amplitude is below the predetermined value.
[0015] Some embodiments relate to at least one non-transitory computer-readable medium encoded with a plurality of computer-executable instructions that, when executed by at least one processor, perform a method for providing a three-dimensional (3D) representation of a portion of a physical world. The 3D representation of the portion of the physical world includes a plurality of voxels corresponding to a plurality of volumes of the portion of the physical world. The plurality of voxels store signed distances and weights. The method includes capturing information about the portion of the physical world as a change occurs within a field of view of a user, computing a depth image based on the captured information, the depth image including a plurality of pixels each indicating a distance to a surface in the portion of the physical world, determining valid pixels and invalid pixels of the plurality of pixels of the depth image based at least in part on the captured information, updating the 3D representation of the portion of the physical world with the valid pixels, and updating the 3D representation of the portion of the physical world with the invalid pixels.
[0016] In some embodiments, the captured information includes a confidence level about the distances indicated by the plurality of pixels. Determining the valid pixels and the invalid pixels includes, for each pixel of the plurality of pixels, determining whether a corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
[0017] In some embodiments, updating the 3D representation of the portion of the physical world with the valid pixels comprises computing signed distance and weights based at least in part on the valid pixels of the depth image, combining the computed weights with respective stored weights in the voxels and storing the combined weights as the stored weights, and combining the computed signed distances with respective stored signed distances in the voxels and storing the combined signed distances as the stored signed distances.
[0018] In some embodiments, updating the 3D representation of the portion of the physical world with the invalid pixels comprises computing signed distance and weights based at least in part on the invalid pixels of the depth image. The computing comprises modifying the computed weights based on a time at which the depth image was captured, combining the modified weights with respective stored weights in the voxels, and for each combined weight, determining whether the combined weight is above a predetermined value.
[0019] In some embodiments, modifying the computed weights comprises, for each computed weight, determining whether there is a difference between the computed signed distance corresponding to the computed weight and a respective stored signed distance.
[0020] In some embodiments, modifying the computed weights comprises, upon determining that there is the difference, decreasing the computed weight.
[0021] In some embodiments, modifying the computed weights comprises, upon determining that there is no difference, assigning the computed weight as the modified weight.
[0022] In some embodiments, updating the 3D representation of the portion of the physical world with the invalid pixels comprises further modifying the computed weights based on the time at which the depth image was captured, upon the combined weight being determined to be above the predetermined value.
[0023] In some embodiments, updating the 3D representation of the portion of the physical world with the invalid pixels comprises, upon the combined weight being determined to be below the predetermined value, storing the combined weight as the stored weight, combining a corresponding computed signed distance with a respective stored signed distance, and storing the combined signed distance as the stored signed distance.
[0024] Some embodiments relate to a method of operating a cross reality (XR) system to reconstruct a three-dimensional (3D) environment. The XR system includes a processor configured to communicate with sensors worn by a user to process image information, the sensors capturing information of regions in a field of view of the sensors. The image information includes a depth image computed from the captured information. The depth image includes a plurality of pixels. Each pixel indicates a distance to a surface in the 3D environment. The method includes determining the plurality of pixels of the depth image as valid pixels and invalid pixels based at least in part on the captured information, updating a representation of the 3D environment with the valid pixels, and updating the representation of the 3D environment with the invalid pixels.
[0025] In some embodiments, updating the representation of the 3D environment with the valid pixels includes modifying a geometry of the representation of the 3D environment based at least in part on the valid pixels.
[0026] In some embodiments, updating the representation of the 3D environment with the invalid pixels includes removing a surface from the representation of the 3D environment based at least in part on the invalid pixels.
[0027] The foregoing summary is provided by way of illustration only and is not intended to limit. BRIEF DESCRIPTION OF DRAWINGS
[0028] The drawings are not intended to be to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component is called out in every drawing. In the drawings:
[0029] Figure 1 is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene in accordance with some embodiments.
[0030] Figure 2 is a sketch of an example simplified AR scene in accordance with some embodiments, illustrating an example 3D reconstruction use case including visual occlusions, physics-based interaction, and environment reasoning.
[0031] Figure 3 is a schematic diagram illustrating a flow of data in an AR system configured to provide an experience of AR content interacting with the physical world in accordance with some embodiments.
[0032] Figure 4 is a schematic diagram illustrating an example of an AR display system in accordance with some embodiments.
[0033] Figure 5A is a schematic diagram illustrating an AR display system in accordance with some embodiments, rendering AR content as a user moves through a physical world environment.
[0034] Figure 5B is a schematic diagram illustrating a viewing optics assembly and ancillary components, according to some embodiments.
[0035] Figure 6 is a schematic diagram illustrating an AR system using a 3D reconstruction system, according to some embodiments.
[0036] Figure 7A is a schematic diagram illustrating a 3D space discretized into voxels, according to some embodiments.
[0037] Figure 7B is a schematic diagram illustrating a reconstruction range relative to a single viewpoint, according to some embodiments.
[0038] Figure 7C is a schematic diagram illustrating a perception range relative to a reconstruction range at a single location, according to some embodiments.
[0039] Figures 8A to 8F is a schematic diagram illustrating reconstructing a surface in the physical world as a voxel model by image sensors viewing the surface from multiple locations and viewpoints, according to some embodiments.
[0040] Figure 9A is a schematic diagram illustrating a scene represented by voxels, a surface in the scene, and a depth sensor capturing the surface in a depth image, according to some embodiments.
[0041] Figure 9B is a schematic diagram illustrating a truncated signed distance function (TSDF) relating to a truncated signed distance and a weight assigned to a voxel Figure 9A based on a distance from a surface.
[0042] Figure 10 is a schematic diagram illustrating an example depth sensor, according to some embodiments.
[0043] Figure 11 is a flowchart illustrating an example method of operating an XR system to reconstruct a 3D environment, according to some embodiments.
[0044] Figure 12 is a flowchart illustrating an example method of determining valid and invalid pixels in a depth image in Figure 11 , according to some embodiments.
[0045] Figure 13 is a flowchart illustrating an example method of updating a 3D reconstruction with valid pixels in Figure 11 , according to some embodiments.
[0046] Figure 14Ais an exemplary depth image showing valid and invalid pixels according to some embodiments.
[0047] Figure 14B is an exemplary depth image without invalid pixels. Figure 14A is an exemplary depth image.
[0048] Figure 15 is a flowchart of an exemplary method of updating a 3D reconstruction with invalid pixels in Figure 11 according to some embodiments.
[0049] Figure 16 is a flowchart of an exemplary method of modifying weights computed in Figure 15 according to some embodiments. DETAILED DESCRIPTION
[0050] Methods and apparatuses for providing a three-dimensional (3D) representation of an X reality (XR or cross reality) environment in an XR system are described herein. In order to provide a realistic XR experience to a user, the XR system must know the user’s physical environment in order to correctly associate the position of virtual objects with respect to real objects.
[0051] However, providing a 3D representation of an environment constitutes a significant challenge. A significant amount of processing can be required to compute the 3D representation. The XR system must know how to correctly position virtual objects with respect to the user’s head, body, etc., and render these virtual objects so that they appear to realistically interact with physical objects. For example, a virtual object can be occluded by a physical object between the user and the location where the virtual object appears. As the user’s position with respect to the environment changes, the relevant portion of the environment also changes, which can require further processing. Furthermore, as objects move in the environment (e.g., a cushion is removed from a sofa), the 3D representation typically needs to be updated. The update of the 3D representation of the environment must be performed quickly without using a significant amount of computing resources of the XR system that generates the XR environment, as the computing resources of the XR system used to update the 3D representation of the environment cannot perform other functions.
[0052] The inventors have recognized and appreciated techniques to accelerate the creation and update of a 3D representation of an XR environment with low computational resource usage by using information captured by sensors. Depth, which represents the distance from a sensor to an object in an environment, can be measured by sensors.
[0053] Using the measured depths, the XR system can maintain a map of objects in the environment. This map can be updated relatively frequently, as the depth sensor can output measurements at a rate of tens of times per second. Moreover, relatively little processing can be required to identify objects from the depths, the map made with the depths can be updated frequently to identify new objects in the vicinity of the user with low computational burden, or conversely, to identify that objects previously in the vicinity of the user have moved.
[0054] However, the inventors have recognized that depths can provide incomplete or ambiguous information about whether a map of objects in the vicinity of the user should be modified. An object previously detected from depths can fail to be detected for various reasons, e.g., the surface disappears, the surface is observed at a different angle and / or under different lighting conditions, the sensor fails to pick up the interposing object, and / or the surface is outside the range of the sensor.
[0055] In some embodiments, a more accurate map of objects can be maintained by selectively removing objects from the map that are not detected in the current depths. For example, an object can be removed based on detecting a surface in the depths that is further from the user than the previous location of the object along a line of sight through the previous location of the object.
[0056] In some embodiments, depths can be associated with different levels of confidence based on information captured by the sensor, e.g., the amplitude of light reflected by a surface. Smaller amplitudes can indicate lower levels of confidence in the associated depth, while larger amplitudes can indicate higher levels of confidence. Various reasons can cause a sensor measurement to be assigned a low level of confidence. For example, a surface closest to the sensor can be outside the operating range of the sensor, failing to gather accurate information about the surface in the environment. Alternatively or additionally, a surface can have poor reflective properties, such that the depth sensor detects too little radiation from the surface, and all measurements are made with a relatively low signal-to-noise ratio. Alternatively or additionally, a surface can be obscured by another surface, such that the sensor fails to acquire information about the surface.
[0057] In some embodiments, the level of confidence of depths in a depth image can be used to selectively update a map of objects. For example, if one or more depth pixels have values indicating with high confidence that a surface detected by the depth sensor is located behind a position in the object map where the object was previously indicated to be present, the object map can be updated to indicate that the object is no longer present at that position. The object map can then be updated to indicate that the object has been removed from the environment or moved to a different position.
[0058] In some embodiments, the confidence threshold for identifying an object in a new location can be different from the threshold for removing an object from a previously detected location. The threshold for removing an object can be lower than the threshold for adding an object. For example, a low confidence measurement can provide enough noisy information about the location of a surface that a surface added based on these measurements would have such an imprecise location that it can introduce more errors than not adding the surface. However, if a surface, regardless of which location within the range of confidence levels it is in, is behind the location of an object, the noisy surface can be sufficient to remove the object from the map of the environment. Similarly, some depth sensors operate according to physical principles that can produce ambiguous depth measurements for depths that are outside the operating range of the sensor. When using depths from these sensors, measurements that are outside the operating range of the sensor can be discarded as invalid. However, when all ambiguous locations of a surface correspond to locations behind the location of an object in the map, those measurements that would otherwise be discarded as invalid can still be used to determine that the object should be removed from the map.
[0059] In some embodiments, the 3D reconstruction can be in a format that facilitates selectively updating the map of objects. The 3D reconstruction can have a plurality of voxels, each voxel representing a volume of the environment represented by the 3D reconstruction. Each voxel can be assigned a value of a signed distance function indicating a distance from the voxel to a detected surface in its respective angle. In embodiments where the signed distance function is a truncated signed distance function, the maximum absolute value for the distance in a voxel can be truncated to a certain maximum value T, such that the signed distance will lie within the interval from -T to T. Furthermore, each voxel can include a weight indicating a certainty that the distance for the voxel accurately reflects the distance to the surface.
[0060] In some embodiments, an object can be added or removed from a map of objects that is part of a 3D representation of an environment based on voxels having a weight above a threshold. For example, if a surface identified as part of an object is located with a high certainty in a particular location, above a certain threshold, the map can be updated to show that the object is now located at that location or that the object has moved to that location. Conversely, if a surface has been detected behind a location indicated in the map with a high certainty of containing an object, the map can be updated to indicate that the object is removed or moved to another location.
[0061] In some embodiments, an object can be added or removed from a map based on a sequence of depth measurements. The weight stored in each voxel can be updated over time. When a surface is repeatedly detected in a location, the weight stored in the voxel with a value defined with respect to that surface can be increased. Conversely, the weight of a voxel indicating that a previously detected surface still exists can be decreased based on new measurements or differences in measurements indicating that the surface no longer exists in that location, thereby failing to confirm the existence of the surface.
[0062] The techniques as described herein can be used with or separately from a variety of types of devices, in a variety of types of scenarios, including wearable or portable devices with limited computing resources that provide cross reality scenarios. In some embodiments, these techniques can be implemented by services that form part of an XR system.
[0063] Figures 1-2 Such scenarios are illustrated. For purposes of illustration, an AR system is used as an example of an XR system. Figure 3 -8 illustrates an example AR system that includes one or more processors, memory, sensors, and user interfaces that can operate in accordance with the techniques described herein.
[0064] Reference is made to Figure 1 depicts an outdoor AR scenario 4 in which a user of AR technology sees a physical world park-like setting 6 featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these items, the user of AR technology also sees what appears to be a robotic statue 10 standing on the physical world concrete platform 8, and a personified cartoon-like avatar character 2 flying by that looks like a big yellow bee, even though these elements (e.g., avatar character 2 and robotic statue 10) do not exist in the physical world. Due to the extreme complexity of the human visual perception and nervous system, it is challenging to produce AR technology that produces a comfortable, natural-feeling, rich presentation of virtual image elements that augments reality in other virtual or physical world image elements.
[0065] Such AR scenarios can be implemented by a system that includes 3D reconstruction components that can construct and update a representation of the surfaces of the physical world around a user. This representation can be used for occlusion rendering in physics-based interactions, placement of virtual objects, and for virtual character path planning and navigation, or for other operations that use information about the physical world. Figure 2 Another example of an indoor AR scenario 200 is depicted, showing example 3D reconstruction use cases, including visual occlusion 202, physics-based interaction 204, and environmental reasoning 206, in accordance with some embodiments.
[0066] The example scene 200 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in a corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of AR technology perceives virtual objects such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking from the bookshelf, and a decoration in the form of a windmill placed on the coffee table. For the image on the wall, the AR technology needs information not only about the surface of the wall, but also about the objects and surfaces within the room that are occluding the image (e.g., the shape of the lamp) to properly render the virtual object. For the flying bird, the AR technology needs information about all the objects and surfaces around the room to render the bird in realistic physics, to avoid or bounce off of objects and surfaces as the bird collides with them. For the deer, the AR technology needs information about the surface (e.g., the floor or the coffee table) to calculate where to place the deer. For the windmill, the system can identify that it is an object separate from the table and can infer that it is movable, while the corner of the bookshelf or the corner of the wall can be inferred to be stationary. Such distinctions can be used to infer which parts of the scene are used or updated in each of various operations.
[0067] A scene can be presented to a user by a system that includes a plurality of components, including a user interface that can stimulate one or more user senses (including visual, sound, and / or touch). In addition, the system can include one or more sensors that can measure parameters of physical portions of the scene, including a user's position and / or motion within a physical portion of the scene. Further, the system can include one or more computing devices with associated computer hardware (e.g., memory). These components can be integrated into a single device or distributed among multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.
[0068] Figure 3 An AR system 302 configured to provide an experience of AR content interacting with a physical world 306 is depicted in accordance with some embodiments. The AR system 302 can include a display 308. In the illustrated embodiment, the display 308 can be worn by a user as part of a headset, such that the user can wear the display over their eyes like a pair of goggles or glasses. At least a portion of the display can be transparent, such that the user can observe see-through reality 310. The see-through reality 310 can correspond to a portion of the physical world 306 within a current viewpoint of the AR system 302, which can correspond to a viewpoint of the user, given that the user is wearing a headset that incorporates a display and sensors of the AR system to acquire information about the physical world.
[0069] AR content can also be presented on display 308, overlaid on see-through reality 310. To provide accurate interaction between AR content and see-through reality 310 on display 308, AR system 302 can include sensors 322 configured to capture information about physical world 306.
[0070] Sensor 322 can include one or more depth sensors that output depth images 312. Each depth image 312 can have a plurality of pixels, each of which can represent a distance to a surface in physical world 306 in a particular direction relative to the depth sensor. Raw depth data can come from the depth sensor to create a depth image. Such a depth image can be updated as fast as the depth sensor can form a new image, up to hundreds of thousands of times per second. However, this data can be noisy and incomplete, and have holes shown as black pixels on the illustrated depth image. In some embodiments, holes can be pixels that are not assigned a value, or pixels that have a low confidence such that any value is below a threshold and ignored.
[0071] The system can include other sensors, such as image sensors. Image sensors can acquire information that can be processed in other ways to represent the physical world. For example, images can be processed in 3D reconstruction component 316 to create a mesh that represents connected portions of objects in the physical world. Metadata about these objects, including for example color and surface texture, can similarly be acquired with sensors and stored as part of the 3D reconstruction.
[0072] The system can also acquire information about a head pose of a user relative to the physical world. In some embodiments, sensor 310 can include an inertial measurement unit that can be used to compute and / or determine head pose 314. Head pose 314 for a depth image can indicate a current viewpoint from which the sensor captured the depth image, for example in six degrees of freedom (6DoF), but head pose 314 can be used for other purposes, such as relating image information to a particular portion of the physical world or relating a position of a display worn on the user's head to the physical world. In some embodiments, head pose information can be derived in other ways than an IMU (e.g., analyzing objects in images).
[0073] 3D reconstruction component 316 can receive depth images 312 and head pose 314 from sensors, as well as any other data, and integrate this data into a reconstruction 318, which can appear to be at least a single combined reconstruction. Reconstruction 318 can be more complete and less noisy than the sensor data. 3D reconstruction component 316 can update reconstruction 318 using spatial and temporal averaging of sensor data from multiple viewpoints over time.
[0074] The reconstruction 318 can include representations of the physical world in one or more data formats including, for example, voxels, meshes, planes, etc. Different formats can represent alternative representations of the same portion of the physical world, or can represent different portions of the physical world. In the example shown, on the left side of the reconstruction 318, a portion of the physical world is presented as a global surface; on the right side of the reconstruction 318, a portion of the physical world is presented as a mesh.
[0075] The reconstruction 318 can be used for AR functions, such as generating a surface representation of the physical world for occlusion processing or physics-based processing. This surface representation can change as the user moves or objects in the physical world change. Aspects of the reconstruction 318 can be used, for example, by a component 320 that produces a global surface representation that changes in world coordinates, which can be used by other components.
[0076] AR content can be generated based on this information, for example, by the AR application 304. The AR application 304 can be, for example, a game program that performs one or more functions based on information about the physical world (e.g., visual occlusions, physics-based interactions, and environmental reasoning). It can perform these functions by querying data in different formats from the reconstruction 318 produced by the 3D reconstruction component 316. In some embodiments, the component 320 can be configured to output updates when the representation in a region of interest of the physical world changes. This region of interest can be set, for example, to approximate a portion of the physical world near the user of the system, such as a portion within the user's field of view, or projected (predicted / determined) to enter the user's field of view.
[0077] The AR application 304 can use this information to generate and update AR content. Virtual portions of the AR content can be presented on the display 308 in conjunction with the see-through reality 310, creating a realistic user experience.
[0078] In some embodiments, an AR experience can be provided to a user through a wearable display system. Figure 4 An example of a wearable display system 80 (hereinafter "system 80") is shown. The system 80 includes a head-mounted display device 62 (hereinafter "display device 62") and various mechanical and electronic modules and systems to support the functionality of the display device 62. The display device 62 can be coupled to a frame 64 that can be worn by a display system user or viewer 60 (hereinafter "user 60") and configured to position the display device 62 in front of the eyes of the user 60. According to various embodiments, the display device 62 can be a sequential display. The display device 62 can be monocular or binocular. In some embodiments, the display device 62 can be a head-mounted display (HMD). Figure 3 An example of a display 308 in a wearable system 80.
[0079] In some embodiments, a speaker 66 is coupled to the frame 64 and positioned proximate an ear canal of the user 60. In some embodiments, another speaker, not shown, is positioned proximate the other ear canal of the user 60 to provide stereo / shapeable sound control. The display device 62 is operatively coupled, e.g., by a wired lead or wireless connectivity 68, to a local data processing module 70, which can be mounted in a variety of configurations, i.e., fixedly attached to the frame 64, fixedly attached to a helmet or hat donned by the user 60, embedded in the ear piece, or otherwise detachably attached to the user 60 (e.g., in a backpack-style configuration, in a belt-coupling style configuration).
[0080] The local data processing module 70 can include a processor, as well as digital memory, such as nonvolatile memory (e.g., flash memory, etc.) both of which can be used to assist in the processing, caching, storage, and / or retrieval of data. Data includes a) data captured from sensors (which can be, e.g., operatively coupled to the frame 64) or otherwise attached to the user 60 (e.g., image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, radio devices, and / or gyros), and / or b) data acquired using remote processing modules 72 and / or remote data repositories 74, possibly for passage to the display device 62 after such processing or retrieval. The local data processing module 70 can be operatively coupled by communication links 76, 78 (e.g., via wired or wireless communication links) to the remote processing modules 72 and remote data repositories 74, respectively, such that these remote modules 72, 74 are operatively coupled to each other and available as resources to the local processing and data module 70. In some embodiments, Figure 3 The 3D reconstruction component 316 in the remote data processing module 72 can be implemented at least partially in the local data processing module 70. For example, the local data processing module 70 can be configured to execute computer-executable instructions to generate a physical world representation based at least in part on at least a portion of the data.
[0081] In some embodiments, the local data processing module 70 can include one or more processors (e.g., graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 70 can include a single processor (e.g., a single core or multi-core ARM processor), which will limit the computational budget of the module 70 but enable a smaller device. In some embodiments, the 3D reconstruction component 316 can use less than the computational budget of a single ARM core to generate a physical world representation in real-time on a non-predefined space, such that the remaining computational budget of the single ARM core can be accessed for other purposes, such as extracting a mesh.
[0082] In some embodiments, the remote data repository 74 can comprise a digital data storage facility, which can be used through the Internet or other network configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, allowing fully autonomous use from a remote module. The 3D reconstruction, for example, can be stored in whole or in part in this repository 74.
[0083] In some embodiments, the local data processing module 70 is operatively coupled to a battery 82. In some embodiments, the battery 82 is a removable power source, such as over a counter battery. In other embodiments, the battery 82 is a lithium ion battery. In some embodiments, the battery 82 comprises both an internal lithium ion battery that can be charged by the user 60 during non-operation times of the system 80 and a removable battery, so that the user 60 can operate the system 80 for longer periods of time without having to connect to a power source to charge the lithium ion battery or having to shut down the system 80 to replace the battery.
[0084] Figure 5A A user 30 is shown wearing an AR display system that renders AR content as the user 30 moves in a physical world environment 32 (hereinafter "environment 32"). The user 30 places the AR display system at a location 34, and the AR display system records passable world information (e.g., a digital representation of real objects in the physical world, which can be stored and updated as real objects in the physical world change) relative to the location 34, such as poses related to mapping features or directional audio inputs. The location 34 is aggregated to a data input 36, and processed at least by a passable world module 38, which can be implemented through processing on a remote processing module 72 of Figure 4 In some embodiments, the passable world module 38 can comprise a 3D reconstruction component 316.
[0085] The passable world module 38 determines where and how AR content 40 can be placed in the physical world as determined from the data input 36. By presenting a representation of the physical world and AR content via a user interface, the AR content is "placed" in the physical world, where the AR content is rendered as if it is interacting with objects in the physical world, and objects in the physical world are rendered as if the AR content is occluding the user's line of sight to those objects when appropriate. In some embodiments, AR content 40 can be placed by determining its shape and position from appropriately selecting portions of a fixed element 42 (e.g., a table) from a reconstruction (e.g., reconstruction 318). As an example, the fixed element can be a table, and virtual content can be placed so that it appears to be on that table. In some embodiments, AR content can be placed within structures in the field of view 44, which can be the current field of view or an estimated future field of view. In some embodiments, AR content can be placed with respect to a mapped mesh model 46 of the physical world.
[0086] As depicted, the fixed element 42 serves as a proxy for any fixed element within the physical world, which can be stored in the passable world module 38 so that the user 30 can perceive content on the fixed element 42 without the system having to map to the fixed element 42 every time the user 30 sees it. The fixed element 42 can thus be a mapped mesh model from a previous modeling session, or determined from a separate user, but both stored on the passable world module 38 for future reference by multiple users. Thus, the passable world module 38 can recognize the environment 32 from a previously mapped environment and display AR content without the user 30's device having to first map the environment 32, saving computational processes and cycles and avoiding any delay in rendered AR content.
[0087] The mapped mesh model 46 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interaction and display of AR content 40 can be mapped and stored in the passable world module 38 for future retrieval by the user 30 or other users without the need to re-map or model. In some embodiments, the data input 36 is input such as a geographic location, user identification, and current activity to indicate to the passable world module 38 which of one or more fixed elements 42 is available, which AR content 40 was last placed on the fixed element 42, and whether to display that same content (the AR content is "persistent" content regardless of whether the user views a particular passable world model or not).
[0088] Even in embodiments where objects are considered to be fixed, the passable world module 38 can be updated from time to time to account for the possibility of changes in the physical world. The model of fixed objects can be updated at a very low frequency. Other objects in the physical world can be moving or can not be considered to be fixed. To render an AR scene that has a realistic feel, the AR system can update the positions of these non-fixed objects at a much higher frequency than is used to update the fixed objects. To be able to accurately track all objects in the physical world, the AR system can extract information from multiple sensors, including one or more image sensors.
[0089] Figure 5B is a schematic view of the viewing optical assembly 48 and ancillary components. In some embodiments, two eye tracking cameras 50 directed at the user’s eyes 49 detect metrics of the user’s eyes 49, such as eye shape, eyelid occlusions, pupil direction, and glints on the user’s eyes 49. In some embodiments, one of the sensors can be a depth sensor 51, such as a time-of-flight sensor, that emits signals toward the world and detects reflections of those signals from nearby objects to determine distances to given objects. For example, the depth sensor can quickly determine whether objects have entered the user’s field of view due to motion of those objects or due to a change in the user’s pose. However, information about the position of objects in the user’s field of view can alternatively or additionally be gathered with other sensors. For example, depth information can be obtained from a stereo vision image sensor or a plenoptic sensor.
[0090] In some embodiments, a world camera 52 records a view larger than the periphery to map the environment 32 and detect inputs that can affect AR content. In some embodiments, the world camera 52 and / or camera 53 can be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. The camera 53 can also capture images of the physical world within the user’s field of view at particular times. The pixels of a frame-based image sensor can also be resampled even if their values do not change. Each of the world camera 52, camera 53, and depth sensor 51 has a respective field of view 54, 55, and 56 to collect data from and record the physical world scene, such as Figure 5A the physical world environment 32 depicted in FIG. 1.
[0091] An inertial measurement unit 57 can determine the motion and orientation of the viewing optical assembly 48. In some embodiments, each component is operatively coupled to at least one other component. For example, the depth sensor 51 can be operatively coupled to the eye tracking cameras 50 as a confirmation of measured accommodation relative to the actual distance at which the user’s eyes 49 are looking.
[0092] It will be appreciated that the viewing optical assembly 48 can include Figure 5BSome of the components shown, and can include components in addition to or instead of those shown. In some embodiments, for example, the viewing optics assembly 48 can include two world cameras 52 instead of four. Alternatively or additionally, the cameras 52 and 53 need not capture visible light images of their full field of view. The viewing optics assembly 48 can include other types of components. In some embodiments, the viewing optics assembly 48 can include one or more dynamic vision sensors (DVS), whose pixels can respond asynchronously to relative changes in light intensity above a threshold.
[0093] In some embodiments, the viewing optics assembly 48 can not include the depth sensor 51 based on time-of-flight information. In some embodiments, for example, the viewing optics assembly 48 can include one or more plenoptic cameras, whose pixels can capture light intensity and angle of incident light from which depth information can be determined. For example, a plenoptic camera can include an image sensor overlaid with a transmissive diffraction mask (TDM). Alternatively or additionally, a plenoptic camera can include an image sensor including angle-sensitive pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors can serve as a source of depth information, instead of or in addition to the depth sensor 51.
[0094] It should also be understood that Figure 5B The configuration of the components in the viewing optics assembly 48 is shown by way of example. The viewing optics assembly 48 can include components having any suitable configuration, which can be arranged to provide the user with the largest field of view practical for a particular set of components. For example, if the viewing optics assembly 48 has one world camera 52, the world camera can be placed in the center region of the viewing optics assembly instead of one side.
[0095] Information from sensors in the viewing optical assembly 48 can be coupled to one or more processors in the system. The processors can generate data that can be rendered to cause a user to perceive virtual content that interacts with objects in the physical world. This rendering can be implemented in any suitable way, including generating image data that depicts both physical and virtual objects. In other embodiments, physical and virtual content can be depicted in a scene by modulating the opacity of a display device through which the user views the physical world. The opacity can be controlled to create the appearance of virtual objects and also to block the user from seeing objects in the physical world that are occluded by virtual objects. In some embodiments, the image data can include only virtual content that can be modified so that when viewed through a user interface, the virtual content is perceived by the user to interact realistically with the physical world (e.g., the content is edited to resolve occlusions). Regardless of how the content is presented to the user, a model of the physical world is needed so that properties of virtual objects that can be influenced by physical objects, including the shape, position, motion, and visibility of virtual objects, can be computed correctly. In some embodiments, the model can include a reconstruction of the physical world, such as the reconstruction 318.
[0096] The model can be created from data collected from sensors on a user’s wearable device. However, in some embodiments, the model can be created from data collected by multiple users, which can be aggregated in a computing device that is remote from all of the users (and can be “in the cloud”).
[0097] The model can be created at least in part by a 3D reconstruction system (e.g., Figure 6 as described in more detail in Figure 3 3D reconstruction component 316 of the system 100. The 3D reconstruction component 316 can include a perception module 160 that can generate, update, and store a representation of a portion of the physical world. In some embodiments, the perception module 160 can represent a portion of the physical world within a reconstruction range of the sensors as a plurality of voxels. Each voxel can correspond to a 3D cube of a predetermined volume in the physical world and include surface information that indicates whether a surface exists in the volume represented by the voxel. The voxels can be assigned a value that indicates whether their corresponding volume has been determined to include a surface of a physical object, to be empty, or to not yet be measured with a sensor and thus have a value that is unknown. It will be appreciated that it is not necessary to explicitly store values for voxels that are determined to be empty or unknown, as the values for the voxels can be stored in computer memory in any suitable way, including not storing information about voxels that are determined to be empty or unknown. In some embodiments, a portion of the computer memory of the XR system can be mapped to represent the grid of voxels and store the values for the individual voxels.
[0098] Figure 7AAn example of a 3D space 100 discretized into voxels 102 is depicted. In some embodiments, the perception module 160 can determine objects of interest and set the volume of voxels in order to capture features of the objects of interest and avoid redundant information. For example, the perception module 160 can be configured to identify larger objects and surfaces, such as walls, ceilings, floors, and large furniture. Thus, the volume of voxels can be set to a relatively large size, such as 4 cm 3 cubes.
[0099] The reconstruction of the physical world including voxels can be referred to as a volumetric model. As a sensor moves in the physical world, information for creating the volumetric model can be created over time. This motion can occur as a user of a wearable device including the sensor moves about. Figure 8A An example of reconstructing a physical world into a volumetric model is depicted. In the example shown, the physical world includes a portion 180 of a surface shown in Figure 8A In Figure 8A , a sensor 182 at a first location can have a field of view 184 in which the portion 180 of the surface is visible.
[0100] The sensor 182 can be any suitable type, such as a depth sensor. However, depth data can be obtained from an image sensor or otherwise. The perception module 160 can receive data from the sensor 182 and then set values of a plurality of voxels 186 as shown in Figure 8B to represent the portion 180 of the surface that is visible by the sensor 182 in the field of view 184.
[0101] In Figure 8C , the sensor 182 can move to a second location and have a field of view 188. As shown in Figure 8D , another set of voxels becomes visible and the values of these voxels can be set to indicate the location of the portion of the surface that has entered the field of view 188 of the sensor 182. The values of these voxels can be added to the volumetric model for the surface.
[0102] In Figure 8E , the sensor 182 can further move to a third location and have a field of view 190. In the example shown, additional portions of the surface become visible in the field of view 190. As shown in Figure 8F , another set of voxels can become visible and the values of these voxels can be set to indicate the location of the portion of the surface that has entered the field of view 190 of the sensor 182. The values of these voxels can be added to the volumetric model for the surface. As shown in Figure 6 , this information can be stored as part of the persisted world as volumetric information 162a. Information about the surface, such as color or texture, can also be stored. Such information can be stored as, for example, volumetric metadata 162b.
[0103] In addition to generating information for a persisted world representation, the perception module 160 can also identify and output indications of changes in the area surrounding a user of the AR system. Such indications of changes can trigger updates to volumetric data stored as part of a persisted world, or trigger other functions, such as triggering a trigger component 304 that generates AR content to update that AR content.
[0104] In some embodiments, the perception module 160 can identify changes based on a signed distance function (SDF) model. The perception module 160 can be configured to receive sensor data such as depth images 160a and head pose 160b, and then fuse the sensor data into an SDF model 160c. Depth images 160a can directly provide SDF information, and images can be processed to obtain SDF information. SDF information represents distances from the sensors used to capture that information. Since those sensors can be part of a wearable unit, the SDF information can represent the physical world from the perspective of the wearable unit, and thus from the perspective of the user. Head pose 160b can enable SDF information to be correlated to voxels in the physical world.
[0105] Returning to Figure 6 In some embodiments, the perception module 160 can generate, update, and store representations of portions of the physical world within a perception range. The perception range can be determined based at least in part on a reconstruction range of a sensor, which can be determined based at least in part on limits of an observation range of the sensor. As a particular example, an active depth sensor operating using active IR pulses can reliably operate within a range of distances, creating an observation range of the sensor that can be from a few centimeters or tens of centimeters to a few meters.
[0106] Figure 7B A reconstruction range is depicted relative to a sensor 104 having a viewpoint 106. A reconstruction of a 3D space within the viewpoint 106 can be constructed based on data captured by the sensor 104. In the example shown, the observation range of the sensor 104 is 40 cm to 5 m. In some embodiments, the reconstruction range of a sensor can be determined to be less than the observation range of the sensor, as sensor output near its observation limits can be more noisy, incomplete, and inaccurate. For example, in the example shown of 40 cm to 5 m, the corresponding reconstruction range can be set to be from 1 to 3 m, and data collected by the sensor indicating surfaces outside of that range can not be employed.
[0107] In some embodiments, the perception range can be larger than the reconstruction range of the sensors. If components 164 that use data about the physical world need data about regions within the perception range that are outside the portion of the physical world that is currently within the reconstruction range, this information can be provided from the persisted world 162. Accordingly, information about the physical world is readily accessible by query. In some embodiments, an API can be provided to respond to such queries, providing information about the user's current perception range. Such techniques can reduce the time required to access existing reconstructions and provide an improved user experience.
[0108] In some embodiments, the perception range can be a 3D space corresponding to a bounding box centered around the user's location. As the user moves, the portion of the physical world within the perception range that is queryable by components 164 can move with the user. Figure 7C A bounding box 110 centered around the location 112 is depicted. It will be appreciated that the size of the bounding box 110 can be set to encompass the observation range of the sensors with a reasonable extension, as the user cannot move at an unreasonable speed. In the illustrated example, the observation limit of the sensors worn by the user is 5m. The bounding box 110 is set to a 20m 3 cube.
[0109] Returning Figure 6 , the 3D reconstruction component 316 can include additional modules that can interact with the perception module 160. In some embodiments, the persisted world module 162 can receive representations of the physical world based on data acquired by the perception module 160. The persisted world module 162 can also include representations of the physical world in various formats. For example, volumetric metadata such as voxels 162b can be stored as well as meshes 162c and planes 162d. In some embodiments, other information such as depth images can be saved.
[0110] In some embodiments, the perception module 160 can include modules that generate representations for the physical world in various formats, including for example meshes 160d, planes, and semantics 160e. These modules can generate representations based on data within the perception range of one or more sensors at the time the representation is generated as well as data captured at previous times and information in the persisted world 162. In some embodiments, these components can operate on depth information captured with depth sensors. However, AR systems can include visual sensors and can generate such representations by analyzing monocular or binocular visual information.
[0111] In some embodiments, the modules can operate on regions of the physical world. When the perception module 160 detects a change in the physical world in other sub-regions, those modules can be triggered to update the sub-regions of the physical world. Such a change can be detected, for example, by detecting a new surface in the SDF model 160c or other criteria (e.g., changing the value of a sufficient number of voxels representing that sub-region).
[0112] The 3D reconstruction components 316 can include components 164 that can receive representations of the physical world from the perception module 160. Information about the physical world can be pulled by the components from the perception module 160 according to, for example, a use request from an application. In some embodiments, information can be pushed to the use components, for example, via an indication of a change in a pre-identified region or a change in the representation of the physical world within the perception range. The components 164 can include, for example, game programs and other components that perform processing for visual occlusion, physics-based interactions, and environmental reasoning.
[0113] In response to a query from a component 164, the perception module 160 can send a representation of the physical world in one or more formats. For example, when the component 164 indicates that the use is for visual occlusion or physics-based interactions, the perception module 160 can send a representation of surfaces. When the component 164 indicates that the use is for environmental reasoning, the perception module 160 can send a mesh, planes, and semantics of the physical world.
[0114] In some embodiments, the perception module 160 can include components that format information to provide to components 164. An example of such a component can be a raycast component 160f. A use component (e.g., a component 164) can query for information about the physical world, for example, from a particular viewpoint. The raycast component 160f can select from one or more representations of the physical world data within the field of view from that viewpoint.
[0115] From the foregoing description, it should be appreciated that the perception module 160 or another component of the AR system can process data to create a 3D representation of a portion of the physical world. The 3D representation of the portion of the physical world can be created by, at least in part, picking portions of a 3D reconstruction volume based on camera frustums and / or depth images, extracting and retaining planar data, capturing, retaining, and updating 3D reconstruction data in tiles that allow for local updates while maintaining neighbor consistency, providing occlusion data to an application that generates such a scene (where the occlusion data is derived from a combination of one or more depth data sources), and / or performing multi-level mesh simplification to reduce data to be processed.
[0116] A 3D reconstruction system can integrate sensor data from multiple viewpoints of the physical world over time. As the device including the sensors moves, the pose (e.g., position and orientation) of the sensors can be tracked. Each of these multiple viewpoints of the physical world can be fused into a single, combined reconstruction, as the frame pose of the sensors is known and how it relates to other poses. By using spatial and temporal averaging (i.e., averaging data from multiple viewpoints over time), the reconstruction can be more complete and less noisy than the original sensor data. The reconstruction can contain different levels of complexity of data, including, for example, raw data (e.g., real-time depth data), fused volumetric data (e.g., voxels), and computed data (e.g., meshes).
[0117] Figure 9A A cross-sectional view of a scene 900 along a plane parallel to the y- and z-coordinates is depicted in accordance with some embodiments. Surfaces in the scene can be represented using truncated signed distance functions (TSDFs), which can map each 3D point in the scene to a distance to its nearest surface. Voxels representing positions on a surface can be assigned a zero depth. Surfaces in the scene can correspond to a range of uncertainty, for example, because the XR system can take multiple depth measurements, e.g., from two different angles or by two different users scanning the surface twice. Each measurement can result in a depth that is slightly different from other measured depths.
[0118] Based on the range of uncertainty of the measured position of the surface, the XR system can assign a weight associated with voxels within that range of uncertainty. In some embodiments, voxels that are more than a certain distance T from the surface (in addition to which) can convey with high confidence that they are not used. These voxels can correspond to positions in front of or behind the surface. These voxels can simply be assigned a size of T to simplify processing. Thus, voxels can be assigned values in the truncated band [-T, T] from the estimated surface, with negative values indicating positions in front of the surface and positive values indicating positions behind the surface. The XR system can compute weights to represent certainty about the computed signed distance to the surface. In the illustrated embodiment, the weights span between “1” and “0,” where “1” represents the most certain and “0” represents the least certain. As different techniques, including, for example, stereo imaging, structured light projection, time-of-flight cameras, sonar imaging, etc., provide different accuracies, the weights can be determined based on the technique used to measure the depth. In some embodiments, voxels corresponding to distances for which no accurate measurement was made can be assigned a weight of zero. In this case, the size of the voxel can be set to any value, e.g., T.
[0119] The XR system can represent the scene 900 by a grid of voxels 902. As described above, each voxel can represent a volume of the scene 900. Each voxel can store a signed distance from the center point of the voxel to its nearest surface. A positive sign can indicate behind the surface, while a negative sign can indicate in front of the surface. The signed distance can be computed as a weighted combination of distances obtained from multiple measurements. Each voxel can store a weight corresponding to the stored signed distance.
[0120] In the illustrated example, the scene 900 includes a surface 904 captured by the depth sensor 906 in a depth image (not shown). The depth image can be stored in computer memory in any convenient way that captures distances between some reference point and surfaces in the scene 900. In some embodiments, the depth image can be represented as values in a plane parallel to the x- and y-coordinates, as Figure 9A illustrated, where the reference point is the origin of the coordinate system. Positions in the X-Y plane can correspond to directions relative to the reference point. Values at those pixel positions can indicate distances from the reference point to the nearest surface in the direction indicated by the coordinates in the plane. Such a depth image can include a grid (not shown) of pixels in a plane parallel to the x- and y-coordinates. Each pixel can indicate a distance from the image sensor 906 to the surface 904 in a particular direction.
[0121] The XR system can update the grid of voxels based on the depth image captured by the sensor 906. TSDFs stored in the grid of voxels can be computed based on the depth image and a corresponding pose of the depth sensor 906. Voxels in the grid can be updated based on one or more pixels in the depth image, depending on, for example, whether the contours of the voxel overlap the one or more pixels.
[0122] In the illustrated example, voxels in front of the surface 904 but outside the cutoff distance -T are assigned the signed distance of the cutoff distance -T and a weight of “1,” because it is certain that the sensor is empty between the surface and any distance. Voxels between the cutoff distance -T and the surface 904 are assigned signed distances between the cutoff distance -T and 0 and weights of “1,” because it is certain that the sensor is outside the object. Voxels between the surface 904 and a predetermined depth behind the surface 904 are assigned signed distances between 0 and the cutoff distance T and weights between “1” and “0,” because the farther a voxel is behind the surface, the less certain it is that it represents the interior of the object or empty space. All voxels located behind the surface after the predetermined depth receive zero updates. Figure 9B depicted in FIG. 9B. The grid of voxels 902 can be updated based on the depth image 908 and a corresponding pose of the depth sensor 906. The grid of voxels 902 can be updated based on one or more pixels in the depth image 908, depending on, for example, whether the contours of the voxel overlap the one or more pixels. Figure 9ATSDF in a row of voxels. In addition, it can not update the partial voxel grid for that depth image, which reduces latency and saves computational power. For example, it does not update all voxels that do not fall within the camera frustum 908 for that depth image. U.S. Patent Application No. 16 / 229,799 describes culling a partial voxel grid for fast volume reconstruction, which is incorporated herein in its entirety.
[0123] In some embodiments, depth images can contain ambiguous data, which causes the XR system to be uncertain whether to update corresponding voxels. In some embodiments, rather than discarding ambiguous data and / or requesting new depth images, these ambiguous data can be used to speed up creation and updating of the 3D representation of the XR environment. The techniques described herein are able to create and update the 3D representation of the XR environment with low computational resource usage. In some embodiments, the techniques can reduce artifacts at the output of the XR system caused by latency due to, for example, delays until update information is available or by latency associated with heavy computation.
[0124] Figure 10 An exemplary depth sensor 1202 according to some embodiments is depicted, which can be used to capture depth information of an object 1204. The sensor 1202 can include a modulator 1206 configured to modulate a signal, for example, in a periodic pattern of a detectable frequency. For example, an IR light signal can be modulated with one or more periodic signals at a frequency between 1 MHz and 100 MHz. A light source 1208 can be controlled by the modulator 1206 to emit light 1210 modulated in a pattern of one or more desired frequencies. Reflected light 1212 reflected by the object 1204 can be collected by a lens 1214 and sensed by a pixel array 1216. The pixel array 1216 can include one or more pixel circuits 1218. Each pixel circuit 1218 can produce data for a pixel of an image output from the sensor 1202, corresponding to light reflected from the object in a direction relative to the sensor 1202.
[0125] The mixer 1220 can receive the signal from the output of the modulator 1206 so that it can act as a down-converter. The mixer 1220 can output one or more phase images 1222 based on, for example, a phase shift between the reflected light 1212 and the emitted light 1210. Each image pixel of the one or more phase images 1222 can have a time-based phase that is used for the emitted light 1210 to travel from the light source to the surface of the object and back to the sensor 1202. The phase of the light signal can be measured by a comparison of the transmitted light and the reflected light, for example, at four points, which can correspond to multiple locations, for example, four locations, over a period of the signal from the modulator 1206. An average phase difference at these points can be calculated. The depth of the point of the reflected light wave from the sensor to the surface of the object can be calculated based on the phase shift of the reflected light and the wavelength of the light.
[0126] The output of the mixer 1220 can be formed into one or more amplitude images 1224 based on, for example, one or more peak amplitudes of the reflected light 1212 measured at each pixel in the array 1216. Some pixels can measure reflected light 1212 with a low peak amplitude, for example, below a predetermined threshold, which can be related to large noise. The low peak amplitude can be caused by one or more of various reasons, including, for example, poor surface reflectivity, long distance between the sensor and the object 1204, etc. Thus, a low amplitude in the amplitude image can indicate a low level of confidence in the depth indicated by the corresponding pixel of the depth image. In some embodiments, these pixels of the depth image associated with the low level of confidence can be determined to be invalid. Other criteria in addition to or instead of the low amplitude can be used as an indication of low confidence. In some embodiments, asymmetry in the four points used for phase measurement can indicate low confidence. For example, the asymmetry can be measured by a standard deviation over a period of time in one or more phase measurements. Other criteria that can be used to assign low confidence can include oversaturation and / or undersaturation of the pixel circuitry. On the other hand, pixels of the depth image with depth values associated with a level of confidence above a threshold can be assigned as valid pixels.
[0127] Figure 11A method 1000 of operating an XR system to reconstruct a 3D environment according to some embodiments is depicted. The method 1000 can begin with determining (act 1002) valid and invalid pixels in a depth image. Invalid pixels can be selectively defined to include ambiguous data in the depth image, for example using heuristic criteria, or otherwise assigning a distance assigned to a voxel such a low confidence that the voxel can not be used in some or all processing operations. In some embodiments, invalid pixels can result from one or more of a variety of causes, including for example shiny surfaces, measurements made on surfaces outside the operating range of the sensor, computational errors due to asymmetry in capturing the data, over- or under-saturation of the sensor, etc. Any or all of the above or other criteria can be used to invalidate pixels in the depth image.
[0128] Figure 12 A method 1002 of determining valid and invalid pixels in a depth image according to some embodiments is depicted. The method 1002 can include capturing (act 1102) depth information (e.g., infrared intensity images) as the field of view of the user is changed by, for example, head pose, user position, and / or motion of physical objects in the environment. The method 1002 can compute (act 1104) one or more amplitude images and one or more phase images based on the captured depth information. The method 1002 can compute (act 1106) a depth image based on the computed one or more amplitude images and one or more phase images such that each pixel of the depth image has an associated amplitude, which can indicate a confidence level of a depth indicated by the pixel of the depth image.
[0129] Returning to Figure 11 , processing can be based on valid and invalid pixels. In some embodiments, pixels having a confidence level below a threshold, or otherwise failing to pass a valid criterion and / or meeting an invalid criterion can be set as invalid pixels. Other pixels can be considered valid. In some embodiments, pixels having a confidence level above a threshold or otherwise passing a valid criterion and / or meeting a valid criterion can be set as valid pixels. Other pixels can be considered invalid. The method 1000 can update (act 1004) the 3D reconstruction of the XR environment based on the valid pixels and / or the invalid pixels. A grid of voxels can be computed from the pixels, for example as shown in Figure 9A . Surfaces in the environment can be computed from the grid of voxels, for example using a marching cubes algorithm. These surfaces can be processed to identify foreground objects and other objects. Foreground objects can be stored in a manner that allows them to be processed and updated relatively quickly. For example, foreground objects can be stored in an object map, as described above.
[0130] In some embodiments, the foreground object map can use different data to update to add objects to the map than to remove objects from the map. For example, only valid pixels can be used to add objects, while some invalid pixels can be used to remove objects. Figure 13 A method 1004 is depicted that updates a voxel grid with valid pixels of a depth image measured by a sensor, in accordance with some embodiments. In Figure 13 In the example of FIG. 13, the signed distance and weight assigned to each voxel can be computed because each new depth sensor measurement is made based on, for example, a running average. The average can be weighted to favor more recent measurements and / or measurements with higher confidence more than previous measurements. Further, in some embodiments, measurements that are considered invalid can not be used at all for the update. The method 1004 can include computing (act 1302) a signed distance and weight based on valid pixels of a depth image, combining (act 1304) the computed weight with a corresponding stored voxel weight, and combining (act 1306) the computed signed distance with a corresponding stored signed distance of a voxel. In some embodiments, act 1306 can be performed after act 1304 and based on the combined weight of act 1304. In some embodiments, act 1306 can be performed before act 1304. Referring back to Figure 11 In some embodiments, after the 3D reconstruction is updated with valid pixels, the method 1000 can update (act 1008) the representation of the 3D reconstruction. As a result of the update, the representation of the world construction can have a different geometry, including, for example, a different mesh model and a global surface with a different shape. In some embodiments, the update can include removing an object from the object map where the updated voxel indicates that a new object is detected and / or a previously detected object is no longer present or has moved, for example, because a surface behind the previously detected object location has been detected with sufficient confidence.
[0131] Some or all of the invalid voxels can also be used in processing to remove a previously detected object. An example depth image 1400A is depicted in FIG. 14, showing valid and invalid pixels. Figure 14B An example depth image 1400B is depicted that is the depth image 1400A with the invalid pixels removed. Figure 14A And Figure 14B A comparison of FIGS. 14A and 14B shows that the image with invalid pixels has more data than the image with the invalid pixels removed. While this data can be noisy, it can be sufficient to identify whether an object is present or, conversely, not present so that a more distant surface can be observed. Thus, data such as depicted in FIG. 13 can be used to update an object map to remove an object. This update can be made with more data, and thus more confidence, than if only data such as depicted in FIG. 14B were available. Figure 14A Figure 14B As shown in the middle, the occurrence is faster. Since the update to remove the object does not involve inaccurately positioning the object in the map, a faster update time can be achieved without the risk of introducing errors.
[0132] The invalid pixels can be used to remove objects from the object map in any suitable manner. For example, a separate voxel grid computed with only valid pixels can be maintained, a separate voxel grid computed with both valid and invalid pixels. Alternatively, the invalid pixels can be processed separately to detect surfaces, and then these surfaces are used in a separate step to identify objects in the object map that are no longer present.
[0133] In some embodiments, to update the voxel grid representing the room 1402 shown in the depth image 1400A, each valid pixel in the depth image 1400B can be used to compute a value for one or more voxels in the grid. For each of the one or more voxels, a signed distance and a weight can be computed based on the depth image. The signed distance stored in association with the voxel can be updated with, for example, a weighted combination of the computed signed distance and the signed distance previously stored in association with the voxel. The weight stored in association with the voxel can be updated along with the voxel. Although this example is described as updating the voxel for each pixel of the depth image, in some embodiments, the voxel can be updated based on multiple pixels of the depth image. In some embodiments, for each voxel in the grid, the XR system can first identify one or more pixels in the depth image that correspond to the voxel, and then update the voxel based on the identified pixels.
[0134] Referring back to Figure 11 A, regardless of how the invalid pixels are processed, at act 1006, the method 1000 can update the 3D reconstruction of the XR environment with the invalid pixels. In the illustrated example, prior to capturing the depth image 1400A, the representation of the room 1402 includes a surface for the cushion on the sofa. In the depth image 1400A, the set of pixels 1404 corresponding to the cushion can be determined to be invalid for various possible reasons. For example, the cushion can have a poor reflectivity because it is covered in sequins. The act 1006 can update the voxels based on the invalid pixels such that if the cushion surface has been removed, it is removed from the representation of the room 1402, and if it is still on the sofa but has a poor reflectivity, it is retained in the representation of the room 1402 because processing only the valid pixels would not indicate or would not quickly or with high confidence indicate that the cushion is no longer present. In some embodiments, the act 1006 can include inferring a state of the surface based on the depth indicated by the invalid pixels, and removing the cushion from the object map when a surface is detected behind where the cushion was previously indicated to be present.
[0135] Figure 15A method 1006 of updating the voxel grid as new depth images are acquired is depicted in accordance with some embodiments. The method 1006 can begin with computing (act 1502) signed distances and weights based on invalid pixels of the depth images. The method 1006 can include modifying (act 1504) the computed weights. In some embodiments, the computed weights can be adjusted based on the time at which the depth images were captured. For example, a greater weight can be assigned to a depth image that was captured more recently.
[0136] Figure 16 A method 1504 of modifying the computed weights is depicted in accordance with some embodiments. The method 1504 can include, for each computed weight, determining (act 1602) whether a discrepancy exists between the corresponding computed signed distance and the respective stored signed distance. When a discrepancy is observed, the method 1504 can reduce (act 1604) the computed weight. When no discrepancy is observed, the method 1504 can assign (act 1606) the computed weight as the modified weight. For example, if the cushion is removed too quickly, the invalid pixels in the depth image can include a depth that is greater than the depth of the cushion surface captured previously, which can indicate that the cushion was removed. On the other hand, if the cushion is still on the sofa but has poor reflectivity, the invalid pixels in the depth image can include a depth that is comparable to the depth of the cushion surface captured previously, which can indicate that the cushion is still on the sofa.
[0137] At act 1506, the method 1006 can combine the modified weights with the respective previously stored weights in the voxels. In some embodiments, for each voxel, the combined weight can be the sum of the previously stored weight and the modified weight computed from the depth image. At act 1508, the method 1006 can determine whether each combined weight is higher than a predetermined value. The predetermined value can be selected based on the confidence level of the invalid pixels, such that pixels with lower confidence levels have smaller weights. When the combined weight is higher than the predetermined value, the method 1006 can further modify the computed weight. When the combined weight is lower than the predetermined value, the method can continue to combine (act 1510) the corresponding computed signed distance with the respective stored signed distance. In some embodiments, act 1510 can be omitted if the combined weight alone indicates that the surface corresponding to the pixel should be removed.
[0138] In some embodiments, each voxel in the voxel grid can have a rolling average of the stored weights when a new depth image is collected. Each new value is weighted to show changes that should warrant an addition or deletion of an object from the object map more quickly.
[0139] In some embodiments, after updating the 3D reconstruction with the invalid pixels, the method 1000 can update (act 1008) the representation of the world construction. In some embodiments, act 1008 can remove surfaces from the 3D representation of the environment based on the signed distance and weight in the updated pixels. In some embodiments, act 1008 can add back surfaces previously removed from the 3D representation of the environment based on the signed distance and weight in the updated pixels.
[0140] In some embodiments, the method is performed in conjunction with Figures 11 to 16 The described methods can be performed in one or more processors of an XR system.
[0141] Thus, having described several aspects of some embodiments, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.
[0142] As one example, embodiments are described in conjunction with an augmented reality (AR) environment. It will be appreciated that some or all of the techniques described herein can be applied in a MR environment or more generally in other XR environments and VR environments.
[0143] As another example, embodiments are described in conjunction with a device such as a wearable device. It will be appreciated that some or all of the techniques described herein can be implemented via a network (e.g., the cloud), discrete applications and / or devices, any suitable combination of a network and discrete applications, and / or the like.
[0144] As a further example, embodiments are described in conjunction with sensors based on time-of-flight techniques. It will be appreciated that some or all of the techniques described herein can be implemented by other sensors based on any suitable techniques, including, for example, stereo imaging, structured light projection, and plenoptic cameras.
[0145] Such alterations, modifications, and improvements are intended to be part of this disclosure, and are intended to be within the spirit and scope of the disclosure. Further, though advantages of the present disclosure have been indicated, it will be appreciated that not all embodiments of the present disclosure will include every indicated advantage. Some embodiments can not implement any or all of the described features. Therefore, the foregoing description and drawings are by way of example only, and are not intended to limit the scope of the disclosure.
[0146] The above-described embodiments of the present disclosure can be implemented in any of numerous ways. For example, the embodiments can be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. Such processors can be implemented as integrated circuits, with one or more processors in an integrated circuit component, including commercially available integrated circuit components such as CPUs, GPUs, microcontrollers, microprocessors, or co-processors. In some embodiments, a processor can be implemented within an ASIC or by configuration of a programmable logic device. As another alternative, a processor can be a portion of a larger circuit or semiconductor device, whether commercially available, semi-custom or custom.
[0147] Also, it should be appreciated that a computer can be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer can be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a PDA, a smart phone or any other suitable portable or fixed electronic device.
[0148] Also, a computer can have one or more input and output devices. These devices can be used, among other things, to present computer output to a user. Examples of output devices that can be used to provide user interfaces include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer can receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output devices are shown as physically separate from the computing device. However, in some embodiments, the input and / or output devices can be physically integrated in the same unit as the processor or other elements of the computing device. For example, a keyboard can be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output devices can be completely detached from the computing device and functionally integrated through a wireless connection.
[0149] Such computers can be interconnected by one or more networks in any suitable form, including as a local area network or a wide area network, such as an enterprise network or the Internet. Such networks can be based on any suitable technology and can operate according to any suitable protocol and can include wireless networks, wired networks or fiber optic networks.
[0150] Furthermore, various methods or processes outlined herein can be encoded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software can be written using any of a number of suitable programming languages and / or programming or scripting tools, and also can be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine.
[0151] In this respect, the disclosure can be embodied as a computer readable storage medium (or multiple computer readable media) (e.g., a computer memory, one or more floppy discs, compact discs (CD), optical discs, digital video disks (DVD), magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other non-transitory medium of a computer readable nature) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement the various embodiments of the disclosure discussed above. As will be evident to one of skill in the art, the computer readable storage medium can be non-transitory in that it does not rely on propagation of a signal per se (as in a propagated signal), but is a non-transitory medium that can store data that is non-transitory for some period of time. The one or more computer readable storage media can be moveable such that the one or more programs stored thereon can be loaded into the one or more different computers or other processors to implement various aspects of the present disclosure as described above. As used herein, the term “computer- readable storage medium” encompasses only a computer-readable medium that can be considered to be a manufacture (i.e., a manuf acture article) or a machine. In some embodiments, the disclosure can be embodied as a computer-readable medium other than a storage medium, such as a propagating signal.
[0152] The terms “program” or “software” are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure as discussed above. Additionally, it should be appreciated that according to one aspect of this embodiment, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor, but can be distributed in a modular fashion amongst a number of different computers or processors to implement various aspects of the present disclosure.
[0153] Computer-executable instructions can be in many forms, such as program modules, executed by one or more computers or other devices. Generally, these program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically the functionality of the program modules can be combined or distributed as desired in various embodiments.
[0154] Moreover, data structures can be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures can be shown to have fields that are related through location in the data structure. Such relationships can likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that are related through the machine's program instructions (e.g., by reference to the fields' names). However, any suitable mechanism can be used that accomplishes this function, including using pointers, tags or any other mechanism for establishing a relation among data elements.
[0155] Various aspects of the disclosure can be used alone, in combination, or in various arrangements not specifically discussed in the foregoing examples, and therefore the description is not limited to the details and arrangements set forth in the foregoing description or illustrated in the accompanying drawings. For example, aspects described in one embodiment can be combined in any manner with aspects described in other embodiments.
[0156] Furthermore, the disclosure can be embodied as a method, of which an example has been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments can be constructed in which acts are performed in an order different than illustrated, which can include performing some acts simultaneously, even though shown as being performed sequentially in illustrative embodiments.
[0157] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another, or of an action performed in one claim element over an action performed in another, but rather merely distinguishes between different claim elements or actions.
[0158] In addition, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," "having," "containing," "involving," and variations thereof herein, is meant to encompass the items listed thereafter and equivalents thereof as well as additional items.
Claims
1. A portable electronic system comprising: a depth sensor configured to capture information about a physical world; and at least one processor configured to execute computer-executable instructions to compute a three-dimensional (3D) representation of a portion of the physical world based at least in part on the captured information about the physical world, wherein the computer-executable instructions include instructions to: compute, from the captured information, a depth image comprising a plurality of pixels, each pixel indicating a distance to a surface in the physical world; determine, based at least in part on the captured information, valid pixels and invalid pixels of the plurality of pixels of the depth image; update the 3D representation of the portion of the physical world with the valid pixels; and update the 3D representation of the portion of the physical world with the invalid pixels, wherein confidence information is associated with pixels of the depth image, the confidence information indicating a confidence in the distance to the surface of the physical world indicated by the respective pixel, and the invalid pixels have a lower confidence than the valid pixels.
2. The portable electronic system of claim 1, wherein: computing the depth image includes computing a confidence level for the distances indicated by the plurality of pixels, and determining the valid pixels and the invalid pixels includes, for each pixel of the plurality of pixels, determining whether the corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
3. The portable electronic system of claim 2, wherein: updating the 3D representation of the portion of the physical world with the valid pixels includes modifying a geometry of the 3D representation of the portion of the physical world with the distances indicated by the valid pixels.
4. The portable electronic system of claim 1, wherein: updating the 3D representation of the portion of the physical world with the valid pixels includes adding objects to an object map.
5. The portable electronic system of claim 4, wherein: updating the 3D representation of the portion of the physical world with the invalid pixels includes removing objects from the object map.
6. The portable electronic system of claim 1, wherein: updating the 3D representation of the portion of the physical world with the invalid pixels includes removing one or more reconstructed surfaces from the 3D representation of the portion of the physical world based at least in part on the distances indicated by the invalid pixels.
7. The portable electronic system of claim 6, wherein: one or more reconstructed surfaces are removed from the 3D representation of the portion of the physical world when the distances indicated by the corresponding invalid pixels are outside an operating range of the sensor.
8. The portable electronic system of claim 6, wherein: one or more reconstructed surfaces are removed from the 3D representation of the portion of the physical world when the distances indicated by the corresponding invalid pixels indicate that the one or more reconstructed surfaces are moving away from the sensor.
9. The portable electronic system of claim 1, wherein: the 3D representation of the portion of the physical world includes information of the invalid pixels.
10. The portable electronic system of claim 1, wherein: the sensor includes: a light source configured to emit light modulated at a frequency; a pixel array including a plurality of pixel circuits and configured to detect light reflected by an object at the frequency; and a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicating an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicating a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array, wherein: the depth image is computed based at least in part on the phase image.
11. A portable electronic system comprising: a depth sensor configured to capture information about a physical world; and at least one processor configured to execute computer-executable instructions to compute a three-dimensional (3D) representation of a portion of the physical world based at least in part on the captured information about the physical world, wherein the computer-executable instructions include instructions to: compute a depth image including a plurality of pixels each indicating a distance to a surface in the physical world from the captured information; determine valid pixels and invalid pixels of the plurality of pixels of the depth image based at least in part on the captured information; update the 3D representation of the portion of the physical world with the valid pixels; and update the 3D representation of the portion of the physical world with the invalid pixels, wherein: updating the 3D representation of the portion of the physical world with the valid pixels includes adding and removing surfaces in the 3D representation, and updating the 3D representation of the portion of the physical world with the invalid pixels includes selectively removing surfaces from the 3D representation.
12. The portable electronic system of claim 11, wherein: computing the depth image includes computing a confidence level about the distance indicated by the plurality of pixels, and determining the valid pixels and the invalid pixels includes, for each pixel of the plurality of pixels, determining whether a corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
13. The portable electronic system of claim 11, wherein: updating the 3D representation of the portion of the physical world with the valid pixels includes adding an object to an object map.
14. The portable electronic system of claim 11, wherein: the 3D representation of the portion of the physical world includes information of the invalid pixels.
15. The portable electronic system of claim 11, wherein: the sensor includes: a light source configured to emit light modulated at a frequency; a pixel array including a plurality of pixel circuits and configured to detect light reflected by an object at the frequency; and a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicating an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicating a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array, wherein: the depth image is computed based at least in part on the phase image. a pixel array comprising a plurality of pixel circuits and configured to detect light at the frequency reflected by an object; and a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicative of an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicative of a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array, wherein: the depth image is computed based at least in part on the phase image.
16. A portable electronic system comprising: a depth sensor configured to capture information about a physical world; and at least one processor configured to execute computer-executable instructions to compute a three-dimensional (3D) representation of a portion of the physical world based at least in part on the captured information about the physical world, wherein the computer-executable instructions comprise instructions to: compute a depth image comprising a plurality of pixels each indicative of a distance to a surface in the physical world from the captured information; determine valid pixels and invalid pixels in the plurality of pixels of the depth image based at least in part on the captured information; update the 3D representation of the portion of the physical world with the valid pixels; and update the 3D representation of the portion of the physical world with the invalid pixels, wherein: the 3D representation of the portion of the physical world comprises information of the invalid pixels; and the information of the invalid pixels is indicative of a distance from the sensor to a surface in the portion of the physical world having a confidence level below a threshold value.
17. The portable electronic system of claim 16, wherein: computing the depth image comprises computing a confidence level with respect to the distance indicated by the plurality of pixels, and determining the valid pixels and the invalid pixels comprises, for each pixel in the plurality of pixels, determining whether a corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
18. The portable electronic system of claim 16, wherein: updating the 3D representation of the portion of the physical world with the valid pixels comprises adding an object to an object map.
19. The portable electronic system of claim 16, wherein: updating the 3D representation of the portion of the physical world with the invalid pixels comprises removing one or more reconstructed surfaces from the 3D representation of the portion of the physical world based at least in part on a distance indicated by the invalid pixels.
20. The portable electronic system of claim 16, wherein: the sensor comprises: a light source configured to emit light modulated at a frequency; a pixel array comprising a plurality of pixel circuits and configured to detect light at the frequency reflected by an object; and a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicative of an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicative of a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array, wherein: the depth image is computed based at least in part on the phase image. a mixer circuit configured to compute an amplitude image of the reflected light and a phase image of the reflected light, the amplitude image indicative of an amplitude of the reflected light detected by the plurality of pixel circuits in the pixel array, the phase image of the reflected light indicative of a phase shift between the emitted light and the reflected light detected by the plurality of pixel circuits in the pixel array, wherein: the depth image is computed based at least in part on the phase image.
21. The portable electronic system of claim 20, wherein: determining the valid pixels and the invalid pixels comprises, for each pixel in the plurality of pixels of the depth image, determining whether a corresponding amplitude in the amplitude image is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding amplitude is below the predetermined value.
22. A non-transitory computer-readable medium encoded with a plurality of computer- executable instructions that, when executed by at least one processor, perform a method for providing a three-dimensional (3D) representation of a portion of a physical world, the 3D representation of the portion of the physical world comprising a plurality of voxels corresponding to a plurality of volumes of the portion of the physical world, the plurality of voxels storing signed distances and weights, the method comprising: capturing information about the portion of the physical world as changes occur in a field of view of a user; computing a depth image based on the captured information, the depth image comprising a plurality of pixels each indicative of a distance to a surface in the portion of the physical world; determining, based at least in part on the captured information, valid pixels and invalid pixels in the plurality of pixels of the depth image; updating the 3D representation of the portion of the physical world with the valid pixels; and updating the 3D representation of the portion of the physical world with the invalid pixels, wherein confidence information is associated with pixels of the depth image, the confidence information indicative of a confidence in a distance to the surface of the physical world indicated by a respective pixel, and a confidence of the invalid pixels is lower than a confidence of the valid pixels.
23. The non-transitory computer-readable medium of claim 22, wherein: the captured information comprises a confidence level about distances indicated by the plurality of pixels, and determining the valid pixels and the invalid pixels comprises, for each pixel in the plurality of pixels, determining whether a corresponding confidence level is below a predetermined value, and designating the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
24. The non-transitory computer-readable medium of claim 23, wherein, updating the 3D representation of the portion of the physical world with the valid pixels comprises: computing signed distances and weights based at least in part on the valid pixels of the depth image, combining the computed weights with respective stored weights in the voxels and storing the combined weights as the stored weights, and combining the computed signed distances with respective stored signed distances in the voxels and storing the combined signed distances as the stored signed distances.
25. The non-transitory computer-readable medium of claim 23, wherein, updating the 3D representation of the portion of the physical world with the invalid pixels comprises: computing signed distance and weights based at least in part on the invalid pixels of the depth image, the computing comprising: modifying the computed weights based on a time at which the depth image was captured, combining the modified weights with respective stored weights in the voxels, and for each combined weight, determining whether the combined weight is above a predetermined value.
26. The non-transitory computer-readable medium of claim 25, wherein, modifying the computed weights comprises, for each computed weight, determining whether there is a difference between a computed signed distance corresponding to the computed weight and a respective stored signed distance.
27. The non-transitory computer-readable medium of claim 26, wherein, modifying the computed weights comprises, upon determining that there is the difference, reducing the computed weight.
28. The non-transitory computer-readable medium of claim 26, wherein, modifying the computed weights comprises, upon determining that there is no difference, assigning the computed weight as the modified weight.
29. The non-transitory computer-readable medium of claim 25, wherein, updating the 3D representation of the portion of the physical world with the invalid pixels comprises, upon the combined weight being determined to be above the predetermined value, further modifying the computed weight based on the time at which the depth image was captured.
30. The non-transitory computer-readable medium of claim 25, wherein, updating the 3D representation of the portion of the physical world with the invalid pixels comprises, upon the combined weight being determined to be below the predetermined value, storing the combined weight as the stored weight, combining a corresponding computed signed distance with a respective stored signed distance, and storing the combined signed distance as the stored signed distance.
31. The non-transitory computer-readable medium of claim 22, wherein: the 3D representation of the portion of the physical world includes information of the invalid pixels.
32. A non-transitory computer-readable medium encoded with a plurality of computer- executable instructions that, when executed by at least one processor, perform a method for providing a three-dimensional (3D) representation of a portion of a physical world, the 3D representation of the portion of the physical world including a plurality of voxels corresponding to a plurality of volumes of the portion of the physical world, the plurality of voxels storing signed distance and weights, the method comprising: capturing information about the portion of the physical world as changes occur within a field of view of a user; computing a depth image based on the captured information, the depth image including a plurality of pixels each indicating a distance to a surface in the portion of the physical world; determining, based at least in part on the captured information, valid pixels and invalid pixels of the plurality of pixels of the depth image; updating the 3D representation of the portion of the physical world with the valid pixels; and updating the 3D representation of the portion of the physical world with the invalid pixels, wherein: updating the 3D representation of the portion of the physical world with the valid pixels comprises adding and removing surfaces in the 3D representation, and updating the 3D representation of the portion of the physical world with the invalid pixels comprises selectively removing surfaces from the 3D representation.
33. The non-transitory computer-readable medium of claim 32, wherein: The captured information includes a confidence level regarding distances indicated by the plurality of pixels, and Determining the valid pixels and the invalid pixels includes, for each pixel in the plurality of pixels, Determining whether a corresponding confidence level is below a predetermined value, and Specifying the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
34. The non-transitory computer-readable medium of claim 32, wherein: The 3D representation of the portion of the physical world includes information of the invalid pixels.
35. A non-transitory computer-readable medium encoded with a plurality of computer- executable instructions that, when executed by at least one processor, perform a method for providing a three-dimensional (3D) representation of a portion of a physical world, the 3D representation of the portion of the physical world including a plurality of voxels corresponding to a plurality of volumes of the portion of the physical world, the plurality of voxels storing signed distances and weights, the method comprising: capturing information regarding the portion of the physical world as changes occur within a field of view of a user; computing a depth image based on the captured information, the depth image including a plurality of pixels, each pixel indicating a distance to a surface in the portion of the physical world; determining, based at least in part on the captured information, valid pixels and invalid pixels in the plurality of pixels of the depth image; updating the 3D representation of the portion of the physical world with the valid pixels; and updating the 3D representation of the portion of the physical world with the invalid pixels, wherein: the 3D representation of the portion of the physical world includes information of the invalid pixels; and the information of the invalid pixels indicates a distance from a sensor to a surface in the portion of the physical world with a confidence below a threshold value.
36. The non-transitory computer-readable medium of claim 35, wherein: The captured information includes a confidence level regarding distances indicated by the plurality of pixels, and Determining the valid pixels and the invalid pixels includes, for each pixel in the plurality of pixels, Determining whether a corresponding confidence level is below a predetermined value, and Specifying the pixel as an invalid pixel when the corresponding confidence level is below the predetermined value.
37. A method of operating a cross reality (XR) system to reconstruct a three-dimensional (3D) environment, the XR system including a processor configured to communicate with a sensor worn by a user to process image information, the sensor capturing information of regions in a field of view of the sensor, the image information including a depth image computed from the captured information, the depth image including a plurality of pixels, each pixel indicating a distance to a surface in the 3D environment, the method comprising: determining, based at least in part on the captured information, the plurality of pixels of the depth image as valid pixels and invalid pixels; updating a representation of the 3D environment with the valid pixels; and updating the representation of the 3D environment with the invalid pixels, wherein the confidence information is associated with pixels of the depth image, the confidence information indicating a confidence in a distance to a surface of the physical world indicated by a respective pixel, and the invalid pixels have a lower confidence than the valid pixels.
38. The method of claim 37, wherein, updating the representation of the 3D environment with the valid pixels comprises modifying a geometry of the representation of the 3D environment based at least in part on the valid pixels.
39. The method of claim 37, wherein, updating the representation of the 3D environment with the invalid pixels comprises removing surfaces from the representation of the 3D environment based at least in part on the invalid pixels.
40. The method of claim 37, wherein: the representation of the 3D environment includes information of the invalid pixels.
41. A method of operating a cross reality (XR) system to reconstruct a three-dimensional (3D) environment, the XR system comprising a processor configured to communicate with sensors worn by a user to process image information, the sensors capturing information of regions in a field of view of the sensors, the image information including a depth image computed from the captured information, the depth image including a plurality of pixels each indicating a distance to a surface in the 3D environment, the method comprising: determining the plurality of pixels of the depth image as valid pixels and invalid pixels based at least in part on the captured information; updating a representation of the 3D environment with the valid pixels; and updating the representation of the 3D environment with the invalid pixels, wherein: updating a 3D representation of a portion of a physical world with the valid pixels comprises adding and removing surfaces in the 3D representation, and updating the 3D representation of the portion of the physical world with the invalid pixels comprises selectively removing surfaces from the 3D representation.
42. The method of claim 41, wherein, updating the representation of the 3D environment with the valid pixels comprises modifying a geometry of the representation of the 3D environment based at least in part on the valid pixels.
43. The method of claim 41, wherein: the representation of the 3D environment includes information of the invalid pixels.
44. A method of operating a cross reality (XR) system to reconstruct a three-dimensional (3D) environment, the XR system comprising a processor configured to communicate with sensors worn by a user to process image information, the sensors capturing information of regions in a field of view of the sensors, the image information including a depth image computed from the captured information, the depth image including a plurality of pixels each indicating a distance to a surface in the 3D environment, the method comprising: determining the plurality of pixels of the depth image as valid pixels and invalid pixels based at least in part on the captured information; updating a representation of the 3D environment with the valid pixels; and updating the representation of the 3D environment with the invalid pixels, wherein: the representation of the 3D environment includes information of the invalid pixels; and the information of the invalid pixels indicates a distance from the sensors to a surface in the 3D environment having a confidence below a threshold value.
45. The method of claim 44, wherein, Updating the representation of the 3D environment with the valid pixel includes modifying a geometry of the representation of the 3D environment based at least in part on the valid pixel.
46. The method of claim 44, wherein, Updating the representation of the 3D environment with the invalid pixel includes removing a surface from the representation of the 3D environment based at least in part on the invalid pixel.
Citation Information
Patent Citations
Viewpoint dependent brick selection for fast volumetric reconstruction
US20190197777A1
3-dimensional shape reconstruction device using depth image and color image and the method
US20140111507A1
3di sensor depth calibration concept using difference frequency approach
US20180106891A1
System for 3D image filtering
US20180211398A1
Multi-stage block mesh simplification
US20190197774A1