A computationally efficient method for computing synthetic representations of 3D environments

By decomposing matrices and using reduced row matrices for pose refinement, the method addresses computational inefficiencies in 3D environment representation, achieving lower latency and resource usage for improved XR systems.

JP7752134B2Active Publication Date: 2025-10-09MAGIC LEAP INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022567781
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-11
Filing Date
2021-05-10
Publication Date
2025-10-09
Estimated Expiration
2041-05-10

AI Technical Summary

Technical Problem

Existing systems for constructing synthetic representations of 3D environments, such as in cross-reality (XR) systems, suffer from accumulated errors and high computational costs due to the need for frequent adjustments of sensor poses and feature alignments, which can lead to inefficiencies and resource constraints.

Method used

A method involving the decomposition of matrices into reduced row matrices and the use of reduced Jacobian and residual vectors to calculate refined poses and parameters for planar features, reducing computational complexity and memory usage.

Benefits of technology

This approach provides an accurate representation of the 3D environment with lower latency, power consumption, and resource requirements, enabling more efficient and immersive XR experiences by minimizing projection errors and optimizing map creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752134000049
    Figure 0007752134000049
  • Figure 0007752134000050
    Figure 0007752134000050
  • Figure 0007752134000051
    Figure 0007752134000051
Patent Text Reader

Abstract

A method and apparatus for providing a representation of an environment, for example, in an XR system and any suitable computer vision and robotics application. The representation of the environment may include one or more planar features. The representation of the environment may be provided by jointly optimizing planar parameters of the planar features and a sensor pose at which the planar features are observed. The joint optimization may be based on a reduced matrix and a reduced residual vector instead of a Jacobian matrix and an original residual vector.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related Applications) This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 023,089, filed May 11, 2020, and entitled "COMPUTATIONALLY EFFICIENT METHOD FOR COMPUTING A COMPOSITE REPRESENTATION OF A 3D ENVIRONMENT," which is incorporated herein by reference in its entirety.

[0002] This application relates generally to providing a representation of an environment in, for example, a cross-reality (XR) system, an autonomous vehicle, or other computer vision system that includes movable sensors. [Background technology]

[0003] Systems that use sensors to obtain information about 3D environments are used in multiple contexts, such as cross-reality (XR) systems or autonomous vehicles. These systems may employ sensors, such as cameras, to obtain information about the 3D environment. The sensors may provide multiple observations of the environment, which may be integrated into a representation of the environment. For example, as an autonomous vehicle moves, its sensors may obtain images of the 3D environment from different poses. Each image provides another observation of the environment and may provide further information about a previously imaged portion of the environment or information about new portions of the environment.

[0004] Such a system may assemble information acquired over multiple observations into a synthetic representation of a 3D environment. New images may be added to the representation of the synthetic environment, filling in information based on the pose from which the image was acquired. Initially, the pose of an image may be estimated based on the output of internal sensors indicating the motion of the sensor acquiring the image, or the correlation of features to features in previous images already incorporated into the synthetic representation of the environment. However, over time, errors in these estimation techniques may accumulate.

[0005] Adjustments to the composite representation may be performed from time to time to compensate for accumulated errors. Such adjustments may involve adjusting the pose of images that may be combined in the composite representation so that features in different images that represent the same object in the 3D environment better align. In some scenarios, the features to be matched may be feature points. In other scenarios, the features may be planes that are identified based on the content of the images. In such scenarios, the adjustments may involve adjusting both the relative pose of each detected image and the relative position of the planes based on the composite representation.

[0006] Such adjustments can provide an accurate representation of the 3D environment, which can be used in any of several ways. In an XR system, for example, a computer can control a human user interface, creating a cross-reality (XR) environment in which part or all of the XR environment is generated by the computer as perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), and mixed reality (MR) environments in which part or all of the XR environment can be generated by the computer, in part, using data that describes the environment. This data can, for example, describe virtual objects that can be rendered so that a user can sense or perceive and interact with them as part of the physical world. The user can experience these virtual objects as a result of data being rendered and presented through a user interface device, such as a head-mounted display device. The data can be displayed to the user so that it is visible or played audibly to the user, can control audio, or can control a tactile (or haptic) interface, allowing the user to experience touch sensations that the user senses or perceives as they feel the virtual objects.

[0007] XR systems can be useful for many applications, ranging from scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more objects in association with real objects in the physical world. The experience of virtual objects interacting with real objects greatly enhances the user's enjoyment when using XR systems and opens up possibilities for a variety of applications that present realistic and easily understandable information about how the physical world can be altered.

[0008] To realistically render virtual content, an XR system may build a representation of the physical world around a user of the system. This representation may be built, for example, by processed images obtained using sensors on a wearable device that forms part of the XR system. In such a system, a user may perform an initialization routine by looking around a room or other physical environment in which the user intends to use the XR system until the system has acquired enough information to build a representation of that environment. As the system operates and the user moves around the environment or into other environments, sensors on the wearable device may acquire additional information and expand or update the representation of the physical world. Summary of the Invention [Means for solving the problem]

[0009] Aspects of the present application relate to methods and apparatus for providing a representation of an environment, for example, in a cross-reality (XR) system, an autonomous vehicle, or other computer vision system that includes a movable sensor. The techniques described herein may be used together, separately, or in any suitable combination.

[0010] Some embodiments relate to a method for operating a computing system to generate a representation of an environment, the method including: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation comprising a first number of initial poses and initial parameters for a second number of planar features based at least in part on the first number of images; calculating, for each of a second number of planar features at each pose corresponding to the images comprising one or more observations of the planar features, a matrix indicative of the one or more observations of the planar features; decomposing the matrix into two or more matrices, the two or more matrices comprising one matrix with reduced rows compared to the matrix; and calculating, at least in part, refined poses and refined parameters for the second number of planar features based on the matrix with reduced rows. The representation of the environment comprises the first number of refined poses and refined parameters for the second number of planar features.

[0011] Some embodiments relate to a method of operating a computing system to generate a representation of an environment, the method including: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation comprising a first number of initial poses and initial parameters for a second number of planar features based at least in part on the first number of images; calculating a matrix having a third number of rows for each of the second number of planar features at each pose corresponding to images comprising one or more observations of the planar features, the third number being less than the number of one or more observations of the planar features; and calculating, at least in part, refined poses and refined parameters for the second number of planar features based on the matrix having the third number of rows. The representation of the environment comprises the first number of refined poses and refined parameters for the second number of planar features.

[0012] The foregoing description is provided by way of illustration and is not intended to be limiting. The present invention provides, for example, the following. (Item 1) 1. A method of operating a computing system to generate a representation of an environment, the method comprising: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation comprising the first number of initial poses and initial parameters of a second number of planar features based at least in part on the first number of images; For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating a matrix indicative of one or more observations of the planar features; decomposing the matrix into two or more matrices, the two or more matrices comprising one matrix having reduced rows compared to the matrix; calculating refined parameters of the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows; Including, The method, wherein the representation of the environment comprises refined parameters of the first number of refined poses and the second number of planar features. (Item 2) Decomposing the matrix into two or more matrices for each of the second number of planar features at each pose corresponding to images with one or more views of the planar features includes: Item 10. The method of item 1, comprising calculating an orthogonal matrix and an upper triangular matrix. (Item 3) Item 3. The method of item 2, wherein calculating refined parameters for the first number of refined poses and the second number of planar features is based, at least in part, on the upper triangular matrix. (Item 4) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: Item 10. The method of item 1, comprising calculating a reduced Jacobian matrix block based at least in part on the matrix having reduced rows. (Item 5) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: Item 5. The method of item 4, comprising stacking the reduced Jacobian matrix blocks to form a reduced Jacobian matrix. (Item 6) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: Item 6. The method of item 5, comprising providing the reduced Jacobian matrix to an algorithm that solves the least squares problem and updates the current estimate. (Item 7) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: Item 10. The method of item 1, comprising calculating a reduced residual block based at least in part on the matrix having reduced rows. (Item 8) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: Item 8. The method of item 7, comprising stacking the reduced residual blocks to form a reduced residual vector. (Item 9) Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: 9. The method of claim 8, comprising providing the reduced residual vector to an algorithm that solves a least-squares problem and updates a current estimate. (Item 10) For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating the matrix indicative of one or more views of the planar feature comprises: calculating, for each of one or more views of the planar feature, a matrix block indicative of the view; stacking said matrix blocks into said matrix representing one or more observations of said planar feature; Item 1. The method according to item 1, comprising: (Item 11) 1. A method of operating a computing system to generate a representation of an environment, the method comprising: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation comprising the first number of initial poses and initial parameters of a second number of planar features based at least in part on the first number of images; For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating a matrix having a third number of rows, the third number being less than a number of one or more observations of the planar feature; calculating refined parameters of the first number of refined poses and the second number of planar features based at least in part on the matrix having the third number of rows; Including, The method, wherein the representation of the environment comprises refined parameters of the first number of refined poses and the second number of planar features. (Item 12) calculating the matrix having the third number of rows for each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature; calculating a matrix indicative of one or more observations of the planar features; decomposing the matrix into two or more matrices, the two or more matrices comprising the matrix having the third number of rows; Item 12. The method according to item 11, comprising: (Item 13) Item 13. The method of item 12, wherein decomposing the matrix into two or more matrices includes calculating an orthogonal matrix and an upper triangular matrix. (Item 14) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 12. The method of item 11, comprising calculating a reduced Jacobian matrix block based at least in part on the matrix having the third number of rows. (Item 15) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 15. The method of item 14, comprising stacking the reduced Jacobian matrix blocks to form a reduced Jacobian matrix. (Item 16) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 16. The method of item 15, comprising providing the reduced Jacobian matrix to an algorithm that solves the least squares problem and updates the current estimate. (Item 17) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 12. The method of item 11, comprising calculating a reduced residual block based at least in part on the matrix having the third number of rows. (Item 18) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 18. The method of item 17, comprising stacking the reduced residual blocks to form a reduced residual vector. (Item 19) Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: Item 19. The method of item 18, comprising providing the reduced residual vector to an algorithm that solves a least-squares problem and updates a current estimate. (Item 20) For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating the matrix indicative of one or more views of the planar feature comprises: calculating, for each of one or more views of the planar feature, a matrix block indicative of the view; stacking said matrix blocks into said matrix representing one or more observations of said planar feature; Item 13. The method according to item 12, comprising: [Brief explanation of the drawings]

[0013] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component is labeled in every drawing.

[0014] [Figure 1] FIG. 1 is a sketch illustrating an example of a simplified augmented reality (AR) scene, according to some embodiments.

[0015] [Figure 2] FIG. 2 is a sketch of an exemplary simplified AR scene illustrating an exemplary use case of an XR system, according to some embodiments.

[0016] [Figure 3] FIG. 3 is a schematic diagram illustrating data flow for a single user in an AR system configured to provide the user with an experience of AR content that interacts with the physical world, according to some embodiments.

[0017] [Figure 4] FIG. 4 is a schematic diagram illustrating an exemplary AR display system displaying virtual content for a single user, according to some embodiments.

[0018] [Figure 5A] FIG. 5A is a schematic diagram illustrating a user wearing an AR display system that renders AR content as the user moves through a physical world environment, according to some embodiments.

[0019] [Figure 5B]FIG. 5B is a schematic diagram illustrating a viewing optics assembly and associated components, according to some embodiments.

[0020] [Figure 6A] FIG. 6A is a schematic diagram illustrating an AR system using a world reconstruction system, according to some embodiments.

[0021] [Figure 6B] FIG. 6B is a schematic diagram illustrating components of an AR system that maintains a model of a passable world, according to some embodiments.

[0022] [Figure 7] FIG. 7 is a schematic illustration of a tracking map formed by a device traversing a path through the physical world, according to some embodiments.

[0023] [Figure 8A] FIG. 8A is a schematic diagram illustrating the viewing of various planes in different poses, according to some embodiments.

[0024] [Figure 8B] FIG. 8B is a schematic diagram illustrating the geometric entities involved in calculating refined pose and plane parameters of a representation of an environment comprising the plane of FIG. 8A as observed in the pose of FIG. 8A, in accordance with some embodiments.

[0025] [Figure 9] FIG. 9 is a flowchart illustrating a method for providing a representation of an environment, according to some embodiments.

[0026] [Figure 10] FIG. 10 is a schematic diagram illustrating a portion of a representation of an environment calculated based on sensor acquisition information using the method of FIG. 9, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0027] Detailed Description Described herein are methods and apparatus for providing a representation of an environment, for example, in XR systems and any suitable computer vision and robotics applications. The inventors have recognized and appreciated methods and apparatus that provide representations of potentially complex environments, such as rooms with many objects therein, with reduced time and reduced computational cost and memory usage. In some embodiments, an accurate representation of an environment can be provided by computational techniques that reduce the amount of information processed when adjusting a synthetic representation of a 3D environment and reduce errors accumulated over time when constructing a synthetic representation from multiple observations. These techniques may be applied in systems that represent an environment by planes detected in a representation of the environment constructed by combining images of the environment obtained at multiple poses. Such techniques may be based on reduced matrices and reduced residual vectors instead of Jacobian matrices and original residual vectors.

[0028] Computationally less intensive processing and / or lower memory usage may be used within a system to provide an accurate representation of a 3D environment with lower latency, lower power consumption, less heat generation, lighter weight, or other advantages. XR systems are examples of systems that may be improved using such techniques. With low computational complexity, for example, a user device for an XR system may be created by programming on a smartphone.

[0029] In some embodiments, to provide a realistic XR experience for multiple users, the XR system must understand the user's physical surroundings to properly correlate the location of virtual objects relative to real objects. The XR system may build a representation of the environment in which virtual objects may be displayed. The representation of the environment may be created from information collected using sensors that are part of the XR device of the XR system. The XR device may be a head-mounted device with an integrated display and sensors, a handheld mobile device (e.g., a smartphone, a smartwatch, a tablet computing device, etc.), or other device with sensors. However, the techniques described herein may be used on other types of devices with sensors, such as autonomous machines with sensors. Such devices may include one or more types of sensors, including, for example, an image sensor, a LiDAR camera, an RGBD camera, an infrared camera, a structured light sensor, an ultrasonic sensor, or a coherent light sensor. These devices may include one or more components that assist in collecting sensor information, such as an infrared emitter, an ultrasonic emitter, or a structured light emitter. Alternatively, or in addition, such devices may include sensors that assist in determining attitude, such as gyroscope sensors, accelerometers, magnetometers, altimeters, proximity sensors, and GPS sensors.

[0030] The representation of the environment, in some embodiments, may be a local map of the environment created by the XR device by integrating information from one or more images surrounding the XR device and collected as the device operates. In some embodiments, the coordinate system of the local map may be tied to the device's position and / or orientation when the device first begins scanning the environment (e.g., starting a new session). That position and / or orientation of the device may change from session to session.

[0031] The local map may include sparse information representing the environment based on a subset of features detected in the sensor-acquired information used in forming the map. Additionally or alternatively, the local map may include dense information representing the environment using information about surfaces in the environment, such as a mesh.

[0032] The subset of features in the map may include one or more types of features. In some embodiments, the subset of features may include point features, such as the corners of a table, which may be detected based on visual information. In some embodiments, the system may alternatively or additionally form a representation of the 3D environment based on higher-level features, such as planar features, such as planes. Planes may be detected, for example, by processing depth information. The device may store planar features instead of or in addition to feature points, but in some embodiments, storing planar features instead of corresponding point features may reduce the size of the map. For example, a plane may be stored in the map represented using a surface normal and a signed distance to the origin of the map's coordinate system. In contrast, the same structure may be stored as multiple feature points.

[0033] Potentially, in addition to providing input for creating a map, sensor-captured information may be used to track the movement of a device within an environment. Tracking may allow the XR system to localize the XR device by estimating the individual device's pose relative to a frame of reference established by the map. Locating the XR device may require performing a comparison to find a match between a set of measurements extracted from images captured by the XR device and a set of features stored in an existing map.

[0034] This codependency between map creation and device localization constitutes a significant challenge. Even with the simplification of using planes to represent surfaces in a 3D environment, substantial processing may be required to accurately create a map and simultaneously localize the device. Processing must be performed quickly as the device moves through the environment. On the other hand, XR devices may have limited computing resources so that the device can move through the environment at a reasonable speed with reasonable flexibility.

[0035] The process may include jointly optimizing plane feature parameters and sensor poses used in forming a map of the 3D environment in which the plane is detected. Conventional techniques for jointly optimizing plane feature parameters and poses may incur high computational costs and memory consumption. Because sensors may record different observations of the plane features at different poses, joint optimization according to conventional approaches may require the solution of a very large nonlinear least-squares problem, even for a small workspace.

[0036] As used herein, "optimizing" (and similar terms) need not result in a perfect or theoretically best solution. Rather, optimization may result from a process that reduces a measure of error in a solution. The process to reduce the error may be performed until a solution with a sufficiently low error is reached. Such a process may be performed iteratively, and the iterations may be performed until a termination criterion indicating a sufficiently low error is detected. The termination criterion may be an indication that the process has converged on a solution, which may be detected by a percentage reduction in error per iteration below a threshold, and / or other criteria indicating that a predetermined number of iterations and / or a solution with a low error has been identified.

[0037] Techniques for efficient map optimization are described herein, and an XR system is used as an example of a system that can use these techniques to optimize a map. In some embodiments, information about each plane in the map is collected in each of multiple images obtained from a separate pose. A view of each plane may be made at each pose at which the images used in creating the 3D representation were captured. This information, representing one or more views of the plane, may be formatted as a matrix. In a conventional approach to optimizing a map, this matrix may be a Jacobian, and the optimization may be based on calculations performed using this Jacobian matrix.

[0038] According to some embodiments, the optimization may be performed based on a decomposed matrix, which may be mathematically related to a matrix as used in conventional approaches, but which may have reduced rows compared to the matrix. The XR system may calculate refined parameters and refined poses of the planar features based on one of the decomposed matrices with reduced rows, which reduces computational costs and memory storage. The refined parameters and refined poses of the planar features may be calculated such that projection errors of observations of the planar features from the respective refined poses onto the respective planar features are minimized.

[0039] The techniques described herein may be used together or separately with many types of devices and for many types of environments, including wearable or portable or autonomous devices with limited computing resources. In some embodiments, the techniques may be implemented by one or more services that form part of an XR system.

[0040] Exemplary System

[0041] 1 and 2 illustrate scenes with virtual content displayed in conjunction with a portion of the physical world. For illustrative purposes, an AR system is used as an example of an XR system. Figures 3-6B illustrate example AR systems including one or more processors, memory, sensors, and a user interface that can operate in accordance with the techniques described herein.

[0042] 1 , an outdoor AR scene 354 is depicted in which a user of the AR technology sees a physical-world park-like setting 356 featuring people, trees, a building in the background, and a concrete platform 358. In addition to these items, the user of the AR technology also perceives as "seeing" a robotic figure 357 standing on the physical-world concrete platform 358 and a flying, cartoon-like avatar character 352, which thereby appears to be an anthropomorphic bumblebee, although these elements (e.g., avatar character 352 and robotic figure 357) do not exist within the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is difficult to produce AR technology that facilitates a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or physical-world image elements.

[0043] Such AR scenes may be achieved using a system that allows a user to place AR content within the physical world, determines the location within a map of the physical world where the AR content is placed, saves the AR scene so that the placed AR content can be reloaded for display within the physical world, for example, between different AR experience sessions, and builds a map of the physical world based on the tracking information, allowing multiple users to share the AR experience. The system may build and update a digital representation of the physical world surfaces around the user. This representation may be used to render virtual content to appear fully or partially occluded by physical objects between the user and the rendered location of the virtual content, for placing virtual objects, in physics-based interactions, and for virtual character path planning and navigation, or for other operations in which information about the physical world is used.

[0044] 2 depicts another example of an indoor AR scene 400, illustrating an exemplary use case of an XR system, according to some embodiments. The exemplary scene 400 is a living room with a wall, a bookshelf on one side of the wall, a floor lamp in the corner of the room, a floor, a sofa, and a coffee table on the floor. In addition to these physical items, a user of the AR technology also perceives virtual objects, such as an image on the wall behind the sofa, a bird flying through the door, a deer peeking out from the bookshelf, and a decoration in the form of a windmill placed on the coffee table.

[0045] For an image on a wall, the AR technology requires information not only about the surface of the wall, but also about objects and surfaces in the room, such as lamp shapes, that occlude the image to properly render the virtual object. For a flying bird, the AR technology requires information about all objects and surfaces around the room to render the bird with realistic physics, such as avoiding objects and surfaces or bouncing off if the bird collides. For a deer, the AR technology requires information about surfaces, such as the floor or coffee table, to calculate where the deer should be placed. For a windmill, the system may identify the table as a separate object and determine that it is movable, while the corners of a shelf or wall may be determined to be stationary. Such specificity may be used in determining which parts of the scene to use or update in each of the various actions.

[0046] The virtual object may be placed within a previous AR experience session. When a new AR experience session begins in a living room, the AR technology requires that the virtual object be displayed exactly where it was previously placed and be realistically visible from a different perspective. For example, a windmill should appear to be standing on a book, rather than drifting above the table, even in a different location without the book. Such drifting may occur if the user's location in the new AR experience session is not accurately located within the living room. As another example, if the user is viewing a windmill from a different perspective than the perspective from which the windmill was placed, the AR technology requires the corresponding side of the windmill to be displayed.

[0047] The scene may be presented to the user via a system that includes multiple components, including a user interface that may stimulate one or more user senses, such as vision, hearing, and / or touch. In addition, the system may include one or more sensors that may measure parameters of the physical portion of the scene, including the user's position and / or movement within the physical portion of the scene. Furthermore, the system may include one or more computing devices with associated computer hardware, such as memory. These components may be integrated into a single device or distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.

[0048] 3 depicts an AR system 502 configured to provide an experience of AR content that interacts with a physical world 506, according to some embodiments. The AR system 502 may include a display 508. In the illustrated embodiment, the display 508 may be worn by a user as part of a headset such that the user may wear the display over their eyes, like a pair of goggles or glasses. At least a portion of the display may be transparent such that the user may observe a see-through reality 510. The see-through reality 510 may correspond to a portion of the physical world 506 within the current viewpoint of the AR system 502, which may correspond to the user's viewpoint when the user is wearing a headset incorporating both the display and sensors of the AR system and obtaining information about the physical world.

[0049] AR content may also be presented on the display 508, overlaid on the see-through reality 510. To provide accurate interaction between the AR content and the see-through reality 510 on the display 508, the AR system 502 may include a sensor 522 configured to capture information about the physical world 506.

[0050] The sensors 522 may include one or more depth sensors that output depth maps 512. Each depth map 512 may have multiple pixels, each of which may represent a distance to a surface in the physical world 506 in a particular direction relative to the depth sensor. Raw depth data may originate from the depth sensors to create depth maps. Such depth maps may be updated as fast as the depth sensors can form new images, which may be hundreds or thousands of times per second. However, the data may be noisy and incomplete and may have holes, shown as black pixels on the illustrated depth maps.

[0051] The sensors 522 may include other sensors, such as image sensors. The image sensors may obtain information, such as monocular or stereoscopic information, which may be processed to otherwise represent the physical world. In some embodiments, the system may include other sensors, such as one or more of the following: LiDAR cameras, RGBD cameras, infrared camera sensors, visible spectrum camera sensors, structured light emitters and / or sensors, infrared light emitters, coherent light emitters and / or sensors, gyroscope sensors, accelerometers, magnetometers, altimeters, proximity sensors, GPS sensors, ultrasonic emitters and detectors, and tactile interfaces. The sensor data may be processed in the world reconstruction component 516 to create a mesh representing connected portions of objects in the physical world. Metadata about such objects, including, for example, color and surface texture, may also be obtained using the sensors and stored as part of the world reconstruction. Metadata about the location of devices, including systems, may be determined or inferred based on the sensor data. For example, magnetometers, altimeters, GPS sensors, and the like may be used to determine or infer the location of a device. The location of a device may be of variable granularity. For example, the accuracy of a determined location of a device may vary from coarse (e.g., accurate to within a 10 meter spherical diameter), to fine (e.g., accurate to within a 3 meter spherical diameter), to very fine (e.g., accurate to within a 1 meter spherical diameter), to ultra-fine (e.g., accurate to within a 0.5 meter spherical diameter), etc.

[0052] The system may also obtain information about the user's head pose (or "pose") relative to the physical world. In some embodiments, a head pose tracking component of the system may be used to calculate head pose in real time. The head pose tracking component may represent the user's head pose in a coordinate frame with six degrees of freedom, including, for example, translation in three perpendicular axes (e.g., forward / back, up / down, left / right) and rotation about three perpendicular axes (e.g., pitch, yaw, and roll). In some embodiments, the sensor 522 may include an inertial measurement unit, which may be used to calculate and / or determine the head pose 514. The head pose 514 for a depth map may indicate, for example, the current viewpoint of the sensor capturing the depth map with six degrees of freedom, although the head pose 514 may also be used for other purposes, such as relating image information to a particular portion of the physical world or relating the position of a display worn on the user's head to the physical world.

[0053] In some embodiments, head pose information may be derived in a manner other than by an IMU, such as from analysis of objects in images. For example, the head pose tracking component may calculate the relative position and orientation of the AR device with respect to a physical object based on visual information captured by a camera and inertial information captured by an IMU. The head pose tracking component may then calculate the head pose of the AR device, for example, by comparing the calculated relative position and orientation of the AR device with features of the physical object. In some embodiments, the comparison may be made by identifying features in images captured using one or more of sensors 522 that are stable over time, such that changes in the position of these features in images captured over time can be associated with changes in the user's head pose.

[0054] Techniques for operating an XR system may provide an XR scene for a more immersive user experience. In such a system, an XR device may estimate head pose at a frequency of 1 kHz with low usage of computational resources. Such a device may be configured, for example, with four video graphics array (VGA) cameras operating at 30 Hz, one inertial measurement unit (IMU) operating at 1 kHz, the computational power of a single advanced RISC machine (ARM) core, less than 1 GB of memory, and less than 100 Mbps of network bandwidth. Techniques such as those described herein may be employed to generate and maintain maps, estimate head pose, and reduce the processing required to provide and consume data with low computational overhead. The XR system may calculate its pose based on matched visual features. U.S. Patent Application Publication No. US2019 / 0188474 describes hybrid tracking and is incorporated herein by reference in its entirety.

[0055] In some embodiments, the AR device may build a map from feature points recognized in successive images in a series of image frames captured as the user moves through the physical world with the AR device. Although each image frame may be obtained from a different pose as the user moves, the system may adjust the orientation of features in each successive image frame by matching features in the successive image frames with previously captured image frames to match the orientation of the initial image frame. Translation of the successive image frames can be used to align each successive image frame and match the orientation of the previously processed image frame so that points representing the same feature will match corresponding feature points from the previously collected image frame. The frames in the resulting map may have a common orientation established when the first image frame was added to the map. This map may be used to determine the user's pose in the physical world by matching features from the current image frame to the map, along with a set of feature points in a common frame of reference. In some embodiments, this map may be referred to as a tracking map.

[0056] Alternatively, or in addition, a map of the 3D environment around the user may be constructed by identifying planes or other surfaces based on image information. The positions of these surfaces per image may be correlated to create a representation. As described herein, techniques for efficiently optimizing such a map may be used to form the map. Such a map may be used to position virtual objects relative to the physical world or for different or additional functions. For example, a plane-based 3D representation may be used for head pose tracking.

[0057] In addition to enabling tracking of the user's pose within the environment, this map may enable other components of the system, such as a world reconstruction component 516, to determine the location of physical objects relative to the user. The world reconstruction component 516 may receive the depth map 512 and head pose 514 and any other data from the sensors and integrate the data into a reconstruction 518. The reconstruction 518 may be more complete and less noisy than the sensor data. The world reconstruction component 516 may update the reconstruction 518 using spatial and temporal averages of the sensor data from multiple viewpoints over time.

[0058] Reconstruction 518 may include a representation of the physical world in one or more data formats, including, for example, voxels, meshes, planes, etc. Different formats may represent alternative representations of the same portion of the physical world or may represent different portions of the physical world. In the illustrated example, on the left side of reconstruction 518, a portion of the physical world is presented as a global surface, and on the right side of reconstruction 518, a portion of the physical world is presented as a mesh.

[0059] In some embodiments, the map maintained by the head pose component 514 may be sparse with respect to other maps that may be maintained of the physical world. Rather than providing information about the location and possibly other characteristics of surfaces, the sparse map may indicate the location of points of interest and / or structures, such as corners or edges. In some embodiments, the map may include image frames as captured by the sensors 522. These frames may be reduced to features that may represent points of interest and / or structures. Along with each frame, information about the user's pose from which the frame was obtained may also be stored as part of the map. In some embodiments, all images acquired by the sensors may or may not be stored. In some embodiments, the system may process images as they are collected by the sensors and select a subset of image frames for further calculation. The selection may be based on one or more criteria that limit the addition of information but ensure that the map contains useful information. The system may add new image frames to the map based on overlap with previous image frames already added to the map, for example, or based on image frames containing a sufficient number of features determined to likely represent stationary objects. In some embodiments, a selected image frame or a group of features from a selected image frame may serve as a keyframe for a map, which is used to provide spatial information.

[0060] In some embodiments, the amount of data processed when constructing a map may be reduced by constructing a sparse map with a set of mapped points and keyframes, and / or by dividing the map into blocks and enabling block-wise updates, etc. The mapped points may be associated with points of interest in the environment. The keyframes may include information selected from camera-captured data. U.S. Patent Application Publication No. US2020 / 0034624 describes determining and / or evaluating a localization map and is incorporated herein by reference in its entirety.

[0061] The AR system 502 may integrate sensor data from multiple perspectives of the physical world over time. The pose (e.g., position and orientation) of the sensor may be tracked as the device containing the sensor is moved. As the sensor's frame pose and how it relates to other poses is understood, each of these multiple perspectives of the physical world may be fused together into a single combined reconstruction of the physical world, which may serve as an abstraction layer for the map and provide spatial information. The reconstruction may be more complete and less noisy than the original sensor data by using spatial and temporal averaging (i.e., averaging data from multiple perspectives over time) or any other suitable method.

[0062] 3, the map represents a portion of the physical world in which a user of a single wearable device resides. In that scenario, head poses associated with frames in the map may be represented as local head poses, indicating orientation relative to an initial orientation for the single device at the start of a session. For example, head poses may be tracked relative to an initial head pose when the device is turned on or otherwise operated to scan the environment and build a representation of that environment.

[0063] In combination with content characterizing that portion of the physical world, the map may include metadata. The metadata may, for example, indicate the capture time of the sensor information used to form the map. Alternatively, or in addition, the metadata may indicate the location of the sensor at the capture time of the information used to form the map. Location may be represented directly, such as using information from a GPS chip, or indirectly, such as using a wireless (e.g., Wi-Fi) signature indicating the strength of signals received from one or more wireless access points while the sensor data was being collected, and / or an identifier such as the BSSID of a wireless access point to which the user device connected while the sensor data was being collected.

[0064] Reconstruction 518 may be used for AR functions, such as producing a surface representation of the physical world for occlusion handling or physics-based processing. This surface representation may change as the user moves or as objects in the physical world change. Aspects of reconstruction 518 may be used by component 520, for example, to produce a changing global surface representation in world coordinates that can be used by other components.

[0065] AR content may be generated, such as by an AR application 504, based on this information. The AR application 504 may be, for example, a game program that performs one or more functions based on information about the physical world, such as visual occlusion, physics-based interactions, and environmental inference. It may perform these functions by querying data in different formats from the reconstruction 518 produced by the world reconstruction component 516. In some embodiments, the component 520 may be configured to output updates as the representation of the physical world within a region of interest changes. The region of interest may be set to approximate a portion of the physical world in the vicinity of a user of the system, such as a portion within the user's field of view, or may be projected (predicted / determined) to be within the user's field of view.

[0066] The AR application 504 may use this information to generate and update AR content, which virtual portions may be presented on the display 508 in combination with see-through reality 510 to create a realistic user experience.

[0067] In some embodiments, the AR experience may be provided to the user through an XR device, which may be a wearable display device that may be part of a system that may include remote processing and / or remote data storage, and / or, in some embodiments, other wearable display devices worn by other users. FIG. 4 illustrates an example of a system 580 (hereinafter referred to as “system 580”) that includes a single wearable device for ease of illustration. System 580 includes a head-mounted display device 562 (hereinafter referred to as “display device 562”) and various mechanical and electronic modules and systems that support the functionality of display device 562. Display device 562 may be coupled to a frame 564, which is wearable by a user or viewer 560 of the display system (hereinafter referred to as “user 560”) and configured to position display device 562 directly in front of the eyes of user 560. According to various embodiments, display device 562 may be a sequential display. Display device 562 may be monocular or binocular. In some embodiments, display device 562 may be an example of display 508 in FIG.

[0068] In some embodiments, a speaker 566 is coupled to the frame 564 and positioned proximate the ear canal of the user 560. In some embodiments, another speaker, not shown, is positioned adjacent another ear canal of the user 560 to provide stereo / adjustable sound control. The display device 562 is operably coupled, such as by wired leads or wireless connectivity 568, to a local data processing module 570, which may be mounted in a variety of configurations, such as fixedly attached to the frame 564, fixedly attached to a helmet or hat worn by the user 560, integrated into headphones, or otherwise removably attached to the user 560 (e.g., in a backpack configuration, in a belt-coupled configuration).

[0069] The local data processing module 570 may include a processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data includes a) data captured from sensors (e.g., which may be operatively coupled to the frame 564 or otherwise attached to the user 560), such as image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes, and / or b) data obtained and / or processed using the remote processing module 572 and / or the remote data repository 574, possibly for passing to the display device 562 after processing or retrieval.

[0070] In some embodiments, the wearable device may communicate with remote components. The local data processing module 570 may be operably coupled to a remote processing module 572 and a remote data repository 574 by communication links 576, 578, respectively, such as via wired or wireless communication links, such that these remote modules 572, 574 are operably coupled to each other and available as resources to the local data processing module 570. In further embodiments, in addition to or as an alternative to the remote data repository 574, the wearable device may access cloud-based remote data repositories and / or services. In some embodiments, the head pose tracking components described above may be implemented at least in part within the local data processing module 570. In some embodiments, the world reconstruction component 516 in FIG. 3 may be implemented at least in part within the local data processing module 570. For example, the local data processing module 570 may be configured to execute computer-executable instructions to generate a map and / or a physical world representation based, at least in part, on at least a portion of the data.

[0071] In some embodiments, processing may be distributed across local and remote processors. For example, local processing may be used to build a map (e.g., a tracking map) on the user device based on sensor data collected using sensors on the user's device. Such a map may be used by applications on the user's device. In addition, a previously created map (e.g., a reference map) may be stored in the remote data repository 574. If a suitable stored or persistent map is available, it may be used instead of or in addition to a tracking map created locally on the device. In some embodiments, the tracking map may be located relative to a stored map such that a correspondence is established between the tracking map, which may be oriented relative to the position of the wearable device at the time the user turned on the system, and the reference map, which may be oriented relative to one or more persistent features. In some embodiments, the persistent map may be loaded onto the user device, enabling the user device to render virtual content without the delay associated with scanning a location to build a tracking map of the user's complete environment from sensor data obtained during the scan. In some embodiments, a user device may access a remote persistent map (eg, stored on the cloud) without having to download the persistent map onto the user device.

[0072] In some embodiments, spatial information may be communicated from the wearable device to a remote service, such as a cloud service, configured to locate the device and store it in a map maintained on the cloud service. According to one embodiment, the localization process may occur in the cloud, matching the device location to an existing map, such as a reference map, and returning a transformation that links virtual content to the wearable device location. In such an embodiment, the system may avoid communicating maps from a remote resource to the wearable device. Other embodiments are configured for both device-based and cloud-based localization and may enable functionality when, for example, network connectivity is unavailable or the user chooses not to enable cloud-based localization.

[0073] Alternatively, or in addition, the tracking map may be merged with previously stored maps to enhance or improve the quality of those maps. Processing to determine whether suitable previously created environment maps are available and / or whether to merge the tracking map with one or more stored environment maps may occur within the local data processing module 570 or the remote processing module 572.

[0074] In some embodiments, the local data processing module 570 may include one or more processors (e.g., graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the local data processing module 570 may include a single processor (e.g., a single-core or multi-core ARM processor), which would limit the computational budget of the local data processing module 570 but enable smaller devices. In some embodiments, the world reconstruction component 516 may generate the physical world representation in real time over a non-predetermined space using a computational budget of less than a single Advanced RISC Machine (ARM) core, such that the remaining computational budget of the single ARM core can be accessed for other uses, such as, for example, extracting meshes.

[0075] The processing as described herein for optimizing the map of the 3D environment may be performed within any processor of the system, however, the reduced computation and reduced memory required for optimization as described herein may allow such operations to be performed with low latency on a local processor that is part of the wearable device.

[0076] In some embodiments, the remote data repository 574 may include a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 570, allowing for fully autonomous use from the remote module. In some embodiments, all data is stored and all or most computations are performed in the remote data repository 574, allowing for smaller devices. World reconstructions may be stored in whole or in part in this repository 574, for example.

[0077] In embodiments in which data is stored remotely and accessible via a network, the data may be shared by multiple users of the augmented reality system. For example, a user device may upload its tracking map and augment it into a database of environmental maps. In some embodiments, uploading of the tracking map occurs at the end of a user session with the wearable device. In some embodiments, uploading of the tracking map may occur continuously, semi-continuously, intermittently, at predefined times, after a predefined period since the previous upload, or when triggered by an event. A tracking map uploaded by any user device may be used to augment or refine a previously stored map, whether based on data from that user device or any other user device. Similarly, a persistent map downloaded to a user device may be based on data from that user device or any other user device. In this way, high-quality environmental maps may be readily available to users to refine their experience with the AR system.

[0078] In further embodiments, downloading of persistent maps may be limited and / or avoided based on localization performed on remote resources (e.g., in the cloud). In such a configuration, a wearable device or other XR device communicates feature information (e.g., positioning information about the device at the time the feature represented in the feature information was sensed) combined with pose information to a cloud service. One or more components of the cloud service may match the feature information with a separate stored map (e.g., a reference map) and generate a transformation between the tracking map maintained by the XR device and the coordinate system of the reference map. Each XR device, having its tracking map localized relative to the reference map, can accurately render virtual content at a location defined relative to the reference map based on its own tracking.

[0079] In some embodiments, local data processing module 570 is operably coupled to battery 582. In some embodiments, battery 582 is a removable power source, such as a commercially available battery. In other embodiments, battery 582 is a lithium-ion battery. In some embodiments, battery 582 includes both an internal lithium-ion battery that is rechargeable by user 560 during periods of non-operation of system 580, and a removable battery, so that user 560 can operate system 580 for longer periods of time without being plugged in and having to charge the lithium-ion battery or shut off system 580 and replace the battery.

[0080] FIG. 5A illustrates a user 530 wearing an AR display system that renders AR content as the user 530 moves through a physical world environment 532 (hereinafter referred to as “environment 532”). Information captured by the AR system along the user's path of movement may be processed into one or more tracking maps. The user 530 positions the AR display system at a location 534, and the AR display system records ambient information of the passable world (e.g., digital representations of real objects in the physical world that may be stored and updated as the real objects change) relative to the location 534. That information may be stored as a pose in combination with images, features, directional audio input, or other desired data. The location 534 is aggregated with data input 536, e.g., as part of a tracking map, and processed by at least a passable world module 538, which may be implemented, for example, by processing on the remote processing module 572 of FIG. 4. In some embodiments, the passable world module 538 may include a head pose component 514 and a world reconstruction component 516 so that the processed information, in combination with other information about physical objects used in the rendered virtual content, may indicate the location of the objects in the physical world.

[0081] The passable world module 538 determines, at least in part, where and how the AR content 540 can be placed in the physical world, as determined from the data input 536. The AR content is “placed” in the physical world by presenting both a representation of the physical world and the AR content via a user interface, where the AR content is rendered as if interacting with objects in the physical world, and the objects in the physical world are presented as if the AR content obscures the user's view of those objects, when appropriate. In some embodiments, the AR content may be placed by appropriately selecting a portion of a fixed element 542 (e.g., a table) from the reconstruction (e.g., reconstruction 518) and determining the shape and position of the AR content 540. As an example, the fixed element may be a table, and the virtual content may be positioned to appear on the table. In some embodiments, the AR content may be placed within a structure in the field of view 544, which may be the current field of view or an estimated future field of view. In some embodiments, the AR content may be persisted to a model 546 (e.g., a mesh) of the physical world.

[0082] As depicted, fixed element 542 serves as a proxy (e.g., a digital copy) for any fixed element in the physical world, which may be stored within passable world module 538 so that user 530 may perceive content on fixed element 542 without the system having to map it to fixed element 542 each time user 530 sees it. Fixed element 542 may thus be a mesh model, determined from a previous modeling session or from a separate user, but stored by passable world module 538 for future reference by multiple users. Thus, passable world module 538 may recognize environment 532 from previously mapped environments and display AR content without user 530's device having to first map all or part of environment 532, saving computational processes and cycles and avoiding latency for any rendered AR content.

[0083] A mesh model 546 of the physical world may be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying the AR content 540 can be stored by the passable world module 538 for future retrieval by the user 530 or other users, without having to recreate the model, in whole or in part. In some embodiments, the data inputs 536 are inputs such as geographic location, user identification, and current activity to indicate to the passable world module 538 which fixed elements 542 of one or more fixed elements are available, the AR content 540 last placed on the fixed elements 542, and whether that same content should be displayed (such AR content is “persistent” content regardless of whether the user is viewing a particular passable world model).

[0084] Even in embodiments where objects are considered fixed (e.g., a kitchen table), the passable world module 538 may update those objects in the model of the physical world from time to time to account for possible changes in the physical world. Models of fixed objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed (e.g., a kitchen chair). To render the AR scene with a realistic feel, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update fixed objects. To enable accurate tracking of all of the objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.

[0085] 5B is a schematic illustration of a viewing optics assembly 548 and associated components. In some embodiments, two eye tracking cameras 550 are pointed towards the user's eye 549 to detect metrics of the user's eye 549, such as eye shape, eyelid occlusion, pupil direction, and phosphenes on the user's eye 549.

[0086] In some embodiments, one of the sensors may be a depth sensor 551, such as a time-of-flight sensor, that emits signals into the world, detects reflections of those signals from nearby objects, and determines the distance to a given object. The depth sensor may, for example, quickly determine whether objects have entered the user's field of view, either as a result of the objects' movement or a change in the user's posture. However, information about the location of objects within the user's field of view may alternatively or additionally be collected using other sensors. Depth information may be obtained, for example, from a stereoscopic image sensor or a plenoptic sensor.

[0087] In some embodiments, world camera 552 records, maps, and / or otherwise models a wider-than-surrounding view of environment 532 and detects inputs that may affect the AR content. In some embodiments, world camera 552 and / or camera 553 may be grayscale and / or color image sensors that output grayscale and / or color image frames at fixed time intervals. Camera 553 may also capture physical world images within the user's field of view at specific times. Pixels of a frame-based image sensor may be repeatedly sampled even if their values ​​remain constant. World camera 552, camera 553, and depth sensor 551 each have separate fields of view 554, 555, and 556, and collect and record data from a physical world scene, such as physical world environment 532 depicted in FIG. 34A .

[0088] The inertial measurement unit 557 may determine the movement and orientation of the viewing optics assembly 548. In some embodiments, the inertial measurement unit 557 may provide an output indicating the direction of gravity. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 551 is operably coupled to the eye tracking camera 550 as a confirmation of the measured accommodation to the actual distance seen by the user's eyes 549.

[0089] It should be understood that viewing optics assembly 548 may include some of the components illustrated in FIG. 34B , or may include components instead of or in addition to the components illustrated. In some embodiments, for example, viewing optics assembly 548 may include two world cameras 552 instead of four. Alternatively, or in addition, cameras 552 and 553 need not capture visible light images of their full field of view. Viewing optics assembly 548 may include other types of components. In some embodiments, viewing optics assembly 548 may include one or more dynamic vision sensors (DVS), whose pixels may asynchronously respond to relative changes in light intensity exceeding a threshold.

[0090] In some embodiments, the viewing optics assembly 548 may not include a depth sensor 551 based on time-of-flight information. In some embodiments, for example, the viewing optics assembly 548 may include one or more plenoptic cameras, the pixels of which may capture the light intensity and angle of incident light, from which depth information can be determined. For example, the plenoptic camera may include an image sensor overlaid with a transmissive diffractive mask (TDM). Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a depth information source instead of, or in addition to, the depth sensor 551.

[0091] 5B is provided as an example. Viewing optics assembly 548 may include components in any suitable configuration, which may be configured to provide the user with the largest field of view practical for a particular set of components. For example, if viewing optics assembly 548 has one world camera 552, the world camera may be located within a central region of the viewing optics assembly instead of on the side.

[0092] Information from sensors in the viewing optics assembly 548 may be coupled to one or more processors in the system. The processor may generate data that can be rendered to cause the user to perceive the virtual content as interacting with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data that depicts both physical and virtual objects. In other embodiments, physical and virtual content may be depicted in a single scene by modulating the opacity of a display device through which the user sees the physical world. The opacity may be controlled to create the appearance of virtual objects and block the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data may include only the virtual content, which may be modified (e.g., clipping content and accounting for occlusion) so that the virtual content is perceived by the user to interact realistically with the physical world when viewed through the user interface.

[0093] The locations on the viewing optics assembly 548 where content can be displayed to create the impression of an object at a particular location may depend on the physics of the viewing optics assembly. Additionally, the user's head posture relative to the physical world and the direction the user's eyes are looking will affect the location within the physical world content displayed at a particular location on the viewing optics assembly where the content will appear. Sensors such as those described above may provide information from which this information can be collected and / or calculated so that a processor receiving the sensor input can calculate where objects should be rendered on the viewing optics assembly 548 to create a desired appearance for the user.

[0094] Regardless of how content is presented to a user, a model of the physical world may be used so that the characteristics of virtual objects that may be affected by physical objects, including the shape, position, movement, and visibility of the virtual objects, may be correctly calculated. In some embodiments, the model may include a reconstruction of the physical world, such as reconstruction 518.

[0095] The model may be created from data collected from sensors on the user's wearable device, although in some embodiments the model may be created from data collected by multiple users, which may be aggregated in a computing device remote from all users (and may be "in the cloud").

[0096] The model may be created, at least in part, by a world reconstruction system, such as the world reconstruction component 516 of FIG. 3 , which is depicted in further detail in FIG. 6A . The world reconstruction component 516 may include a perception module 660, which may generate, update, and store representations for portions of the physical world. In some embodiments, the perception module 660 may represent a portion of the physical world within a sensor's reconstruction range as a plurality of voxels. Each voxel corresponds to a 3D cube of a predetermined volume within the physical world and may include surface information, indicating whether a surface exists within the volume represented by the voxel. A voxel may be assigned a value indicating whether its corresponding volume has been determined to contain a surface of a physical object, has been determined to be empty, or has not yet been measured with a sensor, and therefore its value is unknown. It should be understood that values ​​indicating voxels determined to be empty or unknown need not be explicitly stored, and that voxel values ​​may be stored in computer memory in any suitable manner, including not storing information about voxels determined to be empty or unknown.

[0097] In addition to generating information for the persisted world representation, perception module 660 may identify and output indications of changes in the area surrounding the user of the AR system. Such indications of changes may trigger updates to the stereoscopic data stored as part of the persisted world, or trigger other functionality, such as triggering the generate AR content and update AR content component 604.

[0098] In some embodiments, the perception module 660 may identify changes based on a signed distance function (SDF) model. The perception module 660 may be configured to receive sensor data, such as a depth map 660a and a head pose 660b, and then fuse the sensor data into an SDF model 660c. The depth map 660a may provide SDF information directly, or images may be processed to arrive at the SDF information. The SDF information represents distances from sensors used to capture that information. Because those sensors may be part of the wearable unit, the SDF information may represent the physical world from the perspective of the wearable unit and, therefore, the user. The head pose 660b may allow the SDF information to be associated with voxels in the physical world.

[0099] In some embodiments, perception module 660 may generate, update, and store representations for portions of the physical world that are within its perception range. The perception range may be determined, at least in part, based on the reconstruction range of the sensor, which may be determined, at least in part, based on the limits of the sensor's observation range. As a specific example, an active depth sensor that operates using active IR pulses may operate reliably over a range of distances, creating a sensor's observation range that may be a few centimeters or tens of centimeters to several meters.

[0100] The world reconstruction component 516 may include additional modules that may interact with the perception module 660. In some embodiments, the persisted world module 662 may receive a representation for the physical world based on data obtained by the perception module 660. The persisted world module 662 may also include representations of the physical world in various formats. For example, volumetric metadata 662b, such as voxels, may be stored along with the meshes 662c and planes 662d. In some embodiments, other information, such as depth maps, may also be saved.

[0101] In some embodiments, a representation of the physical world such as that illustrated in FIG. 6A may provide relatively dense information about the physical world compared to a sparse map such as a tracking map based on feature points, as described above.

[0102] In some embodiments, perception module 660 may include modules that generate representations for the physical world in various formats, including, for example, meshes 660d, planes, and semantics 660e. Representations for the physical world may be stored across local and remote storage media. Representations for the physical world may be described in different coordinate frames, for example, depending on the location of the storage media. For example, a representation for the physical world stored in a device may be described in a coordinate frame local to the device. A representation for the physical world may have a counterpart stored in the cloud. The counterpart in the cloud may be described in a coordinate frame shared by all devices in the XR system.

[0103] In some embodiments, these modules may generate representations based on data within the perception range of one or more sensors at the time the representation is generated, data captured at earlier times, and information in the persisted world module 662. In some embodiments, these components may operate on depth information captured using a depth sensor. However, AR systems may also include vision sensors and generate such representations by analyzing monocular or binocular visual information.

[0104] In some embodiments, these modules may operate on regions of the physical world. They may be triggered to update subregions of the physical world when perception module 660 detects a change in the physical world within that subregion. Such a change may be detected, for example, by detecting a new surface in SDF model 660c or by other criteria, such as a change in the values ​​of a sufficient number of voxels representing the subregion.

[0105] The world reconstruction component 516 may include components 664 that may receive representations of the physical world from the perception module 660. Information about the physical world may be pulled by these components, for example, according to usage requests from applications. In some embodiments, information may be pushed to the usage components, such as via indications of changes in pre-identified areas or changes in the physical world representation within the perception range. The components 664 may include, for example, game programs and other components that implement processing for visual occlusion, physics-based interactions, and environmental inference.

[0106] In response to a query from component 664, perception module 660 may transmit a representation for the physical world in one or more formats. For example, when component 664 indicates the use is for visual occlusion or physics-based interaction, perception module 660 may transmit a representation of a surface. When component 664 indicates the use is for environmental inference, perception module 660 may transmit meshes, planes, and semantics of the physical world.

[0107] In some embodiments, perception module 660 may include a component that provides formatting information to component 664. An example of such a component may be ray casting component 660f. A usage component (e.g., component 664) may, for example, query information about the physical world from a particular viewpoint. Ray casting component 660f may select from one or more representations of the physical world data within the field of view from that viewpoint.

[0108] As should be understood from the foregoing description, the perception module 660 or another component of the AR system may process data and create a 3D representation of a portion of the physical world. The data to be processed may be reduced, at least in part, by thinning a portion of the 3D reconstruction volume based on the camera frustum and / or depth image; extracting and persisting planar data; capturing, persisting, and updating the 3D reconstruction data in blocks that enable local updates while maintaining neighborhood consistency; deriving occlusion data from a combination of one or more depth data sources; providing occlusion data to an application that generates such a scene; and / or performing multi-stage mesh simplification. The reconstruction may contain data of different levels of sophistication, including, for example, raw data such as live depth data, fused volumetric data such as voxels, and calculated data such as meshes.

[0109] In some embodiments, components of the passable world model may be distributed, with some parts running locally on the XR device and some parts running remotely, such as on a network connected to a server or otherwise in the cloud. The distribution of information processing and storage between the local XR device and the cloud can affect the functionality and user experience of the XR system. For example, reducing processing on the local device by distributing processing to the cloud may enable longer battery life and reduce heat generated on the local device. However, distributing much more processing to the cloud may create undesirable latency that causes an unacceptable user experience.

[0110] FIG. 6B depicts a distributed component architecture 600 configured for spatial computing, according to some embodiments. The distributed component architecture 600 may include a passable world component 602 (e.g., PW 538 in FIG. 5A ), a Lumin OS 604, an API 606, an SDK 608, and applications 610. The Lumin OS 604 may include a Linux-based kernel with custom drivers compatible with XR devices. The API 606 may include an application programming interface that gives XR applications (e.g., applications 610) access to the spatial computing features of the XR device. The SDK 608 may include a software development kit that enables the creation of XR applications.

[0111] One or more components in architecture 600 may create and maintain a model of the passable world. In this example, sensor data is collected on the local device. Processing of that sensor data may be performed partially locally on the XR device and partially in the cloud. PW 538 may include an environment map created at least in part based on data captured by AR devices worn by multiple users. During a session of an AR experience, individual AR devices (such as the wearable devices described above in connection with FIG. 4) may create a tracking map, which is one type of map.

[0112] In some embodiments, a device may include components that build both sparse and dense maps. The tracking map may serve as a sparse map and may contain information about the head pose of the AR device scanning the environment and the objects detected in the environment at each head pose. These head poses may be maintained locally per device. For example, the head pose on each device may be relative to the initial head pose when the device is turned on for the session. As a result, each tracking map may be local to the device that creates it and may have its own frame of reference defined by its own local coordinate system. However, in some embodiments, the tracking map on each device may be formed such that one coordinate of its local coordinate system is aligned with the direction of gravity as measured by its sensors, such as the inertial measurement unit 557.

[0113] The dense map may include surface information, which may be represented by a mesh or depth information. Alternatively, or in addition, the dense map may include higher level information derived from the surface or depth information, such as the locations and / or properties of planes and / or other objects.

[0114] The creation of the dense map may be independent from the creation of the sparse map in some embodiments. The creation of the dense map and the sparse map may be performed, for example, in separate processing pipelines within the AR system. Separating the processing may allow, for example, the generation or processing of different types of maps to be performed at different rates. The sparse map may, for example, be refreshed at a faster rate than the dense map. However, in some embodiments, the processing of the dense and sparse maps may be related even when performed in different pipelines. Changes in the physical world revealed in the sparse map may, for example, trigger an update of the dense map, or vice versa. Furthermore, even when created independently, the maps may be used together. For example, a coordinate system derived from the sparse map may be used to define the position and / or orientation of objects in the dense map.

[0115] The sparse map and / or the dense map may be persisted for reuse by the same device and / or for sharing with other devices. Such persistence may be achieved by storing the information in a cloud. The AR device may send the tracking map to the cloud and merge it with an environment map selected from persisted maps previously stored in the cloud, for example. In some embodiments, the selected persisted map may be sent from the cloud to the AR device for merging. In some embodiments, the persisted map may be oriented with respect to one or more persistent coordinate frames. Such maps may serve as reference maps because they may be used by any of multiple devices. In some embodiments, a model of the passable world may consist of or be created from one or more reference maps. A device may perform some operations based on a coordinate frame local to the device, but may use the reference map by determining a transformation between its coordinate frame local to the device and the reference map.

[0116] The reference map may originate as a tracking map (TM) (e.g., TM 1102 in FIG. 31A ), which may be promoted to a reference map. The reference map may be persisted so that a device accessing the reference map, once it has determined the transformation between its local coordinate system and the coordinate system of the reference map, can use the information in the reference map to determine the location of objects represented in the reference map in the physical world around the device. In some embodiments, the TM may be a head pose sparse map created by an XR device. In some embodiments, the reference map may be created when an XR device sends one or more TMs to a cloud server to be merged with additional TMs captured by the XR device at different times or by other XR devices.

[0117] In embodiments in which a tracking map is formed on the local device with coordinates in one local coordinate frame aligned with gravity, this orientation with respect to gravity may be preserved upon creation of the reference map. For example, when a tracking map submitted for merging does not overlap with any previously stored map, the tracking map may be promoted to the reference map. Other tracking maps, which may also have orientation with respect to gravity, may then be merged with the reference map. Merging may be performed to ensure that the resulting reference map retains its orientation with respect to gravity. Two maps may not be merged, regardless of the correspondence of feature points in those maps, for example, if the gravity-aligned coordinates of each map do not match each other with a sufficiently close tolerance.

[0118] A reference map or other map may provide information about the portion of the physical world represented by the processed data to create the individual map. FIG. 7 depicts an example tracking map 700, according to some embodiments. The tracking map 700 may provide a top view 706 of a physical object in the corresponding physical world represented by points 702. In some embodiments, the map points 702 may represent features of the physical object, which may include multiple features. For example, each corner of a table may be a feature represented by a point on the map. Features may be derived from processed images, such as those obtained using sensors in a wearable device in an augmented reality system. Features may be derived, for example, by processing image frames output by the sensors and identifying features based on large gradients in the images or other suitable criteria. Further processing may limit the number of features in each frame. For example, processing may select features that are likely to represent persistent objects. One or more heuristics may be applied for this selection.

[0119] The tracking map 700 may include data about points 702 collected by the device. A pose may be stored for each image frame with data points included in the tracking map. The pose may represent the orientation from which the image frame was captured so that feature points in each image frame can be spatially correlated. The pose may be determined by positioning information, such as may be derived from a sensor on the wearable device, such as an IMU sensor. Alternatively, or in addition, the pose may be determined from matching an image frame with another image frame that depicts an overlapping portion of the physical world. By finding such positional correlation, which may be accomplished by matching subsets of feature points in the two frames, a relative pose between the two frames may be calculated. Relative pose may be appropriate for the tracking map because the map may be relative to a coordinate system local to the device, which is established based on the device's initial pose when construction of the tracking map began.

[0120] Because much of the information collected using the sensors is likely redundant, not all of the feature points and image frames collected by the device may be retained as part of the tracking map. Rather, only certain frames may be added to the map. These frames may be selected based on one or more criteria, such as the degree of overlap with image frames already in the map, the number of new features they contain, or a quality metric for the features in the frame. Image frames that are not added to the tracking map may be discarded or used to revise feature locations. As a further alternative, all or most of the image frames, represented as a set of features, may be retained, but a subset of those frames may be designated as key frames, which are used for further processing.

[0121] The keyframes may be processed to produce a keyrig 704. The keyframes may be processed to produce a three-dimensional set of feature points and stored as the keyrig 704. Such processing may involve, for example, comparing image frames derived simultaneously from two cameras and stereoscopically determining the 3D positions of the feature points. Metadata such as pose may be associated with these keyframes and / or keyrigs.

[0122] The environment map may have any of a number of formats, depending on, for example, the storage location of the environment map, including, for example, local storage and remote storage of the AR device. For example, a map in remote storage may have a higher resolution than a map in local storage on the wearable device when memory is limited. To transmit a higher resolution map from remote storage to local storage, the map may be downsampled or otherwise converted to a suitable format, such as by reducing the number of poses per area of ​​the physical world stored in the map and / or the number of feature points stored per pose. In some embodiments, a slice or portion of the high-resolution map from remote storage may be transmitted to local storage, and the slice or portion is not downsampled.

[0123] The database of environment maps may be updated as new tracking maps are created. To determine which of the potentially numerous environment maps in the database should be updated, updating may include efficiently selecting one or more environment maps stored in the database associated with the new tracking map. The selected one or more environment maps may be ranked by relevance, and one or more of the highest-ranked maps may be selected for processing to merge the new tracking map with the higher-ranked selected environment map and create one or more updated environment maps. When a new tracking map represents a portion of the physical world over which there is no existing environment map to update, the tracking map may be stored in the database as a new environment map.

[0124] Techniques for efficient processing of maps using planes.

[0125] 7 illustrates a map based on feature points. Some systems may create maps based on recognized surfaces, such as planes, rather than individual points. In some embodiments, rather than containing map points 702, the map may contain planes. For example, such a map may contain planes to represent the surface of a table, rather than a collection of feature points representing the corners of the table and possibly points on the surface.

[0126] The process of creating / updating such a map may involve calculating features (such as position and orientation) of planes in images acquired from various poses. However, pose determination may depend on estimated features of previously detected planes. Creating and updating the map may therefore involve jointly optimizing the estimated features describing the planes and the sensor poses at which the images used to estimate the features of the planes were captured. The joint optimization may be referred to as plane bundle adjustment.

[0127] The inventors have recognized and appreciated that the conventional approach of planar bundle adjustment, applied to maps with one or more planar features, can result in a large-scale nonlinear least-squares problem, which incurs high computational cost and memory consumption. Unlike traditionally implemented bundle adjustment for maps represented by feature points, in which one recording of a sensor can result in a single observation of a point feature, planar features can be infinite objects, and one image from a sensor can provide multiple observations of the planar feature. Each observation of the planar feature imposes constraints on the planar parameters of the planar feature and the pose parameters of the pose, which can result in a large-scale nonlinear least-squares problem, even for small workspaces.

[0128] The inventors have recognized and appreciated plane bundle adjustment, which reduces computational cost and memory usage. Figure 8A is a schematic diagram illustrating the observation of various planes at different poses, according to some embodiments. In the illustrated example, there are M planes and N sensor poses. The rotation and translation of the ith pose are R i∈SO3 and t i ∈R 3 The j-th plane is shown as π j =[n j ;d j ], where n is the surface normal with ||n||2=1, and d is the negative distance from the coordinate system origin to the plane. The measurement of the jth plane at the ith pose is defined as follows: K ij It is a set of points. [ka] each p ijk ∈P ij provides one constraint on the ith pose and the jth plane, which is illustrated in Figure 8B. ijk is p ijk From the plane π j is the signed distance to , which can be written as: [ka]

[0129] t i Unlike the rotation R i and the plane parameter π j has further constraints. For example, R i can be parameterized by quaternions, angular axes, or Euler angles. The plane parameter π j can be expressed in terms of quaternion-based, homogeneous coordinates, closest point parameters, or minimal parameterization. Regardless of the particular parameterization, the present embodiment i →R(θ i ) and ω j →π(ω j ) and present an arbitrary parameterization with respect to rotation and plane parameters. i and t i is related to the sensor pose, which is i =[θ i ;t i ] may be combined asi may have six or seven unknowns (seven for quaternions and six for minimal representations of rotations such as angular axis parameterizations), which is varied for parameterizations of the rotation matrix. j may have three or four unknowns (three for the minimal representation of the plane and four for the homogeneous coordinates of the plane). Using these notations, the residual δ ijk is ρ i and ω j is a function of

[0130] Planar bundle adjustment solves the problem of minimizing the following nonlinear least-squares problem for all ρ i( i ≠ 1) and ω j It can be a matter of refining both. [ka] Here, the first pose ρ1 is fixed during the optimization, rigidly anchoring the coordinate system.

[0131] The Jacobian matrix of the plane bundle adjustment may be calculated and provided to an algorithm (e.g., the Levenberg-Marquardt (LM) algorithm and the Gaussian-Newton algorithm) that solves the least squares problem. The observation of the jth plane at the ith pose is represented by the point set P ij The Jacobian matrix J ij is the point set P ij The Jacobian matrix J may be derived from ij It can also be a stack of P ij Inside K ij Assume that there are points J ij (i≠1) has the following form: [ka]

[0132] [ka] The partial derivative of is the residual δ ijk This may be defined as follows: [ka] R as defined above i The elements of θ i A function function of d j and n j The elements of ω j Note that this is a function of p ijk may represent the kth measurement of the jth plane at the ith pose. jik , y ijk , and z ijk is the coordinate frame p in which it can be the sensor coordinate frame, device coordinate frame or reference coordinate frame. ijk The residual δ ijk can be calculated by substituting (10) into (7) and expanding it. This gives: [ka] (11) may be rewritten as follows: [ka] In the formula, c ijk and ν ij is a 13-dimensional vector as follows: [ka] c ijk The elements in the observation p ijk or from 1, which is a constant. On the other hand, ν ij The elements in are i and ω j is a function of

[0133] δ ijk The partial derivative of is i But, n ρ There are unknowns, and ω j But, n ω This can be calculated by assuming that we have unknowns, which may be defined as follows: [ka] ζ ij d is ζ ij Assume that the d-th element of ζ is ij d δ for ijk The partial derivative of has the form: [ka] During the ceremony, [ka] is a 13-dimensional vector whose elements are ζ ij d ν for ij is the partial derivative of the element of . According to (15), [ka] has the following form: [ka] V ρi may be a 13x6 or 13x7 matrix (13x7 for quaternions and 13x6 for minimal representations of rotation matrices). Similarly, [ka] has the following form: [ka] V ωj may be a 13x3 or 13x4 matrix (13x3 for the minimal representation of the plane and 13x4 for the homogeneous coordinates of the plane).

[0134] Jacobian matrix J ij may be calculated based on the following: [ka] where the kth row c ijk is defined in (13). C ij is size K ij ×13 matrix. Jacobian matrix J ij is C in (18) ij may be calculated by substituting (16) and (17) into (9) using the definition of [ka]

[0135] P 1j is the j-th plane ω in the first orientation ρ1 j Since ρ1 can be fixed during optimization, P 1j The Jacobian matrix J derived from 1j may have the following form: [ka]

[0136] C ij can be written in the following form: [ka] In the formula, M ij is 4x13 in size, and Q T ij Qij =I4I4 is a 4×4 identity matrix. c in (13) ijk As shown in the definition of x ijk , y ijk , z ijk , and 1 is replicated several times, c ijk Therefore, C ij There are only four individual columns among the 13 columns of P ij It contains the x, y, and z coordinates of the points in the image. This can be shown as follows: [ka] C ij Column 13 of the ij It is a copy of the four columns in E ij can be decomposed as follows: [ka] In the formula, Q T ij Q ij = I4 and U ij is an upper triangular matrix. Q ij is size K ij ×4, and U ij is of size 4x4. Point K ij Since the number of U is generally much larger than 4, QR decomposition may be used. Thin QR decomposition can reduce the calculation time. ij can be partitioned by that column as follows: [ka] Substituting (24) into (22) gives: [ka] Comparing (22) and (23) yields: [ka] C ij The column of E ij Since it is a copy of the column of ijk and E in (22) ij According to the definition of C ij can be written as follows: [ka] Substituting (26) into (27) yields: [ka]

[0137] C ij Decomposition of can be used to significantly reduce computational costs. Although thinQR decomposition is described in the examples, other decomposition methods may also be used, including, for example, singular value decomposition (SVD).

[0138] J ij The reduced Jacobian matrix J of r ij is C ij may be calculated based on, as explained above, C ij =Q ij M ij It may be decomposed as: [ka]

[0139] The reduced Jacobian matrix J r ij is M ij But C ij Since the Jacobian matrix J is much smaller than ij It has fewer rows. C ij is size K ij × 13. In contrast, M ijhas size 4 × 13. Generally, K ij is much larger than 4.

[0140] J ij and J. r ij is the Jacobian matrix J and the reduced Jacobian matrix J for the cost function (8) as follows: r may be stacked to form [ka]

[0141] J r is replaced by J and solves the least squares problem. T We can calculate J. For planar bundle adjustment, J T J=J rT J r J and J r is defined as (30), J ij and J. r ij From the point of view of , it is a block vector. Following block matrix multiplication, [ka] For i ≠ 1, using the expression in (19), J T ij J ij has the following form: [ka] Similarly, using the expression in (29), J r ij T J r ij has the following form: [ka] Fact QT ij Q ij Using =I4, (21) can be [ka] Substitute into [ka] Similarly, [ka] For i=1, according to (20), J T 1j J 1j The only non-zero term in V T ωj C T 1j C 1j V ωj On the other hand, according to (29), J r 1j T J r 1j has only one corresponding non-zero term V T ωj M T 1j M 1j V ωj Similar to the derivation in (34), we have [ka] In short, using (34), (35), and (36), J T ij J ij =J r ij T J r ij According to (31), as a result, J T J=J rT J r This becomes:

[0142] Pij K inside ij The residual vector for points is [ka] According to (12), δ ij can be written as follows: [ka] δ ij The reduced residual vector δ r ij can be defined as follows: [ka] All δ ij and δ r ij Stacking the residual vector δ and the reduced residual vector δ r has the following form: [ka] The reduced residual vector δ r can be substituted for the residual vector δ in the algorithm, which solves the least squares problem. For planar bundle adjustment, J T δ=J rT δ r J, J r , δ and δ r are the elements J as defined in (30) and (39), respectively. ij , J r ij , δ ij , and δ r ij Applying block matrix multiplication gives [ka] For i ≠ 1, J in (19) ij and J in (29) r ij and δ in (37) ij and δ in (38) r ij Using the expression, J T ij δ ij and J. r ij T δ r ij has the following form: [ka] Fact Q T ij Q ij Using =I4, (21) can be [ka] Substitute into [ka] Similarly, [ka] For i=1, (20) and (37) are T 1j δ 1j and apply block matrix multiplication, J T 1j δ 1j The only non-zero term in is V T ω1j C T 1j C 1j ν1 j On the other hand, (29) and (38) can be expressed as r 1j T δ r 1j Substituting into, J r 1jT δ r 1j Only one non-zero term V T ω1j M T 1j M 1j ν1 j Similar to the derivation in (42), we have [ka] In short, from (42), (43), and (44), J T ij δ ij =J r ij T δ ij According to (40), J T δ=J rT It becomes δ.

[0143] It should be understood that in some embodiments, a system using a reduced Jacobian matrix may calculate the reduced Jacobian matrix directly from sensor data or other information, without iteratively forming or calculating the Jacobian matrix to arrive at the reduced matrix form described herein.

[0144] For planar bundle adjustment, the reduced Jacobian matrix J r and the reduced residual vector δ r can be substituted for J and δ in (4) to calculate the steps in the algorithm that solves the least squares problem, r and δ r Each block in J r ij and δ ij r has 4 rows. The algorithm solves the least squares problem by computing the steps at each iteration using (4). T J=J rT J r , J T δ=J rT Since δ, [ka] teeth, [ka] Therefore, J r and δ r can be substituted for J and δ to calculate the steps in the algorithm that solves the least squares problem. r ij and δ in (38) r ij According to the definition of J r ij and δ r ij The number of rows in M ij The number of rows in J is the same as that in J, which has four rows. r and δ r has four rows. Therefore, regardless of the number of possible points in Pij, the reduced J r ij and δ r ij has at most four rows. This significantly reduces the computational cost in the algorithm for solving the least squares problem. The additional cost here is ij =Q ij M ij This is to calculate C ij is calculated once before the iteration, since it is kept constant between iterations.

[0145] In some embodiments, the planar bundle adjustment is performed by adjusting the initial predictions and measurements of N poses and M planar parameters {P ij} can be obtained. Planar bundle adjustment is performed by ijk ∈P ij For each matrix block c ijk Calculate the matrix block C as shown in (16). ij The planar bundle adjustment is as follows: ij is an orthogonal matrix Q ij and upper triangular matrix Mij The upper triangular matrix M ij is provided to the algorithm (e.g., Levenberg-Marquardt (LM) algorithm and Gaussian-Newton algorithm) that solves the least squares problem, and is used to obtain the reduced Jacobian matrix block J as shown in (27). r ij and the reduced residual block δ as in (36). r ij and may be calculated as follows: r and the reduced residual vector δ as in (37). r The least squares problem may be solved by an algorithm which may calculate refined pose and plane parameters until convergence is achieved.

[0146] In some embodiments, planar features may be combined with other features, such as point features, in a 3D reconstruction. In a combined cost function derived from multiple features, the Jacobian matrix from the planar costs has the form of J in (28), and the residual vector also has the form of δ in (37). Thus, the reduced Jacobian matrix J in (28) r and the reduced residual vector δ in (37) can be substituted for the Jacobian matrix and the original residual vector in the bundle adjustment with hybrid features.

[0147] The representation of the environment may be calculated using planar bundle adjustment. Figure 9 is a flowchart illustrating a method 900 for providing a representation of the environment, according to some embodiments. Figure 10 is a schematic diagram illustrating a portion 1000 of a representation of the environment calculated based on sensor acquisition information using method 900, according to some embodiments.

[0148] Method 900 may begin with acquiring (act 902) information captured by one or more sensors at a distinct pose. In some embodiments, the acquired information may be a visual image and / or a depth image. In some embodiments, the image may be a keyframe, which may be a combination of multiple images. The example in FIG. 10 shows three sensors: sensor 0, sensor 1, and sensor 2. In some embodiments, the sensors may belong to a device, which may have a device coordinate frame with an origin O. The device coordinate frame may represent the location and orientation when the device first begins scanning the environment for the session. Each sensor may have a corresponding sensor coordinate frame. Each sensor may be associated with a distinct pose ρ i-1 , ρ i , and ρ i+1 In some embodiments, the coordinate frame with origin O may represent a coordinate frame shared by one or more devices. For example, a sensor may belong to three devices that share a coordinate frame.

[0149] Method 900 may include providing a first representation of the environment (act 904). The first representation may include an initial estimate of feature parameters for features extracted from sensor-acquired information and an initial estimate of a corresponding pose. In the example of FIG. 10, a first plane 1004 is observed in image 1002A and image 1002B. The first plane 1004 is defined by plane parameters π j =[n j ;d j The image 1002A and the image 1002B may each be represented by a plurality of points P (i-1)j and P ijA portion of the first plane 1004 may be observed by both images 1002A and 1002B. A portion of the first plane 1004 may be observed by only image 1002A or only image 1002B. A second plane 1006 is observed in images 1002B and 1002C. The second plane 1006 is observed by plane parameter π j+1 =[n (j+1) ;d (j+1) The images 1002B and 1002C may each be represented by a plurality of points P i(j+1) and P (i+1)j A portion of the second plane 1006 may be observed by both images 1002B and 1002C. A portion of the second plane 1006 may be observed by only image 1002B or only image 1002C. Images 1002B and 1002C also observe point feature 1008. The first representation is represented by a pose ρ i-1 , ρ i , and ρ i+1 , plane parameters for the first plane 1004 and the second plane 1006 , and point feature parameters for the point feature 1008 .

[0150] An initial prediction may be based on first observing features in the image. For the example of FIG. 10, the pose ρ i-1 An initial estimate of ρ and plane parameters for the first plane 1004 may be based on the image 1002A. i The initial estimate of ρ, the plane parameters for the second plane 1006, and the feature parameters for the point features 1008 may be based on the image 1002B. i+1 The initial estimate of may be based on image 1002C.

[0151] As features are observed in subsequent images, the initial prediction may be refined to reduce drift and improve the quality of the presentation. Method 900 includes act 906 and act 908, which calculate refined poses and refined feature parameters, which may include refined plane parameters and refined point feature parameters. Act 906 may be performed for each plane at the corresponding pose. In the example of FIG. 10 , act 906 calculates the poses ρ and ρ, respectively. i-1 and ρ i , respectively, with respect to the first plane 1004. Act 906 also i and ρ i+1 Act 906 is performed with respect to the second plane 1006. Act 906 generates a matrix (e.g., C, such as (16)) that indicates the observation of the individual plane (e.g., plane j) at the individual pose (e.g., pose i). ij ) (Act 906A). Act 906 may include computing a matrix (e.g., an upper triangular matrix M ij ) into two or more matrices, with one matrix having reduced rows compared to (act 906B).

[0152] Method 900 may include providing a second representation of the environment with the refined pose and refined feature parameters (act 910). Method 900 may include determining whether new information is observed such that acts 902-910 should be repeated (act 912).

[0153] While several aspects of several embodiments have been described above, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.

[0154] As an example, embodiments are described in relation to an augmented (AR) environment. It should be understood that some or all of the techniques described herein may be applied in MR environments, or more generally, other XR and VR environments, and any other computer vision and robotics applications.

[0155] As another example, embodiments are described in connection with devices such as wearable devices. It should be understood that some or all of the techniques described herein may be implemented via a network (such as the cloud), a discrete application, and / or any suitable combination of devices, networks, and discrete applications.

[0156] Such alterations, modifications, and improvements are intended to be part of this disclosure, and are intended to be within the spirit and scope of this disclosure. Moreover, while advantages of the present disclosure have been described, it should be understood that not all embodiments of the present disclosure include all described advantages. Some embodiments may not implement any feature described herein as advantageous, and in some instances, the foregoing description and drawings are therefore by way of example only.

[0157] The foregoing embodiments of the present disclosure can be implemented in any of numerous ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. Such a processor may be implemented as an integrated circuit, with one or more processors in an integrated circuit component, including commercially available integrated circuit components known in the art, such as a CPU chip, a GPU chip, a microprocessor, a microcontroller, or a coprocessor, to name a few. In some embodiments, the processor may be implemented in a custom circuit, such as an ASIC, or in a semi-custom circuit resulting from configuring a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores, such that one or a subset of those cores may constitute a processor. However, a processor may be implemented using circuitry in any suitable format.

[0158] Further, it should be understood that a computer may be embodied in any of several forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, etc. Additionally, a computer may be embodied in devices not generally considered computers but with suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or fixed electronic device.

[0159] A computer may also have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound-generating device for audible presentation of output. Examples of input devices that can be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, a computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output devices are illustrated as physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated within the same unit as the processor or other elements of the computing device. For example, a keyboard may be implemented as a soft keyboard on a touchscreen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated through a wireless connection.

[0160] Such computers may be interconnected by one or more networks of any suitable form, including local area networks or wide area networks, such as an enterprise network or the Internet. Such networks may be based on any suitable technology and operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0161] The various methods and processes outlined herein may also be coded as software that is executable on one or more processors employing any one of a variety of operating systems or platforms. In addition, such software may be written using any of a number of suitable programming languages ​​and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine.

[0162] In this aspect, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more floppy disks, compact disks (CDs), optical disks, digital video disks (DVDs), magnetic tape, flash memory, circuitry in a field programmable gate array or other semiconductor device, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods implementing various embodiments of the present disclosure discussed above. As is evident from the foregoing examples, a computer-readable storage medium may retain information for a time sufficient to provide computer-executable instructions in a non-transitory form. Such a computer-readable storage medium or media may be transportable, as described above, such that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present disclosure. As used herein, the term "computer-readable storage medium" encompasses only computer-readable media that may be considered a manufacture (i.e., an article of manufacture) or machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.

[0163] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure, as described above. Additionally, according to one aspect of the present embodiments, it should be understood that one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but can be distributed in a modular manner among several different computers or processors to implement various aspects of the present disclosure.

[0164] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0165] Also, data structures may be stored in a computer-readable medium in any suitable form. For ease of illustration, data structures may be shown to have fields that are related through their locations within the data structure. Such relationships may also be achieved by allocating storage for the fields with locations within the computer-readable medium that convey the relationship between the fields. However, any suitable mechanism may be used to establish relationships between information within fields of a data structure, including through the use of pointers, tags, or other mechanisms that establish relationships between data elements.

[0166] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and therefore, its application is not limited to the details and arrangements of components set forth in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.

[0167] The present disclosure may also be embodied as a method, examples of which are provided. The acts performed as part of the method may be ordered in any suitable manner. Thus, while the illustrative embodiments are shown as sequential acts, embodiments may be constructed in which acts are performed in an order different from that shown, which may include performing some acts simultaneously.

[0168] The use of ordinal terms such as "first," "second," "first," etc. in the claims to modify elements does not, by itself, imply any priority, precedence, or ordering of one claim element relative to another element, or the chronological order in which acts of a method are performed, but rather the ordinal terms are used solely as markers to distinguish between claim elements (due to the use of the ordinal terms) and to distinguish between one claim element having a certain name and another element having the same name.

[0169] Also, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use herein of "including," "comprising," "having," "containing," "with," and variations thereof, is meant to encompass the items listed thereafter and equivalents and additional items.

Claims

1. 1. A method of operating a computing system to generate a representation of an environment, the method comprising: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation based at least in part on the first number of images, comprising the first number of initial poses and initial parameters for a second number of planar features, the initial parameters for the second number of planar features indicating normals of planes represented by the second number of planar features; For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating a matrix indicative of one or more observations of the planar features; and decomposing the matrix into two or more matrices, the two or more matrices comprising one matrix having reduced rows compared to the matrix; and calculating refined parameters for the first number of refined poses and the second number of planar features by adjusting both the first number of initial poses and initial parameters for the second number of planar features based at least in part on the matrix having reduced rows so that a projection error of observations of the second number of planar features from the respective refined poses onto the respective planar features is minimized; Including, The method, wherein the representation of the environment comprises refined parameters of the first number of refined poses and the second number of planar features.

2. Decomposing the matrix into two or more matrices for each of the second number of planar features at each pose corresponding to images comprising one or more views of the planar features comprises: The method of claim 1 , comprising calculating an orthogonal matrix and an upper triangular matrix.

3. The method of claim 2 , wherein calculating the first number of refined poses and the second number of refined parameters of planar features is based at least in part on the upper triangular matrix.

4. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: The method of claim 1 , comprising calculating a reduced Jacobian matrix block based at least in part on the matrix having reduced rows.

5. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: The method of claim 4 , comprising stacking the reduced Jacobian matrix blocks to form a reduced Jacobian matrix.

6. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes:

6. The method of claim 5, comprising providing the reduced Jacobian matrix to an algorithm that solves a least-squares problem to update a current estimate.

7. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: A method according to any preceding claim, comprising calculating a reduced residual block based at least in part on said matrix having reduced rows.

8. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes: The method of claim 7 , comprising stacking the reduced residual blocks to form a reduced residual vector.

9. Calculating refined parameters for the first number of refined poses and the second number of planar features based at least in part on the matrix having reduced rows includes:

9. The method of claim 8, comprising providing the reduced residual vector to an algorithm that solves a least squares problem to update a current estimate.

10. For each of the second number of planar features at each pose corresponding to an image comprising one or more observations of the planar feature, calculating the matrix indicative of one or more observations of the planar feature comprises: calculating, for each of one or more views of the planar feature, a matrix block indicative of the view; stacking the matrix blocks into the matrix representing one or more observations of the planar feature; and 8. The method of claim 1 or claim 7, comprising:

11. 1. A method of operating a computing system to generate a representation of an environment, the method comprising: acquiring sensor acquisition information, the sensor acquisition information comprising a first number of images; providing an initial representation of the environment, the initial representation based at least in part on the first number of images, comprising the first number of initial poses and initial parameters for a second number of planar features, the initial parameters for the second number of planar features indicating normals of planes represented by the second number of planar features; For each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature, calculating a matrix having a third number of rows, the third number being less than a number of one or more observations of the planar feature; calculating refined parameters for the first number of refined poses and the second number of planar features by adjusting both the first number of initial poses and initial parameters for the second number of planar features based at least in part on the matrix having the third number of rows so that a projection error of observations of the second number of planar features from the respective refined poses onto the respective planar features is minimized; Including, The method, wherein the representation of the environment comprises refined parameters of the first number of refined poses and the second number of planar features.

12. calculating the matrix having the third number of rows for each of the second number of planar features at each pose corresponding to an image comprising one or more views of the planar feature; calculating a matrix indicative of one or more observations of the planar features; decomposing the matrix into two or more matrices, the two or more matrices comprising the matrix having the third number of rows; The method of claim 11 , comprising:

13. The method of claim 12 , wherein decomposing the matrix into two or more matrices includes calculating an orthogonal matrix and an upper triangular matrix.

14. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: A method according to any of claims 11 to 13, comprising calculating a reduced Jacobian matrix block based at least in part on the matrix having the third number of rows.

15. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: The method of claim 14 , comprising stacking the reduced Jacobian matrix blocks to form a reduced Jacobian matrix.

16. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes:

16. The method of claim 15, comprising providing the reduced Jacobian matrix to an algorithm that solves a least-squares problem to update a current estimate.

17. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes:

15. The method of claim 11 or claim 14, comprising calculating a reduced residual block based at least in part on the matrix having the third number of rows.

18. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes: The method of claim 17 , comprising stacking the reduced residual blocks to form a reduced residual vector.

19. Calculating the first number of refined poses and refined parameters of the second number of planar features based at least in part on the matrix having the third number of rows includes:

20. The method of claim 18, comprising providing the reduced residual vector to an algorithm that solves a least squares problem to update a current estimate.

20. For each of the second number of planar features at each pose corresponding to an image comprising one or more observations of the planar feature, calculating the matrix indicative of one or more observations of the planar feature comprises: calculating, for each of one or more views of the planar feature, a matrix block indicative of the view; stacking the matrix blocks into the matrix representing one or more observations of the planar feature; and 15. The method of claim 12 or claim 14, comprising:

Citation Information

Patent Citations

  • System, method and program for filtering information

    JP2002269143A

  • Information processing method and device

    JP2007048068A

  • Sensor calibration method and sensor calibration device

    JP2017102942A

  • Method and system for evaluating matching between content item and image based on similarity scores

    JP2017220203A

  • Method and apparatus for obtaining photogrammetric data to estimate impact severity

    US20070288135A1