LiDAR simultaneous localization and mapping

The planar LiDAR SLAM framework addresses indoor mapping challenges by distinguishing planar object sides and using PBA to enhance accuracy and efficiency in LiDAR SLAM systems.

JP7851941B2Active Publication Date: 2026-04-27MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
MAGIC LEAP INC
Filing Date
2022-02-11
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional LiDAR SLAM systems face challenges in indoor environments due to ambiguous plane detection, leading to errors in odometry and mapping, particularly the 'dual-side' problem of planar surfaces, which affects accuracy and computational efficiency.

Method used

A planar LiDAR SLAM framework that distinguishes the two sides of a planar object using the direction of the vector normal and employs Plane Bundle Adjustment (PBA) to reduce attitude errors and drift, incorporating a point-plane cost function to balance accuracy and computational cost.

Benefits of technology

The framework improves the accuracy and reduces computational overhead in indoor LiDAR SLAM by correcting distortions and aligning scans in real-time, enhancing the precision of 3D mapping and localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007851941000184
    Figure 0007851941000184
  • Figure 0007851941000185
    Figure 0007851941000185
  • Figure 0007851941000186
    Figure 0007851941000186
Patent Text Reader

Abstract

Disclosed herein are systems and methods for mapping environmental information. In some embodiments, the systems and methods are configured for mapping information in a mixed reality environment. In some embodiments, the systems are configured to perform a method including scanning an environment, including capturing a plurality of points of the environment with a sensor, tracking a plane of the environment, updating an observation associated with the environment by inserting a key frame in the observation, determining whether the plane is coplanar with a second plane of the environment, performing a plane bundle adjustment on the observation associated with the environment according to a determination that the plane is coplanar with the second plane, and performing a plane bundle adjustment on a portion of the observation associated with the environment according to a determination that the plane is not coplanar with the second plane.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Cross-reference of related applications) This application claims priority to U.S. Provisional Application No. 63 / 149,037, filed on 12 February 2021, the contents of which are both incorporated herein by reference as a whole.

[0002] This disclosure generally relates to systems and methods for mapping environments. In some embodiments, light detection and ranging (LiDAR) sensors are used to map the environment using simultaneous localization and mapping (SLAM). [Background technology]

[0003] Virtual environments are ubiquitous in computing environments, finding use in video games (where the virtual environment can represent the game world), maps (where the virtual environment can represent the terrain to be navigated), simulations (where the virtual environment can simulate the real environment), digital storytelling (where virtual characters can interact with each other within the virtual environment), and many other applications. Modern computer users generally perceive and interact with virtual environments comfortably. However, the user experience with virtual environments can be limited by the technology used to present them. For example, conventional displays (e.g., 2D display screens) and audio systems (e.g., fixed speakers) may be incapable of realizing virtual environments in a way that creates an engaging, realistic, and immersive experience.

[0004] Virtual reality ("VR"), augmented reality ("AR"), mixed reality ("MR"), and related technologies (collectively, "XR") share the ability to present users of XR systems with sensory information corresponding to a virtual environment represented by data within a computer system. This disclosure assumes specificity between VR, AR, and MR systems (however, some systems may be categorized as VR in one aspect (e.g., a visual aspect) and simultaneously as AR or MR in another aspect (e.g., an audio aspect)). As used herein, a VR system presents a virtual environment that replaces the user's real environment in at least one aspect; for example, a VR system may present the user with a view of the virtual environment while simultaneously obscuring that view of the real environment, such as by using a light-blocking head-mounted display. Similarly, a VR system may present the user with audio corresponding to the virtual environment while simultaneously blocking (attenuating) audio from the real environment.

[0005] VR systems can suffer from various drawbacks stemming from their replacement of the user's real environment with a virtual one. One drawback is motion sickness, which can occur when the user's field of vision in the virtual environment no longer corresponds to the state of their inner ear that detects their balance and orientation in the real environment (rather than the virtual one). Similarly, users can suffer disorientation in a VR environment if their body and limbs (on which they rely to feel "grounded" in the real environment) are not directly visible. Another drawback is the computational burden (e.g., memory, processing power) placed on VR systems, especially in real-time applications where the goal is to immerse the user in the virtual environment, as they must present a fully 3D virtual environment. Likewise, such environments may need to reach a very high level of reality in order to be considered immersed, as users tend to be sensitive to even the slightest imperfections in the virtual environment, any of which can disrupt the user's sense of immersion. Furthermore, another drawback of VR systems is that such applications of the system cannot utilize the wide range of sensory data found in real environments, such as various sights and sounds, that are experienced in the real world. A related drawback is that VR systems may struggle to create shared environments in which multiple users can interact, as it may be impossible for users sharing the same physical space in the real environment to directly see or interact with each other within the virtual environment.

[0006] As used herein, an AR system presents a virtual environment that overlaps or overlays the real environment in at least one aspect. For example, an AR system may present a view of the virtual environment overlaid on the user's view of the real environment, for instance, by using a transparent head-mounted display that presents images while allowing light to pass through the display into the user's eyes. Similarly, an AR system may present audio corresponding to the virtual environment while simultaneously mixing it with audio from the real environment. Likewise, as used herein, an MR system, like an AR system, presents a virtual environment that overlaps or overlays the real environment in at least one aspect, and in addition, may allow the virtual environment within the MR system to interact with the real environment in at least one aspect. For example, a virtual character in the virtual environment may toggle a light switch in the real environment, turning a corresponding light bulb in the real environment on or off. In another embodiment, the virtual character may react to audio signals in the real environment (e.g., using facial expressions). By maintaining the presentation of the real environment, AR and MR systems can avoid some of the aforementioned shortcomings of VR systems. For example, motion sickness in users is reduced because visual cues from the real environment (including the user's own body) can remain visible, and such systems, being immersive, do not need to present the user with a fully realized 3D environment. Furthermore, AR and MR systems can create new applications by leveraging and extending real-world sensory input (e.g., scenery, objects, and the views and sounds of other users).

[0007] Presenting a virtual environment that overlaps with or overlays the real environment can be challenging. For example, blending virtual and real environments may require a complex and thorough understanding of the real environment to ensure that objects in the virtual environment do not conflict with objects in the real environment. Furthermore, maintaining persistence in the virtual environment that corresponds to consistency in the real environment may be desirable. For instance, it may be desirable for a virtual object displayed on a physical table to appear in the same location even if the user looks away, moves around, and then returns their gaze to the physical table. To achieve this type of immersion, it may be beneficial to develop accurate and precise estimates of the locations of objects in the real world and the locations of the user in the real world.

[0008] For example, LiDAR sensors can be used to perceive the environment (e.g., in robotic applications). SLAM techniques using LiDAR (or similar sensors) can be used for indoor applications such as autonomous security, cleaning, and delivery robots, including augmented reality (AR) applications. Compared to SLAM using only sensors such as visual cameras, LiDAR SLAM can result in higher-density 3D maps, which can be important for computer vision tasks such as AR, and for 3D object classification and semantic segmentation. In addition, high-density 3D mapping can also have applications in the construction industry. 3D maps of buildings can be used for decorative design, allowing architects to monitor whether the building meets design requirements during the construction process. However, implementing indoor SLAM using sensors such as LiDAR sensors can be challenging.

[0009] Compared to outdoor environments, indoor environments can be more structured. For example, indoor environments are more likely to feature predictable geometric shapes (e.g., planar surfaces, right angles) and consistent lighting. While these characteristics can simplify tasks such as object recognition, more structured indoor environments present their own set of challenges. For instance, GPS signals may not be available in indoor environments, which can make loop closure detection for LiDAR (or similar sensor) SLAM difficult. In addition, when using sensors to detect planar surfaces such as walls or doors in indoor environments, it can be ambiguous which of the two sides of the surface is facing the sensor. This is referred to herein as the “dual-side” problem. For example, in an outdoor environment, LiDAR generally observes the outside of a building. However, in an indoor environment, LiDAR may observe two sides of a planar object. Because two sides of a planar object can typically be close to each other, this can cause erroneous data associations in LiDAR odometry and mapping (LOAM) systems, which are based on a variation of the iterative nearest point (ICP) technique, and this can lead to large errors. In addition, LOAM and its variations can provide low-fidelity odometry poses estimated by aligning the current scan to the last one in real time, and aligning the global pose can be much slower. It is desirable to mitigate these effects in order to obtain faster, less error-prone observations across various sensor environments.

[0010] Planes can be used in conventional LiDAR SLAM methods and generally in variations of the ICP framework. However, in some conventional methods, planes may not be explicitly extracted. For example, planes can be used in the interplane ICP framework, which is a variation of conventional ICP. However, in some embodiments, plane parameters may not be explicitly extracted and are generally estimated from small patches of the point cloud obtained, for example, from K-nearest neighbor search. These methods can lead to increasing error cycles. For example, attitude errors can accumulate in the point cloud, which can reduce the accuracy of attitude estimation, and inaccurate attitudes can further degrade the global point cloud. Bundle adjustment (BA) can provide accurate solutions, but performing BA can be time-consuming and computationally intensive.

[0011] In addition, the cost function (e.g., a function to minimize algebraic errors or geometric or statistical distances associated with the estimation of geometric shape from sensor measurements) can affect the accuracy of the solution. It is desirable that the cost function for geometric problems be invariant with respect to rigid body transformations. With respect to plane correspondences, there may be two ways to construct the cost function: inter-plane cost and point-plane cost. The point-plane cost can be expressed as the square of the distance from a point to a plane, which is invariant with respect to rigid body transformations. The inter-plane cost is a measure of the difference between two plane parameters, which may not be invariant with respect to rigid body transformations. Therefore, the point-plane cost may yield more accurate results compared to the inter-plane cost. However, the point-plane cost may result in a larger least-squares problem than the inter-plane cost, and may not be able to reach a more accurate solution, and may be applicable to smaller problems (e.g., single-scan distortion correction and alignment, scans with fewer points than those of LiDAR or similar sensors). [Overview of the Initiative] [Means for solving the problem]

[0012] Disclosed herein are systems and methods for mapping environmental information. In some embodiments, the systems and methods are configured to map information in a mixed reality environment. In some embodiments, the system is configured to implement a method comprising: scanning an environment, including capturing multiple points of the environment using sensors; tracking a plane of the environment; updating observations associated with the environment by inserting keyframes into the observations; determining whether a plane is coplanar with a second plane of the environment; performing plane bundle adjustments on observations associated with the environment in accordance with the determination that the plane is coplanar with the second plane; and performing plane bundle adjustments on a portion of observations associated with the environment in accordance with the determination that the plane is not coplanar with the second plane.

[0013] In some embodiments, a planar LiDAR SLAM framework for indoor environments is disclosed. In some embodiments, two sides of a planar object are distinguished by the direction of the vector normal toward the observation center (e.g., of a LiDAR sensor), and the undesirable effects of the double-sided problem can be reduced.

[0014] In some embodiments, the system and method include a step of performing a Plane Bundle Adjustment (PBA), which includes a step of jointly optimizing plane parameters and sensor attitudes. Performing a PBA has the advantage of reducing attitude errors and drift associated with planes that are not explicitly extracted, as described above.

[0015] In some embodiments, a LiDAR SLAM system and a method for operating the system are disclosed. The system and method are advantageous in that they can balance the accuracy and computational cost of the base orientation (BA). In some embodiments, the system comprises three parallel components, including localization, local mapping, and PBA. For example, the localization component tracks a plane frame by frame, corrects the distortion of the LiDAR point cloud, and aligns the new scan to a global plane model in real time. In addition, the distortion correction and alignment tasks can be accelerated by using an approximate rotation matrix for small motions between two scans. The local mapping and PBA components can correct for drift, improve the map, and ensure that the localization component can provide accurate orientations. In some embodiments, the point-plane cost includes a special structure that can be utilized to reduce the computational cost related to the PBA. Based on this structure, an integrated point-plane cost can be formed to accelerate local mapping. This specification also provides, for example, the following: (Item 1) It is a method, Scanning the environment, wherein the scanning includes detecting multiple points in the environment using sensors. Identifying a first plane of the environment, wherein the first plane comprises at least a threshold number of points of the plurality of detected points, Obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observations, and not based on a second subset of the observations. Methods that include... (Item 2) The method according to item 1, wherein the sensor comprises a LiDAR sensor. (Item 3) The method according to item 1, wherein the normal of the first plane comprises a vector in the direction toward the observation center associated with the sensor. (Item 4) The method according to item 1, further comprising performing planar bundle adjustments on a first subset of the aforementioned observations, and calculating a cost function associated with keyframe poses in a sliding window associated with the first subset. (Item 5) Determining the distance between the scanned portion of the environment and the first keyframe, A second keyframe is inserted according to the determination that the distance between the scanned portion of the environment and the first keyframe exceeds a threshold distance. The decision to refrain from inserting a second keyframe is made based on the determination that the distance between the scanned portion of the environment and the first keyframe does not exceed the threshold distance. The method described in item 1, further including the method described in item 1. (Item 6) Determining the percentage of points being tracked, wherein the first plane includes the points being tracked. A keyframe is inserted based on the determination that the percentage of the tracked points falls below a threshold percentage. The decision not to insert the keyframes in accordance with the determination that the percentage of the tracked points does not fall below the threshold percentage. The method described in item 1, further including the method described in item 1. (Item 7) The method of item 1, further comprising determining an integral cost matrix and performing a planar bundle adjustment, which comprises calculating a cost function using the integral cost matrix. (Item 8) The method described in item 1, further including estimating keyframe poses. (Item 9) The method according to item 1, further comprising linearly interpolating the points of the keyframe pose. (Item 10) The method according to item 1, further comprising updating the map of the environment based on the planar bundle adjustment. (Item 11) The environment is the method described in item 1, comprising a mixed reality environment. (Item 12) The method according to item 1, wherein the sensor comprises a sensor for a mixed reality device. (Item 13) In accordance with the determination that the first plane is not coplanar with the second plane, the first plane is inserted into the map associated with the environment, In accordance with the determination that the first plane is coplanar with the second plane, refrain from inserting the first plane into the map associated with the environment. The method described in item 1, further including the method described in item 1. (Item 14) The method according to item 1, wherein determining whether the first plane is coplanar with the second plane further includes determining whether a threshold percentage of the point-plane distance is less than a threshold distance. (Item 15) The method according to item 1, further comprising performing a geometric consistency check to determine whether the first plane is coplanar with the second plane. (Item 16) The method according to item 1, further comprising storing the plurality of observation results associated with the environment in memory. (Item 17) The method according to item 1, further comprising transmitting information about the environment, wherein the information about the environment is associated with the plurality of observations associated with the environment. (Item 18) It is a system, Sensors and, One or more processors, wherein the one or more processors in the system Scanning the environment, wherein the scanning includes detecting multiple points in the environment using the sensor, Identifying a first plane of the environment, wherein the first plane comprises at least a threshold number of points of the plurality of detected points, Obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observations, and not based on a second subset of the observations. One or more processors configured to perform a method including A system that includes these features. (Item 19) The system described in item 18 comprises a LiDAR sensor. (Item 20) A non-transient computer-readable storage medium for storing one or more programs, wherein the one or more programs comprises instructions, and when the instructions are executed by an electronic device having one or more processors and memory, the device... Scanning the environment, wherein the scanning includes detecting multiple points in the environment using sensors. Identifying a first plane of the environment, wherein the first plane comprises at least a threshold number of points of the plurality of detected points, Obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observations, and not based on a second subset of the observations. A non-transient computer-readable storage medium that enables the implementation of a method including the following. [Brief explanation of the drawing]

[0016] [Figure 1A] Figure 1A-1C illustrates an exemplary mixed reality environment according to an embodiment of the present disclosure. [Figure 1B] Figure 1A-1C illustrates an exemplary mixed reality environment according to an embodiment of the present disclosure. [Figure 1C] Figure 1A-1C illustrates an exemplary mixed reality environment according to an embodiment of the present disclosure.

[0017] [Figure 2A] Figures 2A-2D illustrate components of an exemplary mixed reality system according to an embodiment of the present disclosure. [Figure 2B] Figures 2A-2D illustrate components of an exemplary mixed reality system according to an embodiment of the present disclosure. [Figure 2C] Figures 2A-2D illustrate components of an exemplary mixed reality system according to an embodiment of the present disclosure. [Figure 2D] Figures 2A-2D illustrate components of an exemplary mixed reality system according to an embodiment of the present disclosure.

[0018] [Figure 3A]Figure 3A illustrates an exemplary mixed reality handheld controller according to an embodiment of the present disclosure.

[0019] [Figure 3B] Figure 3B illustrates an exemplary auxiliary unit according to an embodiment of the present disclosure.

[0020] [Figure 4] Figure 4 illustrates an exemplary functional block diagram relating to an exemplary mixed reality system according to an embodiment of the present disclosure.

[0021] [Figure 5] Figure 5 illustrates an exemplary SLAM system according to an embodiment of the present disclosure.

[0022] [Figure 6] Figure 6 illustrates an exemplary timing diagram of a SLAM system according to an embodiment of the present disclosure.

[0023] [Figure 7] Figure 7 illustrates an exemplary factor graph according to an embodiment of the present disclosure.

[0024] [Figure 8] Figure 8 illustrates an exemplary schematic diagram of updating the integral cost matrix according to an embodiment of the present disclosure.

[0025] [Figure 9] Figure 9 illustrates an exemplary factor graph according to an embodiment of the present disclosure. [Modes for carrying out the invention]

[0026] Detailed explanation In the following description of the embodiments, accompanying drawings, which form part of this specification and illustrate specific embodiments that can be put into practice, are referenced. It should be understood that other embodiments may also be used, and structural modifications may be made without departing from the scope of the disclosed embodiments.

[0027] Like all people, users of a mixed reality system perceive the three-dimensional parts of the real environment, i.e., the "real world," and all of its contents. For example, users perceive the real environment using their normal human senses, namely sight, hearing, touch, taste, and smell, and interact with the real environment by moving their bodies within it. Locations within the real environment can be described as coordinates in coordinate space, for example, coordinates can include latitude, longitude, and altitude relative to sea level, distance in three orthogonal dimensions from a reference point, or other preferred values. Similarly, vectors can describe quantities that have direction and magnitude in coordinate space.

[0028] A computing device can maintain a representation of a virtual environment, for example, in memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. A virtual environment can include representations of any object, actions, signals, parameters, coordinates, vectors, or other properties associated with that space. In some embodiments, the computing device's network (e.g., a processor) can maintain and update the state of the virtual environment; that is, the processor can determine the state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or inputs provided by the user at a first time t0. For example, if an object in the virtual environment is located at a first coordinate at time t0, has some programmed physical parameters (e.g., mass, coefficient of friction), and inputs received from the user indicate that a force should be applied to the object in a certain direction vector, the processor can determine the object's location at time t1 by applying the laws of kinematics and using basic mechanics. The processor can determine the state of the virtual environment at time t1 using any known suitable information about the virtual environment and / or any suitable inputs. When maintaining and updating the state of a virtual environment, the processor may run any suitable software, including software related to creating and deleting virtual objects within the virtual environment, software for defining the behavior of virtual objects or characters within the virtual environment (e.g., scripts), software for defining the behavior of signals within the virtual environment (e.g., audio signals), software for creating and updating parameters associated with the virtual environment, software for generating audio signals within the virtual environment, software for handling inputs and outputs, software for implementing network operations, software for applying asset data (e.g., animation data for moving virtual objects over time), or many other possibilities.

[0029] Output devices such as displays or speakers can present any or all aspects of the virtual environment to the user. For example, the virtual environment may include virtual objects (which may include representations of inanimate objects, people, animals, light, etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" with origin coordinates, visual axes, and frustum) and render a viewable scene of the virtual environment corresponding to that view on the display. Any suitable rendering technique may be used for this purpose. In some embodiments, the viewable scene may include only some virtual objects in the virtual environment and exclude some other virtual objects. Similarly, the virtual environment may include an audio aspect that can be presented to the user as one or more audio signals. For example, virtual objects in the virtual environment may generate sounds resulting from the object's location coordinates (e.g., a virtual character may speak or produce sound effects), or the virtual environment may be associated with musical cues or ambient sounds that may or may not be associated with a specific location. The processor can determine an audio signal that corresponds to the "listener" coordinates, for example, the synthesis of sound in a virtual environment, which is mixed and processed to simulate the audio signal that would be heard by the listener at the listener coordinates, and present the audio signal to the user through one or more speakers.

[0030] Because a virtual environment exists only as a computational structure, users cannot directly perceive it using their normal senses. Instead, users can only perceive the virtual environment indirectly, such as through displays, speakers, or haptic output devices. Similarly, users cannot directly touch, manipulate, or otherwise interact with the virtual environment, but they can provide input data via input devices or sensors to a processor that can use device or sensor data to update the virtual environment. For example, a camera sensor may provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use this data to make the object respond accordingly within the virtual environment.

[0031] A mixed reality system can present a user with a mixed reality environment ("MRE") that combines aspects of the real and virtual environments, for example, using a transparent display and / or one or more speakers (which may be incorporated in a wearable head device, for example). In some embodiments, one or more speakers may be located outside the head-mounted wearable unit. As used herein, the MRE is a simultaneous representation of the real environment and the corresponding virtual environment. In some embodiments, the corresponding real and virtual environments share a single coordinate space, and in some embodiments, the real coordinate space and the corresponding virtual coordinate space are related by a transformation matrix (or other preferred representation). Thus, a single coordinate (in some embodiments, together with the transformation matrix) can define a first location in the real environment and a second corresponding location in the virtual environment, and vice versa.

[0032] In MRE, virtual objects (for example, in a virtual environment associated with MRE) can correspond to real objects (for example, in a real environment associated with MRE). For example, if the real environment of MRE has a real lamppost (real object) at a certain location coordinate, the virtual environment of MRE may have a virtual lamppost (virtual object) at the corresponding location coordinate. As used herein, real objects, together with their corresponding virtual objects, constitute a “mixed reality object.” It is not necessary for a virtual object to perfectly match or be consistent with its corresponding real object. In some embodiments, a virtual object may be a simplified version of its corresponding real object. For example, if the real environment includes a real lamppost, the corresponding virtual object may be a cylinder with approximately the same height and radius as the real lamppost (reflecting that the lamppost may be roughly cylindrical in shape). Simplifying the virtual object in this way can enable computational efficiency and simplify the calculations that would otherwise be performed on such a virtual object. Furthermore, in some embodiments of MRE, not all real objects in the real environment have to be associated with corresponding virtual objects. Similarly, in some embodiments of MRE, not all virtual objects within the virtual environment may be associated with corresponding real-world objects. That is, some virtual objects may exist only within the MRE virtual environment without any real-world counterparts.

[0033] In some embodiments, virtual objects may have characteristics that are sometimes significantly different from those of their corresponding real objects. For example, a real environment within an MRE might comprise a cactus with two green branches, i.e., a thorny, inanimate object, while a corresponding virtual object within an MRE might have the characteristics of a virtual character with two green arms, accompanied by human facial features and an unfriendly demeanor. In this embodiment, the virtual object is similar to its corresponding real object in some characteristics (color, number of arms) but differs from the real object in other characteristics (facial features, personality). Thus, virtual objects have the potential to represent real objects in a creative, abstract, exaggerated, or fictional way, or to confer behavior (e.g., human personality) to otherwise inanimate real objects. In some embodiments, virtual objects may be purely fictional creations with no corresponding real-world counterparts (e.g., a virtual monster in a virtual environment, perhaps in a place corresponding to emptiness in a real environment).

[0034] Compared to VR systems, which present a virtual environment while obscuring the real environment, mixed reality systems that present MRE offer the advantage that the real environment remains perceptible while the virtual environment is presented. Therefore, users of mixed reality systems can experience and interact with the corresponding virtual environment using visual and audio cues associated with the real environment. For example, users of VR systems may struggle to perceive or interact with virtual objects displayed within the virtual environment because they cannot directly perceive or interact with the virtual environment, as described herein. However, users of MR systems may find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching the corresponding real objects in their own real environment. This level of interaction can enhance the user's sense of immersion, connection, and engagement with the virtual environment. Similarly, by presenting the real and virtual environments simultaneously, mixed reality systems can reduce the negative psychological sensations (e.g., cognitive dissonance) and negative physical sensations (e.g., motion sickness) associated with VR systems. Mixed reality systems also offer many possibilities for applications that can extend or modify our real-world experiences.

[0035] Figure 1A illustrates an exemplary real environment 100 in which a user 110 uses a mixed reality system 112. The mixed reality system 112 may comprise a display (e.g., a transmissive display) and one or more speakers, and one or more sensors (e.g., a camera, a LiDAR sensor, a sensor configured to capture multiple points during a single scan) as described herein. The shown real environment 100 comprises a rectangular room 104A in which the user 110 stands, and real objects 122A (a lamp), 124A (a table), 126A (a sofa), and 128A (a painting). Room 104A further comprises a location coordinate 106, which may be considered the origin of the real environment 100. As shown in Figure 1A, an environment / world coordinate system 108 (with x-axis 108X, y-axis 108Y, and z-axis 108Z), with its origin at point 106 (world coordinates), can define a coordinate space for the real environment 100. In some embodiments, the origin 106 of the environment / world coordinate system 108 may correspond to the location where the mixed reality system 112 is powered on. In some embodiments, the origin 106 of the environment / world coordinate system 108 may be reset during operation. In some embodiments, the user 110 may be considered a real object in the real environment 100, and similarly, parts of the user 110's body (e.g., hands, feet) may be considered real objects in the real environment 100. In some embodiments, a user / listener / head coordinate system 114 (with x-axis 114X, y-axis 114Y, and z-axis 114Z), with its origin at point 115 (e.g., user / listener / head coordinates), can define a coordinate space for the user / listener / head on which the mixed reality system 112 is located. The origin 115 of the user / listener / head coordinate system 114 may be defined relative to one or more components of the mixed reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 may be defined relative to the display of the mixed reality system 112, such as during the initial calibration of the mixed reality system 112.A matrix (which may include translation matrices and quaternion matrices or other rotation matrices) or other preferred representation can characterize the transformation between the user / listener / head coordinate system 114 space and the environment / world coordinate system 108 space. In some embodiments, the left ear coordinate 116 and the right ear coordinate 117 may be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which may include translation matrices and quaternion matrices or other rotation matrices) or other preferred representation can characterize the transformation between the left ear coordinate 116 and the right ear coordinate 117 and the user / listener / head coordinate system 114 space. The user / listener / head coordinate system 114 can simplify the representation of location relative to the environment / world coordinate system 108 for the user's head or head-mounted device. The transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real time using SLAM, visual odometry, or other techniques.

[0036] Figure 1B illustrates an exemplary virtual environment 130 corresponding to the real environment 100. The shown virtual environment 130 comprises a virtual rectangular room 104B corresponding to the real rectangular room 104A, a virtual object 122B corresponding to the real object 122A, a virtual object 124B corresponding to the real object 124A, and a virtual object 126B corresponding to the real object 126A. The metadata associated with the virtual objects 122B, 124B, and 126B may include information derived from the corresponding real objects 122A, 124A, and 126A. In addition, the virtual environment 130 comprises a virtual monster 132, which does not correspond to any real object in the real environment 100. The real object 128A in the real environment 100 does not correspond to any virtual object in the virtual environment 130. A persistent coordinate system 133 (with x-axis 133X, y-axis 133Y, and z-axis 133Z), with its origin at point 134 (persistent coordinate), can define a coordinate space for virtual content. The origin 134 of the persistent coordinate system 133 may be defined relative to / with respect to one or more real objects such as real object 126A. Matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other preferred representations can characterize transformations between the persistent coordinate system 133 space and the environment / world coordinate system 108 space. In some embodiments, virtual objects 122B, 124B, 126B, and 132 may each have their own persistent coordinate point relative to the origin 134 of the persistent coordinate system 133. In some embodiments, there may be multiple persistent coordinate systems, and each of the virtual objects 122B, 124B, 126B, and 132 may have its own persistent coordinate points in one or more persistent coordinate systems.

[0037] With respect to Figures 1A and 1B, the environment / world coordinate system 108 defines a shared coordinate space for both the real environment 100 and the virtual environment 130. In the examples shown, the coordinate space has its origin at point 106. Furthermore, the coordinate space is defined by three identical orthogonal axes (108X, 108Y, 108Z). Thus, a first location in the real environment 100 and a second corresponding location in the virtual environment 130 can be described with respect to the same coordinate space. This simplifies the step of identifying and displaying corresponding locations in the real and virtual environments, since the same coordinates can be used to identify both locations. However, in some examples, the corresponding real and virtual environments do not need to use a shared coordinate space. For example, in some examples (not shown), matrices (which may include translation matrices and quaternion matrices or other rotation matrices) or other preferred representations can characterize the transformation between the real environment coordinate space and the virtual environment coordinate space.

[0038] Figure 1C illustrates an exemplary MRE 150 that simultaneously presents aspects of the real environment 100 and the virtual environment 130 to the user 110 via a mixed reality system 112. In the shown embodiment, the MRE 150 simultaneously presents to the user 110 real objects 122A, 124A, 126A, and 128A from the real environment 100 (e.g., via the transparent portion of the display of the mixed reality system 112) and virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 (e.g., via the active display portion of the display of the mixed reality system 112). As described herein, the origin 106 may act as the origin for the coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis for the coordinate space.

[0039] In the embodiments shown, a mixed reality object comprises corresponding pairs of real and virtual objects (i.e., 122A / 122B, 124A / 124B, 126A / 126B) occupying corresponding locations in coordinate space 108. In some embodiments, both the real and virtual objects may be visible to the user 110 simultaneously. This may be desirable, for example, in cases where information is presented in which the virtual object is designed to extend the view of the corresponding real object (e.g., in a museum setting where the virtual object presents a missing part of an ancient, damaged statue). In some embodiments, the virtual objects (122B, 124B, and / or 126B) may be displayed in such a way that they occlude the corresponding real objects (122A, 124A, and / or 126A) (e.g., via active pixelated occlusion, using a pixelated occlusion shutter). This can be desirable, for example, in cases where a virtual object acts as a visual replacement for a corresponding real object (such as in interactive storytelling applications where an inanimate real object becomes a "living" character).

[0040] In some embodiments, real objects (e.g., 122A, 124A, 126A) may be associated with virtual content or helper data, which may not necessarily constitute a virtual object. The virtual content or helper data can facilitate the processing or handling of virtual objects within a mixed reality environment. For example, such virtual content may include a two-dimensional representation of a corresponding real object, a custom asset type associated with the corresponding real object, or statistical data associated with the corresponding real object. This information can enable or facilitate calculations involving real objects without incurring unnecessary computational overhead.

[0041] In some embodiments, the presentations described herein may also incorporate an audio aspect. For example, in MRE150, the virtual monster 132 may be associated with one or more audio signals, such as footsteps, generated as the monster walks around MRE150. As described herein, the processor of the mixed reality system 112 can compute an audio signal corresponding to the mixture and processed synthesis of all such sounds within MRE150 and present the audio signal to the user 110 via one or more speakers contained within the mixed reality system 112 and / or one or more external speakers.

[0042] An exemplary mixed reality system 112 may include a wearable head device (e.g., a wearable augmented reality or mixed reality head device) comprising: a display (which may comprise left and right transmissive displays, which may be eyepiece displays, and associated components for coupling light from the displays to the user's eyes); left and right speakers (e.g., positioned adjacent to the user's left and right ears, respectively); an inertial measurement unit (IMU) (e.g., mounted on the temple arms of the head device); an orthogonal coil electromagnetic receiver (e.g., mounted on the left temple component); left and right cameras (e.g., depth (time-of-flight) cameras) oriented away from the user; and left and right eye cameras oriented towards the user (e.g., for detecting the user's eye movements). However, the mixed reality system 112 may incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LiDAR, EOG, GPS, magnetic, LiDAR sensors, sensors configured to scan multiple points during a single scan). In addition, the mixed reality system 112 may incorporate networking features (e.g., Wi-Fi capability) to communicate with other devices and systems, including other mixed reality systems. The mixed reality system 112 may further include a battery (which may be housed in an auxiliary unit such as a belt pack designed to be worn around the user's waist), a processor, and memory. The wearable head device of the mixed reality system 112 may include a tracking component such as an IMU or other suitable sensor configured to output a set of coordinates of the wearable head device relative to the user's environment. In some embodiments, the tracking component may provide input to the processor and perform SLAM, visual odometry, and / or LiDAR odometry algorithms. In some embodiments, the mixed reality system 112 may also include a handheld controller 300 and / or an auxiliary unit 320, which may be a wearable belt pack, as further described herein.

[0043] Figures 2A-2D illustrate components of an exemplary mixed reality system 200 (which may correspond to mixed reality system 112) that may be used to present an MRE (which may correspond to MRE150) or other virtual environment to a user. Figure 2A shows a perspective view of a wearable head device 2102 included in the exemplary mixed reality system 200. Figure 2B shows a top view of the wearable head device 2102 worn on the user's head 2202. Figure 2C shows a front view of the wearable head device 2102. Figure 2D shows an edge view of an exemplary eyepiece 2110 of the wearable head device 2102. As shown in Figures 2A-2C, the exemplary wearable head device 2102 includes an exemplary left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an exemplary right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 may include a transmissive element through which the real environment can be seen, and a display element for presenting a display that overlaps with the real environment (e.g., via image-modulated light). In some embodiments, such a display element may include a surface diffractive optical element for controlling the flow of image-modulated light. For example, the left eyepiece 2108 may include a left internally coupled grating set 2112, a left orthogonal pupil dilation (OPE) grating set 2120, and a left exit (output) pupil dilation (EPE) grating set 2122. Similarly, the right eyepiece 2110 may include a right internally coupled grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Image-modulated light can be transferred to the user's eye via the internally coupled gratings 2112 and 2118, OPE 2114 and 2120, and EPE 2116 and 2122. Each internally coupled grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to gradually deflect light downward toward its associated EPE 2122, 2116, thereby extending the formed exit pupil horizontally.Each EPE 2122, 2116 can be configured to gradually redirect at least a portion of the light received from its corresponding OPE grating sets 2120, 2114 outward to a user eyebox position (not shown) defined behind the eyepieces 2108, 2110, causing the exit pupil formed in the eyebox to extend vertically. Alternatively, instead of the internally coupled grating sets 2112 and 2118, OPE grating sets 2114 and 2120, and EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 may include other arrays of gratings and / or refractive and reflective features to control the coupling of image-modulated light to the user's eye.

[0044] In some embodiments, the wearable head device 2102 may include a left lance arm 2130 and a right lance arm 2132, the left lance arm 2130 including a left speaker 2134 and the right lance arm 2132 including a right speaker 2136. The orthogonal coil electromagnetic receiver 2138 may be located in the left lance component or in another preferred location within the wearable head unit 2102. The inertial measuring unit (IMU) 2140 may be located in the right lance arm 2132 or in another preferred location within the wearable head device 2102. The wearable head device 2102 may also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142 and 2144 may preferably be oriented in different directions to cover a wider field of view together.

[0045] In the embodiment shown in Figures 2A-2D, the left source 2124 of the per-image modulated light can be optically coupled into the left eyepiece 2108 through the left internal coupling grating set 2112, and the right source 2126 of the per-image modulated light can be optically coupled into the right eyepiece 2110 through the right internal coupling grating set 2118. The per-image modulated light sources 2124 and 2126 may include, for example, an optical fiber scanning device, an electronic optical modulator including a digital light processing (DLP) chip or a liquid crystal on silicon (LCoS) modulator, a projector, or a light-emitting display such as a microlight-emitting diode (μLED) or microorganic light-emitting diode (μOLED) panel coupled into the internal coupling grating sets 2112 and 2118 using one or more lenses per side. The input coupling grid sets 2112 and 2118 can deflect the light from the image-modulated light sources 2124 and 2126 to an angle above the critical angle for total internal reflection (TIR) ​​for the eyepieces 2108 and 2110. The OPE grid sets 2114 and 2120 gradually deflect the propagating light downwards toward the EPE grid sets 2116 and 2122 by TIR. The EPE grid sets 2116 and 2122 gradually combine the light toward the user's face, including the pupil of the user's eye.

[0046] In some embodiments, as shown in Figure 2D, the left eyepiece 2108 and the right eyepiece 2110 each include a plurality of waveguides 2402. For example, each eyepiece 2108, 2110 may include a plurality of individual waveguides, each dedicated to a separate color channel (e.g., red, blue, and green). In some embodiments, each eyepiece 2108, 2110 may include a plurality of sets of such waveguides, each set configured to impart a different wavefront curvature to the emitted light. The wavefront curvature may be convex with respect to the user's eye, for example, to present a virtual object positioned at a certain distance in front of the user (e.g., only a distance corresponding to the reciprocal of the wavefront curvature). In some embodiments, the EPE grating sets 2116, 2122 may include curved grating grooves to produce a convex wavefront curvature by modifying the Poynting vector of the light emitting across each EPE.

[0047] In some embodiments, stereoscopically adjusted left and right eye images can be presented to the user through image-perfect optical modulators 2124, 2126 and eyepieces 2108, 2110 to create the perception that the displayed content is three-dimensional. The perceived reality of the presentation of the three-dimensional virtual object can be enhanced by selecting waveguides (and therefore corresponding wavefront curvatures) such that the virtual object is displayed at a distance approximating the distance indicated by the stereoscopic left and right images. This technique can also reduce motion sickness experienced by some users, which may be caused by the difference between the depth perception cues provided by the stereoscopic left and right eye images and the automatic near and far accommodation of the human eye (e.g., object distance-dependent focus).

[0048] Figure 2D illustrates a top-edge view of the right eyepiece 2110 of an exemplary wearable head device 2102. As shown in Figure 2D, the plurality of waveguides 2402 may include a first subset 2404 of three waveguides and a second subset 2406 of three waveguides. The two subsets of waveguides 2404 and 2406 can be distinguished by different EPE gratings, each featuring different grating line curvatures to impart different wavefront curvatures to the emitted light. Within each of the waveguide subsets 2404 and 2406, each waveguide can be used to couple different spectral channels (e.g., one of the red, green, and blue spectral channels) to the user's right eye 2206. (Although not shown in Figure 2D, the structure of the left eyepiece 2108 is similar to that of the right eyepiece 2110.)

[0049] Figure 3A illustrates an exemplary handheld controller component 300 of the mixed reality system 200. In some embodiments, the handheld controller 300 includes a gripping portion 346 and one or more buttons 350 positioned along the upper surface 348. In some embodiments, the buttons 350 may be configured for use as optical tracking targets for tracking the 6-degree-of-freedom (6DOF) motion of the handheld controller 300 in conjunction with, for example, a camera, LiDAR, or other sensor (which may be mounted within the head unit of the mixed reality system 200 (e.g., a wearable head device 2102)). In some embodiments, the handheld controller 300 includes a tracking component (e.g., an IMU or other suitable sensor) for detecting position or orientation, such as position or orientation relative to the wearable head device 2102. In some embodiments, such a tracking component may be positioned within the handle of the handheld controller 300 and / or mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals corresponding to a button press state, or one or more of the position, orientation, and / or movement of the handheld controller 300 (e.g., via the IMU). Such output signals may be used as inputs to the processor of the mixed reality system 200. Such inputs may correspond to the position, orientation, and / or movement of the handheld controller (by extension, the position, orientation, and / or movement of the user's hand holding the controller). Such inputs may also correspond to the user pressing button 350.

[0050] Figure 3B illustrates an exemplary auxiliary unit 320 of the mixed reality system 200. The auxiliary unit 320 may include a battery for providing energy to operate the system 200 and a processor for running a program to operate the system 200. As shown, the exemplary auxiliary unit 320 includes a clip 2128 for attaching the auxiliary unit 320 to a user's belt, etc. It will also become apparent that other shape factors are suitable for the auxiliary unit 320 and may include shape factors that do not involve mounting the unit on a user's belt. In some embodiments, the auxiliary unit 320 is coupled to a wearable head device 2102 through a multi-tube cable, which may include, for example, electrical wires and optical fibers. Wireless connectivity between the auxiliary unit 320 and the wearable head device 2102 can also be used.

[0051] In some embodiments, the mixed reality system 200 may include one or more microphones that can detect sound and provide corresponding signals to the mixed reality system. In some embodiments, the microphones may be attached to or integrated with a wearable head device 2102 and may be configured to detect the user's voice. In some embodiments, the microphones may be attached to or integrated with a handheld controller 300 and / or an auxiliary unit 320. Such microphones may be configured to detect ambient sounds, background noise, the voice of the user or a third party, or other sounds.

[0052] Figure 4 shows an exemplary functional block diagram that may correspond to an exemplary mixed reality system, such as the mixed reality system 200 described herein (which may correspond to the mixed reality system 112 relating to Figure 1). As shown in Figure 4, the exemplary handheld controller 400B (which may correspond to the handheld controller 300 ("Totem")) includes the Totem / wearable head device 6-degree-of-freedom (6DOF) totem subsystem 404A, and the exemplary wearable head device 400A (which may correspond to the wearable head device 2102) includes the Totem / wearable head device 6DOF subsystem 404B. In embodiments, the 6DOF totem subsystem 404A and the 6DOF subsystem 404B cooperate to determine six coordinates of the handheld controller 400B relative to the wearable head device 400A (e.g., offset in three translation directions and rotation along three axes). The six degrees of freedom may be expressed relative to the coordinate system of the wearable head device 400A. The three translation offsets may be represented as X, Y, and Z offsets in such a coordinate system, as a translation matrix, or as some other representation. The rotational degrees of freedom may be represented as a sequence of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or as some other representation. In some embodiments, a wearable head device 400A, one or more depth cameras 444 (and / or one or more non-depth cameras) contained within the wearable head device 400A, and / or one or more optical targets (e.g., a button 350 on a handheld controller 400B as described herein or a dedicated optical target contained within the handheld controller 400B) can be used for 6DOF tracking. In some embodiments, the handheld controller 400B may include a camera as described herein, and the wearable head device 400A may include an optical target for optical tracking in conjunction with the camera.In some embodiments, the wearable head device 400A and the handheld controller 400B each include a set of three orthogonally oriented solenoids, which are used to wirelessly transmit and receive three distinguishable signals. The 6DOF of the wearable head device 400A relative to the handheld controller 400B can be determined by measuring the relative magnitudes of the three distinguishable signals received in each of the coils used for reception. In addition, the 6DOF totem subsystem 404A may include an inertial measurement unit (IMU), which is useful for providing improved accuracy and / or more timely information regarding the high-speed movement of the handheld controller 400B.

[0053] In some embodiments, for example, it may be necessary to transform the coordinates from a local coordinate space (e.g., a coordinate space fixed relative to the wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment) in order to compensate for the movement of the wearable head device 400A relative to coordinate system 108. For example, such a transformation may be necessary to ensure that the display of the wearable head device 400A presents virtual objects in their expected position and orientation relative to the real environment (e.g., a virtual person seated in a real chair facing forward, regardless of the position and orientation of the wearable head device), rather than in a fixed position and orientation on the display (e.g., the same position at the lower right corner of the display), and to preserve the illusion that the virtual objects exist in the real environment (and do not appear unnaturally positioned in the real environment as the wearable head device 400A shifts and rotates). In some embodiments, compensatory transformations between coordinate spaces can be determined by processing images from a depth camera 444 using SLAM and / or visual odometry procedures to determine the transformation of the wearable head device 400A to coordinate system 108. In the embodiment shown in Figure 4, the depth camera 444 can be coupled to a SLAM / visual odometry block 406 and provide images to the block 406. The SLAM / visual odometry block 406 implementation may include a processor configured to process these images and then determine the position and orientation of the user's head, which can be used to identify transformations between the head coordinate space and another coordinate space (e.g., inertial coordinate space). Similarly, in some embodiments, an additional source of information regarding the user's head pose and location is obtained from an IMU 409. The information from the IMU 409 can be integrated with the information from the SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information regarding the rapid adjustment of the user's head pose and position.

[0054] In some embodiments, the depth camera 444 can supply a 3D image to a hand gesture tracker 411, which may be implemented within the processor of the wearable head device 400A. The hand gesture tracker 411 can identify the user's hand gestures, for example, by matching the 3D image received from the depth camera 444 to a stored pattern representing the hand gesture. Other preferred techniques for identifying the user's hand gestures will also become apparent.

[0055] In some embodiments, one or more processors 416 may be configured to receive data from the wearable head device's 6DOF headgear subsystem 404B, IMU 409, SLAM / visual odometry block 406, depth camera 444, and / or hand gesture tracker 411. The processor 416 can also transmit and receive control signals from the 6DOF totem system 404A. The processor 416 may be coupled wirelessly to the 6DOF totem system 404A, for example, in embodiments where the handheld controller 400B is not tethered. The processor 416 may further communicate with additional components such as an audio / visual content memory 418, a graphical processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 may be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 may include a left channel output coupled to a left source 424 of light modulated for each image, and a right channel output coupled to a right source 426 of light modulated for each image. The GPU 420 can output stereoscopic image data to the source 424, 426 of light modulated for each image, for example, as described herein with respect to Figures 2A-2D. The DSP audio spatialization device 422 can output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatialization device 422 may receive an input from the processor 419 indicating a direction vector from the user to a virtual sound source (e.g., which may be moved by the user via a handheld controller 320). Based on the direction vector, the DSP audio spatialization device 422 can determine the corresponding HRTF (e.g., by accessing the HRTF or by interpolating multiple HRTFs). The DSP audio spatialization device 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object.This can improve the credibility and realism of virtual sounds by incorporating the user's relative position and orientation to virtual sounds within a mixed reality environment; that is, by presenting virtual sounds that match the user's expectations of what they would hear if they were real sounds in a real environment.

[0056] In some embodiments, such as those shown in Figure 4, one or more of the processor 416, GPU 420, DSP audio spatialization device 422, HRTF memory 425, and audio / visual content memory 418 may be contained within an auxiliary unit 400C (which may correspond to the auxiliary unit 320 described herein). The auxiliary unit 400C may include a battery 427 that powers its components and / or supplies power to the wearable head device 400A or handheld controller 400B. Including such components within an auxiliary unit that can be mounted on the user's waist can limit the size and weight of the wearable head device 400A, which in turn can reduce fatigue in the user's head and neck.

[0057] Figure 4 presents elements corresponding to various components of an exemplary mixed reality system, but various other preferred arrangements of these components will also be apparent to those skilled in the art. For example, the elements presented in Figure 4 as associated with the auxiliary unit 400C may instead be associated with a wearable head device 400A or a handheld controller 400B. Furthermore, some mixed reality systems may omit the handheld controller 400B or the auxiliary unit 400C entirely. Such changes and modifications are understood to be within the scope of the disclosed embodiments.

[0058] Displaying virtual content in a mixed reality environment so that it corresponds to real content can be challenging. For example, it may be desirable to display the virtual object 122B in Figure 1C in the same location as the real object 122A. Doing so may involve several capabilities of the mixed reality system 112. For example, the mixed reality system 112 may create a three-dimensional map of the real environment 104A and the real objects within the real environment 104A (e.g., the lamp 122A). The mixed reality system 112 may also establish its location within the real environment 104A (which may correspond to the user's location within the real environment). The mixed reality system 112 may further establish its orientation within the real environment 104A (which may correspond to the user's orientation within the real environment). The mixed reality system 112 may also establish its movement relative to the real environment 104A, e.g., line and / or angular velocity and line and / or angular acceleration (which may correspond to the user's movement relative to the real environment). SLAM can be a method for displaying a virtual object 122B in the same location as the real object 122A, even when the user 110 moves around in room 104A, averts their gaze from the real object 122A, and then returns their gaze to the real object 122A. In some embodiments, SLAM calculations such as those described herein can be performed in the mixed reality system 112 via the SLAM / visual odometry block 406 described above. In some embodiments, SLAM calculations can be performed via the processor 416, the GPU 420, and / or any other suitable component of the mixed reality system 112. SLAM calculations such as those described herein can utilize any suitable sensor of the mixed reality system 112, such as those described herein.

[0059] It may be desirable to initiate SLAM in an accurate, computationally efficient, and low-latency manner. As used herein, latency may refer to the time delay between a change in the position or orientation of a component of the mixed reality system (e.g., rotation of a wearable head device) and the reflection of that change as represented in the mixed reality system (e.g., the display angle of the field of view presented on the display of the wearable head device). Computational inefficiency and / or long latency may negatively impact the user experience using the mixed reality system 112. For example, if user 110 looks around room 104A, virtual objects may appear to "shake" as a result of the user's movement and / or long latency. Accuracy may be important for generating an immersive mixed reality environment; otherwise, virtual content competing with real content may remind the user of the distinction between virtual and real content, reducing the user's sense of immersion. Furthermore, in some cases, latency may result in motion sickness, headaches, or other negative physical experiences for some users. Computational inefficiency can lead to exacerbated problems in embodiments where the mixed reality system 112 is a mobile system that relies on a limited power source (e.g., a battery). The systems and methods described herein can result in an improved user experience as a result of more accurate, computationally efficient, and / or lower latency SLAM.

[0060] Figure 5 illustrates an exemplary method 500 for operating a SLAM system according to embodiments of the present disclosure. Method 500 is advantageous in that it enables the efficient use of a LiDAR or similar sensor (e.g., a sensor configured to capture multiple points during a scan) for a SLAM system (e.g., for environmental mapping applications, for mixed reality mapping applications). While Method 500 is illustrated as including the elements described, it should be understood that elements in a different order, additional elements, or fewer elements may be included without departing from the scope of the present disclosure.

[0061] In some embodiments, Method 500 allows two sides of a planar object to be distinguished by the direction of a vector normal toward the observation center (e.g., of a LiDAR sensor), and the undesirable effects of the double-sided problem can be reduced. For example, the observation center is the origin of the coordinate system of the depth sensor (e.g., a LiDAR sensor, as disclosed herein).

[0062] In some embodiments, Method 500 includes the step of performing PBA, which includes the step of jointly optimizing plane parameters and sensor attitude. Performing PBA has the advantage of reducing attitude errors and drift associated with planes that are not explicitly extracted.

[0063] In some embodiments, Method 500 can, advantageously, balance the accuracy and computational cost of the BA. In some embodiments, Method 500 comprises three parallel components, including localization, local mapping, and PBA. For example, the localization component tracks a plane frame by frame, corrects the distortion of the LiDAR point cloud, and aligns the new scan to a global planar model in real time. In addition, the distortion correction and alignment tasks can be accelerated by using an approximate rotation matrix for small motions between two scans. The local mapping and PBA components can correct for drift, improve the map, and ensure that the localization component can provide accurate orientation. In some embodiments, the point-plane cost includes a special structure that can be utilized to reduce the computational cost related to the PBA. Based on this structure, an integrated point-plane cost can be formed to accelerate local mapping.

[0064] While some embodiments of this specification describe LiDAR systems, it should be understood that these descriptions are illustrative. Although LiDAR sensors are described in detail, it should be understood that other sensors or more specific types of sensors (e.g., sensors configured to capture multiple points during scanning, depth sensors, sensors providing depth information, mechanical spinning LiDAR, solid-state LiDAR, RGBD cameras) may be used to perform the described operations, incorporated into the described systems, and achieve similar described benefits without departing from the scope of this disclosure. In addition, it is assumed that the use of a LiDAR sensor (or a similar sensor) may be complemented, as appropriate, by the use of other sensors such as depth cameras, RGB cameras, acoustic sensors, infrared sensors, GPS units, and / or inertial measurement units (IMUs). For example, such complementary sensors may be used in conjunction with LiDAR to refine the data provided by the LiDAR sensor, to provide redundancy, or to enable reduced-power operation.

[0065] The embodiments described herein may be implemented using a device including a LiDAR or similar sensor, such as a mixed reality device (e.g., mixed reality system 112, mixed reality system 200, mixed reality system described with respect to Figure 4), a mobile device (e.g., a smartphone or tablet), or a robotic device, and / or using a first device (e.g., a cloud computing device, a server, a computing device) configured to communicate with a second device (e.g., a device configured to transmit and / or receive environmental mapping information, a device including a LiDAR or similar sensor, or a device configured to receive environmental mapping information but not including a LiDAR or similar sensor) (e.g., to simplify or reduce the requirements of the second device).

[0066] In some embodiments, results from the systems and methods implemented herein (e.g., PBA results, point clouds, keyframes, poses, calculation results) are stored in a first device that uses the results (e.g., to acquire environmental mapping information) and / or in a second device (e.g., to acquire environmental mapping information) configured to communicate with the second device, which is configured to request and / or receive the results. By storing the results, the time associated with processing and / or acquiring environmental mapping information can be reduced (e.g., the results do not need to be recalculated, the results can be calculated on a faster device), and the steps of processing and / or acquiring environmental mapping information can be simplified (e.g., the results do not need to be calculated on a slower, and / or battery-consuming device).

[0067] In some embodiments, the instructions for carrying out the methods described herein are preloaded onto the device (e.g., pre-installed on the device prior to use). In some embodiments, the instructions for carrying out the methods described herein may be downloaded onto the device (e.g., installed on the device later). In some embodiments, the methods and operations described herein may be performed offline (e.g., mapping is performed offline while the system using the mapping is inactive). In some embodiments, the methods and operations described herein may be performed in real time (e.g., mapping is performed in real time while the system using the mapping is active).

[0068] In some embodiments, LiDAR scanning [ka] This includes a set of points that may or may not be captured simultaneously. [ka] The stance is, [ka] It can be defined as the rigid body transformation of the LiDAR frame to the global frame at the time when the last point in is captured. As described herein, the rigid body transformation can be expressed as rotation and translation. [ka] A transformation can be associated with a transformation matrix. [ka]

[0069] In some embodiments, the rotation matrix R lies within a special orthogonal group SO(3). To represent R, the angle-axis parameter representation ω=[ω1;ω2;ω3] may be used, where T can be parameterized as x=[ω;t] (for example, for the optimizations described herein). The strain matrix of ω may be defined as follows: [ka]

[0070] In the above, [ω] x However, in the identity, it lies within the tangent space so(3) of the manifold SO(3). The exponential map exp: so(3) → SO(3) can be expressed as follows: [ka]

[0071] A planar object may have two sides, and in some cases, these two sides may be in close proximity to each other. As discussed above, the proximity of two sides of a planar object can lead to ambiguity errors, making it ambiguous which side of the planar object is facing the sensor (e.g., the double-sided problem). For example, a robot may enter a room and mistakenly capture the side of a wall that is located opposite to the intended side. The two sides of a planar object can be distinguished by identifying a vector that is normal to the plane and points toward the observation center (e.g., the center of the LiDAR sensor, the center of the sensor).

[0072] A plane can be represented by π = [n; d], where n is the normal vector to the observation center at ||n||=1, and d is the signed distance from the origin to the plane. To parameterize the plane, the nearest-neighbor (CP) vector η = nd can be used.

[0073] As will be discussed, a plane can be represented by π. The plane π can be represented in a global coordinate system, where p is an observation of π in a local coordinate system. The transformation matrix from the local coordinate system to the global coordinate system can be T, as discussed above. [ka] The coordinate system can be homogeneous in p. The point-plane residual δ can be expressed as follows: [ka]

[0074] In some embodiments, method 500 is a new scanning [ka] This includes the step of performing (step 502). For example, a new LiDAR scan is performed. In another embodiment, a scan with multiple points (e.g., tens of points, hundreds of points, thousands of points) is performed.

[0075] In some embodiments, the LiDAR sensor starts in a stationary state so that the first frame is not affected by motion distortion. Planes may be extracted from the LiDAR point cloud using any preferred method. For example, planes may be extracted based on a region growth method, the normal of each point may be estimated, and the point cloud may be segmented by clustering points with similar normals. For each cluster, RANSAC or a similar algorithm may be used to detect the plane. In some embodiments, planes containing more than a threshold number of points (e.g., 50) are retained, and planes containing fewer than a threshold number of points are discarded.

[0076] In some embodiments, Method 500 includes a step of performing localization (step 504). The localization step may include a step of tracking planes (e.g., a frame-by-frame tracking step, a scan-by-scan tracking step). Each local plane may be associated with a global plane, and this relationship may be maintained during tracking. By tracking planes frame by frame or scan by scan, computation time may be saved because the data association between the global plane and local observations may be maintained. As another exemplary benefit, due to the local-global plane correspondence, the point-plane cost is made possible to simultaneously correct for distortion of the point cloud (e.g., the global point cloud) and estimate its orientation. A linear approximation formula for the rotation matrix (as described herein) may be used to simplify the calculation. The localization step may also include a step of aligning the current scan within the global point cloud.

[0077] For example, scanning [ka] These are two consecutive scans, [ka] teeth, [ka] The set of points in the i-th plane detected in the parameter π k-1,i And, [ka] is, m i The second global plane [ka] This is associated with the set of local-global point-plane correspondences. [ka] It can be expressed as follows.

[0078] According to the illustrative method, a KD tree is, [ka] It can be constructed in relation to the point. [ka] Each time, [ka] The n nearest neighbors (e.g., n=2) are found. Redundant points are in the point set. [ka] It may be excluded in order to find the RANSAC algorithm (or another suitable algorithm) [ka] From the plane π k,i Fit it and set the in-liner [Chemistry] may be applied to obtain. [Chemistry] is, π k,i can be extended by collecting neighboring points that are smaller than the threshold distance from π (e.g., 5 cm). [Chemistry] where the number of points is greater than a threshold (e.g., 50) and π k,i and π [[ID=2)]] k-1,i [End]]when the angle between and is smaller than a threshold angle (e.g., 10 degrees), a set of point - plane correspondences [Chemistry] can be obtained with respect to the scan [Chemistry] can be determined.

[0079] Figure 6 illustrates an exemplary timing diagram 600 of a SLAM system according to an embodiment of the present disclosure. The timing diagram can illustrate points captured during a scan. For example, [Chemistry] where the points at can be captured during [t k-1 [End]], t k [End]] with an interval Δt. For points at time t ∈ [t k-1 [End]], t k [End]] (e.g., LiDAR points, sensor points), the pose [Chemistry] is, T k-1 [End]]and Tk Relative posture T between k,k-1 This can be obtained by interpolation (for example, linear interpolation).

[0080] Turning back to the exemplary method 500, the step of performing positioning may also include the step of performing attitude estimation and / or attitude distortion correction. For example, the rigid body transformation between scan k-1 and scan k is T k,k-1 It can be defined as, (R k,k-1 ,t k,k-1 ) can be the corresponding rotation and translation. ω k、k-1 R k,k-1 This can be an angle-axis representation of x k,k-1 =[ω k,k-1 ;t k,k-1 ] but T k,k-1 It can be used to parameterize it.

[0081] The rigid body transformations at the last point of scan k-1 and scan k are, respectively, T k-1 and T k It is possible. k-1 And, T k And, T k,k-1 The relationship between and can be given by the following: [ka] x k,k-1 However, if calculated, T k,k-1 However, it can be requested, T k However, it can be solved.

[0082] LiDAR measurements can be acquired at different times. When the LiDAR is moving between acquisitions, the LiDAR point cloud can be distorted by the motion. This distortion can be reduced by interpolating the orientation of each point (e.g., linear interpolation). k-1 and t k These could be the times of the last point in scan k-1 and k, respectively. t∈(t k-1 ,t k Regarding ], s can be defined as follows: [ka]

[0083] Next, t∈[t k-1 ,t k Rigid body transformation in ] [ka] of [ka] However, it can be estimated by linear interpolation. [ka]

[0084] As discussed above, a set of point-plane correspondences [ka] However, scanning [ka] It may be requested regarding this. [ka] and the j-th point [ka] However, time t k,i,j It can be captured in that. [ka] The residuals related to this can be described as follows: [ka]

[0085] The following equation can be defined. [ka]

[0086] By substituting equation (8) into (7), we obtain the following: [ka]

[0087] N point-plane correspondence k Set of 1 [ka] However, it may exist. [ka] is, N k,i It may have x points. k,k-1 The least squares cost function for can be as follows: [ka]

[0088] The least squares problem presented by equation (10) can be solved using the Levenberg-Marquardt (LM) method.

[0089] In some embodiments, the motion within the scan is small. Therefore, [ka] (For example, from equation (6)) can be approximated by a first-order Taylor expansion. [ka]

[0090] In addition, the following may be defined.

Chem.

[0091] By substituting the first - order approximation formula in Equation (11) and the definition of Equation (12) into Equation (9) and expanding (9), a linear constraint on x k,k-1 can be obtained.

Chem.

Chem.

[0092]

Chem.

Chem.

[0093] Thus, a closed - form solution can be obtained,

Chem.

[0094] In some embodiments,

Chem.

Chem.

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0095] can be used to generate a new linear system for equation (14). The new linear system is

Chemical formula

Chemical formula

[0096] The step of performing localization may also include a step of determining whether a new keyframe needs to be inserted. For example, a new keyframe needs to be inserted if one of the following conditions is met: (1) the distance between the current scan and the last keyframe is greater than a threshold distance (e.g., 0.2 m), or (2) a threshold percentage of points in the current scan (e.g., 20%) are not tracked. In accordance with the decision that a new keyframe needs to be inserted, the new keyframe is inserted. In some embodiments, local mapping is performed in response to the insertion of a new keyframe. In some embodiments, if one of these conditions is met and local mapping has been performed with respect to the last keyframe, local mapping is performed with respect to the newly inserted keyframe.

[0097] In some embodiments, method 500 includes the step of performing local mapping (step 506). During local mapping, keyframe poses within the window and / or planes observed by these keyframes may be optimized, while keyframe poses outside the window remain fixed. The step of performing local mapping may include the step of detecting planes in the remaining point cloud.

[0098] As an example, with respect to a new keyframe, the corresponding scan may be distortion-corrected by applying the transformation of equation (6). Then, the planes with respect to untracked points and the planes with respect to points exceeding a threshold amount (e.g., 50 points) are preserved.

[0099] The step of performing local mapping may include a step of generating an estimated local-global match. For example, a new local plane from this detection may be matched with a global plane using a known orientation. Plane normals may first be used to select candidate global planes. For example, global planes whose normals are approximately parallel to the local planes (e.g., having an angle between them that is below a threshold angle (e.g., 10 degrees)) are considered.

[0100] The step of performing local mapping may include checking for matches and inserting new planes. For example, the distance between local plane points and candidate global planes (e.g., selected from the steps described above) may be calculated, and candidates associated with the minimum mean distance are preserved. For the preserved candidates, the correspondence is accepted if the threshold percentage of the point-plane distance (e.g., 90%) is less than the threshold distance γ. If the threshold percentage of the point-plane distance (e.g., 90%) is less than twice the threshold distance γ, a further geometric consistency check may be performed (GCC). If none of these conditions are met, a new global plane is introduced.

[0101] GCC may be performed to further verify the correspondence between the local plane and the candidate global plane (e.g., identified based on the steps described above). A linear system described by equation (14) is constructed with respect to the original correspondence. [ka] It can be expressed as follows: [ka] A new linear system can be constructed for each new correspondence (e.g., the i-th new correspondence), expressed as follows. The extended linear system for each new correspondence can be solved as follows: [ka]

[0102] [ka] but, [ka] If it is a solution, the following conditions may be required to be satisfied. [ka] (B) The threshold percentage of the point-plane distance (e.g., 90%) is less than the threshold distance γ.

[0103] In some embodiments, λ is set to 1.1. If a new correspondence is added based on the steps above, the attitude is re-estimated using all correspondences, including the new correspondence.

[0104] Equation (16) is a normal equation. [ka] It can be solved efficiently by solving this. [ka] and [ka] and [ka] However, it can be found that these terms can control the calculation and are shared among all new linear systems. Therefore, these terms may, favorably, need to be calculated only once. In addition, the final linear system including all constraints is [ka] Since this can be found for all i, it can be solved efficiently.

[0105] The step of performing local mapping may include the step of performing local PBA and updating the map. For example, if at least one new local-global plane correspondence is found (e.g., from the steps described above, it is determined that the local plane is coplanar with the global plane), or if the SLAM system is initialized, a global PBA (e.g., described with respect to step 508) may be performed. Otherwise, a local PBA may be performed.

[0106] In some embodiments, local mapping optimizes keyframes within a sliding window and the planes observed by these keyframes. For example, the size of the sliding window is N w It can be set to 5. The latest N w Each keyframe is set [ka] It can form N w The plane observed by each keyframe is set [ka] Can form, set [ka] Keyframes outside the sliding window to which the plane in question corresponds are set [ka] It forms. [ka] The pose of keyframes in can be fixed. In some embodiments, the following cost function is minimized during local mapping. [ka] In the formula, the first set of sums is C w It can be defined as, and the second set of sums is C f It can be defined as follows. [ka] This can be the set of planes observed by the i-th keyframe, [ka] This is the j-th plane π j This could be a set of keyframes outside the corresponding sliding window. w and J f These are, respectively, C w Residual δ in w and C f Residual δ in fThis can be the Jacobian matrix. Next, the residual δ based on equation (17) l A vector and its corresponding Jacobian matrix J l It may have the following forms: [ka]

[0107] Figure 7 illustrates an exemplary factor graph according to an embodiment of the present disclosure. For example, exemplary factor graph 700 is a factor graph associated with observations of the environment with respect to local mapping with a window size of 3 (for example, as described in the steps above). During local mapping, the orientations (e.g., x4, x5, x6) and planes (e.g., π4, π5, π6) within the sliding window may be optimized, while other orientations (e.g., x1, x2, x3) and planes (e.g., π1) may be fixed. [ka] is keyframe x i Plane π recorded in j It can be a set of points. The cost function is explained in terms of equation (17), and consists of two parts, namely, C f and C w It may include.

[0108] Turning back to Exemplary Method 500, the LM algorithm solves a system of linear equations. [ka] Solving this equation can be used to calculate the step vector. δ l and J l Calculating it directly can be time-consuming. δ w and J w It can be simplified, δ f and J f This can be calculated as follows:

[0109] Cf δ in f Regarding, and posture T m Plane π in j Regarding this, δ from equation (17) mjk (η j ) can be rewritten as follows: [ka]

[0110] The orientation of keyframes outside the sliding window can be fixed. In equation (19), K mj residual δ mjk (η j By stacking the following, the residual vector can be obtained. [ka] During the ceremony, [ka] That is the case.

[0111] all [ka] By stacking, π outside the sliding window j Restrictions regarding this may be required. [ka] During the ceremony, [ka] That is the case.

[0112] Therefore, C f The residual vector δ in f (η j ) but all [Chemical formula] can be obtained by stacking and can have the following forms. [Chemical formula]

[0113] J f To calculate, the derivative of δ mjk (η) can be calculated first. η j is [Chemical formula] can be defined as. Then, the derivative of δ j with respect to H mjk (η) can have the following form. [Chemical formula] In the formula, <000126​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

[0116] In some embodiments, due to the number of points recorded during a scan (e.g., a LiDAR scan, a scan of a sensor configured to capture multiple points), δ in equation (22) f and J in equation (25) f can be time-consuming to calculate. The integration point-plane cost can significantly reduce this calculation time.

[0117] Using equation (21), at least two sums of C from equation (17) f can be calculated. <#1300><#1301>

Chemical formula

Chemical formula

Chemical formula

Chemical formula

Chemical formula

[0119] [ka] This includes the information required to implement the LM algorithm, and, advantageously, can reduce the computation time associated with point-plane cost calculations. The following lemma may be presented.

[0120] [ka]

[0121] Proof: In some embodiments, J defined in equation (25) f and defined in equation (27) [ka] This is a block vector. Using the block matrix multiplication formula, the following can be obtained: [ka]

[0122] J in equation (24) j According to the definition, [ka] is one non-zero term [ka] H has the following characteristics, defined in equation (26) j According to the definition, the following applies: [ka]

[0123] Similarly, in equation (27) [ka] According to the definition, [ka] is one non-zero term [ka] The equation (26) [ka] Based on this, the following applies: [ka]

[0124] Therefore, according to equations (29) and (30), [ka] Therefore, based on equation (28), [ka] That is the case.

[0125] [ka]

[0126] Proof: From equations (25), (22), and (27), [ka] is an element [ka] This is a block vector. Using the block matrix multiplication formula, the following can be obtained. [ka]

[0127] J in equation (24) j Definition of and equation (21) of δ j According to the definition, [ka] is one non-zero term [ka] H has the following characteristics, defined in equation (26) j According to the definition, the following applies: [ka]

[0128] Similarly, in equation (27) [ka] According to the definition, [ka] is one non-zero term [ka] The equation (26) [ka] Based on this, the following applies: [ka]

[0129] Therefore, according to equations (32) and (33), [ka] Therefore, according to equation (31), [ka] That is the case.

[0130] Lemma 3: In some embodiments, δ is defined in equation (21). j (η j ) in K j Individual constraints exist. [ka] The runtime for calculating each of the original calculations is advantageous, in that each [ka] For terms associated with [ka] It is possible.

[0131] Proof: In equation (27) [ka] J in equation (24) j and δ in equation (21) j By comparing, the difference between the two pairs is P j L to replace j It may be used. j It has four rows, P ij is, K jIt has rows. Therefore, [ka] The runtime for calculating is [ka] For calculating [ka] And, advantageously, it is reduced. In addition, according to the formula for matrix multiplication, [ka] The runtime for calculating each of these is, [ka] For calculating [ka] And, as an advantage, it is reduced.

[0132] As discussed above, the integral cost matrix H from equation (26) j This may be important to simplify the calculations associated with the cost function. In some embodiments, H j This is the plane π in the keyframes outside the sliding window. j Summarize the costs derived from the observations. If a large number of observations exist, H j Calculating this from scratch each time can be time-consuming. Therefore, in some embodiments, H j This avoids redundant calculations and the integral cost matrix H j To reduce the calculation time required to calculate it, it may be calculated gradually.

[0133] For example, the plane π j There are K keyframes outside the corresponding sliding window, and the corresponding keyframes are calculated based on equation (26). [ka] This can result in δ as defined in equation (21). nj (η nj )=P nj π j This is π at the nth keyframe, which is about to move outside the sliding window. j This could be a residual vector derived from the observation results.

[0134] H in equation (26) j Definition of and equation (21) of P j Using the definition, for K+1 observations [ka] However, it may be required. [ka]

[0135] From equation (34), the integral cost matrix H j However, this is maintained for each global plane. [ka] but, [ka] It is updated by. In some embodiments, when a keyframe is moved outside the sliding window, the integral matrix of the plane observed by this keyframe is updated using equation (34).

[0136] Figure 8 illustrates an exemplary schematic diagram 800 for updating the integral cost matrix according to an embodiment of the present disclosure. For example, schematic diagram 800 shows the integral cost matrix H j The diagram illustrates how it is updated gradually. For example, the nth keyframe attempts to move outside the sliding window, and π outside the sliding window j There are K sets of observation results. nj However, in the form defined by equation (20), π in the nth keyframe j This can be determined from observational results (for example, those indicated by shading). [ka] teeth, [ka] It is updated by [this method].

[0137] In the exemplary method 500, specifically, looking again at equation (17), in some embodiments, since equation (17) is a large-scale optimization problem, it may be desirable to minimize equation (17). The residual vector δ from equation (18) l and the corresponding Jacobian matrix J l This can be simplified. J in equation (18) w and δ w This is the reduced Jacobian obtained using the LM algorithm. [ka] and reduced residual vector [ka] It can be replaced by. More specifically, [ka] Therefore, the following can be defined: [ka] Based on this, the following theorem can be derived.

[0138] [ka]

[0139] Proof: J from equation (18) l and δ l From the definition and equation (35) [ka] Based on the definition, the following relationship can be derived. [ka]

[0140] As discussed, [ka] Therefore, [ka] Similarly, [ka] Therefore, [ka] That is the case.

[0141] According to the theorem above, [ka] In the LM algorithm, J l and δ l It can be used to replace. In some embodiments, [ka] J l and δ l It has a lower dimension, and therefore the computation time is favorable. [ka] This can be reduced by deriving and using it.

[0142] In some embodiments, global PBA is performed when it is determined that a loop is closed. For example, a loop is determined to be closed according to a determination that a previous plane has been revisited, and global PBA is performed according to the determination that the loop is closed. In some embodiments, global PBA may jointly optimize LiDAR (or similar sensor) attitude and plane parameters by minimizing the point-plane distance.

[0143] In some embodiments, method 500 includes the step of performing a global mapping (step 508). During the global mapping, keyframe poses and / or plane parameters may be optimized globally (e.g., without windows, compared to local mapping). The step of performing a global mapping may include the step of performing a global PBA and / or the step of updating the map.

[0144] For example, suppose there are M planes and N keyframes, and the transformation matrix for the i-th pose is T i This is possible, and the j-th plane is parameter π j This includes the measurement of the j-th plane in the i-th orientation, which is defined as K. ijIt could be a set of individual points. [ka]

[0145] [ka] This may provide one constraint on the i-th orientation and the j-th plane. Based on equation (3), [ka] Residual δ related to ijk However, it can be expressed as follows: [ka]

[0146] x i and η j These are, respectively, T i and π j It can represent the parameterized expression of δ. ijk is, x i and η j It can be a function of . The PBA process minimizes the following least squares problem (which can be solved efficiently): [ka] This can be refined jointly. [ka]

[0147] The original Jacobian matrix and the original residual vector of equation (39) are, respectively, J g and r g This is possible. (For example, with respect to local PBAs) the reduced Jacobian matrix, as described herein. [ka] and reduced residual vector [ka] However, in the LM algorithm, J g and r g It can be used to replace and may reduce calculation time.

[0148] After the posture is refined, the integral cost matrix H j This may need to be updated. As discussed above, the update can be carried out efficiently. According to equation (34), [ka] H j This can be important when calculating P. Based on equation (20), nj This can be expressed as follows: [ka]

[0149] In some embodiments, [ka] This is a fixed quantity and may only need to be calculated once. From equation (40), after PBA is performed, the refined posture T n teeth, [ka] It can be used to update, then H j This can be updated. [ka] This can be a large fixed quantity and may only need to be calculated once, [ka] Only H may need to be updated. j The calculation time related to this can be reduced, which is advantageous.

[0150] Figure 9 illustrates an exemplary factor graph 900 according to an embodiment of the present disclosure. For example, the factor graph 900 is associated with environmental observations and PBAs with N postures and M planes. i This can represent the i-th keyframe, π j This can represent the j-th global plane, [ka] is, x i π captured in j This could be a set of observations relating to the following. In this embodiment, x1 is fixed during this optimization, and the illustrated PBA problem is keyframe pose [ka] and plane parameters [ka] This includes a step of collaboratively optimizing it.

[0151] According to some embodiments, the method includes the steps of: scanning an environment, the scanning step of including detecting a plurality of points in the environment using a sensor; identifying a first plane of the environment, the first plane comprising at least a threshold number of points of the plurality of detected points; obtaining a plurality of observations associated with the environment, the plurality of observations comprising a first subset of observations and a second subset of observations; determining whether the first plane is coplanar with a second plane of the environment; performing plane bundle adjustment based on the first subset of observations and further on the second subset of observations in accordance with the determination that the first plane is coplanar with the second plane; and performing plane bundle adjustment based on the first subset of observations and not on the second subset of observations in accordance with the determination that the first plane is not coplanar with the second plane.

[0152] According to some embodiments, the sensor comprises a LiDAR sensor.

[0153] According to some embodiments, the normal of the first plane comprises a vector in the direction toward the observation center associated with the sensor.

[0154] According to some embodiments, the step of performing planar bundle adjustments on a first subset of observations further includes the step of calculating a cost function associated with keyframe poses in a sliding window associated with the first subset.

[0155] According to some embodiments, the method further includes the steps of: determining the distance between a portion of the scanned environment and a first keyframe; inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe exceeds a threshold distance; and refraining from inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe does not exceed a threshold distance.

[0156] According to some embodiments, the method further includes the steps of determining the percentage of points being tracked, wherein a first plane includes the points being tracked; inserting a keyframe according to the determination that the percentage of points being tracked is below a threshold percentage; and refraining from inserting a keyframe according to the determination that the percentage of points being tracked is not below a threshold percentage.

[0157] According to some embodiments, the method further includes the step of determining an integral cost matrix, and the step of performing a planar bundle adjustment includes the step of calculating a cost function using the integral cost matrix.

[0158] According to some embodiments, the method further includes the step of updating the integral cost matrix based on the previous integral cost matrix.

[0159] According to some embodiments, the method further includes the step of estimating keyframe poses.

[0160] According to some embodiments, the method further includes the step of linearly interpolating points of keyframe poses.

[0161] According to some embodiments, the method further includes the step of updating a map of the environment based on planar bundle adjustment.

[0162] According to some embodiments, the environment comprises a mixed reality environment.

[0163] According to some embodiments, the sensor comprises a sensor for a mixed reality device.

[0164] According to some embodiments, the method further includes the step of performing the Levenberg-Marquardt method to solve the least squares function.

[0165] According to some embodiments, the method further includes the steps of inserting a first plane into a map associated with an environment, based on a determination that the first plane is not coplanar with a second plane, and refraining from inserting a first plane into a map associated with an environment, based on a determination that the first plane is coplanar with a second plane.

[0166] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of determining whether a threshold percentage of the point-plane distance is less than a threshold distance.

[0167] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of performing a geometric consistency check.

[0168] According to some embodiments, the method further includes the step of storing in memory a plurality of observations associated with the environment.

[0169] According to some embodiments, the method further includes a step of transmitting information about the environment, the information about the environment being associated with a set of observations associated with the environment.

[0170] According to some embodiments, the system includes a sensor and one or more processors configured to cause the system to perform a method comprising: scanning an environment, the scanning step of using the sensor to detect a plurality of points in the environment; identifying a first plane of the environment, the first plane comprising at least a threshold number of points of the plurality of detected points; obtaining a plurality of observations associated with the environment, the plurality of observations comprising a first subset of observations and a second subset of observations; determining whether the first plane is coplanar with a second plane of the environment; performing plane bundle adjustment based on the first subset of observations and further on the second subset of observations in accordance with the determination that the first plane is coplanar with the second plane; and performing plane bundle adjustment based on the first subset of observations and not on the second subset of observations in accordance with the determination that the first plane is not coplanar with the second plane.

[0171] According to some embodiments, the sensor comprises a LiDAR sensor.

[0172] According to some embodiments, the normal of the first plane comprises a vector in the direction toward the observation center associated with the sensor.

[0173] According to some embodiments, the step of performing planar bundle adjustments on a first subset of observations further includes the step of calculating a cost function associated with keyframe poses in a sliding window associated with the first subset.

[0174] According to some embodiments, the method further includes the steps of: determining the distance between a portion of the scanned environment and a first keyframe; inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe exceeds a threshold distance; and refraining from inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe does not exceed a threshold distance.

[0175] According to some embodiments, the method further includes the steps of determining the percentage of points being tracked, wherein a first plane includes the points being tracked; inserting a keyframe according to the determination that the percentage of points being tracked is below a threshold percentage; and refraining from inserting a keyframe according to the determination that the percentage of points being tracked is not below a threshold percentage.

[0176] According to some embodiments, the method further includes the step of determining an integral cost matrix, and the step of performing a planar bundle adjustment includes the step of calculating a cost function using the integral cost matrix.

[0177] According to some embodiments, the method further includes the step of updating the integral cost matrix based on the previous integral cost matrix.

[0178] According to some embodiments, the method further includes the step of estimating keyframe poses.

[0179] According to some embodiments, the method further includes the step of linearly interpolating points of keyframe poses.

[0180] According to some embodiments, the method further includes the step of updating a map of the environment based on planar bundle adjustment.

[0181] According to some embodiments, the environment comprises a mixed reality environment.

[0182] According to some embodiments, the sensor comprises a sensor for a mixed reality device.

[0183] According to some embodiments, the method further includes the step of performing the Levenberg-Marquardt method to solve the least squares function.

[0184] According to some embodiments, the method further includes the steps of inserting a first plane into a map associated with an environment, based on a determination that the first plane is not coplanar with a second plane, and refraining from inserting a first plane into a map associated with an environment, based on a determination that the first plane is coplanar with a second plane.

[0185] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of determining whether a threshold percentage of the point-plane distance is less than a threshold distance.

[0186] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of performing a geometric consistency check.

[0187] According to some embodiments, the method further includes the step of storing in memory a plurality of observations associated with the environment.

[0188] According to some embodiments, the method further includes a step of transmitting information about the environment, the information about the environment being associated with a set of observations associated with the environment.

[0189] According to some embodiments, a non-transient computer-readable storage medium stores one or more programs, and when one or more programs are executed by an electronic device having one or more processors and memory, the device includes the steps of: scanning an environment, the scanning step of which includes detecting a plurality of points of the environment using sensors; identifying a first plane of the environment, the first plane comprising at least a threshold number of points of the plurality of detected points; and obtaining a plurality of observations associated with the environment, the plurality of observations The instructions provide a method for performing a method comprising: a step of determining whether a first plane is coplanar with a second plane of the environment, a step of performing plane bundle adjustment based on the first subset of observations and further on the second subset of observations, in accordance with the determination that the first plane is coplanar with the second plane, and a step of performing plane bundle adjustment based on the first subset of observations and not on the second subset of observations, in accordance with the determination that the first plane is not coplanar with the second plane.

[0190] According to some embodiments, the sensor comprises a LiDAR sensor.

[0191] According to some embodiments, the normal of the first plane comprises a vector in the direction toward the observation center associated with the sensor.

[0192] According to some embodiments, the step of performing planar bundle adjustments on a first subset of observations further includes the step of calculating a cost function associated with keyframe poses in a sliding window associated with the first subset.

[0193] According to some embodiments, the method further includes the steps of: determining the distance between a portion of the scanned environment and a first keyframe; inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe exceeds a threshold distance; and refraining from inserting a second keyframe in accordance with the determination that the distance between the portion of the scanned environment and the last keyframe does not exceed a threshold distance.

[0194] According to some embodiments, the method further includes the steps of determining the percentage of points being tracked, wherein a first plane includes the points being tracked; inserting a keyframe according to the determination that the percentage of points being tracked is below a threshold percentage; and refraining from inserting a keyframe according to the determination that the percentage of points being tracked is not below a threshold percentage.

[0195] According to some embodiments, the method further includes the step of determining an integral cost matrix, and the step of performing a planar bundle adjustment includes the step of calculating a cost function using the integral cost matrix.

[0196] According to some embodiments, the method further includes the step of updating the integral cost matrix based on the previous integral cost matrix.

[0197] According to some embodiments, the method further includes the step of estimating keyframe poses.

[0198] According to some embodiments, the method further includes the step of linearly interpolating points of keyframe poses.

[0199] According to some embodiments, the method further includes the step of updating a map of the environment based on planar bundle adjustment.

[0200] According to some embodiments, the environment comprises a mixed reality environment.

[0201] According to some embodiments, the sensor comprises a sensor for a mixed reality device.

[0202] According to some embodiments, the method further includes the step of performing the Levenberg-Marquardt method to solve the least squares function.

[0203] According to some embodiments, the method further includes the steps of inserting a first plane into a map associated with an environment, based on a determination that the first plane is not coplanar with a second plane, and refraining from inserting a first plane into a map associated with an environment, based on a determination that the first plane is coplanar with a second plane.

[0204] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of determining whether a threshold percentage of the point-plane distance is less than a threshold distance.

[0205] According to some embodiments, the step of determining whether a first plane is coplanar with a second plane further includes the step of performing a geometric consistency check.

[0206] According to some embodiments, the method further includes the step of storing in memory a plurality of observations associated with the environment.

[0207] According to some embodiments, the method further includes a step of transmitting information about the environment, the information about the environment being associated with a set of observations associated with the environment.

[0208] While the disclosed embodiments are fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be obvious to those skilled in the art. For example, one or more elements of an implementation may be combined, removed, modified, or complemented to form a further implementation. Such changes and modifications are understood to fall within the scope of the disclosed embodiments as defined by the appended claims.

Claims

1. It is a method, Scanning the environment, wherein the scanning includes detecting multiple points in the environment using sensors. Identifying a first plane of the environment, wherein the first plane comprises multiple points from the multiple points detected in the environment, and the total number of these multiple points is at least a threshold number. The method involves obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and not based on a second subset of the observation results. Methods that include...

2. The method according to claim 1, wherein the sensor comprises a LiDAR sensor.

3. The method according to claim 1, wherein the normal of the first plane comprises a vector in the direction toward the observation center associated with the sensor.

4. The method according to claim 1, wherein performing planar bundle adjustment on a first subset of the observation results further comprises calculating a cost function associated with the orientation, including the position and orientation of the sensor in keyframes within a sliding window associated with the first subset.

5. Determining the distance between the scanned portion of the environment and the first keyframe, A second keyframe is inserted in accordance with the determination that the distance between the scanned portion of the environment and the first keyframe exceeds a threshold distance. The decision to refrain from inserting a second keyframe is made based on the determination that the distance between the scanned portion of the environment and the first keyframe does not exceed the threshold distance. The method according to claim 1, further comprising:

6. The percentage of points being tracked, wherein the first plane includes the points being tracked. A keyframe is inserted based on the determination that the percentage of the tracked points falls below a threshold percentage. The decision not to insert the keyframes in accordance with the determination that the percentage of the tracked points does not fall below the threshold percentage. The method according to claim 1, further comprising:

7. The method according to claim 1, further comprising determining an integral cost matrix and performing a planar bundle adjustment, which comprises calculating a cost function using the integral cost matrix.

8. The method according to claim 1, further comprising estimating the orientation, including the position and orientation of the sensor in a keyframe.

9. The method according to claim 1, further comprising linearly interpolating points of orientation, including the position and orientation of the sensor in the keyframe.

10. The method according to claim 1, further comprising updating the map of the environment based on the planar bundle adjustment.

11. The method according to claim 1, wherein the environment comprises a mixed reality environment.

12. The method according to claim 1, wherein the sensor comprises a sensor for a mixed reality device.

13. In accordance with the determination that the first plane is not coplanar with the second plane, the first plane is inserted into the map associated with the environment, In accordance with the determination that the first plane is coplanar with the second plane, refrain from inserting the first plane into the map associated with the environment. The method according to claim 1, further comprising:

14. The method according to claim 1, wherein determining whether the first plane is coplanar with the second plane further includes determining whether a threshold percentage of the point-plane distance is less than a threshold distance.

15. The method according to claim 1, further comprising storing the plurality of observation results associated with the environment in memory.

16. The method according to claim 1, further comprising transmitting information about the environment, wherein the information about the environment is associated with a plurality of observations associated with the environment.

17. It is a system, Sensors and, One or more processors, wherein the one or more processors in the system Scanning the environment, wherein the scanning includes detecting multiple points in the environment using the sensor, Identifying a first plane of the environment, wherein the first plane comprises multiple points from the multiple points detected in the environment, and the total number of these multiple points is at least a threshold number. The method involves obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and not based on a second subset of the observation results. One or more processors configured to perform a method including A system that includes these features.

18. The system according to claim 17, wherein the sensor comprises a LiDAR sensor.

19. A non-transient computer-readable storage medium for storing one or more programs, wherein the one or more programs comprises instructions, and when the instructions are executed by an electronic device having one or more processors and memory, the electronic device... Scanning the environment, wherein the scanning includes detecting multiple points in the environment using sensors. Identifying a first plane of the environment, wherein the first plane comprises multiple points from the multiple points detected in the environment, and the total number of these multiple points is at least a threshold number. The method involves obtaining a plurality of observation results associated with the aforementioned environment, wherein the plurality of observation results comprises a first subset of observation results and a second subset of observation results. To determine whether the first plane is coplanar with the second plane of the environment, In accordance with the determination that the first plane is coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and further based on a second subset of the observation results. In accordance with the determination that the first plane is not coplanar with the second plane, plane bundle adjustment is performed based on a first subset of the observation results, and not based on a second subset of the observation results. A non-transient computer-readable storage medium that enables the implementation of a method including the following.

Citation Information

Patent Citations

  • Forecasting of dynamic environmental parameters to optimize operation of a wireless communication system

    US20120202538A1

  • Method and System for Aligning Cameras

    US20120263448A1

  • System and method for augmented and virtual reality

    US20160100034A1

  • Systems and methods for augmented reality

    US20160259404A1

  • Method and device for real-time mapping and localization

    US20180075643A1