Gravity estimation and bundle adjustment for visual-inertial odometry
By using a visual inertial odometry system and bundle adjustment technology on a wearable head device, the problems of motion sickness and disorientation in virtual reality systems have been solved, achieving visibility of the real environment and immersive experience of the virtual environment in mixed reality, and enhancing user interactivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-16
- Publication Date
- 2026-03-27
AI Technical Summary
Existing virtual reality systems often cause motion sickness and disorientation when presenting virtual environments. They also fail to effectively utilize sensory data from the real environment and struggle to create immersive environments shared by multiple users.
By employing a visual inertial odometry system on a wearable head device, and through gravity estimation and bundle adjustment techniques, combined with inertial measurement unit and sensor data, it achieves accurate positioning and presentation of the user's location and virtual objects, maintains the visibility of the real environment, and creates a mixed reality environment by combining audio and visual information.
It enhances the user's immersion and interactivity in the virtual environment, reduces motion sickness, and utilizes sensory data from the real environment to enhance the immersive experience and multi-user interaction capabilities of the virtual environment.
Smart Images

Figure CN114830182B_ABST
Abstract
Description
[0001] Cross-referencing related applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 923,317, filed October 18, 2019, and U.S. Provisional Application No. 63 / 076,251, filed September 9, 2020, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This disclosure generally relates to systems and methods for drawing and displaying visual information, and more particularly, to systems and methods for drawing and displaying visual information in mixed reality environments. Background Technology
[0004] Virtual environments are ubiquitous in computing environments, used in video games (where a virtual environment represents a game world); maps (where a virtual environment represents terrain to navigate); simulations (where a virtual environment simulates a real-world environment); digital narratives (where virtual characters interact with each other within a virtual environment); and many other applications. Modern computer users are generally comfortable perceiving and interacting with virtual environments. However, the user experience with virtual environments can be limited by the technologies used to render them. For example, conventional displays (e.g., 2D displays) and audio systems (e.g., fixed speakers) cannot create virtual environments in a way that produces engaging, realistic, and immersive experiences.
[0005] Virtual reality (“VR”), augmented reality (“AR”), mixed reality (“MR”), and related technologies (collectively, “XR”) share the ability to present sensory information to a user of an XR system, corresponding to a virtual environment represented by data in a computer system. This disclosure considers the distinctions between VR, AR, and MR systems (although some systems may be classified as VR in one aspect (e.g., visual aspect) and simultaneously as AR or MR in another aspect (e.g., audio aspect)). As used herein, a VR system presents a virtual environment that replaces the user’s real environment in at least one aspect; for example, a VR system may present the user with a view of the virtual environment while blurring his or her view of the real environment, such as by using a light-blocking head-mounted display. Similarly, a VR system may present the user with audio corresponding to the virtual environment while blocking (attenuating) audio from the real environment.
[0006] VR systems can suffer from various drawbacks that come with replacing a user's real-world environment with a virtual environment. One drawback is the potential for motion sickness when a user's field of view in the virtual environment no longer corresponds to the state of his or her inner ear (which detects one's balance and orientation in the real-world environment, not the virtual environment). Similarly, users experience disorientation in VR environments where they cannot directly see their body and limbs (the view that users rely on to feel "grounded" in the real-world environment). Another drawback is the computational load (e.g., storage, processing power) on VR systems that must render a full 3D virtual environment, particularly in real-time applications that seek to immerse users in the virtual environment. Similarly, such environments need to reach very high standards of realism to be considered immersive, as users tend to be sensitive to even minor imperfections in the virtual environment - any imperfection can break the user's immersion in the virtual environment. Moreover, another drawback of VR systems is that these applications of the system cannot take advantage of the wide range of sensory data in the real-world environment, such as the various sights and sounds that people experience in the real world. A related drawback is that VR systems have difficulty creating shared environments that multiple users can interact in, as users that share a physical space in the real-world environment can not be able to directly see or interact with each other in the virtual environment.
[0007] As used herein, AR systems present a virtual environment that overlaps or overlays a real-world environment in at least one respect. For example, an AR system can present a view of a virtual environment overlaid on a user's view of the real-world environment, such as using a see-through head-mounted display that presents display images while allowing light to pass through the display into the user's eyes. Similarly, an AR system can present audio corresponding to a virtual environment while mixing in audio from the real-world environment. Similarly, as used herein, MR systems, like AR systems, present a virtual environment that overlaps or overlays a real-world environment in at least one respect, and can additionally allow that virtual environment in the MR system can interact with the real-world environment in at least one respect. For example, a virtual character in the virtual environment can flip a light switch in the real-world environment, causing a corresponding light bulb in the real-world environment to turn on or off. As another example, a virtual character can react to an audio signal in the real-world environment (such as with a facial expression). By maintaining a presentation of the real-world environment, AR and MR systems can avoid some of the above-described drawbacks of VR systems; for example, users' motion sickness is mitigated because visual cues from the real-world environment (including the user's own body) can remain visible, and such systems do not need to present a fully realized 3D environment for users to be immersed in. Moreover, AR and MR systems can take advantage of real-world sensory inputs (e.g., views of scenery, objects, and other users, and sounds) to create new applications that augment that input.
[0008] Presenting a virtual environment that overlaps or overlays a real environment can be difficult. For example, blending a virtual environment with a real environment can require a sophisticated and thorough understanding of the real environment so that objects in the virtual environment do not collide with objects in the real environment. There is also a need to maintain persistence in the virtual environment that corresponds to consistency in the real environment. For example, a virtual object displayed on a physical table needs to appear in the same location even if the user's line of sight moves away from the physical table, walks around the physical table, and then looks back at the physical table. To achieve this level of immersion, it is advantageous to develop accurate and precise estimates of the location of objects in the real world and the location of the user in the real world. SUMMARY
[0009] Examples of the present disclosure describe systems and methods for presenting virtual content on a wearable head device. For example, systems and methods for performing a visual-inertial odometry with gravity estimation and bundle adjustment are disclosed. In some embodiments, first sensor data indicative of a first feature in a first position is received via a sensor of a wearable head device. Second sensor data indicative of the first feature in a second position is received via the sensor. Inertial measurement values are received via an inertial measurement unit on the wearable head device. Based on the inertial measurement values, a velocity is determined. Based on the first position and the velocity, a third position of the first feature is estimated. Based on the third position and the second position, a re-projection error is determined. A weight associated with the re-projection error is reduced. A state of the wearable head device is determined. Determining the state includes minimizing a total error, and the total error is based on the reduced weight associated with the re-projection error. A view reflecting the determined state of the wearable head device is presented via a display of the wearable head device.
[0010] In some embodiments, the wearable head device receives image data via a sensor of the wearable head device. The wearable head device receives first inertial data and second inertial data via a first inertial measurement unit (IMU) and a second IMU, respectively. The wearable head device computes first and second pre-integration terms based on the image data and the inertial data. The wearable head device estimates a position of the device based on the first and second pre-integration terms. Based on the position of the device, the wearable head device presents the virtual content. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figures 1A-1C An example mixed reality environment is shown in accordance with one or more embodiments of the present disclosure.
[0012] Figures 2A-2DComponents of an example mixed reality system are shown in accordance with one or more embodiments of the present disclosure.
[0013] Figure 3A An example mixed reality handheld controller is shown in accordance with one or more embodiments of the present disclosure.
[0014] Figure 3B An example auxiliary unit is shown in accordance with one or more embodiments of the present disclosure.
[0015] Figure 4 An example functional block diagram of an example mixed reality system is shown in accordance with one or more embodiments of the present disclosure.
[0016] Figure 5 An example pipeline of a visual-inertial odometry is shown in accordance with one or more embodiments of the present disclosure.
[0017] Figures 6A-6B An example graph of a visual-inertial odometry is shown in accordance with one or more embodiments of the present disclosure.
[0018] Figures 7A-7B An example decision process for performing bundle adjustment is shown in accordance with one or more embodiments of the present disclosure.
[0019] Figure 8 An example graph of independent gravity estimation is shown in accordance with one or more embodiments of the present disclosure.
[0020] Figure 9 An example IMU configuration is shown in accordance with one or more embodiments of the present disclosure.
[0021] Figure 10 An example IMU configuration is shown in accordance with one or more embodiments of the present disclosure.
[0022] Figure 11 An example graphical representation of SLAM computation is shown in accordance with one or more embodiments of the present disclosure.
[0023] Figure 12 An example process for presenting virtual content is shown in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0024] In the following description of examples, reference is made to the accompanying drawings which form a part hereof, and in which is shown by way of illustration various examples in which can be practiced. It is to be understood that other examples can be used and structural changes can be made without departing from the scope of the disclosed examples.
[0025] Mixed Reality Environment
[0026] Like all people, users of mixed reality systems exist in a real environment - the three-dimensional portion of the "real world" and all that can be perceived by the user. For example, a user perceives a real environment using their ordinary human senses - sight, hearing, touch, taste, smell - and interacts with the real environment by moving their body in the real environment. A location in the real environment can be described as a coordinate in a coordinate space; for example, the coordinate can include a latitude, a longitude, and an elevation relative to sea level; distances in three orthogonal dimensions from a reference point; or other suitable values. Likewise, a vector can describe a quantity with a direction and a magnitude in the coordinate space.
[0027] A computing device may, for example, maintain a representation of a virtual environment in memory associated with the device. As used herein, a virtual environment is a computational representation of a three-dimensional space. The virtual environment can include representations of any objects, actions, signals, parameters, coordinates, vectors, or other features associated with the space. In some examples, circuitry (e.g., a processor) of the computing device can maintain and update a state of the virtual environment; that is, the processor can determine, at a first time t0, a state of the virtual environment at a second time t1 based on data associated with the virtual environment and / or input provided by a user. For example, if an object in the virtual environment is located at a first coordinate at time t0 and has certain programmed physical parameters (e.g., mass, coefficient of friction); and input received from a user indicates that a force should be applied to the object in a direction vector; the processor can apply laws of kinematics to determine a location of the object at time t1 using basic mechanics. The processor can use any suitable information known about the virtual environment and / or any suitable input to determine the state of the virtual environment at time t1. In maintaining and updating the state of the virtual environment, the processor can execute any suitable software, including software related to creating and deleting virtual objects in the virtual environment; software to define behavior of virtual objects or characters in the virtual environment (e.g., scripts); software to define behavior of signals (e.g., audio signals) in the virtual environment; software to create and update parameters associated with the virtual environment; software to generate audio signals in the virtual environment; software to handle input and output; software to implement network operations; software to apply resource data (e.g., animation data to move virtual objects over time); or many other possibilities.
[0028] Output devices, such as displays or speakers, can present any or all aspects of the virtual environment to the user. For example, the virtual environment can include virtual objects (which can include representations of inanimate objects; people; animals; lights; etc.) that can be presented to the user. The processor can determine a view of the virtual environment (e.g., corresponding to a "camera" having an origin coordinate, a view axis, and a frustum); and present a viewable scene of the virtual environment corresponding to the view to the display. Any suitable rendering technique can be used for this purpose. In some examples, the viewable scene can include only some of the virtual objects in the virtual environment, while excluding certain other virtual objects. Similarly, the virtual environment can include audio aspects that can be presented to the user as one or more audio signals. For example, virtual objects in the virtual environment can produce sound originating from a location coordinate of the object (e.g., a virtual character can speak or cause a sound effect); or the virtual environment can be associated with musical cues or environmental sounds that can or can not be associated with particular locations. The processor can determine audio signals corresponding to a "listener" coordinate— e.g., a synthesis corresponding to the sounds in the virtual environment, and mixed and processed to simulate the audio signals that would be heard by a listener located at the listener coordinate— and present the audio signals to the user via one or more speakers.
[0029] Because the virtual environment exists only as a computational construct, the user cannot directly perceive the virtual environment with ordinary senses. Instead, the user can only indirectly perceive the virtual environment, e.g., as presented to the user by a display, speakers, haptic output devices, etc. Similarly, the user cannot directly touch, manipulate, or otherwise interact with the virtual environment; but can provide input data to the processor via input devices or sensors, which the processor can use to update the virtual environment. For example, a camera sensor can provide optical data indicating that the user is attempting to move an object in the virtual environment, and the processor can use that data to cause the object to make a corresponding response in the virtual environment.
[0030] A mixed reality system, for example, can present a mixed reality environment ("MRE") to the user that combines aspects of a real environment and a virtual environment, using a transmissive display and / or one or more speakers (e.g., which can be incorporated into a wearable head device). In some embodiments, the one or more speakers can be located external to the head-mounted wearable unit. As used herein, an MRE is a simultaneous representation of a real environment and a corresponding virtual environment. In some examples, the corresponding real and virtual environments share a coordinate space; in some examples, the real coordinate space and the corresponding virtual coordinate space are related to one another by a transformation matrix (or other suitable representation). Thus, a single coordinate (in some examples, along with a transformation matrix) can define a first location in the real environment, as well as a corresponding second location in the virtual environment; and vice versa.
[0031] In an MRE, virtual objects (e.g., in a virtual environment associated with the MRE) can correspond to real objects (e.g., in a real environment associated with the MRE). For example, if a real environment of an MRE includes a real lamppost (a real object) at a location coordinate, a virtual environment of the MRE can include a virtual lamppost (a virtual object) at a corresponding location coordinate. As used herein, a real object and its corresponding virtual object, taken together, constitute a “mixed reality object.” A virtual object need not perfectly match or align with a corresponding real object. In some examples, a virtual object can be a simplified version of a corresponding real object. For example, if a real environment includes a real lamppost, a corresponding virtual object can include a cylinder (reflecting that the shape of the lamppost can be approximately cylindrical) having approximately the same height and radius as the real lamppost. Simplifying virtual objects in this way can allow for computational efficiencies, and can simplify computations performed on such virtual objects. Moreover, in some examples of an MRE, not all real objects in a real environment can be associated with a corresponding virtual object. Likewise, in some examples of an MRE, not all virtual objects in a virtual environment can be associated with a corresponding real object. That is, some virtual objects can exist only in a virtual environment of an MRE, without any real-world counterpart.
[0032] In some examples, a virtual object can have characteristics that differ, sometimes substantially, from characteristics of a corresponding real object. For example, when a real environment in an MRE can include a green two-armed cactus (an inanimate object with spines), a corresponding virtual object in the MRE includes characteristics of a green two-armed virtual character with human facial features and a gruff demeanor. In this example, the virtual object is similar to its corresponding real object in some characteristics (color, number of arms); but differs from the real object in other characteristics (human facial features, personality). In this way, a virtual object can have the potential to represent a real object in a creative, abstract, exaggerated, or fanciful way; or to attribute behaviors (e.g., human personality) to other inanimate real objects. In some examples, a virtual object can be a purely fanciful creation without a real-world counterpart (e.g., a virtual monster in a virtual environment, possibly at a location corresponding to empty space in a real environment).
[0033] In contrast to VR systems that present virtual environments to users while simultaneously occluding the real environment, mixed reality systems that present MREs provide the advantage of keeping the real environment perceptible while presenting virtual environments. Thus, users of mixed reality systems can use visual and auditory cues associated with the real environment to experience and interact with corresponding virtual environments. For example, while users of VR systems can have some difficulty perceiving or interacting with virtual objects displayed in virtual environments—because, as described herein, users cannot directly perceive or interact with virtual environments—users of MR systems can find it intuitive and natural to interact with virtual objects by seeing, hearing, and touching corresponding real objects in his or her own real environment. This level of interactivity can enhance users' sense of immersion, connection, and engagement with virtual environments. Similarly, by simultaneously presenting real and virtual environments, mixed reality systems can reduce negative psychological feelings (e.g., cognitive dissonance) and negative physical feelings (e.g., motion sickness) associated with VR systems. Mixed reality systems further provide many possibilities for applications that can enhance or change our experience of the real world.
[0034] Figure 1A An example real environment 100 is shown in which a user 110 uses a mixed reality system 112. The mixed reality system 112 can include a display (e.g., a transmissive display) and one or more speakers and one or more sensors (e.g., cameras), for example as described herein. The real environment 100 shown includes a rectangular room 104A in which the user 110 stands; and real objects 122A (a lamp), 124A (a table), 126A (a sofa), and 128A (a painting). The room 104A also includes a location coordinate 106, which can be considered the origin of the real environment 100. As described herein, the mixed reality system 112 can present a virtual environment 102 to the user 110 while simultaneously presenting the real environment 100. The virtual environment 102 includes a virtual object 120A (a virtual lamp), which is positioned at the same location in the virtual environment 102 as the real object 122A in the real environment 100. The virtual object 120A is shown as being in the same position as the real object 122A, but it is understood that the virtual object 120A can be in a different position in the virtual environment 102 than the real object 122A in the real environment 100. For example, the virtual object 120A can be in a different position in the virtual environment 102 than the real object 122A in the real environment 100, but the user 110 can perceive the virtual object 120A as being in the same position as the real object 122A in the real environment 100 because the mixed reality system 112 is configured to present the virtual object 120A in the same position as the real object 122A in the real environment 100. Figure 1AThe environment / world coordinate system 108 (including x-axis 108X, y-axis 108Y, and z-axis 108Z) shown, with origin 106 (world coordinates) can define a coordinate space for the real environment 100. In some embodiments, the origin 106 of the environment / world coordinate system 108 can correspond to the location at which the mixed reality system 112 is opened. In some embodiments, the origin 106 of the environment / world coordinate system 108 can be reset during operation. In some examples, the user / listener / head coordinate system 114 (including x-axis 114X, y-axis 114Y, z-axis 114Z) with origin 115 (e.g., user / listener / head coordinates) can define a coordinate space for the user / listener / head in which the mixed reality system 112 is located. The origin 115 of the user / listener / head coordinate system 114 can be defined relative to one or more components of the mixed reality system 112. For example, the origin 115 of the user / listener / head coordinate system 114 can be defined relative to a display of the mixed reality system 112, such as during initial calibration of the mixed reality system 112. A matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize a transformation between the user / listener / head coordinate system 114 space and the environment / world coordinate system 108 space. In some embodiments, the left ear coordinate 116 and the right ear coordinate 117 can be defined relative to the origin 115 of the user / listener / head coordinate system 114. A matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize a transformation between the left ear coordinate 116, the right ear coordinate 117, and the user / listener / head coordinate system 114 space. The user / listener / head coordinate system 114 can simplify the representation of position relative to the user’s head or wearable head device (e.g., relative to the environment / world coordinate system 108). The transformation between the user coordinate system 114 and the environment coordinate system 108 can be determined and updated in real-time through the use of simultaneous localization and mapping (SLAM), visual odometry, or other techniques.
[0035] Figure 1BAn exemplary virtual environment 130 corresponding to the real environment 100 is shown. The virtual environment 130 shown includes a virtual rectangular room 104B corresponding to the real rectangular room 104A, a virtual object 122B corresponding to the real object 122A, a virtual object 124B corresponding to the real object 124A, and a virtual object 126B corresponding to the real object 126A. Metadata associated with the virtual objects 122B, 124B, 126B can include information derived from the corresponding real objects 122A, 124A, 126A. The virtual environment 130 additionally includes a virtual monster 132 that does not correspond to any real object in the real environment 100. The real object 128A in the real environment 100 does not correspond to any virtual object in the virtual environment 130. A permanent coordinate system 133 (including an x-axis 133X, a y-axis 133Y, and a z-axis 133Z) with an origin at a point 134 (permanent coordinate) can define a coordinate space for virtual content. The origin 134 of the permanent coordinate system 133 can be defined relative to / about one or more real objects, such as the real object 126A. A matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize a transformation between the permanent coordinate system 133 space and the environment / world coordinate system 108 space. In some embodiments, each of the virtual objects 122B, 124B, 126B, and 132 can have their own permanent coordinate point relative to the origin 134 of the permanent coordinate system 133. In some embodiments, there can be multiple permanent coordinate systems and each of the virtual objects 122B, 124B, 126B, and 132 can have their own permanent coordinate point relative to one or more permanent coordinate systems.
[0036] With respect to Figure 1A and Figure 1B The environment / world coordinate system 108 defines a shared coordinate space for the real environment 100 and the virtual environment 130. In the example shown, the origin of the coordinate space is at the point 106. Further, the coordinate space is defined by the same three orthogonal axes (108X, 108Y, 108Z). Thus, a first location in the real environment 100 and a corresponding second location in the virtual environment 130 can be described with respect to the same coordinate space. This simplifies the process of identifying and displaying corresponding locations in the real and virtual environments, as the same coordinates can be used to identify both locations. However, in some examples, the corresponding real and virtual environments need not use a shared coordinate space. For example, in some examples (not shown), a matrix (which can include a translation matrix and a quaternion matrix or other rotation matrix) or other suitable representation can characterize a transformation between the real environment coordinate space and the virtual environment coordinate space.
[0037] Figure 1CAn exemplary MRE 150 is shown that presents aspects of the real environment 100 and the virtual environment 130 to the user 110 simultaneously via the mixed reality system 112. In the illustrated example, the MRE 150 presents the real objects 122A, 124A, 126A, and 128A from the real environment 100 to the user 110 simultaneously (e.g., via a transmissive portion of a display of the mixed reality system 112); and the virtual objects 122B, 124B, 126B, and 132 from the virtual environment 130 to the user 110 simultaneously (e.g., via an active display portion of a display of the mixed reality system 112). As described herein, the origin 106 can serve as an origin of a coordinate space corresponding to the MRE 150, and the coordinate system 108 defines the x-axis, y-axis, and z-axis of the coordinate space.
[0038] In the illustrated example, the mixed reality objects include corresponding real and virtual object pairs (i.e., 122A / 122B, 124A / 124B, 126A / 126B) that occupy corresponding locations in the coordinate space 108. In some examples, the real and virtual objects can be visible to the user 110 simultaneously. This can be desirable, for example, in cases where the virtual object presentation is designed to augment the view of the corresponding real object (e.g., in a museum application, a virtual object presents missing portions of an ancient, damaged sculpture). In some examples, a virtual object (122B, 124B, and / or 126B) can be displayed (e.g., via active pixilated occlusion using a pixilated occlusion shutter) to occlude the corresponding real object (122A, 124A, and / or 126A). This can be desirable, for example, in cases where the virtual object serves as a visual proxy for the corresponding real object (e.g., in an interactive storytelling application where inanimate real objects become "living" characters).
[0039] In some examples, a real object (e.g., 122A, 124A, 126A) can be associated with virtual content or ancillary data that can not necessarily constitute a virtual object. The virtual content or ancillary data can facilitate processing or disposition of the virtual object in the mixed reality environment. For example, such virtual content can include a two-dimensional representation of the corresponding real object; a custom resource type associated with the corresponding real object; or statistical data associated with the corresponding real object. This information can enable or facilitate computations involving the real object with the creation of unnecessary computational overhead.
[0040] In some examples, the presentations described herein can also incorporate audio aspects. For example, in the MRE 150, the virtual monster 132 can be associated with one or more audio signals, such as a footstep sound effect generated as the monster walks around within the MRE 150. As described herein, the processor of the mixed reality system 112 can compute an audio signal corresponding to a mixed synthesis and processed synthesis of all such sounds in the MRE 150 and present the audio signal to the user 110 via one or more speakers included in the mixed reality system 112 and / or one or more external speakers.
[0041] Example Mixed Reality System
[0042] An example mixed reality system 112 can include a wearable head device (e.g., a wearable augmented reality or mixed reality head device) that includes a display (which can include left and right transmissive displays that can be near-eye displays, and associated components for coupling light from the display to the user’s eyes), left and right speakers (e.g., positioned near the user’s left and right ears, respectively), an inertial measurement unit (IMU) (e.g., mounted on the temples of the head device), a quadrature coil electromagnetic receiver (e.g., mounted on the left temple), left and right cameras (e.g., depth (time-of-flight) cameras) that are distal to the user, and left and right eye cameras (e.g., for detecting the user’s eye motion) that are oriented toward the user. However, the mixed reality system 112 can incorporate any suitable display technology and any suitable sensors (e.g., optical, infrared, acoustic, LIDAR, EOG, GPS, magnetic sensors). In addition, the mixed reality system 112 can incorporate network features (e.g., Wi-Fi capabilities) to communicate with other devices and systems, including other mixed reality systems. The mixed reality system 112 can also include a battery (which can be mounted in a secondary unit, such as a belt pack designed to be worn around the user’s waist), a processor, and a memory. The wearable head device of the mixed reality system 112 can include tracking components, such as an IMU or other suitable sensors, configured to output a coordinate set of the wearable head device relative to the user’s environment. In some examples, the tracking components can provide input to a processor that performs simultaneous localization and mapping (SLAM) and / or visual odometry calculations. In some examples, the mixed reality system 112 can also include a handheld controller 300 and / or a secondary unit 320, which can be a wearable belt pack, as described further herein.
[0043] Figures 2A-2D Components of an example mixed reality system 200 (which can correspond to the mixed reality system 112) that can be used to present an MRE (which can correspond to the MRE 150) or other virtual environment to a user are shown. Figure 2AA perspective view of a wearable head device 2102 included in the example mixed reality system 200 is shown. Figure 2B A top view of the wearable head device 2102 worn on a user’s head 2202 is shown. Figure 2C A front view of the wearable head device 2102 is shown. Figure 2D An edge view of an example eyepiece 2110 of the wearable head device 2102 is shown. As Figures 2A-2C As shown, the example wearable head device 2102 includes an example left eyepiece (e.g., a left transparent waveguide set eyepiece) 2108 and an example right eyepiece (e.g., a right transparent waveguide set eyepiece) 2110. Each eyepiece 2108 and 2110 can include transmissive elements for viewing a real environment, and display elements for presenting a display (e.g., via imagewise modulated light) that overlaps the real environment. In some examples, such display elements can include surface diffractive optical elements for controlling a stream of imagewise modulated light. For example, the left eyepiece 2108 can include a left in-coupling grating set 2112, a left orthogonal pupil expansion (OPE) grating set 2120, and a left exit (output) pupil expansion (EPE) grating set 2122. Similarly, the right eyepiece 2110 can include a right in-coupling grating set 2118, a right OPE grating set 2114, and a right EPE grating set 2116. Imagewise modulated light can be transmitted to a user’s eye via the in-coupling gratings 2112 and 2118, the OPEs 2114 and 2120, and the EPEs 2116 and 2122. Each in-coupling grating set 2112, 2118 can be configured to deflect light toward its corresponding OPE grating set 2120, 2114. Each OPE grating set 2120, 2114 can be designed to progressively deflect light downward toward its associated EPE 2122, 2116, thereby horizontally expanding an exit pupil being formed. Each EPE 2122, 2116 can be configured to progressively redirect at least a portion of the light received from its corresponding OPE grating set 2120, 2114 to a user eyebox location (not shown) defined behind the eyepiece 2108, 2110, vertically expanding the exit pupil being formed at the eyebox. Alternatively, instead of the in-coupling grating sets 2112 and 2118, the OPE grating sets 2114 and 2120, and the EPE grating sets 2116 and 2122, the eyepieces 2108 and 2110 can include other arrangements of gratings and / or refractive and reflective features for controlling the coupling of imagewise modulated light to a user’s eye.
[0044] In some examples, the wearable head device 2102 can include a left temple piece 2130 and a right temple piece 2132, with the left temple piece 2130 including a left speaker 2134 and the right temple piece 2132 including a right speaker 2136. A coil electromagnetic receiver 2138 can be located in the left temple piece, or in another suitable location of the wearable head unit 2102. An inertial measurement unit (IMU) 2140 can be located in the right temple piece 2132, or in another suitable location of the wearable head device 2102. The wearable head device 2102 can also include a left depth (e.g., time-of-flight) camera 2142 and a right depth camera 2144. The depth cameras 2142, 2144 can be oriented appropriately in different directions so as to collectively cover a wider field of view.
[0045] In Figures 2A-2D In the example shown, the left imaging modulated light source 2124 can be optically coupled into the left eyepiece 2108 through the left incoupling grating set 2112, and the right imaging modulated light source 2126 can be optically coupled into the right eyepiece 2110 through the right incoupling grating set 2118. The imaging modulated light sources 2124, 2126 can include, for example, a fiber scanner; a projector including an electronic light modulator such as a digital light processing (DLP) chip or a liquid crystal on silicon (LCoS) modulator; or an emissive display such as a micro light emitting diode (pLED) or micro organic light emitting diode (pOLED) panel coupled to the incoupling grating set 2112, 2118 using one or more lenses on each side. The incoupling grating set 2112, 2118 can deflect light from the imaging modulated light sources 2124, 2126 to angles above the total internal reflection (TIR) critical angle of the eyepieces 2108, 2110. The OPE grating set 2114, 2120 gradually deflects light propagating by TIR downward toward the EPE grating set 2116, 2122. The EPE grating set 2116, 2122 gradually couples light toward the user’s face, including the pupils of the user’s eyes.
[0046] In some examples, as Figure 2D As shown, each of the left eyepiece 2108 and the right eyepiece 2110 includes a plurality of waveguides 2402. For example, each eyepiece 2108, 2110 can include a plurality of individual waveguides, each individual waveguide dedicated to a respective color channel (e.g., red, blue, and green). In some examples, each eyepiece 2108, 2110 can include a plurality of such sets of waveguides, each set of waveguides configured to impart a different wavefront curvature to emitted light. The wavefront curvature can be convex with respect to the user’s eye, for example to present a virtual object located a distance in front of the user (e.g., a distance corresponding to the inverse of the wavefront curvature). In some examples, the EPE grating set 2116, 2122 can include curved grating grooves to achieve convex wavefront curvature by changing the Poynting vector of the exiting light through each EPE.
[0047] In some examples, to create the perception that the display content is three-dimensional content, stereoscopically adjusted left-eye and right-eye imagery can be presented to the user through the imaging light modulators 2124, 2126 and the eyepieces 2108, 2110. The presentation of three-dimensional virtual objects can be enhanced to appear more realistic by selecting waveguides (and thus corresponding wavefront curvatures) such that virtual objects are displayed at distances approximating the distances indicated by the stereoscopic left and right images. This technique can also reduce the motion sickness experienced by some users that is caused by discrepancies between depth perception cues provided by stereoscopic left and right eye imagery and the autonomic adjustments of the human eye (e.g., focus depending on object distance).
[0048] Figure 2D An edge view from the top of the right eyepiece 2110 of the example wearable head device 2102 is shown. As Figure 2D indicated, the plurality of waveguides 2402 can include a first subset 2404 of three waveguides and a second subset 2406 of three waveguides. The two subsets 2404, 2406 of waveguides can be distinguished by different EPE gratings that feature different grating line curvatures to impart different wavefront curvatures to the exiting light. Within each subset 2404, 2406 of waveguides, each waveguide can be used to couple a different spectral channel (e.g., one of the red, green, and blue spectral channels) to the user’s right eye 2206. (Although Figure 2D The structure of the left eyepiece 2108 is similar to that of the right eyepiece 2110, although not shown.
[0049] Figure 3AAn example handheld controller component 300 of the mixed reality system 200 is shown. In some examples, the handheld controller 300 includes a handle portion 346 and one or more buttons 350 disposed along a top surface 348. In some examples, the buttons 350 can be configured to function as optical tracking targets, e.g., in conjunction with a camera or other optical sensor that can be mounted in a head unit (e.g., the wearable head device 2102) of the mixed reality system 200, to track six degrees of freedom (6DOF) motion of the handheld controller 300. In some examples, the handheld controller 300 includes tracking components (e.g., an IMU or other suitable sensors) for detecting position or orientation, such as relative to a position or orientation of the wearable head device 2102. In some examples, such tracking components can be positioned in a handle of the handheld controller 300, and / or can be mechanically coupled to the handheld controller. The handheld controller 300 can be configured to provide one or more output signals (e.g., via the IMU) corresponding to one or more of a depressed state of a button; or a position, orientation, and / or movement of the handheld controller 300. Such output signals can be used as input to a processor of the mixed reality system 200. Such input can correspond to a position, orientation, and / or movement of the handheld controller (and, by extension, also to a position, orientation, and / or movement of a hand of a user holding the controller). Such input can also correspond to a user depressing the buttons 350.
[0050] Figure 3B An example auxiliary unit 320 of the mixed reality system 200 is shown. The auxiliary unit 320 can include a battery to provide energy to operate the system 200, and can include a processor to execute programs to operate the system 200. As shown, the example auxiliary unit 320 includes a clip 2128, such as to attach the auxiliary unit 320 to a user's belt. Other form factors are suitable for the auxiliary unit 320 and will be apparent, including form factors that do not involve mounting the unit to a user's belt. In some examples, the auxiliary unit 320 is coupled to the wearable head device 2102 by a multi-conduit optical cable that can include optical wires and optical fibers, for example. Wireless connections can also be used between the auxiliary unit 320 and the wearable head device 2102.
[0051] In some examples, the mixed reality system 200 can include one or more microphones to detect sound and provide corresponding signals to the mixed reality system. In some examples, the microphones can be attached to or integrated with the wearable head device 2102 and can be configured to detect the user's voice. In some examples, the microphones can be attached to or integrated with the handheld controller 300 and / or the auxiliary unit 320. Such microphones can be configured to detect environmental sounds, ambient noise, the voice of the user or a third party, or other sounds.
[0052] Figure 4 An example functional block diagram is shown that can correspond to an example mixed reality system, such as the mixed reality system 200 described herein (which can correspond to the mixed reality system 112 with respect to FIG. 1). As shown, the mixed reality system 200 can include a wearable head device 2102, a handheld controller 300, and an auxiliary unit 320. Figure 4As shown, example handheld controller 400B (which can correspond to handheld controller 300 ("totem")) includes a totem-to-wearable-headset six degrees of freedom (6DOF) totem subsystem 404A, and example wearable headset 400A (which can correspond to wearable headset 2012) includes a totem-to-wearable-headset 6DOF subsystem 404B. In this example, 6DOF totem subsystem 404A and 6DOF subsystem 404B cooperate to determine six coordinates (e.g., offsets in three translational directions and rotations along three axes) of handheld controller 400B relative to wearable headset 400A. The six degrees of freedom can be expressed relative to a coordinate system of wearable headset 400A. The three translational offsets can be expressed as X, Y, and Z offsets in such a coordinate system, as a translation matrix, or as some other representation. The rotational degrees of freedom can be expressed as a sequence of yaw, pitch, and roll rotations, as a rotation matrix, as a quaternion, or as some other representation. In some examples, wearable headset 400A; one or more depth cameras 444 (and / or one or more non-depth cameras) included in wearable headset 400A; and / or one or more optical targets (e.g., buttons 350 of handheld controller 400B as described herein, or dedicated optical targets included in handheld controller 400B) can be used for 6DOF tracking. In some examples, handheld controller 400B can include a camera as described herein; and wearable headset 400A can include optical targets for optical tracking in conjunction with the camera. In some examples, both wearable headset 400A and handheld controller 400B include a set of three orthogonally oriented solenoids that are used to transmit and receive three distinguishable signals wirelessly. By measuring the relative amplitudes of the three distinguishable signals received in each coil used for reception, the 6DOF of wearable headset 400A relative to handheld controller 400B can be determined. Additionally, 6DOF totem subsystem 404A can include an inertial measurement unit (IMU) that can be used to provide improved accuracy and / or more timely information about rapid movements of handheld controller 400B.
[0053] In some examples, it can be necessary to transform coordinates from a local coordinate space (e.g., a coordinate space fixed relative to wearable head device 400A) to an inertial coordinate space (e.g., a coordinate space fixed relative to the real environment), e.g., to compensate for movement of wearable head device 400A relative to coordinate system 108. For example, such a transformation can be necessary for a display of wearable head device 400A to present virtual objects at an intended position and orientation relative to the real environment, rather than at a fixed position and orientation on the display (e.g., at the same position in the lower right corner of the display), e.g., to maintain the illusion that virtual objects exist in the real environment (and do not appear to be unnaturally positioned in the real environment as wearable head device 400A shifts and rotates). In some examples, a compensatory transformation between coordinate spaces can be determined by processing imagery from depth camera 444 using a SLAM and / or visual odometry procedure to determine a transformation of wearable head device 400A relative to coordinate system 108. In Figure 4 In the illustrated example, depth camera 444 is coupled to SLAM / visual odometry block 406 and can provide imagery to block 406. The SLAM / visual odometry block 406 implementation can include a processor configured to process the imagery and determine a position and orientation of the user's head, which can then be used to identify a transformation between a head coordinate space and another coordinate space (e.g., an inertial coordinate space). Similarly, in some examples, an additional source of information about the user's head pose and position is obtained from IMU 409. Information from IMU 409 can be integrated with information from SLAM / visual odometry block 406 to provide improved accuracy and / or more timely information about rapid adjustments in the user's head pose and position.
[0054] In some examples, depth camera 444 can provide 3D imagery to gesture tracker 411, which can be implemented in a processor of wearable head device 400A. Gesture tracker 411 can identify a user's gestures, e.g., by matching 3D imagery received from depth camera 444 to stored patterns representing gestures. Other suitable techniques for identifying a user's gestures will be apparent.
[0055] In some examples, one or more processors 416 can be configured to receive data from a 6DOF headset system 404B of a wearable head device, an IMU 409, a SLAM / visual odometry block 406, a depth camera 444, and / or a gesture tracker 411. The processors 416 can also send and receive control signals from a 6DOF totem system 404A. The processors 416 can be wirelessly coupled to the 6DOF totem system 404A, such as in examples where the handheld controller 400B is detached. The processors 416 can further be in communication with other components, including such as an audiovisual content memory 418, a graphics processing unit (GPU) 420, and / or a digital signal processor (DSP) audio spatializer 422. The DSP audio spatializer 422 can be coupled to a head-related transfer function (HRTF) memory 425. The GPU 420 can include a left channel output coupled to a left imaging modulated light source 424 and a right channel output coupled to a right imaging modulated light source 426. The GPU 420 can output stereoscopic image data to the imaging modulated light sources 424, 426, for example as described herein with respect to Figures 2A-2D The DSP audio spatializer 422 can output audio to the left speaker 412 and / or the right speaker 414. The DSP audio spatializer 422 can receive input from the processor 419 indicating a direction vector from the user to a virtual sound source (which can be moved by the user, for example via the handheld controller 320). Based on the direction vector, the DSP audio spatializer 422 can determine a corresponding HRTF (for example, by accessing an HRTF, or by interpolating between multiple HRTFs). The DSP audio spatializer 422 can then apply the determined HRTF to an audio signal, such as an audio signal corresponding to a virtual sound generated by a virtual object. This can enhance the believability and realism of the virtual sound by incorporating the user’s relative position and orientation with respect to the virtual sound in the mixed reality environment, that is, by making the presented virtual sound match the user’s expectation of how that virtual sound would sound like a real sound in a real environment.
[0056] In some examples, such as Figure 4 As shown, one or more of the processors 416, the GPU 420, the DSP audio spatializer 422, the HRTF memory 425, and the audiovisual content memory 418 can be included in a secondary unit 400C (which can correspond to the secondary unit 320 described herein). The secondary unit 400C can include a battery 427 to power its components and / or to power the wearable head device 400A or the handheld controller 400B. By including these components in a secondary unit that can be mounted to a user’s waist, the size and weight of the wearable head device 400A can be limited, which in turn can reduce fatigue of the user’s head and neck.
[0057] WhileFigure 4 Elements corresponding to various components of an example mixed reality system are presented, but various other suitable arrangements of these components will become apparent to those skilled in the art. For example, Figure 4 Elements presented in the middle as being associated with the auxiliary unit 400C can instead be associated with the wearable head device 400A or the handheld controller 400B. Moreover, some mixed reality systems can forgo the handheld controller 400B or the auxiliary unit 400C altogether. Such changes and modifications will be understood to be within the scope of the disclosed examples.
[0058] Simultaneous localization and mapping
[0059] Displaying virtual content in a mixed reality environment so that the virtual content corresponds to real content is challenging. For example, it is desirable to display a virtual object 122B in the same location as a real object 122A in Figure 1C In the middle, a virtual object 122B is displayed in the same location as a real object 122A. To do this, a number of capabilities of the mixed reality system 112 are involved. For example, the mixed reality system 112 can create a three-dimensional map of the real environment 104A and real objects within the real environment 104A (e.g., the lamp 122A). The mixed reality system 112 can also establish its location in the real environment 104A (which can correspond to the user's location in the real environment). The mixed reality system 112 can further establish its orientation in the real environment 104A (which can correspond to the user's orientation in the real environment). The mixed reality system 112 can also establish its motion relative to the real environment 104A, e.g., linear and / or angular velocity and linear and / or angular acceleration (which can correspond to the user's motion relative to the real environment). SLAM can be one way to display a virtual object 122B in the same location as a real object 122A even as the user 110 walks around the room 104A, looks away from the real object 122A, and looks back at the real object 122A.
[0060] SLAM needs to operate in an accurate but computationally efficient low-latency manner. As used herein, latency can refer to the time delay between a change in position or orientation of a mixed reality system component (e.g., rotation of a wearable head device) and a reflection of that change represented in the mixed reality system (e.g., a display angle of a field of view presented in a display of the wearable head device). Computationally inefficiency and / or high latency can negatively impact the user’s experience of using the mixed reality system 112. For example, if the user 110 looks around the room 104A, the virtual objects can jitter due to the user’s motion and / or high latency. Accuracy is important to produce an immersive mixed reality environment, otherwise virtual content that conflicts with real content can remind the user of the distinction between virtual and real content and decrease the user’s sense of immersion. Moreover, in some cases, latency can cause some users to experience motion sickness, headaches, or other negative physical experiences. In embodiments where the mixed reality system 112 is a mobile system that relies on a limited power source (e.g., a battery), computationally inefficiency can create more serious problems. Due to more accurate, computationally efficient, and / or lower latency SLAM, the systems and methods described herein can produce an improved user experience.
[0061] Visual-inertial odometry
[0062] Figure 5An exemplary pipeline of a visual-inertial odometry (“VIO”) using bundle adjustment and independent gravity estimation is shown. At step 504, sensor input from one or more sensors 502 can be processed. In some embodiments, the sensor 502 can be an IMU, and processing the IMU input at step 504 can include pre-integrating IMU measurements. Pre-integrating IMU measurements can include determining a single relative motion constraint from a series of inertial measurements obtained from an IMU. Pre-integrating IMU measurements is needed to reduce the computational complexity of integrating the entire series of inertial measurements. For example, inertial measurements collected between consecutive frames (may also be keyframes) captured in a video recording include data about the entire path taken by the IMU. Keyframes can be frames specifically selected based on time (e.g., time elapsed since the last keyframe selection), recognized features (e.g., having enough new recognized features compared to the last keyframe), or other criteria. However, in some embodiments, a VIO method only needs data about the starting point (e.g., at the first frame) and the ending point (e.g., at the second frame). In some embodiments, a VIO method only needs data about the current time point (e.g., the most recent frame) and the last state (e.g., the last frame). VIO computations can be simplified by pre-integrating inertial measurement data to produce a single relative motion constraint (e.g., from the first frame to the second frame). Pre-integrated inertial measurements can also be combined with other pre-integrated inertial measurements without repeating the pre-integration across the two sets of measurements. This combination can be useful, for example, when pre-integrating inertial measurements between keyframes, but later determining that a keyframe should no longer be used (e.g., due to the keyframe being deleted for redundancy or obsolescence, or if an optimization is first performed on a first subset of keyframes, then an optimization is performed on a different set of keyframes while reusing the first optimization result). For example, if inertial measurements between frame 1 and frame 2 have been pre-integrated, and inertial measurements between frame 2 and frame 3 have been pre-integrated, then these two sets of pre-integrated measurements can be combined without performing a new pre-integration in the case that frame 2 is deleted from the optimization. It is also contemplated that this can be performed across any number of frames.
[0063] In some embodiments, step 504 can include tracking identified features in camera frames, where the camera frames can be sensor inputs from sensor 502. For example, each image can be fed through a computer vision algorithm to identify features (e.g., corners and edges) within the image. Adjacent or near-neighbor frames can be compared to each other to determine correspondence between features across frames (e.g., a particular corner can be identified in both of two adjacent frames). In some embodiments, adjacent can refer to temporal adjacency (e.g., frames that are captured consecutively) and / or spatial adjacency (e.g., frames that capture similar features that can not have been captured consecutively). In some embodiments, a search radius can be used to search for a corresponding feature within a given radius of an identified feature. The search radius can be fixed or can be based on velocity between frames (e.g., calculated by integrating linear acceleration measured by an IMU).
[0064] At step 506, VIO computations can be run. VIO is a method for SLAM to observe, track, and localize features in an environment. VIO can include information streams from multiple types of sensors. For example, VIO can include information streams from visual sensors (e.g., one or more cameras) and inertial sensors (e.g., an IMU). In one example method of VIO, a camera mounted on a moving mixed reality system (e.g., mixed reality system 112, 200) can record / capture a plurality of images (e.g., frames in a video recording). Each image can be fed through a computer vision algorithm to identify features (e.g., corners and edges) in the image. Adjacent or near-neighbor frames can be compared to each other to determine correspondence between features across frames (e.g., a particular corner can be identified in both of two consecutive frames). In some embodiments, a three-dimensional map can be constructed (e.g., from stereo images), and the identified features can be located in the three-dimensional map.
[0065] In some embodiments, inertial information from an IMU sensor can be coupled with visual information from one or more cameras to verify and / or predict expected locations of identified features across frames. For example, IMU data collected between two captured frames can include linear acceleration and / or angular velocity. In embodiments where the IMU is coupled to the one or more cameras (e.g., both are embedded in a wearable head device), the IMU data can determine movement of the one or more cameras. This information can be used to estimate where the identified feature is seen in a captured image based on an estimated location of the identified feature in a three-dimensional map and based on movement of the one or more cameras. In some embodiments, the newly estimated location of the identified feature in the three-dimensional map can be projected onto a two-dimensional image to replicate how the one or more cameras can have captured the identified feature. This projection of the estimated location of the identified feature can then be compared to a different image (e.g., a subsequent frame). The difference between the projection of the estimated location and the observed location is a re-projection error for the identified feature.
[0066] IMU can also cause VIO optimization errors. For example, sensor noise can cause errors in the IMU data. IMU sensor output can also include bias (e.g., an offset in recorded measurements that can exist even without movement), which is related to physical characteristics of the IMU (e.g., temperature or mechanical stress). In some embodiments, errors in the IMU data can accumulate with integration of IMU data over a period of time. Reprojection errors can also be related to inaccurate estimates of the position of recognized features in the three-dimensional map. For example, if the initially hypothesized position of a recognized feature is incorrect, then the newly estimated position of the feature after movement can also be incorrect, and thus the projection of the newly estimated position onto the two-dimensional image can also be incorrect. VIO optimization is needed to minimize the total error, which can include both reprojection errors and IMU errors.
[0067] Figures 6A-6B An example graph of VIO computation is shown. The example graph shows a non-linear decomposition of functions of several variables intuitively. For example, variables (e.g., 602a and 602b) can be represented as circular nodes, and functions of the variables (e.g., 604, 606, and 610a, also referred to as factors) can be represented as square nodes. Each factor is a function of any additional variables. In the illustrated embodiment, nodes 602a and 602b can represent data associated with an image captured at t = 1. In the example embodiment, data associated with an image captured at t = 2 and an image captured at t = 3 is also shown. Node 602a can represent variable information, including bias (e.g., IMU bias) and velocity, which can be obtained by integrating acceleration measurements over a period of time. Node 602a can represent the bias of the IMU and the velocity of the camera fixed to the IMU at time t = 1 when the camera captured the image. Node 602b can represent a pose estimate at time t = 1. The pose estimate can include an estimate of the position and orientation of the camera in three-dimensional space (which can be caused by VIO). Node 604 can represent a perspective n-point (“PnP”) term at time t = 1. The PnP term can include an estimate of the position and orientation of recognized features in three-dimensional space (e.g., corners of recognized objects in the image). Node 606 can include IMU measurements captured between time t = 1 and time t = 2. The IMU measurements at node 606 can be selectively pre-integrated to reduce computational load. Node 608 can represent a gravity estimate of the real environment. The gravity estimate can include an indication of an estimated direction of gravity (e.g., a vector) based on inertial measurements from the IMU.
[0068] The nodes 610a and 610b can include a marginalization prior. In some embodiments, the marginalization prior can include marginalization information about some or all previously captured frames. In some embodiments, the marginalization prior can contain estimates of some or all previous camera states (which can include pose data), gravity estimates, and / or IMU extrinsics. In some embodiments, the marginalization prior can include an information matrix that can capture the dependence of the state on previous visual and / or inertial measurements. In some embodiments, the marginalization prior can include one or more error vectors that keep residuals at linearization points. Using a marginalization prior can be beneficial for several reasons. One advantage can be that the marginalization term can allow the optimization problem to have a fixed number of variables, even if the optimization does not have pre-determined limits. For example, VIO can run on an arbitrary number of frames, depending on how long a user can use a mixed reality system running VIO. Thus, VIO optimization can need to optimize thousands of frames, which is beyond the computational limits of a portable system. The marginalization prior can provide information about some or all previous frames in one term, enabling the optimization problem to be computed with fewer variables (e.g., three variables: the marginalization prior and two most recent frames). Another advantage can be that the marginalization prior can account for a long-term history of previous frames. For example, computational limits can require that only the most recent frames (e.g., three most recent frames) be optimized. Thus, the optimization can be subject to drift, where an error introduced in one frame can continue to propagate forward. A marginalization prior that approximates information from many previous frames is less susceptible to errors introduced by a single frame. In some embodiments, the marginalization prior 610a can approximate information from all frames captured before time t = 1.
[0069] Figure 6BAn exemplary embodiment is shown in which the frame captured at t = 1 and data associated with the frame captured at t = 1 are marginalized to a marginalization prior node 610b, which can include data about the marginalization prior 610a and data about the newly marginalized frame captured at t = 1. In some embodiments, the VIO can utilize a sliding window size, such as a sliding window size of two frames, in addition to the marginalization prior of the previous frame. As new frames are added to the VIO optimization (which can minimize error), the oldest frame can be marginalized to a marginalization prior term. Any size of sliding window can be used; larger window sizes can result in more accurate VIO estimates, while smaller window sizes can result in more efficient VIO computation. In some embodiments, the VIO optimization can fix nodes representing PnP terms (e.g., node 604) and only optimize nodes representing pose (or state) estimates, gravity estimates, inertial measurements, and / or IMU extrinsics (which can be included in node 602a). In addition to velocity estimates and bias estimates, state estimates can include pose estimates.
[0070] In some embodiments, each frame factored into the VIO state estimate can have an associated weight. For example, minimizing error can result in frames with many identified features having a greater weight in the optimization. Since error can originate from visual information (e.g., re-projection error) and / or inertial information (e.g., IMU bias), the weight of a frame with a large number of identified features can be too large, thereby reducing the relative weight of inertial measurements in the minimization process. It can be desirable to scale the weight of a frame using identified features, as such scaling can reduce the weight of IMU data by comparison, and as scaling can assign too much weight to one frame relative to other frames. One solution to scale weight suppression is to divide the weight of each identified feature by the square root of the sum of all weights of all identified features in the frame. Each suppressed re-projection error is then minimized in the VIO optimization, thereby suppressing scaling of visual information with more identified features.
[0071] In some embodiments, the weight of each identified feature can be associated with a confidence of the identified feature. For example, a re-projection error between an identified feature and its expected position can be assigned a weight based on observing a corresponding feature in a previous frame and estimating motion between frames (e.g., a lower weight can be assigned to a larger re-projection error). In another example, a re-projection error can be removed from computation as an outlier based on a distribution of measured re-projection errors in an image. For example, although any threshold including a dynamic threshold can be used, the top 10% of re-projection errors can be removed from computation as outliers.
[0072] In some embodiments, the way in which the marginalized prior is updated can need to be modified if the mixed reality system (e.g., mixed reality system 112, 200) is stationary. For example, if the mixed reality system is set down on a table but continues to run, the features and inertial measurements can continue to be observed, which can lead to overconfidence in the state estimate due to repeated measurements. Thus, if it is detected that the mixed reality system remains stationary for a threshold time, it can be desirable to stop updating the information matrix (possibly including uncertainty) in the marginalized prior. Stationary position can be determined from IMU measurements (e.g., no significant measurements outside of the range of recorded noise), visual information (e.g., frames continue to show no movement in recognized features), or other suitable methods. In some embodiments, other components of the marginalized prior (e.g., error vector, state estimate) can continue to be updated.
[0073] Bundle adjustment and gravity estimation
[0074] Referring back Figure 5 Keyrig insertion can be determined at step 508. It can be beneficial to use keyrigs for gravity estimation, bundle adjustment, and / or VIO, as the use of keyrigs can result in sparser data over longer time frames without increasing computational load. In some embodiments, a keyrig is a set of keyframes from a multi-camera system (e.g., an MR system with two or more cameras) that can have been captured at a particular time. In some embodiments, a keyrig can be a keyframe. Keyrigs can be selected based on any criteria. For example, keyrigs can be selected in the time domain based on the elapsed time between keyrigs (e.g., one frame can be selected as a keyrig every half second). In another example, keyrigs can be selected in the spatial domain based on recognized features (e.g., if a frame has sufficiently similar or different features compared to a previous keyrig, the frame can be selected as a keyrig). Keyrigs can be stored and saved to memory.
[0075] In some embodiments, dense keyrig insertion can be performed (which can be useful for gravity estimation), e.g., if the data provided by the standard keyrig insertion is too sparse. In some embodiments, threshold conditions need to be met before performing dense keyrig insertion. One threshold condition can be whether the VIO optimization is still accepting gravity input. Another threshold condition can be whether a threshold time (e.g., 0.5 seconds) has elapsed since the last keyrig was inserted. For example, due to IMU error (e.g., bias or drift), inertial measurements are only valid for a short time, so it can be necessary to limit the amount of time between keyrigs when collecting inertial measurements (and optionally pre-integrating). Another threshold condition can be whether there is enough motion between the most recent keyrig candidate and the most recent keyrig (i.e., it can be desirable to have enough motion to facilitate all variable observations). Another threshold condition can be whether the quality of the most recent keyrig candidate is high enough (e.g., if the new keyrig candidate has a high enough ratio of inlier to outlier reprojection values). Some or all of these threshold conditions can need to be met before dense keyrig insertion, and other threshold conditions can be used.
[0076] At step 510, a keyrig buffer can be maintained. The keyrig buffer can include memory configured to store inserted keyrigs. In some embodiments, the keyrigs can be stored with associated inertial measurement data (e.g., raw inertial measurements or pre-integrated inertial measurements) and / or associated timestamps (e.g., times at which the keyrigs were captured). It can also be determined whether the time interval between keyrigs is short enough to maintain the validity of pre-integrated inertial measurements. If it is determined that the time interval is not short enough, the buffer can be reset to avoid bad estimates (e.g., bad gravity estimates). The keyrigs and related data can be stored in a database 511.
[0077] At step 512, bundle adjustment can be performed. Bundle adjustment can optimize estimates of camera poses, states, feature locations, and / or other variables based on repeated observations of identified features across frames to improve accuracy and confidence of estimates (e.g., estimates of locations of identified features). Although the same identified feature can be observed multiple times, each observation can be inconsistent with other observations due to errors (e.g., IMU bias, feature detection inaccuracies, camera calibration errors, computational simplifications, etc.). Thus, multiple observations of an identified feature are needed to estimate a likely location of the identified feature in a three-dimensional map while minimizing errors in the input observations. Bundle adjustment can include optimizing frames (or keyrigs) over a sliding window (e.g., a sliding window of thirty keyrigs). This can be referred to as fixed-lag smoothing and can be more accurate than a Kalman filter that uses cumulative past errors. In some embodiments, bundle adjustment is visual-inertial bundle adjustment, where bundle adjustment is based on visual data (e.g., images from a camera) and inertial data (e.g., measurements from an IMU). Bundle adjustment can output a three-dimensional map of identified features that is more accurate than a three-dimensional map generated by VIO. Bundle adjustment can be more accurate than VIO estimates because VIO estimates can be performed on a frame-by-frame basis, while bundle adjustment can be performed on a keyrig-by-keyrig basis. In some embodiments, bundle adjustment can optimize map points (e.g., identified features) in addition to a position and / or orientation of the MR system (which can approximate a position and / or orientation of a user). In some embodiments, VIO estimates only optimize a position and / or orientation of the MR system (e.g., because VIO estimates take map points as fixed). Using keyrigs can allow input data to span longer time frames without increasing computational load, which can improve accuracy. In some embodiments, bundle adjustment is performed remotely on more powerful processors, allowing for more accurate estimates (due to more optimized frames) compared to VIO.
[0078] In some embodiments, bundle adjustment can include minimizing errors including visual errors (e.g., re-projection errors) and / or inertial errors (e.g., errors resulting from the position of features identified by motion estimation based on a previous keyrig and a next keyrig). Minimizing errors can be achieved by identifying the root mean square of errors and optimizing estimates to achieve the lowest average root mean square of all errors. It is also contemplated that other minimization methods can be used. In some embodiments, individual weights can be assigned to individual measurements (e.g., measurements are assigned weights based on a confidence in the accuracy of the measurement). For example, measurements from images captured by more than one camera can be assigned weights according to the quality of each camera (e.g., higher quality cameras can produce measurements with greater weight). In another example, measurements can be assigned weights according to the temperature of the sensor when the measurement was recorded (e.g., sensors can perform best within a specified temperature range, so measurements taken within that temperature range can be assigned greater weights). In some embodiments, depth sensors (e.g., LIDAR, time-of-flight cameras, etc.) can provide information in addition to visual measurements and inertial measurements. Depth information can also be included in bundle adjustment by also minimizing errors associated with depth information (e.g., when comparing depth information to a three-dimensional map of estimates constructed from visual information and inertial information). In some embodiments, a global bundle adjustment can be performed using the bundle adjustment output as input (e.g., instead of keyrigs) to further improve accuracy.
[0079] Figure 7AAn example decision process for performing bundle adjustment is shown. At step 702, a new keyrig can be added to a sliding window for optimization. At step 704, the time intervals between each keyrig in the sliding window can be evaluated. If none of the time intervals are above a certain static or dynamic time threshold (e.g., 0.5 seconds), then a visual-inertial bundle adjustment can be performed at step 706 using inertial measurements from database 705 and keyrigs (which can be images) from database 712. Database 705 and database 712 can be the same database or separate databases. In some embodiments, if at least one of the time intervals between consecutive keyrigs is above a certain static or dynamic time threshold, then a spatial bundle adjustment can be performed at step 710. If the time intervals are above the static or dynamic time threshold, then the inertial measurements are advantageously excluded because IMU measurements are only valid for short time intervals (e.g., due to sensor noise and / or drift). In contrast to the visual-inertial bundle adjustment, the spatial bundle adjustment selectively relies on a different set of keyrigs. For example, the sliding window can change from being fixed to a certain time or a certain number of keyrigs to being fixed to a certain amount of space (e.g., the sliding window is fixed to a certain amount of movement). At step 708, the oldest keyrig can be removed from the sliding window.
[0080] Figure 7B An example decision process for performing bundle adjustment is shown. At step 720, a new keyrig can be added to a sliding window for optimization. At step 722, the most recent keyrigs can be retrieved from a database 730 of keyrigs. For example, a number (e.g., five) of most recently captured keyrigs can be retrieved (e.g., based on associated timestamps). At step 724, it can be determined whether the time intervals between each of the retrieved keyrigs are below a static or dynamic threshold (e.g., 0.5 seconds). If each of the time intervals between the captured keyrigs is below the static or dynamic threshold, then a visual-inertial bundle adjustment can be performed at step 726, which utilizes inertial measurements in addition to visual information from the keyrigs. If at least one of the time intervals between the captured keyrigs is above the static or dynamic threshold, then at step 728, spatially proximate keyrigs can be retrieved from database 730. For example, keyrigs having a threshold number of identified features corresponding to identified features in the most recent keyrig can be retrieved. At step 728, a spatial bundle adjustment can be performed, which can use visual information from the keyrigs and does not use inertial information to perform bundle adjustment.
[0081] Other decision processes can also be used to perform bundle adjustment. For example, at step 704 and / or 724, it can be determined whether any time interval between consecutive keyrigs exceeds a static or dynamic time threshold. If it is determined that at least one time interval exceeds the threshold, the inertial measurements between the two associated keyrigs can be ignored, but these inertial measurements can still be used for all remaining keyrigs that meet the time interval threshold. In another example, if it is determined that a time interval exceeds the threshold, one or more additional keyrigs can be inserted between the two associated keyrigs, and visual-inertial bundle adjustment can be performed. In some embodiments, the keyrigs in the keyrig buffer or stored in database 511 and / or 730 can be updated with the results of the bundle adjustment.
[0082] Referring back to Figure 5 At step 514, independent gravity estimation can be performed. The independent gravity estimation performed at step 514 can be more accurate than the VIO gravity estimation determined at step 506, for example, because it can be performed on a more powerful processor or over a longer time frame, allowing for optimization over additional variables (e.g., frames or keyrigs). The independent gravity estimation can utilize keyrigs (which can include keyframes) to estimate a gravity (e.g., vector) that can be used for SLAM and / or VIO. Keyrigs are beneficial because they allow independent gravity estimation to be performed over a longer time without exceeding computational limits. For example, an independent gravity estimation over 30 seconds can be more accurate than an independent gravity estimation over 1 second, but if a video is recorded at 30 fps, a 30 second estimation would require optimization over 900 frames. Keyrigs can be obtained at a more sparse interval (e.g., twice per second), such that the gravity estimation only needs to be optimized over 60 frames. In some embodiments, dense keyrig insertion (which can be performed at step 508) can result in. Gravity estimation can also use simple frames. The gravity direction needs to be estimated to anchor the display of virtual content that corresponds to real content. For example, a poor gravity estimation can cause virtual content to appear tilted with respect to real content. Although an IMU can provide inertial measurements that can include gravity, the inertial measurements can also include all other motions (e.g., actual motions or noise), which can obscure the gravity vector.
[0083] Figure 8An example graph showing independent gravity estimation is shown. The example graph can visually show a non-linear decomposition of several variables' functions. For example, variables (e.g., 802a and 802b) can be represented as circular nodes, and variables' functions (e.g., 804, also referred to as factors) can be represented as square nodes. Each factor is a function of any additional variables. Nodes 802a and 802b can represent data associated with a keyrig captured at time i = 0. Node 802a can include data representing an IMU state, which can be an error estimate in IMU measurements. The IMU state is an output of a VIO method. Node 802b can include data representing a keyrig pose associated with a particular keyrig. The keyrig pose can be an output based on bundle adjustment. Node 804 can include an IMU term edge, which can include a pre-integrated inertial motion. Node 804 can also define an error function that relates other additional nodes to each other. Node 806 can include data representing an IMU extrinsic (which can correspond to an accuracy and orientation of the IMU relative to a mixed reality system), and node 808 can include a gravity estimate. The gravity estimate optimization can include minimizing the error function at node 804. In some embodiments, the nodes related to the keyrig pose and the IMU extrinsic (e.g., nodes 802b and 806) can be fixed, and the nodes related to the IMU state and the gravity (e.g., nodes 802a and 808) can be optimized.
[0084] Referring back to Figure 5 At step 515, results from the independent gravity estimation and from the bundle adjustment can be applied to the VIO output. For example, a difference between the VIO gravity estimate and the independent gravity estimate and / or a difference between the VIO pose estimate and the bundle adjustment pose estimate can be represented by a transformation matrix. The transformation matrix can be applied to the VIO pose estimate and / or the marginalized prior pose estimate. In some embodiments, a rotation component of the transformation matrix can be applied to the state estimated by the VIO and / or the gravity estimate of the VIO. In some embodiments, the independent gravity estimation cannot be applied to correct the VIO gravity estimate if the generated map becomes too large (i.e., if the independent gravity estimation must be applied to too many frames, it is computationally infeasible). In some embodiments, the execution location of the blocks in group 516 is different from the execution location of the blocks in group 518. For example, group 516 can execute on a wearable head device that includes a power saving processor and a display. Group 518 can execute on an additional device (e.g., a wearable hip device) that includes a more powerful processor. Certain computations (e.g., in group 516) can need to be performed in near real-time so that the user can obtain low-latency visual feedback of virtual content. More accurate and computationally intensive computations in parallel and backpropagation correction also need to be performed to maintain long-term accuracy of the virtual content.
[0085] Dual-IMU SLAM
[0086] In some embodiments, two or more IMUs can be used for SLAM calculations. Adding a second IMU can improve the accuracy of SLAM calculations, which can reduce jitter and / or drift of virtual content. Adding a second IMU in SLAM calculations can also be advantageous when information associated with map points is insufficient for more accurate SLAM calculations (e.g., low texture (e.g., a wall lacking geometry or visual features, such as a flat wall of one color), low light, low light and low texture). In some examples, using two IMUs to calculate SLAM can achieve a reduction of 14-19% in drift and a reduction of 10-20% in jitter compared to using one IMU. Drift and jitter are measured in units of arcminutes relative to actual position (e.g., actual position of objects in a mixed reality environment).
[0087] In some embodiments, the second IMU can be used for low light or low texture situations, and the second IMU can not be used for situations where the lighting and / or texture of the mixed reality environment is sufficient. In some embodiments, one or more visual metrics and information from sensors of the mixed reality system are used to determine whether the lighting and / or texture of the mixed reality environment is sufficient (e.g., determining sufficient texture from objects of the mixed reality environment; determining sufficient lighting in the mixed reality environment). For example, sensors of the MR system can be used to capture lighting and / or texture information associated with the mixed reality environment, and the captured information can be compared to lighting and / or texture thresholds to determine whether the lighting and / or texture is sufficient. If it is determined that the lighting and / or texture is insufficient, a pre-integrated term calculation can be performed using the second IMU (e.g., to achieve better accuracy, reduce potential jitter and / or drift), as disclosed herein. In some embodiments, the comparison of calculations using one IMU and calculations using two IMUs, if the difference between the calculations is within a threshold, then using one IMU for SLAM calculations in these situations can be sufficient. As an example advantage, the ability to use one IMU in situations where the lighting and / or texture of the mixed reality environment is sufficient can reduce power consumption and reduce computation time.
[0088] In some embodiments, the second IMU can take repeated measurements of the same value, increasing the confidence of the measured value. For example, a first IMU can measure a first angular velocity, while a second IMU can simultaneously measure a second angular velocity.
[0089] The first and second IMUs can be coupled to the same hardware (e.g., the deformation of the hardware is negligible; and the hardware is a frame of a wearable head device). Both the first angular velocity associated with the first IMU and the second angular velocity associated with the second IMU can be measurements of the same true angular velocity. In some embodiments, repeating the measurement of the same value can yield a more accurate estimate of the true value by canceling out noise in each individual measurement. In some embodiments, coupling two IMUs to the same hardware can facilitate SLAM computation in a dual-IMU system, as described herein. Because low latency can be very important for SLAM computation, there is a need to develop systems and methods that incorporate additional inertial information in a computationally efficient manner while preserving the accuracy gain of the additional data.
[0090] Figure 9 An exemplary IMU configuration in a mixed reality system is shown, in accordance with some embodiments. In some embodiments, a MR system 900 (which can correspond to MR system 112, 200) can be configured to include an IMU 902a and an IMU 902b. In some embodiments, the IMUs 902a and 902b are hard-coupled to the mixed reality system (e.g., attached to a frame of the mixed reality system) such that the velocities of the IMUs are coupled. For example, the relationship between the two velocities can be computed using known values such as angular velocities and vectors associated with the respective IMU positions. As described herein, a system that includes two velocity-coupled IMUs can advantageously reduce the computational complexity of the pre-integration term associated with the system while generally observing better computational accuracy than a single-IMU system, particularly in low-light and / or low-texture situations in a mixed reality environment.
[0091] In some embodiments, the MR system 900 can be configured such that the IMU 902a is as close as possible to the camera 904a, and the IMU 902b is as close as possible to the camera 904b. In some embodiments, the cameras 904a and 904b can be used for SLAM (e.g., the cameras 904a and 904b can be used for object / edge recognition and / or the visual component of VIO). It is desirable to configure the MR system 900 such that the IMUs are as close as possible to the SLAM cameras in order to accurately track the motion experienced by the SLAM cameras (e.g., by using the motion of the IMUs as a proxy for the motion of the SLAM cameras).
[0092] Figure 10 An exemplary IMU configuration in a mixed reality system is shown, in accordance with some embodiments. In some embodiments, two IMU sensors (e.g., IMU 1002 and IMU 1004) can generally provide two IMU terms, b and / or pre-integration term. For example, the components of the pre-integration term associated with the IMU 1002 can be represented by equations (1), (2), and (3), where may represent a quaternion corresponding to the IMU 1002 (which can represent a rotation and / or orientation of the MR system), where may represent an angular velocity measured by the IMU 1002 about a point W (which can exist on the MR system and / or be a center of the MR system), where may represent a linear velocity at the IMU 1002 relative to the point W, where may represent a linear acceleration at the IMU 1002 relative to the point W, where may represent a position of the IMU 1002 relative to the point W.
[0093] Equation (1):
[0094]
[0095] Equation (2):
[0096]
[0097] Equation (3):
[0098]
[0099] Similarly, Equations (4), (5), and (6) can represent components of pre-integrated terms associated with the IMU 1004.
[0100] Equation (4):
[0101]
[0102] Equation (5):
[0103]
[0104] Equation (6):
[0105]
[0106] Although t and At are used to describe the rotation, angular velocity, and angular acceleration of individual IMUs at a particular time, it should be understood that the different quantities can not be captured at exactly the same time. For example, due to hardware timing, there can be a delay (e.g., 200 ms) between data capture or sampling between two IMUs. In these cases, the MR system can synchronize the data sets between the two IMUs to account for the delay. As another example, the IMU data associated with different quantities can be sampled or captured at different times during a same clock period of the system. The period of the clock period can be dictated by timing resolution requirements of the system.
[0107] In some embodiments, equations (1)-(6) can be solved to calculate the pre-integration terms associated with the first and second IMUs. For example, the solution can use regression analysis (e.g., least squares) or other suitable methods to reduce errors associated with the solution of the system of equations.
[0108] In some embodiments, if IMU 1002 and IMU 1004 are hard coupled to each other (e.g., by hard body 1006 corresponding to MR system 900), the pre-integration term associated with IMU 1002 and the pre-integration term associated with IMU 1004 can be subject to kinematic constraints. In these cases, the variables associated with equations (1)-(6) can be reduced due to this coupling. In some embodiments, the hard coupling between IMU 1002 and IMU 1004 can allow the two pre-integration terms to be expressed in the same variables. Equation (7) can represent the relationship between the variables measured at IMU 1002 and the variables measured at IMU 1004 (due to the hard coupling), where, may represent the position relationship (e.g., vector) between IMU 1002 and IMU 1004, and ω can represent the angular velocity associated with the mixed reality system (e.g., the angular velocity associated with IMU 1002, the angular velocity associated with IMU 1004, the average angular velocity associated with IMU 1002 and IMU 1004, the noise-reduced angular velocity associated with the system, the bias-eliminated angular velocity associated with the system).
[0109] Equation (7):
[0110]
[0111] By using the kinematic relationship between IMU 1002 and IMU 1004, equations (8), (9), and (10) can replace equations (4), (5), and (6) as components of the pre-integration term for IMU 1004.
[0112] Equation (8):
[0113]
[0114] Equation (9):
[0115]
[0116] Equation (10):
[0117]
[0118] As described above, by hard-coupling IMU 1002 and IMU 1004, the relationship between the two IMUs can be derived (e.g., using equation (7)), and a set of IMU equations associated with the IMU pre-integration terms (e.g., equations (4), (5), (6)) can be advantageously simplified (e.g., simplified to equations (8), (9) and (10)) to reduce the computational complexity of the pre-integration terms associated with the two IMUs, while generally observing better computational accuracy than with a single IMU system.
[0119] Although equations (4), (5), and (6) are simplified for the first IMU, it should be understood that equations (1), (2), and (3) can also be alternatively simplified (e.g., for the second IMU) to perform similar calculations. Although the angular velocity associated with IMU 1002 is used in equations (8), (9), and (10), it should be understood that different angular velocities associated with the mixed reality system can also be used to calculate the pre-integral terms. For example, the average angular velocity between IMU 1002 and IMU 1004 can be used in equations (8), (9), and (10).
[0120] The stiffness of the IMUs coupled to the MR system (e.g., IMU 902a, IMU 902b, IMU 1002, IMU 1004) varies over time. For example, the mechanism used to attach the IMUs to the MR system may undergo plastic or inelastic deformation. These deformations affect the mathematical relationships between the two IMUs. Specifically, these deformations may affect the accuracy of equation (7), and the accuracy of the relationships between a set of equations derived from equation (7) associated with the first IMU (e.g., IMU 902a, IMU 1002) and a set of equations associated with the second IMU (e.g., IMU 902b, IMU 1004). These deformations also affect the accuracy of the pre-integral terms associated with the IMUs. In some embodiments, the position of the IMUs can be calibrated prior to SLAM calculations. For example, prior to SLAM calculations, the current position of the IMU (e.g., obtained using sensors from an MR system) can be compared with predetermined values (e.g., known IMU position, ideal IMU position, manufactured IMU position), and discrepancies (e.g., offsets) in the SLAM calculations can be resolved (e.g., eliminating discrepancies in equation (7)). By calibrating the IMU and resolving potential distortions in IMU coupling prior to SLAM calculations, the accuracy associated with pre-integration terms and SLAM calculations can be advantageously increased.
[0121] Figure 11 An exemplary graphical representation of SLAM according to some embodiments is shown. Figure 11The graphical representation in FIG. 11 can show a non-linear decomposition of a function of several variables. For example, variables (e.g., 1110, 1112, 1114, 1116, and 1122) can be represented as circular nodes, and functions of variables (e.g., 1102, 1118, and 1120, also referred to as factors) can be represented as square nodes. Each factor can be a function of any (or all) of the additional variables.
[0122] In some embodiments, state 1102 can represent a system state (e.g., a state of the MR system) and / or any variables in the system state at time t = 1. Similarly, state 1104 can represent a system state and / or any variables in the system state at time t = 2, and state 1106 can represent a system state and / or any variables in the system state at time t = 3. The system state can correspond to a keyrig captured at that time. The system state can include (and / or be defined by) the state of one or more variables within that state. For example, state 1102 can include PnP item 1108. In some embodiments, state 1102 can include pose estimate 1110. In some embodiments, state 1102 can include bias item 1112 for a first (and / or left) IMU (which can correspond to IMU 902a or IMU 1002). In some embodiments, bias item 1112 can include bias values for linear accelerometers and gyroscopes. In some embodiments, bias item 1112 can include bias values for each linear accelerometer and gyroscope corresponding to respective measurement axes. In some embodiments, state 1102 can include bias item 1114 for a second (and / or right) IMU (which can correspond to IMU 902b or IMU 1004). Bias item 1114 can include corresponding bias values similar to bias item 1112. In some embodiments, state 1102 can include velocity item 1116. Velocity item 1116 can include an angular velocity of the system. In some embodiments, if the system is substantially rigid, the angular velocity can be one of an angular velocity associated with the first IMU, an angular velocity associated with the second IMU, an average angular velocity associated with the first and second IMUs, a noise-reduced angular velocity associated with the system, and a bias-eliminated angular velocity associated with the system. In some embodiments, velocity item 1116 can include one or more linear velocity values corresponding to one or more positions in the system (e.g., a linear velocity value at each IMU in the system). In some embodiments, state 1102 can include stateless item 1122. Stateless item 1122 can include one or more variables that are not dependent on a particular state. For example, stateless item 1122 can include a gravity variable, which can include an estimated direction of gravity. In some embodiments, stateless item 1122 can include IMU extrinsics (e.g., relative positions of IMU 904a or IMU 1002 and / or 904b or 1004 within MR system 900).
[0123] In some embodiments, two pre-integration items can relate two system states to each other. For example, pre-integration item 1118 and pre-integration item 1120 can relate state 1102 and state 1104 to each other. In some embodiments, pre-integration item 1118 (which can correspond to IMU 904a or IMU 1002) and pre-integration item 1120 (which
[0124] The first pre-integration term (e.g., pre-integration term 1118) and the second pre-integration term (e.g., pre-integration term 1120) can be functions of the same set of variables (e.g., corresponding to IMU 904b or IMU 1004). The structure of this nonlinear optimization computation provides several advantages. For example, the second pre-integration term (which is not needed in the nonlinear decomposition using a single IMU) does not need to be proportionally increased in computational power because the second pre-integration term is kinematically limited to the first pre-integration term, as described herein. However, adding data from a second IMU can also improve jitter and / or drift compared to a single IMU decomposition. Figure 8
[0125] An example process for presenting virtual content is shown in accordance with some embodiments. For simplicity, the description of the process is not repeated here with respect to the description of the system. Figure 12 An example correlation step of the described method is shown in step 1202. Image data can be received (e.g., via a sensor of the MR system). The image data can include pictures and / or visual features (e.g., edges). Figures 9 to 11 In step 1204, first inertial data can be received via a first inertial measurement unit (e.g., IMU 902a, IMU 1002) and second inertial data can be received via a second inertial measurement unit (e.g., IMU 902b, IMU 1004). In some embodiments, the first inertial measurement unit and the second inertial measurement unit can be coupled together via a hardware. The hardware can be a body (e.g., composed of hard plastic, metal, etc.) that is not deformed under normal use. The first inertial data and / or the second inertial data can include one or more linear acceleration measurements (e.g., three measurements along each of three measurement axes) and / or one or more angular velocity measurements (e.g., three measurements along each of three measurement axes). In some embodiments, the first inertial data and / or the second inertial data can include linear velocity measurements and / or angular acceleration measurements.
[0126] In step 1206, a first pre-integration term and a second pre-integration term can be computed based on the image data, the first inertial data, and the second inertial data. In some embodiments, the first pre-integration term and the second pre-integration term can be computed using a graph optimization of nonlinearly related variables and functions. For example,
[0127] The factor graph shown can be used to compute the first pre-integration term (e.g., pre-integration term 1118) and the second pre-integration term (e.g., pre-integration term 1120). In some embodiments, optimizing the functions in the factor graph and / or computing the factor graph can involve fixing some variables while optimizing others. Figure 11
[0128] At step 1208, a position of the wearable head device can be estimated based on the first pre-integration term and the second pre-integration term. In some embodiments, the position can be estimated using a graphical optimization of non-linearly related variables and functions. For example, Figure 11 The illustrated factor graph can be used to estimate the position (e.g., the pose term 1110). In some embodiments, optimizing the factor graph and / or computing the functions in the factor graph can involve fixing some variables while optimizing other variables.
[0129] At step 1210, virtual content can be presented based on the position. In some embodiments, the virtual content can be presented via one or more transmissive displays of the wearable head device, as described herein. In some embodiments, the MR system can estimate a field of view of the user based on the position. If it is determined that the virtual content is in the field of view of the user, the virtual content can be presented to the user.
[0130] In some embodiments, the keyrigs are separated by time intervals. For example, as described with respect to Figure 6A 、 Figure 6B 、 Figure 8 and Figure 11 adjacent keyrigs (e.g., t = 0 and t = 1, t = 1 and t = 2, etc.; i = 0 and i = 1, i = 1 and i = 2, etc.; functions of variables 1102 and 1104, functions of variables 1104 and 1106, etc.) are separated by respective time intervals. The respective time intervals can be different (e.g., the times at which the keyrigs are generated can be determined by the system). In some examples, when the time interval between keyrigs is greater than a maximum time interval (e.g., 5 seconds), the advantages of using pre-integrations (e.g., nodes 606, nodes 804, pre-integration terms 1118, 1120) for visual-inertial bundle adjustment can decrease (e.g., accuracy decreases compared to using pre-integrations when the adjacent keyrigs are less than the maximum time interval; accuracy decreases compared to not using pre-integrations when the adjacent keyrigs are greater than the maximum time interval). Thus, when the advantages of pre-integrations decrease, it is desirable not to use pre-integrations for bundle adjustment.
[0131] In some embodiments, the system (e.g., the mixed reality system 112, the mixed reality system 200, Figure 4The mixed reality system in FIG. 1, the mixed reality system 900) determines whether the time interval between keyrigs is greater than a maximum time interval (e.g., 5 seconds). Based on this determination, the system determines whether to use pre-integration for bundle adjustment corresponding to adjacent keyrigs. As an example advantage, by forgoing the use of pre-integration when pre-integration is less accurate (e.g., when the time interval between adjacent keyrigs is greater than the maximum time interval), the accuracy of bundle adjustment is optimized (e.g., improved accuracy compared to using pre-integration when adjacent keyrigs are greater than the maximum time interval, improved accuracy compared to using pre-integration when adjacent keyrigs are greater than the maximum time interval and without adjusting weights). In some embodiments, as another example advantage, by forgoing pre-integration when pre-integration is less accurate, the accuracy of bundle adjustment can be optimized without any weight adjustment and with reduced computation time and resource requirements.
[0132] In some embodiments, in accordance with a determination that the time interval is not greater than the maximum time interval, the system uses pre-integration for bundle adjustment. For example, when the time interval is not greater than the maximum time interval, a visual-inertial bundle adjustment as described herein is performed.
[0133] In some embodiments, in accordance with a determination that the time interval is greater than the maximum time interval, the system forgoes using pre-integration for bundle adjustment. For example, when the time interval is greater than the maximum time interval, a spatial bundle adjustment as described herein is performed (e.g., instead of a visual-inertial bundle adjustment); a method including this step can be referred to as a hybrid visual-inertial bundle adjustment. As another example, when the time interval is greater than the maximum time interval, a visual-inertial bundle adjustment is performed, but without adding a corresponding pre-integration term in the graph (e.g., node 606, node 804, pre-integration terms 1118, 1120) as described with reference to Figure 6A 、 Figure 6B 、 Figure 8 and Figure 11 ; a method including this step can be referred to as a partial visual-inertial bundle adjustment.
[0134] The techniques and methods described herein with respect to FIG. 1 and the dual-IMU can also be applied in other contexts. For example, the dual-IMU data can be used in the VIO method described with respect to Figures 6A-6B . The dual-IMU data can also be used in the gravity estimation method described with respect to Figure 8 . For example, a second pre-integration term and a variable for the bias of each IMU can be added to the graph optimization in Figures 6A-6B and / or Figure 8 .
[0135] While the graph optimization (e.g., the graph optimization in Figures 6A-6B 、 Figure 8 and / or Figure 11The variable nodes (as shown in the optimization) can show variable nodes that include one or more variable values, but it is contemplated that the variable values within the variable nodes can be represented as one or more separate nodes. Although the techniques and methods described herein disclose a dual-IMU graph optimization structure, it is contemplated that similar techniques and methods can be used for other IMU configurations (such as three IMUs, four IMUs, ten IMUs, etc.). For example, the graph optimization model can include a pre-integration term corresponding to each IMU, which can be a function of a pose estimate, one or more bias terms (corresponding to each IMU), a velocity, a gravity direction, and one or more IMU extrinsics corresponding to each IMU.
[0136] While the disclosed examples have been fully described with reference to the accompanying drawings, it is to be noted that various changes and modifications will become apparent to those skilled in the art. For example, elements of one or more implementations can be combined, deleted, modified, or supplemented to form further implementations. Such changes and modifications are to be understood as being included within the scope of the disclosed examples as defined by the appended claims.
Claims
1. A method for drawing and displaying visual information, comprising: First sensor data indicating a first feature at a first location is received via a sensor on a wearable head device; The sensor receives second sensor data indicating the first feature at the second position. Inertial measurement values are received via the inertial measurement unit (IMU) on the wearable head device; Based on the inertial measurement values, the speed of the wearable head device is determined; Based on the first position and the velocity, estimate the third position of the first feature; Based on the third position and the second position, the reprojection error is determined; Reduce the weights associated with the reprojection error; Determine the state of the wearable head device, wherein determining the state includes minimizing the total error, and wherein the total error is based on a reduced weight associated with the reprojection error; and A view reflecting the determined state of the wearable head device is presented via the display of the wearable head device. The method further includes pre-integrating the inertial measurement values. Among them, receiving the data from the first sensor in the first instant, and The pre-integrated inertial measurement value is correlated with the movement of the wearable head device between the first and second times, and the method further includes: Determine whether the time interval between the first time and the second time is greater than the maximum time interval; Based on the determination that the time interval is not greater than the maximum time interval, the pre-integrated inertial measurement value is used; and If the time interval is determined to be greater than the maximum time interval, the pre-integrated inertial measurement value is discarded.
2. The method according to claim 1, wherein, Based on the number of identified features, the weights associated with the reprojection error are reduced.
3. The method according to claim 1, wherein, Reducing the weights associated with the reprojection error includes determining whether the reprojection error includes statistical outliers.
4. The method according to claim 1, wherein, The total error is also based on the error from the inertial measurement unit.
5. The method according to claim 1, further comprising: Determine whether the movement of the wearable head device is below a threshold; The information matrix is updated based on the determination that the movement of the wearable head device is not lower than the threshold. as well as If the wearable head device is determined to be below the threshold, the information matrix is not updated.
6. The method according to claim 1, further comprising: The second inertial measurement value is received via the second IMU on the wearable head device; as well as The second inertial measurement value is pre-integrated.
7. The method according to claim 1, further comprising: Based on the updated status of the wearable head device, a correction is applied to the status.
8. A system for drawing and displaying visual information, comprising: Sensors for wearable head devices; The inertial measurement unit of the wearable head device; The display of the wearable head device; One or more processors are configured to perform a method comprising: The wearable head device receives first sensor data indicating a first feature at a first location via its sensors. The sensor receives second sensor data indicating the first feature at the second position. Inertial measurement values are received via the inertial measurement unit on the wearable head device; Based on the inertial measurement values, the speed of the wearable head device is determined; Based on the first position and the velocity, estimate the third position of the first feature; Based on the third position and the second position, the reprojection error is determined; Reduce the weights associated with the reprojection error; Determining the state of the wearable head device, wherein determining the state includes: minimizing the total error, and wherein the total error is based on a reduced weight associated with the reprojection error; and The wearable head-mounted device displays a view reflecting the determined state of the wearable head-mounted device. The method further includes: pre-integrating the inertial measurement value. Among them, receiving the data from the first sensor in the first instant, and The pre-integrated inertial measurement value is correlated with the movement of the wearable head device between the first and second times, and The method further includes: Determine whether the time interval between the first time and the second time is greater than the maximum time interval; Based on the determination that the time interval is not greater than the maximum time interval, the pre-integrated inertial measurement value is used; and If the time interval is determined to be greater than the maximum time interval, the pre-integrated inertial measurement value is discarded.
9. The system according to claim 8, wherein, Based on the number of identified features, the weights associated with the reprojection error are reduced.
10. The system according to claim 8, wherein, Reducing the weights associated with the reprojection error includes determining whether the reprojection error includes statistical outliers.
11. The system according to claim 8, wherein, The total error is also based on the error from the inertial measurement unit.
12. The system according to claim 8, wherein, The method further includes: Determine whether the movement of the wearable head device is below a threshold; The information matrix is updated based on the determination that the movement of the wearable head device is not lower than the threshold; and If the wearable head device is determined to be below the threshold, the information matrix is not updated.
13. The system according to claim 8, wherein, The method further includes: The second inertial measurement value is received via a second IMU on the wearable head device; and The second inertial measurement value is pre-integrated.
14. The system according to claim 8, wherein, The method further includes: applying a correction to the state based on the updated state of the wearable head device.
15. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising: First sensor data indicating a first feature at a first location is received via a sensor on a wearable head device; The sensor receives second sensor data indicating the first feature at the second position. Inertial measurement values are received via the inertial measurement unit on the wearable head device; Based on the inertial measurement values, the speed of the wearable head device is determined; Based on the first position and the velocity, estimate the third position of the first feature; Based on the third position and the second position, the reprojection error is determined; Reduce the weights associated with the reprojection error; Determine the state of the wearable head device, wherein determining the state includes minimizing the total error, and wherein the total error is based on a reduced weight associated with the reprojection error; and A view reflecting the determined state of the wearable head device is presented via the display of the wearable head device. The method further includes pre-integrating the inertial measurement values. Among them, receiving the data from the first sensor in the first instant, and The pre-integrated inertial measurement value is correlated with the movement of the wearable head device between the first and second times, and the method further includes: Determine whether the time interval between the first time and the second time is greater than the maximum time interval; Based on the determination that the time interval is not greater than the maximum time interval, the pre-integrated inertial measurement value is used; and If the time interval is determined to be greater than the maximum time interval, the pre-integrated inertial measurement value is discarded.
16. The non-transitory computer-readable medium according to claim 15, wherein, The method further includes: pre-integrating the inertial measurement value.
Citation Information
Patent Citations
Mapping optimization in autonomous and non-autonomous platforms
EP3428760A1