Unfused pose-based drift correction of fused pose of totem in user interaction system

The user interaction system uses EM transmitters and IMUs to correct pose drift in augmented reality systems, ensuring virtual objects remain accurately positioned relative to user-held totems, addressing the challenge of sensor model mismatch and enhancing realism.

JP2025116245AActive Publication Date: 2025-08-07MAGIC LEAP INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025094021
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-03-13
Filing Date
2025-06-05
Publication Date
2025-08-07
Estimated Expiration
2039-07-26

AI Technical Summary

Technical Problem

Existing user interaction systems face challenges in maintaining the realistic pose of virtual objects relative to user-held totems due to sensor model mismatch, leading to pose drift over time.

Method used

A user interaction system that incorporates an electromagnetic (EM) transmitter and inertial measurement unit (IMU) on the totem, coupled with a head unit receiver and processor, to generate a fused attitude in the world frame, and includes a non-fused attitude determination modeler and comparator to detect and correct pose drift by resetting the IMU attitude when the difference exceeds a predetermined threshold.

Benefits of technology

The system effectively reduces pose drift, ensuring that virtual objects remain accurately positioned relative to the totem, enhancing the stability and realism of augmented reality experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025116245000001
    Figure 2025116245000001
  • Figure 2025116245000002
    Figure 2025116245000002
  • Figure 2025116245000003
    Figure 2025116245000003
Patent Text Reader

Abstract

To provide a user interaction system.SOLUTION: The invention relates generally to a user interaction system having a head unit for a user to wear and a totem that the user holds in their hand and determines the location of a virtual object that is seen by the user. A fusion routine generates a fused location of the totem in a world frame based on a combination of an EM wave and a totem IMU data. The fused pose may drift over time due to the sensor's model mismatch. An unfused pose determination modeler routinely establishes an unfused pose of the totem relative to the world frame. A drift is declared when a difference between the fused pose and the unfused pose is more than a predetermined maximum distance.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Patent Application No. 62 / 714,609, filed August 3, 2018, and U.S. Provisional Patent Application No. 62 / 818,032, filed March 13, 2019, all of which are incorporated herein by reference in their entireties.

[0002] The present invention relates to a user interaction system with totems that define six degrees of freedom ("6dof") poses, i.e., poses of virtual objects as perceived by a user. [Background technology]

[0003] Modern computing and display technologies have facilitated the development of user interaction systems, including "augmented reality" viewing devices. Such viewing devices typically have a head unit mountable on a user's head, with a head unit body that often includes two waveguides, one in front of each of the user's eyes. The waveguides are transparent so that ambient light from real-world objects can pass through the waveguides, allowing the user to see the real-world objects. Each waveguide also serves to transmit projected light from a projector to a respective eye of the user. The projected light forms an image on the retina of the eye, which thus receives the ambient light and the projected light. The user simultaneously sees the real-world object and one or more virtual objects created by the projected light.

[0004] Such user interaction systems often include a totem. A user may, for example, hold the totem in their right hand and move the totem with six degrees of freedom in three-dimensional space. A virtual object may be attached to the totem and perceived by the user as moving with the totem in three-dimensional space, or the virtual object may be the perception of a light beam hitting a wall or another object the user is moving across a wall.

[0005] It is important that the virtual object remain in its realistic pose relative to the totem. For example, if the totem represents a racket handle and the virtual object represents the racket head, the racket head needs to remain "attached" to the racket handle over time. Summary of the Invention [Means for solving the problem]

[0006] The present invention provides a user interaction system including a totem having a totem body, an electromagnetic (EM) transmitter on the totem body, and a totem inertial measurement unit (IMU) located on the totem and generating a totem IMU signal due to movement of the totem, a head unit having a head unit body and an EM receiver on the head unit body that receives EM waves transmitted by the EM transmitter, the EM waves indicating the location of the totem, a processor, a storage device connected to the processor, and a set of instructions on the storage device that are executable by the processor. The set of instructions includes a fusion routine, coupled to the world frame, the EM receiver, and the totem IMU, that generates a fused attitude of the totem in the world frame based on a combination of EM waves, head unit attitude, and totem IMU data; a non-fused attitude determination modeler that determines the attitude of the totem relative to the head unit and the attitude of the head unit relative to the world frame, and establishes a non-fused attitude of the totem relative to the world frame; and a comparator, coupled to the fused attitude determination modeler and the non-fused attitude determination modeler, that compares the fused attitude and the non-fused attitude. a drift declarer connected to the comparator and declaring drift only if the fused attitude is more than a predetermined distance from the non-fused attitude; a location correction routine connected to the drift declarer and resetting the attitude of the totem IMU to match the non-fused location only if drift is declared; a data source conveying image data; and a display system connected to the data source and using the image data to display a virtual object to a user, the location of the virtual object being based on the fused location of the totem.

[0007] The present invention also provides a method for controlling a head unit body on a user's head, the method comprising the steps of transmitting electromagnetic (EM) waves using an EM transmitter on the totem body, generating a totem inertial measurement unit (IMU) signal using a totem IMU on the totem body due to movement of the totem, locating the head unit body on the user's head, receiving the EM waves transmitted by the EM transmitter with an EM receiver on the head unit body, the EM waves indicating the attitude of the totem, storing a world frame, executing a fusion routine using a processor to generate a fused attitude of the totem in the world frame based on a combination of the EM waves, the head unit attitude, and the totem IMU data, and using a processor to determine the attitude of the totem relative to the head unit and the location of the head unit relative to the world frame, and storing the world frame. a processor for executing a non-fused attitude determination modeler to establish a non-fused attitude of the totem relative to the frame; a processor for executing a comparator to compare the fused attitude with the non-fused attitude; a processor for executing a drift declarer to declare drift only if the fused attitude is more than a predetermined attitude from the non-fused attitude; a processor for executing an attitude correction routine to reset the attitude of the totem IMU to match the non-fused attitude only if drift is declared; receiving image data from a data source; and a display system connected to the data source to display a virtual object to a user using the image data, wherein the location of the virtual object is based on the fused location of the totem. The present invention provides, for example, the following items. (Item 1) 1. A user interaction system, comprising: It is a totem, The totem itself, an electromagnetic (EM) transmitter on the totem body; a totem inertial measurement unit (IMU) located on the totem and generating a totem IMU signal due to movement of the totem; a totem having A head unit, A head unit body, an EM receiver on the head unit body for receiving EM waves transmitted by the EM transmitter, the EM waves indicating the location of the totem; a head unit having a processor; a storage device connected to the processor; A set of instructions, said set of instructions comprising: The world frame and a fusion routine coupled to the EM receiver and the totem IMU, the fusion routine generating a fused attitude of the totem in the world frame based on a combination of the EM waves, the head unit attitude, and the totem IMU data; an unfused attitude determination modeler that determines an attitude of the totem relative to the head unit and an attitude of the head unit relative to the world frame to establish an unfused attitude of the totem relative to the world frame; a comparator coupled to the fused attitude determination modeler and the non-fused attitude determination modeler, the comparator configured to compare the fused attitude with the non-fused attitude; a drift declarer coupled to the comparator for declaring a drift only if the fused attitude is more than a predetermined distance from the non-fused attitude; a location correction routine coupled to the drift declarer for resetting the totem IMU attitude to match the unfused location only if the drift is declared; a data source carrying image data; a display system connected to the data source and using the image data to display a virtual object to a user, the location of the virtual object being based on a fusion location of the totem; and a set of instructions on the storage device executable by the processor, the set of instructions including: A user interaction system comprising: (Item 2) a rendering engine connected to the data channel and having an input for receiving the image data and an output to the display system, the rendering engine providing to the display system a data stream including the virtual object positioned in the fusion pose; Item 1. The user interaction system of item 1, further comprising: (Item 3) 2. The user interaction system of claim 1, wherein the totem IMU reduces jitter of the virtual object. (Item 4) Item 1. The user interaction system of item 1, wherein the totem IMU includes at least one of a gyroscope and an accelerometer. (Item 5) Item 10. The user interaction system of item 1, wherein the fusion routine executes at a first frequency and the non-fused attitude determination modeler executes at a second frequency different from the first frequency. (Item 6) Item 1. The user interaction system of item 1, wherein the EM receiver detects six degrees of freedom ("6dof") movement of the EM transmitter relative to the EM receiver. (Item 7) Item 8. The user interaction system of item 1, wherein the predetermined distance is less than 100 mm. a camera on the head unit body, the camera positioned to capture an image of the totem; a simultaneous localization and mapping (SLAM) system connected to the camera and receiving images of the totem, the SLAM system determining the unfused pose based on the images of the totem; and Item 1. The user interaction system of item 1, further comprising: (Item 9) The display system comprises: a transparent waveguide affixed to the head unit body that allows light from the totem to pass to the eye of a user wearing the head unit body; a projector that converts the image data into light, the light from the projector entering the waveguide at an entrance pupil and exiting the waveguide to the user's eye at an exit pupil; and Item 1. The user interaction system according to item 1, comprising: (Item 10) a head unit detection device for detecting movement of the head-mountable frame; and wherein the set of instructions further comprises: a display adjustment algorithm connected to the head unit detection device, the display adjustment algorithm receiving measurements and calculating installation values based on movements detected by the head unit detection device; a rendering engine that modifies the position of the virtual object within the eye's view based on the setting value; Item 1. The user interaction system of item 1, comprising: (Item 11) The head unit detection device a head unit IMU mounted to the head-mountable frame, the head unit IMU including a motion sensor that detects movement of the head-mountable frame, the head unit IMU including at least one of a gyroscope and an accelerometer. Item 11. The user interaction system of item 10, comprising: (Item 12) The head unit detection device a head unit camera mounted to the head-mountable frame, the head unit camera detecting movement of the head-mountable frame by capturing images of objects within a view of the head unit camera, the set of instructions comprising: a simultaneous localization and mapping (SLAM) system connected to the camera, receiving an image of the object and detecting a pose position of the head unit body, the rendering engine correcting the position of the virtual object based on the pose position; Item 11. The user interaction system of item 10, comprising: (Item 13) The location correction routine storing a rig frame, the rig frame being a mathematical object located between the eyepieces of the head unit and serving as a basis for defining where objects reside relative to the head unit; taking measurements to determine a rig frame attitude that the rig frame is assuming relative to the world frame; deriving a receiver attitude relative to the world using known extrinsic properties, such as those provided by factory level calibration, of the relationship between the EM receiver and the rig frame; estimating an EM relationship between the EM receiver and the EM transmitter to derive a transmitter attitude in the world; deriving an in-world totem IMU attitude as an unfused attitude using known pertinent properties of the totem regarding the relationship between the EM transmitter and the totem IMU; Item 1. A user interaction system according to item 1, which is executable to perform a method comprising: (Item 14) 1. A user interaction system, comprising: transmitting electromagnetic (EM) waves using an EM transmitter on the totem body; generating a totem inertial measurement unit (IMU) signal due to movement of the totem, using a totem IMU on the totem body; Locating a head unit body on a user's head; receiving the EM waves transmitted by the EM transmitter with an EM receiver on the head unit body, the EM waves indicating the posture of the totem; Remembering the world frame and using a processor to execute a fusion routine to generate a fused pose of the totem in the world frame based on a combination of the EM waves, head unit pose, and the totem IMU data; executing, with the processor, an unfused attitude determination modeler to determine an attitude of the totem relative to the head unit and a location of the head unit relative to the world frame and to establish an unfused attitude of the totem relative to the world frame; using the processor to execute a comparator to compare the fused attitude to the non-fused attitude; Implementing a drift declarer with the processor to declare drift only if the fusion attitude is more than a predetermined attitude from the non-fusion attitude; using the processor to execute an attitude correction routine to reset the totem IMU attitude to match the unfused attitude only if the drift is declared; receiving image data from a data source; using the image data to display a virtual object to a user using a display system connected to the data source, the location of the virtual object being based on the fusion pose of the totem; and A user interaction system comprising: [Brief explanation of the drawings]

[0008] The invention will be further described, by way of example only, with reference to the accompanying drawings, in which:

[0009] [Figure 1] FIG. 1 is a perspective view illustrating a user interaction system according to an embodiment of the present invention.

[0010] [Figure 2] FIG. 2 is a block diagram illustrating the components of the user interaction system as they relate to the head unit and the vision algorithms for the head unit.

[0011] [Figure 3] FIG. 3 is a block diagram of a user interaction system as it relates to a totem and the visual algorithms for the totem.

[0012] [Figure 4] FIG. 4 is a front view illustrating how real and virtual objects appear and are perceived by a user.

[0013] [Figure 5] FIG. 5 is a view similar to FIG. 4 after the virtual object has drifted within the user's view.

[0014] [Figure 6] FIG. 6 is a perspective view illustrating the drift of fusion location over time.

[0015] [Figure 7] FIG. 7 is a graph illustrating how drift can be corrected using distance calculations.

[0016] [Figure 8] FIG. 8 is a graph illustrating how drift is corrected by detecting differences between fused and non-fused locations.

[0017] [Figure 9] FIG. 9 is a perspective view illustrating how drift is corrected.

[0018] [Figure 10] FIG. 10 is a block diagram of a machine in the form of a computer that may find use in the system of the present invention, in accordance with one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] Figure 1 of the accompanying drawings illustrates a user 10, a user interaction system 12 according to an embodiment of the present invention, a real-world object 14 in the form of a table, and a virtual object 16 that is invisible from the perspective of the figure but is visible to the user 10.

[0020] The user interaction system 12 includes a head unit 18 , a belt pack 20 , a network 22 , and a server 24 .

[0021] The head unit 18 includes a head unit body 26 and a display system 28. The head unit body 26 has a shape that fits over the user's head 10. The display system 28 is fixed to the head unit body 26.

[0022] The belt pack 20 has a processor and a storage device connected to the processor. Vision algorithms are stored on the storage device and executable by the processor. The belt pack 20 is communicatively connected to a display system 28 using a cable connection 30. The belt pack 20 further includes a network interface device that allows the belt pack 20 to wirelessly connect to a network 22 via a link 32. A server 24 is connected to the network 22.

[0023] In use, the user 10 secures the head unit body 26 to their head. The display system 28 includes an optical waveguide (not shown) that is transparent so that the user 10 can see the real-world objects 14 through the waveguide.

[0024] The belt pack 20 may download image data from the server 24 via the network 22 and link 32. The belt pack 20 provides the image data to the display system 28 through a cable connection 30. The display system 28 has one or more projectors that create light based on the image data. The light propagates to the eyes of the user 10 through one or more optical waveguides. Each waveguide creates light at a specific focal length on the retina of a respective eye so that the eye sees the virtual object 16 at a distance behind the display system 28. The eye therefore sees the virtual object 16 in three-dimensional space. In addition, a slightly different image is created for each eye so that the user's 10 brain perceives the virtual object 16 in three-dimensional space. The user 10 therefore sees the real-world object 14 augmented with the virtual object 16 in three-dimensional space.

[0025] The user interaction system 12 further includes a totem 34. In use, the user 10 holds the totem 34 in one of their hands. The virtual object 16 is positioned in three-dimensional space based on the positioning of the totem 34. As an example, the totem 34 may be the handle of a racket, and the virtual object 16 may include the head of the racket. The user 10 can move the totem 34 in six degrees of freedom in three-dimensional space. The totem 34 thus moves in three-dimensional space relative to the real-world objects 14 and the head unit body 26. Various components within the head unit 18 and belt pack 20 track the movement of the totem 34 and move the virtual object 16 along with the totem 34. The head of the racket thus remains attached to the handle in the view of the user 10.

[0026] FIG. 2 illustrates the display system 28 in more detail, along with the vision algorithm 38. The vision algorithm 38 resides primarily within the belt pack 20 in FIG. 1. In other embodiments, the vision algorithm 38 may reside entirely within the head unit or may be split between the head unit and the belt pack. FIG. 2 also includes a data source 40. In this example, the data source 40 includes image data stored on a storage device of the belt pack 20. The image data may be three-dimensional image data that may be used, for example, to render the virtual object 16. In an alternative embodiment, the image data may be time-series image data that allows for the creation of moving video in two or three dimensions, for example, with a connection to a totem, located on a real-world object, or in a fixed position in front of the user as the user moves their head.

[0027] The vision algorithms 38 include a rendering engine 42 , a stereoscopic analyzer 44 , a display adjustment algorithm 46 , and a simultaneous localization and mapping (SLAM) system 48 .

[0028] The rendering engine 42 is connected to the data source 40 and to the display accommodation algorithm 46. The rendering engine 42 is capable of receiving input from various systems, in this example, the display accommodation algorithm 46, and positioning the image data within a frame to be viewed by the user 10 based on the display accommodation algorithm 46. The display accommodation algorithm 46 is connected to a SLAM system 48. The SLAM system 48 is capable of receiving the image data, analyzing the image data for purposes of determining objects within the image data, and recording the location of the objects within the image data.

[0029] A stereoscopic analyzer 44 is connected to the rendering engine 42. The stereoscopic analyzer 44 is capable of determining left and right image data sets from the data stream provided by the rendering engine 42.

[0030] The display system 28 includes left and right projectors 48A and 48B, left and right waveguides 50A and 50B, and a detection device 52. The left and right projectors 48A and 48B are connected to a power supply. Each projector 48A or 48B has an individual input through which image data is provided to the individual projector 48A or 48B. When powered, the individual projector 48A or 48B generates and emits light in a two-dimensional pattern. The left and right waveguides 50A and 50B are positioned to receive light from the left and right projectors 48A and 48B, respectively. The left and right waveguides 50A and 50B are transparent waveguides.

[0031] The detection device 52 includes a head unit inertial motion unit (IMU) 60 and one or more head unit cameras 62. The head unit IMU 60 includes one or more gyroscopes and one or more accelerometers. The gyroscopes and accelerometers are typically formed in a semiconductor chip and are capable of detecting movement of the head unit IMU 60 and the head unit body 26, including movement along and rotation about three orthogonal axes.

[0032] The head unit camera 62 continuously captures images from the environment around the head unit body 26. The images can be compared to each other to detect movement of the head unit body 26 and the user's head 10.

[0033] The SLAM system 48 is connected to a head unit camera 62. The display accommodation algorithm 46 is connected to a head unit IMU 60. Those skilled in the art will appreciate that the connection between the detection device 52 and the vision algorithm 38 is accomplished through a combination of hardware, firmware, and software. The components of the vision algorithm 38 are linked to each other through subroutines or calls.

[0034] In use, the user 10 places the head unit body 26 on their head. Components of the head unit body 26 may include, for example, a strap (not shown) that wraps around the back of the user's head 10. The left and right waveguides 50A and 50B are then positioned in front of the user's 10 left and right eyes 120A and 120B.

[0035] The rendering engine 42 receives image data from the data source 40. The rendering engine 42 inputs the image data into a stereoscopic analyzer 44. The image data is three-dimensional image data of the virtual object 16 in FIG. 1. The stereoscopic analyzer 44 analyzes the image data and determines left and right image data sets based on the image data. The left and right image data sets are data sets representing two-dimensional images that differ slightly from each other for the purpose of giving the user 10 the perception of a three-dimensional rendering. In this embodiment, the image data is a static data set that does not change over time.

[0036] The stereoscopic analyzer 44 inputs the left and right image data sets into left and right projectors 48A and 48B. The left and right projectors 48A and 48B then create left and right light patterns. While the components of the display system 28 are shown in plan view, it should be understood that the left and right patterns are two-dimensional patterns when shown in a front elevation view. Each light pattern includes multiple pixels. For illustrative purposes, light rays 124A and 126A from two of the pixels are shown exiting the left projector 48A and entering the left waveguide 50A. Light rays 124A and 126A reflect off the sides of the left waveguide 50A. Although rays 124A and 126A are shown propagating from left to right within left waveguide 50A through internal reflection, it should be understood that rays 124A and 126A also propagate in a direction out of the plane of the paper using refractive and reflective systems.

[0037] Light rays 124A and 126A exit left optical waveguide 50A through pupil 128A and then enter left eye 120A through pupil 130A of left eye 120A. Light rays 124A and 126A then impinge on retina 132A of left eye 120A. In this manner, the left light pattern impinges on retina 132A of left eye 120A. User 10 is given the perception that the pixels formed on retina 132A are pixels 134A and 136A, which user 10 perceives as being at a certain distance on the side of left waveguide 50A facing left eye 120A. Depth perception is created by manipulating the focal length of light.

[0038] Similarly, stereoscopic analyzer 44 inputs a right image data set into right projector 48B. Right projector 48B transmits a right light pattern represented by pixels in the form of light rays 124B and 126B. Light rays 124B and 126B reflect within right waveguide 50B and exit through pupil 128B. Light rays 124B and 126B then enter through pupil 130B of right eye 120B and strike retina 132B of right eye 120B. The pixels of light rays 124B and 126B are perceived as pixels 134B and 136B behind right waveguide 50B.

[0039] The patterns created on the retinas 132A and 132B are perceived as left and right images, respectively, which differ slightly from one another due to the function of the stereoscopic analyzer 44. The left and right images are perceived as a three-dimensional rendering in the mind of the user 10.

[0040] As mentioned, the left and right waveguides 50A and 50B are transparent. Light from real objects on the sides of the left and right waveguides 50A and 50B facing the eyes 120A and 120B can be projected through the left and right waveguides 50A and 50B and strike the retinas 132A and 132B. In particular, light from the real-world object 14 in FIG. 1 strikes the retinas 132A and 132B such that the real-world object 14 can be seen by the user 10. In addition, the user 10 can see the totem 34, and an augmented reality is created in which the real-world object 14 and the totem 34 are augmented with a three-dimensional rendering of the virtual object 16 perceived by the user 10 due to the left and right images perceived by the user 10 in combination.

[0041] The head unit IMU 60 detects any movement of the user 10's head. If the user 10, for example, moves their head counterclockwise and simultaneously moves their body with their head toward the right, such movement will be detected by the gyroscope and accelerometer within the head unit IMU 60. The head unit IMU 60 provides measurements from the gyroscope and accelerometer to a display adjustment algorithm 46. The display adjustment algorithm 46 calculates and provides the adjustment values to a rendering engine 42. The rendering engine 42 modifies the image data received from the data source 40 to compensate for the movement of the user 10's head. The rendering engine 42 provides the modified image data to a stereoscopic analyzer 44 for display to the user 10.

[0042] The head unit camera 62 continuously captures images as the user 10 moves their head. The SLAM system 48 analyzes the images and identifies images of objects within the images. The SLAM system 48 analyzes the object's movement and determines the pose position of the head unit body 26. The SLAM system 48 provides the pose position to the display accommodation algorithm 46. The display accommodation algorithm 46 uses the pose position to further refine the positioning values that the display accommodation algorithm 46 provides to the rendering engine 42. The rendering engine 42 therefore modifies the image data received from the data source 40 based on a combination of the motion sensors in the head unit IMU 60 and the images captured by the head unit camera 62. As a practical example, if the user 10 rotates their head to the right, the location of the virtual object 16 rotates to the left in the user's field of view, thus giving the user 10 the impression that the location of the virtual object 16 remains stationary relative to the real-world objects 14 and the totem 34.

[0043] 3 illustrates further details of the head unit 18, the totem 34, and the vision algorithm 38. The head unit 18 further includes an electromagnetic (EM) receiver 150 that is affixed to the head unit body 26. The display system 28, the head unit camera 62, and the EM receiver 150 are mounted in fixed positions relative to the head unit body 26. When the user 10 moves their head, the head unit body 26 moves with the head of the user 10, and the display system 28, the head unit camera 62, and the EM receiver 150 move with the head unit body 26.

[0044] The totem 34 includes a totem body 152, an EM transmitter 154, and a totem IMU 156. The EM transmitter 154 and the totem IMU 156 are mounted in fixed positions relative to the totem body 152. The user 10 holds the totem body 152, and as the user 10 moves the totem body 152, the EM transmitter 154 and the totem IMU 156 move with the totem body 152. The EM transmitter 154 is capable of transmitting EM waves, and the EM receiver 150 is capable of receiving the EM waves. The totem IMU 156 includes one or more gyroscopes and one or more accelerometers. The gyroscopes and accelerometers are typically formed in semiconductor chips and are capable of detecting movement of the totem IMU 156 and the totem body 152, including movement along and rotation about three orthogonal axes.

[0045] The vision algorithms 38 further include a fusion routine 160, a non-fused attitude determination modeler 162, a comparator 164, a drift declarer 166, an attitude correction routine 168, and a sequential controller 170, in addition to the data source 40, the rendering engine 42, the stereoscopic analyzer 44, and the SLAM system 48 described with reference to FIG. 2.

[0046] The head unit camera 62 captures images of the real-world object 14. The images of the real-world object 14 are processed by the SLAM system 48, as described with reference to Figure 2, to establish a world frame 172. Details of how the SLAM system 48 establishes the world frame 172 are not shown in Figure 3 so as not to obscure the drawing.

[0047] EM transmitter 154 transmits EM waves that are received by EM receiver 150. The EM waves received by EM receiver 150 indicate the attitude or change in attitude of EM transmitter 154. EM receiver 150 incorporates the EM wave data into a fusion routine 160.

[0048] The totem IMU 156 continuously monitors the movement of the totem body 152. Data from the totem IMU 156 is incorporated into a fusion routine 160.

[0049] The sequential controller 170 executes the fusion routine 160 at a frequency of 250 Hz. The fusion routine 160 combines data from the EM receiver 150 with data from the totem IMU 156 and the SLAM system 48. The EM waves received by the EM receiver 150 contain data that relatively accurately represent the pose of the EM transmitter 154 relative to the EM receiver 150 in six degrees of freedom (“6 dof”). However, due to EM measurement noise, the measured EM waves may not accurately represent the pose of the EM transmitter 154 relative to the EM receiver 150. The EM measurement noise can result in jitter of the virtual object 16 in FIG. 1. The purpose of combining the data from the totem IMU 156 is to reduce jitter. The fusion routine 160 provides a fused pose 174 in the world frame 172. The fused pose 174 is used by the rendering engine 42 to determine the pose of the virtual object 16 in FIG. 1 using image data from the data source 40.

[0050] 4, the virtual object 16 is shown in the correct pose relative to the totem 34. Furthermore, when the user 10 moves the totem 34, the virtual object 16 moves with the totem 34 with a minimal amount of jitter.

[0051] The totem IMU 156 essentially measures acceleration and angular velocity in six degrees of freedom, which are integrated to determine the location and orientation of the totem IMU 156. Due to integration errors, the fused attitude 174 can drift over time.

[0052] FIG. 5 illustrates that the virtual object 16 has drifted from its correct pose relative to the totem 34. The drift can be caused by so-called “model mismatch,” i.e., an imperfect mathematical model that describes the relationship between physical quantities (e.g., 6 dof, acceleration, and angular velocity) and actual measured signals (such as EM wave measurements and IMU signals). Such drift can also be amplified for highly dynamic motion, which can even lead to a divergence of the fusion algorithm (i.e., the virtual object will appear “blown away” from the real object). In this example, the virtual object 16 has drifted to the right relative to the totem 34. The fused pose 174 in FIG. 3 is based on the system's belief that the totem 34 is located to the right of where it actually is. The fused data has therefore been corrected so that the virtual object 16 is again positioned in its correct location relative to the totem 34, as shown in FIG. 4.

[0053] 3, the sequential controller 170 runs the non-fusion attitude determination modeler 162 at a frequency of 240 Hz. The non-fusion attitude determination modeler 162 therefore runs asynchronously with respect to the fusion routine 160. In this example, the non-fusion attitude determination modeler 162 utilizes the SLAM system 48 to determine the location of the totem 34. Other systems may use other techniques to determine the location of the totem 34.

[0054] The head unit camera 62 routinely captures images of the totem 34 along with images of real-world objects, such as the real-world object 14. The images captured by the head unit camera 62 are fed into the SLAM system 48. The SLAM system 48 also determines the location of the totem 34 in addition to determining the location of real-world objects, such as the real-world object 14. Thus, the SLAM system 48 establishes a relationship 180 of the totem 34 to the head unit 18. The SLAM system 48 also relies on data from the EM receiver 150 to establish the relationship 180.

[0055] The SLAM system 48 also establishes a relationship 182 of the head unit to the world frame 172. As previously mentioned, the fusion routine 60 receives input from the SLAM system 48. The fusion routine uses the head unit to world frame relationship 182, i.e., head pose, as part of computing a fused model of the pose of the totem 34.

[0056] The relative attitude of the totem 34 with respect to the head unit 18 is established by determining an EM dipole model from measurements by the EM receiver 150. Two relationships 180 and 182 therefore establish the attitude of the totem 34 within the world frame 172. The relationship between the totem 34 and the world frame 172 is stored within the world frame 172 as an unfused attitude 184.

[0057] Comparator 164 runs synchronously with unfused attitude determination modeler 162. Comparator 164 compares fused attitude 174 and unfused location 184. Comparator 164 then incorporates the difference between fused attitude 174 and unfused attitude 184 into drift declarer 166. Drift declarer 166 declares drift only if the difference between fused attitude 174 and unfused attitude 184 exceeds a predetermined maximum distance 188 stored in vision algorithm 38. Predetermined maximum distance 188 is typically less than 100 mm, preferably on the order of 30 mm, 20 mm, or more preferably 10 mm, and is determined or adjusted through data analysis of the sensor fusion system. Drift declarer 166 does not declare drift if the difference between fused attitude 174 and unfused attitude 184 is less than predetermined maximum distance 188.

[0058] When drift declarer 166 declares a drift, drift declarer 166 enters attitude reset routine 168. Attitude reset routine 168 resets fused attitude 174 in fusion routine 160 using unfused attitude 184, so the drift is stopped and fusion routine 160 resumes attitude tracking with the drift eliminated.

[0059] 6 illustrates the relationship between the rig frame 196, the world frame 172, and the fused pose 174. The rig frame 196 is a mathematical object that represents the head frame of the head unit 18. The rig frame 196 is located between the waveguides 50A and 50B. In high dynamic motion scenarios, the fused pose 174 may drift over time (T1; T2; T3; T4) due to imperfect modeling of the actual EM receiver measurements. The fused pose 174 initially represents the actual pose of the totem 34, but in such high dynamic motion scenarios, it may gradually fail to represent the actual pose of the totem 34 as it drifts further from the actual pose of the totem 34 over time.

[0060] FIG. 7 illustrates one method of correcting for drift. The method illustrated in FIG. 7 has a distance-based user drift detection threshold. As an example, if the totem 34 is more than 2 meters away from the head unit 18, it is not possible for the user 10 to hold the totem 34 at such a distance, and drift is declared. If the user 10 can extend their arm, for example, 0.5 meters, the system will only declare drift when the drift reaches an additional 1.5 meters. Such large drifts are undesirable. A system in which drift is declared more quickly is more desirable.

[0061] FIG. 8 illustrates the manner in which drift is declared according to the embodiment in FIG. 3 . As described with reference to FIG. 3 , the non-fused attitude determination modeler 162 calculates the non-fused attitude 184 at a frequency of 240 Hz. As mentioned above, drift may be declared if the difference between the fused attitude 174 and the non-fused location 184 is 100 mm or less, as explained above. At t1, for example, a system error detection threshold of 100 mm is reached, and drift is declared. At t2, the drift has been immediately corrected. Drift can therefore be corrected for smaller range errors in the system in FIG. 8 than in the system of FIG. 7 . Additionally, drift may be corrected again at t3. Drift can therefore be corrected more frequently in the system of FIG. 8 than in the system of FIG. 7 .

[0062] FIG. 9 shows how drift is corrected. At A, a relationship is established between world frame 172 and rig frame 196. Rig frame 196 is not co-located with EM receiver 150. Due to factory calibration, the location of EM receiver 150 relative to rig frame 196 is known. At B, an adjustment is made to calculate rig frame 196 relative to the location of EM receiver 150. At C, an estimate is made from the location of EM receiver 150 relative to EM transmitter 154. As mentioned above, such an estimate may be made using SLAM system 48. Due to factory calibration, the location of EM transmitter 154 is known relative to the location of totem IMU 156. At D, an adjustment is made to determine the location of totem IMU 156 relative to EM transmitter 154. The calculations made at A, B, C, and D thus establish the location of totem IMU 156 within world frame 172. The attitude of the totem IMU 156 can then be reset based on the location of the totem IMU 156 in the world frame 172 as calculated.

[0063] 10 shows a schematic representation of a machine in the exemplary form of a computer system 900 upon which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed according to some embodiments. In alternative embodiments, the machine may operate as a stand-alone device or may be connected (e.g., networked) to other machines. Moreover, while only a single machine is illustrated, the term "machine" should also be taken to include any collection of machines that, individually or together, execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.

[0064] The exemplary computer system 900 includes a processor 902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), and a static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), which communicate with each other via a bus 908.

[0065] The computer system 900 may further include a disk drive unit 916 and a network interface device 920 .

[0066] Disk drive unit 916 includes a machine-readable medium 922 on which is stored one or more sets of instructions 924 (e.g., software) that embody any one or more of the methodologies or functions described herein. Software may also reside, completely or at least partially, within main memory 904 and / or processor 902 during its execution by computer system 900, main memory 904, and processor 902, and also constitute a machine-readable medium.

[0067] The software may also be transmitted or received via the network interface device 920 over the network 928 .

[0068] The computer system 900 includes a laser driver chip 950, which is used to drive the projector and generate the laser light. The laser driver chip 950 includes its own data storage device 960 and its own processor 962.

[0069] While machine-readable medium 922 is shown in the exemplary embodiment to be a single medium, the term "machine-readable medium" should be taken to include a single medium or multiple media (e.g., centralized or distributed databases and / or associated caches and servers) that store one or more sets of instructions. The term "machine-readable medium" should also be taken to include any medium capable of storing, encoding, or carrying a set of instructions for execution by a machine, causing the machine to perform any one or more of the methodologies of the present invention. The term "machine-readable medium" should therefore be taken to include, but is not limited to, solid-state memory, optical and magnetic media, and carrier wave signals.

[0070] While certain exemplary embodiments have been described and shown in the accompanying drawings, it is to be understood that such embodiments are merely illustrative of the invention and not limiting, and that the invention is not limited to the specific construction and arrangements shown and described, as modifications may occur to those skilled in the art.

Claims

1. A user interaction system, comprising: It is a totem, The totem itself, a totem inertial measurement unit (IMU) located on the totem body and configured to generate a totem IMU signal due to movement of the totem; a totem having The head unit, a processor; a storage device connected to the processor; a set of instructions on the storage device and executable by the processor, the set of instructions comprising: The world frame and an executable routine coupled to the totem IMU, the executable routine generating a first location of the totem within the world frame based on the totem IMU signal; a location determination modeler that determines a second location of the totem relative to the world frame; a location correction routine coupled to the first location and the second location, the location correction routine resetting the sensors of the totem IMU to match the second location; a set of instructions including: a data source carrying image data; a display system connected to the data source and configured to use the image data to display a virtual object to a user, the location of the virtual object being based on the first location of the totem; and A user interaction system comprising:

2. A comparator connected to the first location and the second location; a drift declarer coupled to the comparator for declaring a drift only if the first location is more than a predetermined distance from the second location; a location correction routine, coupled to the drift declarer, for resetting the totem IMU sensors to match the second location only if the drift is declared; The user interaction system of claim 1 further comprising:

3. A user interaction system as described in claim 1, wherein a location determination modeler determines a location of the totem relative to the head unit and a location of the head unit relative to the world frame, and establishes a second location of the totem relative to the world frame.

4. A rendering engine connected to the data source and having an input for receiving the image data and an output to the display system, the rendering engine providing to the display system a data stream including the virtual object located at the first location. The user interaction system of claim 1 further comprising:

5. The user interaction system of claim 1, wherein the totem IMU reduces jitter of the virtual object.

6. The user interaction system of claim 1, wherein the totem IMU includes at least one of a gyroscope and an accelerometer.

7. The totem is: Electromagnetic (EM) transmitter on the totem body The head unit includes: A head unit body, an EM receiver on the head unit body for receiving EM waves transmitted by the EM transmitter, the EM waves indicating the location of the totem; 2. The user interaction system of claim 1, wherein the executable routine is a fusion routine, the fusion routine connected to the EM receiver and the totem IMU, and generates the first location, the first location being a fused location of the totem in the world frame based on a combination of the EM waves and the totem IMU signal, and the location determination modeler is a non-fused location determination modeler, the non-fused location determination modeler determines a location of the totem relative to the head unit and a location of the head unit relative to the world frame, and establishes the second location, which is an un-fused location of the totem relative to the world frame, and the location of the virtual object is based on the fused location of the totem.

8. A user interaction system as described in claim 7, wherein the fusion routine executes at a first frequency and the non-fusion location determination modeler executes at a second frequency different from the first frequency.

9. A user interaction system as described in claim 7, wherein the EM receiver detects six degrees of freedom ("6dof") movement of the EM transmitter relative to the EM receiver.

10. A user interaction system as described in claim 2, wherein the predetermined distance is less than 100 mm.

11. The location correction routine is executable to perform a method, the method comprising: storing a rig frame, the rig frame being a mathematical object located between the eyepieces of the head unit and serving as a basis for defining where objects reside relative to the head unit; taking measurements to determine a rig frame attitude that the rig frame is assuming relative to the world frame; Deriving a receiver attitude relative to the world using known extrinsic properties, such as those provided by factory level calibration, of the relationship between the EM receiver and the rig frame; estimating an EM relationship between the EM receiver and the EM transmitter to derive a transmitter attitude in the world; deriving an in-world totem IMU attitude as a non-fused attitude using known pertinent properties of the totem regarding the relationship between the EM transmitter and the totem IMU; 8. The user interaction system of claim 7, comprising:

12. A camera on the head unit body, the camera capturing an image of the totem. a camera positioned to a simultaneous localization and mapping (SLAM) system connected to the camera and receiving the image of the totem, the SLAM system determining a non-fusion location based on the image of the totem; and The user interaction system of claim 1 further comprising:

13. The display system, a transparent waveguide secured to the head unit body, the transparent waveguide allowing light from the totem to pass to the eyes of a user wearing the head unit body; a projector that converts the image data into light, the light from the projector entering the waveguide at an entrance pupil and exiting the waveguide to the eye of the user at an exit pupil; and The user interaction system of claim 1 , comprising:

14. A head unit detection device for detecting movement of a head-mountable frame. and wherein the set of instructions further comprises: a display adjustment algorithm connected to the head unit detection device, the display adjustment algorithm receiving measurements based on the movements detected by the head unit detection device and calculating a setpoint value; a rendering engine that modifies the position of the virtual object within the view of the user's eye based on the setting value; 14. The user interaction system of claim 13, comprising:

15. The head unit detection device is a head unit IMU mounted on the head-mountable frame, the head unit IMU including a motion sensor that detects movement of the head-mountable frame; 15. The user interaction system of claim 14, comprising:

16. A user interaction system as described in claim 15, wherein the head unit IMU includes at least one of a gyroscope and an accelerometer.

17. The head unit detection device is a head unit camera mounted on the head-mountable frame, the head unit camera detecting movement of the head-mountable frame by capturing images of objects within the view of the head unit camera; the set of instructions comprising: a simultaneous localization and mapping (SLAM) system connected to the camera, receiving the image of the object and detecting a pose of the head unit body, wherein the rendering engine modifies the position of the virtual object based on the pose.

15. The user interaction system of claim 14, comprising:

Citation Information

Patent Citations

  • Systems and methods for augmented reality

    US20170205903A1

  • Object motion tracking with remote device

    US20170220119A1

  • Systems and methods for drift correction

    US20170221225A1

  • Electromagnetic tracking with augmented reality systems

    US20170307891A1

  • Six DOF mixed reality input by fusing inertial handheld controller with hand tracking

    US20170357332A1