Non-Fusion Posture-Based Drift Correction of the Fusion Posture of Totems in a User Interaction System
The user interaction system addresses the challenge of maintaining virtual object posture by using a totem with EM and IMU components, and a processor to correct drift, achieving stable and realistic alignment in augmented reality.
Patent Information
- Application Number
- JP2024019329
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-13
- Filing Date
- 2024-02-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-07-26
AI Technical Summary
Existing user interaction systems in augmented reality struggle to maintain the realistic posture of virtual objects relative to a totem, leading to potential drift and misalignment over time.
A user interaction system that includes a totem with an electromagnetic transmitter and inertial measurement unit, a head unit with an electromagnetic receiver and IMU, and a processor that executes a fusion routine to generate a fused pose of the totem, comparing it with a non-fused pose to declare and correct drift.
The system effectively reduces jitter and maintains the accurate alignment of virtual objects with the totem, ensuring a stable and realistic user experience by correcting drift and maintaining precise pose tracking.
Smart Images

Figure 0007693866000001 
Figure 0007693866000002 
Figure 0007693866000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 62 / 714,609, filed Aug. 3, 2018, and U.S. Provisional Patent Application No. 62 / 818,032, filed Mar. 13, 2019, which are hereby incorporated by reference in their entirety.
[0002] The present invention relates to a user - interaction system having a totem that defines a six - degree - of - freedom (“6dof”) pose, i.e., the pose of a virtual object perceived by a user.
Background Art
[0003] Modern computing and display technologies have facilitated the development of user - interaction systems, including “augmented reality” viewing devices. Such viewing devices typically have a head - unit body with two waveguides, one in front of each eye of the user, which can be mounted on the user's head. The waveguides are transparent so that ambient light from real - world objects can pass through the waveguides and the user can see the real - world objects. Each waveguide also serves to transmit light projected from a projector to the user's individual eye. The projected light forms an image on the retina of the eye. The retina of the eye thus receives both ambient light and the projected light. The user can see both real - world objects and one or more virtual objects created by the projected light at the same time.
[0004] Such user interaction systems often include a totem. The user may, for example, hold the totem in their right hand and move the totem with six degrees of freedom within three-dimensional space. Virtual objects may be attached to the totem and may be perceived by the user to move with the totem within three-dimensional space, or the virtual objects may be the perception of a light beam hitting a wall or another object that the user moves across the wall.
[0005] It is important that the virtual object remains in its realistic posture relative to the totem. For example, if the totem represents the handle of a racket and the virtual object represents the head of the racket, the head of the racket needs to remain "attached" to the handle of the racket over time. Summary of the Invention Means for Solving the Problems
[0006] The present invention provides a user interaction system including a totem having a totem body, an electromagnetic (EM) transmitter on the totem body, and a totem inertial measurement unit (IMU) located on the totem that generates a totem IMU signal due to the movement of the totem; a head unit having a head unit body and an EM receiver on the head unit body that receives an EM wave transmitted by the EM transmitter and indicates the location of the totem; a processor; a memory device connected to the processor; and a set of instructions on the memory device that are executable by the processor. The set of instructions includes a fusion routine that is connected to a world frame, the EM receiver, and the totem IMU and generates a fused pose of the totem in the world frame based on a combination of the EM wave, the head unit pose, and the totem IMU data; a non-fused pose determination modeler that determines the pose of the totem relative to the head unit and the pose of the head unit relative to the world frame and establishes a non-fused pose of the totem relative to the world frame; a comparator connected to the fusion pose determination modeler and the non-fused pose determination modeler that compares the fused pose and the non-fused pose; a drift declarer connected to the comparator that declares a drift only if the fused pose exceeds a predetermined distance from the non-fused pose; a location correction routine connected to the drift declarer that resets the pose of the totem IMU and matches it to the non-fused location only if a drift is declared; a data source that carries image data; and a display system connected to the data source that uses the image data to display virtual objects to the user, wherein the location of the virtual objects is based on the fused location of the totem.
[0007] The present invention also includes steps of transmitting electromagnetic (EM) waves using an EM transmitter on a totem body; generating a totem inertial measurement unit (IMU) signal using a totem IMU on the totem body due to the movement of the totem; determining the location of a head unit body on a user's head; receiving the EM waves transmitted by the EM transmitter by an EM receiver on the head unit body, where the EM waves indicate the posture of the totem; storing a world frame; using a processor to execute a fusion routine to generate a fused posture of the totem within the world frame based on a combination of the EM waves, head unit posture, and totem IMU data; using a processor to execute a non-fused posture determination model to determine the posture of the totem relative to the head unit and the location of the head unit relative to the world frame, and establish a non-fused posture of the totem relative to the world frame; using a processor to execute a comparator to compare the fused posture and the non-fused posture; using a processor to execute a drift declarer to declare a drift only when the fused posture exceeds a predetermined posture from the non-fused posture; using a processor to execute a posture correction routine to reset the posture of the totem IMU and match it to the non-fused posture only when a drift is declared; receiving image data from a data source; and using a display system connected to the data source to display virtual objects to the user using the image data, where the location of the virtual objects is based on the fused location of the totem. The present invention provides, for example, the following items. (Item 1) A user interaction system, A totem, A totem body, An electromagnetic (EM) transmitter on the totem body, A totem inertial measurement unit (IMU) located on the totem that generates a totem IMU signal due to the movement of the totem, A totem having the above, A head unit, a head unit main body, an EM receiver located on the head unit main body for receiving EM waves transmitted by the EM transmitter, wherein the EM waves indicate the location of the totem, and the EM receiver a head unit having; a processor, a storage device connected to the processor, a set of instructions, the set of instructions comprising: a world frame, a fusion routine connected to the EM receiver and the totem IMU, the fusion routine generating a fused pose of the totem in the world frame based on a combination of the EM waves, the head unit pose, and the totem IMU data, a non-fused pose determination model that determines the pose of the totem relative to the head unit and the pose of the head unit relative to the world frame, and establishes a non-fused pose of the totem relative to the world frame, a comparator connected to the fusion pose determination model and the non-fused pose determination model for comparing the fused pose and the non-fused pose, a drift declarer connected to the comparator for declaring drift only when the fused pose exceeds a predetermined distance from the non-fused pose, a location correction routine connected to the drift declarer for resetting the pose of the totem IMU and matching it to the non-fused location only when drift is declared, a data source for carrying image data, a display system connected to the data source for displaying virtual objects to the user using the image data, wherein the location of the virtual objects is based on the fused location of the totem, a set of instructions stored on the storage device and executable by the processor, A user interaction system comprising (Item 2) A rendering engine connected to the data channel and having an input for receiving the image data and an output to the display system, the rendering engine providing to the display system a data stream including the virtual object positioned in the fused pose The user interaction system according to item 1, further comprising (Item 3) The user interaction system according to item 1, wherein the totem IMU reduces jitter of the virtual object (Item 4) The user interaction system according to item 1, wherein the totem IMU includes at least one of a gyroscope and an accelerometer (Item 5) The user interaction system according to item 1, wherein the fusion routine is executed at a first frequency and the non-fused pose determination model is executed at a second frequency different from the first frequency (Item 6) The EM receiver detects 6 degrees of freedom ("6dof" ) movement of the EM transmitter relative to the EM receiver (Item 7) The user interaction system according to item 1, wherein the predetermined distance is less than 100 mm. (Item 8) A camera on the head unit body, the camera being positioned to capture an image of the totem, and A simultaneous localization and mapping (SLAM) system connected to the camera and receiving an image of the totem, the SLAM system determining the non-fused pose based on the image of the totem The user interaction system according to item 1, further comprising (Item 9) The display system A transparent waveguide, wherein the transparent waveguide is fixed to a head unit main body that passes light from the totem to the eyes of a user wearing the head unit main body, the transparent waveguide; A projector that converts the image data into light, wherein the light from the projector enters the waveguide at an entrance pupil and exits the waveguide to the user's eyes at an exit pupil, the projector; The user interaction system according to item 1, comprising: (Item 10) A head unit detection device that detects movement of the head-mountable frame further comprising, wherein the set of instructions is: A display adjustment algorithm, wherein the display adjustment algorithm is connected to the head unit detection device, receives a measurement value based on the movement detected by the head unit detection device, and calculates an installation value, the display adjustment algorithm; A rendering engine that corrects the position of the virtual object within the view of the eye based on the installation value The user interaction system according to item 1, comprising: (Item 11) The head unit detection device is: A head unit IMU mounted on the head-mountable frame, wherein the head unit IMU includes a motion sensor that detects movement of the head-mountable frame, and the head unit IMU includes at least one of a gyroscope and an accelerometer, the head unit IMU The user interaction system according to item 10, comprising: (Item 12) The head unit detection device is: A head unit camera mounted on the head-mountable frame, wherein the head unit camera includes a head unit camera that detects movement of the head-mountable frame by taking an image of an object within the view of the head unit camera, and the set of instructions is: A simultaneous localization and mapping (SLAM) system connected to the camera, receiving an image of the object, and detecting the pose position of the head unit body, wherein the rendering engine corrects the position of the virtual object based on the pose position. The user interaction system according to item 10, including (Item 13) The location correction routine is to store a rig frame, where the rig frame is a mathematical object located between the eyepieces of the head unit and serves as a basis for defining the location where the object exists relative to the head unit. And to perform measurements to determine the rig frame pose of the rig frame relative to the world frame. And to derive the receiver pose relative to the world using known ancillary properties provided by factory-level calibration regarding the relationship between the EM receiver and the rig frame. And to estimate the EM relationship between the EM receiver and the EM transmitter and derive the in-world transmitter pose. And to use the known ancillary properties of the totem regarding the relationship between the EM transmitter and the totem IMU to derive the in-world totem IMU pose as a non-fused pose. The user interaction system according to item 1, which is executable to perform a method including And to derive the in-world totem IMU pose as a non-fused pose. (Item 14) A user interaction system Transmitting electromagnetic (EM) waves using an EM transmitter on the totem body. Generating a totem inertial measurement unit (IMU) signal using a totem IMU on the totem body due to the movement of the totem. Locating the head unit body on the user's head. Receiving, by an EM receiver on the head unit body, the EM waves transmitted by the EM transmitter, where the EM waves indicate the pose of the totem. Remember the world frame, and Use a processor to execute a fusion routine to generate a fused pose of the totem within the world frame based on a combination of the EM wave, the head unit pose, and the totem IMU data; and Use the processor to execute a non-fused pose determination model that determines the pose of the totem relative to the head unit and the location of the head unit relative to the world frame, and establishes a non-fused pose of the totem relative to the world frame; and Use the processor to execute a comparator to compare the fused pose and the non-fused pose; and Use the processor to execute a drift declarer to declare a drift only if the fused pose exceeds a predetermined pose from the non-fused pose; and Use the processor to execute a pose correction routine to reset the pose of the totem IMU and match it to the non-fused pose only if the drift is declared; and Receive image data from a data source; and Use a display system connected to the data source to display virtual objects to the user using the image data, wherein the location of the virtual objects is based on the fused pose of the totem; and A user interaction system comprising.
Brief Description of the Drawings
[0008] The present invention will be further described by way of example with reference to the accompanying drawings.
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3
[0012]
Figure 4
[0013]
Figure 5
[0014]
Figure 6
[0015]
Figure 7
[0016]
Figure 8
[0017]
Figure 9
[0018]
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0019] FIG. 1 of the accompanying drawings illustrates a user 10, a user interaction system 12 according to an embodiment of the present invention, a real-world object 14 in the form of a table, and a virtual object 16 that is invisible from the perspective of the figure but visible to the user 10.
[0020] The user interaction system 12 includes a head unit 18, a belt pack 20, a network 22, and a server 24.
[0021] The head unit 18 includes a head unit body 26 and a display system 28. The head unit body 26 has a shape that conforms to the user's head 10. The display system 28 is affixed to the head unit body 26.
[0022] The belt pack 20 has a processor and a memory device connected to the processor. Visual algorithms are stored on the memory device and are executable by the processor. The belt pack 20 is communicatively connected to the display system 28 using a cable connection 30. The belt pack 20 further includes a network interface device that enables the belt pack 20 to connect wirelessly via a link 32 to the network 22. The server 24 is connected to the network 22.
[0023] In use, the user 10 secures the head unit body 26 to their head. The display system 28 includes an optical waveguide (not shown) that is transparent so that the user 10 can see the real-world object 14 through the waveguide.
[0024] The belt pack 20 may download image data from the server 24 via the network 22 and the link 32. The belt pack 20 provides the image data to the display system 28 through the cable connection 30. The display system 28 has one or more projectors that create light based on the image data. The light propagates through one or more optical waveguides to the eyes of the user 10. Each waveguide creates the light at a specific focal distance on the retina of an individual eye so that the virtual object 16 appears to the eye at a certain distance behind the display system 28. The user's eyes thus see the virtual object 16 in three-dimensional space. Additionally, slightly different images are created for each eye so that the user 10's brain perceives the virtual object 16 in three-dimensional space. The user 10 thus sees the real-world object 14 extended with the virtual object 16 in three-dimensional space.
[0025] The user interaction system 12 further includes a totem 34. In use, the user 10 holds the totem 34 in one of their hands. The virtual object 16 is positioned in three-dimensional space based on the positioning of the totem 34. As an example, the totem 34 may be the handle of a racket, and the virtual object 16 may include the head of the racket. The user 10 can move the totem 34 in six degrees of freedom in three-dimensional space. The totem 34 thus moves in three-dimensional space relative to the real-world object 14 and the head unit body 26. Various components within the head unit 18 and the belt pack 20 track the movement of the totem 34 and move the virtual object 16 along with the totem 34. The head of the racket thus remains attached to the handle within the view of the user 10.
[0026] Figure 2 illustrates the display system 28 in more detail, together with the visual algorithm 38. The visual algorithm 38 mainly resides in the belt pack 20 in FIG. 1. In other embodiments, the visual algorithm 38 may entirely reside within the head unit or may be split between the head unit and the belt pack. FIG. 2 further includes a data source 40. In this example, the data source 40 includes image data stored on the storage device of the belt pack 20. The image data may be, for example, three-dimensional image data that can be used to render the virtual object 16. In alternative embodiments, the image data may be time-series image data that enables the creation of videos moving in two or three dimensions, for the purpose of which it may be located on a real-world object having an association with the totem, or at a fixed position in front of the user when the user moves their head.
[0027] The visual algorithm 38 includes a rendering engine 42, a stereoscopic analyzer 44, a display adjustment algorithm 46, and a simultaneous localization and mapping (SLAM) system 48.
[0028] The rendering engine 42 is connected to the data source 40 and the display adjustment algorithm 46. The rendering engine 42 can receive inputs from various systems, in this example, from the display adjustment algorithm 46, and position the image data within the frame that will be visually recognized by the user 10 based on the display adjustment algorithm 46. The display adjustment algorithm 46 is connected to the SLAM system 48. The SLAM system 48 can receive the image data, analyze the image data for the purpose of determining the objects within the image of the image data, and record the locations of the objects within the image data.
[0029] The stereoscopic analyzer 44 is connected to the rendering engine 42. The stereoscopic analyzer 44 is capable of determining left and right image data sets from the data stream provided by the rendering engine 42.
[0030] The display system 28 includes left and right projectors 48A and 48B, left and right waveguides 50A and 50B, and a detection device 52. The left and right projectors 48A and 48B are connected to a power supply. Each projector 48A or 48B has an individual input for image data to be provided to the individual projector 48A or 48B. When powered, the individual projector 48A or 48B generates light in a two-dimensional pattern and emits the light therefrom. The left and right waveguides 50A and 50B are positioned to receive light from the left and right projectors 48A and 48B, respectively. The left and right waveguides 50A and 50B are transparent waveguides.
[0031] The detection device 52 includes a head unit inertial motion unit (IMU) 60 and one or more head unit cameras 62. The head unit IMU 60 includes one or more gyroscopes and one or more accelerometers. The gyroscopes and accelerometers are typically formed within a semiconductor chip and are capable of detecting the movement of the head unit IMU 60 and the head unit body 26, including movement along three orthogonal axes and rotation about the three orthogonal axes.
[0032] The head unit camera 62 continuously captures images from the environment around the head unit body 26. The images can be compared with each other to detect the movement of the head unit body 26 and the user's head 10.
[0033] The SLAM system 48 is connected to the head unit camera 62. The display adjustment algorithm 46 is connected to the head unit IMU 60. Those skilled in the art will understand that the connection between the detection device 52 and the vision algorithm 38 is accomplished through a combination of hardware, firmware, and software. The components of the vision algorithm 38 are linked to each other through subroutines or calls.
[0034] In use, the user 10 mounts the head unit body 26 on his head. The components of the head unit body 26 may include, for example, a strap (not shown) that wraps around the perimeter of the back of the user's head 10. The left and right waveguides 50A and 50B are then positioned in front of the user 10's left and right eyes 120A and 120B.
[0035] The rendering engine 42 receives image data from the data source 40. The rendering engine 42 inputs the image data into the stereoscopic analyzer 44. The image data is the three-dimensional image data of the virtual object 16 in FIG. 1. The stereoscopic analyzer 44 analyzes the image data and determines left and right image data sets based on the image data. The left and right image data sets are data sets that represent two-dimensional images that are slightly different from each other for the purpose of giving the user 10 the perception of three-dimensional rendering. In this embodiment, the image data is a static data set that does not change over time.
[0036] The stereoscopic analyzer 44 inputs left and right image datasets into the left and right projectors 48A and 48B. The left and right projectors 48A and 48B then create left and right light patterns. The components of the display system 28 are shown in a plan view, but it should be understood that the left and right patterns are two-dimensional patterns when shown in a front elevation view. Each light pattern includes a plurality of pixels. For illustrative purposes, light rays 124A and 126A from two of the pixels are shown exiting the left projector 48A and entering the left waveguide 50A. Light rays 124A and 126A reflect from the side of the left waveguide 50A. Although light rays 124A and 126A are shown propagating through internal reflection from left to right within the left waveguide 50A, it should be understood that light rays 124A and 126A also propagate in a direction away from the plane of the paper using refractive and reflective systems.
[0037] Light rays 124A and 126A exit the left light waveguide 50A through the pupil 128A and then enter the left eye 120A through the pupil 130A of the left eye 120A. Light rays 124A and 126A then strike the retina 132A of the left eye 120A. Thus, the left light pattern strikes the retina 132A of the left eye 120A. The user 10 is given the perception that the pixels formed on the retina 132A are pixels 134A and 136A, which the user 10 perceives to be at a certain distance on the side of the left waveguide 50A facing the left eye 120A. Depth perception is created by manipulating the focal length of the light.
[0038] Similarly, the stereoscopic analyzer 44 inputs the right image data set into the right projector 48B. The right projector 48B transmits a right light pattern represented by pixels in the form of light rays 124B and 126B. The light rays 124B and 126B are reflected within the right waveguide 50B and exit through the pupil 128B. The light rays 124B and 126B then enter through the pupil 130B of the right eye 120B and strike the retina 132B of the right eye 120B. The pixels of the light rays 124B and 126B are perceived as pixels 134B and 136B behind the right waveguide 50B.
[0039] The patterns created on the retinas 132A and 132B are perceived individually as left and right images. The left and right images are slightly different from each other due to the function of the stereoscopic analyzer 44. The left and right images are perceived by the user 10's brain as a 3D rendering.
[0040] As described above, the left and right waveguides 50A and 50B are transparent. Light from real objects on the sides of the left and right waveguides 50A and 50B facing the eyes 120A and 120B can be projected through the left and right waveguides 50A and 50B and strike the retinas 132A and 132B. In particular, the light from the real-world object 14 in FIG. 1 strikes the retinas 132A and 132B so that the user 10 can see the real-world object 14. In addition, the user 10 can see the totem 34, and augmented reality is created, and the real-world object 14 and the totem 34 are extended by a 3D rendering of the virtual object 16 perceived by the user 10 due to the left and right images that are perceived by the user 10 in combination.
[0041] The head unit IMU 60 detects any movement of the head of the user 10. If the user 10, for example, moves their head counterclockwise and at the same time moves their body to the right along with their head, such movement will be detected by the gyroscope and accelerometer within the head unit IMU 60. The head unit IMU 60 provides the measurement values from the gyroscope and accelerometer to the display adjustment algorithm 46. The display adjustment algorithm 46 calculates the installation values and provides the installation values to the rendering engine 42. The rendering engine 42 corrects the image data received from the data source 40 to compensate for the movement of the head of the user 10. The rendering engine 42 provides the corrected image data to the stereoscopic analyzer 44 for display to the user 10.
[0042] The head unit camera 62 continuously captures images as the user 10 moves their head. The SLAM system 48 analyzes the images and identifies the images of the objects within the images. The SLAM system 48 analyzes the movement of the objects and determines the pose position of the head unit body 26. The SLAM system 48 provides the pose position to the display adjustment algorithm 46. The display adjustment algorithm 46 uses the pose position to further refine the installation values that the display adjustment algorithm 46 provides to the rendering engine 42. The rendering engine 42 thus corrects the image data received from the data source 40 based on the combination of the motion sensors within the head unit IMU 60 and the images captured by the head unit camera 62. As a practical example, when the user 10 rotates their head to the right, the location of the virtual object 16 rotates to the left within the field of view of the user 10, thus giving the user 10 the impression that the location of the virtual object 16 remains stationary with respect to the real-world object 14 and the totem 34.
[0043] Figure 3 illustrates further details of the head unit 18, the totem 34, and the vision algorithm 38. The head unit 18 further includes an electromagnetic (EM) receiver 150 that is fixed to the head unit body 26. The display system 28, the head unit camera 62, and the EM receiver 150 are mounted at fixed positions relative to the head unit body 26. When the user 10 moves his or her head, the head unit body 26 moves with the user 10's head, and the display system 28, the head unit camera 62, and the EM receiver 150 move with the head unit body 26.
[0044] The totem 34 has a totem body 152, an EM transmitter 154, and a totem IMU 156. The EM transmitter 154 and the totem IMU 156 are mounted at fixed positions relative to the totem body 152. The user 10 holds the totem body 152, and when the user 10 moves the totem body 152, the EM transmitter 154 and the totem IMU 156 move with the totem body 152. The EM transmitter 154 is capable of transmitting EM waves, and the EM receiver 150 is capable of receiving EM waves. The totem IMU 156 has one or more gyroscopes and one or more accelerometers. The gyroscopes and accelerometers are typically formed within a semiconductor chip and are capable of detecting the movement of the totem IMU 156 and the totem body 152, including movement along three orthogonal axes and rotation about three orthogonal axes.
[0045] The vision algorithm 38 further includes a fusion routine 160, a non-fused pose determination model 162, a comparator 164, a drift declarer 166, a pose correction routine 168, and a sequential controller 170, in addition to the data source 40, the rendering engine 42, the stereoscopic analyzer 44, and the SLAM system 48, which are described with reference to FIG. 2.
[0046] The head unit camera 62 captures an image of the real-world object 14. The image of the real-world object 14 is processed by the SLAM system 48 to establish the world frame 172 as described with reference to FIG. 2. Details of how the SLAM system 48 establishes the world frame 172 are not shown in FIG. 3 so as not to obscure the drawings.
[0047] The EM transmitter 154 transmits the EM waves received by the EM receiver 150. The EM waves received by the EM receiver 150 indicate the attitude or change in attitude of the EM transmitter 154. The EM receiver 150 incorporates the data of the EM waves into the fusion routine 160.
[0048] The totem IMU 156 continuously monitors the movement of the totem body 152. The data from the totem IMU 156 is incorporated into the fusion routine 160.
[0049] The sequencer 170 executes the fusion routine 160 at a frequency of 250 Hz. The fusion routine 160 combines the data from the EM receiver 150 and the data from the totem IMU 156 and the SLAM system 48. The EM waves received by the EM receiver 150 contain data that relatively accurately represents the attitude of the EM transmitter 154 with respect to the EM receiver 150 in six degrees of freedom (“6dof”). However, due to EM measurement noise, the measured EM waves may not accurately represent the attitude of the EM transmitter 154 with respect to the EM receiver 150. The EM measurement noise can result in the jitter of the virtual object 16 in FIG. 1. The purpose of combining the data from the totem IMU 156 is to reduce the jitter. The fusion routine 160 provides a fusion attitude 174 within the world frame 172. The fusion attitude 174 is used by the rendering engine 42 for the purpose of determining the attitude of the virtual object 16 in FIG. 1 using the image data from the data source 40.
[0050] As shown in FIG. 4, the virtual object 16 is shown in the correct orientation with respect to the totem 34. Further, when the user 10 moves the totem 34, the virtual object 16 moves with the totem 34 with a minimal amount of jitter.
[0051] The totem IMU 156 measures acceleration and angular velocity essentially in six degrees of freedom. The acceleration and angular velocity are integrated to determine the location and orientation of the totem IMU 156. Due to integration errors, the fused pose 174 can drift over time.
[0052] FIG. 5 illustrates that the virtual object 16 has drifted from its correct orientation with respect to the totem 34. The drift can be caused by an imperfect mathematical model that accounts for the so-called "model inconsistency", i.e., the relationship between physical quantities (e.g., 6dof, acceleration, and angular velocity) and the actually measured signals (such as EM wave measurements and IMU signals). Also, such drift can be amplified for highly dynamic motions, which can even lead to a state where the fusion algorithm diverges (i.e., the virtual object will seem to be "blown away" from the actual object). In this embodiment, the virtual object 16 has drifted to the right with respect to the totem 34. The fused pose 174 in FIG. 3 is based on the system's confidence that the totem 34 is located further to the right than where it actually is. The fused data is thus corrected so that the virtual object 16 is repositioned to its correct location with respect to the totem 34 as shown in FIG. 4.
[0053] In FIG. 3, the sequential controller 170 executes the non-fused pose determination modeler 162 at a frequency of 240 Hz. The non-fused pose determination modeler 162 thus executes asynchronously with respect to the fusion routine 160. In this embodiment, the non-fused pose determination modeler 162 uses the SLAM system 48 to determine the location of the totem 34. Other systems may use other techniques to determine the location of the totem 34.
[0054] The head unit camera 62 routinely captures an image of the totem 34, along with an image of a real-world object such as the real-world object 14. The image captured by the head unit camera 62 is incorporated into the SLAM system 48. The SLAM system 48 also determines the location of the totem 34 in addition to determining the location of a real-world object such as the real-world object 14. Thus, the SLAM system 48 establishes the relationship 180 of the totem 34 with respect to the head unit 18. The SLAM system 48 also relies on the data from the EM receiver 150 to establish the relationship 180.
[0055] The SLAM system 48 also establishes the relationship 182 of the head unit with respect to the world frame 172. As described above, the fusion routine 60 receives an input from the SLAM system 48. The fusion routine uses the relationship 182 between the head unit and the world frame, i.e., the head pose, as part of the calculation of the fusion model of the pose of the totem 34.
[0056] The relative pose of the totem 34 with respect to the head unit 18 is established by obtaining an EM dipole model from the measurements by the EM receiver 150. The two relationships 180 and 182 thus establish the pose of the totem 34 within the world frame 172. The relationship between the totem 34 and the world frame 172 is stored as an unfused pose 184 within the world frame 172.
[0057] Comparator 164 runs synchronously with the non-fused pose determination model 162. Comparator 164 compares the fused pose 174 and the non-fused location 184. Comparator 164 then captures the difference between the fused pose 174 and the non-fused pose 184 into the drift declarer 166. Drift declarer 166 declares a drift only when the difference between the fused pose 174 and the non-fused pose 184 exceeds a predetermined maximum distance 188 stored within the visual algorithm 38. The predetermined maximum distance 188 is typically on the order of less than 100 mm, preferably 30 mm, 20 mm, or more preferably 10 mm, and is determined or adjusted through data analysis of the sensor fusion system. Drift declarer 166 does not declare a drift when the difference between the fused pose 174 and the non-fused pose 184 is less than the predetermined maximum distance 188.
[0058] When drift declarer 166 declares a drift, drift declarer 166 enters the pose reset routine 168. Pose reset routine 168 uses the non-fused pose 184 to reset the fused pose 174 within the fusion routine 160. Thus, the drift is stopped, and the fusion routine 160 resumes pose tracking with the drift eliminated.
[0059] Figure 6 illustrates the relationship between the rig frame 196, the world frame 172, and the fused pose 174. Rig frame 196 is a mathematical object representing the head frame of the head unit 18. Rig frame 196 is located between the waveguides 50A and 50B. In a high-dynamic motion scenario, the fused pose 174 can drift over time (T1; T2; T3; T4) due to an incomplete modeling of the actual EM receiver measurements. The fused pose 174 initially represents the actual pose of the totem 34, but in such a high-dynamic motion scenario, it can gradually fail to represent the actual pose of the totem 34 as it drifts further over time from the actual pose of the totem 34.
[0060] FIG. 7 illustrates one way to correct for drift. The method illustrated in FIG. 7 has a distance-based user drift detection threshold. As an example, when the totem 34 is more than 2 meters away from the head unit 18, it is not possible for the user 10 to hold the totem 34 at such a distance, and drift is declared. If the user 10 can extend their arm, for example, 0.5 meters, the system will declare drift only when the drift reaches an additional 1.5 meters. Such a large drift is not desirable. A system that declares drift more quickly is more desirable.
[0061] FIG. 8 illustrates a manner in which drift is declared according to the embodiment in FIG. 3. As described with reference to FIG. 3, the non-fused pose determination model 162 calculates the non-fused pose 184 at a frequency of 240 Hz. As described above, drift can be declared when the difference between the fused pose 174 and the non-fused location 184 is 100 mm or less, as explained above. At t1, for example, the system error detection threshold of 100 mm is reached and drift is declared. At t2, the drift is corrected immediately. Drift can thus be corrected for a smaller distance error in the system of FIG. 8 than in the system of FIG. 7. In addition, the drift may be corrected again at t3. Drift can thus be corrected more frequently in the system of FIG. 8 than in the system of FIG. 7.
[0062] Figure 9 shows how the drift is corrected. In A, a relationship is established between the world frame 172 and the rig frame 196. The rig frame 196 is not located at the same position as the EM receiver 150. Due to factory calibration, the location of the EM receiver 150 relative to the rig frame 196 is known. In B, an adjustment is made to calculate the rig frame 196 relative to the location of the EM receiver 150. In C, an estimation is made from the location of the EM receiver 150 relative to the EM transmitter 154. As described above, such an estimation may be made using the SLAM system 48. Due to factory calibration, the location of the EM transmitter 154 is known relative to the location of the totem IMU 156. In D, an adjustment is made to determine the location of the totem IMU 156 relative to the EM transmitter 154. The calculations performed in A, B, C, and D thus establish the location of the totem IMU 156 within the world frame 172. The orientation of the totem IMU 156 can then be reset based on the location of the totem IMU 156 within the world frame 172 as calculated.
[0063] Figure 10 shows a schematic representation of a machine in an exemplary form of a computer system 900 in which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein can be executed according to some embodiments. In alternative embodiments, the machine may operate as a stand-alone device or may be connected (e.g., networked) to other machines. Further, although only a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set of instructions (or multiple sets of instructions) and perform any one or more of the methodologies discussed herein.
[0064] The exemplary computer system 900 includes a processor 902 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), and a static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and these communicate with each other via a bus 908.
[0065] The computer system 900 may further include a disk drive unit 916 and a network interface device 920.
[0066] The disk drive unit 916 includes a machine-readable medium 922 on which is stored a set 924 of one or more instructions (e.g., software) that embody any one or more of the methodologies or functions described herein. The software may also reside, at least in part, within the main memory 904 and / or the processor 902 during execution by the computer system 900, the main memory 904, and the processor 902, and may also constitute a machine-readable medium.
[0067] The software may also be transmitted or received via the network 928 via the network interface device 920.
[0068] The computer system 900 includes a laser driver chip 950 that is used to drive a projector and generate laser light. The laser driver chip 950 includes its own data storage device 960 and its own processor 962.
[0069] The machine-readable medium 922 is shown as a single medium in the illustrative embodiment, but the term "machine-readable medium" should be construed to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated cache and server) that store a set of one or more instructions. The term "machine-readable medium" should also be construed to include any medium that is capable of storing, encoding, or carrying a set of instructions for machine execution and that causes a machine to perform any one or more of the methodologies of the present invention. The term "machine-readable medium" should therefore be construed to include, without limitation, solid-state memory, optical and magnetic media, and carrier wave signals.
[0070] Although an illustrative embodiment has been described and shown in the accompanying drawings, such embodiments are merely illustrative of the invention and not limiting, and it is to be understood that the invention is not limited to the specific structures and arrangements shown and described since modifications may occur to those skilled in the art.
Claims
1. 1. A user interaction system comprising: It is a totem, The totem body and an electromagnetic (EM) transmitter on the totem body; a totem inertial measurement unit (IMU) located on the totem and generating a totem IMU signal due to movement of the totem; A totem having A head unit, A head unit body, an EM receiver on the head unit body for receiving EM waves transmitted by the EM transmitter, the EM waves indicating the location of the totem; A head unit having A processor; A storage device connected to the processor; A set of instructions on the storage device and executable by the processor, the set of instructions comprising: The world frame and a fusion routine, connected to the EM receiver and the totem IMU, that generates a fusion location of the totem in the world frame based on a combination of the EM waves and the totem IMU signal; A location correction routine, the location correction routine comprising: storing a rig frame, the rig frame being a mathematical object located between the eyepieces of the head unit and serving as a basis for defining where objects reside relative to the head unit; taking measurements to determine a rig frame attitude that the rig frame is assuming relative to the world frame; Deriving a receiver attitude relative to the world using known extrinsic properties as provided by a factory level calibration of the relationship between the EM receiver and the rig frame; estimating an EM relationship between the EM receiver and an EM transmitter to derive a transmitter attitude in the world; deriving an in-world totem IMU attitude as a non-fused attitude using known pertinent properties of the totem regarding the relationship between the EM transmitter and the totem IMU; a location correction routine, the location correction routine being executable to perform a method including: A set of instructions including: a data source carrying image data; a display system coupled to the data source and configured to use the image data to display a virtual object to a user, the location of the virtual object being based on the blended location of the totem; and A user interaction system comprising:
2. The set of instructions includes: a rendering engine connected to the data source and having an input for receiving the image data and an output to the display system, the rendering engine providing to the display system a data stream including the virtual object located at the blending location; The user interaction system of claim 1 further comprising:
3. The user interaction system of claim 1 , wherein the totem IMU reduces jitter of the virtual object.
4. The user interaction system of claim 1 , wherein the totem IMU includes at least one of a gyroscope and an accelerometer.
5. The user interaction system of claim 1 , wherein the EM receiver detects six degrees of freedom ("6dof") movement of the EM transmitter relative to the EM receiver.
6. The set of instructions includes: an unfused location determination modeler that determines a location of the totem relative to the head unit and a location of the head unit relative to the world frame and establishes an unfused location of the totem relative to the world frame; a comparator coupled to the fused location and the non-fused location for comparing the fused location with the non-fused location; a drift declarer coupled to the comparator for declaring a drift only if the fusion location is more than a predetermined distance from the non-fusion location; a location correction routine coupled to the drift declarer for resetting the totem IMU sensors to match the unfused location only if the drift is declared; The user interaction system of claim 1 further comprising:
7. The user interaction system of claim 6 , wherein the fusion routine executes at a first frequency and the non-fusion location determination modeler executes at a second frequency different from the first frequency.
8. A user interaction system according to claim 6 or 7, wherein the predetermined distance is less than 100 mm.
9. a camera on the head unit body, the camera positioned to capture an image of the totem; a simultaneous localization and mapping (SLAM) system connected to the camera and receiving the image of the totem, the SLAM system determining a non-fusion location based on the image of the totem; The user interaction system of claim 1 further comprising:
10. The display system comprises: a transparent waveguide fixed to the head unit body, the transparent waveguide passing light from the totem to the eye of a user wearing the head unit body; a projector that converts the image data into light, the light from the projector entering the transparent waveguide at an entrance pupil and exiting the transparent waveguide to the eye of the user at an exit pupil; and The user interaction system of claim 1 , comprising:
11. Head unit detection device for detecting movement of head mountable frame and wherein the set of instructions further comprises: a display adjustment algorithm, connected to the head unit detection device, for receiving measurements based on movements detected by the head unit detection device and for calculating a setpoint value; a rendering engine that modifies the position of the virtual object within the user's eye view based on the setting value; The user interaction system of claim 10, comprising:
12. The head unit detection device a head unit IMU mounted on the head-mountable frame, the head unit IMU including a motion sensor for detecting movement of the head-mountable frame; The user interaction system of claim 11 , comprising:
13. The user interaction system of claim 12 , wherein the head unit IMU includes at least one of a gyroscope and an accelerometer.
14. The head unit detection device a head unit camera mounted on the head-mountable frame, the head unit camera detecting movement of the head-mountable frame by capturing images of objects within a view of the head unit camera; the set of instructions comprising: a simultaneous localization and mapping (SLAM) system connected to the camera, receiving the image of the object and detecting a pose of the head unit body, the rendering engine modifying the position of the virtual object based on the pose. The user interaction system of claim 11 , comprising:
Citation Information
Patent Citations
Systems and methods for augmented reality
JP2018511122A
Self-referenced tracking
US20040201857A1
Electromagnetic tracking with augmented reality systems
US20170307891A1