Method and system for resolving hemispheric ambiguity using position vectors
The method and system resolve hemispheric ambiguity in AR and VR systems by using magnetic fields and dot product calculations to accurately determine object position and orientation, integrating multiple sensor data efficiently and reducing computational load.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2026-04-02
AI Technical Summary
Existing augmented reality (AR) and virtual reality (VR) systems face challenges in accurately determining the position and orientation of objects, leading to hemispheric ambiguity and inefficiencies in data interpretation, which affect system performance and user experience.
A method and system that utilize magnetic fields emitted by a handheld controller and detected by sensors in a headset to determine the position and orientation of the controller relative to the headset, employing a dot product calculation to resolve hemispheric ambiguity and integrate data from multiple sensors efficiently, with lower-frequency sensors correcting noise data from high-frequency sensors.
This approach enhances the accuracy and efficiency of object positioning in AR and VR systems by resolving hemispheric ambiguity and reducing computational load, thereby improving system performance and user experience.
Smart Images

Figure 0007839822000001 
Figure 0007839822000002 
Figure 0007839822000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims priority to U.S. Provisional Patent Application No. 62 / 702,339, filed Jul. 23, 2018, entitled "SYSTEMS AND METHODS FOR AUGMENTED REALITY", which is incorporated herein by reference in its entirety for all purposes as if fully set forth herein.
[0002] This application is related to U.S. Patent Application No. 15 / 859,277, filed Dec. 29, 2017, entitled "SYSTEMS AND METHODS FOR AUGMENTED REALITY", which is incorporated herein by reference in its entirety.
[0003] The present disclosure relates to systems and methods for determining the position and orientation of one or more objects in the context of an augmented reality system.
Background Art
[0004] Modern computing and display technologies have facilitated the development of systems for so - called "virtual reality" or "augmented reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or is perceived to be real. A virtual reality or "VR" scenario typically involves the presentation of digital or virtual image information without transparency to other actual real - world visual inputs. An augmented reality or "AR" scenario typically involves the presentation of digital or virtual image information as an augmentation to the visualization of the actual world around the user.
[0005] Despite the progress made in AR and VR systems, there is a need for better positioning systems in the context of AR and VR devices.
Summary of the Invention
[0006] The present invention relates to a system and method for optimally interpreting data inputs from multiple sensors. In other words, the embodiments described herein refine multiple inputs into a common coherent output using fewer computational resources than those used to correct data inputs from a single sensor input.
[0007] In one embodiment, a computer implementation method for resolving hemispheric ambiguity in a system comprising one or more sensors includes the step of emitting one or more magnetic fields in a handheld controller of the system. The method further includes the step of detecting one or more magnetic fields by one or more sensors positioned within a headset of the system. The method also includes the step of determining a first position and a first orientation of the handheld controller in a first hemisphere relative to the headset based on one or more magnetic fields, and the step of determining a second position and a second orientation of the handheld controller in a second hemisphere relative to the headset based on one or more magnetic fields. The second hemisphere is diametrically opposed to the first hemisphere relative to the headset. The method also includes the step of determining a normal vector relative to the headset and a position vector that identifies the position of the handheld controller relative to the headset in the first hemisphere. In addition, the method includes the step of calculating the dot product of the normal vector and the position vector, and determining that the first position and first orientation of the handheld controller are accurate when the result of the dot product is positive. The method also includes the step of determining that the second position and second orientation of the handheld controller are accurate when the result of the dot product is negative.
[0008] In one or more embodiments, the position vector is defined in the coordinate frame of the headset. The normal vector originates from the headset and extends at a predetermined angle from the horizontal line from the headset. In some embodiments, the predetermined angle is 45° downward from the horizontal line from the headset. According to some embodiments, when the result of the dot product is positive, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset. In some embodiments, when the result of the dot product is positive, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset. In one or more embodiments, the system is an optical device. According to some embodiments, the method is performed during the headset initialization process.
[0009] In some embodiments, the method also includes the step of delivering virtual content to a display based on a first position and a first orientation when the result of the dot product is positive, or a second position and a second orientation when the result of the dot product is negative.
[0010] In some embodiments, the data input from the first sensor is updated by a correction data input point from the second sensor. As noise data is collected by a high-frequency IMU, etc., it is periodically updated or adjusted to prevent excessive errors or drift from adversely affecting system performance or the interpretation of the data.
[0011] In some embodiments, the input to the first sensor is reset to originate from a correction input point, such as one provided by a lower-frequency and more accurate second sensor, such as a radar or vision system. These more accurate sensors operate at lower frequencies, saving computing cycles that would otherwise be necessary to operate them at full capacity, and their input is only needed to periodically update and correct underlying or noisier data, so lower-frequency operation does not affect system performance.
[0012] In another embodiment, the system includes a handheld controller comprising a magnetic field transmitter configured to emit one or more magnetic fields, a headset comprising one or more magnetic field sensors configured to detect one or more magnetic fields, and a processor coupled to the headset and configured to perform an operation. The operation includes the steps of determining a first position and a first orientation of the handheld controller in a first hemisphere relative to the headset based on one or more magnetic fields, and determining a second position and a second orientation of the handheld controller in a second hemisphere relative to the headset based on one or more magnetic fields. The second hemisphere is diametrically opposed to the first hemisphere relative to the headset. The operation also includes the steps of determining a normal vector relative to the headset and a position vector that identifies the position of the handheld controller relative to the headset in the first hemisphere. In addition, the operation includes the steps of calculating the dot product of the normal vector and the position vector and determining that the first position and first orientation of the handheld controller is accurate when the result of the dot product is positive. The operation also includes the steps of determining that the second position and second orientation of the handheld controller is accurate when the result of the dot product is negative.
[0013] In yet another embodiment, the computer program product is embodied in a non-transient computer-readable medium having a sequence of instructions stored thereon that, when executed by a processor, cause the processor to execute a method for resolving hemispheric ambiguity in a system comprising one or more sensors, which includes the step of emitting one or more magnetic fields in a handheld controller of the system. The method further includes the step of detecting one or more magnetic fields by one or more sensors positioned in a headset of the system. The method also includes the step of determining a first position and a first orientation of the handheld controller in a first hemisphere relative to the headset based on one or more magnetic fields, and the step of determining a second position and a second orientation of the handheld controller in a second hemisphere relative to the headset based on one or more magnetic fields. The second hemisphere is diametrically opposed to the first hemisphere relative to the headset. The method also includes the step of determining a normal vector relative to the headset and a position vector that identifies the position of the handheld controller relative to the headset in the first hemisphere. In addition, the method includes the steps of calculating the dot product of the normal vector and the position vector, and determining that the first position and first orientation of the handheld controller are accurate when the result of the dot product is positive. The method also includes the step of determining that the second position and second orientation of the handheld controller are accurate when the result of the dot product is negative.
[0014] In some embodiments, a computer implementation method includes the step of detecting one or more magnetic fields emitted by electronic devices in the system environment based on data output by one or more sensors positioned within the system's headset. The method also includes the step of determining a first position and a first orientation of the electronic device in a first hemisphere relative to the headset based on the one or more magnetic fields. The method further includes the step of determining a normal vector and a position vector that identifies the position of the electronic device in the first hemisphere relative to the headset. The method also includes the step of calculating the dot product of the normal vector and the position vector, and the step of determining whether the value of the calculated dot product is positive or negative. The method further includes delivering virtual content to a display, at least partially, based on the step of determining whether the value of the calculated dot product is positive or negative.
[0015] In one or more embodiments, the electronic device comprises a handheld controller, a wearable device, or a mobile computing device. In some embodiments, the step of delivering virtual content to the headset display in response to a determination that the value of the calculated dot product is positive includes, at least in part, the step of delivering virtual content to the headset display based on a first position and a first orientation of the electronic device.
[0016] In some embodiments, the method includes the step of determining a second position and a second orientation of an electronic device in a second hemisphere relative to a headset based on one or more magnetic fields. The second hemisphere is diametrically opposed to the first hemisphere relative to the headset. In response to the determination that the calculated dot product value is negative, the step of delivering virtual content to the headset's display includes, at least in part, the step of delivering virtual content to the headset's display based on the second position and second orientation of the electronic device.
[0017] In some embodiments, noise data is adjusted by a coefficient value to preemptively adjust incoming data points provided by the sensor. As the correction data points are received, the system "directs" the incoming noise data toward the correction input point rather than fully adjusting the noise data toward the correction input point. These embodiments are particularly useful when there are large changes in both sensor inputs, as the noise data stream directed toward the correction input will not originate from past correction input points that are substantially different from what the current measurement would show. In other words, the noise data stream will not originate from obsolete correction input points.
[0018] In some embodiments, pose prediction is performed by estimating the user's future location and accessing expected features and points at that future location. For example, if a user is walking around a square table, features such as the corners of the table or lines of objects on the table are "fetched" by the system based on the system's estimated location where the user will be at a future time. When the user is at that location, an image is collected, the fetched features are projected onto the image, correlations are determined, and a specific pose is determined. This is beneficial because it avoids simultaneous feature mapping with image reception and reduces computation cycles by completing preprocessing (such as warping) of fetched features prior to image reception, allowing points to be applied more quickly when an image of the current pose is collected, the estimated pose to be refined rather than generated, and virtual content to be rendered to its new pose more quickly or with less jitter.
[0019] Additional embodiments, advantages, and details are described in more detail below, with specific reference to the following figures as necessary. The present invention provides, for example, the following: (Item 1) A method for resolving hemispherical ambiguity in a system comprising one or more sensors, wherein the method is: In the handheld controller of the system, emitting one or more magnetic fields, detecting the one or more magnetic fields by one or more sensors positioned within the headset of the system, determining a first position and a first orientation of the handheld controller within a first hemisphere with respect to the headset based on the one or more magnetic fields, determining a second position and a second orientation of the handheld controller within a second hemisphere with respect to the headset based on the one or more magnetic fields, wherein the second hemisphere is diametrically opposed to the first hemisphere with respect to the headset, determining a normal vector with respect to the headset and a position vector identifying the position of the handheld controller with respect to the headset within the first hemisphere, calculating a dot product of the normal vector and the position vector, when the result of the dot product is positive, determining that the first position and the first orientation of the handheld controller are accurate, when the result of the dot product is negative, determining that the second position and the second orientation of the handheld controller are accurate A method comprising. (Item 2) The method according to item 1, wherein the position vector is defined in the coordinate frame of the headset. (Item 3) The method according to item 1, wherein the normal vector originates from the headset and extends at a predetermined angle from a horizontal line from the headset. (Item 4) The method according to item 3, wherein the predetermined angle is 45° downward from the horizontal line from the headset. (Item 5) The method according to item 1, wherein when the result of the dot product is positive, the first hemisphere is identified as the front hemisphere with respect to the headset, and the second hemisphere is identified as the rear hemisphere with respect to the headset. (Item 6) The method according to item 1, wherein when the result of the dot product is negative, the second hemisphere is identified as the front hemisphere with respect to the headset, and the first hemisphere is identified as the rear hemisphere with respect to the headset. (Item 7) The first position and the first orientation of the handheld controller within the first hemisphere with respect to the headset when the result of the dot product is positive, or The second position and the second orientation of the handheld controller within the second hemisphere with respect to the headset when the result of the dot product is negative The method according to item 1, further comprising delivering virtual content to a display based on (Item 8) The system is an optical device, the method according to item 1. (Item 9) The method is performed during an initialization process of the headset, the method according to item 1. (Item 10) A system, comprising A handheld controller comprising a magnetic field transmitter configured to emit one or more magnetic fields, and A headset comprising one or more magnetic field sensors configured to detect the one or more magnetic fields, and A processor coupled to the headset, configured to Determine a first position and a first orientation of the handheld controller within a first hemisphere with respect to the headset based on the one or more magnetic fields; and Determine a second position and a second orientation of the handheld controller within a second hemisphere with respect to the headset, the second hemisphere being diametrically opposed to the first hemisphere with respect to the headset, based on the one or more magnetic fields; and Determine a normal vector with respect to the headset and a position vector identifying the position of the handheld controller with respect to the headset within the first hemisphere; and Calculate a dot product of the normal vector and the position vector; and When the result of the dot product is positive, it is determined that the first position and first orientation of the handheld controller are accurate. When the result of the dot product is negative, it is determined that the second position and second orientation of the handheld controller are accurate. A processor and configured to perform operations including A system that includes these features. (Item 11) The position vector is defined in the coordinate frame of the headset, according to the system described in item 10. (Item 12) The system according to item 10, wherein the normal vector originates from the headset and extends at a predetermined angle from the horizontal line from the headset. (Item 13) The system according to item 12, wherein the predetermined angle is at a 45° angle downward from the horizontal line from the headset. (Item 14) The system according to item 10, wherein when the result of the dot product is positive, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset. (Item 15) The system according to item 10, wherein when the result of the dot product is negative, the second hemisphere is identified as the front hemisphere relative to the headset, and the first hemisphere is identified as the rear hemisphere relative to the headset. (Item 16) The aforementioned processor further, The first position and the first orientation when the result of the dot product is positive, or The second position and second orientation when the result of the dot product is negative. Based on this, deliver virtual content to the display. The system described in item 10, configured to perform operations including those described above. (Item 17) The system is an optical device, as described in item 10. (Item 18) The method described above is performed during the initialization process of the headset, according to the system described in item 10. (Item 19) A computer program product embodied in a non-transient computer-readable medium, wherein the computer-readable medium has a sequence of instructions stored thereon, and when the sequence of instructions is executed by a processor, the processor causes the processor to perform a method for resolving hemispheric ambiguity in a system comprising one or more sensors, the method is In the handheld controller of the aforementioned system, one or more magnetic fields are emitted, One or more sensors positioned within the headset of the system detect the one or more magnetic fields, Based on the one or more magnetic fields, the first position and first orientation of the handheld controller within the first hemisphere relative to the headset are determined. Based on the one or more magnetic fields, the second position and second orientation of the handheld controller within the second hemisphere relative to the headset are determined, wherein the second hemisphere is diametrically opposed to the first hemisphere relative to the headset. Determining the normal vector to the headset and the position vector that identifies the position of the handheld controller relative to the headset within the first hemisphere, Calculating the dot product of the normal vector and the position vector, When the result of the dot product is positive, it is determined that the first position and first orientation of the handheld controller are accurate. When the result of the dot product is negative, it is determined that the second position and second orientation of the handheld controller are accurate. Computer program products, including [this]. [Brief explanation of the drawing]
[0020] [Figure 1] Figure 1 illustrates several embodiments of augmented reality scenarios involving a virtual reality object.
[0021] [Figure 2A] Figures 2A-2D illustrate various configurations of components comprising a visual display system according to several embodiments. [Figure 2B] Figures 2A-2D illustrate various configurations of components comprising a visual display system according to several embodiments. [Figure 2C] Figures 2A-2D illustrate various configurations of components comprising a visual display system according to several embodiments. [Figure 2D] Figures 2A-2D illustrate various configurations of components comprising a visual display system according to several embodiments.
[0022] [Figure 3] Figure 3 illustrates remote interaction with cloud computing assets in several embodiments.
[0023] [Figure 4] Figure 4 illustrates an electromagnetic tracking system according to several embodiments.
[0024] [Figure 5] Figure 5 illustrates electromagnetic tracking methods according to several embodiments.
[0025] [Figure 6] Figure 6 illustrates an electromagnetic tracking system coupled to a visual display system, according to several embodiments.
[0026] [Figure 7]Figure 7 illustrates a method for determining the metric of a visual display system coupled to an electromagnetic emitter, according to several embodiments.
[0027] [Figure 8] Figure 8 illustrates a visual display system with various sensing components and accessories according to several embodiments.
[0028] [Figure 9A] Figures 9A-9F illustrate various control modules according to various embodiments. [Figure 9B] Figures 9A-9F illustrate various control modules according to various embodiments. [Figure 9C] Figures 9A-9F illustrate various control modules according to various embodiments. [Figure 9D] Figures 9A-9F illustrate various control modules according to various embodiments. [Figure 9E] Figures 9A-9F illustrate various control modules according to various embodiments. [Figure 9F] Figures 9A-9F illustrate various control modules according to various embodiments.
[0029] [Figure 10] Figure 10 illustrates several embodiments of head-mounted visual displays with minimized shape factors.
[0030] [Figure 11A] Figures 11A and 11B illustrate various configurations of the electromagnetic sensing module. [Figure 11B] Figures 11A and 11B illustrate various configurations of the electromagnetic sensing module.
[0031] [Figure 12] Figures 12A-12E illustrate various configurations for an electromagnetic sensor core according to several embodiments.
[0032] [Figure 13A] Figures 13A-13C illustrate various time-division multiplexing methods for electromagnetic sensing according to several embodiments. [Figure 13B] Figures 13A-13C illustrate various time-division multiplexing methods for electromagnetic sensing according to several embodiments. [Figure 13C] Figures 13A-13C illustrate various time-division multiplexing methods for electromagnetic sensing according to several embodiments.
[0033] [Figure 14] Figures 14-15 illustrate methods for combining various sensor data in response to the activation of a visual display system, according to several embodiments. [Figure 15] Figures 14-15 illustrate methods for combining various sensor data in response to the activation of a visual display system, according to several embodiments.
[0034] [Figure 16A] Figures 16A-16B illustrate a visual display system with various sensing and imaging components and accessories according to several embodiments. [Figure 16B] Figures 16A-16B illustrate a visual display system with various sensing and imaging components and accessories according to several embodiments.
[0035] [Figure 17A] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17B] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17C] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17D]Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17E] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17F] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments. [Figure 17G] Figures 17A-17G illustrate various configurations of transmission coils in an electromagnetic tracking system according to several embodiments.
[0036] [Figure 18A] Figures 18A-18C illustrate the signal interference effects from various system inputs in several embodiments. [Figure 18B] Figures 18A-18C illustrate the signal interference effects from various system inputs in several embodiments. [Figure 18C] Figures 18A-18C illustrate the signal interference effects from various system inputs in several embodiments.
[0037] [Figure 19] Figure 19 illustrates calibration configurations according to several embodiments.
[0038] [Figure 20A] Figures 20A-20C illustrate various summer configurations, such as those between multiple subsystems. [Figure 20B] Figures 20A-20C illustrate various summer configurations, such as those between multiple subsystems. [Figure 20C] Figures 20A-20C illustrate various summer configurations, such as those between multiple subsystems.
[0039] [Figure 21] Figure 21 illustrates signal overlap of multiple inputs with various signal frequencies.
[0040] [Figure 22A] Figures 22A-22C illustrate various arrays of electromagnetic sensing modules according to several embodiments. [Figure 22B] Figures 22A-22C illustrate various arrays of electromagnetic sensing modules according to several embodiments. [Figure 22C] Figures 22A-22C illustrate various arrays of electromagnetic sensing modules according to several embodiments.
[0041] [Figure 23A] Figures 23A-23C illustrate the recalibration of a sensor with a given known input according to several embodiments. [Figure 23B] Figures 23A-23C illustrate the recalibration of a sensor with a given known input according to several embodiments. [Figure 23C] Figures 23A-23C illustrate the recalibration of a sensor with a given known input according to several embodiments.
[0042] [Figure 24A] Figures 24A-24D illustrate the steps for determining the variables in a calibration protocol according to several embodiments. [Figure 24B] Figures 24A-24D illustrate the steps for determining the variables in a calibration protocol according to several embodiments. [Figure 24C] Figures 24A-24D illustrate the steps for determining the variables in a calibration protocol according to several embodiments. [Figure 24D] Figures 24A-24D illustrate the steps for determining the variables in a calibration protocol according to several embodiments.
[0043] [Figure 25A] Figures 25A-25E illustrate potential misreadings based on a given sensor input and applied solution in several embodiments. [Figure 25B]Figures 25A-25E illustrate potential misreadings based on a given sensor input and applied solution in several embodiments. [Figure 25C] Figures 25A-25E illustrate potential misreadings based on a given sensor input and applied solution in several embodiments. [Figure 25D] Figures 25A-25E illustrate potential misreadings based on a given sensor input and applied solution in several embodiments. [Figure 25E] Figures 25A-25E illustrate potential misreadings based on a given sensor input and applied solution in several embodiments.
[0044] [Figure 25F] Figure 25F illustrates a flowchart illustrating the steps of determining the correct orientation of a handheld controller using position vectors relating to a handheld device relative to a headset, according to several embodiments.
[0045] [Figure 26] Figure 26 illustrates feature matching between two images in several embodiments.
[0046] [Figure 27A] Figures 27A and 27B illustrate methods for determining attitude based on sensor input, according to several embodiments. [Figure 27B] Figures 27A and 27B illustrate methods for determining attitude based on sensor input, according to several embodiments.
[0047] [Figure 28A] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28B] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28C]Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28D] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28E] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28F] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments. [Figure 28G] Figures 28A-28G illustrate various sensor fusion corrections according to several embodiments.
[0048] [Figure 29] Figure 29 illustrates several embodiments of single-path multilayer convolutional computing architectures.
[0049] [Figure 30A] Figures 30A-30E illustrate various coil configurations for an electromagnetic tracking system according to several embodiments. [Figure 30B] Figures 30A-30E illustrate various coil configurations for an electromagnetic tracking system according to several embodiments. [Figure 30C] Figures 30A-30E illustrate various coil configurations for an electromagnetic tracking system according to several embodiments. [Figure 30D] Figures 30A-30E illustrate various coil configurations for an electromagnetic tracking system according to several embodiments. [Figure 30E] Figures 30A-30E illustrate various coil configurations for an electromagnetic tracking system according to several embodiments.
[0050] [Figure 31] Figure 31 illustrates a simplified computer system according to an embodiment described herein. [Modes for carrying out the invention]
[0051] For example, referring to Figure 1, an augmented reality scene (4) is depicted, and the user of AR technology sees a real-world park-like setting (6) featuring people, trees, buildings in the background, and a concrete platform (1120). In addition to these items, the user of AR technology also perceives "seeing" a robotic figure (1110) standing on the real-world platform (1120) and a flying cartoon-like avatar character (2) that looks like an anthropomorphic bumblebee, although these elements (2, 1110) do not exist in the real world. In conclusion, because the human visual perception system is complex, producing VR or AR technology that facilitates a comfortable, natural, and rich presentation of virtual image elements among other virtual or real-world image elements is difficult.
[0052] For example, a head-mounted AR display (or helmet-mounted display or smart glasses) is typically at least loosely attached to the user's head and therefore may move as the user's head moves. If the user's head movement is detected by the display system, the displayed data can be updated to take into account the change in head posture.
[0053] As an example, when a user wearing a head-mounted display views a virtual representation of a three-dimensional (3-D) object on the display and walks around the area in which the 3-D object appears, the 3-D object can be re-rendered for each viewpoint, giving the user the perception that they are walking around an object occupying real space. When the head-mounted display is used to present multiple objects (e.g., a rich virtual world) in a virtual space, measurement of head posture (i.e., the location and orientation of the user's head) can be used to re-render the scene to match the user's dynamically changing head location and orientation, thereby providing an increased sense of immersion in the virtual space.
[0054] In AR systems, detecting or calculating head pose can facilitate the rendering of virtual objects so that the display system appears to occupy space in the real world in a manner that makes sense to the user. In addition, detecting the position and / or orientation of real objects such as the user's head or a handheld device (which may also be called a "totem"), a tactile device, or other real physical objects in conjunction with the AR system can also facilitate the display system presenting display information to the user and enabling the user to interact efficiently with certain aspects of the AR system. As the user's head moves around in the real world, virtual objects can be re-rendered as a function of head pose so that the virtual objects appear stable relative to the real world. At least for AR applications, the placement of virtual objects spatially coupled with physical objects (e.g., presented to appear spatially close to physical objects in two or three dimensions) cannot be a trivial matter.
[0055] For example, head movement can significantly complicate the placement of virtual objects from the perspective of the surrounding environment. This applies whether the perspective is captured as an image of the surrounding environment and then projected or displayed to the end user, or whether the end user directly perceives the perspective of the surrounding environment. For instance, head movement is likely to change the end user's field of view, which would likely require updates to the locations where various virtual objects appear in the end user's field of view.
[0056] In addition, head movement can occur at a wide variety of ranges and speeds. Head movement speed can fluctuate not only between different head movements, but also within or across a single head movement. For example, head movement speed may initially increase from the starting point (e.g., linear or nonlinear), decrease as it approaches the endpoint, and reach a maximum speed at some point between the start and end points of the head movement. High-speed head movement may even exceed the capabilities of certain display or projection techniques, rendering images that appear to the end user as uniform and / or smooth motion.
[0057] Head tracking accuracy and latency (i.e., the time elapsed between a user moving their head and the image being updated and displayed to the user) are challenges for VR and AR systems. Particularly for display systems that fill a substantial portion of the user's field of view with virtual elements, high head tracking accuracy and very short overall system latency—from the initial detection of head movement to the update of light delivered to the user's visual system by the display—are crucial. Long latency can lead to mismatches between the user's vestibular and photosensory systems, generating user perception scenarios that may result in motion sickness or 3D sickness. High system latency can also cause the apparent location of virtual objects to appear unstable during rapid head movements.
[0058] In addition to head-mounted display systems, other display systems can also benefit from accurate and low-latency head pose detection. These include head-tracking display systems where the display is not attached to the user's body but mounted, for example, on a wall or other surface. The head-tracking display can act like a window on a scene, and as the user moves their head relative to the "window," the scene is re-rendered to match the user's changing viewpoint. Another system is a head-mounted projection system where the head-mounted display projects light onto the real world.
[0059] In addition, to provide a realistic augmented reality experience, the AR system may be designed to interact with the user. For example, multiple users may play a ball game using a virtual ball and / or other virtual objects. One user may "catch" the virtual ball and throw it back to another user. In another embodiment, the first user may be provided with a totem (e.g., a real bat communicatively coupled to the AR system) for hitting the virtual ball. In yet another embodiment, a virtual user interface may be presented to the AR user to allow the user to select one of many options. The user may interact with the system using a totem, a tactile device, a wearable component, or simply by touching a virtual screen.
[0060] Detecting the user's head posture and orientation, as well as the physical location of real objects in space, enables AR systems to display virtual content in an effective and engaging manner. However, while these capabilities are crucial for AR systems, they are difficult to achieve. In other words, AR systems must recognize the physical location of real objects (e.g., the user's head, totems, haptic devices, wearable components, the user's hands, etc.) and correlate the physical coordinates of these real objects to the virtual coordinates of one or more virtual objects displayed to the user. This requires highly accurate sensors and sensor recognition systems that can track the position and orientation of one or more objects at a high rate. Current approaches do not perform localization at satisfactory speed or accuracy standards.
[0061] Referring to Figures 2A-2D, several common component options are illustrated. In the detailed explanation section following the discussion in Figures 2A-2D, various systems, subsystems, and components are presented to address the objective of providing a high-quality and comfortably perceived display system for human VR and / or AR.
[0062] As shown in Figure 2A, the AR system user (60) is depicted wearing a head-mounted component (58) characterized by a frame (64) structure coupled to a display system (62) positioned in front of the user's eyes. A speaker (66) is coupled to the frame (64) in the depicted configuration and positioned adjacent to the user's ear canal (in one embodiment, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display (62) is operably coupled (68) to a local processing and data module (70) by wired conductors or wireless connectivity, and may be mounted in various configurations, such as being fixedly attached to a frame (64), fixedly attached to a helmet or hat (80) as shown in the embodiment of Figure 2B, built into headphones, detachably attached to the user's (60) torso (82) in a backpack configuration as shown in the embodiment of Figure 2C, or detachably attached to the user's (60) waist (84) in a belt-mounted configuration as shown in the embodiment of Figure 2D.
[0063] The local processing and data module (70) may include a power-efficient processor or controller and digital memory such as flash memory, both of which may be used to assist in processing, caching, and storing data that can be obtained and / or processed using the remote processing module (72) and / or remote data repository (74) for a) data captured from sensors that can be operably coupled to the frame (64), such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope, and / or b) possibly for passage to the display (62) after processing or reading.
[0064] The local processing and data module (70) may be operably coupled (76, 78) to the remote processing modules (72, 74) and remote data repositories (74) via wired or wireless communication links, etc., so that the remote modules (72, 74) are operably coupled to each other and available as resources to the local processing and data module (70).
[0065] In one embodiment, the remote processing module 72 may comprise one or more relatively high-performance processors or controllers configured to analyze and process data and / or image information. In one embodiment, the remote data repository (74) may comprise a relatively large digital data storage facility, which may be available through the Internet or other networking configurations in a “cloud” resource configuration. In one embodiment, all data may be stored, and all calculations may be performed in local processing and data modules, enabling fully autonomous use from any remote module.
[0066] Referring here to Figure 3, the schematic diagram illustrates the coordination between a cloud computing asset (46) and a local processing asset, which may reside, for example, in a head-mounted component (58) coupled to the user's head (120) and a local processing and data module (70) coupled to the user's belt (308). Thus, the component 70 may also be referred to as a “beltpack” 70, as shown in Figure 3. In one embodiment, one or more cloud (46) assets, such as a server system (110), are operably coupled (115) directly to one or both (40, 42) local computing assets, such as a processor and memory configuration, coupled to the user's head (120) and belt (308), as described above, via wired or wireless networking, etc. (wireless is preferred for mobile applications, and wired is preferred for certain high bandwidth or high data transfer). These computing assets local to the user may also be operably coupled to one another via wired and / or wireless connectivity configurations (44), such as wired coupling (68) as discussed below with reference to Figure 8. In one embodiment, in order to maintain a low-inertia and small subsystem mounted on the user's head (120), primary transmission between the user and the cloud (46) may also be via a link between the belt-mounted subsystem and the cloud, and the head-mounted subsystem (120) is data-tethered to the belt-based subsystem (308) primarily using wireless connectivity such as ultra-wideband ("UWB") connectivity, as currently employed in personal computing peripheral connectivity applications.
[0067] Through efficient local and remote processing coordination and the use of appropriate display devices for the user, such as the user interface or user display system (62) shown in Figure 2A or its variations, an aspect of one world relating to the user's current real or virtual location can be transferred or “passed” to the user and updated in an efficient manner. In other words, the world map can be continuously updated in a storage location that may be partially residing on the user’s AR system and partially residing in cloud resources. The map (also referred to as the “passable world model”) may be a large database containing raster images, 3-D and 2-D points, parameter information, and other information about the real world. As more and more AR users continuously capture information about their real environment (e.g., through cameras, sensors, IMUs, etc.), the map becomes increasingly accurate and complete.
[0068] By using the aforementioned configuration, where there exists a single world model residing on cloud computing resources and capable of being delivered from there, such a world can become "passable" to one or more users in a relatively low-bandwidth form, which is preferable for attempting to transmit real-time video data or equivalent. The augmented experience of a person standing near the statue (for example, as shown in Figure 1) may be informed by the cloud-based world model, a subset of which may be passed to them and their local display device to complete the view. A person seated in front of a remote display device, which may be as simple as a personal computer on a desk, can efficiently download the same section of that information from the cloud and have it rendered on their display. In fact, a person actually present in a park near the statue may take a friend, located remotely, for a walk in the park, with the friend joining in through virtual and augmented reality. The system would need to know the locations of the streets, trees, and statues, but using that information in the cloud, the joining friend can download it from the cloud side of the scenario and then begin the walk as local augmented reality for the person actually in the park.
[0069] 3D points may be captured from the environment, and the pose of the camera capturing those images or points (i.e., vectors and / or primordial information relative to the world) may be determined so that these points or images can be “tagged” or associated with this pose information. Points captured by the second camera may then be used to determine the pose of the second camera. In other words, the second camera can be oriented and / or located based on a comparison with the tagged images from the first camera. This knowledge may then be used to extract textures, create maps, and create virtual copies of the real world (since there are now two cameras aligned to the surroundings).
[0070] Therefore, at a basic level, in one embodiment, a wearable system can be used to capture both 3-D points and the 2-D images that generated those points, and these points and images may be sent to cloud storage and processing resources. They may also be cached locally along with built-in pose information (i.e., cache tagged images). Thus, the cloud can be ready (i.e., in the available cache) for tagged 2-D images (e.g., tagged with 3-D pose) along with 3-D points. If the user is observing something dynamic, additional information about the motion may also be sent to the cloud (e.g., if the user is looking at another person's face, the user can capture a texture map of the face and push it at an optimized frequency, even if the surrounding world is otherwise essentially static). Further information relating to object recognition devices and passable world models can be found in the following additional disclosures relating to augmented and virtual reality systems, such as those developed by Magic Leap, Inc. (Fort Lauderdale, Florida), namely U.S. Patent Applications 14 / 641,376, 14 / 555,585, 14 / 212,961, 14 / 690,401, 13 / 663,466, and 13 / 684,489, along with U.S. Patent Application 14 / 205,126, entitled "System and method for augmented and virtual reality" (which is incorporated herein in its entirety by reference), and U.S. GPS and other location information may be used as input to such processing. Highly accurate location of the user's head, totem, hand gestures, tactile devices, etc., is essential for displaying appropriate virtual content to the user.
[0071] One approach to achieving high-precision localization may involve the use of electromagnetic fields coupled with electromagnetic sensors strategically placed on the user's AR headset, beltpack, and / or other auxiliary devices (e.g., totems, haptic devices, gaming devices, etc.).
[0072] An electromagnetic tracking system typically comprises at least an electromagnetic field emitter and at least one electromagnetic field sensor. The sensor may measure an electromagnetic field with a known distribution. Based on these measurements, the position and orientation of the electromagnetic field sensor relative to the emitter are determined.
[0073] Referring to Figure 4, here is an illustrative system diagram of an electromagnetic tracking system (for example, Johnson Examples of such systems are illustrated (developed by organizations such as Biosense (RTM), a subsidiary of Johnson & Johnson Corporation, Polhemus (RTM), Inc. (Colchester, Vermont), and manufactured by Sixense (RTM) Entertainment, Inc. (Los Gatos, California), and other tracking device manufacturers). In one or more embodiments, the electromagnetic tracking system comprises an electromagnetic field emitter 402 configured to emit a known magnetic field. As shown in Figure 4, the electromagnetic field emitter may be coupled to a power source (e.g., electric current, battery, etc.) to provide power to the emitter 402.
[0074] In one or more embodiments, the electromagnetic field emitter 402 comprises several coils (e.g., at least three coils positioned perpendicular to each other and generating fields in the x, y, and z directions) that generate a magnetic field. This magnetic field is used to establish a coordinate space. This allows the system to map the position of the sensor relative to a known magnetic field and helps determine the position and / or orientation of the sensor. In one or more embodiments, electromagnetic sensors 404a, 404b, etc., may be attached to one or more real objects. The electromagnetic sensor 404 may comprise smaller coils that can be induced through an electromagnetic field from which an electric current is emitted.
[0075] Generally, the “sensor” component (404) may comprise a small coil or loop, such as a set of three coils oriented in different ways (i.e., oriented orthogonally to each other, etc.), coupled together within a small structure such as a cube or other container that is positioned / oriented to capture the magnetic flux flowing in from the magnetic field emitted by the emitter (402), and the relative position and orientation of the sensor with respect to the emitter can be calculated by comparing the currents induced through these coils and by determining the relative positions and orientations of the coils with respect to each other.
[0076] One or more parameters relating to the behavior of a coil and an inertial measuring unit ("IMU") component operably coupled to an electromagnetic tracking sensor may be measured to detect the position and / or orientation of the sensor (and the object to which it is attached) relative to the coordinate system to which the electromagnetic field emitter is coupled. In one or more embodiments, multiple sensors may be used in conjunction with the electromagnetic emitter to detect the position and orientation of each sensor in coordinate space. The electromagnetic tracking system may provide position in three directions (i.e., X, Y, and Z directions) and further in two or three orientation angles. In one or more embodiments, IMU measurements may be compared with coil measurements to determine the position and orientation of the sensor. In one or more embodiments, both electromagnetic (EM) data and IMU data may be combined with various other data sources, such as cameras, depth sensors, and other sensors, to determine the position and orientation. This information may be transmitted to a controller 406 (e.g., wireless communication, Bluetooth®, etc.). In one or more embodiments, attitude (or position and orientation) may be reported at a relatively high refresh rate in conventional systems.
[0077] Conventionally, electromagnetic emitters are coupled to relatively stable, large objects such as tables, operating tables, walls, or ceilings, while one or more sensors are coupled to smaller objects such as medical devices, handheld game components, or equivalents. Alternatively, various features of electromagnetic tracking systems may be employed to create a configuration in which changes or deltas in position and / or orientation between two objects moving in space relative to a more stable global coordinate system, as described below with reference to Figure 6. In other words, the configuration is shown in Figure 6, where variations in the electromagnetic tracking system are utilized to track position and orientation deltas between a head-mounted component and a handheld component, while the orientation of the head relative to a global coordinate system (e.g., the user's local indoor environment) is determined separately by simultaneous localization and mapping ("SLAM") techniques, etc., using an outward-facing capture camera that may be coupled to the head-mounted component of the system.
[0078] The controller 406 may control the electromagnetic field generator 402 and may also capture data from various electromagnetic sensors 404. It should be understood that the various components of the system may be coupled to each other through any electromechanical or wireless / Bluetooth® means. The controller 406 may also have data on known magnetic fields and the coordinate space associated with them. This information is then used to detect the position and orientation of the sensors in relation to the coordinate space corresponding to the known electromagnetic field.
[0079] One advantage of electromagnetic tracking systems is that they can produce highly accurate tracking results with minimal latency and high resolution. In addition, electromagnetic tracking systems do not necessarily rely on optical tracking devices, and sensors / objects not within the user's line of sight can be easily tracked.
[0080] It should be understood that the strength of the electromagnetic field v decreases as a cubic function of the distance r from the coil transmitter (e.g., electromagnetic field emitter 402). Therefore, the algorithm may be required based on the distance from the electromagnetic field emitter. The controller 406 may be configured to use such an algorithm to determine the position and orientation of the sensor / object at a variable distance from the electromagnetic field emitter.
[0081] Assuming a rapid decrease in electromagnetic field strength as movement away from the electromagnetic emitter occurs, the best results in terms of accuracy, efficiency, and low latency can be achieved at closer distances. In a typical electromagnetic tracking system, the electromagnetic field emitter has a sensor powered by an electric current (e.g., a plug-in power source) and located within a 20-foot radius of the electromagnetic field emitter. A shorter radius between the sensor and the field emitter may be more desirable in many applications, including AR applications.
[0082] Referring here to Figure 5, an illustrative flowchart illustrating the function of a typical electromagnetic tracking system is briefly shown. In 502, a known electromagnetic field is emitted. In one or more embodiments, the magnetic field emitter may generate a magnetic field, and each coil may generate an electric field in one direction (e.g., x, y, or z). The magnetic field may be generated using an arbitrary waveform.
[0083] In one or more embodiments, each axis may oscillate at a slightly different frequency. In 504, the coordinate space corresponding to the electromagnetic field can be determined. For example, the control 406 in Figure 4 may automatically determine the coordinate space around the emitter based on the electromagnetic field.
[0084] In 506, the behavior of a coil in a sensor (which may be attached to a known object) can be detected. For example, the current induced in the coil may be calculated. In other embodiments, the rotation of the coil or any other quantifiable behavior may be tracked and measured. In 508, this behavior may be used to detect the position and orientation of the sensor and / or a known object. For example, the controller 406 may refer to a mapping table that correlates the behavior of the coil in the sensor to various positions or orientations. Based on these calculations, the position in coordinate space may be determined along with the orientation of the sensor.
[0085] In the context of AR systems, one or more components of an electromagnetic tracking system may need to be modified to facilitate accurate tracking of mobile components. As mentioned above, tracking the user's head posture and orientation is essential in many AR applications. Accurate determination of the user's head posture and orientation allows the AR system to display the correct virtual content to the user. For example, a virtual scene may include a monster hidden behind a real building. Depending on the user's head posture and orientation relative to the building, the view of the virtual monster may need to be modified to provide a realistic AR experience. Alternatively, the position and / or orientation of a totem, haptic device, or any other means of interacting with virtual content may be important in enabling an AR user to interact with the AR system. For example, in many gaming applications, the AR system must detect the position and orientation of real objects relative to the virtual content. Or, when displaying a virtual interface, the position of a totem, the user's hand, a haptic device, or any other real object configured for interaction with the AR system must be known in relation to the displayed virtual interface so that the system can understand commands, etc. Conventional localization methods, including optical tracking and other methods, typically suffer from high latency and low resolution issues, making the rendering of virtual content difficult in many augmented reality applications.
[0086] In one or more embodiments, as discussed in relation to Figures 4 and 5, the electromagnetic tracking system may be adapted to the AR system to detect the position and orientation of one or more objects in relation to an emitted electromagnetic field.
[0087] Typical electromagnetic systems tend to have large and bulky electromagnetic emitters (e.g., 402 in Figure 4), which poses a problem for AR devices. However, smaller electromagnetic emitters (e.g., within the millimeter range) may be used to emit known electromagnetic fields in the context of AR systems.
[0088] Referring here to Figure 6, the electromagnetic tracking system may be incorporated with the AR system, as shown, along with an electromagnetic field emitter 602 incorporated as part of a handheld controller 606. In one or more embodiments, the handheld controller may be a totem for use in a game scenario. In other embodiments, the handheld controller may be a tactile device. In yet another embodiment, the electromagnetic field emitter may simply be incorporated as part of a beltpack 70. The handheld controller 606 may include a battery 610 or other power source for powering the electromagnetic field emitter 602. It should be understood that the electromagnetic field emitter 602 may also include, or be coupled to, an IMU component 650 configured to assist in determining the position and / or orientation of the electromagnetic field emitter 602 relative to other components. This may be particularly important if both the electromagnetic field emitter 602 and the sensor (604) are mobile. As shown in the embodiment of Figure 6, placing the electromagnetic field emitter 602 in the handheld controller rather than in the beltpack ensures that the electromagnetic field emitter does not compete for resources in the beltpack, but rather uses its own battery source in the handheld controller 606.
[0089] In one or more embodiments, the electromagnetic sensor 604 may be installed in one or more locations on the user's headset, along with other sensing devices such as one or more IMUs or additional flux-capturing coils (608). For example, as shown in Figure 6, the sensors (604, 608) may be installed on both sides of the headset (58). Since these sensors can be fabricated to be very small (and therefore potentially low-sensitivity in some cases), having multiple sensors can improve efficiency and accuracy.
[0090] In one or more embodiments, one or more sensors may also be mounted on the beltpack 70 or any other part of the user's body. The sensors (604, 608) may communicate wirelessly or via Bluetooth® with a computing device that determines the attitude and orientation of the sensors (and the AR headset to which they are attached). In one or more embodiments, the computing device may reside on the beltpack 70. In other embodiments, the computing device may reside on the headset itself or further on the handheld controller 606. The computing device may, in turn, include a mapping database (e.g., a passable world model, coordinate space, etc.), detect attitudes, determine the coordinates of real and virtual objects, and in one or more embodiments, may connect to cloud resources and the passable world model.
[0091] As described above, conventional electromagnetic emitters can be too bulky for use in AR devices. Therefore, electromagnetic field emitters may be made more compact by using smaller coils compared to conventional systems. However, assuming that the strength of the electromagnetic field decreases as a cubic function of the distance from the field emitter, the shorter the radius between the electromagnetic sensor 604 and the electromagnetic field emitter 602 (e.g., about 3 to 3.5 feet), the lower the power consumption compared to conventional systems such as those detailed in Figure 4.
[0092] This side may be used in one or more embodiments to extend the life of the battery 610, which can supply power to the controller 606 and the electromagnetic field emitter 602. Alternatively, in other embodiments, this side may be used to reduce the size of the coil that generates the magnetic field in the electromagnetic field emitter 602. However, to obtain the same magnetic field strength, the power may need to be increased. This enables a compact electromagnetic field emitter unit 602 that can be compactly mated in the handheld controller 606.
[0093] Several other modifications may be made when using an electromagnetic tracking system for AR devices. While the attitude reporting rate is very good, AR systems may require an even more efficient attitude reporting rate. To achieve this objective, IMU-based attitude tracking may be used within the sensor. Importantly, the IMU must remain as stable as possible to increase the efficiency of the attitude detection process. The IMU may be machined to remain stable for up to 50–100 milliseconds. It should be understood that some embodiments may utilize an external attitude estimator module, which may allow attitude updates to be reported at a rate of 10–20 Hz (i.e., the IMU may drift over time). By keeping the IMU stable at a reasonable ratio, the attitude update rate can be significantly reduced to 10–20 Hz (compared to higher frequencies in conventional systems).
[0094] If the electromagnetic tracking system could be activated on a 10% duty cycle (e.g., pinging to ground only every 100 milliseconds), this would be another way to conserve power in the AR system. This would mean the electromagnetic tracking system would wake up for 10 milliseconds every 100 milliseconds to generate attitude estimates. This would directly lead to power consumption savings, which in turn could affect the size, battery life, and cost of the AR device.
[0095] In one or more embodiments, this reduction of the duty cycle may be strategically utilized by providing not just one, but two handheld controllers (not shown). For example, a user may play a game that requires two totems, etc. Or, in a multi-user game, two users may each have their own totems / handheld controllers and play the game. When two controllers (e.g., symmetrical controllers for each hand) are used instead of one, the controllers may operate with an offset duty cycle. The same concept may also apply to controllers used by two different users playing a multiplayer game, for example.
[0096] Referring here to Figure 7, an illustrative flowchart illustrating an electromagnetic tracking system in the context of an AR device is described. In 702, a handheld controller emits a magnetic field. In 704, an electromagnetic sensor (installed on a headset, beltpack, etc.) detects the magnetic field. In 706, the position and orientation of the headset / belt are determined based on the behavior of a coil / IMU in the sensor. In 708, the orientation information is transmitted to a computing device (e.g., in a beltpack or headset). In 710, optionally, a mapping database (e.g., a passable world model) may be consulted to correlate real-world coordinates with virtual-world coordinates. In 712, virtual content may be delivered to the user in the AR headset. It should be understood that the flowchart described above is for illustrative purposes only and should not be read as limiting.
[0097] An advantage is that the use of an electromagnetic tracking system similar to that outlined in Figure 6 enables attitude tracking (e.g., head position and orientation, totem and other controller position and orientation). This allows the AR system to project virtual content with higher accuracy and lower latency compared to optical tracking techniques.
[0098] Referring to Figure 8, a system configuration featuring numerous sensing components is illustrated. The head-mounted wearable component (58) is shown here operably coupled (68) to a local processing and data module (70), such as a beltpack, using physical multi-core wires, which also feature a control and rapid release module (86), as described below with reference to Figures 9A-9F. The local processing and data module (70) is here operably coupled (100) to a handheld component (606) by a wireless connection such as Low Power Bluetooth®. The handheld component (606) may also be operably coupled (94) directly to the head-mounted wearable component (58) by a wireless connection such as Low Power Bluetooth®. Generally, when IMU data is passed through for coordinate attitude detection of various components, high-frequency connections, such as hundreds or thousands of cycles / second or higher, are desirable. For electromagnetic positioning and sensing, such as pairing of sensors (604) and transmitters (602), several tens of cycles per second may be appropriate. A global coordinate system (10) is also shown, representing fixed objects in the real world around the user, such as walls (8). Cloud resources (46) may also be operably coupled (42, 40, 88, 90) to local processing and data modules (70), head-mounted wearable components (58), and other items fixed to walls (8) or the global coordinate system (10). Resources coupled to a wall (8) or having a known position and / or orientation with respect to a global coordinate system (10) may include a WiFi transceiver (114), an electromagnetic emitter (602) and / or receiver (604), a beacon or reflector (112) configured to emit or reflect a given type of radiation such as an infrared LED beacon, a cellular network transceiver (110), a radar emitter or detector (108), a lithium-ion emitter or detector (106), a GPS transceiver (118), a poster or marker having a known detectable pattern (122), and a camera (124).The head-mounted wearable component (58) features an optical emitter (130) configured to assist a camera (124) detector, such as an infrared emitter (130) for an infrared camera (124), as well as similar components as shown. The head-mounted wearable component (58) also features one or more strain gauges (116) which may be fixedly coupled to the frame or mechanical platform of the head-mounted wearable component (58) and configured to determine the deflection of such a platform between components such as an electromagnetic receiver sensor (604) or a display element (62), which may be important to understand if bending of the platform occurs in thin parts of the platform, such as the upper part of a protrusion on a spectacle-like platform as depicted in Figure 8. The head-mounted wearable component (58) also features a processor (128) and one or more IMUs (102). Each component is preferably operably coupled to the processor (128). The handheld component (606) and the local processing and data module (70) are illustrated to feature similar components. As shown in Figure 8, using so many sensing and connectivity means, such a system is likely to be heavy, power-hungry, large, and relatively expensive. However, for illustrative purposes, such a system may be utilized to provide very high levels of connectivity, system component integration, and position / orientation tracking. For example, using such a configuration, various key mobile components (58, 70, 606) may be localized in terms of position relative to a global coordinate system using Wi-Fi, GPS, or cellular signal triangulation, and beacons, electromagnetic tracking (as described above), RADAR, and LIDIR systems may further provide location and / or orientation information and feedback. Markers and cameras may also be utilized to provide further information regarding relative and absolute position and orientation.For example, various camera components (124), such as those shown coupled to a head-mounted wearable component (58), may be used to capture data that can be used in a simultaneous localization and mapping protocol, i.e., "SLAM," to determine where and in what state the component (58) is oriented relative to other components.
[0099] Referring to Figures 9A-9F, various aspects of the control and rapid release module 86 are depicted. Referring to Figure 9A, two outer housing components are coupled together using a magnetic coupling configuration, which can be reinforced with mechanical latches. A button (136) for the operation of the associated system may be included. Figure 9B shows a partial cutaway view showing the button (136) and the lower upper printed circuit board (138). Referring to Figure 9C, the button (136) and the lower upper printed circuit board (138) are removed, and the female contact pin array (140) is visible. Referring to Figure 9D, the opposite portion of the housing (134) is removed, and the lower printed circuit board (142) is visible. The lower printed circuit board (142) is removed, and the male contact pin array (144) is visible, as shown in Figure 9E. Referring to the cross-sectional view in Figure 9F, at least one of the male or female pins is configured to be spring-loaded so that it can be pressed along the longitudinal axis of each pin. The pins may be referred to as “pogo pins” and may generally be made of a highly conductive material such as copper or gold. When assembled, the illustrated configuration interlocks 46 male and female pins, and the entire assembly may be rapidly released in half by manually pulling it apart and overcoming a magnetic interface load 146 which can be generated using north and south magnets oriented around the periphery of the pin array (140, 144). In one embodiment, a load of about 2 kg from compressing the 46 pogo pins is counteracted by a closing retention force of about 4 kg. The pins within the array may be separated by only about 1.3 mm, and the pins may be operably coupled to various types of conductive lines such as twisted pairs or other combinations, and may support USB 3.0, HDMI® 2.0, I2S signals, GPIO, and MIPI configurations, and in one embodiment, high-current analog lines and ground configured for up to about 4 amps / 5 volts.
[0100] Referring to Figure 10, it is useful to have a minimized set of components / features, minimizing the weight and size of various components, and enabling relatively slim head-mounted components such as those characterized in Figure 10 (58). Therefore, various permutations and combinations of the various components shown in Figure 8 may be utilized.
[0101] Referring to Figure 11A, an electromagnetic sensing coil assembly (604, i.e., three individual coils coupled to a housing) is shown coupled to a head-mounted component (58). Such a configuration adds additional geometric shape to the overall assembly, which may be undesirable. Referring to Figure 11B, instead of housing the coils in a box or single housing as in the configuration of Figure 11A, the individual coils may be integrated into various structures of the head-mounted component (58), as shown in Figure 11B. Figures 12A–12E illustrate various configurations featuring a ferrite core coupled to an electromagnetic sensor to increase field sensitivity. The embodiments in Figures 12B–12E are lighter in weight than the solid core configuration in Figure 12A and may be used to save mass.
[0102] Referring to Figures 13A-13C, time-division multiplexing ("TDM") may also be used to save mass. For example, referring to Figure 13A, a conventional local data processing configuration is shown for a three-coil electromagnetic receiver sensor, where analog currents emerge from each of the X, Y, and Z coils, flow through a preamplifier, through a band-pass filter, through analog-to-digital conversion, and finally to a digital signal processor. Referring to the transmitter configuration in Figure 13B and the receiver configuration in Figure 13C, time-division multiplexing may be used to share hardware so that each coil sensor chain does not require its own amplifier, etc. In addition to eliminating sensor housings, multiplexing, and saving hardware on the head, the signal-to-noise ratio may be increased by having more than one set of electromagnetic sensors, each set being relatively small compared to a single larger coil set. Also, the low-side frequency limit, which is generally required to have multiple sensing coils in close proximity, can be improved to facilitate improved bandwidth requirements. Furthermore, there is generally a trade-off with multiplexing, as it generally spreads the reception of radio frequency signals over time, resulting in generally noisier signals. Therefore, larger coil diameters may be required for multiplexed systems. For example, a multiplexed system may require a cubic coil sensor box with 9mm side dimensions, while a non-multiplexed system may require only a cubic coil box with 7mm side dimensions for similar performance. Thus, trade-offs may exist when minimizing geometric shape and mass.
[0103] In another embodiment, where a specific system component, such as a head-mounted component (58), features two or more electromagnetic coil sensor sets, the system may be configured to selectively utilize the closest sensor-emitter pairings to each other to optimize system performance.
[0104] Referring to Figure 14, in one embodiment, after the user powers on the wearable computing system (160), the head-mounted component assembly may capture a combination of IMU and camera data (the camera data is used for SLAM analysis, such as a beltpack processor, where more raw processing capability may be present) to determine and update the orientation (i.e., position and orientation) of the head relative to a real-world global coordinate system (162). The user may also activate a handheld component to play, for example, an augmented reality game (164), which may comprise an electromagnetic transmitter operably coupled to one or both of the beltpack and the head-mounted component (166). One or more electromagnetic field coil receiver sets coupled to the head-mounted component (i.e., the set is a set of individual coils oriented in three different ways) capture magnetic flux from the transmitter, which may be used to determine the position or orientation difference (or "delta") between the head-mounted component and the handheld component (168). The combination of a head-mounted component assisting in determining orientation relative to a global coordinate system and a handheld assisting in determining the relative location and orientation of the handheld relative to the head-mounted component allows the system to determine, in general, the location of each component relative to a global coordinate system, and thus the orientation of the user's head, and the handheld orientation can preferably be tracked with relatively short latency for the presentation of augmented reality image features and interaction using the movement and rotation of the handheld component (170).
[0105] Referring to Figure 15, an embodiment is illustrated that is somewhat similar to that of Figure 14, but has more sensing devices and configurations available to assist in determining the postures of both the head-mounted component (172) and the handheld components (176, 178), so that the user's head posture and handheld posture can be tracked with relatively short latency, preferably for the presentation of augmented reality image features and interaction using the movement and rotation of the handheld component (180).
[0106] Referring to Figures 16A and 16B, various aspects of a configuration similar to that of Figure 8 are shown. The configuration of Figure 16A differs from that of Figure 8 in that, in addition to a LIDAR(106) type depth sensor, the configuration of Figure 16A features a general-purpose depth camera or depth sensor (154), which for illustrative purposes may be either a stereo triangulation depth sensor (passive stereo depth sensor, texture projection stereo depth sensor, or structured optical stereo depth sensor, etc.) or a time-of-flight depth sensor (LIDAR depth sensor or modulated emission depth sensor, etc.). Furthermore, the configuration of Figure 16A includes an additional forward-facing "world" camera (124, which may be a grayscale camera with a sensor capable of resolution in the 720p range) and a relatively high-resolution "photographic camera" (156, which may be a full-color camera with a sensor capable of high resolution of, for example, 2 megapixels or more). Figure 16B shows a partial orthogonal view of the configuration of Figure 16A for illustrative purposes, as will be further explained below with reference to Figure 16B.
[0107] Referring back to Figure 16A and the stereo versus time-of-flight depth sensors described above, each of these depth sensor types can be employed with wearable computing solutions such as those disclosed herein, but each has various advantages and disadvantages. For example, many depth sensors have challenges with black surfaces and reflective or glossy surfaces. Passive stereo depth sensing is a relatively simple method of obtaining triangulation to calculate depth using a depth camera or sensor, but this can be a challenge when a wide field of view ("FOV") is required and may require relatively significant computing resources. Furthermore, such sensor types may have challenges with edge detection, which may be important for certain use cases in the future. Passive stereo may have challenges with untextured walls, low-light conditions, and repeating patterns. Passive stereo depth sensors are available from manufacturers such as Intel (RTM) and Aquifi (RTM). Stereo with texture projection (also known as "active stereo") is similar to passive stereo, but the texture projector broadcasts the projection pattern onto the environment, making the texture more broadcast and more accurate, and available for triangulation for depth calculation. Active stereo also presents challenges when a wide FOV is required and may be somewhat suboptimal in edge detection, but it addresses some of the challenges of passive stereo in that it is effective on textureless walls, performs well in low light, and generally does not have problems with repeating patterns.
[0108] Active stereo depth sensors are available from manufacturers such as Intel (RTM) and Aquifi (RTM). Stereo with structured light, such as the system developed by Primesense, Inc. (RTM) and available under the trademark name Kinect (RTM), and generally the system available from Mantis Vision, Inc. (RTM) that utilizes single-camera / projector pairing and a projector, is unique in that it is configured to broadcast a priori known pattern of dots. Essentially, the system understands the pattern to be broadcast and understands that the variable to be determined is depth. Such a configuration can be relatively efficient in terms of computational load and can be challenging in wide-FOV requirement scenarios and scenarios with ambient light and patterns broadcast from other nearby devices, but can be very effective and efficient in many scenarios. Using modulated time-of-flight type depth sensors, such as those available from PMD Technologies (RTM), AG and SoftKinetic Inc. (RTM), the emitter may be configured to transmit waves such as a sine wave of amplitude-modulated light. In some configurations, camera components, which may be positioned in close proximity or even overlapping, may receive return signals onto each of the pixels of the camera components, and depth mapping may be determined / calculated. Such configurations may have a relatively compact geometry, high accuracy, and low computational load, but may present challenges in terms of image resolution (e.g., at the edges of objects) and multi-path errors (e.g., the sensor is aimed at a reflection or glare angle, resulting in the detector receiving more than one return path, such as several depth detection aliasings). Direct time-of-flight sensors, also known as LIDAR as described above, are available from suppliers such as LuminAR (RTM) and Advanced Scientific Concepts, Inc. (RTM). Using these time-of-flight configurations, generally, pulses of light (e.g., picosecond, nanosecond, or femtosecond pulses of light) are transmitted with this light ping to envelop the world oriented around it.Next, each pixel on the camera sensor waits for the pulse to return and, by determining the speed of light, the distance at each pixel can be calculated. Such a configuration offers many of the advantages of modulated time-of-flight sensor configurations (no baseline, relatively wide FOV, high accuracy, relatively low computational load, etc.) and can have relatively high frame rates, such as several million hertz. However, they are also relatively expensive, have relatively low resolution, are sensitive to bright light, and are susceptible to multi-path errors. They can also be relatively large and heavy.
[0109] Referring to Figure 16, a partial top view is shown for illustrative purposes and features a user's eye (12), a camera (14, infrared camera, etc.) with a field of view (28, 30), and a light or radiation source (16, infrared, etc.) directed toward the eye (12) to facilitate eye tracking, observation, and / or image acquisition. Three outward-facing world-capturing cameras (124) are shown with a depth camera (154) and its FOV (24), and a photographic camera (156) and its FOV (26), as well as its FOV (18, 20, 22). Depth information collected from the depth camera (154) may be enhanced by using overlapping FOVs and data from other forward-facing cameras. For example, the system may consist of a sub-VGA image from the depth sensor (154), a 720p image from the world camera (124), and optionally a 2-megapixel color image from the photographic camera (156). Such a configuration would have four cameras sharing a common field of view, two of which would have heterogeneous visible spectral images, one with color, and a third with relatively low-resolution depth. The system might be configured to segment the grayscale and color images, fuse the two images together to create a relatively high-resolution image from them, obtain some stereo correspondence, use the depth sensor to provide a hypothesis about stereo depth, and use the stereo correspondence to obtain a more refined depth map that may be significantly better than what is available from the depth sensor alone. Such a process may be launched on local mobile processing hardware, or, possibly, using cloud computing resources along with data from others in the area (e.g., two people sitting across a table from each other in the vicinity), resulting in a highly refined mapping.
[0110] In another embodiment, all of the sensors described above may be combined into a single integrated sensor to perform such functionality.
[0111] Referring to Figures 17A-17G, aspects of the dynamic transmission coil tuning configuration are shown for electromagnetic tracking, which facilitates the optimal operation of the transmission coils at multiple frequencies per orthogonal axis, allowing multiple users to operate on the same system. Typically, an electromagnetic tracking transmitter would be designed to operate at a fixed frequency per orthogonal axis. Using such an approach, each transmission coil is tuned with a static set of capacitances that create resonance only at the operating frequency. Such resonance allows for the maximum possible current through the coil, which in turn maximizes the magnetic flux produced.
[0112] Figure 17A illustrates a typical resonant circuit used to generate resonance. Element "L1" represents a single-axis transmission coil with a capacitance of 52 nF set at 1 mH, and resonance is generated at 22 kHz, as shown in Figure 17B.
[0113] Figure 17C shows the current through the system plotted against frequency, and it can be seen that the current is maximum at the resonant frequency. If the system is expected to operate at any other frequency, the operating circuit will not be the maximum possible. Figure 17D illustrates an embodiment of a dynamically adjustable configuration. Dynamic frequency adjustment may be set to achieve resonance on the coil and obtain maximum current flow. An embodiment of the adjustable circuit is shown in Figure 17E, and one capacitor ("C4") may be adjusted to produce the simulated data, as shown in Figure 17F. As shown in Figure 17F, one of the orthogonal coils of the electromagnetic tracker is simulated as "Ll", and the static capacitor ("C5") is a fixed high-voltage capacitor. This high-voltage capacitor will suffer a higher voltage due to resonance, and therefore its package size will generally be larger. C4 is a capacitor that is dynamically switched using different values, and therefore will suffer a lower maximum voltage, generally resulting in a smaller geometric package and saving installation space. L3 can also be used to fine-tune the resonant frequency. Figure 17F illustrates the resonance achieved using higher plots (248) versus lower plots (250). It is worth noting that as C4 is varied in the simulation, the resonance changes, and the voltage across C5 (Vmid-Vout) is higher than that across C4 (Vout). This will allow for smaller package portions on C4, as multiple units are generally required for the system, one per operating frequency. Figure 17G illustrates that the maximum current achieved follows the resonance, regardless of the voltage across the capacitor.
[0114] Referring to Figures 18A-18C, the electromagnetic tracking system may be bounded to operate below approximately 30 kHz, slightly above the audible range for human hearing. Referring to Figure 18A, there may be several audio systems that generate noise within the usable frequency range for such an electromagnetic tracking system. Furthermore, audio speakers typically have a magnetic field and one or more coils that can also interfere with the electromagnetic tracking system.
[0115] Referring to Figure 18B, a block diagram is shown for a noise cancellation configuration for electromagnetic tracking interference. Since unintentional interference is a known entity, this knowledge can be used to cancel the interference and improve performance. In other words, audio produced by the system may be used to eliminate the effect received by the receiver coil. The noise cancellation circuit may be configured to receive corrupted signals from the EM amplifier and signals from the audio system, and the noise cancellation system will cancel and remove noise received from the audio speaker. Figure 18C illustrates a plot to show an embodiment of how signals may be inverted and added to cancel out interference. V (Vnoise) (upper plot) is the noise added to the system by the audio speaker. Referring to Figure 19, in one embodiment, a known pattern of light or other emitter (e.g., a circular pattern) may be used to assist in the calibration of the vision system. For example, a circular pattern may be used as a base point. In other words, while the object being joined to the pattern is being reoriented, the orientation of the object, such as a handheld totem device, can be determined as a camera or other capture device with a known orientation captures the shape of the pattern. Such orientation may be compared to that resulting from the associated IMU device for use in error determination and calibration.
[0116] Referring to Figures 20A-20C, the configuration is shown with an add-up amplifier to simplify the circuitry between two subsystems or components of a wearable computing configuration, such as a head-mounted component and a belt-pack component. In a conventional configuration, each coil of an electromagnetic tracking sensor (left in Figure 20A) would be associated with an amplifier, and three distinctly different amplified signals would be transmitted to the other component through cables. In the illustrated embodiment, the three distinctly different amplified signals may also be directed to an add-up amplifier, which produces a single amplified signal directed along advantageously simplified cables, each signal being at a different frequency. The add-up amplifier may be configured to amplified all three incoming signals, which are then separated at the other end by a receiving digital signal processor after analog-to-digital conversion. Figure 20C illustrates frequency-by-frequency filters, so that the signals are separated and returned at such stages.
[0117] Referring to Figure 21, electromagnetic ("EM") tracking updates can be relatively "expensive" from a power standpoint for portable systems, and very high-frequency updates may not be possible. In a "sensor fusion" configuration, more frequently updated localization information or other dynamic inputs (measurable metrics that change over time) from another sensor, such as an IMU, may be combined with data from another sensor, such as an optical sensor (camera or depth camera, etc.), which may or may not be relatively high-frequency. The final effect of fusing all these inputs is to place a lower demand on the EM system and provide faster updates. Furthermore, with respect to "dynamic inputs," other illustrative embodiments include temperature fluctuations, audio volume, size determination such as the dimensions of an object or the distance to it, and not simply the user's position or orientation. A set of dynamic inputs represents the set of those inputs as a function of a given variable (time, etc.).
[0118] Referring back to Figure 11B, a distributed sensor coil configuration is shown. Referring to Figure 22A, a configuration with a single electromagnetic sensor device (604), such as a box containing three orthogonal coils, one for each of the X, Y, and Z directions, may be coupled to a wearable component (58) for six-degree-of-freedom tracking, as described above. Alternatively, as described above, such a device may be disassembled with three sub-parts (i.e., coils) attached to different locations on the wearable component (58), as shown in Figure 22B. Referring to Figure 22C, to provide further design alternatives, each individual coil may be replaced with a group of similarly oriented coils such that the total magnetic flux for any given orthogonal direction is captured by the group (148, 150, 152) rather than by a single coil for each orthogonal direction. In other words, a smaller group of coils, rather than one coil for each orthogonal direction, may be utilized, and their signals may be aggregated to form a signal for that orthogonal direction.
[0119] Referring to Figures 23A-23C, it may be useful to recalibrate wearable computing systems, such as those discussed from time to time herein, and in one embodiment, an ultrasonic signal in the transmitter may be used to determine the sound propagation delay, along with a microphone and acoustic time-of-flight calculations in the receiver. Figure 23A shows that in one embodiment, three coils on the transmitter may be excited with a sinusoidal burst, and simultaneously, an ultrasonic transducer may be excited with a sinusoidal burst of the same frequency as one of the coils, preferably. Figure 23B illustrates that the receiver may be configured to receive three EM waves using a sensor coil and ultrasonic waves using a microphone device. The total distance may be calculated from the amplitudes of the three EM signals. The time-of-flight may then be calculated by comparing the timing of the microphone response with the response of the EM coil (Figure 23C). This may also be used to calculate the distance and calibrate the EM correction coefficient.
[0120] Referring to Figure 24A, in another embodiment, in an augmented reality system featuring a camera, distance may be calculated by measuring the pixel-level size of a known size feature on another device, such as a handheld controller.
[0121] Referring to Figure 24B, in another embodiment, in an augmented reality system featuring a depth sensor such as an infrared ("IR") depth sensor, the distance may be calculated by such depth sensor and reported directly to the controller.
[0122] Referring to Figures 24C and 24D, once the total distance is determined, either a camera or a depth sensor can be used to determine the position in space. The augmented reality system may be configured to project one or more virtual targets onto the user.
[0123] The user may align the controller to the target, and the system calculates the position from both the EM response and the previously calculated distance in addition to the direction of the virtual target. Roll angle calibration may be performed by aligning known features on the controller with a virtual target projected to the user. Yaw and pitch angles may be calibrated by presenting a virtual target to the user and having the user align two features on the controller with the target (as if aiming a rifle).
[0124] Referring to Figures 25A and 25B, an inherent ambiguity exists associated with EM tracking systems for 6-degree-of-freedom devices. Specifically, the receiver will produce similar responses at two diagonally opposite locations around the transmitter. Such challenges are particularly relevant in systems where both the transmitter and receiver may be mobile relative to each other.
[0125] For 6-degree-of-freedom (DOF) tracking, the totem 524, also referred to as a handheld controller (e.g., TX), may generate EM signals modulated on three distinct frequencies, one for each of the X, Y, and Z axes. A wearable 522, as shown in Figure 25A, which may be implemented as an AR headset 528 as shown in Figure 25B, has an EM receiving component configured to receive EM signals on the X, Y, and Z frequencies. The position and orientation (i.e., attitude) of the totem 524 can be derived based on the characteristics of the received EM signals. However, due to the symmetrical nature of the EM signals, it may be impossible to determine the location of the totem (e.g., in front of or behind the user, as illustrated using the afterimage position 526) without using an additional reference frame. That is, the same EM value can be acquired in wearables 522, 528 with respect to two diametrically opposed totem poses, namely the position of the totem 524 in the first (e.g., front) hemisphere 542 and the afterimage position 526 in the second (e.g., rear) hemisphere, along with a selected plane 540 passing through the center of the sphere that divides the two hemispheres. Thus, with respect to a single snapshot of the EM signal received by wearables 522, 528, either pose is valid. However, if the totem 524 is moved, the tracking algorithm will encounter errors, typically due to inconsistencies in the data from various sensors, resulting in the selection of the wrong hemisphere. Hemispheric ambiguity arises in part from the fact that when a 6DOF tracking session is initiated, the initial EM totem data does not have a clear absolute position. Instead, this provides a relative distance, which can be interpreted as one of two positions within a 3D volume divided into two equal hemispheres, with the wearable (e.g., an AR headset mounted on the user's head) at the center between the two hemispheres. Thus, embodiments of the present invention provide a method and system for resolving hemispheric ambiguity to enable successful tracking of the actual position of the totem.
[0126] Such ambiguity has been clearly documented by Kuipers, who describes the EM signal (which is itself an alternating current sinusoidal output) as a 3x3 matrix S as a function of the EM receiver within the EM transmitter coordinate frame. S=f(T,r) In the equation, T is the rotation matrix from the transmitter coordinate frame to the receiver coordinate frame, and r is the EM receiver position within the EM transmitter coordinate frame. However, as the Kuipers solution points out, solving any 6DOF position involves squaring a sine function, and solving the -1 position similarly solves the +1 position, thereby introducing a hemispheric ambiguity problem.
[0127] In one embodiment, an IMU sensor may be used to determine whether the totem 524 (e.g., an EM transmitter) is located on the positive or negative side of the axis of symmetry. In some embodiments, such as those described above, which feature a world camera and a depth camera, this information can be used to detect whether the totem 524 is on the positive or negative side of the reference axis. If the totem 524 is outside the field of view of the camera and / or depth sensor, the system may be configured to determine (or the user may determine) that the totem 524 should be located, for example, within a 180° zone immediately behind the user.
[0128] Conventionally in the art, an EM receiver (provided within the wearable 528) is defined within an EM transmitter coordinate frame (provided within the totem 524) (as illustrated in Figure 25C). To correct for two possible receiver locations (i.e., the actual receiver location 528 and a false receiver location 530 which is symmetric to the actual location of the wearable 528 relative to the transmitter 524), a designated initialization position is established to orient the device (i.e., the wearable 528) relative to a specific reference position and to reject the alternative solution (i.e., the false receiver location 530). However, such a required initialization position may conflict with intent or desired use.
[0129] Instead, in some embodiments, the EM transmitter (provided within totem 524) is defined within an EM receiver coordinate frame (provided within wearable 528) as illustrated in Figure 25D. With respect to such embodiments, the Kuipers solution is modified so that the transpose of S is solved. The result of this S-matrix transpose operation is comparable to the treatment of an EM transmitter measurement, as if it were an EM receiver, and vice versa. It is worth noting that the same-hemisphere ambiguity phenomenon still exists, but is instead represented within the coordinate frame of the EM receiver. That is, possible transmitter locations are identified as afterimage locations 526 (also referred to as mistransmitter locations), which are symmetric to the actual location of totem 524 and the actual location of totem 524 relative to wearable 528.
[0130] In contrast to the ambiguous hemispheres that appear at both the user's location and the misplaced location opposite the transmitter, the EM tracking system here has a hemisphere 542 in front of the user and a hemisphere 544 behind the user, as illustrated in Figure 25E. By setting up a hemispherical boundary (e.g., a switching plane) 540 that is centered on the transposed receiver location (e.g., the location of the wearable 528), the surface normal vector n s The surface normal vector n originates from the origin (i.e., the receiver installed within the wearable 528) and is directed at a predetermined angle α534 from a straight (horizontal) line obtained from the wearable 528. In some embodiments, the predetermined angle α534 may be 45° from the straight line obtained from the wearable 528. In other embodiments, the predetermined angle α534 may be any other angle measured with respect to the straight line obtained from the wearable 528, and the use of 45° is merely illustrative. Generally, when the user holds the totem 524 at waist height, the totem is generally oriented at an angle of about 45° from the straight line obtained from the wearable 528, with respect to the surface normal vector n s It will be in the vicinity of 532. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0131] If totem 524 is located within the hemisphere in front of the user (i.e., hemisphere 542 in Figure 25E), then the surface normal vector n s The position vector p representing the positions of 532 and totem 524 is n s The dot product inequality will be satisfied if p > 0 (wherein p is the position point of the vector defined in the receiver coordinate frame). Vector p536 may be a vector that connects the transmitter provided to totem 524 to the receiver provided to wearable 528. For example, vector p536 is the normal vector n s 532 may be provided at a predetermined angle β538. When the position of the transmitter (in totem 524) is determined within the receiver coordinate frame, two position vectors are identified: a first (positive) position vector p536 and a second (negative) position vector 546 which is symmetric to the first position vector p536 with respect to the receiver (in wearable 528).
[0132] To determine which hemisphere a device (e.g., Totem 524) is in, two possible solutions to the S matrix are applied to the dot product, and a positive number yields the correct hemisphere location. Such a decision is anecdotically valid for the system described and its use cases as well, since it is highly unlikely that a user would hold the device at arm's length behind their back.
[0133] Figure 25F illustrates a flowchart 550 illustrating the steps of determining the correct orientation of a handheld controller using position vectors relating to a handheld device relative to a headset, according to several embodiments. According to various embodiments, the steps of flowchart 550 may be performed during the headset initialization process. In step 552, the handheld controller of the optical device system emits one or more magnetic fields. In step 554, one or more sensors positioned within the system's headset (e.g., wearable by the user) detect the one or more magnetic fields emitted by the handheld controller. In step 556, a processor coupled to the headset determines a first position and first orientation of the handheld controller in a first hemisphere relative to the headset, based on the one or more magnetic fields. In step 558, the processor determines a second position and second orientation of the handheld controller in a second hemisphere relative to the headset, based on the one or more magnetic fields. The second hemisphere is diametrically opposed to the first hemisphere relative to the headset. In step 560, the processor determines (defines) a normal vector for the headset and a position vector that identifies the position of the handheld controller relative to the headset within a first hemisphere. The position vector is defined in the coordinate frame of the headset. The normal vector originates from the headset and extends at a predetermined angle from the horizontal line from the headset. In some embodiments, the predetermined angle is 45° downward from the horizontal line from the headset. In step 562, the processor calculates the dot product of the normal vector and the position vector. In some embodiments, depending on the calculation of the dot product of the normal vector and the position vector, the processor may determine whether the resulting value of the dot product is positive or negative.
[0134] If the result of the dot product is determined to be positive, the processor determines in step 564 that the first position and first orientation of the handheld controller are correct. Thus, when the result of the dot product is positive, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset. The system may then deliver virtual content to the system's display based on the first position and first orientation when the result of the dot product is positive.
[0135] If the result of the dot product is determined to be negative, the processor determines in step 566 that the second position and second orientation of the handheld controller are correct. Thus, when the result of the dot product is negative, the second hemisphere is identified as the front hemisphere relative to the headset, and the first hemisphere is identified as the rear hemisphere relative to the headset. The system may then deliver virtual content to the system's display based on the second position and second orientation when the result of the dot product is negative. In some embodiments, step 566 is not performed if the dot product is determined to be positive (i.e., step 566 may be an optional step). Similarly, in some embodiments, step 558 is not performed if the dot product is determined to be positive (i.e., step 558 may be an optional step).
[0136] Referring back to the embodiments described above, in which outward-facing camera devices (124, 154, 156) are coupled to a system component such as a head-mounted component (58), the position and orientation of the head coupled to such a head-mounted component (58) may be determined using information collected from these camera devices using techniques such as simultaneous localization and mapping, i.e., "SLAM" techniques (also known as parallel tracking and mapping, i.e., "PTAM" techniques).
[0137] Understanding the position and orientation of the user's head, also known as the user's "head posture," in real time or near real time (preferably with short latency determination and updates) is invaluable in determining where the user is located in the real environment around them and how to position and present virtual content relative to the user and the environment related to the user's augmented or mixed reality experience. Typical SLAM or PTAM configurations involve extracting features from incoming image information, using them to triangulate 3-D mapping points, and then tracking to those 3-D mapping points. SLAM techniques are used in many implementations, such as autonomous driving in cars, where computation, power, and sensing resources may be relatively abundant compared to those that may be available onboard in wearable computing devices such as head-mounted components (58).
[0138] Referring to Figure 26, in one embodiment, a wearable computing device such as a head-mounted component (58) may comprise two outward-facing cameras producing two camera images (left-204, right-206). In one embodiment, a relatively lightweight, portable, and power-efficient embedded processor, such as those sold by Movidius (RTM), Intel (RTM), Qualcomm (RTM), or Ceva (RTM), may constitute part of the head-mounted component (58) and be operably coupled to the camera device. The embedded processor may first be configured to extract features (210, 212) from the camera images (204, 206). If the calibration between the two cameras is known, the system can triangulate (214) the 3-D mapping points of those features, resulting in a set of sparse 3-D map points (202). These may be stored as a “map,” and these first frames may be used to establish a “world” coordinate system origin (208). As subsequent image information enters the built-in processor from the camera, the system may be configured to project 3-D map points into the new image information and compare them to the locations of 2-D features detected within the image information. Thus, the system may be configured to attempt to establish 2-D / 3-D correspondences, and using groups of such correspondences, such as about six of them, the pose of the user's head (which is, of course, coupled to the head-mounted device 58) may be estimated. A larger number of correspondences, such as more than six, generally means that it works better for estimating the pose. Naturally, this analysis relies on having some sense of where the user's head was before the current image was considered (i.e., in terms of position and orientation). As long as the system can track without too much latency, the system may use pose estimates from the previous time to estimate where the head was for the most recent data. Thus, the last frame may be the origin, and the system may be configured to estimate that the user's head is not far from it in terms of position and / or orientation, and may search its periphery to find correspondences for the current time interval.This forms the basis of one embodiment of the tracking configuration.
[0139] After moving sufficiently far from the original set of map points (202), one or both camera images (204, 206) may begin to lose map points in the newly incoming images (for example, if the user's head is rotated to the right in space, the original map points may begin to disappear to the left, appearing only in the left image, and then disappearing completely with further rotation). Once the user rotates very far from the original set of map points, the system may be configured to create new map points by using a process similar to those described above (detecting features, creating new map points), in which case the system may be configured to continue inputting data into the map. In one embodiment, this process may be repeated again every 10 to 20 frames, depending on the amount by which the user translates and / or rotates their head relative to the environment, thereby translating and / or rotating the associated cameras. A frame associated with a newly created mapping point may be considered a "keyframe," and the system may be configured to delay the feature detection process using keyframes, or alternatively, feature detection may be performed on each frame, attempting to establish a match, and then, when the system is ready to create a new keyframe, the system has already completed its associated feature detection. Thus, in one embodiment, the basic paradigm is to start creating a map and then continue tracking until the system needs to create another map or an additional part thereof.
[0140] Referring to Figure 27A, in one embodiment, the vision-based attitude calculation is divided into five stages (pre-tracking 216, tracking 218, short latency mapping 220, latency-tolerant mapping 222, and post-mapping / cleanup 224) to assist in accuracy and optimization for an internal processor configuration where computation, power, and sensing resources may be limited.
[0141] With respect to pre-tracking (216), the system may be configured to identify map points that will be projected into an image before the image information arrives. In other words, assuming that the system knows where the user has previously been and has a sense of where the user is moving, the system may be configured to identify map points that will be projected into an image.
[0142] The concept of "sensor fusion" will be discussed further below, but it is worth noting that one of the inputs a system can obtain from a sensor fusion module or functionality may be "post-hoc estimation" information from an inertial measurement unit ("IMU") or other sensors or devices at a relatively high rate, such as 250 Hz (which is a high rate compared to 30 Hz, for example, at which vision-based attitude calculation operations may provide updates). Thus, a much finer temporal resolution of attitude information can be derived from IMUs or other devices compared to vision-based attitude calculations. However, it is also worth noting that, as will be discussed below, data from devices such as IMUs tends to be somewhat noisy and susceptible to attitude estimation drift. For relatively short time windows, such as 10-15 milliseconds, IMU data can be very useful in predicting attitude, and again, when combined with other data in a sensor fusion configuration, an optimized overall result can be determined. For example, as will be explained in more detail below with reference to Figures 28B-28G, the propagation path of the collected points may be adjusted, and from this adjustment, the orientation at a later time may be estimated / informed by the location determined by sensor fusion, where the future points may be based on the adjusted propagation path.
[0143] Pose information arising from a sensor fusion module or functionality may be referred to as “pre-pose,” which can be used by the system to estimate a set of points that will be projected onto the current image. In one embodiment, the system is configured to pre-fetch these map points and perform some pre-processing in a “pre-tracking” step (216), which helps reduce the overall processing latency. Each 3-D map point may be associated with a descriptor so that the system can uniquely identify them and match them to areas in the image. For example, if a given map point is created by using a feature that has a patch around it, the system may be configured to maintain some appearance of the patch with the map point so that when the map point is projected onto another image, the system can return to the original image used to create the map, consider patch correlations, and determine whether they are the same point. In the pre-processing, the system may be configured to fetch a certain amount of map points and perform a certain amount of pre-processing associated with the patches associated with those map points. Therefore, in pre-tracking (216), the system may be configured to pre-fetch map points and pre-warp image patches (the “warping” of images may be done to ensure that the system can match the current image with the patches associated with the map points; this is a way to ensure that the data being compared are compatible).
[0144] Referring back to Figure 27, the tracking stage may comprise several components, including feature detection, optical flow analysis, feature matching, and pose estimation. While detecting features in incoming image data, the system may be configured to save computation time in feature detection by utilizing optical flow analysis to attempt to track features to features from one or more previous images. Once features are identified in the current image, the system may be configured to attempt to match the features to projected map points, which can be considered the "feature matching" part of the configuration. In the pre-tracking stage (216), the system has preferably already identified and fetched the map points of interest. In feature mapping, they are projected into the current image, and the system attempts to match them to the features. The output of feature mapping is a set of 2-D / 3-D correspondings, and with this in mind, the system is configured to estimate the pose.
[0145] As the user moves their head while coupled to the head-mounted component (58), the system is preferably configured to identify whether the user is viewing a new area of the environment and to determine whether a new keyframe is needed. In one embodiment, such analysis of whether a new keyframe is needed may be based largely on purely geometric shapes. For example, the system may be configured to check the distance from the current frame to the remaining keyframes (translation distance, i.e., field-of-view reorientation, where the user's head may be reoriented, for example, to approach in the translation direction but to require a completely new map point). Once the system has determined that a new keyframe should be inserted, the mapping phase may be initiated. As described above, the system may be configured to operate the mapping as three distinct operations (short latency mapping, latency-tolerant mapping, post-mapping, or cleanup), in contrast to a single mapping operation which is more likely to be seen in conventional SLAM or PTAM operations.
[0146] Short-latency mapping (220), which can be considered the simplest form as triangulation and creation of new map points, is a critical step, and the system is preferably configured to perform such a step immediately, since the tracking paradigm discussed herein relies on map points, and the system finds only the location if there are map points available for tracking. The term "short-latency" refers to the concept that there is no tolerance for unacceptable latency (in other words, the main part of the mapping must be done as quickly as possible, otherwise the system will have tracking problems).
[0147] Latency tolerance mapping (222) can be considered the simplest form as an optimization step. The overall process does not absolutely require short latency to perform this operation, known as "bundle tuning," which provides a global optimization to the result. The system may be configured to consider the location of 3-D points and where they were observed. There are many errors that can chain together in the process of creating the map points. The bundle tuning process may, for example, take a particular point observed from two different viewing locations and use all of this information to obtain a better perception of the actual 3-D geometric shape.
[0148] The results may be such that the 3-D points and the calculated trajectory (i.e., the location and path of the capture camera) can be adjusted in small increments. It is desirable to perform these types of processes and avoid accumulating errors throughout the mapping / tracking process.
[0149] The post-mapping / cleanup (224) stage may be configured in which the system removes points on the map that do not provide valuable information in the mapping and tracking analysis. In this stage, these points that do not provide useful information about the scene are removed, and such analysis is useful in keeping the entire mapping and tracking process scalable.
[0150] During the vision pose calculation process, features visible to an outward-facing camera are assumed to be static features (i.e., they do not move frame by frame relative to the global coordinate system). In various embodiments, semantic segmentation and / or object detection techniques may be used to remove moving objects such as people, moving vehicles, and equivalents from the relevant fields so that features for mapping and tracking are not extracted from these regions of various images. In one embodiment, deep learning techniques such as those described below may be used to segment and extract these non-static objects.
[0151] In some embodiments, the pre-fetch protocol leads to attitude estimation according to a method (2710) as depicted in Figure 27B. From 2710 to 2720, attitude data is received over time so that the estimated attitude at future time can be extrapolated at 2730. The extrapolation may be a simple constant value extrapolation from a given input, or an extrapolation corrected based on a corrected input point as described below with reference to sensor fusion.
[0152] Depending on the determination of the estimated future orientation, in some embodiments the system accesses a feature map of that position. For example, if a user is walking while wearing a head-mounted component, the system may extrapolate the future position based on the user's gait and access a feature map of that extrapolated / estimated position. In some embodiments this step is a pre-fetch, as described above with reference to step (216) in Figure 27A. In (2740), specific points of the feature map are extracted, and in some embodiments, patches surrounding the points are similarly extracted. In (2750), the extracted points are processed. In some embodiments the processing includes the steps of warping the points and matching them to the estimated orientation, or finding homography of the extracted points or patches.
[0153] In (2760), a real-time image (or the current view at an estimated time) is received by the system. The processed points are projected onto the received image in (2770). In 2780, the system establishes a correspondence between the received image and the processed points. In some cases, the system makes a perfect estimation, and the received image and the processed points are perfectly aligned, confirming the estimated pose in (2730). In other cases, the processed points do not perfectly align with the features in the received image, and the system performs additional warping or adjustment to determine the correct pose based on the degree of correspondence in (2790). Naturally, it is possible that none of the processed points align with any features in the received image, and the system will need to return to a new tracking and feature mapping as described above with reference to Figure 27A and steps (218)-(224).
[0154] Referring to Figures 28A-28F, a sensor fusion configuration may be used to take advantage of one information source resulting from a sensor with a relatively high update frequency (such as an IMU updating gyroscope, accelerometer, and / or magnetometer data related to head attitude at a frequency such as 250 Hz) and another information source updating at a lower frequency (such as a vision-based head attitude measurement process updating at a frequency such as 30 Hz).
[0155] Referring to Figure 28A, in one embodiment, the system may be configured to track a significant amount of information about the device using an extended Kalman filter (EKF, 232). For example, in one embodiment, 32 states may be considered, such as angular velocity (i.e., from the IMU gyroscope), translational acceleration (i.e., from the IMU accelerometer), calibration information about the IMU itself (i.e., coordinate systems and calibration coefficients for the gyroscope and accelerometer; the IMU may also have one or more magnetometers). Thus, the system may be configured to perform IMU measurements at a relatively high update frequency (226), such as 250 Hz, and data from some other source at a lower update frequency (i.e., calculated vision attitude measurements, odometry data, etc.), here, vision attitude measurements (228), at an update frequency such as 30 Hz.
[0156] Each time the EKF obtains a series of IMU measurements, the system may be configured to integrate the angular velocity information to obtain rotational information (i.e., the integral of angular velocity (change in rotational position over time) is the angular position (change in angular position)). The same applies to translation information (in other words, by performing a double integral of translational acceleration, the system will obtain positional data). Using such calculations, the system is configured to obtain 6-degree-of-freedom (DOF) attitude information from the head from the IMU at a high frequency (i.e., 250 Hz in one embodiment) (translation in X, Y, and Z, orientation with respect to the three rotational axes). Noise accumulates in the data each time integration is performed. Performing a double integral of translational or rotational acceleration can propagate noise.
[0157] Generally, the system is configured not to rely on data that is prone to “drift” due to noise during time windows that are too long, such as slightly longer than about 100 milliseconds in one embodiment. Lower frequency data incoming from vision attitude measurements (228) (i.e., updated at about 30 Hz in one embodiment) may be used with EKF (232) to act as a correction factor to produce a corrected output (230).
[0158] Figures 28B-28F illustrate how data from one source at a higher update frequency may be combined with data from another source at a lower update frequency. As depicted in Figure 28B, for example, a first group (234) of dynamic input points from an IMU at a higher frequency, such as 250 Hz, is shown together with correction input points (238) occurring at a lower frequency, such as 30 Hz, from a vision attitude calculation process. The system may be configured to correct (242) the vision attitude calculation points when such information is available, and then continue forward with a second set of dynamic input points (236), such as points from IMU data, and another correction (244) from another correction input point (240) available from the vision attitude calculation process. In other words, the high-frequency dynamic input collected from the first sensor may be periodically regulated by the low-frequency correction input collected from the second sensor.
[0159] Thus, embodiments of the present invention use EKF to apply a correction "update" in vision attitude data to the "propagation path" of data originating from the IMU. Such an update modifies the propagation path of the dynamic input (234) to a new origin, for a set of future points, at the correction input point (238), again at (240), etc., as depicted in Figure 28B. In some embodiments, instead of modulating the propagation path to originate from the correction input point, the propagation path may be modulated by changing the rate of change with respect to a second set of dynamic inputs by a calculated coefficient (i.e., the slope of point (236) in Figure 28B). These modifiers can reduce jitter in the system by interpreting the new data set as having less difference from the correction input.
[0160] For example, as depicted in Figure 28B-2, the coefficient is applied to the second set of dynamic input points (236) such that the origin of the propagation path is the same as the last point of the propagation path of the first set of dynamic input points (234), but the slope is not as large compared to the rate of change of the first set of dynamic input points (234). As depicted in Figure 28B-2, the size of corrections (242) and (242) changes, but over time T N It is not as large as when there is no correction input, and depending on computing resources and sensors, as shown in Figure 28B-1, T N and T N+M Performing two full corrections in this case is computationally more expensive than the slighter adjustment in Figure 28B-2 and may introduce jitter between sensors. Furthermore, as depicted in Figure 28B, if the correction input points (238) and (240) are in the same location and have also moved, the user can enjoy the advantage of having a dynamic input point tilt adjustment, which places the last point of the propagation path (236) closer to the correction input point (240) compared to origin adjustment.
[0161] Figure 28G illustrates method 2800, in which a first set of dynamic points is collected in (2810), and then correction input points are collected in (2820). In some embodiments, method 2800 proceeds to step (2830a), in which a second set of dynamic points is collected, and the resulting propagation path of those collected points is adjusted in (2830b) based on the correction input (by adjusting the rate of change recorded, or by adjusting the origin of the propagation path, etc.). In some embodiments, after the correction input points are collected in (2820), the adjustment is determined in (2830b) to be applied to subsequently collected dynamic points, and in (2830a), a second set of dynamic input points is collected, and the propagation path is adjusted in real time.
[0162] Regarding systems that employ sensors known to produce time-dependent compound errors / noise and other sensors that produce low errors / noise, such sensor fusion offers more economical computing resource management.
[0163] It is worth noting that in some embodiments, data from a second source (i.e., vision attitude data, etc.) may occur not only at a lower update frequency but also with some latency, which means the system is preferably configured to navigate time-domain adjustments as the information from the IMU and the vision attitude calculations are integrated. In one embodiment, to ensure that the system fuses the vision attitude calculation input to the correct time-domain position in the IMU data, a buffer of the IMU data may be maintained to perform the fusion and return to a time (e.g., "Tx") in the IMU data for calculating the "update" or adjustment at the time related to the input from the vision attitude calculations, and then consider it in forward propagation up to the current time (e.g., "Tcurrent"), leaving a gap between the position and / or orientation data being adjusted and the latest data coming from the IMU. To ensure that there are no too many "jumps" or "jitters" in the presentation to the user, the system may be configured to use smoothing techniques. One way to address this problem is to use a weighted averaging technique, which can be linear, nonlinear, exponential, etc., and can ultimately drive the fused data stream to a tuned path. Referring to Figure 28C, for example, a weighted averaging technique may be used over a time domain between T0 and T1 to drive the signal from an untuned path (252, i.e., originating linearly from the IMU) to a tuned path (254, i.e., based on data originating from the vision attitude calculation process). One embodiment is shown in Figure 28D, where the fused result (260) is shown to increase exponentially from the untuned path (252) to the tuned path (254), starting at time T0 and increasing by T1.Referring to Figure 28E, a series of correction opportunities are shown in each sequence, along with exponential time-domain corrections to the fused result (260) from the upper path to the lower path (the first correction is, for example, from the first path 252 from the IMU to the second path 254 from, for example, vision-based attitude calculation, and then, in this embodiment, using each incoming vision-based attitude calculation point, correcting sequentially toward the corrected lower paths 256, 258 based on the continuity of points from the vision attitude, while continuing forward in a similar pattern using the continuity of IMU data). Referring to Figure 28F, with a sufficiently short time window between "updates" or corrections, the overall fused result (260) may be perceived functionally as a relatively smooth patterned result (262).
[0164] In other embodiments, instead of relying directly on vision attitude measurement, the system may be configured to consider the differential EKF. In other words, instead of directly using the vision attitude calculation results, the system uses the change in vision attitude from the current time to the previous time. Such a configuration may be pursued, for example, when the amount of noise in the vision attitude difference is significantly less than the amount of noise in the absolute vision attitude measurement. It is preferable that all of the outputs are attitudes, which are returned to the vision system as “previous attitude” values, so that instantaneous errors are not removed from the fused results.
[0165] An external system-based “consumer” of pose results may be referred to as a “pose service,” and the system may be configured so that all other system components utilize the pose service when a pose is requested at any given time. The pose service may be configured as a queue or stack (i.e., a buffer) with a sequence of time slices of data, one end of which has the most recent data. If a request for the pose service is for the current pose or any other pose in the buffer, it may be output immediately. In one configuration, the pose service would receive requests for poses that will be taken 20 milliseconds from the present (for example, in video game content rendering a scenario, it may be desirable for the relevant service to know that something needs to be rendered at a given position and / or orientation a little further in the future). In one model for producing future pose values, the system may be configured to use a constant velocity prediction model (i.e., it assumes the user’s head is moving at a constant velocity and / or angular velocity). In another model for producing future pose values, the system may be configured to use a constant acceleration prediction model (i.e., it assumes the user’s head is translating and / or rotating at a constant acceleration).
[0166] The data in the data buffer may be used to extrapolate where the attitude would likely be using such a model. The constant acceleration model uses a slightly longer tail to the buffer data for prediction than the constant velocity model, and it has been found that the system of this subject can predict up to a range of 20 milliseconds ahead with virtually no degradation. Therefore, the attitude service may be configured to have data buffers approximately 20 milliseconds or more ahead in terms of the data that can be used to output the attitude.
[0167] Operationally, content behavior will generally be configured to identify when the next frame retrieval will occur at a given time (for example, attempting to retrieve at either time T or time T+N, where N is the next interval of updated data available from the orientation service).
[0168] The use of a user-facing camera (e.g., inward-facing, such as one depicted in Figure 16B(14)) may be used to perform eye tracking, as described, for example, in U.S. Patent Applications No. 14 / 707,000 and No. 15 / 238,516 (which are incorporated herein by reference as a whole). The system may be configured to first capture an image of the user's eye, then perform several steps in eye tracking, such as segmenting the anatomical structure of the eye using segmentation analysis (e.g., to segment the pupil from the iris, sclera, and surrounding skin), and then the system may be configured to estimate the pupil center using a flash field location identified in the image of the eye, the flash originating from a small illumination source (16), such as an LED, which may be installed around the inward-facing side of a head-mounted component (58). From these steps, the system may be configured to determine an accurate estimate of the location in space where a particular eye is fixating, using geometric relationships. Such processes are highly computationally intensive with respect to two eyes, particularly in light of the resources available on portable systems such as head-mounted components (58) characterized by onboard integrated processors and limited power. Deep learning techniques may be trained and used to address these and other computational challenges.
[0169] For example, in one embodiment, a deep learning network may be used to perform the segmentation portion of the eye-tracking paradigm described above, while all else remains the same (i.e., a deep convolutional network may be used for robust pixel-by-pixel segmentation of the left and right eye images into iris, pupil, sclera, and the rest of the spectrum). Such a configuration occupies one of the larger computationally intensive parts of the process and makes it significantly more efficient. In another embodiment, a single co-deep learning model may be trained and used to perform segmentation, pupil detection, and flash detection (i.e., a deep convolutional network may be used for robust pixel-by-pixel segmentation of the left and right eye images into iris, pupil, sclera, and the rest of the spectrum. Eye segmentation may then be used to narrow down the 2-D flash field area of an active, inward-facing LED light source). Geometric shape calculations for determining the line of sight may then be performed. Such a paradigm also makes the computation more efficient. In a third embodiment, a deep learning model may be trained and used to estimate gaze direction based directly on two images of the eyes resulting from an inward-facing camera (i.e., in such an embodiment, a deep learning model using only a photograph of the user's eyes may be configured to tell the system where the user is gazing in three-dimensional space. A deep convolutional network may be used for robust pixel-by-pixel segmentation of the left and right eye images into iris, pupil, sclera, and rest of the kind. The eye segmentation is then performed on two active inward-facing LED illuminators. -D flashing field locations may be used to narrow down the field of view. 2-D flashing field locations, along with 3-D LED locations, may be used to detect the corneal center in 3-D. Note that all 3-D locations may be within separate camera coordinate systems. Eye segmentation may then also be used to detect the pupil center in the 2-D image using elliptic fitting. Using offline calibration information, the 2-D pupil center may be mapped to the 3-D line of sight point, with the depth determined during calibration. (The line connecting the corneal 3-D location and the 3-D line of sight location is the line of sight vector for that eye.)Such a paradigm may also streamline computation, and the associated deep network may be trained to directly predict 3-D gaze points based on left and right images. The loss function for such a deep network to perform such training may be a simple Euclidean loss, or it may also include well-known geometric constraints of the eye model.
[0170] Furthermore, deep learning models may be included for biometric identification using images of the user's iris from an inward-facing camera. Such models may also be used to determine whether the user is wearing contact lenses, as the model will be evident in the Fourier transform of the image data from the inward-facing camera.
[0171] The use of outward-facing cameras, such as those depicted in Figure 16A (124, 154, 156), may be used to perform SLAM or PTAM analysis to determine the orientation of the user's head, etc., in the environment in which the head-mounted component (58) is worn, as described above. Most SLAM techniques rely on tracking and matching geometric features, as described in the embodiments above. Generally, this is useful in a "textured" world where outward-facing cameras can detect corners, edges, and other features. Furthermore, certain assumptions may be made about the persistence / statics of features detected in the scene, and it is useful to have significant computational and power resources available for all of this mapping and tracking analysis using the SLAM or PTAM process. Such resources may be insufficient in some systems, such as some that are portable or wearable and have limited built-in processing power and power. Deep learning networks can be incorporated into various embodiments to observe differences within image data and, based on their training and configuration, can play a crucial role in SLAM analysis of modified versions of the system described herein (in the context of SLAM, deep networks as used herein may be considered “deep SLAM” networks).
[0172] In one embodiment, a deep SLAM network may be used to estimate the pose between a pair of frames captured by a camera coupled to a component to be tracked, such as a head-mounted component (58) of an augmented reality system. The system may comprise a convolutional neural network configured to learn pose transformations (e.g., the pose of the head-mounted component 58) and apply them in a tracking manner. The system may be configured to start with looking at a specific vector and orientation, such as straight ahead, at a known origin (0,0,0 as X,Y,Z). The user's head may then be moved, for example, slightly to the right, then slightly to the left, between frame 0 and frame 1, with the goal of determining a pose transformation or relative pose transformation. The associated deep network may be trained on a pair of images, for example, to grasp pose A and pose B and image A and image B, which leads to a certain pose transformation. Once an attitude change is determined, associated IMU data (from accelerometers, gyroscopes, etc., as described above) is then integrated into the attitude change, and tracking can continue regardless of the trajectory as the user moves around the room away from the origin. Such a system may be called a “relative attitude network,” which, as described above, is trained on pairs of frames and has known attitude information available (the change is determined frame by frame based on variations in the actual image, and the system learns the attitude change in terms of translation and rotation). Deep homography estimation or relative attitude estimation is discussed, for example, in U.S. Patent Application No. 62 / 339,799 (which is incorporated herein by reference in its entirety).
[0173] When such a configuration is used to perform attitude estimation from frame 0 to frame 1, the results are generally not perfect, and the system must have means to compensate for drift. As the system moves forward from frame 1 to 2, 3, and 4, and estimates the relative attitude, a small amount of error is introduced between each pair of frames. This error generally accumulates and becomes problematic (for example, if this error-based drift is not compensated for, the system may use attitude estimation to place the user and their associated system components in the wrong location and orientation). In one embodiment, the concept of “loop closure” may be applied to solve what may be called the “repositioning” problem. In other words, the system may be configured to determine whether a particular location has been previously seen and, if applicable, whether the predicted attitude information is reasonable in light of previous attitude information with respect to that location. For example, the system may be configured to reposition each time it sees a frame on the map that was previously seen. If a translation is, for example, shifted by 5 mm in the X direction, and a rotation is, for example, shifted by 5 degrees in the theta direction, the system corrects this discrepancy along with those of other associated frames. Thus, the trajectory becomes correct, as opposed to an incorrect one. Repositioning is discussed in U.S. Patent Application No. 62 / 263,529 (which is incorporated herein by reference in its entirety).
[0174] Furthermore, it is known that noise exists within the determined position and orientation data, particularly when attitude is estimated using IMU information (i.e., data from associated accelerometers, gyroscopes, and equivalents, as described above). If such data is used directly by the system without further processing, for example, to present an image, there is a high probability of undesirable jitter and instability being experienced by the user. This is why Kalman filters, sensor fusion techniques, and smoothing functions may be used in some techniques, including some of those described above.
[0175] The smoothing problem can be addressed using a recurrent neural network or RNN, which is similar to a long-short-term memory network, by employing deep network solutions such as those described above, including pose estimation using a convolutional neural network. In other words, the system may be configured to build a convolutional neural network on which an RNN is placed. Conventional neural networks are, by design, feedforward and temporally static. They assume an image or a pair of images, and they produce a response. With an RNN, the output of a layer is added to the next input and fed back into the same layer again. This is typically the only layer in the network and can be thought of as a "passage through time." That is, at each point in time, the same network layer re-examines the input, slightly temporally adjusted, and this cycle is repeated.
[0176] Furthermore, unlike feedforward networks, RNNs can receive sequences of values as input (i.e., sequenced over time) and produce sequences of values as output. The simple structure of the RNN allows it to be built into a feedback loop, which enables it to behave like a predictive engine, and as a result, when combined with the convolutional neural network in this embodiment, the system will take relatively noisy trajectory data from the convolutional neural network, pass it through the RNN, and output a much smoother, much more human-like trajectory, such as the movement of a user's head, which can be coupled to a head-mounted component (58) of a wearable computing system.
[0177] The system may also be configured to determine the depth of an object from a pair of images of a 3D object, and may have a deep network to which left and right images are input. The convolutional neural network may be configured to output differences between left and right cameras (e.g., between the left-eye camera and the right-eye camera on a head-mounted component 58). The determined difference is the reciprocal of the depth, given that the focal lengths of the cameras are known, and therefore the system can be configured to efficiently calculate the depth with the difference information. Meshing and other processes may then be performed without the use of alternative components for sensing depth, such as depth sensors, which may require relatively high computational and power resource loads.
[0178] Regarding the application of deep networks to various embodiments of semantic analysis and augmented reality constructions of this subject, some areas of particular interest and availability include, but are not limited to, gesture and keypoint detection, facial recognition, and 3-D object recognition.
[0179] With regard to gesture recognition, in various embodiments, the system is configured to recognize certain gestures made by the user to control the system. In one embodiment, the built-in processor may be configured to recognize certain gestures made by the user, using something known as a "random forest" along with sensed depth information. The random forest model is a non-deterministic model, which may require a very large library of parameters and may require relatively large processing and, therefore, power consumption.
[0180] Furthermore, due to noise limitations inherent in a particular depth sensor and its inability to accurately determine differences between depths of, for example, 1 or 2 cm, depth sensors may not always be optimally suited for reading hand gestures against a background such as a desk or table surface or a wall near the depth of the hand in question. In some embodiments, random forest-type gesture recognition may be replaced with a deep learning network. One of the challenges in utilizing deep networks for such configurations is the labeling of image information, such as pixels, to distinguish between "hand" and "non-hand." Training and utilizing deep networks with such segmentation challenges may require segmenting using millions of images, which is extremely expensive and time-consuming. To address this, in one embodiment, an infrared camera, such as one available for military or security purposes, may be coupled to a conventional outward-facing camera during training time, so that the infrared camera performs the segmentation of "hand" and "non-hand" itself by essentially indicating the parts of the image that are hot enough to be a human hand and the parts that are not.
[0181] With regard to facial recognition, assuming that the augmented reality system of this subject is configured to be worn in a social setting with other people, understanding people around the user can be relatively useful not only for identifying other nearby people, but also for adjusting the information presented (for example, if the system identifies a nearby person as an adult friend, it may suggest and assist in playing chess; if the system identifies a nearby person as the user's child, it may suggest and assist in going out to play soccer; if the system fails to identify a nearby person or identifies them as a known danger, the user may be prompted to avoid contact with such person).
[0182] In one embodiment, a deep neural network configuration may be used to assist face recognition in a manner similar to that described above in relation to deep relocation. The model may be trained on several different faces related to the user's life, and then, when a face approaches the vicinity of the system, such as the vicinity of a head-mounted component (58), the system can capture an image of the face in pixel space, convert it into, for example, a 128-dimensional vector, and then use the vector as a point in a higher-dimensional space to determine whether the person is in a list of known people. In short, the system may be configured to perform a “nearest neighbor” search in that space, and ultimately, such a configuration can be very accurate, with a false positive rate in the range of 1 in 1000.
[0183] With regard to 3-D object detection, in one embodiment, it is useful to incorporate a deep neural network that will tell the user about the space in which they exist from a three-dimensional perspective (i.e., not just walls, floors, and ceilings, and not merely from conventional two-dimensional perception, but from true three-dimensional perception of objects filling a room such as sofas, chairs, cabinets, and equivalents). For example, in one embodiment, it is desirable for the user to have a model that understands the true three-dimensional boundaries of a sofa in a room, so that the user can grasp the volume occupied by the sofa's volume, for example, if a virtual ball or other object is thrown. The deep neural network model may be used to form a cuboid model with a high level of sophistication.
[0184] In one embodiment, a deep reinforcement network or deep reinforcement learning may be used to effectively learn what an agent should do within a specific context without the user always having to directly communicate it to the agent. For example, if a user always wants a virtual representation of their dog to roam around a room they occupy, but also wants the dog's representation to always be visible (i.e., not hidden behind a wall or cabinet), a deep reinforcement approach may transform the scenario into a kind of game, where the virtual agent (in this case, the virtual dog) is allowed to wander around the physical space near the user, but is rewarded if the dog stays in an acceptable location during training time, for example, from T0 to T1, and penalized if the user's view of the dog is blocked, lost, or hits a wall or object. In such an embodiment, the deep network begins learning what needs to be done to gain points rather than lose them, and quickly grasps what needs to be grasped to provide the desired functionality.
[0185] The system may also be configured to address the lighting of the virtual world in a manner that approximates or matches the lighting of the real world around the user. For example, to blend virtual perception with real perception in augmented reality as best as possible, the lighting color, shading, and lighting vectors are reproduced as realistically as possible along with the virtual object. In other words, if a virtual opaque coffee cup is to be positioned on the surface of a real table in a room, with yellowish-toned light emanating from one particular corner of the room that creates shading from a real-world object on the real-world table, the lighting color and shading of the virtual coffee cup would best match the real-world scenario. In one embodiment, a deep learning model may be used to learn the lighting of a real environment in which the system components are installed. For example, the model may be used to learn the lighting of a room, assuming images or sequences of images from a real environment, and to determine factors such as brightness, hue, and vectors from one or more light sources. Such a model may be trained from synthetic data and images captured from a user device, such as the user's head-mounted component (58).
[0186] Referring to Figure 29, a deep learning network architecture which may be called a "Hydra" architecture (272) is illustrated. In such a configuration, various inputs (270) such as IMU data (from accelerometers, gyroscopes, magnetometers), outward-facing camera data, depth-sensing camera data, and / or sound or speech data may be channeled to a multi-layer central processing resource having a group of lower layers (268) which then passes the results to a group of central layers (266) which in turn pass the results to one or more of a group of associated "heads" (264) that represent various process functionalities such as face recognition, visual search, gesture recognition, semantic segmentation, object detection, daylight detection / determination, SLAM, relocalization, and / or depth estimation (from stereoscopic image information, etc., as described above).
[0187] Traditionally, when using deep networks to accomplish various tasks, algorithms would be built for each task. Therefore, if recognizing a car is desired, an algorithm would be built for that purpose. If recognizing a face is desired, an algorithm would be built for that purpose, and these algorithms may be run simultaneously. Such a configuration would work well and yield results if unlimited or high levels of power and computing resources were available. However, in many scenarios, such as those involving portable augmented reality systems with limited power sources and limited processing power within an internal processor, computing and power resources can be relatively limited, and it may be desirable to handle certain aspects of the task together. Furthermore, there is evidence that if one algorithm possesses knowledge from another, it can make the second algorithm perform better. For example, if one deep network algorithm understands dogs and cats, knowledge transfer from it (also called "domain fitting") may help another algorithm better recognize shoes. Therefore, it is reasonable to have some kind of crosstalk between algorithms during training and estimation.
[0188] Furthermore, there are considerations related to algorithm design and modification. Preferably, if additional capabilities are required for an initial version of the algorithm, it will not be necessary to completely rebuild a new one from scratch. The described hydra architecture (272) may be used to address these challenges and computational and power efficiency challenges, as there may be common aspects of certain computational processes that can be shared, as described above. For example, in the described hydra architecture (272), inputs (270), such as image information from one or more cameras, may be brought into a lower layer (268) where relatively low-level feature extraction may be performed. For example, Gabor functions, Gaussian derivatives, essentially yielding lines, edges, angles, and colors, which are uniform with respect to many problems at the low level. Therefore, regardless of task variation, low-level feature extraction can be identical whether the goal is to extract cats, cars, or cows, and thus the associated computations can be shared. The Hydra architecture (272) is a high-level paradigm that enables knowledge sharing across algorithms, making each one better, and allows feature sharing so that computations are shared, reduced, and not redundant, and makes it possible to extend a set of capabilities without having to rewrite everything, or rather, new capabilities may be stacked on top of existing capabilities.
[0189] Therefore, as described above, in the embodiments described, the Hydra architecture represents a deep neural network with one integrated path. The lower layers (268) of the network are shared, and they extract basic units of visual primitives from the input image and other inputs (270). The system may be configured to extract edges, lines, contours, confluences, and equivalents through several layers of convolution. The basic components that programmers used to feature engineer are now learned by the deep network. Ultimately, these features are useful for many algorithms, whether the algorithm is face recognition, tracking, etc. Therefore, once the lower computational work is done and a shared representation from the image or other inputs exists for all the other algorithms, there may be one individual path per problem. Thus, on this shared representation, for the other "heads" (264) of the architecture (272), there may be a path leading to very specific face recognition for faces, a path leading to very specific tracking for SLAM, and so on. In such embodiments, on the one hand, there is all of this shared calculation which allows for the addition to be essentially increased, and on the other hand, there is a very specific path which allows for fine-tuning based on general knowledge and finding answers to very specific questions.
[0190] Furthermore, a beneficial aspect of such a configuration is the fact that such a neural network is designed so that the lower layers (268) closer to the input (270) require more computation at each layer of computation, as the system takes the original input and transforms it into some other dimensional space where the dimensionality of things is typically reduced. Thus, once the fifth layer from the bottom of the network is reached, the amount of computation can be less than one-twentieth of what was required at the lowest level (i.e., because the input is much larger and much larger matrix multiplications were required). In one embodiment, there is considerable uncertainty about the problem that needs to be solved until the system extracts a shared computation result. The majority of the computation of almost any algorithm is completed in the lower layers, and therefore, when new paths are added for face recognition, tracking, depth, daylighting, and equivalents, these contribute relatively little to the computation constraints, and thus such an architecture offers rich capacity for expansion.
[0191] In one embodiment, pooling may be absent for some first layers in order to reserve the highest resolution data. The middle layer may have a pooling process because ultra-high resolution is not required at that point (for example, ultra-high resolution is not required in the middle layer to determine the location of the car's wheels; in fact, it is only necessary to determine the location of the nuts and bolts from the lower levels of high resolution, and then the image data can be significantly compressed as it passes through to the middle layer regarding the location of the car's wheels).
[0192] Furthermore, once the network has all the learned connections, everything is loosely connected, and the connections are learned favorably through the data. The middle layer (266) may be configured to begin learning some, for example, object parts, facial features, and their equivalents. Thus, rather than a simple Gabor function, the middle layer handles more complex structures (i.e., curved shapes, shading, etc.). Then, as the process moves higher up, there is a division into unique head components (264), some of which may have many layers, and some of which may have few. Again, scalability and efficiency largely owe to the fact that the majority, such as 90% of the processing flops, reside in the lower layers (268), then a smaller portion, such as 5% of the flops, resides in the middle layer (266), and another 5% resides in the head (264).
[0193] Such a network may be pre-trained using already existing information. For example, in one embodiment, images from a large group (up to 10 million) of ImageNets from a large group of classes (up to 1,000) may be used to train all of the classes. In one embodiment, once trained, the upper layers that distinguish classes may be removed, but all weights learned during the training process are preserved.
[0194] Referring to Figure 30A, a pair of coils (302, 304) are shown in a configuration with a specific radius and spacing between them, which may be known as a "Helmholtz coil".
[0195] Helmholtz coils appear in various configurations (here, a pair of rounded coils are shown), such as those depicted in Figure 30B (306), and are known to produce a relatively uniform magnetic field through a given volume. Magnetic field lines are shown with arrows, centered on the cross-sectional view of the coils (302, 304) in Figure 30B. Figure 3 illustrates a three-axis Helmholtz coil configuration, with three pairs (310, 312, 314) oriented orthogonally as shown. Other variations of Helmholtz or Merritt coils, such as those featuring rectangular coils, may also be used to create a predictable and relatively uniform magnetic field through a given volume. In one embodiment, a Helmholtz-type coil may be used to assist in calibrating the orientation, determining the relationship between two sensors operably coupled to a head-mounted component (58), such as those described above. For example, referring to Figure 30D, as described above, a head-mounted component (58), coupled to the IMU (102) and electromagnetic field sensor (604), may be placed within a known magnetic field volume of a Helmholtz coil pair (302, 304). With current applied through the coil pair (302, 304), the coils may be configured to generate a magnetic field at a selectable frequency. In one embodiment, the system may be configured to excite the coils at a DC current level to produce a readable output directly from the magnetometer component of the IMU (102). The coils may then be excited at an alternating current level to produce a readable output directly from, for example, an electromagnetic positioning receiver coil (604). Since the fields they apply in such a configuration are generated by the same physical coils (302, 304), they are aligned with each other and it is understood that the fields should have the same orientation. Therefore, values may be read from the IMU (102) and electromagnetic field sensor (604) and the calibration may be measured directly, which may be used to characterize any difference in orientation readings between two devices (102, 604) in three dimensions and thus provide a usable calibration between the two for runtime. In one embodiment, the head-mounted component (58) may be electromechanically reoriented for further testing against the coil set (302, 304).In another embodiment, the coil sets (302, 304) may be electromechanically reoriented for further testing of the head-mounted component (58). In another embodiment, the head-mounted component (58) and the coil sets (302, 304) may be electromechanically reoriented relative to each other. In another embodiment, a triaxial Helmholtz coil or other more advanced magnetic field production coil, such as the one depicted in Figure 30C, may be used to generate magnetic fields and components for additional test data without the need to reoriented the head-mounted component (58) relative to the coil sets (302, 304).
[0196] Referring to Figure 30E, a system or subsystem used in such a calibration configuration to produce a predictable magnetic field, such as a pair of coils (302, 304) in a Helmholtz-type configuration, may have one or more optical base points (316) coupled to one or more cameras (124), which may comprise a head-mounted component (58), so that such base points can be viewed. Such a configuration provides an opportunity to ensure that the electromagnetic sensing subsystem is aligned with the cameras in a known manner. In other words, using such a configuration results in having optical base points that are physically coupled to or anchored to the magnetic field generating device in a known or measured manner (for example, an articulated coordinate measuring machine may be used to establish precise X, Y, Z coordinates for each base point location 316). The head-mounted component (58) is installed inside the test volume and may be exposed to a magnetic field, while the camera (124) of the head-mounted component (58) observes one or more base points (316) and thus calibrates the extrinsicity of the magnetic field sensor and camera (because the magnetic field generator is mounted on the base point observed by the camera). The optical base points (316) may have flat features such as a checkerboard pattern, an aruco marker, a textured surface, or other three-dimensional features. The optical base points may also be dynamic, such as in configurations where a small display such as an LCD display is used. They may be static and printed. They may be etched into the substrate material by laser or chemical means. They may have coatings or anodizing or other features recognizable by the camera (124). In a factory calibration setup, multiple calibration systems, such as those described herein, may be located adjacent to each other and may be timed so that adjacent systems do not produce magnetic fields that would interfere with readings in adjacent systems. In one embodiment, a group of calibration stations may be time-series, and in another embodiment, they may operate simultaneously to provide functional isolation, such as every other station, every two stations, or every three stations.
[0197] Figure 31 illustrates a simplified computer system 3100 according to embodiments described herein. The computer system 3100 may be incorporated into devices described herein, as shown in Figure 31. Figure 31 provides a schematic illustration of one embodiment of the computer system 3100 that can carry out some or all of the steps of the method provided by various embodiments. It should be noted that Figure 31 is intended only to provide a generalized illustration of various components, and any or all of them may be used as needed. Figure 31 thus illustrates, in a broad sense, how individual system elements may be implemented in a relatively separate or more relatively integrated manner.
[0198] The computer system 3100 is shown to include hardware elements that can be electrically coupled via bus 3105 or otherwise communicate as needed. The hardware elements may include one or more processors 3110, including, but not limited to, one or more general-purpose processors and / or one or more special-purpose processors such as a digital signal processing chip, a graphics accelerator, and / or equivalent; one or more input devices 3115, which may include, but not limited to, a mouse, a keyboard, a camera, and / or equivalent; and one or more output devices 3120, which may include, but not limited to, a display device, a printer, and / or equivalent.
[0199] The computer system 3100 may further include, but is not limited to, local and / or network-accessible storage devices, and / or may include, but is not limited to, one or more non-transient storage devices 3125, which may include disk drives, drive arrays, optical storage devices, solid-state storage devices, for example, random-access memory ("RAM") and / or read-only memory ("ROM"), and / or equivalents, which may be programmable and flash-updatable, and / or communicate with such storage devices. Such storage devices may be configured to implement any suitable data storage, including, but is not limited to, various file systems, database structures, and / or equivalents.
[0200] The computer system 3100 may also include a communication subsystem 3119, which may include, but is not limited to, a modem, a network card (wireless or wired), an infrared communication device, a wireless communication device, and / or a chipset, such as a Bluetooth® device, an 802.11 device, a WiFi device, a WiMAX device, a cellular communication equipment, and / or equivalent. The communication subsystem 3119 may include one or more input and / or output communication interfaces, which may enable data to be exchanged with a network as described below, i.e., a network such as, for example, another computer system, a television, and / or any other device described herein. Depending on the desired functionality and / or other implementation concerns, a portable electronic device or similar device may communicate images and / or other information via the communication subsystem 3119. In other embodiments, a portable electronic device, e.g., the first electronic device, may be incorporated into the computer system 3100, for example, as an input device 3115. In some embodiments, the computer system 3100 will further include a working memory 3135, which may include a RAM or ROM device as described above.
[0201] The computer system 3100 may also include computer programs provided by various embodiments and / or software elements, indicated to be currently located in working memory 3135, including an operating system 3140, device drivers, executable libraries, and / or other code, such as one or more application programs 3145, which may be designed to implement and / or constitute a system, as provided by other embodiments as described herein. Simply as an example, one or more procedures described with respect to the methods discussed above may be implemented as code and / or instructions executable by a computer and / or a processor within a computer. In some respects, such code and / or instructions may then be used to configure and / or adapt a general-purpose computer or other device and perform one or more operations in accordance with the methods described.
[0202] These instruction and / or code sets may be stored on a non-transient computer-readable storage medium, such as the storage device 3125 described above. In some cases, the storage medium may be incorporated into a computer system, such as computer system 3100. In other embodiments, the storage medium may be separate from the computer system, such as a removable medium, such as a compact disk, and / or provided within an installation package, so that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instruction / code stored thereon. These instructions may take the form of executable code that can be executed by computer system 3100, and / or in the form of source and / or installable code, which then take the form of executable code, depending on compilation and / or installation onto computer system 3100 using, for example, one of various generally available compilers, installation programs, compression / decompression utilities, etc.
[0203] Various exemplary embodiments of the present invention are described herein. These embodiments are used for non-limiting purposes only. They are provided to illustrate broader aspects of the present invention. Various modifications may be made to the described invention, and equivalents may be substituted without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt specific situations, materials, compositions, processes, process actions, or steps to the object, spirit, or scope of the invention. Furthermore, as will be understood by those skilled in the art, each of the individual modifications described and illustrated herein has discrete components and features that are readily separable from or combined with features of any of several other embodiments without departing from the scope or spirit of the invention. All such modifications are intended to be within the scope of the claims associated with this disclosure.
[0204] The present invention includes methods that may be carried out using the device of the subject matter. The methods may include the act of providing such a suitable device. Such provision may be made by an end user. In other words, the act of “providing” simply requires the end user to obtain, access, approach, position, configure, activate, power on, or otherwise operate the device that is essential in the method of the subject matter. The methods enumerated herein may be carried out in any logically possible order of the described events, and in the order in which the events are described.
[0205] Exemplary aspects of the present invention, along with details relating to material selection and manufacturing, are described above. Further details of the present invention are understood in relation to the above-referenced patents and publications and, generally, can be grasped or understood by those skilled in the art. The same may apply to the method-based aspects of the present invention in terms of additional effects that may be adopted in general or logically.
[0206] In addition, although the present invention has been described with reference to several embodiments that optionally incorporate various features, the present invention is not limited to those described or illustrated as to be considered with respect to each modification of the invention. Various modifications may be made to the described invention, and equivalents (whether described herein or not, or not included for some brevity) may be substituted without departing from the true spirit and scope of the invention. In addition, where a range of values is provided, it is understood that all intervening values between the upper and lower limits of that range, and any other provisions or intervening values within that defined range, are encompassed within the present invention.
[0207] Furthermore, it is considered that any optional feature of the variations of the invention described herein may be described and claimed independently or in combination with one or more of the features described herein. References to singular items include the possibility of multiple identical items existing. More specifically, as used herein and in the claims associated therewith, the singular forms “a,” “an,” “said,” and “the” include multiple referents unless otherwise specifically stated. In other words, the use of articles allows for “at least one” of the items of reference in the above description and in the claims associated with the invention. Furthermore, it should be noted that such claims may be drafted to exclude any optional elements. Thus, this statement is intended to function as an antecedent for the use of such exclusive terms, or “negative” restrictions, such as “only,” “only,” and equivalents, relating to the description of the claim elements.
[0208] Without using such exclusive technical terms, the term “equipped with” in the claims associated with this disclosure shall allow for the inclusion of any additional elements, whether a given number of elements are enumerated in such claims or whether the addition of features can be considered to transform the nature of the elements described in such claims. Unless otherwise specifically defined herein, all technical and scientific terms used herein should be given the broadest possible generally understood meaning while maintaining the validity of the claims.
[0209] The scope of the present invention should not be limited to the provided examples and / or the specification of this subject matter, but rather should be limited only by the language of the claims associated with this disclosure.
Claims
1. A method for determining the position and orientation of a handheld controller in a system comprising one or more sensors, wherein the method is: Based on detecting the magnetic field emitted by the handheld controller, one or more sensors positioned within the headset of the system determine the first position and first orientation of the handheld controller of the system within a first hemisphere relative to the headset, Based on detecting the magnetic field, one or more sensors determine the second position and second orientation of the handheld controller within the second hemisphere relative to the headset, wherein the second hemisphere faces the first hemisphere relative to the headset, the first and second hemispheres form a sphere centered on the headset, and a reference plane passing through the headset divides the sphere into the first and second hemispheres. Determining a position vector that identifies the position of the handheld controller relative to the headset within the first hemisphere, Determining a surface normal vector that is perpendicular to the reference plane and extends from the headset into the first hemisphere, Based on the relationship between the position vector and the surface normal vector, When the angle between the surface normal vector and the position vector is less than 90 degrees, it is determined that the first position and first orientation of the handheld controller are accurate. When the angle between the surface normal vector and the position vector is greater than 90 degrees, it is determined that the second position and second orientation of the handheld controller are accurate. Methods that include...
2. The method according to claim 1, wherein the position vector is defined in the coordinate frame of the headset.
3. The method according to claim 1, wherein the surface normal vector originates from the headset and extends at an angle downward from a horizontal line from the headset, and the horizontal line is perpendicular to the display of the headset.
4. The method according to claim 1, wherein when the first position and first orientation of the handheld controller are accurate, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset.
5. The method according to claim 1, wherein when the second position and second orientation of the handheld controller are accurate, the second hemisphere is identified as the front hemisphere relative to the headset, and the first hemisphere is identified as the rear hemisphere relative to the headset.
6. The aforementioned method, The first position and first orientation of the handheld controller when the first position and first orientation are accurate, or The second position and second orientation of the handheld controller when the second position and second orientation are accurate. The method according to claim 1, further comprising delivering virtual content to a display based on the above.
7. The method according to claim 1, wherein the system includes an optical device.
8. The method according to claim 1, wherein the method is performed during the initialization process of the headset.
9. A system, wherein the system is A headset equipped with one or more sensors, A handheld controller, A processor coupled to the headset, configured to perform an operation. Equipped with, The aforementioned operation is, Based on detecting the magnetic field emitted by the handheld controller, one or more sensors positioned within the headset determine the first position and first orientation of the handheld controller within the first hemisphere relative to the headset, Based on detecting the magnetic field, one or more sensors determine the second position and second orientation of the handheld controller within the second hemisphere relative to the headset, wherein the second hemisphere faces the first hemisphere relative to the headset, the first and second hemispheres form a sphere with the headset at its center, and a reference plane divides the sphere into the first and second hemispheres. Determining a position vector that identifies the position of the handheld controller relative to the headset within the first hemisphere, Determining a surface normal vector that is perpendicular to the reference plane and extends from the headset into the first hemisphere, Based on the relationship between the position vector and the surface normal vector, When the angle between the surface normal vector and the position vector is less than 90 degrees, it is determined that the first position and first orientation of the handheld controller are accurate. When the angle between the surface normal vector and the position vector is greater than 90 degrees, it is determined that the second position and second orientation of the handheld controller are accurate. A system that includes this.
10. The system according to claim 9, wherein the position vector is defined in the coordinate frame of the headset.
11. The system according to claim 9, wherein the surface normal vector extends at an angle downward from the horizontal line from the headset, and the horizontal line is perpendicular to the display of the headset.
12. The system according to claim 9, wherein when the first position and first orientation of the handheld controller are accurate, the first hemisphere is identified as the front hemisphere relative to the headset, and the second hemisphere is identified as the rear hemisphere relative to the headset.
13. The system according to claim 9, wherein when the second position and second orientation of the handheld controller are accurate, the second hemisphere is identified as the front hemisphere relative to the headset, and the first hemisphere is identified as the rear hemisphere relative to the headset.
14. The aforementioned processor, The first position and first orientation of the handheld controller when the first position and first orientation are accurate, or The second position and second orientation of the handheld controller when the second position and second orientation are accurate. The system according to claim 9, further configured to perform operations including delivering virtual content to a display based on the above.
15. The system according to claim 9, wherein the operation is performed during the initialization process of the headset.
16. A computer program product embodied in a non-transient computer-readable medium, wherein the non-transient computer-readable medium stores a sequence of instructions, which, when executed by a processor, causes the processor to perform a method for determining the position and orientation of a handheld controller in a system comprising one or more sensors. The aforementioned method, Based on detecting the magnetic field emitted by the handheld controller, one or more sensors positioned within the headset of the system determine the first position and first orientation of the handheld controller of the system within a first hemisphere relative to the headset, Based on detecting the magnetic field, one or more sensors determine the second position and second orientation of the handheld controller within the second hemisphere relative to the headset, wherein the second hemisphere faces the first hemisphere relative to the headset, the first and second hemispheres form a sphere with the headset at its center, and a reference plane divides the sphere into the first and second hemispheres. Determining a position vector that identifies the position of the handheld controller relative to the headset within the first hemisphere, Determining a surface normal vector that is perpendicular to the reference plane and extends from the headset into the first hemisphere, Based on the relationship between the position vector and the surface normal vector, When the angle between the surface normal vector and the position vector is less than 90 degrees, it is determined that the first position and first orientation of the handheld controller are accurate. When the angle between the surface normal vector and the position vector is greater than 90 degrees, it is determined that the second position and second orientation of the handheld controller are accurate. Computer program products, including [this].
Citation Information
Patent Citations
Multimodal interface device and multimodal interface method
JP1998301675A
Position-direction measuring apparatus and information processing method
JP2003240532A
Information processing system, information processing device, information processing program, and information processing method
JP2014038588A
Augmented reality display device with deep learning sensors
JP2019532392A
System and method for hemisphere disambiguation in electromagnetic tracking systems
US20050062469A1