Transmodal Input Fusion for Wearable Systems
The wearable system dynamically fuses multiple user input modes to enhance interaction accuracy in VR, AR, and MR environments, addressing the challenge of integrating diverse inputs for precise object selection and action in 3D spaces, thereby improving precision and reducing hardware complexity.
Patent Information
- Application Number
- JP2024094291
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-06-29
- Filing Date
- 2024-06-11
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2039-05-21
AI Technical Summary
Existing VR, AR, and MR technologies face challenges in providing comfortable, natural-feeling, and rich presentations of virtual image elements within complex human visual environments, as they struggle to accurately interpret and integrate multiple user input modes for interaction with both virtual and real-world objects.
A wearable system dynamically fuses multiple user input modes, such as head pose, eye gaze, hand gestures, voice commands, and environmental factors, to enhance interaction accuracy by identifying convergent inputs and disregarding divergent ones, using a hardware processor to analyze and combine these inputs for precise object selection and action in 3D environments.
This approach improves interaction precision and reduces errors by dynamically selecting and weighting relevant input modes, allowing for more accurate and efficient user interactions in dynamic 3D environments, while potentially reducing hardware complexity and cost.
Smart Images

Figure 0007802863000009 
Figure 0007802863000010 
Figure 0007802863000011
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 675,164, filed May 22, 2018, entitled "Electromyographic Sensor Prediction in Augmented Reality," and U.S. Provisional Patent Application No. 62 / 692,519, filed June 29, 2018, entitled "Transmodal Input Fusion for a Wearable System," the disclosures of which are incorporated herein by reference in their entireties.
[0002] (Copyright Notice) A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to anyone copying this patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.
[0003] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly to dynamically blending multiple user input modes to facilitate interaction with virtual objects within a three-dimensional (3D) environment. [Background technology]
[0004] Modern computing and display technologies have facilitated the development of systems for so-called “virtual reality,” “augmented reality,” or “mixed reality” experiences in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality or “VR” scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality or “AR” scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality or “MR” relates to the merging of real and virtual worlds to generate new environments in which physical and virtual objects coexist and interact in real time. Consequently, the human visual perception system is highly complex, making it challenging to produce VR, AR, or MR technologies that facilitate comfortable, natural-feeling, and rich presentations of virtual image elements among other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies. Summary of the Invention [Means for solving the problem]
[0005] Examples of wearable systems and methods described herein can use multiple inputs (e.g., gestures, head pose, eye gaze, voice, from a user input device, or environmental factors (e.g., location)) to determine commands to be performed or objects in a three-dimensional (3D) environment to be acted on or selected. Multiple inputs can also be used by the wearable device to enable a user to interact with physical objects, virtual objects, text, graphics, icons, user interfaces, etc.
[0006] For example, a wearable display device can be configured to dynamically analyze multiple sensor inputs for task execution or object targeting. The wearable device can dynamically use a combination of multiple inputs, such as head pose, eye gaze, hand, arm, or body gestures, voice commands, user input devices, and environmental factors (e.g., the user's location or objects around the user), to determine an object in the user's environment that the user intends to select or an action the wearable device can perform. The wearable device can dynamically select a set of sensor inputs that together indicate the user's intent to select the target object (inputs that provide independent or auxiliary indications of the user's intent to select the target object may be referred to as convergent or converging inputs). The wearable device can combine or fuse inputs from this set (e.g., to improve the quality of user interaction as described herein). If sensor inputs from this set later indicate divergence from the target object, the wearable device can discontinue use of (or reduce the relative weighting given to) the diverging sensor inputs.
[0007] The process of dynamically using convergent sensor inputs while ignoring (or reducing the relative weighting given to) divergent sensor inputs is sometimes referred to herein as transmodal input fusion (or simply transmodal fusion) and can provide substantial advantages over techniques that simply receive inputs from multiple sensors. Transmodal input fusion can anticipate or even predict, on a dynamic, real-time basis, which of many possible sensor inputs will be the appropriate modal input that communicates the user's intent to target or act on real or virtual objects within the user's 3D AR / MR / VR environment.
[0008] Details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, drawings, and claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter. The present specification also provides, for example, the following items: (Item 1) A wearable system, a head pose sensor configured to determine a head pose of a user of the wearable system; an eye gaze sensor configured to determine an eye gaze direction of a user of the wearable system; a gesture sensor configured to determine hand gestures of a user of the wearable system; a hardware processor in communication with the head pose sensor, the eye gaze sensor, and the gesture sensor, the hardware processor comprising: determining a first vergence between the eye gaze direction and the head pose of the user relative to an object; implementing a first interaction command associated with the object based at least in part on input from the head pose sensor and the eye gaze sensor; determining a second vergence of the hand gesture using the eye gaze direction and the head pose of the user relative to the object; and implementing a second interaction command associated with the object based at least in part on input from the hand gesture, the head pose sensor, and the eye gaze sensor; and a hardware processor programmed to perform the A wearable system comprising: (Item 2) Item 1. The wearable system of item 1, wherein the head pose sensor comprises an inertial measurement unit (IMU), the eye gaze sensor comprises an eye tracking camera, and the gesture sensor comprises an outward-facing camera. (Item 3) Item 1, a wearable system according to item 1, wherein, to determine the first vergence, the hardware processor is programmed to determine that the angle between the eye gaze direction associated with the head pose and the head pose direction is less than a first threshold. (Item 4) Item 1, the wearable system of item 1, wherein to determine the second vergence, the hardware processor is programmed to determine that a transmode triangle associated with the hand gesture, the eye gaze direction, and the head pose is less than a second threshold. (Item 5) Item 1. The wearable system of item 1, wherein the first interaction command includes targeting the object. (Item 6) Item 1. The wearable system of item 1, wherein the second interaction command includes selecting the object. (Item 7) Item 10. The wearable system of item 1, wherein the hardware processor is further programmed to determine divergence of at least one of the hand gesture, the eye gaze direction, or the head pose from the object. (Item 8) Item 1. The wearable system of item 1, wherein the first interaction command includes application of a first filter, or the second interaction command includes application of a second filter. (Item 9) Item 9. The wearable system of item 8, wherein the first filter is different from the second filter. (Item 10) Item 9. The wearable system of item 8, wherein the first filter or the second filter comprises a low-pass filter having an adaptive cutoff frequency. (Item 11) Item 11. The wearable system of item 10, wherein the low-pass filter comprises a 1 Euro filter. (Item 12) Item 1, a wearable system according to item 1, wherein, to determine the first vergence, the hardware processor is programmed to determine that the fixation time of the eye gaze direction and head pose directed at the object exceeds a first fixation time threshold. (Item 13) Item 1, a wearable system according to item 1, wherein to determine the second vergence, the hardware processor is programmed to determine that the eye gaze direction, the head pose, and the dwell time of the hand gesture relative to the object exceed a second dwell time threshold. (Item 14) Item 1, the wearable system, wherein the first interaction command or the second interaction command includes providing a stable targeting vector associated with the object. (Item 15) Item 15. The wearable system of item 14, wherein the hardware processor provides the stable targeting vector to an application. (Item 16) Item 10. The wearable system of item 1, wherein the gesture sensor comprises a handheld user input device. (Item 17) Item 17. The wearable system of item 16, wherein the hardware processor is programmed to determine a third vergence between input from the user input device and at least one of the eye gaze direction, the head pose, or the hand gesture. (Item 18) Item 10. The wearable system of item 1, further comprising an audio sensor, wherein the hardware processor is programmed to determine a fourth vergence between input from the audio sensor and at least one of the eye gaze direction, the head pose, or the hand gesture. (Item 19) 1. A system comprising: a first sensor of the wearable system configured to obtain first user input data in a first input mode; a second sensor of the wearable system configured to obtain second user input data in a second input mode, the second input mode being different from the first input mode; and a third sensor of the wearable system configured to obtain third user input data in a third input mode, the third input mode being different from the first input mode and the second input mode; and a hardware processor in communication with the first, second, and third sensors, the hardware processor comprising: receiving a plurality of inputs including the first user input data in the first input mode, the second user input data in the second input mode, and the third user input data in the third input mode; identifying a first interaction vector based on the first user input data; identifying a second interaction vector based on the second user input data; identifying a third interaction vector based on the third user input data; determining a vergence between at least two of the first interaction vector, the second interaction vector, and the third interaction vector; identifying a target virtual object from a set of candidate objects within a three-dimensional (3D) region around the wearable system based at least in part on the vergence; and determining a user interface action to the target virtual object based on at least one of the first user input data, the second user input data, the third user input data, and the vergence; generating a transmode input command that causes the user interface action to be performed on the target virtual object; and a hardware processor programmed to perform the A system comprising: (Item 20) 1. A method comprising: Under the control of the wearable system's hardware processor, accessing sensor data from more than three sensors in different modes; identifying a convergence event of a first sensor and a second sensor from the greater than three plurality of sensors of different modalities; utilizing first sensor data from the first sensor and second sensor data from the second sensor to target an object within a three-dimensional (3D) environment around the wearable system; and A method comprising: (Item 21) 1. A method comprising: Under the control of the wearable system's hardware processor, accessing sensor data from at least first and second sensors of different modalities, the first sensor providing sensor data having a plurality of potential interpretations; identifying a convergence of sensor data from the second sensor with a given one of the potential interpretations of sensor data from the first sensor; generating an input command to the wearable system based on a given one of the potential interpretations; and A method comprising: (Item 22) 1. A method comprising: Under the control of the wearable system's hardware processor, accessing sensor data from a plurality of sensors in different modes; identifying a convergence event in sensor data from a first and a second sensor of the plurality of sensors; selectively applying a filter to sensor data from the first sensor during the convergence event; A method comprising: (Item 23) 1. A method comprising: Under the control of the wearable system's hardware processor, identifying a current transmode state, the current transmode state including a transmode vergence associated with the object; identifying a region of interest (ROI) associated with said transmodal vergence; identifying a corresponding interaction field based at least in part on the ROI; selecting an input fusion method based at least in part on the transformer mode state; selecting a configuration for a primary targeting vector; applying adjustments to the primary targeting vector to provide a stable attitude vector; communicating said stable attitude vector to an application; A method comprising: (Item 24) 1. A method comprising: Under the control of the wearable system's hardware processor, Identifying a transmode fixation point; defining an extended region of interest (ROI) based on the transmode fixation duration or an expected fixation duration near the transmode fixation point; determining that the ROI intersects a rendering element; determining a rendering extension compatible with the ROI, the transmode fixation time, or the predicted fixation time; activating the rendering extension; A method comprising: (Item 25) 1. A method comprising: Under the control of the wearable system's hardware processor, receiving sensor data from a plurality of sensors in different modes; determining that data from a particular subset of the plurality of different modal sensors indicates that the user has initiated execution of a particular motor or sensorimotor control strategy from among a plurality of predetermined motor and sensorimotor control strategies; selecting a particular sensor data processing scheme corresponding to the particular motor or sensorimotor control strategy from among a plurality of different sensor data processing schemes, each corresponding to a different one of the plurality of predetermined motor and sensorimotor control strategies; processing data received from a particular subset of the plurality of different modal sensors according to the particular sensor data processing scheme; A method comprising: (Item 26) 1. A method comprising: Under the control of the wearable system's hardware processor, receiving sensor data from a plurality of sensors in different modes; determining that data from a particular subset of the plurality of sensors of the different modalities is stochastically varying in a particular manner; in response to determining that data from a particular subset of the plurality of sensors of the different modalities is stochastically varying in the particular modalities; processing data received from a particular subset of the plurality of different modal sensors according to a first sensor data processing scheme; processing data received from a particular subset of the plurality of sensors in the different modes according to a second sensor data processing scheme different from the first sensor data processing scheme; Switching between A method comprising: [Brief explanation of the drawings]
[0009] [Figure 1]FIG. 1 depicts an illustration of a mixed reality scenario with a virtual reality object and a physical object viewed by a person.
[0010] [Figure 2A] 2A and 2B diagrammatically illustrate an example of a wearable system that may be configured to use the transmodal input fusion techniques described herein. [Figure 2B] 2A and 2B diagrammatically illustrate an example of a wearable system that may be configured to use the transmodal input fusion techniques described herein.
[0011] [Figure 3] FIG. 3 diagrammatically illustrates aspects of an approach for simulating a three-dimensional image using multiple depth planes.
[0012] [Figure 4] FIG. 4 illustrates diagrammatically an embodiment of a waveguide stack for outputting image information to a user.
[0013] [Figure 5] FIG. 5 shows an exemplary output beam that may be output by a waveguide.
[0014] [Figure 6] FIG. 6 is a schematic diagram showing an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal stereoscopic display, image, or light field.
[0015] [Figure 7] FIG. 7 is a block diagram of an embodiment of a wearable system.
[0016] [Figure 8]FIG. 8 is a process flow diagram of an embodiment of a method for rendering virtual content in relation to recognized objects.
[0017] [Figure 9] FIG. 9 is a block diagram of another embodiment of a wearable system.
[0018] [Figure 10] FIG. 10 is a process flow diagram of an example method for determining user input to a wearable system.
[0019] [Figure 11] FIG. 11 is a process flow diagram of an embodiment of a method for interacting with a virtual user interface.
[0020] [Figure 12A] FIG. 12A diagrammatically illustrates an example of a field of view (FOR), a world camera field of view (FOV), a user's field of view, and a user's fixation field.
[0021] [Figure 12B] FIG. 12B diagrammatically illustrates an example of a virtual object in a user's field of view and in their oculomotor field of view.
[0022] [Figure 13] FIG. 13 illustrates an example of interacting with a virtual object using one mode of user input.
[0023] [Figure 14] FIG. 14 illustrates an example of selecting a virtual object using a combination of user input modes.
[0024] [Figure 15] FIG. 15 illustrates an example of using a combination of direct user inputs to interact with a virtual object.
[0025] [Figure 16] FIG. 16 illustrates an exemplary computing environment for aggregating input modes.
[0026] [Figure 17A] FIG. 17A illustrates an example of using lattice tree analysis to identify a target virtual object.
[0027] [Figure 17B] FIG. 17B illustrates an example of determining a target user interface action based on multimodal input.
[0028] [Figure 17C] FIG. 17C illustrates an example of aggregating confidence scores associated with input modes for a virtual object.
[0029] [Figure 18A] 18A and 18B illustrate an example of calculating confidence scores for objects within a user's FOV. [Figure 18B] 18A and 18B illustrate an example of calculating confidence scores for objects within a user's FOV.
[0030] [Figure 19A] 19A and 19B illustrate an example of using multimodal input to interact with the physical environment. [Figure 19B] 19A and 19B illustrate an example of using multimodal input to interact with the physical environment.
[0031] [Figure 20] FIG. 20 illustrates an example of automatically resizing a virtual object based on multimodal input.
[0032] [Figure 21] FIG. 21 illustrates an example of identifying a target virtual object based on the object's location.
[0033] [Figure 22A] 22A and 22B illustrate another example of interacting with a user's environment based on a combination of direct and indirect input. [Figure 22B] 22A and 22B illustrate another example of interacting with a user's environment based on a combination of direct and indirect input.
[0034] [Figure 23] FIG. 23 illustrates an example process for interacting with a virtual object using multimodal input.
[0035] [Figure 24] FIG. 24 illustrates an example of setting a direct input mode associated with user interaction.
[0036] [Figure 25] FIG. 25 illustrates an example of a user experience using multimodal input.
[0037] [Figure 26] FIG. 26 illustrates an exemplary user interface with various bookmarked applications.
[0038] [Figure 27] FIG. 27 illustrates an exemplary user interface when a search command is issued.
[0039] [Figure 28A] 28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input. [Figure 28B] 28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input. [Figure 28C]28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input. [Figure 28D] 28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input. [Figure 28E] 28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input. [Figure 28F] 28A-28F illustrate an example user experience for composing and editing text based on a combination of voice and eye gaze input.
[0040] [Figure 29] FIG. 29 illustrates an example of selecting words based on input from a user input device and gaze.
[0041] [Figure 30] FIG. 30 illustrates an example of selecting words for editing based on a combination of voice and eye gaze input.
[0042] [Figure 31] FIG. 31 illustrates an example of selecting a word for editing based on a combination of gaze and gesture input.
[0043] [Figure 32] FIG. 32 illustrates an example of word replacement based on a combination of eye gaze and voice input.
[0044] [Figure 33] FIG. 33 illustrates an example of modifying words based on a combination of voice and eye gaze input.
[0045] [Figure 34] FIG. 34 illustrates an example of editing a selected word using a virtual keyboard.
[0046] [Figure 35] FIG. 35 illustrates an exemplary user interface displaying possible actions to apply to a selected word.
[0047] [Figure 36] FIG. 36 illustrates an example of interacting with phrases using multimodal input.
[0048] [Figure 37A] 37A and 37B illustrate additional examples of interacting with text using multimodal input. [Figure 37B] 37A and 37B illustrate additional examples of interacting with text using multimodal input.
[0049] [Figure 38] FIG. 38 is a process flow diagram of an exemplary method for interacting with text using multiple modes of user input.
[0050] [Figure 39A] FIG. 39A illustrates an example of user input received through a controller button.
[0051] [Figure 39B] FIG. 39B illustrates an example of user input received through a controller touchpad.
[0052] [Figure 39C] FIG. 39C illustrates an example of user input received through physical movement of a controller or head-mounted device (HMD).
[0053] [Figure 39D] FIG. 39D illustrates an example of how user inputs can have different durations.
[0054] [Figure 40A] FIG. 40A illustrates an additional example of user input received through a controller button.
[0055] [Figure 40B] FIG. 40B illustrates an additional example of user input received through a controller touchpad.
[0056] [Figure 41A] FIG. 41A illustrates examples of user input received through various user input modes for spatial manipulation of a virtual environment or virtual objects.
[0057] [Figure 41B] FIG. 41B illustrates examples of user input received through various user input modes for interacting with a planar object.
[0058] [Figure 41C] FIG. 41C illustrates examples of user input received through various user input modes for interacting with the wearable system.
[0059] [Figure 42A] 42A, 42B, and 42C illustrate examples of user input in the form of fine finger gestures and hand movements. [Figure 42B] 42A, 42B, and 42C illustrate examples of user input in the form of fine finger gestures and hand movements. [Figure 42C] 42A, 42B, and 42C illustrate examples of user input in the form of fine finger gestures and hand movements.
[0060] [Figure 43A] FIG. 43A illustrates an example of a user's sensory field of a wearable system, including a visual sensory field and an auditory sensory field.
[0061] [Figure 43B] FIG. 43B illustrates an example of a display rendering plane for a wearable system having multiple depth planes.
[0062] [Figure 44A] 44A, 44B, and 44C illustrate examples of different interaction areas whereby the wearable system may receive and respond to user input differently depending on which interaction area the user is interacting with. [Figure 44B] 44A, 44B, and 44C illustrate examples of different interaction areas whereby the wearable system may receive and respond to user input differently depending on which interaction area the user is interacting with. [Figure 44C] 44A, 44B, and 44C illustrate examples of different interaction areas whereby the wearable system may receive and respond to user input differently depending on which interaction area the user is interacting with.
[0063] [Figure 45] FIG. 45 illustrates an example of unimodal user interaction.
[0064] [Figure 46A] 46A, 46B, 46C, 46D, and 46E illustrate examples of multi-modal user interaction. [Figure 46B] 46A, 46B, 46C, 46D, and 46E illustrate examples of multi-modal user interaction. [Figure 46C] 46A, 46B, 46C, 46D, and 46E illustrate examples of multi-modal user interaction. [Figure 46D] 46A, 46B, 46C, 46D, and 46E illustrate examples of multi-modal user interaction. [Figure 46E] 46A, 46B, 46C, 46D, and 46E illustrate examples of multi-modal user interaction.
[0065] [Figure 47A] 47A, 47B, and 47C illustrate examples of cross-mode user interactions. [Figure 47B] 47A, 47B, and 47C illustrate examples of cross-mode user interactions. [Figure 47C] 47A, 47B, and 47C illustrate examples of cross-mode user interactions.
[0066] [Figure 48A] 48A, 48B, and 49 illustrate examples of transmode user interactions. [Figure 48B] 48A, 48B, and 49 illustrate examples of transmode user interactions. [Figure 49] 48A, 48B, and 49 illustrate examples of transmode user interactions.
[0067] [Figure 50] FIG. 50 is a process flow diagram of an embodiment of a method for detecting modal vergence.
[0068] [Figure 51] FIG. 51 illustrates an example of user selection in unimodal, bimodal, and trimodal interactions.
[0069] [Figure 52] FIG. 52 illustrates an example of interpreting user input based on the convergence of multiple user input modes.
[0070] [Figure 53] FIG. 53 illustrates an example of how different user inputs can converge across different interaction regions.
[0071] [Figure 54]54, 55, and 56 illustrate examples of how the system may select among multiple possible input convergence interpretations based, at least in part, on the ranking of different inputs. [Figure 55] 54, 55, and 56 illustrate examples of how the system may select among multiple possible input convergence interpretations based, at least in part, on the ranking of different inputs. [Figure 56] 54, 55, and 56 illustrate examples of how the system may select among multiple possible input convergence interpretations based, at least in part, on the ranking of different inputs.
[0072] [Figure 57A] 57A and 57B are block diagrams of an example wearable system that blends multiple user input modes to facilitate user interaction with the wearable system. [Figure 57B] 57A and 57B are block diagrams of an example wearable system that blends multiple user input modes to facilitate user interaction with the wearable system.
[0073] [Figure 58A] FIG. 58A is a graph of vergence distance for various input pairs and vergence area for user interactions with dynamic transmode input fusion disabled.
[0074] [Figure 58B] FIG. 58B is a graph of vergence distance for various input pairs and vergence area for user interactions with dynamic transmode input fusion enabled.
[0075] [Figure 59A] 59A and 59B illustrate examples of user interaction and feedback during fixation and dwell events. [Figure 59B]59A and 59B illustrate examples of user interaction and feedback during fixation and dwell events.
[0076] [Figure 60A] 60A and 60B illustrate an example of a wearable system that may include at least one neuromuscular sensor, such as, for example, an electromyography (EMG) sensor, and that may be configured to use embodiments of the transmodal input fusion techniques described herein. [Figure 60B] 60A and 60B illustrate an example of a wearable system that may include at least one neuromuscular sensor, such as, for example, an electromyography (EMG) sensor, and that may be configured to use embodiments of the transmodal input fusion techniques described herein. DETAILED DESCRIPTION OF THE INVENTION
[0077] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure.
[0078] (overview) Modern computing systems can possess a variety of user interactions. Wearable devices can present interactive VR / AR / MR environments, which can comprise data elements that can be interacted with by a user through various inputs. Modern computing systems are typically engineered to generate a given output based on a single direct input. For example, a keyboard will relay text input as received from a user's keystrokes. A speech recognition application can create an executable data string based on a user's voice as direct input. A computer mouse can guide a cursor in response to a user's direct manipulation (e.g., a user's hand movements or gestures). The various ways in which a user can interact with a system are sometimes referred to herein as modes of user input. For example, user input via a mouse or keyboard is a hand-gesture-based mode of interaction (as the fingers of a hand press keys on a keyboard or the hand moves a mouse).
[0079] However, in data-rich and dynamic interactive environments (e.g., AR / VR / MR environments), traditional input techniques such as keyboards, user input devices, and gestures may require a high degree of specificity to accomplish a desired task. Otherwise, in the absence of precise input, computing systems may suffer from high error rates and cause incorrect computer actions to be performed. For example, when a user intends to move an object in 3D space using a touchpad, the computing system may be unable to correctly interpret the movement command if the user does not specify a destination or specify an object using the touchpad. As another example, entering a string of text using a virtual keyboard as the sole input mode (e.g., with a user input device or as operated by gestures) can be slow and physically tiring, requiring extended periods of fine motor control to type keys depicted in the air or on a physical surface (e.g., a desk) on which the virtual keyboard is rendered.
[0080] To reduce the degree of specificity required in input commands and to reduce the error rate associated with imprecise commands, wearable systems as described herein can be programmed to dynamically apply multiple inputs to identify an object to be selected or acted upon, e.g., to perform an interaction event associated with the object, such as a task to select, move, resize, or target a virtual object. The interaction event can include causing an application (sometimes referred to as an app) associated with the virtual object to be executed (e.g., if the target object is a media file, the interaction event can include causing a media player to play the media file (e.g., a song or video)). Selecting a target virtual object can include executing the application associated with the target virtual object. As described below, the wearable device can dynamically select any of two or more types of input (or input from multiple input channels or sensors) to generate a command for the performance of a task or to identify a target object on which a command is to be executed.
[0081] The particular sensor inputs used at any one time may change dynamically as the user interacts with the 3D environment. Input modes can be dynamically added (or “fused,” as described further herein) when the device determines that the input mode provides additional information to aid in targeting a virtual object, and input modes can be dynamically removed when the input mode no longer provides relevant information. For example, a wearable device may determine that a user's head pose and eye gaze are directed toward a target object. The device can use these two input modes to assist in selecting the target object. If the device determines that the user is also pointing a totem toward the target object, the device can dynamically add a totem input to the head pose and eye gaze input, which may provide further certainty that the user intends to select the target object. The totem input is said to be “converged” with the head pose input and eye gaze input. Continuing with this example, if the user looks away from the target object such that the user's eye gaze is no longer directed at the target object, the device may cease using eye gaze input while continuing to use totem input and head pose input, in which case the eye gaze input is said to "diverge" from the totem input and head pose input.
[0082] The wearable device can dynamically determine the occurrence of divergence and convergence events among multiple input modes and dynamically select a subset of input modes from among these multiple input modes that is relevant to the user's interaction with the 3D environment. For example, the system can use convergent input modes and not consider divergent input modes. The number of input modes that can be dynamically fused or filtered in response to input convergence is not limited to the three modes described in this example (totem, head pose, eye gaze), but can dynamically switch between one, two, three, four, five, six, or more sensor inputs (as different input modes converge or diverge).
[0083] The wearable device can use convergent inputs by receiving inputs for analysis, by increasing the computing resources available or allocated to convergent inputs (e.g., to an input sensor assembly), by selecting a particular filter to apply to one or more of the convergent inputs, by taking other suitable action, and / or by any combination of these actions. The wearable device may not use, discontinue use of, or reduce the weighting given to divergent or diverging sensor inputs.
[0084] An input mode may be said to be converged, for example, when the variance between the input vectors of the inputs is below a threshold. Once the system recognizes that the inputs have converged, it may filter the converged inputs, fuse them together, and then create a new, adjusted input that can be used to do useful work and utilized to perform tasks with greater confidence and accuracy (than could be accomplished by using the inputs separately). In various embodiments, the system may apply dynamic filtering (e.g., dynamically fuse inputs together) in response to the relative convergence of the inputs. The system may continuously evaluate whether the inputs are convergent. In some embodiments, the system may scale the strength of input fusion (e.g., the strength with which the system fuses inputs together) in relation to the strength of input convergence (e.g., the closeness with which the input vectors of two or more inputs are matched).
[0085] The process of dynamically using convergent sensor inputs while ignoring (or reducing the relative weighting given to) divergent sensor inputs is sometimes referred to herein as transmodal input fusion (or simply transmodal fusion) and can provide substantial advantages over techniques that simply receive inputs from multiple sensors. Transmodal input fusion can anticipate or even predict, on a dynamic, real-time basis, which of many possible sensor inputs will be the appropriate modal input that communicates the user's intent to target or act on real or virtual objects within the user's 3D AR / MR / VR environment.
[0086] As will be further described herein, input modes can include, but are not limited to, hand or finger gestures, arm gestures, body gestures, head poses, eye gaze, body postures, voice commands, environmental input (e.g., the location of the user or an object in the user's environment), a shared posture from another user, etc. Sensors used to detect these input modes can include, for example, an outward-facing camera (e.g., to detect hand or body gestures), an inward-facing camera (e.g., to detect eye gaze), an inertial measurement unit (IMU, e.g., accelerometer, gravimeter, magnetometer), an electromagnetic tracking sensor system, a global positioning system (GPS) sensor, a radar or lidar sensor, etc. (see, e.g., the description of example sensors with reference to FIGS. 2A and 2B ).
[0087] As another example, when a user utters "move it there," the wearable system can use a combination of head pose, eye gaze, and hand gestures, as well as other environmental factors (e.g., the user's location or the location of objects around the user), in combination with the voice command, to determine the object to be moved (e.g., the object corresponding to "it") and the intended destination (e.g., "there") in response to an appropriate dynamic selection of these multiple inputs.
[0088] As will be further described herein, techniques for transmodal input are not simply an aggregation of multiple user input modes. Rather, wearable systems employing such transmodal techniques can advantageously support additional depth dimensions in 3D (compared to traditional 2D interactions) provided to the wearable system. The additional dimensions not only enable additional types of user interaction (e.g., rotation or translation along additional axes in a Cartesian coordinate system), but also require highly precise user input to provide correct results.
[0089] However, user input for interacting with virtual objects is not always accurate due to the user's limitations in motor control. While conventional input techniques can calibrate and adjust for imprecision in a user's motor control in 2D space, such imprecision is magnified in 3D space due to the additional dimension. However, conventional input methods, such as keyboard input, are not well suited to adjusting for such imprecision in 3D space. One advantage (among other advantages) offered by transmodal input techniques is that they adapt the input method for fluid and more precise interaction with objects in 3D space.
[0090] Thus, embodiments of the transmodal input technique dynamically monitor for converging input modes and can more accurately determine or predict that the user intends to use this set of converging input modes to interact with the target. Embodiments of the transmodal input technique dynamically monitor for diverging input modes (e.g., indicating that the input modes are no longer relevant to the potential target) and can discontinue use of these diverging input modes (or reduce the weighting given to the diverging input modes relative to the converging input modes). The group of converging sensor input modes typically changes both temporarily and persistently. For example, different sensor input modes dynamically converge and diverge as the user moves their hands, body, head, or eyes while providing user input on the totem or using voice commands. Thus, a potential advantage of the transmodal input technique is that only the correct set of sensor input modes is used at any particular time or for any particular target object in a 3D environment. In some embodiments, the system may assign a greater weighting (than would normally be assigned) to a given input based on physiological context. As an example, the system may determine that the user is attempting to grasp and move a virtual object. In response, the system may assign more weight to hand posture inputs and less weight to other inputs, such as eye gaze inputs. The system may also shift input weights in a suitable manner from time to time. As an example, the system may shift weights to hand posture inputs as the hand posture converges on the virtual object.
[0091] Additionally, advantageously, in some embodiments, the techniques described herein can reduce the hardware requirements and costs of wearable systems. For example, rather than employing a high-resolution eye-tracking camera alone (which can be expensive and complex to use), a wearable device may use voice commands or head poses to perform a task in conjunction with a low-resolution eye-tracking camera to determine the task (e.g., by determining that some or all of these input modes are converging on a target object). In this example, the user's use of voice commands can compensate for the lower resolution at which eye tracking is performed. Thus, transmodal combination of multiple user input modes, which enables dynamic selection of which of the multiple user input modes should be used, can provide lower cost, lower complexity than the use of a single input mode, and more robust user interaction with an AR / VR / MR device. Additional advantages and examples associated with transmodal sensor fusion techniques for interacting with real or virtual objects are described below with further reference to FIGS. 13-59B.
[0092] Transmode fusion techniques can provide substantial advantages over simple aggregation of multiple sensor inputs, e.g., with respect to functionality such as targeting small objects, targeting objects in a field of view containing many objects, targeting moving objects, managing transitions between close, medium, and long-range targeting methods, manipulating virtual objects, etc. In some implementations, transmode fusion techniques are referred to as providing a TAMDI interaction model for targeting (e.g., defining a cursor vector toward an object), activation (e.g., selecting a specific object or area or volume within a 3D environment), manipulation (e.g., directly moving or changing the selection), deactivation (e.g., deselecting), and integration (e.g., returning a previous selection to the environment when necessary).
[0093] (Example of a 3D display for a wearable system) A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images may be still images, frames of video, or videos, in combination or the like. A wearable system can include a wearable device that can present VR, AR, or MR content, alone or in combination, within an environment for user interaction. The wearable device can be a head-mounted device (HMD) that can include a head-mounted display. In some circumstances, a wearable device is referred to synonymously as an AR device (ARD).
[0094] Figure 1 depicts an illustration of a mixed reality scenario involving a virtual reality object and a physical object viewed by a person. In Figure 1, an MR scene 100 is depicted in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, a building in the background, and a concrete platform 120. In addition to these items, the user of the MR technology also perceives as "seeing" a robotic figure 130 standing on the real-world platform 120 and a flying, cartoon-like avatar character 140 that appears to be an anthropomorphic bumblebee, although these elements do not exist in the real world.
[0095] In order for a 3D display to produce a true depth sensation, and more specifically, a simulated sensation of surface depth, it may be desirable for the display to generate, for each point in its field of view, an accommodation response that corresponds to that point's virtual depth. If the accommodation response to a display point does not correspond to that point's virtual depth as determined by convergence and stereoscopic binocular depth cues, the human eye may experience accommodation conflict, resulting in unstable imaging, adverse eye strain, headaches, and, in the absence of accommodative information, a near-complete lack of surface depth.
[0096] VR, AR, and MR experiences can be provided by a display system having a display in which images corresponding to multiple rendering planes are provided to a viewer. The rendering planes can correspond to a depth plane or multiple depth planes. The images may be different for each rendering plane (e.g., providing slightly different presentations of a scene or object) and may be focused separately by the viewer's eyes, thereby serving to provide depth cues to the user based on the ocular accommodation required to focus on different image features of a scene located on different rendering planes, or based on observing different image features on different rendering planes that are out of focus. As discussed elsewhere herein, such depth cues provide a believable perception of depth.
[0097] FIG. 2A illustrates an example of a wearable system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functionality of the display 220. The display 220 may be coupled to a frame 230, which is wearable by a user, wearer, or viewer 210. The display 220 can be positioned directly in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can comprise a head-mounted display (HMD) worn on the user's head. In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the user's ear canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other ear canal to provide stereo / shapeable sound control). The display 220 can include an audio sensor 232 (e.g., a microphone) to detect audio streams from the environment for performing sound recognition.
[0098] The wearable system 200 may include an outward-facing imaging system 464 (shown in FIG. 4 ) that observes the world in the user's surrounding environment. The wearable system 200 may also include an inward-facing imaging system 462 (shown in FIG. 4 ) that can track the user's eye movements. The inward-facing imaging system may track the movements of either one eye or both eyes. The inward-facing imaging system 462 may be mounted to the frame 230 and may be in electrical communication with a processing module 260 or 270 that may process image information obtained by the inward-facing imaging system and determine, for example, pupil diameter or orientation of the user's 210 eyes, eye movement, or eye posture.
[0099] As an example, wearable system 200 can obtain images of a user's posture (e.g., gestures) using outward-facing imaging system 464 or inward-facing imaging system 462. The images may be still images, frames of video, videos, combinations thereof, or the like. Wearable system 200 can include other sensors, such as electromyography (EMG) sensors, that sense signals indicative of muscle group actions (e.g., see discussion with reference to FIGS. 60A and 60B).
[0100] The display 220 can be operably coupled (250) to a local data processing module 260, which can be mounted in a variety of configurations, such as fixedly attached to the frame 230, by wired or wireless connection, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 210 (e.g., in a backpack configuration, in a belt-coupled configuration).
[0101] The local processing and data module 260 may comprise a hardware processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to aid in processing, caching, and storing data. The data may include a) data captured from environmental sensors (e.g., which may be operably coupled to the frame 230 or otherwise attached to the user 210), audio sensors 232 (e.g., microphones), or b) data obtained or processed using the remote processing module 270 or remote data repository 280, possibly for processing or reading and then passing to the display 220. The local processing and data module 260 may be operably coupled to the remote processing module 270 or remote data repository 280 by communication links 262 or 264, such as via wired or wireless communication links, so that these remote modules are available as resources to the local processing and data module 260. Additionally, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.
[0102] In some embodiments, remote processing module 270 may comprise one or more processors configured to analyze and process data and / or image information. In some embodiments, remote data repository 280 may comprise a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, allowing for fully autonomous use from the remote module.
[0103] 2A or 2B (described below), wearable system 200 may include environmental sensors to detect objects, stimuli, people, animals, places, or other aspects of the world around the user. The environmental sensors may include an image capture device (e.g., a camera, an inward-facing imaging system, an outward-facing imaging system, etc.), a microphone, an inertial measurement unit (IMU) (e.g., an accelerometer, a gyroscope, a magnetometer (compass)), a global positioning system (GPS) unit, a wireless device, an altimeter, a barometer, a chemical sensor, a humidity sensor, a temperature sensor, an external microphone, a light sensor (e.g., a light meter), a timing device (e.g., a clock or calendar), or any combination or subcombination thereof. In one embodiment, the IMU may be a 9-axis IMU, which may include a triple-axis gyroscope, a triple-axis accelerometer, and a triple-axis magnetometer.
[0104] The environmental sensors may also include various physiological sensors. These sensors may measure or estimate the user's physiological parameters, such as heart rate, respiratory rate, galvanic skin response, blood pressure, brainwave state, etc. The environmental sensors may further include emitting devices configured to receive signals, such as lasers, visible light, light of invisible wavelengths, or sound (e.g., audible, ultrasonic, or other frequencies). In some embodiments, one or more environmental sensors (e.g., cameras or light sensors) may be configured to measure the ambient light (e.g., luminance) of the environment (e.g., to capture the lighting conditions of the environment). Physical contact sensors, such as strain gauges, curb detectors, or the like, may also be included as environmental sensors.
[0105] FIG. 2B illustrates another example of a wearable system 200 including a number of example sensors. Input from any of these sensors can be used by the system in the transmodal sensor fusion techniques described herein. The head-mounted wearable component 200 is shown here operably coupled to a local processing and data module (70), such as a belt pack, using a physical multicore conductor that also features a control and quick-release module (86) as connecting the belt pack to a head-mounted display (68). The head-mounted wearable component 200 is also referred to in FIG. 2B and below using the reference numeral 58. The local processing and data module (70) is here operably coupled to a handheld component (606) (100) by a wireless connection, such as low-power Bluetooth®. The handheld component (606) may also be operably coupled directly to the head-mounted wearable component (58) (94), such as by a wireless connection, such as low-power Bluetooth®. Generally, when IMU data is passed to coordinate attitude detection of various components, a high frequency connection, such as in the hundreds or thousands of cycles per second or higher range, is desirable. Tens of cycles per second may be adequate for electromagnetic localization sensing, such as by pairing a sensor 604 and a transmitter 602. Also shown is a global coordinate system 10, which represents fixed objects in the real world around the user, such as a wall 8.
[0106] Cloud resources (46) may also be operatively coupled to resources (42, 40, 88, 90), which may each be coupled to a local processing and data module (70), a head-mounted wearable component (58), a wall (8), or other items fixed relative to the global coordinate system (10). Resources coupled to the wall (8) or having a known position and / or orientation relative to the global coordinate system (10) may include a wireless transceiver (114), an electromagnetic emitter (602) and / or receiver (604), a beacon or reflector (112) configured to emit or reflect a given type of radiation, such as an infrared LED beacon, a cellular network transceiver (110), a radar emitter or detector (108), a lidar emitter or detector (106), a GPS transceiver (118), a poster or marker (122) having a known detectable pattern, and a camera (124).
[0107] System 200 can include a depth camera or depth sensor (154), which can be, for example, either a stereo triangulation depth sensor (such as a passive stereo depth sensor, a texture projection stereo depth sensor, or a structured light stereo depth sensor) or a time-of-flight depth sensor (such as a lidar depth sensor or a modulated emission depth sensor). System 200 can include a forward-facing "world" camera (124, which can be, for example, a grayscale camera with a sensor capable of 720p range resolution) and a relatively high-resolution "photo camera" (156, which can be, for example, a full-color camera with a sensor capable of 2 megapixel or higher resolution).
[0108] The head-mounted wearable component (58) features an illumination emitter (130) as well as similar components configured to assist the camera (124) detector, such as an infrared emitter (130) for the infrared camera (124), as shown. The head-mounted wearable component (58) also features one or more strain gauges (116) thereon, which may be fixedly coupled to the frame or mechanical platform of the head-mounted wearable component (58) and configured to determine deflection of such platform between components such as an electromagnetic receiver sensor (604) or a display element (220), which may be useful to understand when flexion of the platform occurs, such as in a thinned portion of the platform, such as the portion above the nose on the eyeglass-like platform depicted in FIG. 2B .
[0109] The head-mounted wearable component (58) also features a processor (128) and one or more IMUs (102). Each component is preferably operably coupled to the processor (128). The handheld component (606) and the local processing and data module (70) are shown to feature similar components. Using numerous sensing and connectivity mechanisms, as shown in FIG. 2B, such a system can be utilized to provide a very high level of connectivity, system component integration, and position / orientation tracking. For example, using such a configuration, the various primary mobile components (58, 70, 606) can be located in terms of their position relative to a global coordinate system using WiFi, GPS, or cellular signal triangulation. Beacons, electromagnetic tracking, radar, and lidar systems can provide even further location and / or orientation information and feedback. Markers and cameras may also be utilized to provide additional information regarding relative and absolute position and orientation. For example, various camera components (124), such as those shown coupled to the head-mounted wearable component (58), may be utilized to capture data that may be utilized to determine the location of the component (58) and how it is oriented relative to other components in a simultaneous localization and mapping protocol or "SLAM."
[0110] The description with reference to FIGS. 2A and 2B describes an illustrative and non-limiting list of types of sensors and input modes that can be used with wearable system 200. However, not all of these sensors or input modes need to be used in all embodiments. Furthermore, additional or alternative sensors can be used as well. The selection of sensors and input modes for a particular embodiment of wearable system 200 can be based on factors such as cost, weight, size, complexity, etc. Many permutations and combinations of sensors and input modes are contemplated. Wearable systems including sensors such as those described with reference to FIGS. 2A and 2B can advantageously utilize the transmodal input fusion techniques described herein to dynamically select a subset of these sensor inputs to assist a user in selecting, targeting, or interacting with a real or virtual object. The subset of sensor inputs (typically less than the set of all possible sensor inputs) can include sensor inputs that converge on a target object and can exclude (or reduce reliance on) sensor inputs that diverge from the subset or do not converge on the target object.
[0111] The human visual system is complex, making it difficult to provide a realistic perception of depth. Without being limited by theory, it is believed that a viewer of an object may perceive the object as three-dimensional due to a combination of vergence and accommodation. Vergence movement of the two eyes relative to one another (e.g., pupil rotation such that the pupils move toward or away from one another, converging the lines of sight of the eyes and fixating on the object) is closely linked to the focusing (or "accommodation") of the eye's lenses. Under normal conditions, a change in the focus of the eye's lenses or accommodation of the eye to change focus from one object to another at a different distance will automatically produce a matching change in vergence at the same distance, a relationship known as the "accommodation-vergence reflex." Similarly, a change in vergence will induce a matching change in accommodation under normal conditions. A display system that provides better matching between accommodation and vergence may produce a more realistic and comfortable simulation of three-dimensional images.
[0112] FIG. 3 illustrates aspects of an approach for simulating a three-dimensional image using multiple rendering planes. With reference to FIG. 3 , objects at various distances from the eyes 302 and 304 on the z-axis are accommodated by the eyes 302 and 304 so that the objects are in focus. The eyes 302 and 304 assume particular accommodated states, focusing objects at different distances along the z-axis. As a result, a particular accommodated state may be said to be associated with a particular one of the rendering planes 306, having an associated focal length, such that an object or portion of an object in the particular rendering plane is in focus when the eye is in an accommodated state relative to that rendering plane. In some embodiments, a three-dimensional image may be simulated by providing a different representation of an image for each of the eyes 302 and 304, and by providing a different representation of an image corresponding to each of the rendering planes. While shown as separate for clarity of illustration, it should be understood that the fields of view of the eyes 302 and 304 may overlap, for example, as the distance along the z-axis increases. Additionally, while shown as flat for ease of illustration, it should be understood that the contours of the rendering plane may be curved in physical space such that all features within the rendering plane are in focus with the eye in a particular state of accommodation. Without being limited by theory, it is believed that the human eye is typically capable of interpreting a finite number of rendering planes to provide depth perception. As a result, a highly realistic simulation of perceived depth may be achieved by providing the eye with different presentations of images corresponding to each of these limited number of rendering planes.
[0113] (Waveguide stack assembly) FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. Wearable system 400 includes a stack of waveguides or stacked waveguide assembly 480 that can be utilized to provide three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 4400b. In some embodiments, wearable system 400 may correspond to wearable system 200 of FIG. 2A or 2B, and FIG. 4 schematically illustrates some portions of wearable system 200 in more detail. For example, in some embodiments, waveguide assembly 480 may be integrated into display 220 of FIG. 2A or 2B.
[0114] 4, the waveguide assembly 480 may also include multiple features 458, 456, 454, 452 between the waveguides. In some embodiments, the features 458, 456, 454, 452 may be lenses. In other embodiments, the features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers or structures to form air gaps).
[0115] Waveguides 432b, 434b, 436b, 438b, 440b or multiple lenses 458, 456, 454, 452 may be configured to transmit image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular rendering plane and configured to output image information corresponding to that rendering plane. Image injection devices 420, 422, 424, 426, 428 may be utilized to inject image information into waveguides 440b, 438b, 436b, 434b, 432b, respectively, which may be configured to disperse incident light across each individual waveguide for output toward the eye 410. Light exits the output surfaces of image injection devices 420, 422, 424, 426, 428 and is injected into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) may be injected into each waveguide, outputting an entire field of cloned collimated beams directed toward eye 410 at a particular angle (and divergence) corresponding to the rendering plane associated with the particular waveguide.
[0116] In some embodiments, image input devices 420, 422, 424, 426, 428 are discrete displays that generate image information for input into each corresponding waveguide 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, image input devices 420, 422, 424, 426, 428 are outputs of a single multiplexed display that may, for example, send image information to each of image input devices 420, 422, 424, 426, 428 via one or more optical conduits (such as fiber optic cables).
[0117] A controller 460 controls the operation of stacked waveguide assembly 480 and image injection devices 420, 422, 424, 426, 428. Controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that coordinates the timing and provision of image information to waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, controller 460 may be a single integrated device or a distributed system connected by a wired or wireless communication channel. Controller 460 may, in some embodiments, be part of processing module 260 or 270 (shown in FIGS. 2A, 2B).
[0118] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, 432b may each be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between the major top and bottom surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, 432b may each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguides by redirecting the light to propagate within each individual waveguide and outputting image information from the waveguides to the eye 410. The extracted light may also be referred to as out-coupled light, and the light extraction optical element may also be referred to as out-coupling optical element. The extracted light beam is output by the waveguide where the light propagating within the waveguide strikes the light redirecting element. The light extraction optical element (440a, 438a, 436a, 434a, 432a) may be, for example, a reflective or diffractive optical feature. While shown disposed on the bottom major surfaces of the waveguides 440b, 438b, 436b, 434b, 432b for ease of explanation and clarity of the drawings, in some embodiments, the light extraction optical element 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surfaces, or directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed in a layer of material that is attached to a transparent substrate and forms the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within that piece of material.
[0119] Continuing with reference to FIG. 4, as discussed herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular rendering plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is launched into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next upper waveguide 434b may be configured to send collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to create a slight convex wavefront curvature so that the eye / brain interprets light emerging from the next upper waveguide 434b as emerging from a first focal plane closer inward from optical infinity toward the eye 410. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to produce another, increasing amount of wavefront curvature such that the eye / brain interprets the light emerging from the third waveguide 436b as originating from a second focal plane that is even closer inward from optical infinity towards the person than was the light from the next upper waveguide 434b.
[0120] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all of the lenses between it and the eye for a collective focal power representing the focal plane closest to the person. To compensate for the stack of lenses 458, 456, 454, 452 when viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be placed on top of the stack to compensate for the collective power of the lower lens stacks 458, 456, 454, 452. Such a configuration provides as many perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electro-active). In some alternative embodiments, one or both may be dynamic using electro-active features.
[0121] Continuing with reference to FIG. 4, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to both redirect light from its respective waveguide and output this light with the appropriate amount of divergence or collimation for the particular rendering plane associated with the waveguide. As a result, waveguides with different associated rendering planes may have differently configured light extraction optical elements that output light with different amounts of divergence depending on the associated rendering plane. In some embodiments, as discussed herein, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume or surface features that can be configured to output light at specific angles. For example, light extraction optical elements 440a, 438a, 436a, 434a, 432a may be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated herein by reference in its entirety.
[0122] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features, i.e., "diffractive optical elements" (also referred to herein as "DOEs"), that form a diffraction pattern. Preferably, the DOEs have a relatively low diffraction efficiency so that only a portion of the light in the beam is deflected toward the eye 410 at each intersection of the DOE, while the remainder continues traveling through the waveguide via total internal reflection. The light carrying the image information is thus split into several related output beams that exit the waveguide at multiple locations, which can result in a very uniform pattern of output emission toward the eye 304 for this particular collimated beam bouncing within the waveguide.
[0123] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal in which microdroplets comprise a diffractive pattern in a host medium, and the refractive index of the microdroplets can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets can be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).
[0124] In some embodiments, the number and distribution of rendering planes or depth of field may be dynamically varied based on the pupil size or orientation of the viewer's eyes. The depth of field may vary inversely with the viewer's pupil size. As a result, as the size of the viewer's eye pupil decreases, the depth of field increases so that one plane that is indistinguishable because its location exceeds the eye's depth of focus may become distinguishable and appear more focused with a reduction in pupil size and a corresponding increase in depth of field. Similarly, the number of spaced rendering planes used to present different images to the viewer may be reduced with a reduced pupil size. For example, a viewer may not be able to clearly perceive details in both a first rendering plane and a second rendering plane at one pupil size without adjusting their eye's accommodation from one rendering plane to the other. However, these two rendering planes may be sufficient to simultaneously focus on the user at another pupil size without changing accommodation.
[0125] In some embodiments, the display system may vary the number of waveguides receiving image information based on a determination of pupil size or orientation or in response to receiving an electrical signal indicating a particular pupil size or orientation. For example, if a user's eye is unable to distinguish between two rendering planes associated with two waveguides, the controller 460 may be configured or programmed to stop providing image information to one of those waveguides. Advantageously, this may reduce the processing burden on the system, thereby increasing system responsiveness. In embodiments in which the DOE for a waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.
[0126] In some embodiments, it may be desirable to have the output beam satisfy the condition of having a diameter less than the diameter of the viewer's eye. However, meeting this condition may be difficult in light of the variability in the size of the viewer's pupil. In some embodiments, this condition is met over a wide range of pupil sizes by varying the size of the output beam in response to a determination of the size of the viewer's pupil. For example, as the pupil size decreases, the size of the output beam may also decrease. In some embodiments, the output beam size may be varied using a variable aperture.
[0127] The wearable system 400 may include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 may be referred to as the world camera's field of view (FOV), and the imaging system 464 is sometimes referred to as an FOV camera. The entire area available for viewing or imaging by a viewer may be referred to as the ocular field of view (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400, since the wearer can move their body, head, or eyes to perceive virtually any direction in space. In other contexts, the wearer's movement may be more constrained, and accordingly, the wearer's FOR may subtend a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures (e.g., hand or finger gestures) made by the user, detect objects in the world 470 in front of the user, etc.
[0128] The wearable system 400 may also include an inward-facing imaging system 462 (e.g., a digital camera) that observes user movements, such as eye and facial movements. The inward-facing imaging system 462 may be used to capture images of the eyes 410 and determine the size and / or orientation of the pupils of the eyes 304. The inward-facing imaging system 462 may be used to obtain images for use in determining the direction the user is looking (e.g., eye pose) or for biometric identification of the user (e.g., via iris identification). In some embodiments, at least one camera may be utilized for each eye independently to separately determine the pupil size or eye pose of each eye, thereby allowing the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, the pupil diameter or orientation of only a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. Images obtained by inward-facing imaging system 462 may be analyzed to determine the user's eye posture or mood, which may be used by wearable system 400 to determine audio or visual content to be presented to the user. Wearable system 400 may also determine head pose (e.g., head position or head orientation) using sensors such as an IMU, accelerometer, gyroscope, etc.
[0129] The wearable system 400 may include a user input device 466 through which a user may input commands into the controller 460 and interact with the wearable system 400. For example, the user input device 466 may include a trackpad, touchscreen, joystick, multi-degree-of-freedom (DOF) controller, capacitive sensing device, game controller, keyboard, mouse, directional pad (D-pad), wand, tactile device, totem (e.g., functioning as a virtual user input device), etc. A multi-DOF controller may sense user input in possible translation (e.g., left / right, forward / backward, or up / down) or rotation (e.g., yaw, pitch, or roll) of some or all of the controller. A multi-DOF controller that supports translation may be referred to as 3DOF, while a multi-DOF controller that supports translation and rotation may be referred to as 6DOF. In some cases, a user may use a finger (e.g., a thumb) to press or swipe across a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 may communicate with the wearable system 400 via wired or wireless communication.
[0130] FIG. 5 shows an example of an output beam output by a waveguide. While one waveguide is shown, it should be understood that other waveguides in waveguide assembly 480 may function similarly, and that waveguide assembly 480 includes multiple waveguides. Light 520 is launched into waveguide 432b at input edge 432c of waveguide 432b and propagates within waveguide 432b by TIR. At the point where light 520 impinges on DOE 432a, a portion of the light exits the waveguide as output beam 510. Although output beams 510 are shown as approximately parallel, they may also be redirected to propagate to eye 410 at an angle (e.g., forming a diverging output beam) depending on the rendering plane associated with waveguide 432b. It should be understood that a nearly collimated exit beam may refer to a waveguide with light extraction optics that outcouples light to form an image that appears to be set at a rendering plane at a large distance (e.g., optical infinity) from the eye 410. Other waveguides or other sets of light extraction optics may output a more divergent exit beam pattern, which would require the eye 410 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 410 than optical infinity.
[0131] FIG. 6 is a schematic diagram illustrating an optical system including a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem used in generating a multifocal volumetric display, image, or light field. The optical system can include a waveguide device, an optical coupler subsystem for optically coupling light to or from the waveguide device, and a control subsystem. The optical system can be used to generate a multifocal volumetric display, image, or light field. The optical system can include one or more primary planar waveguides 632a (only one is shown in FIG. 6) and one or more DOEs 632b associated with each of at least some of the primary waveguides 632a. The planar waveguides 632b can be similar to the waveguides 432b, 434b, 436b, 438b, and 440b discussed with reference to FIG. 4. The optical system may employ a dispersive waveguide device to relay light along a first axis (the vertical or Y-axis in the illustration of FIG. 6 ) and expand the effective exit pupil of the light along the first axis (e.g., the Y-axis). The dispersive waveguide device may include, for example, a dispersive planar waveguide 622 b and at least one DOE 622 a (illustrated by a double-dashed line) associated with the dispersive planar waveguide 622 b. The dispersive planar waveguide 622 b may be similar or identical in at least some respects to a primary planar waveguide 632 b having a different orientation therefrom. Similarly, the at least one DOE 622 a may be similar or identical in at least some respects to the DOE 632 a. For example, the dispersive planar waveguide 622 b or the DOE 622 a may be made of the same material as the primary planar waveguide 632 b or the DOE 632 a, respectively. The embodiment of the optical display system 600 shown in FIG. 6 can be integrated into the wearable system 200 shown in FIG. 2A or 2B.
[0132] The relayed, exit-pupil-expanded light can be optically coupled from the dispersive waveguide device into one or more primary planar waveguides 632b. The primary planar waveguides 632b can relay the light along a second axis, preferably orthogonal to the first axis (e.g., the horizontal or X-axis in the diagram of FIG. 6). Notably, the second axis can be non-orthogonal to the first axis. The primary planar waveguides 632b expand the effective exit pupil of the light along that second axis (e.g., the X-axis). For example, the dispersive planar waveguide 622b can relay and expand the light along the vertical or Y-axis and pass the light to a primary planar waveguide 632b, which can relay and expand the light along the horizontal or X-axis.
[0133] The optical system may include one or more colored light sources (e.g., red, green, and blue laser light) 610, which may be optically coupled into the proximal end of a single-mode optical fiber 640. The distal end of the optical fiber 640 may be threaded or received through a hollow tube 642 of piezoelectric material. The distal end protrudes from the tube 642 as a free-standing, flexible cantilever 644. The piezoelectric tube 642 may be associated with four quadrant electrodes (not shown). The electrodes may be plated, for example, on the outside, outer surface or outer periphery, or diameter of the tube 642. A core electrode (not shown) may also be located in the core, center, inner periphery, or inner diameter of the tube 642.
[0134] For example, drive electronics 650, electrically coupled via wires 660, drive opposing pairs of electrodes to bend piezoelectric tube 642 independently in two axes. The protruding distal tip of optical fiber 644 has a mechanical resonant mode. The frequency of the resonance may depend on the diameter, length, and material properties of optical fiber 644. By oscillating piezoelectric tube 642 near the first mechanical resonant mode of fiber cantilever 644, fiber cantilever 644 may be caused to oscillate and sweep through a large deflection.
[0135] By stimulating resonant vibrations in two axes, the tip of fiber cantilever 644 is scanned biaxially within an area filling a two-dimensional (2-D) scan. By modulating the intensity of light source 610 synchronously with the scanning of fiber cantilever 644, light emitted from fiber cantilever 644 can form an image. A description of such a setup is provided in U.S. Patent Publication No. 2014 / 0003762, which is incorporated herein by reference in its entirety.
[0136] Components of the optical coupler subsystem can collimate light emitted from the scanning fiber cantilever 644. The collimated light can be reflected by a mirrored surface 648 into a narrow dispersive planar waveguide 622b containing at least one diffractive optical element (DOE) 622a. The collimated light can propagate perpendicularly (with respect to the view of FIG. 6) along the dispersive planar waveguide 622b via TIR, and in doing so, repeatedly intersect with the DOE 622a. The DOE 622a preferably has a low diffraction efficiency. This diffracts a portion of the light (e.g., 10%) toward the edge of the larger primary planar waveguide 632b at each point of intersection with the DOE 622a, allowing a portion of the light to continue on its original trajectory down the length of the dispersive planar waveguide 622b via TIR.
[0137] At each point of intersection with DOE 622a, additional light can be diffracted toward the entrance of primary waveguide 632b. By splitting the incident light into multiple outcoupled sets, the exit pupil of the light can be vertically expanded by DOE 4 within dispersive planar waveguide 622b. This vertically expanded light outcoupled from dispersive planar waveguide 622b can enter the edge of primary planar waveguide 632b.
[0138] Light entering the primary waveguide 632b can propagate horizontally (with respect to the diagram of FIG. 6) along the primary waveguide 632b via TIR. The light propagates horizontally along at least a portion of the length of the primary waveguide 632b via TIR as it intersects the DOE 632a at multiple points. The DOE 632a advantageously has a phase profile that is the sum of a linear diffraction pattern and a radially symmetric diffraction pattern, and may be designed or configured to produce both deflection and focusing of the light. The DOE 632a advantageously may have a low diffraction efficiency (e.g., 10%) so that only a portion of the light in the beam is deflected toward the viewer's eye at each intersection of the DOE 632a, while the remainder of the light continues to propagate through the primary waveguide 632b via TIR.
[0139] At each point of intersection between the propagating light and the DOE 632a, a portion of the light is diffracted toward the adjacent face of the primary waveguide 632b, allowing the light to escape the TIR and emerge from the face of the primary waveguide 632b. In some embodiments, the radially symmetric diffraction pattern of the DOE 632a additionally imparts a focal level to the diffracted light, both shaping (e.g., imparting curvature to) the optical wavefronts of the individual beams and steering the beams to angles that match the designed focal level.
[0140] These different paths can then couple light out of the primary planar waveguide 632b by providing different fill patterns at the DOE 632a's multiplicity, focal level, and / or exit pupil at different angles. Different fill patterns at the exit pupil can be advantageously used to generate light field displays with multiple rendering planes. Each layer in the waveguide assembly or set of layers (e.g., three layers) in the stack may be employed to generate distinct colors (e.g., red, blue, and green). Thus, for example, a first set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a first focal depth. A second set of three adjacent layers may be employed to generate red, blue, and green light, respectively, at a second focal depth. Multiple sets may be employed to generate full 3D or 4D color image light fields with various focal depths.
[0141] (Other components of the wearable system) In many implementations, the wearable system may include other components in addition to or as an alternative to the components of the wearable system described above. The wearable system may include, for example, one or more tactile devices or components. The tactile device or component may be operable to provide a haptic sensation to the user. For example, the tactile device or component may provide a sensation of pressure or texture upon touching virtual content (e.g., a virtual object, virtual tool, other virtual structure). The haptic sensation may replicate the sensation of a physical object represented by the virtual object, or may replicate the sensation of an imaginary object or character (e.g., a dragon) represented by the virtual content. In some implementations, the tactile device or component may be worn by the user (e.g., a user-wearable glove). In some implementations, the tactile device or component may be held by the user.
[0142] A wearable system may include, for example, one or more physical objects that can be manipulated by a user to enable input to or interaction with the wearable system. These physical objects may be referred to herein as totems. Some totems may take the form of inanimate objects, such as, for example, a piece of metal or plastic, a wall, a table surface, etc. In some implementations, a totem may not actually have any physical input structures (e.g., keys, triggers, joysticks, trackballs, rocker switches). Instead, the totem may simply provide a physical surface, and the wearable system may render a user interface to appear to the user on one or more surfaces of the totem. For example, the wearable system may render an image of a computer keyboard and trackpad to appear to reside on one or more surfaces of the totem. For example, the wearable system may render a virtual computer keyboard and virtual trackpad to appear on the surface of a thin rectangular plate of aluminum that serves as the totem. The rectangular plate itself does not have any physical keys or trackpads or sensors. However, the wearable system may detect user manipulation or interaction or touch with the rectangular plate as a selection or input made via a virtual keyboard or virtual trackpad. User input device 466 (shown in FIG. 4) may be an embodiment of a totem, which may include a trackpad, touchpad, trigger, joystick, trackball, rocker or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or another physical input device. A user may use the totem alone or in combination with posture to interact with the wearable system or other users.
[0143] Examples of tactile devices and totems usable with the wearable devices, HMDs, and display systems of the present disclosure are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety.
[0144] Exemplary Wearable Systems, Environments, and Interfaces The wearable system may employ various mapping-related techniques to achieve a high depth of field within the rendered light field. When mapping a virtual world, it is advantageous to capture all features and points in the real world and accurately depict virtual objects in relation to the real world. To achieve this goal, FOV images captured from a user of the wearable system can be added to the world model by including new photos that convey information about various points and features in the real world. For example, the wearable system can collect a set of map points (such as 2D or 3D points), find new map points, and render a more accurate version of the world model. The world model of a first user can be communicated to a second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.
[0145] 7 is a block diagram of an example MR environment 700. The MR environment 700 may be configured to receive inputs (e.g., visual input 702 from a user's wearable system, stationary input 704 such as a room camera, sensory input 706 from various sensors, gestures, totems, eye tracking, user input, etc. from user input device 466) from one or more user-wearable systems (e.g., wearable system 200 or display system 220) or stationary room systems (e.g., room cameras, etc.). The wearable systems can determine the location and various other attributes of the user's environment using various sensors (e.g., accelerometers, gyroscopes, temperature sensors, movement sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.). This information may be further supplemented with information from stationary cameras in the room, which may provide images from different perspectives or various cues. Image data acquired by the cameras (e.g., the room cameras and / or the outward-facing imaging system cameras) may be reduced to a set of mapping points.
[0146] One or more object recognizers 708 can crawl through the received data (e.g., a collection of points), recognize or map the points, tag the images, and associate semantic information with the objects using a map database 710. The map database 710 may comprise various points and their corresponding objects collected over time. The various devices and the map database may be interconnected through a network (e.g., a LAN, a WAN, etc.) and accessible to the cloud.
[0147] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects in the environment. For example, the object recognizers may recognize faces, people, windows, walls, user input devices, televisions, other objects in the user's environment, etc. One or more object recognizers may be specialized for objects with certain characteristics. For example, object recognizer 708a may be used to recognize faces, while another object recognizer may be used to recognize totems, while another object recognizer may be used to recognize hands, fingers, arms, or body gestures.
[0148] Object recognition may be performed using various computer vision techniques. For example, the wearable system may analyze images acquired by the outward-facing imaging system 464 (shown in FIG. 4) and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), etc.
[0149] Object recognition can additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms can be stored by the HMD. Some examples of machine learning algorithms can include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models can be customized for individual datasets. For example, the wearable device can generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user in a telepresence session), a data set (e.g., a set of additional images acquired of a user in a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD can be configured to generate models for analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.
[0150] Based on this information and the set of points in the map database, the object recognizers 708a-708n may recognize objects, complement them with semantic information, and bring them to life. For example, if the object recognizer recognizes that a set of points is a door, the system may associate some semantic information (e.g., a door has a hinge and 90-degree movement around the hinge). If the object recognizer recognizes that a set of points is a mirror, the system may associate the semantic information that a mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database grows as the system (which may reside locally or be accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR environment 700 may contain information about a scene generating in California. The environment 700 may be transmitted to one or more users in New York. Based on the data received from the FOV camera and other inputs, the object recognizer and other software components can map points collected from various images, recognize objects, etc. so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment 700 may also use a topology map for localization purposes.
[0151] 8 is a process flow diagram of an example method 800 for rendering virtual content in relation to recognized objects. Method 800 describes how a virtual scene can be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, a user may be in New York but may want to view a scene currently occurring in California, or may want to go for a walk with a friend who is in California.
[0152] In block 810, the wearable system may receive input from the user and other users regarding the user's environment. This may be accomplished through various input devices and knowledge already held in a map database. The user's FOV camera, sensors, GPS, eye tracking, etc., communicate information to the system in block 810. The system may determine sparse points based on this information in block 820. The sparse points may be used to determine pose data (e.g., head pose, eye pose, body pose, or hand gestures) that may be used in displaying and understanding the orientation and position of various objects in the user's surroundings. The object recognizers 708a-708n may crawl through these collected points and recognize one or more objects using the map database in block 830. This information may then be communicated to the user's respective wearable system in block 840, and the desired virtual scene may be displayed to the user appropriately in block 850. For example, a desired virtual scene (eg, a user in CA) may be displayed in the proper orientation, position, etc., relative to various objects and other surroundings of the user in New York.
[0153] FIG. 9 is a block diagram of another example of a wearable system. In this example, the wearable system 900 includes a map, which may include map data about the world. The map may reside partially locally on the wearable system and partially in a networked storage location (e.g., in a cloud system) accessible by a wired or wireless network. A posture process 910 (e.g., head or eye posture) may run on the wearable computing architecture (e.g., processing module 260 or controller 460) and utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The posture data may be calculated from data collected on the fly as the user experiences the system and moves within its world. The data may include images of objects in the real or virtual environment, data from sensors (such as inertial measurement units, which generally include accelerometer and gyroscope components), and surface information.
[0154] The sparse point representation may be the output of a simultaneous localization and mapping (SLAM or V-SLAM, which refers to configurations where the input is only image / vision) process. The system can be configured to find not only the location of various components in the world, but also what the world is made of. Poses can be building blocks that accomplish many goals, including capturing in and using data from maps.
[0155] In one embodiment, the sparse point locations may not be entirely correct by themselves, and additional information may be required to generate a multifocal AR, VR, or MR experience. A dense representation, generally referring to depth map information, may be utilized to fill in this gap, at least in part. Such information may be calculated from a process referred to as stereoscopic vision 940, where depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and active patterns (such as infrared patterns generated using an active projector) may serve as inputs to the stereoscopic vision process 940. A significant amount of depth map information may be fused together, some of which may be summarized using a surface representation. For example, mathematically definable surfaces may be an efficient (e.g., for large point clouds) and easy-to-summarize input to other processing devices, such as a game engine. Thus, the outputs of the stereoscopic vision process (e.g., depth map) 940 may be combined in a fusion process 930. Pose may likewise be an input to this fusion process 930, the output of which becomes an input to the map capture process 920. Sub-surfaces can interconnect to form larger surfaces, such as in topographic mapping, where the map becomes a large-scale hybrid of points and surfaces.
[0156] Various inputs may be utilized to resolve various aspects of the mixed reality process 960. For example, in the embodiment depicted in Figure 9, game parameters may be inputs for determining that a user of the system is playing a monster battle game with one or more monsters at various locations, whether monsters are dead or fleeing under various conditions (such as when a user shoots a monster), walls or other objects at various locations, and the like. A world map may contain information about the location of such objects relative to one another, which is another useful input for mixed reality. Attitude relative to the world is likewise an input and plays an important role for nearly any interactive system.
[0157] Controls or inputs from the user are another input to the wearable system 900. As described herein, user inputs can include visual inputs, gestures, totems, audio inputs, sensory inputs, etc. To move around or play games, for example, the user may need to command the wearable system 900 with respect to a desired object. There are various forms of user control that can be utilized beyond just moving around in space. In one embodiment, a totem (e.g., a user input device) or an object such as a toy gun may be held by the user and tracked by the system. The system would preferably be configured to know that the user is holding an item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be configured to understand not only the location and orientation, but also whether the user is clicking a trigger or other sensitive button or element, which may be equipped with sensors such as an IMU, which can help determine the situation occurring even when such activity is not within the field of view of any of the cameras).
[0158] Hand gesture tracking or recognition may also provide input information. The wearable system 900 may be configured to track and interpret hand gestures to gesture for button presses, left or right, stop, grasp, hold, etc. For example, in one configuration, a user may wish to flip through email or calendar in a non-gaming environment or perform a “fist bump” with another person or player. The wearable system 900 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, gestures may be simple static gestures, such as extending the hand to indicate stop, thumbs up to indicate OK, thumbs down to indicate not OK, or flipping the hand left and right or up and down to indicate a directional command.
[0159] Eye tracking is another input (e.g., to track where the user is looking and control display technology to render at a specific depth or range). In one embodiment, eye vergence may be determined using triangulation, and then accommodation may be determined using a vergence / accommodation model developed for that particular person.
[0160] Voice recognition may be another input that may be used alone or in combination with other inputs (e.g., totem tracking, eye tracking, gesture tracking, etc.). The system 900 may include an audio sensor 232 (e.g., a microphone) that receives an audio stream from the environment. The received audio stream may be processed (e.g., by the processing modules 260, 270 or the central server 1650) to recognize the user's voice (from other voices or background audio) and extract commands, targets, parameters, etc. from the audio stream. For example, the system 900 may identify from the audio stream that the phrase "move it there" was uttered, identify that this phrase was uttered by the wearer of the system 900 (and not another person in the user's environment), and extract from the phrase an executable command ("move it") and the presence of an object ("it") to be moved to a certain location ("there"). The object to be acted upon by the command may be referred to as the target of the command, and other information may be provided to the command as parameters. In this example, the location where the object should be moved is a parameter for the "move" command. Parameters can include, for example, location, time, other objects to be interacted with (e.g., "move it next to the red chair" or "give Linda the magic wand"), how the command should be executed (e.g., "play a song using the speakers upstairs"), etc.
[0161] As another example, the system 900 can process an audio stream using speech recognition techniques to input a string of text or modify the content of the text. The system 900 can incorporate speaker recognition techniques to determine who is speaking and speech recognition techniques to determine what is being said. Speech recognition techniques can include, alone or in combination, hidden Markov models, Gaussian mixture models, pattern matching algorithms, neural networks, matrix representations, vector quantization, speaker diarization, decision trees, and dynamic time warping (DTW) techniques. Speech recognition techniques can also include anti-speaker techniques such as cohort models and world models. Spectral features can be used to represent speaker characteristics.
[0162] With respect to the camera system, the exemplary wearable system 900 shown in FIG. 9 may include three pairs of cameras: a pair of relatively wide-FOV or passive SLAM cameras arranged on either side of the user's face, and a different pair of cameras oriented in front of the user to handle the stereoscopic imaging process 940 and capture hand gestures and totem / object trajectories in front of the user's face. The FOV cameras and pair of cameras for the stereoscopic process 940 may be part of the outward-facing imaging system 464 (shown in FIG. 4). The wearable system 900 may include an eye-tracking camera (which may be part of the inward-facing imaging system 462 shown in FIG. 4) oriented toward the user's eyes to triangulate eye vectors and other information. The wearable system 900 may also include one or more textured light projectors (such as infrared (IR) projectors) to inject texture into the scene.
[0163] 10 is a process flow diagram of an example embodiment of a method 1000 for determining user input to a wearable system. In this example, a user may interact with a totem. A user may have multiple totems. For example, a user may have one totem designated for social media applications, another totem for playing games, etc. In block 1010, the wearable system may detect movement of the totem. Movement of the totem may be recognized through an outward-facing imaging system or may be detected through sensors (e.g., tactile gloves, image sensors, hand tracking devices, eye tracking cameras, head pose sensors, etc.).
[0164] Based at least in part on the detected gestures, eye poses, head poses, or input through the totem, the wearable system detects the position, orientation, and / or movement of the totem (or the user's eyes or head or gestures) relative to a frame of reference in block 1020. The frame of reference may be a set of map points based on which the wearable system translates the totem's (or the user's) movements into actions or commands. In block 1030, the user's interactions with the totem are mapped. Based on the mapping of the user interactions to the frame of reference 1020, the system determines the user input in block 1040.
[0165] For example, a user may move a totem or physical object back and forth, indicate turning a virtual page, moving to the next page, or moving from one user interface (UI) display screen to another. As another example, a user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's gaze at a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as user input. In some implementations, the user's eye vergence can be tracked, and an accommodation / vergence model can be used to determine the user's eye accommodation state, which provides information about the rendering plane the user is focusing on. In some implementations, the wearable system can use a cone casting technique to determine real or virtual objects that are aligned with the user's head pose or eye pose. Generally described, the cone casting technique projects an invisible cone in the direction the user is looking and can identify any objects that interact with the cone. Cone casting can include casting a thin bundle of light rays with substantially little lateral width, or a light ray with substantial lateral width (e.g., a cone or a truncated cone) from an AR display (of a wearable system) toward a physical or virtual object. Cone casting with a single light ray may also be referred to as ray casting. Detailed examples of cone casting techniques are described in U.S. Application No. 15 / 473,444, entitled "Interactions with 3D Virtual Objects Using Poses and Multiple-DOF Controllers," filed March 29, 2017, the disclosure of which is incorporated herein by reference in its entirety.
[0166] The user interface may be projected by a display system as described herein (such as display 220 in FIG. 2A or 2B ). It may also be displayed using a variety of other techniques, such as one or more projectors. A projector may project an image onto a physical object, such as a canvas or a sphere. Interactions with the user interface may be tracked using one or more cameras external to or part of the system (e.g., using inward-facing imaging system 462 or outward-facing imaging system 464).
[0167] 11 is a process flow diagram of an example method 1100 for interacting with a virtual user interface. Method 1100 may be performed by a wearable system described herein.
[0168] In block 1110, the wearable system may identify a specific UI. The type of UI may be predetermined by the user. The wearable system may identify that a specific UI needs to be captured based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). In block 1120, the wearable system may generate data for the virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system may display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical position, head pose, or eye pose so that a ring UI may be displayed around the user or a planar UI may be displayed on a wall or in front of the user. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensory input, or any other type of collected data.
[0169] In block 1130, the wearable system may send data from the cloud to the display, or data may be sent from a local database to the display component. In block 1140, a UI is displayed to the user based on the sent data. For example, a light field display can project the virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may simply wait for commands from the user and generate more virtual content on the virtual UI in block 1150. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for a command (such as a gesture, head or eye movement, input from a user input device, etc.) and, if recognized (block 1160), the virtual content associated with the command may be displayed to the user (block 1170). As an example, the wearable system may wait for a user's hand gesture before mixing multiple stream tracks.
[0170] Additional examples of wearable systems, UIs, and user experiences (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety.
[0171] (Example of Objects in the Ocular Field (FOR) and Field of View (FOV)) FIG. 12A diagrammatically illustrates examples of a field of view (FOR) 1200, a world camera field of view (FOV) 1270, a user's field of view 1250, and a user's fixation field of view 1290. As described with reference to FIG. 4, the FOR 1200 comprises a portion of the user's surrounding environment that can be perceived by the user via a wearable system. The FOR may encompass a solid angle of 4π steradians surrounding the wearable system, since the wearer can move their body, head, or eyes to perceive virtually any direction in space. In other contexts, the wearer's movement may be more constrained, and thus the wearer's FOR may cover a smaller solid angle.
[0172] The field of view of the world camera 1270 may include the portion of the user's FOR currently being observed by the outward-facing imaging system 464. Referring to FIG. 4 , the field of view of the world camera 1270 may include the world 470 as observed by the wearable system 400 at a given time. The size of the FOV of the world camera 1270 may depend on the optical properties of the outward-facing imaging system 464. For example, the outward-facing imaging system 464 may include a wide-angle camera that may image a 190-degree space around the user. In some implementations, the FOV of the world camera 1270 may be greater than or equal to the natural FOV of the user's eyes.
[0173] The user's FOV 1250 may comprise the portion of the FOR 1200 that the user perceives at a given time. The FOV may depend on the size or optical properties of the wearable device's display. For example, the AR / MR display may include optics that provide AR / MR functionality when the user looks through a particular portion of the display. The FOV 1250 may correspond to the solid angle perceivable by the user when looking through an AR / MR display such as, for example, stacked waveguide assembly 480 ( FIG. 4 ) or planar waveguide 600 ( FIG. 6 ). In some embodiments, the user's FOV 1250 may be smaller than the natural FOV of the user's eyes.
[0174] The wearable system can also determine the user's field of fixation 1290. The field of fixation 1290 can include a portion of the FOV 1250 on which the user's eye may fixate (e.g., maintain visual line of sight). The field of fixation 1290 may correspond to the foveal region of the eye where light strikes. The field of fixation 1290 can be smaller than the user's FOV 1250; for example, the field of fixation can be a few degrees to about 5 degrees across it. As a result, the user can perceive some virtual objects within the FOV 1250 that are not within the field of fixation 1290 but are within the user's peripheral vision.
[0175] FIG. 12B diagrammatically illustrates an example of virtual objects in a user's field of view (FOV) and field of view (FOR). In FIG. 12B, FOR 1200 can contain a group of objects (e.g., 1210, 1220, 1230, 1242, and 1244), which can be perceived by the user via a wearable system. The objects in the user's FOR 1200 can be virtual and / or physical objects. For example, the user's FOR 1200 may include physical objects such as a chair, a sofa, a wall, etc. Virtual objects may include operating system objects such as, for example, a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, text, text editing applications, messaging applications, etc. Virtual objects may also include objects within an application, such as, for example, an avatar, a virtual object within a game, a graphic, or an image. Some virtual objects can be both operating system objects and objects within an application. In some embodiments, the wearable system can add virtual elements to existing physical objects. For example, the wearable system may add a virtual menu associated with a television in a room, and the virtual menu may give the user options to turn on the television or change the channel using the wearable system.
[0176] A virtual object may be a three-dimensional (3D), two-dimensional (2D), or one-dimensional (1D) object. For example, a virtual object may be a 3D coffee mug (which may represent virtual controls for a physical coffee maker). A virtual object may also be a 2D graphical representation of a clock (which displays the current time to the user). In some implementations, one or more virtual objects may be displayed within (or associated with) another virtual object. A virtual coffee mug may be shown inside the user's interface plane, but the virtual coffee mug appears to be 3D within this 2D planar virtual space.
[0177] Objects in a user's FOR can be part of a world map, as described with reference to FIG. 9 . Data associated with objects (e.g., location, semantic information, properties, etc.) can be stored in various data structures, such as arrays, lists, trees, hashes, graphs, etc. The index of each stored object may be determined, for example, by the object's location, if applicable. For example, the data structure may index objects by a single coordinate, such as the object's distance from a base position (e.g., distance to the left or right of the base position, distance from above or below the base position, or distance per depth from the base position). The base position may be determined based on the user's position (e.g., the position of the user's head). The base position may also be determined based on the position of a virtual or physical object (e.g., a target object) in the user's environment. Thus, the 3D space in the user's environment may be represented in a 2D user interface, with virtual objects arranged according to the object's distance from the base position.
[0178] In Figure 12B, FOV 1250 is diagrammatically illustrated by dashed line 1252. A user of the wearable system can perceive multiple objects within FOV 1250, such as object 1242, object 1244, and a portion of object 1230. As the user's posture (e.g., head pose or eye pose) changes, FOV 1250 changes correspondingly, and the objects within FOV 1250 may also change. For example, map 1210 is initially outside the user's FOV in Figure 12B. When the user looks at map 1210, map 1210 may move into the user's FOV 1250, and (for example) object 1230 may move out of the user's FOV 1250.
[0179] The wearable system may track objects within FOR 1200 and objects within FOV 1250. For example, local processing and data module 260 can communicate with remote processing module 270 and remote data repository 280 to retrieve virtual objects within the user's FOR. Local processing and data module 260 can store the virtual objects, for example, in a buffer or temporary storage device. Local processing and data module 260 can use techniques described herein to determine the user's FOV and render a subset of virtual objects within the user's FOV. As the user's pose changes, local processing and data module 260 can update the user's FOV and, accordingly, render a different set of virtual objects corresponding to the user's current FOV.
[0180] (Overview of different user input modes) A wearable system can be programmed to receive various input modes to perform operations. For example, a wearable system can receive two or more of the following types of input modes: voice commands, head pose, body pose (which may be measured, e.g., by an IMU in a beltpack or a sensor external to the HMD), eye gaze (also referred to herein as eye pose), hand gestures (or gestures by other body parts), signals from a user input device (e.g., a totem), environmental sensors, etc. Computing devices are typically engineered to generate a given output based on a single input from a user. For example, a user may enter a text message by typing on a keyboard or use a mouse to guide the movement of a virtual object, which is an example of a hand gesture input mode. As another example, a computing device can receive a stream of audio data from a user's voice and use speech recognition techniques to convert the audio data into executable commands.
[0181] User input modes may, in some cases, be non-exclusively categorized as direct user input or indirect user input. Direct user input may be user interaction directly provided by the user, for example, through voluntary movements of the user's body (e.g., turning the head or eyes, gazing at an object or location, uttering a word, or moving a finger or hand). As an example of direct user input, a user may interact with a virtual object using postures such as head posture, eye posture (also referred to as eye gaze), hand gestures, or another body posture. For example, a user may look at a virtual object (using their head and / or eyes). Another example of direct user input is user voice. For example, a user may say "launch browser" to cause the HMD to open a browser application. As yet another example of direct user input, a user may actuate a user input device, for example, through a touch gesture (e.g., touching a touch-sensitive portion of a totem) or body movement (e.g., rotating a totem that functions as a multi-degree-of-freedom controller).
[0182] In addition to, or as an alternative to, direct user input, a user can also interact with virtual objects based on indirect user input. Indirect user input may be determined from various contextual factors, such as the geographic location of the user or virtual object, the user's environment, etc. For example, the user's geographic location may be in the user's office (rather than the user's home), and different tasks (e.g., work-related tasks) can be performed based on the geographic location (e.g., derived from a GPS sensor).
[0183] Contextual factors can also include the affordances of a virtual object. The affordances of a virtual object can comprise a relationship between the virtual object and the object's environment, which provides an opportunity for an action or use associated with the object. The affordances may be determined based on, for example, the function, orientation, type, location, shape, and / or size of the object. The affordances may also be based on the environment in which the virtual object is located. As an example, the affordance of a horizontal table is that the object can be set on the table, while the affordance of a vertical wall is that the object can be hung from or projected onto the wall. As an example, "Put it there" can be uttered, and a virtual office calendar can be set to appear horizontally on the user's desk in the user's office.
[0184] A single mode of direct user input may create various limitations, and the number or types of available user interface operations may be limited due to the type of user input. For example, a user may not be able to zoom in or out using head pose because the head pose may not be able to provide precise user interaction. As another example, a user may need to move their thumb back and forth across the touchpad (or move their thumb over a long distance) to move a virtual object from the floor to the wall, which may cause fatigue to the user over time.
[0185] However, some direct input modes may be more convenient and intuitive for users to provide. For example, a user can speak to a wearable system and issue voice commands without having to type sentences using gesture-based keyboard input. As another example, a user can use hand gestures to point to a target virtual object rather than moving a cursor and identifying the target virtual object. While they may not be as convenient or intuitive, other direct input modes can increase the accuracy of user interaction. For example, a user can move a cursor to a virtual object and indicate that the virtual object is the target object. However, as described above, if a user wishes to select the same virtual object using direct user input (e.g., head pose or other input that dictates the outcome of the user's action), the user may need to control precise head movements, which may result in muscle fatigue. 3D environments (e.g., VR / AR / MR environments) may add additional challenges to user interaction because user input may also need to be defined with respect to depth (as opposed to a planar surface). This additional depth dimension can create more opportunity for error than in a 2D environment. For example, in a 2D environment, user input can be transformed relative to the horizontal and vertical axes within a coordinate system, whereas in a 3D environment, user input may need to be transformed relative to three axes (horizontal, vertical, and depth). Thus, imprecise translation of user input can result in errors in three axes (rather than two axes in a 2D environment).
[0186] To improve accuracy and reduce user fatigue when interacting with objects in 3D space while leveraging the existing benefits of direct user input, multiple modes of direct input may be used to perform user interface operations. Multimodal input can further enhance existing computing devices (particularly wearable devices) for interacting with virtual objects in data-rich and dynamic environments, such as AR, VR, or MR environments.
[0187] In a multimodal user input technique, one or more of the direct inputs may be used to identify a target virtual object (also referred to as a target) with which the user will interact and to determine a user interface action to be performed on the target virtual object. For example, the user interface action may include a command action such as select, move, zoom, pause, play, etc., and parameters of the command action (e.g., how to perform the action, where or when the action will occur, the object with which the target object will interact, etc.). As an example of identifying a target virtual object and determining an interaction to be performed on the target virtual object, a user may look at a virtual sticky note (head or eye posture input mode), point to a table (gesture input mode), and utter "move it there" (speech input mode). The wearable system can identify that the target virtual object in the phrase "move it there" is the virtual sticky note ("it") and can determine that the user interface action involves moving the virtual sticky note to the table ("there") (an executable command). In this example, the command action may be to "move" the virtual object, while parameters of the command action may include a destination object, which is the table the user is pointing to. Advantageously, in an embodiment, the wearable system may increase the overall accuracy of user interface actions or increase the convenience of user interaction by implementing user interface actions based on multiple modes of direct user input (e.g., in the above example, three modes: head / eye posture, gesture, and voice). For example, instead of uttering "move the leftmost browser 2.5 feet to the right," the user may utter "move it there" (without pointing to the object being moved in the spoken input) while using a head or hand gesture to indicate that the object is the leftmost browser, and using a head or hand movement to indicate the distance of the move.
[0188] (Example interactions within a virtual environment using various input modes) FIG. 13 illustrates an example of interacting with virtual objects using one mode of user input. In FIG. 13, a user 1310 is wearing an HMD and interacting with virtual content in three scenes 1300a, 1300b, and 1300c. The user's head position (and corresponding eye gaze direction) is represented by a geometric cone 1312a. In this example, the user can perceive the virtual content through the HMD's display 220. While interacting with the HMD, the user can input text messages via a user input device 466. In scene 1300a, the user's head is in its natural resting position 1312a, and the user's hands are also in their natural resting position 1316a. However, while a user may find it more comfortable to type text on the user input device 466, the user cannot see the interface on the user input device 466 to ensure that the characters are typed correctly.
[0189] To view text being typed on the user input device, the user can move their hands to position 1316b, as shown in scene 1300b. Thus, the hands would be within the FOV of the user's head when the head is in its natural resting position 1312a. However, position 1316b is not a natural resting position for the hands, which may result in fatigue for the user. Alternatively, as shown in scene 1300c, the user can move their head to position 1312c to maintain the hands in natural resting position 1316a. However, the muscles around the user's neck may fatigue due to the unnatural position of the head, and the user's FOV is directed toward the ground or floor rather than toward the outside world (which may be unsafe if the user is walking in a crowded area). In either scene 1300b or scene 1300c, the user's natural ergonomics are sacrificed to achieve the desired user interface operation when the user performs the user interface operation using the single-input mode.
[0190] The wearable systems described herein can at least partially mitigate the ergonomic limitations depicted in scenes 1300b and 1300c. For example, a virtual interface can be projected into the user's field of view in scene 1300a. The virtual interface can allow the user to observe inputs being typed from a natural position.
[0191] The wearable system can also display and support interaction with virtual content without device constraints. For example, the wearable system can present multiple types of virtual content to a user, and the user can use a touchpad to interact with one type of content while using a keyboard to interact with another type of content. Advantageously, in some embodiments, the wearable system can determine the virtual content (that the user is intended to act on) that is the target virtual object by calculating a confidence score (where a higher confidence score indicates a higher confidence (or likelihood) that the system has identified the correct target virtual object). Detailed examples of identifying the target virtual object are described with reference to FIGS. 15-18B .
[0192] 14 illustrates an example of selecting a virtual object using a combination of user input modes. In scene 1400a, a wearable system can present a user 1410 with multiple virtual objects, represented by a square 1422, a circle 1424, and a triangle 1426.
[0193] A user 1410 can interact with virtual objects using head pose, as shown in scene 1400b. This is an example of a head pose input mode. The head pose input mode may involve cone casting to target or select virtual objects. For example, the wearable system can cast a cone 1430 from the user's head toward the virtual objects. The wearable system can detect whether one or more of the virtual objects fall within the volume of the cone and identify the object that the user intends to select. In this example, the cone 1430 intersects with the circle 1424 and the triangle 1426. Thus, the wearable system can determine that the user intends to select either the circle 1424 or the triangle 1426. However, because the cone 1430 intersects both the circle 1424 and the triangle 1426, the wearable system may not be able to confirm whether the target virtual object is the circle 1424 or the triangle 1426 based on the head pose input alone.
[0194] In scene 1400c, a user 1410 can interact with virtual objects by manually orienting a user input device 466, such as a totem (e.g., a handheld remote control device). This is an example of a gesture input mode. In this scene, the wearable system can determine that either the circle 1424 or the square 1422 is the intended target because these two objects are in the direction that the user input device 466 is pointing. In this example, the wearable system can determine the direction of the user input device 466 by detecting the position or orientation of the user input device 466 (e.g., via an IMU within the user input device 466) or by performing cone casting originating from the user input device 466. Because the circle 1424 and the square 1422 are both candidate target virtual objects, the wearable system cannot confirm, based solely on the gesture input mode, which of them is the object the user actually desires to select.
[0195] In scene 1400d, the wearable system can determine a target virtual object using multi-modal user input. For example, the wearable system can identify the target virtual object using both results obtained from cone casting (head pose input mode) and the orientation of a user input device (gesture input mode). In this example, circle 1424 is an identified candidate in both the results from cone casting and the results obtained from the user input device. Thus, using these two input modes, the wearable system can determine with high confidence that the target virtual object is circle 1424. As further illustrated in scene 1400d, the user can give voice command 1442 (illustrated as “move it”), which is an example of a third input mode (i.e., voice), to interact with the target virtual object. The wearable system can associate the word “it” with the target virtual object and the word “move” with the command to be performed, thus moving circle 1424. However, the voice command 1442 itself (without any indication from the user input device 466 or cone casting 143) may cause confusion to the wearable system because the wearable system may not know the object associated with the word "it."
[0196] Advantageously, in some embodiments, by receiving multiple input modes to identify and interact with a virtual object, the amount of precision required per input mode may be reduced. For example, cone casting may be unable to accurately indicate an object in a distant rendering plane because the diameter of the cone increases as the cone moves farther away from the user. As another example, a user may need to hold an input device in a particular orientation, point toward a target object, and speak using a particular phrase or pace to ensure correct voice input. However, by combining results from voice input and cone casting (either from head pose or gestures using the input device), the wearable system can still identify a target virtual object without requiring precision from either input (e.g., cone casting or voice input). For example, even if cone casting selects multiple objects (e.g., as described with reference to scenes 1400b and 1400c), voice input may help narrow down the selection (e.g., increase a confidence score for the selection). For example, cone casting may capture three objects, of which the first object is to the user's right, the second object is to the user's left, and the third object is in the center of the user's FOV. The user can narrow the selection by uttering, "Select the rightmost object." As another example, two identically shaped objects may be present within the user's FOV. For the user to select the correct object, the user may need to provide further description of the object via voice command. For example, instead of uttering, "Select the square object," the user may need to utter, "Select the red square object." However, with cone casting, the voice command may not need to be precise. For example, the user may look at one of the square objects and utter, "Select the square object" or even, "Select object."The wearable system can automatically select square objects that coincide with the user's line of sight and will not select square objects that are not in the user's line of sight.
[0197] In some embodiments, the system may have a hierarchy of preferences for combinations of input modes. For example, users tend to look in the direction their head is facing. Therefore, eye gaze and head pose may provide similar information to each other. Combining head pose and eye gaze may be less preferred because the combination does not provide much redundant information compared to using eye gaze alone or head pose alone. Thus, the system may use a hierarchy of modal input preferences for selecting a modal input that generally provides contrasting information rather than redundant information. In some embodiments, the hierarchy uses head pose and speech as the primary modal input, followed by eye gaze and gesture.
[0198] Thus, as described further herein, based on the multi-modal input, the system can calculate confidence scores for various objects in the user's environment, with each such object being a target object. The system can select the particular object in the environment with the highest confidence score as the target object.
[0199] 15 illustrates an example of interacting with a virtual object using a combination of direct user input. As depicted in FIG. 15, a user 1510 is wearing an HMD 1502 configured to display virtual content. The HMD 1502 may be part of the wearable system 200 described herein and may include a belt-mounted power and processing pack 1503. The HMD 1502 may be configured to receive user input from a totem 1516. The user 1510 of the HMD 1502 may have a first FOV 1514. The user can observe a virtual object 1512 within the first FOV 1514.
[0200] A user 1510 can interact with a virtual object 1512 based on a combination of direct inputs. For example, the user 1510 can select a virtual object 1512 through a cone casting technique based on the user's head or eye posture, by a totem 1516, by a voice command, or by a combination of these (or other) input modes (e.g., as described with reference to FIG. 14 ).
[0201] The user 1510 may shift their head pose to move a selected virtual object 1512. For example, the user may turn their head to the left, updating the FOV from a first FOV 1514 to a second FOV 1524 (as shown in scene 1500a to scene 1500b). The user's head movement, combined with other direct inputs, may move the virtual object from the first FOV 1514 to the second FOV 1524. For example, the head pose change may be aggregated with other inputs, such as a voice command (“move it there”), guidance from a totem 1516, or eye gaze direction (e.g., as recorded by the inward-facing imaging system 462 shown in FIG. 4). In this example, the HMD 1502 may use the updated FOV 1524 as a general area into which the virtual object 1512 should be moved. The HMD 1502 can further determine a destination for the movement of the virtual object 1512 based on the user's gaze direction. As another example, the HMD may capture a voice command "move it there." The HMD can identify the virtual object 1512 as the object with which the user will interact (because the user has pre-selected the virtual object 1512). The HMD can further determine that the user intends to move the object from the FOV 1514 to the FOV 1524 by detecting a change in the user's head pose. In this example, the virtual object 1512 may initially be in a central portion of the user's first FOV 1514. Based on the voice command and the user's head pose, the HMD may move the virtual object to the center of the user's second FOV 1524.
[0202] Example of Using Multimodal User Input to Identify Target Virtual Object or User Interface Action 14, in some situations, the wearable system may be unable to identify (with sufficient confidence) the target virtual object with which the user intends to interact using a single input mode. Furthermore, even when multiple modes of user input are used, user input in one mode may indicate one virtual object, while user input in another mode may indicate a different virtual object.
[0203] To provide an improved wearable system that resolves ambiguities and supports multi-modal user input, the wearable system can aggregate the modal user inputs, calculate a confidence score, and identify the desired virtual object or user interface action. As discussed above, a higher confidence score indicates a higher probability or likelihood that the system has identified the desired target object.
[0204] FIG. 16 illustrates an exemplary computing environment for aggregating input modes. The exemplary environment 1600 includes, for example, three virtual objects associated with applications A 1672, B 1674, and C 1676. As described with reference to FIGS. 2A, 2B, and 9, the wearable system can include various sensors, receive various user inputs from these sensors, analyze the user inputs, and interact with mixed reality 960, for example, using the transmodal input fusion techniques described herein. In the exemplary environment 1600, a central runtime server 1650 can aggregate direct input 1610 and indirect user input 1630 to produce multimodal interactions for the applications. Examples of direct input 1610 may include gestures 1612, head pose 1614, voice input 1618, totems 1622, eye gaze direction (e.g., eye gaze tracking 1624), other types of direct input 1626, etc. Examples of indirect input 1630 may include environmental information (e.g., environmental tracking 1632) and geographic location 1634. Central runtime server 1650 may include remote processing module 270. In some implementations, local processing and data module 260 (or processor 128) may perform one or more functions of central runtime server 1650. Local processing and data module 260 may also communicate with remote processing module 270 to aggregate input modes.
[0205] The wearable system can track gestures 1612 using an outward-facing imaging system 464. The wearable system can track hand gestures using various techniques described in FIG. 9 . For example, the outward-facing imaging system 464 can obtain images of the user's hands and map the images to corresponding hand gestures. The outward-facing imaging system 464 may also image the user's hand gestures using an FOV camera or a depth camera (configured for depth detection). The central runtime server 1650 can identify the user's head gestures using the object recognizer 708. The gestures 1612 can also be tracked by the user input device 466. For example, the user input device 466 may include a touch-sensitive surface, which can track the user's hand movements, such as swipe gestures or tap gestures.
[0206] The HMD can use an IMU to recognize head pose 1614. The head 1410 may have multiple degrees of freedom, including three types of rotation (e.g., yaw, pitch, and roll) and three types of translation (e.g., pitch, sway, and heave). The IMU can be configured, for example, to measure 3-DOF or 6-DOF movement of the head. Measurements obtained from the IMU may be communicated to a central runtime server 1650 for processing (e.g., to identify head pose).
[0207] The wearable system can implement eye gaze tracking 1624 using an inward-facing imaging system 462. For example, the inward-facing imaging system 462 can include an eye camera configured to capture images of the user's eye region. The central runtime server 1650 can analyze the images (e.g., via the object recognizer 708) to infer the user's gaze direction or track the user's eye movement.
[0208] The wearable system can also receive input from a totem 1622. As described herein, the totem 1622 can be an embodiment of a user input device 466. Additionally, or alternatively, the wearable system can receive voice input 1618 from the user. Input from the totem 1622 and the voice input 1618 can be communicated to a central runtime server 1650. The central runtime server 1650 can use natural language processing, in real time or near real time, to analyze the user's audio data (e.g., obtained from the microphone 232). The central runtime server 1650 can identify the content of the utterance by applying various speech recognition algorithms, such as, for example, hidden Markov models, dynamic time warping (DTW)-based speech recognition, deep learning algorithms such as neural networks, deep feedforward and recurrent neural networks, end-to-end automatic speech recognition, machine learning algorithms (described with reference to FIGS. 7 and 9 ), other algorithms using semantic analysis, acoustic modeling, or language modeling, etc. The central runtime server 1650 can also apply speech recognition algorithms, which can identify the identity of the speaker, such as whether the speaker is the user of the wearable device or a person within the user's background.
[0209] The central runtime server 1650 can also receive indirect input when a user interacts with the HMD. The HMD can include various environmental sensors, as described with reference to FIGS. 2A and 2B. Using data obtained by the environmental sensors (alone or in combination with data related to direct input 1610), the central runtime server 1650 can reconstruct or update the user's environment (e.g., map 920, etc.). For example, the central runtime server 1650 can determine the user's ambient light conditions based on the user's environment. These ambient light conditions can be used to determine virtual objects with which the user can interact. For example, when the user is in a bright environment, the central runtime server 1650 may identify the target virtual object as a virtual object that supports gesture 1612 as an input mode because the camera can observe the user's gesture 1612. However, if the environment is dark, the central runtime server 1650 may determine that the virtual object may be an object that supports voice input 1618 rather than gesture 1612.
[0210] The central runtime server 1650 can perform environment tracking 1632, aggregate direct input modes, and produce multi-modal interaction for multiple applications. As an example, when a user moves from a quiet environment into a noisy environment, the central runtime server 1650 may disable voice input 1618. Additional examples of selecting an input mode based on the environment are described with further reference to FIG. 24.
[0211] The central runtime server 1650 can also identify target virtual objects based on the user's geographic location information. Geographic location information 1634 may also be obtained from environmental sensors (e.g., GPS sensors, etc.). The central runtime server 1650 may identify virtual objects for potential user interaction where the distance between the virtual object and the user is within a threshold distance. Advantageously, in some embodiments, the cone in the cone casting may have a length that is adjustable by the system (e.g., based on the number or density of objects in the environment). By selecting objects within a certain radius of the user, the number of potential objects that may be target objects can be significantly reduced. Additional examples of using indirect input as an input mode are described with reference to FIG. 21 .
[0212] (Example of identifying a target object) The central runtime server 1650 can determine the target object using various techniques. FIG. 17A illustrates an example of identifying a target object using lattice tree analysis. The central runtime server 1650 can derive given values from input sources and produce a lattice of possible values for candidate virtual objects with which a user may potentially interact. In some embodiments, the value can be a confidence score. The confidence score can include a ranking, rating, evaluation, quantitative or qualitative value (e.g., a number in the range of 1 to 10, a percentage or percentile, or qualitative values of "A," "B," "C," etc.), etc. Each candidate object may be associated with a confidence score, and in some cases, the candidate object with the highest confidence score (e.g., higher than the confidence scores of other objects or higher than a threshold score) is selected by the system as the target object. In other cases, objects with confidence scores below a threshold confidence score are eliminated from consideration as target objects by the system, which can improve computational efficiency.
[0213] Many examples herein refer to selecting a target virtual object or selecting from a group of virtual objects. This is intended to illustrate example implementations and is not intended to be limiting. The described techniques can be applied to virtual objects or physical objects in a user's environment. For example, the voice command "move it there" may refer to moving a virtual object (e.g., a virtual calendar) onto a physical object (e.g., a horizontal surface on a user's desk). Or the voice command "move it there" may refer to moving a virtual object (e.g., a virtual word processing application) to another location within another virtual object (e.g., another position within a user's virtual desktop).
[0214] The context of a command may also provide information about whether the system should attempt to identify a virtual object, a physical object, or both. For example, in the command "move it there," the system may recognize that "it" is a virtual object because AR / VR / MR systems cannot move real physical objects. Therefore, the system may eliminate physical objects as candidates for "it." As described in the above example, the target location "there" can be a virtual object (e.g., a user's virtual desktop) or a physical object (e.g., a user's desk).
[0215] The system may also assign confidence scores to objects in the user's environment, which may be in the FOR, FOV, or field of fixation (see, e.g., FIG. 12A ), depending on the context and the system's goal at the time. For example, a user may want to move a virtual calendar to a location on the user's desk, both of which are within the user's FOV. The system may analyze objects within the user's FOV rather than all objects within the user's FOR because the context of the situation suggests that a command to move the virtual calendar would result in a target destination within the user's FOV, which may improve processing speed or efficiency. In another case, a user may be browsing a menu of movie selections in a virtual movie application and may fixate on several movie options. The system may analyze (and, e.g., provide confidence scores for) only movie selections within the user's field of fixation (e.g., based on the user's eye line of sight) rather than the entire FOV (or FOR), which may also increase processing efficiency or speed.
[0216] 17A , a user can interact with a virtual environment using two input modes: head pose 1614 and eye gaze 1624. Based on the head pose 1614, the central runtime server 1650 can identify two candidate virtual objects associated with application A 1672 and application B 1674. The central runtime server 1650 can distribute a 100% confidence score evenly between application A 1672 and application B 1674. As a result, application A 1672 and application B 1674 may each be assigned a confidence score of 50%. The central runtime server 1650 can also identify two candidate virtual objects (application A 1672 and application C 1676) based on the direction of eye gaze 1624. The central runtime server 1650 can also split the 100% confidence between application A 1672 and application C 1676.
[0217] The central runtime server 1650 may implement a lattice compression logic function 1712 to reduce or eliminate mismatched confidence values or those below a certain threshold that are inconsistent among multiple input modes and determine the application with which the user most likely intends to interact. For example, in FIG. 17A , the central runtime server 1650 may eliminate application B 1674 and application C 1676 because these two virtual objects are not identified by both head pose 1614 and eye gaze 1624 analysis. As another example, the central runtime server 1650 may aggregate the values assigned to each application. The central runtime server 1650 may set a threshold confidence value of 80% or greater. In this example, the aggregated value for application A 1672 is 100% (50% + 50%), the aggregated value for application B 1674 is 50%, and the value for application C 1676 is 50%. Because the individual trust values for applications B and C are below the threshold trust value, the central runtime server 1650 may be programmed to select application A 1672 rather than selecting applications B and C because the aggregate trust value of application A (100%) is above the threshold trust value.
[0218] 17A divides the values associated with the input devices (e.g., confidence scores) equally among the candidate virtual objects, in some embodiments, the value distribution may not be equal among the candidate virtual objects. For example, if head pose 1614 has a value of 10, application A 1672 may receive a value of 7, and application B 1674 may receive a value of 3 (because the head pose is pointing more toward A 1672). As another example, if head pose 1614 has a qualitative grade of "A," application A 1672 may be assigned a grade of "A," while applications B 1674 and C 1676 receive neither from head pose 1614.
[0219] The wearable system (e.g., central runtime server 1650) can assign a focus indicator to a target virtual object so that a user may more easily perceive the target virtual object. The focus indicator can be a visual focus indicator. For example, the focus indicator can comprise a halo (substantially surrounding or near the object), a color, a perceived size or depth change (e.g., causing the target object to appear closer and / or larger when selected), or other visual effect that attracts the user's attention. The focus indicator can also include an audible or tactile effect, such as a vibration, a ringing, a beep, etc. The focus indicator can provide useful feedback to the user that the system is “working correctly” by confirming to the user (via the focus indicator) that the system correctly determined the object associated with the command (e.g., correctly determined “it” and “there” in a “move it there” command). For example, the identified target virtual object can be assigned a first focus indicator, and the destination location (e.g., “there” in a command) can be assigned a second focus indicator. In some cases, if the system incorrectly determines the target object, the user may override the system's determination by, for example, gazing (fixating) on the correct object and providing a voice command such as "not that."
[0220] Example of Identifying Target User Interface Actions In addition to, or as an alternative to, identifying a target virtual object, the central runtime server 1650 can also determine a target user interface action based on multiple received inputs. FIG. 17B illustrates an example of determining a target user interface action based on multi-modal input. As depicted, the central runtime server 1650 can receive multiple inputs in the form of head poses 1614 and gestures 1612. The central runtime server 1650 can display multiple virtual objects to the user associated with application A 1672 and application B 1674, for example. However, the head pose input mode itself may be insufficient to determine the desired user interface action because there is a 50% confidence (shown as modification option 1772) that the head pose applies to a user interface action associated with application A 1672 and another 50% confidence (shown as modification option 1774) that the head pose applies to another user interface action associated with application B 1674.
[0221] In various embodiments, particular applications or types of user interface actions may be programmed to be more responsive to certain input modes. For example, the HTML tags or Java script programming of application B 1674 may be set to be more responsive to gesture input than that of application A 1672. For example, application A 1672 may be more responsive to head pose 1672 than to gesture 1612, while a "select" action may be more responsive to gesture 1612 (e.g., a tap gesture) than to head pose 1614 because a user is more likely to select an object using a gesture than a head pose, in some cases.
[0222] 17B , gesture 1612 may be responsive to a certain type of user interface action in application B 1674. As shown, gesture 1612 may have a higher confidence associated with the user interface action for application B, while gesture 1612 may be inapplicable for interface actions in application A 1672. Thus, if the target virtual object is application A 1672, the input received from head pose 1614 may be the target user interface action. However, if the target virtual object is application B 1674, the input received from gesture 1612 (alone or in combination with input based on head pose 1614) may be the target user interface action.
[0223] As another example, because gesture 1612 has a higher confidence level than head pose 1614 when a user interacts with application B, gesture 1612 may be the primary input mode for application B 1674, while head pose 1614 may be the secondary input mode. Thus, input received from gesture 1612 may be associated with a higher weighting than head pose 1614. For example, if the head pose indicates that a virtual object associated with application B 1674 should remain stationary, while gesture 1612 indicates that the virtual object should be moved to the left, the central runtime server 1650 may render the virtual object as moving to the left. In some implementations, the wearable system may allow a user to interact with the virtual object using the primary input mode and may consider a secondary input mode if the primary input mode is insufficient to determine the user's action. For example, a user may primarily interact with application B 1674 using gesture 1612. However, if the HMD is unable to determine the target user interface action (e.g., because multiple candidate virtual objects may exist within application B 1674 or if the gesture 1612 is unclear), the HMD can use the head pose as input to confirm the target virtual object or the target user interface action to be performed on application B 1674.
[0224] The scores associated with each input mode may be aggregated to determine the desired user interface behavior. FIG. 17C illustrates an example of aggregating confidence scores associated with input modes for a virtual object. As illustrated in this example, head pose input 1614 produces a higher confidence score for application A (80% confidence) than for application B (30% confidence), while gesture input 1612 produces a higher confidence score for application B (60% confidence) than for application A (30% confidence). The central runtime server 1650 can aggregate confidence scores for each object based on the confidence scores derived from each user input mode. For example, the central runtime server 1650 can produce an aggregate score 110 for application A 1672 and an aggregate score 90 for application B 1674. The aggregated scores may be a weighted or unweighted average or other mathematical combination. Because application A 1672 has a higher aggregate score than application B 1674, the central runtime server 1650 may select application A as the application to be interacted with. Additionally, or alternatively, due to application A's 1672 higher aggregate score, the central runtime server 1650 can determine that application B is more "responsive" to gesture 1612 than application A, but that head pose 1614 and gesture 1612 are intended to perform a user interface action on application A 1672.
[0225] In this example, the central runtime server 1650 aggregates the resulting confidence scores by adding the confidence scores of various inputs related to a given object. In various other embodiments, the central runtime server 1650 can aggregate confidence scores using techniques other than simple addition. For example, input modes or scores may be associated with weights. As a result, aggregation of confidence scores will take into account the weights assigned to input modes or scores. The weights may be user-adjustable, allowing users to selectively adjust the “responsiveness” of multimodal interactions with the HMD. Weights may also be contextual. For example, weights used in public places may emphasize head or eye postures over hand gestures to avoid potentially socially awkward behaviors that would cause users to frequently gesture while operating the HMD. As another example, on a subway, airplane, or train, voice commands may be weighted lower than head or eye postures because users may not want to speak loudly into their HMD in such environments. Environmental sensors (e.g., GPS) may assist the user in determining the appropriate context in which to operate the HMD.
[0226] While the examples in Figures 17A-17C are illustrated with reference to two objects, the techniques described herein can also be applied when more or fewer objects are present. Additionally, the techniques described with reference to these figures can also be applied to a wearable system's application or to virtual objects associated with one or more applications. Furthermore, the techniques described herein can also be applied to direct or indirect input modes other than head pose, eye gaze, or gesture. For example, voice commands may also be used. Additionally, while a central runtime server 1650 is used throughout as an example to illustrate the processing of various input modes, the HMD's local processing and data module 260 may also perform some or all of the operations in addition to, or instead of, the central runtime server 1650.
[0227] (Example Techniques for Calculating Confidence Scores) The wearable system can calculate the confidence score of an object using one or a combination of various techniques. Figures 18A and 18B illustrate an example of calculating a confidence score for an object within a user's FOV. The user's FOV may be calculated based on the user's head pose or eye gaze, for example, during cone casting. The confidence scores in Figures 18A and 18B may be based on a single input mode (e.g., the user's head pose, etc.). Multiple confidence scores may be calculated (for some or all of the various multimodal inputs) based on multimodal user input and then aggregated to determine a user interface action or a target virtual object.
[0228] 18A illustrates an example in which a confidence score for a virtual object is calculated based on the portion of the virtual object that falls within a user's FOV 1810. In FIG. 18A , the user's FOV has portions of two virtual objects (represented by a circle 1802 and a triangle 1804). The wearable system may assign confidence scores to the circle and the triangle based on the percentage of the objects' projected areas that fall within the FOV 1810. As shown, approximately half of the circle 1802 falls within the FOV 1810, so the wearable system may assign a confidence score of 50% to the circle 1802. As another example, approximately 75% of the triangle is within the FOV 1810, so the wearable system may assign a confidence score of 75% to the triangle 1804.
[0229] The wearable system can use a regression analysis of the content in the FOV and FOR to calculate the percentage of the virtual object within the FOV. As described with reference to FIG. 12B , the wearable system tracks the object within the FOR, but the wearable system may deliver the object (or a portion of the object) within the FOV to a rendering projector (e.g., display 220) for display within the FOV. The wearable system can determine the portion provided for the rendering projector, analyze the proportion delivered to the rendering projector relative to the virtual object as a whole, and determine the percentage of the virtual object within the FOV.
[0230] In addition to, or as an alternative to, calculating a confidence score based on the proportional area that falls within the FOV, the wearable system can also analyze the space near an object within the FOV and determine the object's confidence score. FIG. 18B illustrates an example of calculating a confidence score based on the uniformity of the space surrounding a virtual object within an FOV 1820. The FOV 1820 includes two virtual objects, as depicted by a triangle 1814 and a circle 1812. The space around each virtual object may be represented by a vector. For example, the space around virtual object 1812 may be represented by vectors 1822a, 1822b, 1822c, and 1822d, while the space around virtual object 1814 may be represented by vectors 1824a, 1824b, 1824c, and 1824d. The vectors may originate from the virtual object (or its boundary) and terminate at the edge of the FOV 1820. The system can analyze the distribution of vector lengths from objects to the edges of the FOV to determine which of the objects are located closer to the center of the FOV. For example, objects in the very center of a circular FOV will have a relatively uniform distribution of vector lengths, while objects very close to the edges will have a non-uniform distribution of vector lengths (because some vectors pointing to nearby edges will be shorter, while vectors pointing to the farthest edges will be longer). As depicted in FIG. 18B , the distribution of vector lengths from virtual triangle 1814 to the edges of the field of view 1820 is more variable than the distribution of vector lengths from circle 1812 to the edges of the field of view 1820, indicating that virtual circle 1812 is closer to the center of the FOV 1820 than virtual triangle 1814. The variability of the distribution of vector lengths may be represented by the standard deviation or variance (or other statistical measure) of the lengths. The wearable system can therefore assign a higher confidence score to virtual circle 1812 than to virtual triangle 1814.
[0231] In addition to the techniques described with reference to FIGS. 18A and 18B , the wearable system can assign confidence scores to virtual objects based on a historical analysis of a user's interactions. As an example, the wearable system can assign higher confidence scores to virtual objects with which the user frequently interacts. As another example, one user may tend to move a virtual object using voice commands (e.g., "move it there"), while another user may prefer to use hand gestures (e.g., by reaching out, "grabbing," the virtual object, and moving it to another position). The system can determine such user tendencies from the historical analysis. As yet another example, an input mode may be frequently associated with a particular user interface action or a particular virtual object, and as a result, the wearable system may increase the confidence score for the particular user interface action or virtual object, even though alternative user interface actions or virtual objects may exist based on the same input.
[0232] As depicted in FIG. 18A or 18B , given either field of view 1810 or 1820, the second input mode can facilitate selection of an appropriate virtual object or an appropriate user interface action on the virtual object. For example, a user can utter "enlarge the triangle" to increase the size of the triangle in field of view 1810. As another example, in FIG. 18A , a user may give a voice command such as "double it." The wearable system may determine that the subject (e.g., target object) of the voice command is virtual object 1804 because virtual object 1804 has a higher confidence score based on head pose. Advantageously, in some embodiments, this reduces the specificity of the interaction required to produce a desired result. For example, a user does not need to utter "double the triangle" for the wearable system to achieve the same interaction.
[0233] The triangles and circles in Figures 18A and 18B are for illustrative purposes only. The various techniques described herein can also be applied to virtual content that supports more complex user interactions.
[0234] (Example multimodal interactions in the physical environment) In addition to, or as an alternative to, interacting with virtual objects, wearable systems can also provide a wide range of interactions within a real-world environment. Figures 19A and 19B illustrate an example of interacting with a physical environment using multimodal input. In Figure 19A, three input modes are illustrated: hand gestures 1960, head pose 1920, and input from a user input device 1940. Head pose 1920 can be determined using a posture sensor. The posture sensor may be an IMU, gyroscope, magnetometer, accelerometer, or other type of sensor described with reference to Figures 2A and 2B. Hand gestures 1960 may be measured using an outward-facing imaging system 464, while user input device 1940 may be an embodiment of user input device 466 shown in Figure 4.
[0235] In some embodiments, the wearable system can also measure the user's eye line of sight. The eye line of sight may include a vector extending from each of the user's eyes to the point where the two eye lines of sight converge. The vector can be used to determine the direction the user is looking and to select or identify virtual content at the convergence point or along the vector. Such eye lines of sight may be determined by eye-tracking techniques such as, for example, glint detection, iris or pupil shape mapping, infrared illumination, or binocular imaging with a regression of the intersection point resulting from individual pupil orientations. The eye line of sight or head pose can then be considered as a source point for cone-casting or ray-casting for virtual object selection.
[0236] As described herein, an interaction event for moving selected virtual content within a user's environment (e.g., "put it there") may require determination of a command action (e.g., "put"), a target (e.g., "it" as may be determined from the multimodal selection techniques described above), and a parameter (e.g., "there"). The command action (or command for short) and target (also referred to as a target object or target virtual object) may be determined using a combination of input modes. For example, a command to move a target 1912 may be based on a head pose 1920 change (e.g., a head turn or nod) or a hand gesture 1960 (e.g., a swipe gesture), alone or in combination. As another example, the target 1912 may be determined based on a combination of head pose and eye gaze. Thus, commands based on multimodal user input may also sometimes be referred to as multimodal input commands.
[0237] The parameters may also be determined using a single input or multi-modal input. The parameters may be associated with objects in the user's physical environment (e.g., a table or wall) or objects in the user's virtual environment (e.g., a movie application, an avatar, or a virtual building in a game). Identifying real-world parameters can, in some embodiments, enable faster and more accurate content placement responses. For example, a particular virtual object (or portion of a virtual object) may be approximately planar with a horizontal orientation (e.g., the virtual object's normal is perpendicular to the room floor). When a user initiates an interaction that moves the virtual object, the wearable system can identify a real-world surface with a similar orientation (e.g., the surface of a table) and move the virtual object to the real-world surface. In some embodiments, such movement may be automatic. For example, a user may want to move a virtual book from its location on the floor. The only horizontal surface available in the room may be the user's desk. Thus, the wearable system may automatically move a virtual book to a desk surface in response to a voice command of "move it" without the user having to input any additional commands or parameters, because the desk surface is the most likely location where the user would want to move the book. As another example, the wearable system may identify a real-world surface of a suitable size for given content, thereby providing a better parameter match for the user. For example, if a user is viewing a virtual video screen with a given display size and wishes to move it to a particular surface using a simple voice command, the system may determine the real-world surface that provides the surface area necessary to best support the virtual video's display size.
[0238] The wearable system can identify target parameters (e.g., target surfaces) using techniques described with reference to identifying the target virtual object. For example, the wearable system can calculate confidence scores associated with multiple target parameters based on indirect or direct user input. As an example, the wearable system can calculate a confidence score associated with a wall based on direct input (such as a user's head pose) and indirect input (such as a characteristic of the wall (e.g., vertical surface)).
[0239] (Example Techniques for Identifying Real-World Parameters) The wearable system can use various techniques to determine parameters (such as target locations) of multi-modal input commands. For example, the wearable system can use various depth-sensing techniques, such as applying a SLAM protocol to environmental depth information (e.g., as described with reference to FIG. 9 ) or building or accessing a mesh model of the environment. In some embodiments, depth sensing determines the distance between known points in 3D space (e.g., the distance between sensors on an HMD) and points of interest ("POIs") on the surfaces of objects in the real world (e.g., walls for positioning virtual content). This depth information may be stored in a world map 920. Parameters for interaction may be determined based on a collection of POIs.
[0240] The wearable system can apply these depth-sensing techniques to data obtained from a depth sensor to determine the demarcation and boundaries of the physical environment. The depth sensor may be part of an outward-facing imaging system 464. In some embodiments, the depth sensor is coupled to an IMU. The data obtained from the depth sensor can be used to determine the orientation of multiple POIs relative to one another. For example, the wearable system can calculate a truncated signed distance function ("TSDF") for the POIs. The TSDF may include a numerical value for each POI. The numerical value may be zero when the point is within a given tolerance of a particular plane, positive when the point is spaced apart from the particular plane in a first direction (e.g., upward or outward), and negative when the point is spaced apart from the particular plane in a second (e.g., opposite) direction (e.g., downward or inward). The calculated TSDF can be used to define a 3-D volumetric grid of bricks or boxes aligned within, above, and below a particular plane, along an orientation as determined by the IMU, that constitutes or represents a particular surface.
[0241] POIs outside a given planar tolerance (e.g., with absolute values of TSDFs that exceed the tolerance) may be eliminated, leaving only POIs adjacent to each other within the given tolerance to create a virtual representation of a surface in the real-world environment. For example, the real-world environment may include a conference table. Various other objects (e.g., a telephone, a laptop computer, a coffee mug, etc.) may be present on top of the conference table. With respect to the surface of the conference table, the wearable system may retain the POIs associated with the conference table and remove POIs related to other objects. As a result, a planar map (outlining the surface of the conference table) may represent the conference table with only points belonging to the conference table. The map may exclude points associated with objects on top of the conference table. In one embodiment, the collection of remaining POIs in the planar map may be referred to as the “workable surface” of the environment, because these areas of the planar map represent a space where virtual objects may be placed. For example, when a user desires to move a virtual screen onto a table, the wearable system may identify suitable surfaces in the user's environment (such as a tabletop, a wall, etc.) while excluding objects (e.g., a coffee mug or a pencil or a picture on the wall) or surfaces that are not suitable for placing the screen (e.g., the surface of a bookshelf). In this example, the identified suitable surfaces may be workable surfaces of the environment.
[0242] Referring back to the example shown in FIG. 19A , the environment 1900 may include a physical wall 1950. The HMD or user input device 1940 may house a depth sensor system (e.g., a time-of-flight sensor or a vertical cavity surface-emitting laser (VCSEL) or the like) and an attitude sensor (e.g., an IMU or the like). Data acquired by the depth sensor system may be used to identify various POIs within the user's environment. The wearable system may group POIs that are generally planar together to form a bounding polygon 1910. The bounding polygon 1910 may be an exemplary embodiment of a workable surface.
[0243] In some embodiments, the outward-facing imaging system 464 can identify a user gesture 1960, which may include pointing a finger at an area within the real-world environment 1900. The outward-facing imaging system 464 can identify the pre-measured bounding polygon 1910 by determining a sparse point vector structure of the finger pointing towards the bounding polygon 1910.
[0244] As illustrated in FIG. 19A, a virtual video screen 1930 can exist inside a bounding polygon 1910. A user can interact with a virtual object 1912 inside the virtual video screen 1930 using multimodal input. FIG. 19B depicts interaction using multimodal input of virtual content within a real-world environment. The environment in FIG. 19B includes a vertical surface 1915 (which may be part of a wall) and a tabletop surface 1917. In a first state 1970a, virtual content 1926 is initially displayed within a bounding polygon 1972a on the wall surface 1915. A user can select the virtual object 1926, for example, through cone casting or multimodal input (including two or more of a gesture 1960, a head pose 1920, an eye gaze, or input from a user input device 1940).
[0245] The user can use another input as part of the multimodal input to select surface 1917 as a destination. For example, the user can use head pose combined with hand gestures to indicate that surface 1917 is a destination. The wearable system can recognize surface 1917 (and polygon 1972b) by grouping POIs that appear to be on the same plane. The wearable system can also use other surface recognition techniques to identify surface 1917.
[0246] The user can also use multi-modal input to transport the virtual content 1126 to a bounding polygon 1972b on the surface 1917, as shown in second state 1970b. For example, the user can move the virtual content 1926 through a combination of head pose changes and movement of the user input device 1940.
[0247] As another example, a user may utter "move it there" via the wearable system's microphone 232, which may receive an audio stream and parse the command therefrom (as described herein). The user may combine this voice command with head pose, eye gaze, gesture, or totem actuation. The wearable system may detect virtual object 1926 as the target of the command because virtual object 1926 is the highest-confidence object (see, e.g., the dashed line in scene 1970a indicating the user's finger 1960, with the HMD 1920 and totem 1940 oriented toward object 1926). The wearable system may also identify the command action as "move" and determine that the parameter of the command is "there." The wearable system may further determine that "there" refers to bounding polygon 1972b based on input modes other than voice (e.g., eye gaze, head pose, gesture, totem).
[0248] Commands in an interaction event can involve the adjustment and calculation of multiple parameters. For example, parameters may include a destination, location, orientation, appearance (e.g., size or shape), or animation of a virtual object. The wearable system can automatically calculate parameters when changing them, even when direct input is not explicit. As an example, the wearable system can automatically change the orientation of virtual object 1926 when it is moved from vertical surface 1915 to horizontal surface 1917. In first state 1970a, virtual content 1926 is approximately vertically oriented on surface 1915. When virtual content 1926 is moved to surface 1917 in second state 1970b, the orientation of virtual content 1926 can be kept consistent (e.g., maintain a vertical orientation), as shown by virtual object 1924. The wearable system can also automatically adjust the orientation of the virtual content 1926 to match the orientation of the surface 1917 so that the virtual content 1926 appears to be in a horizontal position, as illustrated by the virtual object 1922. In this example, the orientation may be automatically adjusted based on the environment tracking 1632 as an indirect input. The wearable system can automatically consider the characteristics of the object (e.g., the surface 1917) when the wearable system determines that the object is a target destination object. The wearable system can adjust the parameters of the virtual object based on the characteristics of the target destination object. In this example, the wearable system automatically rotates the orientation of the virtual object 1926 based on the orientation of the surface 1917.
[0249] Additional examples of automatically placing or moving virtual objects are described in U.S. Application No. 15 / 673,135, filed August 9, 2017, entitled "AUTOMATIC PLACEMENT OF A VIRTUAL OBJECT IN A THREE-DIMENSIONAL SPACE," and published as U.S. Patent Publication No. 2018\0045963 (the disclosure of which is incorporated herein by reference in its entirety).
[0250] In some implementations, the input may explicitly modify multiple parameters. A voice command of "Place it flat there" may alter the orientation of virtual object 1926 in addition to identifying surface 1917 as the destination. In this example, the word "flat" and the word "there" can both be parameter values, with "there" causing the wearable system to update the location of the target virtual object, while the word "flat" is associated with the orientation of the target virtual object at the destination location. To implement the parameter "flat," the wearable system can match the orientation of virtual object 1926 to match the orientation of surface 1917.
[0251] In addition to, or as an alternative to, selecting and moving virtual objects, multimodal input can interact with virtual content in other ways. Figure 20 illustrates an example of automatically resizing a virtual object based on multimodal input. In Figure 20, a user 1510 can wear an HMD 1502 and can interact with virtual objects using hand gestures and voice commands 2024. Figure 20 illustrates four scenes 2000a, 2000b, 2000c, and 2000d. Each scene includes a display screen and a virtual object (illustrated by a smiley face).
[0252] In scene 2000a, the display screen has size 2010 and the virtual object has size 2030. The user can change their hand gesture from gesture 2020 to gesture 2022 to indicate that they want to adjust the size of either the virtual object or the display screen. The user can use voice input 2024 to indicate whether the virtual object or the display screen is the subject of manipulation.
[0253] As an example, a user may wish to enlarge both the display screen and the virtual object. Thus, the user may use input gesture 2022 as a command to enlarge. A parameter regarding the degree of expansion may be represented by the degree of spread fingers. Alternatively, the user may use voice input 2024 to determine the target of interaction. As shown in scene 2000b, the user may utter "all" to produce an enlarged display 2012 and an enlarged virtual object 2032. As another example, in scene 2000c, the user may utter "content" to produce an enlarged virtual object 2034, while the size of the display screen remains the same as in scene 2000a. As yet another example, in scene 2000d, the user may utter "display" to produce an enlarged display screen 2016, while the virtual object remains the same size as in scene 2000a.
[0254] (Example of indirect input as an input mode) As described herein, a wearable system can be programmed to allow a user to interact with direct and indirect user input as part of multimodal input. Direct user input may include head pose, eye gaze, voice input, gestures, input from a user input device, or other input directly from the user. Indirect input may include various environmental factors, such as, for example, the user's position, user characteristics / preferences, characteristics of an object, characteristics of the user's environment, etc.
[0255] As described with reference to FIGS. 2A and 2B , the wearable system may include a location sensor, such as a GPS, radar, or lidar. The wearable system may determine a target for a user's interaction as a function of the user's proximity to the object. FIG. 21 illustrates an example of identifying a target virtual object based on the object's location. FIG. 21 diagrammatically illustrates a bird's-eye view 2100 of a user's FOR. The FOR may include multiple virtual objects 2110a-2110q. The user may wear an HMD, which may include a location sensor. The wearable system may determine candidate target objects based on the user's proximity to the object. For example, the wearable system may select virtual objects within a threshold radius (e.g., 1 m, 2 m, 3 m, 5 m, 10 m, or greater) from the user as candidate target virtual objects. In FIG. 21 , virtual objects (e.g., virtual objects 2110o, 2110p, 2110q) fall within a threshold radius (illustrated by dashed circle 2122) from a user's position 2120. As a result, the wearable system can set virtual objects 2110o-2110q as candidate target virtual objects. The wearable system can further refine the selection based on other inputs (e.g., the user's head pose, etc.). The threshold radius may depend on contextual factors such as the user's location. For example, the threshold radius may be smaller if the user is in their office than if the user is outdoors in a park. Candidate objects can be selected from a portion of region 2122 within the threshold radius from the user. For example, only those objects within both circle 2122 and the user's FOV (e.g., generally in front of the user) can be candidates, while objects within circle 2122 but outside the user's FOV (e.g., behind the user) cannot be candidates. As another example, multiple virtual objects may be along a common line of sight, for example, cone casting may select multiple virtual objects.The wearable system can use the user's position as another input to determine target virtual objects or parameters for user interaction. For example, cone casting may select objects corresponding to different depth planes, while the wearable system may be configured to identify target virtual objects as objects within the user's reach.
[0256] Similar to direct input, indirect input may also be assigned a value that may be used to calculate a confidence score for a virtual object. For example, multiple subjects or parameters were within the general confidence of the selection, but the indirect input may also be used as a confidence factor. Referring to FIG. 21 , a virtual object within circle 2122 may have a higher confidence score than a virtual object between circles 2122 and 2124 because an object closer to user's position 2120 is more likely to be an object the user is interested in interacting with.
[0257] In the example shown in Figure 21, the dashed circles 2122, 2124 are conveniently illustrated to represent the projection of a sphere of corresponding radius onto the plane shown in Figure 21. This is for illustration purposes and not limitation. In other implementations, regions of other shapes (e.g., polygons) may be selected.
[0258] 22A and 22B illustrate another example of interacting with a user's environment based on a combination of direct and indirect input. These two figures show two virtual objects, virtual object A 2212 and virtual object B 2214, within the world camera's FOV 1270, which may be larger than the user's FOV 1250. Virtual object A 2212 is also within the user's FOV 1250. For example, virtual object A 2212 may be a virtual document the user is currently viewing, while virtual object B 2214 may be a virtual sticky note on the wall. However, while the user interacts with virtual object A 2212, the user may want to look at virtual object B 2214 to obtain additional information from virtual object B 2214. As a result, the user may turn their head to the right (changing the FOV 1250) to view virtual object B 2214. Advantageously, in some embodiments, rather than turning the head, the wearable system may detect a change in the user's gaze direction (towards virtual object B 2214), so that the wearable system can automatically move virtual object B 2214 within the user's FOV without the user having to change their head pose. Virtual object B may overlay virtual object A (or be contained within object A), or object B may be located within the user FOV 1250 but at least partially spaced apart from object A (so that object A is also at least partially visible to the user).
[0259] As another example, virtual object B2214 may be on another user interface screen. A user may desire to switch between a user interface screen having virtual object A2212 and a user interface screen having virtual object B2214. The wearable system may perform the switching without changing the user's FOV 1250. For example, in response to detecting a change in eye gaze or an actuation of a user input device, the wearable system may automatically move the user interface screen having virtual object A2212 out of the user's FOV 1250, while moving the user interface screen having virtual object B2214 into the user's FOV 1250. As another example, the wearable system may automatically overlay the user interface screen having virtual object B2214 on top of the user interface screen having virtual object A2212. Once the user provides an indication that they are finished interacting with the virtual user interface screen, the wearable system may automatically move the virtual user interface screen out of the FOV 1250.
[0260] Advantageously, in some embodiments, the wearable system can identify virtual object B2214 as a target virtual object to be moved within the FOV based on the multimodal input. For example, the wearable system can make the determination based on the user's eye line of sight and the location of the virtual object. The wearable system can set the target virtual object as the object that is in the user's line of sight and is closest to the user.
[0261] (Illustrative Process for Interacting with Virtual Objects Using Multimodal User Input) 23 illustrates an example process for interacting with a virtual object using multimodal input. Process 2300 can be performed by the wearable system described herein. For example, process 2300 may be performed by local processing and data module 260, remote processing module 270, and central runtime server 1650, alone or in combination.
[0262] In block 2310, the wearable system can optionally detect a start condition. The start can be a user-initiated input, which can provide an indication that the user intends to issue a command to the wearable system. The start condition can be predetermined by the wearable system. The start condition can be a single input or a combination of inputs. For example, the start condition can be a voice input, such as by uttering the phrase "Hey, Magic Leap." The start condition can also be gesture-based. For example, the wearable system can detect the presence of a start condition when it detects the user's hand within the world camera's FOV (or the user's FOV). As another example, the start condition can be a specific hand movement, such as a finger snap. The start condition can also be detected when the user actuates a user input device. For example, the user can click a button on the user input device to indicate that the user will issue a command. In some implementations, the start condition can be based on multimodal input. For example, both a voice command and a hand gesture may be required for the wearable system to detect the presence of a starting condition.
[0263] Block 2310 is optional. In some embodiments, the wearable system may receive and begin analysis of multimodal input without detecting a start condition. For example, when a user is watching a video, the wearable system may capture the user's multimodal input and adjust volume, fast-forward, rewind, skip to the next episode, etc., without requiring the user to first provide a start condition. Advantageously, in some embodiments, the user may not need to wake up the video screen before the user can interact with the video screen using the multimodal input (e.g., so that the video screen can present time or volume adjustment tools).
[0264] In block 2320, the wearable system can receive multimodal input for user interaction. The multimodal input may be direct or indirect input. Exemplary input modes may include voice, head pose, eye gaze, gesture (on the user input device or in the air), input on a user input device (e.g., a totem, etc.), properties of the user's environment, or objects (physical or virtual objects) in 3D space.
[0265] In block 2330, the wearable system can analyze the multimodal input and identify targets, commands, and parameters for user interaction. For example, the wearable system can assign confidence scores to candidate target virtual objects, target commands, and target parameters and select the target, command, and parameter based on the highest confidence score. In some embodiments, one input mode can be a primary input mode, while another input mode can be a secondary input mode. Input from the secondary input mode can complement the input from the primary input mode and confirm the target target, command, or parameter. For example, the wearable system can set head pose as the primary input mode and voice command as the secondary input mode. The wearable system can first interpret the input from the primary input mode as much as possible and then interpret additional input from the secondary input mode. If the additional input is interpreted to suggest a different interaction from that of the primary input, the wearable system can automatically provide a disambiguation prompt to the user. The disambiguation prompt can request the user to select a desired task from alternative options based on an interpretation of the primary input or an interpretation of the secondary input. Although the present embodiment is described with reference to a primary input mode and a secondary input mode, in various circumstances there may be more than two input modes, and the same techniques may also be applicable to a tertiary input mode, a fourth input mode, etc.
[0266] In block 2340, the wearable system can perform the user interaction based on the target, the command, and the parameter. For example, the multimodal input may include eye gaze and a voice command "put it there." The wearable system can determine that the target of the interaction is the object the user is currently interacting with, the command is "put it," and the parameter is the center of the user's field of fixation (determined based on the user's eye gaze direction). Thus, the user can move the virtual object the user is currently interacting with to the center of the user's field of fixation.
[0267] Example of Setting Direct Entry Mode Associated with User Interaction In some situations, such as when a user interacts with a wearable system using posture, gestures, or voice, there is a risk that other people in the user's vicinity may "hijack" the user's interaction by issuing commands using these direct inputs. For example, user A is standing near user B in a park. User A can interact with the HMD using voice commands. User B may hijack user A's experience by uttering "take a photo." This voice command issued by user B may cause user A's HMD to take a photo even if user A did not intend to take a photo. As another example, user B may perform a gesture within the FOV of the world camera of user A's HMD. This gesture may cause user A's HMD to go to a homepage, for example, while user A is playing a video game.
[0268] In some implementations, the input can be analyzed to determine whether the input originated from a user. For example, the system can apply speaker recognition techniques to determine whether the command "take a photo" was uttered by user A or hijacker B. The system may also apply computer vision techniques to determine whether a gesture was made by user A's hand or hijacker B's hand.
[0269] Additionally or alternatively, to prevent security breaches and disruption of a user's interaction with the wearable system, the wearable system can automatically configure available direct input modes based on indirect inputs or require multiple modes of direct input before a command is issued. FIG. 24 illustrates an example of configuring direct input modes associated with user interaction. Three direct inputs are illustrated in FIG. 24: voice 2412, head pose 2414, and hand gesture 2416. As explained further below, slider bars 2422, 2424, and 2426 represent the amount each input is weighted in determining a command. When the slider is at the far right, the input is given full weighting (e.g., 100%); when the slider is at the far left, the input is given zero weighting (e.g., 0%); and when the slider is between these extreme settings, the input is given partial weighting (e.g., 20% or 80% or some other intermediate value, such as a value between 0 and 1). In this example, the wearable system can be configured to require both a voice command 2422 and a hand gesture 2426 (while not using head pose 2414) before a command is executed. Thus, the wearable system cannot execute a command if the voice command 2422 and the gesture 2426 indicate different user interactions (or virtual objects). By requiring both types of input, the wearable system can reduce the likelihood that someone will hijack the user's interaction.
[0270] As another example, one or more input modes may be disabled. For example, when a user interacts with a document processing application, head pose 2414 may be disabled as an input mode, as shown in FIG. 24, where head pose slider 2424 is set to 0.
[0271] Each input may be associated with an authentication level. In FIG. 24 , speech 2412 is associated with authentication level 2422, head pose 2414 is associated with authentication level 2424, and hand gesture 2416 is associated with authentication level 2426. The authentication level may be used to determine whether the input is required for a command to be executed, whether the input is disabled, or whether the input is given partial weighting (between fully enabled and fully disabled). As illustrated in FIG. 24 , the authentication levels for speech 2412 and hand gesture 2416 are set to the far right (associated with the maximum authentication level), suggesting that these two inputs are required to issue a command. As another example, the authentication level for head pose is set to the far left (associated with the minimum authentication level). This suggests that head pose 2414 is not required to issue a command, even though it may still be used to determine a target virtual object or a target user interface action. In some circumstances, by setting the authentication level to minimum, the wearable system may disable head pose 2414 as an input mode.
[0272] In some implementations, the authentication level may also be used to calculate a confidence level associated with the virtual object. For example, the wearable system may assign higher values to input modes with higher authentication levels, while assigning lower values to input modes with lower authentication levels. As a result, when aggregating confidence scores from multiple input modes to calculate an aggregated confidence score for a virtual object, input modes with higher authentication levels may have a greater weight in the aggregated confidence score than input modes with lower authentication levels.
[0273] The authentication level can be set by the user (through input or via a settings panel) or can be set automatically by the wearable system, for example, based on indirect input. The wearable system may require more input modes when the user is in a public place, while requiring fewer input modes when the user is in a private place. For example, the wearable system may require both voice 2412 and hand gestures 2416 when the user is on the subway. However, when the user is at home, the wearable system may require only voice 2412 to issue commands. As another example, the wearable system may disable voice commands when the user is in a public park, thereby providing privacy to the user's interactions. However, voice commands may still be available when the user is at home.
[0274] Although these examples are described with reference to setting a direct input mode, similar techniques can also be applied to setting an indirect input mode as part of multimodal input. For example, when a user is using public transportation (e.g., a bus), the wearable system may be configured to disable geographic location as an input mode because the wearable system may not know exactly where the user is specifically sitting or standing on the public transportation.
[0275] (Additional Example User Experiences) In addition to the examples described herein, this section describes additional user experiences using multimodal input. As a first example, multimodal input can include voice input. For example, a user can issue a voice command such as "Hey, Magic Leap, call her," which is received by the audio sensor 232 on the HMD and analyzed by the HMD system. In this command, the user can say, "Hey, Magic Leap, call her," A user can initiate the task (or provide a starting condition) by uttering "Leap." "Call" can be a pre-programmed word, so the wearable system knows to make a phone call (rather than initiate a video call). In some implementations, these pre-programmed words may also be referred to as "hot words" or "carrier phrases," which the system recognizes as indicating that the user desires to perform a particular action (e.g., "call"), which may alert the system to receive further input to complete the desired action (e.g., identifying the person ("she") or phone number after the word "call"). The wearable system can use the additional input to identify who "she" is. For example, the wearable system can use eye tracking to verify the virtual contact list or contacts on the user's phone that the user is looking at. The wearable system can also use head pose or eye tracking to determine whether the user is looking directly at the person the user desires to call. In some embodiments, the wearable system can utilize facial recognition techniques (e.g., using the object recognizer 708) to determine the identity of the person the user is looking at.
[0276] As a second example, a user can place a virtual browser directly on a wall (e.g., the wearable system's display 220 can project the virtual browser as if it were overlaid on the wall). The user can reach out and provide a tap gesture on a link in the browser. Because the browser appears to be on the wall, the user may tap on the wall or tap in space, such that a projection of the user's finger appears tapping on the wall and provides an indication. The wearable system can use multimodal input to identify a link the user intends to click. For example, the wearable system can use gesture detection (e.g., via data obtained by the outward-facing imaging system 464), head pose-based cone casting, and eye gaze. In this example, gesture detection may be less than 100% accurate. The wearable system can use data obtained from head pose and eye gaze to refine gesture detection and increase the accuracy of gesture tracking. For example, the wearable system can identify the radius where the eyes are most likely focused based on data obtained by the inward-facing imaging system 462. In some embodiments, the wearable system can identify the user's field of fixation based on the eye gaze. The wearable system can also use indirect inputs, such as environmental features (e.g., wall locations, browser or web page characteristics, etc.), to improve gesture tracking. In this example, walls may be represented by a planar mesh (which may be pre-stored in the map of the environment 920), and the wearable system can determine the user's hand location against the planar mesh to determine the link the user targets and selects. Advantageously, in various embodiments, by combining multiple input modes, the accuracy required for one input mode for user interaction may be reduced compared to a single input mode.For example, the FOV camera may not need to have very high resolution for hand gesture recognition, as the wearable system may complement the hand gesture with head pose or eye gaze to determine the intended user interaction.
[0277] Although the multimodal input in the above examples includes audio input, audio input is not required for the multimodal input interactions described above. For example, a user can use a 2D-touch swipe gesture (e.g., on a totem) to move a browser window from one wall to a different wall. The browser may initially be on the left wall. The user can select the browser by activating the totem. The user can then look to the right wall and perform a right swipe gesture on the totem's touchpad. Swiping on a touchpad is unrestrained and imprecise because 2D swipes themselves do not easily / well translate into 3D movement. However, the wearable system can detect the wall (e.g., based on environmental data obtained by an outward-facing imaging system) and detect the specific point on the wall at which the user is looking (e.g., based on eye gaze). Using these three inputs (touch-swipe, gaze, and environmental features), the wearable system can smoothly place the browser wherever the user wants to move the browser window with high reliability.
[0278] Additional Examples of Head Pose as Multimodal Input In various embodiments, multimodal input can support totem-free experiences (or experiences in which totems are used less frequently). For example, multimodal input can include a combination of head pose and voice control, which can be used to share or search for virtual objects. Multimodal input can also use a combination of head pose and gestures to navigate various user interface planes and virtual objects within the user interface planes. A combination of head pose, voice, and gestures can be used to move objects, perform social networking activities (e.g., starting and conducting a telepresence session, sharing a post), browse information about a web page, or control a media player.
[0279] FIG. 25 illustrates an example of a user experience using multimodal input. In example scene 2500a, user 2510 can use head pose to target and select applications 2512 and 2514. The wearable system can display focus indicator 2524a and use head pose to indicate that the user is currently interacting with a virtual object. Once the user selects application 2514, the wearable system may display focus indicator 2524a for application 2514 (e.g., a target graphic as shown in FIG. 25, a halo around application 2514, or causing virtual object 2514 to appear closer to the user, etc.). The wearable system can also change the appearance of the focus indicator from focus indicator 2524a to focus indicator 2524b (e.g., an arrow graphic shown in scene 2500b) to indicate that interaction with user input device 466 also becomes available after the user selects virtual object 2514. Voice and gesture interaction also extends this interaction pattern to head pose plus hand gestures. For example, when a user issues a voice command, an application targeted using head pose may respond to or be manipulated by the voice command. Additional examples of virtual objects interacting with a combination of head pose, hand gestures, and voice recognition, for example, are described in U.S. Application No. 15 / 296,869, filed October 18, 2016, entitled "SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE," and published as U.S. Patent Publication No. 2017 / 0109936 (the disclosure of which is incorporated herein by reference in its entirety).
[0280] Head pose may be integrated with voice control, gesture recognition, and environmental information (e.g., mesh information) to provide hands-free browsing. For example, a voice command of "Search Fort Lauderdale" would be handled by a browser if the user targeted the browser using head pose. If the user did not target a specific browser, the wearable system could also handle this voice command without going through the browser. As another example, if the user uttered "Share this with Karen," the wearable system would perform a share action on the application the user targeted (e.g., using head pose, eye gaze, or gestures). As another example, voice control could perform browser window functions, such as "Go to bookmarks," while gestures could be used to perform basic web page navigation, such as clicking and scrolling.
[0281] Multimodal input can also be used to launch and move virtual objects without the need for a user input device. The wearable system can naturally place content near the user and the environment using multimodal input, such as gestures, voice, and gaze. For example, a user can use voice to open an unlaunched application when the user interacts with the HMD. The user can issue a voice command by saying, "Hey, Magic Leap, launch my browser." In this command, the initiation condition includes the presence of the invocation phrase "Hey, Magic Leap." The command can be interpreted as including "launch" or "open" (which may be interchangeable commands). The target of this command is an application name, e.g., "browser." However, this command does not require parameters. In some embodiments, the wearable system can automatically apply default parameters, such as placing a browser within the user's environment (or the user's FOV).
[0282] Multimodal input can also be used to implement basic browser controls, such as opening bookmarks, opening new tabs, navigating history, etc. The ability to browse web content in hands-free or hands-full multitasking scenarios can make users more informed and productive. For example, user Ada is a radiologist reading films in her office. While reading films, Ada can use voice and gestures to navigate the web and retrieve reference materials, reducing the need to move the mouse back and forth to switch between the film and reference materials on the screen. As another example, user Chris is cooking a new recipe from a virtual browser window. The virtual browser window can be placed on the cabinet. Chris can use voice commands to start shearing food while retrieving the bookmarked recipe.
[0283] 26 illustrates an exemplary user interface with various bookmarked applications. A user can select an application on user interface 2600 by uttering the name of the application. For example, a user can utter "open food" to launch the food application. As another example, a user can utter "open this." The wearable system can determine the user's gaze direction and identify applications on user interface 2600 that intersect with the user's gaze direction. The wearable system can then open the identified application.
[0284] The user can also issue a search command using voice. The search command can be implemented by the application the user is currently targeting. If the object does not currently support the search command, the wearable system may implement a search for the information within the wearable system's data storage or via a default application (e.g., via a browser, etc.). FIG. 27 illustrates an exemplary user interface 2700 when a search command is issued. The user interface 2700 shows both an email application and a media viewing application. The wearable system may determine (based on the user's head pose) that the user is currently interacting with the email application. As a result, the wearable system may automatically convert the user's voice command into a search command within the email application.
[0285] Multimodal input can also be used for media control. For example, a wearable system can use voice and gesture control to issue commands such as play, pause, mute, fast forward, and rewind to control a media player within an application (such as a screen). A user can use voice and gesture control for media applications and omit the totem.
[0286] Multimodal input can also be used in social networking contexts. For example, users can initiate conversations and share experiences (e.g., virtual images, documents, etc.) without using user input devices. As another example, a user can join a telepresence session and set up a private context so that the user may feel comfortable using voice to navigate the user interface.
[0287] Thus, in various implementations, the system may utilize multi-modal inputs such as head pose + voice (e.g., for information sharing and general application search), head pose + gesture (e.g., for navigation within an application), or head pose + voice + gesture (e.g., for "put it there" functionality, media player control, social interaction, or browser applications).
[0288] Additional Examples of Gesture Control as Part of Multimodal Input There may be two non-limiting and non-exclusive classes of gesture interactions: event gestures and dynamic hand tracking. Event gestures may be responsive to events while a user is interacting with the HMD, such as a catcher's signal to the pitcher in a baseball game or a "like" sign in a browser window to cause the wearable system to open a sharing dialog. The wearable system may follow one or more gesture patterns performed by the user and respond to the events accordingly. Dynamic hand tracking may involve tracking of the user's hands with low latency. For example, a user may move their hands across the user's FOV, and a virtual character may follow the movement of the user's fingers.
[0289] The quality of gesture tracking may depend on the type of user interaction. Quality may involve multiple factors, such as robustness, responsiveness, and ergonomics. In some embodiments, event gestures may be nearly perfectly robust. The threshold for minimum acceptable gesture performance may be slightly lower for social experiences, prototype interactions, and third-party applications because the aesthetics of these experiences may tolerate glitches, interruptions, low latency, etc., but gesture recognition can still remain highly performant and responsive in these experiences.
[0290] To increase the likelihood that the wearable system will respond to a user's gesture, the system can reduce or minimize latency for gesture detection (for both event gestures and dynamic hand tracking). For example, the wearable system can reduce or minimize latency by detecting when the user's hand is within the field of view of a depth sensor, automatically switching the depth sensor to an appropriate gesture mode, and then providing feedback to the user when a gesture can be performed.
[0291] As described herein, gestures can be used in combination with other input modes to launch, select, and navigate applications. Gestures can also be used to interact with virtual objects within an application by tapping in the air or on a surface (e.g., on a table or wall), scrolling, etc.
[0292] In one embodiment, the wearable system can implement a social networking tool that can support gestural interaction. Users can perform semantic event gestures to enrich their communications. For example, a user can wave their hand in front of the FOV camera, and a waving animation can be sent appropriately to the person with whom the user is chatting. The wearable system can also provide a virtualization of the user's hand using dynamic hand tracking. For example, a user can lift their hand in front of their FOV and get visual feedback that their hand is being tracked to animate their avatar's hand.
[0293] Hand gestures can also be used as part of multimodal input for media player control. For example, a user can use hand gestures to play or pause a video stream. A user can perform gesture operations away from the device (e.g., a television) playing the video. In response to detecting the user's gestures, the wearable system can remotely control the device based on the user's gestures. The user can also view a media panel, and the wearable system can use the user's hand gestures in combination with the user's gaze direction to update parameters of the media panel. For example, a pick (OK) gesture can suggest a "play" command, and a fist gesture can suggest a "pause" command. A user can also close a menu by waving one of their arms in front of the FOV camera. Examples of hand gestures 2080 are shown in FIG. 20.
[0294] Additional Examples of Interacting with Virtual Objects As described herein, the wearable system can support various multimodal interactions with objects (physical or virtual) in the user's environment. For example, the wearable system can support direct input for interacting with found objects, such as targeting, selecting, or controlling (e.g., movement or properties) the found objects. Interaction with found objects can also include interaction with found object geometry or interaction with found object connection surfaces.
[0295] Direct input is also supported for interaction with flat surfaces, such as targeting and selecting a wall or tabletop. A user can also initiate various user interface events, such as a touch event, a tap event, a swipe event, or a scroll event. A user can manipulate 2D user interface elements (e.g., a panel) using direct interactions, such as panel scrolling, swiping, and selecting elements (e.g., virtual objects or user interface elements such as buttons) within a panel. A user can also move or resize a panel using one or more direct inputs.
[0296] Direct input can also be used to manipulate objects at different depths. The wearable system can set various threshold distances (from the user) to determine the area of the virtual object. Referring to FIG. 21 , objects within dashed circle 2122 may be considered objects in the near distance, objects within dashed circle 2124 (but outside dashed circle 2122) may be considered objects in the middle distance, and objects outside dashed circle 2124 may be considered objects in the far distance. The threshold distance between the near and far distances may be, for example, 1 m, 2 m, 3 m, 4 m, 5 m, or more, and may depend on the environment (e.g., larger in an outdoor park than an indoor office block).
[0297] The wearable system can support various 2D or 3D manipulations of virtual objects within close range. Exemplary 2D manipulations may include moving or resizing. Exemplary 3D manipulations may include placing virtual objects in 3D space by pinching, drawing, moving, or rotating the virtual object, etc. The wearable system can also support interaction with virtual objects within a medium range, such as panning and repositioning objects within the user's environment, performing radial motion of objects, or moving objects within close or far ranges.
[0298] The wearable system can also support continuous fingertip interactions. For example, the wearable system can allow a user's finger to point like an attractor or pinpoint an object and perform a push interaction on the object. The wearable system can further support high-speed pose interactions, such as hand surface interaction or hand contour interaction.
[0299] Additional Examples of Voice Commands in Social Networking and Sharing Contexts The wearable system may support voice commands as input for social networking (or messaging) applications, for example, to share information with contacts or to call contacts.
[0300] As an example of initiating a call with a contact, a user can use a voice command such as "Hey Magic Leap, call Karen." In this command, "Hey Magic Leap" is the invocation phrase, the command is "call," and the parameter of the command is the name of the contact. The wearable system can automatically use a messenger application (as the target) to initiate the call. The command "call" may be associated with a task, such as "start a call with," "start a chat with," etc.
[0301] If the user says "initiate a call" and then a name, the wearable system can attempt to recognize the name. If the wearable system does not recognize the name, it can communicate a message to the user for the user to confirm the name or contact information. If the wearable system recognizes the name, it may present a dialog prompt that allows the user to confirm / reject (or cancel) the call or provide an alternative contact.
[0302] A user can also initiate a call with several contacts with a list of friends. For example, a user can utter, "Hey, Magic Leap, start a group chat with Karen, Cole, and Kojo." The group chat command may be derived from the phrase "start a group chat" or may be from a list of friends provided by the user. While a user is on a call, the user can add another user to the conversation. For example, a user can utter, "Hey, Magic Leap, invite Karen," and the phrase "invite" can be associated with the invite command.
[0303] The wearable system can share virtual objects with contacts using voice commands. For example, a user can utter, "Hey, Magic Leap, share my screen with Karen" or "Hey, Magic Leap, share it with David and Tony." In these examples, the word "share" is the share command. The word "screen" or "it" can refer to an object that the wearable system can determine based on multimodal input. Names such as "Karen," "David and Tony," etc. are parameters of the command. In some embodiments, when a voice command provided by a user includes an application reference and the word "share" with contacts, the wearable system may provide a confirmation dialog, asking the user to confirm whether the user wants to share the application itself or the object referenced through the application. When a user issues a voice command including the word "share," an application reference, and contacts, the wearable system can determine whether the application name is recognized by the wearable system or whether the application exists on the user's system. If the system does not recognize the name or the application does not exist in the user's system, the wearable system may provide a message to the user, which may suggest that the user try the voice command again.
[0304] If the user provides a deictic or anaphoric reference (e.g., "this" or "that") in a voice command, the wearable system can use multimodal input (e.g., the user's head pose) to determine whether the user is interacting with an object that can be shared. If the object cannot be shared, the wearable system may prompt the user with an error message or transition to a second input mode, such as gestures, to determine the object to be shared.
[0305] The wearable system can also determine whether the contact with which the object is to be shared can be recognized (e.g., as part of the user's contact list). If the wearable system recognizes the contact's name, the wearable system can provide a confirmation dialog and the user can confirm that they wish to proceed with sharing. If the user confirms, the virtual object can be shared. In some embodiments, the wearable system can share multiple virtual objects associated with an application. For example, the wearable system can share an entire album of photos or share the most recently viewed photo in response to the user's voice command. If the user declines the share, the share command is canceled. If the user indicates that the contact is incorrect, the wearable system may prompt the user to speak the contact's name again or to select the contact from a list of available contacts.
[0306] In one implementation, if a user utters "share" and an application reference but does not specify a contact, the wearable system may locally share the application with people in the user's environment who have access to the user's files. The wearable system may also reply and prompt the user to enter a name using one or more of the input modes described herein. Similar to social networking examples, the user can issue a voice command to share a virtual object with one contact or a group of contacts.
[0307] A challenge in making a call via voice is when a voice user interface incorrectly recognizes or fails to recognize a contact's name. This can be particularly problematic for less common or non-English names, such as Isi or Ileana. For example, when a user issues a voice command that includes a contact's name (such as "Share my screen with lly"), the wearable system may be unable to identify the name "lly" or its pronunciation. The wearable system may open a contact dialog with a prompt, such as "Who is that?" The user may again use voice to specify "Ily," use voice or a user input device to spell the name "ILY," or attempt to quickly select a name from a panel of available names using a user input device. The name "Ily" may be a nickname for Ileana, who has an entry in the user's contacts. Once the user tells the system that "Ily" is a nickname, the system may be configured to "remember" the nickname by automatically associating the nickname (or the pronunciation or audio pattern associated with the nickname) with the friend's name.
[0308] Additional Examples of Selecting and Moving Virtual Objects Using Voice Commands A user can naturally and quickly manage the location of virtual objects within their environment using multimodal input, such as a combination of eye gaze, gestures, and voice. For example, a user named Lindsay is seated at a table and ready to perform a task. She opens her laptop and starts a desktop monitor app on the computer. As the computer loads, she reaches out her hand above the laptop screen and says, "Hey, Magic Leap, put the monitor here." In response to this voice command, the wearable system can automatically bring up the monitor screens and place them above the laptop. However, when Lindsay looks across to a wall on the other side of the room and says, "Put the screen there," the wearable system can automatically place the screen on the wall across from her. Lindsay can also look at her desk and say, "Put Halcyon here." Halcyon is initially on the kitchen table, but in response to a voice command, the wearable system can automatically move it to the tabletop surface. As she works, she can use her totem to interact with these objects and adjust their scale to her liking.
[0309] A user can use voice to open an application that is not launched at any time within the user's environment. For example, a user can utter, "Hey, Magic Leap, launch a browser." In this command, "Hey, Magic Leap" is the invocation word, the word "launch" is the launch command, and the word "browser" is the target application. The "launch" command may be associated with the words "launch," "open," and "play." For example, the wearable system can still identify a launch command when a user utters "open browser." In one embodiment, the application may be an immersive application, which can provide a 3D virtual environment to the user as if the user were part of the 3D virtual environment. As a result, when an immersive application is launched, the user can be positioned as if they were present within the 3D virtual environment. In one implementation, the immersive application also includes a store application. When the store application is launched, the wearable system can provide a 3D shopping experience for the user, so that the user may feel as if they were shopping in a physical store. In contrast to an immersive application, an application may be a landscape application, which, when launched, may be placed relative to where it would be placed if launched via a totem in the launcher, so that the user can interact with the landscape application but may not feel like they are part of it.
[0310] A user can also use voice commands to launch a virtual application at a specified location within the user's FOV, or to move an already placed virtual application (e.g., a landscape application) to a specific location within the user's FOV. For example, a user can utter, "Hey, Magic Leap, put the browser here," "Hey, Magic Leap, put the browser there," "Hey, Magic Leap, put this here," or "Hey, Magic Leap, put that there." These voice commands include an invocation word, a placement command, an application name (which is the target), and a location cue (which is the parameter). The target may be referenced based on the audio data, for example, based on the name of the application spoken by the user. The target may also be identified based on head pose or eye gaze when the user utters the word "this" or "it" instead. To facilitate this voice interaction, the wearable system can, for example, make two inferences: (1) the application to launch and (2) the location where the application should be placed.
[0311] The wearable system can use the installation command and the application name to infer the application to launch. For example, if the user utters an application name that the wearable system does not recognize, the wearable system may provide an error message. If the user utters an application name that the wearable system recognizes, the wearable system can determine whether the application is already installed in the user's environment. If the application is already shown in the user's environment (e.g., within the user's FOV), the wearable system can determine the number of instances of the application (e.g., the number of open browser windows) present in the user's environment. If only one instance of the target application exists, the wearable system can move the application to a location specified by the user. If more than one instance of the uttered application exists in the environment, the wearable system can move all instances of the application to a specified location or the most recently used instance to a specified location. If the virtual application is not already installed in the user's environment, the system can determine whether the application is a landscape application, an immersive application, or a store application (where the user can download or purchase other applications). If the application is a landscape application, the wearable system can launch a virtual application at a specified location. If the application is an immersive application, the wearable system can place a shortcut to the application at a specified location because immersive applications do not support the ability to launch at a specified location within the user's FOV.If the application is a store application, the system may place a mini-store at a specified location because the store application may require full 3D immersion of the user in the virtual world and therefore does not support launching at a specific location within the user's environment. The mini-store may include an overview or icon of the virtual objects in the store.
[0312] The wearable system can use various inputs to determine where to place an application. The wearable system can analyze syntax in the user's command (e.g., "here" or "there"), determine the intersection of head pose-based ray casting (or cone casting) with virtual objects in the user's environment, determine the user's hand position, determine a planar surface mesh or an environment plane mesh (e.g., a mesh associated with a wall or table), etc. As an example, if the user utters "here," the wearable system can determine the user's hand gesture, such as whether an outstretched hand is within the user's FOV. The wearable system can place an object in a rendering plane near the position of the user's outstretched hand and within the reach of the user's hand. If the outstretched hand is not within the FOV, the wearable system can determine whether the head pose (e.g., the direction of the head pose-based cone casting) intersects with a surface plane mesh within the user's arm's reach. If a surface plane mesh exists, the wearable system can place the virtual object at the intersection of the head pose direction and the surface plane mesh in a rendering plane within the user's arm's reach. The user can place the object flat on the surface. If a surface plane mesh does not exist, the wearable system can place the virtual object in a rendering plane that has a distance between the user's arm's reach and the optimal reading distance. If the user utters "there," the wearable system can perform a similar operation as when the user utters "here," but if a surface plane mesh does not exist within the user's arm's reach, the wearable system may place the virtual object in a rendering plane within a medium distance.
[0313] Once the user utters "Place application...", the wearable system can immediately provide predictive feedback to the user, indicating where the virtual object will be placed based on available input if the user utters either "here" or "there". This feedback can be in the form of a focus indicator. For example, the feedback may include a hand, mesh, or planar surface that intersects with the user's head pose direction in a rendering plane within the user's arm's reach, with a small floating text bubble uttering "here". The planar surface may be located in the near distance if the user's command is "here", while it may be located in the middle or far distance if the user's command is "there". This feedback may be visualized as a shadow or outline of a visual object.
[0314] The user can also cancel an interaction, which may be canceled in two ways in various cases: (1) failure to complete the command due to an n-second timeout, or (2) by entering a cancel command, such as uttering "no," "never mind," or "cancel."
[0315] (Example of interaction with text using a combination of user inputs) Freeform text entry in mixed reality environments, especially for long string sequences using traditional interaction modalities, can be problematic. As an example, particularly in "hands-free" environments lacking input or interface devices such as keyboards, handheld controllers (e.g., totems), or mice, systems that rely entirely on automatic speech recognition (ASR) can have difficulty using text editing (e.g., to correct ASR errors inherent in the speech recognition technology itself, such as incorrect transcription of a user's speech). As another example, a virtual keyboard in a "hands-free" environment can require sophisticated user control and can be tiring if used as the primary form of user input.
[0316] The wearable system 200 described herein can be programmed to allow a user to naturally and quickly interact with virtual text using multimodal input, such as a combination of two or more of voice, eye gaze, gesture, head posture, totem input, etc. The term "text," as used herein, can include type, characters, words, phrases, sentences, paragraphs, or other types of free-form text. Text can also include graphics or animations, such as emojis, ideograms, emoticons, smileys, symbols, etc. Interaction with virtual text can include, alone or in combination, composing, selecting (e.g., selecting some or all of the text), or editing (e.g., modifying, copying, cutting, pasting, deleting, clearing, undoing, redoing, inserting, replacing, etc.) the text. By utilizing a combination of user inputs, the systems described herein offer significant improvements in speed and convenience over single-input systems.
[0317] The multimodal text interaction techniques described herein can be applied in any dictation scenario or application (e.g., the system simply transcribes user utterances rather than applying any semantic evaluation, even if the transcription is part of another task that relies on semantic evaluation). Some example applications can include messaging applications, word processing applications, gaming applications, system configuration applications, etc. Example use cases include a user writing a text message to be sent to a contact who may or may not be in the user's contact list; a user writing a letter, article, or other text content; a user posting and sharing content on a social media platform; and a user completing or otherwise filling out a form using the wearable system 200.
[0318] Systems utilizing a combination of user inputs need not be wearable systems: as desired, such systems may be any suitable computing system, such as a desktop computer, laptop, tablet, smartphone, or another computing device having multiple user input channels, such as a keyboard, trackpad, microphone, eye or gaze tracking system, gesture recognition system, etc.
[0319] (Example of Composing Text Using Multimodal User Input) 28A-28F illustrate an exemplary user experience of composing and editing text based on a combination of inputs, such as voice commands or eye gaze. As described herein, the wearable system can determine the user's gaze direction based on images obtained by the inward-facing imaging system 462 shown in FIG. 4. The inward-facing imaging system 462 may determine the orientation of one or both of the user's pupils and may extrapolate the user's line of sight for one or both eyes or multiple eyes. By determining the user's line of sight for both eyes, the wearable system 200 can determine the three-dimensional location in space where the user is looking.
[0320] The wearable system can also determine voice commands based on data obtained from the audio sensor 232 (e.g., a microphone) shown in Figure 2A. The system may have an automatic speech recognition (ASR) engine that converts the spoken input 2800 to text. The speech recognition engine may use natural language understanding in converting the spoken input 2800 to text, including isolating and extracting the message text from the longer utterance.
[0321] As shown in FIG. 28A, the audio sensor 232 can receive a phrase 2800 spoken by a user. As shown in FIG. 28A, the phrase 2800 may include a command, such as "Send a message to John Smith saying...," and parameters of the command, such as composing and sending the message, and the destination of the message as John Smith. The phrase 2800 may also include the content of the message to be composed. In this example, the message content may include "I'm boarding a flight from Boston and will arrive around 7:00. Meet me on the corner near my office." Such content can be obtained by analyzing the audio data using an ASR engine (which can implement natural language understanding and isolate and extract message content and punctuation (e.g., ".") from the user's utterance). In some examples, the punctuation may be processed for presentation in the context of the transcribed string (e.g., "2 o'clock" may be presented as "2:00," or a "question mark" may be presented as "?"). The wearable system can also tokenize the text string, such as by isolating discrete words within the text string, and display the results, such as by displaying the discrete words in a mixed reality environment.
[0322] However, automatic speech recognition can be prone to errors in some situations. As illustrated in FIG. 28B , a system using an ASR engine may produce results that do not precisely match a user's spoken input for a variety of reasons, including poor or idiosyncratic pronunciation, environmental noise, homonyms and other similar-sounding words, hesitation or stuttering, and vocabulary not within the ASR's lexicon (e.g., foreign words, technical terms, jargon, slang, etc.). In the example of FIG. 28B , the system properly interpreted the command aspect of phrase 2800 and generated a message with header 2802 and body 2804. However, in the message body 2804, the system incorrectly interpreted the user's utterance of "corner" as the somewhat similar-sounding "quarter." A system that relies entirely on voice input may have difficulty quickly substituting the intended word or phrase for the misrecognized word (or phrase) for the user. However, the wearable system 200 described herein can advantageously allow a user to quickly correct the error, as illustrated in FIGS. 28C-28F.
[0323] The ASR engine in the wearable system may produce text results including at least one word associated with the user's utterance and may produce an ASR score associated with each word (or phrase) in the text results. A high ASR score may indicate a high confidence or likelihood that the ASR engine correctly transcribed the user's utterance into text, while a low ASR score may indicate a low confidence or low likelihood that the ASR engine correctly transcribed the user's utterance into text. In some embodiments, the system may display words with low ASR scores (e.g., ASR scores below an ASR threshold) in an emphasized manner (e.g., with a background highlight, italic or bold font, a different color font, etc.), which may make it easier for a user to identify or select an incorrectly recognized word. A low ASR score for a word may indicate that there is a reasonable likelihood that the ASR engine misrecognized the word and therefore that the user is more likely to select the word for editing or replacement.
[0324] 28C and 28D, the wearable system may allow the user to select the misrecognized word (or phrase) using an eye-tracking system, such as inward-facing imaging system 462 of Figure 4. In this example, the selected word may be an example of a target virtual object described above with reference to the preceding figures.
[0325] The wearable system 200 can determine the gaze direction based on the inward-facing imaging system 462 and can cast a cone 2806 or a ray of light in the gaze direction. The wearable system can select one or more words that intersect with the user's gaze direction. In some implementations, a word may be selected when the user's gaze dwells on an incorrect word for at least a threshold time. As described above, an incorrect word may be determined, at least in part, by being associated with a low ASR score. The threshold time may be any amount of time sufficient to indicate that the user desires to select a particular word, but is long enough that it does not unnecessarily delay the selection. The threshold time may also be used to determine a confidence score indicating that the user desires to select a particular virtual word. For example, the wearable system can calculate a confidence score based on how long the user gazes in a certain direction / object, and the confidence score may increase as the time duration for looking in a certain direction / object increases. The confidence score may also be calculated based on multimodal input, as described herein. For example, the wearable system may determine if both the user's hand gestures and eye gaze indicate that a word should be selected, with a higher confidence score (than a confidence score derived from eye gaze alone).
[0326] As another example, the wearable system may calculate a confidence score based in part on the ASR score, which may indicate the relative confidence of the ASR engine in transcribing a particular word, as discussed in more detail herein. For example, a low ASR engine score may indicate that the ASR engine has relatively low confidence that it correctly transcribed the spoken word. Thus, there may be a higher probability that the user will likely select that word for editing or replacement. If the user's gaze lingers on a word with a low ASR score for longer than a threshold time, the system may assign a higher confidence score to reflect that the user selected that word for at least two reasons: first, the length of the eye gaze on the word; and second, the fact that the word was likely incorrectly transcribed by the ASR engine (both of which tend to indicate that the user will want to edit or replace the word).
[0327] A word may be selected if the confidence score reaches a threshold criterion. As an example, the threshold time may be 0.5 seconds, 1 second, 1.5 seconds, 2 seconds, 2.5 seconds, 1-2 seconds, 1-3 seconds, etc. Thus, a user can easily and quickly select the incorrect word, "quarter," simply by looking for a sufficient amount of time. A word may be selected based on a combination of eye gaze (or gesture) time and an ASR score above the ASR threshold, both criteria providing an indication that the user intended to select that particular word.
[0328] As an example, if the results of an ASR engine include a first word with a high ASR score (e.g., a word that the ASR engine is relatively confident that it recognized correctly) and a second word with a low ASR score (e.g., a word that the ASR engine is relatively confident that it recognized incorrectly), and these two words are displayed adjacent to each other by the wearable system, the wearable system may assume that a user's gaze input encompassing both the first and second words is actually an attempt by the user to select the second word based on its relatively low ASR score, because the user is more likely to want to edit the incorrectly recognized second word than the correctly recognized first word. In this way, a word produced by the ASR engine with a low ASR score that is incorrect and more likely to require editing may be significantly easier for the user to select for editing, and thus may encourage editing by the user.
[0329] While this example describes using eye gaze to select a misrecognized word, other multimodal inputs can also be used to select a word. For example, cone casting can identify multiple words such as "around," "7:00," "the," and "quarter" because they also intersect with portions of the virtual cone 2806. As will be further described with reference to FIGS. 29-31 , the wearable system can combine eye gaze input with another input (e.g., gestures, voice commands, or input from the user input device 466) to select the word "quarter" as the word for further editing.
[0330] In response to selecting word 2808, the system may enable editing of the selected word. The wearable system may allow the user to edit the word using various techniques, such as, for example, change, cut, copy, paste, delete, clear, undo, redo, insert, replace, etc. As shown in FIG. 28D , the wearable system may allow the user to change word 2808 to another word. The wearable system may support various user inputs for editing word 2808, such as by receiving additional spoken input through a microphone to replace or delete the selected word, displaying a virtual keyboard to allow the user to type a replacement word, or receiving user input via a user input device. In some implementations, the input may be associated with a specific type of text editing. For example, a hand waving gesture may be associated with deleting selected text, while a gesture involving pointing with a finger at a location in the text may cause the wearable system to insert additional text at that location. The wearable system may also support combinations of user inputs to edit words. As will be further explained with reference to Figures 32-35, the system can support eye gaze in combination with other input modes to edit words.
[0331] 28D and 28E, the system may automatically present the user with an array of suggested alternatives, such as alternatives 2810a and 2810b, in response to selection of word 2808. The suggested alternatives may be generated by an ASR engine or other language processing engine within the system and may be based on original speech input (which may also be referred to in some embodiments as voice input), natural language understanding, context learned from user behavior, or other suitable sources. In at least some embodiments, the suggested alternatives may be alternative hypotheses generated by the ASR engine, hypotheses generated by a predictive text engine (which may attempt to “fill in the blanks” using the context of adjacent words and the user's historical patterns of text), homophones of the original translation, generated using a thesaurus, or generated using other suitable techniques. In the illustrated example, the suggested alternatives for “quarter” include “corner” and “courter,” which may be provided by the language engine as words sounding similar to “quarter.”
[0332] FIG. 28E illustrates how a system may allow a user to select a desired alternative word, such as “corner,” using eye gaze. The wearable system may select an alternative word using a similar technique, such as that described with reference to FIG. 28C. For example, the system may use an inward-facing imaging system 462 to track the user's eyes and determine that the user's gaze 2812 has been focused on a particular alternative, i.e., alternative 2810A or “corner,” for at least a threshold time. After determining that the user's gaze 2812 has been focused on an alternative for the threshold time, the system may revise the text (message) by replacing the originally selected word with the selected alternative word 2814, as shown in FIG. 28F. In some implementations where the wearable system selects words using cone casting, the wearable system can dynamically adjust the size of the cone based on the density of the text. For example, the wearable system may present a cone with a larger opening (and therefore a larger surface area away from the user) to select an alternative word for editing, as shown in Figure 28E, because there are several available options. However, the wearable system may present a cone with a smaller opening to select word 2808 in Figure 28C because word 2808 is surrounded by other words; a smaller cone may reduce the error rate of accidentally selecting another word.
[0333] The wearable system can provide feedback (e.g., visual, auditory, tactile, etc.) to the user throughout the course of the action. For example, the wearable system can present a focus indicator to facilitate user recognition of the target virtual object. For example, as shown in FIG. 28E, the wearable system can provide a contrasting background 2830 around the word "quarter" to indicate that the word "quarter" has been selected and that the user is currently editing the word "quarter." As another example, as shown in FIG. 28F, the wearable system can change the font of the word "corner" 2814 (e.g., to a bold font) to indicate that the wearable system has confirmed the replacement of the word "quarter" with the alternate word "corner." In other implementations, the focus indicator can include crosshairs, a circle or oval surrounding the selected text, or other graphical techniques to highlight or emphasize the selected text.
[0334] Example of Word Selection Using Multimodal User Input The wearable system can be configured to support and utilize multiple modes of user input to select words. Figures 29-31 illustrate an example of selecting a word based on a combination of eye gaze and another input mode. However, in other examples, input other than eye gaze can also be used in combination with another input mode to interact with text.
[0335] FIG. 29 illustrates an example of selecting a word based on input from a user input device and a line of sight. As shown in FIG. 29 , the system may combine the user's line of sight 2900 (which may be determined based on data from the inward-facing imaging system 462) with user input received via the user input device 466. In this example, the wearable system may perform cone casting based on the user's gaze direction. The wearable system may confirm the selection of the word "quarter" based on input from the user input device. For example, the wearable system may identify the word "quarter" as the word closest to the user's gaze direction, and the wearable system may confirm that the word "quarter" was selected based on the user's actuation of the user input device 466. As another example, cone casting may capture multiple words, such as "around," "7:00," "the," and "quarter." The user may select a word from the multiple words for further editing via the user input device 466. By receiving input independent of the user's gaze, the system may not have to wait long before confidently identifying a particular word as one the user desires to edit. After selecting and thus editing a word, the system may present alternatives (as discussed in connection with FIG. 28E) or otherwise allow the user to edit the selected word. The same process of combining user gaze and user input received via the totem may be applied to selecting a desired replacement word (e.g., selecting the word "corner" from among the alternatives to replace the word "quarter"). Some implementations may utilize a confidence score to determine the text being selected by the user. The confidence score may aggregate multiple input modalities and provide a better determination of the selected text.For example, the confidence score may be based on how long the user gazed at the text, whether the user actuated the user input device 466 while gazed at the text, whether the user pointed toward the selected text, etc. If the confidence score reaches a threshold, the wearable system can determine, with increased confidence, that the system correctly selected the user's desired text. For example, to select text using only eye gaze, the system may be configured to select text if the gaze time exceeds 1.5 seconds. However, if the user only gazes at the text for 0.5 seconds but simultaneously actuates the user input device, the system can more quickly and confidently determine the selected text, which may improve the user experience.
[0336] FIG. 30 illustrates an example of selecting a word for editing based on a combination of voice and gaze input. The wearable system can determine a target virtual object based on the user's gaze. As shown in FIG. 30, the system may determine that the user's gaze 3000 is directed toward a particular word (in this case, "quarter"). The wearable system can also determine an action to be performed on the target virtual object based on the user's voice command. For example, the wearable system may receive the user's spoken input 3010 via the audio sensor 232, recognize the spoken input 3010 as a command, combine the two user inputs into a command, and apply the command action ("edit") to the target virtual object (e.g., the word ("quarter") on which the user is focusing their gaze). As described above, the system may present alternative words after the user selects a word for editing. The same process of combining the user's gaze and spoken input may be applied to selecting a desired replacement word from among the alternative words to replace the word "quarter." As described herein, terms such as "edit" represent context-specific wake-up words that serve to invoke a constrained system command library associated with editing for each of one or more different user input modalities. That is, such terms, when received as spoken input by the system, can cause the system to evaluate subsequently received user input against a limited set of references to recognize edit-related commands provided by the user with increased accuracy. For example, in the context of spoken input, the system can consult a limited, command-specific vocabulary of terms and perform speech recognition on the subsequently received spoken input. In another example, in the context of gaze or gesture input, the system can consult a limited, command-specific library of template images and perform image recognition on the subsequently received gaze or gesture input.Terms like "edit" are sometimes referred to as "hot words" or "carrier phrases," and the system may include several pre-programmed (and optionally user-configurable) hot words such as edit (in the editing context), cut, copy, paste, bold, italic, delete, move, etc.
[0337] 31 illustrates an example of selecting a word for editing based on a combination of gaze and gesture input. As illustrated in the example of FIG. 31, the system may select a word for editing using gesture input 3110 in conjunction with eye gaze input 3100. In particular, the system may determine eye gaze input 3100 (e.g., based on data obtained by inward-facing imaging system 462) and identify gesture input 3110 (e.g., based on images obtained by outward-facing imaging system 464). An object recognizer, such as recognizer 708, may be used to detect a part of the user's body, such as their hand, making a gesture associated with identifying the word for editing.
[0338] Gestures may be used alone or in combination with eye gaze to select a word. For example, cone casting may capture multiple words, but the wearable system may nevertheless identify the word “quarter” as the target virtual object because it is identified from both the cone casting and the user's hand gestures (e.g., a confidence score based on the eye gaze cone casting plus the hand gestures exceeds a confidence threshold indicating that the user selected the word “quarter”). As another example, cone casting may capture multiple words, but the wearable system may nevertheless identify the word “quarter” as the target virtual object because it is identified from the cone casting and is the word with the lowest ASR score from the ASR engine that is within (or near) the cone casting. In some implementations, gestures may be associated with commands, such as “edit” or other hot words described herein, and thus may be associated with command actions. As an example, the system may recognize when users are pointing at the same word with their gaze and interpret these user inputs as a request to edit the same word. If desired, the system may also utilize additional user input, such as a voice command for "edit," at the same time as determining that the user wishes to edit a particular word.
[0339] Example of Editing Words Using Multimodal User Input Once a user selects a word for editing, the system can utilize any desired mode of user input to edit the selected word. The wearable system can allow the user to modify or replace the selected word by displaying a list of potential alternatives and receiving user gaze input 2812 to select an alternative word to replace the original word (see the example illustrated in FIG. 28E). Figures 32-34 illustrate additional examples of editing a selected word, where the selected word can be edited using multimodal input.
[0340] FIG. 32 illustrates an example of replacing a word based on a combination of eye gaze and speech input. In FIG. 32, the system receives speech input 3210 from a user (through audio sensor 232 or other suitable sensor). The speech input 3210 may contain a desired replacement word (which may or may not be a replacement word from the list of suggested replacements 3200). In response to receiving the speech input 3210, the wearable system can analyze the input (e.g., to extract a carrier phrase such as "change this to..."), identify the word spoken by the user, and replace the word "corner" as spoken by the user with the selected word "quarter." In this example, the replacement is a word, but in some implementations, the wearable system can be configured to replace the word "quarter" with a phrase or sentence or some other element (e.g., an emoji). In embodiments where multiple words are contained within the eye gaze cone casting, the wearable system may automatically select the word within the eye gaze cone that is closest to the replacement word (e.g., "quarter" is closer to "corner" than "the" or "7:00").
[0341] FIG. 33 illustrates an example of word modification based on a combination of voice and gaze input. In this example, the wearable system can receive speech input 3310 and determine a user's gaze direction 3300. As shown in FIG. 33, the speech input 3310 includes the phrase "change it to 'corner'." The wearable system can analyze the speech input 3310 and determine that the speech input 3310 includes a command action "modify" (which is an example of a carrier phrase), a target "it," and parameters of the command (e.g., the resulting word "corner"). This speech input 3310 can be combined with the eye gaze 3300 to determine the target of the action. As described with reference to FIGS. 28A and 28B, the wearable system can identify the word "quarter" as the target of the action. Thus, the wearable system can modify the target ("quarter") to the resulting word "corner."
[0342] FIG. 34 illustrates an example of editing a selected word 3400 using a virtual keyboard 3410. The virtual keyboard 3410 can be controlled by user gaze input, gesture input, input received from a user input device, etc. For example, a user may type a replacement word by moving their eye gaze direction 3420 across the virtual keyboard 3410 displayed to the user by the display of the wearable system 200. The user may type each letter in the replacement word by pausing their gaze over individual keys for a threshold time period, or the wearable system may recognize a change in the direction of the user's gaze 3420 over a particular key as an indication that the user desires to select that key (thereby eliminating the need for the user to hold their focus stationary on each individual key as they type the word). As described with reference to FIG. 28D, in some implementations, the wearable system may vary the size of the cone based on the size of the key. For example, in a virtual keyboard 3410 where the size of each key is relatively small, the wearable system may reduce the size of the cone to allow the user to more accurately identify letters in a replacement word (so that the cone casting will not accidentally capture multiple possible keys). If the size is relatively large, the wearable system may increase the size of the keys appropriately so that the user does not need to precisely direct their gaze (which may reduce fatigue).
[0343] In some implementations, after a word is selected, the wearable system can present a set of possible actions in addition to, or as an alternative to, displaying a list of suggested alternative words for replacing the selected word. The user 210 can select an action and edit the selected word using techniques described herein. FIG. 35 illustrates an exemplary user interface displaying possible actions to apply to the selected word. In FIG. 35, in response to selecting a word 3500 for editing, the wearable system may present a list 3510 of options for editing, including (in this example) options to: (1) change the word (using any of the techniques described herein for editing), (2) cut the word and, optionally, store it in a clipboard, or copy the word and store it in a clipboard, or (3) paste the word or phrase from the clipboard. Additional or alternative options that may be presented include an option to delete the selection, an option to redo, an option to undo, an option to select all, an option to insert here, and an option to replace. The various options may be selected using gaze input, totem input, gesture input, etc., as described herein.
[0344] Example of Interacting with Phrases Using Multimodal User Input While the preceding examples have described steps for selecting and editing words using multimodal input, this is intended for illustrative purposes, and the same or similar processes and inputs may be used in selecting and editing phrases or sentences or paragraphs, generally containing multiple words or characters.
[0345] 36(i)-36(iii) illustrate examples of interacting with words and phrases using multimodal input. In FIG. 36(i), the wearable system can determine the user's gaze 3600 direction and perform cone casting based on the user's gaze direction. In FIG. 36(ii), the system can recognize that the user's 210 gaze 3600 is focused on a first word 3610 (e.g., "I'm"). The system may make such a determination of the first word 3610 using any of the techniques discussed herein, including, but not limited to, recognizing that the user's gaze 3600 rests (e.g., dwells) on a particular word for a threshold time period, that the user is providing voice, gesture, or totem input simultaneously with the user's gaze 3600 being on a particular word, etc. The wearable system may also display a focus indicator (e.g., a contrasting background, as shown) over the selected word "I'm" 3610 to indicate that the word was determined from eye gaze cone casting. While viewing the first word 3610, the user may activate a totem 3620 (an example of a user input device 466). This activation may indicate that the user intends to select a phrase or sentence starting with the first word 3610.
[0346] In Figure 36(iii), after actuation of user input device 466, the user can view the last intended word (e.g., the word "there") and indicate that the user desires to select a phrase starting with the word "I'm" and ending with the word "there." The wearable system can also detect that the user has stopped actuating totem 3620 (e.g., the user releases a previously pressed button), thus selecting the entire range 3630 of the phrase "I'm flying in from Boston and will be there." The system can indicate the selected phrase using a focus indicator (e.g., by extending a contrasting background to all words in the phrase).
[0347] The system may use various techniques to determine that a user desires to select a phrase rather than another word for editing. As an example, the system may determine that a user desires to select a phrase rather than undo their selection of the first word when the user selects a second word shortly after selecting a first word. As another example, the system may determine that a user desires to select a phrase when the user selects a second word that appears after the first word, and that the user has not yet edited the first selected word. As yet another example, a user may press a button on the totem 3620 when their gaze is focused on the first word 3610 and then hold the button until their gaze settles on the last word. If the system recognizes that the button was pressed when the gaze 3610 was focused on the first word, but released only after the user's gaze 3610 shifted to the second word, the system may recognize the multimodal user input as a selection of a phrase. The system may then identify all of the words in the phrase, including the first word, the last word, and all words in between, and may allow editing of the phrase as a whole. The system may use a focus indicator to highlight the selected phrase so that it stands out from unselected text (e.g., highlight, emphasized text (e.g., bold or italic or a different color), etc.). The system may then contextually display appropriate options for editing the selected phrase, such as options 3510, a virtual keyboard such as keyboard 3410, alternative phrases, etc. The system may receive additional user input, such as spoken input, totem input, gesture input, etc., to determine how to edit the selected phrase 3630.
[0348] 36 illustrates a user selecting the first word 3610 at the beginning of a phrase, but the system may also allow the user to select backward from the first word 3610. In other words, a user may select a phrase by selecting the last word of the phrase (e.g., "there") and then the first word of the desired phrase (e.g., "I'm").
[0349] FIGS. 37A-37B illustrate another example of interacting with text using multimodal input. In FIG. 37A, the user 210 utters a sentence (“I want to sleep”). The wearable system can capture the user's utterance as speech input 3700. For this speech input, the wearable system can display both primary and secondary results from an automatic speech recognition (ASR) engine, on a word-by-word basis, as shown in FIG. 37B. The primary results for each word may represent the ASR engine's best guess for the word spoken by the user in the speech input 3700 (e.g., the word with the highest ASR score for representing the word actually spoken by the user), while the secondary results may represent similar-sounding alternatives or words with lower ASR scores than the ASR engine's best guess. In this FIG. 37B, the primary results are displayed as a sequence 3752. In some embodiments, the wearable system may present alternative results or hypotheses as alternative phrases and / or entire sentences, as opposed to alternative words. As an example, the wearable system may provide the primary result "four score and seven years ago" along with the secondary result "force caring seven years to go," where there is no one-to-one correspondence between the discrete words in the primary and secondary results. In such an embodiment, the wearable system may support input from the user (in any of the manners described herein) to select alternative or secondary phrases and / or sentences.
[0350] As shown in Figure 37B, each word from a user's speech input 3700 may be displayed as a collection of primary and secondary results 3710, 3720, 3730, 3740. This type of arrangement may allow a user to quickly swap between incorrect primary results and any correct errors introduced by the ASR engine. Primary results 3752 may be highlighted with a focus indicator (e.g., each word is bold text surrounded by a bounding box in the example in Figure 37B) to distinguish them from secondary results.
[0351] The user 210 can pause on a secondary result, such as a secondary word, phrase, or sentence, if the primary word is not the one intended by the user. As an example, the primary result of an ASR engine in collection 3740 may be "slip," while the correct transcript is actually the first secondary result, "sleep." To correct this error, the user can focus their gaze on the correct secondary result, "sleep," and the system may recognize that the user's gaze remains on the secondary result for a threshold time period. The system may interpret user gaze input as a request to replace the selected secondary result, "sleep," with the primary result, "slip." Additional user input may be received in conjunction with selecting a desired secondary result, such as user speech input (e.g., the user may ask the system to "edit," "use this," or "replace" while viewing the desired secondary result).
[0352] Once the user has finished editing the phrase "I want to sleep" or has verified that the transcript is correct, the phrase can be added to the body of text using any mode of user input described herein. For example, the user can utter a hot word such as "end" to cause the edited phrase to be added back to the body of text.
[0353] (An exemplary process for interacting with text using a combination of user inputs) 38 is a process flow diagram of an example method 3800 of interacting with text using multiple modes of user input. The process 3800 can be implemented by the wearable system 200 described herein.
[0354] In block 3810, the wearable system may receive spoken input from the user. The spoken input may include a user utterance containing one or more words. In one embodiment, the user may dictate a message, and the wearable system may receive the dictated message. This may be accomplished through any suitable input device, such as the audio sensor 232.
[0355] In block 3820, the wearable system may convert the spoken input to text. The wearable system may utilize an automatic speech recognition (ASR) engine to convert the user's spoken input into text (e.g., a written transcript) and may further utilize natural language processing techniques to convert such text into semantic representations that indicate intent and concepts. The ASR engine may be optimized for free-form text input.
[0356] In block 3830, the wearable system may tokenize the text into discrete actionable elements, such as words, phrases, or sentences. The wearable system may also display the text for the user using a display system, such as display 220. In some embodiments, the wearable system does not need to understand the meaning of the text during the tokenization process. In other embodiments, the wearable system is equipped with the capability to understand the meaning of the text (e.g., one or more natural language processing models or other probabilistic-statistical models), or simply the ability to distinguish between (i) words, phrases, and sentences that represent a user-configured message or portion thereof, and (ii) words, phrases, and sentences that do not represent a user-configured message or portion thereof, but instead correspond to commands to be executed by the wearable system. For example, the wearable system may need to grasp the meaning of the text to recognize command actions or command parameters spoken by the user. Examples of such text may include context-specific wake-up words, also referred to herein as hot words, that serve to invoke one or more constrained system command libraries associated with editing for one or more different user input modalities.
[0357] A user can interact with one or more of the actionable elements using multimodal user input. In block 3840, the wearable system can select one or more elements in response to a first indication. The first indication can be a user input or a combination of user inputs, as described herein. The wearable system may receive input from the user selecting one or more of the elements of the text string for editing. The user may select a single word or multiple words (e.g., a phrase or a sentence). The wearable system may receive user input selecting the elements for editing in any desired form, including, but not limited to, speech input, gaze input (e.g., via the inward-facing imaging system 462), gesture input (e.g., as captured by the outward-facing imaging system 464), totem input (e.g., via actuation of the user input device 466), or any combination thereof. As an example, the wearable system may receive user input in the form of a user's gaze lingering on a particular word for a threshold time period, or may receive user input via a microphone or totem simultaneously with the user's gaze on a particular word indicating selection of that particular word for editing.
[0358] In block 3850, the wearable system can edit the selected element in response to the second indication. The second indication can be received via a single input mode or a combination of input modes, as described with the preceding figures, including, but not limited to, user gaze input, spoken input, gesture input, and totem input. The wearable system may receive user input indicating how the selected element should be edited. The wearable system may edit the selected element according to the user input received in block 3850. For example, the wearable system can replace the selected element based on spoken input. The wearable system can also present a list of suggested alternatives and select from among the selected alternatives based on the user's eye gaze. The wearable system can also receive input via user interaction with a virtual keyboard or via a user input device 466 (e.g., a physical keyboard or a handheld device, etc.).
[0359] In block 3860, the wearable system may display the results of editing the selected element. In some implementations, the wearable system may provide a focus indicator on the element being edited.
[0360] As indicated by arrow 3870, the wearable system may repeat blocks 3840, 3850, and 3860 if the user provides additional user input to edit additional elements of the text.
[0361] Additional details related to multimodal task execution and text editing for wearable systems are provided in U.S. Patent Application No. 15 / 955,204, filed April 17, 2018, entitled "MULTIMODAL TASK EXECUTION AND TEXT EDITING FOR A WEARABLE SYSTEM," and published as U.S. Patent Publication No. 2018 / 0307303, which is incorporated herein by reference in its entirety.
[0362] (Example of Transmode Input Fusion Technique) As described above, transmode input fusion techniques that provide dynamic selection of an appropriate input mode can advantageously enable users to target real or virtual objects more accurately and confidently, providing a more robust and user-friendly AR / MR / VR experience.
[0363] The wearable system can advantageously support the timely fusion of multiple user input modes to facilitate user interaction within a three-dimensional (3D) environment. The system can detect when a user provides two or more inputs via two or more separate input modes that may converge together. As an example, a user may point at a virtual object with their finger while also directing their eye gaze toward the virtual object. The wearable system can detect this convergence of the eye gaze and finger gesture input and apply timely fusion of the eye gaze and finger gesture input, thereby determining the virtual object the user is pointing at with greater accuracy and / or speed. The system thus enables users to select smaller elements (or more rapidly moving elements) by reducing the uncertainty of the primary input targeting method. The system can also be used to accelerate and simplify element selection. The system can enable users to improve the success rate of targeting moving elements. The system can be used to accelerate rich rendering of display elements. The system can be used to prioritize and accelerate local (and cloud) processing of object point cloud data, dense meshing, and plane acquisition, and to improve the inset fidelity of found entities and surfaces along with grasped objects of interest. Embodiments of the transmode technique described herein enable the system to establish variable transmode focus from the user's viewpoint while still preserving subtle movements of the head, eyes, and hands, thereby significantly improving the system's understanding of the user's intent.
[0364] As will be further described herein, identification of a transmodal state may be performed through an analysis of the relative convergence of some or all of the available input vectors. For example, this may be achieved by considering the angular distance between pairs of targeting vectors (e.g., the vector from the user's eyes to the target object and the vector pointing from the totem held by the user toward the target object). The relative variance of each pair of inputs may then be considered. If the distance or variance is below a threshold, a bimodal state (e.g., bimodal input convergence) may be associated with the pair of inputs. If the triplet of inputs has targeting vectors with angular distances or variances below a threshold, a trimodal state may be associated with the triplet of inputs. Convergence of four or more inputs is also possible. In an embodiment in which a triplet of head pose targeting vector (head gaze), eye vergence targeting vector (eye gaze), and tracked controller or tracked hand (hand pointer) targeting vector is identified, the triplet may be referred to as a transmodal triangle. The relative size of the triangle's sides, the triangle's area, or its associated variance may present characteristic features that the system may use to predict targeting and activation intent. For example, if the area of the transmodal triangle is less than a threshold area, a trimodal state may be associated with a triplet of input. Examples of vergence calculations are provided herein and in Appendix A. As an example of predictive targeting, the system may recognize that a user's eye and head inputs tend to converge prior to the convergence of the user's hand inputs. For example, when attempting to grasp an object, eye and head movements may be implemented rapidly and may converge on the object just before the hand movement inputs converge (e.g., approximately 200 ms). The system can detect head-eye convergence and predict that the hand movement inputs will converge shortly thereafter.
[0365] Convergence of pairs (bi-mode), triplet (tri-mode), or quadruplet (quad-mode) of targeted vectors, or a greater number of inputs (e.g., 5, 6, 7, or more) can be used to further define subtypes of trans-mode coordination. In at least some embodiments, the desired fusion method is identified based on the detailed trans-mode state of the system. The desired fusion method may also be determined, at least in part, by the trans-mode type (e.g., the converged inputs), the motion type (e.g., the way the converged inputs are moving), and the interaction field type (e.g., the interaction field into which the inputs are focused, such as the mid-range region, task space, and workspace interaction region described in connection with Figures 44A, 44B, and 44C). In at least some embodiments, the selected fusion method may determine which of the available input modes (and associated input vectors) are selectively converged. The motion type and field type may determine the settings of the selected fusion method, such as the relative weighting or filtering of one or more of the converged inputs.
[0366] Additional advantages and examples of techniques related to transmode input and the timely fusion of multiple user input modes for interacting with virtual objects are further described with reference to Figures 39A-60B (and below and Appendix A).
[0367] (An explanation of some transformer mode terminology) A description of certain terms used for the transmode input fusion technique is provided below. These descriptions are intended to illustrate the transmode terminology, not to limit its scope. The transmode terminology should be understood from the perspective of one skilled in the art in light of the entire description set forth in the specification, claims, and accompanying figures.
[0368] The term "IP region" may include a volume associated with an interaction point (IP). For example, a pinch IP region may include a volume (e.g., a sphere) created by the posed separation of the index fingertip and thumb fingertip. The term "region of interest (ROI)" may include a volume constructed from overlapping uncertainty regions (e.g., volumes) associated with a set of targeting vectors for an intended target object. The ROI may represent a volume within which the intended target object is likely to be found.
[0369] The term "modal input" may refer to input from any of the sensors of a wearable system. For example, common modal input includes input from six degrees of freedom (6DOF) sensors (e.g., for head pose or totem position or orientation), eye-tracking cameras, microphones (e.g., for voice commands), outward-facing cameras (for hand or body gestures), etc. Transmodal input may refer to multiple, simultaneously used, dynamically combined modal inputs.
[0370] The term "vergence" may include the convergence of multiple input vectors associated with multiple user input modes on a common interaction point (e.g., when both eye gaze and hand gestures point to the same spatial location). The term "fixation" may include a localized deceleration and pausing of a vergence point or a single input vector point. The term "fixation" may include a fixation that extends in time for at least a given duration. The term "eye tracking" may include target (e.g., projectile-like) movement of a vergence point toward a target or other object. The term "smooth tracking" may include smooth (e.g., low acceleration or low jerk) movement of a vergence point toward a target or other object.
[0371] The term "sensor convergence" may include convergence of sensor data (e.g., convergence of data from gyroscopes and accelerometers forming an inertial measurement unit (IMU), convergence of data from an IMU with a camera, convergence of multiple cameras for a SLAM process, etc.). The term "feature convergence" may include spatial convergence of inputs (e.g., convergence of input vectors from multiple input modes) and temporal convergence of inputs (e.g., parallel or sequential timing of multiple input modes).
[0372] The term "bimodal" convergence may include the convergence of two input modes (e.g., the convergence of two input modes as input vectors converge onto a common interaction point). The terms "trimodal" and "quadmodal" convergence may include the convergence of three and four input modes, respectively. The term "transmodal" convergence may include the transitional convergence of multiple inputs (e.g., the detection of the temporary convergence of multiple user input modes and the corresponding integration or fusion of those user inputs to improve the overall input experience).
[0373] The term "divergence" may refer to at least one input mode that was previously converged (and consequently merged) with at least one other input mode, but is no longer converged with the other input modes (e.g., tri-modal divergence may refer to a transition from a tri-modal convergent state to a bi-modal convergent state as an initially convergent third input vector diverges from the converged first and second input vectors).
[0374] The term "head-hand-vergence" may include the convergence of head and hand ray casting vectors or the convergence of head pose and hand interaction points. The term "head-eye vergence" may include the convergence of head pose and eye gaze vectors. The term "head-eye-hand vergence" may include the convergence of head pose, eye gaze, and hand direction input vectors.
[0375] The term "passive transmodal intent" may include pre-targeting, targeting, head-eye fixation, and fixation. The term "active transmodal intent" may include head-eye-hand fixation interactions or head-eye-hand manipulation interactions. The term "transmodal triangle" may include the area created by the convergence of three modal input vectors (and may refer to the area of uncertainty of trimodal convergent inputs). This area may also be referred to as the vergence area or modal vergence area. The term "transmodal quadrangle" may include the area created by the convergence of four modal input vectors.
[0376] (Example of user input) 39A and 39B illustrate examples of user input received through controller buttons or input areas on a user input device. In particular, FIGS. 39A and 39B illustrate a controller 3900 that may be part of a wearable system disclosed herein and may include a home button 3902, a trigger 3904, a bumper 3906, and a touchpad 3908. The user input device 466 or totem 1516, described with reference to FIGS. 4 and 15, respectively, can serve as the controller 3900 in various embodiments of the wearable system 200.
[0377] Potential user inputs that may be received through the controller 3900 include, but are not limited to, pressing and releasing the home button 3902, partially and fully (and other partial) pressing of the trigger 3904, releasing the trigger 3904, pressing and releasing the bumper 3906, touching, moving while touching, releasing the touch, increasing or decreasing the pressure applied to the touch, touching a specific portion such as the edge of the touchpad 3908, or performing a gesture on the touchpad 3908 (e.g., by drawing a shape with your thumb).
[0378] FIG. 39C illustrates an example of user input received through physical movement of a controller or head-mounted device (HMD). As shown in FIG. 39C, physical movement of the controller 3900 and head-mounted display 3910 (HMD) can form user input into the system. The HMD 3910 can comprise the head-mounted components 220, 230 shown in FIG. 2A or the head-mounted wearable component 58 shown in FIG. 2B. In some embodiments, the controller 3900 provides three degrees of freedom (3DOF) input by recognizing rotation of the controller 3900 in any direction. In other embodiments, the controller 3900 provides six degrees of freedom (6DOF) input by also recognizing translation of the controller in any direction. In still other embodiments, the controller 3900 may provide input with less than 6DOF or less than 3DOF. Similarly, the head-mounted display 3910 may recognize and receive input with 3DOF, 6DOF, less than 6DOF, or less than 3DOF.
[0379] FIG. 39D illustrates an example of how user inputs can have different durations. As shown in FIG. 39D, a user input may have a short duration (e.g., a duration less than a fraction of a second, such as 0.25 seconds) or a long duration (e.g., a duration greater than a fraction of a second, such as more than 0.25 seconds). In at least some embodiments, the duration of the input itself may be recognized and utilized as an input by the system. Short and long duration inputs can be treated differently by the wearable system 200. For example, a short duration input may represent a selection of an object, while a long duration input may represent an activation of the object (e.g., causing an app associated with the object to run).
[0380] 40A, 40B, 41A, 41B, and 41C illustrate various examples of user input that may be received and recognized by the system. User input may be received through one or more user input modes (individually, as shown, or in combination). User input may include input through controller buttons such as the home button 3902, trigger 3904, bumper 3906, and touchpad 3908, physical movement of the controller 3900 or HMD 3910, eye gaze direction, head pose direction, gestures, voice input, etc.
[0381] 40A, a short press and release of the home button 3902 may indicate a home tap action, while a long press of the home button 3902 may indicate a home press and hold action. Similarly, a short press and release of the trigger 3904 or bumper 3906 may indicate a trigger tap action or a bumper tap action, respectively, while a long press of the trigger 3904 or bumper 3906 may indicate a trigger and hold action or a bumper and hold action, respectively.
[0382] As shown in FIG. 40B , a touch on the touchpad 3908 that moves across the touchpad may indicate a touch-drag action. A short touch and release of the touchpad 3908, with substantially no touch movement, may indicate a light tap action. If such a short touch and release of the touchpad 3908 occurs with a force above a certain threshold level (which may be a predetermined threshold, a dynamically determined threshold, a learned threshold, or some combination thereof), the input may indicate a forceful tap input. A touch on the touchpad 3908 with a force above the threshold level may indicate a forceful press action, while a long touch with such force may indicate a forceful press and hold input. A touch near the edge of the touchpad 3908 may indicate an edge press action. In some embodiments, an edge press action may also involve an edge touch above a threshold level of pressure. FIG. 40B also shows that a touch on the touchpad 3908 that moves in an arc may indicate a touch-circular action.
[0383] The example of FIG. 41A illustrates that interaction with (e.g., by moving a thumb across the touchpad) or physical movement (6 DOF) of the touchpad 3908 of the controller 3900 can be used to rotate a virtual object (e.g., by making a circular gesture on the touchpad), move a virtual object in the z-direction toward or away from the user (e.g., by making a gesture on the touchpad, e.g., in the y-direction), and increase or decrease the size of a virtual object (e.g., by making a gesture on the touchpad in a different direction, e.g., in the x-direction).
[0384] FIG. 41A also shows that combinations of inputs can represent actions. In particular, FIG. 41 illustrates that interacting with bumper 3906 and a user turning and tilting their head (e.g., adjusting their head pose) can indicate initiating and / or terminating manipulation actions. As an example, a user may provide an indication of initiating manipulation of an object by double-tapping or holding bumper 3906, then move the object by providing additional input, and then provide an indication of ending manipulation of the object by double-tapping or releasing bumper 3906. In at least some embodiments, a user may provide additional input to move an object in the form of physical movement of controller 3900 (6 or 3 DOF) or by adjusting their head pose (e.g., tilting and / or rotating their head).
[0385] FIG. 41B illustrates additional examples of user interactions. In at least some embodiments, the interaction in FIG. 41B involves two-dimensional (2D) content. In still other embodiments, the interaction in FIG. 41B may be used for three-dimensional content. As shown in FIG. 41B, a head pose (which may be directed at 2D content) combined with a touch movement on the touchpad 3908 may indicate a set selection action or a scrolling action. A head pose combined with a forceful press on the edge of the touchpad 3908 may indicate a scrolling action. A head pose combined with a light and short tap on the touchpad 3908, or a short tap with pressure on the touchpad 3908, may indicate an active action. A forceful press and holding a forceful press on the touchpad 3908 combined with a head pose (which may be a specific head pose) may indicate a context menu action.
[0386] 41C, the wearable device can use a head pose indicating that the user's head is pointing toward a virtual app along with a home tap action to open a menu associated with the app, or a head pose along with a home press and hold action to open a launcher application (e.g., an app that allows multiple apps to run). In some embodiments, the wearable device can use a home tap action (e.g., a single or double tap on the home button 3902) to open a launcher application associated with a pre-targeted application.
[0387] 42A, 42B, and 42C illustrate examples of user input in the form of fine finger gestures and hand movements. The user inputs illustrated in FIGS. 42A, 42B, and 42C may sometimes be referred to herein as microgestures and may take the form of fine finger movements such as pinching the thumb and index finger together, pointing with a single finger, grasping with an open and closed hand, pointing with the thumb, tapping with the thumb, etc. The microgestures may be detected by the wearable system using, as one example, a camera system. In particular, the microgestures may be detected using one or more cameras (which may include a pair of cameras in a stereo configuration) that may be part of the outward-facing imaging system 464 (shown in FIG. 4). The object recognizer 708 can analyze images from the outward-facing imaging system 464 and recognize the example microgestures shown in FIGS. 42A-42C. In some implementations, a microgesture is activated by the system when the system determines that the user has focused on the target object for a sufficiently long fixation or dwell time (e.g., the convergence of multiple input modes is deemed robust).
[0388] (Examples of perceptual fields, display rendering planes, and interaction areas) 43A illustrates the visual and auditory perceptual fields of a wearable system. As shown in FIG. 43A, a user may have a primary field of view (FOV) and a peripheral FOV within their field of view. Similarly, a user may sense direction within the auditory perceptual field, including at least forward, backward, and peripheral directions.
[0389] FIG. 43B illustrates the display rendering planes of a wearable system having multiple depth planes. In the example of FIG. 43B, the wearable system has at least two rendering planes, one displaying virtual content at a depth of approximately 1.0 meter and the other displaying virtual content at a depth of approximately 3.0 meters. The wearable system may display virtual content at a given virtual depth on the depth plane that has the closest display depth. FIG. 43B also illustrates a 50-degree field of view for this exemplary wearable system. In addition, FIG. 43B illustrates a near clipping plane at approximately 0.3 meters and a far clipping plane at approximately 4.0 meters. Virtual content closer than the near clipping plane may be clipped (e.g., not displayed) or shifted away from the user (e.g., to at least the distance of the near clipping plane). Similarly, virtual content farther from the user than the far clipping plane may be clipped or shifted toward the user (e.g., to at least the distance of the far clipping plane).
[0390] 44A, 44B, and 44C illustrate examples of different interaction regions around a user, including a mid-distance region, an extended workspace, a task space, a manipulation space, a probing space, and a head space. These interaction regions represent spatial regions within which a user can interact with real and virtual objects; the type of interaction may differ within different regions, and the appropriate set of sensors used for transmodal fusion may differ within different regions. F...
Claims
1. 1. A method, comprising: Under the control of the wearable system's hardware processor, accessing sensor data from a plurality of sensors in different modes; identifying a convergence event of sensor data from a first sensor of the plurality of sensors and sensor data from a second sensor of the plurality of sensors; selectively applying a filter to the sensor data from the first sensor during the convergence event; including: Selectively applying the filter to the sensor data from the first sensor during the convergence event includes: detecting convergence of a vector from the first sensor and a vector from the second sensor based on the sensor data from the first and second sensors; applying the filter to the sensor data from the first sensor based on the convergence; detecting a divergence of the vector from the first sensor and the vector from the second sensor based on the sensor data from the first and second sensors; Disabling application of the filter to the sensor data from the first sensor based on the divergence; and A method comprising:
2. The method of claim 1 , wherein the filter comprises a low-pass filter with an adaptive cutoff frequency.
3. 10. The method of claim 1, further comprising targeting an object within a three-dimensional (3D) environment around the wearable system by utilizing first sensor data from the first sensor and second sensor data from the second sensor.
4. 2. The method of claim 1 , wherein the first sensor includes an electromyography (EMG) sensor that senses hand movement during movement, and the second sensor includes a camera-based hand gesture sensor, and identifying the sensor data convergence event includes using the EMG sensor to determine that a user's muscles are flexing in a manner consistent with a non-verbal symbol, and using the camera-based hand gesture sensor to determine that at least a portion of the user's hand is positioned in a manner consistent with the non-verbal symbol.
5. a plurality of sensors of different modes; Hardware processor and A wearable system comprising: The hardware processor includes: accessing sensor data from the plurality of sensors; identifying a convergence event of sensor data from a first sensor of the plurality of sensors and sensor data from a second sensor of the plurality of sensors; selectively applying a filter to the sensor data from the first sensor during the convergence event; It is programmed to To selectively apply the filter to the sensor data from the first sensor during the convergence event, the hardware processor: detecting convergence of a vector from the first sensor and a vector from the second sensor based on the sensor data from the first and second sensors; applying the filter to the sensor data from the first sensor based on the convergence; detecting a divergence of the vector from the first sensor and the vector from the second sensor based on the sensor data from the first and second sensors; Disabling application of the filter to the sensor data from the first sensor based on the divergence; and A wearable system that is programmed to:
6. The wearable system of claim 5 , wherein the filter comprises a low-pass filter with an adaptive cutoff frequency.
7. The hardware processor includes:
6. The wearable system of claim 5, programmed to target objects within a three-dimensional (3D) environment surrounding the wearable system by utilizing first sensor data from the first sensor and second sensor data from the second sensor.
8. the first sensor includes an electromyography (EMG) sensor that senses hand movement during movement; the second sensor includes a camera-based hand gesture sensor; To identify a convergence event in the sensor data, the hardware processor: using the EMG sensor to determine that the user's muscles are flexing in a manner consistent with non-verbal symbols; determining, with the camera-based hand gesture sensor, that at least a portion of the user's hand is positioned in a manner consistent with the non-verbal symbol; The wearable system of claim 5 , programmed to:
Citation Information
Patent Citations
User-directed personal information assistant
JP2015135674A