Integration of Virtual Reality Interaction Modes

The method and system automatically adapt virtual reality interaction modes based on context, addressing user frustration by switching between modes like no-hands 3-DOF, no-hands 6-DOF, gaze-gesture, and ray-casting, enhancing usability and efficiency in virtual reality systems.

JP7745575B2Active Publication Date: 2025-09-29META PLATFORMS TECHNOLOGIES LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022579982
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-29
Filing Date
2021-06-24
Publication Date
2025-09-29
Estimated Expiration
2041-06-24

AI Technical Summary

Technical Problem

Existing virtual reality systems lack the ability to automatically transition between interaction modes based on context, leading to user frustration and inefficiency, as current modes are often inappropriate for the current situation, requiring manual selection or failing to adapt to changes in user availability or environmental conditions.

Method used

A method and system that automatically transitions between interaction modes in a virtual reality environment by identifying context-specific conditions, such as hand tracking availability and lighting, to switch between no-hands 3-DOF, no-hands 6-DOF, gaze-gesture, and ray-casting modes, using a combination of head, hand, and eye tracking, and providing visual affordances like gaze cursors.

Benefits of technology

Enhances user experience by providing intuitive and efficient interaction modes that adapt to user context, improving usability and reducing the need for manual mode switching, especially in situations where traditional modes are inadequate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007745575000001
    Figure 0007745575000001
  • Figure 0007745575000002
    Figure 0007745575000002
  • Figure 0007745575000003
    Figure 0007745575000003
Patent Text Reader

Abstract

Aspects of the present disclosure are directed to an interaction mode system that provides multiple interaction modes in a virtual reality environment and has automatic context-specific transitions between the interaction modes. The interaction modes can specify how the interaction mode system determines orientation and movement within the virtual reality environment and interacts to make selections or perform other actions. In some implementations, the interaction mode system can control at least four interaction modes, including a no-hands 3DoF mode, a no-hands 6DoF mode, a gaze-and-gesture mode, and a ray-casting mode. The interaction mode system can control transitions between specific interaction modes using a mapping of interaction mode context elements (e.g., enabled components, mode settings, lighting or other environmental conditions, current body position, etc.) to the interaction modes. The interaction mode system can also provide affordances for signaling the current interaction mode and interaction mode transitions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure is directed to a virtual reality environment that can switch between different interaction modes depending on the context of the virtual reality device. [Background technology]

[0002] Various virtual reality systems can be embodied in wearable devices, systems with external displays and sensors, pass-through systems, etc. Many interaction modes have been created for these virtual reality systems, each of which typically provides a user with a way to command a direction and / or position and take an action relative to the commanded direction and / or position. For example, some interaction modes allow the user to move in three degrees of freedom (“3DoF”—typically allowing command of direction in the pitch, roll, and yaw axes, but no movement in the X, Y, or Z axes), while other interaction modes allow movement in six degrees of freedom (“6DoF”—typically allowing movement in the pitch, roll, yaw, X, Y, and Z axes). Some interaction modes use controllers, others track hand positions and gestures, and still others track only head movement. Interaction modes may be based on motion sensors, various types of cameras, time-of-flight sensors, timers, etc. However, despite all the options and possible combinations available for interaction modes, many virtual reality systems use interaction modes that are only suited to certain situations and are difficult to use or inappropriate for others. For example, users are frustrated when using a head-tracking-only interaction mode when conducting complex interactions. Furthermore, users often do not want to use a controller or hand-tracking input when conducting simple interactions or in public settings. On the other hand, some systems use multiple interaction modes that often require manual user selection to switch modes or do not automatically switch to an interaction mode appropriate for the current context. Summary of the Invention

[0003] Accordingly, the present invention is directed to a method, a computer-readable storage medium and a computing system according to the appended claims.

[0004] In an aspect, the present invention is directed to a method for automatically transitioning between interaction modes to interpret user input in a virtual reality system, the method including: identifying a first interaction mode context indicating that user position tracking input is unavailable; enabling a no-hands 3-DOF interaction mode in response to identifying the first interaction mode context; identifying a second interaction mode context indicating that user position tracking input is available and that hand tracking input is unavailable or that the first tracked hand pose is not consistent with a hand-ready state; enabling a no-hands 6-DOF interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that the second tracked hand pose is consistent with a hand-ready state; enabling a gaze-gesture interaction mode in response to identifying the third interaction mode context; identifying a fourth interaction mode context indicating that the third tracked hand pose is consistent with a ray state; and enabling a ray-casting interaction mode in response to identifying the fourth interaction mode context.

[0005] In an embodiment of the method according to the present invention, the step of enabling at least one of the plurality of interaction modes may further be responsive to an interaction mode change trigger that occurs periodically or in response to a change in interaction mode context identified as a result of monitoring the interaction mode context elements.

[0006] In an embodiment of the method according to the present invention, the method may further comprise the step of transitioning to an interaction mode upon receiving a user command to change from a current interaction mode.

[0007] In an embodiment of a method according to the present invention, the second tracked hand pose consistent with a hand ready state may include a hand pose identified as the user's palm pointing up by at least a threshold amount.

[0008] In an embodiment of the method according to the present invention, the third tracked hand pose consistent with the lighting conditions may include a hand pose identified as the user's palm pointing down by at least a threshold amount.

[0009] In an embodiment of the method according to the present invention, the no-hands three degrees of freedom interaction mode may include receiving a directional user instruction based at least on a determined orientation of the user's head, and receiving a user instruction of an action for one of the directional user instructions based on a dwell timer, wherein while in the no-hands three degrees of freedom interaction mode, user movements in the X, Y and Z axes are not automatically translated into movements of the field of view in the X, Y and Z axes in the virtual reality environment.

[0010] In an embodiment of the method according to the present invention, the no-hands six degrees of freedom interaction mode may include steps of receiving a directional user instruction based at least on a determined orientation of the user's head, receiving a user instruction of an action for one of a plurality of directional user instructions based on a dwell timer, and translating the user's movements in the X, Y and Z axes into movements of a field of view in the X, Y and Z axes in the virtual reality environment.

[0011] In an embodiment of the method according to the present invention, the gaze-gesture interaction mode may include receiving a user instruction of a direction based at least on a determined orientation of the user's head and based on a determined position of the user's head relative to the virtual reality environment, the determined position of the user's head being based on tracked movements of the user along the X-axis, Y-axis and Z-axis; and receiving a user instruction of an action for one of a plurality of user instructions of direction by tracking the user's hand posture and matching the hand posture with a specified action.

[0012] In an embodiment of the method according to the present invention, the ray casting interaction mode may include steps of receiving a user indication of a direction based at least on a ray cast by the virtual reality system, the ray being cast from a location relative to a tracked position of at least one of the user's hands, and receiving a user indication of an action relative to the ray by tracking a posture of the user's hand and matching the posture of the hand with a specified action.

[0013] In embodiments of the method according to the present invention, the no-hands 3-DOF interaction mode, the no-hands 6-DOF interaction mode, and the gaze-gesture interaction mode can each provide a visual affordance including a gaze cursor, which is presented in a field of view for the user, positioned at least in part based on the tracked position of the user's head.

[0014] In embodiments of the method according to the present invention, the no-hands 3 DOF interaction mode, the no-hands 6 DOF interaction mode, and the gaze and gesture interaction mode can each provide a visual affordance that includes a gaze cursor, and the gaze cursor is presented in the field of view for the user that is positioned based at least in part on the tracked position of the user's head and based on the user's tracked gaze direction.

[0015] In one aspect, the present invention is further directed to a computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform the above-described method, i.e., a process for transitioning between interaction modes to interpret user input in a virtual reality system, the process including: identifying a first interaction mode context indicating that hand tracking input is unavailable or that a first tracked hand pose is not consistent with a hand-ready state; enabling a no-hands interaction mode in response to identifying the first interaction mode context; identifying a second interaction mode context indicating that a second tracked hand pose is consistent with a hand-ready state; enabling a gaze-gesture interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that a third tracked hand pose is consistent with a ray state; and enabling a ray-casting interaction mode in response to identifying the third interaction mode context.

[0016] In an embodiment of the computer-readable storage medium, the no-hands interaction mode may include receiving a directional user instruction based at least on a determined orientation of the user's head, and receiving a user instruction of an action for one of the directional user instructions based on a dwell timer, wherein if user position tracking input is unavailable, the no-hands interaction mode does not translate user movement in the X, Y, and Z axes into X, Y, and Z field of view movement in the artificial reality environment, and when user position tracking input is available, the no-hands interaction mode automatically translates user movement in the X, Y, and Z axes into X, Y, and Z field of view movement in the artificial reality environment.

[0017] In an embodiment of the computer-readable storage medium, the gaze-gesture interaction mode can provide a visual affordance including a sphere presented in the user's field of view positioned between the user's thumb and one of the other fingers, which can also be presented as resized or distorted according to the determined distance between the user's thumb and one of the other fingers.

[0018] In an embodiment of the computer-readable storage medium, the ray casting interaction mode may provide a visual affordance including a shape shown in the user's field of view positioned between the user's thumb and one of the other fingers, which shape may also be shown resized or distorted according to a determined distance between the user's thumb and one of the other fingers.

[0019] In an embodiment of the computer-readable storage medium, the second tracked hand pose consistent with the hand ready state can include a hand pose identified as the user's palm facing up.

[0020] In an embodiment of the computer-readable storage medium, the gaze-gesture interaction mode may include receiving a user instruction of a direction based at least on a determined orientation of the user's head and based on a determined position of the user's head relative to the artificial reality environment, where the determined position of the user's head is based on tracked movement of the user along an X-axis, a Y-axis, and a Z-axis; and receiving a user instruction of an action for one of a plurality of user instructions of a direction by tracking a posture of the user's hand and matching the hand posture with a specified action.

[0021] In one aspect, the present invention is directed to a computing system for interpreting user input in a virtual reality system, the computing system comprising: The computing system includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the computing system to implement the method described above, i.e., a process for transitioning between interaction modes, the process including: identifying a first interaction mode context indicating that hand tracking input is unavailable or that a first tracked hand pose is not consistent with a hand-ready state; enabling a no-hand interaction mode in response to identifying the first interaction mode context; identifying a second interaction mode context indicating that a second tracked hand pose is consistent with a hand-ready state; enabling a gaze-gesture interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that a third tracked hand pose is consistent with a ray state; and enabling a ray-casting interaction mode in response to identifying the third interaction mode context.

[0022] In an embodiment of the computing system, the third tracked hand pose consistent with the lighting conditions may include a hand pose identified as the user's palm pointing down by at least a threshold amount.

[0023] In embodiments of the computing system, the no-hands interaction mode and the gaze-gesture interaction mode can provide a visual affordance including a gaze cursor, which is presented in a field of view for the user, positioned at least in part based on the tracked position of the user's head. [Brief explanation of the drawings]

[0024] [Figure 1] FIG. 1 is a block diagram illustrating an overview of a device in which some implementations of the present technology may operate. [Figure 2A] FIG. 1 is a diagram illustrating a virtual reality headset that can be used in some implementations of the present technology. [Figure 2B] FIG. 1 is a diagram illustrating a mixed reality headset that can be used in some implementations of the present technology. [Figure 3] FIG. 1 is a block diagram illustrating an overview of an environment in which some implementations of the present technology may operate. [Figure 4] FIG. 1 is a block diagram illustrating components that can be used in a system employing the disclosed technology in some implementations. [Figure 5] 1 is a flow diagram illustrating a process for identifying a trigger for switching between specified interaction modes and enabling the triggered interaction mode, used in some implementations of the present technology. [Figure 6] FIG. 10 is a conceptual diagram illustrating an exemplary gaze cursor. [Figure 7] FIG. 10 is a conceptual diagram illustrating an example gaze cursor with a dwell timer. [Figure 8] FIG. 10 is a conceptual diagram illustrating an example gaze cursor with hand gesture selection. [Figure 9] FIG. 10 is a conceptual diagram illustrating an example of ray-casting directional input with hand gesture selection. [Figure 10A] FIG. 1 is a conceptual diagram illustrating an example interaction between device positioning and object selection using an eye-tracking based gaze cursor with a dwell timer. [Figure 10B] FIG. 1 is a conceptual diagram illustrating an example interaction between device positioning and object selection using an eye-tracking based gaze cursor with a dwell timer. DETAILED DESCRIPTION OF THE INVENTION

[0025] The technology introduced herein can be better understood by reference to the following detailed description in conjunction with the accompanying drawing figures, in which like reference numerals indicate identical or functionally similar elements.

[0026] Aspects of the present disclosure are directed to an interactive mode system that provides multiple interactive modes in a virtual reality environment and has automatic, context-specific transitions between the interactive modes. The interactive modes can specify how the interactive mode system (a) determines orientation and movement within the virtual reality environment and (b) determines interactions for selecting or performing other actions. The interactive mode system also provides affordances for signaling the current interactive mode and interactive mode transitions. The interactive mode system can control multiple interactive modes and can control transitions between interactive modes using a mapping of interactive mode context elements to interactive modes. A user can also effect additional interactive mode transitions by manually activating software or hardware controls, performing interactive mode transition gestures, or other explicit commands.

[0027] Interaction mode context elements may include, for example, hardware components that are enabled or not (via hardware or software control); software modules that are enabled or not; current privacy, power, or accessibility settings; environmental conditions (e.g., lighting, location detection elements such as surface type or marker availability, region-based camera limitations, determined region type, number of surrounding people or identified specific surrounding people, etc.); current body position (e.g., whether hands are visible, hand orientation, gestures being performed); current controller position or state, etc.

[0028] The interaction mode system can determine the interaction mode context by integrating inputs from various sub-systems, such as head tracking, position tracking, eye tracking, hand tracking, other camera-based systems, hardware controllers, simultaneous localization and mapping (SLAM) systems, or other data sources. The head tracking sub-system can identify the orientation of the user's hands based on one or more data from, for example, an internal motion unit (IMU—discussed below), tracked light emitted by a head-mounted device, a camera or other external sensor aimed at capturing a representation of the user's head, or a combination thereof. The position tracking sub-system can identify the user's position within the artificial reality environment based on, for example, tracked light emitted by a wearable device, time-of-flight sensors, conventional camera data and / or depth camera data, IMU data, environmental mapping based on GPS data, etc. The eye tracking sub-system can model one or both of the user's eyes and often identify the gaze direction relative to the head-mounted device by determining a gaze direction vector along a line connecting the user's fovea and the center of the user's pupil, based on factors such as an image capturing the circle of light reflected by the user's eye (i.e., "glint"). The hand tracking sub-system can determine the position, pose, and / or movement of a user's hands based on images depicting one or both of the user's hands, e.g., from one or more cameras, and / or sensor data from devices worn by the user, such as gloves, rings, wristbands, etc. Information from other camera systems can gather additional information, such as lighting conditions, environmental mapping, etc. Each of these sub-systems can provide result data (e.g., head orientation, position, eye orientation, hand pose and position, body pose and position, environmental data, mapping data, etc.), as well as subsystem meta-data, such as whether aspects of the sub-system are enabled and whether the sub-system has access to particular types of data (e.g., whether a particular set of cameras or other sensors is enabled, battery levels, accuracy predictions for associated result data, etc.).

[0029] In some implementations, the interaction mode system can control at least four interaction modes, including a no-hands 3DoF mode, a no-hands 6DoF mode, a gaze-and-gesture mode, and a ray-casting mode. The no-hands 3DoF interaction mode can track only three degrees of freedom (roll, yaw, and pitch) (however, the user can direct X, Y, and Z movement via other means, such as a joystick on a controller). In essence, tracking in 3DoF is oriented as a vector originating from the virtual reality device, extending to a point on a sphere surrounding the origin, which is identified based on IMU data. In some implementations, the vector can be modified based on the user's identified gaze direction. Based on this vector, a "gaze cursor" can be provided in the user's field of view, either at the center of the field of view where eye tracking is unavailable, or positioned according to the user's eye focus, where eye tracking is available. While in the no-hands 3DoF interaction mode, the interaction mode system can use a dwell timer that begins counting down if the direction vector does not change by a threshold within a certain time period and is reset when a threshold change in the direction vector is detected. When the dwell timer counts down to zero, it can indicate a user action. Such an action, combined with the vector direction, can result in selection or other actuation of the object being pointed at by the gaze cursor. For additional details regarding gaze cursors and dwell timers configured in accordance with the present technology, see Figures 6 and 10 and the associated discussion below. Interaction mode contextual factors that can trigger a transition to the no-hands 3DoF interaction mode can include factors such as when the position tracking camera is off or otherwise unavailable (e.g., in low-power mode, privacy mode, public mode, etc.), or when lighting conditions are such that hand tracking cannot be reliably performed.

[0030] The no-hands 6DoF interaction mode may be similar to the no-hands 3DoF interaction mode, except that the no-hands 6DoF interaction mode system not only tracks a direction vector for the gaze cursor, but also tracks the movement of the user's position within the virtual reality environment. As with the no-hands 3DoF interaction mode, actions in the no-hands 6DoF interaction mode may include steering the gaze cursor (which may now be accomplished by the user changing their orientation and position) and using a dwell timer to perform actions on the gaze cursor. Interaction mode context elements that can trigger a transition to the no-hands 6DoF interaction mode include an indication that the position tracking system is enabled, but either (a) the hand tracking system is not enabled or lighting conditions are insufficient to track hand position, or (b) there is no hand "ready" detection. The ready state may occur when the user's hand is present in the user's field of view above a threshold position and / or when the user's hand is in a particular configuration (e.g., palm up).

[0031] The gaze-gesture interaction mode can be a version of the 6DoF or 3DoF interaction mode, where direction is based on the gaze cursor, but instead of (or in addition to) using a dwell timer for actions, the user can use hand gestures to indicate actions. For example, if a hand is present in the user's field of view and palm facing up, a "pinching" gesture with the thumb and fingers can indicate an action. Other gestures, such as an air tap, air swipe, grab, or other identifiable movement or posture, can also be used. Interaction mode contextual elements that can trigger a transition to the gaze-gesture interaction mode include an enabled hand tracking system and identification of the user's hand in a ready state. In some implementations, transitioning to the gaze-gesture interaction mode may further require the availability of a position tracking system for 6DoF movement. See Figure 8 for additional details regarding the gaze-gesture interaction mode.

[0032] The ray-casting interaction mode can replace the gaze cursor with a ray emanating from the user's hand into the virtual reality environment. For example, the ray can be specified along a line connecting an origin (e.g., the center of mass of the hand or the user's eye, shoulder, or hip) and a control point (e.g., a point relative to the user's hand). Hand gestures similar to those used in the gaze-gesture interaction mode can be used to indicate actions. Interaction mode context elements that can trigger a transition to the ray-casting interaction mode can include an available hand tracking system and the user's hand being in a ray-casting pose (e.g., both (a) a hand in the user's field of view or positioned above a threshold, and (b) a hand positioned palm-down). For additional details regarding the ray-casting interaction mode, see Figure 9 and the related discussion below.

[0033] In some implementations, other controls may augment or replace direction and / or selection decisions in any of the above interaction modes. For example, a user may use the controller and / or push buttons on the controller to indicate a direction, or may use a paired device to indicate a direction or action (e.g., by swiping gestures or tapping on a touchscreen, performing a scrolling action on a capacitive sensor such as on a watch or ring, etc.) to indicate a selection or other action.

[0034] Embodiments of the disclosed technology may include or be implemented in connection with a virtual reality system. Virtual reality or extra reality (XR) is a form of reality that is adjusted in some way prior to presentation to a user, and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Virtual reality content may include fully generated content or generated content combined with captured content (e.g., real-world photography). Virtual reality content may include video feedback, audio feedback, haptic feedback, or some combination thereof, any of which may be provided in a single channel or multiple channels (e.g., stereoscopic video to create a three-dimensional effect for the viewer). Furthermore, in some embodiments, the virtual reality may be associated with applications, products, accessories, services, or some combination thereof, for example, used to create content in the virtual reality and / or used in the virtual reality (e.g., used to perform activities in the virtual reality). Virtual reality systems that provide virtual reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, "cave" environments or other projection systems, or any other hardware platform capable of providing virtual reality content to one or more observers.

[0035] As used herein, "virtual reality" or "VR" refers to an immersive experience in which a user's visual input is controlled by a computing system. "Augmented reality" or "AR" refers to a system in which a user sees an image of the real world after the image passes through a computing system. For example, a tablet with a camera on its back can capture an image of the real world and then display the image on a screen on the opposite side of the tablet from the camera. The tablet can process and adjust, or "augment," the image as it passes through the system, such as by adding virtual objects. "Mixed reality" or "MR" refers to a system in which light entering a user's eye is partially generated by a computing system and partially composed of light reflected from real-world objects. For example, an MR headset can be configured as glasses with a pass-through display that allow light from the real world to pass through a waveguide that simultaneously emits light from the MR headset's projector, allowing the MR headset to present virtual objects mixed with the actual objects the user can see. As used herein, "artificial reality," "extra reality," or "XR" can refer to any of VR, AR, MR, or any combination thereof, i.e., hybrid.

[0036] Existing XR systems have developed many interaction modes, including features such as ray casting, air tapping, dwell timers, and gaze cursors. However, existing XR systems do not provide automatic transitions between interaction modes using these features, requiring the user to manually change the interaction mode, either using the same interaction mode for all interactions with a particular application, or requiring the application to specify the interaction mode regardless of context. However, this can create problems when the current interaction mode is inappropriate for the current context, for example, when conditions do not allow aspects of the current interaction mode to function properly (e.g., the interaction mode relies on hand detection and the hands are not visible, lighting conditions are poor, and / or hardware is not enabled to capture and interpret hand poses), when an alternative interaction mode is unobtrusive or does not offer as much performance benefit (e.g., when the user's hands are otherwise occupied or it is socially inappropriate to make hand gestures, when the device power is low and it is desirable to use an interaction mode with lower power demands, etc.), or when the current interaction mode is suboptimal for the interaction to be performed (e.g., when the current interaction mode uses a time-oriented dwell timer and continuous selection is to be performed).

[0037] The interaction mode system and process described herein are expected to overcome these problems associated with conventional XR interaction techniques and are expected to (a) provide greater functionality while being more natural and intuitive than interactions in existing XR systems, (b) provide new efficiencies by selecting interaction modes consistent with current power and environmental conditions, and (c) make virtual reality interactions less intuitive for the user and other bystanders. Despite being natural and intuitive, the systems and processes described herein are rooted in computerized virtual reality systems instead of being analogs of traditional object interaction. For example, existing object interaction techniques do not provide a set of interaction modes that smoothly and automatically transition between various modes, as described herein. Furthermore, the interaction mode system and process described herein provide an improvement by enabling interaction in additional situations, such as when an existing XR system provides an interaction mode that cannot operate in the current context or in which the user cannot interact (e.g., by switching to a more appropriate interaction mode in low lighting conditions or when the user's hands are otherwise occupied).

[0038] Some implementations are discussed in more detail below with reference to the figures. For example, FIG. 1 is a block diagram illustrating an overview of a device in which some implementations of the disclosed technology may operate. The device may include hardware components of a computing system 100 that switches interaction modes based on elements in an identified interaction mode context or in response to user control. In various implementations, the computing system 100 may include a single computing device 103 or multiple computing devices (e.g., computing device 101, computing device 102, and computing device 103) that communicate via wired or wireless channels to distribute processing and share input data. In some implementations, the computing system 100 may include a standalone headset that can provide a computer-generated or computer-augmented experience to a user without requiring external processing or sensors. In other implementations, the computing system 100 may include multiple computing devices, such as a headset and a core processing component (such as a console, mobile device, or server system), with some processing operations performed on the headset and other processing operations offloaded to the core processing component. An exemplary headset is described below in connection with FIGS. 2A and 2B. In some implementations, location and environmental data may be collected solely by sensors built into the headset device, while in other implementations, one or more of the multiple non-headset computing devices may include sensor components capable of tracking environmental or location data.

[0039] Computing system 100 may include one or more processors 110 (e.g., a central processing unit (CPU), a graphical processing unit (GPU), a holographic processing unit (HPU), etc.). Processor 110 may be a single processing unit, or multiple processing units within a single device or distributed across multiple devices (e.g., distributed across two or more of computing devices 101-103).

[0040] Computing system 100 may include one or more input devices 120 that provide input to processor 110 that informs processor 110 of actions. Actions may be mediated by a hardware controller that interprets signals received from the input devices and communicates that information to processor 110 using a communication protocol. Individual input devices 120 may include, for example, a mouse, a keyboard, a touchscreen, a touchpad, a wearable input device (e.g., haptic gloves, bracelets, rings, earrings, necklaces, watches, etc.), a camera (or other light-based input device, e.g., an infrared sensor), a microphone, or other user input device.

[0041] The processor 110 can be coupled to other hardware devices using an internal or external bus, such as a PCI bus, a SCSI bus, or a wireless connection. The processor 110 can communicate with a hardware controller for devices such as a display 130. The display 130 can be used to display text and graphics. In some implementations, the display 130 includes an input device as part of the display, such as when the input device is a touchscreen or when wearing an eye-direction monitoring system. In some implementations, the display is separate from the input device. Examples of display devices are LCD display screens, LED display screens, projected holographic or augmented reality displays (such as head-up display devices or head-mounted devices), etc. Other I / O devices 140 can also be coupled to the processor, such as a network chip or card, a video chip or card, an audio chip or card, a USB, Firewire or other external device, a camera, a printer, speakers, a CD-ROM drive, a DVD drive, a disk drive, etc.

[0042] The computing system 100 may include a communication device capable of wireless or hardwired communication with other local computing devices or network nodes. The communication device may communicate with another device or server over a network using, for example, the TCP / IP protocol. The computing system 100 may utilize the communication device to distribute operations across multiple network devices.

[0043] The processor 110 may have access to memory 150, which may be included in one of multiple computing devices of the computing system 100 or distributed across multiple computing devices of the computing system 100 or other multiple external devices. Memory includes one or more hardware devices for volatile or non-volatile storage and may include both read-only and writable memory. For example, memory may include one or more of random access memory (RAM), various caches, CPU registers, read-only memory (ROM), and writable non-volatile memory such as flash memory, hard drives, floppy disks, CDs, DVDs, magnetic storage devices, tape drives, etc. Memory is not a propagating signal separate from the underlying hardware; therefore, memory is non-transitory. Memory 150 may include program memory 160, which stores programs and software such as an operating system 162, an interactive mode system 164, and other application programs 166. Memory 150 may also include data memory 170, which may include, for example, mapping of interaction mode contextual elements to interaction modes for triggering interaction mode switching, configuration data, settings, user options or preferences, etc., which may be provided to program memory 160 or any element of computing system 100.

[0044] Some implementations can operate using many other computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with the present technology include, but are not limited to, XR headsets, personal computers, server computers, handheld or laptop devices, cellular telephones, wearable electronics, gaming consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and so forth.

[0045] 2A is a diagram of a virtual reality head-mounted display (HMD) 200 according to some embodiments. The HMD 200 includes a front rigid body 205 and a band 210. The front rigid body 205 includes one or more electronic display elements of an electronic display 245, an inertial motion unit (IMU—also called an inertial measurement unit) 215, one or more position sensors 220, a locator 225, and one or more computational units 230. The position sensor 220, the IMU 215, and the computational units 230 may be internal to the HMD 200 and may be invisible to the user. In various implementations, the IMU 215, the position sensor 220, and the locator 225 can track the movement and location of the HMD 200 in real-world and virtual environments with three degrees of freedom (3DoF) or six degrees of freedom (6DoF). For example, the locator 225 can emit infrared light beams that generate light points on real objects around the HMD 200. As another example, the IMU 215 may include, for example, one or more accelerometers, gyroscopes, magnetometers, other non-camera-based position, force, or orientation sensors, or a combination thereof. One or more cameras (not shown) integrated with the HMD 200 may detect light spots. A computing unit 230 in the HMD 200 may use the detected light spots to capture the position and movement of the HMD 200 and to identify the shape and position of actual objects surrounding the HMD 200.

[0046] Electronic display 245 may be integrated with front rigid body 205 and may provide image light to the user as directed by computing unit 230. In various embodiments, electronic display 245 may be a single electronic display or multiple electronic displays (e.g., one display for each user's eye). Examples of electronic display 245 include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, an active matrix organic light emitting diode display (AMOLED), a display including one or more quantum dot light emitting diode (QOLED) sub-pixels, a projector unit (e.g., microLED, LASER, etc.), some other display, or some combination thereof.

[0047] In some implementations, the HMD 200 can be coupled to a core processing component, such as a personal computer (PC) (not shown), and / or one or more external sensors (not shown). The external sensors, in conjunction with output from the IMU 215 and position sensor 220, can monitor the HMD 200 (e.g., via light emitted from the HMD 200) which the PC can use to determine the location and movement of the HMD 200.

[0048] In some implementations, the HMD 200 can communicate with one or more other external devices, such as controllers (not shown), which the user can hold in one or both hands. The controllers can have their own IMU units, position sensors, and / or emit other light points. The HMD 200 or external sensors can track these controller light points. The computing unit 230 or core processing component in the HMD 200 can use this tracking in conjunction with the IMU and position outputs to monitor the position and movement of the user's hands. The controllers can also include various buttons that the user can actuate to provide input and interaction with virtual objects. In various implementations, the HMD 200 can also include additional subsystems, such as an eye tracking unit, an audio system, various network components, and so on. In some implementations, instead of or in addition to the controllers, one or more cameras included within or external to the HMD 200 can monitor the position and pose of the user's hands, thereby determining gestures and other hand and body movements.

[0049] 2B is a diagram of a mixed reality HMD system 250 including a mixed reality HMD 252 and a core processing component 254. The mixed reality HMD 252 and the core processing component 254 can communicate via a wireless connection (e.g., a 60 GHz link) as indicated by link 256. In other implementations, the mixed reality system 250 lacks an external computing device and includes only a headset or other wired or wireless connection between the mixed reality HMD 252 and the core processing component 254. The mixed reality HMD 252 includes a pass-through display 258 and a frame 260. The frame 260 can house various electronic components (not shown), such as a light projector (e.g., a laser, an LED, etc.), a camera, an eye tracking sensor, MEMS components, networking components, etc.

[0050] The projector can be coupled to a pass-through display 258, for example, via optical elements, to display the media to the user. The optical elements can include one or more waveguide assemblies, reflectors, lenses, mirrors, collimators, gratings, etc., to direct light from the projector to the user's eyes. Image data can be transmitted from the core processing component 254 to the HMD 252 via link 256. A controller within the HMD 252 can convert the image data into light pulses from the projector, which can be transmitted via optical elements to the user's eyes as output light. This output light can be mixed with light passing through the display 258, such that this light output can present virtual objects that appear as if they existed in the real world.

[0051] Like HMD 200, HMD system 250 may also include motion and position tracking units, cameras, light sources, etc., which enable HMD system 250 to, for example, track itself with 3DoF or 6DoF, track parts of the user (e.g., hands, feet, head or other body parts), map virtual objects so they appear stationary as HMD 252 moves, and make virtual objects responsive to gestures and other real-world objects.

[0052] 3 is a block diagram illustrating an overview of an environment 300 in which some implementations of the disclosed technology may operate. The environment 300 may include one or more client computing devices 305A-D, where an example of a client computing device may include computing system 100. In some implementations, some of the client computing devices (e.g., client computing device 305B) may be HMD 200 or HMD system 250. The client computing devices 305 may operate in a networked environment using logical connections via a network 330 to one or more remote computers, such as a server computing device.

[0053] In some implementations, server 310 may be an edge server that receives client requests and coordinates fulfillment of those requests through other servers, such as servers 320A-C. Server computing devices 310 and 320 may comprise computing systems, such as computing system 100. Although individual server computing devices 310 and 320 are logically represented as single servers, each of these server computing devices may be a distributed computing environment encompassing multiple computing devices located in the same location or in disparate geographical physical locations.

[0054] Client computing device 305 and server computing devices 310 and 320 can each act as a server or client to other server / client devices. Server 310 can be connected to database 315. Servers 320A-C can be connected to corresponding databases 325A-C, respectively. As discussed above, each server 310 or 320 can correspond to a group of multiple servers, each of which can share a database or have its own database. Although databases 315 and 325 are logically viewed as a single unit, databases 315 and 325 may each be located within their respective servers, or in a distributed computing environment encompassing multiple computing devices that can be located in the same location or in completely different geographical physical locations.

[0055] Network 330 may be a local area network (LAN), wide area network (WAN), mesh network, hybrid network, or other wired or wireless network. Network 330 may also be the Internet or some other public or private network. Client computing device 305 may connect to network 330 via a network interface, such as by wired or wireless communication. Although the connections between server 310 and server 320 are shown as separate connections, these connections may be any type of local, wide area, wired, or wireless network, including network 330 or a separate public or private network.

[0056] 4 is a block diagram illustrating components 400 that can be used in a system employing the disclosed technology in some implementations. Components 400 can be included within one device of computing system 100 or distributed across multiple devices of computing system 100. Components 400 include hardware 410, a mediator 420, and multiple dedicated components 430. As discussed above, systems implementing the disclosed technology can use a variety of hardware, including a processing unit 412, working memory 414, input and output devices 416 (e.g., cameras, displays, IMU units, network connections, etc.), and storage memory 418. In various implementations, storage memory 418 can be one or more of a local device, an interface to a remote storage device, or a combination thereof. For example, storage memory 418 can be one or more hard drives or flash drives accessible via a system bus, or can be a cloud storage provider (such as in storage device 315 or 325) or other network storage accessible via one or more communication networks. In various implementations, component 400 can be realized in a client computing device, such as client computing device 305, or on a server computing device, such as server computing device 310 or 320.

[0057] The mediator 420 may include a component that mediates resources between the hardware 410 and multiple dedicated components 430. For example, the mediator 420 may include an operating system, a service, a driver, a basic input / output system (BIOS), a controller circuit, or other hardware or software system.

[0058] The dedicated components 430 may include software or hardware configured to perform operations to switch interaction modes (a) based on elements in an identified interaction mode context or (b) in response to user control. The dedicated components 430 may include components and APIs such as a mode selector 434, a no-hands 3DoF mode controller 436, a no-hands 6DoF mode controller 438, a gaze-gesture mode controller 440, a ray-casting mode controller 442, and an interface 432 that can be used to provide a user interface, transfer data, and control the dedicated components. In some implementations, the components 400 may be in a computing system distributed across multiple computing devices or may interface to a server-based application that executes one or more of the dedicated components 430. While depicted as separate components, the dedicated components 430 may represent logically, i.e., non-physically, distinct functions and / or submodules or code blocks of one or more applications.

[0059] The mode selector 434 can identify interaction mode switching conditions and switch user interaction control to a corresponding one of the no-hands 3DoF mode controller 436, the no-hands 6DoF mode controller 438, the gaze-and-gesture mode controller 440, or the ray-casting mode controller 442. In various implementations, the mode selector 434 can identify such conditions in response to an explicit user selection of an interaction mode or by mapping an interaction mode context element to the corresponding interaction mode. For example, the interaction mode context element of no useful tracking input can be mapped to the no-hands 3DoF interaction mode, the interaction mode context element of useful tracking input without a ready hand identification can be mapped to the no-hands 6DoF interaction mode, the interaction mode context element of useful tracking input and a ready hand identification can be mapped to the gaze-and-gesture interaction mode, and the interaction mode context element of useful tracking input and a ray-state hand can be mapped to the ray-casting interaction mode. Interaction Mode Additional details regarding selection of an interaction mode based on contextual factors or user selection are provided below in connection with blocks 502-510 and 552-562 of FIG.

[0060] The no-hands 3DoF mode controller 436 can provide user interaction where direction is indicated using a gaze cursor and actions are indicated using a dwell timer. The no-hands 3DoF mode controller 436 does not translate user movement in the X, Y, or Z axes into changes in the user's field of view. The no-hands 3DoF mode controller 436 can interface with an IMU unit of the I / O 416 via the mediator 420 to determine head position and can also interface with an eye tracker of the I / O 416 via the mediator 420 to determine eye gaze direction, the combination of which can specify the location of the gaze cursor. Additional details regarding the no-hands 3DoF interaction mode are provided below in connection with block 580 of FIG. 5 .

[0061] The no-hands 6DoF mode controller 438 can provide user interaction where direction is indicated using a gaze cursor and actions are indicated using a dwell timer. The no-hands 6DoF mode controller 438 can interface with the position tracking system of the I / O 416 via the mediator 420 to determine user movement along the X, Y, and Z axes and convert these user movements into changes in the user's field of view. The no-hands 6DoF mode controller 438 can interface with the IMU unit of the I / O 416 via the mediator 420 to determine head position and can interface with the eye tracker of the I / O 416 via the mediator 420 to determine eye gaze direction, the combination of which can specify the location of the gaze cursor. Additional details regarding the no-hands 6DoF interaction mode are provided below in connection with block 582 of FIG. 5 .

[0062] The gaze-gesture mode controller 440 can provide user interaction where direction is indicated using a gaze cursor and actions are indicated using hand gestures. The gaze-gesture mode controller 440 can interface with the position tracking system of the I / O 416 via the mediator 420 to determine user movement along the X, Y, and Z axes and translate these user movements into changes in the user's field of view. The gaze-gesture mode controller 440 can interface with the IMU unit of the I / O 416 via the mediator 420 to determine head position and with the eye tracker of the I / O 416 via the mediator 420 to determine eye gaze direction, the combination of which can specify the location of the gaze cursor. The gaze-gesture mode controller 440 can further interface with the hand tracking system via the mediator 420 to identify hand posture, which the gaze-gesture mode controller 440 translates into actions. Additional details regarding the gaze-gesture interaction mode are provided below in connection with block 584 of FIG.

[0063] The ray casting mode controller 442 can provide user interaction in which direction is indicated using a user-controlled light beam and actions are indicated using hand gestures. The ray casting mode controller 442 can interface with the position tracking system of the I / O 416 via the mediator 420 to determine the user's movement along the X-, Y-, and Z-axes, and can also interface with the IMU unit of the I / O 416 via the mediator 420 to determine head position, a combination of which the ray casting mode controller 442 can use to determine changes in the user's field of view. The ray casting mode controller 442 can further interface with the hand tracking system via the mediator 420 to identify hand poses, which the ray casting mode controller 442 converts into control points for casting rays and identifying actions. Additional details regarding the ray casting interaction mode are provided below in connection with block 584 of FIG. 5 .

[0064] 1-4 described above, and in each of the flow charts discussed below, can be modified in various ways. For example, the order of logic can be rearranged, substeps can be performed simultaneously, illustrated logic can be omitted, other logic can be included, etc. In some implementations, one or more of the components described above can perform one or more of the processes described below.

[0065] FIG. 5 is a flow diagram illustrating a process 500 for identifying a trigger for switching a specified interaction mode and enabling the triggered interaction mode, used in some implementations of the present technology. Process 500 is depicted as having two possible starting points 502 and 552. Process 500 can begin at start point 552 when an explicit user command to manually switch interaction modes is received. Such a command can be implemented, for example, using a physical button or other control included in the virtual reality system, using a virtual control (e.g., a user interface element) displayed by the virtual reality system, using a voice command, or via a gesture or other input mapped to switch interaction modes. Process 500 can begin at start point 502 in response to various prompts for checking an interaction mode change trigger, such as a periodic check (e.g., every 1 ms, 5 ms, or 10 ms), or in response to detecting a change in an interaction mode context element. In some implementations, the element that, when changed, will cause process 500 to be implemented may differ depending on the current interaction mode. For example, if the current interaction mode is no-hands 3DoF mode because the interaction mode system is in a low power state, a change in lighting conditions will not trigger a mode change check, whereas if the current interaction mode is ray casting mode, which requires the camera to capture hand positions, a change in lighting conditions will trigger a mode change check.

[0066] Process 500 may proceed from block 502 to block 504. At block 504, process 500 may determine whether there is any useful input for tracking position and tracking hands. If the camera or corresponding processing component, such as a machine learning model, used for one of these systems is not enabled (e.g., because the interactive mode system is in low power or privacy mode, the system has not yet been initialized, etc.), or if lighting conditions are insufficient to capture images of sufficient quality for one of these systems, there may be no useful input for tracking position and tracking hands. In some implementations, hand or position tracking can be based on systems other than cameras, such as wearable gloves, rings, and / or bracelets for a hand tracking system, or time-of-flight sensors or GPS data for a position tracking system, and the determination of useful hand tracking input can be based on whether these systems are enabled and receive sufficient information to determine hand pose and user position. If there is no such useful tracking input, process 500 may continue to block 580. If there is such useful tracking input, process 500 may continue to block 506.

[0067] At block 580, process 500 can enable a no-hands 3DoF interaction mode. In the no-hands 3DoF interaction mode, a user moving parallel to or perpendicular to the floor (i.e., on the X, Y, and Z axes) does not result in corresponding movement in the virtual reality environment (although such movement can be implemented using other controls, such as a joystick or directional pad on a controller), but the user can indicate direction by moving their head to points at various locations in the virtual reality environment, for example, using a gaze cursor. In some implementations, the gaze cursor position can be based on both head tracking (e.g., via IMU data indicating pitch, roll, and yaw movement) and eye tracking (e.g., a camera capturing images of the user's eyes to determine the direction of the user's gaze). In some instances in the no-hands 3DoF interaction mode, an action can be implemented using a dwell timer that begins a countdown (e.g., from 3 seconds) if the gaze cursor does not move more than a threshold amount for at least a threshold amount of time (e.g., if the user's gaze remains relatively fixed for at least 1 second). In some implementations, the action can instead or additionally be indicated in other ways, such as by pressing a button on the controller, by activating a UI control (i.e., a "soft" button), by using a voice command, etc. In some implementations, the action can then be performed on the gaze cursor, e.g., the system performs a default action on one or more objects at which the gaze cursor is pointed. In some implementations, an indication of the gaze cursor and / or the user's gaze direction can be displayed in the user's field of view to indicate that the user is in a no-hands 3DoF interaction mode (or a no-hands 6DoF interaction mode, discussed below) and as a visual affordance to assist the user in making selections using the gaze cursor. Additional details regarding the gaze cursor and dwell timer are provided below in connection with FIGS. 6, 7, and 10.Once the interactive mode is enabled at block 580 , the process 500 may end at block 590 .

[0068] At block 506, process 500 may determine whether position information is available while the hands are not identified as ready. Position information is available when the positioning system is enabled and receives sufficient data to determine the user's movements in 6 DoF. In various implementations, such a positioning system may include a system attached to a wearable portion of the virtual reality system ("inside-out tracking") or external sensors pointed at the user to track body position ("outside-in tracking"). Furthermore, these systems may use one or more technologies, such as coded points of light (e.g., infrared) emitted by the virtual reality system, use a camera to track the movement of these points, identify features in captured images and track their relative movement across frames (e.g., "motion vectors"), identify parts of the user in captured images and track the movement of these parts (e.g., "skeleton tracking"), obtain time-of-flight sensor readings from objects in the environment, acquire Global Positioning System (GPS) data, perform triangulation using multiple sensors, and so on. Determining whether sufficient data exists may include determining whether a location tracking system is enabled (such systems may be disabled, for example, in privacy mode, low power mode, public mode, do not disturb mode, etc.), and whether such systems can properly determine the user's location (which may not be possible in some cases when some of these systems are in low lighting conditions, for example, when there are not enough nearby surfaces to illuminate a light spot, when environmental mapping has not yet occurred, when the user is out of frame of the capture device, etc.).

[0069] A hand can be identified as being in a ready state if the hand tracking system is enabled and the current hand posture matches the ready state hand posture. Hand tracking systems can include a variety of technologies, such as a camera that captures and interprets images of the user's hand (e.g., using a trained model) or a wearable device such as a glove, ring, or bracelet that captures measurements that can be interpreted as the user's hand posture. Examples of possible ready state hand postures include when the user's palm is facing substantially up (e.g., rotated toward the sky by at least a threshold amount), when the user's fingers are spread a certain amount, when certain fingers are touching (e.g., thumb and index finger), when the user's hand is tilted above the horizon by a certain angle, when the user is making a fist, etc. In some implementations, the hand ready state may require the hand to be in the user's field of view, elevated above a certain threshold (e.g., above hip level), or when the user's elbow is bent at least a certain amount. Additional details regarding the hand ready state are provided below in connection with FIG. 8. In some implementations, block 506 may also determine whether the hand is not in a state mapped to another interaction mode, such as the ray state discussed below in connection with blocks 508 and 510. If position information is available and the hand is not identified as being in a ready state (i.e., another interaction mode map state), process 500 may continue at block 582; otherwise, process 500 may continue at block 508.

[0070] At block 582, process 500 may enable a no-hands 6DoF interaction mode. The no-hands 6DoF interaction mode may provide similar interaction to the no-hands 3DoF interaction mode described above, except that user movement parallel and perpendicular to the floor (i.e., along the X, Y, and Z axes) is detected and translated into gaze movement in the virtual reality environment. Gaze movement and head and / or eye tracking may be used to position a gaze cursor. As with the no-hands 3DoF interaction mode, actions in the no-hands 6DoF interaction mode may be performed using a dwell timer and / or other user controls (e.g., physical or soft buttons, voice commands, etc.), and these actions may be performed on the gaze cursor. Once the interaction mode is enabled at block 582, process 500 may end at block 590.

[0071] At block 508, process 500 may determine whether the hand is in a ready state, as discussed above. If the hand is in a ready state, process 500 may continue at block 584; if the hand is not in a ready state, process 500 may continue at block 510.

[0072] At block 584, process 500 may enable a gaze-and-gesture interaction mode. The gaze-and-gesture interaction mode may include a gaze cursor, as discussed above. In some implementations, the gaze-and-gesture interaction mode always uses position tracking to allow a user to control their position in the virtual reality environment by moving through the real world, while in other cases, position tracking may or may not be enabled for the gaze-and-gesture interaction mode. When position tracking is enabled at block 584, the gaze cursor may be controlled with 6 DoF by the user, and when position tracking is not enabled at block 584, the gaze cursor may be controlled with 3 DoF by the user. A dwell timer may or may not be enabled in the gaze-and-gesture interaction mode.

[0073] When hand tracking is enabled and the hand is in a ready state, the gaze-gesture interaction mode can enable actions based on user gestures. For example, when the user performs a particular gesture, such as bringing their thumb and index finger together, a default action can be performed. In some implementations, different actions can be mapped to different gestures. For example, when the user brings their thumb and index finger together, a first action can be performed, and when the user brings their thumb and middle finger together, a second action can be performed. The gaze-gesture interaction mode can use one or more of many different gestures, such as a swipe gesture, an air tap gesture, a grip gesture, a finger extension gesture, a finger curl gesture, etc. In some cases, a visual affordance can be used to signal to the user that the gaze-gesture interaction mode is enabled and / or that a gesture is recognized. For example, the hand-ready state may be a user's palm facing substantially upward, and the visual affordance for the gaze-gesture interaction mode may include placing a sphere between the user's thumb and index finger (see FIG. 8 ). This may both signal to the user that the gaze-gesture interaction mode is enabled and provide feedback for a pinch gesture between the fingers (e.g., by distorting or resizing the sphere as the user brings the fingers together). Other visual affordances may also be used, such as by showing a trace of the user's hand movement and placing indicators of different actions that can be performed near the tips of each finger, or by showing an indicator on one wrist that can be tapped by a finger from the other hand. Additional details regarding the gaze-gesture interaction mode using visual affordances are provided below in connection with FIG. 8 . Once the interaction mode is enabled in block 584, process 500 may end in block 590.

[0074] At block 510, process 500 may determine whether the user's hands are in a ray state. Similar to the processes described above in connection with blocks 506 and 508, this determination may include determining whether hand tracking is enabled, whether sufficient data exists to determine hand pose, and / or whether one or both hands are visible, whether they are raised above a threshold, or whether the user's elbows are bent beyond a threshold amount. Identifying a ray state may also include determining whether one or both of the user's hands are in a pose mapped for performing ray casting. Examples of such poses may include the user's hand being rotated so that the palm is facing substantially downward (e.g., rotated toward the floor by at least a threshold amount), fingers (e.g., index finger) are extended, certain fingers are touching (e.g., thumb, middle finger, and ring finger), etc. If the user's hand is identified as being in a ray state, process 500 may continue to block 586. If the user's hand is identified as not being in a ray state, process 500 may continue to end at block 590 without implementing a change to the current interaction mode.

[0075] At block 586, process 500 may enable a ray casting interactive mode. The ray casting interactive mode allows a user to make selections using a “ray,” which may be a curved or straight line segment, a cone, a pyramid, a cylinder, a sphere, or another geometry whose position is guided by the user (often based on the position of at least one of the user's hands). Additional details regarding ray casting are provided in U.S. patent application Ser. No. 16 / 578,221, which is incorporated herein by reference in its entirety. For example, the ray may be a line extending from a first point between the tip of the user's index finger and the tip of the user's thumb along a line connecting the first point to a second point at a pivot point on the user's palm between the user's thumb and index finger. As with the gaze-and-gesture interaction mode, in various implementations, the ray-casting interaction mode always uses positional tracking to allow the user to control their position in the virtual reality environment by moving along the X, Y, and Z axes in the real world, while in other cases, positional tracking may or may not be enabled. In either case, the ray-casting interaction mode allows the user to make selections based on the user's ray-casts from a single position (in 3 DoF) or as the user moves around the virtual reality environment (in 6 DoF). In the ray-casting interaction mode, the user can perform actions related to the ray and / or objects intersected by the ray by making gestures such as pinching, air tapping, grabbing, etc. As with all interaction modes, in some implementations, other actions can also be performed using, for example, a controller or soft buttons, voice commands, etc. In some implementations, the ray-casting interaction mode can include a visual affordance of a certain shape (e.g., a teardrop shape) at the point between the user's fingertips from which the ray emanates. In some implementations, the shape can be resized, distorted, or otherwise changed as the user gestures, for example, the teardrop shape can compress as the user brings their thumb and fingertips closer together.Additional details regarding the ray casting interaction mode using visual affordances are provided below in connection with Figure 9. Once the interaction mode is enabled at block 586, process 500 may end at block 590.

[0076] As previously mentioned, in some implementations, process 500 may begin at block 552 instead of block 502. This may occur, for example, when a user provides an instruction to manually transition interaction modes. In various implementations, the instruction may be performed by activating a physical control, a soft control, speaking a command, performing a gesture, or some other command mapped to a particular interaction mode or cycling through interaction modes. The instruction may be received at block 554. At block 556, process 500 may determine whether the interaction mode change instruction corresponds to a no-hands 3DoF interaction mode (e.g., whether the interaction mode change instruction is mapped to that mode or is the next mode from the current mode in a cycle of modes). If the interaction mode change instruction corresponds to a no-hands 3DoF interaction mode, process 500 continues at block 580 (discussed above); if the interaction mode change instruction does not correspond to a no-hands 3DoF interaction mode, process 500 continues at block 558. At block 558, process 500 may determine whether the interaction mode change instruction corresponds to a no-hands 6 DoF interaction mode (e.g., whether the interaction mode change instruction is mapped to that mode or is the next mode from the current mode in a cycle of modes). If the interaction mode change instruction corresponds to a no-hands 6 DoF interaction mode, process 500 continues at block 582 (discussed above); if the interaction mode change instruction does not correspond to a no-hands 6 DoF interaction mode, process 500 continues at block 560. At block 560, process 500 may determine whether the interaction mode change instruction corresponds to a gaze-gesture interaction mode (e.g., whether the interaction mode change instruction is mapped to that mode or is the next mode from the current mode in a cycle of modes).If the interaction mode change instruction corresponds to the gaze-and-gesture interaction mode, process 500 continues at block 584 (discussed above); if the interaction mode change instruction does not correspond to the gaze-and-gesture interaction mode, process 500 continues at block 562. At block 562, process 500 may determine whether the interaction mode change instruction corresponds to the ray-casting interaction mode (e.g., whether the interaction mode change instruction maps to that mode or is the next mode from the current mode in a cycle of modes). If the interaction mode change instruction corresponds to the ray-casting interaction mode, process 500 continues at block 586 (discussed above); if the interaction mode change instruction does not correspond to the ray-casting interaction mode, process 500 may not implement the interaction mode change and continue to block 590, ending.

[0077] FIG. 6 is a conceptual diagram illustrating an example 600 having a gaze cursor 606. The gaze cursor 606 is a selection mechanism controlled at least by the position of the user's head and, in some implementations, further controlled by the gaze direction of the user's eyes. The example 600 further includes a user's field of view 608, which is the field of view the user can see through their virtual reality device 602. The field of view 608 includes a gaze cursor 606 projected by the virtual reality device 602 into the field of view 608 as seen by the user. The gaze cursor 606 may follow a direction indicated by a line 604 (which may or may not be visible to the user). If the line 604 is not based on eye tracking, the line 604 may be perpendicular to the coronal plane of the user's head and may intersect a point between the user's eyes. Typically, in the absence of eye tracking, the gaze cursor 606 will be centered in the user's field of view 608. In some implementations, the gaze cursor 606 can be further based on the user's tracked eye gaze direction. In eye tracking cases, the user's field of view 608 is a region based on the user's head position, but the position of the gaze cursor 606 within the field of view 608 can be controlled based on the direction of the user's gaze and / or a determination of the point within the field of view at which the user is looking. In some implementations, the gaze cursor 606 can appear at a specified distance from the user, at the user's determined focal plane, or over the nearest object (real or virtual) that the line 604 intersects. Another example of a gaze cursor is provided in connection with FIGS. 10A and 10B .

[0078] FIG. 7 is a conceptual diagram illustrating an example 700 having a dwell timer 702. The dwell timer 702 may be for a set amount of time (e.g., 3 seconds, 4 seconds, etc.) that starts when the gaze cursor 606 does not move more than a threshold amount within a threshold time (e.g., 1 second) (e.g., line 604 does not move more than 3 degrees). The dwell timer can be reset if the gaze cursor 606 moves more than the threshold amount before the dwell timer expires. In some implementations, the dwell timer 702 starts only when the gaze cursor 606 is directed at an object (e.g., object 704) that can take a default action when the dwell timer expires. In some instances, when the dwell timer starts, a visual affordance such as a ring around the gaze cursor 606 can be presented, with the ring changing color or being shaded depending on the percentage of time remaining in the dwell timer.

[0079] FIG. 8 is a conceptual diagram illustrating an example 800 with a gaze cursor 606, further utilizing hand gesture selection. In example 800, a gaze-gesture interaction mode is enabled in response to determining that the user's hand 802 is ready based on the user's hand 802 being in the user's field of view 608 and in a palm-up pose. In this gaze-gesture interaction mode, the user continues to use the gaze cursor 606 (see FIG. 6) to indicate a direction. However, instead of using the dwell timer 702 (FIG. 7) to perform an action, the user can perform a pinch gesture between their thumb and index finger to trigger an action. As indicated by line 806, such a pinch action is interpreted relative to the gaze cursor 606. For example, if the default action is to select an object, then because the gaze cursor 606 is over the object 704, the user's pinch gesture selects the object 704. When enabling gaze-gesture interaction mode, a visual affordance is displayed in the user's field of view 608 as a sphere 804 between the thumb and index finger on the user's hand 802. When the user begins to bring their thumb and index finger together, the sphere 804 may compress to an ellipsoid to indicate to the user that part of a pinch gesture is being recognized.

[0080] FIG. 9 is a conceptual diagram illustrating an example 900 using ray-casting direction input with hand gesture selection. In example 900, a ray-casting interaction mode is enabled in response to determining that a user's hand 802 is in a ray state based on the user's hand 802 being in the user's field of view 608 and in a palm-down pose. In this ray-casting interaction mode, a ray 902 is cast from a point between the user's thumb and index finger. The user can perform a pinch gesture between their thumb and index finger to trigger an action. Such a pinch action is interpreted with respect to one or more objects (e.g., object 704) intersected by the ray. For example, if the default action is to select an object, then because the ray 902 intersects with object 704, when the user performs a pinch gesture, object 704 is selected. Upon enabling the ray-casting interaction mode, a first visual affordance is displayed in the user's field of view 608 as a teardrop shape 906 between the thumb and index finger on the user's hand 802. As the user begins to bring their thumb and index finger together, the teardrop 906 may compress to indicate to the user that part of a pinch gesture is being recognized. Also, when enabling the ray casting interaction mode, a visual affordance is displayed showing the line of the ray 902 and the point 904 where the line 902 intersects with the object.

[0081] 10A and 10B are conceptual diagrams illustrating examples 1000 and 1050 of interaction between device positioning and object selection using an eye-tracking-based gaze cursor with a dwell timer. Example 1000 includes a pair of virtual reality glasses 1002 having waveguide lenses 1004A and 1004B that project light into a user's eyes to cause the user to see a field of view 1014. In some embodiments, the virtual reality glasses 1002 may be similar to the glasses 252 discussed above with reference to FIG. 2B . The orientation in the virtual reality environment for the field of view 1014 is controlled by a head tracking unit 1016 of the virtual reality glasses 1002. The head tracking unit 1016 can determine the orientation and movement of the virtual reality glasses 1002 using, for example, a gyroscope, magnetometer, and accelerometer, which can then be translated into a camera position in the virtual reality environment from which content is displayed. In example 1000, field of view 1014 displays virtual objects 1012A, 1012B, and 1012C. Virtual reality glasses 1002 include eye tracking module 1006, which includes a camera and a light source (e.g., infrared light) that illuminates a glint on the user's eyes, captures an image of the user's eyes, and uses a trained machine learning model to interpret the image into a gaze direction (shown by line 1008). Gaze cursor 1010 resides at the end of eye gaze direction line 1008. In example 1000, the user looks at the center of field of view 1014 and positions gaze cursor 1010 at the center of field of view 1014.

[0082] In example 1050, the user moves their head to the left, as indicated by arrow 1058. This movement is tracked by head tracking unit 1016, causing field of view 1014 to shift to the left, which in turn causes virtual objects 1012A, 1012B, and 1012C to shift to the right of field of view 1014. The user also shifts their gaze from the center of field of view 1014 to the lower left corner of field of view 1014. This change in eye gaze is tracked by eye tracking unit 1006, as indicated by line 1008. Thus, gaze cursor 1010 resides over virtual object 1012C. When the user holds their gaze over object 1012C for one second, dwell timer 1052 begins counting down from three seconds. This countdown is indicated to the user via visual affordances 1054 and 1056, with the portion of the shaded ring 1056 corresponding to the amount of time remaining in the dwell timer. When the dwell timer reaches zero, the virtual reality glasses 1002 take the default action of selecting object 1012C and resetting the gaze cursor 1010 on object 1012C.

[0083] References herein to "implementations" (e.g., "some implementations," "various implementations," "one implementation," "implementation," etc.) mean that a particular feature, structure, or characteristic described in connection with an implementation is included in at least one implementation of the present disclosure. The appearances of these phrases in various places herein do not necessarily all refer to the same implementation, nor are they separate or alternative implementations that are mutually exclusive of other implementations. Furthermore, various features are described that may be exhibited by some implementations and not by other implementations. Similarly, various requirements are described that may be requirements for some implementations, but are not requirements for other implementations.

[0084] As used herein, "exceeding a threshold" means that the value for the item under comparison is greater than a specified other value, that the item under comparison is between a certain specified number of items with the highest values, or that the item under comparison has a value within a specified highest percentage. As used herein, "below a threshold" means that the value for the item under comparison is less than a specified other value, that the item under comparison is between a certain specified number of items with the lowest values, or that the item under comparison has a value within a specified lowest percentage. As used herein, "within a threshold" means that the value for the item under comparison is between two other specified values, that the item under comparison is between a specified intermediate number of items, or that the item under comparison has a value within a specified intermediate percentage. Relative terms, such as "high" or "insignificant," unless otherwise defined, can be understood as assigning a value and determining how to compare that value against an established threshold. For example, the phrase "selecting a high-speed connection" can be understood to mean selecting a connection having an assigned value corresponding to its connection speed that is greater than a threshold.

[0085] Various aspects of the above techniques are described as being implemented using a hand tracking system. In individual cases, these techniques may use controller tracking instead of hand tracking. For example, if a hand tracking system is used to determine hand gestures or for an origin for casting rays, the tracked position of the controller can be used instead.

[0086] As used herein, the word "or" means any possible permutation of a set of items. For example, the phrase "A, B, or C" means at least one of A, B, C, or any combination thereof, such as a plurality of any of the following items: A; B; C; A and B; A and C; B and C; A, B, and C; or A and A; B, B, and C; A, A, B, C, and C, etc.

[0087] Although the present subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the present subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. While specific embodiments and implementations have been described herein for illustrative purposes, various modifications may be made without departing from the scope of the embodiments and implementations. The specific features and acts described above are disclosed as example forms for implementing the following claims. Accordingly, the embodiments and implementations are not limited, except as limited by the appended claims.

[0088] Any patents, patent applications, and other references mentioned above are incorporated herein by reference. Aspects can be modified as necessary to use the systems, functionality, and concepts of the various references described above, thereby providing yet other implementations. In the event that statements or subject matter in a document incorporated by reference conflict with statements or subject matter of this application, this application will control.

Claims

1. 1. A method for automatically transitioning between interaction modes for interpreting user input in a virtual reality system, comprising: identifying a first interaction mode context indicating that user location tracking input is unavailable; enabling a no-hands three-degree-of-freedom interaction mode in response to identifying the first interaction mode context; The availability of user location tracking input; and Hand tracking input is unavailable or the first tracked hand pose does not match the hand ready state identifying a second interaction mode context indicating enabling a no-hands six degrees of freedom interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that a second tracked hand pose is consistent with the hand ready state; and enabling a gaze-and-gesture interaction mode in response to identifying the third interaction mode context; identifying a fourth interaction mode context indicating that the third tracked hand pose is consistent with the lighting condition; and enabling a ray casting interactive mode in response to identifying the fourth interactive mode context; A method comprising:

2. Enabling at least one of the interaction modes further comprises: Occurs periodically or Occurs in response to changes in interaction mode context identified as a result of monitoring interaction mode context elements The method of claim 1 , responsive to an interaction mode change trigger.

3. 10. The method of claim 1, further comprising transitioning to an interaction mode based on receiving a user command to change from a current interaction mode.

4. 2. The method of claim 1 , wherein the second tracked hand pose consistent with the hand ready state includes a hand pose identified as having a user's palm facing up by at least a threshold amount, and / or the third tracked hand pose consistent with the light state includes a hand pose identified as having the user's palm facing down by at least a threshold amount.

5. The no-hands three-degree-of-freedom interactive mode is receiving a user indication of a direction based at least on the determined orientation of the user's head; receiving a user indication of an action relative to one of the directional user indications based on a dwell timer; Including, While in the no-hands three-degree-of-freedom interaction mode, the user's movements in the X, Y and Z axes are not automatically translated into movements of the field of view in the X, Y and Z axes in the virtual reality environment; or The no-hands 6-DOF interactive mode is receiving a user indication of a direction based at least on the determined orientation of the user's head; receiving a user indication of an action relative to one of the directional user indications based on a dwell timer; translating the user's movements in the X, Y and Z axes into movements of the X, Y and Z fields of view in the virtual reality environment; Including, The method of claim 1.

6. The gaze and gesture interaction mode receiving a directional user instruction based on at least a determined orientation of a user's head and based on a determined position of the user's head relative to a virtual reality environment, the determined position of the user's head being based on tracked movement of the user along an X-axis, a Y-axis, and a Z-axis; receiving a user indication of an action for one of the directional user indications by tracking a hand posture of the user and matching the hand posture with a designated action; The method of claim 1 , comprising:

7. The ray casting interactive mode includes: receiving a user indication of a direction based on ray casting by at least the virtual reality system, the ray being cast from a location associated with the tracked position of at least one hand of the user; receiving a user indication of an action on the light ray by tracking the user's hand pose and matching the hand pose with a specified action; The method of claim 1 , comprising:

8. each of the no-hands 3-DOF interaction mode, the no-hands 6-DOF interaction mode, and the gaze-and-gesture interaction mode provides a visual affordance including a gaze cursor, the gaze cursor being presented in a user's field of view and positioned at least in part based on a tracked position of the user's head; or each of the no-hands 3-DOF interaction mode, the no-hands 6-DOF interaction mode, and the gaze and gesture interaction mode provides a visual affordance including a gaze cursor; the gaze cursor is presented in the user's field of view and positioned based at least in part on the tracked position of the user's head and on the tracked gaze direction of the user; The method of claim 1.

9. 9. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform the method of any one of claims 1 to 8, i.e., to perform a process for transitioning between interaction modes for interpreting user input in a virtual reality system. A computer-readable storage medium.

10. The no-hands 3-DOF interactive mode and the no-hands 6-DOF interactive mode are receiving a user indication of a direction based at least on the determined orientation of the user's head; receiving a user indication of an action relative to one of the directional user indications based on a dwell timer; Including, if user position tracking input is unavailable, the no-hands three degrees of freedom interaction mode does not translate the user's movements in the X, Y and Z axes into X, Y and Z field of view movements in the virtual reality environment; If user position tracking input is available, the no-hands six degrees of freedom interaction mode automatically translates the user's movements in the X, Y and Z axes into X, Y and Z field of view movements in the virtual reality environment.

10. The computer-readable storage medium of claim 9.

11. the gaze-gesture interaction mode provides a visual affordance including a sphere presented in the user's field of view positioned between the user's thumb and one of the other fingers; the sphere is shown resized or distorted according to the determined distance between the user's thumb and the other finger, or the ray casting interaction mode provides a visual affordance including a shape shown in the user's field of view positioned between the user's thumb and one of the other fingers; the shape is shown resized or distorted according to the determined distance between the user's thumb and the other fingers; 10. The computer-readable storage medium of claim 9.

12. 10. The computer-readable storage medium of claim 9, wherein the second tracked hand pose consistent with the hand ready state comprises a hand pose identified as a user's palm facing up.

13. The gaze and gesture interaction mode receiving a directional user indication based on at least a determined orientation of a user's head and based on a determined position of the user's head relative to the virtual reality environment, the determined position of the user's head being based on tracked movement of the user along an X-axis, a Y-axis, and a Z-axis; receiving a user indication of an action for one of the directional user indications by tracking a hand posture of the user and matching the hand posture with a designated action; 10. The computer-readable storage medium of claim 9, comprising:

14. 1. A computing system for interpreting user input in a virtual reality system, comprising: one or more processors; one or more memories for storing instructions; which instructions, when executed by the one or more processors, cause the computing system to perform the method of any one of claims 1 to 8, i.e., to perform a process for transitioning between interaction modes. Computational system.

15. the third tracked hand pose consistent with the lighting condition includes a hand pose identified as having the user's palm pointing down by at least a threshold amount; or the no-hands 3-DOF interaction mode, the no-hands 6-DOF interaction mode, and the gaze-and-gesture interaction mode provide a visual affordance including a gaze cursor, the gaze cursor being presented in the user's field of view and positioned at least in part based on the tracked position of the user's head; 15. The computing system of claim 14.

Citation Information

Patent Citations

  • Control equipment operation gesture recognition device; control equipment operation gesture recognition system, and control equipment operation gesture recognition program

    JP2009037434A

  • Enhanced gesture-based image manipulation

    JP2014099184A

  • Terminal device and program

    JP2018198075A

  • Gesture control system and method for smart home

    JP2018516422A

  • Interaction mode selection based on detected distance between user and machine interface

    JP2019533846A