Integration of artificial reality interaction modes
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-24
- Publication Date
- 2026-08-14
AI Technical Summary
虽然一些系统使用多种交互模式,但是它们通常需要用户手动选择以在多种模式之间进行切换,或无法自动切换到适合于当前上下文的交互模式
Smart Images

Figure CN116097209B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an artificial reality environment that can switch between different interaction modes based on the context of the artificial reality device. Background Technology
[0002] Various artificial reality systems can be embodied in wearable devices, systems with external displays and sensors, and passthrough systems. Numerous interaction modes have been created for these artificial reality systems. Each of these modes typically provides the user with a way to indicate direction and / or position, and a way to take actions related to that indicated direction and / or position. For example, some interaction modes allow users to move with three degrees of freedom (3DoF – typically allowing directions representing the pitch, roll, and yaw axes, but no X, Y, or Z axis movement), while others allow six degrees of freedom (6DoF – typically allowing movement along the pitch, roll, yaw, X, Y, and Z axes). Some interaction modes use controllers, others track hand position and posture, and still others track only head movements. Interaction modes can be based on motion sensors, various types of cameras, time-of-flight sensors, timers, etc. However, many AI systems use only one interaction mode, suitable for specific situations but difficult or unsuitable for others, despite all available options and possible combinations of interaction modes. For example, users find it frustrating to use only head-tracking interaction modes when performing complex interactions. Additionally, users generally do not want to use controllers or hand-tacking input when performing simple interactions or in public places. While some systems use multiple interaction modes, they often require users to manually select and switch between modes, or do not automatically switch to the appropriate interaction mode for the current context. Summary of the Invention
[0003] Therefore, the present invention relates to methods, computer-readable storage media, and computing systems.
[0004] In one aspect, the present invention relates to a method for automatically switching between multiple interaction modes in an artificial reality system to interpret user input, the method comprising: identifying a first interaction mode context, the first interaction mode context indicating that user position tracking input is unavailable; in response to identifying the first interaction mode context, enabling a handless three-degree-of-freedom interaction mode; identifying a second interaction mode context, the second interaction mode context indicating that user position tracking input is available and hand tracking input is unavailable, or that a tracked first hand gesture does not match a hand-ready state; in response to identifying the second interaction mode context, enabling a handless six-degree-of-freedom interaction mode; identifying a third interaction mode context, the third interaction mode context indicating that a tracked second hand gesture matches a hand-ready state; in response to identifying the third interaction mode context, enabling a gaze and gesture interaction mode; identifying a fourth interaction mode context, the fourth interaction mode context indicating that a tracked third hand gesture matches a ray state; and in response to identifying the fourth interaction mode context, enabling a raycasting interaction mode.
[0005] In embodiments of the method according to the invention, enabling at least one of a plurality of interaction modes may also be in response to an interaction mode change trigger, which occurs periodically or in response to a change in the interaction mode context, the change in the interaction mode context being identified as a result of monitoring interaction mode context factors.
[0006] In an embodiment of the method according to the present invention, the method may further include: switching to an interaction mode based on receiving a user instruction to change the current interaction mode.
[0007] In an embodiment of the method according to the invention, the tracked second hand gesture matching the hand-ready state may include: a hand gesture identified as the user's palm facing upward by at least a threshold amount.
[0008] In an embodiment of the method according to the invention, the tracked third hand gesture matched with the ray state may include: a hand gesture identified as the user's palm facing down by at least a threshold amount.
[0009] In an embodiment of the method according to the invention, the handless three-degree-of-freedom interaction mode may include: receiving a plurality of user orientation indications based at least on the determined head orientation of the user; and receiving a user action indication associated with one of the plurality of user orientation indications based on a dwell timer, wherein, when in the handless three-degree-of-freedom interaction mode, user motion on the X, Y, and Z axes is not automatically converted into field-of-view motion on the X, Y, and Z axes in an artificial reality environment.
[0010] In an embodiment of the method according to the invention, the handless six-degree-of-freedom interaction mode may include: receiving multiple user orientation indications based at least on the determined user's head orientation; receiving a user action indication associated with one of the multiple user orientation indications based on a dwell timer; and converting user motion on the X, Y, and Z axes into field-of-view motion on the X, Y, and Z axes in an artificial reality environment.
[0011] In an embodiment of the method according to the invention, the gaze and gesture interaction mode may include: receiving user orientation instructions based at least on the determined head orientation of the user and based on the determined position of the user's head relative to the artificial reality environment, wherein the determined head position of the user is based on the tracked movement of the user along the X-axis, Y-axis and Z-axis; and receiving user action instructions associated with one of the plurality of user orientation instructions by tracking the user's hand gestures and matching the hand gestures with specified actions.
[0012] In an embodiment of the method according to the invention, the ray-projection interaction mode may include: receiving multiple user orientation indications based at least on a ray projected by an artificial reality system, wherein the ray is projected from a position associated with the tracking position of at least one of the user's hands; and receiving ray-related user action indications by tracking the user's hand gestures and matching the hand gestures with specified actions.
[0013] In embodiments of the method according to the invention, each of the handless three-degree-of-freedom interaction mode, the handless six-degree-of-freedom interaction mode, and the gaze and gesture interaction mode can provide visual functional visibility including a gaze cursor, wherein the gaze cursor is displayed in the user's field of vision and located at least in part based on the tracking position of the user's head.
[0014] In embodiments of the method according to the invention, each of the handless three-degree-of-freedom interaction mode, the handless six-degree-of-freedom interaction mode, and the gaze and gesture interaction mode can provide visual functional visibility including a gaze cursor, which can be displayed in the user's field of vision, located at least in part based on the tracking position of the user's head and based on the user's tracking gaze direction.
[0015] In one aspect, the invention also relates to a computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform the method described above, or to perform a process for switching between multiple interaction modes in an artificial reality system to interpret user input, the process comprising: identifying a first interaction mode context indicating that hand-tracking input is unavailable or that a tracked first hand gesture does not match a hand-ready state; enabling a handless interaction mode in response to identifying the first interaction mode context; identifying a second interaction mode context indicating that a tracked second hand gesture matches a hand-ready state; enabling a gaze and gesture interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that a tracked third hand gesture matches a ray state; and enabling a raycasting interaction mode in response to identifying the third interaction mode context.
[0016] In an embodiment of a computer-readable storage medium, the hands-free interaction mode may include: receiving a plurality of user orientation indications at least based on a determined user head orientation; and receiving a user action indication associated with one of the plurality of user orientation indications based on a dwell timer, wherein: when user position tracking input is unavailable, the hands-free interaction mode does not convert user motion on the X, Y, and Z axes into field-of-view motion on the X, Y, and Z axes in an artificial reality environment; and when user position tracking input is available, the hands-free interaction mode automatically converts user motion on the X, Y, and Z axes into field-of-view motion on the X, Y, and Z axes in an artificial reality environment.
[0017] In embodiments of a computer-readable storage medium, gaze and gesture interaction modes can provide visual functional visibility including a sphere shown in the user's field of vision, positioned between the user's thumb and another finger; and the sphere can be shown to adjust its size or shape according to a defined distance between the user's thumb and the other finger.
[0018] In embodiments of a computer-readable storage medium, the ray-projection interaction mode can provide visual functional visibility, which includes a shape shown in the user's field of vision, located between the user's thumb and the other fingers; and the shape can be shown to adjust its size or shape according to a defined distance between the user's thumb and the other fingers.
[0019] In an embodiment of a computer-readable storage medium, the tracked second hand gesture matching the hand-ready state may include: a hand gesture identified as the user's palm facing upwards.
[0020] In an embodiment of a computer-readable storage medium, the gaze and gesture interaction mode may include: receiving multiple user orientation indicators based at least on a determined user head orientation and based on the determined position of the user's head relative to the artificial reality environment, wherein the determined user head position is based on the tracked movement of the user along the X, Y, and Z axes; and receiving a user action indicator associated with one of the multiple user orientation indicators by tracking the user's hand gestures and matching the hand gestures with a specified action.
[0021] In one aspect, the present invention relates to a computing system for interpreting user input in an artificial reality system, the computing system comprising: one or more processors; and one or more memories storing instructions that, when executed by the one or more processors, cause the computing system to perform the methods described above or to perform a process for switching between multiple interaction modes, the process comprising: identifying a first interaction mode context indicating that hand-tracking input is unavailable or that a tracked first hand gesture does not match a hand-ready state; enabling a handless interaction mode in response to identifying the first interaction mode context; identifying a second interaction mode context indicating that a tracked second hand gesture matches a hand-ready state; enabling a gaze and gesture interaction mode in response to identifying the second interaction mode context; identifying a third interaction mode context indicating that a tracked third hand gesture matches a ray state; and enabling a raycasting interaction mode in response to identifying the third interaction mode context.
[0022] In an embodiment of the computing system, the tracked third hand gesture matched with the ray state may include: a hand gesture identified as the user's palm facing down by at least a threshold amount.
[0023] In embodiments of the computing system, handless interaction modes and gaze and gesture interaction modes can provide visual functional visibility including a gaze cursor, which is displayed in the user's field of view and located at least in part based on the tracking position of the user's head. Attached Figure Description
[0024] Figure 1 This is a block diagram illustrating an overview of a device on which some embodiments of the present technology can be operated.
[0025] Figure 2A This is a wireframe diagram illustrating a virtual reality head-mounted viewer that can be used in some embodiments of the present technology.
[0026] Figure 2B This is a wireframe diagram illustrating a mixed reality head-mounted viewer that can be used in some embodiments of the present technology.
[0027] Figure 3 This is a block diagram illustrating an overview of the environment in which some embodiments of the present technology can be operated.
[0028] Figure 4 This is a block diagram illustrating components that can be used in a system employing the disclosed technology in some embodiments.
[0029] Figure 5 This is a flowchart illustrating a process used in some embodiments of the present technology to identify triggers for switching between specified interaction modes and to enable the triggered interaction mode.
[0030] Figure 6 This is a conceptual diagram showing an example of a gaze cursor.
[0031] Figure 7 This is a conceptual diagram showing an example gaze cursor and a dwell timer.
[0032] Figure 8 This is a conceptual diagram showing an example of cursor gaze and hand gesture selection.
[0033] Figure 9 This is a conceptual diagram illustrating an example of ray projection direction input and gesture selection.
[0034] Figure 10A and Figure 10B This is a conceptual diagram illustrating an example interaction between device localization and object selection using a gaze cursor based on eye tracking and a dwell timer.
[0035] The technology described herein can be better understood by referring to the following detailed description in conjunction with the accompanying drawings, in which the same reference numerals denote the same or similarly functional elements. Detailed Implementation
[0036] This disclosure relates to an interaction mode system that provides multiple interaction modes in an artificial reality environment, with automatic, context-specific transitions between these modes. The interaction modes can specify how the system determines (a) directional indication and movement within the artificial reality environment, and (b) interactions for making selections or performing other actions. The system also provides functional visibility indicating the current interaction mode and interaction mode transitions. The system can control multiple interaction modes and can employ a mapping from interaction mode context factors to interaction modes to control transitions between them. Additional interaction mode transitions can also be triggered by the user manually activating software or hardware controls, performing interaction mode transition gestures, or other explicit commands.
[0037] Interaction mode context factors may include, for example, which hardware components are enabled or disabled (via hardware or software control); which software modules are enabled or disabled; current privacy, power, or accessibility settings; environmental conditions (e.g., lighting, location detection factors such as surface type or marker availability, area-based camera restrictions, the type of area identified, the number of people in the vicinity, or specific people in the vicinity identified); current body position (e.g., whether the hand is in the field of vision, hand orientation, gesture being performed); current controller position or state; and so on.
[0038] This interaction mode system determines the interaction mode context by integrating inputs from various subsystems, such as head tracking systems, position tracking systems, eye tracking systems, hand tracking systems, other camera-based systems, hardware controllers, simultaneous localization and mapping (SLAM) systems, or other data sources. The head tracking subsystem can identify the user's head orientation, for example, based on one or more of the following: data from an inertial motion unit (IMU), tracking light emitted from the head-mounted device, a camera or other external sensors guided to acquire representations of the user's head, or a combination thereof. The position tracking subsystem can identify the user's location in an artificial reality environment, for example, based on tracking light emitted from a wearable device, a time-of-flight sensor, environmental mapping based on conventional and / or depth camera data, IMU data, Global Positioning System (GPS) data, etc. An eye-tracking subsystem can identify the gaze direction relative to a head-mounted device by determining the gaze direction vector, for example, based on factors such as an image of a ring of light reflected from the user's eyes (i.e., a "flash"), through modeling one or both eyes (typically along a line connecting the user's eye socket and the center of the user's pupil). A hand-tracking subsystem can determine the position, posture, and / or movement of the user's hands, for example, based on images depicting one and / or both hands from one or more cameras and / or sensor data from devices worn by the user (e.g., gloves, rings, wristbands, etc.). Information from other camera systems can collect additional information, such as lighting conditions, environmental mapping, etc. Each of these subsystems can provide outcome data (e.g., head orientation, position, eye orientation, hand posture and position, body posture and position, environmental data, mapping data, etc.) and subsystem metadata, such as whether various aspects of the subsystem are enabled and whether the subsystem has access to certain types of data (e.g., whether certain sets of cameras or other sensors are enabled, battery level, accuracy estimates of the relevant outcome data, etc.).
[0039] In some implementations, the interaction mode system can control at least four interaction modes, including a hands-free 3DoF mode, a hands-free 6DoF mode, a gaze and gesture mode, and a raycasting mode. The hands-free 3DoF interaction mode can only track three degrees of freedom indicating roll, yaw, and pitch (although the user can indicate X, Y, and Z positional movement via other devices such as a joystick on a controller). 3DoF tracking can be equivalent to the direction of a vector having an origin at the artificial reality device and extending to a point on a sphere surrounding that origin, where the point is identified based on IMU data. In some implementations, this vector can be modified based on the identified user gaze direction. Based on this vector, a "gaze cursor" can be set in the user's field of view, either centered in the field of view when eye tracking is unavailable, or positioned according to the user's eye focus when eye tracking is available. In a hands-free 3DoF interaction mode, the system can use a dwell timer that begins counting down if the direction vector does not change a threshold within a certain time period and resets when a change in the threshold is detected. When the dwell timer counts down to zero, this can indicate a user action. This action, combined with the vector direction, can induce selection of an object represented by a gaze cursor or other actuations. For additional details on gaze cursors and dwell timers configured according to this technology, see [link to relevant documentation]. Figure 6 , Figure 10A and Figure 10B And the related description below. Interaction mode context factors that can trigger a switch to a hands-free 3DoF interaction mode may include factors such as when the location tracking camera is off or otherwise unavailable (e.g., when in low power mode, privacy mode, public mode, etc.), or when lighting conditions make reliable location tracking and hand tracking impossible.
[0040] Hands-free 6DoF interaction modes can be similar to hands-free 3DoF interaction modes, except that the hands-free 6DoF interaction mode system tracks not only the direction vector of the gaze cursor but also the user's positional movement within the artificial reality environment. As in hands-free 3DoF interaction mode, actions in hands-free 6DoF interaction mode can include guiding the gaze cursor (which can now be done by the user changing its direction and position) and using a dwell timer to perform actions associated with the gaze cursor. Interaction mode context factors that can trigger a transition to hands-free 6DoF interaction mode can include a representation that the position tracking system is enabled but (a) the hand tracking system is not enabled or the lighting conditions are insufficient to track hand position, or (b) no representation of a hand "ready state" is detected. A ready state can occur when the user's hand is within the user's field of vision, raised above a threshold position, and / or in a specific configuration (e.g., palm up).
[0041] The gaze and gesture interaction mode can be a form of 6DoF or 3DoF interaction mode, where orientation is based on the gaze cursor, but instead of (or in addition to) using a dwell timer for actions, the user can also use hand gestures to indicate actions. For example, when the hand is in the user's field of vision and the palm is facing up, an action can be indicated by making a thumb and finger "pinch" gesture. Other gestures can also be used, such as air taps, air swipes, grabs, or other recognizable movements or gestures. Interaction mode context factors that can trigger a transition to a gaze and gesture interaction mode can include enabling the hand tracking system and recognizing that the user's hand is in a ready state. In some implementations, the transition to a gaze and gesture interaction mode may also require the availability of a position tracking system for 6DoF motion. See [link to details on gaze and gesture interaction modes] for further information. Figure 8 .
[0042] Raycasting interaction modes can use a ray instead of a gaze cursor, extending from the user's hand into the artificial reality environment. For example, this ray can be specified along a line connecting the origin (e.g., the center of mass of the hand or the user's eye, shoulder, or hip) to a control point (e.g., a point relative to the user's hand). Similar hand gestures used in gaze and gesture interaction modes can be used to indicate actions. Interaction mode context factors that can trigger a transition to raycasting interaction modes may include the availability of a hand tracking system and the user's hand being in a raycasting pose (e.g., (a) located within the user's field of vision or positioned above a threshold, and (b) positioned with the palm facing down). For additional details on raycasting interaction modes, see [link to relevant documentation]. Figure 9 And the related description below.
[0043] In some implementations, additional controls or other controls may be added to replace the direction determination and / or selection determination in any of the above interaction modes. For example, a user may use a controller and / or buttons on the controller to indicate direction, to indicate a selection action or other action, or a paired device may be used to indicate direction or action (e.g., by swiping gestures or taps on a touchscreen, or by performing a scrolling action on a capacitive sensor such as a watch or ring).
[0044] Embodiments of the disclosed technology may include artificial reality systems or combinations thereof. Artificial reality, or extra reality (XR), is a form of reality that has been adjusted in some way before being presented to a user; it may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., photographs of the real world). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that provides a three-dimensional effect to the viewer). Additionally, in some embodiments, artificial reality may be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or for use in artificial reality (e.g., performing activities in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, stand-alone HMDs, mobile devices or computing systems, "cave-like" environments or other projection systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.
[0045] As used herein, “virtual reality” or “VR” refers to an immersive experience in which the user’s visual input is controlled by a computing system. “Augmented reality” or “AR” refers to a system from which a user views images of the real world after they have passed through a computing system. For example, a tablet with a camera on the back can capture images of the real world and then display them on a screen on the side of the tablet opposite the camera. The tablet can process and adjust or “enhance” these images as they pass through the system (e.g., by adding virtual objects). “Mixed reality” or “MR” refers to a system in which light entering the user’s eyes is partly generated by a computing system and partly composed of light reflected from objects in the real world. For example, an MR headset can be shaped as a pair of glasses with a pass-through display, allowing light from the real world to pass through a waveguide that simultaneously emits light from a projector in the MR headset, allowing the MR headset to present virtual objects that blend with real objects that the user can see. As used herein, “artificial reality,” “additional reality,” or “XR” refers to any of VR, AR, MR, or any combination or hybrid thereof.
[0046] Existing XR systems have developed numerous interaction modes, including features such as raycasting, air clicks, dwell timers, and gaze cursors. However, existing XR systems cannot use these features to provide automatic switching between interaction modes. Either the user needs to manually change the interaction mode, using the same mode for all interactions with a particular application, or the application needs to specify the interaction mode regardless of context. Problems arise when the current interaction mode is unsuitable for the current context, such as when conditions do not allow for proper functioning of various aspects of the current interaction mode (e.g., when the interaction mode relies on manual detection and the hand is not in the field of vision, lighting conditions are poor, and / or the hardware used to acquire and interpret hand gestures is disabled), when another interaction mode is less disruptive or provides performance benefits (e.g., when the user's hand is otherwise occupied, or when making hand gestures is socially awkward, when the device is low-power and a lower-power-requirement interaction mode is desired, etc.), or when the current interaction mode is suboptimal for the interaction to be performed (e.g., when the current interaction mode uses a time-intensive dwell timer and continuous selection is to be performed).
[0047] The interaction mode system and process described herein promise to overcome these problems associated with traditional XR interaction technologies and promises to (a) provide more powerful functionality while being more natural and intuitive than interactions in existing XR systems; (b) provide new efficiencies by selecting interaction modes that match current power and environmental conditions; and (c) make artificial reality interactions less disruptive to users and others in the surrounding environment. Despite being natural and intuitive, the systems and processes described herein are rooted in computerized artificial reality systems, rather than simulations of traditional object interactions. For example, existing object interaction technologies cannot provide the set of interaction modes described herein, as well as the smooth and automatic transitions between the various modes. Furthermore, the interaction mode system and process described herein provide improvements by making interactions available in additional situations, such as when existing XR systems already provide interaction modes that are inoperable or that the user cannot interact with in the current context (e.g., when under low light conditions or when the user's hands are otherwise occupied, by switching to a more suitable interaction mode).
[0048] Several implementation methods are described in more detail below with reference to the accompanying drawings. For example, Figure 1 This is a block diagram illustrating an overview of devices on which some embodiments of the disclosed technology may operate. These devices may include hardware components of a computing system 100 that switch between interaction modes based on factors in the context of an identified interaction mode or in response to user control. In various embodiments, the computing system 100 may include a single computing device 103 or multiple computing devices (e.g., computing devices 101, 102, and 103) that communicate via wired or wireless channels to distribute processing and share input data. In some embodiments, the computing system 100 may include a standalone head-mounted viewer capable of providing a user with a computer-created or computer-enhanced experience without requiring external processing or sensors. In other embodiments, the computing system 100 may include multiple computing devices, such as a head-mounted viewer and a core processing unit (e.g., a console, mobile device, or server system), wherein some processing operations are performed on the head-mounted viewer, while other processing operations are offloaded to the core processing unit. The following discusses… Figure 2A and Figure 2B An example head-mounted viewer is described. In some implementations, location and environmental data may be collected solely by sensors included in the head-mounted viewer device, while in other implementations, one or more non-head-mounted viewer computing devices may include sensor components capable of tracking environmental or location data.
[0049] The computing system 100 may include one or more processors 110 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a holographic processing unit (HPU), etc.). The processor 110 may be a single processing unit or multiple processing units in a device, or it may be distributed across multiple devices (e.g., distributed across two or more computing devices 101 to 103).
[0050] The computing system 100 may include one or more input devices 120 that provide input to the processor 110, informing them of actions. These actions may be coordinated by a hardware controller that interprets the signals received from the input devices and transmits information to the processor 110 using a communication protocol. Each input device 120 may include, for example, a mouse, keyboard, touchscreen, touchpad, wearable input device (e.g., haptic gloves, bracelets, rings, earrings, necklaces, watches, etc.), camera (or other light-based input device, such as an infrared sensor), microphone, or other user input device.
[0051] Processor 110 may be coupled to other hardware devices, such as a Peripheral Component Interconnect (PCI) bus or a Small Computer System Interface (SCSI) bus, using an internal bus, an external bus, or a wireless connection. Processor 110 may communicate with a hardware controller used by a device (e.g., display 130). Display 130 may be used to display text and graphics. In some embodiments, display 130 includes an input device as part of the display, such as when the input device is a touchscreen or equipped with an eye-tracking orientation monitoring system. In some embodiments, the display is separate from the input device. Examples of such display devices are: Liquid Crystal Display (LCD) screens, Light-Emitting Diode (LED) screens, projectors, holographic or augmented reality displays (e.g., heads-up displays or head-mounted displays), etc. Other input / output (I / O) devices (also referred to as other I / O) 140 may also be coupled to the processor. These other I / O devices include, for example, network chips or network cards, video chips or graphics cards, audio chips or sound cards, Universal Serial Bus (USB), FireWire or other external devices, cameras, printers, speakers, Compact Disc Read-only memory (CD-ROM) drives, Digital Video Disc (DVD) drives, disk drives, etc.
[0052] The computing system 100 may include a communication device capable of wireless or wired communication with other local computing devices or network nodes. This communication device may use, for example, a Transmission Control Protocol / Internet Protocol (TCP / IP) protocol to communicate with another device or server over a network. The computing system 100 may utilize the communication device to distribute multiple operations across multiple network devices.
[0053] Processor 110 can access memory 150, which may be located on one of the computing devices of computing system 100, or may be distributed across multiple computing devices or other external devices of computing system 100. Memory includes one or more hardware devices for volatile or non-volatile storage, and may include read-only and writable memory. For example, memory may include one or more of random access memory (RAM), various cache memories, CPU registers, read-only memory (ROM), and writable non-volatile memory, such as flash memory, hard disk drive, floppy disk, optical disc (CD), digital video disc (DVD), magnetic storage devices, magnetic tape drives, etc. Memory is not a propagating signal detached from the underlying hardware; therefore, memory is non-transitory. Memory 150 may include program memory 160, which stores programs and software, such as operating system 162, interactive mode system 164, and other application programs 166. The memory 150 may also include a data memory 170, which may include, for example, a mapping of interaction mode context factors to interaction modes (for triggering interaction mode switching), configuration data, settings, user options or preferences, etc., which may be provided to the program memory 160 of the computing system 100 or any other component.
[0054] Some implementations can operate with many other computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with this technology include, but are not limited to: XR head-mounted displays, personal computers, server computers, handheld or laptop devices, cellular phones, wearable electronic devices, game consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, networked personal computers (PCs), microcomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.
[0055] Figure 2AThis is a block diagram of a virtual reality head-mounted display (HMD) 200 according to some embodiments. The HMD 200 includes a front rigid body 205 and a strap 210. The front rigid body 205 includes one or more electronic display elements of an electronic display 245, an inertial motion unit (IMU—also called an inertial measurement unit) 215, one or more position sensors 220, a locator 225, and one or more computing units 230. The position sensors 220, IMU 215, and computing units 230 may be located inside the HMD 200 and may be invisible to the user. In various embodiments, the IMU 215, position sensors 220, and locator 225 may track the motion and position of the HMD 200 in the real world and in a virtual environment with three degrees of freedom (3DoF) or six degrees of freedom (6DoF). For example, the locator 225 may emit an infrared beam that produces light spots on real objects around the HMD 200. As another example, IMU 215 may include, for example, one or more accelerometers, gyroscopes, magnetometers, other non-camera-based position sensors, force sensors, or orientation sensors, or combinations thereof. One or more cameras (not shown) integrated with HMD 200 can detect light spots. Computational unit 230 in HMD 200 can use the detected light spots to infer the position and motion of HMD 200, as well as to identify the shape and location of real objects around HMD 200.
[0056] The electronic display 245 may be integrated with the front rigid body 205 and may provide image light to the user as instructed by the computing unit 230. In various embodiments, the electronic display 245 may be a single electronic display or multiple electronic displays (e.g., one display for each eye of the user). Examples of the electronic display 245 include: liquid crystal displays (LCDs), organic light-emitting diode (OLED) displays, active-matrix organic light-emitting diode (AMOLED) displays, displays comprising one or more quantum dot light-emitting diode (QOED) subpixels, projector units (e.g., micro-LEDs, lasers, etc.), some other display, or some combination thereof.
[0057] In some implementations, HMD 200 may be coupled to a core processing unit, such as a personal computer (PC) (not shown), and / or one or more external sensors (not shown). The external sensors may monitor HMD 200 (e.g., via light emitted from HMD 200), and the PC may use the HMD, in conjunction with outputs from IMU 215 and position sensor 220, to determine the position and movement of HMD 200.
[0058] In some embodiments, the HMD 200 can communicate with one or more other external devices, such as controllers (not shown) that can be held by the user with one or both hands. Multiple controllers may have their own IMU units, position sensors, and / or emit additional light spots. The HMD 200 or external sensors can track these controller light spots. The computing unit 230 or core processing unit in the HMD 200 can use this tracking, combined with the IMU and position output, to monitor the user's hand position and movement. The controller may also include various buttons that the user can actuate to provide input and interact with virtual objects. In various embodiments, the HMD 200 may also include additional subsystems, such as eye-tracking units, audio systems, various networking components, etc. In some embodiments, alternative to the controller, or in addition to the controller, one or more cameras, either in or outside the HMD 200, can monitor the user's hand position and posture to determine gestures and other hand and body movements.
[0059] Figure 2B This is a block diagram of a mixed reality HMD system (also referred to as an HMD system) 250, including a mixed reality HMD 252 and a core processing unit 254. The mixed reality HMD 252 and the core processing unit 254 can communicate via a wireless connection, such as link 256 (e.g., a 60 GHz link). In other embodiments, the mixed reality HMD system 250 may include only a head-mounted display without an external computing device, or may include other wired or wireless connections between the mixed reality HMD 252 and the core processing unit 254. The mixed reality HMD 252 includes a pass-through display 258 and a frame 260. The frame 260 may accommodate various electronic components (not shown), such as projectors (e.g., lasers, LEDs, etc.), cameras, eye-tracking sensors, micro-electro-mechanical system (MEMS) components, network components, etc.
[0060] A projector can be coupled to a pass-through display 258, for example, via optical elements, to display media to a user. The optical elements may include one or more waveguide components, reflectors, lenses, mirrors, collimators, gratings, etc., for guiding light from the projector toward the user's eyes. Image data can be transmitted from the core processing unit 254 to the HMD 252 via link 256. A controller in the HMD 252 can convert the image data into light pulses from the projector, which can be transmitted as output light to the user's eyes via the optical elements. The output light can be mixed with light passing through the pass-through display 258, allowing the output light to render virtual objects that appear as if they exist in the real world.
[0061] Similar to HMD 200, HMD system 250 may also include motion and position tracking units, cameras, light sources, etc., which allows HMD system 250 to track itself, for example in 3DoF or 6DoF, track various parts of the user (e.g., hands, feet, head or other body parts), map virtual objects to appear stationary when HMD 252 moves, and make virtual objects react to gestures and other real-world objects.
[0062] Figure 3 This is a block diagram illustrating an overview of an environment 300 in which some embodiments of the disclosed technology may operate. Environment 300 may include one or more client computing devices 305A to 305D, examples of which may include computing system 100. In some embodiments, some of the client computing devices (e.g., client computing device 305B) may be HMD 200 or HMD system 250. Client computing device 305 may operate in a network environment using a logical connection via network 330 to one or more remote computers (e.g., server computing devices).
[0063] In some implementations, server (also referred to as server computing device) 310 may be an edge server that receives client requests and coordinates the fulfillment of these requests through other servers (e.g., servers (also referred to as server computing devices) 320A to 320C). Server computing devices 310 and 320 may include computing systems, such as computing system 100. Although each server computing device 310 and 320 is logically presented as a single server, each server computing device may be a distributed computing environment comprising multiple computing devices located in geographically similar or different physical locations.
[0064] Client computing device 305 and server computing devices 310 and 320 can each act as a server or client for one or more other server / client devices. Server 310 can connect to database (DB) 315. Servers 320A to 320C can each connect to their respective databases (DBs) 325A to 325C. As described above, each server 310 or 320 can correspond to a set of servers, and each of these servers can share a database or have its own database. Although databases 315 and 325 are logically presented as a single unit, databases 315 and 325 can each be a distributed computing environment comprising multiple computing devices, which can reside within their respective servers or in geographically similar or different physical locations.
[0065] Network 330 can be a Local Area Network (LAN), Wide Area Network (WAN), Mesh Network, Hybrid Network, or other wired or wireless network. Network 330 can be the Internet or some other public or private network. Client computing device 305 can connect to network 330 via a network interface (e.g., via wired or wireless communication). Although the connection between server 310 and server 320 is shown as a separate connection, these connections can be any kind of LAN, WAN, wired network, or wireless network, including network 330 or a separate public or private network.
[0066] Figure 4This is a block diagram illustrating multiple components 400 that can be used in systems employing the disclosed technology in some embodiments. The multiple components 400 may be included in a single device of the computing system 100 or distributed across multiple devices within various devices of the computing system 100. The multiple components 400 include hardware 410, intermediaries 420, and dedicated components 430. As described above, systems implementing the disclosed technology can use various hardware, including a processing unit 412, working memory 414, input and output (I / O) devices (also referred to as I / O) 416 (e.g., a camera, display, IMU unit, network connection, etc.), and storage memory 418. In various embodiments, storage memory 418 may be one or more of a local device, an interface to a remote storage device, or a combination thereof. For example, storage memory 418 may be one or more hard disk drives or flash drives accessible via a system bus, or it may be a cloud storage provider (e.g., in storage device 315 or 325) or other network storage devices accessible via one or more communication networks. In various implementations, the multiple components 400 may be implemented in a client computing device (e.g., client computing device 305) or in a server computing device (e.g., server computing device 310 or 320).
[0067] The mediator 420 may include a component that mediates resources between the hardware 410 and the dedicated component 430. For example, the mediator 420 may include an operating system, services, drivers, a Basic Input Output System (BIOS), controller circuitry, or other hardware or software systems.
[0068] Dedicated component 430 may include software or hardware configured to perform operations for switching between multiple interaction modes based on (a) factors in the context of an identified interaction mode or (b) in response to user control. Dedicated component 430 may include a mode selector 434, a hands-free 3DoF mode controller 436, a hands-free 6DoF mode controller 438, a gaze and gesture mode controller 440, a raycasting mode controller 442, and a component and application programming interface (API) (e.g., interface 432) that can be used to provide a user interface, transfer data, and control the dedicated component. In some embodiments, the multiple components 400 may reside in a computing system distributed across multiple computing devices, or may be an interface to a server-based application that executes one or more of the multiple dedicated components 430. Although depicted as separate components, dedicated components 430 may be logically or otherwise non-physical functional distinctions, and / or may be submodules or blocks of code of one or more applications.
[0069] The mode selector 434 can identify conditions for switching interaction modes, such that control over user interaction is switched to a corresponding one of the following: a hands-free 3DoF mode controller 436, a hands-free 6DoF mode controller 438, a gaze and gesture mode controller 440, or a raycasting mode controller 442. In various embodiments, the mode selector 434 can identify such conditions in response to an explicit user selection of an interaction mode or by mapping interaction mode context factors to corresponding interaction modes. This mapping can, for example, map an interaction mode context factor with no available tracking input to a hands-free 3DoF interaction mode; map an interaction mode context factor with available tracking input and an identification that the hand is not in a ready state to a hands-free 6DoF interaction mode; map an interaction mode context factor with available tracking input and an identification that the hand is in a ready state to a gaze and gesture interaction mode; and map an interaction mode context factor with available tracking input and a hand in a raycasting state to a raycasting interaction mode. Additional details regarding the selection of interaction modes based on interaction mode context factors or user selection are described below. Figure 5 Boxes 502 to 510 and boxes 552 to 562 are provided.
[0070] The hands-free 3DoF mode controller 436 can provide user interaction, where a gaze cursor indicates direction and a dwell timer indicates action. The hands-free 3DoF mode controller 436 does not translate user movement on the X, Y, or Z axes into changes in the user's field of view. The hands-free 3DoF mode controller 436 can be connected to the IMU unit of I / O 416 via mediator 420 to determine head position, and can be connected to the eye tracker of I / O 416 via mediator 420 to determine eye gaze direction; their combination can specify the position of the gaze cursor. Additional details about the hands-free 3DoF interaction mode are described below. Figure 5 The frame 580 is provided.
[0071] The hands-free 6DoF mode controller 438 can provide user interaction, where a gaze cursor indicates direction and a dwell timer indicates action. The hands-free 6DoF mode controller 438 can connect to the position tracking system of I / O 416 via mediator 420 to determine the user's movements along the X, Y, and Z axes and convert these movements into changes in the user's field of view. The hands-free 6DoF mode controller 438 can connect to the IMU unit of I / O 416 via mediator 420 to determine head position and can connect to the eye tracker of I / O 416 via mediator 420 to determine eye gaze direction; their combination can specify the position of the gaze cursor. Additional details about the hands-free 6DoF interaction mode are described below. Figure 5 The 582 box is provided.
[0072] The gaze and gesture pattern controller 440 can provide user interaction, where a gaze cursor indicates direction and gestures indicate actions. The gaze and gesture pattern controller 440 can connect to the position tracking system of I / O 416 via mediator 420 to determine the user's movements along the X, Y, and Z axes and translate these movements into changes in the user's field of view. The gaze and gesture pattern controller 440 can connect to the IMU unit of I / O 416 via mediator 420 to determine head position and can connect to the eye tracker of I / O 416 via mediator 420 to determine eye gaze direction; their combination can specify the position of the gaze cursor. The gaze and gesture pattern controller 440 can also connect to a hand tracking system via mediator 420 to recognize hand gestures, which are then translated into actions. Additional details regarding gaze and gesture interaction modes are described below. Figure 5 The 584th frame is provided.
[0073] The raycasting mode controller 442 can provide user interaction, where a user-controlled ray is used to represent direction, and hand gestures are used to represent actions. The raycasting mode controller 442 can be connected via an intermediary 420 to the position tracking system of I / O 416 to determine the user's movement along the X, Y, and Z axes, and can also be connected via the intermediary 420 to the IMU unit of I / O 416 to determine head position. The raycasting mode controller 442 can use a combination of these two to determine changes in the user's field of view. The raycasting mode controller 442 can also be connected via the intermediary 420 to the hand tracking system to recognize hand gestures, which are then translated into control points for raycasting and for action recognition. Additional details regarding raycasting interaction modes are described below. Figure 5 The 586th frame is provided.
[0074] Those skilled in the art will understand that the above Figures 1 to 4 The components shown in each of the flowcharts discussed below can be modified in various ways. For example, the order of the logic can be rearranged, sub-steps can be executed in parallel, the logic shown can be omitted, and other logic can be included, etc. In some embodiments, one or more of the above-described components can perform one or more of the processes described below.
[0075] Figure 5 This is a flowchart illustrating a process 500 used in some embodiments of the present technology, which is used to identify triggers for switching between specified interaction modes and to enable the triggered interaction mode. Process 500 is depicted as having two possible starting points 502 and 552. Process 500 may begin at starting point 552 when an explicit user command for manually switching interaction modes is received. Such a command may be executed, for example, using physical buttons or other controls included in an artificial reality system, using virtual controls displayed by the artificial reality system (e.g., user interface elements), using voice commands, or by gestures or other inputs mapped to switch interaction modes. Process 500 may begin at starting point 502 in response to checking various indications that trigger an interaction mode change, such as periodic checks (e.g., performed every 1 ms, 5 ms, or 10 ms) or in response to detecting changes in interaction mode context factors. In some embodiments, the factors that cause process 500 to be executed when changed may vary depending on the current interaction mode. For example, if the current interaction mode is a 3DoF mode that does not require hands because the interaction mode system is in a low-power state, then a change in lighting conditions will not trigger a mode change check. However, if the current interaction mode is a raycasting mode that requires the camera to capture the position of the hands, then a change in lighting conditions can trigger a mode change check.
[0076] Process 500 can proceed from box 502 to box 504. At box 504, process 500 can determine whether there is no available input for position tracking and hand tracking. Available input for position tracking and hand tracking may not exist when a camera used by one of these systems or a corresponding processing component such as a machine learning model is disabled (e.g., because the interaction mode system is in low power or privacy mode, these systems have not yet been initialized, etc.), or when lighting conditions are insufficient for one of these systems to acquire images of sufficient quality. In some implementations, hand tracking or position tracking may be based on systems other than a camera (e.g., wearable gloves, rings, and / or bracelets for hand tracking systems, or time-of-flight sensor or GPS data for position tracking systems), and the determination of available hand tracking input may be based on whether these systems are enabled and whether sufficient information is received to determine hand gestures and user position. If no such available tracking input exists, process 500 can proceed to box 580. If such available tracking input exists, process 500 can proceed to box 506.
[0077] At box 580, process 500 may enable a hands-free 3DoF interaction mode. In the hands-free 3DoF interaction mode, user movements parallel to or perpendicular to the floor (i.e., X, Y, and Z axes) will not elicit corresponding movements in the artificial reality environment (although such movements can be performed using other controls such as joysticks or directional keys on a controller), but the user can indicate direction by moving their head to points in the artificial reality environment, for example, using a gaze cursor. In some implementations, the gaze cursor position may be based on both head tracking (e.g., via IMU data representing pitch, roll, and yaw movements) and eye tracking (e.g., where a camera captures images of one or both of the user's eyes to determine the user's gaze direction). In some cases, in the hands-free 3DoF interaction mode, a dwell timer may be used to perform actions, which begins a countdown (e.g., starting from three seconds) when the gaze cursor has not moved beyond a threshold amount of time for at least a threshold amount of time (e.g., when the user's gaze remains relatively fixed for at least one second). In some implementations, actions may be represented alternatively or otherwise in other ways, such as by pressing a button on a controller, activating a user interface (UI) control (i.e., a "soft" button), using a voice command, etc. In some implementations, an action may then be performed in association with a gaze cursor; for example, the system performs a default action associated with one or more objects pointed to by the gaze cursor. In some implementations, the gaze cursor and / or an indication of the user's gaze direction may be displayed in the user's field of vision as visual affordance to indicate that the user is in a hands-free 3DoF interaction mode (or a hands-free 6DoF interaction mode, as described below) and to assist the user in making selections using the gaze cursor. Additional details regarding the gaze cursor and dwell timer are provided below. Figure 6 , Figure 7 , Figure 10A and Figure 10B To provide. After enabling interactive mode at box 580, process 500 can end at box 590.
[0078] At box 506, process 500 can determine whether location information is available when the hand is not recognized as being in a ready state. Location information can be available when the positioning system is enabled and is receiving sufficient data to determine the user's 6DoF motion. In various embodiments, such a positioning system may include a system mounted on a wearable portion of the artificial reality system (“inside-out tracking”) or an external sensor for tracking the user's body position (“outside-in tracking”). Furthermore, these systems may use one or more techniques such as coded light spots (e.g., infrared) emitted by the artificial reality system and using a camera to track the motion of those spots; identifying features in acquired images and tracking their relative motion between frames (e.g., “motion vectors”); identifying portions of the user in acquired images and tracking the motion of those portions (e.g., “skeleton tracking”); acquiring time-of-flight sensor readings of objects in the environment; Global Positioning System (GPS) data; triangulation using multiple sensors; and so on. Determining whether sufficient data exists can include determining whether location tracking systems are enabled (e.g., such systems can be disabled in privacy mode, low power mode, public mode, do not disturb mode, etc.) and whether they are capable of properly determining the user's location (which may be impossible for some of these systems, for example, in low lighting conditions, when there are not enough nearby surfaces to illuminate the light spot, when environmental mapping has not yet occurred, when the user is outside the frame of the acquisition device, etc.).
[0079] When a hand tracking system is enabled and the current hand posture matches a ready hand posture, the hand can be identified as being in a ready state. The hand tracking system can include, for example, various technologies or wearable devices, such as cameras capturing and interpreting images of one or both of a user's hands (e.g., using trained models), and wearable devices such as gloves, rings, or bracelets used for measurements that can be interpreted as user hand postures. Examples of possible ready hand postures include when the user's palm is generally facing upwards (e.g., turned towards the sky by at least a threshold amount), when the user's fingers are extended by a specific amount, when certain fingers are in contact (e.g., thumb and index finger), when the user's hand is tilted beyond a specific horizontal angle, when the user makes a fist, etc. In some implementations, a hand ready state may require the hand to be in the user's field of vision, the hand to be raised above a specific threshold (e.g., above hip level), or the user's elbows to be bent by at least a specific amount. Additional details regarding hand ready states are provided below. Figure 8To provide. In some implementations, block 506 may also determine whether the hand is not in a state mapped to another interaction mode, such as the ray state discussed below with respect to blocks 508 and 510. If the position information is available and it is recognized that the hand is not in a ready state (or another interaction mode mapping state), then process 500 may continue to block 582; otherwise, process 500 may continue to block 508.
[0080] At box 582, process 500 can enable a hands-free 6DoF interaction mode. In addition to detecting user movements parallel to and perpendicular to the floor (i.e., along the X, Y, and Z axes) and translating these movements into viewpoint movements in an artificial reality environment, the hands-free 6DoF interaction mode can provide an interaction similar to the hands-free 3DoF interaction mode described above. Viewpoint movements and head tracking and / or eye tracking can be used to locate the gaze cursor. Similar to the hands-free 3DoF interaction mode, actions in the hands-free 6DoF interaction mode can be executed using a dwell timer and / or other user controls (e.g., physical or soft buttons, voice commands, etc.), and the dwell timer and / or other controls can be executed in association with the gaze cursor. After enabling the interaction mode at box 582, process 500 can terminate at box 590.
[0081] At box 508, process 500 can determine whether the hand is in a ready state, as discussed above. If yes, process 500 can proceed to box 584; if no, process 500 can proceed to box 510.
[0082] At box 584, process 500 may enable a gaze and gesture interaction mode. The gaze and gesture interaction mode may include a gaze cursor as described above. In some implementations, the gaze and gesture interaction mode will always use position tracking, allowing the user to control their position in the artificial reality environment by moving in the real world; in other cases, position tracking may or may not be enabled for the gaze and gesture interaction mode. When position tracking is enabled at box 584, the user can control the gaze cursor at 6DoF; when position tracking is disabled at box 584, the user can control the gaze cursor at 3DoF. A dwell timer may or may not be enabled in the gaze and gesture interaction mode.
[0083] When hand tracking is enabled and the hand is in a ready state, gaze and gesture interaction modes can enable actions based on user gestures. For example, a default action can be performed when the user performs a specific gesture (e.g., putting their thumb and index finger together). In some implementations, different actions can be mapped to different gestures. For example, a first action can be performed when the user puts their thumb and index finger together, and a second action can be performed when the user puts their thumb and middle finger together. Gaze and gesture interaction modes can use one or more of many different gestures, such as swipe gesture, air tap gesture, grip gesture, finger extension gesture, finger curlgesture, etc. In some cases, visual functional visibility can be used to indicate to the user that gaze and gesture interaction modes are enabled and / or gestures are being recognized. For example, a hand-ready state can be the user's palm generally facing upwards, and visual functional visibility of gaze and gesture interaction modes can include placing a sphere between the user's thumb and index finger (see...). Figure 8 This can both indicate to the user that gaze and gesture interaction modes are enabled and provide feedback on these finger-pinching gestures (e.g., by deforming or resizing a sphere when the user puts their fingers together). Other visual visibility can also be used, such as by showing the trajectory of the user's hand movements, placing indicators of different actions that can be performed near the fingertips of each finger, or displaying an indicator on a wrist that can be clicked by the fingers of another hand. Additional details about gaze and gesture interaction modes with visual visibility are discussed below. Figure 8 Provided. After interactive mode is enabled at box 584, process 500 can end at box 590.
[0084] At box 510, process 500 may determine whether the user's hand is in a ray state. Similar to the processes described above with respect to boxes 506 and 508, the process may include determining whether hand tracking is enabled, whether there is sufficient data to determine hand posture, and / or whether one or more hands are in the field of vision, whether the one or more hands are raised above a threshold, or whether the user's elbow is bent above a threshold amount. Identifying the ray state may also include determining that one or both of the user's hands are in a posture that leads to the ray projection mapping. Examples of such postures may include the user's hand being rolled so that the user's palm is generally facing down (e.g., rolled towards the floor by at least a threshold amount), one finger (e.g., index finger) being extended, some fingers (e.g., thumb, middle finger, and ring finger) being touched, etc. If one or both of the user's hands are identified as being in a ray state, process 500 may proceed to box 586. If one or both of the user's hands are identified as not being in a ray state, process 500 does not change the current interaction mode and may proceed to box 590, where it ends.
[0085] At box 586, process 500 may enable a raycasting interaction mode. This raycasting interaction mode allows the user to select using a “ray,” which can be a curved or straight line segment, a cone, pyramid, cylinder, sphere, or other geometric shape, the position of which is indicated by the user (typically based on the position of at least one of the user’s hands). Additional details regarding raycasting are provided in U.S. Patent Application No. 16 / 578,221, the entire contents of which are incorporated herein by reference. For example, a ray may be a line extending from a first point along a line connecting a first point and a second point, the first point being located between the tip of the user’s index finger and the tip of their thumb, and the second point being located on the user’s palm at a pivot point between the user’s thumb and index finger. Similar to gaze and gesture interaction modes, in various embodiments, the raycasting interaction mode may consistently use position tracking, allowing the user to control their position in an artificial reality environment by moving along the X, Y, and Z axes in the real world, while in other cases, position tracking may or may not be enabled. In either case, the raycasting interaction mode allows the user to make selections based on a ray cast by the user, which is projected from a single location (in 3DoF) or when the user moves within an artificial reality environment (in 6DoF). In the raycasting interaction mode, the user can perform actions related to the ray and / or objects intersecting the ray, for example, by making gestures such as pinching, air clicking, or grabbing. As with all interaction modes, in some implementations, other actions can also be performed, for example, using controllers or soft buttons, voice commands, etc. In some implementations, the raycasting interaction mode may include visual functional visibility, wherein a shape (e.g., a teardrop shape) is present at a point between the tips of the user's fingers from which the ray emanates. In some implementations, the shape can be resized, deformed, or otherwise altered when the user makes a gesture. For example, the teardrop shape can be compressed when the user brings the tips of their thumb and forefinger closer together. Additional details regarding raycasting interaction modes with visual functional visibility will be provided below. Figure 9 Provided. After interactive mode is enabled at box 586, process 500 can end at box 590.
[0086] As previously described, in some embodiments, process 500 may begin at box 552 instead of box 502. This can occur, for example, when a user provides an instruction to manually switch interaction modes. In various embodiments, this instruction may be executed by activating physical controls, soft controls, speaking a command, performing a gesture, or some other command mapped to a specific interaction mode or cycling through all interaction modes. This instruction may be received at box 554. At box 556, process 500 may determine whether the interaction mode change instruction corresponds to a hands-free 3DoF interaction mode (e.g., mapped to that mode or the next mode from the current mode in the mode cycle). If yes, process 500 proceeds to box 580 (as described above); if no, process 500 proceeds to box 558. At box 558, process 500 may determine whether the interaction mode change instruction corresponds to a hands-free 6DoF interaction mode (e.g., mapped to that mode or the next mode from the current mode in the mode cycle). If yes, process 500 continues to box 582 (as described above); if no, process 500 continues to box 560. At box 560, process 500 may determine whether the interaction mode change indication corresponds to a gaze and gesture interaction mode (e.g., mapped to that mode or the next mode from the current mode in the mode cycle). If yes, process 500 continues to box 584 (as described above); if no, process 500 continues to box 562. At box 562, process 500 may determine whether the interaction mode change indication corresponds to a raycasting interaction mode (e.g., mapped to that mode or the next mode from the current mode in the mode cycle). If yes, process 500 continues to box 586 (as described above); if no, process 500 does not perform an interaction mode change and continues to box 590, where it ends.
[0087] Figure 6This is a conceptual diagram illustrating Example 600 with a gaze cursor 606. The gaze cursor 606 is a selection mechanism controlled at least by the user's head position, and in some embodiments, the selection mechanism is also controlled by the user's eye gaze direction. Example 600 also includes a user's field of view 608, which is visible to the user through their artificial reality device 602. The field of view 608 includes the gaze cursor 606 as seen by the user, projected into the field of view 608 by the artificial reality device 602. The gaze cursor 606 may be along a direction indicated by a line 604 (which may or may not be shown to the user). Without eye tracking, the line 604 may be perpendicular to the coronal plane of the user's head and may intersect at a point between the user's eyes. Typically, when there is no eye tracking, the gaze cursor 606 will be located at the center of the user's field of view 608. In some embodiments, the gaze cursor 606 may also be based on the tracked user's eye gaze direction. In eye-tracking scenarios, the user's field of view 608 is an area based on the user's head position, but the position of the gaze cursor 606 within the field of view 608 can be controlled based on the determination of the user's gaze direction and / or points within the field of view the user is viewing. In some implementations, the gaze cursor 606 may appear to be located at a distance specified by the user, at the determined focal plane of the user, or on the nearest object (real or virtual) intersecting line 604. Another example of a gaze cursor... Figure 10A and Figure 10B To provide.
[0088] Figure 7 This is a conceptual diagram of an example 700 with a dwell timer 702. The dwell timer 702 can be continuously set for a duration (e.g., three seconds, four seconds, etc.) that begins when the gaze cursor 606 has not moved beyond a threshold amount within a threshold time (e.g., one second) (e.g., when line 604 has not moved more than three degrees). If the gaze cursor 606 moves beyond the threshold amount before the dwell timer expires, the dwell timer can be reset. In some embodiments, the dwell timer 702 only starts when the gaze cursor 606 is pointing at an object (e.g., object 704), for which a default action can be taken when the dwell timer expires. In some cases, when the dwell timer starts, visual function visibility can be shown, such as a ring 702 around the gaze cursor 606, wherein the percentage of the ring's color or shadow changes based on the remaining percentage of the dwell timer.
[0089] Figure 8This is a conceptual diagram illustrating an example 800 with a gaze cursor 606 and also utilizing hand gesture selection. In example 800, a gaze and gesture interaction mode is enabled in response to determining that the user's hand 802 is in a ready state based on the user's hand 802 being in the user's field of vision 608 and in a palm-up posture. In this gaze and gesture interaction mode, the user continues to use the gaze cursor 606 (see...). Figure 6 ) to indicate direction. However, instead of using a dwell timer 702 ( Figure 7 The user can perform a pinch gesture between their thumb and forefinger to trigger the action. As indicated by line 806, such a pinch gesture is interpreted as being associated with the gaze cursor 606. For example, if the default action is object selection, object 704 is selected when the user performs the pinch gesture because the gaze cursor 606 is located on object 704. With gaze and gesture interaction mode enabled, the visual function visibility in the user's field of view 608 is displayed as a sphere 804 located between the user's thumb and forefinger on their hand 802. When the user begins to put their thumb and forefinger together, the sphere 804 can be compressed into an ellipsoid to indicate to the user that a part of the pinch gesture is being recognized.
[0090] Figure 9 This is a conceptual diagram illustrating Example 900 with ray projection direction input and hand gesture selection. In Example 900, a ray projection interaction mode is enabled in response to determining that the user's hand 802 is in a ray state based on its position within the user's field of view 608 and its palm facing down. In this ray projection interaction mode, a ray 902 is projected from a point between the user's thumb and forefinger. The user can perform a pinch gesture between their thumb and forefinger to trigger an action. This pinch gesture is interpreted as being associated with one or more objects (e.g., object 704) that intersect with the ray. For example, if the default action is object selection, object 704 is selected when the user performs the pinch gesture because ray 902 intersects with object 704. When the ray projection interaction mode is enabled, the primary visual function visibility in the user's field of view 608 is displayed as a teardrop shape 906 located between the user's thumb and forefinger on their hand 802. As the user begins to place their thumb and forefinger together, the teardrop shape 906 can be compressed to indicate to the user that a portion of the pinch gesture is being recognized. Furthermore, when the ray projection interaction mode is enabled, the visual functionality of the line showing ray 902 and the point 904 where ray 902 intersects with the object is displayed.
[0091] Figure 10A and Figure 10BThese are conceptual diagrams illustrating examples 1000 and 1500 that use an eye-tracking-based gaze cursor to interact between device positioning and object selection using a dwell timer. Example 1000 includes artificial reality glasses 1002 with waveguide lenses 1004A and 1004B that project light into the user's eyes to allow the user to see a field of view 1014. In some embodiments, artificial reality glasses 1002 may be associated with the above references. Figure 2B The discussion is similar to the mixed reality HMD252. The orientation of the field of view 1014 in the artificial reality environment is controlled by the head tracking unit 1016 of the artificial reality glasses 1002. The head tracking unit 1016 can use, for example, a gyroscope, magnetometer, and accelerometer to determine the orientation and movement of the artificial reality glasses 1002, so that the artificial reality glasses 1002 can translate this into the camera position of the displayed content in the artificial reality environment. In example 1000, the field of view 1014 displays virtual objects 1012A, 1012B, and 1012C. The artificial reality glasses 1002 includes an eye-tracking unit 1006, which includes a camera and a light source (e.g., infrared light) that illuminates a flash on the user's eyes, takes an image of the user's eyes, and uses a trained machine learning model to interpret the image as a gaze direction (represented by line 1008). At the end of the eye gaze direction line 1008 is a gaze cursor 1010. In Example 1000, the user is looking at the center of the field of view 1014, causing the gaze cursor 1010 to be positioned there.
[0092] In Example 1050, the user has moved their head to the left, as indicated by arrow 1058. This movement is tracked by head tracking unit 1016, causing the field of view 1014 to shift to the left, thereby shifting virtual objects 1012A, 1012B, and 1012C to the right within the field of view 1014. The user also shifts their gaze from the center of the field of view 1014 to the lower left corner. This change in gaze is tracked by eye tracking unit 1006, as indicated by line 1008. This causes the gaze cursor 1010 to be positioned on virtual object 1012C. After the user holds their gaze on object 1012C for one second, dwell timer 1052 begins a countdown from three seconds. This is displayed to the user via visual function visibility 1054 and 1056, where the shaded portion of the ring 1056 corresponds to the remaining amount of the dwell timer. When the dwell timer reaches zero, the artificial reality glasses 1002 take the default action, which is to select the object 1012C where the gaze cursor 1010 is resting.
[0093] In this specification, references to "implementation" (e.g., "some embodiments," "various embodiments," "one embodiment," "implementation," etc.) mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this disclosure. The appearance of these phrases in different places in the specification does not necessarily refer to the same embodiment, nor are they necessarily separate or alternative embodiments mutually exclusive with other embodiments. Furthermore, various features that can be presented by some embodiments but not by others are described. Similarly, various requirements are described, which may be requirements for some embodiments but not for others.
[0094] As used herein, above the threshold means that the value of the item being compared is higher than any other specified value, that the item being compared is among a specified number of items with a maximum value, or that the value of the item being compared is within a specified highest percentage value. As used herein, below the threshold means that the value of the item being compared is lower than any other specified value, that the item being compared is among a specified number of items with a minimum value, or that the value of the item being compared is within a specified lowest percentage value. As used herein, within the threshold means that the value of the item being compared is between two other specified values, that the item being compared is among a specified number of items in the middle, or that the value of the item being compared is within a specified percentage range in the middle. Relative terms, such as high or unimportant, when not otherwise defined, can be understood as assigning a value and determining how that value is compared to a given threshold. For example, the phrase “select fast connection” can be understood as meaning selecting a connection with a value assigned to it corresponding to a connection speed above the threshold.
[0095] The various aspects of the above techniques are described as being accomplished using a hand tracking system. In each case, these techniques can alternatively use controller tracking. For example, in cases where the hand tracking system is used to determine hand posture or as the origin for projecting rays, the tracking position of the controller can be used instead.
[0096] As used in this article, the word “or” refers to any possible permutation of the set of items. For example, the phrase “A, B, or C” means at least one of A, B, C, or any combination thereof, such as any of the following: A; B; C; A and B; A and C; B and C; A, B, and C; or multiples of any of the items, such as A and A; B, B, and C; A, A, B, C, and C; and so on.
[0097] Although the subject matter has been described in language specific to structural features and / or methodological actions, it will be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. For illustrative purposes, specific embodiments and implementations have been described herein, but various modifications may be made without departing from the scope of the embodiments and implementations. The specific features and actions described above are disclosed as exemplary forms for implementing the appended claims. Therefore, the embodiments and implementations are not limited beyond the scope of the appended claims.
[0098] All of the foregoing patents, patent applications, and other references are incorporated herein by reference. If necessary, aspects may be modified to utilize the systems, functions, and concepts of the foregoing references to provide alternative implementations. In the event of any conflict between statements or subject matter in the referenced documents and statements or subject matter of this application, this application shall prevail.
Claims
1. A method for automatically switching between multiple interaction modes in an artificial reality system to interpret user input, the method comprising: Identify a first interaction mode context that indicates user location tracking input is unavailable; In response to recognizing the first interaction mode context, a hands-free three-degree-of-freedom interaction mode is enabled, the hands-free three-degree-of-freedom interaction mode indicating the roll, yaw and pitch directions; Identify the second interaction mode context, which indicates: User location tracking input is available, and Hand tracking input is unavailable, or the first hand gesture being tracked does not match the hand ready state; In response to recognizing the second interaction mode context, a hands-free six-degrees-of-freedom interaction mode is enabled, which tracks the user's position movement and the direction vector of the gaze cursor; Identify a third interaction mode context, which indicates that the tracked second hand gesture matches the hand ready state; In response to recognizing the third interaction mode context, gaze and gesture interaction modes are enabled; Identify a fourth interaction mode context, which indicates that the tracked third hand gesture matches the ray state, the third hand gesture matching the ray state including one or both of the user's hands being in a gesture toward performing ray projection mapping; and In response to recognizing the fourth interaction mode context, a raycasting interaction mode is enabled, which uses a ray to replace the gaze cursor, extending from the user's hand into the artificial reality environment.
2. The method according to claim 1, wherein, Enabling at least one of the multiple interaction modes also responds to an interaction mode change trigger, the interaction mode change trigger being: Occurs periodically; or This occurs in response to a change in the interaction mode context, which is identified as a result of monitoring interaction mode context factors.
3. The method according to claim 1, further comprising: Based on receiving a user command to change the current interaction mode, switch to a new interaction mode.
4. The method according to claim 1, wherein, The tracked second hand gesture matching the hand-ready state includes: a hand gesture identified as the user's palm facing upward by at least a threshold amount, and / or, wherein the tracked third hand gesture matching the ray state includes: a hand gesture identified as the user's palm facing downward by at least a threshold amount.
5. The method according to claim 1, wherein, The hands-free three-degree-of-freedom interaction mode includes: Based at least on the determined head orientation of the user, receive multiple user orientation indicators; and Based on a dwell timer, a user action indication related to one of the plurality of user orientation indications is received; Specifically, in the handless three-degree-of-freedom interaction mode, user movements on the X, Y, and Z axes are not automatically converted into field-of-view movements on the X, Y, and Z axes in an artificial reality environment, or The handless six-degree-of-freedom interaction mode includes: At least based on the determined head orientation of the user, multiple user orientation indications are received; Based on a dwell timer, a user action indication related to one of the plurality of user orientation indications is received; and... User motion on the X, Y, and Z axes will be converted into field-of-view motion on the X, Y, and Z axes in an artificial reality environment.
6. The method according to claim 1, wherein, The gaze and gesture interaction modes include: Multiple user orientation indicators are received based at least on the determined user's head orientation and on the determined position of the user's head relative to the artificial reality environment, wherein the determined user's head position is based on the tracked movement of the user along the X, Y, and Z axes; and By tracking the user's hand gestures and matching those gestures with a specified action, a user action instruction associated with one of the plurality of user orientation instructions is received.
7. The method according to claim 1, wherein, The ray projection interaction mode includes: Based at least on rays projected by the artificial reality system, multiple user orientation indications are received, wherein the rays are projected from positions associated with the tracking position of at least one of the user's hands; and By tracking the user's hand gestures and matching those gestures with a specified action, the system receives user action instructions associated with the ray.
8. The method according to claim 1, wherein, Each of the handless three-DOF interaction mode, the handless six-DOF interaction mode, and the gaze and gesture interaction mode provides visual functional visibility including a gaze cursor, wherein the gaze cursor is displayed in the user's field of vision and located at least in part based on the tracking position of the user's head, or Each of the handless three-degree-of-freedom interaction mode, the handless six-degree-of-freedom interaction mode, and the gaze and gesture interaction mode provides visual functionality visibility including a gaze cursor, and The gaze cursor is displayed in the user's field of view and is located at least in part based on the tracking position of the user's head and based on the user's tracking gaze direction.
9. A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform the method according to any one of claims 1 to 8, or to perform a process for switching between multiple interaction modes in an artificial reality system to interpret user input, the process comprising: Identify a first interaction mode context, which indicates that hand tracking input is unavailable or that the tracked first hand gesture does not match the hand ready state; In response to recognizing the first interaction mode context, a hands-free interaction mode is enabled; Identify a second interaction mode context, which indicates that the tracked second hand gesture matches the hand ready state; In response to recognizing the second interaction mode context, gaze and gesture interaction modes are enabled; Identify a third interaction mode context, which indicates that the tracked third hand gesture matches the ray state; as well as In response to recognizing the third interaction mode context, the raycasting interaction mode is enabled.
10. The computer-readable storage medium according to claim 9, wherein, The hands-free interaction modes include: Based at least on the determined head orientation of the user, receive multiple user orientation indicators; and Based on a dwell timer, a user action indication related to one of the plurality of user orientation indications is received; in: When user position tracking input is unavailable, the hands-free interaction mode does not translate user motion on the X, Y, and Z axes into field-of-view motion on the X, Y, and Z axes in an artificial reality environment; and When user position tracking input is available, the hands-free interaction mode automatically translates user motion on the X, Y, and Z axes into field-of-view motion on the X, Y, and Z axes in an artificial reality environment.
11. The computer-readable storage medium according to claim 9, in, The gaze and gesture interaction mode provides visual function visibility, which includes a sphere shown in the user's field of vision, located between the user's thumb and one other finger; and The sphere is shown to adjust its size or shape according to a predetermined distance between the user's thumb and the other fingers, or The ray-projection interaction mode provides visual functional visibility, which includes a shape shown in the user's field of vision, located between the user's thumb and another finger; and The shape is shown to be adjusted in size or deformed according to a certain distance between the user's thumb and the other fingers.
12. The computer-readable storage medium according to claim 9, wherein, The tracked second hand gesture that matches the hand-ready state includes: a hand gesture identified as the user's palm facing upwards.
13. The computer-readable storage medium according to claim 9, wherein, The gaze and gesture interaction modes include: Multiple user orientation indicators are received based at least on the determined user's head orientation and on the determined position of the user's head relative to the artificial reality environment, wherein the determined user's head position is based on the tracked movement of the user along the X, Y, and Z axes; and By tracking the user's hand gestures and matching those gestures with a specified action, a user action instruction associated with one of the plurality of user orientation instructions is received.
14. A computing system for interpreting user input in an artificial reality system, the computing system comprising: One or more processors; as well as One or more memories storing instructions that, when executed by the one or more processors, cause the computing system to perform the method according to any one of claims 1 to 8, or to perform a process for switching between multiple interaction modes, the process comprising: Identify a first interaction mode context, which indicates that hand tracking input is unavailable or that the tracked first hand gesture does not match the hand ready state; In response to recognizing the first interaction mode context, a hands-free interaction mode is enabled; Identify a second interaction mode context, which indicates that the tracked second hand gesture matches the hand ready state; In response to recognizing the second interaction mode context, gaze and gesture interaction modes are enabled; Identify a third interaction mode context, which indicates that the tracked third hand gesture matches the ray state; and In response to recognizing the third interaction mode context, the raycasting interaction mode is enabled.
15. The computing system according to claim 14, wherein, The tracked third hand gesture matched with the ray state includes: a hand gesture identified as the user's palm facing down by at least a threshold amount, or The hands-free interaction mode and the gaze and gesture interaction mode provide visual functionality visibility including a gaze cursor, which is displayed in the user's field of vision and located at least in part based on the tracking position of the user's head.
Citation Information
Patent Citations
Projection casting in virtual environments
US11176745B2
Dynamic switching and merging of head, gesture and touch input in virtual reality
CN107533374A
Context-sensitive hand interaction
US20190050062A1