Management of gaze to gesture control transitions
Transition modes in head-worn devices switch between gaze and gesture controls based on hand location and orientation, improving user input efficiency by predicting optimal input methods.
Patent Information
- Application Number
- PCT/US2025/021144
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-30
- Filing Date
- 2025-03-24
- Publication Date
- 2025-09-25
AI Technical Summary
Existing systems face challenges in determining when to transition between gaze and gesture controls in head-worn devices, leading to inefficiencies in user input methods.
Implementing transition modes that switch between gaze and gesture controls based on factors such as the location of the user's hand relative to the field of view, orientation, and gaze direction, using sensors and machine learning models to predict optimal input methods.
Enables seamless and efficient user input by automatically switching between gaze and gesture controls, enhancing user experience and interaction efficiency in head-worn devices.
Smart Images

Figure US2025021144_25092025_PF_FP_ABST
Abstract
Description
MANAGEMENT OF GAZE TO GESTURE CONTROLTRANSITIONSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority- to U.S. Provisional Patent Application No. 63 / 569,040, filed on March 22, 2024, entitled ‘MANAGEMENT OF GAZE TO GESTURE CONTROL TRANSITIONS,” and U.S. Provisional Patent Application No. 63 / 740,108, filed on December 30, 2024, entitled “MANAGEMENT OF GAZE TO GESTURE CONTROL TRANSITIONS.” the disclosures of which are incorporated herein by reference in their entirety-.BACKGROUND
[0002] A head-wom device is a wearable technology designed to be worn on or around the head, including smart glasses, augmented reality (AR) and virtual reality (VR) headsets, and head-mounted displays. These devices typically feature advanced sensors, displays, and communication interfaces to support immersive and interactive experiences. Users can provide input through multiple methods, including physical buttons, touch- sensitive surfaces, voice commands, gesture recognition, eye detection, and external controllers. Additionally, some head-wom devices incorporate spatial tracking, accelerometers, and gyroscopes to detect head movements, enabling hands-free control.SUMMARY
[0003] This disclosure relates to systems and methods for managing gaze and gesture input transitions. In some implementations, a device can be configured to transition between gaze control and gesture control based on the movement of the hand and the gaze of the user. In some implementations, a device can be configured to identify a first action provided using gaze control that identifies eye movements to enable interaction with digital interfaces or physical systems. The device can further be configured to identify a portion of a user relative to a field of view associated with the gaze of the user. In some implementations, the portion corresponds to a hand or finger of the user. In some examples, the device can further be configured to change from gaze control to gesture control based on the portion of the user relative to the field of view. In some examples, the device can transition from gaze to gesturecontrol when the portion enters or nears the field of view. In some implementations, the device can determine additional values associated with the gaze and the body portion and use the values to determine whether to transition from gaze to gesture control. The device can further be configured to identify a second action using gesture control.
[0004] In some implementations, the device can be configured to transition from gesture to gaze control. The device can identify the portion of the user relative to the field of view associated with the gaze of the user and change to gaze control from gesture control based on the portion of the user relative to the field of view.
[0005] In some aspects, the techniques described herein relate to a method including: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing between a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
[0006] In some aspects, the techniques described herein relate to a computing system including: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer- readable storage medium that, when executed by the at least one processor, direct the computing system to perform a method, the method including: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing betw een a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
[0007] In some aspects, the techniques described herein relate to a computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method including: identifying a first action based on a gaze control associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; changing, based on the portion of the user relative to the field of view, from the gaze control to a gesture control associated with the portion of the user; and identifying a second action based on the gesture control.
[0008] The accompanying drawings and the description below outline the details of one or more implementations. Other features will be apparent from the description, drawings, and claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG. 1 illustrates a computing environment for transitioning between gaze and gesture control according to an implementation.
[0010] FIG. 2 illustrates a method of transitioning between gaze and gesture control according to an implementation.
[0011] FIG. 3 illustrates an operational scenario of transitioning from gaze to gesture control according to an implementation.
[0012] FIG. 4 illustrates an operational scenario of transitioning from gesture to gaze control according to an implementation.
[0013] FIG. 5 illustrates a method of transitioning between gaze and gesture control according to an implementation.
[0014] FIG. 6 illustrates an operational scenario of selecting gesture control according to an implementation.
[0015] FIG. 7 illustrates a computing system that displays an indicator according to an implementation.DETAILED DESCRIPTION
[0016] Examples herein support managing transitions between gesture controls on a device and gaze controls on the device. In some implementations, gesture control is used to interact with digital devices and systems without touching hardware interfaces, such as mice, keyboards, and screens. These products may include wearable devices, gaming consoles, virtual reality (VR) or augmented reality (VR or AR) headsets (examples of extended reality7(XR) headsets), smart televisions and streaming devices, computers, and laptops, amongst other types of devices. To support gesture control, a digital device uses cameras, infrared sensors, motion sensors, or some other sensor to capture the motion of the user. Once captured, image or signal processing is applied to the motion to determine a desired action for the user. The actions may include selecting an item in a menu, scrolling a menu, manipulating a menu, manipulating an object presented by the device, or some other action. For example, a user of an XR headset may use a hand motion to select an option offered on the display of the XR headset.
[0017] In addition to the gesture controls, the same device can be configured to permit gaze controls using gaze detection. Gaze control systems can enable interaction with devices based on where the user is looking. Gaze control systems can use cameras and infrared light to determine where a user is looking by analyzing the position of the pupil andreflections in the eyes. The system then translates these eye movements into commands that control the device. This allows for hands-free interaction by mapping the user’s gaze to cursor movements or selections on the screen. For example, the user can use their gaze to focus on displayed content and use a voice command to select the content.
[0018] However, at least one technical problem exists in combining gesture and gaze controls. Specifically, systems and devices encounter difficulties determining whether to use gaze or gesture controls during different conditions. As a technical solution, the systems and methods herein provide transition modes that trigger transitions between gaze and hand controls. The technical benefit of this permits the user to provide input using gaze when desired and efficiently transition to gesture input based on extremity' or body portion location relative to the field of view or content and vice versa.
[0019] In at least one implementation, devices like XR devices may transition between gaze and gesture control based on various factors. In some examples, a device can be configured to transition from gaze controls to hand controls based on the location of the user’s hand to the user’s field of view. For example, while the hand is not located in the user’s field of view, the device may use gaze controls to implement actions in association with the device (e.g., scrolling, selecting, and the like). The device may determine the field of view by detecting the user’s gaze and determining the extent of the observable environment (physical and virtual) as seen from the device. The field of view can be determined by the optical and display limitations related to the user’s perspective of the content and the physical environment (i.e., the extent visible to the user). The device then determines when a hand enters the field of view associated with the user’s gaze and automatically switches to using gesture controls associated with the hand in place of the gaze controls. Advantageously, the device can switch from gaze controls to hand controls based on whether a hand is in (or near) the field of view.
[0020] In some technical solutions, the device will use attributes or values associated with the hand (or other body portion) and / or gaze to determine whether to transition from gaze control to hand control. The attributes may include the location of the hand relative to the field of view, the device location relative to the content, the hand location relative to the content, the hand orientation (or aim) relative to the field of view, or the location of the palm to the user. From the attributes, the device determines whether the attributes satisfy the criteria for using hand controls, and the device switches to hand controls when the requirements are satisfied. Similarly, the attributes associated with the hand and / or gaze can be monitored to determine when the device should transition back from hand gestures to gazecontrol. In at least one example, the device may determine when one or more attributes satisfy second criteria to transition from hand gestures to gaze control. Once transitioned, the device will monitor the gaze of the user to identify and implement the requested actions of the user.
[0021] In some implementations, the device can be configured to determine when one or more factors satisfy at least one criterion to transition from gaze control to gesture control. In some examples, the device can determine whether the user s hand or body portion is within a threshold degree of the user’s field of view, which may indicate a user’s desire for gesture control. In some examples, the device can be configured to determine whether the device (or user) is facing content within a threshold degree. For instance, with an XR device, the content can depend on the direction the user faces, where content can be visible from a first direction and not from a second direction. If the device (or user) faces content, the device can infer the user may interact with the content using gesture control. In some implementations, the device can be configured to determine whether the body portion (e.g., hand) is aiming at the content or is within a degree threshold of the content. The aim of the body portion can be determined using a ray. a virtual pointing mechanism that extends from a user’s body portion (e.g., hand), allowing them to interact with UI elements or objects at a distance, like a laser pointer. In some examples, the device is configured to determine whether the body portion (e.g., hand or finger) aims w ithin a threshold degree value of a field of view or center gaze point for the user. In some implementations, the device can be configured to determine whether the body portion (e.g., hand or palm) faces outward away from the user’s body (or toward the displayed content), indicating an attempt to interact with content. In some implementations, the device can be configured to determine the velocity or rate of movement of the body portion, where the velocity can indicate an attempt for the body portion to provide gesture control. In some implementations, the device can be configured to use any of the factors above to determine whether to transition from gaze to gesture control.
[0022] In some implementations, the device can be configured to determine one or more of the above values and determine whether the one or more values satisfy at least one criterion. If the one or more values do not satisfy the at least one criterion, then the device can be configured to remain using gaze control. If the one or more values do satisfy the at least one criterion, then the device can be configured to transition from gaze control to gesture control. In some implementations, the device can be configured to iteratively test to determine whether a transition occurs between gaze and gesture control.
[0023] In at least one example, a machine learning model can be applied to the factors listed above to determine when to change from gaze to gesture control. The model can be trained or configured using patterns from labeled data that associate various factors with one of the input modes. Using algorithms like decision trees, neural networks, or logistic regression, the model can identify patterns and generalize rules to predict which input model (gaze or gesture) to use based on new input data (gaze and body location).
[0024] In some examples, the device can be configured to determine when the device should transition from gesture to gaze controls. In some implementations, the device can be configured to determine when the angle between the gaze (e.g., gaze point or field of view) and the ray from the body portion exceeds a threshold. These can be considered vectors, where the line associated with the user’s gaze is a first vector and the hand ray is a second vector. The ray (or hand ray) is a virtual pointing mechanism that extends from a user’s body portion, allowing them to interact with interface elements or objects at a distance, like a laser pointer. The angle exceeding the threshold (angle between the gaze vector and hand ray vector) can indicate that the user is transitioning to gaze control on the device. In some implementations, the device can be configured to determine when the body portion exits a field of view for the user. This can indicate the user no longer intends to use gesture control to provide input. In some implementations, the device can be configured (when the body portion is a user’s hand) to determine when the palm of the hand is no longer facing outwards or directed toward the content on the display. In some implementations, the device can be configured to determine a velocity associated with the user’s body portion, wherein a quick movement can indicate that the user no longer requires gesture controls. In some examples, the device can be configured to determine whether the user executes a defined gesture, and based on the gesture, the device can revert to gaze control. In at least one example, the gesture can include pulling backward and turning the hand (e.g., palm toward the user). The user can initially interact with content using gestures with their fingers pointing up. To return to gaze-based control, the user can provide a gesture that rotates the hand to fingers pointing down and will pull the hand backward. The sensors on the device can identify the specific gesture and return the user to gaze-based control. Although demonstrated as using a gesture to return to gaze control, the device can also monitor for a gesture that transitions from gaze to gesture controls. For example, when the user provides the inverse gesture to the one described above, the device can transition from gaze to gesture control. The device can be configured to use any combination of the above elements to determine whether to transition from gesture to gaze control.
[0025] FIG. 1 illustrates a computing environment 100 for transitioning between gaze and gesture control according to an implementation. Computing environment 100 includes user 110, device 130, and display view 105. Device 130 includes display 131, sensors 132, camera 133, and input application 126. Display view 105, representative of the view provided by display 131, includes indicator 112, field of view 120, body portion 170, and content 140, 141, and 142.
[0026] In computing environment 100, device 130 includes display 131, a screen or projection surface that presents visual content to user 110. Display 131 can include optical see-through displays (e.g., AR headsets) or video pass-through (e.g., MR / VR devices). Display 131 can use projectors and / or waveguides to display content for user 110. Device 130 further includes sensors 132, such as accelerometers, gyroscopes, magnetometers, depth, infrared, and proximity sensors. The sensors can be used to monitor the user’s physical movement, identify depth information for other objects, identify eye movement for the user, or provide some other operation. Device 130 also includes camera 133, which can capture the real or physical environment to overlay virtual objects (e.g., content 140, 141, and 142) and identify the movements of user 110 and surroundings to enable accurate interaction within the augmented or virtual space. In some examples, camera 133 can be positioned as an outward view to capture the physical world associated with the user’s gaze. Display 131 can receive updates from input application 126 to implement actions (e.g., associated with user inputs) and display content for user 110. Sensors 132 and camera 133 provide data to input application 126 that can be used to identify input using gestures or the user’s gaze. In some examples, input application 126 can be configured to determine w hen to transition between gaze and gesture input.
[0027] In some implementations, user 110 can use gesture controls to provide input and implement actions on device 130. Gesture control is a technology that allows users to interact with devices through hand or body movements without physical contact, using sensors and cameras to interpret gestures. It enhances user experience by providing intuitive, touch-free operation in applications like gaming, smart devices, and XR. Gaze control is a technology that enables users to interact with a device using their eye movements, which sensors or cameras can determine. It allows for hands-free navigation, making it useful in accessibility tools, wearable devices, and assistive technology for individuals with mobility impairments. In some examples, gaze control systems use cameras and infrared light to determine where a user is looking by analyzing the position of the pupil and reflections in the eyes, then translating these eye movements into commands that control the device (i.e..actions). This allows for hands-free interaction by mapping the user's gaze to cursor movements or selections on the screen.
[0028] In some implementations, device 130 can be configured to transition between gaze and gesture control. The transition can be determined based on a variety’ of factors related to the position of the user gaze and the position of a body portion 170 (e.g., hand or finger of user 110). In some examples, input application 126 can determine the location of body portion 170 relative to field of view 120 to determine the transition. In some implementations, while body portion 170 is not located in the user’s field of view, device 130 may use gaze controls to implement actions associated with the device (e.g., scrolling, selecting, and the like). The device may determine the field of view by determining the gaze of the user and determining the extent of the observable environment (physical and virtual) as seen from the device. The device then determines when body portion 170 enters or approaches the field of view associated with the gaze and switches to using gesture controls related to the hand in place of the gaze controls. Advantageously, the device can switch from gaze controls to hand controls based on whether a hand is within a portion of the field of view.
[0029] In some examples, device 130 can consider various factors in determining whether criteria are met to transition from gaze to gesture control. The additional factors can include the device facing content (e.g., within threshold angle of content), the body portion (e.g.. hand) aiming toward the content (e.g.. measured by a threshold degree), the body portion aimed within a threshold angle of the field of view or center gaze point, the body portion (e.g., hand) facing outward, or some other factor. In some implementations, the device can consider one or more of the above factors to determine when to transition from gaze to gesture control. Similar factors can also be considered in some implementations when reverting from gesture to gaze control. The factors can include the location of the body portion relative to the field of view, the aim of the body portion facing away from the content, or some other factor.
[0030] FIG. 2 illustrates method 200 of transitioning between gaze and gesture control according to an implementation. Method 200 is described below with reference to systems and elements of computing environment 100 of FIG. 1. In some examples, device 130 of FIG. 1 can implement method 200.
[0031] Method 200 includes identifying a first action associated with a user of a device at step 201. In some implementations, the first action corresponds to an input. In some implementations, the first action is implemented based on gaze control. Gaze control refers tousing a user’s eye movements to interact with content provided by the device. It is implemented through eye-tracking sensors that detect gaze direction and enable hands-free navigation, object selection, and user interface interactions (i.e., actions). In some implementations, the first action is implemented based on gesture control. Gesture control permits users to interact with content using body movements, such as hand and finger movements. It is implemented through cameras, infrared sensors, or depth sensors that track hand positions and gestures, enabling actions like grabbing, pointing, or swiping to manipulate content.
[0032] Method 200 further includes identifying a portion of the user relative to a field of view associated with a gaze of the user at step 202. In some implementations, the device identifies the body portion relative to the field of view using inside-out tracking with cameras and depth sensors mounted on the headset. These sensors detect the position, shape, and movement of the user's hands, mapping them within a 3D coordinate space relative to the headset’s viewpoint. In some examples, the field of view can be defined as a region that corresponds to a cone with a tip from the gaze focus dispersed on the display of the device (e.g.. a circular region around the gaze focus). For example, device 130 can track body portion 170 relative to field of view 120. In some examples, when body portion 170 is outside of field of view 120, device 130 can be configured to use gaze control. In some examples, when body portion 170 is inside the field of view 120, device 130 can be configured to use gesture control.
[0033] Method 200 further includes, based on the portion of the user relative to the field of view, changing between a gaze control and gesture control associated with the portion of the user at step 203. For example, when body portion 170 enters field of view 120, the device can be configured to transition to gesture control. Gesture control can permit user 110 to move indicator 112 using movement of body portion 170. In another example, when body portion 170 leaves field of view 120, the user can provide gaze control. Method 200 further includes identifying a second action at step 204, the second action implemented using the selected gaze or gesture control.
[0034] In some implementations, in addition to using the position of the body portion 170 relative to the field of view 120, device 130 can be configured to identify additional factors associated with the gaze and / or body portion. In some examples, an additional factor can include the degree to which the user or device faces content provided by the device. This can be calculated using a vector for the gaze relative to the content, wherein an angle is defined based on a vector for the displayed content (e.g., one direction) and the vector for theuser’s gaze. When the device points within a threshold angle of the content, it can indicate that the user is attempting to use gesture control. When the device points away from the content (e.g., exceeds the threshold angle), it can indicate that the user is not attempting to use gesture control.
[0035] In some examples, another factor includes determining whether the body portion (e.g., hand) points at or near content (e.g., within a threshold angle). The determination can be based on a ray (i.e., vector) that extends from the hand relative to the content. When within the threshold angle (defined by the hand ray vector and a vector for the displayed content), it can indicate that the user is attempting to use gesture control. When outside of the threshold angle, it can indicate that the user is not trying to use gesture control.
[0036] In some examples, an additional factor includes determining whether the body portion points at or near the gaze center of the user. Further away from the gaze center can indicate that the user does not intend to use gesture control. For example, when body portion 170 is further from the center of field of view 120 (e.g., outside a threshold distance or angle), device 130 can use the factor to predict that user 110 is not attempting to use gesture control. In some examples, the angles are defined based on vectors associated with the user gaze, the user’s hand ray, and / or the content displayed by the device. The vectors and angles can be used to define relationships between the user’s gaze, the pointing direction associated with the body portion, and the content provided by the device.
[0037] In some examples, an additional factor, when the body portion is the user’s hand, includes determining whether the hand is facing outward or away from the user’s body. Facing away from the user can indicate that the user is attempting to interact with content using gesture control. The device can determine whether the vector or ray associated with the user’s palm satisfies a threshold of pointing away from the user’s body. In some examples, any combination of the factors above can be combined to determine whether the device should use gaze or gesture control. The user or device type can define different thresholds and criteria in some implementations. For example, the user can indicate the factors, threshold angles, and / or values that trigger the transitions between gaze and gesture control. In some examples, a model or machine learning model can be applied to the values associated with the various factors to determine whether to use gaze-based input or gesture-based input. The machine learning model can be trained or configured using patterns from labeled data that associate various factors with one of the input modes. Using algorithms like decision trees, neural networks, or logistic regression, the model can identify patterns and generalizerules to make predictions about which input model to use based on new input data (gaze and body location).
[0038] In some implementations, device 130 can be configured to identify a gesture or gestures to transition between gaze and gesture control. Device 130 can determine a unique gesture by capturing hand motion data (i.e., body portion data) using optical cameras and sensors and analyzing the data through machine learning or rule-based algorithms. Features such as hand shape, finger positions, velocity, and trajectory are extracted and compared against a predefined gesture dataset using techniques like convolutional neural networks (CNNs) or hidden Markov models (HMMs) to recognize and classify the gesture. For example, device 130 can be configured to identify when the user turns their palm toward the user and pulls their hand toward their body. The gesture can be indicative of a request to change from gesture-based control to gaze-based control.
[0039] FIG. 3 illustrates an operational scenario 300 of transitioning from gaze to gesture control according to an implementation. Operational scenario 300 includes displayview 305, content 307. gaze selector 310, gesture selector 311, field of view 320. body portion 330, and steps 350, 351, and 352 for transitioning from gaze control to gesture control. In some implementations, the steps of operational scenario 300 can be performed by a device, such as device 130 of FIG. 1.
[0040] In operational scenario 300, a device is configured to identify selection using gaze control at step 350. Here, gaze control is managed using gaze selector 310. A user can provide input using gaze on a device by looking at specific areas on the screen, where an eye system detects and interprets their gaze. This technology7enables interaction through dwell time (staring at an option for a set duration), blink activation, or combining gaze with other inputs like voice, button, and the like. For example, a user may gaze at an input object and provide voice input to select the object and implement an action.
[0041] Operational scenario 300 further determines that a body portion and / or gaze satisfies at least one criterion at step 351 and changes to gesture control when the at least one criterion is satisfied at step 352. The at least one criterion can include the user or device aimed within a threshold angle of content, a threshold distance of the body portion from the center of the user’s gaze (or field of view), an orientation criterion for the body portion (e g., a hand facing away from the user), a threshold angle for the body portion relative to the user’s gaze point (e.g., hand ray aiming within a threshold angle of the user’s gaze), or some other criterion related to the body portion or the user’s gaze. In some implementations, the system can use a set of criteria. In some implementations, the device can identify valuesassociated with the user’s gaze and body portion (e.g., the angular distance between the gaze and aim of the body portion) and apply a machine learning model. The various values can be used to provide an output indicating whether gaze control or gesture control should be used. In some examples, the model can be configured using patterns from labeled data that associate various factors with one of the input modes (gesture or gaze). Using algorithms like decision trees, neural networks, or logistic regression, the model can identify patterns and generalize rules to predict which input model (gesture or gaze) to use based on new input data.
[0042] Once transitioned, the device can be configured to identify gesture control or actions caused from gesture input. Gesture input can refer to recognizing and interpreting hand and finger movements (i.e., body movements) as commands for interacting with a computing system. It can be identified using computer vision and depth-sensing technologies, such as infrared cameras, time-of-flight sensors, or LiDAR, which capture hand position, motion, and shape. Machine learning algorithms can process these inputs to classify7gestures based on predefined patterns, enabling real-time interaction with various content.
[0043] FIG. 4 illustrates an operational scenario 400 of transitioning from gesture to gaze control according to an implementation. Operational scenario 400 includes display 405, content 407, gesture selector 410, gaze selector 411, body portion 430, and steps 450, 451, and 452. In some implementations, the steps of operational scenario 400 can be performed by a device, such as device 130 of FIG. 1.
[0044] Operational scenario 400 includes identifying selections using gesture control at step 450. Operational scenario 400 further comprises determining that a body portion and / or gaze satisfies at least one criterion at step 451 and changing to gaze control when the at least one criterion is satisfied. In some implementations, the at least criterion can include an angle of the ray from the body portion exceeding an angle from the gaze center, the hand exiting the field of view (or a defined area for the field of view), a determined orientation for the body portion, or some other criterion. In some implementations, the system can use multiple criteria to determine when to transition. For example, the device can determine whether body portion 430 leaves the area defined by field of view 420. In some examples, the user may not view body portion 430, but the device can determine the body portion location relative to the user’s field of view in 3D space. The device can determine the location of the user’s body portion (e.g., hand) relative to the gaze even when a screen (e.g., XR device) may interrupt the gaze from viewing the body portion. In some implementations, the device can employ a machine learning model with values associated with the body portion and gaze todetermine whether to transition between gesture and gaze inputs. The values can include angles, locations, or some other value associated with the gaze and / or body portion.
[0045] In some implementations, the user can provide a unique gesture to transition from gesture control to gaze control. The gesture can include rotating the hand or body portion of the user, changing the location of the body portion (e.g., pulling the hand toward the user), or some other gesture unique to the transition. In some examples, a gesture can also transition from gaze to gesture control. The gesture can be the same or a different gesture as the transition from the gesture to gaze control.
[0046] FIG. 5 illustrates method 500 of transitioning between gaze and gesture control according to an implementation. In some examples, device 130 of FIG. 1 can implement method 500.
[0047] Method 500 includes identifying a first input using a first input mode at step501. In some implementations, the first input mode comprises a gesture input mode. In some implementations, the first input mode comprises a gaze input mode. Method 500 further includes determining that a body portion and / or a gaze satisfy at least one criterion at step502. In some implementations, the at least one criterion comprises a threshold angle between the body portion and the field of view associated with the gaze. In some implementations, the at least one criterion comprises a threshold angle for the device facing content. In some implementations, the at least one criterion comprises a threshold angle for the body portion to be facing content. In some implementations, the at least one criterion comprises a threshold angle between the gaze point and the pointing direction of the body portion. In some implementations, the at least one criterion comprises an orientation of the body portion relative to the user’s body. In some implementations, the at least one criterion can comprise a set of the examples described above.
[0048] Method 500 includes changing from the first input mode to a second input mode based on the body portion and / or the gaze satisfying the at least one criterion at step503. In some implementations, the first mode is gaze control, and the second mode is gesture control. In some implementations, the first mode is gesture control, and the second mode is gaze control. Method 500 further includes identifying a second input using the second input mode at step 504.
[0049] FIG. 6 illustrates an operational scenario 600 of selecting gesture control according to an implementation. Operational scenario 600 includes user 605, gesture ray 610, gaze 620, angle 630. and operational steps 650 and 651. A wearable device can perform steps650 and 651 in some examples. In some examples, steps 650 and 651 can be performed by device 130 of FIG. 1.
[0050] Operational scenario 600 provides an example of transitioning or changing from gaze control to gesture control based on the angular proximity of gesture ray 610 and gaze 620. Gesture ray 610 is representative of an aim point associated with the user’s body portion. The ray extends from the user's body portion (e.g., hand) into 3D space. Gaze 620 represents an invisible line extending from the user’s eyes or head direction into the environment. It can be determined using cameras and other sensors on the device. Operational scenario 600 determines that the gaze and gesture are within a threshold angle at step 650 and changes to gesture control when the gaze and gesture are within the threshold angle. For example, the body portion can comprise a hand that corresponds to gesture ray 610. When user 605 raises their hand to interact with the content, the device can determine that gesture ray 610 for the hand and gaze 620 satisfy threshold angle 630. This causes the transition from gaze control to gesture control. When the gesture fails to satisfy threshold angle 630, then the device can be configured to return to gaze control. In some examples, the device can consider additional factors in determining when to change between gaze and gesture control. The factors are related to the user’s gaze, the body portion location, the direction of the device, or some other factor.
[0051] In some implementations, the device can be configured to identify one or more angles based on a set of vectors. In some examples, a vector corresponds to the gaze of the user. In some examples, a field of view can be determined from the user gaze, where the field of view corresponds to the extent of an observable area that the user can see. In some examples, a vector corresponds to a hand ray. In some examples, a vector is defined for the content displayed by the device. The device can be configured to determine angles relative to the various angles. The device can also determine when a hand or other body portion is visible in the field of view .
[0052] FIG. 7 illustrates a computing system that displays an indicator according to an implementation. Computing system 700 represents any apparatus, computing system, or systems with which the various operational architectures, processes, scenarios, and sequences are disclosed herein for managing gaze and gesture control. Computing system 700 can be an example of an XR device, w earable device, or other computing device capable of the operations described herein. Computing system 700 can be an example of device 130 of FIG. 1 in some implementations. Computing system 700 can be a system of devices, such as a wearable device and a companion device (e.g., smartphone, tablet, etc.) in some examples.Computing system 700 includes storage system 745, processing system 750, communication interface 760, and input / output (I / O) device(s) 770. Processing system 750 is operatively linked to communication interface 760, I / O device(s) 770, and storage system 745. In some implementations, communication interface 760 and / or I / O device(s) 770 may be communicatively linked to storage system 745. Computing system 700 may further include other components such as a battery and enclosure that are not shown for clarity.
[0053] Communication interface 760 comprises components that communicate over communication links, such as network cards, ports, radio frequency, processing circuitry (and corresponding software), or some other communication devices. Communication interface 760 may be configured to communicate over metallic, wireless, or optical links. Communication interface 760 may be configured to use Time Division Multiplex (TDM), Internet Protocol (IP), Ethernet, optical networking, wireless protocols, communication signaling, or some other communication format - including combinations thereof. Communication interface 760 may be configured to communicate with external devices, such as servers, user devices, or other computing devices.
[0054] I / O device(s) 770 may include peripherals of a computer that facilitate the interaction between the user and computing system 700. Examples of I / O device(s) 770 may include keyboards, mice, trackpads, monitors, displays, printers, cameras, microphones, external storage devices, sensors, and the like. In some implementations, I / O device(s) 770 include at least one outward-facing camera configured to capture images associated with the physical environment. In some implementations, I / O device(s) 770 can include cameras and sensors to identify the user gaze and the gestures provided by a body portion of the user.
[0055] Processing system 750 comprises microprocessor circuitry (e.g., at least one processor) and other circuitry that retrieves and executes operating software (i.e.. program instructions) from storage system 745. Storage system 745 may include volatile and nonvolatile, removable, and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. Storage system 745 may be implemented as a single storage device but may also be implemented across multiple storage devices or sub-systems. Storage system 745 may comprise additional elements, such as a controller to read operating software from the storage systems. Examples of storage media (also referred to as computer-readable storage media or a computer-readable storage medium) include random access memory', readonly memory, magnetic disks, optical disks, and flash memory, as well as any combination or variation thereof, or any other type of storage media. In some implementations, the storagemedia may be non- transitory. In some instances, at least a portion of the storage media may be transitory. In no case is the storage media a propagated signal.
[0056] Processing system 750 is ty pically mounted on a circuit board that may also hold the storage system. The operating software of storage system 745 comprises computer programs, firmware, or some other form of machine-readable program instructions. The operating software of storage system 745 comprises input application 724. The operating software on storage system 745 may further include an operating system, utilities, drivers, network interfaces, applications, or some other type of software. When read and executed by processing system 750 the operating software on storage system 745 directs computing system 700 to operate as described herein. In at least one implementation, the operating software can provide method 200 described in FIG. 2 and method 500 described in FIG. 5. The operating software can provide or cause the at least one processor to change between gaze and gesture control in some implementations.
[0057] Example clauses are provided below. Although these are examples, these clauses should not be considered exhaustive.
[0058] Clause 1. A method comprising: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing between a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
[0059] Clause 2. The method of clause 1, wherein the portion of the user comprises a hand, and the method further comprising: determining an orientation of the hand relative to content displayed by the device, wherein changing between the gaze control and the gesture control is further based on the orientation.
[0060] Clause 3. The method of clause 1 or 2, further comprising: determining an angle between a direction of the device and content displayed by the device, wherein changing between the gaze control and the gesture control is further based on the angle.
[0061] Clause 4. The method of any one of clauses 1 to 3. wherein the portion of the user comprises a hand, and the method further comprising: determining an angle between the hand and a center of the field of view, wherein changing between the gaze control and the gesture control is further based on the angle.
[0062] Clause 5. The method of any one of clauses 1 to 4. wherein the portion of the user comprises a hand, and the method further comprising: determining a direction of a palmof the hand relative to a body of the user, wherein changing between the gaze control and the gesture control is further based on the direction of the palm.
[0063] Clause 6. The method of any one of clauses 1 to 5, further comprising: determining a movement velocity associated with the portion of the user, wherein changing between the gaze control and the gesture control is further based on the movement velocity.
[0064] Clause 7. The method of any one of clauses 1 to 6. wherein identifying the first action comprises identifying the first action via gaze control, wherein identifying the second action comprises identify ing the second action via gesture control.
[0065] Clause 8. The method of any one of clauses 1 to 6, wherein identify ing the first action comprises identifying the first action via gesture control, wherein identifying the second action comprises identifying the second action via gaze control.
[0066] Clause 9. A computing system comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing system to perform a method, the method comprising: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing between a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
[0067] Clause 10. The computing system of clause 9, wherein the portion of the user comprises a hand, and the method further comprising: determining an orientation of the hand relative to content on the device, wherein changing between the gaze control and the gesture control is further based on the orientation.
[0068] Clause 11. The computing system of clause 9 or 10, wherein the method further comprises: determining an angle between a direction of the device and content, wherein changing between the gaze control and the gesture control is further based on the angle.
[0069] Clause 12. The computing system of any one of clauses 9 to 11, wherein the portion of the user comprises a hand, and the method further comprising: determining an angle between the hand and a center of the field of view, wherein changing between the gaze control and the gesture control is further based on the angle.
[0070] Clause 13. The computing system of any one of clauses 9 to 12, wherein the portion of the user comprises a hand, and the method further comprising: determining that adirection of a palm of the hand relative to a body of the user, wherein changing between the gaze control and the gesture control is further based on the direction of the palm.
[0071] Clause 14. The computing system of any one of clauses 9 to 13, wherein identifying the first action comprises identifying the first action via gaze control, wherein identify ing the second action comprises identifying the second action via gesture control.
[0072] Clause 15. The computing system of any one of clauses 9 to 13, wherein identifying the first action comprises identifying the first action via gaze control, wherein identify ing the second action comprises identifying the second action via gesture control.
[0073] Clause 16. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising: identifying a first action based on a gaze control associated with a user of a device; identify ing a portion of the user relative to a field of view associated with a gaze of the user; changing, based on the portion of the user relative to the field of view, from the gaze control to a gesture control associated with the portion of the user; and identifying a second action based on the gesture control.
[0074] Clause 17. The computer-readable storage medium of clause 16, wherein the method further comprises: changing, based on the portion of the user relative to the field of view, from the gesture control associated wi th the portion of the user to the gaze control; and identifying a third action based on the gaze control.
[0075] Clause 18. The computer-readable storage medium of clause 16 or 17, wherein the portion of the user comprises a hand, and the method further comprises: determining an orientation of the hand relative to content on the device, wherein changing from the gaze control to the gesture control is further based on the orientation.
[0076] Clause 19. The computer-readable storage medium of any one of clauses 16 to18, wherein the method further comprises: determining an angle between a direction of the device and content, wherein changing from the gaze control to the gesture control is further based on the angle.
[0077] Clause 20. The computer-readable storage medium of any one of clauses 16 to19. wherein the method further comprises: determining a velocity associated with the portion of the user, wherein changing from gaze control to the gesture control is further based on the velocity.
[0078] In this specification and the appended claims, the singular forms "a." “an"’ and ‘"the” do not exclude the plural reference unless the context dictates otherwise. Further, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictatesotherwise. For example, “A and / or B” includes A alone, B alone, and A with B. Further, connecting lines or connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physical connections, or logical connections may be present in a practical device. Moreover, no item or component is essential to the practice of the implementations disclosed herein unless the element is specifically described as ‘"essential” or “critical.”
[0079] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that a precise value or range thereof is not required and need not be specified. As used herein, the terms discussed above will have ready and instant meaning to one of ordinary skill in the art.
[0080] Moreover, the use of terms such as up, down, top, bottom, side, end, front, back, etc. herein are used concerning a currently considered or illustrated orientation. If they are considered concerning another orientation, such terms must be correspondingly modified.
[0081] Further, in this specification and the appended claims, the singular forms “a,” “an” and “the” do not exclude the plural reference unless the context dictates otherwise. Moreover, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B.
[0082] Although certain example methods, apparatuses, and articles of manufacture have been described herein, the scope of coverage of this patent is not limited thereto. It is to be understood that the terminology employed herein is to describe aspects and is not intended to be limiting. On the contrary7, this patent covers all methods, apparatus, and articles of manufacture fairly falling within the scope of the claims of this patent.
Claims
WHAT IS CLAIMED IS:
1. A method comprising: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing between a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
2. The method of claim 1, wherein the portion of the user comprises a hand, and the method further comprising: determining an orientation of the hand relative to content displayed by the device, wherein changing between the gaze control and the gesture control is further based on the orientation.
3. The method of claim 1 or 2, further comprising: determining an angle between a direction of the device and content displayed by the device, wherein changing between the gaze control and the gesture control is further based on the angle.
4. The method of any one of claims 1 to 3, wherein the portion of the user comprises a hand, and the method further comprising: determining an angle between the hand and a center of the field of view, wherein changing between the gaze control and the gesture control is further based on the angle.
5. The method of any one of claims 1 to 4, wherein the portion of the user comprises a hand, and the method further comprising: determining a direction of a palm of the hand relative to a body of the user, wherein changing between the gaze control and the gesture control is further based on the direction of the palm.
6. The method of any one of claims 1 to 5, further comprising: determining a movement velocity associated with the portion of the user, wherein changing between the gaze control and the gesture control is further based on the movement velocity.
7. The method of any one of claims 1 to 6, wherein identifying the first action comprises identifying the first action via gaze control, wherein identifying the second action comprises identifying the second action via gesture control.
8. The method of any one of claims 1 to 6, wherein identifying the first action comprises identifying the first action via gesture control, wherein identifying the second action comprises identifying the second action via gaze control.
9. A computing system comprising: a computer-readable storage medium; at least one processor operatively coupled to the computer-readable storage medium; and program instructions stored on the computer-readable storage medium that, when executed by the at least one processor, direct the computing system to perform a method, the method comprising: identifying a first action associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; based on the portion of the user relative to the field of view, changing between a gaze control and a gesture control associated with the portion of the user; and identifying a second action.
10. The computing system of claim 9, wherein the portion of the user comprises a hand, and the method further comprising: determining an orientation of the hand relative to content on the device.wherein changing between the gaze control and the gesture control is further based on the orientation.
11. The computing system of claim 9 or 10, wherein the method further comprises: determining an angle betw een a direction of the device and content, wherein changing between the gaze control and the gesture control is further based on the angle.
12. The computing system of any one of claims 9 to 11, wherein the portion of the user comprises a hand, and the method further comprising: determining an angle between the hand and a center of the field of view, wherein changing between the gaze control and the gesture control is further based on the angle.
13. The computing system of any one of claims 9 to 12, wherein the portion of the user comprises a hand, and the method further comprising: determining that a direction of a palm of the hand relative to a body of the user, wherein changing between the gaze control and the gesture control is further based on the direction of the palm.
14. The computing system of any one of claims 9 to 13, wherein identifying the first action comprises identifying the first action via gaze control, wherein identifying the second action comprises identifying the second action via gesture control.
15. The computing system of any one of claims 9 to 13, wherein identifying the first action comprises identifying the first action via gaze control, wherein identifying the second action comprises identifying the second action via gesture control.
16. A computer-readable storage medium having program instructions stored thereon that, when executed by at least one processor, direct the at least one processor to perform a method, the method comprising: identifying a first action based on a gaze control associated with a user of a device; identifying a portion of the user relative to a field of view associated with a gaze of the user; changing, based on the portion of the user relative to the field of view, from the gaze control to a gesture control associated with the portion of the user; and identifying a second action based on the gesture control.
17. The computer-readable storage medium of claim 16, wherein the method further comprises: changing, based on the portion of the user relative to the field of view, from the gesture control associated with the portion of the user to the gaze control; and identifying a third action based on the gaze control.
18. The computer-readable storage medium of claim 16 or 17, wherein the portion of the user comprises a hand, and the method further comprises: determining an orientation of the hand relative to content on the device. wherein changing from the gaze control to the gesture control is further based on the orientation.
19. The computer-readable storage medium of any one of claims 16 to 18, wherein the method further comprises: determining an angle between a direction of the device and content, wherein changing from the gaze control to the gesture control is further based on the angle.
20. The computer-readable storage medium of any one of claims 16 to 19, wherein the method further comprises: determining a velocity associated with the portion of the user, wherein changing from gaze control to the gesture control is further based on the velocity.
Citation Information
Patent Citations
Artificial reality multi-modal input switching model
US11294475B1
Integration of Artificial Reality Interaction Modes
US20230244321A1