Gesture input by hand for wearable system
Hand gesture recognition in augmented reality systems addresses the challenge of user interaction by detecting feature points to align interaction points and form multi-DOF controllers, enhancing precision and reducing errors in virtual reality environments.
Patent Information
- Application Number
- JP2025136614
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-19
- Filing Date
- 2025-08-19
- Publication Date
- 2025-12-23
AI Technical Summary
Existing augmented reality systems lack efficient and intuitive methods for user interaction, particularly in virtual reality and mixed reality environments, leading to high error rates and difficulty in interpreting complex 3D movements.
The use of hand gestures is employed to interact with virtual environments by detecting and analyzing feature points on a user's hand to determine specific gestures, aligning interaction points, and forming multi-DOF controllers, which can include tracking a subset of feature points for computational efficiency.
This approach provides precise and intuitive user interaction, reducing system complexity and error rates, allowing for accurate manipulation of virtual objects in 3D space.
Smart Images

Figure 2025186251000001_ABST
Abstract
Description
[Background technology]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 981,934, filed February 26, 2020, entitled "HAND GESTURE INPUT FOR WEARABLE SYSTEM," and U.S. Provisional Patent Application No. 63 / 027,272, filed May 19, 2020, entitled "HAND GESTURE INPUT FOR WEARABLE SYSTEM," the entire contents of which are incorporated herein by reference for all purposes.
[0002] Modern computing and display technology has facilitated the development of systems for so-called "virtual reality" or "augmented reality" experiences, in which digitally reproduced images, or portions thereof, are presented to a user in a manner that appears or can be perceived as real. Virtual reality, or "VR," scenarios typically involve the presentation of digital or virtual image information without transparency to other actual, real-world visual input, while augmented reality, or "AR," scenarios typically involve the presentation of digital or virtual image information as an extension to the user's visualization of the real world around them.
[0003] Despite the advances made in these display technologies, there remains a need in the art for improved methods, systems, and devices relating to augmented reality systems, and particularly display systems. Summary of the Invention [Means for solving the problem]
[0004] The present disclosure relates generally to techniques for improving the performance and user experience of optical systems. More specifically, embodiments of the present disclosure provide methods for operating an augmented reality (AR), virtual reality (VR), or mixed reality (MR) wearable system in which a user's hand gestures are used to interact within a virtual environment.
[0005] A description of various embodiments of the invention is provided below as a list of examples. As used below, any reference to a series of examples shall be understood as a disjunctive reference to each of those examples (e.g., "Examples 1-4" shall be understood as "Examples 1, 2, 3, or 4").
[0006] Example 1 is a method for interacting with a virtual object, the method including receiving an image of a user's hand; analyzing the image to detect a plurality of feature points associated with the user's hand; determining, based on the image analyzing step, whether the user's hand is making or transitioning to make a gesture from a plurality of gestures; in response to determining that the user's hand is making or transitioning to make a gesture, determining a specific location for the plurality of feature points, the specific location determined based on the plurality of feature points and the gesture; aligning an interaction point at the specific location; and forming a multi-DOF controller for interacting with the virtual object based on the interaction point.
[0007] Example 2 is a system configured to implement the method described in Example 1.
[0008] Example 3 is a non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of Example 1.
[0009] Example 4 is a method of interacting with a virtual object, the method including receiving an image of a user's hand from one or more image capture devices of a wearable system; analyzing the image to detect a plurality of feature points associated with the user's hand; determining, based on the analyzing the image, whether the user's hand is performing or transitioning to perform a particular gesture from a plurality of gestures; selecting a subset of the plurality of feature points that corresponds to the particular gesture in response to determining that the user's hand is performing or transitioning to perform the particular gesture; determining a particular location for the subset of the plurality of feature points, the particular location determined based on the subset of the plurality of feature points and the particular gesture; aligning an interaction point with the particular location; aligning a proximal point with a location along the user's body; casting a light ray from the proximal point through the interaction point; and forming a multi-DOF controller for interacting with the virtual object based on the light ray.
[0010] Example 5 is the method of example 4, wherein the plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.
[0011] Example 6 is the method of example 4, wherein the subset of the plurality of feature points is selected from a plurality of subsets of the plurality of feature points, each of the plurality of subsets of the plurality of feature points corresponding to a different gesture from a plurality of gestures.
[0012] Example 7 is the method of example 4, further comprising displaying a graphical representation of the multi-DOF controller.
[0013] Example 8 is the method of example 4, wherein the proximal point is aligned at an estimated location of the user's shoulder, an estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.
[0014] Example 9 is the method of example 4, further comprising capturing, with the image capture device, an image of the user's hand.
[0015] Example 10 is the method of example 9, wherein the image capture device is an element of a wearable system.
[0016] Example 11 is the method of example 9, wherein the image capture device is mounted in a headset of the wearable system.
[0017] Example 12 is the method of example 4, further comprising determining whether the user's hand is performing an action event based on analyzing the image.
[0018] Example 13 is the method of example 12, further comprising modifying the virtual object based on the multi-DOF controller and the action event in response to determining that the user's hands are performing the action event.
[0019] Example 14 is the method of example 13, wherein the user's hand is determined to be performing an action event based on a particular gesture.
[0020] Example 15 is the method of example 4, wherein the user's hand is determined to be performing or transitioning to perform a particular gesture based on a plurality of feature points.
[0021] Example 16 is the method of example 15, wherein the user's hand is determined to be performing or transitioning to performing a particular gesture based on neural network estimation using multiple feature points.
[0022] Example 17 is the method of example 4, wherein the user's hand is determined to be performing or transitioning to perform a particular gesture based on neural network inference using the image.
[0023] Example 18 is the method described in example 4, wherein the plurality of feature points are on the user's hand.
[0024] Example 19 is the method of example 4, wherein the multi-DOF controller is a 6-DOF controller.
[0025] Example 20 is a system configured to perform the method described in any of Examples 4-19.
[0026] Example 21 is a non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method according to any of Examples 4-19.
[0027] Example 22 is a method, comprising the steps of: receiving a sequence of images of a user's hand; analyzing each image in the sequence of images to detect a plurality of feature points on the user's hand; determining, based on the analyzing one or more images in the sequence of images, whether the user's hand is performing or transitioning to perform any of a plurality of different gestures; and, in response to determining that the user's hand is performing or transitioning to perform a particular one of the plurality of different gestures, selecting a particular location for the plurality of feature points corresponding to the particular gesture from among a plurality of locations for the plurality of feature points, each corresponding to the plurality of different gestures; and selecting a particular location for the plurality of feature points from among a plurality of different subsets of the plurality of feature points, each corresponding to the plurality of different gestures. selecting a particular subset of the plurality of feature points corresponding to a particular gesture; aligning an interaction point to a particular location on the user's hand relative to the plurality of feature points while the user's hand is determined to be performing or transitioning to perform the particular gesture; aligning a proximal point to an estimated location of the user's shoulder, an estimated location of the user's elbow, or a location along the user's upper arm between the estimated location of the user's shoulder and the estimated location of the user's elbow; casting a light ray from the proximal point through the interaction point; displaying a graphical representation of the multi-DOF controller corresponding to the light ray; and repositioning and / or reorienting the multi-DOF controller based on the locations of the interaction point, the proximal point, and the particular subset of the plurality of feature points.
[0028] Example 23 is the method of example 22, wherein the sequence of images is received from one or more outward-facing cameras on the headset.
[0029] Example 24 is the method of example 22, wherein the plurality of different gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.
[0030] Example 25 includes aligning an interaction point with a feature point along the user's index finger while the user's hand is determined to be performing a grasping gesture, and determining, at least in part, the specific location of the interaction point, the location of at least a portion of the user's body other than the user's hand, and / or the feature point I. m , T m , M m and determining an orientation or direction of the ray based on the relative positions of a subset of the plurality of feature points, the subset including three or more of H.
[0031] Example 26 includes, while determining that the user's hand is performing a pointing gesture, aligning the interaction point with a feature point at the tip of the user's index finger, and determining, at least in part, the specific location of the interaction point, the location of at least a portion of the user's body other than the user's hand, and / or the location of the feature point. t , I d , I p , I m , T t , T i , T m , M m determining an orientation or direction of the ray based, at least in part, on the relative positions of a subset of the plurality of feature points, the subset including three or more of H; [ka] The angle θ measured between [ka] [ka] 23. The method of example 22, further comprising detecting an action event based on the
[0032] Example 27 is the method of example 26, wherein a hover action event is detected when θ is determined to be greater than a predetermined threshold.
[0033] Example 28 is the method of example 26, wherein a touch action event is detected when θ is determined to be less than a predetermined threshold.
[0034] Example 29 is a diagram illustrating a method for pinching an interaction point while determining that the user's hand is performing a pinch gesture. [ka] and aligning, at least in part, with a particular location of the interaction point, a location of at least a part of the user's body other than the user's hand, and / or a feature point I. t , I d , I p , I m , T t , T i , T m , M m determining an orientation or direction of the ray based, at least in part, on the relative positions of a subset of the plurality of feature points, the subset including three or more of H; [ka] The angle θ measured between [ka] [ka] 23. The method of example 22, further comprising detecting an action event based on the
[0035] Example 30 is the method of example 29, wherein a hover action event is detected when θ is determined to be greater than a predetermined threshold.
[0036] Example 31 is the method of example 29, wherein a touch action event is detected when θ is determined to be less than a predetermined threshold.
[0037] Example 32 is the method of example 29, in which the tap action event is detected based on the duration of time over which θ is determined to be less than a predetermined threshold.
[0038] Example 33 is the method of example 29, in which a holding action event is detected based on the duration of time over which θ is determined to be less than a predetermined threshold.
[0039] Example 34 is a method for determining whether a user's hand is transitioning between performing a grasp gesture and a point gesture by using a gesture to move the interaction point. [ka] and determining the orientation or direction of the ray in the same manner as is done for point gestures, at least in part based on the specific location of the interaction point, the location of at least a part of the user's body other than the user's hand, and / or the feature point I. t , I d , I p , I m , T t , T i , T m , M m and determining an orientation or direction of the ray based on the relative positions of a subset of the plurality of feature points, the subset including three or more of H.
[0040] Example 35 is the method of example 34, wherein the user's hand is determined to be transitioning between performing a grasp gesture and a point gesture when the user's index finger is partially extended outward while the other fingers of the user's hand are curled inward.
[0041] Example 36 illustrates a method for determining whether a user's hand is transitioning between performing a point gesture and a pinch gesture by using a gesture to move an interaction point. [ka] and determining the orientation or direction of the light beam in the same manner as is done for point gestures and / or pinch gestures, and determining, at least in part, the specific location of the interaction point, the location of at least a part of the user's body other than the user's hand, and / or the feature point I. t , I d , I p , I m , T t , T i , T m , M m and determining an orientation or direction of the ray based on the relative positions of a subset of the plurality of feature points, the subset including three or more of H.
[0042] Example 37 is the method of example 36, wherein the user's hand is determined to be transitioning between performing a point gesture and a pinch gesture when the user's thumb and index finger are at least partially extended outward and at least partially curled toward each other.
[0043] Example 38 is a method for determining whether a user's hand is transitioning between performing a pinch gesture and performing a grasp gesture, and for determining whether the user's hand is transitioning between performing a pinch gesture and a grasp gesture. [ka] and determining the orientation or direction of the light beam in the same manner as is done for a pinch gesture, at least in part based on the specific location of the interaction point, the location of at least a part of the user's body other than the user's hand, and / or the feature point I. t , I d , I p , I m , T t , T i , T m , M m and determining an orientation or direction of the ray based on the relative positions of a subset of the plurality of feature points, the subset including three or more of H.
[0044] Example 39 is the method of example 38, wherein the user's hand is determined to be transitioning between performing a pinch gesture and performing a grasp gesture when the user's thumb and index finger are at least partially extended outward and at least partially curled toward each other.
[0045] Example 40 is a system configured to perform the method of any of Examples 22-39.
[0046] Example 41 is a non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method described in any of Examples 22-39.
[0047] Example 42 is a method of interacting with a virtual object, the method including receiving one or more images of a first hand and a second hand of a user; analyzing the one or more images to detect a plurality of feature points associated with each of the first hand and the second hand; determining an interaction point for each of the first hand and the second hand based on the plurality of feature points associated with each of the first hand and the second hand; generating one or more bimanual deltas based on the interaction points for each of the first hand and the second hand; and interacting with the virtual object using the one or more bimanual deltas.
[0048] Example 43 is the method of example 42, further comprising determining, for each of the first hand and the second hand, a two-hand interaction point based on the interaction point.
[0049] Example 44 is the method described in Example 42, wherein the interaction points for the first hand are determined based on a plurality of feature points associated with the first hand, and the interaction points for the second hand are determined based on a plurality of feature points associated with the second hand.
[0050] Example 45 is the method of Example 42, in which the step of determining an interaction point for each of the first hand and the second hand includes the steps of: determining whether the first hand is performing or transitioning to performing a first specific gesture from a plurality of gestures based on analyzing the one or more images; selecting a subset of the plurality of feature points associated with the first hand that corresponds to the first specific gesture in response to determining that the first hand is performing or transitioning to perform the first specific gesture; determining a first specific location for the subset of the plurality of feature points associated with the first hand, wherein the first specific location is determined based on the subset of the plurality of feature points associated with the first hand and the first specific gesture; and aligning the interaction point for the first hand to the first specific location.
[0051] Example 46 is the method of Example 45, wherein, for each of the first hand and the second hand, determining an interaction point further includes: determining, based on analyzing the one or more images, whether the second hand is performing or transitioning to perform a second specific gesture from the plurality of gestures; selecting, in response to determining that the second hand is performing or transitioning to perform the second specific gesture, a subset of the plurality of feature points associated with the second hand that corresponds to the second specific gesture; determining a specific location for the second subset of the plurality of feature points associated with the second hand, wherein the second specific location is determined based on the subset of the plurality of feature points associated with the second hand and the second specific gesture; and aligning the interaction point for the second hand to the second specific location.
[0052] Example 47 is the method of example 46, wherein the plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.
[0053] Example 48 is the method of example 42, wherein the one or more images include a first image of a first hand and a second image of a second hand.
[0054] Example 49 is the method of example 42, wherein the one or more images include a single image of a first hand and a second hand.
[0055] Example 50 is the method of example 42, wherein the one or more images include a series of time-series images.
[0056] Example 51 is the method of example 42, wherein one or more two-hand deltas are determined based on frame-to-frame movements of the interaction points for each of the first hand and the second hand.
[0057] Example 52 is the method of example 51, wherein the one or more two-hand deltas include a translation delta corresponding to a frame-to-frame translation of the interaction point for each of the first hand and the second hand.
[0058] Example 53 is the method of example 51, wherein the one or more two-hand deltas include a rotational delta corresponding to a frame-to-frame rotational movement of the interaction point for each of the first hand and the second hand.
[0059] Example 54 is the method of example 51, wherein the one or more two-hand deltas include a sliding delta corresponding to frame-to-frame separation movement of the interaction point for each of the first hand and the second hand.
[0060] Example 55 is a system configured to perform the method described in any of Examples 42-54.
[0061] Example 56 is a non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method described in any of Examples 42-54. [Brief explanation of the drawings]
[0062] The accompanying drawings, which are included to provide a further understanding of the present disclosure, are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure, and together with the detailed description, serve to explain the principles of the present disclosure. No attempt is made to show structural details of the present disclosure in more detail than may be necessary for a fundamental understanding of the disclosure and various ways in which it may be practiced.
[0063] [Figure 1] FIG. 1 illustrates an example operation of a wearable system that provides hand gesture input for interacting with virtual objects.
[0064] [Figure 2] FIG. 2 illustrates a schematic diagram of an exemplary AR / VR / MR wearable system.
[0065] [Figure 3] FIG. 3 illustrates an exemplary method for interacting with a virtual user interface.
[0066] [Figure 4A] FIG. 4A illustrates an example of a ray and cone projection.
[0067] [Figure 4B] FIG. 4B illustrates an example of a cone projection onto a group of objects.
[0068] [Figure 5] FIG. 5 illustrates examples of various features that may be detected or tracked by a wearable system.
[0069] [Figure 6-1] 6A-6F illustrate examples of possible subsets of feature points that may be selected based on gestures identified by the wearable system. [Figure 6-2] 6A-6F illustrate examples of possible subsets of feature points that may be selected based on gestures identified by the wearable system. [Figure 6-3] 6A-6F illustrate examples of possible subsets of feature points that may be selected based on gestures identified by the wearable system.
[0070] [Figure 7] 7A-7C illustrate examples of ray casting for various gestures while the user's arm is extended outward.
[0071] [Figure 8] 8A-8C illustrate examples of ray casting for various gestures while the user's arms are retracted inward.
[0072] [Figure 9] FIG. 9 illustrates an example of how action events can be detected using feature points.
[0073] [Figure 10] 10A-10C illustrate an example interaction with a virtual object using light rays.
[0074] [Figure 11] FIG. 11 illustrates an exemplary scheme for managing pointing gestures.
[0075] [Figure 12] FIG. 12 illustrates an exemplary scheme for managing pinch gestures.
[0076] [Figure 13] FIG. 13 illustrates an exemplary scheme for detecting action events while a user's hand is making a grasping gesture.
[0077] [Figure 14] FIG. 14 illustrates an exemplary scheme for detecting action events while a user's hand is performing a pointing gesture.
[0078] [Figure 15] FIG. 15 illustrates an exemplary scheme for detecting action events while a user's hands are performing a pinch gesture.
[0079] [Figure 16] FIG. 16 illustrates exemplary experimental data for detecting action events while a user's hand is performing a pinch gesture.
[0080] [Figure 17]17A-17D illustrate exemplary experimental data for detecting action events while a user's hand is performing a pinch gesture.
[0081] [Figure 18] FIG. 18 illustrates an exemplary scheme for detecting action events while a user's hands are performing a pinch gesture.
[0082] [Figure 19] 19A-19D illustrate example noise experiment data for detecting action events while a user's hand is performing a pinch gesture.
[0083] [Figure 20] 20A-20C illustrate an exemplary scheme for managing a grasping gesture.
[0084] [Figure 21] 21A-21C illustrate an exemplary scheme for managing pointing gestures.
[0085] [Figure 22] 22A-22C illustrate an exemplary scheme for managing pinch gestures.
[0086] [Figure 23] FIG. 23 illustrates various activation types for point and pinch gestures.
[0087] [Figure 24] FIG. 24 illustrates various gestures and transitions between gestures.
[0088] [Figure 25] FIG. 25 illustrates an example of two-handed interaction.
[0089] [Figure 26]FIG. 26 illustrates an example of two-handed interaction.
[0090] [Figure 27] FIG. 27 illustrates various examples of collaborative two-handed interaction.
[0091] [Figure 28] FIG. 28 illustrates an example of controlled two-handed interaction.
[0092] [Figure 29] FIG. 29 illustrates exemplary one-handed and two-handed interaction fields.
[0093] [Figure 30] FIG. 30 illustrates a method for forming a multi-DOF controller associated with a user's hand to allow the user to interact with virtual objects.
[0094] [Figure 31] FIG. 31 illustrates a method for forming a multi-DOF controller associated with a user's hand to allow the user to interact with virtual objects.
[0095] [Figure 32] FIG. 32 illustrates how two-handed input can be used to interact with virtual objects.
[0096] [Figure 33] FIG. 33 illustrates a simplified computer system. DETAILED DESCRIPTION OF THE INVENTION
[0097] Detailed Description of Specific Embodiments Wearable systems can present interactive augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) environments in which virtual data elements are interacted with by a user through various inputs. While many modern computing systems are engineered to generate a given output based on a single direct input in data-rich and dynamic interactive environments such as AR / VR / MR environments (e.g., a computer mouse can guide a cursor in response to a user's direct manipulation), a high degree of specificity may be desirable to accomplish certain tasks. Otherwise, in the absence of precise input, the computing system may suffer from a high error rate, resulting in incorrect computer actions being performed. For example, when a user intends to move an object in three-dimensional (3D) space using a touchpad, the computing system may have difficulty interpreting the desired 3D movement using a device with an essentially two-dimensional (2D) input space.
[0098] The use of hand gestures as input within AR / VR / MR environments has several attractive features. First, in AR environments, in which virtual content is overlaid on the real world, hand gestures provide an intuitive interaction method that bridges both worlds. Second, a wide range of expressive hand gestures exists that can potentially be mapped to various input commands. For example, hand gestures can simultaneously exhibit several distinct parameters, such as hand shape (e.g., distinct configurations the hand can assume), orientation (e.g., distinct degrees of relative hand rotation), location, and movement. Third, with recent hardware improvements in imaging devices and processing units, hand gesture input offers sufficient accuracy so that system complexity can be reduced compared to other inputs, such as handheld controllers, which employ various sensors, such as electromagnetic tracking emitters / receivers.
[0099] One approach to recognizing hand gestures is to track the positions of various feature points on one or both of a user's hands. In one implementation, a hand tracking system may identify the 3D positions of more than 20 feature points on each hand. Gestures associated with the hand may then be recognized by analyzing the feature points. For example, the distance between different feature points may indicate whether the user's hand is fisted (e.g., a short average distance) or open and relaxed (e.g., a long average distance). As another example, various angles formed by three or more feature points (e.g., including at least one feature point along the user's index finger) may indicate whether the user's hand is pointing or pinching.
[0100] Once a gesture is recognized, an interaction point, through which the user may interact with the virtual object, can be determined. The interaction point may be aligned with one of the feature points or a location between the feature points, with each gesture having a unique algorithm for determining the interaction point. For example, when performing a point gesture, the interaction point may be aligned with the feature point at the tip of the user's index finger. As another example, when performing an open pinch gesture, the interaction point may be aligned with the midpoint between the tip of the user's index finger and the tip of the user's thumb. Certain gestures may further allow a radius associated with the interaction point to be determined. As an example, for a pinch gesture, the radius may be related to the distance between the tip of the user's index finger and the tip of the user's thumb.
[0101] After a gesture is recognized and / or interaction points are determined, continuing to track the entire network of feature points can be computationally burdensome. Therefore, in some embodiments of the present disclosure, a subset of the total number of feature points can continue to be tracked once a gesture is recognized. This subset of feature points can be used to periodically update the interaction points at a more manageable computational burden than would be the case if the total number of feature points were used. In some examples, this subset of feature points can be used to periodically update the orientation of a virtual multi-DOF controller (e.g., a virtual cursor or pointer associated with the interaction points) with relatively high computational efficiency, as described in further detail below. Furthermore, the subset of feature points can be analyzed to determine whether the user's hand is no longer performing a gesture or is transitioning, for example, from performing a first gesture to a second gesture, or from a first gesture to an unrecognized gesture.
[0102] In addition to determining the interaction point, a proximal point along the user's body (or in space) can be determined such that a control ray (or simply, a "ray") can be formed between two points extending. The ray (or a portion thereof) can serve as a cursor or pointer (e.g., as part of a multi-DOF controller) for interacting with virtual content in 3D space. In some instances, the proximal point may be aligned with the user's shoulder, the user's elbow, or along the user's arm (e.g., between the user's shoulder and elbow). The proximal point may alternatively be aligned with one or more other locations within or along the surface of the user's body, such as a knuckle, hand, wrist, forearm, elbow, arm (e.g., upper arm), shoulder, scapula, neck, head, eye, face (e.g., cheek), chest, torso (e.g., navel area), or a combination thereof. The ray may then extend a particular distance from the proximal point through the interaction point. The interaction point, proximity point, and light ray may each be dynamically updated to provide a responsive and comfortable user experience.
[0103] Embodiments herein relate to both single-hand interaction, referred to as single-hand interaction, and two-hand interaction, referred to as bimanual interaction. Tracking a single-hand posture may include tracking a single hand's interaction point (e.g., its position, orientation, and radius), optionally its corresponding proximal point and ray, and any gestures the hand is performing. For bimanual interaction, the interaction point (e.g., position, orientation, and radius) for each hand of the user, optionally a corresponding proximal point, ray, and gesture, may be tracked. Bimanual interaction may further involve tracking bimanual interaction points between the two hands, which may have positions (e.g., average positions), orientations (e.g., average orientations), and radii (e.g., average radii). Frame-to-frame movement of bimanual interaction points can be captured through bimanual deltas, which may be calculated based on the deltas for the two hands, as described below.
[0104] The two-hand deltas may include a translation component referred to as a translation delta and a rotation component referred to as a rotation delta. The translation delta may be determined based on the translation deltas for the two hands. For example, the translation delta may be determined based on a left translation delta (e.g., an average thereof) corresponding to the inter-frame translation movement of the user's left hand and a right translation delta corresponding to the inter-frame translation movement of the user's right hand. Similarly, the rotation delta may be determined based on the rotation deltas for the two hands. For example, the rotation delta may be determined based on a left rotation delta (e.g., an average thereof) corresponding to the inter-frame rotation movement of the user's left hand and a right rotation delta corresponding to the inter-frame rotation movement of the user's right hand.
[0105] Alternatively, or in addition, the rotation delta may be determined based on the rotation of a line formed between the locations of the interaction points. For example, a user may rotate a digital cube by pinching two corners of the cube and rotating the locations of the interaction points of the two hands. This rotation may occur independently of whether each hand's interaction point is rotated independently, or in some embodiments, the rotation of the cube may be further facilitated by the rotation of the interaction points. In some instances, the two-hand delta may include other components, such as a separation component, referred to as a separation delta (or scaling delta), which is determined based on the distance between the locations of the interaction points, with a positive separation delta corresponding to the hands moving apart and a negative separation delta corresponding to the hands moving closer together.
[0106] Various types of two-handed interactions can fall into one of three categories: the first category is independent two-handed interaction, in which each hand independently interacts with a virtual object (e.g., a user is typing on a virtual keyboard, and the configuration of each hand is independent of the other); the second category is collaborative two-handed interaction, in which both hands collaboratively interact with a virtual object (e.g., using both hands to resize, rotate, and / or translate a virtual cube by pinching opposite corners); and the third category is managed two-handed interaction, in which one hand manages how the other hand is interpreted (e.g., the right hand is the cursor, while the left hand is a modifier that toggles the cursor between a pen and an eraser).
[0107] In the following description, various examples will be described. For purposes of explanation, specific configurations and details are set forth to provide a thorough understanding of the examples. However, it will also be apparent to those skilled in the art that the examples may be practiced without the specific details. Additionally, well-known features may be omitted or simplified so as not to obscure the embodiments being described.
[0108] 1 illustrates an example operation of a wearable system providing hand gesture input for interacting with a virtual object 108, according to some embodiments of the present disclosure. The wearable system may include a wearable device 102 (e.g., a headset) worn by a user and including at least one forward-facing camera 104 that includes the user's hand 106 within its field of view (FOV). Images captured from the camera 104 may thus include the hand 106, allowing subsequent processing of the image to be performed by the wearable system, e.g., to detect feature points associated with the hand 106. In some embodiments, the wearable system and wearable device 102 described with reference to FIG. 1 may correspond to wearable system 200 and wearable device 201, respectively, as described in further detail below with reference to FIG. 2.
[0109] The wearable system may maintain a frame of reference within which the position and orientation of elements in the AR / VR / MR environment may be determined. In some embodiments, the wearable system may be configured to calculate a time domain (X WP , Y WP , Z WP ) relative to the reference frame (X WO , Y WO , Z WO), and an orientation ("wearable orientation"), defined as (X, Y, and Z). The position of the wearable device 102 may be expressed in X, Y, and Z Cartesian values, or longitude, latitude, and altitude values, among other possibilities. The orientation of the wearable device 102 may be expressed in X, Y, and Z Cartesian values, or pitch, yaw, and roll angle values, among other possibilities. The frame of reference for each position and orientation may be a world reference frame, or alternatively or additionally, the position and orientation of the wearable device 102 may be used as a frame of reference, such that, for example, the position of the wearable device 102 may be set as (0, 0, 0) and the orientation of the wearable device 102 may be set as (0°, 0°, 0°).
[0110] The wearable system may perform one or more processing steps 110 using images captured by camera 104. In some embodiments, one or more processing steps 110 may be performed by one or more processors, and may be performed at least in part by one or more processors of the wearable system, one or more processors communicatively coupled to the wearable system, or a combination thereof. In step 110-1, multiple feature points (e.g., nine or more feature points) are detected or tracked based on the captured image. In step 110-2, the tracked feature points are used to determine whether the hand 106 is performing, or is transitioning to performing, one of a set of predetermined gestures. In the illustrated embodiment, the hand 106 is determined to be performing a pinch gesture. Alternatively, or in addition, the gesture may be predicted directly from the image without the intermediate step of detecting feature points. Thus, steps 110-1 and 110-2 may be performed in parallel or sequentially in either order. In response to determining that the user's hand is performing or transitioning to perform a particular gesture (e.g., a pinch gesture), a subset of the plurality of feature points (e.g., eight or fewer feature points) associated with the particular gesture may be selected and tracked.
[0111] In step 110-3, an interaction point 112 is determined by aligning the interaction point 112 to a particular location relative to a selected subset of feature points based on the predicted gesture (or predicted gesture transition) from step 110-2. Also in step 110-3, a proximal point 114 is determined by aligning the proximal point 114 to a location along the user's body based, at least in part, on one or more of a variety of factors. Also in step 110-3, a light ray 116 is cast from the proximal point 114 through the interaction point 112. In step 110-4, an action event to be performed by the hand 106 is predicted based on the feature points (e.g., based on movement of the feature points over time). In the illustrated example, the hand 106 is determined to be performing a targeted action, which can be recognized by the wearable system when the user performs a dynamic opening pinch gesture.
[0112] 2 illustrates a schematic diagram of an exemplary AR / VR / MR wearable system 200 according to some embodiments of the present disclosure. The wearable system 200 may include a wearable device 201 and at least one remote device 203 that is remote from (e.g., separate hardware but communicatively coupled to) the wearable device 201. As noted above, in some embodiments, the wearable system 200 and the wearable device 201, as described with reference to FIG. 2, may correspond to the wearable system and the wearable device 102, respectively, as described above with reference to FIG. 1. The wearable device 201 is worn by the user (generally as a headset), while the remote device 203 may be held by the user (e.g., as a handheld controller) or mounted in a variety of configurations, such as fixedly attached to a frame, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user (e.g., in a backpack configuration, belt-attached configuration, etc.).
[0113] The wearable device 201 may include a left eyepiece 202A and a left lens assembly 205A arranged in a side-by-side configuration, and a right eyepiece 202B and a right lens assembly 205B also arranged in a side-by-side configuration. In some embodiments, the wearable device 201 includes one or more sensors, including, but not limited to, a left-front-facing world camera 206A mounted directly or near the left eyepiece 202A, a right-front-facing world camera 206B mounted directly or near the right eyepiece 202B, a left-side-facing world camera 206C mounted directly or near the left eyepiece 202A, and a right-side-facing world camera 206D mounted directly or near the right eyepiece 202B. The wearable device 201 may include one or more image projection devices, such as a left projector 214A optically linked to the left eyepiece 202A and a right projector 214B optically linked to the right eyepiece 202B.
[0114] The wearable system 200 may include a processing module 250 for collecting, processing, and / or controlling data within the system. Components of the processing module 250 may be distributed between the wearable device 201 and the remote device 203. For example, the processing module 250 may include a local processing module 252 on the wearable portion of the wearable system 200 and a remote processing module 256 that is physically separate from and communicatively linked to the local processing module 252. The local processing module 252 and the remote processing module 256 may each include one or more processing units (e.g., a central processing unit (CPU), a graphics processing unit (GPU), etc.) and one or more storage devices, such as non-volatile memory (e.g., flash memory).
[0115] The processing module 250 may collect data captured by various sensors of the wearable system 200, such as the camera 206, depth sensor 228, remote sensor 230, ambient light sensor, eye tracker, microphone, inertial measurement unit (IMU), accelerometer, compass, global navigation satellite system (GNSS) unit, wireless device, and / or gyroscope. For example, the processing module 250 may receive images 220 from the camera 206. Specifically, the processing module 250 may receive a left front image 220A from a left front-facing world camera 206A, a right front image 220B from a right front-facing world camera 206B, a left side image 220C from a left side-facing world camera 206C, and a right side image 220D from a right side-facing world camera 206D. In some embodiments, the images 220 may include a single image, a pair of images, a video comprising a stream of images, a video comprising a stream of paired images, and the like. Images 220 may be generated and transmitted to processing module 250 periodically while wearable system 200 is powered on, or may be generated in response to instructions transmitted by processing module 250 to one or more of the cameras.
[0116] The cameras 206 may be configured at various positions and orientations along the exterior of the wearable device 201 to capture images of the user's surroundings. In some instances, the cameras 206A and 206B may be positioned to capture images that substantially overlap the FOV of the user's left and right eyes, respectively. Thus, the placement of the cameras 206 may be near the user's eyes, but not so close as to obscure the user's FOV. Alternatively, or in addition, the cameras 206A and 206B may be positioned to align with the internal coupling locations of the virtual image lights 222A and 222B, respectively. The cameras 206C and 206D may be positioned to capture images to the side of the user, e.g., within or outside the user's peripheral vision. The images 220C and 220D captured using the cameras 206C and 206D do not necessarily overlap with the images 220A and 220B captured using the cameras 206A and 206B.
[0117] In various embodiments, processing module 250 may receive ambient light information from an ambient light sensor. The ambient light information may indicate a brightness value or a range of spatially resolved brightness values. Depth sensor 228 may capture depth image 232 in a front-facing direction of wearable device 201. Each value in depth image 232 may correspond to the distance between depth sensor 228 and the nearest detected object in a particular direction. As another example, processing module 250 may receive gaze information from one or more eye trackers. As another example, processing module 250 may receive projected image brightness values from one or both projectors 214. Remote sensor 230, located in remote device 203, may include any of the sensors described above with similar functionality.
[0118] Virtual content is delivered to a user of the wearable system 200 primarily using the projector 214 and the eyepieces 202. For example, the eyepieces 202A, 202B may each include a transparent or semi-transparent waveguide configured to direct and outcouple light generated by the projectors 214A, 214B. Specifically, the processing module 250 may cause the left projector 214A to output left virtual image light 222A onto the left eyepiece 202A and the right projector 214B to output right virtual image light 222B onto the right eyepiece 202B. In some embodiments, the eyepieces 202A, 202B may each include multiple waveguides corresponding to different colors. In some embodiments, lens assemblies 205A, 205B may be coupled to and / or integrated with the eyepieces 202A, 202B. For example, lens assemblies 205A, 205B may be incorporated into a multi-layer eyepiece and may form one or more layers that make up one of the eyepieces 202A, 202B.
[0119] During operation, the wearable system 200 can support various user interactions with objects within the field of view (FOR) (i.e., the overall area available for viewing or imaging) based on the contextual information. For example, the wearable system 200 can adjust the size of the cone opening through which the user interacts with objects using cone projection. As another example, the wearable system 200 can adjust the amount of virtual object movement associated with actuation of a user input device based on the contextual information. Detailed examples of these interactions are provided below.
[0120] The user's FOR can contain a group of objects, which can be perceived by the user via the wearable system 200. The objects in the user's FOR can be virtual and / or physical objects. Virtual objects may include operating system objects, such as a trash can for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, etc. Virtual objects may also include objects within an application, such as an avatar, a virtual object, graphic, or image within a game. Some virtual objects can be both operating system objects and objects within an application. In some embodiments, the wearable system 200 can add virtual elements to existing physical objects. For example, the wearable system 200 may add a virtual menu associated with a television in a room, and the virtual menu may give the user options to turn on the television or change its channel using the wearable system 200.
[0121] Objects in the user's FOR can be part of a world map. Data associated with objects (e.g., location, semantic information, properties, etc.) can be stored in various data structures, such as arrays, lists, trees, hashes, graphs, etc. An index for each stored object may be determined, if applicable, by the object's location. For example, the data structure may index objects by a single coordinate, such as the object's distance from a base position (e.g., distance to the left (or right) of the base position, distance from the top (or bottom) of the base position, or distance by depth from the base position). In some implementations, the wearable system 200 can display virtual objects to the user at different depth planes, such that interactable objects can be organized into multiple arrays such that they are located at different fixed depth planes.
[0122] A user can interact with a subset of the objects in the user's FOR. This subset of objects may sometimes be referred to as interactable objects. A user can interact with the objects using various techniques, such as by selecting the object, moving the object, opening a menu or toolbar associated with the object, or by selecting a new set of interactable objects. A user may interact with the interactable objects by actuating a user input device using hand gestures or postures, such as clicking on a mouse, tapping on a touchpad, swiping on a touchscreen, hovering over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a five-way D-pad), pointing a joystick, wand, or totem toward the object, pressing a button on a remote control, or other user interaction with an input device. A user may also interact with the interactable objects using head, eye, or body postures, such as gazing or pointing at the object for a period of time. These user hand gestures and postures can cause the wearable system 200 to initiate a selection event in which, for example, a user interface action is performed (associated with a target interactable object, a menu is displayed, a gaming action is performed on an avatar in a game, etc.).
[0123] FIG. 3 illustrates an exemplary method 300 for interacting with a virtual user interface according to some embodiments of the present disclosure. In step 302, the wearable system may identify a particular user interface (UI). The type of UI may be predetermined by the user. The wearable system may identify that a particular UI needs to be captured based on user input (e.g., gestures, visual data, audio data, sensory data, direct commands, etc.). In step 304, the wearable system may generate data for the virtual UI. For example, data associated with the UI's boundaries, general structure, shape, etc. may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system may display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine coordinates of the user's physical position, head pose, or eye pose so that a ring UI may be displayed around the user or a planar UI may be displayed on a wall or in front of the user. If the UI is hand-centered, map coordinates of the user's hand may be determined. These map points may be derived through data received through a FOV camera, sensory input, or any other type of collected data.
[0124] In step 306, the wearable system may transmit data from the cloud to the display, or the data may be transmitted from a local database to the display component. In step 308, a UI is displayed to the user based on the transmitted data. For example, a light field display may project the virtual UI into one or both of the user's eyes. Once the virtual UI is created, the wearable system may simply wait for a command from the user to generate additional virtual content on the virtual UI in step 310. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for a command (gesture, head or eye movement, input from a user input device, etc.), and if it is recognized (step 312), virtual content associated with the command may be displayed to the user (step 314). As an example, the wearable system may wait for a user's hand gesture before blending multiple stem tracks.
[0125] As described herein, a user can interact with objects in their environment using hand gestures or postures. For example, a user may look around a room and see a table, a chair, a wall, and a virtual television display on one of the walls. To determine the object the user is looking at, wearable system 200 may use a cone projection technique, generally described as projecting a cone toward it in the direction the user is looking and identifying any object that intersects with the cone. Cone projection may involve projecting a single ray of light, with no lateral thickness, from a headset (of wearable system 200) toward a physical or virtual object. Cone projection with a single ray of light may also be referred to as ray projection.
[0126] Ray casting can use a collision detection agent to trace along the ray and identify whether and where any objects intersect with the ray. The wearable system 200 can use an IMU (e.g., an accelerometer), an eye-tracking camera, etc. to track the user's pose (e.g., body, head, or eye direction) and determine the direction the user is looking toward. The wearable system 200 can use the user's pose to determine the direction in which to cast the ray. Ray casting techniques can also be used in conjunction with user input devices, such as handheld, multi-degree-of-freedom (DOF) input devices. For example, a user can actuate the multi-DOF input device to anchor the size and / or length of the ray while the user moves around. As another example, rather than casting a ray from a headset, the wearable system 200 can cast the ray from the user input device. In an embodiment, rather than casting a ray with negligible thickness, the wearable system can cast a cone with a non-negligible aperture (lateral to the central ray).
[0127] FIG. 4A illustrates an example of ray and cone projection, according to some embodiments of the present disclosure. The cone projection can project a conical (or other shaped) volume 420 with an adjustable aperture. The cone 420 can be a geometric cone having an interaction point 428 and a surface 432. The size of the aperture can correspond to the size of the cone's surface 432. For example, a large aperture can correspond to a large surface area of the surface 432. As another example, the large aperture can correspond to a large diameter 426 of the surface 432, while a small aperture can correspond to a small diameter 426 of the surface 432. As illustrated in FIG. 4A , the interaction point 428 of the cone 420 can have its origin at various locations, such as the center of the user's ARD (e.g., between the user's eyes), a point on one of the user's limbs (e.g., a hand, such as a finger), a user input device, or a totem (e.g., a toy weapon) being held or manipulated by the user. It should be understood that interaction point 428 represents one example of an interaction point that may be generated using one or more of the systems and techniques described herein, and that other interaction point arrangements are possible and within the scope of the present invention.
[0128] The central ray 424 can represent the direction of the cone. The direction of the cone can correspond to the user's body posture (head posture, hand gestures, etc.) or the user's gaze direction (also referred to as eye posture). Example 406 in FIG. 4A illustrates cone projection with posture, and the wearable system can use the user's head posture or eye posture to determine the cone direction 424. This example also illustrates a coordinate system for head posture. The head 450 may have multiple degrees of freedom. As the head 450 moves toward different directions, the head posture will change relative to the natural resting orientation 460. The coordinate system in FIG. 4A shows three angular degrees of freedom (e.g., yaw, pitch, and roll) that can be used to measure the head posture relative to the head's natural resting state 460. As illustrated in FIG. 4A, the head 450 can tilt forward and backward (e.g., pitch), turn left and right (e.g., yaw), and tilt laterally (e.g., roll). In other implementations, other techniques or angle representations for measuring head pose can be used, for example, any other type of Euler angle method. The wearable system may use an IMU to determine the user's head pose.
[0129] Example 404 shows another example of cone projection with pose, where the wearable system can determine the direction 424 of the cone based on the user's hand gestures. In this example, the interaction point 428 of the cone 420 is at the fingertip of the user's hand 414. As the user points their finger at a different location, the position of the cone 420 (and central ray 424) can be moved accordingly.
[0130] The direction of the cone can also correspond to the position or orientation of the user input device or the actuation of the user input device. For example, the direction of the cone can be based on a trajectory traced by the user on the touch surface of the user input device. The user can move their finger forward on the touch surface to indicate that the direction of the cone is forward. Example 402 illustrates another cone projection using a user input device. In this example, the interaction point 428 is located at the tip of the weapon-shaped user input device 412. As the user input device 412 is moved around, the cone 420 and central ray 424 can also move with the user input device 412.
[0131] The wearable system can initiate a cone projection when the user activates the user input device 466, for example, by clicking on a mouse, tapping on a touchpad, swiping on a touchscreen, hovering over or touching a capacitive button, pressing a key on a keyboard or game controller (e.g., a 5-way D-pad), pointing a joystick, wand, or totem towards an object, pressing a button on a remote control, or other interaction with the user input device 466.
[0132] The wearable system may also initiate cone projection based on the user's posture, such as a long-term gaze in one direction or a hand gesture (e.g., waving a hand in front of an outward-facing imaging system). In some implementations, the wearable system can automatically initiate cone projection events based on context information. For example, the wearable system can automatically initiate cone projection when the user is facing the main page of the AR display. In another example, the wearable system can determine the relative position of objects in the user's gaze direction. If the wearable system determines that objects are located relatively far from each other, the wearable system can automatically initiate cone projection, so that the user does not need to move with high precision to select an object within a group of sparsely located objects.
[0133] The direction of the cone can further be based on the position or orientation of the headset. For example, the cone may be projected in a first direction when the headset is tilted and in a second direction when the headset is not tilted.
[0134] The cone 420 may have various properties, such as, for example, size, shape, or color. These properties may be displayed to the user so that the cone is perceptible to the user. In some cases, a portion of the cone 420 may be displayed (e.g., the end of the cone, the surface of the cone, the central ray of the cone, etc.). In other embodiments, the cone 420 may be a rectangular prism, a polyhedron, a pyramid, a frustum, etc. The distal end of the cone can have any cross-section, for example, circular, oval, polygonal, or irregular.
[0135] 4A and 4B, the cone 420 can have an apex located at an interaction point 428 and a distal end formed at a plane 432. The interaction point 428 (also referred to as the zero point of the central ray 424) can be associated with a location from which the cone projection originates. The interaction point 428 may be anchored to a location in 3D space so that a virtual cone appears to emanate from that location. The location may be a position on the user's head (such as between the user's eyes), a user input device (such as a 6DOF handheld controller or a 3DOF handheld controller) that functions as a pointer, the tip of a finger (which may be detected by gesture recognition), etc. With respect to a handheld controller, the location to which the interaction point 428 is anchored may depend on the form factor of the device. For example, for a weapon-shaped controller 412 (for use in a shooting game), the interaction point 428 may be at the tip of the muzzle of the controller 412. In this example, the cone interaction point 428 may occur at the center of the gun barrel, and the cone 420 (or central ray 424) of the cone 420 may be projected forward such that the center of the cone projection will be concentric with the barrel of the weapon-shaped controller 412. The cone interaction point 428 may, in various embodiments, be anchored anywhere within the user's environment.
[0136] Once the interaction point 428 of the cone 420 is anchored to a location, the orientation and movement of the cone 420 may be based on the movement of an object associated with that location. For example, as described with reference to example 406, if the cone is anchored to a user's head, the cone 420 can move based on the user's head pose. As another example, in example 402, if the cone 420 is anchored to a user input device, the cone 420 can be moved based on actuation of the user input device, such as a change in the position or orientation of the user input device. As another example, in example 404, if the cone 420 is anchored to a user's hand, the cone 420 can be moved based on movement of the user's hand.
[0137] The surface 432 of the cone can extend until it reaches an exit threshold. The exit threshold may involve a collision between the cone and a virtual or physical object (e.g., a wall) in the environment. The exit threshold may also be based on a threshold distance. For example, the surface 432 can continue to extend away from the interaction point 428 until the cone collides with the object or until the distance between the surface 432 and the interaction point 428 reaches a threshold distance (e.g., 20 centimeters, 1 meter, 2 meters, 10 meters, etc.). In some embodiments, the cone can extend beyond the object even if a collision could occur between the cone and the object. For example, the surface 432 can extend through a real-world object (a table, chair, wall, etc.) and terminate upon hitting the exit threshold. Assuming the exit threshold is a wall of a virtual room located outside the user's current room, the wearable system can extend the cone beyond the current room until it reaches the surface of the virtual room. In one embodiment, a world mesh can be used to define the extents of one or more rooms. The wearable system can detect the presence of an exit threshold by determining whether the virtual cone intersects with a portion of the world mesh. In some embodiments, the user can easily target a virtual object once the cone extends through the real-world object. As an example, the headset can present a virtual hole in a physical wall through which the user can remotely interact with virtual content in other rooms, even when the user is not physically present in those rooms.
[0138] The cone 420 may have a depth. The depth of the cone 420 may be represented by the distance between the interaction point 428 and the surface 432. The depth of the cone may be adjusted automatically by the wearable system, by the user, or a combination. For example, if the wearable system determines that an object is located far from the user, the wearable system may increase the depth of the cone. In some implementations, the depth of the cone may be anchored to a depth plane. For example, a user may select the depth of the cone to be anchored to a depth plane within one meter of the user. As a result, during the cone projection, the wearable system will not capture objects outside the one-meter boundary. In one embodiment, if the depth of the cone is anchored to a depth plane, the cone projection will only capture objects at that depth plane. Thus, the cone projection will not capture objects closer to or farther from the user than the anchored depth plane. In addition to, or as an alternative to, setting the depth of cone 420, the wearable system can set surface 432 to a depth plane so that the cone projection can enable user interaction with objects at or below the depth plane.
[0139] The wearable system can anchor the depth, interaction point 428, or surface 432 of the cone in response to detecting certain hand gestures, body posture, gaze direction, actuation of a user input device, voice commands, or other techniques. In addition to or in place of the examples described herein, the anchor location of interaction point 428, surface 432, or anchored depth can be based on contextual information, such as the type of user interaction or the function of the object to which the cone is anchored. For example, interaction point 428 can be anchored to the center of the user's head due to user availability and sensation. As another example, when a user points to an object using a hand gesture or a user input device, interaction point 428 can be anchored to the tip of the user's finger or the tip of the user input device to increase the accuracy of the direction the user points.
[0140] The wearable system can generate a visual representation of at least a portion of the cone 420 or the ray 424 for display to the user. Properties of the cone 420 or the ray 424 may be reflected in the visual representation of the cone 420 or the ray 424. The visual representation of the cone 420 may correspond to at least a portion of the cone, such as the cone's opening, the cone's surface, or a central ray. For example, if the virtual cone is a geometric cone, the visual representation of the virtual cone may include a gray geometric cone extending from a position between the user's eyes. As another example, the visual representation may include a portion of the cone that interacts with real or virtual content. Assuming the virtual cone is a geometric cone, the visual representation may include a circular pattern representing the base of the geometric cone, because the base of the geometric cone may be used to target and select virtual objects. In an embodiment, the visual representation is triggered based on a user interface action. As an example, the visual representation may be associated with a state of the object. The wearable system can present a visual representation when the object changes from a resting or hovering state (the object can be moved or selected). The wearable system can further hide the visual representation when the object changes from a hovering state to a selected state. In some implementations, when the object is in the hovering state, the wearable system can receive input from a user input device (in addition to or as an alternative to the cone projection) and can allow a user to select a virtual object using the user input device when the object is in the hovering state.
[0141] In some embodiments, the cone 420, the beam of light 424, or portions thereof may be invisible to the user (e.g., may not be displayed for the user). The wearable system may assign a focus indicator to one or more objects that indicates the direction and / or location of the cone. For example, the wearable system may assign a focus indicator to an object that is in front of the user and intersects the user's line of sight. The focus indicator may comprise a halo, a change in color, a change in perceived size or depth (e.g., causing a target object to appear closer and / or larger when selected), a change in the shape of a cursor sprite graphic (e.g., the cursor changes from a circle to an arrow), or other audible, tactile, or visual effect that attracts the user's attention. The cone 420 may have an opening laterally relative to the beam of light 424. The size of the opening may correspond to the size of the surface 432 of the cone. For example, the large opening may correspond to the large diameter 426 on the surface 432, while the small opening may correspond to the small diameter 426 on the surface 432.
[0142] 4B , the aperture can be adjusted by the user, the wearable system, or a combination. For example, the user may adjust the aperture through a user interface action, such as selecting an aperture option shown on the AR display. The user may also adjust the aperture by actuating a user input device, for example, by scrolling the user input device or by pressing a button to anchor the aperture size. In addition to or as an alternative to input from the user, the wearable system can update the aperture size based on one or more contextual factors.
[0143] Cone projection can be used to increase accuracy when interacting with objects in the user's environment, especially when those objects are located at a distance where small amounts of movement from the user can translate into large movements of light rays. Cone projection can also be used to reduce the amount of movement required from the user to overlap the cone with one or more virtual objects. In some implementations, the user can manually update the cone opening to improve the speed and accuracy of selecting a target object, for example, by using a narrower cone when many objects are present and a wider cone when fewer objects are present. In other implementations, the wearable system can determine contextual factors associated with objects in the user's environment and allow automatic cone updates, in addition to or as an alternative to manual updates, which can advantageously make it easier for the user to interact with objects in the environment because less user input is required.
[0144] FIG. 4B illustrates an example of a cone or ray being cast onto a group of objects 430 (e.g., objects 430A, 430B) in a user's FOR 400. The objects may be virtual and / or physical objects. During cone or ray casting, the wearable system casts a cone 420 or ray 424 (visible or invisible to the user) in a direction and can identify any objects that intersect with the cone 420 or ray 424. For example, object 430A (shown in bold) intersects with the cone 420. Object 430B is outside of the cone 420 and does not intersect with the cone 420.
[0145] The wearable system can automatically update the aperture based on the context information. Context information may include information related to the user's environment (e.g., lighting conditions in the user's virtual or physical environment), user preferences, the user's physical conditions (e.g., whether the user is nearsighted), information associated with objects in the user's environment, such as the type of objects (e.g., physical or virtual) or layout of the objects in the user's environment (e.g., object density, object location and size, etc.), characteristics of the objects with which the user is interacting (e.g., object function, type of user interface action supported by the object, etc.), combinations thereof, or the like. Density can be measured in various ways, such as, for example, the number of objects per projected area, the number of objects per solid angle, etc. Density can also be represented in other ways, such as, for example, the spacing between nearby objects (smaller spacing reflects increased density). The wearable system can use the object location information to determine the layout and density of objects in an area. As shown in FIG. 4B, the wearable system can determine that a group of objects 430 is dense. A wearable system may therefore use a cone 420 with a smaller opening.
[0146] The wearable system can dynamically update the aperture (e.g., size or shape) based on the user's pose. For example, the user may initially point toward group of objects 430 in FIG. 4B, but as the user moves their hand, the user may now point to a group of objects that are sparsely positioned relative to one another. As a result, the wearable system may increase the size of the aperture. Similarly, if the user moves their hand back toward group of objects 430, the wearable system may decrease the size of the aperture.
[0147] Additionally or alternatively, the wearable system can update the aperture size based on the user's preferences. For example, if the user prefers to select a large number of items at the same time, the wearable system may increase the size of the aperture.
[0148] As another example of dynamically updating the aperture based on context information, if the user is in a dark environment or is nearsighted, the wearable system may increase the size of the aperture so that the user can more easily capture objects. In one implementation, a first cone projection can capture multiple objects. The wearable system can perform a second cone projection to further select a target object among the captured objects. The wearable system can also allow the user to select a target object from the captured objects using a body posture or a user input device. The object selection process can be a recursive process, and one, two, three, or more cone projections may be performed to select a target object.
[0149] FIG. 5 illustrates an example of various feature points 500 associated with a user's hand that may be detected or tracked by a wearable system according to some embodiments of the present disclosure. For each feature point, the capital letter corresponds to a region of the hand as follows: "T" corresponds to the thumb, "I" corresponds to the index finger, "M" corresponds to the middle finger, "R" corresponds to the ring finger, "P" corresponds to the little finger, "H" corresponds to the hand, and "F" corresponds to the forearm. The lowercase letters correspond to a more specific location within each region of the hand as follows: "t" corresponds to the tip (e.g., fingertip), "i" corresponds to the interphalangeal joint ("IP joint"), "d" corresponds to the distal interphalangeal joint ("DIP joint"), "p" corresponds to the proximal interphalangeal joint ("PIP joint"), "m" corresponds to the metacarpophalangeal joint ("MCP joint"), and "c" corresponds to the carpometacarpal joint ("CMC joint").
[0150] 6A-6F illustrate examples of possible subsets of feature points 500 that may be selected based on gestures identified by a wearable system, according to some embodiments of the present disclosure. In each example, feature points included in the selected subset are outlined in bold, feature points not included in the selected subset are outlined in dashed lines, and optional feature points that may be selected to facilitate subsequent determinations are outlined in solid lines. In each example, in response to selecting a subset of feature points, each feature point in the subset may be used to determine an interaction point, an orientation of a virtual multi-DOF controller (e.g., a virtual cursor or pointer associated with the interaction point), or both.
[0151] 6A illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is making or transitioning to making a grasping gesture (e.g., the user's fingers are all curled inward). In the illustrated example, feature point I m , T m , M m , and H may be used to determine a particular location that is included in the subset and to which the interaction point 602A is aligned. For example, the interaction point 602A is aligned with the feature point I m In some examples, the subset of feature points may also be used to determine, at least in part, the orientation of the virtual multi-DOF controller associated with the interaction point 602A. In some implementations, the subset of feature points associated with the grasp gesture may be aligned with feature point I. m , T m , M m , and H. In some embodiments, the particular location and / or orientation of the virtual multi-DOF controller to which the interaction point 602A will be aligned may be determined without regard to some or all feature points excluded from the subset of feature points associated with the grasp gesture.
[0152] 6B illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is making or transitioning to making a pointing gesture (e.g., the user's index finger is extended fully outward while the other fingers of the user's hand are curled inward). In the illustrated example, feature point I t , I d , I p , I m , T t , T i , T m , M m , and H may be used to determine a particular location that is included in the subset and to which interaction point 602B is aligned. For example, interaction point 602B is aligned with feature point I t In some examples, the subset of feature points may also be used to determine, at least in part, the orientation of a virtual multi-DOF controller associated with interaction point 602B. In some implementations, the subset of feature points associated with a point gesture may be aligned with feature point I. t , I d , I p , I m , T t , T i , T m , M m , and H. In some embodiments, the feature points I, I, and H may be included. d , M m , and one or more of H may be excluded from the subset of feature points associated with the point gesture. In some embodiments, the particular location and / or orientation of the virtual multi-DOF controller to which interaction point 602B will be aligned may be determined without regard to some or all of the feature points excluded from the subset of feature points associated with the point gesture.
[0153] 6C illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is making or transitioning to making a pinch gesture (e.g., the user's thumb and index finger are at least partially extended outward and brought into close proximity to one another). In the illustrated example, feature point I t , I d , I p , I m , T t , T i , T m , M m , and H may be used to determine a particular location that is included in the subset and to which interaction point 602C is aligned. For example, interaction point 602C may be [ka] Locations along the [ka] Alternatively, the interaction point may be aligned with the midpoint ("α") of [ka] Locations along the [ka] the midpoint of ("β"), or [ka] Locations along the [ka] Alternatively, the interaction point may be aligned with the midpoint ("γ") of [ka] Locations along the [ka] the midpoint of, or [ka] Locations along the [ka] In some examples, the subset of feature points may also be used to determine, at least in part, the orientation of the virtual multi-DOF controller associated with interaction point 602C. In some implementations, the subset of feature points associated with the pinch gesture may be aligned with the midpoint of feature point I. t , I d , I p , I m , T t , T i , T m , M m , and H. In some embodiments, the feature points I, I, and H may be included. d , M m , and one or more of H may be excluded from the subset of feature points associated with the pinch gesture. In some embodiments, the particular location and / or orientation of the virtual multi-DOF controller to which interaction point 602C will be aligned may be determined without regard to some or all of the feature points excluded from the subset of feature points associated with the pinch gesture.
[0154] 6D illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is transitioning between performing a grasping gesture and performing a pointing gesture (e.g., the user's index finger is partially extended outward while the other fingers of the user's hand are curled inward). In the illustrated example, feature point I t , I d , I p , I m , T t , T i , T m , M m , and H may be used to determine a particular location that is included in the subset and to which interaction point 602D is aligned. For example, interaction point 602D may be [ka] Additionally or alternatively, the interaction points may be aligned to locations along [ka] In some embodiments, the interaction point 602D is aligned with the user's hand, the location of which may change as the user transitions between the grasp and point gestures. [ka] along (or [ka] , and the visual representation (e.g., a light ray) of interaction point 602D displayed for the user may reflect that. That is, in these embodiments, the location to which interaction point 602D is aligned with respect to the user's hand is aligned with feature point I as the user transitions between a grasp gesture and a point gesture. m and I t Rather than making abrupt transitions between, it may glide along one or more paths between such feature points to provide a smoother and more intuitive user experience.
[0155] In some examples, when a user transitions between a grasp gesture and a point gesture, a visual representation of the interaction point 602D is displayed for the user's hand, the location of which may intentionally track that of the actual interaction point 602D according to the current positions of the subset of feature points at a given time. For example, when a user transitions between a grasp gesture and a point gesture, a visual representation of the interaction point 602D in the nth frame is displayed for the user, the location of which may correspond to the location of the actual interaction point 602D according to the positions of the subset of feature points in the (nm)th frame, where m is a predetermined number of frames (e.g., a fixed time delay). In another example, when a user transitions between a grasp gesture and a point gesture, the visual representation of the interaction point 602D displayed for the user may be configured to move at a fraction (e.g., a predetermined percentage) of the velocity of the actual interaction point 602D according to the current positions of the subset of feature points at a given time. In some embodiments, one or more filters or filtering techniques may be employed to achieve one or more of these behaviors. In some implementations, when a user does not transition between gestures or otherwise maintains a particular gesture, there may be little or no variation in where the visual representation of interaction point 602D is displayed relative to the user's hand and in the location of the actual interaction point 602D according to the current positions of the subset of feature points at any given time. Other configurations are possible.
[0156] 6E illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is transitioning between performing a point gesture and performing a pinch gesture (e.g., the user's thumb and index finger are at least partially extended outward and at least partially curled toward each other). In the illustrated example, feature point I t , I d , I p , I m , T t , T i , T m , M m , and H may be used to determine a particular location that is included in the subset and to which the interaction point 602E is aligned. For example, the interaction point 602E may be [ka] In some embodiments, when a user transitions between a point gesture and a pinch gesture, a visual representation of the interaction point 602E may be displayed for the user (e.g., a ray of light), and / or the actual interaction point 602E according to the current position of the subset of feature points at a given time may behave similarly or comparable to that described above with reference to FIG. 6D , which may serve to enhance the user experience.
[0157] 6F illustrates an example of a subset of feature points that may be selected when it is determined that the user's hand is transitioning between performing a pinch gesture and performing a grasp gesture (e.g., the user's thumb and index finger are at least partially extended outward and at least partially curled toward each other). In the illustrated example, feature point I t , I d , I p , I m , T t , T i , Tm , M m , and H may be used to determine a particular location that is included in the subset and to which interaction point 602F is aligned. For example, interaction point 602F may be [ka] In some embodiments, when a user transitions between a pinch gesture and a grasp gesture, a visual representation of the interaction point 602F may be displayed for the user (e.g., a ray of light), and / or the actual interaction point 602F according to the current positions of the subset of feature points at a given time may behave similarly or comparable to that described above with reference to FIGS. 6D-6E, which may serve to enhance the user experience.
[0158] 7A-7C illustrate examples of ray casting for various gestures while a user's arm is extended outward, according to some embodiments of the present disclosure. FIG. 7A illustrates a user performing a grasping gesture while their arm is extended outward. Interaction point 702A is positioned at feature point I. m (as described with reference to FIG. 6A ), and a proximal point 704A is aligned with a location on the user's shoulder (labeled “S”). A light ray 706A may be projected from the proximal point 704A through the interaction point 702A.
[0159] 7B illustrates a user making a pointing gesture while their arm is extended outward. Interaction point 702B is positioned at feature point I t6B ), with proximal point 704B aligned with a location at the user's shoulder (labeled "S"). A light ray 706B may be projected from proximal point 704B through interaction point 702B. FIG. 7C illustrates a user making a pinch gesture while their arms are extended outward. Interaction point 702C is aligned with location α (as described with reference to FIG. 6C ), with proximal point 704C aligned with a location at the user's shoulder (labeled "S"). A light ray 706C may be projected from proximal point 704C through interaction point 702C. The range of locations to which the interaction point may be aligned as the user transitions between the gestures of Figures 7A and 7B, 7B and 7C, and 7A and 7C are described in further detail above with reference to Figures 6D, 6E, and 6F.
[0160] 8A-8C illustrate examples of ray casting for various gestures while the user's arms are retracted inward, according to some embodiments of the present disclosure. FIG. 8A illustrates a user performing a grasp gesture while their arms are retracted inward. Interaction point 802A is positioned at feature point I. m (as described with reference to FIG. 6A ), and a proximal point 804A is aligned with a location at the user's elbow (labeled "E"). A light ray 806A may be cast from the proximal point 804A through the interaction point 802A.
[0161] FIG. 8B illustrates a user making a pointing gesture while their arm is retracted inward. Interaction point 802B is at feature point I t6B ), with proximal point 804B aligned with a location at the user's elbow (labeled "E"). A light ray 806B may be cast from proximal point 804B through interaction point 802B. FIG. 8C illustrates a user making a pinch gesture while their arm is pulled inward. Interaction point 802C is aligned with location α (as described with reference to FIG. 6C ), with proximal point 804C aligned with a location at the user's elbow (labeled "E"). A light ray 806C may be cast from proximal point 804C through interaction point 802C. The range of locations to which the interaction point may be aligned as the user transitions between the gestures of Figures 8A and 8B, to the gestures of Figures 8B and 8C, to the gestures of Figures 8A and 8C, are also described in further detail above with reference to Figures 6D, 6E, and 6F, respectively.
[0162] It can be seen that the location where proximal points 704A-704C in FIGS. 7A-7C are aligned with respect to the user's body differs from the location where proximal points 804A-804C in FIGS. 8A-8C are aligned with respect to the user's body. Such differences in location may be the result of, among other things, differences between the position and / or orientation of one or more portions of the user's arms in FIGS. 7A-7C (e.g., the user's arms are extended outward) and the position and / or orientation of one or more portions of the user's arms in FIGS. 8A-8C (e.g., the user's arms are retracted inward). Thus, when transitioning between the position and / or orientation of one or more portions of the user's arms in FIGS. 7A-7C and the position and / or orientation of one or more portions of the user's arms in FIGS. 8A-8C, the location where the proximal points are aligned may transition between a location at the user's shoulder (“S”) and a location at the user's elbow (“E”). In some embodiments, when the position and / or orientation of one or more portions of a user's arm transitions between that of Figures 7A-7C and that of Figures 8A-8C, the proximal point and one or more visual representations associated therewith may behave in a manner similar or comparable to that described above with reference to Figures 6D-6F, which may serve to enhance the user experience.
[0163] In some embodiments, the system may align the proximal point to one or more estimated locations within or along the surface of the user's knuckle, hand, wrist, forearm, elbow, arm (e.g., upper arm), shoulder, scapula, neck, head, eye, face (e.g., cheek), chest, torso (e.g., navel area), or combinations thereof. In at least some of these embodiments, the system may dynamically shift the location to which the proximal point is aligned between such one or more estimated locations based on at least one of a variety of different factors.For example, the system may detect (a) a gesture (e.g., grasp, point, pinch, etc.) that the user's hand is determined to be making or transitioning to making, (b) the position and / or orientation of a subset of feature points associated with a gesture that the user's hand is determined to be making or transitioning to making, (c) the location of the interaction point, (d) the estimated position and / or orientation (e.g., pitch, yaw, and / or roll) of the user's hand, (e) one or more measures of wrist flexion and / or extension, (f) one or more measures of wrist adduction and / or abduction, (g) the estimated position and / or orientation (e.g., pitch, yaw, and / or roll) of the user's forearm, (h) one or more measures of forearm supination and / or pronation, (i) one or more measures of elbow flexion and / or extension, (j) the position and / or orientation of the user's arm (e.g., upper arm), The location to which the proximal point will be aligned may be determined based on at least one of a variety of different factors, including: (k) an estimated position and / or orientation (e.g., pitch, yaw, and / or roll), (k) one or more measures of shoulder internal and / or external rotation, (l) one or more measures of shoulder flexion and / or extension, (m) one or more measures of shoulder adduction and / or abduction, (n) an estimated position and / or orientation of the user's head, (o) an estimated position and / or orientation of the wearable device, (p) an estimated distance between the user's hand or interaction point and the user's head or wearable device, (q) an estimated length or span of the user's entire arm (e.g., from shoulder to fingertips) or at least a portion thereof, (r) one or more measures of the user's visually coordinated attention, or (s) a combination thereof.
[0164] In some embodiments, the system may determine or otherwise evaluate one or more of the aforementioned factors based, at least in part, on data received from one or more outward-facing cameras, data received from one or more inward-facing cameras, data received from one or more other sensors of the system, data received as user input, or a combination thereof. In some embodiments, as one or more of the above-mentioned factors vary, the proximal point and one or more visual representations associated therewith may behave in a manner similar or comparable to that described above with reference to FIGS. 6D-8C , which may serve to enhance the user experience.
[0165] In some embodiments, the system may be configured to: (i) wrist adduction may serve to urge the location to which the proximal point is determined to be aligned along the user's arm toward the user's knuckles, while wrist abduction may serve to urge the location to which the proximal point is determined to be aligned along the user's arm toward the user's shoulder, neck, or other location closer to the center of the user's body; (ii) elbow flexion may serve to urge the location to which the proximal point is aligned downward toward the navel region of the user's body, while elbow extension may serve to urge the location to which the proximal point is aligned downward toward the user's head, shoulder, or other location on the upper portion of the user's body; and (iii) shoulder internal rotation may serve to urge the location to which the proximal point is determined to be aligned along the user's arm toward the user's knuckles, while wrist abduction may serve to urge the location to which the proximal point is aligned toward the user's shoulder, neck, or other location closer to the center of the user's body. (iv) shoulder adduction may serve to bias the location to which the proximal point is determined to be aligned toward the user's head, neck, chest, or other location closer to the center of the user's body, while shoulder abduction may serve to bias the location to which the proximal point is determined to be aligned toward the user's shoulder, arm, or other location further along the user's arm and away from the center of the user's body, or (v) combinations thereof. Thus, in these embodiments, the location to which the system determines the proximal point is aligned may dynamically change over time as the user repositions and / or reorients one or more of their hand, forearm, and arm. In some examples, the system may assign different weights to different factors and determine a location to which the proximal point will be aligned based on one or more such factors and their assigned weights.For example, the system may be configured to give more weight to one or more measures of a user's visually coordinated attention than to some or all of the other aforementioned factors. Other configurations are possible.
[0166] For examples in which the system is configured to dynamically shift the location to which the proximal point is aligned between such one or more estimated locations based, at least in part, on one or more measures of the user's visually coordinated attention, such one or more measures may be determined by the system based, at least in part, on the user's eye gaze, one or more characteristics of virtual content being presented to the user, hand position and / or orientation, one or more cross-modality convergence and / or divergence, or combinations thereof. Examples of cross-modality convergence and divergence, and systems and techniques for detecting and responding to the occurrence of such cross-modality convergence and divergence, are provided in U.S. Patent Publication No. 2019 / 0362557, which is incorporated herein by reference in its entirety. In some embodiments, the system may utilize one or more of the systems and / or techniques described in the aforementioned patent applications to detect one or more occurrences of cross-modality convergence and / or divergence, and may further determine the location of the proximal point based, at least in part, on the detected occurrences of one or more cross-modality convergence and / or divergence. Other configurations are also possible.
[0167] 9 illustrates an example of how an action event (e.g., hover, touch, tap, hold, etc.) may be detected using feature points, according to some embodiments of the present disclosure. In some embodiments, an action event may be detected, at least in part, by: [ka] The angle θ measured between [ka] and γ may be detected based on [ka] represents the midpoint of θ. For example, if θ is determined to be above a predetermined threshold, a "hover" action event may be detected, while if θ is determined to be below the predetermined threshold, a "touch" action event may be detected. As another example, "tap" and "hold" action events may be detected based on the duration of time over which θ is determined to be below the predetermined threshold. In the illustrated example, I t and T t may represent feature points included within a subset of feature points selected in response to determining that a user is performing or transitioning to performing a particular gesture (e.g., a pinch gesture).
[0168] 10A-10C illustrate example interactions with a virtual object using light rays, according to some embodiments of the present disclosure. Figures 10A-10C demonstrate how some of the paradigms described above can be employed within a wearable system and leveraged by a user for totem-less interactions (e.g., interactions without using physical handheld controllers). Figures 10A-10C each include renderings of what a user of a wearable system might see at various times while interacting with a virtual object 1002 using their hands. In this example, a user can manipulate the position of a virtual object by (1) making a pinch gesture with their hand to create a virtual 6DoF ray 1004, (2) positioning their hand so that the virtual 6DoF ray intersects with the virtual object, (3) bringing the tip of their thumb and the tip of their index finger closer together while maintaining the position of their hand so that the value of angle θ transitions from above a threshold to below a threshold while the virtual 6DoF ray intersects with the virtual object, and (4) guiding their hand to a new location while keeping their thumb and index finger pinched together so that angle θ remains below the threshold.
[0169] 10A illustrates an interaction point 1006 that is aligned with the α location while the user's hand is determined to be making a pinch gesture. The α location is a feature point (e.g., I) associated with the pinch gesture that is selected in response to determining that the user is making or transitioning to making a pinch gesture. t , I p , I m , T t , T i , and T m ) This selected subset of feature points may be tracked and utilized to determine a location (e.g., α-location) relative to which to align the interaction point 1006, and may further be utilized to determine an angle θ, similar or comparable to that described above with reference to FIG.
[0170] In the illustrated example of FIG. 10A , a ray of light 1004 is projected from a location near the user's right shoulder or upper arm through the interaction point. A graphical representation of the portion of the ray of light forward from the interaction point is displayed through the headset and utilized by the user as a sort of pointer or cursor with which to interact with the virtual object 1002. In FIG. 10A , the user has positioned their hand so that the virtual 6DoF ray of light intersects with the virtual object. Here, the angle θ presumably exceeds a threshold such that the user is considered to be simply “hovering” over the virtual object with the virtual 6DoF ray of light. Thus, the system may compare the angle θ to one or more thresholds and, based on the comparison, determine whether the user is considered to be touching, grasping, or otherwise selecting the virtual content. In the illustrated example, the system may determine that the angle θ exceeds one or more thresholds and, therefore, determine that the user is not considered to be touching, grasping, or otherwise selecting the virtual content.
[0171] 10B illustrates the user's hand with the virtual 6DoF ray intersecting the virtual object and still positioned as if performing a pinch gesture (note that the interaction point is still aligned at the α location). However, in FIG. 10B, the user has brought the tip of their thumb and the tip of their index finger closer together. Thus, in FIG. 10B, the angle θ is likely below one or more thresholds such that the user is now considered to be touching, grasping, or otherwise selecting virtual content with the virtual 6DoF ray.
[0172] FIG. 10C illustrates the user still making the same pinch gesture as in the previous image; therefore, the angle θ is likely below the threshold. However, in FIG. 10C , the user is moving their arm while keeping their thumb and index finger pinched together, effectively dragging the virtual object to a new location. Note that the interaction point advances with the user's hand by aligning with the α location. Although not shown in FIGS. 10A-10C , instead of or in addition to adjusting the position of the virtual object by adjusting the position of the interaction point relative to the headset while “holding” the virtual object, the user may also be able to adjust the orientation of the virtual object (e.g., the yaw, pitch, and / or roll of at least one vector and / or at least one plane defined by at least two and / or at least three feature points, respectively, included in the subset of selected feature points) by adjusting the orientation of the system of feature points associated with the pinch gesture relative to the headset while “holding” the virtual object. 10A-10C, after manipulating the position and / or orientation of the virtual object, the user may "release" the virtual object by separating their thumb and index finger. In such an embodiment, the system may determine that the angle θ again exceeds one or more thresholds and therefore determine that the user again is not considered to be touching, grasping, or otherwise selecting the virtual content.
[0173] 11 illustrates an example scheme for managing a point gesture, according to some embodiments of the present disclosure. The interaction point 1102 is preferably the index finger tip feature point (e.g., I t When the index finger tip is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1102 is aligned with the next nearest index finger PIP feature point (e.g., Ip When the index finger PIP is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1102 is moved to the index finger MCP feature point (e.g., I m In some embodiments, a filter is applied to smooth the transitions between different possible feature points.
[0174] 12 illustrates an exemplary scheme for managing a pinch gesture according to some embodiments of the present disclosure. The interaction point 1202 is preferably aligned at the midpoint between the index finger tip feature point and the thumb tip feature point (e.g., the α location described above with reference to FIG. 6C ). If the index finger tip feature point is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1202 is moved to the midpoint between the index finger PIP feature point and the thumb tip feature point. If the thumb tip feature point is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1202 is moved to the midpoint between the index finger tip feature point and the thumb IP feature point.
[0175] If both the index finger tip feature point and the thumb tip feature point are unavailable, the interaction point 1202 is moved to the midpoint between the index finger PIP feature point and the thumb IP feature point (e.g., the β location described above with reference to FIG. 6C ). If the index finger PIP feature point is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1202 is moved to the midpoint between the index finger MCP feature point and the thumb IP feature point. If the thumb IP feature point is unavailable (e.g., occluded or below a critical confidence level), the interaction point 1202 is moved to the midpoint between the index finger PIP feature point and the thumb MCP feature point. If both the index finger PIP feature point and the thumb IP feature point are unavailable, the interaction point 1202 is moved to the midpoint between the index finger MCP feature point and the thumb MCP feature point (e.g., the γ location described above with reference to FIG. 6C ).
[0176] 13 illustrates an exemplary scheme for detecting action events while a user's hand is performing a grasping gesture, according to some embodiments of the present disclosure. Relative angular distance and relative angular velocity may be tracked based on the angle between the index finger vector and the thumb vector. If the index finger tip feature point is unavailable, the index finger PIP feature point may be used to form the angle. If the thumb tip feature point is unavailable, the thumb IP feature point may be used to form the angle. Additional description regarding a subset of feature points that may be selectively tracked while a user is determined to be performing the grasping gesture of FIG. 13 is provided with reference to FIG. 6A above.
[0177] A first relative maximum angular distance (with its timestamp) may be detected at 1302. A relative minimum angular distance (with its timestamp) may be detected at 1304. A second relative maximum angular distance (with its timestamp) may be detected at 1306. Based on the difference in angular distance and the difference in time between the data detected at 1302, 1304, and 1306, it may be determined that an action event is being performed.
[0178] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more first thresholds (e.g., upper and lower thresholds), the difference between the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more second thresholds (e.g., upper and lower thresholds), the difference between the timestamps of the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more third thresholds (e.g., upper and lower thresholds), and the difference between the timestamps of the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more fourth thresholds (e.g., upper and lower thresholds).
[0179] FIG. 14 illustrates an example scheme for detecting an action event while a user's hand is performing a pointing gesture, according to some embodiments of the present disclosure. Relative angular distances may be tracked based on the angle between the index finger vector and the thumb vector. At 1402, a first relative maximum angular distance (with its timestamp) may be detected. At 1404, a relative minimum angular distance (with its timestamp) may be detected. At 1406, a second relative maximum angular distance (with its timestamp) may be detected. Based on the difference in angular distance and the difference in time between the data detected at 1402, 1404, and 1406, it may be determined that an action event is being performed. In some examples, such angular distances may be at least similar to the angle θ described above with reference to FIGS. 9 and 10A-10C. Additional description regarding a subset of feature points that may be selectively tracked while a user is determined to be performing the pointing gesture of FIG. 14 is provided with reference to FIG. 6B above.
[0180] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more first thresholds (e.g., upper and lower thresholds), the difference between the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more second thresholds (e.g., upper and lower thresholds), the difference between the timestamps of the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more third thresholds (e.g., upper and lower thresholds), and the difference between the timestamps of the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more fourth thresholds (e.g., upper and lower thresholds).
[0181] FIG. 15 illustrates an example scheme for detecting an action event while a user's hand is performing a pinch gesture, according to some embodiments of the present disclosure. Relative angular distances may be tracked based on the angle between the index finger vector and the thumb vector. At 1502, a first relative maximum angular distance (with its timestamp) may be detected. At 1504, a relative minimum angular distance (with its timestamp) may be detected. At 1506, a second relative maximum angular distance (with its timestamp) may be detected. Based on the difference in angular distance and the difference in time between the data detected at 1502, 1504, and 1506, it may be determined that an action event is being performed. In some examples, such angular distances may be at least similar to the angle θ described above with reference to FIGS. 9 and 10A-10C. Additional description regarding a subset of feature points that may be selectively tracked while a user is determined to be performing the pinch gesture of FIG. 15 is provided with reference to FIG. 6C above.
[0182] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more first thresholds (e.g., upper and lower thresholds), the difference between the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more second thresholds (e.g., upper and lower thresholds), the difference between the timestamps of the first relative maximum angular distance and the relative minimum angular distance may be compared to one or more third thresholds (e.g., upper and lower thresholds), and the difference between the timestamps of the relative minimum angular distance and the second relative maximum angular distance may be compared to one or more fourth thresholds (e.g., upper and lower thresholds).
[0183] FIG. 16 illustrates exemplary experimental data for detecting action events while a user's hand is making a pinch gesture, according to some embodiments of the present disclosure. The experimental data illustrated in FIG. 16 may correspond to the depicted movement of the user's hand in FIG. 15. In FIG. 16, the user's hand movement is characterized by the smoothed distance between the thumb and index finger. Noise is removed during low-latency smoothing so that the remaining signal shows an inflection of the normalized relative separation between paired finger features. Inflection, such as that seen by a local minimum followed by a local maximum, followed immediately by another local minimum, can be used to recognize a tap action. In addition, the same inflection pattern can also be seen in feature pose states. Feature pose A followed by feature pose B, followed by A, can also be used to recognize a tap action. Feature pose inflection can be robust when hand feature points have low confidence. When feature poses have low confidence, relative distance inflection can be used. If the confidence is high for both feature changes, then both inflections can be used to recognize the tap action.
[0184] 17A-17D illustrate exemplary experimental data for detecting action events while a user's hand is performing a pinch gesture according to some embodiments of the present disclosure. The experimental data illustrated in FIGS. 17A-17D may correspond to a user's hand repeatedly performing the movements illustrated in FIG. 15. FIG. 17A illustrates the distance between the tip of a user's index finger and the target content as the user's hand repeatedly approaches the target content. FIG. 17B illustrates the angular distance between the tip of the user's index finger and the tip of the user's thumb. FIG. 17C illustrates the angular velocity corresponding to the angle formed using the tip of the user's index finger and the tip of the user's thumb. FIG. 17D illustrates a feature pose change determined based on various data, which may optionally include the data illustrated in FIGS. 17A-17C. The experimental data illustrated in FIGS. 17A-17D may be used to identify a tap action. In some embodiments, all feature inflections can be utilized in parallel or simultaneously to reduce the false positive recognition rate.
[0185] 18 illustrates an example scheme for detecting an action event while a user's hand is performing a pinch gesture, according to some embodiments of the present disclosure. FIG. 18 differs from FIG. 15 in that the user's middle, ring, and pinky fingers are curled inward.
[0186] 19A-19D illustrate exemplary noise experiment data for detecting action events while a user's hand is making a pinch gesture, according to some embodiments of the present disclosure. The experiment data illustrated in FIGS. 19A-19D may correspond to a user's hand repeatedly performing the movements shown in FIG. 18. FIG. 19A illustrates the distance between the tip of a user's index finger and target content. FIG. 19B illustrates the angular distance between the tip of a user's index finger and the tip of a user's thumb. FIG. 19C illustrates the angular velocity corresponding to the angle formed using the tip of a user's index finger and the tip of a user's thumb. FIG. 19D illustrates a characteristic pose change determined based on various data, which may optionally include the data shown in FIGS. 19A-19C. The noise experiment data illustrated in FIGS. 19A-19D may be used to identify a tap action, which is determined to occur within window 1902. This represents an edge case scenario that utilizes at least a medium-confidence determination in all of the inflections to qualify as a recognized tap action.
[0187] 20A-20C illustrate an example scheme for managing a grasp gesture according to some embodiments of the present disclosure. A ray 2006 is cast from a proximal point 2004 (aligned with a location on the user's shoulder) through an interaction point 2002 (aligned with a location on the user's hand), as described herein. FIG. 20A illustrates a grasp gesture that enables a collective pointing mechanical action. This can be used for robust far-field targeting. FIG. 20B illustrates the size of the interaction point relative to the calculated hand radius as characterized by the relative distance between fingertip features. FIG. 20C illustrates that as the hand changes from an open feature pose to a fist feature pose, the hand radius decreases, and therefore the size of the interaction point decreases proportionally.
[0188] 21A-21C illustrate an exemplary scheme for managing a point gesture according to some embodiments of the present disclosure. A ray of light 2106 is projected from a proximal point 2104 (aligned with a location on the user's shoulder) through an interaction point 2102 (aligned with a location on the user's hand), as described herein. FIG. 21A illustrates a mechanical action of pointing and selecting that leverages finger joint movement for refined mid-field targeting. FIG. 21B illustrates a relaxed (open) point hand posture. The interaction point is placed at the index fingertip. The relative distance between the thumb tip and index finger tip is at a maximum value, proportionally increasing the size of the interaction point. FIG. 21C illustrates a (closed) point hand posture in which the thumb is curled under the index finger. The relative distance between the thumb tip and index finger tip is at a minimum value, resulting in a proportionally smaller interaction point size, but still placed at the index fingertip.
[0189] 22A-22C illustrate an exemplary scheme for managing a pinch gesture according to some embodiments of the present disclosure. A ray of light 2206 is projected from a proximal point 2204 (aligned with a location on the user's shoulder) through an interaction point 2202 (aligned with a location on the user's hand), as described herein. FIG. 22A illustrates a point-and-select mechanical action leveraging finger joint movement to target a refined mid-field of view. FIG. 22B illustrates an open pinch (OK) posture. The interaction point is located at the midpoint between the index fingertip and thumb as one of several pinch styles enabled by the managed pinch posture. The relative distance between the thumb tip and index fingertip is at a maximum value, proportionally increasing the size of the interaction point. FIG. 22C illustrates a (closed) pinch hand posture in which the middle, ring, and pinky fingers are curled inward and the index finger and thumb fingertips are touching. The relative distance between the thumb tip and index finger tip is at a minimum, resulting in a proportionally small interaction point size, but still located at the midpoint between the fingertips.
[0190] 23 illustrates various activation types for point and pinch gestures according to some embodiments of the present disclosure. For point gestures, activation types include touch (close), hover (open), tap, and hold. For pinch gestures, activation types include touch (close), hover (open), tap, and hold.
[0191] FIG. 24 illustrates various gestures and transitions between gestures according to some embodiments of the present disclosure. In the illustrated example, the set of gestures includes a grasp gesture, a point gesture, and a pinch gesture, along with their respective transition states. Each gesture also includes subgestures (or subposes) whose resulting gestures may be further defined by the wearable system. A grasp gesture may include a fist subpose, a control subpose, and a stylus subpose, among other possibilities. A point gesture may include a single finger subpose and an "L" shape subpose, among other possibilities. A pinch gesture may include an open subpose, a closed subpose, and an "OK" subpose, among other possibilities.
[0192] 25 illustrates examples of two-handed interaction in which both of a user's hands are used to interact with a virtual object, according to some embodiments of the present disclosure. In each of the illustrated examples, the user's hands are each determined to be making a pointing gesture based on each individual hand feature point. Interaction points 2510 and 2512 for both of the user's hands are determined based on the individual hand feature points and the determined gesture. Interaction points 2510 and 2512 are used to determine two-handed interaction point 2514, which may facilitate selecting and targeting the virtual object for two-handed interaction. Two-handed interaction point 2514 may be aligned to a location along a line (e.g., a midpoint) formed between interaction points 2510 and 2512.
[0193] In each of the illustrated embodiments, delta 2516 is generated based on one or both of the movement of interaction points 2510 and 2512. In 2502, delta 2516 is a translation delta corresponding to the inter-frame translation of one or both of interaction points 2510 and 2512. In 2504, delta 2516 is a scaling delta corresponding to the inter-frame separation movement of one or both of interaction points 2510 and 2512. In 2506, delta 2516 is a rotation delta corresponding to the inter-frame rotation movement of one or both of interaction points 2510 and 2512.
[0194] 26 illustrates an example of two-handed interaction that differs from FIG. 26 in that each of the user's hands is determined to be making a pinch gesture based on each individual hand feature point. Interaction points 2610 and 2612 for both of the user's hands are determined based on the individual hand feature points and the determined gesture. Interaction points 2610 and 2612 are used to determine two-handed interaction point 2614, which may facilitate selecting and targeting a virtual object for two-handed interaction. Two-handed interaction point 2614 may be aligned to a location along a line (e.g., a midpoint) formed between interaction points 2610 and 2612.
[0195] In each of the illustrated embodiments, delta 2616 is generated based on one or both of the movement of interaction points 2610 and 2612. In 2602, delta 2616 is a translation delta, which corresponds to the inter-frame translation of one or both of interaction points 2610 and 2612. In 2604, delta 2616 is a scaling delta, which corresponds to the inter-frame separation movement of one or both of interaction points 2610 and 2612. In 2606, delta 2616 is a rotation delta, which corresponds to the inter-frame rotation movement of one or both of interaction points 2610 and 2612.
[0196] 27 illustrates various examples of collaborative two-handed interactions in which both hands work together to interact with a virtual object, according to some embodiments of the present disclosure. The examples shown include a pinch gesture, a point gesture, a flat gesture, a hook gesture, a fist gesture, and a trigger gesture.
[0197] 28 illustrates examples of managed two-hand interactions in which one hand manages how the other hand may be interpreted, according to some embodiments of the present disclosure. The examples shown include index-thumb-pinch+index-point, middle-thumb-pinch+index-point, index-middle-point+index-point, and index-trigger+index-point.
[0198] 29 illustrates exemplary single-handed and two-handed interaction fields 2902 and 2904, according to some embodiments of the present disclosure. The interaction fields 2902 and 2904 each include a peripheral space, an extended workspace, a workspace, and a task space. The camera of the wearable system may be oriented to capture one or both of the user's hands while operating within the various spaces based on whether the system supports single-handed or two-handed interaction.
[0199] FIG. 30 illustrates a method 3000 of forming a multi-DOF controller associated with a user's hand to enable the user to interact with a virtual object, according to some embodiments of the present disclosure. One or more steps of method 3000 may be omitted during implementation of method 3000, and the steps of method 3000 need not be performed in the order presented. One or more steps of method 3000 may be performed by one or more processors of a wearable system, such as those included within processing module 250 of wearable system 200. Method 3000 may be implemented as a computer-readable medium or computer program product comprising instructions, which, when executed by one or more computers, cause the one or more computers to perform the steps of method 3000. Such a computer program product can be transmitted via a wired or wireless network within a data carrier signal carrying the computer program product.
[0200] In step 3002, an image of a user's hand is received. The image may be captured by an image capture device, which may be mounted on a wearable device. The image capture device may be a camera (e.g., a wide-angle lens camera, a fisheye lens camera, an infrared (IR) camera), or a depth sensor, among other possibilities.
[0201] In step 3004, the image is analyzed to detect feature points associated with the user's hand, which may be on or near the user's hand (within a threshold distance of the user's hand).
[0202] In step 3006, based on analyzing the image, it is determined whether the user's hands are performing or transitioning to perform any gesture from a plurality of gestures. The plurality of gestures may include, among other possibilities, a grasp gesture, a point gesture, and / or a pinch gesture. If it is determined that the user's hands are performing or transitioning to perform any gesture, method 3000 proceeds to step 3008. Otherwise, method 3000 returns to step 3002.
[0203] In step 3008, a specific location for the plurality of feature points is determined. The specific location may be determined based on the plurality of feature points and the gesture. As an example, if the user's hand is determined to be performing a first gesture of the plurality of gestures, the specific location may be set to the location of a first feature point of the plurality of feature points, and if the user's hand is determined to be performing a second gesture of the plurality of gestures, the specific location may be set to the location of a second feature point of the plurality of feature points. Continuing with the above example, if the user's hand is determined to be performing a third gesture of the plurality of gestures, the specific location may be set to the midpoint between the first and second feature points. Alternatively, or in addition, if the user's hand is determined to be performing a third gesture, the specific location may be set to the midpoint between the third and fourth feature points.
[0204] In step 3010, the interaction point is aligned to a specific location. Aligning the interaction point to a specific location may include setting and / or moving the interaction point to the specific location. The interaction point (and similarly, the specific location) may be a 3D value.
[0205] In step 3012, a multi-DOF controller for interacting with the virtual object is formed based on the interaction point. The multi-DOF controller may correspond to a light ray projected from a proximal point through the interaction point. The light ray may be used to perform various actions such as targeting, selecting, grasping, scrolling, extracting, hovering, touching, tapping, and holding.
[0206] FIG. 31 illustrates a method 3100 of forming a multi-DOF controller associated with a user's hand to enable the user to interact with a virtual object, according to some embodiments of the present disclosure. One or more steps of method 3100 may be omitted during implementation of method 3100, and the steps of method 3100 need not be performed in the order shown. One or more steps of method 3100 may be performed by one or more processors of a wearable system, such as those included within processing module 250 of wearable system 200. Method 3100 may be implemented as a computer-readable medium or computer program product comprising instructions, which, when executed by one or more computers, cause the one or more computers to perform the steps of method 3000. Such a computer program product can be transmitted via a wired or wireless network within a data carrier signal carrying the computer program product.
[0207] In step 3102, an image of a user's hand is received. Step 3102 may be similar to step 3002 described with reference to FIG.
[0208] In step 3104, the image is analyzed to detect a number of feature points associated with the user's hand. Step 3104 may be similar to step 3004 described with reference to FIG.
[0209] In step 3106, based on analyzing the image, it is determined whether the user's hands are performing or transitioning to perform any gesture from the plurality of gestures. Step 3106 may be similar to step 3006 described with reference to FIG. 30. If it is determined that the user's hands are performing or transitioning to perform any gesture, method 3100 proceeds to step 3108. Otherwise, method 3100 returns to step 3102.
[0210] In step 3108, a subset of the plurality of feature points corresponding to a particular gesture is selected. For example, a first subset of feature points may correspond to a first gesture of the plurality of gestures, and a second subset of feature points may correspond to a second gesture of the plurality of gestures. Continuing with the above example, if the user's hands are determined to be performing a first gesture, the first subset of feature points may be selected, or if the user's hands are determined to be performing a second gesture, the second subset of feature points may be selected.
[0211] In step 3110, a specific location for a subset of the plurality of feature points is determined. The specific location may be determined based on the subset of the plurality of feature points and the gesture. As an example, the specific location may be set to a location of a first feature point of the first subset of the plurality of feature points if the user's hand is determined to be performing a first gesture of the plurality of gestures. As another example, the specific location may be set to a location of a second feature point of the second subset of the plurality of feature points if the user's hand is determined to be performing a second gesture of the plurality of gestures.
[0212] In step 3112, the interaction point is aligned to a particular location. Step 3112 may be similar to step 3010 described with reference to FIG.
[0213] In step 3114, the proximal point is aligned to a location along the user's body. The location to which the proximal point is aligned may be the estimated location of the user's shoulder, the estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.
[0214] In step 3116, a ray is cast from the near point through the interaction point.
[0215] In step 3118, a multi-DOF controller for interacting with the virtual object is formed based on the light rays. The multi-DOF controller may correspond to light rays that are cast from a proximal point through the interaction point. The light rays may be used to perform various actions such as targeting, selecting, grasping, scrolling, extracting, hovering, touching, tapping, and holding.
[0216] In step 3120, a graphical representation of the multi-DOF controller is displayed by the wearable system.
[0217] 32 illustrates a method 3200 of interacting with a virtual object using two-handed input according to some embodiments of the present disclosure. One or more steps of method 3200 may be omitted during implementation of method 3200, and the steps of method 3200 need not be performed in the order presented. One or more steps of method 3200 may be performed by one or more processors of a wearable system, such as those included within processing module 250 of wearable system 200. Method 3200 may be implemented as a computer-readable medium or computer program product comprising instructions, which, when executed by one or more computers, cause the one or more computers to perform the steps of method 3200. Such a computer program product can be transmitted via wired or wireless networks within a data carrier signal carrying the computer program product.
[0218] In step 3202, one or more images of a user's first hand and second hand are received. Some of the one or more images may include both the first hand and the second hand, and some may include only one of the hands. The one or more images may include a series of time-sequential images. The one or more images may be captured by an image capture device, which may be mounted on a wearable device. The image capture device may be a camera (e.g., a wide-angle lens camera, a fisheye lens camera, an infrared (IR) camera), or a depth sensor, among other possibilities.
[0219] In step 3204, one or more images are analyzed to detect feature points associated with each of the first hand and the second hand. For example, one or more images may be analyzed to detect two separate sets of feature points: a set of feature points associated with the first hand and a set of feature points associated with the second hand. Each of the feature points may be on or near (within a threshold distance of) a respective hand. In some embodiments, a different set of feature points may be detected for each time series of images or image frames.
[0220] In step 3206, an interaction point is determined for each of the first and second hands based on a plurality of feature points associated with the first and second hands. For example, an interaction point for the first hand may be determined based on a plurality of feature points associated with the first hand, and an interaction point for the second hand may be determined based on a plurality of feature points associated with the second hand. In some embodiments, it may be determined whether the first and second hands are performing (or transitioning to performing) a particular gesture from a plurality of gestures. Based on the particular gesture, an interaction point for each hand may be positioned at a particular location as described herein.
[0221] In step 3208, a two-hand interaction point is determined based on the interaction points for the first hand and the second hand. In some embodiments, the two-hand interaction point may be the average position of the interaction points. For example, a line may be formed between the interaction points, and the two-hand interaction point may be aligned to a point (e.g., the midpoint) along the line. The location to which the two-hand interaction point is aligned may also be determined based on the gesture each hand is making (or transitioning to making). For example, if one hand is making a pointing gesture and the other hand is making a grasping gesture or a pinch gesture, the two-hand interaction point may be aligned to that hand regardless of which hand is making the pointing gesture. As another example, if both hands are making the same gesture (e.g., a pinch gesture), the two-hand interaction point may be aligned to the midpoint between the interaction points.
[0222] In step 3210, one or more two-hand deltas may be generated for each of the first and second hands based on the interaction points. In some embodiments, the one or more two-hand deltas may be generated based on movement (e.g., frame-to-frame movement) of the interaction points. For example, the one or more two-hand deltas may include a translation delta, a rotation delta, and / or a scaling delta. The translation delta may correspond to a translational movement of one or both of the interaction points, the rotation delta may correspond to a rotational movement of one or both of the interaction points, and the scaling delta may correspond to a separation movement of one or both of the interaction points.
[0223] In one example, a set of time-series images may be analyzed to determine that the interaction points for the first hand and the second hand are moving closer together, and in response, a scaling delta may be generated with a negative value to indicate that the interaction points are moving closer together. In another example, a set of time-series images may be analyzed to determine that the interaction points are moving farther apart, and a scaling delta may be generated with a positive value to indicate that the interaction points are moving farther apart.
[0224] In another example, a set of time-series images may be analyzed to determine that the interaction points for the first hand and the second hand are both moving in the positive x-direction. In response, a translation delta may be generated to indicate that the interaction points are moving in the positive x-direction. In another example, a set of time-series images may be analyzed to determine that the interaction points for the first hand and the second hand are rotated relative to one another (e.g., a line formed between the interaction points is rotated). In response, a rotation delta may be generated to indicate that the interaction points are rotated relative to one another.
[0225] In some embodiments, two-handed deltas may be generated based on one of the interaction points and the established plane. For example, the plane may be established based on the user's hands, head pose, the user's hips, a real-world object, or a virtual object, among other possibilities. In response to establishing the plane, a translation delta may be generated based on a projection of the interaction point on the plane, a rotation delta may be generated based on a rotation of the interaction point relative to the plane, and a scaling delta may be generated based on the distance between the interaction point and the plane. In some examples, these deltas may be referred to as planar deltas.
[0226] The above-described examples of two-handed deltas may be generated with respect to a set of the same time-series images. For example, two-handed deltas including translation deltas, rotation deltas, and scaling deltas may be generated with respect to a set of single time-series images. In some examples, only specific types of two-handed deltas may be generated based on the requirements of a particular application. For example, a user may initiate a scaling operation while keeping the position and orientation of a virtual object fixed. In response, only scaling deltas may be generated, while translation and rotation deltas may not be generated. As another example, a user may initiate a translation operation and a rotation operation while keeping the size of a virtual object fixed. In response, only translation and rotation deltas may be generated, while scaling deltas may not be generated. Other possibilities are also contemplated.
[0227] In step 3212, the virtual object is interacted with using one or more two-handed deltas. The virtual object may be interacted with by applying one or more two-handed deltas to the virtual object, such as by moving the virtual object using the one or more two-handed deltas. For example, applying a translation delta to the virtual object may translate the virtual object by a particular amount indicated by the translation delta, applying a rotation delta to the virtual object may rotate the virtual object by a particular amount indicated by the rotation delta, and applying a scaling delta to the virtual object may scale / resize the virtual object by a particular amount indicated by the scaling delta.
[0228] In some embodiments, prior to interacting with a virtual object, it may be determined whether the virtual object is targeted. In some instances, it may be determined whether a two-handed interaction point overlaps with or is within a threshold distance of the virtual object. In some embodiments, it may be determined whether the virtual object is currently selected or was previously selected by using one-handed interaction, for example, as described herein. In one example, a virtual object may be initially selected using one-handed interaction and subsequently interacted with using two-handed interaction.
[0229] FIG. 33 illustrates a simplified computer system 3300 according to some embodiments of the present disclosure. A computer system 3300 such as that illustrated in FIG. 33 may be incorporated into a device as described herein. FIG. 33 provides a schematic illustration of one embodiment of a computer system 3300 that may perform some or all of the steps of the methods provided by various embodiments. Note that FIG. 33 is intended only to provide a generalized illustration of various components, any or all of which may be utilized as desired. FIG. 33 therefore broadly illustrates situations in which individual system elements may be implemented in a relatively separate manner or a relatively more integrated manner.
[0230] Computer system 3300 is shown to include hardware elements that may be electrically coupled via a bus 3305, or may otherwise communicate as needed. The hardware elements may include one or more processors 3310, including one or more general-purpose processors and / or one or more special-purpose processors, such as, but not limited to, digital signal processing chips, graphics acceleration processors, and / or the like, one or more input devices 3315, which may include, but are not limited to, a mouse, keyboard, camera, and / or the like, and one or more output devices 3320, which may include, but are not limited to, a display device, printer, and / or the like.
[0231] The computer system 3300 may further include and / or communicate with one or more non-transitory storage devices 3325, which may include, but are not limited to, local and / or network-accessible storage devices, and / or may include, but are not limited to, disk drives, drive arrays, optical storage devices, solid-state storage devices such as random access memory ("RAM"), and / or read-only memory ("ROM"), which may be programmable, flash-updateable, and / or the like. Such storage devices may be configured to implement any suitable data storage, including, but not limited to, various file systems, database structures, and / or the like.
[0232] The computer system 3300 may also include a communications subsystem 3319, which may include, but is not limited to, a modem, a network card (wireless or wired), an infrared communications device, a wireless communications device, and / or a chipset, such as, but not limited to, a Bluetooth® device, an 802.11 device, a WiFi device, a WiMax device, a cellular communications facility, etc., and / or the like. The communications subsystem 3319 may include one or more input and / or output communications interfaces, allowing data to be exchanged with networks, such as those described below by way of example, other computer systems, televisions, and / or any other devices described herein. Depending on desired functionality and / or other implementation concerns, a portable electronic device or similar device may communicate images and / or other information via the communications subsystem 3319. In other embodiments, a portable electronic device, e.g., a first electronic device, may be incorporated into the computer system 3300, e.g., an electronic device, as an input device 3315. In some embodiments, the computer system 3300 further includes a working memory 3335, which may include a RAM or ROM device as described above.
[0233] Computer system 3300 may also include software elements shown as currently residing in working memory 3335, including an operating system 3340, device drivers, executable libraries, and / or other code, such as one or more application programs 3345, which may include computer programs provided by various embodiments and / or may be designed to implement methods and / or configure systems provided by other embodiments as described herein. By way of example only, one or more procedures described with respect to the methods discussed above may be implemented as code and / or instructions executable by a computer or a processor within a computer, and in certain aspects such code and / or instructions can then be used to configure and / or adapt a general-purpose computer or other device to perform one or more operations in accordance with the described methods.
[0234] A set of these instructions and / or code may be stored on a non-transitory computer-readable storage medium, such as storage device 3325 described above. In some cases, the storage medium may be incorporated within a computer system, such as computer system 3300. In other embodiments, the storage medium is separate from the computer system, e.g., a removable medium such as a compact disc, and / or may be provided in an installation package such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer with the instructions / code stored thereon. These instructions may take the form of executable code that is executable by computer system 3300 and / or may take the form of source and / or installable code that, upon compilation and / or installation on computer system 3300 using, for example, any of various commonly available compilers, installation programs, compression / decompression utilities, etc., then takes the form of executable code.
[0235] It will be apparent to those skilled in the art that substantial variations may be made according to particular requirements. For example, customized hardware may also be used, and / or particular elements may be implemented in hardware, software, including portable software such as applets, or both. Furthermore, connection to other computing devices, such as network input / output devices, may also be employed.
[0236] As noted above, in one aspect, some embodiments may employ a computer system, such as computer system 3300, to perform methods according to various embodiments of the present technology. According to one set of embodiments, some or all of the procedures of such methods are performed by computer system 3300 in response to processor 3310 executing one or more sequences of one or more instructions, which may be embedded in operating system 3340, and / or other code, such as application program 3345, contained in working memory 3335. Such instructions may be read into working memory 3335 from another computer-readable medium, such as one or more of storage devices 3325. By way of example only, execution of a sequence of instructions contained in working memory 3335 may cause processor 3310 to perform one or more procedures of the methods described herein. Additionally or alternatively, some of the methods described herein may be performed through specialized hardware.
[0237] The terms “machine-readable medium” and “computer-readable medium,” as used herein, refer to any medium that participates in providing data that causes a machine to operate in a tangible fashion. In an embodiment implemented using computer system 3300, various computer-readable media may be involved in providing instructions / code to processor 3310 for execution and / or may be used to store and / or carry such instructions / code. In many implementations, computer-readable media are physical and / or tangible storage media. Such media may take the form of non-volatile or volatile media. Non-volatile media include, for example, optical and / or magnetic disks, such as storage device 3325. Volatile media include dynamic memory, such as, but not limited to, working memory 3335.
[0238] Common forms of physical and / or tangible computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape or any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, or any other medium from which a computer can read instructions and / or code.
[0239] Various forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to processor 3310 for execution. By way of example only, the instructions may initially be carried on a magnetic and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions as signals over a transmission medium to be received and / or executed by computer system 3300.
[0240] The communications subsystem 3319 and / or its components generally receive signals and the bus 3305 may then convey the signals and / or the data, instructions, etc. carried by the signals to the working memory 3335, from which the processor 3310 retrieves and executes the instructions. The instructions received by the working memory 3335 may optionally be stored on a non-transitory storage device 3325 either before or after execution by the processor 3310.
[0241] The methods, systems, and devices discussed above are examples. Various configurations may omit, substitute, or add various procedures or components, as appropriate. For example, in alternative configurations, the methods may be performed in a different order than described, and / or various steps may be added, omitted, and / or combined. Also, features described with respect to one configuration may be combined in various other configurations. Different aspects and elements of the configurations may be combined in a similar manner. Also, technology evolves, and therefore, many of the elements are examples and do not limit the scope of the disclosure or the claims.
[0242] Specific details are given in the description to provide a thorough understanding of example configurations, including implementations. However, the configurations may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary detail to avoid obscuring the configurations. This description provides only example configurations and does not limit the scope, applicability, or configuration of the claims. Rather, the foregoing description of the configurations will provide those skilled in the art with an effective description for implementing the described techniques. Various changes may be made in the function and arrangement of elements without departing from the spirit or scope of the present disclosure.
[0243] Configurations may also be described as processes, depicted as schematic flowcharts or block diagrams. While each operation may be described as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process may have additional steps not included in the diagram. Furthermore, embodiments of the method may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments to perform the necessary tasks may be stored in a non-transitory computer-readable medium, such as a storage medium. A processor may perform the described tasks.
[0244] While several example configurations have been described, various modifications, alternative constructions, and equivalents may be used without departing from the spirit of the present disclosure. For example, the elements described above may be components of a larger system, and other rules may take precedence over or otherwise modify the application of the present technology. Also, some steps may occur before, during, or after the elements described above are discussed. Therefore, the foregoing description does not constrain the scope of the claims.
[0245] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Thus, for example, reference to a "user" includes one or more such users, reference to a "processor" includes reference to one or more processors and equivalents thereof known to those skilled in the art, and so forth.
[0246] Additionally, the words "comprise," "comprising," "contains," "containing," "include," "including," and "includes," when used in this specification and the claims that follow, are intended to specify the presence of stated features, integers, components, or steps, but they do not exclude the presence or addition of one or more other features, integers, components, steps, acts, or groups.
[0247] It is also to be understood that the examples and embodiments described herein are for illustrative purposes only, and that various modifications or changes in light thereof will be suggested to those skilled in the art and are within the spirit and scope of the present application and the appended claims.
Claims
1. 1. A method for interacting with a virtual object, the method comprising: receiving images of a user's hands from one or more image capture devices of a wearable system; analyzing the image to detect a plurality of feature points associated with the user's hands; determining whether the user's hand is performing or transitioning to perform a particular gesture from a plurality of gestures based on analyzing the image; In response to determining that the user's hand is performing or transitioning to perform a particular gesture, selecting a subset of the plurality of feature points corresponding to the particular gesture; determining a specific location for the subset of the plurality of feature points, the specific location being determined based on the subset of the plurality of feature points and the specific gesture; and aligning an interaction point to said particular location; aligning a proximal point to a location along the user's body; projecting a light ray from the proximal point through the interaction point; forming a multi-DOF controller for interacting with the virtual object based on the light rays; A method comprising:
2. The method of claim 1 , wherein the plurality of gestures includes at least one of a grasp gesture, a point gesture, or a pinch gesture.
3. The method of claim 1 , wherein the subset of feature points is selected from a plurality of subsets of the plurality of feature points, each of the plurality of subsets of feature points corresponding to a different gesture from the plurality of gestures.
4. The method of claim 1 , further comprising displaying a graphical representation of the multi-DOF controller.
5. 2. The method of claim 1, wherein the location to which the proximal point is aligned is at an estimated location of the user's shoulder, an estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.
6. The method of claim 1 , further comprising capturing an image of the user's hand with an image capture device of the one or more image capture devices.
7. The method of claim 6 , wherein the image capture device is mounted in a headset of a wearable system.
8. 1. A system comprising: one or more processors; A machine-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to: receiving images of a user's hands from one or more image capture devices of a wearable system; analyzing the image to detect a plurality of feature points associated with the user's hands; determining whether the user's hand is performing or transitioning to perform a particular gesture from a plurality of gestures based on analyzing the image; In response to determining that the user's hand is performing or transitioning to perform a particular gesture, selecting a subset of the plurality of feature points corresponding to the particular gesture; determining a specific location for the subset of the plurality of feature points, the specific location being determined based on the subset of the plurality of feature points and the specific gesture; and aligning an interaction point to said particular location; aligning a proximal point to a location along the user's body; projecting a light ray from the proximal point through the interaction point; forming a multi-DOF controller for interacting with a virtual object based on the light rays; a machine-readable medium for performing operations including: A system comprising:
9. The system of claim 8 , wherein the plurality of gestures includes at least one of a grasp gesture, a point gesture, or a pinch gesture.
10. The system of claim 8 , wherein the subset of feature points is selected from a plurality of subsets of the plurality of feature points, each of the plurality of subsets of feature points corresponding to a different gesture from the plurality of gestures.
11. The system of claim 8 , wherein the operations further include displaying a graphical representation of the multi-DOF controller.
12. 9. The system of claim 8, wherein the location to which the proximal point is aligned is at an estimated location of the user's shoulder, an estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.
13. The system of claim 8 , wherein the actions further include capturing an image of the user's hand with an image capture device of the one or more image capture devices.
14. The system of claim 13 , wherein the image capture device is mounted in a headset of a wearable system.
15. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to: receiving images of a user's hands from one or more image capture devices of a wearable system; analyzing the image to detect a plurality of feature points associated with the user's hands; determining whether the user's hand is performing or transitioning to perform a particular gesture from a plurality of gestures based on analyzing the image; In response to determining that the user's hand is performing or transitioning to perform a particular gesture, selecting a subset of the plurality of feature points corresponding to the particular gesture; determining a specific location for the subset of the plurality of feature points, the specific location being determined based on the subset of the plurality of feature points and the specific gesture; and aligning an interaction point to said particular location; aligning a proximal point to a location along the user's body; projecting a light ray from the proximal point through the interaction point; forming a multi-DOF controller for interacting with a virtual object based on the light rays; A non-transitory machine-readable medium for performing operations including:
16. 16. The non-transitory machine-readable medium of claim 15, wherein the plurality of gestures comprises at least one of a grasp gesture, a point gesture, or a pinch gesture.
17. 16. The non-transitory machine-readable medium of claim 15, wherein the subset of feature points is selected from a plurality of subsets of the plurality of feature points, each of the plurality of subsets of feature points corresponding to a different gesture from the plurality of gestures.
18. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise displaying a graphical representation of the multi-DOF controller.
19. 16. The non-transitory machine-readable medium of claim 15, wherein the location to which the proximal point is aligned is an estimated location of the user's shoulder, an estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.
20. 16. The non-transitory machine-readable medium of claim 15, wherein the actions further include capturing an image of the user's hand with an image capture device of the one or more image capture devices.