Hand gesture input for wearable systems

CN115443445BActive Publication Date: 2026-09-01MAGIC LEAP INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180030873.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-19
Filing Date
2021-02-25
Publication Date
2026-09-01
Estimated Expiration
2041-02-25

Smart Images

  • Figure CN115443445B_ABST
    Figure CN115443445B_ABST
Patent Text Reader

Abstract

A technique for allowing a user's hand to interact with a virtual object is disclosed. An image of at least one hand can be received from an image capture device. Multiple keypoints associated with at least one hand can be detected. In response to determining that the hand is making or is transitioning to making a specific gesture, a subset of the multiple keypoints can be selected. Interaction points can be registered to specific locations relative to the subset of multiple keypoints based on a specific gesture. Neighboring points can be registered to locations along the user's body. Light rays from neighboring points can be projected through the interaction points. A multi-DOF controller for interacting with the virtual object can be formed based on the light rays.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference of related applications

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 981,934, filed February 26, 2020, entitled “HAND GESTURE INPUT FOR WEARABLE SYSTEM,” and U.S. Provisional Patent Application No. 63 / 027,272, filed May 19, 2020, entitled “HAND GESTURE INPUT FOR WEARABLE SYSTEM,” the contents of which are incorporated herein by reference in their entirety for all purposes. Background Technology

[0003] Modern computing and display technologies have facilitated the development of systems for so-called “virtual reality” or “augmented reality” experiences, in which digitally reproduced images or portions thereof are presented to the user in a way that appears real or can be perceived as real. Virtual reality (or “VR”) scenarios typically involve the presentation of digital or virtual image information that is opaque to other actual visual input from the real world; augmented reality (or “AR”) scenarios typically involve the presentation of digital or virtual image information as an enhancement of the visualization of the real world surrounding the user.

[0004] Despite the progress made in these display technologies, there is a need in the art for improved methods, systems and apparatuses related to augmented reality systems, particularly display systems. Summary of the Invention

[0005] This disclosure generally relates to techniques for improving the performance of optical systems and user experience. More specifically, embodiments of this disclosure provide methods for operating augmented reality (AR), virtual reality (VR), or mixed reality (MR) wearable systems, wherein user hand gestures are used to interact within a virtual environment.

[0006] An overview of various embodiments of the invention is provided below as a list of examples. As used below, any reference to a series of examples will be understood as a reference to each of those separate examples (e.g., "Examples 1-4" will be understood as "Examples 1, 2, 3, or 4").

[0007] Example 1 is a method for interacting with a virtual object, the method comprising: receiving an image of a user's hand; analyzing the image to detect a plurality of keypoints associated with the user's hand; determining, based on the analysis of the image, whether the user's hand is making or transitioning to making a gesture among a plurality of gestures; and in response to determining that the user's hand is making or transitioning to making the gesture: determining a specific position relative to the plurality of keypoints, wherein the specific position is determined based on the plurality of keypoints and the gesture; registering interaction points to the specific position; and forming a multi-DOF controller for interacting with the virtual object based on the interaction points.

[0008] Example 2 is a system configured to perform the method described in Example 1.

[0009] Example 3 is a non-transitory machine-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to Example 1.

[0010] Example 4 is a method for interacting with a virtual object, the method comprising: receiving an image of a user's hand from one or more image capture devices of a wearable system; analyzing the image to detect a plurality of keypoints associated with the user's hand; determining, based on the analysis of the image, whether the user's hand is making or transitioning to make a specific gesture among a plurality of gestures; in response to determining that the user's hand is making or transitioning to make the specific gesture: selecting a subset of the plurality of keypoints corresponding to the specific gesture; determining a specific position relative to the subset of the plurality of keypoints, wherein the specific position is determined based on the subset of the plurality of keypoints and the specific gesture; registering an interaction point to the specific position; registering neighboring points to positions along the user's body; projecting light from the neighboring points through the interaction point; and forming a multi-DOF controller for interacting with the virtual object based on the light.

[0011] Example 5 is the method according to Example 4, wherein the plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.

[0012] Example 6 is the method according to Example 4, wherein the subset of the plurality of key points is selected from a plurality of subsets of the plurality of key points, wherein each subset of the plurality of key points corresponds to a different gesture among the plurality of gestures.

[0013] Example 7 is the method according to Example 4, further comprising: displaying a graphical representation of the multi-DOF controller.

[0014] Example 8 is the method according to Example 4, wherein the location to which the neighboring point is registered is located at an estimated location of the user's shoulder, an estimated location of the user's elbow, or between the estimated location of the user's shoulder and the estimated location of the user's elbow.

[0015] Example 9 is the method according to Example 4, further comprising: capturing the image of the user's hand by an image capturing device.

[0016] Example 10 is the method according to Example 9, wherein the image capturing device is an element of a wearable system.

[0017] Example 11 is the method according to Example 9, wherein the image capture device is mounted on a headset of a wearable system.

[0018] Example 12 is the method according to Example 4, further comprising: determining, based on analyzing the image, whether the user's hand is performing a motion event.

[0019] Example 13 is the method according to Example 12, further comprising: modifying the virtual object based on the multi-DOF controller and the action event in response to determining that the user's hand is performing the action event.

[0020] Example 14 is the method according to Example 13, wherein the user's hand is determined to be performing the action event based on the specific gesture.

[0021] Example 15 is the method according to Example 4, wherein the user’s hand is making or is transitioning to make the specific gesture based on the plurality of key points.

[0022] Example 16 is the method according to Example 15, wherein the user’s hand is making or is transitioning to making the specific gesture based on neural network inference using the plurality of key points.

[0023] Example 17 is the method according to Example 4, wherein the user's hand is making or is transitioning to making the specific gesture based on neural network inference using the image.

[0024] Example 18 is the method according to Example 4, wherein multiple key points are located on the user's hand.

[0025] Example 19 is the method according to Example 4, wherein the multi-DOF controller is a 6DOF controller.

[0026] Example 20 is a system configured to perform the method described according to any one of Examples 4-19.

[0027] Example 21 is a non-transitory machine-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of Examples 4-19.

[0028] Example 22 is a method comprising: receiving a sequence of images of a user's hand; analyzing each image in the image sequence to detect a plurality of keypoints on the user's hand; determining, based on the analysis of one or more images in the image sequence, whether the user's hand is making or transitioning to make any of a plurality of different gestures; in response to determining that the user's hand is making or transitioning to make a specific gesture among the plurality of different gestures: selecting, from a plurality of positions relative to the plurality of keypoints corresponding to the plurality of different gestures, a specific position relative to the plurality of keypoints corresponding to the specific gesture; selecting, from a plurality of different subsets of the plurality of keypoints corresponding to the plurality of different gestures, a position corresponding to the specific gesture. The system involves: a specific subset of the plurality of keypoints; when it is determined that the user's hand is making or transitioning to making a specific gesture: registering an interaction point to the specific position on the user's hand relative to the plurality of keypoints; registering a neighboring point to an estimated position on the user's shoulder, an estimated position on the user's elbow, or a position along the user's upper arm between the estimated position on the user's shoulder and the estimated position on the user's elbow; projecting light from the neighboring point through the interaction point; displaying a graphical representation of the multi-DoF controller corresponding to the light; and repositioning and / or reorienting the multi-DoF controller based on the specific subset of the plurality of keypoints, the positions of the interaction point, and the neighboring point.

[0029] Example 23 is the method according to Example 22, wherein the image sequence is received from one or more outward-facing cameras on a headset.

[0030] Example 24 is the method according to Example 22, wherein the plurality of different gestures include at least one of grasping gestures, pointing gestures, or pinching gestures.

[0031] Example 25 is the method according to Example 22, further comprising: when it is determined that the user's hand is making a grasping gesture: registering the interaction point to a key point along the user's index finger; determining the orientation or direction of the light based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or including key point I. m T m M mThe relative positions of the subset of the plurality of key points in H, consisting of three or more key points.

[0032] Example 26 is the method according to Example 22, further comprising: when it is determined that the user's hand is making a pointing gesture: aligning the interaction point to a key point at the tip of the user's index finger; determining the orientation or direction of the light based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or including I t I d I p I m T t T i T m M m The relative positions of subsets of the plurality of key points in H and three or more key points; and at least in part based on and The angle θ between them (i.e., ) to detect action events, where γ represents The midpoint.

[0033] Example 27 is the method according to Example 26, wherein a hovering action event is detected if θ is determined to be greater than a predetermined threshold.

[0034] Example 28 is the method according to Example 26, wherein a touch action event is detected if θ is determined to be less than a predetermined threshold.

[0035] Example 29 is the method according to Example 22, further comprising: when it is determined that the user's hand is making a pinch gesture: registering the interaction point to along... or or The location; the orientation or direction of the light is determined based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or includes I t I d I p I m T t T i T m M m The relative positions of subsets of the plurality of key points in H and three or more key points; and at least in part based on and The angle θ between them (i.e., ) to detect action events, where γ represents T m I m The midpoint.

[0036] Example 30 is the method according to Example 29, wherein a hovering action event is detected if θ is determined to be greater than a predetermined threshold.

[0037] Example 31 is the method according to Example 29, wherein a touch action event is detected if θ is determined to be less than a predetermined threshold.

[0038] Example 32 is the method according to Example 29, wherein a tapping action event is detected based on a duration for which θ is determined to be less than a predetermined threshold.

[0039] Example 33 is the method according to Example 29, wherein a hold action event is detected based on a duration for which θ is determined to be less than a predetermined threshold.

[0040] Example 34 is the method according to Example 22, further comprising: when it is determined that the user's hand is transitioning between a grasping gesture and a pointing gesture: registering the interaction point to along... or The location; determining the orientation or direction of the light in the same manner as for the pointing gesture; and determining the orientation or direction of the light based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or including keypoint I. t I d I p I m T t T i T m M m The relative positions of the subset of the plurality of key points in H, consisting of three or more key points.

[0041] Example 35 is the method according to Example 34, wherein when the user's index finger is partially extended outward while the other fingers of the user's hand are curled inward, it is determined that the user's hand is transitioning between making the grasping gesture and making the pointing gesture.

[0042] Example 36 is the method according to Example 22, further comprising: when it is determined that the user's hand is transitioning between a pointing gesture and a pinching gesture: registering the interaction point to along... The location; determining the orientation or direction of the light in the same manner as for the pointing gesture and / or the pinching gesture; and determining the orientation or direction of the light based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or including keypoint I. t I d I p I m T t T i T m M m The relative positions of the subset of the plurality of key points in H, consisting of three or more key points.

[0043] Example 37 is the method according to Example 36, wherein the user’s hand is determined to be transitioning between making a pointing gesture and making a pinching gesture when the user’s thumb and index finger extend at least partially outward and curl at least partially toward each other.

[0044] Example 38 is the method according to Example 22, further comprising: when it is determined that the user's hand is transitioning between making a pinching gesture and making a grasping gesture: registering the interaction point to along... The location; determining the orientation or direction of the light in the same manner as for a pinch gesture; and determining the orientation or direction of the light based at least in part on: the specific location of the interaction point; the location of at least one part of the user's body other than the user's hand; and / or including keypoint I. t ,I d ,I p ,I m ,T t ,T i ,T m M m The relative positions of subsets of multiple keypoints of three or more keypoints in H.

[0045] Example 39 is the method according to Example 38, wherein it is determined that the user's hand is transitioning between making the pinching gesture and making the grasping gesture when the user's thumb and index finger are at least partially extended outward and at least partially curled toward each other.

[0046] Example 40 is a system configured to perform the method described in any of Examples 22-39.

[0047] Example 41 is a non-transitory machine-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of Examples 22-39.

[0048] Example 42 is a method for interacting with a virtual object, the method comprising: receiving one or more images of a user's first hand and a second hand; analyzing the one or more images to detect a plurality of keypoints associated with each of the first hand and the second hand; determining an interaction point for each of the first hand and the second hand based on the plurality of keypoints associated with each of the first hand and the second hand; generating one or more hand increments based on the interaction points for each of the first hand and the second hand; and using the one or more hand increments to interact with the virtual object.

[0049] Example 43 is the method according to Example 42, further comprising: determining a two-handed interaction point based on the interaction point of each of the first hand and the second hand.

[0050] Example 44 is the method according to Example 42, wherein: the interaction point of the first hand is determined based on the plurality of key points associated with the first hand; and the interaction point of the second hand is determined based on the plurality of key points associated with the second hand.

[0051] Example 45 is the method according to Example 42, wherein determining the interaction point of each of the first hand and the second hand includes: determining, based on analyzing the one or more images, whether the first hand is making or transitioning to make a first specific gesture among a plurality of gestures; and in response to determining that the first hand is making or transitioning to make the first specific gesture: selecting a subset of the plurality of keypoints associated with the first hand corresponding to the first specific gesture; determining a first specific location relative to the subset of the plurality of keypoints associated with the first hand, wherein the first specific location is determined based on the subset of the plurality of keypoints associated with the first hand and the first specific gesture; and registering the interaction point of the first hand to the first specific location.

[0052] Example 46 is the method according to Example 45, wherein determining the interaction point of each of the first hand and the second hand further includes: determining, based on analyzing the one or more images, whether the second hand is making or transitioning to make a second specific gesture among the plurality of gestures; in response to determining that the second hand is making or transitioning to make the second specific gesture: selecting a subset of the plurality of keypoints associated with the second hand corresponding to the second specific gesture; determining a second specific position relative to the subset of the plurality of keypoints associated with the second hand, wherein the second specific position is determined based on the subset of the plurality of keypoints associated with the second hand and the second specific gesture; and registering the interaction point of the second hand to the second specific position.

[0053] Example 47 is the method according to Example 46, wherein the plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.

[0054] Example 48 is the method according to Example 42, wherein the one or more images include a first image of the first hand and a second image of the second hand.

[0055] Example 49 is the method according to Example 42, wherein the one or more images include a single image of the first hand and the second hand.

[0056] Example 50 is the method according to Example 42, wherein the one or more images comprise a series of time-series imaging.

[0057] Example 51 is the method according to Example 42, wherein the one or more hand increments are determined based on frame-to-frame movement of the interaction point of each of the first hand and the second hand.

[0058] Example 52 is the method according to Example 51, wherein the one or more hand increments include a translation increment corresponding to a frame-to-frame translation movement of the interaction point of each of the first hand and the second hand.

[0059] Example 53 is the method according to Example 51, wherein the one or more hand increments include a rotation increment corresponding to a frame-to-frame rotational movement of the interaction point of each of the first hand and the second hand.

[0060] Example 54 is the method according to Example 51, wherein the one or more hand increments include a sliding increment of frame-to-frame separation movement corresponding to the interaction point of each of the first hand and the second hand.

[0061] Example 55 is a system configured to perform the methods described according to any of Examples 42-54.

[0062] Example 56 is a non-transitory machine-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any of Examples 42-54. Attached Figure Description

[0063] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this specification, illustrate embodiments of the disclosure and, together with the detailed description, serve to explain the principles of the disclosure. No attempt is made to compare the basic understanding of the disclosure with the various ways in which the disclosure may be practiced, and the structural details of the disclosure may be shown in more detail as needed.

[0064] Figure 1 An example operation of a wearable system that provides hand gesture input for interacting with virtual objects is shown.

[0065] Figure 2 A schematic diagram of an example AR / VR / MR wearable system is shown.

[0066] Figure 3 An example method for interacting with a virtual user interface is shown.

[0067] Figure 4A Examples of light rays and cone casting are shown.

[0068] Figure 4B An example of a cone projection over a set of objects is shown.

[0069] Figure 5 Examples of various key points that can be detected or tracked by wearable systems are shown.

[0070] Figures 6A-6F Examples of possible subsets of key points that can be selected based on gestures recognized by the wearable system are shown.

[0071] Figures 7A-7C Examples of light projections for various gestures are shown when the user's arm is extended.

[0072] Figures 8A-8C Examples of light projection for various gestures are shown when the user's arm retracts.

[0073] Figure 9 An example of how to use keypoints to detect motion events is shown.

[0074] Figures 10A-10C An example interaction with a virtual object using light is shown.

[0075] Figure 11 An example scheme for managing pointing gestures is shown.

[0076] Figure 12 An example scheme for managing pinch gestures is shown.

[0077] Figure 13 An example scheme for detecting motion events when a user's hand is making a grasping gesture is shown.

[0078] Figure 14 An example scheme for detecting motion events when a user's hand is making a pointing gesture is shown.

[0079] Figure 15 An example scheme for detecting motion events when a user's hand is making a pinching gesture is shown.

[0080] Figure 16 Example experimental data is shown for detecting motion events when a user's hand is making a pinching gesture.

[0081] Figures 17A-17D Example experimental data is shown for detecting motion events when a user's hand is making a pinching gesture.

[0082] Figure 18 An example scheme for detecting motion events when a user's hand is making a pinching gesture is shown.

[0083] Figures 19A-19D The experiment data presented is noisy and is used to detect motion events when a user's hand is making a pinching gesture.

[0084] Figures 20A-20C An example scheme for managing grasping gestures is shown.

[0085] Figures 21A-21C An example scheme for managing pointing gestures is shown.

[0086] Figures 22A-22C An example scheme for managing pinch gestures is shown.

[0087] Figure 23 The various activation types for pointing and pinching gestures are shown.

[0088] Figure 24 It shows various gestures and transitions between gestures.

[0089] Figure 25 An example of bimanual interaction is shown.

[0090] Figure 26 An example of two-handed interaction is shown.

[0091] Figure 27 Various examples of collaborative two-handed interaction are shown.

[0092] Figure 28 An example of managed two-handed interaction is shown.

[0093] Figure 29 Example one-handed (manual) interactive fields and two-handed interactive fields are shown.

[0094] Figure 30 A method is shown to form a multi-DOF controller associated with the user's hand to allow the user to interact with virtual objects.

[0095] Figure 31 A method is shown to form a multi-DOF controller associated with the user's hand to allow the user to interact with virtual objects.

[0096] Figure 32 This demonstrates a method for interacting with virtual objects using two-handed input.

[0097] Figure 33 A simplified computer system according to some embodiments described herein is shown. Detailed Implementation

[0098] Wearable systems can present interactive augmented reality (AR), virtual reality (VR), and / or mixed reality (MR) environments, where virtual data elements interact with the user through various inputs. While many modern computing systems are designed to generate a given output based on a single direct input (e.g., a computer mouse guides the cursor in response to direct user manipulation), high specificity may be required to achieve specific tasks in data-rich and dynamically interactive environments such as AR / VR / MR environments. Otherwise, without precise input, computing systems may suffer from high error rates and potentially lead to incorrect computer operations being performed. For example, when a user intends to move an object in three-dimensional (3D) space using a touchpad, the computing system may struggle to interpret the desired 3D movement using a device with an inherent two-dimensional (2D) input space.

[0099] Using hand gestures as input in AR / VR / MR environments has many appealing features. First, in AR environments where virtual content is overlaid on the real world, hand gestures provide an intuitive way to interact between the two worlds. Second, there is a wide range of expressive hand gestures that can potentially be mapped to various input commands. For example, hand gestures can simultaneously represent multiple different parameters, such as hand shape (e.g., different configurations the hand can take), orientation (e.g., different relative rotations of the hand), position, and movement. Third, with recent hardware improvements in imaging devices and processing units, hand gesture input provides sufficient accuracy so that the complexity of the system can be reduced by employing additional inputs from various sensors, such as electromagnetic tracking transmitters / receivers (e.g., handheld controllers).

[0100] One method for recognizing hand gestures is to track the positions of various key points on one or both of a user's hands. In one implementation, a hand-tracking system can identify the 3D positions of more than 20 key points on each hand. The gestures associated with the hand can then be identified by analyzing these key points. For example, the distance between different key points can indicate whether the user's hand is clenched (e.g., low average distance) or open and relaxed (e.g., high average distance). As another example, various angles formed by three or more key points (e.g., including at least one key point along the user's index finger) can indicate whether the user's hand is pointing or pinching.

[0101] Once a gesture is identified, the interaction point through which the user can interact with the virtual object can be determined. This interaction point can be registered to a keypoint or a location between keypoints, with each gesture having a unique algorithm for determining the interaction point. For example, when making a pointing gesture, the interaction point can be registered to the keypoint at the tip of the user's index finger. As another example, when making an open pinch gesture, the interaction point can be registered to the midpoint between the tip of the user's index finger and the tip of the user's thumb. Some gestures may also allow for a radius associated with the interaction point to be determined. As an example, for a pinch gesture, the radius could be related to the distance between the tip of the user's index finger and the tip of the user's thumb.

[0102] Continuing to track the entire network of keypoints after a gesture has been identified and / or after an interaction point has been determined can be computationally intensive. Therefore, in some embodiments of this disclosure, once a gesture has been identified, a subset of the total number of keypoints can be tracked. This subset of keypoints can be used to periodically update the interaction point with a more manageable computational burden than using the total number of keypoints. In some examples, this subset of keypoints can be used to periodically update the orientation of a virtual multi-DOF controller (e.g., a virtual cursor or pointer associated with an interaction point) with a relatively high degree of computational efficiency, as described in further detail below. Furthermore, the subset of keypoints can be analyzed to determine whether the user's hand is no longer making a gesture, or, for example, has changed from making a first gesture to making a second gesture or has changed from a first gesture to an unrecognized gesture.

[0103] In addition to identifying the interaction point, neighboring points along the user's body (or in space) can be identified, allowing control rays (or simply "rays") to extend between these two points. A ray (or a portion thereof) can be used as a cursor or pointer (e.g., as part of a multi-DOF controller) to interact with virtual content in 3D space. In some cases, neighboring points can be registered to the user's shoulder, elbow, or along the user's arm (e.g., between the user's shoulder and elbow). Alternatively, neighboring points can be registered to one or more other locations within or along the surface of the user's body, such as knuckles, hands, wrists, forearms, elbows, arms (e.g., upper arms), shoulders, scapulae, neck, head, eyes, face (e.g., cheeks), chest, torso (e.g., the naval area), or combinations thereof. A ray can then extend a specific distance from a neighboring point through the interaction point. Each of the interaction point, neighboring points, and ray can be dynamically updated to provide a responsive and comfortable user experience.

[0104] The embodiments described herein relate to one-handed interactions, referred to as hand-based interactions, and two-handed interactions, referred to as two-handed interactions. Tracking a one-handed pose may include tracking the interaction point of the one hand (e.g., its position, orientation, and radius) and optionally its corresponding neighboring points and rays, as well as any gestures the hand is making. For two-handed interactions, the interaction point of each of the user's hands (e.g., position, orientation, and radius) and optionally its corresponding neighboring points, rays, and gestures may be tracked. Two-handed interactions may also require tracking the two-handed interaction point between the two hands, which may have a position (e.g., average position), orientation (e.g., average orientation), and radius (e.g., average radius). Frame-to-frame movement of the two-handed interaction point can be captured by two-handed increments (delta), which can be calculated based on the increments of the two hands, as described below.

[0105] The two-hand increments can include a translation component called a translation increment and a rotation component called a rotation increment. The translation increment can be determined based on the translation increments of both hands. For example, the translation increment can be determined based on the left translation increment corresponding to the user's left hand and the right translation increment corresponding to the user's right hand (e.g., their average). Similarly, the rotation increment can be determined based on the rotation increments of both hands. For example, the rotation increment can be determined based on the left rotation increment corresponding to the user's left hand and the right rotation increment corresponding to the user's right hand (e.g., their average).

[0106] Alternatively or additionally, the rotation increment can be determined based on the rotational movement of a line formed between the positions of the interaction points. For example, a user can pinch two corners of a digital cube and rotate the cube by rotating the positions of the interaction points of both hands. This rotation can occur independently of whether the interaction points of each hand rotate themselves, or in some embodiments, the rotation of the cube can be further facilitated by the rotation of the interaction points. In some cases, the hand increment may include other components, such as a separation component, referred to as a separation increment (or scaling increment), which is determined based on the distance between the positions of the interaction points. This includes positive separation increments corresponding to the hands moving apart and negative separation increments corresponding to the hands moving closer together.

[0107] Various types of two-handed interactions can fall into one of three categories. The first category is independent two-handed interaction, where each hand interacts with a virtual object independently (e.g., a user is typing on a virtual keyboard, and each hand is configured independently of the other). The second category is collaborative two-handed interaction, where both hands interact with a virtual object collaboratively (e.g., resizing, rotating, and / or translating a virtual cube by pinching opposite corners with both hands). The third category is managed two-handed interaction, where one hand manages how the other hand interprets the interaction (e.g., the right hand is the cursor, while the left hand is the qualifier that switches the cursor between the pen and eraser).

[0108] In the following description, various examples will be described. Specific configurations and details are set forth for illustrative purposes to provide a thorough understanding of the examples. However, it will also be apparent to those skilled in the art that the examples can be practiced without specific details. Furthermore, well-known features may be omitted or simplified so as not to obscure the described embodiments.

[0109] Figure 1Example operation of a wearable system providing hand gesture input for interaction with a virtual object 108, according to some embodiments of this disclosure, is illustrated. The wearable system may include a wearable device 102 (e.g., a headset) worn by a user and including at least one forward-facing camera 104 that includes the user's hand 106 within its field of view (FOV). Therefore, captured images from the camera 104 may include the hand 106, thereby allowing the wearable system to perform subsequent image processing, such as detecting key points associated with the hand 106. In some embodiments, references... Figure 1 The described wearable system and wearable device 102 can be respectively referred to in the following references. Figure 2 The wearable system 200 and wearable device 201 are described in further detail.

[0110] The wearable system can maintain a reference frame within which the position and orientation of elements in an AR / VR / MR environment can be determined. In some embodiments, the wearable system can determine the position relative to the reference frame defined as (X... WP ,Y WP Z WP The position of the wearable device 102 (“wearable position”) and its relative position to the reference frame are defined as (X). WO ,Y WO Z WO The orientation of the wearable device 102 (“wearable orientation”) can be represented by X, Y, and Z Cartesian values ​​or by longitude, latitude, and altitude values, among other possibilities. The orientation of the wearable device 102 can be represented by X, Y, and Z Cartesian values ​​or by pitch, yaw, and roll angle values, among other possibilities. The reference frame for each of the position and orientation can be a world reference frame, or alternatively or additionally, the position and orientation of the wearable device 102 can be used as a reference frame such that, for example, the position of the wearable device 102 can be set to (0,0,0), and the orientation of the wearable device 102 can be set to (0°,0°,0°).

[0111] The wearable system can use images captured by camera 104 to perform one or more processing steps 110. In some examples, one or more processing steps 110 may be performed by one or more processors, and may be performed at least in part by one or more processors of the wearable system, one or more processors communicatively coupled to the wearable system, or a combination thereof. At step 110-1, multiple keypoints (e.g., nine or more keypoints) are detected or tracked based on the captured images. At step 110-2, the tracked keypoints are used to determine whether hand 106 is making or transitioning to make one of a predetermined set of gestures. In the example shown, hand 106 is determined to be making a pinch gesture. Alternatively or additionally, the gesture can be predicted directly from the image without the intermediate step of detecting keypoints. Therefore, steps 110-1 and 110-2 can be performed simultaneously or sequentially in any order. In response to determining that the user's hand is making or transitioning to make a specific gesture (e.g., a pinch gesture), a subset of multiple keypoints (e.g., eight or fewer keypoints) associated with the specific gesture can be selected and tracked.

[0112] At step 110-3, interaction point 112 is determined by registering interaction point 112 to a specific location relative to a subset of selected keypoints, based on the predicted gesture (or predicted gesture change) from step 110-2. Similarly, at step 110-3, neighboring point 114 is determined by registering neighboring point 114 to a location along the user's body, at least in part based on one or more of various factors. Further at step 110-3, light ray 116 is projected from neighboring point 114 through interaction point 112. At step 110-4, a motion event performed by hand 106 is predicted based on keypoints (e.g., based on the movement of keypoints over time). In the illustrated example, hand 106 is determined to perform a targeting action that can be recognized by the wearable system when the user performs a dynamic pinch-to-open gesture.

[0113] Figure 2 A schematic diagram of an example AR / VR / MR wearable system 200 according to some embodiments of the present disclosure is shown. The wearable system 200 may include a wearable device 201 and at least one remote device 203 remote from the wearable device 201 (e.g., separate hardware but communicatively coupled). As described above, in some embodiments, as referenced... Figure 2 The described wearable system 200 and wearable device 201 can be respectively referred to in the above reference. Figure 1The wearable system and wearable device 102 are described. When the wearable device 201 is worn by a user (typically as a headset), the remote device 203 can be held by the user (e.g., as a handheld controller) or mounted in various configurations, such as being fixedly attached to a frame, fixedly attached to a helmet or hat worn by the user, embedded in a headset, or otherwise removably attached to the user (e.g., in a backpack configuration, in a belt-coupled configuration, etc.).

[0114] Wearable device 201 may include a left eyepiece 202A and a left lens assembly 205A arranged side-by-side, and a right eyepiece 202B and a right lens assembly 205B also arranged side-by-side. In some embodiments, wearable device 201 includes one or more sensors, including but not limited to: a left-facing forward-facing world camera 206A directly attached to or near the left eyepiece 202A; a right-facing forward-facing world camera 206B directly attached to or near the right eyepiece 202B; a left-facing side-facing world camera 206C directly attached to or near the left eyepiece 202A; and a right-facing side-facing world camera 206D directly attached to or near the right eyepiece 202B. Wearable device 201 may include one or more image projection devices, such as a left projector 214A optically linked to the left eyepiece 202A and a right projector 214B optically linked to the right eyepiece 202B.

[0115] Wearable system 200 may include a processing module 250 for collecting, processing, and / or controlling data within the system. Components of processing module 250 may be distributed between wearable device 201 and remote device 203. For example, processing module 250 may include a local processing module 252 on the wearable portion of wearable system 200 and a remote processing module 256 physically separated from and communicatively linked to local processing module 252. Each of local processing module 252 and remote processing module 256 may include one or more processing units (e.g., central processing unit (CPU), graphics processing unit (GPU), etc.) and one or more storage devices such as non-volatile memory (e.g., flash memory).

[0116] Processing module 250 can collect data captured by various sensors of wearable system 200, such as camera 206, depth sensor 228, remote sensor 230, ambient light sensor, eye tracker, microphone, inertial measurement unit (IMU), accelerometer, compass, Global Navigation Satellite System (GNSS) unit, radio device, and / or gyroscope. For example, processing module 250 can receive image 220 from camera 206. Specifically, processing module 250 can receive left-front image 220A from left-facing world camera 206A, right-front image 220B from right-facing world camera 206B, left-side image 220C from left-side camera 206C, and right-side image 220D from right-side camera 206D. In some embodiments, image 220 may include a single image, a pair of images, video containing image streams, video containing paired image streams, etc. Image 220 may be generated periodically and sent to processing module 250 when wearable system 200 is powered on, or may be generated in response to instructions sent by processing module 250 to one or more cameras.

[0117] Camera 206 can be configured in various positions and orientations along the outer surface of the wearable device 201 to capture images of the user's surroundings. In some cases, cameras 206A and 206B can be positioned to capture images substantially overlapping with the user's left and right eye FOVs, respectively. Therefore, the placement of cameras 206 can be near the user's eyes, but not so close as to blur the user's FOV. Alternatively or additionally, cameras 206A and 206B can be positioned to align with the coupling positions of virtual image lights 222A and 222B, respectively. Cameras 206C and 206D can be positioned to capture images of the user's side, for example, images within or outside the user's peripheral vision. Images 220C and 220D captured using cameras 206C and 206D do not necessarily overlap with images 220A and 220B captured using cameras 206A and 206B.

[0118] In various embodiments, processing module 250 may receive ambient light information from an ambient light sensor. The ambient light information may indicate a luminance value or a range of spatially resolved luminance values. Depth sensor 228 may capture a depth image 232 in a forward-facing direction of wearable device 201. Each value of depth image 232 may correspond to the distance between depth sensor 228 and the nearest detected object in a particular direction. As another example, processing module 250 may receive gaze information from one or more eye trackers. As another example, processing module 250 may receive projected image luminance values ​​from one or both projectors 214. Remote sensor 230 located within remote device 203 may include any of the aforementioned sensors with similar functionality.

[0119] Virtual content is primarily delivered to the user of the wearable system 200 using projector 214 and eyepiece 202. For example, eyepieces 202A and 202B may include transparent or semi-transparent waveguides configured to guide and couple light generated by projectors 214A and 214B, respectively. Specifically, processing module 250 may cause left projector 214A to output left virtual image light 222A to left eyepiece 202A, and may cause right projector 214B to output right virtual image light 222B to right eyepiece 202B. In some embodiments, each of eyepieces 202A and 202B may include multiple waveguides corresponding to different colors. In some embodiments, lens assemblies 205A and 205B may be coupled to and / or integrated with eyepieces 202A and 202B. For example, lens assemblies 205A and 205B can be incorporated into multilayer eyepieces and can form one or more layers constituting one of the eyepieces 202A and 202B.

[0120] During operation, the wearable system 200 can support various user interactions with objects within the field of regard (FOR) (i.e., the entire area available for viewing or imaging) based on contextual information. For example, the wearable system 200 can adjust the size of the aperture of the cone used by the user to interact with the object using cone projection. As another example, the wearable system 200 can adjust the amount of movement of a virtual object associated with actuation of the user input device based on contextual information. Detailed examples of these interactions are provided below.

[0121] The user's FOR (Foreign Object) may contain a set of objects that the user can perceive via the wearable system 200. These objects can be virtual and / or physical. Virtual objects may include operating system objects such as a recycle bin for deleted files, a terminal for entering commands, a file manager for accessing files or directories, icons, menus, applications for audio or video streaming, notifications from the operating system, etc. Virtual objects may also include objects from applications, such as avatars, virtual objects in games, graphics, or images. Some virtual objects may be both operating system objects and application objects. In some embodiments, the wearable system 200 may add virtual elements to existing physical objects. For example, the wearable system 200 may add a virtual menu associated with a television in a room, where the virtual menu provides the user with the option to turn on the television or change television channels using the wearable system 200.

[0122] The objects in the user's FOR can be part of a world map. Data associated with the objects (e.g., location, semantic information, attributes, etc.) can be stored in various data structures such as arrays, lists, trees, hashes, and graphs. The index of each applicable stored object can be determined, for example, by the object's location. For instance, a data structure can index objects using a single coordinate (such as the object's distance from a reference location—e.g., how far to the left (or right), top (or bottom), or depth from the reference location). In some implementations, the wearable system 200 can display virtual objects at different depth planes relative to the user, allowing interactive objects to be organized into multiple arrays located at different fixed depth planes.

[0123] Users can interact with a subset of objects in their FOR (Forward) list. This subset of objects is sometimes referred to as interactive objects. Users can interact with objects using various techniques, such as selecting objects, moving objects, opening menus or toolbars associated with objects, or selecting a new set of interactive objects. Users can interact with interactive objects by using hand gestures or actuating user input devices, such as clicking a mouse, tapping a touchpad, swiping on a touchscreen, hovering over or touching a capacitive button, pressing keys on a keyboard or game controller (e.g., a 5-way d-pad), pointing a joystick, wand, or totem at an object, pressing a button on a remote control, or other interactions with the user input device. Users can also interact with interactive objects using head, eye, or body gestures, such as gazing at or pointing at an object for a period of time. These hand gestures and hand gestures from the user can cause the wearable system 200 to initiate selection events, such as performing user interface operations (displaying menus associated with the target interactive object, performing game operations on the avatar in a game, etc.).

[0124] Figure 3An example method 300 for interacting with a virtual user interface according to some embodiments of the present disclosure is illustrated. At step 302, the wearable system can identify a specific user interface (UI). The type of UI can be predetermined by the user. The wearable system can identify the specific UI that needs to be populated based on user input (e.g., gestures, visual data, audio data, sensor data, direct commands, etc.). At step 304, the wearable system can generate data for the virtual UI. For example, data associated with boundaries, general structure, shape of the UI, etc., can be generated. Furthermore, the wearable system can determine the map coordinates of the user's physical location, allowing the wearable system to display the UI relative to the user's physical location. For example, if the UI is body-centric, the wearable system can determine the coordinates of the user's physical location, head pose, or eye pose, allowing a circular UI to be displayed around the user, or a planar UI to be displayed on a wall or in front of the user. If the UI is hand-centric, the map coordinates of the user's hand can be determined. These map points can be derived using data received by a FOV camera, sensor input, or any other type of collected data.

[0125] At step 306, the wearable system can send data from the cloud to the display, or it can send data from a local database to the display component. At step 308, the UI is displayed to the user based on the sent data. For example, a light field display can project a virtual UI onto one or both of the user's eyes. Once the virtual UI has been created, at step 310, the wearable system can simply wait for a command from the user to generate more virtual content on the virtual UI. For example, the UI could be a body-centric ring around the user's body. The wearable system can then wait for a command (gesture, head or eye movement, input from a user input device, etc.), and if the command is recognized (step 312), it can display the virtual content associated with the command to the user (step 314). As an example, the wearable system can wait for a user's hand gesture before blending multiple stream tracks.

[0126] As described herein, a user can interact with objects in their environment using hand gestures or hand gestures. For example, a user can look into a room and see a table, chairs, walls, and a virtual television display on one of the walls. To determine which objects the user is looking at, the wearable system 200 can use cone projection technology, which typically projects a cone in the direction the user is looking and identifies any objects that intersect the cone. Cone projection can involve projecting a single ray of light with no lateral thickness from the headset (of the wearable system 200) toward a physical or virtual object. Cone projection with a single ray can also be referred to as ray projection.

[0127] Ray projection can use a collision detection agent to trace along the ray and identify whether and where any object intersects the ray. Wearable system 200 can use an IMU (e.g., an accelerometer), an eye-tracking camera, etc., to track the user's posture (e.g., body, head, or eye orientation) to determine the direction the user is looking. Wearable system 200 can use the user's posture to determine the direction of the projected ray. Ray projection technology can also be used in conjunction with user input devices such as handheld, multi-degree-of-freedom (DOF) input devices. For example, as the user moves around, the user can actuate the multi-DOF input device to anchor the size and / or length of the ray. As another example, wearable system 200 can project light from a user input device instead of from a headset. In some embodiments, the wearable system can project a cone with a non-negligible aperture (transverse to the central ray) instead of projecting a ray with negligible thickness.

[0128] Figure 4A Examples of light and conical projection according to some embodiments of the present disclosure are shown. Conical projection can project a conical (or other shaped) volume 420 having an adjustable aperture. The cone 420 can be a geometric cone having an interaction point 428 and a surface 432. The size of the aperture can correspond to the size of the surface 432 of the cone. For example, a large aperture can correspond to a large surface area of ​​surface 432. As another example, a large aperture can correspond to a large diameter 426 of surface 432, while a small aperture can correspond to a small diameter 426 of surface 432. Figure 4A As shown, the interaction point 428 of the cone 420 can have its origin at various locations, such as the center of the user's ARD (e.g., between the user's eyes), one of the user's limbs (e.g., the hand, such as the fingers of the hand), or a point on a totem or user input device (e.g., a toy weapon) held or manipulated by the user. It should be understood that interaction point 428 represents an example of an interaction point that can be generated using one or more of the systems and techniques described herein, and other interaction point arrangements are possible and within the scope of the invention.

[0129] The central ray 424 can represent the direction of the cone. The direction of the cone can correspond to the user's body posture (such as head posture, hand gestures, etc.) or the user's gaze direction (also known as eye posture). Figure 4A Example 406 illustrates a cone projection with posture, where the wearable system can determine the cone's orientation 424 using the user's head or eye posture. The example also shows a coordinate system for the head posture. The head 450 can have multiple degrees of freedom. As the head 450 moves in different directions, the head posture will change relative to the natural rest direction 460. Figure 4AThe coordinate system in the diagram illustrates three angular degrees of freedom (e.g., yaw, pitch, and roll), which can be used to measure head posture relative to the head's natural stationary state. For example... Figure 4A As shown, the head 450 can tilt forward and backward (e.g., pitch), turn left and right (e.g., yaw), and tilt side to side (e.g., roll). In other implementations, other techniques or angle representations for measuring head posture can be used, such as any other type of Euler angle system. Wearable systems can use an IMU to determine the user's head posture.

[0130] Example 404 illustrates another example of a cone projection with posture, where the wearable system can determine the orientation 424 of the cone based on the user's hand gestures. In this example, the interaction point 428 of the cone 420 is at the fingertip of the user's hand 414. When the user points his finger to another location, the position of the cone 420 (and the central ray 424) can move accordingly.

[0131] The direction of the cone can also correspond to the position or orientation of the user input device or the actuation of the user input device. For example, the direction of the cone can be based on a user-drawn trajectory on the touch surface of the user input device. The user can move his finger forward on the touch surface to indicate that the direction of the cone is forward. Example 402 shows another cone projection using a user input device. In this example, the interaction point 428 is located at the tip of the weapon-shaped user input device 412. As the user input device 412 moves around, the cone 420 and the central ray 424 can also move with the user input device 412.

[0132] The wearable system can initiate cone projection when the user input device 466 is actuated by the user through actions such as clicking a mouse, tapping a touchpad, swiping on a touchscreen, hovering over or touching a capacitive button, pressing a button on a keyboard or game controller (e.g., a 5-way directional pad), pointing a joystick, baton, or totem at an object, pressing a button on a remote control, or other interactions with the user input device 466.

[0133] Wearable systems can also initiate cone projection based on the user's posture (e.g., an extended period of gazing in one direction) or hand gestures (e.g., waving in front of an outward-facing imaging system). In some implementations, the wearable system can automatically initiate a cone projection event based on contextual information. For example, the wearable system can automatically initiate cone projection when the user is at the home screen of an AR display. In another example, the wearable system can determine the relative position of objects in the user's gaze direction. If the wearable system determines that objects are positioned relatively far apart from each other, the wearable system can automatically initiate cone projection, so the user does not have to move with precision to select objects from a sparsely positioned group of objects.

[0134] The direction of the cone can also be based on the position or orientation of the headphones. For example, the cone can be projected in a first direction when the headphones are tilted, and in a second direction when the headphones are not tilted.

[0135] The cone 420 can have various attributes, such as size, shape, or color. These attributes can be displayed to the user, allowing the user to perceive the cone. In some cases, portions of the cone 420 can be displayed (e.g., the ends of the cone, the surfaces of the cone, the central rays of the cone, etc.). In other embodiments, the cone 420 can be a cuboid, polyhedron, pyramid, frustum, etc. The distal end of the cone can have any cross-section, such as circular, elliptical, polygonal, or irregular cross-section.

[0136] exist Figure 4A and Figure 4B In this design, the cone 420 may have a vertex positioned at an interaction point 428 and a distal end formed at a plane 432. The interaction point 428 (also referred to as the zero point of the central ray 424) may be associated with the location from which the cone projection originates. The interaction point 428 may be anchored to a location in 3D space such that the virtual cone appears to emanate from that location. This location may be a position on the user's head (such as between the user's eyes), a position on a user input device used as a pointer (e.g., a 6DOF or 3DOF handheld controller), the tip of a finger (which can be detected by gesture recognition), etc. For handheld controllers, the position to which the interaction point 428 is anchored may depend on the shape factor of the device. For example, in a weapon-shaped controller 412 (used in shooting games), the interaction point 428 may be at the tip of the muzzle of the controller 412. In this example, the interaction point 428 of the cone can originate from the center of the barrel, and the cone 420 (or the central ray 424) of the cone 420 can be projected forward such that the center of the cone projection will be concentric with the barrel of the weapon-shaped controller 412. In various embodiments, the interaction point 428 of the cone can be anchored to any location in the user environment.

[0137] Once the interaction point 428 of cone 420 is anchored to a position, the orientation and movement of cone 420 can be based on the movement of an object associated with that position. For example, as described with reference to Example 406, when cone 420 is anchored to a user's head, cone 420 can move based on the user's head posture. As another example, in Example 402, when cone 420 is anchored to a user input device, cone 420 can move based on actuation of the user input device (e.g., based on changes in the position or orientation of the user input device). As another example, in Example 404, when cone 420 is anchored to a user's hand, cone 420 can move based on movement of the user's hand.

[0138] The surface 432 of the cone can extend until it reaches a termination threshold. The termination threshold can relate to a collision between the cone and a virtual or physical object in the environment (e.g., a wall). The termination threshold can also be based on a threshold distance. For example, surface 432 can extend away from interaction point 428 until the cone collides with an object or until the distance between surface 432 and interaction point 428 reaches a threshold distance (e.g., 20 cm, 1 m, 2 m, 10 m, etc.). In some embodiments, the cone can extend beyond the object even if a collision between the cone and the object is possible. For example, surface 432 can extend through a real-world object (such as a table, chair, wall, etc.) and terminate when it hits a termination threshold. Assuming the termination threshold is a wall of a virtual room located outside the user's current room, the wearable system can extend the cone beyond the current room until it reaches the surface of the virtual room. In some embodiments, a world grid can be used to define the extent of one or more rooms. The wearable system can detect the presence of a termination threshold by determining whether the virtual cone intersects with a portion of the world grid. In some embodiments, the user can easily aim at the virtual object as the cone extends through a real-world object. As an example, a headset can display a virtual hole on a physical wall, allowing the user to remotely interact with virtual content in other rooms even if the user is not physically in another room.

[0139] The cone 420 may have a depth. The depth of the cone 420 may be represented by the distance between the interaction point 428 and the surface 432. The depth of the cone may be automatically adjusted by the wearable system, the user, or a combination thereof. For example, when the wearable system determines that an object is far from the user's location, the wearable system may increase the depth of the cone. In some implementations, the depth of the cone may be anchored to a depth plane. For example, the user may choose to anchor the depth of the cone to a depth plane within 1 meter of the user. As a result, during cone projection, the wearable system will not capture objects beyond the 1-meter boundary. In some embodiments, if the depth of the cone is anchored to a depth plane, the cone projection will only capture objects at that depth plane. Therefore, the cone projection will not capture objects closer to or farther from the user than the anchored depth plane. In addition to setting the depth of the cone 420 or as an alternative, the wearable system may set the surface 432 to a depth plane so that the cone projection allows the user to interact with objects at or below that depth plane.

[0140] The wearable system can anchor the depth of the cone, interaction point 428, or surface 432 upon detecting a hand gesture, body posture, gaze direction, actuation of a user input device, voice command, or other technology. In addition to or as an alternative to the examples described herein, the anchoring position of interaction point 428, surface 432, or anchoring depth can be based on contextual information such as the type of user interaction, the function of the object to which the cone is anchored, etc. For example, interaction point 428 can be anchored to the center of the user's head due to user availability and perception. As another example, when the user points at an object using a hand gesture or user input device, interaction point 428 can be anchored to the tip of the user's finger or the tip of the user input device to increase the accuracy of the direction the user is pointing.

[0141] The wearable system can generate a visual representation of at least a portion of cone 420 or ray 424 to display to a user. Properties of cone 420 or ray 424 can be reflected in the visual representation of cone 420 or ray 424. The visual representation of cone 420 can correspond to at least a portion of the cone, such as a hole in the cone, a surface of the cone, a central ray, etc. For example, in the case where the virtual cone is a geometric cone, the visual representation of the virtual cone can include a gray geometric cone extending from a position between the user's eyes. As another example, the visual representation can include a portion of the cone that interacts with real or virtual content. Assuming the virtual cone is a geometric cone, the visual representation can include a circular pattern representing the base of the geometric cone, since the base of the geometric cone can be used to aim and select virtual objects. In some embodiments, the visual representation is triggered based on user interface operations. As an example, the visual representation can be associated with the state of an object. The wearable system can present the visual representation when the object changes from a stationary state or a hovering state (where the object can be moved or selected). The wearable system can further hide the visual representation when the object changes from a hovering state to a selected state. In some implementations, when the object is hovering, the wearable system can receive input from the user input device (in addition to or as an alternative to cone projection), and can allow the user to select the virtual object using the user input device when the object is hovering.

[0142] In some embodiments, the cone 420, the ray 424, or a portion thereof may be invisible to the user (e.g., may not be displayed to the user). The wearable system may assign a focus indicator indicating the direction and / or position of the cone to one or more objects. For example, the wearable system may assign a focus indicator to an object in front of the user and intersecting the user's gaze direction. The focus indicator may include a halo, color, perceptual size or depth variation (e.g., making the target object appear closer and / or larger when selected), a change in the shape of a cursor sub-image graphic (e.g., the cursor changes from a circle to an arrow), or other auditory, tactile, or visual effects to attract the user's attention. The cone 420 may have an aperture transverse to the ray 424. The size of the aperture may correspond to the size of the surface 432 of the cone. For example, a large aperture may correspond to a large diameter 426 on surface 432, while a small aperture may correspond to a small diameter 426 on surface 432.

[0143] For reference Figure 4BFurther described, the aperture can be adjusted by a user, the wearable system, or a combination thereof. For example, a user can adjust the aperture through user interface operations such as selecting options for the aperture shown on an AR display. A user can also adjust the aperture by actuating a user input device, for example, by scrolling the user input device or by pressing a button to anchor the size of the aperture. In addition to or alternative to user input, the wearable system can update the size of the aperture based on one or more contextual factors.

[0144] Conical projection can be used to increase accuracy when interacting with objects in a user's environment, especially when these objects are located at distances where small movements from the user can translate into large movements of light. Conical projection can also be used to reduce the amount of movement required by the user to make the cone overlap with one or more virtual objects. In some implementations, for example by using a narrower cone when many objects are present and a wider cone when fewer objects are present, the user can update the cone's aperture with one hand and improve the speed and accuracy of selecting target objects. In other implementations, the wearable system can determine the contextual factors associated with objects in the user's environment and allow automatic cone updates as an adjunct to or alternative to one-handed updates, which can advantageously make it easier for the user to interact with objects in the environment because less user input is required.

[0145] Figure 4B An example of a cone or ray projection on a group of objects 430 (e.g., objects 430A, 430B) in the user's FOR 400 is shown. The objects can be virtual and / or physical objects. During the cone or ray projection, the wearable system can project a cone 420 or ray 424 (visible or invisible to the user) in one direction and identify objects that intersect with the cone 420 or ray 424. For example, object 430A (shown in bold) intersects with cone 420. Object 430B is outside cone 420 and does not intersect with cone 420.

[0146] Wearable systems can automatically update the aperture based on contextual information. Contextual information can include information related to the user's environment (e.g., lighting conditions in the user's virtual or physical environment), user preferences, the user's physical condition (e.g., whether the user is nearsighted), information associated with objects in the user's environment (such as the type of objects in the user's environment (e.g., physical or virtual) or the layout of objects (e.g., object density, object position, and size, etc.)), characteristics of the objects the user is interacting with (e.g., object functionality, the type of user interface operations supported by the object, etc.), and combinations thereof. Density can be measured in various ways, such as the number of objects per projection area, the number of objects per solid angle, etc. Density can also be represented in other ways, such as the spacing between adjacent objects (having smaller intervals reflecting increasing density). Wearable systems can use object position information to determine the layout and density of objects in a region. Figure 4B As shown, the wearable system can determine that the density of the object group 430 is high. The wearable system can accordingly use a cone 420 with a smaller aperture.

[0147] Wearable systems can dynamically update the aperture (e.g., size or shape) based on the user's posture. For example, the user can initially point... Figure 4B The device is initially grouped into object group 430, but when the user moves their hand, they can now point to object groups that are sparsely positioned relative to each other. Therefore, the wearable system can increase the size of the aperture. Similarly, if the user moves their hand back to object group 430, the wearable system can decrease the size of the aperture.

[0148] Additionally or alternatively, wearable systems can update the aperture size based on user preferences. For example, if a user prefers to select large groups of items at once, the wearable system can increase the aperture size.

[0149] As another example of dynamically updating the aperture based on contextual information, the wearable system can increase the aperture size if the user is in a dark environment or if the user is nearsighted, making it easier for the user to capture objects. In some implementations, the first cone projection can capture multiple objects. The wearable system can perform a second cone projection to further select a target object from the captured objects. The wearable system can also allow the user to select a target object from the captured objects using body posture or user input devices. The object selection process can be recursive, where one, two, three, or more cone projections can be performed to select a target object.

[0150] Figure 5Examples of various key points 500 associated with a user's hand that can be detected or tracked by a wearable system according to some embodiments of the present disclosure are shown. For each key point, uppercase characters correspond to regions of the hand as follows: "T" corresponds to the thumb, "I" to the index finger, "M" to the middle finger, "R" to the ring finger, "P" to the little finger, "H" to the hand, and "F" to the forearm. Lowercase characters correspond to more specific locations within each region of the hand as follows: "t" corresponds to the tip (e.g., fingertip), "i" to the interphalangeal joint ("IP joint"), "d" to the distal interphalangeal joint ("DIP joint"), "p" to the proximal interphalangeal joint ("PIP joint"), "m" to the metacarpophalangeal joint ("MCP joint"), and "c" to the carpometacarpophalangeal joint ("CMC joint").

[0151] Figures 6A-6F Examples of possible subsets of key points 500 selectable based on gestures recognized by a wearable system, according to some embodiments of this disclosure, are shown. In each example, key points included in the selected subset are outlined in bold, key points not included in the selected subset are outlined in dashed lines, and optional key points that can be selected to facilitate subsequent determination are outlined in solid lines. In each example, when selecting a subset of key points, each key point in the subset can be used to determine the orientation of an interaction point, a virtual multi-DOF controller (e.g., a virtual cursor or pointer associated with the interaction point), or both.

[0152] Figure 6A Examples of a subset of keypoints can be selected when it is determined that the user's hand is making or transitioning to a grasping gesture (e.g., all the user's fingers curl inward). In the example shown, keypoint I... m T m M m H can be included in a subset and used to determine the specific location to which interaction point 602A is registered. For example, interaction point 602A can be registered to keypoint I. m In some examples, a subset of keypoints can also be used to at least partially determine the orientation of the virtual multi-DOF controller associated with interaction point 602A. In some implementations, the subset of keypoints associated with grasping gestures may include keypoint I. m T m M m And three or more of H. In some embodiments, the specific location to which the interaction point 602A will be registered and / or the orientation of the virtual multi-DOF controller can be determined, regardless of some or all of the key points excluded from the subset of key points associated with the grasp gesture.

[0153] Figure 6BExamples of a subset of keypoints that can be selected when it is determined that the user's hand is making or is transitioning to making a pointing gesture (e.g., the user's index finger is fully extended outward while the other fingers of the user's hand are curled inward). In the example shown, keypoint I t I d I p I m T t T i T m M m H can be included in a subset and used to determine the specific location to which interaction point 602B is registered. For example, interaction point 602B can be registered to keypoint I. t In some examples, a subset of keypoints can also be used to at least partially determine the orientation of the virtual multi-DOF controller associated with interaction point 602B. In some implementations, the subset of keypoints associated with pointing gestures may include keypoint I. t I d I p I m T t T i T m M m And three or more of H. For example, by Figure 6B The key points are represented by their outlines; in some embodiments, key point I... d M m One or more of H can be excluded from a subset of keypoints associated with the pointing gesture. In some embodiments, the specific location to which the interaction point 602B will be registered and / or the orientation of the virtual multi-DOF controller can be determined, regardless of some or all of the keypoints excluded from the subset of keypoints associated with the pointing gesture.

[0154] Figure 6C Examples of a subset of keypoints are shown when it is determined that the user's hand is making or transitioning to a pinching gesture (e.g., the user's thumb and forefinger are at least partially extended outward and close to each other). In the example shown, keypoint I... t I d I p I m T t T i T m M m H can be included in a subset and used to determine the specific location to which interaction point 602C is registered. For example, interaction point 602C can be registered along... The location, for example, The midpoint (“α”). Alternatively, the interaction point can be registered along... The location, for example, The midpoint (“β”), or along The location, for example, The midpoint (“γ”). Alternatively, the interaction point can be registered along... The location, for example, The midpoint, or along The location, for example, The midpoint. In some examples, a subset of keypoints can also be used to at least partially determine the orientation of the virtual multi-DOF controller associated with interaction point 602C. In some implementations, the subset of keypoints associated with the pinch gesture may include keypoint I. t I d I p I m T t T i T m M m And three or more of H. For example... Figure 6C The key points are represented by their outlines; in some embodiments, key point I... d M m One or more of H can be excluded from a subset of keypoints associated with the pinch gesture. In some embodiments, the specific location to which the interaction point 602C will be registered and / or the orientation of the virtual multi-DOF controller can be determined, regardless of some or all of the keypoints excluded from the subset of keypoints associated with the pinch gesture.

[0155] Figure 6D Examples of a subset of keypoints can be selected when it is determined that the user's hand is transitioning between a grasping gesture and a pointing gesture (e.g., the user's index finger is partially extended outward while the other fingers of the user's hand are curled inward). In the example shown, keypoint I t I d I p I m T t T i T m M m H can be included in a subset and used to determine the specific location to which interaction point 602D is registered. For example, interaction point 602D can be registered along... or The location. Additionally or alternatively, interaction points can be registered along... or The position relative to the user's hand, registered to the interaction point 602D in some embodiments, can change as the user switches between grasping and pointing gestures. and (or along) and / or The visual representation (e.g., light) of the interaction point 602D displayed to the user can reflect the same movement. That is, in these embodiments, when the user changes between grasping and pointing gestures, the position relative to the user's hand registered to the interaction point 602D may not suddenly appear at the key point I. m and I t Instead of snapping between points, it slides along one or more paths between such key points to provide a smoother and more intuitive user experience.

[0156] In some examples, as the user transitions between grasping and pointing gestures, the position of the visual representation of the interaction point 602D displayed relative to the user's hand can intentionally track the position of the actual interaction point 602D based on the current position of a subset of keypoints at a given time. For example, when the user transitions between grasping and pointing gestures, the position of the visual representation of the interaction point 602D displayed to the user in frame n can correspond to the position of the actual interaction point 602D based on the position of a subset of keypoints in frame (nm), where m is a predetermined number of frames (e.g., a fixed time delay). In another example, when the user transitions between grasping and pointing gestures, the visual representation of the interaction point 602D displayed to the user can be configured to move at a portion (e.g., a predetermined percentage) of the speed of the actual interaction point 602D based on the current position of a subset of keypoints at a given time. In some embodiments, one or more filters or filtering techniques may be employed to implement one or more of these behaviors. In some implementations, when the user is not switching between gestures or otherwise maintaining a particular gesture, the position of the visual representation of the interaction point 602D displayed relative to the user's hand and the position of the actual interaction point 602D based on the current position of a subset of key points at any given point in time may have little or no difference. Other configurations are also possible.

[0157] Figure 6E Examples of a subset of keypoints are shown when determining a shift in a user's hand between a pointing gesture and a pinching gesture (e.g., the user's thumb and forefinger extending at least partially outward and curling at least partially toward each other). In the example shown, keypoint I... t I d I p I m T t Ti T m M m H can be included in a subset and used to determine the specific location to which interaction point 602E is registered. For example, interaction point 602E can be registered along... The location. In some embodiments, as the user transitions between pointing and pinching gestures, a visual representation (e.g., light) of the interaction point 602E may be displayed to the user and / or may be referenced above. Figure 6D The described method is similar or equivalent to representing the actual interaction points 602E based on the current location of a subset of key points at a given point in time, which can be used to enhance the user experience.

[0158] Figure 6F Examples of a subset of key points are shown when determining a shift in a user's hand between a pinching gesture and a grasping gesture (e.g., the user's thumb and forefinger extending at least partially outward and curling at least partially toward each other). In the example shown, key point I... t I d I p I m T t T i T m M m H can be included in a subset and used to determine the specific location to which interaction point 602F is registered. For example, interaction point 602F can be registered along... The location. In some embodiments, as the user transitions between pinching and gripping gestures, a visual representation (e.g., light) of the interaction point 602F may be displayed to the user and / or may be referenced above. Figure 6D-6E The described method is similar or equivalent to the actual interaction point 602F based on the current location of a subset of key points at a given point in time, which can be used to enhance the user experience.

[0159] Figures 7A-7C Examples of light projection for various gestures when a user's arm is outstretched, according to some embodiments of this disclosure, are shown. Figure 7A This demonstrates a user making a grasping gesture with their arm extended outwards. Interaction point 702A is registered to key point I. m (as referenced) Figure 6A As described, and the neighboring point 704A is registered to the position of the user's shoulder (marked "S"). A ray 706A can be projected from the neighboring point 704A through the interaction point 702A.

[0160] Figure 7B This demonstrates a user pointing gesture while extending their arm outward. Interaction point 702B is registered to keypoint I.t (as referenced) Figure 6B As described, the neighboring point 704B is registered to the user's shoulder (marked "S"). A ray 706B can be projected from the neighboring point 704B through the interaction point 702B. Figure 7C This illustrates a user making a pinch gesture while their arm is extended outward. Interaction point 702C is registered to position α (as referenced). Figure 6C As described, the nearest point 704C is registered to the user's shoulder (marked "S"). A ray 706C can be projected from the nearest point 704C through the interaction point 702C. When the user is... Figure 7A and 7B Between gestures Figure 7B and 7C Between gestures and Figure 7A and 7C The range of positions where the interaction point can be registered when transitioning between gestures is shown in the references above. Figure 6D , Figure 6E and Figure 6F It was described in more detail.

[0161] Figures 8A-8C Examples of light projection for various gestures when a user retracts their arm, according to some embodiments of this disclosure, are shown. Figure 8A This demonstrates a user making a grasping gesture as their arm retracts inward. Interaction point 802A is registered to keypoint I. m (as referenced) Figure 6A As described, the neighboring point 804A is registered to the user's elbow (labeled "E"). A ray 806A can be projected from the neighboring point 804A through the interaction point 802A.

[0162] Figure 8B This demonstrates a user pointing gesture as their arm retracts inward. Interaction point 802B is registered to keypoint I. t (as referenced) Figure 6B As described, the nearest point 804B is registered to the user's elbow (labeled "E"). A ray 806B can be projected from the nearest point 804B through the interaction point 802B. Figure 8C This illustrates a user making a pinch gesture as their arm retracts inward. Interaction point 802C is registered to position α (as referenced). Figure 6C As described, the nearest point 804C is registered to the user's elbow (labeled "E"). A ray 806C can be projected from the nearest point 804C through the interaction point 802C. When the user is... Figure 8A and 8B Between gestures Figure 8B and 8C Between gestures and Figure 8A and 8C The range of positions where the interaction point can be registered when transitioning between gestures is also shown in the reference above. Figure 6D , Figure 6E and Figure 6F It was described in more detail.

[0163] It can be seen that, relative to the user's body being registered to Figures 7A-7C The locations of the neighboring points 704A-704C differ from those registered relative to the user's body. Figures 8A-8C The location of the neighboring points 804A-804C. This difference in location could be... Figures 7A-7C The position and / or orientation of one or more parts of the user's arm (e.g., the user's arm extended outwards) and Figures 8A-8C The results are the difference between the position and / or orientation of one or more parts of the user's arm (e.g., the user's arm is retracted), etc. Therefore, in Figures 7A-7C The position and / or orientation of one or more parts of the user's arm in the image. Figures 8A-8C When the position and / or orientation of one or more portions of the user's arm changes between different locations, the location to which the neighboring point is registered can change between a location at the user's shoulder ("S") and a location at the user's elbow ("E"). In some embodiments, when the position and / or orientation of one or more portions of the user's arm changes between different locations, the neighboring point is registered to a location that changes between different locations. Figures 7A-7C Those in Figures 8A-8C When transforming between those, it can be referenced above. Figure 6D-6F The methods described are similar or equivalent in representing neighboring points and one or more visual representations associated with neighboring points, which can be used to enhance the user experience.

[0164] In some embodiments, the system may register neighbor points to one or more estimated locations within or along the surface of the user's following areas: knuckles, hand, wrist, forearm, elbow, arm (e.g., upper arm), shoulder, scapula, neck, head, eyes, face (e.g., cheek), chest, torso (e.g., navel area), or combinations thereof. In at least some of these embodiments, the system may dynamically shift the location to which the neighbor point is registered between one or more such estimated locations based on at least one of a variety of different factors. For example, the system may determine the location to which neighboring points should be registered based on at least one of a variety of different factors including: (a) the user's hand is identified as making or transitioning to a gesture (e.g., grasping, pointing, pinching, etc.), (b) the location and / or orientation of a subset of keypoints associated with the user's hand being identified as making or transitioning to a gesture, (c) the location of the interaction point, (d) the location and / or orientation of the user's hand (e.g., pitch, yaw, and / or roll), (e) one or more measurements of wrist flexion and / or extension, (f) one or more measurements of wrist adduction and / or abduction, (g) the estimated location and / or orientation of the user's forearm (e.g., pitch, yaw, and / or roll), and (h) [other factors]. One or more measurements of forearm supination and / or pronation, (i) one or more measurements of elbow flexion and / or extension, (j) estimated position and / or orientation (e.g., pitch, yaw, and / or roll) of the user's arm (e.g., upper arm), (l) one or more measurements of shoulder flexion and / or extension, (m) one or more measurements of shoulder adduction and / or abduction, (n) estimated position and / or orientation of the user's head, (o) estimated position and / or orientation of the wearable device, (p) estimated distance between the user's hand or interaction point and the user's head or wearable device, (q) estimated length or span of the user's entire arm (e.g., from shoulder to fingertips) or at least a portion thereof, (r) one or more measurements of the user's visual coordination attention, or (s) a combination thereof.

[0165] In some embodiments, the system may determine or otherwise evaluate one or more of the foregoing factors based at least in part on data received from one or more outward-facing cameras, data received from one or more inward-facing cameras, data received from one or more other sensors of the system, data received as user input, or a combination thereof. In some embodiments, when one or more of the foregoing factors change, the system may refer to the above-mentioned factors. Figure 6D-8C The methods described are similar or equivalent in representing neighboring points and one or more visual representations associated with them, which can be used to enhance the user experience.

[0166] In some embodiments, the system may be configured such that (i) wrist adduction can be used to bias the adjacent point as registered along the user's arm toward the user's knuckles, while wrist abduction can be used to bias the adjacent point as registered along the user's arm toward the user's shoulder, neck, or other locations closer to the center of the user's body; (ii) elbow flexion can be used to bias the adjacent point as registered downward toward the user's navel area, while elbow extension can be used to bias the adjacent point as registered downward toward the user's head, shoulder, or other locations in the upper part of the user's body; and (iii) medial shoulder rotation can be used to bias the adjacent point as registered along the user's arm toward the user's elbow, hand, or knuckles, while lateral shoulder rotation... (iv) Shoulder adduction can be used to offset the location where the neighboring point is determined to be registered, towards the user's shoulder, neck, or other location closer to the center of the user's body; shoulder abduction can be used to offset the location where the neighboring point is determined to be registered, towards the user's head, neck, chest, or other location closer to the center of the user's body; and shoulder abduction can be used along the user's arm towards the user's shoulder, arm, or other location further away from the center of the user's body; or (v) combinations thereof. Thus, in these embodiments, the location where the neighboring point will be registered, determined by the system, can dynamically change over time as the user repositions and / or reorients one or more of their hand, forearm, and arm. In some examples, the system can assign different weights to different factors and determine the location where the neighboring point will be registered based on one or more such factors and their assigned weights. For example, the system can be configured to give more weight to one or more measurements of the user's visual coordination attention than some or all of the other aforementioned factors. Other configurations are also possible.

[0167] For an example in which the system is configured to dynamically shift the registered location of a neighboring point between one or more estimated locations based at least in part on one or more measurements of the user's visual coordinated attention, such one or more measurements may be determined by the system at least in part on the user's eye gaze, one or more characteristics of the virtual content being presented to the user, hand position and / or orientation, one or more transmodal convergence and / or divergence, or a combination thereof. Examples of transmodal convergence and divergence, as well as systems and techniques for detecting and responding to the occurrence of such transmodal convergence and divergence, are provided in U.S. Patent Publication No. 2019 / 0362557, the entire contents of which are incorporated herein by reference. In some embodiments, the system may utilize one or more of the systems and / or techniques described in the preceding patent application to detect the occurrence of one or more transmodal convergence and / or divergence, and may further determine the location of the neighboring point at least in part based on the detected occurrence of one or more transmodal convergence and / or divergence. Other configurations are also possible.

[0168] Figure 9 Examples of how keypoints can be used to detect motion events (e.g., hover, touch, tap, hold, etc.) according to some embodiments of this disclosure are shown. In some embodiments, this can be based at least in part on... and The angle θ between them (i.e., ) to detect action events, where γ represents The midpoint. For example, if θ is determined to be greater than a predetermined threshold, a "hover" action event can be detected, while if θ is determined to be less than the predetermined threshold, a "touch" action event can be detected. As another example, "tap" and "hold" action events can be detected based on the duration for which θ is determined to be less than the predetermined threshold. In the example shown, I t and T t It can represent key points included in a subset of key points selected in response to determining that a user is making or is transitioning to making a specific gesture (e.g., a pinch gesture).

[0169] Figures 10A-10C Examples of interaction between light and virtual objects are shown according to some embodiments of this disclosure. Figures 10A-10C Some of the paradigms conveyed above can be used in wearable systems, and some can be used by users for totemless interaction (e.g., interaction without using a physical handheld controller). Figures 10A-10CEach of these includes rendering what a user of the wearable system might see at various points in time while interacting with the virtual object 1002 using their hands. In this example, the user is able to manipulate the position of the virtual object by: (1) making a pinching gesture with their hands to magically conjure a virtual 6DoF ray 1004; (2) positioning their hands so that the virtual 6DoF ray intersects the virtual object; (3) while maintaining the position of their hands, bringing the tips of their thumbs and index fingers closer together so that the value of angle θ changes from greater than a threshold to less than the threshold while the virtual 6DoF ray is intersecting the virtual object; and (4) while keeping their thumbs and index fingers tightly pinched together to maintain angle θ at a value below the threshold, guiding their hands to a new position.

[0170] Figure 10A This illustrates that when a user's hand is identified as making a pinch gesture, interaction point 1006 is registered to position α. ​​Key points associated with the pinch gesture (e.g., I) can be selected based on the determination that the user is making or transitioning to making a pinch gesture. t I p I m T t T i and T m The location of a subset of keypoints is used to determine the α location. This selected subset of keypoints can be tracked, used to determine the location to which interaction point 1006 is registered (e.g., the α location), and also used to determine the location relative to the reference above. Figure 9 An angle θ that is similar to or equivalent to the previously described angle θ.

[0171] exist Figure 10A In the example shown, a ray 1004 has been projected from a location near the user's right shoulder or upper arm through an interaction point. A graphical representation of a portion of the ray forward from the interaction point is displayed via headphones and used by the user as a pointer or cursor, which is then used to interact with the virtual object 1002. Figure 10A In this example, the user has positioned their hand so that the virtual 6DoF ray intersects with the virtual object. Here, the angle θ is approximately greater than a threshold, causing the user to be perceived as simply hovering over the virtual object using the virtual 6DoF ray. Therefore, the system can compare the angle θ to one or more thresholds and, based on that comparison, determine whether the user is perceived as touching, grasping, or otherwise selecting the virtual content. In the example shown, the system can determine that the angle θ is greater than one or more thresholds, and therefore determine that the user is not perceived as touching, grasping, or otherwise selecting the virtual content.

[0172] Figure 10BThis shows that the user's hand is still positioned so that the virtual 6DoF ray intersects with the virtual object, and the pinch gesture is still made (note that the interaction point is still registered to the α position). However, in Figure 10B In this context, users have brought the tips of their thumbs and index fingers closer together. Therefore, in Figure 10B In this case, the angle θ is approximately below one or more thresholds, making the user now perceived as touching, grabbing, or otherwise selecting virtual objects using virtual 6DoF light.

[0173] Figure 10C This shows that the user is still making the same pinch gesture as they did in the previous images, and therefore the angle θ is probably below the threshold. However, in Figure 10C In this scenario, the user has moved their arm while keeping their thumb and forefinger tightly clenched together to effectively drag the virtual object to a new position. It should be noted that the interaction point has moved along with the user's hand by being registered to the α position. Although not in... Figures 10A-10C As shown, but instead of adjusting the position of the virtual object by adjusting the position of the interaction point relative to the headset while "holding" the virtual object, the user can also adjust the orientation (e.g., yaw, pitch, and / or roll of the virtual object) of the keypoint system associated with the pinch gesture relative to the headset while "holding" the virtual object (e.g., yaw, pitch, and / or roll of at least one vector and / or at least one plane defined by at least two and / or at least three keypoints included in a subset of selected keypoints). Figures 10A-10C Not shown, but after manipulating the position and / or orientation of a virtual object, a user can "let go" of the virtual object by separating their thumb and forefinger. In such an example, the system can determine that the angle θ is again greater than one or more thresholds, and therefore determine that the user is no longer considered to be touching, grasping, or otherwise selecting the virtual content.

[0174] Figure 11 Example schemes for managing pointing gestures according to some embodiments of the present disclosure are shown. Preferably, the interaction point 1102 is registered to a key point at the tip of the index finger (e.g., I...). t Key points). When the index finger tip is unavailable (e.g., occluded or below a critical confidence level), interaction point 1102 moves to the next nearest neighbor index finger PIP key point (e.g., I). p Key points). When the index finger PIP is unavailable (e.g., occluded or below the critical confidence level), interaction point 1102 moves to the index finger MCP key point (e.g., I). m (Key points). In some embodiments, filters are applied to smooth transitions between different possible key points.

[0175] Figure 12 Example schemes for managing pinch gestures according to some embodiments of the present disclosure are shown. Preferably, the interaction point 1202 is registered to the midpoint between the key point at the tip of the index finger and the key point at the tip of the thumb (e.g., referenced above). Figure 6C (The α position is described). If the index finger tip keypoint is unavailable (e.g., occluded or below the critical confidence level), the interaction point 1202 moves to the midpoint between the index finger PIP keypoint and the thumb tip keypoint. If the thumb tip keypoint is unavailable (e.g., occluded or below the critical confidence level), the interaction point 1202 moves to the midpoint between the index finger tip keypoint and the thumb IP keypoint.

[0176] If neither the index finger tip keypoint nor the thumb tip keypoint is available, then move interaction point 1202 to the midpoint between the index finger PIP keypoint and the thumb IP keypoint (e.g., see reference above). Figure 6C (Described β position). If the index finger PIP keypoint is unavailable (e.g., occluded or below the critical confidence level), the interaction point 1202 moves to the midpoint between the index finger MCP keypoint and the thumb IP keypoint. If the thumb IP keypoint is unavailable (e.g., occluded or below the critical confidence level), the interaction point 1202 moves to the midpoint between the index finger PIP keypoint and the thumb MCP keypoint. If neither the index finger PIP keypoint nor the thumb IP keypoint is available, the interaction point 1202 moves to the midpoint between the index finger MCP keypoint and the thumb MCP keypoint (e.g., as described above). Figure 6C The location of γ described).

[0177] Figure 13 Example schemes for detecting motion events while a user's hand is making a grasping gesture, according to some embodiments of this disclosure, are illustrated. The relative angular distance and relative angular velocity can be tracked based on the angle between the vectors of the index finger and thumb. If the index finger tip keypoint is unavailable, the index finger PIP keypoint can be used to form the angle. If the thumb tip keypoint is unavailable, the thumb IP keypoint can be used to form the angle. (See above reference) Figure 6A Provide information on what can be done by a determined user. Figure 13 Additional description of a subset of key points selectively tracked when grasping a hand gesture.

[0178] At 1302, the first relative maximum angular distance (and its timestamp) can be detected. At 1304, the relative minimum angular distance (and its timestamp) can be detected. At 1306, the second relative maximum angular distance (and its timestamp) can be detected. The action event has been determined to have been executed based on the angular distance difference and time difference between the data detected at 1302, 1304, and 1306.

[0179] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance can be compared with one or more first thresholds (e.g., upper threshold and lower threshold), the difference between the relative minimum angular distance and the second relative maximum angular distance can be compared with one or more second thresholds (e.g., upper threshold and lower threshold), the difference between the timestamp of the first relative maximum angular distance and the timestamp of the relative minimum angular distance can be compared with one or more third thresholds (e.g., upper threshold and lower threshold), and the difference between the timestamp of the relative minimum angular distance and the timestamp of the second relative maximum angular distance can be compared with one or more fourth thresholds (e.g., upper threshold and lower threshold).

[0180] Figure 14 An example scheme for detecting a motion event while a user's hand is making a pointing gesture, according to some embodiments of this disclosure, is illustrated. The relative angular distance can be tracked based on the angle between the vectors of the index finger and thumb. At 1402, a first relative maximum angular distance (with its timestamp) can be detected. At 1404, a relative minimum angular distance (with its timestamp) can be detected. At 1406, a second relative maximum angular distance (with its timestamp) can be detected. The difference in angular distance and time difference between the data detected at 1402, 1404, and 1406 can be used to determine that a motion event has been performed. In some examples, such angular distances can be at least similar to those referenced above. Figure 9 and Figures 10A-10C The angle θ described. (See above for reference.) Figure 6B Provide information on what can be done by a determined user. Figure 14 Additional description of a subset of key points selectively tracked during pointing gestures.

[0181] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance can be compared with one or more first thresholds (e.g., upper threshold and lower threshold), the difference between the relative minimum angular distance and the second relative maximum angular distance can be compared with one or more second thresholds (e.g., upper threshold and lower threshold), the difference between the timestamp of the first relative maximum angular distance and the timestamp of the relative minimum angular distance can be compared with one or more third thresholds (e.g., upper threshold and lower threshold), and the difference between the timestamp of the relative minimum angular distance and the timestamp of the second relative maximum angular distance can be compared with one or more fourth thresholds (e.g., upper threshold and lower threshold).

[0182] Figure 15An example scheme for detecting a motion event while a user's hand is making a pinching gesture, according to some embodiments of this disclosure, is shown. The relative angular distance can be tracked based on the angle between the vectors of the index finger and thumb. At 1502, a first relative maximum angular distance (with its timestamp) can be detected. At 1504, a relative minimum angular distance (with its timestamp) can be detected. At 1506, a second relative maximum angular distance (with its timestamp) can be detected. The motion event being performed can be determined based on the difference in angular distance and time difference between the data detected at 1502, 1504, and 1506. In some examples, such angular distances can be at least similar to those referenced above. Figure 9 and Figures 10A-10C The angle θ described. (See above for reference.) Figure 6C Provide information on what can be done by a determined user. Figure 15 Additional description of a subset of key points selectively tracked during the pinch gesture.

[0183] For example, the difference between the first relative maximum angular distance and the relative minimum angular distance can be compared with one or more first thresholds (e.g., upper threshold and lower threshold), the difference between the relative minimum angular distance and the second relative maximum angular distance can be compared with one or more second thresholds (e.g., upper threshold and lower threshold), the difference between the timestamp of the first relative maximum angular distance and the timestamp of the relative minimum angular distance can be compared with one or more third thresholds (e.g., upper threshold and lower threshold), and the difference between the timestamp of the relative minimum angular distance and the timestamp of the second relative maximum angular distance can be compared with one or more fourth thresholds (e.g., upper threshold and lower threshold).

[0184] Figure 16 Example experimental data for detecting motion events while a user's hand is making a pinching gesture, according to some embodiments of this disclosure, are shown. Figure 16 The experimental data shown can correspond to Figure 15 The depiction of the user's hand movements. Figure 16In the model, the user's hand movement is characterized by the smoothed distance between the thumb and index finger. Noise removal during low-latency smoothing leaves the remaining signal exhibiting normalized relative separation inflections between paired finger features. Inflections such as those consisting of local minima, followed by local maxima, and then immediately after another local minima can be used to identify tapping actions. Additionally, the same inflection pattern can be seen in key pose states. Key pose A, followed by key pose B, and then back to A can also be used to identify tapping actions. Key pose inflections may be robust when hand keypoints have low confidence. Relative distance inflections can be used when key poses have low confidence. With high confidence, both inflections can be used to identify tapping actions for two feature variations.

[0185] Figures 17A-17D Example experimental data for detecting motion events while a user's hand is making a pinching gesture, according to some embodiments of this disclosure, are shown. Figures 17A-17D The experimental data shown can correspond to the user's hand repeatedly making... Figure 15 The movement is shown. Figure 17A This shows the distance between the tip of a user's index finger and the target content as the user's hand repeatedly approaches the target content. Figure 17B The angular distance between the tip of the user's index finger and the tip of the user's thumb is shown. Figure 17C The angular velocity corresponding to the angle formed by the tip of the user's index finger and the tip of the user's thumb is shown. Figure 17D The key attitude changes determined based on various data, which may optionally include... Figures 17A-17C The data shown. Figures 17A-17D The experimental data shown can be used to identify tapping motions. In some embodiments, all feature inflections can be utilized simultaneously or concurrently to reduce the false positive rate.

[0186] Figure 18 An example scheme for detecting motion events when a user's hand is making a pinching gesture, according to some embodiments of the present disclosure, is shown. Figure 18 and Figure 15 The difference lies in the fact that the user's middle, ring, and little fingers curl inwards.

[0187] Figures 19A-19D The following are examples of noisy experimental data for detecting motion events while a user's hand is making a pinching gesture, according to some embodiments of the present disclosure. Figures 19A-19D The experimental data shown can correspond to the user's hand repeatedly making... Figure 18 The movement shown. Figure 19A It shows the distance between the tip of the user's index finger and the target content. Figure 19B The angular distance between the tip of the user's index finger and the tip of the user's thumb is shown. Figure 19C The angular velocity corresponding to the angle formed by the tip of the user's index finger and the tip of the user's thumb is shown. Figure 19D The key attitude changes determined based on various data, which may optionally include... Figures 19A-19C The data shown. Figures 19A-19D The noisy experimental data shown can be used to identify tapping actions determined to have occurred within window 1902. This represents an edge case scenario that utilizes at least moderate confidence in determining all turns to qualify as a tapping action.

[0188] Figures 20A-20C An example scheme for managing hand gestures according to some embodiments of the present disclosure is shown. As described herein, a ray 2006 is projected from a proximity point 2004 (a position registered to the user's shoulder) through an interaction point 2002 (a position registered to the user's hand). Figure 20A This demonstrates a grip gesture capable of achieving gross pointing mechanical action. This can be used for robust far-field aiming. Figure 20B The size of the interaction point relative to the calculated hand radius is shown, characterized by the relative distance between fingertip features. Figure 20C It shows that when the hand changes from an open to a clenched fist posture, the hand radius decreases, and therefore the size of the interaction point decreases proportionally.

[0189] Figures 21A-21C An example scheme for managing pointing gestures according to some embodiments of the present disclosure is shown. As described herein, a ray 2106 is projected from a proximity point 2104 (a position registered to the user's shoulder) through an interaction point 2102 (a position registered to the user's hand). Figure 21A The diagram illustrates the selection of mechanical movements for pointing and utilizing finger joints for refined midfield aiming. Figure 21B This illustrates the release (opening) of the pointing hand gesture. The interaction point is placed at the tip of the index finger. The relative distance between the tip of the thumb and the tip of the index finger is the greatest, making the size of the interaction point proportionally large. Figure 21C This illustrates a pointing hand posture with the thumb curled below the index finger (closed). The relative distance between the tips of the thumb and index finger is minimal, resulting in a proportionally small interaction point size, but this interaction point is still positioned at the tip of the index finger.

[0190] Figures 22A-22CAn example scheme for managing pinch gestures according to some embodiments of the present disclosure is shown. As described herein, a ray 2206 is projected from a proximity point 2204 (registered to a position on the user's shoulder) through an interaction point 2202 (registered to a position on the user's hand). Figure 22A The diagram illustrates the selection of mechanical movements for pointing and utilizing finger joints for refined midfield aiming. Figure 22B This shows the open (OK) pinch gesture. The interaction point is placed at the midpoint between the tip of the index finger and the thumb, as one of several pinch styles enabled by the managed pinch gesture. The relative distance between the tip of the thumb and the tip of the index finger is the largest, making the size of the interaction point proportionally large. Figure 22C The image shows a (closed) pinched hand gesture, with the middle, ring, and little fingers curled inwards and the tips of the index finger and thumb touching. The relative distance between the tips of the thumb and index finger is minimal, resulting in a proportionally small interaction point, but this interaction point is still positioned at the midpoint between the fingertips.

[0191] Figure 23 Various activation types for pointing and pinching gestures according to some embodiments of the present disclosure are illustrated. For pointing gestures, activation types include touch (close), hover (open), tap, and hold. For pinching gestures, activation types include touch (close), hover (open), tap, and hold.

[0192] Figure 24 Various gestures and transitions between gestures are illustrated according to some embodiments of this disclosure. In the illustrated examples, the set of gestures includes grasping gestures, pointing gestures, and pinching gestures, as well as transition states between each gesture. Each gesture also includes sub-gestures (or sub-poses) that can be further specified by the wearable system. Grasping gestures may include fist sub-poses, control sub-poses, and stylus sub-poses, among other possibilities. Pointing gestures may include single-finger sub-poses and “L”-shaped sub-poses, among other possibilities. Pinching gestures may include open sub-poses, closed sub-poses, and “OK” sub-poses, among other possibilities.

[0193] Figure 25Examples of two-handed interaction according to some embodiments of the present disclosure are shown, wherein both of a user's hands are used to interact with a virtual object. In each of the examples shown, the pointing gesture of each of the user's hands is determined based on key points of each corresponding hand. Interaction points 2510 and 2512 of the user's two hands are determined based on the key points of the respective hands and the determined gesture. Interaction points 2510 and 2512 are used to determine a two-handed interaction point 2514, which can facilitate the selection and aiming of the virtual object for two-handed interaction. The two-handed interaction point 2514 can be registered to a position (e.g., a midpoint) along a line formed between interaction points 2510 and 2512.

[0194] In each of the examples shown, increment 2516 is generated based on the movement of one or both of interaction points 2510 and 2512. At 2502, increment 2516 is a translation increment corresponding to a frame-to-frame translation movement of one or both of interaction points 2510 and 2512. At 2504, increment 2516 is a scaling increment corresponding to a frame-to-frame separation movement of one or both of interaction points 2510 and 2512. At 2506, increment 2516 is a rotation increment corresponding to a frame-to-frame rotation movement of one or both of interaction points 2510 and 2512.

[0195] Figure 26 It shows a difference Figure 26 An example of two-handed interaction, differing in that each of the user's hands is determined to make a pinching gesture based on key points of each corresponding hand. Interaction points 2610 and 2612 for the user's two hands are determined based on the key points of each hand and the determined gesture. Interaction points 2610 and 2612 are used to determine a two-handed interaction point 2614, which facilitates the selection and aiming of virtual objects for two-handed interaction. The two-handed interaction point 2614 can be registered to a position (e.g., a midpoint) along the line formed between interaction points 2610 and 2612.

[0196] In each of the examples shown, increment 2616 is generated based on the movement of one or both of interaction points 2610 and 2612. At 2602, increment 2616 is a translation increment corresponding to a frame-to-frame translation movement of one or both of interaction points 2610 and 2612. At 2604, increment 2616 is a scaling increment corresponding to a frame-to-frame separation movement of one or both of interaction points 2610 and 2612. At 2606, increment 2616 is a rotation increment corresponding to a frame-to-frame rotation movement of one or both of interaction points 2610 and 2612.

[0197] Figure 27Various examples of collaborative two-handed interaction according to some embodiments of the present disclosure are shown, in which two hands collaboratively interact with a virtual object. The examples shown include pinch manipulation, pointing manipulation, flat-manipulate manipulation, hook-manipulate manipulation, fist-manipulate manipulation, and trigger-manipulate manipulation.

[0198] Figure 28 Examples of managed two-handed interactions according to some embodiments of the present disclosure are shown, wherein one hand manages how the other hand is interpreted. Examples shown include index finger-thumb-pinch + index finger-pointing, middle finger-thumb-pinch + index finger-pointing, index finger-middle finger-pointing + index finger-pointing, and index finger-trigger pull + index finger-pointing.

[0199] Figure 29 Example single-handed interaction field 2902 and two-handed interaction field 2904 are shown according to some embodiments of the present disclosure. Each of interaction fields 2902 and 2904 includes a peripheral space, an extended workspace, a workspace, and a task space. The camera of the wearable system can be oriented to capture one or both of the user's hands while operating in various spaces, depending on whether the system supports single-handed or two-handed interaction.

[0200] Figure 30 A method 3000 is illustrated for forming a multi-DOF controller associated with a user's hand to allow the user to interact with virtual objects, according to some embodiments of this disclosure. One or more steps of method 3000 may be omitted during execution of method 3000, and the steps of method 3000 need not be performed in the order shown. One or more steps of method 3000 may be executed by one or more processors of a wearable system, such as those included in the processing module 250 of wearable system 200. Method 3000 may be implemented as a computer-readable medium or computer program product comprising instructions that, when executed by one or more computers, cause the one or more computers to perform the steps of method 3000. Such a computer program product may be transmitted via a wired or wireless network as a data carrier signal carrying the computer program product.

[0201] At step 3002, an image of the user's hand is received. The image can be captured by an image capturing device that can be mounted on the wearable device. The image capturing device can be a camera (e.g., a wide-angle lens camera, a fisheye lens camera, an infrared (IR) camera) or a depth sensor, among other possibilities.

[0202] At step 3004, the image is analyzed to detect multiple keypoints associated with the user's hand. These keypoints may be on or near the user's hand (within a threshold distance of the user's hand).

[0203] At step 3006, based on the analyzed image, it is determined whether the user's hand is making or transitioning to any of a plurality of gestures. These plurality of gestures may include grasping gestures, pointing gestures, and / or pinching gestures, as well as other possibilities. If it is determined that the user's hand is making or transitioning to any gesture, method 3000 proceeds to step 3008; otherwise, method 3000 returns to step 3002.

[0204] At step 3008, a specific position relative to multiple keypoints is determined. This specific position can be determined based on the multiple keypoints and the gesture. For example, if it is determined that the user's hand is making a first gesture among multiple gestures, the specific position can be set to the position of the first keypoint among the multiple keypoints; and if it is determined that the user's hand is making a second gesture among multiple gestures, the specific position can be set to the position of the second keypoint among the multiple keypoints. Continuing with the above example, if it is determined that the user's hand is making a third gesture among the multiple gestures, the specific position can be set to the midpoint between the first and second keypoints. Alternatively or additionally, if it is determined that the user's hand is making a third gesture, the specific position can be set to the midpoint between the third and fourth keypoints.

[0205] At step 3010, the interaction point is registered to a specific location. Registering the interaction point to a specific location may include setting and / or moving the interaction point to the specific location. The interaction point (and similarly, the specific location) may be a 3D value.

[0206] At step 3012, a multi-DOF controller for interacting with virtual objects is formed based on the interaction points. The multi-DOF controller can correspond to rays of light projected from neighboring points through the interaction points. Rays can be used to perform various actions, such as: aiming, selecting, grabbing, scrolling, extracting, hovering, touching, tapping, and holding.

[0207] Figure 31A method 3100 is illustrated for forming a multi-DOF controller associated with a user's hand to allow the user to interact with virtual objects, according to some embodiments of this disclosure. One or more steps of method 3100 may be omitted during execution of method 3100, and the steps of method 3100 need not be performed in the order shown. One or more steps of method 3100 may be executed by one or more processors of a wearable system, such as those included in the processing module 250 of wearable system 200. Method 3100 may be implemented as a computer-readable medium or computer program product comprising instructions that, when executed by one or more computers, cause the one or more computers to perform the steps of method 3000. Such a computer program product may be transmitted via a wired or wireless network as a data carrier signal carrying the computer program product.

[0208] At step 3102, an image of the user's hand is received. Step 3102 can be similar to reference [reference needed]. Figure 30 Step 3002 is described.

[0209] At step 3104, the image is analyzed to detect multiple key points associated with the user's hand. Step 3104 can be similar to the reference... Figure 30 Step 3004 is described.

[0210] At step 3106, based on the analyzed image, it is determined whether the user's hand is making or transitioning to any of a plurality of gestures. Step 3106 may be similar to reference [reference needed]. Figure 30 Step 3006 is described. If it is determined that the user's hand is making or is changing into making any gesture, then method 3100 proceeds to step 3108. Otherwise, method 3100 returns to step 3102.

[0211] At step 3108, a subset of multiple key points corresponding to a specific gesture is selected. For example, a first subset of key points may correspond to a first gesture among multiple gestures, and a second subset of key points may correspond to a second gesture among multiple gestures. Continuing the example above, if it is determined that the user's hand is making a first gesture, a first subset of key points can be selected; or if it is determined that the user's hand is making a second gesture, a second subset of key points can be selected.

[0212] At step 3110, a specific position is determined relative to a subset of multiple keypoints. This specific position can be determined based on both the subset of keypoints and the gesture. For example, if it is determined that the user's hand is making a first gesture among multiple gestures, the specific position can be set to the position of the first keypoint within the first subset of the multiple keypoints. As another example, if it is determined that the user's hand is making a second gesture among multiple gestures, the specific position can be set to the position of the second keypoint within the second subset of the multiple keypoints.

[0213] In step 3112, the interaction points are registered to specific locations. Step 3112 can be similar to the reference... Figure 30 Step 3010 is described.

[0214] In step 3114, neighboring points are registered to positions along the user's body. The registered positions of neighboring points can be located at the estimated positions of the user's shoulder, the estimated positions of the user's elbow, or between the estimated positions of the user's shoulder and the estimated positions of the user's elbow.

[0215] At step 3116, light rays from the neighboring point are projected through the interaction point.

[0216] At step 3118, a multi-DOF controller for interacting with virtual objects is formed based on light rays. The multi-DOF controller can correspond to light rays projected from neighboring points through the interaction point. Light rays can be used to perform various actions, such as: aiming, selecting, grabbing, scrolling, extracting, hovering, touching, tapping, and holding.

[0217] At step 3120, a graphical representation of the multi-DOF controller is displayed via the wearable system.

[0218] Figure 32 A method 3200 for interacting with a virtual object using two-handed input according to some embodiments of the present disclosure is illustrated. One or more steps of method 3200 may be omitted during execution of method 3200, and the steps of method 3200 need not be performed in the order shown. One or more steps of method 3200 may be executed by one or more processors of a wearable system, such as those included in the processing module 250 of wearable system 200. Method 3200 may be implemented as a computer-readable medium or a computer program product comprising instructions that, when executed by one or more computers, cause the one or more computers to perform the steps of method 3200. Such a computer program product may be transmitted via a wired or wireless network as a data carrier signal carrying the computer program product.

[0219] At step 3202, one or more images of the user's first and second hands are received. Some of the one or more images may include both the first and second hands, and some may include only one of the two hands. The one or more images may include a series of time-series images. The one or more images may be captured by an image capture device that can be mounted on the wearable device. The image capture device may be a camera (e.g., a wide-angle lens camera, a fisheye lens camera, an infrared (IR) camera) or a depth sensor, among other possibilities.

[0220] At step 3204, one or more images are analyzed to detect multiple keypoints associated with each of the first and second hands. For example, one or more images may be analyzed to detect two separate sets of keypoints: multiple keypoints associated with the first hand and multiple keypoints associated with the second hand. Each set of multiple keypoints may be on or near the corresponding hand (within a threshold distance of the corresponding hand). In some embodiments, different sets of multiple keypoints may be detected for each time-series image or each image frame.

[0221] At step 3206, interaction points with respect to each of the first and second hands are determined based on multiple key points associated with each of the first and second hands. For example, the interaction points of the first hand may be determined based on multiple key points associated with the first hand, and the interaction points of the second hand may be determined based on multiple key points associated with the second hand. In some embodiments, it may be determined whether the first and second hands are making (or transitioning to making) a specific gesture among multiple gestures. Based on the specific gesture of each hand, the interaction points of each hand can be registered to specific locations, as described herein.

[0222] At step 3208, a two-handed interaction point is determined based on the interaction points of the first and second hands. In some embodiments, the two-handed interaction point can be the average position of the interaction points. For example, a line can be formed between the interaction points, and the two-handed interaction point can be registered to a point along the line (e.g., the midpoint). The location to which the two-handed interaction point is registered can also be determined based on the gesture that each hand is making (or transitioning to making). For example, if one hand is making a pointing gesture and the other hand is making a grasping or pinching gesture, the two-handed interaction point can be registered to the hand making the pointing gesture. As another example, if both hands are making the same gesture (e.g., a pinching gesture), the two-handed interaction point can be registered to the midpoint between the interaction points.

[0223] At step 3210, one or more hand increments may be generated based on the interaction point of each of the first hand and the second hand. In some embodiments, one or more hand increments may be generated based on the movement of the interaction point (e.g., frame-to-frame movement). For example, one or more hand increments may include translation increments, rotation increments, and / or scaling increments. Translation increments may correspond to translational movements of one or both interaction points, rotation increments may correspond to rotational movements of one or both interaction points, and scaling increments may correspond to separation movements of one or both interaction points.

[0224] In one example, a set of time-series images can be analyzed to determine if the interaction points of the first and second hands have moved closer together. In response, a scaling increment with a negative value can be generated to indicate that the interaction points have moved closer together. In another example, a set of time-series images can be analyzed to determine if the interaction points have moved further apart, and a scaling increment with a positive value can be generated to indicate that the interaction points have moved further apart.

[0225] In another example, a set of time-series images can be analyzed to determine that the interaction points of the first and second hands are both moving in the positive X direction. In response, a translation increment can be generated to indicate that the interaction points are moving along the positive X direction. In yet another example, a set of time-series images can be analyzed to determine that the interaction points of the first and second hands are rotating relative to each other (e.g., the line formed between the interaction points is rotating). In response, a rotation increment can be generated to indicate that the interaction points are rotating relative to each other.

[0226] In some embodiments, hand increments can be generated based on either the established plane or the interaction point. For example, the plane can be established based on the user's hand, head pose, hip position, real-world objects, virtual objects, and other possibilities. When establishing the plane, translation increments can be generated based on the projection of the interaction point onto the plane, rotation increments can be generated based on the rotation of the interaction point relative to the plane, and scaling increments can be generated based on the distance between the interaction point and the plane. In some examples, these increments may be referred to as plane increments.

[0227] The hand increments described above can be generated for the same set of time-series images. For example, hand increments including translation, rotation, and scaling increments can be generated for a single set of time-series images. In some examples, only specific types of hand increments can be generated based on the requirements of a particular application. For example, a user can initiate a scaling operation while keeping the position and orientation of a virtual object fixed. In response, scaling increments can be generated only, without translation and rotation increments. As another example, a user can initiate translation and rotation operations while keeping the size of a virtual object fixed. In response, translation and rotation increments can be generated only, without scaling increments. Other possibilities are envisioned.

[0228] At step 3212, one or more hand increments are used to interact with the virtual object. Interaction with the virtual object can be achieved by applying one or more hand increments (e.g., moving the virtual object using one or more hand increments). For example, applying a translation increment to the virtual object causes it to translate by a specific amount indicated by the translation increment, applying a rotation increment to the virtual object causes it to rotate by a specific amount indicated by the rotation increment, and applying a scaling increment to the virtual object causes it to scale proportionally by a specific amount indicated by the scaling increment / resize by that specific amount.

[0229] In some embodiments, it can be determined whether the virtual object is being targeted before interacting with it. In some cases, it can be determined whether the two-handed interaction points overlap with or are within a threshold distance of the virtual object. In some embodiments, it can be determined whether the virtual object is currently or previously selected by, for example, using two-handed interaction as described herein. In one example, a virtual object can be selected first using one-handed interaction, and then interacted with using two-handed interaction.

[0230] Figure 33 A simplified computer system 3300 according to some embodiments of the present disclosure is shown. Figure 33 The computer system 3300 shown can be incorporated into the apparatus described herein. Figure 33 A schematic diagram of one embodiment of a computer system 3300 capable of performing some or all of the steps of the methods provided by various embodiments is provided. It should be noted that... Figure 33 This is intended only to provide a general illustration of the various components; any or all of them may be used appropriately. Therefore, Figure 33 It extensively demonstrates how individual system components can be implemented in a relatively separate or relatively more integrated manner.

[0231] A computer system 3300 is shown, which includes hardware elements that can be electrically coupled via a bus 3305 or can be otherwise communicated as appropriate. The hardware elements may include one or more processors 3310, including but not limited to one or more general-purpose processors and / or one or more special-purpose processors, such as digital signal processing chips, graphics accelerators, etc.; one or more input devices 3315, which may include but are not limited to mice, keyboards, cameras, etc.; and one or more output devices 3320, which may include but are not limited to display devices, printers, etc.

[0232] The computer system 3300 may also include one or more non-transitory storage devices 3325 and / or communicate with them, which may include, but are not limited to, locally and / or network-accessible storage devices, and / or may include, but are not limited to, disk drives, drive arrays, optical storage devices, solid-state storage devices (such as random access memory (“RAM”) and / or read-only memory (“ROM”, which may be programmable, updatable, etc.). Such storage devices may be configured to implement any suitable data storage, including but not limited to various file systems, database structures, etc.

[0233] Computer system 3300 may also include communication subsystem 3330, which may include, but is not limited to, modems, network interface cards (wireless or wired), infrared communication devices, wireless communication devices and / or devices such as Bluetooth. TM Chipsets for devices such as 802.11 devices, WiFi devices, WiMax devices, cellular communication facilities, etc. The communication subsystem 3319 may include one or more input and / or output communication interfaces to allow data exchange with networks such as those described herein, including, for example, other computer systems, televisions, and / or any other devices described herein. Depending on the desired functionality and / or other implementation, portable electronic devices or similar devices may transmit images and / or other information via the communication subsystem 3319. In other embodiments, a portable electronic device (e.g., a first electronic device) may be incorporated into the computer system 3300, for example, as an input device 3315. In some embodiments, as described above, the computer system 3300 will also include a working memory 3335, which may include RAM or ROM devices.

[0234] Computer system 3300 may also include software elements shown as currently residing within working memory 3335, including operating system 3340, device drivers, executable libraries and / or other code, such as one or more applications 3345, which, as described herein, may include computer programs provided by various embodiments and / or may be designed to implement methods and / or configure systems provided by other embodiments. By way of example only, one or more processes described with respect to the above methods may be implemented as code and / or instructions executable by a computer and / or a processor within a computer; thus, in one respect, according to the described methods, such code and / or instructions may be used to configure and / or adapt a general-purpose computer or other device to perform one or more operations.

[0235] This collection of instructions and / or code may be stored on a non-transitory computer-readable storage medium, such as the storage device 3325 described above. In some cases, the storage medium may be incorporated into a computer system, such as computer system 3300. In other embodiments, the storage medium may be separate from the computer system (e.g., a removable medium such as an optical disc) and / or contained in an installation package, such that the storage medium can be used to program, configure, and / or adapt a general-purpose computer and the instructions / code stored thereon. These instructions may take the form of executable code that can be executed by computer system 3300, and / or may take the form of source code and / or installable code, which, for example, is compiled and / or installed on computer system 3300 using various commonly available compilers, installers, compression / decompression utilities, etc., and then takes the form of executable code.

[0236] It will be apparent to those skilled in the art that substantial changes can be made to suit specific requirements. For example, custom hardware may be used, and / or specific elements may be implemented in the hardware, software including portable software (such as applets), or both. Furthermore, connections to other computing devices, such as network input / output devices, may be employed.

[0237] As described above, in one aspect, some embodiments may employ a computer system, such as computer system 3300, to perform the methods according to various embodiments of the present technology. According to one set of embodiments, some or all of the processes of such methods are executed by computer system 3300 in response to processor 3310 executing one or more sequences of one or more instructions, which may be incorporated into operating system 3340 and / or other code, such as application program 3345, included in working memory 3335. Such instructions may be read into working memory 3335 from another computer-readable medium, such as one or more storage devices 3325. By way of example, execution of the sequence of instructions included in working memory 3335 may cause processor 3310 to perform one or more processes of the methods described herein. Additionally or alternatively, portions of the methods described herein may be executed by dedicated hardware.

[0238] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any medium that participates in providing data that enables a machine to operate in a particular manner. In embodiments implemented using computer system 3300, various computer-readable media may involve providing instructions / code to processor 3310 for execution and / or being available for storing and / or carrying such instructions / code. In many embodiments, computer-readable media are physical and / or tangible storage media. Such media may take the form of non-volatile or volatile media. Non-volatile media include, for example, optical discs and / or magnetic disks, such as storage devices 3325. Volatile media include, but are not limited to, dynamic memory, such as working memory 3335.

[0239] Common forms of physical and / or tangible computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, CD-ROMs, any other optical media, punched cards, paper tapes, any other physical media with perforated patterns, RAM, PROMs, EPROMs, FLASH-EPROMs, any other memory chips or cassette tapes or any other media from which a computer can read instructions and / or code.

[0240] Various forms of computer-readable media may involve carrying one or more sequences of one or more instructions to processor 3310 for execution. By way of example only, the instructions may initially be carried on a disk and / or optical disk of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit the instructions as signals via a transmission medium for reception and / or execution by computer system 3300.

[0241] The communication subsystem 3319 and / or its components typically receive signals, and then the bus 3305 may carry the signals and / or the data, instructions, etc., carried by the signals to the working memory 3335, from which the processor 3310 retrieves and executes the instructions. The instructions received by the working memory 3335 may optionally be stored on the non-transitory storage device 3325 before or after execution by the processor 3310.

[0242] The methods, systems, and apparatus discussed above are examples. Various configurations may appropriately omit, substitute, or add various processes or components. For example, in alternative configurations, the method may be performed in a different order than that described, and / or various stages may be added, omitted, and / or combined. Furthermore, features described with respect to certain configurations may be combined in various other configurations. Different aspects and elements of the configurations may be combined in a similar manner. Moreover, technology is evolving, and therefore many elements are examples and do not limit the scope of this disclosure or the claims.

[0243] Specific details are set forth in the specification to provide a thorough understanding of exemplary configurations, including implementations. However, configurations may be practiced without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary details to avoid obscuring the configuration. This description provides only exemplary configurations and does not limit the scope, applicability, or configuration of the claims. Rather, the prior description of the configuration will provide those skilled in the art with enabling descriptions for implementing the described techniques. Various changes may be made to the function and arrangement of the elements without departing from the spirit or scope of this disclosure.

[0244] Furthermore, the configuration can be described as a process, which is depicted as a schematic flowchart or block diagram. Although each operation can be described as a sequential process, many operations can be executed in parallel or simultaneously. Additionally, the order of operations can be rearranged. The process may have additional steps not included in the diagram. Furthermore, examples of the method can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments used to perform the necessary tasks can be stored in a non-transitory computer-readable medium such as a storage medium. The processor can execute the described tasks.

[0245] Several example configurations have been described, and various modifications, alternative constructions, and equivalents may be used without departing from the spirit of this disclosure. For example, the above elements may be components of a larger system, where other rules may take precedence over or otherwise modify the application of the technology. Similarly, numerous steps may be taken before, during, or after considering the above elements. Therefore, the above description does not limit the scope of the claims.

[0246] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly indicates otherwise. Thus, for example, a reference to “user” includes one or more such users, while a reference to “processor” includes one or more processors and their equivalents known to those skilled in the art, and so on.

[0247] Furthermore, when used in this specification and the appended claims, the words “comprising,” “including,” “including,” “included,” “already included,” and “currently including” are intended to specify the presence of the stated features, integers, components, or steps, but they do not exclude the presence or addition of one or more other features, integers, components, steps, actions, or groups.

[0248] It should also be understood that the examples and embodiments described herein are for illustrative purposes only, and various modifications or changes thereof will be suggested by those skilled in the art and will be included within the spirit and scope of this application and the scope of the appended claims.

Claims

1. A method for interacting with a virtual object, the method comprising: Receive images of the user's hand from one or more image capture devices of the wearable system; The image is analyzed to detect multiple key points associated with the user's hand; Based on the analysis of the image, it is determined whether the user's hand is making or transitioning to make a specific gesture among multiple gestures; In response to determining that the user's hand is making or is changing into making the specific gesture: Select a subset of the multiple key points corresponding to the specific gesture; Determine a specific position relative to a subset of the plurality of key points, wherein the specific position is determined based on the subset of the plurality of key points and the specific gesture; Register the interaction point to the specific location; Register the nearest point to a location along the user's body; Projecting light rays from the neighboring point through the interaction point; and A multi-DOF controller for interacting with the virtual object is formed based on the light rays; Determine angles associated with at least three of the plurality of key points, wherein the angles lie between (i) a first line connecting the tip of the thumb to the midpoint between the two key points on the thumb and index finger, and (ii) a second line connecting the tip of the index finger to the midpoint between the two key points on the thumb and index finger; and Based on the change of the angle from outside the specific range to within the specific range, the motion events associated with the multi-DOF controller are detected.

2. The method according to claim 1, wherein, The plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.

3. The method according to claim 1, wherein, Select a subset of the plurality of key points from a plurality of subsets of the plurality of key points, wherein each subset of the plurality of key points corresponds to a different gesture among the plurality of gestures.

4. The method according to claim 1, further comprising: A graphical representation of the multi-DOF controller is shown.

5. The method according to claim 1, wherein, The location to which the neighboring point is registered is located at the estimated position of the user's shoulder, the estimated position of the user's elbow, or between the estimated position of the user's shoulder and the estimated position of the user's elbow.

6. The method according to claim 1, further comprising: The image of the user's hand is captured by one or more of the image capturing devices.

7. The method according to claim 6, wherein, The image capture device is mounted on the headset of the wearable system.

8. A system for interacting with virtual objects, comprising: One or more processors; as well as A machine-readable medium including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including the following: Receive images of the user's hand from one or more image capture devices of the wearable system; The image is analyzed to detect multiple key points associated with the user's hand; Based on the analysis of the image, it is determined whether the user's hand is making or transitioning to make a specific gesture among multiple gestures; as well as In response to determining that the user's hand is making or is changing into making the specific gesture: Select a subset of the multiple key points corresponding to the specific gesture; Determine a specific position relative to a subset of the plurality of key points, wherein the specific position is determined based on the subset of the plurality of key points and the specific gesture; Register the interaction point to the specific location; Register the nearest point to a location along the user's body; Projecting light rays from the neighboring point through the interaction point; and The light source forms a multi-DOF controller for interacting with the virtual object. Determine angles associated with at least three of the plurality of key points, wherein the angles lie between (i) a first line connecting the tip of the thumb to the midpoint between the two key points on the thumb and index finger, and (ii) a second line connecting the tip of the index finger to the midpoint between the two key points on the thumb and index finger; and Based on the change of the angle from outside the specific range to within the specific range, the motion events associated with the multi-DOF controller are detected.

9. The system according to claim 8, wherein, The plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.

10. The system according to claim 8, wherein, Select a subset of the plurality of key points from a plurality of subsets of the plurality of key points, wherein each subset of the plurality of key points corresponds to a different gesture among the plurality of gestures.

11. The system according to claim 8, wherein, The operation also includes: A graphical representation of the multi-DOF controller is shown.

12. The system according to claim 8, wherein, The location to which the neighboring point is registered is located at the estimated position of the user's shoulder, the estimated position of the user's elbow, or between the estimated position of the user's shoulder and the estimated position of the user's elbow.

13. The system according to claim 8, wherein, The operation also includes: The image of the user's hand is captured by one or more of the image capturing devices.

14. The system according to claim 13, wherein, The image capture device is mounted on the headset of the wearable system.

15. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method of interacting with a virtual object, the method comprising the following operations: Receive images of the user's hand from one or more image capture devices of the wearable system; The image is analyzed to detect multiple key points associated with the user's hand; Based on the analysis of the image, it is determined whether the user's hand is making or transitioning to make a specific gesture among multiple gestures; In response to determining that the user's hand is making or is changing into making the specific gesture: Select a subset of the multiple key points corresponding to the specific gesture; Determine a specific position relative to a subset of the plurality of key points, wherein the specific position is determined based on the subset of the plurality of key points and the specific gesture; Register the interaction point to the specific location; Register the nearest point to a location along the user's body; Projecting light rays from the neighboring point through the interaction point; and A multi-DOF controller for interacting with the virtual object is formed based on the light rays; Determine angles associated with at least three of the plurality of key points, wherein the angles lie between (i) a first line connecting the tip of the thumb to the midpoint between the two key points on the thumb and index finger, and (ii) a second line connecting the tip of the index finger to the midpoint between the two key points on the thumb and index finger; and Based on the change of the angle from outside the specific range to within the specific range, the motion events associated with the multi-DOF controller are detected.

16. The non-transitory machine-readable medium according to claim 15, wherein, The plurality of gestures includes at least one of a grasping gesture, a pointing gesture, or a pinching gesture.

17. The non-transitory machine-readable medium according to claim 15, wherein, Select a subset of the plurality of key points from a plurality of subsets of the plurality of key points, wherein each subset of the plurality of key points corresponds to a different gesture among the plurality of gestures.

18. The non-transitory machine-readable medium according to claim 15, wherein, The operation also includes: A graphical representation of the multi-DOF controller is shown.

19. The non-transitory machine-readable medium according to claim 15, wherein, The location to which the neighboring point is registered is located at the estimated position of the user's shoulder, the estimated position of the user's elbow, or between the estimated position of the user's shoulder and the estimated position of the user's elbow.

20. The non-transitory machine-readable medium according to claim 15, wherein, The operation also includes: The image of the user's hand is captured by one or more of the image capturing devices.

Citation Information

Patent Citations

  • Transmodal input fusion for a wearable system

    US20190362557A1

  • Motion detection system

    JP2018169720A

  • Information processing device, information processing method, and program

    WO2019163372A1