Hand posture and hand gesture determination system
The VR/AR/MR system efficiently determines hand postures and gestures using wearable input devices, reducing computational load and enhancing interaction efficiency.
Patent Information
- Application Number
- PCT/IB2025/054186
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Hand posture and gesture determination systems in VR/AR/MR systems consume significant computational resources, limiting the performance of these systems.
A VR/AR/MR system with wearable I/O devices, including depth and visual input devices, and a processor that acquires and processes datasets to identify hand postures and gestures, generating visual notifications in response to event triggering gestures.
Enhances interaction efficiency by reducing computational demands and improving immersive user interaction without additional equipment.
Smart Images

Figure IB2025054186_30102025_PF_FP_ABST
Abstract
Description
HAND POSTURE AND HAND GESTURE DETERMINATION SYSTEM
[0001] This application claims priority to US provisional patent application serial No. 63 / 638217 (filed Apr. 24, 2024), all of which is incorporated herein by reference.FIELD
[0002] The specification relates generally to Virtual Reality, Augmented Reality, or Mixed Reality (VR / AR / MR) systems, and specifically to VR / AR / MR systems using input devices to determine hand postures and hand gestures.BACKGROUND
[0003] Virtual Reality (VR) refers to a computer-generated simulated three- dimensional environment that can be interacted with in a seemingly real or physical way by a user using electronic equipment such as, for example, a VR headset comprising input and output devices. Augmented Reality (AR) refers to overlaying digital information or virtual objects onto a real-world environment, that can be viewed through electronic devices such as smartphones, smart glasses, AR headsets, etc. Mixed Reality (MR) refers to a continuum that spans between VR and AR, allowing users to interact with both physical and digital elements. Hand posture and hand gesture determination systems can aid a user to interact with a VR / AR / MR system without using additional equipment such as gaming controllers, keyboards, etc., thereby enhancing an immersive quality of the interaction with the VR / AR / MR system. However, hand posture and hand gesture determination systems demand a significant percentage of the computational resources of commercial VR / AR / MR systems, severely limiting the performance of these VR / AR / MR systems.SUMMARY
[0004] An aspect of the specification provides a Virtual Reality, Augmented Reality, or Mixed Reality (VR / AR / MR) system comprising: an Input / Output (I / O) device wearable on a head of a user including: at least one of a depth input device and a visual input devicemounted on the I / O device, the at least one of the depth input device and the visual input device configured to sense data of the user’s hands when the user’s hands are within a volume within a field-of-view of the at least one of the depth input device and the visual input device; and a visual output device mounted on the I / O device; and a processor configured to: acquire a first dataset from the at least one of the depth input device and the visual input device; identify a first posture of at least one of the user’s hands based on the first dataset; acquire a second dataset from the at least one of the depth input device and the visual input device; identify a second posture based on the second dataset; determine whether the first and second postures correspond to an event triggering gesture; and in response to the determination of the event triggering gesture, generate a visual notification in a Graphic User Interface (GUI) displayed in a display of the visual output device.
[0005] Another aspect of the specification provides a method for input detection in a VR / AR / MR system by a processor, the method comprising: acquiring a first dataset from at least one of a depth input device and a visual input device mounted on an I / O device, the I / O device being wearable on a head of a user; determining whether a user hand is within a volume within a field-of-view of the at least one of the depth input device and the visual input device based on the first dataset; identifying a first posture of the user hand based on a portion of the first dataset; acquiring a second dataset from the at least one of the depth input device and the visual input device; identifying a second posture based on a portion of the second dataset; determining whether the first and second postures correspond to an event triggering gesture; and in response to the determination of the event triggering gesture, generating a visual notification in a GUI displayed in a display of the visual output device.BRIEF DESCRIPTIONS OF THE DRAWINGS
[0006] Embodiments are described with reference to the following figures.
[0007] FIG. 1 depicts a block diagram of example internal components of an example VR / AR / MR system.
[0008] FIG. 2 depicts a schematic diagram of a user interacting with the example VR / AR / MR system.
[0009] FIG. 3 shows a flowchart depicting an example method of detecting postures and gestures of an object by a processor of the example VR / AR / MR system.
[0010] FIG. 4 shows a flowchart depicting an example method of identifying postures of an object by the processor of the example VR / AR / MR system.
[0011] FIG. 5 depicts a diagram of example depth data points measured by a depth input device of the example VR / AR / MR system.
[0012] FIGS. 6 and 7 depict diagrams of example silhouettes of an object detected within a volume by an input device of the example VR / AR / MR system.
[0013] FIG. 8 depicts a diagram of example depth data within the example silhouette depicted in FIG. 7.
[0014] FIG. 9 depicts a diagram of example visual data within the example silhouette depicted in FIG. 7.
[0015] FIG. 10 shows a flowchart depicting an example method of identifying subsequent postures of an object by the processor of the example VR / AR / MR system.
[0016] FIGS. 1 1 and 12 depict further schematic diagrams of a user interacting with the example VR / AR / MR system.
[0017] FIG. 13 depicts a schematic diagram of an example Graphic User Interface (GUI) displayed in an output device of the example VR / AR / MR system.
[0018] FIG. 14 depicts schematic diagrams of virtual representations of example initial and final hand postures of a squeeze gesture.
[0019] FIG. 15 depicts schematic diagrams of virtual representations of example initial and final hand postures of a fingertip pinch gesture.
[0020] FIG. 16 depicts schematic diagrams of virtual representations of example initial, intermediate, and final hand postures of a fingertip forward slide gesture.
[0021] FIG. 17 depicts schematic diagrams of virtual representations of example initial, intermediate, and final hand postures of a fingertip backward slide gesture.
[0022] FIG. 18 depicts a schematic diagram of another example GUI displayed in an output device of the example VR / AR / MR system.
[0023] FIG. 19 shows a flowchart depicting an example setup method for detecting postures and gestures of an object by the processor of the example VR / AR / MR system.
[0024] FIG. 20 depicts another schematic diagram of a user interacting with the example VR / AR / MR system.
[0025] FIG. 21 depicts a schematic diagram of an example surface divided into a grid.
[0026] FIG. 22 depicts schematic diagrams of virtual representations of example initial and final hand postures of a tapping gesture.
[0027] FIG. 23 depicts schematic diagrams of virtual representations of example initial and final hand postures of a gliding gesture.DETAILED DESCRIPTION
[0028] FIG. 1 shows a schematic diagram of a non-limiting example of a VR / AR / MR system 100. The system 100 comprises a computing device 104 such as a general- purpose computer, or a dedicated-purpose computer such as a gaming computer, for example.
[0029] The computing device 104 comprises a processor 108 that may be implemented as a plurality of processors or one or more multi-core processors. The processor 108 may include one or more of a Central Processing Unit (CPU), a microcontroller, a microprocessor, a processing core, a Field-Programmable Gate Array (FPGA) or the like, and combinations thereof. The processor 108 may additionally include a built-in Graphics Processing Unit (GPU), however, in other embodiments a GPU may be provided separately from the processor 108.
[0030] The computing device 104 further includes one or more memory units, including a volatile memory 112 and a non-volatile memory 116, communicatively coupled to the processor 108. The volatile memory 112 is based on any random-access memory(RAM) technology. For example, the volatile memory 112 can be based on a Double Data Rate (DDR) Synchronous Dynamic Random-Access Memory (SDRAM). Other types of volatile memory 112 are contemplated.
[0031] The non-volatile memory 116 can be based on any persistent memory technology, such as an Erasable Electronic Programmable Read Only Memory (“EEPROM”), flash memory, solid-state hard disk (SSD), other type of hard-disk, or combinations of them. The non-volatile memory 116 may also be described as a non- transitory computer readable media. Also, more than one type of non-volatile memory 116 may be provided.
[0032] Programming instructions in the form of applications 120 are typically maintained, persistently, in non-volatile memory 116 and used by the processor 108 which reads from and writes to volatile memory 112 during the execution of applications 120. Various methods discussed herein can be coded as one or more applications 120. One or more tables or databases 124 are maintained in non-volatile memory 1 16 for use by applications 120. Note that applications 120 are individually labelled as 120-1 ... 120- m. Collectively, they are referred to as applications 120, and generically, as application 120. The nomenclature is used elsewhere such as for databases 124.
[0033] The system 100 can be connected to a network 128 such as the Internet. The network 128 may interconnect the system 100 with one or more additional VR / AR / MR systems and / or non-VR / AR / MR systems.
[0034] The computing device 104 can further include a network interface 232 through which the processor 108 is communicatively connected to the network 128. The network interface 232 may further connect an additional device communicatively coupled to the computing device 104 to the network 128. In other embodiments, the network interface 232 may be provided separate from the computing device 104.
[0035] The system 100 further includes Input / Output (I / O) devices 136 such as a VR Headset I / O device 136-1 but may further comprise other kinds of I / O devices such as, for example, a VR glove. The VR Headset I / O device 136-1 includes input devices 140 including a visual input device 140-1 such as a visible light camera, an eye-tracking camera, an IR light camera, etc., a depth input device 140-2 such as a time-of-flight inputdevice, a structured light input device, a stereoscopic vision input device, a laser range finder, an ultrasonic input device, etc., a motion input device 140-3 such as an accelerometer, a gyroscope, etc., an audio input device 140-4 such as a microphone, a musical instrument input device, etc. (Note that in variants, one or more stereoscopic cameras, one or more light field range cameras, etc. for the visual input device 140-1 may also be used for depth sensing, in lieu of or supplementary to the depth input device 140- 2). The VR Headset I / O device 136-1 further includes output devices 144 including a visual output device 144-1 such as a VR display, an AR display, a set of Light Emitting Diodes (LEDs), etc., an audio output device 144-2 such as a speaker, a set of headphones, etc., a haptic output device 144-3 such as a vibrotactile actuator, etc. The input and output devices 140 and 144 may be preferably physically mounted on the I / O devices 136, such as, for example, along a frame of the VR Headset I / O device 136-1 and are communicatively connected to the processor 108 via wired or wireless connections. The I / O devices 136 may include two or more input or output devices 140 or 144 of a same kind as well as other kinds of input and output devices 140 and 144 not listed above. In other embodiments the I / O devices 136 may additionally include their own processors such as, for example, a dedicated GPU.
[0036] FIG. 2 shows a top plan view of a user 200 interacting with the system 100 where the user 200 is wearing the VR Headset I / O device 136-1. FIG. 2 also shows a volume 204 within a field-of-view of one or more of the input devices 140 of the VR Headset I / O device 136-1 , for example, within a field-of view of the visual input device 140-1 or of the depth input device 140-2, so that the one or more of the input devices 140 may detect a presence of physical objects 208 such as, for example, user hands 208-1 or 208-2 when the physical objects 208 are within the volume 204. The physical objects 208 may further include other physical objects such as, for example, a pen, a sheet of paper, a surface, etc., or a combination of physical objects, such as, for example, a hand holding a pen, a hand holding a sheet of paper, a hand resting on a surface, etc.
[0037] By analyzing input from the input devices 140, the processor 108 can determine at least one of a type of physical object 208 within the volume 204, a three-dimensional location of the physical object 208 within the volume 204 and an orientation of the physical object 208 with respect to a point of reference. Furthermore, wherein the physical object208 is at least one of a flexible and an articulated physical object 208, for example, the user hand 208-1 , the processor 108 may additionally determine a pose or posture of the physical object, for example, a “thumbs-up” posture where a thumb of the user hand 208- 1 is extended away from a palm of the user hand 208-1 and other fingers of the user hand 208-1 are closed towards the palm. Additionally, by tracking changes in at least one of the three-dimensional location, the orientation and the posture of the physical object 208, the processor 108 can additionally determine a gesture of the physical object 208. Furthermore, the processor 108 may be configured to trigger events of an application 120 in response to specific postures and / or gestures of the physical object 208, for example, running a specific function or sub-routine of the application 120. The processor 108 may be further configured to generate an indication in response to the triggered events, for example, a visual, an audio, or a haptic indication, that an output device 144 may output to the user 200. Additionally, the processor 108 may be further configured to generate a virtual representation of the physical object 208, for example, a mesh or a skeleton representation, such as a virtual skeleton hand, and display the virtual representation to the user 200, for example through the visual output device 144-1. FIG 3 depicts a flowchart of an example method for the determination of a gesture of a physical object 208 and the triggering of events responsive to the gesture determination by the processor 108.
[0038] Starting at block 305, the processor 108 identifies the posture of the physical object 208. FIG. 4 depicts a flowchart of an example method 400 by which the processor 108 may identify postures of the physical object 208.
[0039] Starting at block 405, the processor 108 acquires data from at least one input device 140, for example, from the visual input device 140-1 or from the depth input device 140-2.
[0040] After acquiring the data, the processor 108 proceeds to block 410, where the processor 108 analyzes the data and determines whether a physical object 208 is within the volume 204. If the system is configured so that the processor 108 may acquire data from more than one input device 140, the processor 108 may be configured to select a first type of data as primary data and a second type of data as secondary data for thedetermination, for example, based on detection / reaction speed, power consumption, the computational resources required for the determination, or combinations thereof. If, for example, the processor 108 selects depth data as the primary data, the processor 108 may compare the depth data to a threshold value to determine that a physical object 208 is within the volume 204. Alternatively, the processor 108 may select the visual data as the primary data and the processor 108 may compare changes in current color or light values of the visual data with respect to previous color or light values of the visual data to threshold change values to determine the presence of the physical object 208 within the volume 204. If the primary data is not available, the processor 108 may use the secondary data for the determination. Furthermore, the processor 108 may use both the primary and the secondary data for a more robust determination.
[0041] If the determination at block 410 is negative, the processor 108 returns to block405 for subsequent data acquisition and iteration through the method 400.
[0042] If the determination at block 410 is affirmative, the processor 108 proceeds to block 415 where the processor 108 identifies a silhouette of the physical object 208 within the volume 204 based on the acquired data at block 405. The processor 108 may be configured to select the primary and the secondary data for the silhouette identification. As shown in FIG. 5, if the depth data is selected as the primary data, the processor 108 may identify an object silhouette 500 by identifying, within the volume 204, continuously adjacent depth data points 504x,yalong a plane (X ,Y) orthogonal to a depth direction (Z) with respect to the orientation of the depth input device 140-2, the adjacent depth data points 504x,y having depth data values within a predetermined depth range, for example, of about ±20 cm. Alternatively, if, for example, the identification of the object silhouette 500 may be more efficiently performed by using the visual data, the processor 108 may select the visual data as the primary data at this stage, and, for example, identify continuously adjacent pixels from the visual data based on, for example, specific pixel light or color value ranges, differences in pixel light or color values from neighboring pixels, etc. Similarly to block 410, if the primary data is not available the processor 108 may use the secondary data for the identification. Furthermore, the processor 108 may use both the primary and the secondary data for a more robust identification of the object silhouette 500.
[0043] Returning to FIG. 4, after the identification at block 415, the processor 108 proceeds to block 420, where the processor 108 inputs data such as the shape of the object silhouette 500 and its position within the volume 204 to a first deep learning model, for example a neural network such as a convolutional neural network or a residual neural network that is configured to identify at least one type of physical object 208 and a first set of features of the at least one type of physical object 208 based on the object silhouette data. If, for example, if the first learning model is configured to detect human hands such as the user hands 208-1 and 208-2, the first deep learning model may output a series of flags, including a flag indicating that the detected physical object 208 within the volume 204 is a human hand, a flag indicating whether the hand is a left or a right hand, a series of flags indicating whether each of the hand’s fingers is visually discernable from the object silhouette 500, a series of flags indicating whether each of the fingers are fully extended, partially extended, etc. Additionally, the first deep learning model may output a three-dimensional direction that the palm of the hand is facing, a three- dimensional direction that each of the fingertips of the fingers are facing, etc. Additionally, the first deep learning model may further output a confidence percentage value for each of its outputs. FIG. 6 depicts a first example object silhouette 600 within the volume 204 representative of a left hand with the palm extending upwards and with all of its fingers fully extended and separated from each other. Based on the shape of the first example object silhouette 600 and on its position within the volume 204, the first deep learning model may output with high confidence percentages (for example > 80%), for example, that the object is a left hand, that all of its fingers are discernable and fully extended, a position of a center of gravity of the palm with respect to the volume 204, positions of ends of all of the fingertips with respect to the volume 204, and the directions that the palm and each of the fingertips are facing. FIG. 7 depicts a second example object silhouette 700 within the volume 204 representative of a left hand with the palm extending upwards and with the thumb fully extended and separated from the index finger, and with the index, middle, ring, and little finger fully closed towards the palm of the hand. Based on the shape of the second example object silhouette 700 and on its position within the volume 204, the first deep learning model may output with high confidence percentages (for example > 80%), for example, that the object is a left hand, that the thumb isdiscernable and fully extended, and the direction of the palm and of the thumb fingertip. The first deep learning model may be trained, for example, with a library of previously extracted silhouettes from depth and visual data representative of a variety of hands from different users of different ages, heights, body mass indices, etc., in different postures.
[0044] Returning to FIG. 4, after the identification at block 420, the processor 108 proceeds to block 425, where a determination is made of whether the first set of features are sufficient for an identification of the posture of the physical object 208. This determination may be made, for example, based on whether all of the confidence percentage values are above a first confidence percentage threshold, whether at least a number of confidence percentage values are above a second confidence percentage threshold, etc.
[0045] If the determination at block 425 is affirmative, the processor 108 may proceed to block 430 where the processor 108 may identify the posture of the physical object 208 based on the first set of features, by for example, using a postures reference library included in a database 124, such as the database 124-1 , to match the first set of features to a posture. Alternatively, the postures reference library may be included in a remote database, for example in the network 128, that the processor 108 may access through the network interface 132.
[0046] If the determination at block 425 is negative, the processor 108 proceeds to block 435, where the processor may input the first set of features and at least one of the visual data and the depth data within the object silhouette 500 to a second deep learning model, for example a neural network such as a convolutional neural network or a residual neural network that is configured to identify a second set of features of the at least one type of physical object 208, the second set of features being supplementary to the first set of features. For example, the second deep learning model may output flags indicating whether each of the fingers is fully closed towards the palm, a flag confirming whether the hand is a left or a right hand, the position of the center of gravity of the palm with respect to the volume 204, the positions of the ends of all of the fingertips with respect to the volume 204, and the directions that the palm and each of the fingertips are facing. FIG. 8 depicts example depth data 800 within the second example object silhouette 700 thatmay be input to the second deep learning model. FIG. 9 depicts example visual data 900 within the second example object silhouette 700 that may be input to the second deep learning model. Based on the example depth data 800 or the example visual data 900, the second deep learning model may confirm that the hand is indeed a left hand, identify that the index, middle, ring, and little finger are fully closed towards the palm, and may additionally identify or confirm the direction of all of the fingers. The second deep learning model may use the first set of features to at least one of reduce the number of possible values or value ranges of its outputs and increase their corresponding confidence percentage values. Additionally, if both visual and depth data are available, the second deep learning model may further use the first set of features to determine whether to use the visual or the depth data, or a combination of the visual and the depth data to perform the identification of the second set of features more efficiently or more accurately, based on the outputs and respective confidence values associated to the first set of features. The second deep learning model may be trained, for example, with a library of previously extracted depth and visual data within the silhouettes used to train the first deep learning model.
[0047] After the identification of the second set of features at block 435, the processor 108 proceeds to block 430 where the processor 108 may identify the posture of the physical object 208 based on the first and the second set of features, by for example, using the postures reference library to match the first and the second set of features to a posture.
[0048] The method 400 is just an example method by which the processor 108 may identify postures of the physical object 208. Different methods may be implemented for said identification. For example, the processor 108 may continuously feed raw depth or visual data to an alternative deep learning model that, for example may output the most likely positions of a set of joints of the physical object 208, and the processor 108 may convert the positions of the set of joints into the posture of the physical object 208 using, for example, inverse kinematics principles based on typical positions of joints of the physical object 208, such as the physical joints in a human hand, and their corresponding restrictions such as their possible size ranges. However, by performing two different feature identification blocks such as the block 420 and the block 435, the processor 108may perform the identification of the posture of the physical object 208 more efficiently. Additionally, while all of the blocks of the example method 400 have been described as performed by the processor 108, any of the blocks may be alternatively performed by a different computing device to which the processor 108 may be communicatively connected to.
[0049] Returning to FIG. 3, once the processor 108 has identified the posture of the physical object 208, the processor 108 may proceed to block 310, where a determination is made of whether the identified posture matches an event triggering posture from a triggering posture reference library contained in a database 124, for example in the database 124-1 , Alternatively, the triggering posture reference library may be located in a remote database, for example in the network 128, that the processor 108 may access through the network interface 132.
[0050] If the determination at block 310 is affirmative, the processor 108 proceeds to block 335, where the processor triggers an event associated with the identified matching posture. The processor 108 may then return to block 305 for a subsequent identification of a posture and iteration through the method 300.
[0051] If the determination at block 310 is negative the processor 108 proceeds to block 315, where the processor 108 can access a triggering gesture reference library containing gestures, each of the gestures composed by at least two gesture postures ordered in a gesture sequence. The triggering gesture reference library may be further included in the database 124-1 , in a different database 124, or in a remote database. At this point, the processor 108 can compare the identified posture to the gesture postures to determine whether the identified posture matches at least one of the gesture postures.
[0052] If the determination at block 315 is negative, the processor 108 may then return to block 305 for a subsequent identification of a posture and iteration through the method 300.
[0053] If the determination at block 315 is affirmative, the processor 108 proceeds to block 320 where the processor may identify the next (subsequent) posture of the physical object 208. FIG. 10 depicts a flowchart of an example method 1000 by which the processor 108 may identify the next posture of the physical object 208.
[0054] Starting at block 1005, the processor 108 may acquire new depth or visual data from at least one input device 140. The new depth or visual data may be acquired at a predetermined rate. The rate may be fixed or variable, for example, based on rate requirements associated with an application that the system 100 is performing, such as, for example, a typing application, a video game application, a web browsing application, etc. The rate of new depth data may be different to the rate of new visual data. The rate may be further determined based on a previously determined gesture speed score associated to the user 200.
[0055] After the new data has been acquired at block 1005, the processor 108 proceeds to block 1010, where a determination is made whether at least one significant change has happened based on a comparison of the new data to previously acquired data; for example, by comparing changes in data values to different thresholds, for example, a difference in depth values of at least 5% for at least 10% of the depth values.
[0056] If the determination at block 1010 is negative, the processor 108 returns to block 1005 for a subsequent new data acquisition and iteration through the method 1000. The rate of acquisition of new depth or visual data may be adjusted, for example, lowered, based on the negative determination at block 1010, on a number of successive negative determinations at block 1010 reaching a predetermined threshold, etc.
[0057] If the determination at block 1010 is affirmative, the processor 108 may proceed to block 1015, where the processor may input information pertaining to the previous posture such as an identifier of the posture and the first and the second set of features as well as at least a portion of the new data, such as the values of the new data that have significantly changed with respect to values of the previously acquired data, the change in the values of the new data with respect to the values of the previously acquired data, etc. into a third deep learning model, for example a neural network such as a convolutional neural network or a residual neural network that is configured to identify a new first and second set of features of the at least one type of physical object 208, The third deep learning model may use the information pertaining to the previous posture to at least one of reduce the number of possible values or value ranges of its outputs and increase their corresponding confidence percentages. Additionally, if the new data comprises bothvisual and depth data, the third deep learning model may further use the information pertaining to the previous posture to determine whether to use the new visual or the new depth data, or a combination of the new visual and the new depth data to perform the identification of the new first and second set of features more efficiently or more accurately. The third deep learning model may be trained, for example, with a library of gestures composed of sequences of postures. With the new first and second set of features, the processor 108 may identify the next posture of the physical object 208 by for example, using the postures reference library to match the new first and second set of features to a posture. The library of gestures may be specific to the application that the system 100 is performing, for example, a typing application, a game application, a web browsing application, etc. Furthermore, the rate of acquisition of new depth or visual data may be adjusted, for example, raised, based on the positive determination at block 1010, on a number of successive positive determinations at block 1010 reaching a predetermined threshold, etc.
[0058] The method 1000 is just an example method by which the processor 108 may identify subsequent postures of the physical object 208. Different methods may be implemented for said identification, for example, by performing a combination of the method 1000 and the method 400 that implements blocks 1005 and 1010, and then performs the blocks 410 to 440 when the determination at block 1010 is affirmative. However; by using the information pertaining to the previous posture to identify the next posture and by adjusting the rate of acquisition of new data for subsequent posture determinations, the processor 108 may perform the identification of the next posture more efficiently. Additionally, while all of the blocks of the example method 1000 have been described as performed by the processor 108, any of the blocks may be alternatively performed by a different computing device to which the processor 108 may be communicatively connected to.
[0059] Returning to FIG. 3, after the processor 108 has identified the next posture of the physical object 208 at block 320, the processor proceeds to block 325 where the processor 108 determines whether a sequence of the identified postures match a gesture sequence in the triggering gesture reference library.
[0060] If the determination at block 325 is negative, the processor 108 returns to block 305 for a subsequent identification of a posture and iteration through the method 300.
[0061] If the determination at block 325 is affirmative, the processor 108 proceeds to block 330, where a determination is made whether the most recently identified posture matches a final posture in the matching gesture sequence.
[0062] If the determination at block 330 is negative, the processor 108 returns to block 320 for a subsequent determination of the next posture of the physical object 208.
[0063] If the determination is affirmative, the processor 108 proceeds to block 335 where the processor 108 triggers an event associated with the matching final posture. The processor 108 may then return to block 305 for a subsequent identification of a posture and iteration through the method 300.
[0064] The method 300 is just an example method by which the processor 108 may track postures of the physical object 208 and determine event triggering gestures. Different methods may be implemented for said event triggering gesture determination. For example, the processor 108 may continuously identify postures of the physical object 208 and, for example, transmit the identified postures and a time tag to a different subprocess also performed by the processor 108 or by a different computing device to which the processor 108 may be communicatively connected to. The different sub-process may then determine whether a subset of the postures identified within a moving time window and their sequence correspond to an event triggering gesture.
[0065] Identification of certain hand postures demands more computational resources from the processor 108 than identification of other hand postures. For example, with reference to the example method 400, identification of hand postures that is based solely on the first set of features is more computationally efficient than identification of hand postures based on the first and the second set of features. Furthermore, identification of hand postures based on the first and the second set of features where the first set of features significantly reduces the possible values or value ranges of the second set of features may be more computationally efficient than identification of hand postures based on the first and the second set of features where the first set of features does not significantly reduce the possible values or value ranges of the second set of features.Similarly, with reference to the example method 1000, identification of subsequent postures where information pertaining to a previous posture significantly reduces the possible new values or new value ranges of a new set of features may be more computationally efficient than identification of subsequent postures that does not use information pertaining to the previous posture. Furthermore, with reference to the example method 300, event triggering determination that requires the identification of a single hand posture is more computationally efficient than event triggering determination that requires the identification of a gesture comprised of two hand postures or more in a sequence.
[0066] The location and orientation of the input devices 140 with respect to the hands 208-1 and 208-2 of the user 200 can also play a significant role in the computational resources needed by the processor 108 to identify the hand postures and gestures, as well as in confidence percentages of said hand posture and gesture identification. For example, with reference to FIGS. 11 and 12, the user 200 is depicted wearing the VR headset I / O device 136-1 and performing example hand postures 1100 and 1200. While the position of the fingers with respect to the palm in both of the example hand postures 1100 and 1200 may be the same, the processor 108 may perform a more efficient identification of the example hand posture 1100 than of the example hand posture 1200, as the orientation of the palm in the example hand posture 1200 results in at least a partial occlusion of the ring and the little fingers from the fields-of-view of input devices 140-1 and 140-2, and therefore, in order to identify their position with a sufficient degree of confidence based on data from the input devices 140-1 and 140-2, the processor 108 may require extra processing steps such as, for example, comparing the hand posture 1200 to previous hand postures.
[0067] As previously discussed, the processor 108 may optionally generate a virtual representation of the physical object 208, for example, a mesh or a skeleton representation, such as a virtual skeleton hand, and display the virtual representation to the user 200, for example through the visual output device 144-1. By controlling the position of the visual representation and its orientation with respect to other visual elements displayed by the visual output device 144-1 , the processor 108 can provide visual feedback to the user 200 and can further direct the user 200 to perform a specificset of postures and gestures, such as, for example, a set of postures and gestures with a high computational efficiency score so that the controller 108 may dedicate a smaller amount of its resources to a posture and gesture identification task and a larger amount of its resources to other tasks. For example, FIG. 13, depicts an example Graphic User Interface (GUI) 1300 displayed by a display 1304 of the visual output device 144-1 , the GUI presenting a virtual environment rendered within a field-of-view that mimics a natural field of vision of the user 200. Virtual hands 1308-1 and 1308-2, corresponding to the user hands 208-1 and 208-2, respectively, are displayed in the GUI 1300. Virtual interaction objects 1312-1 and 1312-2 are further displayed in the GUI 1300. The virtual interaction objects 1312-1 and 1312-2 are positioned been the virtual hands 1308-1 and 1308-2, respectively, and a starting point (i.e. the center) of the field-of-view of the user 200 wearing the VR headset I / O device 136-1 so that each of the virtual interaction objects 1312 may appear at least partially overlayed on each of the virtual hands 1308, thereby directing (i.e. limiting) the user 200 to perform a specific set of event triggering postures and gestures with the user hands 208-1 and 208-2 facing the input devices 140-1 and 140-2 to interact with the virtual interaction objects 1312. The processor 108 may further continuously update the position and orientation of the virtual interaction objects 1312 with respect to the virtual hands 1308 to redirect the orientation of the user hands 208-1 and 208-2 with respect of the field-of-view of the user 200 wearing the VR headset I / O device 136-1 to maintain a minimal occlusion of elements and distinguishing features of the user hands 208-1 and 208-2 from the input devices 140-1 and 140-2. When the processor 108 identifies an event triggering gesture, the processor 108 may change the size, shape, or color of one or both of the virtual interaction objects 1312 to provide the user 200 with an indication of the event. Alternatively, the processor 108 may display at least one of a virtual hand 1308 and an interaction object 1312 to provide visual feedback to the user 200. As a further alternative, the processor 108 need not display the virtual interaction objects 1312 or the virtual hands 1308 and may simply trigger events when the processor 108 identifies the specific set of postures and gestures.
[0068] Additional details regarding example specific postures and gestures are described below with reference to FIGS. 14 to 17.
[0069] Squeeze Gesture
[0070] FIG 14. depicts example initial and final virtual representations 1404 and 1408, respectively, of possible initial and final hand postures of a squeeze gesture 1400. FIG. 14 further depicts an example virtual object in a relaxed state 1412 and a stressed state 1416, respectively, the example virtual object to be squeezed upon identification of a possible final hand posture of the squeeze gesture 1400. The squeeze gesture 1400 is composed of at least one possible initial and at least one possible final hand posture and may optionally further be composed of at least one possible intermediate hand posture. The at least one possible initial hand posture indicates a hand with the index, middle, ring, and little fingers extended away from a palm of the hand, and the at least one possible final hand posture indicates the hand with the hand with the index, middle, ring, and little fingers closed toward the palm of the hand. Additionally, the at least one possible initial and final hand postures may further indicate the palm being oriented towards a face of the user.
[0071] Fingertip Pinch Gesture
[0072] FIG. 15 depicts example initial and final virtual representations 1504 and 1508, respectively, of possible initial and final hand postures of a fingertip pinch gesture 1500. FIG. 15 further depicts example virtual objects in a relaxed state 1512 and a stressed state 1516, respectively, the example virtual objects to be pinched upon identification of a possible final hand posture of the fingertip pinch gesture 1500. The fingertip pinch gesture 1500 is composed of at least one possible initial and at least one possible final hand posture and may optionally further be composed of at least one possible intermediate hand posture. The at least one possible initial hand posture indicates a hand with a thumb fingertip extending away from the index, middle, ring, and little fingers, the index, middle, ring and little fingers extending away from the palm of the hand, and the at least one possible final hand posture indicates the thumb fingertip meeting one of an index fingertip, a middle fingertip, a ring fingertip, or a little fingertip. Additionally, the at least one possible initial and final hand postures may further indicate the palm being oriented towards the face of the user.
[0073] Fingertip Forward Slide Gesture
[0074] FIG. 16 depicts example initial, intermediate, and final virtual representations 1604, 1608 and 1612, respectively, of possible initial, intermediate, and final hand postures of a fingertip forward slide gesture 1600. FIG. 16 further depicts example virtual objects in an initial, intermediate, and final state 1616, 1620 and 1624, respectively, the example virtual objects to be slid upon identification of a possible intermediate or a possible final hand posture of the fingertip forward slide gesture 1600. The fingertip forward slide gesture 1600 is composed of at least one possible initial, at least one possible intermediate, and at least one possible final hand posture. The at least one possible initial hand posture indicates a hand with the thumb fingertip extending away from the index, middle, ring, and little fingers, the index, middle, ring, and little fingers extending away from the palm, the at least one possible intermediate hand posture indicates the thumb fingertip meeting the index fingertip, the index and the middle fingers being adjacent to each other, and the at least one possible final hand posture indicates the thumb fingertip meeting the middle fingertip, the index and the middle fingers remaining adjacent to each other. Additionally, the at least one possible initial, intermediate and final hand postures may further indicate the palm being oriented towards the face of the user.
[0075] Fingertip Backward Slide Gesture
[0076] FIG. 17 depicts example initial, intermediate, and final virtual representations 1704, 1708 and 1712, respectively, of possible initial, intermediate, and final hand postures of a fingertip backward slide gesture 1700. FIG. 17 further depicts example virtual objects in an initial, intermediate, and final state 1716, 1720 and 1724, respectively, the example virtual objects to be slid upon identification of a possible intermediate or a possible final hand posture of the fingertip backward slide gesture 1700. The fingertip backward slide gesture 1700 is composed of at least one possible initial, at least one possible intermediate, and at least one possible final hand posture. The at least one possible initial hand posture indicates a hand with the thumb fingertip extending away from the index, middle, ring, and little fingers, the index, middle, ring, and little fingers extending away from the palm, the at least one possible intermediate hand posture indicates the thumb fingertip meeting the middle fingertip, the index and the middle fingers being adjacent to each other, and the at least one possible final hand posture indicatesthe thumb fingertip meeting the index fingertip, the index and the middle fingers remaining adjacent to each other. Additionally, the at least one possible initial, intermediate, and final hand postures may further indicate the palm being oriented towards the face of the user.
[0077] The processor 108 may trigger different events such as, for example, an item selection event, a zoom-in or zoom-out event, a menu display event, a menu scroll event, etc. upon detection of the at least one possible final hand postures of the squeeze, the fingertip pinch, the fingertip forward slide, or the fingertip backward slide gestures 1400, 1500, 1600 or 1700, respectively. By triggering the different events, the processor 108 may run special functions or subroutines, generate event notifications such as visual, audio, or haptic notifications, etc.
[0078] While a VR / AR / MR system, such as the example system 100, may computationally efficiently determine hand postures and gestures where the palms of the hands are oriented towards the center or origin of the fields-of-view of the system’s input devices 140, certain applications 120 executed by the processor 108 may benefit from a series of hand postures and / or gestures where the orientation of the hand’s palms is facing away from the center or origin of the fields-of-view of the input devices 140. FIG. 18 depicts an example GUI 1800 displayed the display 1304, where the virtual hands 1308 are shown overlayed on a virtual keyboard 1804 which is depicted as a musical keyboard with virtual keys 1808. The virtual keyboard 1804 can alternatively be, for example, a typewriting keyboard, a gaming keyboard, a trackpad, a sketchpad, a controls console, etc. To play the virtual keys 1808 of the virtual keyboard 1804, the elements displayed by the GUI 1800, specifically, the virtual hands 1308 direct the user 200 to orient the palms of their hands 208 towards a plane within the volume of detection of the input devices 140, the desired orientation of the user hands 208 defined by the position and orientation of a virtual plane 1812 in the GUI 1800 (i.e. along virtual coordinates X’,Y’) on which the virtual keyboard 1804 is rendered with respect to the position and orientation of the virtual hands 1308. Setting up the dimensions of the virtual keyboard 1804 and the position and orientation of the virtual plane 1812 with respect to reference dimensions of, for example a real object, and with respect to the position and orientation of a real reference plane within the field-of-view of the input devices 140 on which the user 200can position their hands 208 to perform keyboard playing postures and gestures to play the virtual keyboard 1804 can be done by a setup method such as example setup method 1900, depicted in FIG. 19.
[0079] In the example setup method 1900, starting at block 1905, a determination is made by the processor 108 of whether a substantially planar surface, for example, a tabletop, is present within the volume 204 within which the visual and / or depth input devices 140-1 and / or 140-2 can detect an object’s three-dimensional position with sufficient precision (for example, within about 5 mm). The determination can include the three-dimensional coordinates of the planar surface within the volume 204, which can include the position and orientation of a plane on which the planar surface is located and the distance from at least one of input devices 140-1 and 140-2 or from the center of the field-of-view of the user 200 to the center of the planar surface. The determination of the planar substance can be made, for example, based on a plane detection algorithm such as a Random Sample Consensus (RANSAC) algorithm which can fit a plane model to data from the visual and / or depth input devices 140-1 and / or 140-2 where if a significant number of neighboring data points (for example, greater than a percentage threshold of the total datapoints available from the visual and / or depth input devices 140-1 and / or 140- 2) align to a plane within a first tolerance threshold, the processor 108 can determine that the planar surface is present. Alternatively, the determination can be made based on a calculation of surface normals for local patches of a depth map where a consistently flat normal across a large region can indicate the planar surface, based on a depth gradient and variation analysis, an edge detection and contour analysis, a Hough transform for determining dominant plane structures in a depth image, a Convolutional Neural Network (CNN) or deep learning model trained to detect a planar surface based on visual and / or depth data of different planar surfaces at different angles and other non-planar objects to classify and segment the planar region from the rest of the measuring data from the visual and / or depth input devices 140-1 and / or 140-2, for example, based on a depth-aware segmentation model, etc.
[0080] If the determination at block 1905 is negative, the processor 108 can proceed to block 1910 at which a further determination is made by the processor 108 of whether the user’s lap is within the volume 204 so that the processor 108 can use the user’s lapas setup reference in lieu of a substantially planar surface in the volume 204. The determination can include the three-dimensional coordinates of the lap surface within the volume 204, which can include the position and orientation of a plane on which the lap surface is located and the distance from at least one of input devices 140-1 and 140-2 or from the center of the field-of-view of the user 200 to the center of the lap surface. The determination of the lap surface can be made, for example, based on a plane detection algorithm such as a RANSAC algorithm where a significant number of neighboring data points align to a plane within a second tolerance threshold for example, where the second tolerance threshold is greater than the first tolerance threshold for the determination at block 1905 to account for greater surface irregularities of the user’s lap than of a planar surface such as a table top. Alternatively, the determination can be made based on a calculation of surface normals for local patches of a depth map where a consistently flat normal across a large region can indicate the user’s lap, based on a depth gradient and variation analysis, an edge detection and contour analysis, a Hough transform for determining dominant plane structures in a depth image, a Convolutional Neural Network (CNN) or deep learning model trained to detect a the user’s lap based on visual and / or depth data of diverse laps of diverse users at different angles and other planar and non- planar objects to classify and segment a user lap region from the rest of the measuring data from the visual and / or depth input devices 140-1 and / or 140-2, for example, based on a depth-aware segmentation model, etc.
[0081] If a determination at block 1910 is negative, the processor 108 can proceed to block 1915, at which an indication is generated by the processor 108 to the user 200 through one or more of the output devices 144 for the user 200 to provide at least two reference points to the processor 108, for example, points 2000-1 and 2000-2 as shown in FIG. 20, via hand postures or gestures such as postures 2004-1 and 2004-2, for the processor 108 to determine the three-dimensional coordinates of an air or imaginary surface such as air surface 2008 within the volume 204, which can include the position and orientation of the plane 2012 on which the air surface 2008 is located and the distance from at least one of the visual input device 140-1 and the depth input device 140-2 to the center 2016 of the air surface 2012.
[0082] With reference to FIG. 19, after the indication at block 1915 is generated, the processor 108 can proceed to block 1920, where a determination is made by the processor 108 of whether the reference points 2000 have been indicated by the user 200 within the volume 204, the determination further including the three-dimensional location of the reference points 2000, for example, with respect to the origin of the user’s field-of- view. The determination can further include the three-dimensional coordinates of the air surface 2008, based on the three-dimensional location of the reference points 2000.
[0083] If the determination at block 1920 is negative, the processor 108 may loop back to block 1915 for maintaining the indication to the user 200 until the user 200 indicates the reference points 2000 to the processor 108.
[0084] If a determination by the processor 108 at block 1905 is positive, a determination by the processor 108 at block 1910 is positive, or if the processor 108 has determined the at least two reference points 2000 based on one or more hand postures or gestures 2004 at block 1920 and the associated three-dimensional coordinates of an air surface 2008, the processor 108 can proceed to block 1925, at which the processor 108 can determine the size of the planar surface, the lap surface or of the air surface 2008, respectively. The determination of the size of the (planar, lap or air) surface can be made by the processor 108, for example, based on a predetermined dimension or set of dimensions corresponding to the distance from at least one of the visual input device 140- 1 and the depth input device 140-2 or the center of the field-of-view of the user 200 to the center of the surface. The size of the surface can alternatively be set by the processor 108 to be equal to the boundaries of the surface, for example, to the perimeter of a rectangle with a diagonal line defined by the distance between the two reference points 2000 or the perimeter (i.e. the boundaries) of the planar or lap surfaces. The size of the surface can alternatively be set by the processor 108 to be equal to a predetermined fraction of the boundaries of the surface.
[0085] After the size of the surface is determined by the processor 108 at block 1925, the processor 108 can proceed to block 1930, at which the processor 108 can generate a grid such as grid 2100 of surface 2012 shown in FIG. 21. The grid 2100 divides the surface 2012 into cells 2104. The cells 2104 define the limits within the surface 2012 fordetection of a specific action by the user 200 interacting with the surface 2012, for example, for detecting that the user 200 is playing a specific virtual key 1808 of the virtual keyboard 1804 if the user 200 touches a specific cell 2104 of the surface 2012, for example with one of his fingers. The number and dimensions of the cells 2104 can be selected by the processor 108, based on the size of the surface 2012, on the number of virtual keys 1808 of the virtual keyboard 1804. and on minimum detection dimensions of the visual and / or depth input devices 140 for the distance from at least one of the visual input device 140-1 and the depth input device 140-2 or the center of the field-of-view of the user 200 to the center of the surface 2012. If the number of virtual keys 1808 for a specific virtual keyboard 1804 is higher than the number of cells 2104 of minimum detection dimensions that fit within the surface 2012, the size of the surface 2012 can be updated (e.g. enlarged) accordingly by the processor 108 if the updated size still falls within the volume 204. Alternatively, the virtual keyboard 1804 can be modified or updated to have a lower number of virtual keys 1808.
[0086] The example setup method 1900 or variations thereof can be implemented for the processor 108 to set up the three-dimensional coordinates and size of a surface with which the user 200 can interact with the system 100, creating an immersive VR / AR / MR experience for the user 200. By using a plane detection algorithm such as a RANSAC algorithm where a significant number of neighboring data points align to a plane within a first and a second tolerance thresholds, the method 1900 or variations thereof can distinguish between a flat surface and a lap surface (defined by a user’s lap) being present within the volume 204 in a computationally efficient manner. By determining the size of the surface based on the distance from at least one of the visual input device 140-1 and the depth input device 140-2 or the center of the field-of-view of the user 200 to the center of the surface, the processor 108 can generate a virtual keyboard 1808 spanning the surface with virtual keys 1804 of a sufficient size for the processor 108 to detect hand gestures or postures without requiring extensive input from the user 200 and in a computationally efficient manner. By performing the setup, the processor 108 is calibrated to detect a series of postures and / or gestures of the user 200 interacting with the surface (e.g. with surface 2012). The setup method 1900 or variations thereof can be further re-run by the processor 108 to update the coordinates and / or size of the surface, for example, upon receiving a setup update request, for example, from the user 200.
[0087] FIGS. 22 and 23 depict example gestures of the user’s hands interacting with the surface that the processor 108 can detect after the three-dimensional coordinates and size of the surface have been set up, for example, by method 1900 or variations thereof.
[0088] Tapping Gesture
[0089] FIG. 22 depicts initial and final virtual representations 2204 and 2208, respectively, of example initial and final hand postures of a tapping gesture 2200, with virtual finger 2212 interacting with virtual key 2216 on virtual surface 2220, illustrating the initial and final postures of the tapping gesture 2200 on an example GUI 2224 that can be displayed to the user 200 as the user 200 performs the tapping gesture 2200 with his real hand 208 on a real surface within the volume 204 (for example, a planar surface, a lap surface or an air surface such as air surface 2012) to provide the user 200 with visual feedback of the tapping gesture 2200. The tapping gesture 2200 is composed of at least one possible initial and at least one possible final hand posture. The at least one possible initial hand posture indicates a hand with a fingertip extending away from the palm, the fingertip located proximal to (i.e. hovering over) a specific grid cell 2104 in the surface (for example, air surface 2012). The at least one possible final hand posture indicates the fingertip contacting the specific grid cell 2104. The processor 108 may determine that the fingertip is contacting a grid cell 2104 by, for example, determining the three-dimensional position of the fingertip within the volume 204, for example from at least one of a visual input device 140-1 and at least one of a depth input device 140-2 with respect to the three-dimensional coordinates of the surface. The processor 108 may further determine that the fingertip is contacting the grid cell 2104 by, for example, tracking the velocity of the finger as it taps the grid cell 2104 (for example, as the fingertip approaches the grid cell 2104), and determining that the fingertip has contacted the grid cell 2104 when the velocity of the finger is zero or below a threshold. The processor 108 can use a determination of the tapping gesture 2200 to trigger a click event of a specific virtual key 1808 on a virtual keyboard 1804. Determinations of tapping gestures 2200 with different fingertips performing the tapping (for example, a middle fingertip tap or an index fingertiptap) on a specific grid cell 2104 can be used by the processor 108 to trigger different events, for example, a left-click or a right-click of the virtual key 1808.
[0090] Gliding Gesture
[0091] FIG. 23 depicts initial and final virtual representations 2304 and 2308, respectively, of example initial and final hand postures of a gliding gesture 2300, with virtual finger 2312 tracing virtual path 2316 on with virtual surface 2320, illustrating the initial and final postures of the gliding gesture 2300 on an example GUI 2324 that can be displayed to the user 200 as the user 200 performs the gliding gesture 2300 with his real hand 208 on a real surface within the volume 204 (for example, a planar surface, a lap surface or an air surface such as air surface 2012) to provide the user 200 with visual feedback of the gliding gesture 2300. The gliding gesture 2300 is composed of at least one possible initial and at least one possible final hand posture. The at least one possible initial hand posture indicates a hand with a fingertip contacting a first grid cell 2104 in the surface. The processor 108 may determine that the fingertip is contacting the first grid cell 2104 by, for example, determining the three-dimensional position of the fingertip within the volume 204, for example from at least one of a visual input device 140-1 and at least one of a depth input device 140-2 with respect to the three-dimensional coordinates of the surface. The at least one possible final hand posture indicates the hand with the fingertip contacting a second grid cell 2104 in the surface, the fingertip having glided from the first grid cell 2104 to the second grid cell 2104 by maintaining the fingertip at the surface. The processor 108 may determine that the fingertip glided from the first grid cell 2104 to the second grid cell 2104 by maintaining the fingertip at the surface by, for example, tracking the three-dimensional position of the fingertip within the volume 204 with reference to the three-dimensional coordinates of the surface as the fingertip moves to the second grid cell 2104 and determining that the position of the fingertip is maintained within a threshold distance range from the surface. The processor 108 can use a determination of the gliding gesture 2300 to trigger a scroll event. The processor 108 can further use a determination of the gliding gesture 2300 to move a virtual element in a GUI, for example, a pointer, from a first position to a second position, the changed in position based on the relative coordinates of the first grid cell 2104 to the second grid cell 2104. The processor 108 can further use a determination of the gliding gesture 2300 or a seriesof consecutive determinations of the gliding gesture 2300 to trigger a drawing or typing event, for example, to enable the user 200 to draw a virtual element by gliding a fingertip through different cells 2104 in the surface.
[0092] The processor 108 may perform determinations of the tapping gesture 2200 or the gliding gesture 2300 by performing example methods 300, 400 and / or 1000 (discussed above) or variations thereof.
[0093] Although the invention has been described with reference to certain specific embodiments, various modifications thereof will be apparent to those skilled in the art without departing from the spirit and scope of the invention as outlined in the claims appended thereto. For example, while the input devices 140 have been presented as preferably physically mounted on I / O devices 136, such as on a VR headset, the input devices 140 may alternatively be distal to the I / O devices 136 and may additionally be positioned distal to the user 200, for example, in front of the user 200 and facing the user 200. A modified set of hand postures and gestures could be used for said alternative location of the input devices 140, with postures indicating hand palms oriented towards the distal input devices 140 and facing away from the user 200. Furthermore, the location of the virtual hands could be modified with respect to the virtual interaction objects in a thereby modified GUI, placing the virtual hands between the virtual interaction objects and the starting point of the field-of-view in the modified GUI. Additionally, while the Graphic User Interface (GUI) 1300 has been discussed as presenting a virtual environment on which virtual hands 1308 and virtual interaction objects 1312 may be displayed, an alternative GUI may, for example, display a mixed environment, composed of both real and virtual elements. For example, the user 200 may be able to see his own hands 208 on the alternative GUI, and the virtual interaction objects 1312 may be, for example, overlayed on the hands 208.
[0094] A person skilled in the art will now appreciate that the teachings herein can improve the technological efficiency and computational and resource utilization across system 100 (and its variants) by making more efficient use of its processing resources, as well as more efficient use of its input resources. At least one technical problem addressed by the present teachings includes the continuous identification of handpostures and gestures of a user interacting with a VR / AR / MR system that consumes processing resources. Such continuous identification consumes significant computational resources of commercial VR / AR / MR systems, severely limiting their capabilities of performing additional processes concurrent to a hand posture and gesture identification process. It should now be apparent to the person skilled in the art that the provided example methods or variants thereof can be used to perform a more computationally efficient hand posture and gesture identification that also optimizes data acquisition from specific input devices for said hand posture and gesture identification. It should now also be apparent that the system 100 and its variants can dedicate a smaller amount of computational resources for hand posture and gesture identification by identifying a specific set of hand postures with a high computational efficiency score, such as the example hand postures hereby discussed or variants thereof.
[0095] It should be recognized that features and aspects of the various examples provided above can be combined into further examples that also fall within the scope of the present disclosure. In addition, the figures are not to scale and may have size and shape exaggerated for illustrative purposes.
[0096] The scope of the claims should not be limited by the embodiments set forth in the above examples but should be given the broadest interpretation consistent with the description as a whole.
Claims
CLAIMS1 . A Virtual Reality, Augmented Reality, or Mixed Reality (VR / AR / MR) system comprising: an Input / Output (I / O) device wearable on a head of a user including: at least one of a depth input device and a visual input device mounted on the I / O device, the at least one of the depth input device and the visual input device configured to sense data associated to the user’s hands when the user’s hands are within a volume in a field-of-view of the at least one of the depth input device and the visual input device; and a visual output device mounted on the I / O device; and a processor configured to: acquire a first dataset from the at least one of the depth input device and the visual input device; identify a first posture of at least one of the user’s hands based on the first dataset; acquire a second dataset from the at least one of the depth input device and the visual input device; identify a second posture based on the second dataset; determine whether the first and second postures correspond to an event triggering gesture; and in response to the determination of the event triggering gesture, generate a visual notification in a Graphic User Interface (GUI) displayed in a display of the visual output device.
2. The VR / AR / MR system of claim 1 , the I / O device further comprising: both a depth input device and a visual input device mounted on the I / O device, the depth and the visual input devices configured to sense data of the user’s hands within the volume, wherein the processor is further configured to: acquire the first and the second datasets from one or both of the depth and the visual input devices.
3. The VR / AR / MR system of claim 1 wherein the first posture indicates the at least one of the user’s hands with an index finger, a middle finger, a ring finger, and a little finger extended away from a palm of the at least one of the user’s hands, and wherein the second posture indicates the at least one of the user’s hands with the index finger, the middle finger, the ring finger, and the little finger closed toward the palm.
4. The VR / AR / MR system of claim 3 wherein the first and the second posture further indicate that the palm is oriented toward the user head.
5. The VR / AR / MR system of claim 1 wherein the first posture indicates the at least one of the user’s hands with a thumb fingertip extending away from an index finger, a middle finger, a ring finger, and a little finger, the index finger, the middle finger, the ring finger, and the little finger extending away from a palm of the at least one of the user’s hands, and wherein the second posture indicates the at least one of the user’s hands with the thumb fingertip meeting one of an index fingertip, a middle fingertip, a ring fingertip, or a little fingertip.
6. The VR / AR / MR system of claim 5 wherein the first and the second posture further indicate that the palm is oriented toward the user head.
7. The VR / AR / MR system of claim 1 wherein the processor is further configured to: determine, based on at least one of the first dataset and an additional dataset from the at least one of the depth input device and the visual input device, that a surface is present within the volume; determine a size of the surface; and divide the surface into a plurality of grid cells.
8. The VR / AR / MR system of claim 7 wherein the first posture indicates a finger of the at least one of the user’s hands with a fingertip extending away from a palm of the at least one of the user’s hands, the fingertip proximal to one of the plurality of grid cells, andwherein the second posture indicates the fingertip contacting the one of the plurality of grid cells.
9. The VR / AR / MR system of claim 7 wherein the first posture indicates a finger of the at least one of the user’s hands with a fingertip extending away from a palm of the at least one of the user’s hands, the fingertip contacting one of the plurality of grid cells, and wherein the second posture indicates the fingertip contacting another one of the plurality of grid cells.
10. The VR / AR / MR system of claim 1 the processor further configured to: acquire a third dataset from the at least one of the depth input device and the visual input device; identify a third posture based on the third dataset; and determine whether the first, the second and the third postures correspond to the event triggering gesture.
11. The VR / AR / MR system of claim 10 wherein the first posture indicates a thumb fingertip extending away from an index, a middle, a ring, and a little finger, the index finger, the middle finger, the ring finger, and the little finger extending away from a palm of the at least one of the user’s hands, wherein the second posture indicates the thumb fingertip meeting an index fingertip, the index and the middle fingers being adjacent to each other, and wherein the third posture indicates the thumb fingertip meeting a middle fingertip, the index and the middle fingers remaining adjacent to each other.
12. The VR / AR / MR system of claim 11 wherein the first, the second, and the third postures further indicate that the palm is oriented toward the user head.
13. The VR / AR / MR system of claim 10 wherein the first posture indicates a thumb fingertip extending away from an index, a middle, a ring, and a little finger, the index finger, the middle finger, the ring finger, and the little finger extending away from a palm of the at least one of the user’s hands, wherein the second posture indicates the thumbfingertip meeting a middle fingertip, the index and the middle fingers being adjacent to each other, and wherein the third posture indicates the thumb fingertip meeting an index fingertip, the index and the middle fingers remaining adjacent to each other.
14. The VR / AR / MR system of claim 13 wherein the first, the second, and the third postures further indicate that the palm is oriented toward the user head.
15. The VR / AR / MR system of claim 1 wherein the processor is further configured to: generate a virtual representation of the at least one of the user’s hands; display the virtual representation in the GUI; and display a virtual interaction object overlayed on the virtual representation in the GUI, and wherein the visual notification is a change in one of a shape, a size, and a color of the virtual interaction object.
16. A method for input detection in a Virtual Reality, Augmented Reality, or Mixed Reality (VR / AR / MR) system by a processor, the method comprising: acquiring a first dataset from at least one of a depth input device and a visual input device mounted on an Input / Output (I / O) device, the I / O device being wearable on a head of a user; determining whether a user hand is within a volume within a field-of-view of the at least one of the depth input device and the visual input device based on the first dataset; identifying a first posture of the user hand based on a portion of the first dataset; acquiring a second dataset from the at least one of the depth input device and the visual input device; identifying a second posture based on a portion of the second dataset; determining whether the first and second postures correspond to an event triggering gesture; and in response to the determination of the event triggering gesture, generating a visual notification in a Graphic User Interface (GUI) displayed in a display of the visual output device.
17. The method of claim 16 further comprising: generating a virtual representation of the user hand; displaying the virtual representation in the GUI; and displaying a virtual interaction object overlayed on the virtual representation, and wherein the visual notification is a change in one of a shape, a size, and a color of the virtual interaction object.
18. The method of claim 16 wherein the first posture indicates the user hand with an index finger, a middle finger, a ring finger, and a little finger extended away from a palm of the user hand, and wherein the second posture indicates the user hand with the index finger, the middle finger, the ring finger, and the little finger closed toward the palm.
19. The method of claim 18 wherein the first and the second posture further indicate that the palm is oriented toward the user head.
20. The method of claim 16 further comprising: determining, based on at least one of the first dataset and an additional dataset from the at least one of the depth input device and the visual input device, that a surface is present within the volume; determining a size of the surface; and dividing the surface into a plurality of grid cells.21 . The method of claim 20 wherein the first posture indicates a finger of the at least one of the user’s hands with a fingertip extending away from a palm of the at least one of the user’s hands, the fingertip proximal to one of the plurality of grid cells, and wherein the second posture indicates the fingertip contacting the one of the plurality of grid cells.
22. The method of claim 20 wherein the first posture indicates a finger of the at least one of the user’s hands with a fingertip extending away from a palm of the at least one of theuser’s hands, the fingertip contacting one of the plurality of grid cells, and wherein the second posture indicates the fingertip contacting another one of the plurality of grid cells.
Citation Information
Patent Citations
Sensory eyewear
US20180075659A1
Virtual user input controls in a mixed reality environment
US20180157398A1
Keyboards for virtual, augmented, and mixed reality display systems
US20180350150A1
Hand gesture input for wearable system
US20210263593A1
Gesture recognition device, operation method for gesture recognition device, and operation program for gesture recognition device
US20220375268A1