Head-mounted display device, command sensing method and non-transitory computer readable storage medium
The head-mounted display device uses a combination of camera and inertial measurement data to accurately track user gestures, addressing inaccuracies in existing systems by confirming user inputs, thus improving command execution and reducing power consumption.
Patent Information
- Application Number
- TW114111943
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2025-03-28
- Publication Date
- 2026-07-11
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Existing immersive systems face challenges in accurately tracking user gestures due to inaccuracies in gesture detection based solely on streaming images or inertial measurement data, leading to potential erroneous command execution.
A head-mounted display device combines camera-based streaming images with inertial measurement data from wearable devices to track user gestures, using pre-defined patterns and hand motion tracking to confirm user inputs, ensuring accurate command execution.
This approach enhances gesture tracking accuracy by double-confirming user inputs, reducing power consumption and resource usage, and preventing erroneous command execution.
Smart Images

Figure IMG-2_DRAW_114111943-A0101-14-0001-1 
Figure IMG-2_DRAW_114111943-A0101-14-0001-2 
Figure IMG-2_DRAW_114111943-A0101-14-0002-3
Abstract
Description
Technical Field
[0001] This disclosure relates to a command sensing method, and more particularly to a command sensing method based on gestures or hand dynamics for use in head-mounted display devices. Prior Technology
[0002] Various virtual reality (VR), augmented reality (AR), substitutional reality (SR), and / or mixed reality (MR) devices are designed to provide users with an immersive experience. When a user wears a head-mounted display (HMD) device, their field of vision is covered by immersive content displayed on the HMD device. This immersive content can display virtual backgrounds and virtual objects within an immersive scene.
[0003] Immersive systems typically track user gestures, allowing users to perform interactive actions on virtual objects (such as touching, clicking, pressing, and slamming). Therefore, accurately tracking user gestures is crucial for providing a truly immersive experience. Summary of the Invention
[0004] This disclosure document discloses a head-mounted display device including a camera unit and a processor. The camera unit is used to capture multiple streaming images. The processor is coupled to the camera unit. The processor is used to track gestures based on the streaming images. The processor is used to monitor whether the gesture matches a pre-defined pattern. When the gesture matches the pre-defined pattern at a first time point, the processor initiates hand motion tracking from the first time point to a sensing period that continues until a second time point. During this sensing period, the processor monitors whether the hand motion matches an instruction pattern corresponding to the pre-defined pattern. When the hand motion matches the instruction pattern, the processor executes an operation corresponding to that instruction pattern.
[0005] Another aspect of this disclosure discloses a command sensing method, comprising the following steps: capturing multiple streaming images; tracking gestures based on the streaming images; monitoring whether the gestures match a pre-defined pattern; when the gestures match the pre-defined pattern at a first time point, activating hand motion tracking during a sensing period starting at the first time point; during the sensing period, monitoring whether the hand motions match a command pattern corresponding to the pre-defined pattern; and when the hand motions match the command pattern, executing an operation corresponding to the command pattern.
[0006] Another aspect of this disclosure document discloses a non-transitory computer-readable medium having a computer program for executing the aforementioned instruction sensing method.
[0007] It should be noted that the above description and the following detailed description are illustrative of this case by way of embodiments, and are used to assist in the explanation and understanding of the invention content claimed in this case. Simple Explanation of the Diagram
[0008] To make the above and other objects, features and embodiments of this disclosure more apparent and understandable, the accompanying drawings are described below: Figure 1A illustrates a schematic diagram of an immersive system according to some embodiments of this disclosure; Figure 1B illustrates a functional block diagram of the immersive system in Figure 1A according to some embodiments of this disclosure; Figure 2 illustrates a flowchart of an instruction sensing method according to some embodiments of this disclosure; Figure 3A illustrates a diagram of a streaming video containing gestures; Figure 3B illustrates a schematic diagram of a streaming image related to another gesture; Figure 3C illustrates a schematic diagram of a streaming image related to another gesture; Figure 4A illustrates a hand movement in an example; Figure 4B illustrates a hand movement in another example; Figure 5 illustrates a flowchart of an instruction sensing method according to some embodiments of this disclosure; Figure 6A illustrates a schematic diagram of an immersive environment displayed on a display according to some embodiments; and Figure 6B illustrates a schematic diagram of another immersive environment displayed on a display according to some other embodiments. Implementation
[0009] The following disclosure provides numerous different embodiments or examples for implementing various features of this disclosure. Elements and arrangements in the specific examples are used in the following discussion to simplify this disclosure. Any examples discussed are for illustrative purposes only and do not in any way limit the scope or meaning of this disclosure or its examples. Where appropriate, the same reference numerals are used between figures and in corresponding text descriptions to represent the same or similar elements.
[0010] Please refer to Figures 1A and 1B. Figure 1A illustrates a schematic diagram of an immersive system 100 according to some embodiments of this disclosure. Figure 1B illustrates a functional block diagram of the immersive system 100 in Figure 1A according to some embodiments of this disclosure. The immersive system 100 includes a head-mounted display (HMD) device 120 and at least one wearable device 140.
[0011] In the embodiment shown in Figure 1A, the at least one wearable device 140 comprises four smart rings worn on four different fingers of the user. However, this disclosure is not limited to a specific number of wearable devices 140. In some other embodiments, the immersive system 100 may comprise K wearable devices 140 worn on K fingers of the user, where K is a positive integer between 1 and 10.
[0012] As shown in Figure 1B, in some embodiments, the head-mounted display device 120 includes a camera unit 122, a processor 124, a transceiver circuit 126, and a display 128. The processor 124 is coupled to the camera unit 122, the transceiver circuit 126, and the display 128. In some embodiments, the camera unit 122 may be disposed on the front surface of the head-mounted display device 120. The camera unit 122 is used to capture a series of streaming images. The camera unit 122 may include a lens, an optical sensor, and / or a graphics processing unit. The processor 124 may include a central processing unit, a microcontroller unit (MCU), or an application-specific integration circuit (ASIC). The transceiver circuit 126 may include local communication circuitry (e.g., Bluetooth transceiver circuitry, WiFi transceiver circuitry, Zigbee transceiver circuitry) or telecommunications circuitry (e.g., 4G transceiver circuitry, 5G transceiver circuitry). The display 128 may include a display panel for displaying an immersive environment corresponding to the user's field of vision.
[0013] Based on the streaming images captured by camera unit 122, head-mounted display device 120 can track the user's gestures (and / or hand movements) to detect user input commands and execute corresponding functions. If the user's gestures / hand movements are detected solely based on the streaming images captured by camera unit 122, the detected gestures / hand movements may be inaccurate in certain extreme cases (e.g., the hand moves only slightly within the field of view of camera unit 122, the user is in bright light, or the user's hand leaves the field of view of camera unit 122).
[0014] In some embodiments, wearable device 140 may provide information other than streaming video to more accurately track gestures / hand movements. As shown in Figure 1B, each wearable device 140 may include an inertial measurement unit 142 and a transceiver circuit 144. The inertial measurement unit 142 generates inertial measurement data (DIMU). The DIMU represents the acceleration and / or orientation of each wearable device 140 along the X / Y / Z axes. The DIMU can be transmitted from each wearable device 140 to a head-mounted display device 120 by the transceiver circuit 144. The inertial measurement unit 142 may include a gyroscope sensor and an accelerometer. The transceiver circuit 144 may include local communication circuitry (e.g., Bluetooth transceiver circuitry, WiFi transceiver circuitry, Zigbee transceiver circuitry) or telecommunications circuitry (e.g., 4G transceiver circuitry, 5G transceiver circuitry). The head-mounted display device 120 can further track gestures / hand movements based on the received DIMU. The inertial measurement unit 142 is interconnected with the transceiver circuit 144.
[0015] On the other hand, if the user's gestures / hand movements are detected solely based on the inertial measurement data (DIMU) detected by the wearable device 140, it may cause the head-mounted display device 120 to erroneously trigger some unexpected commands / functions.
[0016] In some embodiments, the head-mounted display device 120 performs a command sensing method to sense the user's gestures / hand movements based on a combination of streaming images captured by the camera unit 122 and inertial measurement data (DIMU) collected from the wearable device 140, thereby enabling the head-mounted display device 120 to execute corresponding commands based on the gestures / hand movements.
[0017] Please also refer to Figure 2, which illustrates a flowchart of a command sensing method 200 according to some embodiments of this disclosure. The command sensing method 200 can be performed by the head-mounted display device 120 shown in Figure 1B. In step S210 of the command sensing method 200, the camera unit 122 is used to capture streaming images.
[0018] In step S220, the processor 124 tracks the gesture based on the streaming image captured by the camera unit 122. Please refer to Figures 3A, 3B, and 3C. Figure 3A illustrates a schematic diagram of the streaming image IMGa for gesture HGa. Figure 3B illustrates a schematic diagram of the streaming image IMGb for another gesture HGb. Figure 3C illustrates a schematic diagram of the streaming image IMGc for another gesture HGc.
[0019] As shown in Figure 3A, processor 124 executes computer vision algorithms to identify and locate the knuckle positions KN of the hand in the streaming image IMGa. Based on the positional distribution of the knuckle positions KN, processor 124 is able to track the gesture HGa appearing in the streaming image IMGa.
[0020] Similarly, processor 124 is used to execute computer vision algorithms to identify and locate the knuckle positions KN of the hand in the streaming images IMGb (as shown in Figure 3B) and IMGc (as shown in Figure 3C). Based on the positional distribution of the knuckle positions KN, processor 124 is able to track the gesture HGb appearing in the streaming image IMGb and the gesture HGc appearing in the streaming image IMGc.
[0021] Since the finger joint positions KN of the hand are distributed differently in Figures 3A, 3B and 3C, the processor 124 can identify different hand gestures HGa, HGb and HGc based on the streaming images IMGa, IMGb and IMGc.
[0022] In step S230, the processor 124 detects whether the gestures HGa, HGb, or HGc appearing in the streaming images IMGa~IMGc match a pre-defined pattern. The pre-defined pattern is a predetermined gesture shape that indicates that the user may or is about to input a command.
[0023] For example, the preparatory style includes a click preparatory style P PRE1 (indicating that the user is about to perform a click input), as shown in Figure 3A. The click preparatory style P PRE1 is in the form of at least one finger hovering diagonally in front of the head-mounted display device 120.
[0024] For example, a prep style can also include a pinch prep style P PRE2 (indicating that the user is about to perform a pinch input), as shown in Figure 3B. The shape of the pinch prep style P PRE2 is two fingers suspended with a certain distance between them GP.
[0025] On the other hand, the gesture HGc shown in Figure 3C (e.g., a scissor gesture) is not similar to either the click preparatory style P PRE1 or the pinch preparatory style P PRE2. In this case, if the processor 124 receives the streaming image IMGc from the camera unit 122 and the processor 124 determines that the gesture HGc in the streaming image IMGc does not match either preparatory style, the instruction sensing method 200 will proceed to step S280.
[0026] In the first example, assuming that processor 124 receives streaming image IMGa from camera unit 122, and at a first time point T1, processor 124 determines that the gesture HGa in streaming image IMGa matches the click preparation pattern P PRE1. In this case, step S240 will be executed to initiate hand dynamic tracking during the sensing period SP from the first time point T1 to the second time point T2.
[0027] In some embodiments, the second time point T2 can be set at an appropriate time point after the first time point T1. For example, the second time point T2 can be set 500 milliseconds after the first time point T1 (i.e., the duration of SP during the sensing period is equal to 500 milliseconds).
[0028] In some embodiments, in step S240, hand movement tracking can be based on inertial measurement data (DIMU) from wearable device 140. In this case, as shown in Figure 1B, processor 124 is used to transmit a trigger signal TR to each wearable device 140. The trigger signal TR is used to trigger the inertial measurement unit 142 in each wearable device 140. In response to the trigger signal TR, wearable device 140 transmits the sensed inertial measurement data (DIMU) back to processor 124 of head-mounted display device 120. In step S240, processor 124 is used to track hand movements based on inertial measurement data (DIMU) received from at least one wearable device 140.
[0029] Please refer to Figures 4A and 4B together. Figure 4A shows a schematic diagram of the hand dynamic HMa in one example. Figure 4B shows a schematic diagram of the hand dynamic HMb in another example.
[0030] In some embodiments shown in Figure 4A, hand dynamics HMa can be determined as at least one finger pressing down or moving downward, based on inertial measurement data (DIMU) received from the wearable device 140. For example, hand dynamics HMa can be determined primarily based on the vertical acceleration along the Z-axis in the inertial measurement data (DIMU).
[0031] In step S250, during the sensing period SP, the processor 124 monitors whether the hand dynamic HMa matches the instruction pattern corresponding to the pre-selection pattern (e.g., the click pre-selection pattern P PRE1 determined in step S230).
[0032] For example, as shown in Figure 4A, the instruction pattern includes the click instruction pattern P CMD1 (indicating that the user is performing a click input). The click instruction pattern P CMD1 is in the form of at least one finger pressing down or moving down.
[0033] In the first example, if a click preparation pattern P PRE1 is detected in step S230 and a hand motion HMa is subsequently detected in step S250, and the processor 124 detects that the hand motion HMa matches the click instruction pattern P CMD1 corresponding to the click preparation pattern P PRE1, in this case, the processor 124 executes step S260 to perform the operation corresponding to the click instruction pattern P CMD1 on the head-mounted display device 120 (e.g., clicking or confirming on a button or icon).
[0034] In another embodiment shown in Figure 4B, hand dynamics HMb can be determined as two fingers moving closer to each other based on inertial measurement data (DIMU) received from the wearable device 140. For example, hand dynamics HMb can be determined primarily based on the horizontal acceleration from the inertial measurement data (DIMU) from the two wearable devices 140 worn on the two fingers.
[0035] In the first example, if a click preparation pattern P PRE1 is detected in step S230 and a hand dynamic HMb as shown in Figure 4B is subsequently detected in step S250, the hand dynamic HMb detected by the processor 124 in step S250 cannot match the click instruction pattern P CMD1 (see Figure 4A) corresponding to the click preparation pattern P PRE1 detected in step S230. In this case, the hand dynamic HMb will be considered an invalid instruction, and the processor 124 will not perform the corresponding operation (because no reliable instruction input was detected). The instruction sensing method 200 proceeds to step S270. In step S270, the processor 124 checks whether the sensing period SP has expired. If the sensing period SP has not expired, the instruction sensing method 200 returns to step S250 and continues to monitor the hand dynamic.
[0036] If the sensing period SP has expired, the instruction sensing method 200 proceeds to step S280, whereby the processor 124 stops tracking the hand's dynamics. In some embodiments, in step S280, the processor 124 may ignore the inertial measurement unit (DIMU) data from the wearable device 140. In other embodiments, in step S280, the processor 124 may send a stop signal (not shown) to each wearable device 140 to shut down the inertial measurement unit 142 in each wearable device 140. In still other embodiments, in step S280, the processor 124 may turn off the transceiver circuit 126 to stop transmitting the inertial measurement unit (DIMU) data.
[0037] In the first example in the preceding paragraph, it is presupposed that processor 124 receives streaming image IMGa from camera unit 122, and processor 124 detects at a first time point T1 that gesture HGa in streaming image IMGa matches click preparation pattern P PRE1. However, this disclosure is not limited thereto.
[0038] In the second example, assuming that processor 124 receives the streaming image IMGb from camera unit 122, and in step S230 at a first time point T1, processor 124 determines that the gesture HGb in the streaming image IMGb matches the pinch preparation pattern P PRE2. Then, step S240 is executed to initiate hand dynamic tracking.
[0039] In the second example, if a pinch preparation pattern P PRE2 is detected in step S230 and subsequently a hand dynamic HMa as shown in Figure 4A is detected in step S250, the processor 124 detects that the hand dynamic HMa fails to match the pinch instruction pattern P CMD2 (refer to Figure 4B) corresponding to the pinch preparation pattern P PRE2. In this case, the hand dynamic HMa will be considered an invalid instruction, and the processor 124 will not perform the corresponding operation (because no reliable instruction input was detected). The instruction sensing method 200 proceeds to step S270.
[0040] On the other hand, in the second exemplary example, if a pinch preparation pattern P PRE2 is detected in step S230 and subsequently a hand dynamic HMb as shown in Figure 4B is detected in step S250, the processor 124 detects that the hand dynamic HMb matches the pinch instruction pattern P CMD2 (refer to Figure 4B) corresponding to the pinch preparation pattern P PRE2 (refer to Figure 3B). As shown in Figure 4B, the shape of the pinch instruction pattern P CMD2 is two fingers moving closer to each other. In this case, the processor 124 executes step S260, performing an operation corresponding to the pinch instruction pattern P CMD2 on the head-mounted display device 120 (e.g., a pinch operation thereby collecting, holding, or deforming a virtual object).
[0041] Based on the above embodiments, the pre-defined pattern and the corresponding instruction pattern are used to double-confirm the correct intent of the user's input. If the detected gesture matches a specific pre-defined pattern but the subsequently detected hand movement does not match the corresponding instruction pattern, the operation will not be executed, thereby improving the accuracy of the instruction sensing method 200. If no gesture matching the pre-defined pattern is detected, the tracking of hand movement can be paused, thereby reducing the power consumption of the head-mounted display device 120 and / or the wearable device 140 and saving the computing resources of the head-mounted display device 120 and / or the wearable device 140.
[0042] In this disclosure, the preparatory patterns and instruction patterns are not limited to the click and pinch mentioned in the above embodiments. The head-mounted display device 120 and the instruction sensing method 200 can handle other similar preparatory patterns and instruction patterns (e.g., tap, grasp, clap, hold, etc.).
[0043] In the foregoing embodiments, processor 124 is used to track hand movements based on inertial measurement data (DIMU) received from wearable device 140. However, this disclosure is not limited thereto.
[0044] In some other embodiments, the processor 124 is used in step S240 to locate the knuckle positions of the hand in the streaming image by executing a computer vision algorithm (similar to the embodiments shown in Figures 3A, 3B, and 3C), and to track hand movements based on the knuckle positions. In this case, both the gesture (corresponding to the preparatory pattern) and the hand movements (corresponding to the instruction pattern) are tracked based on the computer vision algorithm in the streaming image captured by the camera unit 122. In this case, the head-mounted display device 120 does not rely on the wearable device 140; the head-mounted display device 120 alone (without the assistance of the wearable device 140) can perform the matching of gestures with preparatory patterns and the matching of hand movements with instruction patterns, thereby doubly confirming the correct intent of the user's input operation.
[0045] Please also refer to Figure 5, which illustrates a flowchart of a command sensing method 500 according to some embodiments of this disclosure. The command sensing method 500 in Figure 5 can be performed by the head-mounted display device 120 shown in Figure 1B. Steps S510, S520, S530, S540, S550, S560, S570, and S580 of the command sensing method 500 in Figure 5 are similar to the aforementioned steps S210, S220, S230, S240, S250, S260, S270, and S280 of the command sensing method 200 in Figure 2, and the details of these steps will not be repeated here.
[0046] As shown in Figure 5, after the gesture is determined to match a certain pre-selected pattern in step S530, the instruction sensing method 500 further includes steps S531, S532 and S533 before starting hand dynamic tracking (i.e. step S540).
[0047] As shown in Figure 1B, the display 128 can display an immersive environment to the user's field of vision. In some embodiments, the display 128 is used to display virtual objects in the immersive environment. Steps S531, S532, and S533 can be used to verify whether a gesture (matching a pre-defined pattern) is close to a virtual object.
[0048] Please refer to Figures 6A and 6B together. Figure 6A illustrates a schematic diagram of an immersive environment IMa displayed on display 128 according to some embodiments. Figure 6B illustrates a schematic diagram of another immersive environment IMb displayed on display 128 according to other embodiments.
[0049] Assume that in step S530, the processor 124 receives the streaming image IMGa (see Figure 3A) from the camera unit 122, and the processor 124 determines that the gesture HGa in the streaming image IMGa matches the click preparation style P PRE1.
[0050] In this scenario, the virtual avatar V HAND of the gesture can be displayed in the immersive environment IMa / IMb shown in Figure 6A or Figure 6B. In step S531, the processor 124 locates the virtual position of the virtual avatar V HAND in the immersive environment IMa / IMb shown in Figure 6A or Figure 6B. In step S532, the processor 124 detects the distance between the virtual position of the virtual avatar V HAND and the virtual object V OBJ in the immersive environment IMa / IMb shown in Figure 6A or Figure 6B.
[0051] In the embodiments shown in Figures 5 and 6A, in step S532, the distance GD1 between the virtual position of the gesture's virtual avatar V HAND and the virtual object V OBJ in the immersive environment IMa is detected. In step S533, the distance GD1 is determined to be shorter than a threshold GTH, which indicates that the gesture's virtual avatar V HAND and the virtual object V OBJ are relatively close. The processor 124 can determine that the user is about to (or has a high probability of) interacting with the virtual object V OBJ. In this case, the instruction sensing method 500 proceeds to step S540 to initiate hand movement tracking.
[0052] On the other hand, in the embodiments shown in Figures 5 and 6B, in step S532, the distance GD2 between the virtual position of the gesture's virtual avatar V HAND and the virtual object V OBJ in the immersive environment IMa is detected. In step S533, the distance GD2 is determined to exceed a threshold GTH, which means that the gesture's virtual avatar V HAND and the virtual object V OBJ are relatively far apart. The processor 124 can determine that the user will not (or only has a low probability) interact with the virtual object V OBJ. In this case, the gesture can be ignored (in the case where the virtual avatar V HAND is far away from the virtual object V OBJ), and in some embodiments, the instruction sensing method 500 may return to step S520 to continue tracking and updating the gesture's position.
[0053] In other words, hand motion tracking is only initiated (step S540) when the gesture matches the pre-defined pattern and the spacing distance (e.g., spacing distance GD1 in Figure 6A) is less than the threshold GTH. In this case, when the user waves their hand in an empty area, the processor 124 can continue to track the gesture in the first stage (based on the camera unit 122) and disable hand motion tracking in the second stage to reduce the power consumption of the head-mounted display device 120 and / or the wearable device 140 and save computing resources on the head-mounted display device 120 and / or the wearable device 140.
[0054] Another embodiment of this disclosure includes a non-transitory computer-readable storage medium for storing at least one program instruction executed by a processing unit (see processor 124 in the embodiment shown in Figure 1B) to perform the instruction sensing method 200 shown in Figure 2 or the instruction sensing method 500 shown in Figure 5.
[0055] While specific embodiments of the present disclosure have been disclosed in relation to the above embodiments, these embodiments are not intended to limit the present disclosure. Various alternatives and modifications can be made by those skilled in the art in accordance with the present disclosure without departing from the principles and spirit of the present disclosure. Therefore, the scope of protection of the present disclosure is determined by the appended claims.
[0056] 100: Immersive System 120: Head-mounted display device 122: Camera Unit 124: Processor 126: Transceiver Circuit 128: Monitor 140: Wearable devices 142: Inertial Measurement Unit 144: Transceiver Circuit 200, 500: Command Sensing Method S210, S220, S230, S240: Steps S250, S260, S270, S280: Steps S510, S520, S530, S540: Steps S531, S532, S533: Steps S550, S560, S570, S580: Steps TR: Trigger signal DIMU: Inertial Measurement Data IMGa, IMGb, IMGc: Streaming video GP: Spacing HGa, HGb, HGc: Gestures KN: Knuckle position P PRE1: Click Preparing Style P PRE2: Kneading Preparation Style P CMD1: Click command style P CMD2: Pinch command style SP: Sensing period T1: First Time Point T2: Second Time Point HMa, HMb: Hand dynamics V HAND: Virtual Avatar IMa,IMb: Immersive Environment VOBJ: Virtual Object GD1, GD2: Spacing distance G TH: Threshold
[0057] Domestic storage information (please note in order of storage institution, date, and number) none Overseas storage information (please note in the order of storage country, institution, date, and number) none
Claims
1. A head-mounted display device comprising: a camera unit for capturing a plurality of streaming images; a display for displaying a virtual object in an immersive environment; and a processor coupled to the camera unit and the display, the processor being configured to: track a gesture based on the streaming images; monitor whether the gesture matches a pre-defined pattern; locate a virtual position of the gesture in the immersive environment; detect a distance between the virtual position of the gesture and the virtual object in the immersive environment; when the gesture matches the pre-defined pattern at a first time point and the distance is less than a threshold, initiate a hand motion tracking function during a sensing period from the first time point to a second time point; during the sensing period, monitor whether the hand motion matches a command pattern corresponding to the pre-defined pattern; and when the hand motion matches the command pattern, execute an operation corresponding to the command pattern.
2. The head-mounted display device as claimed in claim 1, wherein the processor is configured to track the gesture by: executing a computer vision algorithm to locate the positions of a plurality of knuckles of a hand in the streaming images; and tracking the gesture based on the knuckle positions.
3. The head-mounted display device as claimed in claim 1 further includes: a transceiver circuit for communicating with a wearable device, wherein the wearable device includes an inertial measurement unit for generating inertial measurement data.
4. The head-mounted display device as described in claim 3, wherein when the gesture matches the pre-defined pattern, the processor sends a trigger signal to the wearable device to trigger the inertial measurement unit of the wearable device.
5. The head-mounted display device as described in claim 3, wherein when the gesture matches the pre-defined pattern, the processor tracks the hand movement based on the inertial measurement data received from the wearable device.
6. The head-mounted display device as claimed in claim 3, wherein the wearable device includes at least one smart ring worn on at least one finger of a user.
7. The head-mounted display device as claimed in claim 1, wherein the processor is configured to track the hand movements by: executing a computer vision algorithm to locate the positions of a plurality of knuckles of a hand in the streaming images; and tracking the hand movements based on the knuckle positions.
8. The head-mounted display device as claimed in claim 1, wherein the processor stops the hand movement tracking function when the gesture fails to match the pre-set pattern or when the sensing period has expired.
9. The head-mounted display device as claimed in claim 1, wherein the preparation style includes a click preparation style corresponding to at least one finger hovering diagonally in front of the head-mounted display device, and the instruction style includes a click instruction style corresponding to pressing down or moving the at least one finger.
10. The head-mounted display device as claimed in claim 1, wherein the preparation mode includes a pinch preparation mode corresponding to two fingers hovering and maintaining a distance between them, and the instruction mode includes a pinch instruction mode corresponding to the two fingers moving closer to each other.
11. A command sensing method, comprising: capturing a plurality of streaming images; tracking a gesture based on the streaming images; monitoring whether the gesture matches a pre-defined pattern; displaying a virtual object in an immersive environment; locating a virtual position of the gesture in the immersive environment; detecting a distance between the virtual position of the gesture and the virtual object in the immersive environment; when the gesture matches the pre-defined pattern at a first time point and the distance is less than a threshold, activating a hand motion tracking function during a sensing period from the first time point to a second time point; during the sensing period, monitoring whether the hand motion matches a command pattern corresponding to the pre-defined pattern; and when the hand motion matches the command pattern, performing an operation corresponding to the command pattern.
12. The instruction sensing method as described in claim 11, wherein the step of tracking the gesture comprises: executing a computer vision algorithm to locate the positions of a plurality of knuckles of a hand in the streaming images; and tracking the gesture based on the knuckle positions.
13. The command sensing method as described in claim 11, wherein the step of activating the hand dynamics tracking function when the gesture matches the preparatory pattern comprises: sending a trigger signal to trigger an inertial measurement unit of a wearable device.
14. The command sensing method as described in claim 13, wherein the step of activating the hand dynamic tracking function when the gesture matches the preparatory pattern comprises: tracking the hand dynamic based on an inertial measurement data received from the wearable device.
15. The command sensing method as described in claim 11 further includes: stopping the hand movement tracking function when the gesture fails to match the preparatory pattern or when the sensing period has expired.
16. The instruction sensing method as described in claim 11, wherein the preparatory pattern includes a click preparatory pattern corresponding to at least one finger hovering in a forward direction, and the instruction pattern includes a click instruction pattern corresponding to at least one finger pressing down or moving.
17. The command sensing method as described in claim 11, wherein the preparatory pattern includes a pinch preparatory pattern corresponding to two fingers hovering and maintaining a distance between them, and the command pattern includes a pinch command pattern corresponding to the two fingers moving closer to each other.
18. A non-transitory computer-readable medium having a computer program for executing an instruction sensing method, the instruction sensing method comprising: capturing a plurality of streaming images; tracking a gesture based on the streaming images; monitoring whether the gesture matches a pre-defined pattern; displaying a virtual object in an immersive environment; locating a virtual position of the gesture in the immersive environment; detecting a distance between the virtual position of the gesture and the virtual object in the immersive environment; when the gesture matches the pre-defined pattern at a first time point and the distance is less than a threshold, activating a hand motion tracking function during a sensing period from the first time point to a second time point; during the sensing period, monitoring whether the hand motion matches an instruction pattern corresponding to the pre-defined pattern; and when the hand motion matches the instruction pattern, performing an operation corresponding to the instruction pattern.