Head-mounted display and click input signal generation method

Through the image capturer and inertial sensing device of the head-mounted display, the real-time image and inertial data of the finger are analyzed, the contact mode is judged and the click input signal is generated, which solves the problem of identifying errors in the prior art and improves the accuracy of the input signal.

CN120428847APending Publication Date: 2025-08-05HTC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411148338.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-02
Filing Date
2024-08-21
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the recognition of user's hand movements, existing head-mounted displays are prone to errors or incoherent recognition of click input signals due to image obstruction and difficulty in identifying small hand parts.

Method used

Through the image capturer and inertial sensing device in the head-mounted display, the real-time image and inertial sensing data of the user's finger are analyzed to determine whether the finger complies with the contact mode of the solid plane, and to generate the target finger trajectory and click input signal when complying with the contact mode.

Benefits of technology

It improves the accuracy of the click input signal and solves the problem of misjudgment in computer visual recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428847A_ABST
    Figure CN120428847A_ABST
Patent Text Reader

Abstract

The invention discloses a head-mounted display and a click input signal generation method. The method determines, based on a plurality of real-time images including a plurality of fingers of a user and inertial sensing data, whether the plurality of fingers conform to a contact pattern corresponding to a physical plane. The method generates a target finger trajectory based on the plurality of real-time images and the inertial sensing data in response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane. The method generates click input signals corresponding to the plurality of fingers based on a target input pattern and the target finger trajectory. According to the click input signal generation technology provided by the invention, the accuracy of the click input signal is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a head-mounted display, a method for generating a click input signal, and a non-transitory computer-readable storage medium thereof. Specifically, the present invention relates to a head-mounted display capable of accurately generating a click input signal, a method for generating a click input signal, and a non-transitory computer-readable storage medium thereof. Background Art

[0002] In recent years, various technologies related to virtual reality have developed rapidly, and various head-mounted display technologies and applications have been proposed one after another.

[0003] In the prior art, when a user wears an inside-out tracking head-mounted display, the head-mounted display can directly recognize input signals input by the user's hand movements through computer vision.

[0004] However, recognizing the user's hand movements solely through computer vision may result in incorrect recognition of click input signals or discontinuous click input signals due to image occlusion and other issues.

[0005] In addition, since computer vision is less able to recognize the movements of smaller hand parts (eg, fingers), the possibility of misjudging click input signals is increased.

[0006] In view of this, how to provide a click input signal generating technology that can accurately generate click input signals is a goal that the industry urgently needs to work hard on. Summary of the Invention

[0007] One object of the present invention is to provide a head-mounted display. The head-mounted display includes an image capturer and a processor, and the processor is coupled to the image capturer. The image capturer is used to capture a plurality of real-time images including a plurality of fingers of a user, wherein the user wears at least one wearable device on at least one of the plurality of fingers, and the at least one wearable device is used to generate inertial sensing data. The processor determines whether the plurality of fingers conform to a contact pattern corresponding to a physical plane based on the plurality of real-time images and the inertial sensing data. In response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, the processor generates a target finger trajectory based on the plurality of real-time images and the inertial sensing data. The processor generates a click input signal corresponding to the plurality of fingers based on a target input type and the target finger trajectory.

[0008] Another object of the present invention is to provide a click input signal generating method for an electronic device. The click input signal generating method comprises the following steps: determining whether the multiple fingers of a user conform to a contact pattern corresponding to a physical plane based on multiple real-time images and inertial sensing data, wherein the user wears at least one wearable device on at least one of the multiple fingers, and the at least one wearable device is used to generate the inertial sensing data; in response to the multiple fingers conforming to the contact pattern corresponding to the physical plane, generating a target finger trajectory based on the multiple real-time images and the inertial sensing data; and generating a click input signal corresponding to the multiple fingers based on a target input type and the target finger trajectory.

[0009] Another object of the present invention is to provide a non-transitory computer-readable storage medium, which stores a computer program, wherein the computer program includes multiple program instructions. After being loaded into an electronic device, the computer program executes a click input signal generating method, wherein the click input signal generating method includes the following steps: based on multiple real-time images and inertial sensing data including multiple fingers of a user, determining whether the multiple fingers conform to a contact pattern corresponding to a physical plane, wherein the user wears at least one wearable device on at least one of the multiple fingers, and the at least one wearable device is used to generate the inertial sensing data; in response to the multiple fingers conforming to the contact pattern corresponding to the physical plane, generating a target finger trajectory based on the multiple real-time images and the inertial sensing data; and generating a click input signal corresponding to the multiple fingers based on a target input type and the target finger trajectory.

[0010] In one embodiment of the present invention, determining whether the multiple fingers conform to the contact pattern corresponding to the physical plane includes the following operations: determining whether the multiple fingers of the user are located on the physical plane based on the multiple real-time images; in response to the multiple fingers being located on the physical plane, determining whether the multiple fingers correspond to a finger-down tapping action based on the inertial sensing data; and in response to the multiple fingers corresponding to the finger-down tapping action, determining that the multiple fingers conform to the contact pattern corresponding to the physical plane.

[0011] In one embodiment of the present invention, determining whether the multiple fingers of the user are located on the physical plane includes the following operations: calculating multiple distances between the multiple fingers of the user and the physical plane based on the multiple real-time images; and in response to the multiple distances being less than a preset threshold, determining that the multiple fingers are located on the physical plane.

[0012] In one embodiment of the present invention, the processor further performs the following operations: in response to the multiple fingers not being located on the physical plane, determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane; and in response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, not generating the click input signal corresponding to the multiple fingers.

[0013] In one embodiment of the present invention, the processor further performs the following operations: in response to the multiple fingers not corresponding to the finger tapping action, determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane; and in response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, not generating the click input signal corresponding to the multiple fingers.

[0014] In one embodiment of the present invention, generating the target finger trajectory includes the following operations: in response to the multiple fingers conforming to the contact pattern corresponding to the physical plane, generating a first finger trajectory corresponding to the multiple fingers based on the multiple real-time images; and generating the target finger trajectory based on the first finger trajectory corresponding to the multiple fingers and the inertial sensing data.

[0015] In one embodiment of the present invention, generating the click input signal corresponding to the multiple fingers includes the following operations: in response to the target input type being a writing type, calculating a displacement path corresponding to the target finger trajectory on the physical plane to generate the click input signal corresponding to the multiple fingers.

[0016] In one embodiment of the present invention, generating the click input signal corresponding to the multiple fingers includes the following operations: in response to the target input type being a cursor type, selecting a target cursor action from multiple cursor actions based on the target finger trajectory; and generating the click input signal corresponding to the multiple fingers based on the target finger trajectory and the target cursor action.

[0017] In one embodiment of the present invention, generating the click input signal corresponding to the multiple fingers includes the following operations: in response to the target input type being a keyboard type, calculating a click position corresponding to the target finger trajectory on the physical plane to generate the click input signal corresponding to the multiple fingers.

[0018] The click input signal generating technology provided by the present disclosure (at least including a head-mounted display, a method and a non-transitory computer-readable storage medium thereof) determines whether the multiple fingers conform to the contact mode corresponding to the physical plane by analyzing the real-time images and inertial sensing data corresponding to the user's multiple fingers. Then, the click input signal generating technology provided by the present disclosure can start the operation of generating the target finger trajectory based on the multiple real-time images and the inertial sensing data only when in the contact mode, and generate the click input signal corresponding to the multiple fingers based on the target input type and the target finger trajectory. Since the click input signal generating technology provided by the present disclosure can further assist in determining whether it is in the contact mode through the inertial sensing data of the wearable device and assist in generating the corresponding target finger trajectory in conjunction with the results of computer vision recognition, it solves the problem of misjudgment that may occur only through computer vision recognition. Therefore, the click input signal generating technology provided by the present disclosure improves the accuracy of the click input signal.

[0019] The detailed technology and implementation methods of the present invention are described below in conjunction with the accompanying drawings so that a person having ordinary knowledge in the technical field to which the present invention belongs can understand the technical features of the invention for which protection is sought. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram showing an application environment of the head mounted display according to the first embodiment;

[0021] Figure 2 A schematic diagram illustrating a head mounted display according to certain embodiments;

[0022] Figure 3 A schematic diagram illustrating a wearable device according to certain embodiments;

[0023] Figure 4 A schematic diagram illustrating a wearable device according to certain embodiments;

[0024] Figure 5 A schematic diagram illustrating the operation of certain embodiments;

[0025] Figure 6 A schematic diagram illustrating location information integration operations according to certain embodiments; and

[0026] Figure 7 A partial flow chart showing a method for generating a click input signal according to a second embodiment.

[0027] Explanation of symbols:

[0028] 100: Application environment diagram

[0029] C: User

[0030] 1: Head-mounted display

[0031] 2: Wearable devices

[0032] EP: Solid Plane

[0033] 11: Image Capture

[0034] 13: Processor

[0035] 15: Display

[0036] 21: Transceiver interface

[0037] 23: Processor

[0038] 25: Inertial Measurement Unit

[0039] 27: Electromyography measurement unit

[0040] 500: Operation diagram

[0041] S501, S503, S505, S507, S509, S511: Operation

[0042] 600: Location Information Integration Operation Diagram

[0043] S601, S603, S605, S607, S609: Operation

[0044] ITP: Iterative Process

[0045] 700: Method for generating click input signal

[0046] S701, S703, S705: Steps DETAILED DESCRIPTION

[0047] The following will explain a head-mounted display, a click input signal generating method, and a non-transitory computer-readable storage medium provided by the present invention through an embodiment. However, the multiple embodiments are not intended to limit the present invention to any environment, application, or method as described in the multiple embodiments. Therefore, the description of the embodiments is only for the purpose of illustrating the present invention, and is not intended to limit the scope of the present invention. It should be understood that in the following embodiments and drawings, elements that are not directly related to the present invention have been omitted and are not shown, and the size of each element and the size ratio between elements are only examples, and are not intended to limit the scope of the present invention.

[0048] First, the applicable scenario of this embodiment is described, and its schematic diagram is shown in Figure 1 .like Figure 1As shown in the application environment diagram 100 of the present disclosure, a user C can use a head-mounted display 1, and the user C wears at least one wearable device 2 on the hand (for example, the user C wears a smart ring on the index finger of the right hand) to perform a click input operation corresponding to the display screen of the head-mounted display 2.

[0049] In some embodiments, a system for implementing a click input signal generating method includes a head-mounted display 1 and a wearable device 2 , where the head-mounted display 1 is communicatively connected to the wearable device 2 .

[0050] In this embodiment, the schematic diagram of the structure of the head mounted display 1 is shown in Figure 2 The head-mounted display 1 includes an image capturer 11, a processor 13, and a display 15. The processor 13 is electrically connected to the image capturer 11 and the display 15. The image capturer 11 may include multiple image capture units (e.g., multiple depth camera lenses) for capturing multiple real-time images of the wearable device 2 worn on the hand of the user C.

[0051] In this embodiment, the structure diagram of the wearable device 2 is shown in FIG. Figure 3 The wearable device 2 includes a transceiver interface 21, a processor 23, and an inertial measurement unit 25. The processor 23 is electrically connected to the transceiver interface 21 and the inertial measurement unit 25. The inertial measurement unit 25 can be used to detect the inertial sensing data corresponding to the hand of the user C wearing the wearable device 2.

[0052] Specifically, the inertial measurement unit 25 continuously generates a series of inertial sensing data (e.g., a stream of inertial sensing data generated at a frequency of 1000 times per second), and each of the inertial sensing data may include an acceleration and an angular velocity. During operation, the head-mounted display 1 may periodically receive the inertial sensing data from the wearable device 2.

[0053] It should be noted that the inertial sensing data generated by the wearable device 2 can correspond to body parts (e.g., fingers) of the user C. For example, the user C can wear the wearable device 2 on any finger to collect data. For ease of explanation, this embodiment will illustrate the user C wearing the wearable device 2 on their index finger.

[0054] It should be noted that the transceiver interface 21 is an interface capable of receiving and transmitting data, or any other interface capable of receiving and transmitting data known to those skilled in the art. The transceiver interface may receive data from, for example, an external device, an external webpage, an external application, or the like. The processors 13 and 23 may be various processing units, central processing units (CPUs), microprocessors, or other computing devices known to those skilled in the art.

[0055] It should be noted that Figure 1 This is merely an example, and the present invention does not limit the scope of the system implementing the click input signal generation method. For example, the present invention does not limit the number of wearable devices 2 connected to the head-mounted display 1. The head-mounted display 1 can be connected to multiple wearable devices via a network simultaneously, depending on the scale of the system and actual needs.

[0056] For ease of understanding, we first briefly describe an operation process of this disclosure. Figure 5 , see operation diagram 500 in FIG. In this example, the processor 13 first executes operation S501 to receive data from the image capture device 11 and the wearable device 2. Next, based on this data, the processor 13 executes operation S503 to determine whether the current user action is in contact mode. If the determination is negative, the processor 13 returns to continue executing operation S501.

[0057] If the determination result is yes, the processor 13 executes operation S505 to generate a finger trajectory. Next, the processor 13 executes operation S507 to match the input type and executes operation S507 to generate a click input signal.

[0058] In addition, after the processor 13 executes operation S505, it executes S511 to determine whether the current user's action is not in the contact mode. If the determination result is yes, the processor 13 ends the current continuous action determination. If the determination result is no, the processor 13 continues to execute operation S505 to generate a finger trajectory.

[0059] Next, the following paragraphs will explain the specific details of the operation, please refer to Figure 1 In this embodiment, the processor 13 determines whether the plurality of fingers conform to a contact pattern corresponding to a physical plane EP based on the plurality of real-time images and the inertial sensing data.

[0060] It should be noted that the present disclosure does not limit the shape or size of the physical plane EP. The physical plane EP can be any flat surface in the physical space (for example, a desktop, a wall, the surface of the user's legs, etc.).

[0061] In some embodiments, the processor 13 may first perform a plane detection operation (e.g., through a trained deep learning model) based on the multiple real-time images to search for a physical plane EP in the multiple real-time images.

[0062] In some embodiments, when determining whether the touch mode is in progress, the processor 13 may first analyze the user's finger position using computer vision, and only when the user's finger position is determined to be within the physical plane EP, determine whether the touch mode is in progress using the inertial sensing data.

[0063] Specifically, the processor 13 determines whether the user's multiple fingers are located on the physical plane EP based on the multiple real-time images. Then, in response to the multiple fingers being located on the physical plane EP, the processor 13 determines whether the multiple fingers correspond to a finger tap-down action based on the inertial sensing data. Finally, in response to the multiple fingers corresponding to the finger tap-down action, the processor 13 determines that the multiple fingers meet the contact pattern corresponding to the physical plane EP.

[0064] In some embodiments, the processor 13 can determine whether the user's fingers are located on the physical plane EP by calculating the distances between the user's fingers and the physical plane EP. Specifically, the processor 13 calculates multiple distances between the user's fingers and the physical plane EP based on the multiple real-time images. Then, in response to the multiple distances being less than a predetermined threshold, the processor 13 determines that the fingers are located on the physical plane EP.

[0065] In some embodiments, the processor 13 may determine that the multiple fingers are located on the physical plane EP when only some of the fingers meet the distance condition (for example, the index finger and middle finger of the user are within a preset distance of the physical plane EP).

[0066] In some embodiments, the processor 13 can determine the finger tapping action through a trained deep learning model (e.g., a neural network model), wherein the deep learning model performs deep learning on determining the finger tapping action through a large amount of historical inertial sensing data.

[0067] In some embodiments, in order to save judgment and calculation resources, the processor 13 does not continue to perform the operation of generating the click input signal when it is determined that the contact mode is not met.

[0068] For example, in response to the multiple fingers not being located on the physical plane EP, the processor 13 determines that the multiple fingers do not conform to the contact pattern corresponding to the physical plane EP. Then, in response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane EP, the processor 13 does not generate the click input signals corresponding to the multiple fingers.

[0069] For another example, in response to the multiple fingers not corresponding to the finger tapping action, the processor 13 determines that the multiple fingers do not meet the contact pattern corresponding to the physical plane EP. Then, in response to determining that the multiple fingers do not meet the contact pattern corresponding to the physical plane EP, the processor 13 does not generate the click input signals corresponding to the multiple fingers.

[0070] Next, in this embodiment, in response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane EP, the processor 13 generates a target finger trajectory based on the plurality of real-time images and the inertial sensing data.

[0071] In some embodiments, the processor 13 may adjust the finger trajectory generated based on computer vision using the inertial sensing data. Specifically, in response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane EP, the processor 13 generates a first finger trajectory corresponding to the plurality of fingers based on the plurality of real-time images. Next, the processor 13 generates the target finger trajectory based on the first finger trajectory corresponding to the plurality of fingers and the inertial sensing data.

[0072] For example, see Figure 6 FIG6 is a schematic diagram of the position information integration operation 600. In this example, the processor 13 may execute operations S601 and S603 to obtain the three-axis position information corresponding to the hand tracking (i.e., the position corresponding to the X, Y, and Z axes) and the three-axis acceleration corresponding to the inertial sensing data. Next, after integrating the two types of data, the processor 13 executes Kalman filter prediction in operation S605 and Kalman filter correction in operation S607, and performs multiple iterations of the ITP process. Finally, the processor 13 executes operation S609 to output the enhanced three-axis position information.

[0073] Finally, in this embodiment, the processor 13 generates a click input signal corresponding to the plurality of fingers based on a target input type and the target finger trajectory.

[0074] In some embodiments, the target input type is determined by the processor 13 by determining a target object displayed on the display 15. Specifically, the head-mounted display 1 further includes the display 15 for displaying a target object. The processor 13 determines an input type corresponding to the target object and selects the target input type from a plurality of candidate input types.

[0075] For example, when the target object is a canvas, a signature area, or the like, the processor 13 determines that the appropriate target input type is a writing type (i.e., the click input signal is a finger movement trajectory on the physical plane EP). Specifically, in response to the target input type being a writing type, the processor 13 calculates a displacement path corresponding to the target finger trajectory on the physical plane EP to generate the click input signals corresponding to the multiple fingers.

[0076] In some embodiments, the processor 13 may also determine when a finger (e.g., thumb) is detected clicking on the wearable device 2 and switch a corresponding function (e.g., switching brushes, switching colors, undoing, or redoing). For example, when the processor 13 determines a single click, the brush switching function is executed, and when the processor 13 determines a second click, the color switching function is executed.

[0077] For another example, when the target object is a menu, a drop-down field, or the like, the processor 13 determines that the appropriate target input type is a cursor type (i.e., the click input signal is a finger operation on the physical plane EP). Specifically, in response to the target input type being a cursor type, the processor 13 selects a target cursor action from a plurality of cursor actions based on the target finger trajectory. Then, based on the target finger trajectory and the target cursor action, the processor 13 generates the click input signals corresponding to the plurality of fingers.

[0078] In some embodiments, the processor 13 may switch to an appropriate cursor action (eg, move cursor, left click, right click, scroll up / down, hold left click) based on the target finger trajectory.

[0079] For example, when the processor 13 determines that a finger has moved, the cursor movement function is executed. When the processor 13 determines that a finger has clicked once, the left-click function is executed. When the processor 13 determines that two fingers have clicked simultaneously (for example, the index finger and the middle finger), the right-click function is executed. When the processor 13 determines that two fingers have clicked and slid up and down simultaneously, the up / down slide function is executed. When the processor 13 determines that a finger has clicked twice, the left-click press function is executed.

[0080] For another example, when the target object is an input field (e.g., for entering an account or password), the processor 13 determines that the appropriate target input type is a keyboard type (i.e., the click input signal is a finger operation input on the virtual keyboard corresponding to the physical plane EP). Specifically, in response to the target input type being a keyboard type, the processor 13 calculates a click position corresponding to the target finger trajectory on the physical plane EP to generate the click input signals corresponding to the multiple fingers (i.e., generates an output signal based on the key press content of the virtual keyboard).

[0081] In some embodiments, the processor 13 may further refer to an electromyography (EMG) signal to generate the target finger trajectory. Figure 4 As shown, the wearable device 2 may further include an electromyography (EMG) measurement unit 27 electrically connected to the processor 23. The EMG measurement unit 27 may be configured to detect an EMG signal corresponding to the hand of the user C wearing the wearable device 2.

[0082] In some embodiments, when the wearable device 2 includes an EMG measurement unit 27 , the processor 13 compares the EMG signal with multiple gesture EMG signals (eg, recorded gesture EMG signals of each finger) to identify the finger movement corresponding to the EMG signal.

[0083] As can be seen from the above description, the head-mounted display 1 provided by the present invention determines whether the multiple fingers meet the contact mode corresponding to the physical plane by analyzing the real-time images and inertial sensing data corresponding to the user's multiple fingers. Then, the head-mounted display 1 provided by the present disclosure can start the operation of generating the target finger trajectory based on the multiple real-time images and the inertial sensing data only when it is in the contact mode, and generate the click input signal corresponding to the multiple fingers based on the target input type and the target finger trajectory. Since the head-mounted display 1 provided by the present disclosure can further assist in determining whether it is in the contact mode through the inertial sensing data of the wearable device and assist in generating the corresponding target finger trajectory in conjunction with the results of computer vision recognition, it solves the problem of misjudgment that may occur only through computer vision recognition. Therefore, the head-mounted display 1 provided by the present disclosure improves the accuracy of the click input signal.

[0084] The second embodiment of the present invention is a method for generating a click input signal, the flow chart of which is shown in FIG. Figure 7 The click input signal generating method 700 is applicable to an electronic device, such as the head mounted display 1 described in the first embodiment. The click input signal generating method 700 generates a click input signal through steps S701 to S705.

[0085] In step S701, an electronic device determines whether the multiple fingers of a user conform to a contact pattern corresponding to a physical plane based on multiple real-time images and inertial sensing data, wherein the user wears at least one wearable device on at least one of the multiple fingers, and the at least one wearable device is used to generate the inertial sensing data.

[0086] Next, in step S703 , in response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, the electronic device generates a target finger trajectory based on the plurality of real-time images and the inertial sensing data.

[0087] Finally, in step S705 , the electronic device generates a click input signal corresponding to the plurality of fingers based on a target input type and the target finger trajectory.

[0088] In some embodiments, the step of determining whether the multiple fingers conform to the contact pattern corresponding to the physical plane includes the following steps: determining whether the multiple fingers of the user are located on the physical plane based on the multiple real-time images; in response to the multiple fingers being located on the physical plane, determining whether the multiple fingers correspond to a finger-down tapping action based on the inertial sensing data; and in response to the multiple fingers corresponding to the finger-down tapping action, determining that the multiple fingers conform to the contact pattern corresponding to the physical plane.

[0089] In some embodiments, the step of determining whether the multiple fingers of the user are located on the physical plane includes the following steps: calculating multiple distances between the multiple fingers of the user and the physical plane based on the multiple real-time images; and in response to the multiple distances being less than a preset threshold, determining that the multiple fingers are located on the physical plane.

[0090] In some embodiments, the click input signal generating method 700 further includes the following steps: in response to the multiple fingers not being located on the physical plane, determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane; and in response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, not generating the click input signal corresponding to the multiple fingers.

[0091] In some embodiments, the click input signal generating method 700 further includes the following steps: in response to the multiple fingers not corresponding to the finger tapping action, determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane; and in response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, not generating the click input signal corresponding to the multiple fingers.

[0092] In some embodiments, the step of generating the target finger trajectory includes the following steps: in response to the multiple fingers conforming to the contact pattern corresponding to the physical plane, generating a first finger trajectory corresponding to the multiple fingers based on the multiple real-time images; and generating the target finger trajectory based on the first finger trajectory corresponding to the multiple fingers and the inertial sensing data.

[0093] In some embodiments, the step of generating the click input signals corresponding to the multiple fingers includes the following steps: in response to the target input type being a writing type, calculating a displacement path corresponding to the target finger trajectory on the physical plane to generate the click input signals corresponding to the multiple fingers.

[0094] In some embodiments, the step of generating the click input signal corresponding to the multiple fingers includes the following steps: in response to the target input type being a cursor type, selecting a target cursor action from multiple cursor actions based on the target finger trajectory; and generating the click input signal corresponding to the multiple fingers based on the target finger trajectory and the target cursor action.

[0095] In some embodiments, the step of generating the click input signal corresponding to the multiple fingers includes the following steps: in response to the target input type being a keyboard type, calculating a click position corresponding to the target finger trajectory on the physical plane to generate the click input signal corresponding to the multiple fingers.

[0096] In addition to the aforementioned steps, the second embodiment can also perform all of the operations and steps of the head-mounted display 1 described in the first embodiment, having the same functions and achieving the same technical effects. A person skilled in the art will readily understand how the second embodiment performs these operations and steps based on the first embodiment, achieving the same functions and technical effects, and therefore, a detailed description thereof will not be provided.

[0097] The click input signal generating method described in the second embodiment can be implemented by a computer program having multiple instructions. Each computer program can be a file that can be transmitted over a network, or can be stored in a non-transitory computer-readable storage medium. For each computer program, after the multiple instructions contained therein are loaded into an electronic device (e.g., head-mounted display 1), the computer program executes the click input signal generating method described in the second embodiment. The non-transitory computer-readable storage medium can be an electronic product, such as a read-only memory (ROM), a flash memory, a floppy disk, a hard disk, a compact disk (CD), a portable disk, a database accessible by a network, or any other storage medium known to a person of ordinary skill in the art to which the present invention belongs and having the same function.

[0098] In summary, the click input signal generating technology provided by the present disclosure (at least including a head-mounted display, a method and a non-transitory computer-readable storage medium thereof) determines whether the multiple fingers conform to the contact mode corresponding to the physical plane by analyzing the real-time images and inertial sensing data corresponding to the user's multiple fingers. Then, the click input signal generating technology provided by the present disclosure can start the operation of generating the target finger trajectory based on the multiple real-time images and the inertial sensing data only when in the contact mode, and generate the click input signal corresponding to the multiple fingers based on the target input type and the target finger trajectory. Since the click input signal generating technology provided by the present disclosure can further assist in determining whether it is in the contact mode through the inertial sensing data of the wearable device and assist in generating the corresponding target finger trajectory in conjunction with the results of computer vision recognition, it solves the problem of misjudgment that may occur only through computer vision recognition. Therefore, the click input signal generating technology provided by the present disclosure improves the accuracy of the click input signal.

[0099] The above embodiments are intended only to illustrate some of the embodiments of the present invention and to illustrate the technical features of the present invention, and are not intended to limit the scope and extent of protection of the present invention. Any modifications or equivalent arrangements that can be easily accomplished by a person having ordinary skill in the art to which the present invention belongs fall within the scope claimed by the present invention, and the scope of protection of the present invention is subject to the claims.

Claims

1. A head-mounted display, characterized in that: Include: an image capturer for capturing a plurality of real-time images including a plurality of fingers of a user, wherein the user wears at least one wearable device on at least one of the plurality of fingers, and the at least one wearable device is configured to generate inertial sensing data; and A processor is electrically connected to the image capture device and performs the following operations: determining, based on the plurality of real-time images and the inertial sensing data, whether the plurality of fingers conform to a contact pattern corresponding to a physical plane; In response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, generating a target finger trajectory based on the plurality of real-time images and the inertial sensing data; as well as A click input signal corresponding to the plurality of fingers is generated based on a target input type and the target finger trajectory.

2. The head-mounted display according to claim 1, wherein Determining whether the plurality of fingers conform to the contact pattern corresponding to the physical plane includes the following operations: determining, based on the plurality of real-time images, whether the plurality of fingers of the user are located on the physical plane; In response to the plurality of fingers being located on the physical plane, determining, based on the inertial sensing data, whether the plurality of fingers corresponds to a finger tap-down action; as well as In response to the plurality of fingers corresponding to the finger tapping action, it is determined that the plurality of fingers meet the contact pattern corresponding to the physical plane.

3. The head-mounted display according to claim 2, wherein: Determining whether the multiple fingers of the user are located on the physical plane includes the following operations: Calculating a plurality of distances between the plurality of fingers of the user and the physical plane based on the plurality of real-time images; as well as In response to the multiple distances being smaller than a preset threshold, it is determined that the multiple fingers are located on the physical plane.

4. The head-mounted display according to claim 2, wherein: The processor further performs the following operations: In response to the plurality of fingers not being located on the physical plane, determining that the plurality of fingers do not conform to the contact pattern corresponding to the physical plane; as well as In response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, the click input signals corresponding to the multiple fingers are not generated.

5. The head-mounted display according to claim 2, wherein: The processor further performs the following operations: In response to the plurality of fingers not corresponding to the finger tapping action, determining that the plurality of fingers do not conform to the contact pattern corresponding to the physical plane; as well as In response to determining that the multiple fingers do not conform to the contact pattern corresponding to the physical plane, the click input signals corresponding to the multiple fingers are not generated.

6. The head-mounted display according to claim 1, wherein Generating the target finger trajectory includes the following operations: In response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, generating a first finger trajectory corresponding to the plurality of fingers based on the plurality of real-time images; as well as The target finger trajectory is generated based on the first finger trajectories corresponding to the plurality of fingers and the inertial sensing data.

7. The head-mounted display according to claim 1, wherein Generating the click input signals corresponding to the plurality of fingers includes the following operations: In response to the target input type being a writing type, a displacement path corresponding to the target finger trajectory on the physical plane is calculated to generate the click input signals corresponding to the multiple fingers.

8. The head-mounted display according to claim 1, wherein Generating the click input signals corresponding to the plurality of fingers includes the following operations: In response to the target input type being a cursor type, selecting a target cursor action from a plurality of cursor actions based on the target finger trajectory; and The click input signals corresponding to the plurality of fingers are generated based on the target finger trajectory and the target cursor movement.

9. The head-mounted display according to claim 1, wherein Generating the click input signals corresponding to the plurality of fingers includes the following operations: In response to the target input type being a keyboard type, a click position corresponding to the target finger trajectory on the physical plane is calculated to generate the click input signals corresponding to the multiple fingers.

10. A method for generating a click input signal, characterized in that: For an electronic device, the click input signal generation method includes the following steps: determining, based on a plurality of real-time images of a plurality of fingers of a user and inertial sensing data, whether the plurality of fingers conform to a contact pattern corresponding to a physical plane, wherein the user wears at least one wearable device on at least one of the plurality of fingers, and the at least one wearable device is configured to generate the inertial sensing data; In response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, generating a target finger trajectory based on the plurality of real-time images and the inertial sensing data; as well as A click input signal corresponding to the plurality of fingers is generated based on a target input type and the target finger trajectory.

11. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a computer program, the computer program including a plurality of program instructions. After being loaded into an electronic device, the computer program executes a click input signal generating method, the click input signal generating method including the following steps: determining, based on a plurality of real-time images of a plurality of fingers of a user and inertial sensing data, whether the plurality of fingers conform to a contact pattern corresponding to a physical plane, wherein the user wears at least one wearable device on at least one of the plurality of fingers, and the at least one wearable device is configured to generate the inertial sensing data; In response to the plurality of fingers conforming to the contact pattern corresponding to the physical plane, generating a target finger trajectory based on the plurality of real-time images and the inertial sensing data; as well as A click input signal corresponding to the plurality of fingers is generated based on a target input type and the target finger trajectory.