Input device and method

By recognizing user gestures to generate a virtual keyboard, the problem of limited virtual object generation environment and complex operation in virtual reality and augmented reality technologies is solved, and intuitive virtual keyboard operation and editing functions are realized.

CN120653100APending Publication Date: 2025-09-16HTC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411596943.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2024-11-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In existing virtual reality and augmented reality technologies, the environment for generating virtual objects is limited by physical reference objects, and the operations for inputting or editing text are complex and unintuitive.

Method used

By using cameras and processors, it recognizes multiple hand images of the user, determines gestures and generates a virtual keyboard on a virtual plane, generating input instructions based on gesture operations, including typing and editing functions.

Benefits of technology

It enables intuitive operation of the virtual keyboard on a virtual plane, reduces learning costs, improves the convenience of text editing, and is independent of the physical environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653100A_ABST
    Figure CN120653100A_ABST
Patent Text Reader

Abstract

An input device is used for executing the following operations: determining a first gesture of a user based on a plurality of first hand images in a plurality of hand images; responding to the condition that the first gesture accords with a starting gesture, a virtual keyboard is generated on a virtual plane at a first time point, and the virtual plane is generated based on a palm position corresponding to the first gesture; determining a second gesture of the user based on a plurality of second hand images corresponding to a second time point in the plurality of hand images, the first time point being earlier than the second time point; and generating an input instruction corresponding to a typewriting gesture based on a displacement between the second gesture and the virtual keyboard in response to the second gesture conforming to the typewriting gesture. According to the input technology provided by the invention, the virtual keyboard is generated on the virtual plane by judging the gesture, the input technology is not limited by an entity environment, and more intuitive operation experience is provided for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an input device and method, and more particularly to an input device and method based on user gestures. Background Art

[0002] In current virtual reality and / or augmented reality technologies, if a virtual object needs to be generated at a specific location in a real environment, it is necessary to rely on a specific pattern or physical object as a reference, and the generated virtual object will move with the reference object.

[0003] However, existing technologies limit the environment for generating virtual objects, and in applications of virtual reality and / or augmented reality technology, the operation method of inputting or editing text is more complicated and unintuitive than using a physical keyboard.

[0004] In view of this, how to provide an intuitive input technology that is not restricted by the physical environment is a goal that the industry urgently needs to work hard on. Summary of the Invention

[0005] In order to solve the above problems, the present disclosure proposes an input device comprising a camera and a processor. The camera is used to capture multiple hand images of a user. The processor is communicatively connected to the camera and is used to perform the following operations: based on multiple first hand images among the multiple hand images, determine a first gesture of the user; in response to the first gesture being consistent with a start gesture, generate a virtual keyboard on a virtual plane at a first time point, wherein the virtual plane is generated based on a palm position corresponding to the first gesture; based on multiple second hand images corresponding to a second time point among the multiple hand images, determine a second gesture of the user, wherein the first time point is earlier than the second time point; and in response to the second gesture being consistent with a typing gesture, generate an input instruction corresponding to the typing gesture based on a displacement between the second gesture and the virtual keyboard.

[0006] In one embodiment of the present invention, the operation of generating the virtual keyboard further includes: generating the virtual plane below the palm position based on the palm position corresponding to the first gesture; and generating the virtual keyboard on the virtual plane.

[0007] In one embodiment of the present invention, the operation of generating the input instruction corresponding to the typing gesture further includes: calculating a first movement path of each of the multiple fingertips based on the multiple second hand images; and generating the input instruction of a key corresponding to one of the multiple fingertips in response to the first movement path of one of the multiple fingertips being perpendicular to the virtual plane.

[0008] In one embodiment of the present invention, the processor is further configured to perform the following operations: in response to the second gesture being consistent with one of a plurality of editing gestures, executing an editing function corresponding to one of the plurality of editing gestures.

[0009] In one embodiment of the present invention, the processor is further configured to perform the following operations: calculate a plurality of hand joint points in the plurality of hand images; and determine the first gesture and the second gesture based on the plurality of hand joint points.

[0010] In one embodiment of the present invention, the processor is further used to perform the following operations: calculate multiple fingertip positions in the multiple second hand images; and based on the multiple fingertip positions and the virtual keyboard, calculate a key corresponding to each of the multiple fingertip positions in the virtual keyboard.

[0011] In one embodiment of the present invention, the processor is further configured to perform the following operations: in response to the second gesture being a close gesture, close the virtual keyboard.

[0012] In one embodiment of the present invention, the processor is further configured to perform the following operations: in response to the second gesture indicating that the user changes from a hands-out gesture to a hands-closed gesture, determining that the second gesture matches the closing gesture.

[0013] In one embodiment of the present invention, the processor is further configured to perform the following operations: select a cursor position based on one of a plurality of fingertip positions in the plurality of second hand images; and generate input content based on the cursor position and the input instruction.

[0014] In one embodiment of the present invention, the processor is further configured to perform the following operations: in response to the second gesture being consistent with a selection gesture, calculate a second movement path of one of the plurality of fingertips in the plurality of second hand images; and select a plurality of texts based on the second movement path.

[0015] The present disclosure also provides an input method applicable to an electronic device, the steps of which include: capturing multiple hand images of a user; judging a first gesture of the user based on multiple first hand images among the multiple hand images; generating a virtual keyboard on a virtual plane at a first time point in response to the first gesture being consistent with a start gesture, wherein the virtual plane is generated based on a palm position corresponding to the first gesture; judging a second gesture of the user based on multiple second hand images corresponding to a second time point among the multiple hand images, wherein the first time point is earlier than the second time point; and generating an input instruction corresponding to the typing gesture based on a displacement between the second gesture and the virtual keyboard in response to the second gesture being consistent with a typing gesture.

[0016] In one embodiment of the present invention, the step of generating the virtual keyboard further includes: generating the virtual plane below the palm position based on the palm position corresponding to the first gesture; and generating the virtual keyboard on the virtual plane.

[0017] In one embodiment of the present invention, the step of generating the input instruction corresponding to the typing gesture further includes: calculating a first movement path of each of the multiple fingertips based on the multiple second hand images; and generating the input instruction of a key corresponding to one of the multiple fingertips in response to the first movement path of one of the multiple fingertips being perpendicular to the virtual plane.

[0018] In one embodiment of the present invention, the method further comprises: in response to the second gesture being consistent with one of a plurality of editing gestures, executing an editing function corresponding to the one of the plurality of editing gestures.

[0019] In one embodiment of the present invention, the method further includes: calculating a plurality of hand joint points in the plurality of hand images; and determining the first gesture and the second gesture based on the plurality of hand joint points.

[0020] In one embodiment of the present invention, the method further comprises: calculating multiple fingertip positions in the multiple second hand images; and calculating a key corresponding to each of the multiple fingertip positions in the virtual keyboard based on the multiple fingertip positions and the virtual keyboard.

[0021] In one embodiment of the present invention, the method further comprises: closing the virtual keyboard in response to the second gesture being consistent with a closing gesture.

[0022] In one embodiment of the present invention, the method further includes: in response to the second gesture indicating that the user changes from a hands-out gesture to a hands-closed gesture, determining that the second gesture matches the closing gesture.

[0023] In one embodiment of the present invention, the method further includes: selecting a cursor position based on one of a plurality of fingertip positions in the plurality of second hand images; and generating input content based on the cursor position and the input instruction.

[0024] In one embodiment of the present invention, the method further includes: calculating a second movement path of one of the plurality of fingertips in the plurality of second hand images in response to the second gesture being consistent with a selection gesture; and selecting a plurality of texts based on the second movement path.

[0025] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are intended to provide further explanation of the disclosure as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To make the above and other objects, features, advantages and embodiments of the present disclosure more apparent and understandable, the accompanying drawings are described as follows:

[0027] Figure 1 is a schematic diagram of an input device in a first embodiment of the present disclosure;

[0028] Figure 2 This is a diagram showing how the input device is used in a head-mounted display in some embodiments of the present disclosure;

[0029] Figure 3 This is a flowchart of the operation of the input device in some embodiments of the present disclosure;

[0030] Figure 4 A schematic diagram of a start gesture in some embodiments of the present disclosure;

[0031] Figure 5 A detailed flowchart of determining whether a user's gesture meets the requirements of a start gesture in some embodiments of the present disclosure;

[0032] Figure 6 A detailed flow chart of generating a virtual keyboard in some embodiments of the present disclosure;

[0033] Figure 7A 、 7B , 8, 9A and 9B are schematic diagrams of editing gestures in some embodiments of the present disclosure;

[0034] Figure 10 A detailed flow chart of executing the typing function in some embodiments of the present disclosure;

[0035] Figure 11 A schematic diagram of marking keyboard keys corresponding to fingers on a virtual keyboard in some embodiments of the present disclosure;

[0036] Figure 12 A schematic diagram of fingers typing on a virtual keyboard in some embodiments of the present disclosure;

[0037] Figures 13A to 13C A schematic diagram of a closing gesture in some embodiments of the present disclosure; and

[0038] Figure 14 Flowchart of the input method in the second embodiment of the present disclosure.

[0039] Explanation of symbols:

[0040] 1: Input device

[0041] 12: Processor

[0042] 14: Camera

[0043] U: User

[0044] HMD: Head-mounted display

[0045] OP1~OP9,OP11~OP14,OP21,OP22,OP71~OP73: Operation

[0046] VK: Virtual Keyboard

[0047] G1~G8: Gestures

[0048] D1: Screen

[0049] IR:Indicator

[0050] SB:Search box

[0051] H:Hand

[0052] MV:Moving path

[0053] X, Y, Z: axis

[0054] 200: Input method

[0055] S201~S205: Steps DETAILED DESCRIPTION

[0056] In order to make the description of the present disclosure more detailed and complete, reference may be made to the accompanying drawings and various embodiments described below, in which the same numbers in the drawings represent the same or similar elements.

[0057] Please refer to Figure 1 , which is a schematic diagram of an input device 1 in a first embodiment of the present disclosure. The input device 1 includes a processor 12 and a camera 14. The input device 1 is used to generate a virtual keyboard and execute corresponding functions based on user gestures.

[0058] In some embodiments, the processor 12 may include a central processing unit (CPU), a graphics processing unit (GPU), multiple processors, a distributed processing system, an application specific integrated circuit (ASIC), and / or a suitable computing unit.

[0059] The camera 14 is used to capture images in space, allowing the input device 1 to determine the position of an object in three-dimensional space based on the images. In some embodiments, the camera 14 may include a depth camera for capturing depth images or a camera for capturing multiple planar images. This allows the input device 1 to determine the position of an object in three-dimensional space based on the depth images or a combination of multiple planar images. More specifically, the input device 1 can determine a user's gestures based on the images.

[0060] In some embodiments, the processor 12 calculates a plurality of hand joint points in the plurality of hand images; and the processor 12 determines the first gesture and the second gesture based on the plurality of hand joint points.

[0061] For example, the processor 12 of the input device 1 can utilize an image recognition model to recognize a user's gesture based on the image captured by the camera 14. For example, the image recognition model can identify the locations of hand joints, such as the palm, knuckles, and fingertips, in an image of the hand and construct the user's gesture accordingly.

[0062] Please refer to Figure 2 , which is a diagram illustrating a usage scenario of the input device 1 applied to a head-mounted display (HMD) in some embodiments of the present disclosure. In some embodiments, the input device 1 can be installed in the head-mounted display (HMD). In this way, the user U can control the input device 1 in the head-mounted display (HMD) through specific gestures to display a virtual keyboard and perform functions related to the virtual keyboard. It should be noted that the virtual keyboard can be displayed by the display element of the head-mounted display (HMD).

[0063] It should be noted that the input device 1 can be applied to other scenarios such as a computer host, and for the convenience of explanation, the present disclosure takes a head-mounted display HMD as an example.

[0064] In order to complete the aforementioned functions, the processor 12 of the input device 1 is used to perform the following operations: judging a first gesture of the user based on multiple first hand images among the multiple hand images; generating a virtual keyboard on a virtual plane at a first time point in response to the first gesture being consistent with a start gesture, wherein the virtual plane is generated based on a position corresponding to the first gesture; judging a second gesture of the user based on multiple second hand images corresponding to a second time point among the multiple hand images, wherein the first time point is earlier than the second time point; and generating an input instruction corresponding to the typing gesture based on a displacement between the second gesture and the virtual keyboard in response to the second gesture being consistent with a typing gesture.

[0065] For example, after the processor 12 recognizes the user's hand making a start gesture, it generates a virtual keyboard beneath the user's palm (e.g., the processor 12 controls the head-mounted display (HMD) to display a keyboard image). Next, when the processor 12 recognizes the user's hand making a typing gesture on the virtual keyboard, it determines the triggered key function based on the position of the user's hand movement.

[0066] For details on the operation, please refer to Figure 3 , which is an operation flow chart of the input device 1 in some embodiments of the present disclosure, wherein the input device 1 is used to perform operations OP1 to OP9. In order to complete the above functions, as Figure 3 As shown, first, the processor 12 of the input device 1 executes operation OP1 to determine whether the hands of the user U meet the start gesture based on the first hand image captured by the camera 14 (i.e., the hand image before the virtual keyboard is generated), where the start gesture can be a pre-set specific gesture.

[0067] When the user U's hands present a start gesture, the processor 12 executes operation OP2 to generate a virtual keyboard. Conversely, if the user U's hands do not present a start gesture, the processor 12 continues to execute operation OP1.

[0068] After generating the virtual keyboard, the processor 12 further performs operation OP3 to determine a subsequent gesture of the user U based on the second hand image captured by the camera 14 (ie, the hand image after generating the virtual keyboard).

[0069] In some embodiments, in response to the second gesture meeting one of a plurality of editing gestures, the processor 12 executes an editing function corresponding to one of the plurality of editing gestures. Specifically, if the gesture of the user U meets the editing gesture (i.e., operation OP4), the processor 12 executes operation OP5 to execute the editing function corresponding to the editing gesture. Specifically, the editing gesture may include specific gestures corresponding to editing functions such as copy, paste, and move the cursor, and when one or both hands of the user U meet the specific gesture, the processor 12 executes the corresponding editing function (for example: copy, paste, move the cursor). Further, after operation OP5, the input device 1 returns to operation OP3 to continue to judge the subsequent gestures of the user U.

[0070] On the other hand, if the user U's gesture matches a typing gesture (i.e., operation OP6), the processor 12 executes operation OP7 to enable the typing function of the virtual keyboard. Specifically, the input device 1 can detect the interaction between the user U's gesture and the virtual keyboard to determine which key on the virtual keyboard the user U has touched. Furthermore, after operation OP7, the input device 1 returns to operation OP3 to continue determining subsequent gestures of the user U.

[0071] In some embodiments, in response to the second gesture being a close gesture, the processor 12 closes the virtual keyboard. Specifically, when the processor 12 determines in operation OP8 that the gesture of the user U is a specific close gesture, the processor 12 executes operation OP9 to close the virtual keyboard to end text editing.

[0072] For the operation of the startup gesture mentioned by OP1, please refer to Figure 4 , which is a schematic diagram of the start gesture G1 in some embodiments of the present disclosure. Figure 4 As shown, the start gesture G1 can be set to maintain the palms of both hands roughly in the same plane and present a posture ready to type. In other words, in response to the input device 1 determining that the two planes formed by the two palms of the user U roughly overlap, it is determined that the gesture of the user U meets the start gesture. When the user U presents the start gesture and maintains it for a period of time (for example: 1 second), the input device 1 can generate a virtual keyboard VK under the hands of the user U. Accordingly, the input device 1 does not need to be based on a specific pattern or physical plane, that is, it can generate the virtual keyboard VK on a virtual plane.

[0073] Please refer to Figure 5 In some embodiments, operation OP1 further includes operations OP11 to OP14.

[0074] First, the processor 12 of the input device 1 executes operation OP11 to set a world coordinate system based on the device's posture. For example, the processor 12 may determine the posture of the input device 1 (which may also be the head-mounted display HMD) using information detected by components such as a gyroscope and inertial sensors in the head-mounted display HMD, and set the world coordinate system with the input device 1 as the origin.

[0075] Next, the processor 12 performs operation OP12 to determine whether the hands of the user U are detected based on the image captured by the camera 14. When the processor 12 detects the hands of the user U, the processor 12 performs operation OP13 to calculate the gesture of the user U based on the world coordinate system.

[0076] Finally, the processor 12 executes operation OP14 to determine whether the gesture of the user U meets the activation gesture (for example: Figure 4 If the user U's gesture matches the start gesture, the processor 12 performs operation OP2. Conversely, if the user U's gesture does not match the start gesture, the processor 12 returns to operation OP13 to determine the user U's subsequent gesture.

[0077] In this way, the processor 12 can determine whether the gesture of the user U meets the activation gesture by operating OP1.

[0078] Please refer to Figure 6 In some embodiments, operation OP2 further includes operations OP21 to OP22.

[0079] First, in operation OP21 , the processor 12 generates the virtual plane located below the palm position based on the palm position corresponding to the first gesture.

[0080] Finally, in operation OP22 , the processor 12 generates the virtual keyboard on the virtual plane.

[0081] For example, when the hands of user U are as shown Figure 4 The starting gesture G1 shown, the processor 12 can calculate the position of the palms of the user U and generate a virtual plane below the palms of the hands (for example, 5 cm below the palms of the hands). It should be noted that the virtual plane can be a horizontal plane, and the tilt angle can be adjusted based on the angle of the user's gesture. For example: a virtual plane parallel to the plane formed by the user's palms is generated. Furthermore, the processor 12 generates a virtual keyboard VK on the virtual plane, so that the virtual keyboard VK is located under the user's hands. In this way, the input device 1 can simulate the usage scenario of typing on a physical keyboard.

[0082] For the editing gestures mentioned in OP4, please refer to Figure 7A 、 7B , 8, 9A and 9B are usage diagrams of editing gestures G2 to G5 in some embodiments of the present disclosure.

[0083] In some embodiments, the processor 12 selects a cursor position based on one of the multiple fingertip positions of a fingertip in the multiple second hand images; and the processor 12 generates an input content based on the cursor position and an operation of the input instruction user on the virtual keyboard.

[0084] First, if Figure 7A As shown, when the user U makes an editing gesture G2 by extending their index finger, the processor 12 can move the cursor to the position of the index fingertip in the article displayed on screen D1, allowing the user U to further enter text at that position. In some embodiments, the input device 1 can also generate an indicator IR at the position pointed by the editing gesture G2 to indicate to the user U the location of the cursor movement.

[0085] In some embodiments, in response to the second gesture being a selection gesture, the processor 12 calculates a second movement path of one of the fingertips in the second hand images; and the processor 12 selects a plurality of texts based on the second movement path.

[0086] like Figure 7B As shown, when the user U makes an editing gesture G3 (ie, a selection gesture) of extending the thumb and index finger, the processor 12 can frame the range of the selected text based on the path of movement of the user U's index finger.

[0087] Then, if Figure 8 As shown, after selecting text, when the user U makes an editing gesture G4 with the palm facing the camera 14 , the processor 12 can copy the previously selected text.

[0088] Next, if Figure 9A As shown, after copying the text, similarly, when the user U performs the editing gesture G2, the processor 12 can move the cursor to the search box SB.

[0089] Finally, if Figure 9B As shown, after moving the cursor, when the user U makes an editing gesture G5 with the back of his hand facing the camera 14 , the processor 12 can paste the previously selected text in the search box SB.

[0090] In this way, the input device 1 can execute the corresponding editing function by recognizing a specific gesture of the user U. It should be noted that the editing gestures described in the above embodiment are only examples, and the present disclosure is not limited thereto. In practice, the input device 1 can be configured with one or more other gestures to trigger the above functions, or further configured with more gestures to execute other functions.

[0091] In some embodiments, the operation of generating the input instruction corresponding to the typing gesture further includes the processor 12 calculating a first movement path of each of the multiple fingertips based on the multiple second hand images; and in response to the first movement path of one of the multiple fingertips being perpendicular to the virtual plane, the processor 12 generates the input instruction of a key corresponding to one of the multiple fingertips.

[0092] For details on typing gestures, please refer to Figure 10 In some embodiments, operation OP7 further includes operations OP71 to OP73.

[0093] First, in operation OP71 , the processor 12 calculates the movement path of the user's U fingertip in the second hand image.

[0094] Next, in operation OP72, the processor 12 determines whether the movement paths of the fingertips are perpendicular to the virtual plane. If the processor 12 determines that the movement path of one of the fingertips is perpendicular to the virtual plane, operation OP73 is executed. Conversely, if the processor 12 determines that the movement path of the fingertips is not perpendicular to the virtual plane, the process returns to operation OP71.

[0095] Finally, in operation OP73, the processor 12 generates an input instruction for the key corresponding to the fingertip.

[0096] Specifically, if Figure 11 As shown, in the three-dimensional space defined by the X, Y, and Z axes, the virtual keyboard VK is located in the XY plane (i.e., a virtual plane). In operation OP71, the processor 12 tracks the position of each fingertip of hand H and calculates the movement path MV of the index fingertip accordingly. When the index fingertip of hand H moves back and forth once along the movement path MV, the processor 12 determines in operation OP72 that the movement path MV is parallel to the Z axis and perpendicular to the XY plane. Based on this, the processor 12 can then execute operation OP73 to trigger the key function corresponding to the index finger of hand H.

[0097] Please refer to Figure 12In some embodiments, the input device 1 may further indicate on the virtual keyboard VK the keyboard key corresponding to each finger of the user U. Specifically, the processor 12 calculates a plurality of fingertip positions in the plurality of second hand images; and based on the plurality of fingertip positions and the virtual keyboard, the processor 12 calculates a corresponding key on the virtual keyboard for each of the plurality of fingertip positions.

[0098] For example, the processor 12 can use an image recognition model to track the fingertip position of each finger of the user U, and calculate the projection point of the fingertip position on the virtual plane (i.e., the virtual keyboard VK), and then determine the keyboard key corresponding to each finger based on the projection point.

[0099] like Figure 12 As shown, the four fingers of the right hand H of the user U are respectively located above the H, U, I and L keys, and the input device 1 marks the above four keyboard keys in the virtual keyboard VK to prompt the user U.

[0100] In some embodiments, in response to the second gesture indicating that the user changes from an open-hand gesture to a closed-hand gesture, the processor 12 determines that the second gesture meets the closing gesture.

[0101] For details on how to disable gestures, see Figures 13A to 13C , which is a schematic diagram of closing gestures G6 to G8 in some embodiments of the present disclosure.

[0102] First, if Figure 13A In the gesture G6 shown in FIG, the user U places both hands flat on both sides of the virtual keyboard VK with the palms facing the camera 14. Figure 13B In the gesture G7 shown in FIG, the user U gradually closes his hands, and the input device 1 correspondingly closes the virtual keyboard VK. Figure 13C In the gesture G8 shown in FIG, the user U brings his hands together to complete the closing gesture, and the input device 1 correspondingly closes the virtual keyboard VK to end text editing.

[0103] In this way, the input device 1 can close the virtual keyboard VK by recognizing a specific gesture of the user U. It should be noted that the closing gesture described in the above embodiment is only for example, and the present disclosure is not limited thereto. In practice, the input device 1 can be configured with one or more other gestures to close the virtual keyboard VK.

[0104] In summary, the input device 1 proposed in the present disclosure can generate and close a virtual keyboard on a virtual plane to provide text editing functions by recognizing user gestures, without the need to pre-set a specific pattern or physical plane. Correspondingly, the input device 1 can also perform the key functions of the virtual keyboard by recognizing gestures similar to operating a physical keyboard to provide an intuitive operating experience and reduce the user's learning cost. In addition, the input device 1 can also perform corresponding editing functions by recognizing user gestures to enhance the convenience of text editing.

[0105] Please refer to Figure 14 , which is a flow chart of the input method 200 in the second embodiment of the present disclosure. The input method 200 includes steps S201 to S205. The input method 200 is used to generate a virtual keyboard based on the user's gestures and execute corresponding functions. The input method 200 can be executed by an electronic device (for example: Figure 1 The input device 1 ) is shown.

[0106] First, in step S201 , the electronic device captures a plurality of hand images of a user.

[0107] Next, in step S202 , the electronic device determines a first gesture of the user based on a plurality of first hand images among the plurality of hand images.

[0108] Next, in step S203 , in response to the first gesture being consistent with a start gesture, the electronic device generates a virtual keyboard on a virtual plane at a first time point, wherein the virtual plane is generated based on a palm position corresponding to the first gesture.

[0109] Next, in step S204 , the electronic device determines a second gesture of the user based on a plurality of second hand images corresponding to a second time point among the plurality of hand images, wherein the first time point is earlier than the second time point.

[0110] Finally, in step S205 , in response to the second gesture being consistent with a typing gesture, the electronic device generates an input instruction corresponding to the typing gesture based on a displacement between the second gesture and the virtual keyboard.

[0111] In some embodiments, step S203 further includes the electronic device generating the virtual plane below the palm position based on the palm position corresponding to the first gesture; and the electronic device generating the virtual keyboard on the virtual plane.

[0112] In some embodiments, step S205 further includes the electronic device calculating a first movement path of each of the multiple fingertips based on the multiple second hand images; and in response to the first movement path of one of the multiple fingertips being perpendicular to the virtual plane, the electronic device generating the input instruction of a key corresponding to one of the multiple fingertips.

[0113] In some embodiments, the input method 200 further includes, in response to the second gesture being consistent with one of a plurality of editing gestures, the electronic device executing an editing function corresponding to one of the plurality of editing gestures.

[0114] In some embodiments, the input method 200 further includes the electronic device calculating a plurality of hand joint points in the plurality of hand images; and the electronic device determining the first gesture and the second gesture based on the plurality of hand joint points.

[0115] In some embodiments, the input method 200 further includes the electronic device calculating multiple fingertip positions in the multiple second hand images; and the electronic device calculating a key corresponding to each of the multiple fingertip positions in the virtual keyboard based on the multiple fingertip positions and the virtual keyboard.

[0116] In some embodiments, the input method 200 further includes closing the virtual keyboard by the electronic device in response to the second gesture being a close gesture.

[0117] In some embodiments, the input method 200 further includes, in response to the second gesture indicating that the user changes from an open-hand gesture to a closed-hand gesture, the electronic device determining that the second gesture complies with the closing gesture.

[0118] In some embodiments, the input method 200 further includes the electronic device selecting a cursor position based on one of a plurality of fingertip positions in the plurality of second hand images; and the electronic device generating an input content based on the cursor position and the input instruction.

[0119] In some embodiments, the input method 200 further includes, in response to the second gesture being consistent with a selection gesture, the electronic device calculating a second movement path of one of the plurality of fingertips in the plurality of second hand images; and the electronic device selecting a plurality of texts based on the second movement path.

[0120] In some embodiments, the input method 200 further includes the electronic device generating an indicator at the cursor position to prompt the user.

[0121] In summary, the input method 200 proposed in the present disclosure can generate and close a virtual keyboard on a virtual plane to provide the function of editing text by recognizing the user's gestures, and does not require pre-setting a specific pattern or physical plane. Correspondingly, the input method 200 can also perform the key functions of the virtual keyboard by recognizing gestures similar to operating a physical keyboard to provide an intuitive operating experience and reduce the user's learning cost. In addition, the input method 200 can also perform corresponding editing functions by recognizing the user's gestures to enhance the convenience of text editing.

[0122] Although several embodiments have been described above as examples, the input device and method proposed in this disclosure may also be implemented using other systems, hardware, software, storage media, or any combination thereof. Therefore, the scope of protection of this disclosure should not be limited to the specific implementations described in the embodiments of this disclosure, but should be determined by the appended claims.

[0123] It is obvious to those skilled in the art that various modifications and variations can be made to the structure of the present disclosure without departing from the scope or spirit of the present disclosure. In view of the foregoing, the scope of protection of the present disclosure also covers modifications and variations made within the appended claims.

Claims

1. An input device, characterized in that: Include: a camera for capturing a plurality of hand images of a user; and a processor, communicatively connected to the camera, configured to perform the following operations: determining a first gesture of the user based on a plurality of first hand images among the plurality of hand images; In response to the first gesture being consistent with a start gesture, generating a virtual keyboard on a virtual plane at a first time point, wherein the virtual plane is generated based on a palm position corresponding to the first gesture; determining a second gesture of the user based on a plurality of second hand images corresponding to a second time point among the plurality of hand images, wherein the first time point is earlier than the second time point; and In response to the second gesture being consistent with a typing gesture, an input instruction corresponding to the typing gesture is generated based on a displacement between the second gesture and the virtual keyboard.

2. The input device according to claim 1, wherein The operation of generating the virtual keyboard further comprises: generating the virtual plane below the palm position based on the palm position corresponding to the first gesture; and The virtual keyboard is generated on the virtual plane.

3. The input device according to claim 1, wherein The operation of generating the input instruction corresponding to the typing gesture further includes: Calculating a first movement path of each of the plurality of fingertips based on the plurality of second hand images; as well as In response to the first movement path of one of the plurality of fingertips being perpendicular to the virtual plane, the input instruction of a key corresponding to the one of the plurality of fingertips is generated.

4. The input device according to claim 1, wherein The processor is further configured to perform the following operations: In response to the second gesture being consistent with one of a plurality of editing gestures, an editing function corresponding to the one of the plurality of editing gestures is executed.

5. The input device according to claim 1, wherein The processor is further configured to perform the following operations: Calculating a plurality of hand joint points in the plurality of hand images; and The first gesture and the second gesture are determined based on the multiple hand joints.

6. The input device according to claim 1, wherein The processor is further configured to perform the following operations: calculating a plurality of fingertip positions in the plurality of second hand images; and Based on the multiple fingertip positions and the virtual keyboard, a key corresponding to each of the multiple fingertip positions in the virtual keyboard is calculated.

7. The input device according to claim 1, wherein The processor is further configured to perform the following operations: In response to the second gesture being a close gesture, closing the virtual keyboard.

8. The input device according to claim 7, wherein: The processor is further configured to perform the following operations: In response to the second gesture indicating that the user changes from a hands-out gesture to a hands-closed gesture, it is determined that the second gesture meets the closing gesture.

9. The input device according to claim 1, wherein The processor is further configured to perform the following operations: selecting a cursor position based on one of a plurality of fingertip positions in the plurality of second hand images; and An input content is generated based on the cursor position and the input instruction.

10. The input device according to claim 1, wherein The processor is further configured to perform the following operations: In response to the second gesture being consistent with a selection gesture, calculating a second movement path of one of the plurality of fingertips in the plurality of second hand images; and A plurality of characters are selected based on the second movement path.

11. An input method, characterized in that: Applicable to an electronic device, the steps include: capturing multiple hand images of a user; determining a first gesture of the user based on a plurality of first hand images among the plurality of hand images; In response to the first gesture being consistent with a start gesture, generating a virtual keyboard on a virtual plane at a first time point, wherein the virtual plane is generated based on a palm position corresponding to the first gesture; determining a second gesture of the user based on a plurality of second hand images corresponding to a second time point among the plurality of hand images, wherein the first time point is earlier than the second time point; and In response to the second gesture being consistent with a typing gesture, an input instruction corresponding to the typing gesture is generated based on a displacement between the second gesture and the virtual keyboard.

12. The input method according to claim 11, wherein: The step of generating the virtual keyboard further comprises: generating the virtual plane below the palm position based on the palm position corresponding to the first gesture; and The virtual keyboard is generated on the virtual plane.

13. The input method according to claim 11, wherein: The step of generating the input instruction corresponding to the typing gesture further comprises: Calculating a first movement path of each of the plurality of fingertips based on the plurality of second hand images; and In response to the first movement path of one of the plurality of fingertips being perpendicular to the virtual plane, the input instruction of a key corresponding to the one of the plurality of fingertips is generated.

14. The input method according to claim 11, wherein: Further including: In response to the second gesture being consistent with one of a plurality of editing gestures, an editing function corresponding to the one of the plurality of editing gestures is executed.

15. The input method according to claim 11, wherein: Further including: Calculating a plurality of hand joint points in the plurality of hand images; and The first gesture and the second gesture are determined based on the multiple hand joints.

16. The input method according to claim 11, wherein: Further including: calculating a plurality of fingertip positions in the plurality of second hand images; and Based on the multiple fingertip positions and the virtual keyboard, a key corresponding to each of the multiple fingertip positions in the virtual keyboard is calculated.

17. The input method according to claim 11, wherein: Further including: In response to the second gesture being a close gesture, closing the virtual keyboard.

18. The input method according to claim 17, wherein: Further including: In response to the second gesture indicating that the user changes from a hands-out gesture to a hands-closed gesture, it is determined that the second gesture meets the closing gesture.

19. The input method according to claim 11, wherein: Further including: selecting a cursor position based on one of a plurality of fingertip positions in the plurality of second hand images; and An input content is generated based on the cursor position and the input instruction.

20. The input method according to claim 11, wherein: Further including: In response to the second gesture being consistent with a selection gesture, calculating a second movement path of one of the plurality of fingertips in the plurality of second hand images; and A plurality of characters are selected based on the second movement path.