Information input method and device, head-mounted display equipment and storage medium
By combining hand images and finger motion data to identify and determine the spatial position of fingertips, the problem of low efficiency and accuracy of information input of existing head-mounted display devices is solved, and more efficient and accurate information input is achieved.
Patent Information
- Application Number
- CN202311812515.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
When entering information, the input efficiency and accuracy of existing head-mounted display devices are low, making it difficult to meet functional requirements.
By obtaining the hand image and finger action data of the user collected by the head-mounted display device, it is possible to identify whether the gesture action of the finger belongs to the preset input action, and determine the spatial position of the fingertip relative to the virtual input interface based on the hand image to realize information input.
It improves the information input efficiency and accuracy of head-mounted display devices, and achieves fast and accurate information input.
Smart Images

Figure CN120215685A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of human-computer interaction of head-mounted display devices, and particularly relates to a method and device for information input of a head-mounted display device, the head-mounted display device, and a storage medium. Background Art
[0002] With the progress of technology, the technology of head-mounted display devices in the virtual display field has become increasingly mature and is widely used in life, such as in scenarios like games, theme parks, movies, education and training, and office work. The human-computer interaction method of head-mounted display devices is crucial for the virtual reality experience.
[0003] Taking an AR (Augmented Reality) glasses as an example, currently, the human-computer interaction of AR glasses can be achieved through a camera provided on the AR glasses. Specifically, the camera captures the gesture depth image of the wearer of the AR glasses in real time, and based on the depth information in the depth image, the gestures of the wearer are recognized, especially focusing on the recognition of the common gesture click action, so as to achieve the information input of the AR glasses. However, since the click action is a very small action, it has high requirements for the response time of the camera and the accuracy of the depth information. In fact, the camera performs poorly in both the response time and the accuracy requirements of the depth information, making it difficult to meet the functional requirements of the information input of the AR glasses, resulting in a decrease in input efficiency and input accuracy.
[0004] Therefore, how to improve the information input efficiency and accuracy of head-mounted display devices has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method and device for information input of a head-mounted display device, the head-mounted display device, and a storage medium, aiming to improve the information input efficiency and accuracy of the head-mounted display device.
[0006] To achieve the above object, this application provides a method for information input of a head-mounted display device, and the method for information input of the head-mounted display device includes:
[0007] Obtain the hand image of the user collected by the head-mounted display device, and detect the motion data of the fingers of the user's hand;
[0008] According to the motion data, determine whether the gesture action of the finger belongs to a preset input action;
[0009] When the gesture action belongs to a preset input action, determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position.
[0010] In addition, to achieve the above object, the present application further provides an information input device for a head-mounted display device, and the information input device for the head-mounted display device includes;
[0011] An acquisition module, configured to acquire a hand image of a user collected by the head-mounted display device, and detect motion data of fingers of the user's hand;
[0012] A first determination module, configured to determine whether the gesture action of the finger belongs to a preset input action according to the motion data;
[0013] A second determination module, configured to, when the gesture action belongs to a preset input action, determine a spatial position of the finger tip relative to a virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position.
[0014] In addition, to achieve the above object, the present application further provides a head-mounted display device, and the head-mounted display device includes a processor, a memory, and a computer program stored on the memory and executable by the processor, where when the computer program is executed by the processor, the steps of the above information input method for the head-mounted display device are implemented.
[0015] In addition, to achieve the above object, the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program, where when the computer program is executed by a processor, the steps of the above information input method for the head-mounted display device are implemented.
[0016] The present application discloses an information input method, device, head-mounted display device, and storage medium for a head-mounted display device. The information input method for the head-mounted display device acquires a hand image of a user collected by the head-mounted display device, and detects motion data of fingers of the user's hand; determines whether the gesture action of the finger belongs to a preset input action according to the motion data; and when the gesture action belongs to a preset input action, determines a spatial position of the finger tip relative to a virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position. In this way, by combining the hand image of the user and the motion data of the fingers of the user's hand, accurate positioning of the finger tip performing the input action is achieved, thereby quickly and accurately implementing information input of the head-mounted display device, and thus improving the information input efficiency and accuracy of the head-mounted display device. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of the steps of an information input method for a head-mounted display device provided by an embodiment of the present application;
[0019] Figure 2 It is an example diagram of setting an image acquisition device on an AR glasses provided by an embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of hand key point recognition provided by an embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of gesture classification provided by an embodiment of the present application;
[0022] Figure 5 It is an example diagram of a palm input scenario of a head-mounted display device provided by an embodiment of the present application;
[0023] Figure 6 It is a schematic block diagram of an information input device of a head-mounted display device provided by an embodiment of the present application;
[0024] Figure 7 It is a schematic block diagram of the structure of a head-mounted display device provided by an embodiment of the present application. Detailed implementation manners
[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0026] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.
[0027] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0028] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0029] For AR glasses, the gesture depth image of the AR glasses wearer can be captured in real time through a camera, and the gestures of the wearer can be recognized based on the depth information in the depth image, especially focusing on the recognition of the common gesture click action to achieve information input of the AR glasses. However, since the click action is a very small action, it has high requirements for the response time of the camera and the accuracy of the depth information. In fact, the camera performs poorly in terms of response time and accuracy requirements of the depth information, making it difficult to meet the functional requirements of information input of the AR glasses, resulting in a decrease in input efficiency and input accuracy.
[0030] To solve the above problems, embodiments of this application provide a method and device for information input of a head-mounted display device, a head-mounted display device, and a storage medium. The method for information input of the head-mounted display device combines the hand image of the user and the motion data of the user's hand fingers, realizes accurate positioning of the fingertips of the fingers performing the input action, and thus quickly and accurately realizes information input of the head-mounted display device, thereby improving the information input efficiency and accuracy of the head-mounted display device.
[0031] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for information input of a head-mounted display device provided by an embodiment of this application. This method can be applied to a head-mounted display device, where the head-mounted display device includes but is not limited to devices such as AR glasses and VR (Virtual Reality) glasses.
[0032] As Figure 1 shown, the method for information input of the head-mounted display device specifically includes steps S101 to S104.
[0033] S101. Obtain the hand image of the user collected by the head-mounted display device and detect the motion data of the user's hand fingers.
[0034] Embodiments of this application are applicable to the information input application scenario of a head-mounted display device. The application environment of embodiments of this application includes a head-mounted display device and a wearable device.
[0035] Among them, the head-mounted display device can be built with an image acquisition device, which is used to collect images of the user's (i.e., the wearer's) hand in real time. It should be noted that the image acquisition device can be a monocular camera. Due to advantages such as small size, low weight, and low power consumption, the monocular camera is more suitable for integration in the head-mounted display device, and it has a low cost, fast response speed, and can meet the visual function requirements in the input application scenario of the head-mounted display device in the embodiments of the present application.
[0036] Exemplarily, as Figure 2 shown, taking the AR glasses as an example, the image acquisition device can be arranged at the frame of the AR glasses.
[0037] The wearable device is an input interaction external device of the head-mounted display device, and can include intelligent finger devices such as finger rings and finger clips, which are worn on the user's hand on the finger for input operations in the virtual input interface of the head-mounted display device.
[0038] For the virtual input interface of the head-mounted display device, in the virtual input interface, interactive display interfaces for information input such as a virtual keyboard, a chat interface, and an information search interface can be displayed. For the sake of understanding, the following virtual input interface can be understood as the interface where the virtual keyboard is located.
[0039] The wearable device can be built with an inertial measurement device. The inertial measurement device can be an IMU (Inertial Measurement Unit), and the IMU is a sensor used to measure the angular acceleration and linear acceleration of an object and can identify the pose of the object. Therefore, the inertial measurement device is used to detect the angular acceleration and linear acceleration of the finger wearing the wearable device in real time.
[0040] It can be understood that the head-mounted display device and the wearable device are established with a communication connection. When the head-mounted display device starts the information input mode, the image acquisition device continuously tracks and collects the user's hand image in real time, and the inertial measurement device continuously detects the angular acceleration and linear acceleration of the wearable device in real time. It should be noted that the image acquisition device continuously tracks and collects the user's hand image in real time, and the inertial measurement device continuously detects the angular acceleration and linear acceleration of the wearable device in real time, which are carried out synchronously.
[0041] In the embodiments of the present application, through the image acquisition device and the inertial measurement device, the information input of the head-mounted display device is quickly realized, and the accuracy of information input is improved.
[0042] The technical solutions provided in the embodiments of the present application will be described in detail below.
[0043] Exemplarily, when the head-mounted display device starts the information input mode, the head-mounted display device continuously and real-time tracks and captures the user's hand image through a monocular camera, and at the same time detects the angular acceleration and linear acceleration of the finger through the IMU built in the wearable device worn on the user's finger, and uses the angular acceleration and linear acceleration as the motion data of the finger.
[0044] S102. Determine whether the gesture action of the finger belongs to a preset input action according to the motion data.
[0045] After that, according to the motion data of the finger wearing the wearable device detected by the IMU, it is determined whether the gesture action of the finger belongs to a preset input action.
[0046] In some embodiments, the preset input action includes a click action. Step S102 may be to determine whether the gesture action of the finger belongs to a click action according to the angular acceleration and linear acceleration; in the case where the gesture action belongs to a click action, it is determined that the gesture action belongs to a preset input action.
[0047] Exemplarily, according to the angular acceleration and linear acceleration detected by the IMU, it is determined whether the finger of the user wearing the wearable device clicks on the virtual input interface of the head-mounted display device.
[0048] Exemplarily, for example, the angular acceleration and linear acceleration detected by the IMU can be integrated once to obtain the speeds of the finger wearing the wearable device at different times, and the angular acceleration and linear acceleration can also be integrated twice to obtain the spatial displacements of the finger wearing the wearable device at different times. The obtained speeds and spatial displacements are also used as motion data, that is, the motion data may include angular acceleration, linear acceleration, speed and spatial displacement. Finally, according to the angular acceleration, linear acceleration, speed and spatial displacement, it is determined whether the gesture action of the finger wearing the wearable device belongs to a click action. In the case of belonging to a click action, it is determined that the gesture action of the finger wearing the wearable device belongs to an input action on the virtual input interface of the head-mounted display device.
[0049] In this way, based on the angular acceleration and linear acceleration detected by the IMU, it can be quickly determined that the finger wearing the wearable device has performed an input action on the virtual input interface of the head-mounted display device.
[0050] S103. In the case where the gesture action belongs to a preset input action, determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position.
[0051] When the gesture action of the finger wearing the wearable device belongs to an input action, determine the spatial position of the fingertip of the finger wearing the wearable device relative to the virtual input interface of the head-mounted display device according to the user's hand image, so as to input information according to the spatial position.
[0052] In some embodiments, S103 may be to determine the position area of the finger from the hand image; perform key point recognition processing on the position area of the finger to determine the spatial position of the fingertip of the finger relative to the virtual input interface of the head-mounted display device; use the input button in the virtual input interface indicated by the spatial position as the input of the head-mounted display device.
[0053] The position area of the finger wearing the wearable device can be determined from the user's hand image, and then the position area of the finger wearing the ring is subjected to key point recognition processing through a preset key point recognition network, so as to determine the spatial position of the fingertip of the finger wearing the ring relative to the virtual input interface of the head-mounted display device. Furthermore, the input button in the virtual input interface indicated by the spatial position is used as the input of the head-mounted display device, and the character corresponding to the input button is the input content of the head-mounted display device.
[0054] Exemplarily, please refer to Figure 3 , Figure 3 which is a schematic diagram of hand key point recognition. Taking the finger wearing the wearable device as the index finger as an example, as Figure 3 shown, the position area of the index finger can be subjected to key point recognition processing through a pre-trained MediaPipe gesture recognition model, so as to extract the 2D coordinate information of key point 8 in Figure 3 as the position of the fingertip of the finger wearing the wearable device. The position of the fingertip is the spatial position relative to the virtual input interface of the head-mounted display device, and the character of the input button corresponding to the fingertip position is the input content of the head-mounted display device.
[0055] In this way, based on the hand image collected in real time by the image acquisition device and the finger motion data of the finger wearing the ring detected in real time by the inertial measurement device, the information input of the head-mounted display device is realized jointly, quickly and accurately.
[0056] The information input method of the head-mounted display device provided in the above embodiments acquires the hand image of the user collected by the head-mounted display device and detects the motion data of the fingers of the user's hand; according to the motion data, it is determined whether the gesture motion of the finger belongs to a preset input motion; when the gesture motion belongs to the preset input motion, according to the hand image, the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device is determined, so as to perform information input according to the spatial position. In this way, by combining the hand image of the user and the motion data of the fingers of the user's hand, accurate positioning of the finger tip performing the input motion is achieved, thereby quickly and accurately realizing the information input of the head-mounted display device, and thus improving the information input efficiency and accuracy of the head-mounted display device.
[0057] In some embodiments, after step S101, it may further include: anchoring a virtual input interface at the palm position of the user's hand according to the hand image.
[0058] The embodiments of the present application can be specifically applied to the palm keyboard input scenario of the head-mounted display device.
[0059] Palm keyboard input refers to the way that the head-mounted display device recognizes and tracks the single hand palm of the user through the image acquisition device, and then anchors the virtual input interface on the palm through the rendering projection method, so that the user uses the finger tips of the other hand to click and touch the virtual input interface anchored on the palm to achieve text input, which can be simply referred to as palm input.
[0060] Exemplarily, for the convenience of distinction, the hand that uses the palm as the medium for presenting the virtual input interface is defined as the first hand (also called the keyboard hand); the other hand that performs finger tip clicking and touching is defined as the second hand.
[0061] In order to implement the palm input of the head-mounted display device, when the head-mounted display device starts the information input mode, the head-mounted display device determines the palm position of the user's first hand according to the hand image of the user collected by the monocular camera in real time, and anchors the virtual input interface at the palm position for the user's second hand to touch.
[0062] Specifically, first determine the palm position of the user's first hand according to the hand image of the user.
[0063] Exemplarily, gesture detection can be performed on the user's hand image to determine the position area of the user's hand in the hand image. Then, gesture classification and recognition are performed on the position area of the user's hand in the hand image, and a hand with the palm completely exposed in the gesture form in the hand image is determined as the first hand. After that, key point recognition processing is performed on the position area of the first hand in the hand image to determine the palm position of the first hand. In this way, by performing gesture detection, gesture classification, and key point recognition on the user's hand image, the palm position of the first hand as the keyboard hand can be quickly and accurately located.
[0064] In some embodiments, a preset gesture detection network can be used to perform gesture detection on the user's hand image to determine the position area of the user's hand in the hand image. It can be understood that the preset gesture detection network is a gesture detection network pre-trained with hand image samples.
[0065] Exemplarily, the preset gesture detection network can be an object detection network such as the YOLO (You Only Look Once) network or the SSD network pre-trained with hand image samples. The training process will not be elaborated here.
[0066] Taking the preset gesture detection network as the pre-trained YOLO network as an example, YOLO is an object detection algorithm. The YOLO network structure is composed of a convolutional neural network (abbreviated as CNN), and can directly perform predictions on the entire image, thereby directly predicting the category and position of the target.
[0067] Specifically, the user's hand image can be directly input into the pre-trained YOLO network. The pre-trained YOLO network extracts the features of the hand image, uses feature maps of multiple scales to detect the user's hand, then converts the feature map into the bounding box and class probability prediction of the user's hand, and applies non-maximum suppression to filter overlapping bounding boxes, and finally directly outputs the position area of the user's hand.
[0068] In this way, through the preset gesture detection network, the position area of the user's hand can be quickly and accurately recognized from the user's hand image.
[0069] In some embodiments, a preset gesture classification network can be used to perform gesture classification and recognition on the position area of the user's hand in the hand image, so as to determine a hand with the palm completely exposed in the gesture form in the hand image as the first hand.
[0070] It can be understood that the preset gesture classification network is a gesture classification network pre-trained with gesture image samples. The training process will not be elaborated here.
[0071] Exemplarily, the preset gesture classification network can be a pre-trained mobilenet network, shufflenet network, etc. using gesture image samples. The training process will not be elaborated here.
[0072] Taking the preset gesture classification network as the pre-trained mobilenet network as an example, the mobilenet network is a convolutional neural network with a small size, less computational complexity, higher accuracy, and faster speed, which is very suitable for head-mounted display devices and can achieve target classification.
[0073] Specifically, the position area of the user's hand in the hand image can be input into the pre-trained mobilenet network. The pre-trained mobilenet network extracts features from the position area of the user's hand, and then performs gesture classification prediction on the extracted features, directly outputting the gesture form corresponding to the position area of the user's hand.
[0074] In some embodiments, in order to improve the accuracy of recognizing the first hand, when the gesture form of one hand of the user is recognized from the user's hand image and the palm is completely exposed and lasts for a preset duration (for example, lasts for two seconds), this hand is determined as the first hand.
[0075] Exemplarily, please refer to Figure 4 , Figure 4 as a gesture classification schematic diagram. When the gesture form recognized from the user's hand image is Figure 4 palm in
[0076] and lasts for a period of time, the corresponding hand is determined as the first hand.
[0077] In some embodiments, the position area of the first hand in the hand image can be recognized for key points by a preset key point recognition network to determine the palm position of the first hand.
[0078] Exemplarily, the preset hand key point recognition network can be a pre-trained MediaPipe gesture recognition model using gesture image samples. The training process will not be elaborated here.
[0079] Specifically, the position area of the first hand in the hand image is input into a pre-trained MediaPipe gesture recognition model to extract the square bounding box of the position where the palm of the first hand is located; then, with the square bounding box of the position where the palm of the first hand is located as the fixed position, the bounding box of the entire first hand is cut out; then, according to the bounding box of the entire first hand, precise key point positioning of 21 2D knuckle coordinates within the first hand area is obtained, and finally, the 2D coordinate information of the key points belonging to the palm of the first hand is extracted to determine the palm position of the first hand.
[0080] Exemplarily, as Figure 3 shown, Figure 3 is a schematic diagram of hand key point recognition. Finally, the 2D coordinate information of key points 0, 1, 2, 5, 9, 13, and 17 in Figure 3 is extracted to determine the palm position of the first hand.
[0081] In this way, through the preset key point recognition network, the palm position of the first hand as a keyboard hand can be quickly and accurately located.
[0082] After determining the palm position of the user's first hand, the virtual input interface is anchored at the palm position of the user's first hand through the rendering projection method, realizing the anchoring of the virtual input interface at the palm position of the user's first hand, thereby increasing the display area of the virtual input interface in the user's vision and enhancing the display stability of the virtual input interface in the user's vision. In this way, the user's second hand can perform touch input operations on the virtual input interface.
[0083] Exemplarily, according to the angular acceleration and linear acceleration detected by the IMU, it is determined whether the finger wearing the ring on the user's second hand clicks on the virtual input interface. For example, the angular acceleration and linear acceleration detected by the IMU can be integrated once to obtain the speed of the finger wearing the ring at different times, and the angular acceleration and linear acceleration can also be integrated twice to obtain the spatial displacement of the finger wearing the ring at different times. The obtained speed and spatial displacement are also used as finger motion data, that is, the finger motion data can include angular acceleration, linear acceleration, speed, and spatial displacement. Finally, according to the angular acceleration, linear acceleration, speed, and spatial displacement, it is determined whether the finger wearing the ring clicks on the virtual input interface.
[0084] In this way, based on the angular acceleration and linear acceleration detected by the IMU, the click action of the finger wearing the ring on the virtual input interface can be quickly determined.
[0085] When it is determined that the finger wearing the ring makes a fingertip click action on the virtual input interface, the input button touched by the fingertip of the finger wearing the ring is determined. The position area of the finger wearing the ring can be recognized from the position area of the user's hand in the user's hand image, and then through a preset key point recognition network, key point recognition processing is performed on the position area of the finger wearing the ring, so as to determine the input button touched by the fingertip of the finger wearing the ring, and the character corresponding to the input button is the input content of the head-mounted display device.
[0086] In this way, based on the hand image collected in real time by the image acquisition device and the finger motion data of the finger wearing the ring detected in real time by the inertial measurement device, the palm input of the head-mounted display device is realized jointly, quickly and accurately.
[0087] In some embodiments, in order to determine the input button more quickly and accurately, the designated finger of the user's second hand can also be defaulted as the finger wearing the ring. Exemplarily, for example, the head-mounted display device can pre-generate a prompt message for the user to wear the ring on the index finger of the second hand to prompt the user to wear the ring on the index finger of the second hand. When it is determined that the user has completed wearing the ring on the index finger of the second hand, the head-mounted display device starts the input mode. In this way, subsequently, the head-mounted display device can directly perform key point recognition processing on the position area of the user's hand in the user's hand image through a preset key point recognition network, so as to determine the input button touched by the fingertip of the user's index finger, simplifying the process of determining the input button, and thus further improving the efficiency and accuracy of determining the input button.
[0088] For a better understanding of the above embodiments, please refer to Figure 5 , Figure 5 which is an example diagram of the palm input scenario of the head-mounted display device. As shown in Figure 5 , taking the virtual input interface as a virtual keyboard as an example, the user's left hand is the first hand as the keyboard hand, and the user's right hand is the second hand, and its index finger wears a ring as the clicking finger. Figure 5 It is defaulted that the user's left hand is open and the palm is facing up, and the index finger of the right hand wears a ring and makes a clicking gesture (not shown in the figure). The following will describe the implementation process of the palm input of the head-mounted display device in combination with Figure 5 .
[0089] When the head-mounted display device starts the palm input mode, the head-mounted display device collects images of the user's two hands in real time through a fast-response monocular camera; then performs gesture detection on the images of the user's two hands through a preset gesture detection network to determine the position areas of the user's two hands in the images of the user's two hands; and then performs gesture classification on the position areas of the user's two hands through a preset gesture classification network. Among them, the gesture categories are as Figure 4As shown, it includes gestures such as call, dislike, like, palm, etc., and thus determines the left hand with the gesture shape of palm; then, through a preset hand key point recognition network, key points of the user's left hand are recognized to determine the position of the palm of the user's left hand. After that, the virtual keyboard is anchored at the position of the palm of the user's left hand through the rendering projection method, realizing the anchoring of the virtual keyboard at the position of the palm of the user's left hand. Further, the IMU in the ring worn on the right index finger of the user detects the angular acceleration and linear acceleration of the right index finger, and based on the angular acceleration and linear acceleration of the right index finger, it is determined whether the fingertip of the right index finger clicks on the virtual keyboard; in the case where the fingertip of the right index finger clicks on the virtual keyboard, the position area of the fingertip of the right index finger is determined from the images of the user's two hands, and through the preset hand key point recognition network, key points of the position area of the fingertip of the right index finger are recognized to determine the virtual keyboard touched by the fingertip of the right index finger, and the character corresponding to the input key is the input content of the head-mounted display device. Thus, the palm input of the head-mounted display device is realized quickly and accurately.
[0090] The information input method of the head-mounted display device provided in the above embodiment realizes the anchoring of the virtual keyboard at the position of the palm of the user's first hand through the hand image, improving the display stability of the virtual keyboard in the user's vision. Then, combined with the finger movement data of the second hand detected by the inertial measurement device, the information input of the head-mounted display device is realized quickly and accurately, thereby improving the information input efficiency and accuracy of the head-mounted display device.
[0091] Please refer to Figure 6 , Figure 6 An information input device for a head-mounted display device is further provided in an embodiment of the present application.
[0092] As Figure 6 shown, the device 200 includes: an acquisition module 201, a first determination module 202, and a second determination module 203.
[0093] The acquisition module 201 is configured to acquire the hand image of the user collected by the head-mounted display device and detect the movement data of the fingers of the user's hand;
[0094] The first determination module 202 is configured to determine whether the gesture action of the finger belongs to a preset input action according to the movement data;
[0095] The second determination module 203 is configured to, when the gesture action belongs to a preset input action, determine the spatial position of the finger fingertip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to input information according to the spatial position.
[0096] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module and unit can refer to the corresponding processes in the information input method embodiment of the foregoing head-mounted display device, and will not be elaborated herein.
[0097] The device provided by the above embodiment can be implemented in the form of a computer program, and the computer program can run on the head-mounted display device as shown in Figure 7 .
[0098] Please refer to Figure 7 , Figure 7 , which is a schematic block diagram of the structure of a head-mounted display device provided by an embodiment of the present application.
[0099] As shown in Figure 7 , the head-mounted display device includes a processor 310 and a memory 320. Among them, the processor 310 is connected to the memory 320 through a bus, and the bus is, for example, an I2C (Inter-integrated Circuit) bus.
[0100] Specifically, the processor 310 may be a micro-control unit (MCU), a central processing unit (CPU), a digital signal processor (DSP), or the like.
[0101] Specifically, the memory 320 may be a Flash chip, a read-only memory (ROM), a magnetic disk, an optical disc, a USB flash drive, a mobile hard disk, or the like. Various computer programs for the processor 310 to execute are stored in the memory 320.
[0102] Among them, the processor 310 is used to run the computer program stored in the memory, and when executing the computer program, the following steps are implemented:
[0103] Obtain the hand image of the user collected by the head-mounted display device, and detect the motion data of the fingers of the user's hand;
[0104] Determine whether the finger gesture action belongs to a preset input action according to the motion data;
[0105] When the gesture action belongs to a preset input action, determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to input information according to the spatial position.
[0106] In some embodiments, the motion data is detected in real time by an inertial measurement device disposed in the wearable device worn on the finger.
[0107] In some embodiments, the motion data includes angular acceleration and linear acceleration. When the processor 310 determines whether the gesture motion of the finger belongs to a preset input motion according to the motion data, it is configured to:
[0108] Judge whether the gesture motion of the finger belongs to a click motion according to the angular acceleration and the linear acceleration;
[0109] In the case where the gesture motion belongs to a click motion, determine that the gesture motion belongs to a preset input motion.
[0110] In some embodiments, when the processor 310 determines the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position, it is configured to:
[0111] Determine the position area of the finger from the hand image;
[0112] Perform key point recognition processing on the position area of the finger to determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device;
[0113] Use the input button in the virtual input interface indicated by the spatial position as the input of the head-mounted display device.
[0114] In some embodiments, after the processor 310 acquires the hand image of the user collected by the head-mounted display device and detects the motion data of the fingers of the user's hand, it is further configured to:
[0115] Anchoring the virtual input interface at the palm position of the user's hand according to the hand image.
[0116] In some embodiments, when the processor 310 anchors the virtual input interface at the palm position of the user's hand according to the hand image, it is configured to:
[0117] Perform gesture detection processing, gesture classification processing and key point recognition processing on the hand image to anchor the virtual input interface at the palm position of the hand where the user's palm is fully exposed.
[0118] In some embodiments, the hand image is collected in real time by an image acquisition device disposed in the head-mounted display device.
[0119] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the information input method of the head-mounted display device as described above are implemented.
[0120] Among them, the computer-readable storage medium may be the information input device of the head-mounted display device or the internal storage unit of the head-mounted display device described in the foregoing embodiments, such as the information input device of the head-mounted display device or the hard disk or memory of the head-mounted display device. The computer-readable storage medium may also be an external storage device of the information input device of the head-mounted display device or the head-mounted display device, such as a plug-in hard disk, a smart media card (SMC), a secure digital card (SD card), a flash card, etc. equipped on the information input device of the head-mounted display device or the head-mounted display device.
[0121] Since the computer program stored in this storage medium can execute any information input method of the head-mounted display device provided by the embodiments of the present application, the beneficial effects achievable by any information input method of the head-mounted display device provided by the embodiments of the present application can be realized. For details, refer to the foregoing embodiments and will not be elaborated herein.
[0122] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0123] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application.
Claims
1. An information input method for a head-mounted display device, characterized in that The method includes: Obtaining a hand image of the user collected by the head-mounted display device, and detecting motion data of the fingers of the user's hand; Determining whether the gesture motion of the finger belongs to a preset input motion according to the motion data; When the gesture motion belongs to the preset input motion, determining the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position.
2. The information input method of the head-mounted display device according to claim 1, wherein The motion data is obtained by real-time detection of an inertial measurement device disposed in a wearable device worn on the finger.
3. The information input method of the head-mounted display device according to claim 2, wherein The motion data includes angular acceleration and linear acceleration; determining whether the gesture motion of the finger belongs to the preset input motion according to the motion data includes: Judging whether the gesture motion of the finger belongs to a click motion according to the angular acceleration and the linear acceleration; When the gesture motion belongs to the click motion, determining that the gesture motion belongs to the preset input motion.
4. The information input method of the head-mounted display device according to claim 1, characterized in that, Determining the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position, includes: Determining the position area of the finger from the hand image; Performing key point recognition processing on the position area of the finger to determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device; Using the input button in the virtual input interface indicated by the spatial position as the input of the head-mounted display device.
5. The information input method of the head-mounted display device according to claim 1, wherein After obtaining the hand image of the user collected by the head-mounted display device and detecting the motion data of the fingers of the user's hand, it further includes: Anchoring the virtual input interface at the palm position of the user's hand according to the hand image.
6. The information input method of the head-mounted display device according to claim 5, characterized in that, Anchoring the virtual input interface at the palm position of the user's hand according to the hand image includes: Performing gesture detection processing, gesture classification processing and key point recognition processing on the hand image to anchor the virtual input interface at the palm position of the hand where the user's palm is completely exposed.
7. The information input method of the head-mounted display device according to any one of claims 1 to 6, characterized in that The hand image is obtained by real-time collection of an image collection device disposed in the head-mounted display device.
8. An information input device for a head-mounted display device, characterized in that, The information input device of the head-mounted display device includes: An obtaining module, configured to obtain a hand image of the user collected by the head-mounted display device, and detect motion data of the fingers of the user's hand; A first determining module, configured to determine whether the gesture motion of the finger belongs to a preset input motion according to the motion data; A second determining module, configured to, when the gesture motion belongs to the preset input motion, determine the spatial position of the finger tip relative to the virtual input interface of the head-mounted display device according to the hand image, so as to perform information input according to the spatial position.
9. A head-mounted display device, characterized in that, The head-mounted display device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, the steps of the information input method of the head-mounted display device according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the information input method of the head-mounted display device according to any one of claims 1 to 7 are implemented.