Input recognition method, device, and storage medium in a virtual scene

The input recognition method using a stereo vision algorithm for calculating fingertip positions in virtual reality systems allows users to interact directly with virtual interfaces, enhancing immersion and reducing hardware requirements.

JP2026063038APending Publication Date: 2026-04-10CHIMETA LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing virtual reality and augmented reality systems require users to interact with controllers or special sensor devices, reducing user immersion and presence.

Method used

An input recognition method using a stereo vision algorithm to calculate the coordinates of fingertips based on hand keypoints in binocular images, allowing interaction with virtual interfaces without additional hardware.

Benefits of technology

Enhances user immersion and reduces hardware costs by enabling input operations directly in virtual scenes using smart devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063038000001_ABST
    Figure 2026063038000001_ABST
Patent Text Reader

Abstract

This invention provides a method, device, and storage medium for input recognition in a virtual scene. [Solution] The input recognition method calculates the coordinates of the fingertips using a stereo vision algorithm based on the position of the recognized hand keypoints, compares the fingertips' coordinates with at least one virtual input interface in the virtual scene, and determines that the user is performing an input operation through the target virtual input interface if the fingertips' position and the target virtual input interface in at least one virtual input interface satisfy the set position rule. With this configuration, the position of the user's fingertips can be calculated using a stereo vision algorithm, eliminating the need for the user to interact with real-world controllers or special sensor devices, further enhancing the immersion and realism of the virtual scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] <Cross - reference to Related Applications> This application incorporates by reference in its entirety the Chinese Patent Application No. 202210261992.8, titled "Input Recognition Method, Device, and Storage Medium in Virtual Scenes", filed on March 16, 2022.

[0002] Embodiments of this application relate to the technical field of virtual reality or augmented reality, and in particular, to an input recognition method, device, and storage medium in virtual scenes.

Background Art

[0003] With the rapid development of related technologies such as virtual reality, augmented reality, and mixed reality, head - mounted smart devices, including smart glasses such as head - mounted virtual reality glasses and head - mounted mixed reality glasses, have been continuously developed, and the user experience has been gradually improved.

[0004] In the prior art, virtual interfaces such as holographic keyboards and holographic screens can be generated using smart glasses, and a controller or a special sensor device can be used to determine whether a user interacts with the virtual interface. Thereby, the user can use a keyboard or a screen in the virtual world.

[0005] However, in this method, since the user needs to interact with a controller or a special sensor device in the real world, the immersion and presence of the user are reduced. Urgent measures are needed to solve this problem.

Summary of the Invention

Problems to be Solved by the Invention

[0006] Embodiments of the present invention provide an input recognition method, device, and storage medium in a virtual scene that enable users to perform input operations without using additional hardware, thereby reducing hardware costs. [Means for solving the problem]

[0007] The embodiments of this application are as follows: A method for recognizing input in a virtual scene applied to smart devices, The present invention provides an input recognition method in a virtual scene, which includes the steps of: recognizing key points of the user's hand in a binocular image captured by a binocular camera; calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular image; comparing the coordinates of the fingertips with at least one virtual input interface in the virtual scene; and determining that the user is performing an input operation through the target virtual input interface if the position of the fingertips and the target virtual input interface in the at least one virtual input interface satisfy a set position rule.

[0008] More preferably, the step of recognizing keypoints of a user's hand in a binocular image captured by a binocular camera includes: detecting the hand region from one of the monocular images in the binocular image using a target detection algorithm; dividing the monocular image into a foreground image corresponding to the hand region; and recognizing the foreground image and obtaining the keypoints of the hand in the monocular image using a pre-configured hand keypoint recognition model.

[0009] More preferably, the step of calculating the coordinates of the fingertips using a stereo vision algorithm based on the position of the key points of the hand in the binocular image includes the steps of determining whether the fingertip joint point of any of the user's fingers is included in the recognized key points of the hand, and if the fingertip joint point of the finger is included in the key points of the hand, calculating the position of the fingertip joint point in the virtual scene using a stereo vision algorithm based on the position of the fingertip joint point in the binocular image and using this as the coordinates of the fingertip.

[0010] More preferably, if the keypoints of the hand do not include the fingertip joint points of the fingers, the method further includes the steps of calculating the bending angle of the finger based on the position of the visible keypoints of the fingers in the binocular image and the phalangeal association features when performing an input operation, and calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the position of the visible keypoints of the fingers in the binocular image.

[0011] More preferably, the finger includes a first phalangeal segment near the palm, a second phalangeal segment connected to the first phalangeal segment, and a fingertip segment connected to the second phalangeal segment, and the step of calculating the bending angle of the finger based on the position of the visible key point of the finger in the binocular image and the phalangeal association features when performing an input operation includes the steps of determining the actual lengths of the first phalangeal segment, the second phalangeal segment, and the fingertip segment of the finger; calculating the observed lengths of the first phalangeal segment, the second phalangeal segment, and the fingertip segment based on the coordinates of the recognized key point of the hand; determining that the bending angle of the finger is less than 90 degrees if the observed length of the second phalangeal segment and / or the fingertip segment is less than the corresponding actual length, and calculating the bending angle of the finger from the observed length and actual length of the second phalangeal segment and / or the observed length and actual length of the fingertip segment; and determining that the bending angle of the finger is 90 degrees if the observed length of the second phalangeal segment and / or the fingertip segment is 0.

[0012] More preferably, the step of calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the position of the visible key point of the finger in the binocular image includes, if the bending angle of the finger is less than 90 degrees, the step of calculating the coordinates of the fingertip of the finger from the position of the starting joint point of the second phalange, the bending angle of the finger, the actual length of the second phalange, and the actual length of the fingertip phalange, and if the bending angle of the finger is 90 degrees, the step of calculating the position of the fingertip from the position of the starting joint point of the second phalange and the distance the first phalange moves to the at least one virtual input interface.

[0013] More preferably, the step of determining that the user is performing an input operation through the target virtual input interface if the position of the fingertip and the target virtual input interface in the at least one virtual input interface satisfy a set position rule includes the steps of determining that the user is touching the target virtual input interface if the position of the fingertip is on the target virtual input interface, and / or determining that the user is clicking the target virtual input interface if the position of the fingertip is on the side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold.

[0014] More preferably, an infrared sensor is attached to the smart device, and the method further includes the steps of: collecting distance values ​​between the infrared sensor and the keypoint of the hand using the infrared sensor; and performing position correction on the calculated position of the user's fingertip using the distance values.

[0015] The embodiments of this application are as follows: It includes memory and a processor. The memory stores one or more computer instructions. The processor further provides a terminal device for performing steps of an input recognition method in a virtual scene by executing one or more computer instructions.

[0016] Embodiments of the present invention further provide a computer-readable storage medium storing a computer program. When executed by a processor, the computer program causes the processor to perform steps of an input recognition method in a virtual scene. [Effects of the Invention]

[0017] In the input recognition method, device, and storage medium in a virtual scene according to the embodiment of the present invention, the coordinates of the fingertips are calculated using a stereo vision algorithm based on the position of the recognized hand keypoints, and these fingertips are compared with at least one virtual input interface in the virtual scene. If the fingertips' position satisfies the position rules set for the target virtual input interface in at least one virtual input interface, it is determined that the user is performing an input operation through the target virtual input interface. With this configuration, the position of the user's fingertips can be calculated using a stereo vision algorithm, eliminating the need for the user to interact with real-world controllers or special sensor devices, and further enhancing the immersion and presence of the virtual scene. [Brief explanation of the drawing]

[0018] To more clearly describe embodiments of the present invention or technical solutions in the prior art, the drawings necessary for describing embodiments or the prior art will be briefly described below. The drawings in the following description show some embodiments of the present invention, and it is clear that those skilled in the art can obtain other drawings based on these drawings without expending any creative effort. [Figure 1] This is a schematic diagram illustrating the flow of an input recognition method as an example of the present invention. [Figure 2]It is a schematic diagram of the key points of the hand shown as an example of the present application. [Figure 3] It is a schematic diagram of target detection shown as an example of the present application. [Figure 4] It is a schematic diagram of foreground image segmentation shown as an example of the present application. [Figure 5] It is a schematic diagram of a stereo vision algorithm shown as an example of the present application. [Figure 6] It is a diagram showing the image formation principle of the stereo vision algorithm shown as an example of the present application. [Figure 7] It is a schematic diagram of the knuckle shown as an example of the present application. [Figure 8] It is a schematic diagram of the calculation of the key point positions of the hand shown as an example of the present application. [Figure 9] It is a schematic diagram of the virtual input interface shown as an example of the present application. [Figure 10] It is a schematic diagram of the parallax of the binocular camera shown as an example of the present application. [Figure 11] It is a schematic diagram of the terminal device shown as an example of the present application.

Embodiments for Carrying out the Invention

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, hereinafter, referring to the drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. However, it is obvious that the described embodiments are only a part of the present invention and do not cover all the embodiments. Based on the embodiments of the present invention, all other embodiments that can be obtained by those skilled in the art without creative efforts are included in the protection scope of the present invention.

[0020] Conventional technologies use smart glasses to generate virtual interfaces such as holographic keyboards and holographic screens, and controllers or special sensor devices are used to determine whether the user has interacted with the virtual interface. This allows the user to use the keyboard or screen in the virtual world. However, even with this method, the user still needs to interact with controllers or special sensor devices in the real world, which can reduce the user's sense of immersion and presence.

[0021] To address the technical challenges described above, several embodiments of this application provide solutions. The technical solutions provided by each embodiment of this application will be described in detail below with reference to the drawings.

[0022] Figure 1 is a schematic flowchart of an input recognition method in a virtual scene shown as an example of the present invention, and as shown in Figure 1, the method includes the following steps 11 to 14.

[0023] Step 11: In the binocular images of the hand captured by the binocular cameras, key points of the user's hand are recognized.

[0024] Step 12: Based on the positions of key points in the binocular images, the coordinates of the fingertips are calculated using a stereo vision algorithm.

[0025] Step 13: Compare the coordinates of the fingertip to at least one virtual input interface in the virtual scene.

[0026] Step 14: If the position of the fingertip and the target virtual input interface in at least one virtual input interface satisfy the established position rule, it is determined that the user is performing an input operation through the target virtual input interface.

[0027] This embodiment can be implemented by a smart device, and examples of smart devices include, but are not limited to, wearable devices such as virtual reality (VR) glasses, mixed reality (MR) glasses, and VR head-mounted displays (HMDs). Taking VR glasses as an example, when the VR glasses display a virtual scene, they can generate at least one virtual input interface within the virtual scene, including at least one virtual input interface such as a virtual keyboard and / or a virtual screen. The user can interact with these virtual input interfaces in the virtual scene.

[0028] In this embodiment, the smart device can use binocular images captured by a binocular camera of the hands. Here, the binocular camera may be mounted on the smart device or in another location capable of capturing both hands, but this embodiment is not limited to this. The binocular camera includes two monocular cameras, and the binocular image consists of two monocular images.

[0029] A smart device may recognize key points on the user's hand in binocular images captured by the binocular cameras. These key points may be at the joints of each of the user's fingers, at the fingertips, or at any other location on the hand, as shown in the schematic diagram in Figure 2.

[0030] After recognizing the keypoints of the user's hand, the coordinates of the fingertips may be calculated using a stereo vision algorithm based on the positions of the keypoints in the binocular images. Here, the stereo vision algorithm, also known as the binocular vision algorithm, is an algorithm that simulates the principles of human vision and passively detects distance using a computer. Its main principle is as follows: Observe a single object from two points, acquire images at various viewing angles, and calculate the object's position using the pixel matching relationship between the images and the principle of triangulation.

[0031] After calculating the fingertip position, the fingertip coordinates may be compared to at least one virtual input interface in the virtual scene. If the fingertip position and the target virtual input interface in at least one virtual input interface satisfy the established position rule, it is determined that the user is performing an input operation through the target virtual input interface. Here, the user's input operation includes at least clicks, long presses, or touches.

[0032] In this embodiment, the smart device calculates the coordinates of the fingertips using a stereo vision algorithm based on the position of the hand's recognized keypoints. These fingertips are then compared to at least one virtual input interface within the virtual scene. If the fingertips' position satisfies the positional rules set for the target virtual input interface within the virtual input interface, it is determined that the user is performing an input operation through the target virtual input interface. This configuration allows the stereo vision algorithm to calculate the position of the user's fingertips, eliminating the need for the user to interact with real-world controllers or special sensor devices, further enhancing the immersion and realism of the virtual scene.

[0033] Furthermore, this embodiment enables a reduction in hardware costs by using a smart device or an existing binocular camera installed in the environment to capture images of the hands, allowing the user to perform input operations without the use of additional hardware.

[0034] In some preferred embodiments, the operation described in the above embodiment, "recognizing key points of the user's hand in binocular images captured by a binocular camera," may be achieved by the following steps.

[0035] As shown in Figure 3, for either monocular image in a binocular image, the smart device may use an object detection algorithm to detect the hand region from the monocular image. Here, the object detection algorithm may be implemented based on a region-convolutional neural network (R-CNN).

[0036] The target detection algorithm will be explained further below.

[0037] For a given picture, the algorithm generates approximately 2000 candidate regions based on the picture, then transforms each candidate region to a predetermined size, and sends the modified candidate regions to a Convolutional Neural Network (CNN) model. Furthermore, a feature vector corresponding to each candidate region can be obtained through this model. Subsequently, the feature vector is sent to a classifier containing multiple categories, and the probability value of the image within the candidate region belonging to each category can be predicted. For example, if the classifier predicts that the images within a total of 10 candidate regions (1-10) have a 95% probability of belonging to the hand region and a 20% probability of belonging to the face region, then candidate regions 1-10 can be detected as the hand region. With such a configuration, a smart device can accurately detect the hand region in any monocular image.

[0038] In real-world scenarios, when a user interacts with a virtual scene using a smart device, the user's hand is often the most prominent object to the user. Therefore, the foreground image in a monocular image captured by one of the cameras is typically the area of ​​the user's hand. As shown in Figure 4, a smart device can separate the foreground image corresponding to the hand area from the monocular image. According to such embodiments, by separating the hand area, the smart device can reduce interference from other areas to subsequent recognition and recognize the hand area as a target, thereby increasing recognition efficiency.

[0039] Based on the steps described above, as shown in Figure 2, the smart device can use a pre-configured hand keypoint recognition model to recognize the foreground image and obtain the hand keypoints in the monocular image. The hand keypoint recognition model may be pre-trained. For example, a single hand image is input to the model, the model recognition result for the hand keypoints is obtained, the model parameters are further adjusted according to the error between the model recognition result and the expected result, and the hand keypoints are recognized again using the model with the adjusted parameters. By repeating this process, the hand keypoint recognition model can accurately recognize the foreground image corresponding to the hand region and obtain the hand keypoints in the monocular image.

[0040] In actual scenarios, when a user performs an input operation, their fingertips may be obstructed by other parts of their hand. As a result, the binocular cameras may not be able to capture the user's fingertips, and the fingertip joint points will be missing from the recognized hand keypoints. Here, the fingertip joint points are shown in Figure 2, 4, 8, 12, 16, and 20. On the other hand, if the user's fingertips can be captured by either of the binocular cameras, the recognized hand keypoints may include the fingertip joint points.

[0041] Preferably, after recognizing the key points of the user's hand as described in the above embodiment, the step of calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular image may be implemented based on the following steps S1 and S2.

[0042] Step S1: For any of the user's fingers, determine whether the fingertip joint point of the finger is included in the recognized keypoints of the hand.

[0043] Step S2: If the keypoints of the hand include the fingertip joint points, the position of the fingertip joint points in the virtual scene is calculated using a stereo vision algorithm based on the position of the fingertip joint points in the binocular images, and this is used as the coordinate of the fingertip.

[0044] The stereo vision algorithm will be explained in detail below with reference to Figures 5 and 6.

[0045] The two rectangles on the left and right sides of Figure 5 represent the camera planes of the two cameras, the left and right, respectively. Point P represents the target object (the fingertip joint point of the user's finger), and P1 and P2 are projections of point P on the two camera planes. The image points on the image-forming planes of the two cameras, located to the left and right of a single point P(X, Y, Z) in world space, are P1(ul, vl) and P2(ur, vr), respectively. These two image points are images of the same object point P in world space (world coordinate system) and are called "conjugate points." The two conjugate image points are connected by projection lines PlOl and P2Or, which connect the respective image points to the optical centers Ol and Or of the cameras. The intersection of these lines becomes the object point P(X, Y, Z) in world space (world coordinate system).

[0046] Specifically, Figure 6 illustrates the principle of simple head-up binocular stereoscopic image formation. Let T be the baseline distance, which is the distance between the connecting lines between the projection centers of the two cameras. The origin of the camera coordinate system is at the optical center of the camera lens, and this coordinate system is shown in Figure 6. The image formation origin of the camera is behind the optical center of the lens, and the left and right image formation planes are located f in front of the optical center of the lens. The u and v axes of this virtual image plane coordinate system O1uv are in the same direction as the x and y axes of the camera coordinate system. This simplifies the calculation process. The origins of the left and right image coordinate systems are at the optical axis of the camera and the focals O1 and O2 of the plane, respectively, and the corresponding coordinates of point P in the left and right images are xl(u1,v1) and xr(u2,v2), respectively. Assuming that the images from the two cameras lie on the same plane, the Y coordinates of the image coordinates of point P are the same, i.e., v1=v2, and from the geometric relationship of the triangle, the following equation is obtained. JPEG2026063038000002.jpg18128

[0047] The above (x,y,z) are the coordinates of point P in the left camera coordinate system, T is the baseline distance, f is the focal length of the two cameras, and (u1,v1) and (u1,v2) are the coordinates of point P in the left and right images, respectively.

[0048] Parallax is defined as the difference d in position between a specific point in two images and its corresponding point. JPEG2026063038000003.jpg20125

[0049] Therefore, the coordinates of point P in the left camera coordinate system are calculated as follows: JPEG2026063038000004.jpg20115

[0050] Based on the process described above, the points on the image-forming planes of the two cameras on the left and right, corresponding to the fingertip joint points (i.e., the positions of the fingertip joint points in the binocular images), are identified. By obtaining the intrinsic and extrinsic parameters of the cameras through camera calibration, the three-dimensional coordinates of the fingertip joint points in the world coordinate system can be determined using the above formula.

[0051] Preferably, a correspondence may be pre-defined between the coordinate system of the virtual scene generated by the smart device and the 3D coordinates in the world coordinate system. Based on this correspondence, the 3D coordinates of the fingertip joint obtained above are converted to the coordinate system of the virtual scene to obtain the position of the fingertip joint in the virtual scene, which is then used as the coordinates of the fingertip.

[0052] According to the above embodiment, even if the fingertip joint is obstructed, the smart device can accurately calculate the coordinates of the user's fingertip using a stereo vision algorithm.

[0053] As shown in Figure 7, a finger includes the first phalangeal segment near the palm, the second phalangeal segment connected to the first phalangeal segment, and the phalangeal segment connected to the second phalangeal segment. Each phalangeal segment of a human finger follows certain bending rules. For example, most people cannot bend the phalangeal segment without moving the second and first phalangeal segments. Also, when the phalangeal segment bends downward by 20°, the second phalangeal segment usually bends by a certain angle in conjunction with the phalangeal segment. The reason for these bending rules is that each phalangeal segment of a human finger is related to others; in other words, there is a characteristic called phalangeal association.

[0054] Based on the above, in some preferred embodiments, if the user's fingertip joint points are obscured, the step of calculating the coordinates of the fingertips using a stereo vision algorithm based on the position of keypoints of the hand in the binocular image may be implemented based on the following steps.

[0055] Step S3: If the keypoints of the hand do not include the fingertip joint points, calculate the finger flexion angle from the binocular image position of the visible keypoints of the fingers and the phalangeal association features when performing input operations.

[0056] Step S4: Calculate the coordinates of the fingertip from the finger bending angle and the position of the visible keypoint of the finger in the binocular image.

[0057] In such embodiments, even when the fingertip phalangeal cover of the finger is obstructed, the coordinates of the fingertip may be calculated using visible keypoints and phalangeal association features.

[0058] In step S3, a visible keypoint refers to a keypoint that can be detected in the binocular image. For example, if the user's little finger is bent at a certain angle and the fingertip joint of the little finger is obscured by the palm, the fingertip joint point of the user's little finger cannot be recognized in this state, meaning that the fingertip joint point of the little finger is an invisible keypoint. Other keypoints of the hand other than this fingertip joint point are recognized normally, meaning that the other keypoints of the hand are visible keypoints. Here, the bending angle of the finger includes the bending angles of one or more phalanges.

[0059] In some preferred embodiments, step S3 described above may be implemented based on the following embodiments.

[0060] The actual lengths of the first, second, and fingertips of each finger are determined, and the observed lengths of the first, second, and fingertips are calculated based on the coordinates of the recognized keypoints of the hand.

[0061] As shown in Figure 8, the observed length is the length of the finger observed from the angle of the binocular camera. This observed length is the length of each phalangeal segment calculated from the key points of the hand, i.e., the projected length relative to the camera. For example, if two key points R1 and R2 of the hand corresponding to the first phalangeal segment of the finger are recognized, the observed length of the first phalangeal segment can be calculated from the coordinates of these two key points.

[0062] Preferably, if the observed length of the second phalanges is less than the actual length corresponding to the second phalanges, or the observed length of the fingertips is less than the actual length corresponding to the fingertips, or if the observed lengths of both the second phalanges and the fingertips are less than the actual lengths corresponding to each, the finger bending angle is determined to be less than 90 degrees. In this case, the finger bending angle can be calculated from the observed length and actual length of the second phalanges, or from the observed length and actual length of the fingertips, or from the observed length of the second phalanges, the actual length of the second phalanges, the observed length of the fingertips, and the actual length of the fingertips.

[0063] The following example illustrates how to calculate the bending angle from the observed length and the actual length, with reference to Figure 8.

[0064] Figure 8 illustrates the phalangeal state when a finger is bent and the corresponding key points R1, R2, and R3, where the first phalangeal is R1R2, the second phalangeal is R2R5, and the fingertip phalangeal is R5R6. As shown in Figure 8, in the triangle formed by R1, R2, and R3, the bending angle a of the first phalangeal can be determined if R2R3 (observed length) and R1R2 (actual length) are known. Similarly, in the triangle formed by R4, R2, and R5, the bending angle b of the second phalangeal can be determined if R2R5 (actual length) and R2R4 (observed length) are known. Similarly, the bending angle c of the third phalangeal can be determined.

[0065] Preferably, if the observed length of the second phalangeal and / or fingertip phalangeal is 0, the binocular camera cannot observe the second phalangeal and / or fingertip phalangeal, and in this case, it may be assumed that the finger bending angle is 90 degrees based on the finger bending characteristics.

[0066] Preferably, based on the bending angle calculation process described above, the step of "calculating the coordinates of the fingertip from the bending angle of the finger and the position of the visible key point of the finger in the binocular image" described in the above embodiment may be implemented based on the following embodiment. Embodiment 1

[0067] If the finger flexion angle is less than 90 degrees, the coordinates of the fingertip may be calculated from the position of the starting joint of the second phalangeal joint, the finger flexion angle, the actual length of the second phalangeal joint, and the actual length of the fingertip phalangeal joint.

[0068] As shown in Figure 8, if the starting joint points R2 and R2 of the second phalanges can be observed, the position of R2 can be calculated using the stereo vision algorithm. If the position of R2, the actual length of the second phalanges, and the bending angle b of the second phalanges are known, the position of the starting joint point R5 of the fingertips can be determined. Furthermore, the position R6 of the fingertips can be calculated from the position of R5, the bending angle c of the fingertips, and the actual length of the fingertips. Embodiment 2

[0069] If the finger is bent at a 90-degree angle, the position of the fingertip is calculated from the position of the starting joint of the second phalangeal segment and the distance the first phalangeal segment moves to at least one virtual input interface.

[0070] Furthermore, if the finger is bent at a 90-degree angle, the user's fingertip may move in the same way as the first phalanges. For example, if the first phalanges move 3 cm downwards, the fingertip will also move 3 cm. Based on this, if the position of the starting joint of the second phalanges and the distance the first phalanges move to at least one virtual input interface are known, the fingertip position is calculated. This problem of calculating the fingertip position is then transformed into a geometric problem of calculating the endpoint position when the starting point position, the direction of movement of the starting point, and the distance of movement are known. This will not be explained in detail here.

[0071] In some preferred embodiments, after calculating the fingertip position, the fingertip position is compared with at least one virtual input interface in the virtual scene, and the user is determined from the comparison result whether or not to perform an input operation. The following will be explained using one of the at least one virtual input interfaces as an example.

[0072] Embodiment 1 If the fingertip is located on the target virtual input interface, it is determined that the user is touching the target virtual input interface.

[0073] Embodiment 2 If the fingertip is positioned away from the user on the target virtual input interface, and the distance to the target virtual input interface is greater than a preset distance threshold, it is determined that the user has clicked the target virtual input interface. Here, the distance threshold may be preset to 1cm, 2cm, 5cm, etc., but this embodiment is not limited to these.

[0074] The two embodiments described above may be carried out individually or in combination, but this embodiment is not limited thereto.

[0075] Preferably, an infrared sensor may be attached to the smart device. After calculating the position of the user's fingertip, the smart device may use the infrared sensor to collect distance values ​​between the infrared sensor and key points on the hand. The calculated position of the user's fingertip can be corrected using these distance values.

[0076] This position correction method reduces the error between the calculated fingertip position and the actual fingertip position, thereby further improving the recognition accuracy of recognizing user input operations.

[0077] The input recognition method described above will be further explained below with reference to Figures 9 and 10 and actual application scenarios.

[0078] As shown in Figure 9, the virtual screen and virtual keyboard generated by VR glasses (smart device) are virtual three-dimensional surfaces, and these surfaces also function as boundaries (i.e., virtual input interfaces). Users can interact with the virtual surfaces and virtual keyboard by appropriately adjusting the position of the virtual surfaces and virtual keyboard and clicking or pressing and pulling action buttons. If the user's fingertip crosses the boundary, it is determined that the user has clicked. If the user's fingertip is on this boundary, it is determined that the user has touched.

[0079] To determine whether the user's fingertips are crossing a boundary, it is necessary to calculate the position of the user's fingertips. This calculation may use at least two cameras located outside the VR glasses. The following describes the case where there are two cameras.

[0080] When a user interacts with VR glasses, their hands are generally in the position closest to the camera. We assume that the two hands of the user are the closest objects to the camera, and that there are no obstacles between the camera and the hands. In addition, as shown in Figure 10, the provision of binocular cameras creates parallax between monocular images of different visual perspectives captured by the two monocular cameras. The VR glasses can then calculate the position of an object based on the pixel matching relationship between the two monocular images and the principle of triangulation.

[0081] If the user's fingertips are not obstructed by other parts of their hand, the VR glasses use a stereo vision algorithm to directly calculate the position of the user's fingertips and determine if the user's fingertips are on a screen, keyboard, or other virtual input interface.

[0082] Typically, each finger has three lines (hereinafter simply referred to as the three-joint line), and the finger is divided into three parts by this three-joint line: the first phalanx near the palm, the second phalanx connected to the first phalanx, and the fingertip phalanx connected to the second phalanx. In addition, there is a bending correlation (phalanx association feature) between each phalanx of the user.

[0083] Based on the above, if the user's fingertips are obscured by other parts of the hand, the terminal device can determine the actual lengths of the user's first, second, and fingertips, which may be pre-measured based on the three-joint lines of the hand. In a real-world scenario, the back of the user's hand and the first joint are usually visible, and the VR glasses can calculate the positions of the second and fingertips from the bending correlation, the observed length of the first joint, and the actual length, and further calculate the coordinates of the fingertips.

[0084] For example, if only the first phalangeal segment is visible, and the second and fingertips are not visible at all, we can assume the user's finger is bent at a 90° angle. This means that the distance the first phalangeal segment moves downward is equal to the distance the fingertips move downward. Based on this, the position of the fingertip can be calculated.

[0085] After calculating the fingertip position, it may be compared to the position of the screen, keyboard, and other virtual input interfaces. If the fingertip extends beyond the virtual input interface but does not exceed a preset touch depth, the user is determined to be performing a click operation. If the fingertip is on the virtual input interface, the user is determined to be performing a touch operation. If the fingertip extends beyond the virtual input interface but does not exceed a preset touch depth, the user is determined to be performing a cancel operation.

[0086] Preferably, an infrared sensor is attached to the VR glasses. The infrared sensor obtains the distance between the infrared sensor and the fingertip position using the time-of-flight (ToF) method, and by further correcting the fingertip position, the error between the calculated fingertip position and the actual fingertip position can be reduced.

[0087] Figure 11 is a schematic diagram of the structure of a terminal device shown as an example of the present invention, and as shown in Figure 11, the terminal device includes a memory 1101 and a processor 1102.

[0088] Memory 1101 is used to store computer programs and is configured to store various other types of data to support operation on the terminal device. This includes instructions for any applications or methods running on the terminal device, contact data, phonebook data, messages, photos, videos, etc.

[0089] Here, memory 1101 implements any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0090] The processor 1102 is coupled to the memory 1101 and executes a computer program in the memory 1101 to recognize key points of the user's hand in the binocular image captured by the binocular camera, calculates the coordinates of the fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular image, compares the coordinates of the fingertips with at least one virtual input interface in the virtual scene, and determines that the user is performing an input operation through the target virtual input interface if the position of the fingertips and the target virtual input interface in the at least one virtual input interface satisfy the set position rules.

[0091] More preferably, when the processor 1102 recognizes keypoints of the user's hand in a binocular image captured by a binocular camera, it is used to detect the hand region from one of the monocular images in the binocular image using a target detection algorithm, to split the monocular image into a foreground image corresponding to the hand region, and to recognize the foreground image and obtain the keypoints of the hand in the monocular image using a pre-configured hand keypoint recognition model.

[0092] More preferably, when the processor 1102 calculates the coordinates of the fingertips using a stereo vision algorithm based on the position of the key points of the hand in the binocular image, it specifically determines whether the fingertip joint point of any of the user's fingers is included in the recognized key points of the hand, and if the fingertip joint point of the finger is included in the key points of the hand, it uses the stereo vision algorithm to calculate the position of the fingertip joint point in the virtual scene according to the position of the fingertip joint point in the binocular image, and uses this as the coordinates of the fingertip.

[0093] More preferably, if the keypoints of the hand do not include the fingertip joint points of the fingers, the processor 1102 is further used to calculate the bending angle of the fingers based on the position of the visible keypoints of the fingers in the binocular image and the phalangeal association features when performing an input operation, and to calculate the coordinates of the fingertips of the fingers from the bending angle of the fingers and the position of the visible keypoints of the fingers in the binocular image.

[0094] More preferably, the finger includes a first phalangeal segment near the palm, a second phalangeal segment connected to the first phalangeal segment, and a fingertip segment connected to the second phalangeal segment, and the processor 1102 calculates the bending angle of the finger based on the position of the visible keypoint of the finger in the binocular image and the segment association features when performing an input operation, specifically by determining the actual lengths of the first phalangeal segment, the second phalangeal segment, and the fingertip segment of the finger, calculating the observed lengths of the first phalangeal segment, the second phalangeal segment, and the fingertip segment based on the coordinates of the recognized keypoint of the hand, and determining that the bending angle of the finger is less than 90 degrees if the observed length of the second phalangeal segment and / or the fingertip segment is shorter than the corresponding actual length. The bending angle of the finger is calculated from the observed length and actual length of the second phalangeal segment and / or the observed length and actual length of the fingertip segment, and further, if the observed length of the second phalangeal segment and / or the fingertip segment is 0, it is determined that the bending angle of the finger is 90 degrees.

[0095] More preferably, when the processor 1102 calculates the coordinates of the fingertip from the bending angle of the finger and the position of the visible key point of the finger in the binocular image, specifically, if the bending angle of the finger is less than 90 degrees, it calculates the coordinates of the fingertip from the position of the starting joint of the second phalange, the bending angle of the finger, the actual length of the second phalange, and the actual length of the fingertip phalange, and if the bending angle of the finger is 90 degrees, it calculates the position of the fingertip from the position of the starting joint of the second phalange and the distance the first phalange moves to the at least one virtual input interface.

[0096] More preferably, the processor 1102 determines that the user is performing an input operation through the target virtual input interface if the position of the fingertip and the target virtual input interface within the at least one virtual input interface satisfy a set position rule. Specifically, if the position of the fingertip is on the target virtual input interface, the processor determines that the user is touching the target virtual input interface, and / or, if the position of the fingertip is on the side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold, the processor determines that the user is clicking the target virtual input interface.

[0097] More preferably, an infrared sensor is attached to the smart device. The processor 1102 can further collect distance values ​​between the infrared sensor and the key points of the hand using the infrared sensor, and perform position correction on the calculated position of the user's fingertips using those distance values.

[0098] The memory in Figure 11 implements all types of volatile or non-volatile storage devices, or combinations thereof, including static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0099] The display 1103 in Figure 11 refers to a screen that includes a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen that receives input signals from the user. The touch panel includes one or more touch sensors that detect touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundary of a touch or slide action, but also the duration and pressure associated with the touch or slide action.

[0100] The audio component 1104 in Figure 11 may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device on which the audio component is located is in an operating mode such as calling mode, recording mode, or speech recognition mode. The received audio signals may further be stored in memory or transmitted via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0101] Furthermore, as shown in Figure 11, the electronic device further includes other components such as a communication component 1105 and a power supply component 1106. Figure 11 shows only some of the components schematically and does not imply that the electronic device includes only the components shown in Figure 11.

[0102] The communication component 1105 in Figure 11 is configured to facilitate wired or wireless communication between the device on which the communication component is located and other devices. The device on which the communication component is located can connect to a wireless network based on a communication standard such as WiFi, 2G, 3G, 4G, or 5G, or a combination thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0103] Of these, the power supply component 1106 supplies power to various components of the device in which the power supply component is located. The power supply component may also include a power management system, one or more power supplies, and other components related to the generation, management, and distribution of power to the device in which the power supply component is located.

[0104] In this embodiment, the terminal device calculates the coordinates of the fingertips using a stereo vision algorithm based on the position of the recognized hand keypoints, compares the fingertips' coordinates with at least one virtual input interface in the virtual scene, and determines that the user is performing an input operation through the target virtual input interface if the fingertips' position and the target virtual input interface in at least one virtual input interface satisfy the set position rules. With this configuration, the position of the user's fingertips is calculated using a stereo vision algorithm, further enhancing the immersion and realism of the virtual scene without the user having to interact with real-world controllers or special sensor devices.

[0105] Accordingly, embodiments of the present application further provide a computer-readable storage medium storing a computer program that implements each step that a terminal device can execute in the embodiments of the above method.

[0106] Those skilled in the art will understand that embodiments of the present invention may be provided as methods, systems, or computer program products. Accordingly, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. The present invention may also take the form of a computer program product implemented on one or more computer-compatible storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-compatible program code.

[0107] The present invention will be described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products relating to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions are provided to the processor of a general-purpose computer, a dedicated computer, an embedded processor, or other programmable data processing device, and the instructions executed by the computer or other programmable data processing device processor can generate means to realize a function specified in one flow of a flowchart or one or more blocks of a set of flows and / or block diagrams.

[0108] These computer program instructions may also be stored in computer-readable memory that can operate a computer or other programmable data processing device in a particular way, thereby generating a product that includes an instruction unit that implements a function specified for one or more flows in a flowchart and / or one or more blocks in a block diagram.

[0109] These computer program instructions may be loaded onto a computer or other programmable data processing device, thereby executing a series of operational steps on the computer or other programmable device to generate a computer implementation process, and thereby the instructions executed on the computer or other programmable device provide steps to realize a function specified in one or more flows of a flowchart and / or one or more blocks of a block diagram.

[0110] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0111] Memory can take various forms, including non-persistent memory on a computer-readable medium, random-access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0112] Computer-readable media include persistent and non-persistent, removable and non-removable media, and information can be stored by any method or technique. Information may be computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital multifunction discs (DVDs) or other optical storage devices, magnetic cartridges, magnetic disks or other magnetic storage devices, or other non-transmission media, which can be used to store information accessible by computing devices. Computer-readable media, as defined herein, do not include transient computer-readable media such as modulated data signals or carriers.

[0113] The terms “includes,” “incorporates,” or any other variations thereof are intended to include non-exclusive inclusions such that a process, method, product, or device containing a set of elements also includes other elements not expressly listed, or elements specific to such a process, method, product, or device. Unless further limited, the elements limited by the phrase “includes…” do not preclude the existence of other identical elements in the process, method, product, or device containing that element.

[0114] The above description is merely an embodiment of the present application and does not limit it. Various modifications and changes are possible for those skilled in the art. Any amendments, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the scope of the claims.

Claims

1. A method for recognizing input in a virtual scene applied to smart devices, The process involves recognizing key points on the user's hand in binocular images captured by a binocular camera, and The steps include: calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of key points in the binocular images; The steps include comparing the coordinates of the fingertip with at least one virtual input interface in the virtual scene, A method for recognizing input in a virtual scene, characterized by comprising the step of determining that the user performs an input operation through the target virtual input interface if the position of the fingertip and the target virtual input interface within the at least one virtual input interface satisfy a set position rule.

2. In binocular images of the hand captured by binocular cameras, the step of recognizing key points on the user's hand is: The steps include: detecting the hand region from one of the monocular images in the aforementioned binocular images using a target detection algorithm; The steps include: dividing the monocular image into a foreground image corresponding to the hand region, The method according to claim 1, comprising the steps of: recognizing the foreground image using a pre-configured hand keypoint recognition model and obtaining the hand keypoints in the monocular image.

3. The step of calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of key points in the hand in the aforementioned binocular images is as follows: The steps include determining whether the fingertip joint point of any of the user's fingers is included in the recognized keypoints of the hand, The method according to claim 1, characterized in that, if the key points of the hand include the fingertip joint points of the fingers, the method includes the step of calculating the position of the fingertip joint points of the fingers in the virtual scene using a stereo vision algorithm based on the position of the fingertip joint points in the binocular image and using this as the coordinate of the fingertip.

4. If the keypoints of the hand do not include the fingertip joint points of the fingers, the steps include calculating the bending angle of the fingers based on the position of the visible keypoints of the fingers in the binocular image and the phalangeal association features when performing an input operation, The method according to claim 3, further comprising the step of calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the position of the visible key point of the finger in the binocular image.

5. The aforementioned finger includes a first phalanx near the palm, a second phalanx connected to the first phalanx, and a fingertip phalanx connected to the second phalanx. The step of calculating the bending angle of the finger based on the position of the visible key point of the finger in the binocular image and the phalangeal association features when performing an input operation is: The steps include determining the actual lengths of the first phalanges, second phalanges, and fingertips of the aforementioned finger, A step of calculating the observed lengths of the first phalanges, the second phalanges, and the fingertips based on the coordinates of the recognized key points of the hand, If the observed length of the second phalangeal segment and / or the fingertip segment is shorter than the corresponding actual length, it is determined that the bending angle of the finger is less than 90 degrees, and the bending angle of the finger is calculated from the observed length and actual length of the second phalangeal segment and / or the observed length and actual length of the fingertip segment. The method according to 4, comprising the step of determining that the bending angle of the finger is 90 degrees if the observed length of the second phalangeal and / or the fingertip phalangeal is 0.

6. The step of calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the position of the visible key point of the finger in the binocular image is: If the bending angle of the finger is less than 90 degrees, the steps include calculating the coordinates of the fingertip from the position of the starting joint of the second phalangeal joint, the bending angle of the finger, the actual length of the second phalangeal joint, and the actual length of the fingertip joint, The method according to 5, characterized in that, when the bending angle of the finger is 90 degrees, the step of calculating the position of the fingertip from the position of the starting joint of the second phalangeal and the distance the first phalangeal is moved to the at least one virtual input interface.

7. The step of determining that the user will perform an input operation through the target virtual input interface if the position of the fingertip and the target virtual input interface within the at least one virtual input interface satisfy the set position rule is as follows: If the position of the fingertip is on the target virtual input interface, it is determined that the user is touching the target virtual input interface, and / or, The method according to claim 1, further comprising the step of determining that the user is clicking the target virtual input interface if the position of the fingertip is on the side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold.

8. An infrared sensor is attached to the aforementioned smart device. The steps include: collecting distance values ​​between the infrared sensor and the key points of the hand using the infrared sensor; The method according to any one of claims 1 to 7, further comprising the step of performing position correction on the calculated position of the user's fingertip using the distance value.

9. A terminal device, including memory and a processor, The memory mentioned above stores one or more computer instructions. The terminal device is characterized in that the processor is for performing the steps of the method according to any one of claims 1 to 8 by executing one or more computer instructions.

10. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, causes the processor to perform the steps of the method described in any one of claims 1 to 8.