Method, device, and storage medium for input recognition in a virtual scene
The method uses a stereo vision algorithm to calculate fingertip coordinates in virtual reality environments, enabling input recognition without additional hardware, thereby enhancing immersion and realism.
Patent Information
- Application Number
- JP2024555086
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-16
- Filing Date
- 2022-08-10
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2042-08-10
AI Technical Summary
Conventional methods for input recognition in virtual reality environments require users to interact with controllers or special sensor devices, reducing the sense of immersion and realism.
A method for input recognition using a stereo vision algorithm to calculate fingertip coordinates based on hand key points in binocular images, allowing interaction with virtual interfaces without additional hardware.
Enhances immersion and realism by eliminating the need for controllers or special sensors, reducing hardware costs and improving interaction accuracy.
Smart Images

Figure 0007808374000004 
Figure 0007808374000005 
Figure 0007808374000006
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application cites Chinese patent application No. 202210261992.8, entitled "Method, device, and storage medium for input recognition in virtual scenes," filed on March 16, 2022, the entire contents of which are incorporated herein by reference.
[0002] TECHNICAL FIELD Embodiments of the present application relate to the technical field of virtual reality or augmented reality, and in particular to a method, device, and storage medium for input recognition in a virtual scene. [Background technology]
[0003] With the rapid development of related technologies such as virtual reality, augmented reality, and mixed reality, head-mounted smart devices, including smart glasses such as head-mounted virtual reality glasses and head-mounted mixed reality glasses, are being developed one after another, gradually improving user experience.
[0004] Conventional technology uses smart glasses to generate virtual interfaces such as holographic keyboards and holographic screens, and uses controllers or special sensor devices to determine whether a user has interacted with the virtual interface, allowing the user to use the keyboard or screen in the virtual world.
[0005] However, this method requires users to interact with controllers and special sensor devices in the real world, which reduces the sense of immersion and realism. Measures to resolve this issue are urgently needed. Summary of the Invention [Problem to be solved by the invention]
[0006] The embodiments of the present application provide a method, device, and storage medium for input recognition in a virtual scene that enable a user to perform input operations without using additional hardware, thereby reducing hardware costs. [Means for solving the problem]
[0007] The present embodiment is 1. A method for input recognition in a virtual scene applied to a smart device, comprising: Provided is a method for recognizing input in a virtual scene, the method including the steps of: recognizing key points of a user's hand in a binocular image of the hand captured by a binocular camera; calculating the coordinates of fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular image; comparing the coordinates of the fingertips with at least one virtual input interface in the virtual scene; and determining that the user will perform an input operation through the target virtual input interface if the positions of the fingertips and a target virtual input interface in the at least one virtual input interface satisfy a set position rule.
[0008] More preferably, the step of recognizing key points of a user's hands in a binocular image of the hands taken with a binocular camera includes the steps of: detecting a hand area from any monocular image in the binocular image using a target detection algorithm; segmenting a foreground image corresponding to the hand area from the monocular image; and recognizing the foreground image using a preset hand key point recognition model and obtaining hand key points in the monocular image.
[0009] More preferably, the step of calculating the coordinates of the fingertip using a stereo vision algorithm based on the positions of the hand key points in the binocular image includes the steps of: determining, for any of the user's fingers, whether or not the recognized hand key points include a fingertip joint point of the finger; and, if the hand key points include the fingertip joint point of the finger, calculating the position of the fingertip joint point of the finger in the virtual scene using a stereo vision algorithm based on the position of the fingertip joint point in the binocular image, and setting the calculated position as the coordinate of the fingertip.
[0010] More preferably, when the key points of the hand do not include the fingertip joint points, the method further includes the steps of: calculating a bending angle of the finger based on the positions of visible key points of the finger in the binocular image and the phalanx association features when performing an input operation; and calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the positions of visible key points of the finger in the binocular image.
[0011] More preferably, the finger includes a first phalange close to the palm, a second phalange connected to the first phalange, and a tip phalange connected to the second phalange, and the step of calculating a bending angle of the finger based on positions of visible key points of the finger in the binocular image and phalange association features when performing an input operation includes the steps of: determining actual lengths of the first phalange, the second phalange, and the tip phalange of the finger; calculating observed lengths of the first phalange, the second phalange, and the tip phalange according to coordinates of the recognized key points of the hand; determining that the bending angle of the finger is less than 90 degrees if the observed lengths of the second phalange and / or the tip phalange are less than the corresponding actual lengths; and calculating the bending angle of the finger from the observed length and actual length of the second phalange and / or the observed length and actual length of the tip phalange; and determining that the bending angle of the finger is 90 degrees if the observed length of the second phalange and / or the tip phalange is 0.
[0012] More preferably, the step of calculating the coordinates of the fingertip from the bending angle of the finger and the positions of visible key points of the finger in the binocular image includes: when the bending angle of the finger is less than 90 degrees, calculating the coordinates of the fingertip from the position of a start joint point of the second phalanx, the bending angle of the finger, the actual length of the second phalanx, and the actual length of the last phalanx; and when the bending angle of the finger is 90 degrees, calculating the position of the fingertip from the position of a start joint point of the second phalanx and a moving distance of the first phalanx to the at least one virtual input interface.
[0013] More preferably, when the position of the fingertip and a target virtual input interface in the at least one virtual input interface satisfy a set position rule, the step of determining that the user performs an input operation through the target virtual input interface includes the step of determining that the user is touching the target virtual input interface if the position of the fingertip is on the target virtual input interface, and / or determining that the user is clicking on the target virtual input interface if the position of the fingertip is on the side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold.
[0014] More preferably, an infrared sensor is attached to the smart device, and the method further includes a step of collecting distance values between the infrared sensor and key points on the hand using the infrared sensor, and a step of performing position correction on the calculated positions of the user's fingertips using the distance values.
[0015] The present embodiment is a memory and a processor, the memory stores one or more computer instructions; The present invention further provides a terminal device, wherein the processor is for performing steps of a method for input recognition in a virtual scene by executing the one or more computer instructions.
[0016] An embodiment of the present application further provides a computer-readable storage medium having stored thereon a computer program that, when executed by a processor, causes the processor to perform steps of a method for recognizing inputs in a virtual scene. [Effects of the Invention]
[0017] In the method, device, and storage medium for input recognition in a virtual scene according to the embodiments of the present application, a stereo vision algorithm is used to calculate the coordinates of a fingertip based on the positions of the recognized key points of the hand, and the coordinates of the fingertip are compared with at least one virtual input interface in the virtual scene. If the position of the fingertip and the target virtual input interface in the at least one virtual input interface satisfy a set position rule, it is determined that the user is performing an input operation through the target virtual input interface. With this configuration, the position of the user's fingertip can be calculated using a stereo vision algorithm, eliminating the need for the user to interact with a controller or special sensor device in the real world, thereby further improving the sense of immersion and realism of the virtual scene. [Brief explanation of the drawings]
[0018] In order to more clearly describe the embodiments of the present invention or the technical solutions in the prior art, the following will briefly describe the drawings necessary for describing the embodiments or the prior art. The drawings in the following description show some embodiments of the present invention, and it is obvious that those skilled in the art can derive other drawings based on these drawings without exerting creative efforts. [Figure 1] 1 is a flow diagram of an example input recognition method of the present application; [Figure 2]FIG. 1 is a schematic diagram of key points of a hand shown as an example of the present application. [Figure 3] FIG. 1 is a schematic diagram of target detection as an example of the present application. [Figure 4] FIG. 1 is a schematic diagram of foreground image segmentation as an example of the present application. [Figure 5] FIG. 1 is a schematic diagram of a stereo vision algorithm shown as an example in the present application. [Figure 6] 1 illustrates the image formation principle of a stereo vision algorithm presented as an example in the present application; [Figure 7] FIG. 1 is a schematic diagram of a phalanx shown as an example of the present application. [Figure 8] FIG. 1 is a schematic diagram illustrating calculation of keypoint positions of a hand, shown as an example of the present application. [Figure 9] FIG. 1 is a schematic diagram of a virtual input interface shown as an example in the present application. [Figure 10] FIG. 2 is a schematic diagram of the parallax of a binocular camera shown as an example of the present application. [Figure 11] FIG. 1 is a schematic diagram of a terminal device shown as an example of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0019] In order to clarify the objectives, technical solutions and advantages of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the drawings in the embodiments of the present invention. However, it is clear that the described embodiments are only a part of the present invention and do not cover all the embodiments. All other embodiments that can be obtained by those skilled in the art based on the embodiments of the present invention without any creative efforts are included in the protection scope of the present invention.
[0020] Conventional technology uses smart glasses to generate virtual interfaces such as holographic keyboards and screens, and then uses controllers or special sensor devices to determine whether a user has interacted with the virtual interface, allowing the user to use the keyboard or screen in the virtual world. However, this method still requires the user to interact with controllers or special sensor devices in the real world, which reduces the user's sense of immersion and realism.
[0021] In order to solve the above technical problems, some embodiments of the present application provide solutions. The technical solutions according to each embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0022] FIG. 1 is a flow diagram of an input recognition method in a virtual scene shown as an example of the present application. As shown in FIG. 1, the method includes the following steps 11 to 14.
[0023] Step 11: Recognize key points of the user's hand in the binocular image taken with the binocular camera.
[0024] Step 12: Based on the positions of the key points of the hand in the binocular images, calculate the coordinates of the fingertips using a stereo vision algorithm.
[0025] Step 13: Compare the coordinates of the fingertip with at least one virtual input interface in the virtual scene.
[0026] Step 14: If the position of the fingertip and the target virtual input interface in the at least one virtual input interface satisfy the set position rule, it is determined that the user performs an input operation through the target virtual input interface.
[0027] The present embodiment may be executed by a smart device, including, but not limited to, wearable devices such as virtual reality (VR) glasses, mixed reality (MR) glasses, and a VR head-mounted display (HMD). Taking VR glasses as an example, when the VR glasses display a virtual scene, at least one virtual input interface may be generated within the virtual scene, including at least one virtual input interface such as a virtual keyboard and / or a virtual screen. A user may interact with the virtual input interface in the virtual scene.
[0028] In this embodiment, the smart device can use a binocular image of the hands captured by a binocular camera. Here, the binocular camera may be mounted on the smart device or in another position capable of capturing both hands, but this embodiment is not limited thereto. The binocular camera includes two monocular cameras, and the binocular image is composed of two monocular images.
[0029] The smart device may recognize key points of the user's hand in a binocular image captured by a binocular camera, which may be the knuckles of the user's fingers, the fingertips, or any other location on the user's hand, as shown in the schematic diagram of FIG.
[0030] After recognizing the key points on the user's hand, a stereo vision algorithm can be used to calculate the coordinates of the fingertips based on the positions of the key points on the hand in the binocular images. Here, a stereo vision algorithm, also known as a binocular vision algorithm, is an algorithm that simulates the principles of human vision to passively detect distances using a computer. Its main principle is as follows: an object is observed from two points, images are obtained at various viewing angles, and the position of the object is calculated based on the pixel matching relationship between the images and the triangulation principle.
[0031] After calculating the position of the fingertip, the coordinates of the fingertip may be compared with at least one virtual input interface in the virtual scene. If the position of the fingertip and a target virtual input interface in the at least one virtual input interface satisfy a set position rule, it is determined that the user performs an input operation through the target virtual input interface. Here, the input operation by the user includes at least a click, a long press, or a touch.
[0032] In this embodiment, the smart device uses a stereo vision algorithm to calculate the coordinates of the fingertips based on the recognized positions of the key points of the hand, compares the coordinates of the fingertips with at least one virtual input interface in the virtual scene, and determines that the user will perform an input operation through the target virtual input interface if the position of the fingertip and the target virtual input interface in the at least one virtual input interface satisfy a set position rule. This configuration allows the position of the user's fingertips to be calculated using the stereo vision algorithm, eliminating the need for the user to interact with controllers or special sensor devices in the real world, further improving the immersive and realistic feel of the virtual scene.
[0033] Furthermore, in this embodiment, an image of the hand is captured using an existing binocular camera installed in the smart device or the environment, and the user can perform input operations without using additional hardware, thereby reducing hardware costs.
[0034] In some preferred embodiments, the operation of "recognizing key points of a user's hand in a binocular image taken with a binocular camera" described in the previous embodiment may be realized by the following steps.
[0035] As shown in Fig. 3, for any one of the monocular images in the binocular image, the smart device may detect a hand region from the monocular image using a target detection algorithm, where the target detection algorithm may be implemented based on a Region-Convolutional Neural Network (R-CNN).
[0036] The target detection algorithm is further described below.
[0037] For one picture, the algorithm generates approximately 2,000 candidate regions based on the picture, then transforms each candidate region to a predetermined size, and sends the modified candidate regions to a convolutional neural network (CNN) model. A feature vector corresponding to each candidate region can then be obtained through the model. The feature vector is then sent to a classifier containing multiple categories, and the probability that the image in the candidate region belongs to each category can be predicted. For example, if the classifier predicts that the image in a total of 10 candidate regions, candidate regions 1 to 10, has a 95% probability of belonging to a hand region and a 20% probability of belonging to a face region, candidate regions 1 to 10 can be detected as hand regions. This configuration allows the smart device to accurately detect hand regions in any monocular image.
[0038] In a real scene, when a user interacts with a virtual scene using a smart device, the user's hands are often the most visible object to the user. Therefore, the foreground image in a monocular image captured by one of the cameras is typically the user's hand region. Therefore, as shown in FIG. 4, the smart device can segment the foreground image corresponding to the hand region from the monocular image. According to this embodiment, by segmenting the hand region, the smart device can reduce interference from other regions in subsequent recognition and target the hand region, thereby improving recognition efficiency.
[0039] Based on the above steps, as shown in Figure 2, the smart device can use a preset hand keypoint recognition model to recognize the foreground image and obtain hand keypoints in the monocular image. The hand keypoint recognition model can also be pre-trained. For example, a hand image is input into the model to obtain a model recognition result for the hand keypoints. Depending on the error between the model recognition result and the expected result, the model parameters are further adjusted, and the hand keypoints are again recognized using the model with the adjusted parameters. Through this repeated iteration, the hand keypoint recognition model can accurately recognize the foreground image corresponding to the hand region and obtain hand keypoints in the monocular image.
[0040] In a real situation, when a user performs an input operation, the fingertip may be blocked by other parts of the hand. As a result, the binocular camera cannot capture the user's fingertip, and the fingertip joint points are missing from the recognized hand key points. Here, the fingertip joint points are shown as 4, 8, 12, 16, and 20 in Figure 2. On the other hand, if the user's fingertip can be captured by one of the binocular cameras, the fingertip joint points may be included in the recognized hand key points.
[0041] Preferably, after recognizing the key points of the user's hand as described in the above embodiment, the step of calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular images may be realized by the following steps S1 and S2.
[0042] Step S1: For any of the user's fingers, it is determined whether the recognized key points of the hand include the fingertip joint points of the finger.
[0043] Step S2: If the key points of the hand include the fingertip joint points, calculate the positions of the fingertip joint points in the virtual scene using a stereo vision algorithm based on the positions of the fingertip joint points in the binocular images, and use this as the coordinates of the fingertip.
[0044] The stereo vision algorithm will now be described in detail with reference to FIGS.
[0045] The two rectangles on the left and right sides of Figure 5 represent the camera planes of the two cameras on the left and right, respectively. Point P indicates the target object (the knuckle point of the user's finger), and P1 and P2 are the projections of point P on the two camera planes. The image points on the imaging planes of the two cameras on the left and right of a point P(X, Y, Z) in world space are P1(ul, vl) and P2(ur, vr), respectively. These two image points are images of the same object point P in world space (global coordinate system), and are called "conjugate points." The two conjugate image points are connected by projection lines P1Ol and P2Or, which connect the respective image points to the optical centers Ol and Or of the cameras, and their intersection is the object point P(X, Y, Z) in world space (global coordinate system).
[0046] Specifically, Figure 6 illustrates the principle of simple head-up binocular stereoscopic imaging. Let T be the baseline distance, which is the distance between the projection centers of the two cameras. The origin of the camera coordinate system is located at the optical center of the camera lens, and the coordinate system is shown in Figure 6. The camera's image-forming plane is behind the optical center of the lens, and the left and right image-forming planes are located in front of the optical center, f. The u- and v-axes of this virtual image plane coordinate system O1uv are aligned with the x- and y-axes of the camera coordinate system. This simplifies the calculation process. The origins of the left and right image coordinate systems are located at the optical axis of the camera and the plane foci O1 and O2. The corresponding coordinates of point P in the left and right images are xl(u1,v1) and xr(u2,v2), respectively. Assuming the images from the two cameras are on the same plane, the Y-coordinate of point P in the image coordinates is the same, i.e., v1 = v2. The following equation can be obtained from the geometric relationship of a triangle: JPEG0007808374000001.jpg18128
[0047] In the above, (x, y, z) are the coordinates of point P in the left camera coordinate system, T is the baseline distance, f is the focal length of the two cameras, and (u1, v1) and (u1, v2) are the coordinates of point P in the left and right images, respectively.
[0048] Disparity is defined as the difference in position d between a particular point and a corresponding point in the two images. JPEG0007808374000002.jpg20125
[0049] As a result, the coordinates of point P in the left camera coordinate system are calculated as follows: JPEG0007808374000003.jpg20115
[0050] Based on the above process, the points on the image forming planes of the two cameras on the left and right that correspond to the fingertip joint points (i.e., the positions of the fingertip joint points in the binocular images) are identified, and the internal and external parameters of the cameras are obtained through camera calibration, so that the 3D coordinates of the fingertip joint points in the world coordinate system can be determined using the above equations.
[0051] Preferably, a correspondence relationship may be preset between the coordinate system of the virtual scene generated by the smart device and the three-dimensional coordinates in the world coordinate system. Based on this correspondence relationship, the three-dimensional coordinates of the fingertip joint points obtained above are converted into the coordinate system of the virtual scene to obtain the positions of the fingertip joint points in the virtual scene, which are then used as the coordinates of the fingertip of the finger.
[0052] According to the above embodiment, even if the fingertip joint points are occluded, the smart device can accurately calculate the coordinates of the fingertip of the user's finger using a stereo vision algorithm.
[0053] As shown in FIG. 7, a finger includes a first phalanx closest to the palm, a second phalanx connected to the first phalanx, and a third phalanx connected to the second phalanx. Each phalanx of a human finger has a certain bending rule. For example, most people cannot bend the third phalanx without moving the second and first phalanxes. When the third phalanx bends downward by 20°, the second phalanx also bends by a certain angle as the third phalanx bends. The reason for this bending rule is that the third phalanx of a human finger is related to each other, i.e., there is a characteristic of phalanx association.
[0054] Based on the above, in some preferred embodiments, when the user's fingertip joint points are occluded, the step of calculating the coordinates of the fingertip using a stereo vision algorithm based on the positions of the hand key points in the binocular images may be realized based on the following steps:
[0055] Step S3: If the hand key points do not include the fingertip joint points, calculate the finger bending angle from the positions of the visible key points of the fingers in the binocular image and the phalanx association features when performing the input operation.
[0056] Step S4: The coordinates of the fingertip are calculated from the bending angle of the finger and the positions of the visible key points of the finger in the binocular image.
[0057] In such an embodiment, the coordinates of the fingertip may be calculated by visible keypoint and phalanx-associated features even when the finger's phalanx is occluded.
[0058] The visible key points in step S3 refer to key points that can be detected in the binocular image. For example, if a user's little finger is bent at a certain angle and the tip phalanx of the little finger is blocked by the palm, the tip phalanx of the little finger cannot be recognized in this state, i.e., the tip phalanx of the little finger is an invisible key point. Other key points on the hand other than the tip phalanx are recognized normally, i.e., the other key points on the hand are visible key points. Here, the bending angle of the finger includes the bending angle of each of one or more phalanxes.
[0059] In some preferred embodiments, the above step S3 may be realized according to the following embodiments.
[0060] The actual lengths of each of the first, second and tip phalanges of the fingers are determined, and the observed lengths of each of the first, second and tip phalanges are calculated based on the coordinates of the recognized hand keypoints.
[0061] As shown in Figure 8, the observed length is the length of the finger observed from the angle of the binocular camera, and this observed length is the length of each phalanx calculated from the hand keypoints, i.e., the projected length relative to the camera. For example, if two hand keypoints R1 and R2 corresponding to the first phalanx of a finger are recognized, the observed length of the first phalanx can be calculated from the coordinates of these two hand keypoints.
[0062] Preferably, the finger bend angle is determined to be less than 90 degrees if the observed length of the second phalanx is less than the actual length corresponding to the second phalanx, or if the observed length of the distal phalanx is less than the actual length corresponding to the distal phalanx, or if the observed lengths of both the second phalanx and the distal phalanx are less than their respective actual lengths. In this case, the finger bend angle can be calculated from the observed length and actual length of the second phalanx, or from the observed length and actual length of the distal phalanx, or from the observed length of the second phalanx, the actual length of the second phalanx, the observed length of the distal phalanx, and the actual length of the distal phalanx.
[0063] Hereinafter, calculation of the bending angle from the observed length and the actual length will be described by way of example with reference to FIG.
[0064] FIG. 8 exemplarily illustrates the state of the phalanges when a finger is bent and the corresponding key points R1, R2, and R3, where R1 is the first phalange, R2 is the second phalange, and R5 is the third phalange. As shown in FIG. 8, in the triangle formed by R1, R2, and R3, if R2R3 (the observed length) and R1R2 (the actual length) are known, the bending angle a of the first phalange can be calculated. Similarly, in the triangle formed by R4, R2, and R5, if R2R5 (the actual length) and R2R4 (the observed length) are known, the bending angle b of the second phalange can be calculated. Similarly, the bending angle c of the third phalange can be calculated.
[0065] Preferably, if the observation length of the second phalanx and / or the tip phalanx is 0, the binocular camera cannot observe the second phalanx and / or the tip phalanx, and in this case, it may be assumed that the finger bending angle is 90 degrees based on the finger bending characteristics.
[0066] Preferably, based on the above bending angle calculation process, the step of "calculating the coordinates of the fingertip from the bending angle of the finger and the position of the visible key point of the finger in the binocular image" described in the above example may be realized based on the following embodiment. Embodiment 1
[0067] If the bending angle of the finger is less than 90 degrees, the coordinates of the fingertip may be calculated from the position of the starting joint point of the second phalanx, the bending angle of the finger, the actual length of the second phalanx, and the actual length of the tip phalanx.
[0068] As shown in Figure 8, when the starting joint points R2 and R2 of the second phalanx can be observed, the position of R2 can be calculated using a stereo vision algorithm. When the position of R2, the actual length of the second phalanx, and the bending angle b of the second phalanx are known, the position of the starting joint point R5 of the tip phalanx can be determined. Furthermore, the position of the fingertip R6 can be calculated from the position of R5, the bending angle c of the tip phalanx, and the actual length of the tip phalanx. Embodiment 2
[0069] When the finger bending angle is 90 degrees, the position of the fingertip is calculated from the position of the starting joint point of the second phalanx and the distance of movement of the first phalanx to at least one virtual input interface.
[0070] Note that when the finger is bent at a 90-degree angle, the user's fingertip may move in the same manner as the first phalanx. For example, when the first phalanx moves downward by 3 cm, the fingertip also moves 3 cm. Based on this, when the position of the starting joint point of the second phalanx and the movement distance of the first phalanx to at least one virtual input interface are known, the position of the fingertip is calculated. The problem of calculating the fingertip position is transformed into a geometric problem of calculating the end point position when the position of the starting point, the movement direction of the starting point, and the movement distance are known. This will not be described in detail here.
[0071] In some preferred embodiments, after calculating the fingertip position, the fingertip position is compared with at least one virtual input interface in the virtual scene, and whether the user performs an input operation is determined based on the comparison result. The following description will take one of the at least one virtual input interface as an example.
[0072] Embodiment 1 If the position of the fingertip is on the target virtual input interface, it is determined that the user is touching the target virtual input interface.
[0073] Embodiment 2 If the position of the fingertip is on the side of the target virtual input interface away from the user and the distance from the target virtual input interface is greater than a preset distance threshold, it is determined that the user has clicked on the target virtual input interface, where the distance threshold may be preset to 1 cm, 2 cm, 5 cm, etc., but is not limited thereto in this embodiment.
[0074] The above two embodiments may be implemented individually or in combination, but the present embodiment is not limited to this.
[0075] Preferably, the smart device may be equipped with an infrared sensor. After calculating the position of the user's fingertip, the smart device may use the infrared sensor to collect distance values between the infrared sensor and key points on the hand. The calculated position of the user's fingertip can be corrected using the distance values.
[0076] Such a position correction method can reduce the error between the calculated fingertip position and the actual fingertip position, thereby further improving the recognition accuracy for recognizing input operations by the user.
[0077] The above input recognition method will be further described below with reference to FIGS. 9 and 10 and actual application scenes.
[0078] As shown in Figure 9, the virtual screen and virtual keyboard generated by the VR glasses (smart device) are virtual three-dimensional surfaces, which also function as a boundary (i.e., a virtual input interface). The user can interact with the virtual surface and virtual keyboard by appropriately adjusting their positions and clicking or pressing or pulling action buttons. If the user's fingertip crosses the boundary, it is determined that the user is clicking. If the user's fingertip is on the boundary, it is determined that the user is touching.
[0079] To determine whether the user's fingertip is beyond the boundary, the position of the user's fingertip must be calculated. This calculation may be performed using at least two cameras outside the VR glasses. The following describes the case where there are two cameras.
[0080] When a user interacts with the VR glasses, the user's hands are generally closest to the camera. Assume that the user's two hands are the closest objects to the camera, and there are no obstacles between the camera and the hands. In addition, as shown in Figure 10, the binocular camera provides parallax between the monocular images of different perspectives captured by the two monocular cameras. The VR glasses can calculate the position of the object based on the pixel matching relationship between the two monocular images and the triangulation principle.
[0081] If the user's fingertips are not occluded by other parts of the hand, the VR glasses use a stereo vision algorithm to directly calculate the position of the user's fingertips and determine whether the user's fingertips are over the screen, keyboard, or other virtual input interface.
[0082] Typically, each human finger has three lines (hereinafter simply referred to as triad lines), which divide the finger into three parts: the first phalanx closest to the palm, the second phalanx connected to the first phalanx, and the tip phalanx connected to the second phalanx. In addition, there is a bending correlation (phalanx association feature) between each phalanx of the user's finger.
[0083] Based on the above, when the user's fingertip is blocked by other parts of the hand, the terminal device can determine the actual lengths of the user's first phalanx, second phalanx, and first phalanx, which may be pre-measured based on the triad of the hand. In a real scene, the back of the user's hand and the first phalanx are usually visible, and the VR glasses can calculate the positions of the second phalanx and first phalanx from the bending correlation and the observed and actual length of the first phalanx, and further calculate the coordinates of the fingertip.
[0084] For example, if only the first phalanx is visible, and the second and third phalanxes are completely invisible, we can assume that the user's finger is bent 90 degrees. This means that the distance the first phalanx of the finger moves downward is equal to the distance the third phalanx moves downward. Based on this, the position of the fingertip can be calculated.
[0085] After calculating the fingertip position, the fingertip position may be compared with the positions of the screen, keyboard, and other virtual input interfaces. If the fingertip is over the virtual input interface but does not exceed a preset touch depth, it is determined that the user is performing a click operation; if the fingertip is on the virtual input interface, it is determined that the user is performing a touch operation; if the fingertip is over the virtual input interface but does not exceed a preset touch depth, it is determined that the user is performing a cancel operation.
[0086] Preferably, an infrared sensor is attached to the VR glasses, which uses a time-of-flight (ToF) method to obtain the distance between the infrared sensor and the fingertip position, thereby further correcting the fingertip position to reduce the error between the calculated fingertip position and the actual fingertip position.
[0087] FIG. 11 is a structural schematic diagram of a terminal device shown as an example of the present application. As shown in FIG. 11, the terminal device includes: a memory 1101 and a processor 1102.
[0088] The memory 1101 is used to store computer programs and is configured to store various other data to support operation of the terminal device, including instructions for any applications or methods running on the terminal device, contact data, phone book data, messages, photos, videos, etc.
[0089] Here, the memory 1101 may be implemented with any type of volatile or non-volatile storage device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk, or a combination thereof.
[0090] The processor 1102 is coupled to the memory 1101 and executes a computer program in the memory 1101 to recognize key points of a user's hand in a binocular image captured by a binocular camera, calculate the coordinates of the fingertips using a stereo vision algorithm based on the positions of the key points of the hand in the binocular image, compare the coordinates of the fingertips with at least one virtual input interface in the virtual scene, and determine that the user will perform an input operation through the target virtual input interface if the positions of the fingertips and a target virtual input interface in the at least one virtual input interface satisfy a set position rule.
[0091] More preferably, when recognizing key points of a user's hands in a binocular image taken with a binocular camera, the processor 1102 is used to, for any monocular image in the binocular image, detect a hand region from the monocular image using a target detection algorithm, segment a foreground image corresponding to the hand region from the monocular image, and recognize the foreground image using a preset hand key point recognition model to obtain hand key points in the monocular image.
[0092] More preferably, when calculating the coordinates of the fingertips using a stereo vision algorithm based on the positions of the hand key points in the binocular image, the processor 1102 is used to determine, for any of the user's fingers, whether the recognized hand key points include the fingertip joint point of the finger, and if the hand key points include the fingertip joint point of the finger, calculate the position of the fingertip joint point of the finger in the virtual scene using a stereo vision algorithm according to the position of the fingertip joint point in the binocular image, and use the calculated position as the coordinate of the fingertip of the finger.
[0093] More preferably, the processor 1102 is further used for, when the key points of the hand do not include the fingertip joint points of the fingers, calculating a bending angle of the fingers based on the positions of the visible key points of the fingers in the binocular image and the phalangeal association features when performing an input operation, and calculating the coordinates of the fingertips of the fingers from the bending angle of the fingers and the positions of the visible key points of the fingers in the binocular image.
[0094] More preferably, the finger includes a first phalange close to the palm, a second phalange connected to the first phalange, and a tip phalange connected to the second phalange. When calculating the bending angle of the finger based on the positions of visible keypoints of the finger in the binocular image and the phalange association feature when performing the input operation, the processor 1102 specifically determines the actual lengths of the first phalange, the second phalange, and the tip phalange of the finger, calculates the observed lengths of the first phalange, the second phalange, and the tip phalange based on the coordinates of the recognized keypoints of the hand, and determines that the bending angle of the finger is less than 90 degrees if the observed lengths of the second phalange and / or the tip phalange are shorter than the corresponding actual lengths. The bending angle of the finger is calculated from the observed length and actual length of the second phalange and / or the observed length and actual length of the tip phalange, and determines that the bending angle of the finger is 90 degrees if the observed lengths of the second phalange and / or the tip phalange are 0.
[0095] More preferably, when calculating the coordinates of the fingertip of the finger from the bending angle of the finger and the positions of the visible key points of the finger in the binocular image, specifically, if the bending angle of the finger is less than 90 degrees, the processor 1102 calculates the coordinates of the fingertip from the position of the start joint point of the second phalanx, the bending angle of the finger, the actual length of the second phalanx, and the actual length of the first phalanx, and if the bending angle of the finger is 90 degrees, the processor 1102 calculates the position of the fingertip from the position of the start joint point of the second phalanx and the movement distance of the first phalanx to the at least one virtual input interface.
[0096] More preferably, the processor 1102 determines that the user performs an input operation through the target virtual input interface when the position of the fingertip and a target virtual input interface in the at least one virtual input interface satisfy a set position rule. Specifically, when the position of the fingertip is on the target virtual input interface, the processor 1102 determines that the user is touching the target virtual input interface, and / or when the position of the fingertip is on the side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold, the processor 1102 determines that the user is clicking the target virtual input interface.
[0097] More preferably, the smart device is equipped with an infrared sensor, and the processor 1102 can further use the infrared sensor to collect distance values between the infrared sensor and key points on the hand, and perform position correction on the calculated positions of the user's fingertips using the distance values.
[0098] The memory in Figure 11 may be implemented with any type of volatile or non-volatile storage device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.
[0099] The display 1103 in FIG. 11 refers to a screen including a liquid crystal display (LCD) and a touch panel (TP). When the screen includes a touch panel, the screen may be implemented as a touch screen that receives input signals from a user. The touch panel includes one or more touch sensors that detect touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide motion, but also the duration and pressure associated with the touch or slide motion.
[0100] The audio component 1104 of Figure 11 may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device in which the audio component is located is in an operational mode such as a call mode, a recording mode, a voice recognition mode, etc. The received audio signals may be further stored in memory or transmitted via a communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.
[0101] 11, the electronic device further includes other components such as a communication component 1105 and a power supply component 1106. In FIG. 11, only some components are schematically shown, and this does not mean that the electronic device includes only the components shown in FIG.
[0102] The communication component 1105 of FIG. 11 is configured to facilitate wired or wireless communication between a device in which the communication component is located and another device. The device in which the communication component is located can be connected to a wireless network based on a communication standard such as WiFi, 2G, 3G, 4G, or 5G, or a combination thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may be implemented based on near field communication (NFC) technology, radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0103] Of these, the power supply component 1106 provides power to various components of the device in which it is located, and may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which it is located.
[0104] In this embodiment, the terminal device calculates the coordinates of the fingertips using a stereo vision algorithm based on the positions of the recognized key points of the hand, compares the coordinates of the fingertips with at least one virtual input interface in the virtual scene, and determines that the user will perform an input operation through the target virtual input interface if the positions of the fingertips and a target virtual input interface in the at least one virtual input interface satisfy a set position rule. With this configuration, by calculating the positions of the user's fingertips using a stereo vision algorithm, the user does not need to interact with a controller or special sensor device in the real world, and the immersion and realism of the virtual scene are further improved.
[0105] Therefore, an embodiment of the present application further provides a computer-readable storage medium storing a computer program for implementing each step executable by a terminal device in the above method embodiment.
[0106] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as a method, a system, or a computer program product. Accordingly, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. The present invention may also take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, magnetic disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0107] The present invention will be described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, and the instructions executed by the processor of the computer or other programmable data processing device can generate a machine to implement the function(s) specified in one or more flows in the flowcharts and / or one or more blocks in the block diagrams.
[0108] These computer program instructions may also be stored in a computer-readable memory that can cause a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction apparatus that implements the functions specified in one or more flows of the flowcharts and / or one or more blocks of the block diagrams.
[0109] These computer program instructions may be loaded into a computer or other programmable data processing device, whereby a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, whereby the instructions executing on the computer or other programmable device provide steps for realizing the functions specified in one or more flows of the flowcharts and / or one or more blocks of the block diagrams.
[0110] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0111] Memory may include forms of non-persistent memory, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), in a computer-readable medium. Memory is one example of a computer-readable medium.
[0112] Computer-readable media include persistent and non-persistent, removable and non-removable media, and may be implemented by any method or technology to store information. Information may be computer-readable instructions, data structures, program modules, or other data. Computer storage media include, for example, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital multifunction disk (DVD) or other optical storage devices, magnetic cartridges, magnetic disks or other magnetic storage devices, or other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals or carriers.
[0113] It should be noted that the terms "comprise," "include," or any variation thereof, are intended to encompass a non-exclusive inclusion, such that a process, method, article, or device that includes a set of elements includes not only those elements but also other elements not expressly listed or that are inherent in such process, method, article, or device. Unless further limited, an element qualified by the phrase "comprises..." does not exclude the presence of other identical elements in the process, method, article, or device that includes that element.
[0114] The above description is merely an example of the present application and does not limit the present application. Various modifications and variations are possible for those skilled in the art to the present application. Any amendments, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
Claims
1. 1. A method for input recognition in a virtual scene applied to a smart device, comprising: Recognizing key points of a user's hand in a binocular image of the hand captured by a binocular camera; determining, for any finger of the user, whether the recognized hand key points include a fingertip joint point of the finger; If the key points of the hand include a fingertip joint point of the finger, calculating the position of the fingertip joint point of the finger in the virtual scene using a stereo vision algorithm based on the position of the fingertip joint point in the binocular image, and setting the calculated position as the coordinate of the fingertip of the finger; comparing the coordinates of the fingertip with at least one virtual input interface in the virtual scene; and determining that the user will perform an input operation through the target virtual input interface if the position of the fingertip and a target virtual input interface in the at least one virtual input interface satisfy a set position rule.
2. The step of recognizing key points of a user's hand in a binocular image of the hand taken by a binocular camera includes: detecting a hand region from any one of the monocular images in the binocular images using a target detection algorithm; Segmenting a foreground image corresponding to the hand region from the monocular image; and recognizing the foreground image using a preset hand keypoint recognition model to obtain hand keypoints in the monocular image.
3. If the key points of the hand do not include the fingertip joint points of the fingers, calculating a bending angle of the fingers based on the positions of the visible key points of the fingers in the binocular image and phalange association features when performing an input operation; 2. The method of claim 1, further comprising: calculating coordinates of the fingertip from the bend angle of the finger and the positions of visible key points of the finger in the binocular image.
4. the finger includes a first phalanx close to the palm, a second phalanx connected to the first phalanx, and a tip phalanx connected to the second phalanx; Calculating a bending angle of the finger based on the positions of the visible key points of the finger in the binocular image and phalange association features when performing an input operation, determining the actual length of each of the first phalanges, the second phalanges, and the tip phalanges of the finger; calculating observed lengths of the first phalange, the second phalange, and the tip phalange based on the coordinates of the recognized key points of the hand; determining that the bending angle of the finger is less than 90 degrees when the observed length of the second phalanx and / or the tip phalanx is shorter than the corresponding actual length, and calculating the bending angle of the finger from the observed length and actual length of the second phalanx and / or the observed length and actual length of the tip phalanx; and determining that the bend angle of the finger is 90 degrees if the observed length of the second phalanx and / or the tip phalanx is zero.
5. The step of calculating the coordinates of the fingertip from the bending angle of the finger and the position of the visible key point of the finger in the binocular image includes: If the bending angle of the finger is less than 90 degrees, calculating the coordinates of the fingertip from the position of the start joint point of the second phalanx, the bending angle of the finger, the actual length of the second phalanx, and the actual length of the tip phalanx; and calculating a position of the fingertip from a position of a start joint point of the second phalanx and a moving distance of the first phalanx to the at least one virtual input interface when the bending angle of the finger is 90 degrees.
6. determining that the user performs an input operation through the target virtual input interface when the position of the fingertip and a target virtual input interface in the at least one virtual input interface satisfy a set position rule; If the position of the fingertip is on the target virtual input interface, determining that the user is touching the target virtual input interface; and / or 2. The method of claim 1, further comprising: determining that the user is clicking on the target virtual input interface if the position of the fingertip is on a side of the target virtual input interface away from the user and the distance to the target virtual input interface is greater than a preset distance threshold.
7. an infrared sensor attached to the smart device; using the infrared sensor to collect distance values between the infrared sensor and key points on the hand; 7. The method according to claim 1, further comprising: correcting the calculated position of the user's fingertip using the distance value.
8. A terminal device, comprising: a memory; and a processor; the memory stores one or more computer instructions; A terminal device, characterized in that said processor is for carrying out the steps of the method according to any one of claims 1 to 6 by executing said one or more computer instructions.
9. A computer-readable storage medium storing a computer program, the computer program causing the processor to perform the steps of the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Virtual keyboard interaction method and system
CN113238705A
User interface device and computer program
JP2011133942A
Method of input with virtual keyboard, program, storage medium, and virtual keyboard system
JP2014165660A
Detection device and detection method
JP2015170206A
Information processing device, information processing method, and program
WO2019163372A1