Method for achieving typing or touch control with tactile feedback
The method uses a hand joint detection model and binocular disparity to accurately determine if the trigger fingertip touches the functional area in XR devices, addressing misjudgment and providing tactile feedback without additional sensors, enhancing the typing experience in virtual reality.
Patent Information
- Application Number
- JP2024187674
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-10-24
- Publication Date
- 2025-07-09
AI Technical Summary
Existing virtual keyboards and touch systems for XR extended reality devices face challenges in accurately determining whether the trigger fingertip touches the functional area, leading to misjudgment and lack of tactile feedback, and often require additional sensors like gloves or rings, which are cumbersome.
A method utilizing a hand joint detection model to output time-series position information, where a trigger fingertip is defined on the joint connection line of the palm, and binocular disparity is used to determine if the fingertip touches a functional area by comparing the relative positional relationships of the fingertip with left and right trigger judgment points across multiple images, without the need for additional sensors.
Accurately confirms the presence or absence of touch in a virtual space, providing a realistic typing experience by ensuring the trigger fingertip actually touches the functional area, eliminating misjudgment and the need for additional devices, and offering tactile feedback through sound, vibration, or mechanical sensation.
Smart Images

Figure 2025104251000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual keyboards and touch, and specifically to a method for realizing tactile typing or touch applicable to XR extended reality wearable devices and head-mounted display devices.
Background Art
[0002] Extended Reality (XR) extended reality refers to a combination of one reality and virtuality by computer technology and wearable devices, meaning an environment where human-computer interaction is possible, and is a general term for various forms such as augmented reality AR, virtual reality VR, and mixed reality MR. With the popularization and development of (XR) extended reality in all industries and business formats, various XR smart glasses have emerged, and user-system interaction is realized through virtual keyboard and touch inputs.
[0003] Currently, there are two types of virtual keyboards and touch: (1) In a 1 / 3 / 6DoF three-dimensional environment, a virtual keyboard is drawn, typing or touching in the air with both hands, and using a joint identification model to calculate a threshold position for determining whether the fingertip or radiation position touches a virtual key; (2) Virtual keys are drawn on the palm of the hand. Usually, the fingertip of the thumb (or any finger) (or any location that can focus on one cursor point) is defined as the "trigger fingertip", and virtual keys are drawn on the three finger joints of each of the other fingers and / or different regions of the palm. Each of these virtual keys is defined as the corresponding different numeric key, alphabet key, or function key, and a human hand joint detection model is used to estimate the threshold position for determining whether the trigger fingertip touches a virtual key.
[0004] The input method of the virtual keyboard (including the realization of various functions such as buttons, links, drawing, etc., hereinafter collectively referred to as the "function area") described above in (1) is similar to the typing and cursor click trigger methods of conventional keyboards. However, (a) since the back of the hand and fingers shield the function area, it is difficult to distinguish whether the invisible trigger fingertip actually touches the threshold position of the function area during visual calculation. (b) There is no feeling of touching a physical key by the user. When typing in the air, the user can only rely on their own eyes to judge whether the trigger fingertip touches the correct character key. Therefore, there are two problems that the possibility of blind touch / blind writing cannot be realized.
[0005] The trigger form of the function area on the palm described above in (2) is similar to the conventional Taoist finger folding movement. The function area is on the visible palm surface. When inputting, the palm (hereinafter, the definition of "palm" includes all parts of the palm and fingers that need to be judged) faces the imaging lens of the XR glasses. By touching the function area on the palm with the trigger fingertip, the trigger operation is realized, and the problem of the feeling of touch and the shielding by the back of the hand is solved. However, in such a way, the problem of shielding of the function area by the trigger fingertip still exists. When it is seen that the trigger fingertip is held in front of the function area during visual calculation, it is impossible to know whether the trigger fingertip touches the function area or is in a floating non-touch state. Therefore, it may be misjudged that the trigger fingertip touches the corresponding function area, and the corresponding content of the function area is accidentally triggered. In order to solve the problem that the actual presence or absence of touch cannot be confirmed by the visual calculation or gesture recognition model, many patents have also tried to accurately recognize whether the trigger fingertip actually touches the function area by using a ring or glove sensor. However, sensors such as gloves or rings deviate from the intention of not desiring the wearing of any device or sensor, and the experience and practicality are not high.
Summary of the Invention
Problems to be Solved by the Invention
[0006] The present invention aims to provide a method for realizing a tangible marking or touch, which can accurately confirm whether the trigger fingertip actually touches the functional area only by the visual calculation of the camera for the problems of misjudgment of the presence or absence of an actual touch in the conventional gesture recognition and visual calculation, without the need for an auxiliary device such as a physical sensor, with a small amount of calculation. Also, since the trigger fingertip touches the palm or the surface of an object, it does not trace without a sense of touch in the virtual space, but has a sense of reality during typing or touching, providing a good experience and enabling blind touch / blind writing.
Means for Solving the Problem
[0007] The present invention is a method for realizing a tangible marking or touch, which is applicable to a system of an XR extended reality wearable device and a head-mounted display device. The system outputs time-series position information of several joint points in the video screen of the hand by a hand joint detection model. The palm includes the palm base or fingers, and the trigger fingertip includes triggerable characters / digital buttons, function keys, and shortcut keys. Typing and touch are realized by touching a functional area bound to a preset point marked on the joint connection line of the palm. A preset point is marked on the joint connection line of the palm. The user can see the functional area bound to the preset point on the palm through glasses. Let the width of the functional area be W, the preset point be the center point of the functional area, and at the left W / 2 and right W / 2 parallel to the X-axis, take the left trigger judgment point WL and the right trigger judgment point WR. Based on the joint points, step 1 of estimating the position information of the corresponding preset point and the left trigger judgment point WL and the right trigger judgment point WR of its bound functional area. By default, the tip of the thumb is set as the trigger fingertip. If the thumb does not enter the area within the palm and any other finger tries to touch the palm or the functional area bound thereto, it is determined that the fingertip of the finger is the trigger fingertip, and the position of the trigger fingertip is defined as P. Step 2. The system acquires N image video streams with binocular disparity, where N is an integer and N≥2. For N images in the same time series, it tracks and determines whether the position of the trigger fingertip P falls between the left trigger determination point WL and the right trigger determination point WR corresponding to the left and right boundaries of any functional area in all the images. If so, it calculates the position values of three target points in each image. The target points include the left trigger determination point WL, the trigger fingertip P, and the right trigger determination point WR. It takes the X-axis values (WRX, PX, WLX) among the position values of the three target points, and calculates the ratios (PX - WRX):(WLX - PX) of the differences between WL and P and between P and WR respectively. Only when all the ratios of the N images are the same, it indicates that the trigger fingertip P touches the functional area, and includes step 3 of outputting or triggering the corresponding content of the functional area.
[0008] In step 3, two cameras are used as the left and right cameras, the connecting line between the two center points L and R of the cameras is used as the X-axis. In the field of view of the left camera, the included angle in radians between the connecting line between the center point L of the left camera and the target point T and the X-axis is TθL. In the field of view of the right camera, the included angle in radians between the connecting line between the center point R of the right camera and the target point T and the X-axis is TθR. If the binocular disparity between the two center points L and R of the left and right cameras is d, the respective positions (X, Z) of the target point T in the image are calculated. When the target point T is between the two center points L and R of the left and right cameras, it becomes Equation 1.
Equation
Equation
Equation
[0009] The functional area is circular, with a preset point provided at any position on the joint connection line of the palm as the center of the circle, and a circle is drawn with W as the diameter.
[0010] A functional area is drawn at a position within the phalangeal region, between phalangeal regions, outside the phalangeal region, or between the wrist of the palm and a certain finger.
[0011] In step 3, the system processes N image video streams, where N is an integer and N ≥ 2. When displayed on the screen, one matrix grid including several grids at the same position of the palm is drawn. Each grid includes several sides, and each grid is used as a functional area. The system tracks and determines whether the trigger fingertip P(X, Y) appears in a certain functional area of the matrix grid on all screens simultaneously. If so, taking the left trigger judgment point WL, the trigger fingertip P, and the right trigger judgment point WR, which are the left and right sides of the functional area, as three target points, taking the X-axis values (WRX, PX, WLX) in the position information of the three target points, calculating the ratios (PX - WRX):(WLX - PX) of the differences between WL and P and between P and WR respectively. If the ratios of all images are equal, it indicates that the trigger fingertip P touches the functional area, draws a plot point at the position P(X, Y) of the trigger fingertip, and sequentially connects the plot points in time series to realize the touch function of the tablet or touch pad with the fingertip of one hand as the trigger fingertip on the palm of the other hand.
[0012] The matrix grid has the connection joint between the little finger and the palm as the right vertex, the connection joint between the index finger and the palm as the left vertex, and the connection part between the palm and the wrist as the lower boundary.
[0013] The matrix grid is an invisible setting not displayed on the screen.
[0014] The grids are square or rectangular.
[0015] Another method for realizing tactile typing or touching according to the present invention is applied to a system of XR extended reality wearable devices and head-mounted display devices. The system outputs time-series position information on the video screen of the target point. Typing and touching are realized by the trigger fingertip touching the functional area. Step 1: The system calibrates one touch interface image at the same position on the surface of the same preset object on each screen. Several functional areas are provided in the touch interface image. The functional areas in the same time-series frame are parallel to both sides of the X-axis, and the left trigger judgment point WL and the right trigger judgment point WR are determined. Step 2: Designate the fingertip of any finger attempting to touch the functional area as the trigger fingertip P. Step 3: The system acquires N image video streams with binocular disparity, where N is an integer and N≥2. It tracks and determines whether the trigger fingertip P(X, Y) appears within the functional area on all screens simultaneously. If so, using the trigger fingertip P(X, Y), the corresponding left trigger judgment point WL and right trigger judgment point WR of the functional area as three target points, take the X-axis values (WRX, PX, WLX) in the position information of the three target points, and calculate the ratios (PX - WRX):(WLX - PX) of the differences between WL and P and between P and WR respectively. Only when all the ratios of the N images are the same, it indicates that the trigger fingertip P is touching the functional area, and the corresponding content of the functional area is output or triggered.
[0016] The touch interface image is a conventional calculator diagram and a conventional keyboard diagram.
[0017] The surface of the preset object is any physical surface.
[0018] The surface of the preset object is the surface of a virtual object. When the trigger fingertip touches the functional area, feedback is given to the user through sound, vibration, electric shock, or machinery to give the feeling of touching a physical object.
[0019] A head-mounted display device, comprising at least two cameras for imaging a target image of a target area, further comprising a memory for storing a computer program, and a processor for executing the computer program to implement the method according to any one of the above items.
[0020] When the technical solution of the present invention is used, a parallax image video stream is acquired by at least two cameras of the smart glass. When touch judgment is performed on images in the same time series in the image video stream, if the connection line of the two cameras is used as the X-axis or is parallel to the X-axis, the trigger fingertip P, the corresponding left trigger judgment point WL and the right trigger judgment point WR of the function area are used as three target points, the X-axis values in the position information of the three target points are taken, and the ratios of the X-axis difference value between WL and P and the X-axis difference value between P and WR are calculated respectively. Only when all the ratios of N images are the same, it indicates that the trigger fingertip touches the function area. The formula for calculating the depth Z of the spatial position of the target point in the visual field can ignore the calculation of the Y-axis. Since the cameras have a fixed parallax interval, when the trigger fingertip actually touches the function area without changing the binocular parallax interval, based on the principle that the relative positions of the positions of the trigger fingertip seen from the left and right eyes on the X-axis of the function area are the same, that is, when the three points are all on one line, based on the principle that the relative positions of the three points are the same when viewed from different angles of left and right (or more cameras), the present invention proves that when judging whether the trigger fingertip actually touches the function area, it is not necessary to know the depth Z value in advance, and the parallax distance value d between the cameras can be any value. The touch error in the present invention is Z / αd, where Z is the depth distance between the target point and the camera, d is the distance between any two comparison cameras, and α is the threshold setting. The present invention converts the judgment of whether the trigger fingertip in the visual field actually touches the function area into only calculating the relative position relationships between the trigger fingertip positions in the images captured by different cameras and the two trigger judgment points at the left and right edges of the corresponding function area. If the relative position relationships of the three target points in the images captured by all cameras are consistent, it is judged that the trigger fingertip touches the function area, otherwise, it is judged that it does not touch. Therefore, it does not require information on X, Y, Z or d, but only requires the X-axis pixel values of N≧2 imaging displays. The present invention solves the problem of misjudgment of the presence or absence of actual touch by gesture recognition and visual calculation, and provides a method for implementing a realistic typing or touch in a virtual space.
[0021] When the present invention makes a touch determination, it is not necessary to calculate the depth Z value of the trigger fingertip. Therefore, a matrix grid is drawn at the central part of the palm, and each grid is used as a functional area. At least two cameras of the smart glass are used to obtain left and right image video streams with a disparity distance d, and the relative positional relationship between the trigger fingertip position P(X, Y) in the acquired images of the same time series and two trigger determination points corresponding to the grid at the same Y height is determined. If the relative positional relationships of the above three target points in the images captured by all cameras are consistent, it is determined that the trigger fingertip touches the functional area; otherwise, it is determined that there is no touch. When there is a touch, a plot point is drawn at the position P(X, Y) of the trigger fingertip, and by sequentially connecting the plot points in time series, the drawing and drag functions can be realized so that the fingertip of one hand is used as the trigger fingertip to realize the touch function of a tablet or a touch pad on the palm of the other hand. The touch function of a multi-finger tablet or touch pad may be realized with a plurality of trigger fingertips. The depth Z value of the trigger fingertip may be triangulated to draw a three-dimensional plot point P(X, Y, Z).
[0022] In addition to calibrating keyboard buttons and touch panels with the palm, the present invention can also allow typing and touching outside the palm. The smart glass can calibrate a simple calculator image or a keyboard image on the surface of an object such as a wall or a desktop. It may be other real objects or virtual objects, and may be an irregular surface rather than a flat surface. At least two cameras of the smart glass are used to obtain image video streams with a disparity. When the trigger fingertip enters the functional area of the above image, the relative positional relationship between the trigger fingertip position in the acquired images of the same time series and two trigger determination points corresponding to the functional area is determined. If the relative positional relationships of the above three target points in the images captured by all cameras are consistent, it is determined that the trigger fingertip touches the functional area; otherwise, it is determined that there is no touch. In this way, instead of touching virtual keys in the air, the user touches the surface of a real object during typing or touching, so there is a sense of reality.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Embodiments for Carrying Out the Invention
[0024] Hereinafter, while referring to the drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. It is obvious that the following embodiments are only a part of the present application, not all of the embodiments. According to the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts are included in the scope of the present invention.
[0025] Also, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, and may include other steps or units not explicitly listed or specific to those processes, methods, products or apparatuses.
[0026] Regarding the explanation of the principle of technical implementation of the present invention (1) Regarding the recognition model used to obtain the position information of the palm, the present invention will be described by taking Mediapipe as an example with respect to the open-source software of the pre-trained human hand joint detection model that can obtain the planar position of the human hand joints on the market. Mediapipe is an open-source item of Google and is a tool library of machine learning and mainly visual algorithms. A large number of models such as face detection, face keypoints, gesture recognition, avatar segmentation, and pose recognition are integrated. As shown in FIG. 1, it can output the time-series position information of 21 joint points (also called keypoints) on the video screen of the human hand. Generally, the human hand joint detection model outputs the joint position information with the (X, Y) pixels of the imaging screen as the X and Y axes. The present invention may also use a human hand joint detection model trained by itself. The present invention uses an artificial intelligence chip such as a GPU graphics processor or an NPU neural network processor to learn by convolution KNN and RNN of tags, or to learn by a pre-training method with a Transformer model, and to identify whether the trigger fingertip enters the functional area.
[0027] (2) Regarding the setting of the functional area on the palm, using the conventional human hand joint detection model, the time-series (X, Y) position information of 21 joint points (also called key points) on the video screen of the human hand can be output. In the present invention, preset points (for example, the midpoint of the joint connection line) are calibrated on the joint connection line of the palm. The user can see, through the smart glasses, a functional area including triggerable character / numeric buttons, function keys or shortcut keys bound to the preset points on the palm. Let the width of the functional area be W, with the preset point as the center point of the functional area. In the direction parallel to the X-axis, two trigger judgment points WL and WR are taken at left W / 2 and right W / 2 respectively. Based on the joint points, the position information of the corresponding preset point and the two trigger judgment points WL and WR of the functional area bound thereto can be estimated. The functional area may have any shape, preferably circular, because regardless of the rotation direction of the palm, it does not affect the display effect of the functional area on the palm. Like the circular functional area shown in FIG. 4, a circle is drawn with a preset point provided at any position on the joint connection line of the palm as the center and W as the diameter. In the present invention, the functional area is drawn at a position within the finger joint area, between the finger joint areas, outside the finger joint area, or between the wrist of the palm and a certain finger.
[0028] (3) Regarding the setting of the trigger fingertip, by default, the fingertip of the thumb is set as the trigger fingertip. If the thumb does not enter the area within the palm or the fingertip of the thumb is not used as the trigger fingertip, when any other finger tries to touch the palm or the functional area bound thereto, it is determined that the fingertip of that finger is the trigger fingertip.
[0029] Taking the arrangement of the single-handed keypad function area in FIG. 8 as an example, when the trigger fingertip touches a function area other than the fingertip, as shown in FIG. 10, the characters "C" of the index finger, " / " of the middle finger, "X" of the ring finger, and the function key "Delete" of the little finger are triggered respectively. When the trigger fingertip touches the function area of the distal phalanx of another finger, as shown in FIG. 11, the characters "1" of the index finger, "2" of the middle finger, "3" of the ring finger, and "-" of the little finger are triggered respectively. When the trigger fingertip touches the function area of the middle phalanx of another finger, as shown in FIG. 12, the characters "4" of the index finger, "5" of the middle finger, "6" of the ring finger, and "+" of the little finger are triggered respectively. As shown in FIG. 13, when the trigger fingertip touches the function area of the proximal phalanx of another finger, the characters "7" of the index finger, "8" of the middle finger, "9" of the ring finger, and "=" of the little finger are triggered respectively. When the trigger fingertip touches the function area at the lower end of the proximal phalanx of another finger, as shown in FIG. 14, the "%" of the index finger, the character "0" of the middle finger, the "." of the ring finger, and the "=" of the little finger are triggered respectively. As can be seen from this, when the function area is provided at the tip of the fingertip, each finger joint or a location close to the finger joint on the palm, if the thumb is used as the trigger fingertip to touch the function area, the output of the function area can be triggered in any case. However, for the function area provided at a location close to the wrist in the palm area, the thumb is difficult to touch. The present invention sets the corresponding fingers to trigger respectively. At this time, the trigger fingertip is not the fingertip of the thumb, but the fingertip of the corresponding finger. By touching the corresponding function area with the fingertip of the corresponding finger, the corresponding function area is triggered to output characters / functions. As shown in FIG. 15, function keys such as "MC" of the index finger, "M+" of the middle finger, "M-" of the ring finger, and "MR" of the little finger are triggered respectively.
[0030] (4) Regarding the calculation of the spatial position of the target point, what can be seen from the XR smart glasses is a three-dimensional XYZ space. However, when calculating the positions of the left and right trigger judgment points in the X-axis direction of the trigger fingertip and the functional area, the calculation of the Y-axis can be ignored and simplified to two-dimensional position calculation. As shown in Fig. 2, taking the connection line of the center points L / R of the left and right cameras as the X-axis, in the field of view of the left camera, referring to Fig. 2, if the included angle between the connection line of the center point L of the left camera and the target point T of the calculation target spatial position and the X-axis is TθL, then the included angle of the trigger judgment point WR is WRθL, and the included angle of the trigger judgment point WL is WLθL. Similarly, as shown in Fig. 3, in the field of view of the right camera, if the included angle between the connection line of the center point R of the right camera and the target point T of the calculation target spatial position and the X-axis is TθR, then the included angle of the trigger judgment point WR is WRθR, and the included angle of the trigger judgment point WL is WLθR.
[0031] Regarding the trigger fingertip P point, the left and right trigger judgment points WL and WR as the three target points T to be calculated, and the visual distance difference d between the two center points L and R of the left and right cameras, calculate the position of any target point T(X, Z). Specifically, If the target point T is between the two center points L and R of the left and right cameras, it becomes Equation 4
Equation
Equation
Equation
[0032] The above examples are calculated using TAN and COT, but the present invention can be realized using any trigonometric calculation method.
[0033] (5) Regarding the method for determining whether the trigger fingertip touches the functional area, The system acquires left and right (or more angles) image video streams with binocular disparity, makes judgments on the left and right (multiple) images in the same time series respectively. When the trigger fingertip P enters between the corresponding two trigger judgment points WL and WR in any functional area, it compares the radian ratio value (PθL - WRθL):(WLθL - PθL) in the left image with the radian ratio value (PθL - WRθR):(WLθR - PθL) in the right image. If the two ratio values are not equal, it indicates that the trigger fingertip does not touch the functional area. Referring to the upper two figures in FIGS. 5 and 7, if the two ratio values are the same, it indicates that the trigger fingertip touches the functional area. Referring to the lower two figures in FIGS. 6 and 7, it outputs the corresponding content of the functional area.
[0034] In the present invention, when comparing two or more numerical values, those with an error within the threshold range are all regarded as the same, equal, or consistent. Generally, the error threshold can be set to about Z / 5d.
[0035] Since the fields of view (FOVs) captured by different cameras are different, the X-axis pixel value X obtained by the human hand joint detection model can be directly converted into θ radians / angles in all the above formulas. Assuming that the total X-axis resolution of the image is 1800 pixels, the camera FOV is 180 degrees, and the X of the target point T(X, Y) fed back by the human hand joint detection model is 900 pixels, then the θ radian of the target point is π / 2 (the angle is 90°). Since the present invention only needs to compare the relative radian ratio values of the three target points (WL, P, WR) in the left and right (multiple) cameras, there is no need to convert the absolute θ radians or angles. Using the X value of the target point fed back by the human hand joint detection model directly, the relative radian ratio values of the three target points can be calculated. Therefore, assuming that θ is the X value output by the human hand joint detection model, the radian ratio value in the left image is (PX L -WRX L ):(WLX L -PX L) and the radian proportional value in the right image is (PX R -WRX R ):(WLX R -PX R ).
[0036] (6) Regarding the arrangement example of the functional areas of the palm FIG. 8 is an arrangement example of the one-handed numeric keypad functional area, and by folding fingers with one hand, typing within the palm in the virtual space can be realized.
[0037] FIG. 9 is an arrangement example of the two-handed 26-character functional area, and by folding fingers with both hands respectively, typing within the palm in the virtual space is realized.
[0038] According to the technical solution of the present invention, based on typing habits and convenience of use, the positions of each functional area and the characters (or function keys / shortcut keys) bound correspondingly can be set by oneself. As long as it is any position on the joints of the palm or on the joint connection lines, the position information of each functional area can be obtained according to the time-series position information of each joint point output by the human hand joint detection model, and the position information of the two corresponding trigger judgment points of the functional area is obtained and used for judging whether the trigger fingertip touches the functional area.
[0039] (7) Regarding the principle of realizing tablet touch on the palm Since there is two-dimensional pixel data of X and Y on the screen of the camera in the present invention, the touch function of a two-dimensional tablet such as functions of two-dimensional operations such as drawing, writing, dragging, and pulling on the plane of the palm can be realized.
[0040] On each screen, a matrix grid (which may be in a hidden form) is drawn at the same position on the palm. Taking the left hand as an example (see Fig. 17), the connection point between the little finger and the palm is set as the right vertex of the matrix grid, the connection point between the index finger and the palm is set as the left vertex of the matrix grid, and the connection point between the palm and the wrist is set as the lower boundary of the matrix grid. When the palm rotates and moves arbitrarily, since it is bound to the joint points of the palm, the position of the matrix grid relative to the palm is fixed. Each grid in the matrix grid has upper, lower, left, and right lines corresponding to the four sides. The grid shape does not necessarily have to be a square, and any shape can be used. If it is a triangle, it has three sides; if it is a hexagon, it has six sides; it may also be an irregular shape. If it is an irregular shape, each grid may have a different number of sides, or each grid may have a different texture. Each grid can be used as a functional area and is judged according to the "judgment method for whether the trigger fingertip touches the functional area" in point (5).
[0041] If the system tracks and determines whether the trigger fingertip P(X, Y) appears simultaneously in a certain functional area of the matrix grid on the left and right (or multiple) screens respectively, then taking the left and right sides of the functional area and the trigger fingertip P(X, Y) as three target points (left trigger determination point WL, trigger fingertip P, right trigger determination point WR), taking the X-axis values (WRX, PX, WLX) in the position information of the three target points, calculating the ratios of the differences between WL and P and between P and WR respectively, i.e., (PX - WRX):(WLX - PX), if the ratios of all images are equal, it indicates that the trigger fingertip P is touching the functional area. By plotting points at the position P(X, Y) of the trigger fingertip and sequentially connecting the plotted points in time series, the drawing and drag functions can be realized to achieve the touch function of a tablet or touch pad with the fingertip of one hand as the trigger fingertip on the palm of the other hand. The touch function of a multi-finger tablet or touch pad can also be realized with multiple trigger fingertips. Triangulation can be performed to obtain the depth Z value of the trigger fingertip and draw a 3D plot point P(X, Y, Z). When adopting the above-mentioned formula of the radian ratio, the radian-converted pixel ratio value in the left image is (PX L -WRX L ):(WLX L -PX L ), and the radian-converted pixel ratio value in the right image is (PX R -WRX R ):(WLX R -PX R ). If there are N cameras (N is an integer and N ≥ 2), if the radian ratios of all cameras (PX N -WRX N ):(WLX N -PX N ) are the same (within a predetermined threshold error), it can be determined that there is an actual touch; otherwise, it can be determined that there is no touch.
[0042] Since the palm is rotatable, the corresponding grid is also rotatable along with the palm. Therefore, if the trigger fingertip is simultaneously within a certain grid (functional area), both the WLX and WRX on the left and right sides at the same Y height may change in real time. Thus, the calculation formula for confirming whether the touch is using the present invention must be a comparison of the same time series frames.
[0043] FIG. 18 is a schematic diagram of an XY matrix grid displayed on the palm and shortcut keys displayed in the knuckle area, and has the function of a combination of tablet touch and shortcut keys.
[0044] (8) Since all current smart glasses have an IMU chip, it is possible to calibrate a fixed position in a 3D environment near any image with 1 / 3 / 6 DoF (1 / 3 / 6 free dimensions). In addition to calibrating keyboard buttons and touch panels with the palm, the present invention can also allow typing and touching outside the palm. FIG. 19 shows a simple calculator image calibrated on the surface of an object such as a wall or a desktop. FIG. 20 is a keyboard image, which can also be calibrated and used on a wall or a desktop in a similar manner. The object surface may be a concave-convex surface. The present invention acquires a parallax image video stream by at least two cameras of the smart glasses. When the trigger fingertip enters the functional area of the above image, it determines the relative positional relationship between the trigger fingertip position in the acquired images of the same time series and the two corresponding trigger judgment points of the functional area. If the relative positional relationships of the above three target points in the images captured by all cameras match, it is determined that the trigger fingertip touches the functional area; otherwise, it is determined that there is no touch. In this way, since the user touches the surface of the real object during typing or touching instead of touching virtual keys in the air, there is a sense of reality.
[0045] The surface of the object may be the surface of a virtual object. When the trigger fingertip touches the functional area, feedback is provided to the user in the form of sound, vibration, electric shock, or a mechanical sensation of touching a real object.
[0046] The present invention further includes different depth and velocity sensors, which may be combined with conventional imaging sensors or used independently. Since the present invention is for determining whether actual touching is occurring based on the relative positional relationship between the trigger fingertip and two trigger judgment points, the computer does not need to perform triangulation of the depth position, and monitors and executes the relative distances and ratios of the positions of three target points using depth sensors such as sensors like laser SLAM, IR infrared tracking, and moving Motion. For example, a Motion Velocity sensor outputs pixels during movement, and such pixels can also be used. SLAM gives the Z value for each X-axis pixel, but can also give the X value. IR and other ToF sensors can give the depth Z value, but the X and Y values can also be calculated in the present invention.
[0047] The present invention can be applied to any interaction command that requires not only typing within the palm but also combining typing within the palm or touch operations. For example, the user A. When projected far away along a certain anchor position through a firing position of the hand to form a single radiation line, and the radiation line is directed at a certain virtual key or link target position in the distance, the touch command of the combined fingertip and knuckle within the palm can be executed according to the method of the present invention. B. When the user's index finger clicks on a virtual screen or virtual key link, it is necessary to touch the virtual key at the end joint of the middle finger by combining the fingertip within the palm, for example, the thumb, to trigger a short press or long press command. The touch command of the combined fingertip and knuckle within the palm can be executed according to the method of the present invention. C. Some smart glasses employ an eye tracker to calculate, based on the angles of the left and right eye pupils, at which angle the user is looking to form a single three-dimensional ray, and when the ray points to the target position of a virtual key or link function area in the distance, the touch commands of the fingertips and finger joints in the palm can be executed according to the method of the present invention. D. Some smart glasses form a single three-dimensional vertical ray by using the central position as a simple vertical ray, and when the ray points to the target position of a virtual key or link function area in the distance, the touch commands of the fingertips and finger joints in the palm can be executed according to the method of the present invention.
[0048] Embodiment 1 The embodiment of the present invention relates to a method for realizing a tactile typing or touch, which is applicable to an XR extended reality wearable device and a system of a head-mounted display device. The system outputs time-series position information of a target point on a video screen. The palm includes the palm bottom or fingers, and when the trigger fingertip touches the function area, realizing typing and touch A preset point is calibrated on the joint connection line of the palm. The user can see, through the glasses, a function area including a triggerable character / number button, function key or shortcut key bound to the preset point on the palm. Let the width of the function area be W, take the preset point as the center point of the function area, and take two trigger judgment points WL and WR at left W / 2 and right W / 2 parallel to the X axis respectively. Based on the joint points, step 1 of estimating the position information of the corresponding preset point and the two trigger judgment points WL and WR of the function area bound thereto is included. The function area may be of any shape, preferably circular, and a circle is drawn with the preset point provided at any position on the joint connection line of the palm as the center and W as the diameter.
[0049] In the present invention, the function area is drawn at a position within the finger joint area, between finger joint areas, outside the finger joint area, or between the wrist and a finger in the palm center.
[0050] In step 2, if the thumb is defaulted as the trigger finger and the thumb does not enter the area within the palm, and any other finger tries to touch the palm, it is determined that the finger is the trigger finger.
[0051] In step 3, the system acquires N image video streams with binocular disparity. For N images in the same time series, it is tracked and determined whether the position of the trigger fingertip P in all images falls between two trigger judgment points WL and WR corresponding to the left and right boundaries of any functional area. If so, for the three target points (left trigger judgment point WL, trigger fingertip P, and right trigger judgment point WR) in each image, the X-axis values (WRX, PX, WLX) among the position information of the three target points are taken, and the ratios (PX - WRX):(WLX - PX) of the difference between WL and P and the difference between P and WR are calculated respectively. Only when all the ratios of the N images are the same, it indicates that the trigger fingertip P touches the functional area, and the corresponding content of the functional area is output or triggered.
[0052] Taking two cameras as the left and right cameras, taking the connecting line between the two center points L and R of the cameras as the X-axis. In the field of view of the left camera, taking the included angle radian between the connecting line between the center point L of the left camera and the target point T and the X-axis as TθL, and in the field of view of the right camera, taking the included angle radian between the connecting line between the center point R of the right camera and the target point T and the X-axis as TθR. Assuming the binocular disparity between the two center points L and R of the left and right cameras is d, the respective positions (X, Z) of the target point T in the image are calculated. If the target point T is between the two center points L and R of the left and right cameras, Equation 7 is obtained
Equation
Equation
Equation
[0053] The system includes several equally divided grids at the same position on the palm in at least two screens for the left and right eyes. Taking the connection joint between the little finger and the palm as the right vertex, the connection joint between the index finger and the palm as the left vertex, and the connection point between the palm and the wrist as the lower boundary, one matrix grid is drawn. Each grid includes several sides, and each grid is used as a functional area. The system monitors whether the trigger fingertip P(X, Y) appears in a certain functional area of the matrix grid on the above screen at the same time. If so, taking the trigger fingertip P(X, Y) and the two corresponding trigger judgment points of the functional area as three target points (left trigger judgment point WL, trigger fingertip P, right trigger judgment point WR), taking the X-axis values (WRX, PX, WLX) in the position information of the three target points, and calculating the ratios (PX - WRX):(WLX - PX) of the difference between WL and P and the difference between P and WR respectively. If the ratios of all images are equal, it indicates that the trigger fingertip P touches the functional area, draws a plot point at the position P(X, Y) of the trigger fingertip, and by sequentially connecting the plot points in time series, the touch function of the tablet or touch pad can be realized with the fingertip of one hand as the trigger fingertip on the palm of the other hand.
[0054] The matrix grid is an invisible setting that is not displayed on the screen.
[0055] Another method for realizing realistic typing or touch according to the present invention is applied to the systems of XR extended reality wearable devices and head-mounted display devices. The system outputs the time-series position information of the target points on the video screen. By the trigger fingertip touching the functional area, typing and touch can be realized. The system calibrates one touch interface image at the same position on the surface of the same preset object on each screen respectively. Several functional areas are provided in the touch interface image. The functional areas in the same time-series frame are parallel to both sides of the X-axis. Step 1 is to take the left trigger judgment point WL and the right trigger judgment point WR. Step 2 of designating the fingertip of any finger attempting to touch the functional area as the trigger fingertip P; Step 3 of the system acquiring N image video streams with binocular disparity, where N is an integer and N ≥ 2, tracking and determining whether the trigger fingertip P(X, Y) appears within the functional area that is present in all the screens simultaneously. If so, using the trigger fingertip P(X, Y), the corresponding left trigger determination point WL and right trigger determination point WR of the functional area as three target points, taking the X-axis values (WRX, PX, WLX) in the position information of the three target points, and calculating the ratios (PX - WRX):(WLX - PX) of the differences between WL and P and between P and WR respectively. Only when all the ratios of the N images are the same, it indicates that the trigger fingertip P touches the functional area, and outputting or triggering the corresponding content of the functional area.
[0056] The touch interface image is a conventional calculator diagram and a conventional keyboard diagram.
[0057] The surface of the preset object may be a wall, a desktop, etc.
[0058] Those skilled in the art can combine the units and algorithm steps of each example according to the embodiments disclosed in the present invention, and can be realized by electronic hardware, computer software or a combination of both. For the sake of clearly explaining the compatibility between hardware and software, in the above description, the configurations and steps of each example are generally described according to functions, and it is further understood that. Whether these functions are executed by hardware or software depends on the specific application of the technical solution and the design constraint conditions. Those skilled in the art can realize the functions described in different ways for each specific application, but such realizations should not be considered as exceeding the scope of the present invention.
[0059] Specifically, each step of the method embodiment in the embodiments of the present application may be completed by the integrated logic circuit of the hardware in the processor and / or the instructions in the form of software. The steps combined with the method disclosed in the embodiments of the present application can be directly embodied as the completion of the execution of the hardware decoding processor, or can be executed and completed by the combination of the hardware and software modules in the decoding processor. Preferably, the software module is stored in a storage medium mature in this field, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines it with its hardware to complete the steps in the embodiments of the above method.
[0060] Embodiment 2 As shown in FIG. 16, Embodiment 2 of the present invention includes a memory 710 and a processor 720. The memory 710 stores a computer program and provides a head-mounted display device 700 that transmits the program code to the processor 720. In other words, the processor 720 can call and execute the computer program from the memory 710 to implement the method in the embodiments of the present invention. For example, the processor 720 is for executing the processing steps described in the method of Embodiment 1 according to the instructions in the computer program.
[0061] In some embodiments of the present invention, the computer program may be divided into one or more modules stored in the memory 710 and executed by the processor 720 to complete the method of Embodiment 1 according to the present invention. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are for describing the execution process of the computer program in the head-mounted display device 700.
[0062] As shown in FIG. 16, the head-mounted display device may further include a transceiver 730 connected to the processor 720 or the memory 710. Here, the processor 720 can control the transceiver 730 to communicate with another device. Specifically, it can transmit information or data to another device or receive information or data transmitted by another device. The transceiver 730 may be at least two cameras for imaging a target image of a target area.
[0063] It should be understood that each component of the head-mounted display device 700 is connected via a bus system, and the bus system further includes a power bus, a control bus, and a status signal bus in addition to a data bus.
[0064] The above specific embodiments have further elaborated on the object, technical solution, and beneficial effects of the present invention. However, the above are only specific embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should all be included within the scope of the present invention.
Explanation of Reference Numerals
[0065] 26 Both hands, 700 Head-mounted display device, 710 Memory, 720 Processor, 730 Transceiver, 900 Pixel, L Center point, P Trigger fingertip, Trigger fingertip position, R Center point, T Target point, W Functional area, WL Left trigger determination point, WR Right trigger determination point, Z Depth, d Disparity distance, Disparity distance value
Claims
1. An XR extended reality wearable device, applicable to a system of a head-mounted display device, the system outputs time-series position information of several joint points on a video screen of a hand by a hand joint detection model, and the palm is a method for realizing a tactile typing or touch including the palm bottom or fingers, The trigger fingertip includes triggerable characters / digital buttons, function keys, and shortcut keys, and realizes typing and touch by touching a function area bound to a preset point marked on the joint connection line of the palm. This is, A preset point is marked on the joint connection line of the palm. The user can view the function area bound to the preset point on the palm through glasses. Let the width of the function area be W, and the preset point be the center point of the function area. At the left W / 2 and right W / 2 parallel to the X-axis, take the left trigger judgment point WL and the right trigger judgment point WR. Based on the joint points, estimate the position information of the corresponding preset point and the left trigger judgment point WL and the right trigger judgment point WR of its bound function area in step 1, By default, the tip of the thumb is the trigger fingertip. If the thumb does not enter the area within the palm and any other finger tries to touch the palm or the function area bound thereto, it is determined that the fingertip of the finger is the trigger fingertip, and define the position of the trigger fingertip as P in step 2, The system acquires N image video streams with binocular disparity, where N is an integer and N≥2. For N images in the same time series, track and determine whether the position of the trigger fingertip P in all images falls between the left trigger judgment point WL and the right trigger judgment point WR corresponding to the left and right boundaries of any function area. If so, calculate the position values of three target points in each image. The target points include the left trigger judgment point WL, the trigger fingertip P, and the right trigger judgment point WR. Take the X-axis values (WRX, PX, WLX) among the position values of the three target points, and calculate the ratios of the differences between WL and P and between P and WR, respectively, (PX - WRX):(WLX - PX). Only when all the ratios of the N images are the same, it indicates that the trigger fingertip P is touching the function area, and output or trigger the corresponding content of the function area in step 3, A method for realizing a tactile typing or touch, characterized by including the above.
2. In the step 3, two cameras are used as the left and right cameras. The connecting line between the two center points L and R of the cameras is taken as the X-axis. In the field of view of the left camera, the included angle in radians between the connecting line between the center point L of the left camera and the target point T and the X-axis is defined as TθL. In the field of view of the right camera, the included angle in radians between the connecting line between the center point R of the right camera and the target point T and the X-axis is defined as TθR. If the disparity distance between the two center points L and R of the left and right cameras is d, the respective positions (X, Z) of the target point T in the image are calculated. If the target point T is between the two center points L and R of the left and right cameras, the formula 1 is obtained. 【Number 1】 If the target point T is on the left side of the center point L of the left camera, the formula 2 is obtained. 【Number 2】 If the target point T is on the right side of the center point R of the right camera, the formula 3 is obtained. 【Number 3】 A method for realizing realistic typing or touch according to claim 1, characterized in that.
3. The functional area is circular, with a preset point provided at any position on the joint connecting line of the palm as the center of the circle, and a circle is drawn with W as the diameter. A method for realizing realistic typing or touch according to claim 1 or 2, characterized in that.
4. A method for realizing realistic typing or touch according to claim 1 or 2, characterized in that a functional area is drawn at a position within the finger joint area, between the finger joint areas, outside the finger joint area, or between the wrist of the palm and a certain finger.
5. In the said step 3, the system processes N image video streams, where N is an integer and N ≥ 2. When it is displayed on the screen, one matrix grid including several grids at the same position on the palm is drawn. Each grid includes several sides, and each grid is used as a functional area. The system tracks and determines whether the trigger fingertip P(X, Y) appears in a certain functional area of the matrix grid on all screens at the same time. If so, taking the left trigger judgment point WL, the trigger fingertip P, and the right trigger judgment point WR, which are the left and right sides of the functional area, as three target points, taking the X-axis values (WRX, PX, WLX) in the position information of the three target points, calculating the ratios (PX - WRX):(WLX - PX) of the difference between WL and P and the difference between P and WR respectively. If the ratios of all images are equal, it indicates that the trigger fingertip P touches the functional area. A plotting point is drawn at the position P(X, Y) of the trigger fingertip, and by sequentially connecting the plotting points in time series, the touch function of the tablet or touch pad is realized with one fingertip of one hand as the trigger fingertip on the palm of the other hand. The method for realizing tactile typing or touch according to claim 1 or 2, characterized in that...
6. The matrix grid is characterized in that the connecting joint between the little finger and the palm is the right vertex, the connecting joint between the index finger and the palm is the left vertex, and the connecting position between the palm and the wrist is the lower boundary. The method for realizing tactile typing or touch according to claim 1 or 2, characterized in that...
7. The matrix grid is characterized in that it is an invisible setting not displayed on the screen. The method for realizing tactile typing or touch according to claim 1 or 2, characterized in that...
8. The grid is characterized in that it is square or rectangular. The method for realizing tactile typing or touch according to claim 1 or 2, characterized in that...
9. Applied to the system of the XR extended reality wearable device and the head-mounted display device, the system outputs time-series position information of the target point on the video screen, and by the trigger fingertip touching the functional area, it is a method for realizing tactile typing or touch that realizes typing and touch. The system calibrates one touch interface image at the same position on the surface of the same preset object on each screen, several functional areas are provided in the touch interface image, and the functional areas in the same time-series frame are parallel to both the left and right sides of the X-axis. Step 1 is to obtain the left trigger judgment point WL and the right trigger judgment point WR. Step 2 is to use the fingertip of any finger attempting to touch the functional area as the trigger fingertip P. The system acquires N image video streams with binocular disparity, where N is an integer and N≥2. It tracks and determines whether the trigger fingertip P(X, Y) appears within the functional area that exists simultaneously on all screens. If so, taking the trigger fingertip P(X, Y), the corresponding left trigger judgment point WL and right trigger judgment point WR of the functional area as three target points, obtaining the X-axis values (WRX, PX, WLX) in the position information of the three target points, and calculating the ratios of the differences between WL and P and between P and WR, respectively, (PX−WRX):(WLX−PX). Only when all the ratios of all N images are the same, it indicates that the trigger fingertip P is touching the functional area, and outputs or triggers the corresponding content of the functional area. Step 3, which is a method for realizing a tactile typing or touch with a sense of reality, characterized by including these steps.
10. The method for realizing a tactile typing or touch according to claim 9, characterized in that the touch interface image is a conventional calculator diagram and a conventional keyboard diagram.
11. The method for realizing a tactile typing or touch according to claim 9, characterized in that the surface of the preset object is any physical surface.
12. The method for realizing a tactile typing or touch according to claim 9, characterized in that the surface of the preset object is the surface of a virtual object, and when the trigger fingertip touches the functional area, it feeds back to the user by means of sound, vibration, electric shock, or a machine to give a feeling of touching a physical object.
13. A head-mounted display device, characterized by including at least two cameras for imaging the target image of the target area, further including a memory for storing a computer program, and a processor for executing the computer program to realize the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Visual assistance method and AR glasses
CN116560089A
Virtual keyboard-based text input method and device
JP2023538687A
System and method for human computer interaction
US20150244911A1
Input device interaction
US20170090747A1