METHOD FOR OBTAINING A KEYSTICK OR TOUCH COMMAND WITH TACTILE FEEDBACK
The method uses parallax video streams and relative positional calculations to accurately determine virtual key touches in XR devices, addressing visual blocking and tactile feedback issues, enabling blind typing and enhancing user experience.
Patent Information
- Application Number
- FR2024011270
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-10-17
- Publication Date
- 2025-07-04
AI Technical Summary
Existing virtual keyboards and touch controls in extended reality (XR) devices face challenges in accurately determining whether a trigger fingertip has touched a function region due to visual blocking and lack of tactile feedback, leading to inaccurate input and the inability to perform blind typing.
A method using a wearable XR device with at least two cameras to capture parallax video streams, determining the relative positional relationship of the trigger fingertip to trigger determination points, and calculating a ratio to confirm accurate touch without needing depth information, enabling tactile feedback through virtual or real object interaction.
Enables accurate determination of virtual key touches with tactile feedback, allowing blind typing and improving user experience by eliminating the need for auxiliary sensors and providing tactile sensation during typing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: METHOD FOR OBTAINING A KEYSTICK OR TOUCH COMMAND WITH TACTILE FEEDBACK Technical field
[0001] The present invention relates to the fields of virtual keyboards and touch control technologies, and more particularly relates to a method for achieving touch typing or control with tactile feedback applicable to a wearable extended reality (XR) device or more particularly to an extended reality headset.
[0002] BACKGROUND OF THE INVENTION
[0003] Extended reality (XR) refers to a combined reality and virtuality environment enabling human-machine interaction realized by computer technologies and human-wearable devices, and is a general term that encompasses augmented reality (AR), virtual reality (VR), and mixed reality (MR). With the popularization and development of XR in various industries, various XR glasses have emerged, which enable interaction between a user and a system through inputs via a virtual keyboard and touch control.
[0004] There are currently two types of virtual keyboards and touch controls: (1) a virtual keyboard anchored in a 1 / 3 / 6 degree of freedom (1 / 3 / 6DoF) three-dimensional environment, where a keystroke or touch control is performed in the air by both hands, and positions of fingertips or rays are calculated using a joint recognition model to determine whether threshold positions of virtual keys have been touched or not;and (2) virtual keys assigned to the palm and fingers, where one end (or any part that can be focused as a cursor point) of a thumb (or any finger) is generally defined as "the trigger fingertip", and the virtual keys are assigned to three phalangeal joints of each of the other fingers and / or to different regions of the palm, and the virtual keys are defined to correspond to different number keys, letter keys or function keys, respectively, and a human hand joint detection model is used to determine whether or not the trigger fingertip touches threshold positions of the virtual keys. ;
[0005] The first type of virtual keyboard mentioned above (which allows various functions such as pressing a key, accessing a link and computer drawings, the relevant sections of the virtual keyboard performing these functions are collectively referred to as "function regions" hereinafter) has an input method similar to that of typing on a conventional keyboard and triggering a cursor pointer, but there are two problems: (a) since function regions are often blocked by the backs of hands and fingers, it is difficult to determine whether or not a non-visible trigger fingertip actually touches a threshold position of a certain function region during visual detection and calculation; and (b) tactile feedback of using a physical keyboard is lacking during use, and the user can only determine whether or not a trigger fingertip touches a correct character key by their own visual determination when typing in the air, so that blind typing / touch typing is impossible.
[0006] The second type of virtual keyboard triggering the function regions of the palm and fingers is similar to finger gestures in various ways, as in traditional Chinese Taoism. The function regions are assigned to the visible and detectable palms and fingers. Since the palms (for simplicity, the term "palm" used in this specification from here on refers to all parts of the palm, including the fingers, where detection and determination are required, and the part of the "palm" without the fingers is specifically referred to as "palm center") are oriented toward a camera of the XR glasses during input, and the trigger fingertips are used to touch the function regions on the palms to trigger input, the problems of tactile feedback and blocking by the backs of the hands can be solved.However, the problem of visual blocking of function regions by trigger fingertips still exists. When the trigger fingertip is positioned above a certain function region, it is not possible to determine by visual detection and calculation whether the trigger fingertip is touching the function region or is still far from the function region in an untouched state; therefore, the trigger fingertip may be inaccurately considered to be touching the corresponding function region and thus a function of that function region may be triggered by mistake.In order to solve the problem of uncertain determination of a touch by visual detection and calculation or by a gesture recognition model, many patents have attempted to accurately determine whether a trigger fingertip actually touches a certain function region or not by means of sensor-equipped rings or sensor-equipped gloves. However, wearing sensors in the form of gloves or rings is unwelcome to the user and inconvenient, as users generally do not want to wear devices or sensors during use.
[0007] BRIEF SUMMARY OF THE INVENTION
[0008] The present invention addresses the problem of inaccurate determination of whether a virtual key is touched or not in existing gesture recognition and visual calculation technologies by providing a method for achieving a touch keystroke or control with tactile feedback. The present invention can accurately confirm whether the trigger fingertip actually touches a function region or not simply by calculation performed on visual images captured by cameras without using any auxiliary physical devices such as sensors, and the present invention requires less calculation.Additionally, when the trigger fingertip touches a palm or object surface, instead of gesturing in the air without tactile feedback, a tactile sensation is obtained during a touch typing or command, which improves the user experience and enables blind typing / touch typing.
[0009] The present invention relates to a method for obtaining a keystroke or touch command with tactile feedback, implemented by a system configured in a wearable extended reality (XR) device or an extended reality headset; the system outputting position information, which carries a time sequence, of articulation points of a human hand captured in a video stream of each camera of the system via a human hand articulation detection model; characterized in that a keystroke and touch command are obtained via a trigger fingertip touching virtually assigned function regions on a palm of the human hand, the palm being defined to include both a palm center without the fingers, and also the fingers; each of the function regions being a character or number button, a function key or a shortcut key which is capable of being triggered;and a corresponding function region being assigned and fixed to a predefined point marked on a virtual joint line of each pair of two adjacent joints of the palm; said method comprising the following steps: ;
[0010] Step 1: Marking the predefined point on the virtual joint line of each pair of two adjacent joints of the palm, the corresponding function region assigned to each predefined point on the palm being capable of being visually perceived using a pair of smart glasses; setting a width of each of the function regions to W; a corresponding predefined point of each of the function regions being determined as a center point of this function region; setting a left trigger determination point WL and a right trigger determination point WR at positions W / 2 to the left and W / 2 to the right relative to the center point of each of the function regions respectively along a direction parallel to an X axis; determining the position information of each predefined point and the trigger determination point left trigger WL and right trigger determination point WR of the corresponding function region assigned to each predefined point based on the articulation points;
[0011] Step 2: consider by default one end of a thumb as the trigger fingertip; if the thumb does not have access to areas on the palm, but one end of any finger among the other fingers is intended to touch the palm and the function regions assigned to the palm, the end of said any finger among the other fingers is determined to be the trigger fingertip; the trigger fingertip is identified as P;
[0012] Step 3: The system acquires a number N of parallax video streams from at least two cameras, N being an integer, and N>2, tracking and determining whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to left and right sides of a corresponding feature region, respectively, in all of the corresponding number N of images carrying a same time from said number N of video streams, respectively, and, if so, calculating the position information of three target points T which are the left trigger determination point WL, the trigger fingertip P and the right trigger determination point WR for each of said number N of images carrying the same time;then, in each of said number N of images carrying the same time, X-axis values (including WRX of the right trigger determination point WR, PX of the trigger fingertip P, and WLX of the left trigger determination point WL) of the position information of the three target points T in that image are used to calculate a ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all the ratios calculated in all of said number N of images carrying the same time are the same, the trigger fingertip P is determined to have touched the corresponding function region, and then a function corresponding to the corresponding function region is output or triggered. ;
[0013] Advantageously, in step 3 above, said at least two cameras comprise two cameras which are a left camera and a right camera; a connecting line passing through two central points L and R of the left camera and the right camera respectively being considered as the X axis; assuming that, in a field of view of the left camera, an included angle defined as T0L is formed between the X axis and a connecting line connecting the central point L of the left camera and one of the three target points T; assuming that, in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and one of the three target points T; assuming that a length of a parallax baseline between the two center points L and R of the left camera and the right camera is d, calculating a position (X, Z) of any one of the three target points T in each of said number N of images carrying the same time, on the basis of the following:
[0014] if the target point T whose position is to be calculated is located between the two central points L and R of the left camera and the right camera:
[0015] Z=d / [TAN(T0L)-TAN(T0R-Jt / 2)], X=Z*TAN(T0L);
[0016] if the target point T whose position is to be calculated is located on a left side of the central point L of the left camera:
[0017] Z=d / [TAN(T0R-Jt / 2)-TAN(T0L-Jt / 2)], X=-Z*TAN(T0L);
[0018] if the target point T whose position is to be calculated is located on a right side of the central point R of the right camera:
[0019] Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L).
[0020] Advantageously, each of the function regions has a circular shape; a circle is drawn for each of the function regions with the predefined point which is defined at any position on the joint line between each pair of two adjacent joints of the palm as the center of the circle and the width W of each of the function regions as the diameter.
[0021] Advantageously, each of the function regions is assigned to the interior of a phalangeal region, between two phalangeal regions, on an outer side of the phalangeal region, or to a certain area of the center of the palm between a wrist and a certain finger.
[0022] Advantageously, in step 3 above, the system virtually assigns a matrix grid to a same position on the palm center when a number N of video streams are processed, N being an integer, and N>2; the matrix grid comprises a plurality of grid units each having a plurality of sides, and each grid unit is considered as a function region; the system tracks and determines whether or not the trigger fingertip P (X, Y) is located in a same function region of the matrix grid across all of said number N of viewer screens and, if so, the trigger fingertip P (X, Y), and the left trigger determination point WL and the right trigger determination point WR on the left and right sides of the function region concerned are determined as the three target points T;then X-axis values (including WRX of the right trigger determination point WR, PX of the trigger fingertip P and WLX of the left trigger determination point WL) of the position information of the three target points are used to calculate a ratio (PX-WRX): (WLX-PX) which is a difference in value between PX and WRX with respect to one; difference in value between WLX and PX; if calculated ratios in all corresponding frames carrying a same time in said number N of video streams are identical, the trigger fingertip P is determined to have touched the function region, and then a point or line is drawn corresponding to a position of the trigger fingertip P (X, Y) and, as the trigger fingertip P moves, a series of determinations by the system over time carrying a time sequence will determine points or lines drawn successively in successive locations over time and thus determine that a line is drawn where all the drawn points or lines are determined to be joined together; therefore, tablet control or touch control functions can be implemented in the palm of one hand by using a fingertip of another hand as the trigger fingertip.
[0023] Advantageously, a connecting point between a little finger and the palm center is determined as an upper right corner of the matrix grid; a connecting point between an index finger and the palm center is determined as an upper left corner of the matrix grid; and a connecting line between the palm center and a wrist is determined as a lower edge of the matrix grid.
[0024] Advantageously, the matrix grid is invisible and not displayed on said number N of screens of the viewer.
[0025] Advantageously, each grid unit has a square or rectangular shape.
[0026] The present invention also relates to another method for obtaining a keystroke or touch command with tactile feedback, implemented by a system configured in a wearable extended reality (XR) device or an extended reality headset; the system outputting position information, carrying a time sequence, of target points captured by videos; characterized in that a keystroke and touch command are obtained through a trigger fingertip touching function regions; said method comprising the following steps:
[0027] Step 1: The system anchors a touch control interface image on each screen of the viewer at a same position of a same predefined object surface; a plurality of said function regions are assigned to the touch control interface image as viewed from any one of the number N of screens of the viewer; corresponding images from all videos carrying a same time are each determined to have a left trigger determination point WL and a right trigger determination point WR on the left and right sides of a respective corresponding function region respectively along a direction parallel to an axis X of the image from a corresponding video;
[0028] Step 2: One end of any finger intended to touch the function regions is determined to be the trigger fingertip;
[0029] Step 3: The system acquires a number N of parallax video streams, where N is an integer, and N>2; tracking and determining whether the trigger fingertip P(X,Y) is located in a same function region or not in all of the corresponding number N of images carrying a same time from said number N of video streams, respectively, and if so, the trigger fingertip P(X,Y) and the left trigger determination point WL and the right trigger determination point WR corresponding to the relevant function region are used as three target points T;X-axis values (including WRX of the right trigger determination point WR, PX of the trigger fingertip T, and WLX of the left trigger determination point WL) in the position information of the three target points T are used to calculate a ratio (PX-WRX):(WLX-PX) which is a value difference between PX and WRX relative to a value difference between WLX and PX; only when all ratios of said number N of images bearing the same time are the same, the trigger fingertip P is determined to have touched the function region, and then a function corresponding to the function region is output or triggered. ;
[0030] Advantageously, the touch control interface image is an image of a conventional numeric keypad or a conventional keyboard.
[0031] Advantageously, the predefined object surface is any surface of a real object.
[0032] Advantageously, the predefined object surface is a surface of a virtual object and, when the trigger fingertip touches a corresponding function region, feedback in the form of sound, vibration, electric shock or other mechanical feedback is provided to create a sensation of touching a real object.
[0033] The present invention further relates to a head-mounted display device, comprising at least two cameras configured to take videos or images of a targeted region; the head-mounted display device also comprising a memory and a processor; the memory being configured to store a computer program; the processor being configured to execute the computer program to perform one of the methods described above.
[0034] According to the technical solutions provided by the present invention, video streams with parallax are captured by at least two cameras of a pair of smart glasses, respectively, and corresponding images from all the video streams having a same time are used to determine whether a function region is affected or not, a connecting line which passes through the center point of said at least two cameras being considered as an X-axis or as being parallel to an X-axis, the tip of trigger finger P, the left trigger determination point WL and the right trigger determination point WR of the function region that the trigger fingertip P intends to touch are regarded as the three target points T, and the X-axis values of the position information of the three target points T are used to calculate the ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all ratios of the number N of frames carrying the same time are the same, the trigger fingertip P is determined to have touched the function region. Accordingly, formulas determining a depth Z of a spatial position of any of the target points to be calculated in a field of view of each camera may omit the Y-axis from the calculation.Since a distance defined by the parallax baseline between each pair of adjacent cameras is fixed, a distance defined by the parallax baseline between two eyes of a viewer is unchanged, such that, when the trigger fingertip actually touches a feature region, both the left eye and the right eye of the viewer will perceive a same relative position of the trigger fingertip on or along the X-axis of the feature region, more precisely, when three target points lie on a same line, the relative positioning between the three target points will be the same, even if they are viewed from different angles by the left and right cameras (or even more).Thus, it is proved that it is not necessary to know an exact value for a depth Z before determining whether the trigger fingertip has touched the function region or not, and a value d which is a distance of the parallax baseline between two cameras can be any value during calculation. In the present invention, the touch control tolerance is Z / ad, where Z is a depth between a target point and a corresponding camera, d is a distance of the parallax baseline between two cameras, and a is a threshold value.According to the present invention, in the field of view of the smart glasses, whether or not the trigger fingertip has actually touched a function region is determined by the relative positional relationship of the trigger fingertip with respect to the two trigger determination points on the left and right sides of the function region, respectively, in each of the set of corresponding frames carrying the same time of all the video streams captured by the cameras, and if the determined relative positional relationships in the set of corresponding frames carrying the same time are the same, it is determined that the trigger fingertip has touched the function region, otherwise, it is determined that there has been no touch, so it is not necessary to know the exact values of X, Y, Z and d before making the determination; only the X-axis pixel value of each of the number N>2 of cameras. is required. The present invention solves the problem of inaccurate determination of whether a virtual key has been touched or not in existing gesture recognition and visual computing technologies by providing a method for obtaining a virtual keystroke or touch command with tactile feedback.
[0035] Since the present invention does not need to know the exact value of the depth Z of the trigger fingertip before determining whether the trigger fingertip has touched the function region or not, the present invention can also anchor a matrix grid virtually on a palm center, and treat each grid unit as a function region. At least two video streams are captured by at least two cameras of a pair of smart glasses corresponding to at least one left eye and one right eye of the user respectively, and a parallax baseline of distance d exists between said at least two cameras.Determining a relative positional relationship between the trigger fingertip P (X,Y) and two trigger determination points of a grid unit corresponding to a same height Y where the trigger fingertip P is positioned in each of the set of corresponding images from said at least two video streams carrying a same time; if the relative positional relationship of the three target points is the same in all of said corresponding images from said at least two video streams carrying the same time, it is determined that the trigger fingertip has touched the corresponding function region, otherwise, it is determined that there has been no touch.If a successful touch is determined, a point or stroke is drawn corresponding to a position of the trigger fingertip P (X,Y) and, as the trigger fingertip P moves, a series of determinations in time carrying a time sequence will determine successively drawn points or strokes in successive locations in time and thus determine that a line is drawn where all drawn points or strokes are determined to be joined together. Therefore, drawing, writing and dragging functions can be implemented in the palm of one hand using a fingertip of another hand as the trigger fingertip, just like implementing a touch control function on a tablet or touch screen.In addition, the touch control function of a tablet or touch screen with multiple fingers can also be implemented by using multiple trigger fingertips. In addition, a three-dimensional point or line P (X, Y, Z) can also be drawn by calculating a depth Z of the trigger fingertip by trigonometry.
[0036] In addition to anchoring a numeric keypad, keyboard, and drawing pad virtually on the palm, the present invention also enables virtual typing and touch control outside of the palm. More specifically, the Smart glasses can project an image of a numeric keypad or simple keyboard onto a certain object surface, such as a wall surface or a table surface, or any surface of another real or virtual object, and the object surface may not be a flat surface but a wavy surface.At least two video streams having parallax are captured by at least two cameras of the smart glasses, respectively; when the trigger fingertip accesses a feature region of the projected image, a relative positional relationship between the trigger fingertip and two trigger determination points of the corresponding feature region is determined in each of the set of corresponding images carrying a same time from the at least two video streams; if the positional relationship of the three target points is the same in the set of corresponding images carrying the same time, it is determined that the trigger fingertip has touched the feature region; otherwise, it is determined that there has been no touch.As a result, a user can perform a keystroke or touch command on a real object surface instead of tapping or controlling virtual keys projected in the air, thereby obtaining tactile feedback during the keystroke or touch command. Brief description of the drawings
[0037] [Fig.l] represents 21 recognizable points of articulation of a human hand and their numerical identifiers given by the official Mediapipe website;
[0038] [Fig.2] is a schematic diagram of calculating a spatial position of a target point T by a left camera of a pair of smart glasses according to the present invention;
[0039] [Fig.3] is a schematic diagram of calculating a spatial position of a target point T by a right camera of the pair of smart glasses according to the present invention;
[0040] [Fig.4] are schematic illustrations of a W function region disposed on a palm when the palm is in different orientations according to the present invention;
[0041] [Fig.5] is a combined image from a left camera and a right camera showing relative positioning of the trigger fingertip and the two trigger determination points when the trigger fingertip is not touching a function region;
[0042] [Fig.6] is a combined image from a left camera and a right camera showing relative positioning of the trigger fingertip and the two trigger determination points when the trigger fingertip touches a function region;
[0043] [Fig.7] illustrates four separate images showing relative positioning of the tip of trigger finger and the two trigger determination points, the upper two images showing images from the left camera and the right camera, respectively, when the trigger fingertip does not touch a function region, and the lower two images showing images from the left camera and the right camera, respectively, when the trigger fingertip touches a function region;
[0044] [Fig.8] is a schematic diagram of an arrangement of function regions in the form of a numeric keypad on a hand according to the present invention;
[0045] [Fig.9] is a schematic diagram of an arrangement of function regions of a conventional QWERTY keyboard on both hands according to the present invention;
[0046] [Fig. 10] is a schematic diagram of triggering a function region at a tip end of an index finger by a trigger fingertip according to the present invention;
[0047] [Fig. 11] is a schematic diagram of triggering a function region at a distal phalangeal region of an index finger by a trigger fingertip according to the present invention;
[0048] [Fig. 12] is a schematic diagram of triggering a function region at an intermediate phalangeal region of an index finger by a trigger fingertip according to the present invention;
[0049] [Fig. 13] is a schematic diagram of triggering a function region at a proximal phalangeal region of an index finger by a trigger fingertip according to the present invention;
[0050] [Fig. 14] is a schematic diagram of triggering a function region at a lower end of the proximal phalangeal region of an index finger by a trigger fingertip according to the present invention;
[0051] [Fig. 15] is a schematic diagram of triggering a function region at a position of a palm near the wrist by an end of an index finger serving as a trigger fingertip according to the present invention;
[0052] [Fig. 16] is a structural block diagram of a head-mounted display device according to the present invention;
[0053] [Fig. 17] is a schematic diagram of an XY matrix grid assigned to a palm center for implementing touch control functions of sliding, writing and drawing on the palm according to the present invention;
[0054] [Fig. 18] is a schematic diagram of an XY matrix grid assigned to a palm center and shortcut keys assigned to phalangeal regions according to the present invention;
[0055] [Fig. 19] is an image of a numeric keypad anchored to any object surface in a 1 / 3 / 6DoF three-dimensional environment according to the present invention; and
[0056] [Fig.20] is an image of a keyboard anchored to any object surface in a three-dimensional environment of 1 / 3 / 6DoF according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0057] The technical solutions in the embodiments of the present invention will be described clearly and in detail below, with reference to the accompanying drawings of the embodiments of the present invention. It is obvious that the described embodiments illustrate only some of the embodiments of the present invention, but not all. Based on the embodiments of the present invention, all other embodiments obtainable by the person skilled in the art without any inventive effort are within the scope of the present invention.
[0058] Further, the terms "include" and "comprising," and any variations thereof, are intended to be non-exclusive inclusion. For example, a process, method, system, product, or server comprising a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are known for such a process, method, product, or device.
[0059] The principles of the technical solutions of the present invention are as follows:
[0060] (1) Recognition model used to acquire position information on a palm: Currently commercially available open source software for a pre-trained human hand joint detection model that can acquire two-dimensional positions of human hand joints can be used. The present invention uses Mediapipe as an illustrative example. As an open project of Google, Mediapipe is a tool library for machine learning that mainly concerns visual algorithms, and integrates a large number of models relating to facial detection, facial keypoints, gesture recognition, head segmentation, and posture recognition. As shown in [Fig.l], position information of 21 joint points (also called keypoints) of a human hand in a video carrying a time sequence can be output.In general, a human hand joint detection model outputs joint position information in the form of pixels (X, Y) which are coordinates of an X axis and a Y axis of the video. The present invention may also use a self-trained human hand joint detection model. The present invention also includes training to recognize whether a trigger fingertip is within a function region using an artificial intelligence chip such as . than a graphics processing unit (GPU) or a neural network processing unit (NPU) by label convolution KNN or RNN or by a Transformer model plus pre-training processes.
[0061] (2) Definition of function regions on the palm: the position information of the 21 articulation points (also called key points) of a human hand in a video in the form of pixels (X, Y) carrying a time sequence can be output using the existing human hand articulation detection model. The present invention marks a predefined point on a joint line of each pair of two adjacent joints of the palm (said predefined point may be a center point of the joint line). The user can see a corresponding function region assigned to each predefined point on the palm via a pair of smart glasses, and each function region may be a character button, a function key or a shortcut key that can be triggered.A width of each function region is defined as W, a corresponding preset point of each function region is determined as a center point of that function region, and two trigger determination points WL and WR are defined at positions W / 2 to the left and W / 2 to the right relative to the center point respectively along a direction parallel to the X axis. The position information of each preset point and the two trigger determination points WL and WR of the function region assigned to each preset point can be determined based on the articulation points. Each function region can have any shape. Preferably, each function region has a circular shape, since the display effect of the function region on the palm is not affected regardless of the rotation of the palm. A circular function region is shown in [Fig.4], in which a circle is drawn with a predefined point set at any position on the joint line between two adjacent joints as the center of the circle and a width W of the function region as the diameter. The present invention may assign each function region to the inside of a phalangeal region, between two phalangeal regions, on an outer side of a phalangeal region, or to a certain area of the palm center between the wrist and a certain finger.
[0062] (3) Definition of trigger fingertip: one end of a thumb serves as the tip Default trigger fingertip. If the thumb does not have access to areas on the palm or the tip of the thumb does not serve as a trigger fingertip, any fingertip of any other finger intended to touch the palm and the function regions assigned to the palm will be determined as the trigger fingertip.
[0063] An illustrative example of an arrangement of function regions on a palm in the form of a numeric keypad is shown in [Fig.8]. Also, with reference in [Fig. 10], when the trigger fingertip touches any of the function regions on the tip end of any of the other fingers, a character "C" corresponding to the index finger, a symbol " / " corresponding to the middle finger, a character "X" corresponding to the ring finger, or a function key "delete" corresponding to the little finger will be triggered. As shown in [Fig. 11], when the trigger fingertip touches any of the function regions on a distal phalangeal region of any of the other fingers, a character "1" corresponding to the index finger, a character "2" corresponding to the middle finger, a character "3" corresponding to the ring finger, or a symbol "-" corresponding to the little finger will be triggered. As shown in [Fig.12], when the trigger fingertip touches any of the function regions on a middle phalangeal region of any of the other fingers, a character "4" corresponding to the index finger, a character "5" corresponding to the middle finger, a character "6" corresponding to the ring finger, or a symbol "+" corresponding to the little finger will be triggered. As shown in [Fig. 13], when the trigger fingertip touches any of the function regions on a proximal phalangeal region of any of the other fingers, a character "7" corresponding to the index finger, a character "8" corresponding to the middle finger, a character "9" corresponding to the ring finger, or a symbol "=" corresponding to the little finger will be triggered. As shown in [Fig.14], when the trigger fingertip touches any of the function regions on a lower end of the proximal phalangeal region of any of the other fingers, a "%" symbol corresponding to the index finger, a "0" character corresponding to the middle finger, a "." symbol corresponding to the ring finger, and an "=" symbol corresponding to the little finger will be triggered. It is seen that, if the function regions are assigned to the apices (tip ends) of the fingertips, to different phalangeal regions, or to positions on the palm center close to corresponding phalangeal regions, the function regions can be triggered when the thumb serves as the trigger fingertip to touch the function regions. However, the thumb cannot easily touch the function regions assigned to positions on the palm center close to the wrist.Therefore, the present invention will assign corresponding fingers to trigger these function regions, wherein the triggering fingertips are the fingertips corresponding to said assigned corresponding fingers instead of the tip of the thumb, so that the corresponding function regions can be triggered to output characters / functions by touching the corresponding function regions with the fingertips of said assigned corresponding fingers. As shown in Figures 8 and 15, a function key "MC" can be triggered by the index finger, a function key "M+" can be triggered by the middle finger, a function key "M-" can be . triggered by the ring finger, and the “MR” function key can be triggered by the little finger.
[0064] (4) Calculation of a spatial position of a target point: although smart glasses XR smart glasses allow viewing a three-dimensional space (X-axis, Y-axis, and Z-axis), the Y-axis can actually be omitted when calculating positions of the trigger fingertip and the left and right trigger determination points along the X-axis direction of the function regions, so that the calculation is simplified to a two-dimensional position calculation. As illustrated in [Fig.2], by treating a connecting line passing through L / R center points of both left and right cameras of the XR smart glasses as the X-axis, and again referring to [Fig.2] in which a field of view of the left camera is shown, an included angle defined as T0L is formed between the X-axis and a connecting line connecting the center point L of the left camera and a target point T whose spatial position is yet to be calculated; more precisely, an included angle between the X-axis and a connecting line connecting the center point L of the left camera and the trigger determination point WL is defined as WL0L, an included angle between the X-axis and a connecting line connecting the center point L of the left camera and the trigger determination point WR is defined as WR0L; similarly, as illustrated in [Fig.3] in which a field of view of the right camera is shown, an included angle defined as T0R is formed between the X-axis and a connecting line connecting the center point R of the right camera and a target point T whose spatial position is yet to be calculated; more specifically, an included angle between the X-axis and a connecting line connecting the center point R of the right camera and the trigger determination point WL is defined as WL0R, an included angle between the X-axis and a connecting line connecting the center point R of the right camera and the trigger determination point WR is defined as WR0R, and an included angle between the X-axis and a connecting line connecting the center point R of the right camera and the trigger fingertip T is defined as T0R.
[0065] The trigger fingertip P and the left and right trigger determination points WL and WR are three target points T whose spatial positions are to be calculated. A length of a parallax baseline between the two center points L and R of the left and right cameras is assumed to be d. A position (X, Z) of any target point T can be calculated based on the following:
[0066] If the target point T is located between the two central points L and R of the left and right cameras:
[0067] Z=d / [TAN(T0L)-TAN(T0R-Jt / 2)], X=Z*TAN(T0L).
[0068] If the target point T is located on a left side of the central point L of the left camera:
[0069] Z=d / [TAN(T0R-jt / 2)-TAN(T0L-jt / 2)], X=-Z*TAN(T0L).
[0070] If the target point T is located on a right side of the central point R of the right camera:
[0071] Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L).
[0072] The above examples are calculated using TAN (tangent) and COT (cotangent), but any other trigonometric calculation method can be used by the present invention.
[0073] (5) Method for determining whether the trigger fingertip touches the region or not of function:
[0074] A system of the present invention acquires video streams from the left and right cameras (or from more than these two angles) having parallax, and then makes a determination on a left image and a right image (or a plurality of images in the case of more than these two images) corresponding to a same time in the video streams from the left and right cameras (or from more than these two angles), respectively. If the trigger fingertip P is located between two trigger determination points WL and WR of any function region, a radian ratio (P0L - WR0L):(WL0L - P0L) of the left image is compared with a radian ratio (P0L - WR0R):(WL0R - P0L) of the right image. If the two ratios are not equal, it is determined that the trigger fingertip does not touch the function region, as illustrated in [Fig. 5] and the upper two images of [Fig. 7].If the two ratios are equal, it is determined that the trigger fingertip touches the function region, as shown in [Fig.6] and the lower two images of [Fig.7], and then the function corresponding to the function region is output.
[0075] In the present invention, when at least two numerical values are compared, values that deviate within a threshold range are always considered to be the same, equal, or consistent. A general threshold value for the tolerance may be set to about Z / 5d.
[0076] Since the fields of view (FOV) captured by different cameras are different, a pixel value X of the X-axis acquired by the human hand joint detection model can be directly converted to radian / angle 0 in all the above formulas. Assuming that a total resolution of the X-axis of an image is 1800 pixels, the FOV of a corresponding camera is 180 degrees, and the pixel value X of the X-axis of (X, Y) of the target point T returned by the human hand joint detection model is pixel 900, then the radian 0 of the target point is ir / 2 (the angle is 90). Since the present invention only needs to compare relative radian ratios formed by three target points (WL, P, WR) of the images of the left and right cameras (plurality of cameras), the relative radian ratios formed by the three target points T can be calculated directly. using the X-axis pixel value X of the target point T returned by the human hand joint detection model without the need to convert to an absolute radian or angle 0. Therefore, assuming 0 is the X-axis pixel value X output by the human hand joint detection model, the radian ratio of the left image is (PX L - WRX L ):(WLX L - PX L ), and the radian ratio of the right image is (PX R - WRX R ):(WLX R - PX R ).
[0077] (6) Examples of arrangements of function regions on the palm:
[0078] [Fig.8] is an example of an arrangement of function regions in the form of a numeric keypad on a palm, where virtual typing can be performed by tapping different areas of the palm.
[0079] [Fig.9] is an example of an arrangement of function regions on both palms in the form of a conventional QWERTY keyboard, where virtual typing can be performed by tapping on different areas of both palms.
[0080] By adopting the technical solutions of the present invention, the position of each function region and the corresponding character (or function key / shortcut key) assigned to the function region can be set by a user based on typing habit and convenience of use.As long as the function region is defined at a joint of the palm or any position of a joint line between two adjacent joints of the palm, the position information of the function region can be obtained from the position information of the joint points carrying a time sequence as output by the human hand joint detection model, and the position information of the two trigger determination points corresponding to the function region can also be obtained, so that it is possible to determine whether or not the trigger fingertip touches the function region.
[0081] (7) Principles of implementing touch control on the palm center:
[0082] Since images captured by a camera comprise two-dimensional pixel data (X and Y), the present invention can implement two-dimensional touch control functions on a flat surface of the palm center, such as drawing, writing, sliding, pulling and other two-dimensional actions.
[0083] According to a system of the present invention, a matrix grid (which may be visible or invisible) is virtually assigned to a same position on the palm center on a viewer screen of each camera of the smart glasses. [Fig. 17] shows an illustrative example of a left hand, where a connecting point between the little finger and the palm center is determined as the upper right corner of the matrix grid, a connecting point between the index finger and the palm center is determined as the upper right corner left of the matrix grid, and a connecting line between the palm center and the wrist is determined as the bottom edge of the matrix grid. When the palm rotates and moves, a position of the matrix grid is always fixed relative to the palm center, because the matrix grid is fixed at the palm joint points. Each grid unit of the matrix grid has four lines corresponding to four sides which are the top, bottom, left and right sides. The matrix grid is not limited to a square shape, but can have any shape, for example, a triangular matrix grid has three sides, and a hexagonal matrix grid has six sides.Alternatively, the matrix grid may have an irregular shape, where grid units of the irregular matrix grid may have different numbers of sides or different patterns, and each of the grid units may be regarded as a corresponding feature region, and the "method for determining whether or not the trigger fingertip touches the feature region" as mentioned in item (5) above may be used to determine whether or not the trigger fingertip touches the feature region.
[0084] According to the system of the present invention, the trigger fingertip P (X, Y) is tracked and determined whether or not it is located in a same function region of the matrix grid when it is tracked and determined by both left and right cameras (or a plurality of cameras), and then the trigger fingertip P (X, Y) and the left and right sides of the function region are determined as three target points T (the left trigger determination point WL, the trigger fingertip P and the right trigger determination point WR). X-axis values (WRX, PX and WLX) of the position information of the three target points T are used to calculate a ratio (PX-WRX).(WLX-PX) (a value difference between PX and WRX with respect to a value difference between WLX and PX).If ratios of corresponding images determined by all cameras are identical, it is determined that the trigger fingertip P has touched the function region, and a point or line is then drawn corresponding to a position of the trigger fingertip P, and, as the trigger fingertip P moves, a series of determinations by the system in time carrying a time sequence will determine points or lines drawn successively in successive locations in time and thus determine that a line is drawn at the location where all the drawn points or lines are determined to be joined together. Therefore, drawing, writing, and dragging functions can be implemented in the palm of one hand by using a fingertip of another hand as the trigger fingertip, just like implementing a touch control function on a tablet or touch screen.In addition, the touch control function of a tablet or touch screen with multiple fingers can also be implemented. implements using multiple trigger fingertips. Furthermore, a three-dimensional point or line P(X,Y,Z) can also be drawn by calculating a depth Z of the trigger fingertip by trigonometry, the radian ratio formulas being in accordance with what has been disclosed above, a pixel ratio converted to radian in the left image (from the left camera) being (PX L - WRX L ):(WLX L - PX L ), and a pixel ratio converted to radian in the right image (from the right camera) being (PX R - WRX R ):(WLX R - PX R ); if there are N cameras (N is an integer and N>2), the radian ratios (PX N - WRX N ):(WLX N - PX N ) of all cameras must be the same (or within a predefined threshold difference) to determine an actual touch, otherwise there is no touch.
[0085] It should be noted that, since the palm can rotate at will, the corresponding matrix grid also rotates with the palm. Therefore, when the trigger fingertip is in the same grid unit (function region) determined, for example, by a left camera and a right camera, values of WLX and WRX on the left and right sides may change in real time since the same height Y remains unchanged, so that the calculation for confirming the touch according to the present invention must perform a comparison of corresponding images carrying the same time.
[0086] [Fig. 18] is a schematic diagram of an XY matrix grid displayed in a palm center and shortcut keys displayed in phalangeal regions, thereby integrating tablet touch control functions and shortcut keys.
[0087] (8) Existing smart glasses with IMU chips, 1 / 3 / 6 degrees of freedom (1 / 3 / 6DoF) can be used to anchor any image to a fixed position in a three-dimensional space around the user. In addition to the use of keyboards and keypads anchored to the palm, the present invention can also enable typing and touch control outside the palm. [Fig. 19] shows an image of a simple keypad anchored to an object surface, such as a wall or desk. [Fig. 20] shows an image of a keyboard that can be anchored to a wall or desk in the same way for use. The object surface can be an irregular surface. According to the present invention, video streams of images with parallax are acquired by at least two cameras of the smart glasses.When viewed from the at least two cameras that the trigger fingertip enters a function region, a relative positional relationship between the position of the trigger fingertip and two trigger determination points corresponding to the function region is determined for all corresponding images from all cameras carrying . at the same time. If the relative positional relationships of the three target points are consistent across said corresponding images acquired by all cameras, it is determined that the trigger fingertip has touched the function region; otherwise, the trigger fingertip has not touched the function region. Thus, a user touches a real object surface during typing or touch operation instead of performing touch operation on a virtual button in the air, thereby achieving tactile feedback during typing or touch operation.
[0088] The object surface may also be a surface of a virtual object and, when the trigger fingertip touches the function region, the user may receive feedback such as sound, vibration, electric shock, or other mechanical feedback, to create a sensation of touching a real object.
[0089] The present invention further includes various depth and speed sensors that can be used in conjunction with conventional camera sensors or used independently. Since the present invention relies on the relative positional relationship between the triggering fingertip and the two trigger determination points to determine whether or not there is an actual touch, a trigonometric calculation of a depth position is not necessary; however, monitoring and implementation can be performed by the present invention based on relative distances and ratios of the positions of three target points obtained by depth sensors such as laser SLAM, IR (infrared) tracking, and motion sensors. For example, a motion speed sensor outputs a moving pixel, which can be used by the present invention.The SLAM, while providing a Z value for each pixel in the X axis, can also provide an X value. The IR sensor and other time-of-flight (ToF) sensors, while providing the depth Z value, can also provide X and Y values for calculation by the present invention.
[0090] The present invention is suitable not only for palm tapping, but also for any interactive instructions that need to be combined with palm tapping or touch control. For example, the user can perform the following actions:
[0091] A. A ray is projected along a designated anchor position from a designated launching position of a hand, when the ray points to a target position which is a virtual key or link at a certain distance from the user, the cooperative touch control instruction by tapping the palm (including the fingers and their phalangeal regions) using the trigger fingertip can be executed according to the method of the present invention.
[0092] B. When the user touches a virtual screen or a virtual button link with the index finger, a short press instruction or a long press instruction may be cooperatively required by triggering a virtual key, for example, by touching a distal phalangeal region of the middle finger with the tip of the thumb, and thus the cooperative touch control instruction by tapping the palm (including the fingers and their phalangeal regions) using the trigger fingertip can be executed according to the method of the present invention.
[0093] C. Some smart glasses are equipped with eyeball tracking devices to determine the user's gaze angle according to the angle of the pupils of the left and right eyes, so as to project a ray in three dimensions; when the ray points to a target position which is a virtual key or a function region corresponding to a link at a certain distance from the user, the cooperative touch control instruction by tapping the palm (including the fingers and their phalangeal regions) using the trigger fingertip can be executed according to the method of the present invention.
[0094] D. Some smart glasses project a three-dimensional ray perpendicular to a central position of the glasses; when the ray points to a target position which is a virtual key or a function region corresponding to a link at a certain distance from the user, the cooperative touch control instruction by tapping the palm (including the fingers and their phalangeal regions) using the trigger fingertip can be executed according to the method of the present invention.
[0095] Embodiment 1
[0096] Embodiment 1 of the present invention relates to a method for achieving a keystroke or touch control with tactile feedback, applicable to a system with a wearable extended reality (XR) device or in particular an extended reality headset; wherein the system outputs position information, which carries a time sequence, of articulation points of a human hand captured in a video of each camera of the system via a human hand articulation detection model. As defined herein, a "palm" is understood to include both the palm and the fingers; a keystroke and touch control are achieved by touching function regions with a trigger fingertip; said method comprises the following steps:
[0097] Step 1: Marking a predefined point on a joint line of each pair of two adjacent joints of the palm, the user being able to visually perceive a corresponding function region assigned to each predefined point on the palm through a pair of smart glasses, and each function region being a character button, a function key or a shortcut key which is able to be triggered; setting a width of each function region as W, a corresponding predefined point of each function region function is determined as the center point of this function region, and setting two trigger determination points WL and WR at positions W / 2 to the left and W / 2 to the right relative to the center point respectively along a direction parallel to an X axis; determining the position information of each preset point and the two trigger determination points WL and WR of the corresponding function region assigned to each preset point based on the articulation points;
[0098] Each feature region has any shape; preferably, each feature region has a circular shape; a circle is drawn with a predefined point which is defined at any position on the hinge line between two adjacent hinges as the center of the circle and the width W of the feature region as the diameter.
[0099] The present invention may assign each function region within a phalangeal region, between two phalangeal regions, on an outer side of a phalangeal region, or to a certain area of the palm center between the wrist and a certain finger.
[0100] Step 2: Consider one end of a thumb as the default trigger fingertip; if the thumb does not have access to areas on the palm, one end of any of the other fingers intended to touch the palm is determined as the trigger fingertip;
[0101] Step 3: The system acquiring a number N of video streams with parallax from at least two cameras, tracking and determining whether or not the trigger fingertip P is located between the two trigger determination points WL and WR corresponding to the left and right sides of a corresponding function region, respectively, in all of the corresponding number N of images carrying a same time from said number N of video streams, respectively, and, if so, X-axis values (WRX, PX and WLX) of position information of three target points T (a left trigger determination point WL of the two trigger determination points, the trigger fingertip P and a right trigger determination point WR of the two trigger determination points) of each of said number N of images are used to calculate a ratio (PX-WRX).(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all ratios of said number N of images are the same, the trigger fingertip P is determined to have touched the function region, and then a function corresponding to the function region is output or triggered.
[0102] Said at least two cameras comprise two cameras, which are a left camera and a right camera; a connecting line passing through central points L and R of the left camera and the right camera is considered as an X axis; assuming that, in a field of view of the left camera, an included angle defined as T0L is formed between the X axis and a connecting line connecting the center point L of the left camera and one of the three target points T; similarly, assuming that, in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and one of the three target points T; assuming that a length of a parallax baseline between the two center points L and R of the left camera and the right camera is d; calculating a position (X, Z) of any one of the three target points T based on the following:
[0103] If the target point T whose position is to be calculated is located between the two central points L and R of the left camera and the right camera:
[0104] Z=d / [TAN(T0L)-TAN(T0R-Jt / 2)], X=Z*TAN(T0L);
[0105] If the target point T whose position is to be calculated is located on a left side of the central point L of the left camera:
[0106] Z=d / [TAN(T0R-jt / 2)-TAN(T0L-jt / 2)], X=-Z*TAN(T0L);
[0107] If the target point T whose position is to be calculated is located on a right side of the central point R of the right camera:
[0108] Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L).
[0109] The system virtually assigns a matrix grid to a same position on a palm center captured on at least two viewer screens corresponding to a user's left and right eyes; a connecting point between a little finger and the palm center is determined as an upper right corner of the matrix grid, a connecting point between an index finger and the palm center is determined as an upper left corner of the matrix grid, and a connecting line between the palm center and the wrist is determined as a lower edge of the matrix grid; the matrix grid includes a plurality of uniformly arranged grid units, each of which has a plurality of sides, and each grid unit is considered a feature region;the system tracks and determines whether or not the trigger fingertip P (X, Y) is located in a same function region of the matrix grid in the at least two screens of the viewer, and if so, the trigger fingertip P (X, Y) and the two trigger determination points corresponding to the function region concerned are determined as the three target points T (the left trigger determination point WL, the trigger fingertip P and the right trigger determination point WR); the X-axis values (WRX, PX and WLX) of the position information of the three target points T are then used to calculate a ratio (PX-WRX).(WLX-PX) which is a difference in value between PX and WRX with respect to a difference of; value between WLX and PX; if ratios of all corresponding images carrying the same time on said at least two screens of the viewer are equal, it is determined that the trigger fingertip has touched the function region, and a point or line is then drawn corresponding to a position of the trigger fingertip P (X, Y) and, as the trigger fingertip P moves, a series of determinations by the system in time carrying a time sequence will determine points or lines successively drawn in successive locations in time and thus determine that a line is drawn at the location where all the drawn points or lines are determined to be joined together. Therefore, tablet control or touch control functions can be implemented in the palm of one hand by using a fingertip of another hand as the trigger fingertip.
[0110] The matrix grid is invisible and is not displayed on said at least two screens of the viewer.
[0111] The present invention also relates to another method for achieving touch typing or control with tactile feedback, applicable to a system provided with a wearable Extended Reality (XR) device or in particular an Extended Reality headset; the system outputs position information, carrying a time sequence, of target points captured by videos, and achieves touch typing and control by touching function regions using a trigger fingertip; said method comprises the following steps:
[0112] Step 1, the system anchors a touch control interface image on each of the viewer screens, respectively, at a same position of a same predefined object surface; a plurality of function regions are assigned on the touch control interface image as viewed from any one of the N number of viewer screens; corresponding images from all videos carrying a same time are each determined to have a left trigger determination point WL and a right trigger determination point WR on the left and right sides of a respective corresponding function region respectively along a direction parallel to an X axis of the image from a corresponding video.
[0113] Step 2: One end of any finger intended to touch the function regions is determined to be a trigger fingertip P.
[0114] Step 3: The system acquires a number N of parallax video streams from a number N of cameras, where N is an integer, and N>2; track and determine whether or not the trigger fingertip P(X,Y) is located in a same feature region in all of the corresponding number N of images carrying a same time from said number N of video streams, respectively, and, if so, the fingertip trigger P (X, Y) and the left trigger determination point WL and the right trigger determination point WR corresponding to the function region concerned are used as the three target points; X-axis values (WRX, PX and WLX) in the position information of the three target points are used to calculate a ratio (PX-WRX).(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all ratios of said number N of images are the same, the trigger fingertip P is determined to have touched the function region, and then a function corresponding to the function region is output or triggered.
[0115] The touch control interface image is an image of a classic numeric keypad or a classic keyboard.
[0116] The predefined object surface is a wall surface, a desk surface or other.
[0117] It should also be understood by those skilled in the art that portions and algorithmic steps of various examples described with reference to the embodiments disclosed herein may be implemented using electronic hardware, computer software, or a combination thereof. In order to clearly illustrate the interchangeability of implementation using hardware and software, the components and steps of various examples have been described generally in terms of operational characteristics in the above description. Whether these features are performed using hardware or software depends on the constraints and conditions of the technical solutions proposed in the context of a particular use or design.One skilled in the art may implement the described features in various ways for each particular use example, and these various ways of implementation should not be considered as departing from the scope of the present invention.
[0118] More specifically, the steps of the method disclosed in embodiments of the present invention may be performed by a processor using hardware integrated logic circuits and / or software instructions. The steps of the method described with reference to embodiments of the present invention may be directly performed by being performed by a hardware encoding processor or performed by a combination of hardware and software modules in the encoding processor. Optionally, the software modules may be located in a storage medium well known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or a register.The storage medium is located in a storage device; the processor reads information from the storage device and implements the method steps of the above embodiments in combination with the processor hardware.
[0119] Embodiment 2
[0120] Embodiment 2 of the present invention provides a head-mounted display device. As illustrated in [Fig. 16], the head-mounted display device 700 comprises: a memory 710 and a processor 720; the memory 710 being configured to store a computer program and transmit program codes to the processor 720. In other words, the processor 720 can call and execute the computer program from the memory 710 in order to implement the methods described in the embodiments of the present invention. For example, the processor 720 can be configured to execute the processing steps of the method described according to an embodiment of the present invention based on instructions configured in the computer program.
[0121] In some embodiments of the present invention, the computer program may be divided into one or more modules, and the one or more modules are stored in memory 710 and executed by processor 720 to perform the method of an embodiment provided by the present invention. The one or more modules may be a series of computer program instruction segments adapted to perform specific functions, and the instruction segments are defined to describe the execution of the computer program on the head-mounted display device 700.
[0122] As illustrated in [Fig.16], the head-mounted display device further comprises a transceiver 730, which is connected to the processor 720 or the memory 710. The processor 720 may control the transceiver 730 to communicate with other devices, in particular to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 730 may consist of at least two cameras configured to take videos / images of a targeted region.
[0123] It should be noted that the various components in the head-mounted display device 700 are connected via a bus system, the bus system including a power bus, a control bus, and a status signal bus in addition to a data bus.
[0124] The above specific embodiments further illustrate the objects, technical solutions and beneficial effects of the present invention, and it should be noted that the above description only shows specific embodiments of the present invention and is not intended to limit the scope of the present invention. Any modifications, equivalent configurations, improvements and the like may be made, without departing from the scope of the present invention.
Claims
1. Claims A method for achieving a keystroke or touch command with tactile feedback, implemented by a system configured in a wearable extended reality device or extended reality headset; the system outputting position information, which carries a time sequence, of articulation points of a human hand captured in a video stream of each camera of the system through a human hand articulation detection model; characterized in that: a keystroke and touch command are achieved through a trigger fingertip touching virtually assigned function regions on a palm of the human hand, the palm being defined to include both a palm center without fingers, and also the fingers; each of the function regions being a character or number button, a function key or a shortcut key that is capable of being triggered;and a corresponding function region being assigned and fixed to a predefined point marked on a virtual joint line of each pair of two adjacent joints of the palm; said method comprising the following steps:; Step 1: Marking the predefined point on the virtual joint line of each pair of two adjacent joints of the palm, the corresponding function region assigned to each predefined point on the palm being capable of being visually perceived using a pair of smart glasses; setting a width of each of the function regions at W; a corresponding predefined point of each of the function regions being determined as the center point of this function region; setting a left trigger determination point WL and a right trigger determination point WR at positions W / 2 to the left and W / 2 to the right relative to the center point of each of the function regions respectively along a direction parallel to an X axis;determining the position information of each preset point and the left trigger determination point WL and the right trigger determination point WR of the corresponding function region assigned to each preset point based on the articulation points;
2. Step 2: Consider by default one end of a thumb as the trigger fingertip; if the thumb does not have access to areas on the palm, but one end of any finger among the other fingers is intended to touch the palm and the function regions assigned to the palm, the end of said any finger among the other fingers is determined to be the trigger fingertip; the trigger fingertip is identified as P; Step 3: The system acquires a number N of parallax video streams from at least two cameras, N being an integer, and N>2, tracks and determines whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to left and right sides of a corresponding feature region, respectively, in all of the corresponding number N of same-time frames from said number N of video streams, respectively, and, if so, calculates the position information of three target points which are the left trigger determination point WL, the trigger fingertip P and the right trigger determination point WR for each of said number N of same-time frames;then, in each of said number N of images bearing the same time, X-axis values, including WRX of the right trigger determination point WR, PX of the trigger fingertip P, and WLX of the left trigger determination point WL, position information of the three target points in that image is used to calculate a ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all the ratios calculated in all of said number N of images bearing the same time are the same, the trigger fingertip P is determined to have touched the corresponding function region, and then a function corresponding to the corresponding function region is output or triggered.; Method according to claim 1, characterized in that, in step 3, said at least two cameras comprise two cameras which are a left camera and a right camera; a connecting line passing through two central points L and R of the left camera and the right camera respectively being considered as the X axis; assuming that, in a field of view of the left camera, an included angle defined as T0L is formed between the X axis and a connecting line connecting the center point L of the left camera and one of the three target points T; assuming that, in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and one of the three target points T;assuming that a length of a parallax baseline between the two center points L and R of the left camera and the right camera is d, calculating a position (X, Z) of any one of the three target points T in each of said number N of images carrying the same time, based on the following: if the target point T whose position is to be calculated is located between the two center points L and R of the left camera and the right camera: Z=d / [TAN(T0L)-TAN(T0R-jt / 2)], X=Z*TAN(T0L); if the target point T whose position is to be calculated is located on a left side of the center point L of the left camera: Z=d / [TAN(T0R-jr / 2)-TAN(T0L-jr / 2)], X=-Z*TAN(T0L); if the target point T whose position is to be calculated is located on a right side of the central point R of the right camera: Z=d / [COT(T0L)-COT(T0R)],X=Z / Tan(T0L).;
3. A method according to claim 1 or 2, characterized in that each of the function regions has a circular shape; a circle is drawn for each of the function regions with the predefined point which is defined at any position on the joint line between each pair of two adjacent joints of the palm as the center of the circle and the width W of each of the function regions as the diameter.
4. A method according to claim 1 or 2, characterized in that each of the function regions is assigned to the inside of a phalangeal region, between two phalangeal regions, on an outer side of the phalangeal region, or to a certain area of the palm center between a wrist and a certain finger.
5. Method according to claim 1 or 2, characterized in that, in step 3, the system virtually assigns a matrix grid to the same position on the palm center when a number N of flows
6. video is processed, wherein N is an integer, and N>2; the matrix grid comprises a plurality of grid units each having a plurality of sides, and each grid unit is considered a function region; the system tracks and determines whether or not the trigger fingertip P (X, Y) is located in a same function region of the matrix grid across all of said number N of viewer screens and, if so, the trigger fingertip P (X, Y), and the left trigger determination point WL and the right trigger determination point WR on the left and right sides of the relevant function region are determined as the three target points T;then from the X-axis values, including WRX of the right trigger determination point WR, PX of the trigger fingertip P and WLX of the left trigger determination point WL, position information of the three target points is used to calculate a ratio (PX-WRX):(WLX-PX) which is a value difference between PX and WRX relative to a value difference between WLX and PX;if ratios calculated in all corresponding frames carrying a same time in said number N of video streams are identical, the trigger fingertip P is determined to have touched the function region, and then a point or line is drawn corresponding to a position of the trigger fingertip P (X, Y) and, as the trigger fingertip P moves, a series of determinations by the system over time carrying a time sequence will determine points or lines drawn successively in successive locations over time and thus determine that a line is drawn where all the drawn points or lines are determined to be joined together; therefore, tablet control or touch control functions can be implemented in the palm of one hand by using a fingertip of another hand as a trigger fingertip.; A method according to claim 5, characterized in that a connecting point between a little finger and the center of the palm is determined as an upper right corner of the matrix grid; a connecting point between an index finger and the center of the palm is determined as an upper left corner of the matrix grid; and a connecting line
7.
8.
9. between the center of palm and a wrist is determined as a lower edge of the matrix grid. Method according to claim 5, characterized in that the matrix grid is invisible and not displayed on said number N of screens of the viewer. Method according to claim 5, characterized in that each grid unit has a square or rectangular shape. A method for obtaining a keystroke or touch command with tactile feedback, implemented by a system configured in a wearable extended reality device or an extended reality headset; the system outputting position information, carrying a time sequence, of target points captured by videos; characterized in that a keystroke and touch command are obtained by means of a trigger fingertip touching function regions; said method comprising the following steps: Step 1: The system anchors a touch control interface image on each screen at a same position of a same predefined object surface; a plurality of said function regions are assigned to the touch control interface image as viewed from any one of the number N of screens of the viewer; corresponding images from all videos carrying a same time are each determined to have a left trigger determination point WL and a right trigger determination point WR on the left and right sides of a respective corresponding function region respectively along a direction parallel to an X axis of the image from a corresponding video; Step 2: One end of any finger intended to touch the function regions is determined as the trigger fingertip P; Step 3: The system acquires a number N of parallax video streams, where N is an integer, and N>2; tracking and determining whether the trigger fingertip P(X,Y) is located in a same feature region or not in all of the corresponding number N of same-time frames from said number N of video streams, respectively, and, if so, the trigger fingertip P(X,Y) and the trigger determination point left WL and the right trigger determination point WR corresponding to the relevant function region are used as three target points T; X-axis values, including WRX of the right trigger determination point WR, PX of the trigger fingertip T and WLX of the left trigger determination point WL, in the position information of the three target points T are used to calculate a ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all ratios of said number N of images carrying the same time are the same, the trigger fingertip P is determined to have touched the function region, and then a function corresponding to the function region is output or triggered.
10. Method according to claim 9, characterized in that the touch control interface image is an image of a conventional numeric keypad or a conventional keyboard.
11. Method according to claim 9, characterized in that the predefined object surface is any surface of a real object.
12. A method according to claim 9, characterized in that the predefined object surface is a surface of a virtual object and, when the trigger fingertip touches a corresponding function region, feedback in the form of sound, vibration, electric shock or other mechanical feedback is provided to create a sensation of touching a real object.
13. A head-mounted display device (700), comprising at least two cameras configured to take videos or images of a targeted region; the head-mounted display device (700) also comprising a memory (710) and a processor (720); the memory (710) being configured to store a computer program; the processor (720) being configured to execute the computer program to perform the method according to any one of claims 1 to 12.