Method for obtaining virtual touch control with a virtual three-dimensional cursor, storage medium and chip implementing said method
The method allows users to interact with three-dimensional virtual spaces using a bare hand by defining a control aiming point and projecting an interactive control line, addressing the lack of hand-based interaction in ER systems and enabling precise virtual touch control and drawing.
Patent Information
- Application Number
- FR2024012010
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2024-11-04
- Publication Date
- 2025-08-29
AI Technical Summary
Existing extended reality (ER) systems lack the ability for users to interact with three-dimensional virtual spaces using a bare hand to control a cursor, as they typically require remote controls or sensors, and there is no method to calculate a three-dimensional cursor position based on visual images.
A method for achieving virtual touch control using a three-dimensional cursor involves defining a control aiming point with a light ray passing through a joint or fingertip, projecting an interactive control line, and using a trigger region and fingers to activate and click the cursor, with calculations based on camera parallax and trigonometry to determine spatial positions.
Enables touch control operations, writing, and drawing in a three-dimensional virtual space by forming an interactive control line with a bare hand, allowing precise interaction with virtual objects.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for obtaining a virtual touch control with a virtual three-dimensional cursor, storage medium and chip implementing said method Technical field
[0001] The present invention relates to the field of virtual touch control technology, and more particularly relates to a method for achieving virtual touch control with a virtual three-dimensional cursor, a storage medium and a chip implementing said method. The present invention applies to an extended reality (ER) wearable device, or more particularly to an extended reality headset.
[0002] BACKGROUND OF THE INVENTION
[0003] Extended reality (ER) refers to a combined reality and virtuality environment enabling human-machine interaction realized by computer technologies and devices that can be worn by humans, and is a general term that encompasses augmented reality (AR), virtual reality (VR), and mixed reality (MR). With the popularization and development of ER in various industries, various ER glasses have emerged, enabling interaction between a user and a system by inputs through a virtual keyboard and touch control.
[0004] When using a smart terminal with RE glasses, the user sees the world on two screens with two eyes, and the world he sees is different from two-dimensional images seen from a mobile phone, a tablet computer, and a conventional display screen. The world seen through binocular display screens of RE glasses is three-dimensional. The user can move and click a simple cursor on a conventional two-dimensional screen based on (X, Y) coordinates. However, in a three-dimensional space, the user cannot move and click a conventional cursor in a three-dimensional (X, Y, Z) space, especially with the depth dimension.Typically, an RE glasses smart terminal uses a remote control, game controller, mobile phone, or other similar sensors to express a "straight line" like that of a laser projected by a laser pen or a "curve" like that of a cast fishing rod to command a cursor position in three-dimensional space. A closer virtual object can be virtually touched or operated with a finger or gesture. At present, there are no disclosed patents or non-patent literatures in which a bare hand can be used to direct a cursor to a position. remote does not disclose a method of calculating a three-dimensional cursor position based on visual images.
[0005] BRIEF SUMMARY OF THE INVENTION
[0006] The present invention aims to provide a method for achieving virtual touch control with a three-dimensional cursor, as well as a storage medium and a chip implementing said method. The present invention makes it possible to achieve touch control operations, writing or drawing in a three-dimensional virtual space by forming an interactive control line of a defined or indefinite length with a bare hand.
[0007] The subject of the invention is a method for achieving virtual touch control with a three-dimensional cursor, applicable to a system using an extended reality wearable device or an extended reality headset; wherein, in a virtual space, a weighted average position of a certain joint or fingertip or a plurality of joints of a human hand that a light ray projected from a predefined light source passes through before being projected further to a distant position is defined as a control aiming point; the light ray projected from the predefined light source, passing through the control aiming point, and which is projected onto the distant position, being defined as an interactive control line; and the three-dimensional cursor being displayed at a distal end of the interactive control line; the method comprising the following steps:
[0008] Step 1: assigning a trigger region for the control aiming point, and assigning at least one switch finger to activate a projection of the three-dimensional cursor and at least one click finger to touch the trigger region;
[0009] Step 2: When any one of said at least one switching finger touches the trigger region, spatial positions of the control aiming point and the predefined light source are acquired, then the light ray projected from the predefined light source, passing through the control aiming point and which is projected onto the remote position, forms said interactive control line; before said at least one clicking finger has already clicked the trigger region, the interactive control line disappears as soon as said at least one switching finger leaves the trigger region and no longer touches it; when there is the interactive control line, changing a direction to which the control aiming point points can guide a movement of the interactive control line;when the interactive command line and a virtual object or a virtual model of a real object in virtual space intersect, the three-dimensional cursor is displayed at that point of intersection and if, at that time, the click finger also touches; the trigger region, a virtual touch command is performed on the virtual object or virtual model touched by the three-dimensional cursor.
[0010] Advantageously, the spatial positions of the control aiming point and the predefined light source are acquired during the following steps:
[0011] if the predefined light source is positioned within a visible range of a pair of extended reality (ER) smart glasses and is also a certain predefined joint, without yet being said control aiming point, on any of the two hands, then the spatial positions of the control aiming point and the predefined light source are acquired using a same calculation method, as follows: considering a connecting line passing through center points L / R of two left and right cameras of the ER smart glasses as an X-axis, and in a field of view of the left camera, an included angle defined as T0L is formed between the X-axis and a connecting line connecting the center point L of the left camera and a targeted joint point T whose spatial position is yet to be calculated;similarly, in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and a targeted articulation point T whose spatial position is yet to be calculated; a parallax distance between the two center points L and R of the left and right cameras is defined as d, and a position (X, Z) of the targeted articulation point T is calculated as follows: ;
[0012] if the targeted articulation point T is located between the two central points L and R of the left and right cameras:
[0013] Z=d / [TAN(T0L)-TAN(T0R-jt / 2)], X=Z*TAN(T0L);
[0014] if the targeted articulation point T is located on a left side of the central point L of the left camera:
[0015] Z=d / [TAN(T0R-jt / 2)-TAN(T0L-jt / 2)], X=-Z*TAN(T0L);
[0016] if the targeted articulation point T is located on a right side of the central point R of the right camera:
[0017] Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L);
[0018] a reference point for a determination of a Y-axis value is any point on a lower side of video images perceived by the left and right cameras; a pixel value Y of the targeted articulation point T counted upwards from the reference point is the Y-axis value of a spatial position (X, Y, Z) of the targeted articulation point T;
[0019] if the preset light source is not positioned within the visible range of the RE smart glasses, the spatial position of the preset light source is calculated by a method different from the method for calculating the spatial position of the control aiming point as described above; the spatial position of the source of predefined light is calculated as follows: a relative position with respect to a central point (Xocenter, Ycenter) of the RE smart glasses between the left cameras and right is used as a preset light source; if a knuckle or a fingertip of a thumb of a right hand is used as a control aiming point, a preset light source position (Xlightsource, Ylightsource), which is the relative position with respect to the center point of the RE smart glasses, is defined as Xlightsource — Xcenter-!- [>x, and Ylightsource — Ycenter- Py , if a knuckle or a fingertip of a thumb of a left hand is used as a control aiming point, a light source position (Xlightsource, Ylightsource), which is the relative position with respect to the center point of the RE smart glasses, is defined as Xlightsource — Xcenter - Px, and Ylightsource — Ycenter ” Py, where Px and Py are offset values.
[0020] Advantageously, steps for determining whether or not a trigger fingertip P, which is said at least one switching finger or said at least one clicking finger, touches the trigger region comprise:
[0021] the trigger region assigned to the control aiming point is defined with a width W; a left trigger determination point WL and a right trigger determination point WR are defined at positions W / 2 to the left and W / 2 to the right of a center point of the trigger region respectively along a direction parallel to an X axis, in other words, the left trigger determination point WL and the right trigger determination point WR are points corresponding to left and right boundaries of the trigger region, respectively;the system acquires a number N of parallax video streams from at least two cameras of the RE smart glasses, where N is an integer, and N>2, tracks and determines whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to the trigger region in all of the corresponding N number of same-time images from said number N of video streams, respectively, and if so, calculates position information of three target points which are the left trigger determination point WL, the trigger fingertip P, and the right trigger determination point WR for each of said number N of same-time images;then, in each of said number N of images carrying the same time, X-axis values (including WRX of the right trigger determination point WR, PX of the trigger fingertip P and WLX of the left trigger determination point WL) of the position information of the three target points in this image are used to calculate a ratio (PX- ; WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all the ratios calculated in all of said number N of images bearing the same time are identical, the trigger fingertip P is determined to have touched the trigger region.
[0022] Advantageously, the virtual touch control refers to the fact that the fingertip of a thumb is defined as a control aiming point; one of the remaining four fingers is selected as a click finger; and at least one of the still remaining three fingers is defined as a switch finger to activate a projection of the three-dimensional cursor; if a plurality of switch fingers are defined, the plurality of switch fingers are defined to activate different functions.
[0023] Advantageously, the virtual touch control refers to the fact that the fingertip of a thumb is defined as a control aiming point; one of the remaining four fingers is selected as a switching finger to activate a projection of the three-dimensional cursor; and two of the still remaining three fingers are respectively defined as a mouse right button click finger and a mouse left button click finger.
[0024] Advantageously, the interactive control line is a projected light ray of indefinite length, which is a straight light ray or a parabolic light ray.
[0025] Advantageously, the interactive command line is a virtual brush with a predefined length; while displaying the virtual brush, by touching the trigger region with the click finger, a brush tip of the virtual brush draws dots or strokes in the virtual space, thereby achieving a drawing or writing function.
[0026] The invention also relates to a head-mounted display device, comprising at least two cameras configured to take videos and / or images of a targeted region; the head-mounted display device comprising a memory and a processor; the memory being configured to store a computer program; the processor being configured to execute the computer program in order to carry out one of the methods for obtaining virtual touch control with a three-dimensional cursor according to the present invention.
[0027] The invention further relates to a computer-readable storage medium, in which a computer program is stored, the computer program, when executed by a processor, carrying out one of the methods for obtaining virtual touch control with a three-dimensional cursor according to the present invention.
[0028] The invention also relates to a chip for executing instructions, the chip comprising an integrated circuit substrate encapsulated therein, and the substrate of integrated circuits being configured to perform one of the methods for obtaining virtual touch control with a three-dimensional cursor according to the present invention.
[0029] According to the present invention, a weighted average position of a certain joint or fingertip or a plurality of joints that a light ray projected from a predefined light source passes through before being projected to a distant position is defined as a control aiming point of the three-dimensional cursor; a position of the predefined light source is predefined on a hand joint or any other body part whose position is calculated relative to a center point of the RE smart glasses;the light ray of a definite or indefinite length is projected from the predefined light source, passing through the control aiming point and finally projected onto the remote position, so that a remote three-dimensional cursor or a virtual brush controllable by a bare hand can be formed, thereby achieving virtual touch control or writing or drawing in a three-dimensional virtual space. The present invention has the following technical effects: ;
[0030] 1. The RE smart glasses of the present invention calculate information of three-dimensional position (X, Y, Z) of a finger joint based on images from a plurality of cameras and a parallax or physical distance between the cameras, thereby quickly obtaining the three-dimensional position information of a control aiming point and a predefined light source, thereby projecting an interactive control line from the predefined light source through the control aiming point.
[0031] 2. A trigger region is assigned to the control aiming point, and to the at least one switching finger to activate a projection of the three-dimensional cursor as well as at least one clicking finger to touch the trigger region are also affected. When the switching finger touches the trigger region, the spatial positions of the control aiming point and the predefined light source are acquired, then a light ray projected from the predefined light source, passing through the control aiming point, forms the interactive control line; before the clicking finger has already clicked the trigger region, the interactive control line disappears as soon as the switching finger leaves the trigger region and no longer touches it.When the interactive command line exists, changing a direction to which the command aiming point points guides a movement of the interactive command line; when the interactive command line and a virtual object or a virtual model of a real object in a virtual space intersect, a three-dimensional cursor is displayed at that intersection point, and if, at that time, the click finger also touches the trigger region, a virtual touch command is . performed on the virtual object or virtual model touched by the three-dimensional cursor. According to the above technical solutions, the present invention can achieve touch control operations, writing or drawing in a three-dimensional virtual space by forming an interactive control line with a defined or indefinite length through a bare hand. Brief description of the drawings
[0032] [Fig.l] represents 21 recognizable points of articulation on a human hand and their names given on the official Mediapipe website;
[0033] [Fig.2] is a schematic diagram of calculating spatial positions of targeted articulation points via a left camera of a pair of RE smart glasses according to the present invention;
[0034] [Fig.3] is a schematic diagram of calculating spatial positions of targeted articulation points via a right camera of a pair of RE smart glasses according to the present invention;
[0035] [Fig.4] is a schematic diagram of complete three-dimensional information of two articulation points of a single finger through the use of corresponding Y-axis values according to the present invention;
[0036] [Fig.5] is a schematic diagram of two alternative interactive control lines formed by projecting light rays from predefined light sources whose relative positions are calculated with respect to a central point of the RE smart glasses and where the light rays are projected to distant positions by passing through two thumb fingertips respectively on their way according to the present invention;
[0037] [Fig.6] is a schematic diagram of a ring finger as a switching finger when implementing left and right mouse button control according to the present invention;
[0038] [Fig.7] is a schematic diagram of an index finger as a right button click finger when implementing left and right mouse button control according to the present invention;
[0039] [Fig.8] is a schematic diagram of a middle finger as a left button clicking finger when implementing a left and right mouse button control according to the present invention;
[0040] [Fig.9] is a schematic diagram of a relative positioning of the trigger fingertip and two trigger determination points in combined left and right images of the RE smart glasses when the trigger fingertip does not actually touch the trigger region according to the present invention;
[0041] [Fig. 10] is a schematic diagram of a relative positioning of the trigger fingertip and two trigger determination points in combined left and right images of the RE smart glasses when the trigger fingertip touches the trigger region according to the present invention;
[0042] [Fig. 11] illustrates four separate images showing relative positioning of the trigger fingertip and the two trigger determination points, the upper two images showing images from the left camera and the right camera, respectively, when the trigger fingertip is not touching a trigger region, and the lower two images showing images from the left camera and the right camera, respectively, when the trigger fingertip is touching a trigger region; and
[0043] [Fig. 12] is a structural block diagram of a head-mounted display device according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] The technical solutions in the embodiments of the present invention will be described clearly and in detail below, with reference to the accompanying drawings of the embodiments of the present invention. It is obvious that the described embodiments illustrate only some of the embodiments of the present invention, but not all. Based on the embodiments of the present invention, all other embodiments obtainable by the person skilled in the art without any inventive effort are within the scope of the present invention.
[0045] Furthermore, the terms "include" and "comprising," and any variations thereof, are intended to be non-exclusive inclusion. For example, a process, method, system, product, or server comprising a series of steps or units is not necessarily limited to the steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are known for such a process, method, product, or device.
[0046] In embodiments of the present invention, the use of the terms "by way of example" or "for example" is intended to present relevant concepts in a specific manner.
[0047] The principles of the technical solutions of the present invention are as follows:
[0048] (1) Recognition model used to acquire position information on a palm: Currently commercially available open source software for a pre-trained human hand joint detection model that can acquire two-dimensional positions of human hand joints can be used. The present invention uses Mediapipe as an illustrative example. As an open project of Google®, Mediapipe is a library of tools for learning which mainly concerns visual algorithms and integrates a large number of models relating to facial detection, facial keypoints, gesture recognition, head segmentation and posture recognition. As shown in [Fig.l], position information of 21 articulation points (also called keypoints) of a human hand in a video carrying a time sequence can be output. In general, a human hand articulation detection model outputs articulation position information in the form of pixels (X, Y) which are coordinates of an X axis and a Y axis of the video. The present invention can also use a self-trained human hand articulation detection model.The present invention also includes training and recognition using an artificial intelligence chip such as a graphics processing unit (GPU) or a neural network processing unit (NPU) by label convolutional KNN or RNN or by a Reinforced Transformer Plus model or any reinforcement pretraining methods.
[0049] (2) Calculation of the spatial position of an articulation point:
[0050] As illustrated in [Fig.2], considering a connecting line passing through center points L / R of two left and right cameras of a pair of RE smart glasses as an X-axis, and in a field of view of the left camera, an included angle defined as T0L is formed between the X-axis and a connecting line connecting the center point L of the left camera and a targeted articulation point T whose spatial position is yet to be calculated. Similarly, as illustrated in [Fig.3], in a field of view of the right camera, an included angle defined as T0R is formed between the X-axis and a connecting line connecting the center point R of the right camera and a targeted articulation point T whose spatial position is yet to be calculated.
[0051] A parallax distance between the two center points L and R of the left and right cameras is defined as d, and a position (X, Z) of the targeted articulation point T is calculated as follows:
[0052] If the targeted articulation point T is located between the two central points L and R of the left and right cameras:
[0053] Z=d / [TAN(T0L)-TAN(T0R-jt / 2)], X=Z*TAN(T0L);
[0054] If the targeted articulation point T is located on a left side of the central point L of the left camera:
[0055] Z=d / [TAN(T0R-jt / 2)-TAN(T0L-jt / 2)], X=-Z*TAN(T0L);
[0056] If the targeted articulation point T is located on a right side of the central point R of the right camera:
[0057] Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L).
[0058] The above examples are calculated using TAN (tangent) and COT (cotangent), but any other trigonometric calculation method can be used within the scope of the present invention.
[0059] Since the X-axis is defined as a straight connecting line passing through L / R center points of two left and right cameras of the RE smart glasses, parallax exists only on the X-axis. Therefore, there is no parallax on the Y-axis. In other words, the Y-axis value of the images perceived by the left eye and the right eye must be the same. A reference point for a determination of the Y-axis value can be defined as a lowest point or other definable positions of the video images perceived by the left and right cameras. Therefore, a number of Y pixels (or any other distance unit converted thereto) counted upward from the reference point is the Y-axis value. By adding this Y-axis value to the target articulation point position (X, Z), a complete target articulation point position (X, Y, Z) is obtained, as shown in [Fig.4].
[0060] (3) Acquisition of a spatial position of a predefined light source:
[0061] A light ray is required to form a three-dimensional cursor in a virtual space, and an emission of the light ray requires an emission source or a predefined light source. A weighted average position of a predefined joint or a certain predefined fingertip or a plurality of joints that a light ray projected from the light source passes through before being projected further to a distant position is defined as a control aiming point. An emission direction of the light ray is defined by the light ray emitted from the predefined light source passing through the control aiming point; such a light ray emitted from the light source is defined as an interactive control line; a "shadow" or a three-dimensional cursor is displayed at a distal end of the interactive control line projected onto a surface of a certain virtual object.
[0062] If the predefined light source is positioned within a visible range of the RE smart glasses, and is also a certain predefined joint (but not a control aiming point) on both hands, then a spatial position of the predefined light source can be acquired using the calculation method for the spatial position of the joint point as described in the aforementioned item (2).
[0063] If the preset light source is not positioned within the visible range of the RE smart glasses, the light source is generally positioned on the RE smart glasses or at a certain position relative to the RE smart glasses (an offset position). In general, if the light source is placed at a position in the middle of the RE smart glasses, between the left and right cameras, the light ray is projected from the light source, in passing through a control aiming point, which is usually a fingertip of a thumb on its way, and then finally is projected onto a distant object to display a three-dimensional cursor, but unfortunately, the three-dimensional cursor is always visually blocked by the thumb and thus the user cannot see the three-dimensional cursor through the RE smart glasses. Alternatively, prior art patents and non-patent literatures always mention the use of the shoulder or crotch as the position of the light source; however, the prior art patents and non-patent literatures do not indicate how to calculate the three-dimensional spatial position of the shoulder or crotch.In the present invention, a relative position (an offset position) with respect to a center point (Xcenter, Ycenter) of the RE smart glasses between the left and right cameras is used as a light source.If the fingertip of the thumb of the right hand is used as the control aiming point, the position of the light source (X light source, Y light source) which is the relative position (offset position) with respect to the center point of the RE smart glasses is defined by X light source = X center + [3x, and Y light source = Y center - [3y; if the fingertip of the thumb of the left hand is used as the control aiming point, the position of the light source (X light source, Y light source) which is the relative position (offset position) with respect to the center point of the RE smart glasses is defined by X light source = X center - |3x, and Y light source = Y center - Py, where offset values Px and Py can be preset as desired, for example, can be 20 cm and 30 cm. As shown in [Fig.5], in this case, the relative position of the light source is not in the middle of the RE smart glasses, but beside and below the RE smart glasses; therefore, the three-dimensional cursor projected from the light source onto a certain virtual object passing through the thumb fingertip which is the control aiming point on its way will not be blocked by the user's hand and can be clearly seen.
[0064] (4) Determining whether a trigger fingertip P touches or not a trigger region:
[0065] The trigger region is defined with a width W; a left trigger determination point WL and a right trigger determination point WR are defined at positions W / 2 to the left and W / 2 to the right of a center point of the trigger region respectively along a direction parallel to an X axis, in other words, the left trigger determination point WL and the right trigger determination point WR are points corresponding to left and right boundaries of the trigger region, respectively. The system acquires a number N of video streams with parallax from at least two cameras, where N is an integer, and N>2, tracks and determines whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to the trigger region in all of the corresponding N number of images carrying a same time from said N number of video streams, respectively, and, if so, calculates position information of three target points which are the left trigger determination point WL, the trigger fingertip P and the right trigger determination point WR for each of said N number of images carrying the same time;then, in each of said number N of images carrying the same time, X-axis values (WRX, PX and WLX) of the position information of the three target points in that image are used to calculate a ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX;as shown in Figures 9 to 11, only when all the calculated ratios in all of said number N of images bearing the same time are identical, it is determined that the trigger fingertip has touched the trigger region, which is the thumb fingertip. [Fig.4] shows alternative positions of the light source placed on hinge points within the visible range of the RE smart glasses. The function of projecting the shadow or cursor can be achieved as long as both the light source and the control aiming point are within the visible range of the cameras of the RE smart glasses. ;
[0066] (5) Method for obtaining virtual touch control with the cursor three-dimensional:
[0067] A trigger region, at least one click finger, and at least one switch finger are assigned to the control aiming point. In the embodiment, the fingertip of a thumb is defined as the control aiming point, an index finger is defined as the at least one click finger, and at least one of the other three fingers, which are the middle finger, the ring finger, and the little finger, is defined as a switch finger for activating the projection of the three-dimensional cursor.When the switching finger touches the trigger region, the spatial positions of the control aiming point and the preset light source are acquired, and then a light ray projected from the preset light source is projected toward the control aiming point as a rectilinearly projected light ray or a non-rectilinearly projected light ray, and the rectilinearly projected light ray or the non-rectilinearly projected light ray projects further outward from the control aiming point, resembling a laser projected by a laser pen (straight line) or a cast fishing rod (parabola), and this light ray projected from the light source. predefined is said interactive command line; before the click finger has already clicked the trigger region, the interactive command line disappears as soon as the switching finger leaves the trigger region and no longer touches it. When there is the interactive command line, changing a direction to which the control aiming point points can guide a movement of the interactive command line; when the interactive command line and a virtual object (or a virtual model of a real object) in a virtual space intersect, a three-dimensional cursor is displayed at that intersection point and if, at that time, the click finger also touches the trigger region, various mouse operations such as clicking, dragging, selecting and drawing are performed on the virtual object touched by the three-dimensional cursor.If left and right mouse button operations are desired, two different click fingers can be defined; for example, as shown in Figures 6 to 8, the thumb fingertip is defined as the control aiming point, the index finger is defined as the right button click finger, the middle finger is defined as the left button click finger, and the ring finger is defined as the switching finger to activate the three-dimensional cursor. When the ring finger touches the thumb fingertip to display the interactive command line, if the index finger also touches the trigger region, the right mouse button click action is performed, and if instead the middle finger also touches the trigger region, the left mouse button click action is performed.
[0068] (6) Obtaining virtual command such as virtual drawing and writing with a three-dimensional brush:
[0069] The interactive control line is a virtual brush with a preset length, and the distal end of the interactive control line is a pen tip position. By touching the trigger region with the click finger, the pen tip of the virtual brush can draw dots or strokes in a virtual space, so that writing or drawing in a virtual space is realized to simulate real drawing and writing.
[0070] Embodiment 1
[0071] Embodiment 1 of the present invention relates to a method for achieving virtual touch control with a three-dimensional cursor, applicable to a system using an extended reality (ER) wearable device or an extended reality headset; wherein, in a virtual space, a weighted average position of a certain predefined joint or a certain predefined fingertip or a plurality of joints of a human hand that a light ray projected from a predefined light source passes through before being projected further to a distant position is defined as a control aiming point; the light ray projected from the predefined light source, passing through the aiming point control aim, and which is projected onto the remote position being defined as an interactive control line, and the three-dimensional cursor being displayed at a distal end of the interactive control line; the method comprising the following steps:
[0072] Step 1: assigning a trigger region for the control aiming point, and assigning at least one switch finger for activating a projection of the three-dimensional cursor and at least one click finger for touching the trigger region; in this embodiment, a fingertip of a thumb is defined as the control aiming point; an index finger is defined as the at least one click finger; and one of the middle finger, the ring finger, and the little finger is defined as the at least one switch finger for activating a projection of the three-dimensional cursor, or the middle finger, the ring finger, and the little finger are respectively defined as different switch fingers, said different switch fingers being defined to activate different functions;
[0073] Step 2: When any one of said at least one switching finger touches the trigger region, spatial positions of the control aiming point and the predefined light source are acquired, then the light ray projected from the predefined light source, passing through the control aiming point and which is projected onto the remote position, forms said interactive control line; before said at least one clicking finger has already clicked the trigger region, the interactive control line disappears as soon as said at least one switching finger leaves the trigger region and no longer touches it; when the interactive control line exists, the change of a direction to which the control aiming point points guides a movement of the interactive control line;when the interactive command line and a virtual object or a virtual model of a real object in a virtual space intersect, the three-dimensional cursor is displayed at that intersection point and if, at that time, the click finger also touches the trigger region, a virtual touch command is performed on the virtual object or virtual model touched by the three-dimensional cursor. ;
[0074] A specific calculation for acquiring the spatial positions of the control aiming point and the predefined light source comprises the following steps:
[0075] If the predefined light source is positioned within a visible range of a pair of RE smart glasses, and is also a certain predefined joint, without yet being said control aiming point, on two hands, then the spatial positions of the control aiming point and the predefined light source are acquired using a same calculation method, as follows: considering a connecting line passing through L / R center points of both left and right cameras of the RE smart glasses as an X axis, and in a field of view
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084] of the left camera, an included angle defined as T0L is formed between the X axis and a connecting line connecting the center point L of the left camera and a targeted articulation point T whose spatial position has yet to be calculated; similarly, as illustrated in [Fig.3], in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and a targeted articulation point T whose spatial position has yet to be calculated; a parallax distance between the two center points L and R of the left and right cameras is defined as d, and a position (X, Z) of the targeted articulation point T is calculated as follows: if the targeted articulation point T is located between the two central points L and R of the left and right cameras: Z=d / [TAN(T0L)-TAN(T0R-ir / 2)], X=Z*TAN(T0L); if the targeted articulation point T is located on a left side of the center point L of the left camera: Z=d / [TAN(T0R-Jt / 2)-TAN(T0L-Jt / 2)], X=-Z*TAN(T0L); if the targeted articulation point T is located on a right side of the central point R of the right camera: Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L); a reference point for a determination of a Y-axis value is any point on a lower side of video frames perceived by the left and right cameras; a pixel value Y of the targeted articulation point T counted upwards from the reference point is the Y-axis value of a spatial position (X, Y, Z) of the targeted articulation point T; If the preset light source is not positioned within the visible range of the RE smart glasses, the spatial position of the preset light source is calculated by a method different from the method for calculating the spatial position of the control aiming point as described above: a relative position with respect to a center point (Xcenter, Ycenter) of the RE smart glasses between the left and right cameras is used as the preset light source; if a knuckle or the fingertip of the thumb of a right hand is used as the control aiming point, the position of the preset light source (Xlightsource, Ylightsource), which is the relative position with respect to the center point of the RE smart glasses, is defined by Xsource light -^-center' light source Ycenter |3y , SI UUC Hit iculat ÎOU or the fingertip of the thumb of a left hand is used as the control aiming point, the position of the light source (Xlightsource, Ylightsource), which is the relative position to the center point of the RE smart glasses, is defined by X light source -center light source 1 center' [3y, where [3x and [3y are offset values;
[0085] Steps for determining whether a trigger fingertip P, which is said at least one switching finger or said at least one clicking finger, touches the trigger region or not, comprise:
[0086] the trigger region assigned to the control aiming point is defined with a width W; a left trigger determination point WL and a right trigger determination point WR are defined at positions W / 2 to the left and W / 2 to the right of a center point of the trigger region respectively along a direction parallel to an X axis, in other words, the left trigger determination point WL and the right trigger determination point WR are points corresponding to left and right boundaries of the trigger region, respectively;the system acquires a number N of parallax video streams from at least two cameras of the RE smart glasses, where N is an integer, and N>2, tracks and determines whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to the trigger region in all of the corresponding N number of same-time images from said number N of video streams, respectively, and if so, calculates position information of three target points which are the left trigger determination point WL, the trigger fingertip P, and the right trigger determination point WR for each of said number N of same-time images;then, in each of said number N of images bearing the same time, X-axis values (WRX, PX and WLX) of the position information of the three target points in that image are used to calculate a ratio (PX-WRX):(WLX-PX), which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all the ratios calculated in all of said number N of images bearing the same time are the same is it determined that the trigger fingertip P has touched the trigger region. ;
[0087] Virtual touch control refers to the fact that the fingertip of the thumb is set as a control aiming point; one of the remaining four fingers, for example, the index finger, is selected as a click finger; and at least one of the remaining three fingers is set as a switch finger to activate a projection of the three-dimensional cursor; if a plurality of switch fingers are set, the plurality of switch fingers are set to activate different functions. For example, the middle finger corresponds to an activation of a three-dimensional cursor, and the ring finger corresponds to an activation of a virtual brush.
[0088] Virtual touch control refers to the fact that the fingertip of the thumb is set as the control aiming point; one of the remaining four fingers, for example the ring finger, is selected as the switching finger to activate a projection of the three-dimensional cursor; and two of the still remaining three fingers, for example the index finger and the middle finger, are respectively set as the right mouse button click finger and the left mouse button click finger.
[0089] The interactive command line is a projected light beam of indefinite length, which is a straight light beam resembling a laser projected from a laser pen or a parabolic light beam resembling a cast fishing rod.
[0090] The interactive control line is a virtual brush with a predefined length, and a distal end of the interactive control line is a pen tip position. While displaying the virtual brush, by also touching the trigger region with the click finger, the pen tip draws dots or strokes in a virtual space, thereby achieving a drawing or writing function.
[0091] It should also be understood by those skilled in the art that portions and algorithmic steps of various examples described with reference to the embodiments disclosed herein may be implemented using electronic hardware, computer software, or a combination thereof. In order to clearly illustrate the interchangeability of implementation using hardware and software, the components and steps of various examples have been described generally in terms of operational characteristics in the above description. Whether these characteristics are performed using hardware or software depends on the constraints and conditions of the technical solutions proposed in the context of a particular use or design.One skilled in the art may implement the described features in various ways for each particular example of use, and these various ways of implementation should not be considered as departing from the scope of the present invention.
[0092] More specifically, the steps of the method disclosed in embodiments of the present invention may be executed by a processor using hardware integrated logic circuits and / or software instructions. The steps of the method described with reference to embodiments of the present invention may be directly executed by being performed by a hardware encoding processor or performed by a combination of hardware and software modules in the encoding processor. Optionally, the software modules may be located in a storage medium well known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or a register. The storage medium is located in a storage device; the processor reads information from the storage device; storage and implements the method steps of the above embodiments in combination with processor hardware.
[0093] Embodiment 2
[0094] Embodiment 2 of the present invention relates to a head-mounted display device. As illustrated in [Fig. 12], the head-mounted display device 700 comprises: a memory 710 and a processor 720; the memory 710 being configured to store a computer program and transmit program codes to the processor 720. In other words, the processor 720 can call and execute the computer program from the memory 710 in order to implement the methods described in the embodiments of the present invention. For example, the processor 720 can be configured to execute the processing steps in the method described according to an embodiment of the present invention based on instructions configured in the computer program.
[0095] In some embodiments of the present invention, the computer program may be divided into one or more modules, and the one or more modules are stored in memory 710 and executed by processor 720 to implement the method of an embodiment provided by the present invention. The one or more modules may be a series of computer program instruction segments adapted to perform specific functions, and the instruction segments are defined to describe the execution of the computer program on the head-mounted display device 700.
[0096] As illustrated in [Fig.12], the head-mounted display device further comprises a transceiver 730, which is connected to the processor 720 or the memory 710. The processor 720 may control the transceiver 730 to communicate with other devices, in particular to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 730 may comprise at least two cameras configured to take videos / images of a targeted region.
[0097] It should be noted that the various components in the head-mounted display device 700 are connected by a bus system, the bus system including a power bus, a control bus, and a status signal bus, in addition to a data bus.
[0098] Embodiment 3
[0099] Embodiment 3 of the present invention further relates to a computer storage medium in which a computer program is stored, the computer program, when executed by a computer, enabling the computer to execute the processing steps described in the method according to embodiment 1 above.
[0100] Embodiment 4
[0101] Embodiment 4 of the present invention further relates to a chip for executing instructions, the chip comprising an integrated circuit substrate encapsulated therein, and the integrated circuit substrate being configured to perform the processing steps described in the method according to embodiment 1 above.
[0102] The above specific embodiments further illustrate the objects, technical solutions and beneficial effects of the present invention, and it should be noted that the above description only shows specific embodiments of the present invention and is not intended to limit the scope of the present invention. Any modifications, equivalent configurations, improvements and the like may be made without departing from the scope of the present invention.
Claims
1. Claims A method for achieving virtual touch control with a three-dimensional cursor, applicable to a system using an extended reality wearable device or an extended reality headset; characterized in that, in a virtual space, a weighted average position of a certain joint or fingertip or a plurality of joints of a human hand that a light ray projected from a predefined light source passes through before being projected further to a distant position is defined as a control aiming point; the light ray projected from the predefined light source, passing through the control aiming point, and which is projected onto the distant position, is defined as an interactive control line, and the three-dimensional cursor is displayed at a distal end of the interactive control line; the method comprising the following steps: Step 1: Assign a trigger region for the control aiming point, and assign at least one switch finger to activate a projection of the three-dimensional cursor and at least one click finger to touch the trigger region; Step 2: When any one of the at least one switching finger touches the trigger region, spatial positions of the control aiming point and the preset light source are acquired, then the light ray projected from the preset light source, passing through the control aiming point and being projected onto the remote position, forms the interactive control line; before the at least one clicking finger has already clicked the trigger region, the interactive control line disappears as soon as the at least one switching finger leaves the trigger region and no longer touches it; when there is the interactive control line, changing a direction to which the control aiming point points can guide a movement of the interactive control line;when the interactive command line and a virtual object or a virtual model of a real object in virtual space intersect, the three-dimensional cursor is displayed at that point of intersection and if, at that time, the click finger also touches;
2. the trigger region, a virtual touch command is performed on the virtual object or virtual model touched by the three-dimensional cursor. A method according to claim 1, characterized in that the spatial positions of the control aiming point and the predefined light source are acquired in the following steps: if the predefined light source is positioned in a visible range of a pair of extended reality, RE, smart glasses, and is also a certain predefined joint, without yet being said control aiming point, on any of the two hands, then the spatial positions of the control aiming point and the predefined light source are acquired using a same calculation method, as follows: considering a connecting line passing through L / R center points of two left and right cameras of the RE smart glasses as an X-axis, and in a field of view of the left camera,an included angle defined as T0L is formed between the X axis and a connecting line connecting the center point L of the left camera and a targeted articulation point T whose spatial position has yet to be calculated; similarly, in a field of view of the right camera, an included angle defined as T0R is formed between the X axis and a connecting line connecting the center point R of the right camera and the targeted articulation point whose spatial position has yet to be calculated; a parallax distance between the two center points L and R of the left and right cameras is defined as d, and a position (X, Z) of the targeted articulation point T is calculated as follows: if the targeted articulation point T is located between the two center points L and R of the left and right cameras:, Z=d / [TAN(T0L)-TAN(T0R-Jt / 2)], X=Z*TAN(T0L); if the targeted articulation point T is located on a left side of the center point L of the left camera: Z=d / [TAN(T0R-Jt / 2)-TAN(T0L-Jt / 2)], X=-Z*TAN(T0L); if the targeted articulation point T is located on a right side of the central point R of the right camera: Z=d / [COT(T0L)-COT(T0R)], X=Z / Tan(T0L); a reference point for a determination of a Y-axis value is any point on a lower side of video frames perceived by the left and right cameras; a Y pixel value of the
3. targeted articulation point T counted upwards from the reference point is the Y-axis value of a spatial position (X, Y, Z) of the targeted articulation point T; if the preset light source is not positioned within the visible range of the RE smart glasses, the spatial position of the preset light source is calculated by a method different from the method for calculating the spatial position of the control aiming point as described above; the spatial position of the preset light source is calculated as follows: a relative position with respect to a center point (Xcenter, Ycenter) of the RE smart glasses between the left and right cameras is used as the preset light source;if a knuckle or a fingertip of a thumb of a right hand is used as a control aiming point, a predefined light source position (Xlight source, Ylight source), which is the relative position with respect to the center point of the RE smart glasses, is defined as Xlight source Xcenter + P^, and Ylight source Ycenter- Py , If a knuckle OR a fingertip of a thumb of a left hand is used as a control aiming point, a light source position (Xlight source, Ylight source), which is the relative position with respect to the center point of the RE smart glasses, is defined as Xlight source Xcenter ” Px, and Ylight source Ycenter- Py, OR Px and Py are offset values.; A method according to claim 1, characterized in that steps for determining whether a trigger fingertip P, which is said at least one switching finger or said at least one clicking finger, touches the trigger region or not include: the trigger region assigned to the control aiming point is defined with a width W; a left trigger determining point WL and a right trigger determining point WR are defined at positions W / 2 to the left and W / 2 to the right of a center point of the trigger region respectively along a direction parallel to an X axis, in other words, the left trigger determining point WL and the right trigger determining point WR are points corresponding to left and right boundaries of the trigger region, respectively; the system acquires a number N of parallax video streams from at least two cameras of the RE smart glasses, where N is an integer, and N>2, tracks and determines whether or not the trigger fingertip P is located between the left trigger determination point WL and the right trigger determination point WR corresponding to the trigger region in all of the corresponding N number of same-time images from said number N of video streams, respectively, and if so, calculates position information of three target points which are the left trigger determination point WL, the trigger fingertip P, and the right trigger determination point WR for each of said number N of same-time images;then, in each of said number N of images bearing the same time, X-axis values, including WRX of the right trigger determination point WR, PX of the trigger fingertip P and WLX of the left trigger determination point WL, position information of the three target points in that image is used to calculate a ratio (PX-WRX):(WLX-PX) which is a difference in value between PX and WRX relative to a difference in value between WLX and PX; only when all the ratios calculated in all of said number N of images bearing the same time are the same, the trigger fingertip P is determined to have touched the trigger region.;
4. A method according to claim 1, characterized in that the virtual touch control refers to the fact that the fingertip of a thumb is set as a control aiming point; one of the remaining four fingers is selected as a click finger; and at least one of the still remaining three fingers is set as a switch finger to activate a projection of the three-dimensional cursor; if a plurality of switch fingers are set, the plurality of switch fingers are set to activate different functions.
5. A method according to claim 1, characterized in that the virtual touch control refers to the fact that the fingertip of a thumb is set as a control aiming point; one of the remaining four fingers is selected as a switching finger to activate a projection of the three-dimensional cursor; and two of the still remaining three fingers are respectively set as a finger right mouse button click and left mouse button click finger.
6. A method according to claim 1, characterized in that the interactive control line is a projected light ray of indefinite length, which is a straight light ray or a parabolic light ray.
7. A method according to claim 1, characterized in that the interactive command line is a virtual brush with a predefined length; while displaying the virtual brush, by touching the trigger region with the click finger, a brush tip of the virtual brush draws dots or strokes in the virtual space, thereby achieving a drawing or writing function.
8. A head-mounted display device (700), comprising at least two cameras configured to take videos and / or images of a targeted region; characterized in that the head-mounted display device (700) comprises a memory (710) and a processor (720); the memory (710) being configured to store a computer program; the processor (720) being configured to execute the computer program to perform the method according to any one of claims 1 to 7.
9. A computer-readable storage medium in which a computer program is stored, characterized in that the computer program, when executed by a processor (720), performs the method of any one of claims 1 to 7.
10. A chip for executing instructions, characterized in that the chip comprises an integrated circuit substrate encapsulated therein, and the integrated circuit substrate is configured to perform the method according to any one of claims 1 to 7.