Human-computer interaction method and device, augmented reality equipment and storage medium
By identifying the coordinates of the fingertips of the user's hand and fitting the rays, detecting the collision intersection points with the screen interface, the problems of low interaction accuracy and delay in the prior art are solved, and efficient and accurate human-computer interaction is achieved.
Patent Information
- Application Number
- CN202411975315.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
The existing human-computer interaction technology that extends reality equipment has problems such as low recognition accuracy, fingertip positioning delay and error, resulting in low interaction efficiency.
By identifying the fingertip coordinates of the user's hand, fitting the target ray, detecting the collision intersection point between the ray and the screen interface, mapping the intersection coordinates to the screen coordinates, and identifying the target controls pointed to by the screen coordinates, to achieve high-precision interactive operation.
It improves the accuracy and efficiency of interaction, reduces interaction delay, enhances immersion and user experience, making interaction in an augmented reality environment more convenient and efficient.
Smart Images

Figure CN119987541A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to a human-computer interaction method, device, extended reality device and storage medium. Background Art
[0002] In the human-computer interaction technology of extended reality devices (such as AR, VR, MR and other devices), the commonly used interaction methods are controllers, touch screens or gesture recognition. However, the interaction technology based on gesture recognition generally has the problem of low recognition accuracy, especially in the positioning of fingertips and the interaction with the virtual environment. There are delays and errors.
[0003] Therefore, how to improve the interaction efficiency of extended reality devices has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present application provides a human-computer interaction method, apparatus, device and storage medium, aiming to improve the interaction efficiency of an extended reality device.
[0005] In a first aspect, the present application provides a human-computer interaction method, the method comprising:
[0006] Identify the coordinates of the fingertips of the user's hand;
[0007] Fitting a target ray based on a ray direction from the extended reality device to the user's hand and the fingertip coordinates;
[0008] Detecting the collision intersection point between the target ray and the screen interface to obtain the coordinates of the collision intersection point;
[0009] Mapping the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection point;
[0010] Identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operation on the target control.
[0011] In a second aspect, the present application further provides a human-computer interaction device, the human-computer interaction device comprising:
[0012] A fingertip coordinate recognition module, used to recognize the fingertip coordinates of the user's hand;
[0013] A ray fitting module, used for fitting a target ray based on a ray direction from an extended reality device to the user's hand and the fingertip coordinates;
[0014] A collision intersection detection module is used to detect the collision intersection between the target ray and the screen interface to obtain the coordinates of the collision intersection;
[0015] A coordinate mapping module, used to map the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection point;
[0016] The target control identification module is used to identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operations on the target control.
[0017] In a third aspect, the present application also provides an extended reality device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the human-computer interaction method as described above are implemented.
[0018] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the human-computer interaction method as described above are implemented.
[0019] The present application provides a human-computer interaction method, apparatus, extended reality device and storage medium. The method of the present application includes identifying the fingertip coordinates of a user's hand; fitting a target ray based on the ray direction from the extended reality device to the user's hand and the fingertip coordinates; detecting the collision intersection of the target ray and the screen interface to obtain the collision intersection coordinates; mapping the collision intersection coordinates to the plane coordinate system where the screen interface is located to obtain the screen coordinates corresponding to the collision intersection; identifying the target control pointed to by the screen coordinates in the screen interface to facilitate interactive operations on the target control. Through the above method, the present application accurately identifies the fingertip coordinates of the user's hand through a fingertip tracking algorithm, provides accurate input points for subsequent interactions, and ensures the accuracy of the interaction; uses the ray direction from the user's eyes to the fingertips and the fingertip coordinates to fit the target ray to determine the user's interaction object, making the interaction more intuitive and natural; determines the collision intersection of the target ray and the screen interface through a collision detection algorithm, and converts the collision intersection coordinates into screen coordinates through a coordinate mapping algorithm to determine the specific location of the user's intended operation, and then identifies the target control of the user's intended interaction according to the screen coordinates, achieving accurate correspondence between user gestures and controls, thereby improving the accuracy and efficiency of the interaction, and making the interaction in the augmented reality environment more convenient and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of a flow chart of a first embodiment of a human-computer interaction method provided in an embodiment of the present application;
[0022] Figure 2 A schematic diagram of a collision detection structure provided in an embodiment of the present application;
[0023] Figure 3 A schematic diagram of a flow chart of a second embodiment of a human-computer interaction method provided in an embodiment of the present application;
[0024] Figure 4 It is a structural schematic diagram of a first embodiment of a human-computer interaction device provided by the present application;
[0025] Figure 5 It is a schematic block diagram of the structure of an extended reality device provided in an embodiment of the present application.
[0026] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0028] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0029] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0030] Please refer to Figure 1 , Figure 1 A flowchart of a first embodiment of a human-computer interaction method provided in an embodiment of the present application.
[0031] like Figure 1 As shown, the human-computer interaction method includes steps S101 to S104.
[0032] S101, identifying the coordinates of the fingertips of the user's hand;
[0033] In one embodiment, when a user wears an extended reality device (such as AR glasses, VR glasses, etc.), an image or video stream of the user's hand is collected through a camera configured in the extended reality device, and the fingertip coordinates of the user's hand are identified through a fingertip tracking algorithm.
[0034] Among them, extended reality devices (Extended Reality, referred to as XR) refer to wearable devices that combine reality and virtuality through computer technology to create a virtual environment for human-computer interaction, including augmented reality (AR) devices, virtual reality (VR) devices, and mixed reality (MR) devices.
[0035] In one embodiment, the fingertip coordinates may be the coordinates of the fingertip of a specified finger, such as specifying the fingertip of an index finger, and the fingertip tracking algorithm identifies the fingertip coordinates of the index finger of the user's hand. The fingertip coordinates may also be the fingertip coordinates of all recognized fingers of the user's hand, for example, if only three fingers of the user's hand can be recognized in the video stream captured by the camera, the fingertip tracking algorithm identifies the fingertip coordinates of the three fingers.
[0036] In one embodiment, the fingertip coordinates are three-dimensional coordinates (x, y, z) in the imaging space of the augmented reality device.
[0037] S102, fitting a target ray based on the ray direction from the extended reality device to the user's hand and the fingertip coordinates;
[0038] In one embodiment, in the imaging space of the extended reality device, such as Figure 2 As shown, the identified fingertip coordinates are used as the starting point coordinates of the target ray, and the ray direction of the target ray is the direction in which the extended reality device points to the fingertip, and the target ray is fitted. For example, in the imaging space of the extended reality device, the coordinates of the position of the extended reality device (such as the device fitting center coordinates) pointing to the direction of the fingertip can be used as the ray direction.
[0039] In one embodiment, the ray equation of the target ray is calculated based on the fingertip coordinates and the ray direction.
[0040] In one embodiment, the calculation of the ray equation is usually a vector from one point to another point. In the imaging space of the extended reality device, the ray direction of the target ray can be expressed as a direction vector from the user's viewing angle coordinates (which can be the midpoint of the user's binocular line) to the user's fingertip coordinates. The ray extending along the direction vector with the fingertip coordinates as the starting point is the target ray.
[0041] For example, assuming that the user's viewing angle coordinate is E=(E x , E y , E z ), the coordinate of the user’s fingertip is F = (F x , Fy , F z ). The direction vector of the target ray It can be calculated by subtracting the vector result of the eye coordinates from the fingertip coordinates:
[0042]
[0043] In order to get the unit direction vector, you need to transform the direction vector Normalize it so that its length is 1.
[0044] First calculate Length:
[0045]
[0046] Then, the unit direction vector for:
[0047]
[0048] In the ray equation, the unit direction vector D unit Multiplying by the parameter t to get the position of any point on the ray, the ray equation can be expressed as:
[0049] P(t)=P+t*D unit
[0050] Among them, P is the starting point of the target ray, that is, the fingertip coordinate F, P(t) represents the position coordinate corresponding to any parameter t on the target ray, and the parameter t represents the unit direction vector D from the starting point P unit The distance. Among them, because D unit represents a unit vector, and P is the starting point of the target ray. Therefore, the coordinate position of any point on the target ray can be expressed by setting the value of the parameter t. In other words, any point on the target ray uniquely corresponds to a parameter t.
[0051] S103, detecting the collision intersection point between the target ray and the screen interface, and obtaining the coordinates of the collision intersection point;
[0052] Generally, the collision detection algorithm can be used to determine whether two or more objects intersect or touch. In the embodiment of the present application, the collision detection algorithm is used to calculate whether there is a collision intersection between the target ray and the plane where the screen interface is located, and calculate the coordinates of the collision intersection.
[0053] For example, in the embodiment of the present application, the ray equation can be substituted into the plane equation of the plane where the screen interface is located, and then the ray equation P(t)=P+t*D can be solved. unitThe parameter t in is the parameter value of the target ray from the fingertip coordinates to the plane where the screen interface is located. Substituting the parameter t into the ray equation to solve the unique point coordinates is the collision intersection coordinates.
[0054] It is understandable that the collision detection algorithm can also adopt other calculation methods that can be applied to the embodiments of the present application. For example, OpenGL's ray collision detection function can be used to calculate the coordinates of the collision point, or the collision detection process can be optimized through a ray tracing algorithm to further reduce delays and improve the interactive experience.
[0055] Specifically, given the ray equation and the plane equation, the ray equation P(t) = P + t*D unit Substitute the parameter t in the ray equation into the plane equation to solve it. If the value of the parameter t obtained is within a reasonable range (usually a non-negative number), it means that the target ray intersects with the plane. Specifically, if the ray equation is substituted into the plane equation, a non-negative parameter t can be solved, which can be understood as extending from the fingertip coordinates along the ray direction. A point on the target ray will intersect with the plane where the screen interface is located. The intersection point is the coordinate of a point on the target ray solved by substituting the parameter t into the ray equation. Because the coordinate of this point is the intersection of the target ray and the plane where the screen interface is located, the coordinate of this point is also on the plane where the screen interface is located, which is the coordinate of the collision intersection.
[0056] This embodiment combines a fingertip tracking algorithm to identify the fingertip coordinates of the user's hand, fits a target ray based on the fingertip coordinates, and then performs collision detection with the screen interface to obtain the coordinates of the collision intersection, thereby improving the accuracy and response speed of the interaction, allowing the user to directly interact with the virtual interface through gestures, enhancing the sense of immersion and user experience, while reducing interaction delays, improving interaction efficiency, and providing users with a more natural and intuitive interaction method.
[0057] Furthermore, the plane equation of the screen interface is obtained; based on the ray equation corresponding to the target ray and the plane equation, the coordinates of the collision intersection point between the target ray and the screen interface are calculated.
[0058] Generally speaking, the screen interface can be regarded as a plane, which can usually be expressed as:
[0059] n·(p-p0)=0
[0060] Where n is the normal vector of the plane, p is a point on the plane, and p0 is an arbitrary point on the plane.
[0061] In one embodiment, the ray equation is substituted into the plane equation to solve for the parameter t. That is, assuming that there is an intersection point P(t0) between the target ray and the plane where the screen interface is located, the coordinates of the intersection point on the target ray are expressed as:
[0062] P(t0)=P+t0*D unit
[0063] At the same time, the intersection point is also located on the plane where the screen interface is located. The coordinates of the intersection point are substituted into the plane equation, and the parameter t0 is solved by the plane equation. Specifically, the solution process is as follows:
[0064] n·((P+t0·D unit )-p0)=0
[0065] n·(P-p0+t0·D unit )=0
[0066] n·(P-p0)+t0·(n·D unit )=0
[0067] t0·(n·D unit )=-n·(P-p0)
[0068]
[0069] Where n is the normal vector of the plane, P is the coordinate of the finger tip, p0 is the coordinate of any point on the plane, and D unit is the unit direction vector of the target ray; if the calculated result of the parameter t0 is non-negative, then the target ray intersects the plane of the screen.
[0070] Substitute the t0 value into the ray equation to calculate the coordinates of the collision intersection point:
[0071] Intersection Point=P+t0*D unit
[0072] Among them, Intersection Point represents the coordinates of the collision intersection point.
[0073] In one embodiment, because the collision detection algorithm detects the collision intersection between the target ray and the plane where the screen interface is located, but in the imaging space of the extended reality device, the screen interface has a boundary. Therefore, after calculating the coordinates of the collision intersection, it is also necessary to detect whether the X-axis coordinates and Y-axis coordinates of the collision intersection coordinates are within the valid range of the screen interface.
[0074] If the X-axis coordinates and Y-axis coordinates of the collision intersection coordinates are within the valid range, it means that there is effective interaction between the user's fingertips and the screen interface, and the controls in the screen interface can be effectively manipulated; if the X-axis coordinates and Y-axis coordinates of the collision intersection coordinates are not within the valid range, it means that there is no effective interaction between the user's fingertips and the screen interface, and the controls in the screen interface cannot be manipulated.
[0075] Furthermore, based on the collision detection algorithm, the collision intersection of the target ray and the screen interface is detected to obtain the coordinates of the intersection to be determined; the valid coordinate area corresponding to the screen interface is obtained; when the coordinates of the intersection to be determined are within the valid coordinate area, the coordinates of the intersection to be determined are determined to be the collision intersection coordinates.
[0076] In one embodiment, the valid coordinate area refers to an area on the screen interface with which the user can effectively interact, and the valid coordinate area defines the valid coordinate range of the X-axis and the Y-axis of the screen interface in the plane coordinate system.
[0077] Exemplarily, the effective coordinate area can be determined according to the physical size of the screen or the user interface. For example, if the screen interface is a screen projection interface of an external display device (such as a mobile phone, computer, etc.), the regional coordinates of the effective coordinate area are calculated according to the projection ratio and the screen size of the external display device; if the screen interface is a visual display of the user interface within the field of view of the extended reality device, the regional coordinates of the effective coordinate area are determined according to the display setting parameters of the extended reality device, such as the screen starting coordinates, width size, and height size.
[0078] Exemplarily, the valid coordinate area can be a rectangular frame area (or other shapes) in the plane where the screen interface is located. When calculating the coordinates of the pending intersection point where the target ray and the plane where the screen interface is located, it is determined whether the coordinates of the pending intersection point are located in the valid coordinate area.
[0079] If the coordinates of the undetermined intersection point are outside the valid coordinate area, it means that the user's fingertips cannot effectively interact with the screen interface at the current position; if the coordinates of the undetermined intersection point are within the valid coordinate area, it means that the user's fingertips can effectively interact with the screen interface at the current position, and the coordinates of the undetermined intersection point are determined as the collision intersection point coordinates.
[0080] Specifically, the coordinates of the intersection to be determined are three-dimensional coordinates, and the X-axis coordinates and Y-axis coordinates of the intersection to be determined are compared with the convenience of the valid coordinate area. If the X-axis coordinates and Y-axis coordinates of the intersection to be determined are both within the valid coordinate area of the screen interface, the coordinates of the intersection to be determined are determined to be valid and determined as the collision intersection coordinates; conversely, if any of the X-axis coordinates and Y-axis coordinates of the intersection to be determined are outside the valid coordinate area of the screen interface, the coordinates of the intersection to be determined are invalid.
[0081] This embodiment ensures effective interaction between user gestures and the virtual interface by calculating the coordinates of the collision intersection of the target ray and the screen interface, and verifies whether the coordinates of the collision intersection are within the valid coordinate area of the screen interface, thereby improving the accuracy and reliability of gesture control, allowing users to more accurately manipulate controls on the screen, thereby improving the efficiency and practicality of the interaction.
[0082] S104, mapping the coordinates of the collision intersection to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection;
[0083] In one embodiment, after the collision detection is completed, the three-dimensional coordinates of the collision intersection are mapped to two-dimensional screen coordinates in the plane coordinate system where the screen interface is located through a coordinate mapping algorithm.
[0084] Exemplarily, the mapping process of the coordinate mapping algorithm may include: converting the collision intersection coordinates into clipping coordinates through view transformation and projection transformation, converting the clipping coordinates into Normalized Device Coordinates (NDC) coordinates, and then converting the NDC coordinates into screen coordinates.
[0085] Specifically, the view transformation refers to the conversion of the collision intersection coordinates from the world coordinate system to the view coordinate system, that is, the conversion of the collision intersection coordinates in the world coordinate system to the point coordinates in the camera coordinate system. The projection transformation converts the three-dimensional view coordinates into two-dimensional clipping coordinates through perspective projection or orthogonal projection. The clipping coordinates are then converted to NDC coordinates, whose range is [-1,1]. The NDC coordinates are multiplied by the screen size (or viewport size) and added to the viewport offset, and mapped from the range of [-1,1] to the actual pixel coordinates of the screen.
[0086] Generally, in the process of converting the clip coordinates to NDC coordinates, for the X-axis coordinates and the Y-axis coordinates, the clip coordinates are usually divided by the W coordinates (in homogeneous coordinates) and the results are scaled to the range of [-1, 1].
[0087] It is understandable that the coordinate mapping process of the above-mentioned coordinate mapping algorithm is an existing feasible solution in the art. In short, the algorithm maps points in one coordinate system to another coordinate system through a series of mathematical transformations to meet specific application requirements. In different application fields, although the specific implementation details may vary due to different application scenarios and performance requirements, the core principles and methodology are similar. The specific implementation process can refer to the relevant prior art in the art, and the specific calculation process of the present application embodiment is not described in detail.
[0088] S105: Identify the target control pointed to by the screen coordinates in the screen interface, so as to perform interactive operations on the target control.
[0089] In one embodiment, after the screen coordinates corresponding to the collision intersection are calculated, corresponding UI interaction events such as Hover, Move, Click, etc. can be triggered according to the screen coordinates in the corresponding controls in the control system (such as Android system) interface. The control system responds to the current user's operation according to the triggered UI interaction event, such as selecting a button, sliding a list, etc.
[0090] Furthermore, based on the fingertip tracking algorithm, the fingertip coordinate movement of the user's hand is calculated; based on the coordinate mapping relationship between the collision intersection coordinates and the screen coordinates, the fingertip coordinate movement is mapped to the plane coordinate system where the screen interface is located to obtain the screen coordinate movement; based on the screen coordinate movement, a corresponding cursor movement operation is performed in the screen interface.
[0091] During coordinate mapping, the mapped screen coordinates can be displayed in the form of a cursor in the screen interface to intuitively indicate the position pointed to by the screen coordinates mapped to the user's fingertip coordinates in the plane coordinate system where the screen interface is located, making it easier for users to accurately interact with the controls in the screen interface.
[0092] In one embodiment, the coordinate distance change of the fingertip coordinates collected twice in succession is calculated by a fingertip tracking algorithm to obtain the fingertip coordinate movement amount during the two consecutive fingertip coordinate collection processes. For example, if the fingertip coordinates collected for the first time are (x1, y1, z1) and the fingertip coordinates collected for the second time are (x2, y2, z2), then the fingertip coordinate movement amount is (x2-x1, y2-y1, z2-z1). Based on the coordinate mapping relationship between the collision intersection coordinates and the screen coordinates, the fingertip coordinate movement amount is mapped to the plane coordinate system where the screen interface is located and converted into the screen coordinate movement amount in the screen interface.
[0093] In one embodiment, the fingertip coordinate movement of the fingertip coordinates may be calculated first, and then the fingertip coordinate movement may be mapped to the screen coordinate movement. Alternatively, the fingertip coordinates may be mapped to the screen coordinates first, and then the screen coordinate movement between the two mapped screen coordinates may be calculated.
[0094] In one embodiment, because the screen coordinates are two-dimensional coordinates, the screen coordinate movement can be composed of two parts: the X-axis coordinate movement and the Y-axis coordinate movement. Therefore, it is only necessary to map the X-axis coordinate movement and the Y-axis coordinate movement in the fingertip coordinates to the plane coordinate system where the screen interface is located to obtain the screen coordinate movement. Alternatively, the X-axis coordinates and the Y-axis coordinates in the fingertip coordinates collected twice are mapped to the plane coordinate system to obtain the screen coordinates corresponding to the fingertip coordinates collected twice, and then the coordinate movement of the two screen coordinates on the X-axis and the coordinate movement on the Y-axis are calculated to obtain the screen coordinate movement.
[0095] Among them, the mapping of the coordinate movement amount can be converted according to the size conversion ratio between the pixel size of the three-dimensional coordinate system in the extended reality device and the pixel size of the plane coordinate system where the screen interface is located. For example, if the size conversion ratio between the pixel size of the three-dimensional coordinate system in the extended reality device and the pixel size of the plane coordinate system where the screen interface is located is 2:1, that is, a point in the three-dimensional coordinate system moves two pixel coordinate distances along the X-axis, which is mapped to a pixel coordinate distance along the X-axis in the plane coordinate system where the screen interface is located. Correspondingly, when mapping the X / Y axis coordinate movement amount, it is also necessary to multiply the coordinate movement amount of the X-axis and the coordinate movement amount of the Y-axis in the fingertip coordinates by the size conversion ratio to adapt to the coordinate size in the plane coordinate system where the screen interface is located, and map it to the X-axis coordinate movement amount and the Y-axis coordinate movement amount in the plane coordinate system.
[0096] This embodiment converts the three-dimensional collision intersection coordinates into two-dimensional coordinates in the screen coordinate system through a coordinate mapping algorithm, thereby achieving precise interaction between user gestures and screen interface controls, simplifying the connection between gesture recognition and screen operations, and enabling the control system to respond to user gesture operations more quickly and accurately, which not only improves the accuracy of gesture control, but also enhances the intuitiveness and convenience of the user interaction experience.
[0097] The present embodiment provides a human-computer interaction method, which uses a fingertip tracking algorithm to accurately identify the fingertip coordinates of the user's hand, provides accurate input points for subsequent interactions, and ensures the accuracy of the interaction; uses the ray direction from the user's eyes to the fingertips and the fingertip coordinates to fit the target ray, which is used to determine the user's interaction object, making the interaction more intuitive and natural; determines the collision intersection between the target ray and the screen interface through a collision detection algorithm, and determines the specific location of the user's intended operation, which helps to achieve faster response and more precise control, enhance the real-time nature of the interaction, and improve the interaction efficiency; converts the collision intersection coordinates into screen coordinates through a coordinate mapping algorithm, achieves accurate correspondence between user gestures and controls, thereby improving the accuracy and efficiency of the interaction, and making the interaction in an augmented reality environment more convenient and efficient.
[0098] Please refer to Figure 3 , Figure 3 A flowchart of a second embodiment of a human-computer interaction method provided in an embodiment of the present application.
[0099] like Figure 3 As shown, based on the above Figure 1 In the illustrated embodiment, after step S104, the following steps are further included:
[0100] S201, obtaining a coordinate area corresponding to at least one control in a screen interface;
[0101] In one embodiment, the screen interface can be used as a coordinate plane according to the resolution of the screen interface. For example, the upper left corner of the screen interface can be used as the coordinate origin (0,0), the horizontal direction is the X axis, and the vertical direction is the Y axis. The width of the screen interface defines the maximum value of the X axis, and the height of the screen interface defines the maximum value of the Y axis. For example, if the screen resolution is 1920*1080, the range of the X axis is 0 to 1920, and the range of the Y axis is 0 to 1080.
[0102] Among them, parameters such as the coordinate origin of the plane coordinate system where the screen interface is located can be set according to actual needs or user habits. For example, the plane coordinate system can be constructed with the center of the screen as the origin, or other points in the screen interface can be used as the coordinate origin.
[0103] In one embodiment, in the plane coordinate system where the screen interface is located, the coordinate area of each control is recorded according to the position of the area where each control is located in the screen interface relative to the coordinate origin. For example, for an icon control of an application APP, the coordinates of the four vertex coordinates of the icon control can be collected, and the area formed by the four vertex coordinates is used as the coordinate area corresponding to the icon control of the application APP. For example, for an input method keyboard, the coordinates of the intersection of the edges of the input method keyboard can be collected to determine the coordinate area of the input method keyboard. For the keys in the input method keyboard, the coordinate area corresponding to each key control can be calculated based on the area size of the input method keyboard, the coordinate area, the size and distribution of the keys.
[0104] S202, determining a target control based on a coordinate area corresponding to the screen coordinates;
[0105] In one embodiment, coordinate matching is performed between the screen coordinates and the coordinate regions of each control to determine whether the screen coordinates are located in the coordinate region corresponding to a control. If the screen coordinates are located in the coordinate region corresponding to a control, the control is determined as the target control, so that the user can perform the next interactive operation on the control.
[0106] In one embodiment, if the screen coordinates are located in a blank area, that is, there is no control at the position corresponding to the screen coordinates, interactive operations can be performed on the entire screen interface, such as performing an interface sliding operation on the screen interface, switching the current display page of the screen interface to the next display page hidden in the sliding direction, etc.
[0107] This embodiment achieves precise control positioning and convenient user interaction by converting the screen interface into a coordinate plane and recording the coordinate area of the control, thereby improving the accuracy and efficiency of the operation and enhancing the user experience.
[0108] S203, obtaining a control operation gesture library corresponding to the target control, wherein the control operation gesture library includes at least one control operation gesture;
[0109] In one embodiment, the control operation gesture refers to a gesture used by a user when interacting with a control, such as clicking, sliding, long pressing, etc. These control operation gestures can be recognized and converted into corresponding control operation instructions.
[0110] In one embodiment, the control operation gesture library may include predefined control operation gestures; the user may also be allowed to customize the control operation gestures and save them in the control operation gesture library, or modify and replace the predefined control operation gestures.
[0111] Furthermore, at least one control operation instruction corresponding to a control in the screen interface is obtained; and the control operation gestures corresponding to each of the control operation instructions are preset to construct a control operation gesture library corresponding to the control.
[0112] In one embodiment, the control operation gesture library includes at least one set of predefined control operation gestures, and each control operation gesture corresponds to a control operation instruction. Different control operation instructions can correspond to different control operation gestures for distinction; different control operation instructions can also correspond to the same control operation gesture. For example, if a certain control corresponds to a unique detail page, the user can click once to display the unique detail page, and click again to hide the unique detail page.
[0113] For example, the interactive operations that can be performed by each control can be obtained in advance, such as selection operation, editing operation, etc. Then, the control operation gesture corresponding to each interactive operation is defined, and then the control operation instruction corresponding to each control operation gesture is determined, and a control operation gesture library corresponding to each control is created. For example, a single-click gesture corresponds to a selection operation instruction, and a long-press gesture corresponds to an editing operation instruction.
[0114] Specifically, for an input method keyboard, lightly touching a key on the keyboard can be defined as an operation gesture for inputting a corresponding character. The extended reality device can recognize the lightly touching operation gesture through a gesture recognition algorithm and convert it into a corresponding character input operation instruction.
[0115] This embodiment can make the interaction between the user and the control more intuitive and personalized by building a control operation gesture library. The control operation gesture library not only contains predefined gesture operations, but also supports user-defined gestures, which improves the flexibility and convenience of operation. At the same time, by matching gestures with control operation instructions, the user's intention can be more accurately identified, thereby improving the user experience and interaction efficiency.
[0116] S204: Identify the control operation gesture corresponding to the user's gesture action, generate a control operation instruction, and perform a corresponding control operation on the target control according to the control operation instruction.
[0117] In one embodiment, gesture recognition algorithms such as machine learning and deep learning can be used to implement gesture recognition. For example, the gesture recognition API of MediaPipe can be used to analyze the coordinates of key points of the hand and recognize gestures, and then the control operation gesture corresponding to the user's gesture can be recognized by analyzing the gesture.
[0118] For example, a single click can be recognized as a single touch without movement within a short period of time; a double click as two quick clicks; a long press as touching and holding for a period of time; a slide as touching and moving the finger; a drag as long press followed by moving the finger; and a pinch as touching and moving two fingers at the same time.
[0119] Furthermore, based on a gesture recognition algorithm, a first gesture feature of the gesture action is identified; based on a feature matching algorithm, feature matching is performed on the first gesture feature and the second gesture feature of each of the control operation gestures to generate the control operation instruction corresponding to the control operation gesture with the highest matching degree.
[0120] Exemplarily, the MediaPipe gesture detection model can be used to recognize gestures, provide 3D coordinate points of finger key points (such as fingertips, joints, etc.), and identify the first gesture feature of the gesture action by acquiring hand key point (landmarks) information.
[0121] In one embodiment, in the control operation gesture library, a control operation gesture corresponding to each control operation instruction may be constructed, and a second gesture feature such as a key point pattern or a coordinate sequence corresponding to each control operation gesture may be defined.
[0122] In one embodiment, a feature matching algorithm is used to calculate the feature similarity between the first gesture feature and the second gesture feature of each of the control operation gestures, and the second gesture feature with the highest matching degree is screened out, and then the gesture action made by the user is identified as the control operation gesture corresponding to the second gesture feature, and a control operation instruction corresponding to the control operation gesture is generated, and the corresponding operation is performed on the target control according to the generated control operation instruction.
[0123] Exemplarily, a feature matching algorithm such as a Brute-Force matching algorithm is used to match the first gesture feature (i.e., the feature descriptor of the user gesture) with the second gesture feature (i.e., the feature descriptor of the preset gesture) in the control operation gesture library. The Brute-Force matching algorithm finds the most matching feature point pair by calculating the similarity of each pair of feature points. For each pair of feature point pairs obtained by matching, the distance between the feature descriptors is calculated to express the feature similarity, wherein the feature similarity calculation can be performed using distance measurement algorithms such as Euclidean distance and Hamming distance. The matching results can be filtered using K nearest neighbor matching (KNN) and comparison threshold screening. For example, for each feature point, the two closest matching points are found and the ratio of their distances is calculated. According to the filtered matching results, the similarity scores of all matching pairs are compared to select the feature point pair with the highest matching degree, and then determine the control operation gesture corresponding to the gesture action.
[0124] In one embodiment, after the gesture action recognition is completed, a corresponding control operation instruction is generated, and a control operation, such as clicking, sliding, etc., is performed on the target control according to the generated control operation instruction.
[0125] For example, suppose a user switches application interfaces by sliding gestures. The camera can capture the sliding motion of the hand, and the gesture recognition algorithm can be used to identify that it is a sliding gesture. Then, the control operation instruction for switching interfaces is generated, and the interface switching operation is performed. In this way, the user can control the switching of application interfaces by gestures.
[0126] This embodiment maps fingertip coordinates to screen coordinates through a combination of coordinate mapping and gesture recognition, accurately identifies target controls, and uses a gesture recognition algorithm to recognize user gestures and convert them into control operation instructions, thereby achieving efficient control of target controls, improving the naturalness and intuitiveness of user interaction, and enhancing the flexibility and response speed of operations, providing users with a more convenient and intuitive interactive experience.
[0127] See also Figure 4 , Figure 4 It is a structural schematic diagram of a first embodiment of a human-computer interaction device provided by the present application, and the human-computer interaction device is used to execute the aforementioned human-computer interaction method.
[0128] like Figure 4 As shown, the human-computer interaction device 300 includes: a fingertip coordinate recognition module 301, a ray fitting module 302, a collision intersection detection module 303, a coordinate mapping module 304 and a target control recognition module 305.
[0129] A fingertip coordinate recognition module 301 is used to recognize the fingertip coordinates of the user's hand;
[0130] A ray fitting module 302, configured to fit a target ray based on a ray direction from an extended reality device to the user's hand and the fingertip coordinates;
[0131] A collision intersection detection module 303 is used to detect the collision intersection between the target ray and the screen interface to obtain the collision intersection coordinates;
[0132] A coordinate mapping module 304 is used to map the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, so as to obtain the screen coordinates corresponding to the collision intersection point;
[0133] The target control identification module 305 is used to identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operations on the target control.
[0134] In one embodiment, the collision intersection detection module 303 includes:
[0135] A plane equation obtaining unit, used to obtain the plane equation of the screen interface;
[0136] A collision intersection coordinate calculation unit is used to calculate the collision intersection coordinates of the target ray and the screen interface based on the ray equation corresponding to the target ray and the plane equation.
[0137] In one embodiment, the collision intersection detection module 303 further includes:
[0138] A unit for detecting the coordinates of the intersection to be determined, used to detect the collision intersection between the target ray and the screen interface based on a collision detection algorithm, and obtain the coordinates of the intersection to be determined;
[0139] An effective coordinate area acquisition unit, used to acquire the effective coordinate area corresponding to the screen interface;
[0140] The collision intersection coordinate determining unit is used to determine the undetermined intersection coordinates as the collision intersection coordinates when the undetermined intersection coordinates are within the valid coordinate area.
[0141] In one embodiment, the human-computer interaction device 300 further includes a coordinate movement mapping module, including:
[0142] A coordinate movement amount calculation unit, used to calculate the coordinate movement amount of the fingertip of the user's hand based on the fingertip tracking algorithm;
[0143] A screen coordinate movement amount acquisition unit, configured to map the fingertip coordinate movement amount to the plane coordinate system where the screen interface is located based on the coordinate mapping relationship between the collision intersection coordinates and the screen coordinates, so as to obtain the screen coordinate movement amount;
[0144] A movement operation execution unit is used to execute a corresponding cursor movement operation in the screen interface based on the screen coordinate movement amount.
[0145] In one embodiment, the human-computer interaction device 300 further includes a control operation instruction generating module, including:
[0146] A control coordinate area acquisition unit, used to acquire a coordinate area corresponding to at least one control in the screen interface;
[0147] A target space determination unit, configured to determine a target control based on a coordinate area corresponding to the screen coordinates;
[0148] A control operation gesture library acquisition unit, used to acquire a control operation gesture library corresponding to the target control, wherein the control operation gesture library includes at least one control operation gesture;
[0149] The control operation instruction generating unit is used to identify the control operation gesture corresponding to the user's gesture action, generate a control operation instruction, and perform the corresponding control operation on the target control according to the control operation instruction.
[0150] In one embodiment, the control operation instruction generation module further includes:
[0151] A control operation instruction acquisition unit, used to acquire at least one control operation instruction corresponding to a control in the screen interface;
[0152] The control operation gesture library construction unit is used to preset the control operation gesture corresponding to each control operation instruction and construct a control operation gesture library corresponding to the control.
[0153] In one embodiment, the coordinate mapping module 304 includes:
[0154] A clipping coordinate conversion unit, used to convert the collision intersection coordinates into clipping coordinates through view transformation and projection transformation;
[0155] NDC coordinate conversion unit, used to convert the clipping coordinates into standardized device NDC coordinates;
[0156] The screen coordinate conversion unit is used to convert NDC coordinates into screen coordinates.
[0157] It should be noted that technicians in the relevant field can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described device and each module can refer to the corresponding process in the aforementioned human-computer interaction method embodiment, and will not be repeated here.
[0158] The apparatus provided in the above embodiment may be implemented in the form of a computer program. The computer program may be Figure 5 Running on an extended reality device as shown.
[0159] See also Figure 5 , Figure 5 1 is a schematic block diagram of the structure of an extended reality device provided in an embodiment of the present application. The extended reality device may be a server.
[0160] See also Figure 5 The extended reality device includes a processor, a memory, and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0161] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any human-computer interaction method.
[0162] The processor is used to provide computing and control capabilities to support the operation of the entire extended reality device.
[0163] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any human-computer interaction method.
[0164] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present application, and does not constitute a limitation on the extended reality device to which the scheme of the present application is applied. The specific extended reality device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0165] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0166] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:
[0167] Identify the coordinates of the fingertips of the user's hand;
[0168] Fitting a target ray based on a ray direction from the extended reality device to the user's hand and the fingertip coordinates;
[0169] Detecting the collision intersection point between the target ray and the screen interface to obtain the coordinates of the collision intersection point;
[0170] Mapping the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection point;
[0171] Identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operation on the target control.
[0172] In one embodiment, when the processor implements the collision detection algorithm to detect the collision intersection between the target ray and the screen interface and obtain the coordinates of the collision intersection, it is used to implement:
[0173] Obtaining the plane equation of the screen interface;
[0174] Based on the ray equation corresponding to the target ray and the plane equation, the coordinates of the collision intersection point between the target ray and the screen interface are calculated.
[0175] In one embodiment, when the processor implements the collision detection algorithm to detect the collision intersection between the target ray and the screen interface and obtain the coordinates of the collision intersection, it is also used to implement:
[0176] Based on the collision detection algorithm, the collision intersection point between the target ray and the screen interface is detected to obtain the coordinates of the intersection point to be determined;
[0177] Obtaining a valid coordinate area corresponding to the screen interface;
[0178] When the coordinates of the undetermined intersection point are located in the valid coordinate area, the coordinates of the undetermined intersection point are determined as the collision intersection point coordinates.
[0179] In one embodiment, after implementing the coordinate mapping algorithm to map the coordinates of the collision intersection to the plane coordinate system where the screen interface is located and obtaining the screen coordinates corresponding to the collision intersection, the processor is further used to implement:
[0180] Based on the fingertip tracking algorithm, calculate the coordinate movement of the fingertip of the user's hand;
[0181] Based on the coordinate mapping relationship between the collision intersection coordinates and the screen coordinates, mapping the fingertip coordinate movement amount to the plane coordinate system where the screen interface is located to obtain the screen coordinate movement amount;
[0182] Based on the screen coordinate movement amount, a corresponding cursor movement operation is performed in the screen interface.
[0183] In one embodiment, after implementing the coordinate mapping algorithm to map the coordinates of the collision intersection to the plane coordinate system where the screen interface is located and obtaining the screen coordinates corresponding to the collision intersection, the processor is further used to implement:
[0184] Get the coordinate area corresponding to at least one control in the screen interface;
[0185] Determining a target control based on a coordinate area corresponding to the screen coordinates;
[0186] Acquire a control operation gesture library corresponding to the target control, wherein the control operation gesture library includes at least one control operation gesture;
[0187] The control operation gesture corresponding to the user's gesture action is identified, and a control operation instruction is generated to perform a corresponding control operation on the target control according to the control operation instruction.
[0188] In one embodiment, before the processor implements the acquiring of the control operation gesture library corresponding to the target control, wherein the control operation gesture library includes at least one control operation gesture, the processor is further configured to implement:
[0189] Obtaining at least one control operation instruction corresponding to a control in the screen interface;
[0190] The control operation gestures corresponding to the control operation instructions are preset to construct a control operation gesture library corresponding to the controls.
[0191] In one embodiment, when the processor implements mapping the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located to obtain the screen coordinates corresponding to the collision intersection point, it is used to implement:
[0192] The collision intersection coordinates are converted into clipping coordinates through view transformation and projection transformation;
[0193] Convert the clipping coordinates to normalized device NDC coordinates;
[0194] The NDC coordinates are converted to the screen coordinates.
[0195] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program. The computer program includes program instructions. The processor executes the program instructions to implement any human-computer interaction method provided in the embodiment of the present application.
[0196] The computer-readable storage medium may be an internal storage unit of the extended reality device described in the foregoing embodiment, such as a hard disk or memory of the extended reality device. The computer-readable storage medium may also be an external storage device of the extended reality device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the extended reality device.
[0197] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A human-computer interaction method, characterized in that: The method comprises: Identify the coordinates of the fingertips of the user's hand; Fitting a target ray based on a ray direction from the extended reality device to the user's hand and the fingertip coordinates; Detecting the collision intersection point between the target ray and the screen interface to obtain the coordinates of the collision intersection point; Mapping the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection point; Identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operation on the target control.
2. The human-computer interaction method according to claim 1, characterized in that: The method of detecting the collision intersection point between the target ray and the screen interface based on the collision detection algorithm and obtaining the coordinates of the collision intersection point includes: Obtaining the plane equation of the screen interface; Based on the ray equation corresponding to the target ray and the plane equation, the coordinates of the collision intersection point between the target ray and the screen interface are calculated.
3. The human-computer interaction method according to claim 1, characterized in that: The method of detecting the collision intersection between the target ray and the screen interface based on the collision detection algorithm and obtaining the coordinates of the collision intersection also includes: Based on the collision detection algorithm, the collision intersection point between the target ray and the screen interface is detected to obtain the coordinates of the intersection point to be determined; Obtaining a valid coordinate area corresponding to the screen interface; When the coordinates of the undetermined intersection point are located in the valid coordinate area, the coordinates of the undetermined intersection point are determined as the collision intersection point coordinates.
4. The human-computer interaction method according to claim 1, characterized in that: After mapping the coordinates of the collision intersection to the plane coordinate system of the screen interface based on the coordinate mapping algorithm and obtaining the screen coordinates corresponding to the collision intersection, the method further includes: Based on the fingertip tracking algorithm, calculate the coordinate movement of the fingertip of the user's hand; Based on the coordinate mapping relationship between the collision intersection coordinates and the screen coordinates, mapping the fingertip coordinate movement amount to the plane coordinate system where the screen interface is located to obtain the screen coordinate movement amount; Based on the screen coordinate movement amount, a corresponding cursor movement operation is performed in the screen interface.
5. The human-computer interaction method according to claim 1, characterized in that: After mapping the coordinates of the collision intersection to the plane coordinate system of the screen interface based on the coordinate mapping algorithm and obtaining the screen coordinates corresponding to the collision intersection, the method further includes: Get the coordinate area corresponding to at least one control in the screen interface; Determining a target control based on a coordinate area corresponding to the screen coordinates; Acquire a control operation gesture library corresponding to the target control, wherein the control operation gesture library includes at least one control operation gesture; The control operation gesture corresponding to the user's gesture action is identified, and a control operation instruction is generated to perform a corresponding control operation on the target control according to the control operation instruction.
6. The human-computer interaction method according to claim 5, characterized in that: The acquiring of the control operation gesture library corresponding to the target control, wherein before the control operation gesture library includes at least one control operation gesture, further includes: Obtaining at least one control operation instruction corresponding to a control in the screen interface; The control operation gestures corresponding to the control operation instructions are preset to construct a control operation gesture library corresponding to the controls.
7. The human-computer interaction method according to claim 1, characterized in that: Mapping the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located to obtain the screen coordinates corresponding to the collision intersection point includes: The collision intersection coordinates are converted into clipping coordinates through view transformation and projection transformation; Convert the clipping coordinates to normalized device NDC coordinates; The NDC coordinates are converted to the screen coordinates.
8. A human-computer interaction device, characterized in that: The human-computer interaction device comprises: A fingertip coordinate recognition module, used to recognize the fingertip coordinates of the user's hand; A ray fitting module, used for fitting a target ray based on a ray direction from an extended reality device to the user's hand and the fingertip coordinates; A collision intersection detection module is used to detect the collision intersection between the target ray and the screen interface to obtain the coordinates of the collision intersection; A coordinate mapping module, used to map the coordinates of the collision intersection point to the plane coordinate system where the screen interface is located, to obtain the screen coordinates corresponding to the collision intersection point; The target control identification module is used to identify the target control pointed to by the screen coordinates in the screen interface, so as to facilitate interactive operations on the target control.
9. An extended reality device, characterized in that: The extended reality device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the human-computer interaction method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the human-computer interaction method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Human-computer interaction method and apparatus, extended reality device, and storage medium
WO2026144015A1