A method for guiding human-computer interaction
By collecting and switching user gaze and hand information in real time, identifying user confusion, and proactively providing on-screen interface guidance, the problem of high learning costs and low efficiency caused by the complexity of in-vehicle central control interface operation is solved, thereby improving user interaction efficiency and experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FORYOU GENERAL ELECTRONICS
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-26
Smart Images

Figure CN122086239A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and in particular to a human-computer interaction guidance method. Background Technology
[0002] With the development of "software-defined vehicles," smart cockpits are becoming increasingly complex, integrating a large amount of traditional and implicit interaction logic. Currently, users face two main challenges when encountering complex in-vehicle control interfaces: first, difficulty in quickly locating the required function; and second, after finding interface elements, uncertainty about their specific functions or interaction methods. Existing user guidance mainly relies on offline methods such as paper manuals, in-vehicle electronic manuals, or sales presentations, requiring users to interrupt their current operation to look up information, resulting in a heavy memory burden and low efficiency. Furthermore, in pursuit of interface simplicity, many operational clues (such as edge swiping and long-press functions) are deliberately hidden, lacking visual cues, forcing users to explore through trial and error, leading to high learning costs and frustration. Therefore, there is an urgent need for an intelligent interaction solution that can provide timely, proactive, and non-intrusive operational guidance. Summary of the Invention
[0003] The purpose of this invention is to disclose a human-computer interaction guidance method that solves the technical problems of heavy memory burden and low efficiency in existing human-computer interaction systems.
[0004] To achieve the above objectives, the present invention discloses a human-computer interaction guidance method, comprising: Acquire the user's gaze and hand information, and convert them into coordinates in the same screen coordinate system; Based on the coordinates within a predetermined time window, calculate the eye-hand coordination distance, hand position stability, and tortuosity of the hand movement trajectory; Based on the judgment results of the eye-hand coordination distance, the stability of the hand position, and the tortuosity of the hand movement trajectory, it is identified whether the user is in a state of confusion and hesitation regarding the screen interface elements; When the confused and hesitant state is identified, interactive guidance information is actively triggered and displayed on the screen interface elements before the user performs an actual touch operation.
[0005] As an optional implementation, the step of acquiring the user's gaze information and hand information and converting coordinates includes: Capture the user's facial and hand images through a camera; The gaze direction is determined based on the facial image, and the gaze direction is mapped to the screen coordinate system through a pre-stored extrinsic parameter matrix from the camera coordinate system to the screen coordinate system to obtain the coordinates of the gaze point. Based on the hand image, the coordinates of key hand points are determined, and then transformed to the screen coordinate system using the extrinsic parameter matrix to obtain the three-dimensional coordinates of the fingertip in the screen coordinate system.
[0006] As an optional implementation, the extrinsic parameter matrix is obtained and stored by calibration during system initialization. The calibration includes: displaying a preset calibration pattern on the screen, capturing images through a camera, and calculating the rotation matrix and translation vector of the camera relative to the screen based on the known physical coordinates of the pattern in the screen coordinate system and the pixel coordinates in the image to form the extrinsic parameter matrix.
[0007] As an optional implementation, the determination based on eye-hand coordination distance includes: Calculate the Euclidean distance between the coordinates of the gaze point in the current frame and the coordinates of the fingertip's projection point on the screen plane; If the Euclidean distance is less than the first threshold, then the eye-hand coordination condition is satisfied.
[0008] As an optional implementation, the determination based on hand position stability includes: Calculate the standard deviation of the fingertip projection point coordinates within the predetermined time window; The standard deviation is compared with a dynamic threshold. If the standard deviation is less than the dynamic threshold, the hand is determined to be in a stable hovering state. The dynamic threshold is determined based on a preset physical radius and a correction coefficient related to the environmental state.
[0009] As an optional implementation, the dynamic threshold σ_th is calculated using the following formula:
[0010] Among them, R phy The physical radius is a preset value, and k is an environmental correction coefficient. When the vehicle is stationary, k = 1.0. When the vehicle is moving, k increases with the vehicle speed v, and there is an upper limit value.
[0011] As an optional implementation, the determination based on the curvature of the hand movement trajectory includes: Calculate the total path length L of the fingertip projection point movement within the predetermined time window. total ; Calculate the net displacement L between the start and end points within the window. net ; Calculate tortuosity ; If τ is greater than the second threshold, the hand movement trajectory is determined to conform to the hesitation feature.
[0012] As an optional implementation, the step of actively triggering and displaying interactive guidance information includes: within a predetermined delay time after recognizing a state of confusion or hesitation, displaying a guidance layer overlaid on the screen interface element, the guidance layer including a visual emphasis effect on the target interface element, a visual weakening effect on non-target areas, and text or graphic prompts for explaining functions or operation methods.
[0013] As an optional implementation, the visual emphasis effect includes proportionally enlarging the target interface element and adding a highlighted outline.
[0014] As an optional implementation, the guide layer is displayed for a predetermined time. If a valid user operation on the target interface element is detected during this period, the guide layer is turned off in advance.
[0015] The beneficial effects of this invention are as follows: By collecting and fusing user gaze and hand information in real time, and mapping it to a unified screen coordinate system using a pre-calibrated extrinsic parameter matrix, the system establishes a foundation for eye-hand coordination perception. Within a sliding time window, it sequentially calculates and judges eye-hand coordination distance, hand position stability based on dynamic thresholds, and the tortuosity of hand movement trajectory. This allows the system to accurately distinguish between unconscious hovering, purposeful clicking, and hesitation caused by confusion. When hesitation is detected, the system proactively triggers a guidance module before the user actually touches the screen, providing immediate and clear operational guidance. This invention achieves a shift from "response after contact" to "perception and proactive assistance before contact," effectively reducing the learning cost and cognitive load for users in complex in-vehicle interfaces, minimizing frustration caused by blind trial and error, and improving the intuitiveness and efficiency of interaction. Simultaneously, the dynamic threshold design enables the system to adapt to environmental changes such as vehicle vibration, ensuring robustness of judgment and consistency with user experience. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the human-computer interaction guidance method of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] In this invention, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing the invention and its embodiments, and are not intended to limit the indicated devices, elements, or components to having a specific orientation, or to be constructed and operated in a specific orientation.
[0020] Furthermore, in addition to indicating direction or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in certain situations to indicate a dependency or connection. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0021] Furthermore, the terms "installation," "setup," "equipped with," "connection," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of these terms in this invention based on the specific circumstances.
[0022] Furthermore, the terms "first," "second," etc., are primarily used to distinguish different devices, components, or parts (which may be the same or different in specific type and construction), and are not intended to indicate or imply the relative importance or quantity of the indicated devices, components, or parts. Unless otherwise stated, "a plurality of" means two or more.
[0023] The technical solution of the present invention will be further described below with reference to the embodiments and accompanying drawings.
[0024] like Figure 1As shown, this embodiment provides a human-computer interaction guidance method applicable to in-vehicle intelligent cockpit systems equipped with a central control touchscreen display, a camera assembly, and a corresponding processing unit. The camera assembly includes at least one infrared or RGB camera facing the driver or passenger, mounted on the A-pillar or the upper edge of the central control screen, for capturing images of the user's face and hands. The central control display has a capacitive touch layer for detecting touch operations. The system also includes a memory and a processor for executing the steps of the method described in this embodiment.
[0025] This embodiment provides a human-computer interaction guidance method, applicable to single-finger operation, including the following steps: Step 1: System calibration and parameter fixing.
[0026] During vehicle assembly or equipment manufacturing, spatial calibration between the camera coordinate system and the screen coordinate system is pre-commissioned to obtain and solidify the extrinsic parameter matrix describing the pose relationship between the two. The specific calibration process includes: Step 101: Establish the calibration environment. Ensure that the physical installation positions of the camera and the central control display screen are relatively fixed. Establish a screen coordinate system with the display plane of the central control display screen as the XY plane and the positive Z-axis direction perpendicular to the screen outward as the positive Z-axis, and a camera coordinate system with the optical center of the camera as the origin.
[0027] Step 102: Acquire calibration feature points. Display a specific calibration pattern at a preset position in the screen display area, or place a physical calibration plate with known geometric dimensions on the screen surface. Control the camera to capture image frames containing the calibration pattern, and extract the coordinates of the calibration feature points in the camera pixel coordinate system.
[0028] Step 103: Solve the extrinsic parameter matrix. Based on the PnP algorithm or the principles of stereo vision geometry, calculate the rotation matrix and translation vector of the camera relative to the screen based on the known physical coordinates of the calibrated feature points in the screen coordinate system and their pixel coordinates in the image frame. This rotation matrix and translation vector together constitute the extrinsic parameter matrix.
[0029] Step 104: Solidify parameters. The calculated extrinsic parameter matrix is stored as a fixed system parameter in the non-volatile memory of the vehicle terminal for direct access during subsequent system operation.
[0030] Step 2: Real-time data acquisition and coordinate transformation.
[0031] When the system runs, the following sub-steps are executed: Step 201: Collect facial and hand image frames of the operator in real time through the camera, and collect the touch signals on the screen surface through the capacitive touch layer of the central control screen as a reference.
[0032] Step 202: Based on the facial image frame, identify the pupil center and corneal reflection point to construct the gaze vector. Using the extrinsic parameter matrix fixed in Step 1, transform the gaze direction vector from the camera coordinate system to the screen coordinate system, and generate the coordinates of the gaze point on the screen plane using geometric intersection. .
[0033] Step 203: Based on the hand image frame, identify the key points of the hand skeleton and extract the coordinates of the index fingertip. Similarly, call the aforementioned extrinsic parameter matrix to transform the fingertip coordinates from the camera coordinate system to the screen coordinate system, obtaining the three-dimensional coordinates of the fingertip in the screen coordinate system. Where x and y represent the projected positions of the fingertip on the screen plane, and z represents the vertical distance from the fingertip to the screen plane. The screen coordinate system is established with reference to the display plane of the central control display screen, and its origin and coordinate axis directions are preset according to the screen geometric parameters.
[0034] Step 3: Determine user intent features.
[0035] The system determines user intent based on data within a sliding time window, where the time window length T is set to 500ms, the sampling frequency f is set to 60Hz, and the data corresponds to the most recent N=30 frames. The system continuously stores and updates these N frames of data, each frame containing: coordinates of the gaze point. and fingertip three-dimensional coordinates , where i=N represents the latest frame.
[0036] The judgment process includes the following sub-steps in sequence: Step 301: Calculate the eye-hand coordination distance D ehc .
[0037] Specifically, it involves calculating the coordinates of the current frame's viewpoint. The projection point of the fingertip on the screen plane Euclidean distance between them:
[0038] If D ehc If the distance is less than 60mm, it is determined that the user's gaze is focused near the hand, meeting the "eye-hand coordination" condition, and step 302 is executed; otherwise, it is determined to be a high risk of unconscious hovering or accidental touch, and guidance is not triggered, and monitoring continues.
[0039] The distance threshold (60mm) in this step is set based on research into the characteristics of human visual-hand coordination. In a typical in-vehicle cockpit (approximately 50-70 cm) viewing distance, when an adult user performs a precise operation requiring visual feedback, the focus of their gaze and the projection point of their operating finger typically overlap within a 40mm range. Considering minor head movements, inherent noise in gaze-tracking algorithms, and slight gaze shifts that may occur as users search for cues, setting the threshold to 60mm effectively filters out unintentional fixations (high accuracy) caused by scanning other areas of the screen, while maintaining a high intent recognition rate (high recall). This threshold can be fine-tuned through calibration experiments to adapt to the cockpit layout of different vehicle models.
[0040] Step 302: Calculate the stability of the fingertip position.
[0041] Specifically, it is based on the coordinates of the fingertip projection points in the N frames. Calculate its standard deviation within the window. :
[0042] Where, x i y i Let X and Y be the projection points of the fingertip in the i-th frame, respectively. , These are the average values of the X and Y coordinates, respectively.
[0043] The calculated standard deviation σ xy With dynamic judgment threshold σ th The dynamic threshold σ is compared. th The calculation is obtained through a screen pixel density adaptive algorithm, and the formula is as follows:
[0044] Among them, R phy The preset effective hovering determination physical radius is set to 5mm in this embodiment; k is the environmental correction coefficient. Considering the influence of vehicle driving vibration, k is set to 1.0 when the vehicle is stationary, k = 1.0 + 0.01 * v when the vehicle speed v > 0 km / h, and k is set to the maximum value of 1.5 when v ≥ 50 km / h.
[0045] If σ xy <σ th If the condition is met, the system is considered to be in a "hovering" state, and step 303 is executed; otherwise, the system is not guided and monitoring continues.
[0046] Step 303: Calculate the curvature of the fingertip movement trajectory.
[0047] Specifically, this includes: calculating the total path length L of the fingertip projection point coordinates in the N frames.total With net displacement L net :
[0048]
[0049] Then, the tortuosity τ is calculated:
[0050] If τ > 1.5, the user is determined to be in a "hesitant" state, and step 4 is executed; otherwise, it is determined to be a purposeful click preparation, the guidance is not triggered, and silence is maintained.
[0051] The threshold τ was determined based on experimental data analysis of user click operation patterns in an in-vehicle environment. Experiments show that when users click icons with a clear intention, the τ value of the fingertip trajectory is mainly distributed between 1.1 and 1.3 (due to the uncertainty of micro-hand movements); when the τ value is consistently higher than 1.5, the behavior pattern deviates significantly from purposeful clicking and is more consistent with the repetitive exploration characteristics caused by confusion.
[0052] Step 4: Wake up the boot module and execute the boot process.
[0053] When step 3 determines that the user is in a hesitant state, the system proactively triggers the guidance module within 300ms of the determination being successful, without requiring any actual user clicks. The guidance is implemented by rendering an auxiliary interaction layer on top of the original user interface; this layer does not participate in the event dispatching of the original UI. The guidance content includes: a) Visually Emphasize Target Controls: Scale the target controls (such as icons and buttons) that are the focus of the user's gaze and hand proportionally, by a magnification factor of 1.20 times the original size, and add a highlighted outline. The outline width is 4 pixels, using a preset accent color (such as orange-red), and the outline transparency is 80%.
[0054] b) Soften non-target areas: Centered on the enlarged target control, visually soften the surrounding areas, for example, by reducing the opacity by 30%, to highlight the object of focus. The softening effect only applies to the display layer.
[0055] c) Display operation prompts: Display prompt bubbles or text descriptions near the target control (such as at the top or bottom, with an 18-pixel gap) to indicate the function or operation method of the control (e.g., "Click to open the air conditioner interface, long press to turn on the air conditioner").
[0056] The guided display lasts for 2.5 seconds. If the user performs a valid action (click, long press, or explicit swipe) during the guided display, the guided module will immediately terminate and disappear. If the user does not perform any action after the guided display times out, the guided module will automatically disappear, and the system will return to silent monitoring mode.
[0057] The following specific example illustrates the execution process of the method described in this embodiment.
[0058] The following conditions are set: the screen coordinate system is based on physical dimensions (millimeters), with the origin located at the lower left corner of the screen; the vehicle is stationary, the environmental correction factor k=1.0; the sampling window T=500ms, the sampling rate f=60Hz, and the corresponding most recent N=30 frames of data.
[0059] The system execution steps are as follows: Step 1: Calculate the eye-hand distance D in the current frame. ehc .
[0060] The data for the current frame (frame N=30) is: Coordinates of the line of sight: P gaze ( x gaze[N] , y gaze[N] = (100.0, 150.0) Finger tip projection point coordinates: P hand ( x hand[N] , y hand[N] = (105.5, 152.1)
[0061] Since 5.89 mm < 60 mm, the "eye-hand coordination" condition is met, proceed to the next step.
[0062] Step 2: Calculate the standard deviation of the fingertip coordinates s xy and with dynamic threshold s th Compare.
[0063] The coordinates (xi, yi) of the fingertip projection point in the most recent 30 frames are shown in Table 1 (unit: mm):
[0064] Table 1 Calculate the statistic: Mean value of X-coordinate:
[0065] Mean value of Y-coordinate:
[0066] Calculate the standard deviation s xy :
[0067] Dynamic threshold s th calculate: Preset physical radius R phy =5mm, k =1.0 s th = R phy × k =5.0mm Since 2.2 mm < 5.0 mm, the "hovering" condition is met, proceed to the next step.
[0068] Step 3: Calculate the path tortuosity τ.
[0069] Calculate the total path length L total (The sum of the distances between adjacent frames):
[0070] Based on the data in the table above, calculate: Frames 1-2:
[0071] Frames 2-3:
[0072] Frames 3-4:
[0073] ... Frames 28-29:
[0074] Frames 29-30:
[0075] Adding all 29 segments together, we get L total ≈9.8mm.
[0076] Calculate net displacement L net (The straight-line distance from frame 1 to frame 30):
[0077] Calculate tortuosity t :
[0078] because t ≈13.4>1.5, which matches the "hesitation" characteristic.
[0079] Step 4: Trigger the bootloader.
[0080] Once the system determines that the user is in a hesitant state, it will trigger the guidance module within 300ms.
[0081] Assuming the user's gaze and hand focus are on a fan icon, the system proportionally enlarges the fan icon (magnification ratio 1.20x) and adds a 4px wide, orange-red (80% opacity) stroke. Simultaneously, the visual impact of the adjacent areas to the left and right of the icon is reduced (opause decreased by 30%). A pop-up notification bubble appears above the icon with the text: "Click to open the air conditioner interface, long press to turn on the air conditioner." This guidance screen remains displayed for 2.5 seconds. If the user clicks or long-presses the icon during this time, the guidance disappears immediately; if no action is taken within the time limit, the guidance ends automatically.
[0082] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.
Claims
1. A human-computer interaction guidance method, characterized in that, include: Acquire the user's gaze and hand information, and convert them into coordinates in the same screen coordinate system; Based on the coordinates within a predetermined time window, calculate the eye-hand coordination distance, hand position stability, and tortuosity of the hand movement trajectory; Based on the judgment results of the eye-hand coordination distance, the stability of the hand position, and the tortuosity of the hand movement trajectory, it is identified whether the user is in a state of confusion and hesitation regarding the screen interface elements; When the confused and hesitant state is identified, interactive guidance information is actively triggered and displayed on the screen interface elements before the user performs an actual touch operation.
2. The method according to claim 1, characterized in that, The steps of acquiring the user's gaze and hand information and converting coordinates include: Capture the user's facial and hand images through a camera; The gaze direction is determined based on the facial image, and the gaze direction is mapped to the screen coordinate system through a pre-stored extrinsic parameter matrix from the camera coordinate system to the screen coordinate system to obtain the coordinates of the gaze point. Based on the hand image, the coordinates of key hand points are determined, and then transformed to the screen coordinate system using the extrinsic parameter matrix to obtain the three-dimensional coordinates of the fingertip in the screen coordinate system.
3. The method according to claim 2, characterized in that, The extrinsic parameter matrix is obtained and stored by calibration during system initialization. The calibration includes: displaying a preset calibration pattern on the screen, capturing images through a camera, and calculating the rotation matrix and translation vector of the camera relative to the screen based on the known physical coordinates of the pattern in the screen coordinate system and the pixel coordinates in the image to form the extrinsic parameter matrix.
4. The method according to claim 1, characterized in that, The judgment based on eye-hand coordination distance includes: Calculate the Euclidean distance between the coordinates of the gaze point in the current frame and the coordinates of the fingertip's projection point on the screen plane; If the Euclidean distance is less than the first threshold, then the eye-hand coordination condition is satisfied.
5. The method according to claim 4, characterized in that, The judgment based on hand position stability includes: Calculate the standard deviation of the fingertip projection point coordinates within the predetermined time window; The standard deviation is compared with a dynamic threshold. If the standard deviation is less than the dynamic threshold, the hand is determined to be in a stable hovering state. The dynamic threshold is determined based on a preset physical radius and a correction coefficient related to the environmental state.
6. The method according to claim 5, characterized in that, The dynamic threshold σ_th is calculated using the following formula: Among them, R phy The physical radius is a preset value, and k is an environmental correction coefficient. When the vehicle is stationary, k = 1.
0. When the vehicle is moving, k increases with the vehicle speed v, and there is an upper limit value.
7. The method according to claim 1, characterized in that, The judgment based on the curvature of the hand movement trajectory includes: Calculate the total path length L of the fingertip projection point movement within the predetermined time window. total ; Calculate the net displacement L between the start and end points within the window. net ; Calculate tortuosity ; If τ is greater than the second threshold, the hand movement trajectory is determined to conform to the hesitation feature.
8. The method according to any one of claims 1 to 7, characterized in that, The step of actively triggering and displaying interactive guidance information includes: within a predetermined delay time after recognizing the confused and hesitant state, displaying a guidance layer overlaid on the screen interface element, the guidance layer including a visual emphasis effect on the target interface element, a visual weakening effect on non-target areas, and text or graphic prompts for explaining functions or operation methods.
9. The method according to claim 8, characterized in that, The visual emphasis effect includes proportionally enlarging the target interface elements and adding highlighted outlines.
10. The method according to claim 8, characterized in that, The guide layer is displayed for a predetermined time. If a valid user interaction with the target interface element is detected during this period, the guide layer is turned off in advance.