A three-dimensional control method of eye-hand coordination in augmented reality

By combining eye-tracking and gesture-based collaborative interaction methods in augmented reality, and utilizing eye-hand collaborative interaction cues and real-time updated eye-tracking parameters, the efficiency and accuracy issues of 3D object manipulation in occluded environments are solved, achieving efficient target selection and object movement while reducing the burden on the user's hands.

CN116185194BActive Publication Date: 2026-03-17BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

In augmented reality environments, existing technologies struggle to efficiently manipulate 3D objects in partially or completely occluded conditions. In particular, eye-tracking interaction suffers from insufficient accuracy and efficiency, failing to meet the needs of real-world scenarios.

Method used

By collecting calibration sample datasets through nine-point explicit calibration of multiple users, initial general eye-tracking parameters are calculated. Implicit calibration is performed on new users to obtain the origin of pupil coordinates. Combined with eye-hand coordination interaction cues, implicit calibration sample points in the three-dimensional interactive space are collected. The eye-tracking estimation model is used to update the user's eye-tracking parameters in real time. Users control the indicator to select the target by pinching gestures to move over obstructions and control the movement of objects by adjusting and fixing the gaze direction.

Benefits of technology

It enables one-time selection of occluded targets in occluded environments, improving interaction efficiency, reducing user arm fatigue, and providing an efficient interactive experience and accurate object manipulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116185194B_ABST
    Figure CN116185194B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method for eye-hand coordination in three-dimensional control in augmented reality. A specific embodiment of the method comprises: calculating initial general eye movement parameters, and performing implicit calibration on a new user to obtain a user pupil coordinate origin and calculate a user gaze point; collecting implicit calibration sample points in a three-dimensional interaction space, and using the latest sample point set and the initial collected general calibration data set to substitute into an eye movement estimation model; a user controls an indicator to pass through an occlusion by a pinching gesture, and the user releases the gesture, thereby selecting an object within a certain range that is closest to the indicator; the user first adjusts a line of sight direction to align with a target position at which an object is to be placed by a pinching gesture, and after releasing the gesture, the line of sight direction is fixed, and then the user maintains the pinching gesture to control the object to move to the target position along the fixed line of sight direction and releases the gesture. The embodiment can effectively solve the occlusion problem, and does not require multiple repeated operations to move the occlusion away, thereby improving interaction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method for eye-hand coordinated three-dimensional manipulation in augmented reality. Background Technology

[0002] In recent years, with the continuous development of the augmented reality field, eye-tracking-based human-computer interaction has received increasing attention both domestically and internationally, and has excellent application prospects in various fields. Gaze, as a crucial clue revealing how humans understand their external environment, accounts for over 80% of the information received by the brain. The location of the gaze directly indicates the user's interests. Eye movements are fast, natural, and effortless, but they also have limitations. Eye-tracking interaction often suffers from low accuracy due to insufficient eye-tracking precision and the "Midas Touch" problem. Gestures, on the other hand, are a natural and refined interaction method, but require slightly more effort. The combination of eyes and hands is also the core of our daily interaction with surrounding objects. Combining these two interaction methods through complementarity can bring a more natural and accurate interactive experience.

[0003] Multimodal interaction combining eye-tracking and gestures in augmented reality (AR) environments has undergone significant development over the years and has been applied to various fields, with extensive exploration and research in diverse scenarios. However, most AR environments are uncomplicated, unobstructed, and open, with little discussion on how to efficiently manipulate 3D objects using eye-tracking cues in partially or even completely occluded environments. This still falls short of the needs of real-world scenarios. Therefore, solving this problem will provide new clues for the future development and application of eye-hand interaction technology in the AR field. Summary of the Invention

[0004] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0005] Some embodiments of this disclosure propose a method for eye-hand coordinated three-dimensional manipulation in augmented reality to solve one or more of the technical problems mentioned in the background section above.

[0006] Some embodiments of this disclosure propose an eye-hand coordinated 3D manipulation method in augmented reality. The method includes: performing a nine-point explicit calibration on multiple users to collect calibration sample datasets and calculate initial general eye-tracking parameters; performing implicit calibration on new users to obtain the user's pupil coordinate origin and calculate the user's gaze point; utilizing eye-hand coordinated interaction cues, collecting implicit calibration sample points in the 3D interactive space; using the latest sample point set and the initially collected general calibration dataset to input into the eye-tracking estimation model, and updating the user's eye-tracking parameters online in real time; the user controls an indicator to move past an obstruction by pinching a gesture, making the obstruction transparent and revealing the obstructed target object; releasing the gesture selects the object closest to the indicator within a certain range; the user first adjusts their gaze direction to align with the target position of the desired object by pinching a gesture, at which point the gaze direction is displayed as a white beam; after releasing the gesture, the gaze direction is fixed; then the user maintains the pinching gesture again to control the object to move along the fixed gaze direction to the target position, and releases the gesture.

[0007] Eye-hand coordination interaction strategies can effectively solve the occlusion problem, allowing users to select occluded objects in one go without having to repeatedly remove the obstruction, thus improving user interaction efficiency and delivering a highly efficient interactive experience. Furthermore, during object movement, the system ensures that the object moves to the designated position along a fixed line of sight, eliminating the need for constant hand control and effectively utilizing eye-tracking cues to reduce arm fatigue. Attached Figure Description

[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0009] Figure 1 This is a flowchart of some embodiments of the augmented reality eye-hand coordinated three-dimensional manipulation method according to the present disclosure;

[0010] Figure 2 This is a target object selection diagram for the target selection stage according to some embodiments of the augmented reality eye-hand coordinated three-dimensional manipulation method disclosed herein;

[0011] Figure 3 This is a schematic diagram of the target object movement during the object movement stage according to some embodiments of the augmented reality eye-hand cooperative three-dimensional manipulation method disclosed herein;

[0012] Figure 4 This is an interactive flowchart of some embodiments of the eye-hand coordinated three-dimensional manipulation method in augmented reality according to the present disclosure. Detailed Implementation

[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0014] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0015] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0018] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1 A flow 100 of some embodiments of an augmented reality eye-hand coordinated 3D manipulation method according to the present disclosure is shown. This augmented reality eye-hand coordinated 3D manipulation method includes the following steps:

[0020] Step 101 involves collecting calibration sample datasets from multiple users using nine-point explicit calibration, calculating initial general eye movement parameters, and performing implicit calibration on new users to obtain the origin of the user's pupil coordinates and calculate the user's fixation point.

[0021] Here, a universal calibration dataset was formed by collecting calibration samples from n users (n>=15) using a nine-point explicit calibration method. Each data point includes the user's pupil center coordinates and calibration point position. When each user looks at the center coordinates of the screen, the user's pupil coordinates at that moment are the origin, and the pupil center coordinates are the relative values ​​of the pupil coordinates when the user looks at any point on the screen with respect to that origin.

[0022] Implicit calibration for new users incorporates the center point of the screen as the calibration point into the user interaction.

[0023] Substitute the general dataset and implicit calibration points into the eye-tracking calibration parameter calculation model to obtain the parameter solution, and use the parameters to initialize the eye-tracking algorithm.

[0024] Step 102: Utilizing eye-hand coordination interaction cues, implicit calibration sample points are collected in the 3D interactive space. The latest sample point set and the initially collected general calibration dataset are then substituted into the eye-tracking estimation model to update the user's eye-tracking parameters online in real time. Here, a preset number of the latest three types of implicit calibration sample points can be collected and substituted into the eye-tracking estimation model along with the general dataset to update the eye-tracking calibration parameters online. The preset number can be a pre-defined quantity. For example, the preset number could be 60.

[0025] Step 103: The user controls the indicator to move past the obstruction by pinching the gesture, making the obstruction transparent and thus revealing the obscured target object. When the user releases the gesture, the object closest to the indicator within a certain range will be selected.

[0026] In this process, the gaze point estimated by the eye-tracking model is given a depth-based three-dimensional gaze point.

[0027] The gaze point is visualized as an indicator. When the user adjusts the position of the indicator, any obstruction located between the indicator and the user, that is, closer to the user than the indicator, becomes transparent, revealing the object that was previously obscured.

[0028] During the target selection phase, when the user releases the pinch gesture, the nearest object within a 0.3m radius of the indicator will be selected.

[0029] Step 104: The user first adjusts the direction of their gaze to align with the target position where they want to place the object by pinching the gesture. At this time, the direction of the gaze is displayed in the form of a white beam. After releasing the gesture, the direction of the gaze is fixed. Then, the user holds the pinch gesture again to control the object to move along the fixed direction of the gaze to the target position and releases the gesture.

[0030] In the first instance, when the user holds the pinch gesture and adjusts the direction of their gaze, the object is bound to the white beam of vision. The object will also move when the user adjusts the direction of their gaze, but the depth of the object remains unchanged.

[0031] Once the user releases the pinch gesture for the first time, the direction of their gaze remains fixed and no longer changes with the movement of their eyes.

[0032] When the user holds the pinch gesture for the second time, the object will move further or closer to the user in a fixed line of sight as the hand moves further or closer. The direction of the object's movement is determined by the line of sight, and the depth of the object is controlled by the distance between the hand and the user, without requiring the specific direction of the hand's movement.

[0033] The implementation method mainly consists of the following steps:

[0034] The first step is eye-tracking algorithm initialization. Two infrared cameras installed on the Microsoft HoloLens 2 device are used to capture near-eye images of the user. First, a nine-point explicit calibration is performed on n users (n>=15), acquiring the coordinates of the left and right pupil centers {x, y} when the user looks at the corresponding screen coordinate point (X, Y). p y p When a user is detected looking towards the center of the screen, the pupil coordinates at that moment are taken as the origin. The pupil center coordinates are then defined as the relative values ​​of the pupil coordinates when the user is looking at any point on the screen and that origin. For each user, excluding the origin, there are 8 pairs of left and right pupil center coordinates corresponding to 8 calibration points excluding the center of the screen. The entire dataset contains 8n pairs of data. The data is then substituted into the following polynomial eye-tracking estimation model for calculation:

[0035]

[0036]

[0037] Among them, a i and b i The unknown coefficients can be obtained by solving the above system of polynomials, which yields a set of a. i and b i The result, as a general eye-tracking parameter, is that p represents the number of pupil center coordinates, {x p y p} represents the coordinates of the center of the p-th pupil.

[0038] Then, for new users who need to interact, implicit calibration is used to hide the calibration point in the center of the screen into the eye-hand interaction that is not easily noticed by the user. For example, a button that is always located in the middle of the screen. When the user clicks the button in the middle of the screen, they will often look at the button first. The average pupil coordinates in the 0.1 seconds before the user clicks the button can be collected as the new pupil coordinate origin. This completely avoids the standard calibration process and reduces the time spent.

[0039] Universal eye-tracking parameters and the user's individual pupil coordinate origin can be used to initialize the eye-tracking algorithm, providing users with a relatively high level of eye-tracking accuracy. During interaction, for each frame of near-eye image from the user, the pupil center coordinates {x} are obtained using the user's new pupil coordinate origin. p y p}, and the general parameter a i b i As is also known, substituting into the above formula yields the screen point (X, Y) that the user is looking at in the current frame.

[0040] The second step involves 3D object manipulation using eye tracking and gesture coordination. The entire interactive operation is divided into two stages: the target selection stage and the object movement stage.

[0041] Target selection phase: The user first looks at the target object, then pinches their thumb and forefinger together. Once the pinch gesture is detected during the selection phase, a small red circle indicator appears at the user's gaze point at a fixed distance (2m). As the user maintains the pinch gesture, the red indicator moves in tandem with the hand. The user can then remotely adjust the indicator in the 3D interactive space to skip obstructions and directly select the occluded target. Once the pinch gesture is released, the nearest object within a 0.3m radius of the indicator is selected. As the user moves their hand forward, the indicator depth changes, and occluding objects located between the indicator and the user's eyes (closer to the user than the indicator) become transparent, allowing the user to observe the occluded object. The target object selection phase diagram of some embodiments of the eye-hand coordinated 3D manipulation method in augmented reality disclosed herein is shown below. Figure 2 As shown.

[0042] Object Movement Phase: Object movement is also divided into two steps, each ending with the release of the pinch gesture. Once the target is selected, it is bound to the user's line of sight. The user first looks at the desired location and then makes a pinch gesture. At this point, the object appears at the location the user was just looking at, and the user's current line of sight is represented by a white beam connecting the user's eyes and the object. The user maintains the pinch gesture and adjusts their line of sight (the direction of the white beam) to align with the target location. Then, the pinch gesture is released. The user's line of sight (the direction of the white beam) is now fixed, and the first sub-step ends. The user then uses the pinch gesture again to enter the second sub-step. This time, the object's position is bound to the user's hand. The user maintains the pinch gesture and moves the object closer or further away along the fixed line of sight by moving their hand away or closer. When the object reaches the accurate position, the user releases the gesture, ending the second sub-step and the entire operation phase. In the second sub-step of eye-guided interaction, since the gaze direction is fixed, the object's movement always follows the direction of the gaze beam. The object's depth is controlled only by the distance between the pinch gesture and the user, without considering the specific direction of hand movement. This reduces hand fatigue, especially when the target location is high and far away. A schematic diagram of the target object movement stage in some embodiments of the augmented reality eye-hand coordinated 3D manipulation method disclosed herein is shown below. Figure 3 As shown.

[0043] The third step is online implicit calibration updates. Implicit calibration points collected during the user's eye-hand interaction are used to continuously update the eye-tracking calibration online, thereby improving accuracy. Specifically, implicit calibration samples are collected based on the following assumptions: First, in the target selection phase, after the user adjusts the indicator, it is assumed that the user is looking at the indicator, thus confirming that the indicator is in the correct position; second, in the first sub-step of the object movement phase, it is assumed that the user is looking at the center of the object, so that the user knows that the gaze beam is aligned with the target position; third, in the second sub-step of the object movement phase, when the user completes the object translation by releasing the pinch gesture, the user is assumed to be looking at the object, thus confirming that the object has reached the correct position. Therefore, three samples are collected each time the user selects and moves an object. Using 60 up-to-date sample calibration points and an initial 8n pairs of data points from the general parameter set, the individual eye-tracking parameters are updated and optimized using a multinomial eye-tracking estimation model. Therefore, as user interaction time increases, the accuracy of the multinomial eye-tracking estimation model will become increasingly higher, and when the device slides, the individual's eye-tracking parameters will also be updated quickly, without affecting the interactive experience. The interaction flowcharts of some embodiments of the augmented reality eye-hand coordinated 3D manipulation method disclosed herein are as follows: Figure 4 As shown.

[0044] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for eye-hand coordination in three-dimensional augmented reality, comprising: collecting a general calibration dataset by nine-point explicit calibration for a plurality of users, calculating initial general eye movement parameters, and performing implicit calibration for a new user to obtain a user pupil coordinate origin, and calculating a user gaze point; using eye-hand coordination interaction cues, collecting implicit calibration sample points in a three-dimensional interaction space, using the implicit calibration sample point set and the initial collected general calibration dataset to substitute into an eye movement estimation model, and updating the user's eye movement parameters in real time, wherein the eye movement estimation model is represented by the following polynomial eye movement estimation model: , , wherein, and represent unknown coefficients, which can be obtained by solving the above polynomial set to obtain a set of and as initial generic eye movement parameters, represents the number of pupil center coordinates, represents the th pupil center coordinate; a user controls an indicator to pass through an occlusion by a pinch gesture, making the occlusion transparent to display the occluded target object, and the user releases the gesture to select the object closest to the indicator within a certain range, wherein the indicator is a visualized gaze point; a user first adjusts the line of sight direction to align with the target position of the object to be placed by a pinch gesture, at which time the line of sight direction is displayed in the form of a white beam, and after releasing the gesture, the line of sight direction is fixed, and then the user controls the object to move to the target position along the fixed line of sight direction by maintaining the pinch gesture, and releases the gesture; wherein the collecting a general calibration dataset by nine-point explicit calibration for a plurality of users, calculating initial general eye movement parameters, and performing implicit calibration for a new user to obtain a user pupil coordinate origin, and calculating a user gaze point, comprises: collecting calibration sample data sets for each user by a nine-point explicit calibration method for n users, forming a general calibration dataset, each general calibration data including a user's pupil center coordinate and a calibration point position, when each user looks at the center point of the screen, the pupil coordinate of the user at this time is the coordinate origin, and the pupil center coordinate is the relative value of the pupil coordinate when the user looks at any point on the screen to the origin; implicit calibration for a new user is to introduce the center point of the screen as a calibration point into user interaction; substituting the general calibration dataset and the calibration point into the eye movement estimation model to obtain a parameter solution, and initializing the eye movement estimation model with the parameter; for each frame of near-eye image of the user in the interaction, the user's own user pupil coordinate origin is used to obtain the pupil center coordinate; inputting the pupil center coordinate into the initialized eye movement estimation model to generate a user gaze point (X, Y).

2. The method of claim 1, wherein, the using eye-hand coordination interaction cues, collecting implicit calibration sample points in a three-dimensional interaction space, using the implicit calibration sample point set and the initial collected general calibration dataset to substitute into an eye movement estimation model, and updating the user's eye movement parameters in real time, comprises: collecting a preset number of latest three implicit calibration sample points, and substituting them into the eye movement estimation model together with the general calibration dataset to update the eye movement calibration parameters online.

3. The method of claim 1, wherein, the user controls an indicator to pass through an occlusion by a pinch gesture, making the occlusion transparent to display the occluded target object, and the user releases the gesture to select the object closest to the indicator within a certain range, wherein the indicator is a visualized gaze point; the gaze point estimated by the eye movement estimation model is assigned a depth to become a three-dimensional gaze point; The gaze point is materialized in the form of a pointer, when the user adjusts the position of the pointer, the occluder between the pointer and the user, i.e. closer to the user than the pointer, will be transparentized, leaking the objects behind that are occluded; In the target selection phase, when the user releases the pinch gesture, the closest object within a 0.3m radius from the pointer will be selected.

4. The method of claim 1, wherein, The user first adjusts the gaze direction to align with the target position of the object to be placed by a pinch gesture, at this time the gaze direction is presented in the form of a white beam, after releasing the gesture, the gaze direction is fixed, then the user maintains the pinch gesture again to control the object to move to the target position along the fixed gaze direction, and releases the gesture, including: When the user adjusts the gaze direction for the first time, the selected object is bound to the white gaze beam, when the user adjusts the gaze direction, the selected object will also move, but the depth of the selected object remains unchanged; After the user releases the first pinch gesture, the gaze direction is fixed and no longer changes with the movement of the user's eyes; When the user maintains the pinch gesture for the second time, the selected object will move away or move close to the fixed gaze direction as the hand moves away or moves close to the user, the movement direction of the selected object is determined by the gaze direction, and the depth of the object is controlled by the distance between the hand and the user, without the specific movement direction of the hand.

Citation Information

Patent Citations

  • Target tracking mechanical arm method, system and equipment and storage medium

    CN114406985A

  • Methods and systems for selection of objects

    US20220198756A1