Eye-tracking-driven collaborative robotic arm teleoperation control system and method

By using eye-tracking feature extraction to extract action intent and a shared control system for robotic arms, the problems of low efficiency in collaborative robotic arm control and eye saccade interference have been solved, enabling natural and convenient robotic arm control and improving the independent living ability of people with upper limb dysfunction.

CN119871409BActive Publication Date: 2025-11-14SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510115199.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-11-14
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

Existing collaborative robotic arms are inefficient, difficult to operate, and have a heavy cognitive load. Furthermore, in the case of eye-tracking robotic arms, involuntary eye saccades can cause uncontrollable movement of the mechanical grippers, affecting the daily assistive tasks of people with disabilities.

Method used

An action intent recognition and robotic arm sharing control system based on eye movement feature extraction is adopted, including a binocular vision eye tracker, a global camera, a collaborative robotic arm, and a tactile force feedback device. The action intent recognition model is trained by a Gaussian mixture model to achieve real-time correlation between eye movement features and robotic arm movements, reducing eye saccade interference.

Benefits of technology

It enables natural and convenient robotic arm control, improves control efficiency and convenience, reduces uncontrollable movement of the mechanical gripper, and enhances the ability of people with upper limb dysfunction to independently complete daily activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119871409B_ABST
    Figure CN119871409B_ABST
Patent Text Reader

Abstract

This invention discloses a collaborative robotic arm teleoperation control system and method driven by eye-tracking features. The method includes: calibrating an eye tracker and a global camera; constructing and training a motion intention recognition model; establishing a relationship model between the input and output of the motion intention recognition model using eye-tracking features as data input and motion intention as data output; performing online recognition and prediction of motion intentions based on eye-tracking features, and extracting the recognized motion intentions; and commanding the robotic arm to perform actions based on the recognized motion intentions. This method allows operators to control the robotic arm in a natural and intuitive way, without frequently switching control modes, and effectively avoids uncontrollable movements of the mechanical gripper caused by involuntary eye saccades. Simultaneously, this method improves the detailed posture control of the mechanical gripper during grasping and placement, which enhances the efficiency and convenience of operation for people with disabilities, especially those with upper limb dysfunction, thereby strengthening their ability to independently complete daily activities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control and human-computer interaction technology, and particularly relates to a collaborative robotic arm teleoperation control system and method driven by eye-tracking features. Background Technology

[0002] People with upper limb dysfunction, such as those suffering from spinal cord injury, stroke, multiple sclerosis, amyotrophic lateral sclerosis, Duchenne muscular dystrophy, etc., have difficulty controlling their upper limbs to complete daily living activities (such as grasping and placing objects). Using collaborative robotic arms to assist people in completing daily living activities can improve the independent living ability of these special groups.

[0003] Currently, collaborative robotic arms are primarily controlled via joysticks or teach pendants. This process requires the operator to frequently switch between translation, rotation, and gripper opening / closing modes in the Cartesian space of the mechanical gripper, resulting in low efficiency, high difficulty, and heavy cognitive load. In recent years, eye-tracking technology, as a non-invasive control method, has seen initial applications in intelligent robotics and medical assistance. It provides users with a more natural and intuitive interaction method. By controlling the movement of the robotic arm and the opening / closing of the mechanical gripper through eye-tracking signals, the operator can significantly improve the control efficiency of collaborative robotic arms. The core of eye-tracking control of robotic arm movement lies in the control strategy and control algorithm.

[0004] Universities such as UCLA and Imperial College London in the US, and Southeast University in China, have already implemented eye-tracking signals for robotic arm control. The operator gazes at the target object, and when the gaze duration exceeds a certain threshold, a pre-programmed robot action is triggered. However, this method is not true eye-tracking-driven robotic arm assistance; the robotic arm's movement relies entirely on pre-programmed movements, is only effective in structured environments, and the operator lacks a strong sense of control during operation, making it unsuitable for free-intention tasks.

[0005] A major limitation of eye-tracking-driven robotic arm teleoperation technology is that operators cannot voluntarily control their eye movements to ensure that their gaze steadily guides the gripper's movement. Eye movements involve numerous involuntary saccades (the gaze point jumps from one location to another). For example, during teleoperation, the gaze frequently saccades between the target object and the gripper. If the gripper is simply allowed to follow the point of intersection of gaze, it will exhibit uncontrollable movements following these involuntary saccades, which will have a significant negative impact on the daily assistive tasks of people with disabilities.

[0006] Therefore, there is an urgent need for a robotic arm teleoperation control system and method that can achieve natural, convenient and intuitive operation, perfect detailed posture control, and avoid eye sac interference. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a collaborative robotic arm teleoperation control system and method driven by eye-tracking features.

[0008] The technical solution provided by this invention is as follows:

[0009] Firstly, it provides a motion intent recognition and robotic arm sharing control system based on eye-tracking feature extraction, including:

[0010] A binocular vision eye tracker, comprising a scene camera, a pupil-capturing camera, and a wireless module, is used to capture the three-dimensional coordinates of the gaze focus in the scene camera and transmit the data.

[0011] A QR code positioning board is connected to a binocular vision eye tracker;

[0012] A global camera system, including a regular camera and a depth camera, is used to capture videos of a positioning board with a QR code.

[0013] Collaborative robotic arms are used for grasping, moving, and placing objects.

[0014] Tactile force feedback devices are used to collect tactile information from operators;

[0015] The host computer connects to the global camera, collaborative robotic arm, and haptic force feedback device; it is used to analyze and process the data, identify the operator's intentions online and convert them into control commands for the robotic arm, which are then sent to the collaborative robotic arm to execute the actions.

[0016] Secondly, a method for action intent recognition and robotic arm shared control based on eye-tracking feature extraction is provided. This method is implemented based on the eye-tracking feature extraction-based action intent recognition and robotic arm shared control system described in the first aspect, and includes the following steps:

[0017] Eye tracker and global camera calibration;

[0018] Construct and train a motion intention recognition model; the motion intention recognition model uses eye movement features as data input and motion intention as data output, and establishes a relationship model between data input and output;

[0019] Based on eye movement features, online recognition and prediction of action intent are performed, and action intent recognition is extracted.

[0020] The robotic arm is instructed to perform actions based on the intent of the action.

[0021] In one possible implementation, both the scene camera and the global camera of the eye tracker are calibrated using a chessboard.

[0022] In one possible implementation, the method for constructing and training the action intent recognition model is as follows:

[0023] Operators remotely control the collaborative robotic arm through tactile force feedback devices to perform actions such as grasping, moving, and placing objects;

[0024] A training dataset is obtained, including eye movement feature data captured by a binocular vision eye tracker and real-time motion data of a collaborative robotic arm collected by a host computer; the eye movement feature data includes gaze focus position, gaze focus movement speed, gaze vector angular velocity, gaze duration, saccade amplitude and duration;

[0025] The training dataset is divided into a large-range movement dataset and a fine-tuning dataset. The large-range movement dataset is the dataset in which the robotic arm completes a relatively large displacement movement. The fine-tuning dataset is the dataset in which the robotic arm has completed a relatively large displacement movement and is close to the pre-grasping position or pre-placement position, and then performs small-range precise position and posture adjustments.

[0026] Using eye-tracking feature data as input and real-time motion data as output, a Gaussian mixture model is used as the basic model for fitting and training, resulting in action intention recognition models based on large-scale movement datasets and fine-tuning datasets, respectively.

[0027] Furthermore, the method for calculating the focal point position of the line of sight is as follows:

[0028] The two-dimensional image coordinates of the QR code vertices are obtained by using image recognition and corner detection to capture the image of the QR code positioning board captured by the global camera.

[0029] Calculate the pose of the QR code positioning plate {B} relative to the global camera {C} C T B ;

[0030] Based on the position transformation matrix between the eye tracker scene camera {S} and the QR code calibration board {B} B T S Calculate the pose of the scene camera relative to the global camera using an eye tracker. C T S ;

[0031] Based on the transformation matrix between the global camera {C} and the robotic arm's base coordinate system {R} R T C Calculate the pose of the eye tracker scene camera relative to the global coordinate system. R T S And the direction of the three-dimensional line-of-sight vector and the coordinates of the line-of-sight focus in the global coordinate system are obtained.

[0032] Furthermore, in the action intent recognition model, the input includes: state variable x i =(τ i ,δ i )∈R N , where τ i =(τ x,i ,τ y,i ,τ z,i )∈R N δ represents the coordinate position of the end effector gripper of the collaborative robotic arm within data i in the total data volume N during the collection process. i Euler angle parameters of the end effector gripper of the collaborative robotic arm within data i representing the total data volume N; time observation sequence o i =(g i ,ε i )∈R N , where g i =(g x,i ,g y,i ,g z,i ) represents the focal position of the operator's eye in the global coordinate system within data i, and ε represents the focal position of the operator's eye within the data i. i This represents eye movement patterns involving fixation and gaze shifting;

[0033] By using the joint probability distributions P1(o,x) and P2(o,x) of time observation sequences and state variables in the large-scale mobile dataset and the fine-tuned dataset, the conditional probability distributions P1(x|o) and P2(x|o) based on the large-scale mobile dataset and the fine-tuned dataset are solved by Gaussian conditions respectively.

[0034] The Gaussian mixture model for a large-scale mobile dataset is represented as follows:

[0035]

[0036] Where, θ k It is the parameter set of the Gaussian mixture model, including the weights ω of the k-th Gaussian distribution. k Mean vector μ k The sum of the covariance matrix ∑ k Each component It is a multivariate Gaussian distribution, where K is the number of Gaussian distributions;

[0037] Given eye-tracking data o i Robotic arm end effector data x i The probability under the given conditions is:

[0038]

[0039] Due to the observed variable o i If it is a continuous variable, then the joint distribution probability can be calculated to calculate the marginal probability P1(o).i |θ k ),for:

[0040] P1(o i |θ k )=∫P1(o i ,x i |θ k )dx i

[0041] During the training process, for θ k Each Gaussian component is randomly initialized. The Expectation-Maximization (EM) algorithm is used to calculate the posterior probability of each sample belonging to a Gaussian component. The parameters of each Gaussian component are updated using the posterior probability. This process is repeated iteratively until convergence, thus obtaining the observation conditions at a given real-time t under the range-shifting mode. t The state variable x under ' t ′;

[0042] Similarly, the observation conditions o″ at a given real-time time t in the precise adjustment mode are obtained. t The lower state variable x″ t .

[0043] Furthermore, the method for online recognition and prediction of action intent based on eye movement features, and for extracting action intent recognition, includes the following steps:

[0044] Collect operator eye movement data and update time observation sequence in real time. t ;

[0045] The conditional distribution probability P1(x) is calculated based on a large-scale mobile dataset and a finely adjusted dataset. t |o t ) and P2(x t |o t The system determines which of two datasets to use based on a conditional threshold. It then determines the relationship between the distance between the mechanical gripper's position and the intended object to be grasped, and a distance threshold. If the distance is greater than the threshold, the eye-tracking signal is input into a conditional distribution model trained on a large-scale movement dataset to calculate P1(x). t If the distance is less than a distance threshold, the eye-tracking signal is input into a conditional distribution model trained on a finely adjusted dataset to calculate P2(x). t );

[0046] After capture, the relationship between the variance of the eye-tracking fixation focus within the window and the fixation threshold is continuously calculated; if the variance is greater than the fixation threshold, the eye-tracking signal is input into P1(x) trained based on a large-scale motion dataset. t If the variance is less than the fixation threshold, then input P2(x) trained on the fine-tuned dataset. tBased on the maximum a posteriori probability principle, the predicted values ​​x of the robot arm's end effector speed and end effector posture at time t are inferred from the host computer's prediction of the operator's intention. t .

[0047] Thirdly, it provides an eye-tracking feature extraction-based action intent recognition and robotic arm shared control device, including:

[0048] Calibration module for eye tracker and global camera calibration;

[0049] The model building and training module is used to build and train the action intention recognition model; the action intention recognition model uses eye movement features as data input and action intention as data output, and establishes a relationship model between data input and output.

[0050] The extraction module is used to perform online recognition and prediction of action intent based on eye movement features, and to extract action intent recognition.

[0051] The execution module is used to identify commands to the robotic arm to perform actions based on the action intent.

[0052] Fourthly, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the aforementioned method for action intent recognition based on eye-tracking feature extraction and shared control of a robotic arm.

[0053] Fifthly, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the aforementioned method for action intent recognition and robotic arm shared control based on eye-tracking feature extraction.

[0054] The present invention has the following beneficial effects:

[0055] (1) The method provided by this invention proposes a classification of “large-scale movement dataset” and fine adjustment dataset to address the differences in actions such as moving, placing and grasping. This satisfies the high efficiency of the robotic arm’s range of movement and can adapt to the shape of the object and environmental features during grasping and placing. At the same time, the RGB-D depth camera can provide more accurate environmental information and provides a conditional threshold for switching between the “large-scale movement dataset” and the fine adjustment dataset, reducing the complexity of operation for the operator.

[0056] (2) The method provided by this invention distinguishes and trains a model to associate eye-tracking features with the robotic arm's motion intentions, achieving real-time recognition of motion intentions. This eye-tracking-based motion intention recognition method allows operators to control the robotic arm in a natural and intuitive way, without frequently switching control modes, and effectively avoids uncontrollable movements of the mechanical gripper caused by involuntary eye saccades. Simultaneously, this method improves the detailed posture control of the mechanical gripper during grasping and placing, which enhances the efficiency and convenience of operation for people with disabilities, especially those with upper limb dysfunction, thereby strengthening their ability to independently complete daily activities. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of the system composition of the action intention recognition and robot shared control method based on eye movement feature extraction provided in the embodiment of the present invention; in the figure: 1-QR code positioning plate {B}, 2-pupil capture camera, 3-scene camera {S}, 4-RGB sensor, 5-depth sensor, 6-global camera {C}, 7-mechanical gripper, 8-robotic arm {R}, 9-host computer;

[0058] Figure 2 This is a flowchart of the action intent recognition method provided in an embodiment of the present invention;

[0059] Figure 3 This is a flowchart of the shared control system for a robotic arm driven by intent recognition, provided in an embodiment of the present invention.

[0060] Figure 4 This is a flowchart of the teleoperation collection process for the action intent recognition model dataset provided in this embodiment of the invention;

[0061] Figure 5 This is a schematic diagram of the teleoperation "move-grab-place" process provided in an embodiment of the present invention;

[0062] Figure 6 This is a positional diagram of the teleoperation "move-grab-place" process provided in an embodiment of the present invention;

[0063] Figure 7 This is a schematic diagram of the Euler angle definition of the control lever of the tactile force feedback device provided in an embodiment of the present invention;

[0064] Figure 8 This is a schematic diagram of the structure of the action intent recognition and robot shared control device based on eye movement feature extraction provided by the present invention;

[0065] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0066] The present invention will now be described in detail with reference to the accompanying drawings. It should be noted that the described embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0067] like Figure 1 As shown, this is a motion intent recognition and robot sharing control system based on eye-tracking feature extraction, including:

[0068] A binocular vision eye tracker, comprising a scene camera, a pupil-capturing camera, and a wireless module, is used to capture the three-dimensional coordinates of the gaze focus in the scene camera and transmit the data.

[0069] A QR code positioning board is connected to a binocular vision eye tracker;

[0070] A global camera system, including a regular camera and a depth camera, is used to capture videos of a positioning board with a QR code.

[0071] Collaborative robotic arms are used for grasping, moving, and placing objects.

[0072] Tactile force feedback devices are used to collect tactile information from operators;

[0073] The host computer connects to the global camera, collaborative robotic arm, and haptic force feedback device; it is used to analyze and process the data, identify the operator's intentions online and convert them into control commands for the robotic arm, which are then sent to the collaborative robotic arm to execute the actions.

[0074] Specifically, the binocular vision eye tracker is the Tobii Glass 3 or a similar product;

[0075] Specifically, the global camera is a Kinect Azure depth camera or a similar product;

[0076] Specifically, the tactile force feedback is provided by Geomagic Touch or a similar product.

[0077] In one possible implementation, the wireless module is a WIFI module that transmits eye-tracking data to a host computer in real time via the TCP / IP protocol.

[0078] like Figure 2 The image shows a method for action intent recognition and robot shared control based on eye-tracking feature extraction, which includes the following steps:

[0079] Step S100: Eye tracker and global camera calibration.

[0080] In one possible implementation, the eye tracker calibration method in step S100 is as follows:

[0081] Pupil-capturing camera calibration: Before actual operation, the operator wears an eye tracker and stares at the target on a calibration board or calibration card at a distance of half a meter to one meter. The eye tracker's software simultaneously performs relevant calculations for calibration. After calibration, the eye tracker can calculate the three-dimensional coordinates of the gaze focus in the eye tracker's scene camera coordinate system based on the pupil and corneal reflection positions detected by the pupil-capturing camera.

[0082] Scene camera calibration: A chessboard calibration was performed on the eye-tracking scene camera. More than twenty photos of the chessboard calibration board at different positions and poses were taken using the eye-tracking scene camera. These photos were then imported into MATLAB to calculate the position transformation matrix between the eye-tracking scene camera {S} and the QR code calibration board {B}. B T S This allows the operator to calculate the position of the line of sight focus in the global coordinate system (robotic arm base coordinate system) by using the pose of the calibration board during actual shared control operations.

[0083] It should be noted that the global coordinate system refers to the coordinate system of the space in which the collaborative robotic arm is located.

[0084] Step S200: Construct and train a motion intention recognition model; the motion intention recognition model uses eye movement features as data input and motion intention as data output, and establishes a relationship model between data input and output.

[0085] In one possible implementation, step S200 includes the following steps:

[0086] S210: The operator remotely operates the collaborative robotic arm through a tactile force feedback device to complete the grasping, moving and placing of objects;

[0087] S220, Obtain the training dataset, including eye movement feature data captured by the binocular vision eye tracker and real-time motion data of the collaborative robotic arm collected by the host computer; the eye movement feature data includes gaze focus position, gaze focus movement speed, gaze vector angular velocity, gaze duration, saccade amplitude and duration;

[0088] S230 divides the training dataset into a "large-scale moving dataset" and a "fine-tuning dataset";

[0089] S240 takes eye-tracking feature data as input and real-time motion data as output, and uses a Gaussian mixture model as the basic model for fitting and training to obtain action intention recognition models based on "large-scale movement dataset" and "fine-tuning dataset" respectively.

[0090] Furthermore, in step S220, the method for calculating the position of the line of sight focus is as follows:

[0091] The two-dimensional image coordinates of the QR code vertices are obtained by using image recognition and corner detection to capture the image of the QR code positioning board captured by the global camera.

[0092] Calculate the pose of the QR code positioning plate {B} relative to the global camera {C} C T B ;

[0093] Based on the position transformation matrix between the eye tracker scene camera {S} and the QR code calibration board {B} B T S Calculate the pose of the scene camera relative to the global camera using an eye tracker. C T S ;

[0094] Based on the transformation matrix between the global camera {C} and the robotic arm's base coordinate system {R} R T C Calculate the pose of the eye tracker scene camera relative to the global coordinate system. R T S And the direction of the three-dimensional line-of-sight vector and the coordinates of the line-of-sight focus in the global coordinate system are obtained.

[0095] Furthermore, in step S220, the calculation methods for the velocity of the gaze focus movement, the angular velocity of the gaze vector, the gaze duration, the saccade amplitude, and the duration are as follows:

[0096] The speed of the gaze focus movement can be obtained by dividing the distance between the coordinates of the gaze focus in two consecutive frames by the time difference between the two frames; the angular velocity of the gaze vector can be obtained by dividing the angle between the gaze vectors in two consecutive frames by the time difference between the two frames; the fixation duration is the length of a single fixation, the saccade amplitude is the angle swept by the gaze vector during a single gaze saccade, and the saccade duration is the length of a single gaze saccade. These eye movement features are directly provided in real time by the eye tracker.

[0097] Furthermore, in the action intent recognition model, the input includes: (1) state variable x i =(τ i ,δ i )∈R N , τ i =(τ x,i ,τ y,i ,τ z,i )∈R N This represents the coordinate position of the end effector gripper of the collaborative robotic arm within data i in the total data volume N collected during the data collection process. θ represents the Euler angle parameter of the end effector gripper of the collaborative robotic arm within data i in the total data volume N. i Represents pitch angle, Represents the roll angle, β iRepresents the yaw angle; since this invention uses a gripping method with a horizontal posture commonly found in robotic arm grippers, the pitch and roll angles of the end effector are fixed, and the yaw angle, θ, is released. i , The size is a suitable constant, β i As variables, discard pitch and roll angles, retain yaw angle, and take δ. i =β i (2) Time observation sequence i =(g i ,ε i )∈R N g t =(g x,i ,g y,i ,g z,i ) represents the focal position of the operator's eye in the global coordinate system within data i, and ε represents the focal position of the operator's eye within the data i. i This represents two eye movement patterns: fixation and saccade. Fixation refers to the eye movement pattern in which the eyes remain at a specific point, while saccade refers to the eye movement pattern in which the eyeballs move rapidly from one fixation point to another. The two are distinguished by the angular velocity of the line of sight vector: a line of sight vector angular velocity less than 30° / s indicates fixation, while a line of sight vector angular velocity greater than 30° / s indicates saccade.

[0098] After collecting the training dataset for the action intent recognition model, the training dataset was divided into two parts by labeling through replaying the experiment: a large-scale movement dataset and a fine-tuning dataset. Using the joint probability distributions P1(o,x) and P2(o,x) of the time observation sequences and state variables in the large-scale movement dataset and the fine-tuning dataset, the conditional probability distributions P1(x|o) and P2(x|o) obtained from training on the large-scale movement dataset and the fine-tuning dataset were solved using Gaussian conditions.

[0099] The Gaussian mixture model calculated based on a large-scale mobile dataset is represented as follows:

[0100]

[0101] Where, θ k It is the parameter set of the Gaussian mixture model, including the weights ω of the k-th Gaussian distribution. k Mean vector μ k The sum of the covariance matrix ∑ k Each component It is a multivariate Gaussian distribution, where K is the number of Gaussian distributions;

[0102] Given a time observation sequence o i The end effector state variable x of the robotic arm i The probability under the given conditions is:

[0103]

[0104] Due to the time observation sequence o i If it is a continuous variable, then the joint distribution probability can be calculated to calculate the marginal probability P1(o). i |θ k ),for:

[0105] P1(o i |θ k )=∫P1(o i ,x i |θ k )dx i

[0106] During the training process, for θ k Each Gaussian component is randomly initialized, and the Exceptional-Maximization (EM) algorithm is used to calculate the posterior probability of each sample belonging to a Gaussian component. The parameters of each Gaussian component are updated using the posterior probability. This process is repeated iteratively until convergence, obtaining the temporal sequence o′ at a given real-time time t under a conditional distribution model trained on a large-scale mobile dataset. t State variable x′ listed below t Similarly, we obtain the time observation sequence o″ at a given real-time time t under a conditional distribution model trained on a precise mobile dataset. t The lower state variable x″ t .

[0107] Specifically, the steps of the EM algorithm are as follows:

[0108] The E-step (Expectation Step) involves taking each time observation sequence for each observed variable. i and each state point x in each state variable i Calculate the posterior probability γ of each Gaussian distribution k. ik :

[0109]

[0110] The M-step (Maximization Step) updates the parameters of each Gaussian distribution, including the weights ω. k Mean vector μ k The sum of the covariance matrix ∑ k :

[0111]

[0112] Repeat the E-step and M-step iteratively until the parameters converge.

[0113] In the above Gaussian mixture model, each Gaussian distribution Specifically, it is expressed as follows:

[0114]

[0115] Where X is the joint vector, D refers to the dimension of the vector, and T is the transpose matrix. In this embodiment, D is set to 8.

[0116] Based on the conditional distribution probability P1(x) i |o i ,θ k The maximum a posteriori probability (MAP) criterion is used to infer the state variable x. t :

[0117] x t =argmax(P1(x i |o i ,θ k ))

[0118] This allows us to obtain the observation condition o′ at a given real-time time t under the conditional distribution model trained on a large-scale mobile dataset. t The lower state variable x′ t Furthermore, corresponding controls are made according to requirements, and the same applies to conditional distribution models trained based on precisely adjusted datasets.

[0119] Furthermore, in step S220, the method for obtaining the training dataset includes the following steps:

[0120] (1) The operator remotely operates the robotic arm through a tactile force feedback device to complete actions such as grasping, moving and placing objects.

[0121] (2) The host computer records the real-time motion data of the robotic arm at a sampling frequency of 50Hz, including the end-effector movement speed (including speed direction and amplitude), end-effector posture (Euler angle), and gripper opening and closing status. In addition, during the remote operation, the operator wears a binocular vision eye tracker throughout the process, and the host computer records the real-time eye movement data synchronously, including the position of the gaze focus and the state of eye closure.

[0122] (3) The training dataset was divided into two parts by replaying the experimental process and manually labeling the data: a large-scale movement dataset and a precise adjustment dataset. The reason for dividing the training set into two parts is that during the large-scale movement of the robotic arm and the adjustment process of the robotic arm end effector to complete the grasping and moving actions, there are significant differences between human eye movement characteristics and the real-time motion data of the robotic arm. After the dataset is divided, two pairs of eye movement data and robotic arm motion data are formed. Group training is beneficial to improve the matching degree between the conditional distribution model trained on different datasets and the corresponding eye movement characteristics.

[0123] It should be noted that the haptic force feedback device includes a base, connecting rods, a stylus capable of six degrees of freedom of movement / rotation, and buttons on the stylus. The operator holds the stylus and moves it freely. The host computer continuously detects the stylus's position and mapping it to the robotic arm's end effector, controlling the translation and rotation of the end effector in three-dimensional space. Simultaneously, the customizable buttons on the stylus allow the operator to control the opening and closing of the robotic arm's end effector grippers, achieving omnidirectional remote operation of the robotic arm.

[0124] For example, such as Figure 4 As shown, the operator holds the joystick of the haptic force feedback device. The host computer continuously reads the XYZ coordinates of the joystick tip and the Euler angles ABC describing the joystick's attitude at a sampling frequency of 50Hz. A direct proportional mapping is used for the XYZ coordinates, multiplying the joystick tip's XYZ coordinates by a scaling factor and mapping them to the robotic arm's end effector. An equal proportional mapping is used for the Euler angles ABC. The end effector attitude of the joystick tip can be represented by the pitch angle θ. H Yaw angle γ H Roll angle To represent, the definitions of the three perspectives are as follows: Figure 5 As shown. At time t0, both the robot end effector and the haptic force feedback device are placed in an initial posture, i.e. and At time t, the attitude control signal of the robot's end effector can be calculated using the following formula:

[0125] θ R (t)=θ R (t0)+(θ H (t)-θ H (t0))

[0126] γ R (t)=γ R (t0)+(γ H (t)-γ H (t0))

[0127]

[0128] Since the final grasping process uses a suitable horizontal orientation, the attitude control ultimately mapped to the host computer only uses the yaw angle γ. R That's all.

[0129] The three-dimensional positions of the tactile force feedback device's joystick and the robotic arm's end effector at time t can be represented by [X]. H (t)Y H (t)Z H [(t)] and [X] R (t)Y R (t)Z R [t] indicates that the position control signal of the robot's end effector can be calculated using the following formula:

[0130] X R (t)=k1X H (t),

[0131] Y R (t)=k2Y H (t),

[0132] Z R (t)=k3Z H (t),

[0133] Where k1, k2, and k3 are proportionality coefficients, [D XR D YR D ZR ] and [D XH D YH D ZH [ ] represents the dimensions of the motion space of the robotic arm end effector and the control lever end effector of the tactile force feedback device in the X, Y, and Z axes, respectively.

[0134] After establishing the teleoperation mapping relationship, the host computer continuously sends translational and rotational speed commands to the robotic arm at a control frequency of 50Hz. The operator then completes the "move-grab-place" process, and... Figure 6 The diagram shows six different gripping and placement positions for the robotic arm's gripper. For each of these different gripping and placement positions... Figure 7 As shown, two suitable horizontal poses are selected to complete the grasping and placing actions. The robotic arm gripper is in the same initial pose when it starts from the starting position.

[0135] Step S300: Based on eye movement features, perform online recognition and prediction of action intent, and extract action intent recognition.

[0136] In one possible implementation, step S300 includes the following steps:

[0137] S310, collects the operator's eye movement characteristics and updates the time observation sequence in real time. t ;

[0138] S320, a conditional distribution model is trained based on a large-scale mobile dataset and a fine-tuning dataset respectively, to obtain the conditional distribution probability P1(x). t |o t ) and P2(x t |o t Before the grasping is completed, the conditional distribution model to be used is determined based on a distance threshold. The relationship between the distance between the mechanical gripper position and the intended object to be grasped and the distance threshold is assessed. If the distance is greater than the distance threshold, the eye-tracking signal is used as a feature value input to the conditional distribution model trained on a large-scale mobile dataset to calculate P1(x). t If the distance is less than a distance threshold, the eye-tracking signal is used as a feature value input to a conditional distribution model trained on a finely adjusted dataset to calculate P2(x). t );

[0139] S330, after the capture is completed, the conditional distribution model to be used is determined by the fixation threshold; the relationship between the variance of the eye-tracking fixation focus within the window and the fixation threshold is continuously calculated. If the fixation variance is greater than the fixation threshold, the eye-tracking signal is used as a feature value and input into the conditional distribution model trained on a large-scale mobile dataset to calculate P1(x). t If the fixation variance is less than the fixation threshold, the eye movement signal is used as a feature value input to the conditional distribution model trained on the finely adjusted dataset to calculate P2(x). t Based on the maximum a posteriori probability principle, the predicted value x of the host computer's understanding of the operator's intention at time t is inferred. t (Including gripper Euler angle attitude, velocity direction and amplitude).

[0140] For example, see Figure 3 As shown, in this embodiment, based on the recognition result of the action intent, the direction and amplitude of the translational and rotational speeds of the robotic arm end effector with the highest probability expectation distribution are selected to form relevant control commands, which are sent to the robotic arm via USB data cable and USB communication protocol at a control frequency of 50Hz. Simultaneously, the host computer acquires eye movement data through an eye tracker to determine whether the operator has closed their eyes. Since the blinking time is typically between 0.2s and 0.4s, if the eye-closing time exceeds 0.5s and the gripper is not closed, the robotic arm gripper performs a "1-second pause - closing" operation; if the eye-closing time exceeds 0.5s and the gripper is already closed, the robotic arm gripper performs a "1-second pause - opening" operation.

[0141] The operator first uses eye-tracking signals to input a large-scale motion dataset to move the robotic arm gripper to near the pre-grabbing position. A global camera determines the distance between the gripper and the intended object. If the distance is less than a set threshold, the operator inputs eye-tracking signals into a fine-tuning dataset for precise position and posture adjustments to adapt the gripper to the object's shape. Finally, the operator enters the gripping position, closes their eyes, and triggers a "1-second pause-close" gripping action to grasp the object. After grasping, the operator re-inputs eye-tracking signals into the large-scale motion dataset to guide the robotic arm to the vicinity of the pre-placement position. The host computer continuously calculates whether the variance of the eye-tracking gaze focus within the most recent 3-second window is less than a gaze threshold. If the condition is met, the operator inputs eye-tracking signals into a fine-tuning dataset for precise position and posture adjustments to adapt to the surrounding environment. Finally, the operator enters the placement position, closes their eyes, and triggers a "1-second pause-open" gripper to place the object.

[0142] Step S400: Based on the action intent recognition, command the robotic arm to perform the action.

[0143] In one possible implementation, step S400 includes: obtaining the host computer's prediction of the operator's intention at time t; transmitting the predicted end-effector movement speed (including speed direction and amplitude) and end-effector posture (Euler angles) as an execution command to the cooperating robot arm via the host computer and USB communication protocol; and the gripper robot arm executing the host computer command and performing the action under the guidance of the obtained operator's action intention.

[0144] The following describes the action intent recognition and robotic arm sharing control device for eye movement feature extraction provided by the present invention. The action intent recognition and robotic arm sharing control device for eye movement feature extraction described below can be referred to in correspondence with the action intent recognition and robotic arm sharing control method for eye movement feature extraction described above.

[0145] Figure 8 This is a schematic diagram of the structure of the eye-tracking feature extraction action intent recognition and robotic arm shared control device provided in an embodiment of the present invention, as shown below. Figure 8 As shown, it includes: calibration module 81, model building and training module 82, extraction module 83, and extraction module 84, wherein:

[0146] Calibration module 81 is used for eye tracker and global camera calibration;

[0147] The model building and training module 82 is used to build and train the action intention recognition model; the action intention recognition model uses eye movement features as data input and action intention as data output to establish a relationship model between data input and output.

[0148] Extraction module 83 is used to perform online recognition and prediction of action intent based on eye movement features, and to extract action intent recognition;

[0149] Extraction module 84 is used to identify commands to the robotic arm to perform actions based on the action intent.

[0150] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 910, a communication interface 920, a memory 930, and a communication bus 940. The processor 910, communication interface 920, and memory 930 communicate with each other via the communication bus 940. The processor 910 can call logical instructions from the memory 930 to execute a motion intent recognition and robot-shared control method based on eye-tracking feature extraction.

[0151] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0152] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the action intention recognition and robot shared control method based on eye movement feature extraction provided by the above methods.

[0153] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the action intent recognition and robot shared control method based on eye movement feature extraction provided by the above methods.

[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for action intent recognition and shared control of a robotic arm based on eye-tracking feature extraction, characterized in that, The method is based on action intent recognition extracted from eye movement features and a shared control system for robotic arms. The eye-tracking feature extraction action intent recognition and robotic arm sharing control system includes: A binocular vision eye tracker, comprising a scene camera, a pupil-capturing camera, and a wireless module, is used to capture the three-dimensional coordinates of the gaze focus in the scene camera and transmit the data. A QR code positioning board is connected to a binocular vision eye tracker; A global camera system, including a regular camera and a depth camera, is used to capture videos of a positioning board with a QR code. Collaborative robotic arms are used for grasping, moving, and placing objects. Tactile force feedback devices are used to collect tactile information from operators; The host computer connects to the global camera, collaborative robotic arm, and haptic force feedback device; it is used to analyze and process the data, identify the operator's intentions online and convert them into control commands for the robotic arm, which are then sent to the collaborative robotic arm to execute the actions. The method for action intent recognition and shared control of the robotic arm based on eye-tracking feature extraction includes the following steps: Eye tracker and global camera calibration; both the scene camera and the global camera of the eye tracker are calibrated using a chessboard. Building and training a motion intent recognition model includes: Operators remotely control the collaborative robotic arm through tactile force feedback devices to perform actions such as grasping, moving, and placing objects; A training dataset is obtained, including eye movement feature data captured by a binocular vision eye tracker and real-time motion data of a collaborative robotic arm collected by a host computer; the eye movement feature data includes gaze focus position, gaze focus movement speed, gaze vector angular velocity, gaze duration, saccade amplitude and duration; The training dataset is divided into a large-range movement dataset and a fine-tuning dataset. The large-range movement dataset is the dataset in which the robotic arm completes a relatively large displacement movement. The fine-tuning dataset is the dataset in which the robotic arm has completed a relatively large displacement movement and is close to the pre-grasping position or pre-placement position, and then performs small-range precise position and posture adjustments. Using eye-tracking feature data as input and real-time motion data as output, Gaussian mixture model is used as the basic model for fitting and training, resulting in action intention recognition models based on large-scale movement datasets and fine-tuning datasets, respectively. The action intent recognition model uses eye movement features as data input and action intent as data output to establish a relationship model between data input and output; it performs online recognition and prediction of action intent based on eye movement features and extracts action intent recognition data. The robotic arm is instructed to perform actions based on the intent of the action.

2. The method for action intent recognition and robotic arm shared control based on eye-tracking feature extraction according to claim 1, characterized in that, The method for calculating the focal point position of the line of sight is as follows: The two-dimensional image coordinates of the QR code vertices are obtained by using image recognition and corner detection to capture the image of the QR code positioning board captured by the global camera. Calculate the pose of the QR code positioning plate {B} relative to the global camera {C} C T B ; Based on the position transformation matrix between the eye tracker scene camera {S} and the QR code calibration board {B} B T S Calculate the pose of the scene camera relative to the global camera using an eye tracker. C T S ; Based on the transformation matrix between the global camera {C} and the robotic arm's base coordinate system {R} R T C Calculate the pose of the eye tracker scene camera relative to the global coordinate system. R T S And the direction of the three-dimensional line-of-sight vector and the coordinates of the line-of-sight focus in the global coordinate system are obtained.

3. The method for action intent recognition and robotic arm shared control based on eye-tracking feature extraction according to claim 1, characterized in that, In the action intent recognition model, the inputs include: state variables. ,in Represents the total amount of data collected. Data in The coordinate position of the end effector gripper of the internally cooperative robotic arm. Represents the total amount of data Data in Euler angle parameters of the end effector gripper of the internally cooperating robotic arm; time observation sequence ,in Representative data The focal position of the operator's eye in the global coordinate system. This represents eye movement patterns involving fixation and gaze shifting; By moving the dataset over a large area and fine-tuning the joint probability distribution of time observation sequences and state variables in the dataset. and The conditional probability distributions for training on a large-scale mobile dataset and a fine-tuned dataset are solved using Gaussian conditions, respectively. and ; The Gaussian mixture model for a large-scale mobile dataset is represented as follows: in, It is the parameter set of the Gaussian mixture model, including the first... The weights of the Gaussian distribution Mean vector Covariance Matrix Each component It is a multivariate Gaussian distribution. The number of Gaussian distributions; Given eye movement feature data End-effector data The probability under the given conditions is: Due to observed variables If the variable is continuous, the joint distribution probability can be calculated to determine the marginal probabilities. ,for: During the training process, Each Gaussian component is randomly initialized. The Expectation-Maximization (EM) algorithm is used to calculate the posterior probability of each sample belonging to a Gaussian component. The parameters of each Gaussian component are updated using the posterior probability. This process is repeated iteratively until convergence, obtaining a given real-time value under range-shifting mode. Observation conditions at that time Lower state variables ; Similarly, to obtain a given real-time value in precise adjustment mode Observation conditions at that time Lower state variables .

4. The method for action intent recognition and robotic arm shared control based on eye-tracking feature extraction according to claim 3, characterized in that, The method for online recognition and prediction of action intent based on eye movement features, and for extracting action intent recognition, includes the following steps: Collect operator eye movement data and update time observation sequences in real time. ; The conditional distribution probability was calculated based on a large-scale mobile dataset and a finely adjusted dataset. and The system determines which of two datasets to use based on a conditional threshold; it then determines the relationship between the distance between the mechanical gripper's position and the intended object to be grasped, and a distance threshold. If the distance is greater than the distance threshold, the eye-tracking signal is input into a conditional distribution model trained on a large-scale movement dataset to calculate... If the distance is less than a distance threshold, the eye-tracking signal is input into a conditional distribution model trained on a finely adjusted dataset to calculate... ; After capture, the relationship between the variance of the eye-tracking fixation focus within the window and the fixation threshold is continuously calculated; if the variance is greater than the fixation threshold, the eye-tracking signal is input into a dataset trained on a large-scale motion dataset. If the variance is less than the fixation threshold, then the input is trained based on a finely tuned dataset. Based on the maximum a posteriori probability principle, the end effector's moving speed and end effector posture are inferred. The host computer's prediction of the operator's intention at any given time. .

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for action intent recognition and robotic arm shared control based on eye movement feature extraction as described in any one of claims 1 to 4.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for action intent recognition and robotic arm shared control based on eye movement feature extraction as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Flying mechanical arm grabbing operation teleoperation method based on operator intention recognition

    CN112959342A

  • Eye movement selection interaction intention recognition method

    CN117742490A