Machine learning and deep learning combined calibration-free eye movement tracking system and method
Through the combination of machine learning and deep learning, users' personalized parameters are automatically extracted, which solves the problems of complex calibration process, time-consuming and expensive equipment in the existing calibration-free eye tracking technology, and achieves high real-time, low cost and ease of use, adapts to the differences in eye structures of different users, and improves the accuracy of gaze point prediction.
Patent Information
- Application Number
- CN202510310873.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing calibration-free eye tracking technology has problems such as complex calibration process and time-consuming, complex model parameters, poor real-time effects, expensive equipment and complex layout, and cannot meet the requirements of high real-time and low cost.
Using a combination of machine learning and deep learning, eye images are collected through near-eye cameras and infrared light sources, AdaBoost integrated learning algorithm and convolutional neural network are used to automatically extract user personalized kappa angles, and combined with lightweight deep learning models, a calibration process without the active participation of users is realized.
It significantly lowers the threshold for use, improves the real-time and flexibility of the system, reduces equipment costs, adapts to the differences in eye structures of different users, and improves the accuracy of gaze point prediction and the ease of use of the system.
Smart Images

Figure CN120260106A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of eye movement tracking and image processing, and particularly relates to an uncalibrated eye movement tracking system and method combining machine learning and deep learning. Background Art
[0002] Modern eye movement tracking technology mainly relies on high-precision sensors and deep learning algorithms to infer the user's line of sight direction and fixation point by capturing subtle changes in eye movement. The uncalibrated technology is an important innovation for eye movement tracking technology, aiming to improve the user experience and system applicability. The uncalibrated eye movement tracking technology aims to reduce or completely eliminate the calibration step through advanced algorithms and hardware design, enabling the eye movement tracking system to accurately track the user's eye movement in real time without prior calibration.
[0003] The existing technologies are mainly divided into two categories, both of which can solve the uncalibrated problem to a certain extent. One method is the 3D-based method that uses near-eye image features to estimate eye parameters, so as to simulate and adjust the eyeball model, and then infers the fixation point direction based on the established model and real-time near-eye images; the other method is the implicit calibration method that uses the change of eye movement trajectory and binocular visual axis parameters, combined with specific light sources and hardware settings, and can continuously optimize through algorithms during the user's use process to improve the accuracy and reliability of the predicted fixation point position.
[0004] However, the 3D-based method does not consider the requirements of high real-time performance and low cost in actual application scenarios. Although it can achieve uncalibrated to a certain extent, the model parameters are too complex, and at the same time, the requirements for near-eye devices are too high. The implicit calibration method can be calibrated during the user's natural reading behavior, but the actual calibration, performance optimization effect, and calibration time vary from person to person, and the uncalibrated problem cannot be completely solved.
[0005] In view of this, it is very meaningful to propose an uncalibrated eye movement tracking system and method combining machine learning and deep learning. Summary of the Invention
[0006] To solve the problems existing in the existing calibration-free eye tracking technology, including complex and time-consuming calibration processes, complex model parameters, poor real-time performance, expensive equipment and great influence by layout, etc., the present invention provides a calibration-free eye tracking system and method combining machine learning and deep learning. Firstly, in the training stage, the Adaboost algorithm is used to construct a decision maker to replace the complex calibration process in the real machine stage. Secondly, since each part of the designed system is separated, after calculating the corresponding calibration angle and weight combination, only the deep learning model needs to be run, thus improving the real-time performance of the system. In addition, in terms of equipment, the system only requires a scene camera and a near-eye camera, with low equipment cost and simple layout, improving the flexibility and usability of the system.
[0007] In a first aspect, the present invention proposes a calibration-free eye tracking system combining machine learning and deep learning, which system comprises the following modules:
[0008] A pre-calibration module, configured to collect near-eye images of a user's eyes through a near-eye camera and an infrared light source, extract the pupil center coordinates and corneal reflection point coordinates, and calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates;
[0009] A decision maker module, adopting the AdaBoost ensemble learning algorithm, classifying the user according to the calculated kappa angle, and outputting the corresponding weight combination; and
[0010] A deep learning module, configured to receive the output weight combination and real-time near-eye images, and predict the coordinates of the user's fixation point through a convolutional neural network and a residual connection structure.
[0011] Preferably, the pre-calibration module specifically comprises the following steps: graying and binarizing the collected near-eye images, and removing noise through morphological opening operation; identifying the closed regions in the images through a contour detection algorithm, screening the closed regions with the second largest area and the number of contour points ≥ 5, and determining the pupil shape and its center coordinates (x p , y p ) by ellipse fitting; according to the vector relationship v = (x l - x p , y l - y p ) between the pupil center and the reflection point coordinates, calculating the average kappa angle within the valid frames, and using the arctan2 function to calculate the angle angle between the vector and the x-axis, with the formula: angle = arctan2(y l - y p , x l - x p) Add the calculated angle and distance to the [angles] and [distances] lists respectively, calculate the average value of all valid angles to obtain the calibration angle kappa angle, and the formula is: where M is the number of valid frames that satisfy the Euclidean distance threshold of 80; finally, return the calculated calibration angle kappa angle.
[0012] Preferably, the decision maker module is constructed through the following steps: Set T optimal weak classifiers and the thresholds of the weak classifiers, input the obtained kappa angle into the weak classifiers for training, and finally obtain a group of classifiers, and calculate the classification error rate error of the weak classifiers on the training set t , and the formula is: where is the weight of the i-th sample in the t-th iteration, y i is the true label of the i-th sample, f t (x i ) is the predicted value of the t-th weak classifier for the i-th sample, is the indicator function, which is 1 when the prediction is incorrect and 0 otherwise; update the weights of the training samples according to the output of the weak classifier, so that the misclassified samples obtain higher weights in the next iteration, and the specific formula is: To ensure that the sum of the sample weights is 1, normalize the updated sample weights: In each iteration, if the calculated classification error rate error t is less than the preset threshold of 0.1, then include this weak classifier and its weight α t in the final classifier group; otherwise, retrain this weak classifier or adjust its parameters until the threshold requirement is met; the final decision maker consists of all qualified weak classifiers and their weights, and its output is the weighted sum of the outputs of all weak classifiers, and the specific formula is: where F(x) is the output of the classifier group, α t is the weight of the t-th weak classifier, and f t (x) is the output of the t-th weak classifier.
[0013] Preferably, the deep learning module includes: Input layer: Receive grayscale images with a resolution of 120×90; Convolution module: Sequentially include 3 convolutional layers, and each convolutional layer is followed by batch normalization, ReLU activation function and average pooling operations; Residual connection: Add the output of the third convolutional module to the output feature map of the first convolutional module; Fully connected layer: Map the features to the fixation point coordinates (x, y) through a 5-layer fully connected network.
[0014] Second aspect, embodiments of the present invention provide a calibration-free eye movement tracking method combining machine learning and deep learning, including the system described in any one of the first aspects, and further including the following steps:
[0015] Pre-calibration stage: collect near-eye images of the user's eyes through a near-eye camera and an infrared light source, and extract the pupil center coordinates and corneal reflection point coordinates; calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates;
[0016] Decision-making stage: use the AdaBoost ensemble learning algorithm to classify the user according to the calculated kappa angle and output the corresponding weight combination;
[0017] Real-time tracking stage: input the real-time near-eye image and the weight combination into the deep learning model to predict the fixation point coordinates.
[0018] Further preferably, in the pre-calibration stage, pupil center detection includes: performing binarization processing on the image with a threshold set to 80; removing noise through morphological opening operation, and then screening the region with the second largest area and the number of contour points ≥ 5 for ellipse fitting.
[0019] Preferably, in the decision-making stage, the training of the weak classifier satisfies: the classification error rate threshold is set to 0.1; the weak classifier weights are determined by iteratively updating the sample weights.
[0020] Preferably, the input of the deep learning model is a grayscale image of 120×90, the output is the fixation point coordinates, and the network structure includes residual connections.
[0021] Third aspect, embodiments of the present invention provide an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the second aspect.
[0022] Fourth aspect, embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the second aspect.
[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0024] (1) Calibration-free operation: Automatically extract the user's personalized parameters (kappa angle) through the pre-calibration module, without the user's active participation in the cumbersome calibration process, significantly reducing the usage threshold, especially suitable for children, visually impaired people, and scenarios that require rapid deployment.
[0025] (2) High real-time performance: By adopting the AdaBoost decision maker and a lightweight deep learning model (including residual connections), the algorithm complexity is reduced, and the response speed is increased by 30%, which can meet the application requirements with strict real-time performance such as VR / AR.
[0026] (3) Low cost and easy deployment: The hardware only requires a monocular near-eye camera and an ordinary scene camera, combined with an infrared light source. The device cost is reduced by 50%, and the layout is simple, which is suitable for consumer products.
[0027] (4) High-precision tracking: By integrating machine learning classification and deep learning feature extraction, the fixation point prediction error is less than 2°, and the accuracy is better than traditional methods based on 3D models or implicit calibration.
[0028] (5) Strong adaptability: By dynamically adjusting the classification weights through AdaBoost, it can adapt to the differences in the eye structures of different users, reduce the dependence on a large amount of personal data, and improve the generalization ability of the system. Description of the Drawings
[0029] The drawings are included to provide a further understanding of the embodiments and are incorporated into and form a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding like parts.
[0030] Figure 1 It is a schematic diagram of the architecture of the machine learning and deep learning combined calibration-free eye movement tracking system according to the embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of the process of the machine learning and deep learning combined calibration-free eye movement tracking method according to the embodiment of the present invention;
[0032] Figure 3 It is a general scheme flowchart of a specific embodiment of the present invention;
[0033] Figure 4 It is a schematic diagram of the pre-calibration module of a specific embodiment of the present invention;
[0034] Figure 5 It is the pre-processed image of a specific embodiment of the present invention;
[0035] Figure 6 It is a calculation result diagram of pupil and highlight coordinates of a specific embodiment of the present invention;
[0036] Figure 7Schematic diagram of the AdaBoost algorithm classifier according to a specific embodiment of the present invention;
[0037] Figure 8 Schematic diagram of the deep learning model according to a specific embodiment of the present invention;
[0038] Figure 9 One of the real - machine diagrams according to a specific embodiment of the present invention, the fixation point following the finger diagram;
[0039] Figure 10 Two of the real - machine diagrams according to a specific embodiment of the present invention, the fixation point observing a fixed object diagram;
[0040] Figure 11 Schematic diagram of the structure of the computer device of the electronic device suitable for implementing the embodiments of the present invention. Detailed implementation manners
[0041] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only the parts related to the invention are shown in the drawings.
[0042] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.
[0043] In recent years, eye - tracking technology has developed rapidly. In particular, the progress in hardware and algorithms has continuously expanded its application scope. Modern eye - tracking technology mainly relies on high - precision sensors and deep - learning algorithms to infer the user's line of sight direction and fixation point by capturing the subtle changes in eye movements. With the maturity of computer vision and artificial intelligence technologies, the accuracy and response speed of eye - tracking systems have been significantly improved, and the devices have become more portable and user - friendly. Currently, eye - tracking technology is widely used in fields such as human - computer interaction, virtual reality (VR), augmented reality (AR), gaming, market research, medical diagnosis, and assistive technologies.
[0044] However, most existing eye - tracking devices still require users to perform a calibration step before use to ensure the accuracy of the system. These methods either require viewing multiple explicitly fixed points for calibration or require a large amount of time to collect and label personal data sets. Typical methods mainly include the pupil - corneal reflection method (PCCR), which obtains the gaze direction by comparing the relative positions of the corneal reflection point and the pupil center on the eyeball in an infrared camera; and appearance - based methods, which mainly obtain the mapping relationship function of the fixation direction through a large number of collected and labeled personal near - eye image data sets to achieve eye - tracking.
[0045] This process not only increases the user's burden of use, but also may lead to poor experiences in certain application scenarios (such as VR and AR), especially in environments where the head needs to be moved frequently or quick responses are required. Frequent calibration will interrupt the user's immersion and smooth experience. Therefore, developing a head-mounted calibration-free eye-tracking technology that can automatically adapt to different users without the user's active participation in the calibration process has become an important direction of current research.
[0046] Calibration-free technology is an important innovation for eye-tracking technology, aiming to improve the user experience and system applicability. Currently, eye-tracking technology has shown great potential in many fields such as virtual reality, augmented reality, and human-computer interaction. However, the requirement for the calibration process and the threshold of knowledge about it limit its current development. The significant progress of calibration-free technology has significantly reduced the usage threshold, enhanced the flexibility and reliability of the technology, met the urgent need for efficient and natural interaction methods in various fields, and can achieve high-precision tracking in different environments and conditions without the user's participation in the calibration process. In addition, the emergence of calibration-free technology has greatly simplified the operation process. Eye-tracking can be quickly started without cumbersome calibration, which enables eye-tracking technology to be more widely applied in various scenarios, improves its usability and practicality, and brings more possibilities for its collaborative development with other related fields.
[0047] Calibration-free eye-tracking technology aims to reduce or completely eliminate the calibration steps through advanced algorithms and hardware designs, enabling the eye-tracking system to track the user's eye movements in real time and accurately without prior calibration. Existing technologies are mainly divided into two categories, both of which can solve the calibration-free problem to a certain extent. Currently, the main types of calibration-free head-mounted eye-tracking are as follows:
[0048] 1. 3D-based methods utilize near-eye image features to estimate eye parameters, enabling the simulation and adjustment of the eyeball model, and then inferring the gaze direction based on the established model and real-time near-eye images.
[0049] Working principle: Use high-resolution and high-frame-rate cameras to capture images of the user's eyeballs. These cameras are usually equipped with infrared illumination functions and can clearly image under different lighting conditions to obtain high-quality eyeball images. Then, visual features (such as pupils, irises, and glints) are obtained through image processing techniques to calibrate the invariant eye parameters (such as corneal radius and Kappa angle) based on the eyeball structure and geometric imaging model. Next, the variable eye parameters (such as the center of the eyeball, corneal center, pupil center, and iris center) are estimated to reconstruct the optical axis (OA) of the eyeball, and then the visual axis (VA) is determined using OA and the Kappa angle, thereby simulating and approximating the eyeball model to infer the line of sight direction and the fixation point.
[0050] The specific process is as follows: First, a camera coordinate system is established with the optical center Co of the infrared camera as the origin, and a head coordinate system is established with the midpoint Ho of the line between the two inner canthi as the origin. The 3D position of the eye center Ec in the camera coordinate system can be calculated as: E C = R t ·E h + H0, where Rt is the rotation matrix of the head at different times, and Eh is the 3D coordinate of the eye center in the head coordinate system.
[0051] Calculation of the horizontal and vertical angles of the optical axis: The 3D iris center Ic in the camera coordinate system can be calculated based on the depth information Iz obtained from the Kinect sensor and the 2D coordinates of the iris center in the image plane through the projection relationship. The unit vector along the optical axis Ve can be determined by the iris center Ic and the eye center Ec, and re is the eye radius. The specific calculation is as follows:
[0052]
[0053] Then, the angle kappa is applied to the optical axis, so the unit vector along the visual axis can be written as:
[0054]
[0055] The unit vector Vg is defined as the predicted gaze direction. For N groups of calibration data, personal calibration is achieved by minimizing the sum of the angles between the predicted gaze direction and the true gaze direction. The cost function is as follows:
[0056]
[0057] In this method, the average value of human eye parameters is set as the initial personal parameters, and the iterative calculation of eye parameters is carried out using the inlier algorithm.
[0058] Main disadvantages and deficiencies: The 3D-based method does not take into account the requirements of high real-time performance and low cost in actual application scenarios. Although it can achieve calibration-free to a certain extent, the model parameters are too complex, and at the same time, the requirements for near-eye devices are too high.
[0059] 2. The method of implicit calibration uses the change of eye movement trajectory and binocular visual axis parameters, combined with specific light sources and hardware settings, and can continuously optimize through algorithms during the user's use process to improve the accuracy and reliability of the predicted fixation point position.
[0060] Working principle: Calibration is achieved by leveraging human natural behaviors, especially automatic calibration during the user's natural reading process. During reading, the user's line of sight naturally moves to different positions in the text, forming a stable reading trajectory. This natural eye movement trajectory can be captured and analyzed by algorithms, taking into account binocular visual axis parameters. Based on this, by combining saliency map transformation and averaging techniques, key information is extracted and integrated to obtain parameters closely related to calibration. With the help of these parameters, the algorithm can dynamically adjust the calibration parameters, enabling the system to continuously learn the eye movement pattern during the user's continuous use, optimize its own performance, and achieve a more accurate and efficient calibration effect.
[0061] Main disadvantages and deficiencies: The implicit calibration method can perform calibration during the user's natural reading behavior. However, the actual calibration, performance optimization effect, and calibration time vary from person to person, and the problem of calibration-free has not been completely solved.
[0062] Therefore, the present invention aims to solve several deficiencies existing in the existing calibration-free eye movement tracking technology. The main problems include:
[0063] 1. The calibration process is complex and time-consuming: Existing calibration methods require the user to perform a series of complex operations, such as fixating on multiple specific calibration points. This not only increases the user's burden of use but also prolongs the system preparation time. Moreover, it is particularly difficult for users with visual impairments or children to complete the calibration process.
[0064] 2. The model parameters are complex and the real-time effect is poor: The 3D-based calibration-free eye movement tracking model contains various complex parameters, such as the geometry of the eyeball, corneal radius, center point, etc. The real-time adjustment of these parameters and the approximation of the model are difficult, resulting in a slow response speed of the system during operation and unable to meet the application scenarios with high real-time requirements.
[0065] 3. The equipment is expensive and greatly affected by the layout: Existing calibration-free methods have certain requirements for the layout, quantity, and performance of cameras and light sources, which leads to high costs and complex equipment layouts.
[0066] To address these problems, the present invention first constructs a decision-making device using the Adaboost algorithm during the training stage to replace the complex calibration process in the actual machine stage. Secondly, since each part of the designed system is separated, after calculating the corresponding calibration angle and weight combination, only the deep learning model needs to be run, thereby improving the real-time effect of the system. In addition, in terms of equipment, the system only requires one scene camera and one near-eye camera, with low equipment costs and simple layouts, improving the flexibility and usability of the system.
[0067] In a first aspect, an embodiment of the present invention discloses a calibration-free eye movement tracking system that combines machine learning and deep learning. As Figure 1 shown, the system includes the following modules:
[0068] A pre-calibration module 11, configured to collect near-eye images of the user's eyes through a near-eye camera and an infrared light source, extract the pupil center coordinates and corneal reflection point coordinates, and calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates;
[0069] In this embodiment, the pre-calibration module 11 specifically includes the following steps:
[0070] Perform grayscale and binary processing on the collected near-eye images, and remove noise through morphological opening operation;
[0071] Identify the closed regions in the image through a contour detection algorithm, filter the closed region with the second largest area and the number of contour points ≥ 5, and use ellipse fitting to determine the pupil shape and its center coordinates (x p , y p );
[0072] According to the vector relationship v = (x l - x p , y l - y p ) between the pupil center and the reflection point coordinates, calculate the average kappa angle within the valid frame, and use the arctan2 function to calculate the angle angle between the vector and the x-axis. The formula is: angle = arctan2(y l - y p , x l - x p );
[0073] Add the calculated angle and distance to the [angles] and [distances] lists respectively, calculate the average value of all valid angles, and obtain the calibration angle kappa angle. The formula is: where M is the number of valid frames that meet the Euclidean distance threshold of 80; finally, return the calculated calibration angle kappa angle.
[0074] A decision maker module 12, which adopts the AdaBoost ensemble learning algorithm to classify the user according to the calculated kappa angle and output the corresponding weight combination; and
[0075] Specifically, the decision maker module 12 is constructed through the following steps:
[0076] Set T optimal weak classifiers and the thresholds of the weak classifiers, input the obtained kappa angles into the weak classifiers for training, and finally obtain a classifier group. Calculate the classification error rate error of the weak classifier on the training set t , and the formula is: where is the weight of the i-th sample in the t-th iteration, y i is the true label of the i-th sample, f t (x i ) is the predicted value of the t-th weak classifier for the i-th sample, is the indicator function, which is 1 when the prediction is incorrect and 0 otherwise;
[0077] Update the weights of the training samples according to the output of the weak classifier, so that the misclassified samples obtain higher weights in the next iteration. The specific formula is: To ensure that the sum of the sample weights is 1, normalize the updated sample weights:
[0078] In each iteration, if the calculated classification error rate error t is less than the preset threshold of 0.1, then include this weak classifier and its weight α t in the final classifier group; otherwise, retrain this weak classifier or adjust its parameters until the threshold requirement is met;
[0079] The final decision maker consists of all qualified weak classifiers and their weights, and its output is the weighted sum of the outputs of all weak classifiers. The specific formula is: where F(x) is the output of the classifier group, α t is the weight of the t-th weak classifier, and f t (x) is the output of the t-th weak classifier.
[0080] The deep learning module 13 is used to receive the output weight combination and the real-time near-eye image, and predict the coordinates of the user's fixation point through a convolutional neural network and a residual connection structure.
[0081] Specifically, the deep learning module 13 includes:
[0082] Input layer: Receive grayscale images with a resolution of 120×90;
[0083] Convolution module: Sequentially includes 3 convolutional layers, and each convolutional layer is followed by batch normalization, ReLU activation function, and average pooling operations;
[0084] Residual connection: Add the output of the third convolution module to the output feature map of the first convolution module;
[0085] Fully connected layer: Map the features to the fixation point coordinates (x, y) through a five-layer fully connected network.
[0086] In a second aspect, an embodiment of the present invention also discloses a calibration-free eye movement tracking method combining machine learning and deep learning. As Figure 2 shown, it includes the system described in the first aspect, and further includes the following steps:
[0087] S1. Pre-calibration stage: Collect near-eye images of the user's eyes through a near-eye camera and an infrared light source, and extract the pupil center coordinates and corneal reflection point coordinates; calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates.
[0088] Specifically, in this step, pupil center detection includes:
[0089] S11. Perform binary processing on the image, and set the threshold to 80.
[0090] S12. After removing noise through morphological opening operation, select the region with the second largest area and the number of contour points ≥ 5 for ellipse fitting.
[0091] S2. Decision-making stage: Use the AdaBoost ensemble learning algorithm to classify the user according to the calculated kappa angle, and output the corresponding weight combination.
[0092] Specifically, in the decision-making stage, the training of the weak classifier satisfies:
[0093] The classification error rate threshold is set to 0.1.
[0094] The weights of the weak classifier are determined by iteratively updating the sample weights.
[0095] S3. Real-time tracking stage: Input the real-time near-eye image and the weight combination into the deep learning model to predict the fixation point coordinates.
[0096] In this embodiment, the input of the deep learning model is a grayscale image of 120×90, the output is the fixation point coordinates, and the network structure includes residual connections.
[0097] In a specific embodiment, the technical solution of the present invention addresses the calibration-free problem in eye movement tracking technology and proposes a method combining machine learning and deep learning, which can significantly reduce the calibration time and improve the real-time performance of the system. As Figure 3 shown is the overall flowchart of the present invention. The left side is the training stage, and the right side is the actual machine operation stage. The specific solution is as follows:
[0098] 1. Training stage: The dataset mainly includes the collected near-eye images and calibration labels. Then, first, the corresponding calibration angles are calculated through the pre-calibration module 11. The specific process is as Figure 4 shown.
[0099] 1.1. Pre-calibration module
[0100] 1.1.1 Image acquisition and preprocessing
[0101] The system adopts a glasses hardware structure similar to the PCCR corneal reflection method, which includes a near-eye camera and an infrared light source to ensure that the corneal reflection point and the pupil center point are included in the image. At the beginning stage, the user needs to look straight ahead, and then the 10 collected pictures are processed into the PIL format and grayscale. The acquired and processed images are as Figure 5 shown.
[0102] 1.1.2 Calculation of pupil coordinates and bright point coordinates
[0103] When the glasses are worn stably, the relative position relationship between the pupil center and the bright point (corneal reflection) of each person remains basically unchanged, and this invariance helps the personal calibration of the eye movement tracking system.
[0104] The calculation of pupil coordinates is to binarize the grayscale image (threshold set to 120) and use morphological opening operation to remove noise. Then, the contour detection algorithm is used to identify the closed regions in the image. If more than two contours are detected, the contour with the second largest area is selected (the condition is that the number of points of this contour needs to be greater than 5), and the pupil shape and its center coordinates (x p , y p ) are determined through the ellipse fitting algorithm. When the number of contour points is less than 5 or the number of contours is less than 2, the pupil center coordinates are set to (0, 0), indicating that the detection is not successful.
[0105] The calculation of bright point coordinates is to first scale the collected grayscale image to the size of (640 * 480), then binarize the image (threshold set to 80) to distinguish the bright point area from the background; then, perform a morphological opening operation with a (3 * 3) kernel on the image to remove small noise and optimize the bright point contour; finally, use the contour detection algorithm to determine the centroid coordinates of the closed contour within a specific sub-region (rows 100:350, columns 190:400) of the image. If the number of detected contours is greater than 0, the centroid coordinates are determined by calculating the contour moments, and its calculation formula is:
[0106]
[0107] where, M 00 represents the area of the contour, M 10 and M 01Respectively represent the moments in the x and y directions of the contour. To avoid division-by-zero errors, when M 00 is not equal to 0, calculate the centroid coordinates and convert these coordinates back to the coordinate system of the original image. Finally, select the leftmost centroid from all the centroid coordinates as the highlight center coordinates (x l , y l ), where x l and y l are the abscissa and ordinate of the centroid respectively. If no contour is detected or the calculation of centroid coordinates fails, set the highlight center coordinates to (0, 0), and the final valid calculation results are as shown in Figure 6 .
[0108] 1.1.3 Calculate the calibration angle
[0109] To improve the reliability of the calibration angle and filter out unexpected situations caused by the lighting environment at the same time. Therefore, set the total number of valid frames M to 10, and calculate and judge the Euclidean distance between highlights. So, initialize two lists [angles] and [distances] to store the valid angles and distances for each frame respectively. The specific process is as follows:
[0110] Calculate the Euclidean distance distance between the pupil center and the highlight center, and check whether the calculated distance exceeds the set threshold distance_threshold of 80. If it exceeds, skip this frame. The calculation formula is:
[0111]
[0112] Calculate the vector v = (x l - x p , y l - y p ) between the pupil center and the highlight center, and use the arctan2 function to calculate the angle angle between the vector and the x-axis. The formula is:
[0113] angle = arctan2(y l - y p , x l - x p )
[0114] Add the calculated angle and distance to the [angles] and [distances] lists respectively. Calculate the average value of all valid angles to obtain the calibration angle (kappa angle). The formula is:
[0115]
[0116] Among them, M is the number of valid angles in the [angles] list. Finally, the calculated calibration angle (kappaangle) is returned.
[0117] 1.2 Calculate the calibration weight
[0118] According to the obtained calibration angle, the collected data set is classified. When training, the Adam algorithm is selected as the optimizer, and the mean square error between the predicted coordinates and the true coordinates is used as the loss function.
[0119] 1.3 Decision maker module
[0120] In eye tracking technology, the kappa angle is of great significance for achieving high-precision tracking and calibration-free. Since different eye structures of people will result in different kappa angles, but these differences have clustering characteristics within a certain range, that is, calibration-free can be achieved through the classification of the kappa angle. The AdaBoost ensemble learning algorithm forms a decision maker by combining multiple weak classifiers, as Figure 7 shown, and can effectively handle this individual difference and clustering characteristic. The specific training process is as follows:
[0121] First, set T optimal weak classifiers and the thresholds of the weak classifiers. Input the obtained kappa angle into the weak classifiers for training, and finally obtain a classifier group. Among them, calculate the classification error rate error of the weak classifier on the training set t , and the specific formula is:
[0122]
[0123] Among them, is the weight of the i-th sample in the t-th iteration, y i is the true label of the i-th sample, f t (x i ) is the predicted value of the t-th weak classifier for the i-th sample, is the indicator function, which is 1 when the prediction is incorrect and 0 otherwise.
[0124] Update the weights of the training samples according to the output of the weak classifier, so that the misclassified samples will obtain higher weights in the next iteration. The specific formula is:
[0125]
[0126] To ensure that the sum of the sample weights is 1, normalize the updated sample weights:
[0127]
[0128] In each iteration, if the calculated classification error rate errort If it is less than the preset threshold of 0.1, then this weak classifier and its weight α t will be incorporated into the final classifier group. Otherwise, retrain this weak classifier or adjust its parameters until the threshold requirement is met.
[0129] The final decision maker consists of all qualified weak classifiers and their weights, and its output is the weighted sum of the outputs of all weak classifiers. The specific formula is:
[0130]
[0131] where F(x) is the output of the classifier group, α t is the weight of the t-th weak classifier, and f t (x) is the output of the t-th weak classifier.
[0132] 2. Real machine operation stage
[0133] During use, as Figure 3 shown on the right, the near-eye camera takes an image and enters the pre-calibration module 11 to calculate the calibration angle, which is then input into the decision maker to obtain the corresponding weight combination. Finally, the deep learning model outputs the predicted coordinates to the scene camera based on the weight combination and the real-time near-eye image, where the deep learning model is as Figure 8 shown.
[0134] The specific network process is as Figure 8 shown. The processed 120x90 grayscale image is input into the network, and feature extraction is performed through three convolutional modules (conv1, conv2, conv3). Each module contains a convolutional layer, a batch normalization layer, a ReLU activation function, and an average pooling layer. Then, the output feature map of conv3 is added to the output feature map of conv1 to form a residual connection, which helps to retain low-level feature information. Then, the added feature map is flattened into a one-dimensional vector and input into a 5-layer fully connected layer. These fully connected layers gradually reduce the feature dimension, and the final output is the (x, y) coordinates of the fixation point.
[0135] Figure 9 For the real machine Figure 1 : Fixation point follows finger diagram; Figure 10 For the real machine Figure 2 : Fixation point observes fixed object diagram.
[0136] The solution of the present invention mainly consists of a pre-calibration module 11, a decision maker module 12, and a deep learning module 13. The calibration module uses the hardware design of the PCCR reflection method and advanced image processing technology to accurately measure the individual kappa angle within a short time, greatly improving the speed and accuracy of data acquisition. The decision maker module 12 overcomes the problems of the complex traditional model structure and cumbersome calibration procedure. It places the calibration link in the pre-training stage, combines with the AdaBoost machine learning algorithm to establish a decision maker, and can intelligently evaluate and determine the best weight combination according to the individual kappa angle and its distribution characteristics without additional calibration, significantly optimizing the calibration process and improving the system intelligence level. At the same time, the decision maker module 12 and the deep learning module 13 are highly independent in design, which not only simplifies the overall system architecture but also greatly improves the processing speed and efficiency of eye movement tracking applications, comprehensively enhancing the system performance. In summary, while ensuring the accuracy, the present invention greatly enhances the application flexibility and user-friendliness of eye movement tracking technology.
[0137] Reference is made below to Figure 11 , which shows a schematic structural diagram of a computer device 600 suitable for use in implementing an embodiment of the present invention (such as Figure 1 the server or terminal device shown). Figure 11 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0138] As Figure 11 shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 603 or the programs loaded from the storage section 609 into the random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other through a bus 605. The input / output (I / O) interface 606 is also connected to the bus 605.
[0139] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 may also be connected to the I / O interface 606 as required. A removable medium 612 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 611 as required so that a computer program read therefrom is installed into the storage section 609 as required.
[0140] Specifically, according to the embodiments disclosed by the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed by the present invention include a computer program product which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 610, and / or installed from the removable medium 612. When the computer program is executed by a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, the above functions defined in the methods of the present invention are executed.
[0141] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or components, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution device, apparatus, or component. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0142] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0144] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor.
[0145] As another aspect, the present invention also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately and not be assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: perform the methods and steps described in the second aspect.
[0146] The above description is only the preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present invention.
Claims
1. An uncalibrated eye movement tracking system combining machine learning and deep learning, characterized in that, The system includes the following modules: A pre-calibration module, which is used to collect near-eye images of the user's eyes through a near-eye camera and an infrared light source, extract the pupil center coordinates and corneal reflection point coordinates, and calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates; A decision maker module, which adopts the AdaBoost ensemble learning algorithm, classifies the user according to the calculated kappa angle, and outputs the corresponding weight combination; and A deep learning module, which is used to receive the output weight combination and real-time near-eye images, and predict the coordinates of the user's fixation point through a convolutional neural network and a residual connection structure.
2. The calibration-free eye movement tracking system combining machine learning and deep learning according to claim 1, characterized in that, The pre-calibration module specifically includes the following steps: Gray-scale and binary-process the collected near-eye images, and remove noise through morphological opening operation; Identify the closed regions in the image through the contour detection algorithm, filter out the closed region with the second largest area and the number of contour points ≥ 5, and use ellipse fitting to determine the pupil shape and its center coordinates (x p , y p ); According to the vector relationship v=(x l -x p ,y l -y p ) between the pupil center and the reflection point coordinates, calculate the average kappa angle within the valid frame, and use the arctan2 function to calculate the angle angle between the vector and the x-axis. The formula is: angle = arctan2(y l -y p ,x l -x p ); Add the calculated angles and distances to the [angles] and [distances] lists respectively, and calculate the average value of all valid angles to obtain the calibration angle kappa angle. The formula is: where M is the number of valid frames that satisfy the Euclidean distance threshold of 80; Finally, return the calculated calibration angle kappa angle.
3. The calibration-free eye movement tracking system combining machine learning and deep learning according to claim 1, characterized in that, The decision maker module is constructed through the following steps: Set T optimal weak classifiers and the thresholds of the weak classifiers, input the obtained kappa angle into the weak classifiers for training, and finally obtain a classifier group, and calculate the classification error rate error of the weak classifier on the training set t , and the formula is: where is the weight of the i-th sample in the t-th iteration, y i is the true label of the i-th sample, f t (x i ) is the predicted value of the t-th weak classifier for the i-th sample, is an indicator function, which is 1 when the prediction is incorrect and 0 otherwise; Update the weights of the training samples according to the output of the weak classifier, so that the misclassified samples will obtain higher weights in the next iteration. The specific formula is as follows: To ensure that the sum of the sample weights is 1, normalize the updated sample weights: In each iteration, if the calculated classification error rate error t is less than the preset threshold of 0.1, then include this weak classifier and its weight α t in the final classifier group; otherwise, retrain this weak classifier or adjust its parameters until the threshold requirement is met; The final decision maker consists of all eligible weak classifiers and their weights, and its output is the weighted sum of the outputs of all weak classifiers. The specific formula is as follows: where F(x) is the output of the classifier group, and α t is the weight of the t-th weak classifier, and f t (x) is the output of the t-th weak classifier.
4. The machine learning and deep learning combined calibration-free eye movement tracking system according to claim 1, characterized in that, The deep learning module includes: Input layer: Receives grayscale images with a resolution of 120×90; Convolution module: Sequentially includes 3 convolutional layers, and each convolutional layer is followed by batch normalization, ReLU activation function, and average pooling operations; Residual connection: Add the output of the third convolutional module to the output feature map of the first convolutional module; Fully connected layer: Map the features to the fixation point coordinates (x, y) through a 5-layer fully connected network.
5. A calibration-free eye movement tracking method combining machine learning and deep learning, characterized in that, Including the system according to any one of claims 1-4, further comprising the following steps: Pre-calibration stage: Collect near-eye images of the user's eyes through a near-eye camera and an infrared light source, extract the pupil center coordinates and corneal reflection point coordinates; calculate the user's personalized kappa angle based on the extracted pupil center coordinates and corneal reflection point coordinates; Decision stage: Adopt the AdaBoost ensemble learning algorithm, classify the user according to the calculated kappa angle, and output the corresponding weight combination; Real-time tracking stage: Input the real-time near-eye images and the weight combination into the deep learning model to predict the fixation point coordinates.
6. The calibration-free eye movement tracking method combining machine learning and deep learning according to claim 5, wherein In the pre-calibration stage, the pupil center detection includes: Perform binary processing on the image, and set the threshold to 80; After removing noise through morphological opening operation, select the region with the second largest area and the number of contour points ≥ 5 for ellipse fitting.
7. The calibration-free eye movement tracking method combining machine learning and deep learning according to claim 5, characterized in that, In the decision stage, the training of the weak classifier satisfies: The classification error rate threshold is set to 0.1; The weak classifier weights are determined by iteratively updating the sample weights.
8. The calibration-free eye movement tracking method combining machine learning and deep learning according to claim 5, characterized in that, The input of the deep learning model is a grayscale image of 120×90, the output is the fixation point coordinates, and the network structure includes a residual connection.
9. An electronic device, comprising: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 5 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 5 to 8.
Citation Information
Cited By
Non-contact sight line estimation method based on deep learning model
CN121564784A
Brightness control method and display device based on ai and eye tracking
CN122598581A