Artificial intelligence-based gaze tracking system and method

The AI-based eye tracking system enhances accuracy and efficiency by using a U-Net model to detect and correct both eyes' gaze positions, addressing the limitations of conventional methods.

WO2026084200A1PCT designated stage Publication Date: 2026-04-23KOREA INST OF MEDICAL MICROROBOTICS
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KOREA INST OF MEDICAL MICROROBOTICS
Filing Date
2025-07-29
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Conventional eye-tracking methods face challenges such as high computational resource consumption, low accuracy, difficulty in detecting pupils when eyes are squinted, and inability to capture both eyes simultaneously, especially under varying lighting conditions and with high costs.

Method used

An AI-based eye tracking system using a U-Net model to extract binocular pupil information, generating gaze targets at random locations, and calculating convergence coordinates for accurate visualization and correction of both eyes.

Benefits of technology

Improves detection accuracy and reduces computational overhead while simultaneously correcting both eyes' gaze positions, minimizing errors and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025011288_23042026_PF_FP_ABST
    Figure KR2025011288_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention comprises the steps of: acquiring an eyeball image from an optical device; inputting the acquired eyeball image into a trained artificial intelligence model so as to extract binocular pupil information; generating a single gaze target at a random location on a display screen and collecting the binocular pupil information corresponding to the gaze target; and acquiring convergence coordinate information corresponding to a two-dimensional plane for visualization according to the gaze on the basis of the collected binocular pupil information.
Need to check novelty before this filing date? Find Prior Art

Description

AI-based eye-tracking system and method

[0001] The present invention was carried out under project number RS-2023-00302153 with the support of the Ministry of Health and Welfare, the research management agency for the said project is the Korea Health Industry Development Institute, the research project name is "Development of medical products based on micro medical robots", the research task name is "Development of a catheter active guiding medical device for coronary artery intervention procedures", the lead organization is the Korea Institute of Micro Medical Robots, and the research period is August 1, 2023 to December 31, 2027.

[0002] The present invention claims priority to Korean Patent Application No. 10-2024-0140498, titled "Artificial Intelligence-based Eye Tracking System and Method," filed with the Korean Intellectual Property Office on October 15, 2024, the contents of which are incorporated herein by reference in their entirety.

[0003] The present invention relates to an artificial intelligence-based eye tracking system and method, and more particularly to an artificial intelligence-based eye tracking system and method that uses artificial intelligence to track a user's gaze, visualizes interaction information through the tracked gaze, and evaluates eye tracking accuracy based on the visualized information.

[0004] Eye tracking technology is a technology that detects and analyzes a user's eye movements to determine the direction and focus of their gaze. Eye tracking technology is utilized in various fields, including Human-Computer Interaction (HCI), User Experience (UX) research, medical diagnosis, and marketing.

[0005] One eye-tracking method involves detecting eye movements using high-resolution cameras or infrared light sources and cameras. By analyzing the pattern of infrared light reflected from the eyes, the position and movement of the pupils are accurately tracked. Recent technological advancements have significantly improved the accuracy and ease of use of eye-tracking devices. In particular, advancements in head movement compensation technology enable the collection of accurate data without interfering with the user's natural movements.

[0006] One of the prior art inventions discloses technologies for detecting and tracking pupils within binocular eye images acquired based on an image sensor, and in particular, discloses a technology for detecting pupil regions based on a digital image processing algorithm within acquired binocular eye images. However, it has the problem that errors in pupil detection continuously occur when the user looks to the side or squints their eyes, and the detection results may be shaken in consecutive frames.

[0007] Another conventional invention tracks the pupil based on a stereo infrared camera and infrared illumination, and obtains pupil information after detecting the face and eye region based on digital image processing technology. This technology has limitations, such as consuming a large amount of computation and resources, having low accuracy, and making it difficult to detect the pupil when the eyes are squinted.

[0008] Furthermore, conventional methods for visualizing and evaluating eye tracking results correct and analyze eye positions by projecting red dots onto a wall. Commercially available products have the problem of being unable to capture both eyes simultaneously and incurring high costs.

[0009] Therefore, there is a need for an AI-based eye tracking system and method that detects and tracks pupil information within eye images while simultaneously correcting the gaze of both eyes at a low cost.

[0010] The present invention aims to provide an AI-based eye tracking system and method that detects and tracks pupil information within eye images based on artificial intelligence, and can simultaneously correct the gaze of both eyes at a low cost.

[0011] To achieve the above objective, the present invention is characterized by comprising: a step of acquiring an eye image from an optical device; a step of inputting the acquired eye image into a learned artificial intelligence model to extract binocular pupil information; a step of generating a single gaze target at a random location on a display screen and collecting the binocular pupil information corresponding to the gaze target; and a step of acquiring convergence coordinate information corresponding to a two-dimensional plane for visualization based on the collected binocular pupil information.

[0012] Preferably, the pupil information may be at least one of the pupil size, the pupil center coordinates, and the pupil contour information.

[0013] Preferably, the artificial intelligence model may be a U-Net model.

[0014] Preferably, the step of collecting the above-mentioned binocular pupil information can generate the single gaze target at a random location based on computer graphics.

[0015] Preferably, the step of collecting the binocular pupil information may involve repeatedly generating the single gaze target at a random location to collect multiple binocular pupil information corresponding to multiple gaze targets.

[0016] Preferably, the step of collecting binocular pupil information may collect multiple binocular pupil information for a certain period of time for a single gaze target.

[0017] Preferably, the step of collecting the binocular pupil information can calculate the average value of the binocular pupil information for a single gaze target.

[0018] Preferably, the step of obtaining the convergence coordinate information can project the average value of the binocular pupil information onto a two-dimensional plane through the calculation of a homography matrix.

[0019] Preferably, the step of obtaining the convergence coordinate information can be performed by averaging the coordinates of the left eye and the right eye projected onto a two-dimensional plane to obtain the convergence coordinate information.

[0020] Preferably, the method may further include the step of creating a virtual object at a random location on a display screen and interacting with a user based on whether the location coordinates of the virtual object match the convergence coordinate information.

[0021] Preferably, the step of interacting with the user can evaluate the accuracy of the user-specific eye tracking and correction results through interaction.

[0022] In addition, the present invention is further characterized by comprising: an eye image acquisition unit for acquiring an eye image from an optical device; an extraction unit for extracting binocular pupil information by inputting the acquired eye image into a learned artificial intelligence model; a collection unit for generating a single gaze target at a random location on a display screen and collecting the binocular pupil information corresponding to the gaze target; and a convergence coordinate information acquisition unit for acquiring convergence coordinate information corresponding to a two-dimensional plane for visualization based on the collected binocular pupil information.

[0023] The present invention has the advantage of being able to increase detection accuracy and improve computational speed by using artificial intelligence to detect pupil information within eye images.

[0024] In addition, the present invention has the advantage of being able to simultaneously correct pupil information of both eyes.

[0025] In addition, the present invention has the advantage of reducing the error between the actual gaze position and the tracked gaze position by correcting the pupil information of both eyes through the generation of random positions of the gaze target.

[0026] Figure 1 shows a flowchart of an artificial intelligence-based eye tracking method according to an embodiment of the present invention.

[0027] Figure 2 shows the structure of an artificial intelligence model according to an embodiment of the present invention.

[0028] FIG. 3 is a figure illustrating the step of collecting binocular pupil information according to an embodiment of the present invention.

[0029] Figure 4 shows a figure illustrating the process of obtaining and visualizing convergence coordinate information according to an embodiment of the present invention.

[0030] Figure 5 shows an example of the simulation result of the interaction between a virtual object and convergence coordinate information according to an embodiment of the present invention.

[0031] Figure 6 shows a configuration diagram of an artificial intelligence-based eye tracking system according to an embodiment of the present invention.

[0032] A step of acquiring an eye image from an optical device;

[0033] A step of extracting binocular pupil information by inputting acquired eye images into a trained artificial intelligence model;

[0034] A step of generating a single gaze target at a random location on a display screen and collecting the binocular pupil information corresponding to the gaze target; and

[0035] A step of obtaining convergence coordinate information corresponding to a two-dimensional plane for visualization based on the collected binocular pupil information;

[0036] An artificial intelligence-based eye tracking method that includes

[0037] The present invention will be described in detail below with reference to the contents described in the attached drawings. However, the present invention is not limited or restricted by exemplary embodiments. Identical reference numerals in each drawing indicate components that perform substantially the same function.

[0038] The purpose and effects of the present invention may be naturally understood or become clearer through the following description, and the purpose and effects of the present invention are not limited solely to the description below. Furthermore, in describing the present invention, if it is determined that a detailed description of known technology related to the present invention may unnecessarily obscure the essence of the present invention, such detailed description will be omitted.

[0039] The terms used in this invention are used merely to describe specific embodiments and are not intended to limit the invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the description of the invention, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0040] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.

[0041] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which this invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this invention.

[0042] In interpreting the components, they are interpreted to include a margin of error even without a separate explicit indication. In the case of descriptions regarding temporal relationships, for example, where the temporal sequence is described using 'after,' 'following,' 'next,' 'before,' etc., cases that are not continuous are included unless 'immediately' or 'directly' is used.

[0043] Hereinafter, the technical configuration of the present invention will be described in detail with reference to the attached drawings.

[0044] FIG. 1 shows a flowchart of an artificial intelligence-based eye tracking method according to an embodiment of the present invention. Referring to FIG. 1, the artificial intelligence-based eye tracking method may include a step of acquiring an eye image (S100), a step of extracting binocular pupil information (S300), a step of collecting binocular pupil information (S500), a step of acquiring convergence coordinate information (S700), and a step of interacting with a user (S900).

[0045] Conventional eye-tracking methods using pupil information first track the entire face, isolate the eye region, and then detect the pupil within that region; however, this method consumes a significant amount of computational power and resources, and in particular, suffers from poor pupil detection accuracy. Furthermore, RGB images are heavily affected by ambient light and light reflection, and tracking failures frequently occur, especially in dim lighting conditions. While IR images are less affected by ambient light, the pupil detection results are influenced by the position and distance of infrared lighting. Specifically, when the pupil moves laterally relative to the camera, the shape of the pupil in the image is acquired as an ellipse; in such cases, pupil detection and tracking frequently fail, resulting in the pupil being missed. Additionally, pupil detection is difficult when the eyes are squinted (like a double eye). Moreover, there is a problem where already tracked pupil coordinate information is significantly shaken or errors occur in consecutive video frames due to ambient light or various other factors.

[0046] Conventional pupil information correction methods perform calibration by printing or printing a calibrator on paper or a hard plate, but because it is not a perfect flat surface, correction errors inevitably occur, resulting in relatively lower accuracy compared to the proposed method. Additionally, there is the inconvenience of having to create or print new calibration targets or points of various shapes or sizes to configure or change them. Furthermore, Interacoustics’ EyeSeeCam vHIT product cannot capture both eyes simultaneously and performs gaze correction using only five points displayed on a wall, clearly has limitations in scalability, and in particular, has the disadvantage of being expensive.

[0047] The step of acquiring an eye image (S100) may acquire an eye image from an optical device. Here, the optical device may be an RGB sensor or an IR camera.

[0048] The step of extracting binocular pupil information (S300) can extract binocular pupil information by inputting the acquired eye image into a trained artificial intelligence model.

[0049] Pupil information may be at least one of pupil size, pupil center coordinates, and pupil contour information. Preferably, pupil information may be pupil center coordinates.

[0050] The artificial intelligence model may be a deep learning-based artificial intelligence model. The artificial intelligence model can be trained using a dataset consisting of data labeled with pupil information detected in eye images. Although training the artificial intelligence model may take a considerable amount of time, once training is complete, pupil information can be extracted quickly, allowing for the real-time extraction of pupil information from eye images captured in real time.

[0051] FIG. 2 shows the structure of an artificial intelligence model according to an embodiment of the present invention. Referring to FIG. 2, the artificial intelligence model may be a U-Net model. The U-Net model is a Convolutional Neural Network (CNN) structure for image segmentation, exhibits high performance in medical image analysis, and can be used in the present invention to detect binocular pupil information in eye images. The U-Net model has a U-shaped network structure and may be composed of a contraction path (encoder) on the left and an expansion path (decoder) on the right.

[0052] The Contracting Path of the U-Net model can be composed of a series of convolutional layers and max pooling layers, capable of compressing spatial information of an image and extracting features. In the Contracting Path of the U-Net model, the spatial dimension of the feature map can be reduced and the number of channels can be increased at each step. The Expanding Path of the U-Net model can progressively restore the spatial dimension of the feature map through upsampling and convolution operations, and can preserve information by concatenating feature maps in the corresponding layers of the Contracting Path. In this process, the feature maps of the Contracting Path can be concatenated into the corresponding layers of the Expanding Path, thereby maintaining high-resolution features and preserving accurate positional information. The final output of the U-Net model can generate segmentation maps corresponding to the desired number of classes (in this case, pupil contours) through 1x1 convolution.

[0053] The U-Net model can achieve high performance even with a small amount of training data, effectively detect objects of various sizes, and enable accurate segmentation even in complex structures such as medical images. The U-Net model improves the accuracy of binocular pupil information detection and stabilizes tracking results, while demonstrating robust pupil detection performance regardless of diverse lighting conditions or eye shapes.

[0054] FIG. 3 illustrates a step (S500) for collecting binocular pupil information according to an embodiment of the present invention. Referring to FIG. 3, the step (S500) for collecting binocular pupil information may generate a single gaze target at a random location on a display screen and collect binocular pupil information corresponding to the gaze target. Here, the collected binocular pupil information may be collected from among the binocular pupil information extracted in the step (S300) for extracting binocular pupil information.

[0055] The step of collecting binocular pupil information (S500) can correct binocular pupil information based on a single gaze target displayed on a specific two-dimensional plane (two-dimensional monitor screen) based on computer graphics.

[0056] The step of collecting binocular pupil information (S500) can generate the single gaze target at a random location based on computer graphics. Since the step of collecting binocular pupil information (S500) generates the gaze target based on computer graphics, it is easy to change it to a desired shape or size.

[0057] The step of collecting binocular pupil information (S500) can collect multiple binocular pupil information corresponding to multiple gaze targets by repeatedly generating single gaze targets at random locations. Therefore, the step of collecting binocular pupil information (S500) can repeatedly correct binocular pupil information by changing the location of the single gaze targets generated at random locations. Since the step of collecting binocular pupil information (S500) uses single gaze targets generated at random locations rather than using targets at specific locations for correction, the accuracy of the correction can be increased. Additionally, since the step of collecting binocular pupil information (S500) does not use fixed gaze targets, it is possible to perform iterative correction by generating as many single gaze targets as desired. The accuracy of the step of collecting binocular pupil information (S500) can be improved as the iterative correction is repeated.

[0058] The step of collecting binocular pupil information (S500) can collect multiple binocular pupil information for a certain period of time for a single gaze target. That is, the step of collecting binocular pupil information (S500) can collect multiple binocular pupil information that is being extracted in real time while the user is looking at a single gaze target for a certain period of time. The step of collecting binocular pupil information (S500) can calculate the average value of the binocular pupil information for a single gaze target. The step of collecting binocular pupil information (S500) can obtain binocular pupil information with minimized error by collecting multiple binocular pupil information for one (specific) gaze target and averaging them.

[0059] Looking at a specific embodiment, the step (S500) of collecting binocular pupil information involves the user for a certain period of time (Δ s While looking at a gaze target during the ) period, left eye pupil information (G_Left) and right eye pupil information (G_Right) can be collected. Subsequently, the step of collecting binocular pupil information (S500) can repeatedly collect binocular pupil information corresponding to each single gaze target created at various locations. For example, the step of collecting binocular pupil information (S500) can collect (G_Left_1, G_Right_1, T_1), (G_Left_2, G_Right_2, T_2), ···, (G_Left_n, G_Right_n, T_n), where G_Left_n represents left eye pupil information when gazing at the nth single gaze target, G_Right_n represents right eye pupil information when gazing at the nth single gaze target, and T_n represents the coordinates of the nth single gaze target. The step (S500) of collecting binocular pupil information involves collecting binocular pupil center coordinate information from n gaze targets (T1, T2, ...T n ) at a certain time (Δ sAfter sequentially collecting binocular gaze coordinate information corresponding to each during the period, the average coordinate of the coordinates collected for each gaze target can be calculated (Equation 1).

[0060] [Mathematical Formula 1]

[0061]

[0062] The step of obtaining convergence coordinate information (S700) can obtain convergence coordinate information corresponding to a two-dimensional plane for visualization according to the gaze based on collected binocular pupil information.

[0063] The step of acquiring convergence coordinate information (S700) calculates the correlation between the two eyes and a two-dimensional plane (P1) using the coordinates (T_n) of a single gaze target and the collected binocular pupil information (G_Left_n, G_Right_n), and a matrix (H L , H R ) can be calculated (calibrated). Specifically, the step of obtaining convergence coordinate information (S700) calculates the matrix (H) through the calculation of the homography matrix. L , H R Can calculate ) (Mathematical Formula 2), and matrix (H L , H R Using ), the average value of binocular pupil information can be projected onto a two-dimensional plane. That is, the coordinates of the left eye projected onto a two-dimensional plane ( _Left) is G_Left* H L The coordinates of the right eye projected onto a two-dimensional plane, which can be calculated through the operation formula ( _Right) is G_Right* H R It can be calculated through the operation formula.

[0064] [Mathematical Formula 2]

[0065]

[0066] The step of obtaining convergence coordinate information (S700) obtains convergence coordinate information by averaging the coordinates of the left eye and the right eye projected onto a two-dimensional plane ( You can obtain ).

[0067] FIG. 4 illustrates a process for obtaining and visualizing convergence coordinate information according to an embodiment of the present invention. Referring to FIG. 4, the step of interacting with a user (S900) may create a virtual object at a random location on a display screen and interact with the user based on whether the location coordinates of the virtual object match the convergence coordinate information. The step of interacting with a user (S900) may create an arbitrary virtual object at a random location and trigger an interaction when the coordinate information of the virtual object and the convergence coordinate information collide. Here, the interaction can take any form as long as it provides a notification to the user, and is not limited to a specific method.

[0068] FIG. 5 shows an example of a simulation result of interaction between a virtual object and convergence coordinate information according to an embodiment of the present invention. Referring to FIG. 5, the step of interacting with a user (S900) can evaluate the accuracy of the eye tracking and correction results for each user through the interaction. That is, the step of interacting with a user (S900) can evaluate the accuracy of the device through the results of the interaction with the user.

[0069] FIG. 6 shows a configuration diagram of an artificial intelligence-based eye tracking system (10) according to an embodiment of the present invention. Referring to FIG. 6, the artificial intelligence-based eye tracking system (10) may include an eye image acquisition unit (100), an extraction unit (300), a collection unit (500), an acquisition unit (700), and an interaction unit (900).

[0070] The eye image acquisition unit (100) can acquire an eye image from an optical device. The eye image acquisition unit (100) can perform the aforementioned eye image acquisition step (S100).

[0071] The extraction unit (300) can extract binocular pupil information by inputting the acquired eye image into a trained artificial intelligence model. The extraction unit (300) can perform the step (S300) of extracting the aforementioned binocular pupil information.

[0072] The collection unit (500) can generate a single gaze target at a random location on the display screen and collect binocular pupil information corresponding to the gaze target. The collection unit (500) can perform the step (S500) of collecting the aforementioned binocular pupil information.

[0073] The acquisition unit (700) can acquire convergence coordinate information corresponding to a two-dimensional plane for visualization according to the gaze based on binocular pupil information. The acquisition unit (700) can perform the step (S700) of acquiring the aforementioned convergence coordinate information.

[0074] The interaction unit (900) can create a virtual object at a random location on the display screen and interact with the user based on whether the convergence coordinate information and the location coordinates of the virtual object match. The interaction unit (900) can perform the aforementioned interaction step (S900).

[0075] Although the present invention has been described in detail above through representative embodiments, those skilled in the art will understand that various modifications can be made to the above-described embodiments within the scope of the present invention. Therefore, the scope of the present invention should not be limited to the described embodiments, but should be determined by the claims set forth below as well as all modifications or variations derived from the claims and equivalent concepts.

[0076] The present invention aims to provide an AI-based eye tracking system and method that detects and tracks pupil information within eye images based on artificial intelligence, and can simultaneously correct the gaze of both eyes at a low cost.

Claims

1. A step of acquiring an eye image from an optical device; A step of extracting binocular pupil information by inputting acquired eye images into a trained artificial intelligence model; A step of generating a single gaze target at a random location on a display screen and collecting the binocular pupil information corresponding to the gaze target; and A step of obtaining convergence coordinate information corresponding to a two-dimensional plane for visualization based on the collected binocular pupil information; An artificial intelligence-based eye tracking method that includes 2. In Paragraph 1, The above pupil information is, An AI-based eye tracking method comprising at least one of pupil size, pupil center coordinates, and pupil contour information.

3. In Paragraph 1, The above artificial intelligence model is, An AI-based eye tracking method that is a U-Net model.

4. In Paragraph 1, The step of collecting the above-mentioned binocular pupil information is An artificial intelligence-based eye tracking method that generates the single gaze target at a random location based on computer graphics.

5. In Paragraph 4, The step of collecting the above-mentioned binocular pupil information is An AI-based eye tracking method that repeatedly generates a single gaze target at a random location and collects multiple binocular pupil information corresponding to multiple gaze targets.

6. In Paragraph 5, The step of collecting the above-mentioned binocular pupil information is An AI-based eye tracking method that collects multiple binocular pupil information for a certain period of time for a single gaze target.

7. In Paragraph 1, The step of collecting the above-mentioned binocular pupil information is An AI-based eye tracking method that calculates the average value of binocular pupil information for a single gaze target.

8. In Paragraph 7, The step of obtaining the above convergence coordinate information is to project the average value of the above-mentioned binocular pupil information onto a two-dimensional plane through the calculation of a homography matrix, an artificial intelligence-based eye tracking method.

9. In Paragraph 8, The step of obtaining the above convergence coordinate information is, An artificial intelligence-based eye tracking method that obtains convergent coordinate information by averaging the coordinates of the left and right eyes projected onto a two-dimensional plane.

10. In Paragraph 1, An AI-based eye tracking method further comprising the step of generating a virtual object at a random location on a display screen and interacting with a user based on whether the convergence coordinate information and the location coordinates of the virtual object match.

11. In Paragraph 10, The step of interacting with the above user is, An AI-based eye tracking method that evaluates the accuracy of user-specific eye tracking and correction results through interaction.

12. An eye image acquisition unit that acquires an eye image from an optical device; An extraction unit that extracts binocular pupil information by inputting acquired eye images into a trained artificial intelligence model; A collection unit that generates a single gaze target at a random location on a display screen and collects the binocular pupil information corresponding to the gaze target; and A convergence coordinate information acquisition unit that acquires convergence coordinate information corresponding to a two-dimensional plane for visualization according to gaze based on collected binocular pupil information; An artificial intelligence-based eye-tracking system that includes

Citation Information

Patent Citations

  • Eye gaze tracking based upon adaptive homography mapping

    KR1020160138062A

  • Low viscosity oil-in-water cosmetic composition with excellent moisturizing effect, feeling of use, and transparency

    KR1020250125464A

  • Eye gaze tracking

    US20110182472A1

  • Training a neural network model

    US20190156204A1

  • KR20190082688A