Eye Tracking Deep Learning Model Face Vector Input
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current eye tracking technologies face challenges in improving accuracy and efficiency, particularly in providing effective advertising services on user terminals, where precise gaze detection is necessary for targeted advertising.
Innovation Solution
A user terminal equipped with an imaging device and an eye tracking unit that utilizes a deep learning model to track user gaze by inputting face and ocular images, along with a vector representing the direction the user's face is facing, to enhance accuracy and reliability, and collects training data for model training through actions like screen touches and utterances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional eye tracking methods (video analysis, contact lens, sensor attachment) are used, then eye tracking can be performed, but the accuracy is insufficient for effective advertising service provision
Solution Approach 1:
The patent replaces traditional mechanical and optical eye tracking systems (contact lenses, sensor attachments, complex camera setups) with a software-based deep learning model that processes standard face images. This substitution achieves high measurement precision through algorithmic analysis of facial features and eye movements without requiring specialized hardware, thereby resolving the contradiction between accuracy and device complexity
Solution Approach 2:
The patent transforms the eye tracking approach by changing the input parameters from specialized sensor data to standard face images that can be captured by conventional cameras. The deep learning model processes these images to extract gaze information, enabling accurate eye tracking while maintaining simplicity in the hardware configuration and reducing the overall system complexity
2Measurement precision
If deep learning model is used for eye tracking, then accuracy is improved, but training data collection and model training time are required
Solution Approach 1:
The patent implements preliminary action by collecting training data and training the deep learning model in advance, before actual eye tracking is needed. The trained model is then ready for immediate deployment, allowing the system to achieve high gaze detection accuracy without incurring training time delays during actual advertising service operations. This separates the time-consuming training phase from the operational phase
Solution Approach 2:
The system performs self-service by automatically collecting training data from user interactions and autonomously training the deep learning model. This self-training capability eliminates the need for manual data collection and model training interventions, reducing the time loss associated with setup and enabling the system to improve its own accuracy over time without external assistance
3Measurement precision
If multiple input parameters (face image, ocular image, face direction vector) are provided to deep learning model, then gaze tracking accuracy is enhanced, but data processing complexity increases
Solution Approach 1:
The patent merges multiple data sources (face image, ocular image, and face direction vector) into a unified deep learning model input framework. By combining these parameters simultaneously, the system achieves enhanced gaze direction accuracy while managing processing complexity through integrated model architecture that handles all inputs in a coordinated manner, resolving the contradiction between precision and complexity
Data Source
AI summary
A user terminal according to an embodiment of the present invention includes a capturing device for capturing a face image of a user, and an eye tracking unit for, on the basis of a configured rule, acquiring, from the face image, a vector representing the direction that the face of the user is facing, and a pupil image of the user, and performing eye tracking of the user by inputting, in a configured deep learning model, the face image, the vector and the pupil image.


