Single-Camera Driver Gaze Tracking Without Calibration Markers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing driver monitoring systems require additional hardware, such as second cameras, which are expensive and difficult to install, to determine driver gaze direction, and often necessitate the driver to look at specific indicators, complicating the calibration process.
Innovation Solution
A system that uses a single image sensor, such as a camera, integrated with a processor to capture and process images of the driver's face, allowing for the determination of gaze direction without the need for a second camera or specific calibration points, by analyzing eye features and correlating them with predefined directions based on vehicle motion and behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single image sensor is used to capture driver face images, then hardware cost and installation complexity are reduced, but the ability to accurately determine driver gaze direction is compromised
Solution Approach 1:
The system segments the driver monitoring task into distinct components: face detection, eye region identification, pupil center detection, and gaze direction calculation. By processing different facial features separately and combining their information, the system achieves accurate gaze determination using only a single image sensor.
Solution Approach 2:
The system transitions from two-dimensional image coordinates to three-dimensional spatial relationships by incorporating head pose estimation (pitch, yaw, roll angles) and applying geometric transformations. This dimensional transformation enables accurate gaze direction calculation on the road surface despite using a single camera viewpoint.
2Measurement precision
If existing eye tracking systems are calibrated using predefined indicators, then gaze measurement accuracy is improved, but the calibration process becomes complex and time-consuming
Solution Approach 1:
The system performs self-calibration by automatically detecting facial landmarks (eyes, pupils, nose, mouth) and establishing the relationship between eye position and gaze direction without requiring the driver to follow calibration instructions or look at specific indicators. The calibration data is extracted directly from natural driver behavior during normal operation.
Solution Approach 2:
The system performs preliminary facial feature detection and head pose estimation to establish the geometric relationship between the camera, eyes, and road surface before calculating gaze direction. This preliminary setup creates a transformation model that enables subsequent gaze measurements without requiring real-time calibration adjustments.
3Device complexity
If a single image sensor is used without additional cameras, then system cost is reduced, but the capability to track driver attention on the road is insufficient
Solution Approach 1:
The single image sensor performs multiple functions: capturing driver face images for eye tracking, detecting head pose, and providing geometric reference for road surface mapping. By making the single sensor multi-functional through sophisticated image processing and geometric modeling, the system achieves reliable driver attention tracking without requiring multiple specialized sensors.
Solution Approach 2:
The system introduces geometric transformation models and head pose estimation as intermediaries between the single image sensor and the final gaze direction output. These computational intermediaries bridge the gap between the limited sensor data and the comprehensive driver attention information required for reliable monitoring.
Data Source
AI summary
Systems and methods are disclosed for driver monitoring. In one implementation, one or more images are received, e.g., from an image sensor. Such image(s) can reflect a at least a portion of a face of a driver. Using the images, a direction of a gaze of the driver is determined. A set of determined driver gaze directions is identified using at least one predefined direction. One or more features of one or more eyes of the driver are extracted using information associated with the identified set.


