Driver Gaze Mapping for Road Object Importance Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current dashcam applications require manual labeling of video data to understand road object relevance, which is inefficient and resource-intensive due to the scale of data and the need for manual training of computer vision models.
Innovation Solution
A video system that processes both forward-facing and driver-facing video data to determine road object importance by using transformation matrices to estimate image coordinates, generating heat maps, and training a machine learning model to identify relevant objects without manual labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling of video data is used to understand road object relevance, then measurement precision of road object importance is improved, but loss of time and productivity deteriorate due to the scale of data and manual training requirements
Solution Approach 1:
The system uses driver-facing camera data to automatically generate heat maps that label forward-facing video data without human intervention. The driver's gaze patterns serve as self-generated labels, enabling the system to process vast amounts of data at scale while maintaining accuracy in determining road object importance.
Solution Approach 2:
The patent replaces the mechanical process of manual data labeling with an automated computer vision system. The system uses transformation matrices to map driver gaze coordinates to forward-facing video coordinates, automatically generating heat maps that identify important road objects without human labor.
2Measurement precision
If manual labeling of video data is performed, then measurement precision of road object relevance is improved, but loss of energy and computing resources worsen due to the intensive processing required
Solution Approach 1:
The system leverages existing driver-facing camera footage to automatically generate labels for forward-facing video data. This self-service approach eliminates the need for expensive manual annotation processes while maintaining high accuracy in identifying relevant road objects, significantly reducing computing resource consumption.
Solution Approach 2:
The system performs preliminary processing of driver-facing video data to generate heat maps before applying them to forward-facing video analysis. This preliminary action creates reusable labels that can be applied across multiple datasets, reducing the overall computing resources needed for training and analysis.
3Manufacturing precision
If manual training of computer vision models is performed, then manufacturing precision of the model is improved, but loss of time and productivity deteriorate
Solution Approach 1:
The patent replaces manual model training with an automated system that uses transformation matrices to map driver gaze data to forward-facing video coordinates. This mechanical substitution enables rapid generation of training data and model iteration, dramatically reducing training time while maintaining or improving model accuracy through systematic heat map generation.
Solution Approach 2:
The system performs preliminary generation of heat maps from driver-facing video data before model training begins. This preliminary action creates a ready-to-use labeled dataset that accelerates the model training process, allowing for faster iteration and deployment without sacrificing manufacturing precision of the computer vision model.
Data Source
AI summary
A device may receive driver facing video data associated with a driver of a vehicle and forward facing video data associated with the vehicle, and may process the driver facing video data, with a face model, to identify driver head orientation and driver gaze. The device may generate a first transformation matrix mapping the driver facing video data, the driver head orientation, and the driver gaze, and may generate a second transformation matrix mapping the driver facing video data and the forward facing video data. The device may utilize the first transformation matrix and the second transformation matrix to estimate image coordinates, and may aggregate the image coordinates to generate aggregated coordinates. The device may generate heat maps based on the aggregated coordinates, may train machine learning model, with the heat maps, to generate a trained machine learning model, and may perform actions based on the trained machine learning model.


