Vehicle Gaze Detection with Virtual Camera Ground Truth

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches for training machine learning models to detect objects in vehicle environments, such as determining a driver's gaze direction, face challenges due to the need for extensive labeled training data and the difficulty in obtaining ground truth data when cameras are placed at arbitrary locations within a vehicle, leading to uncertainty in relative positions of objects to the camera.

Innovation Solution

The use of a virtual camera space and a coordinate propagation mechanism to bridge physical world points to virtual camera space, enabling the creation of ground truth data for arbitrary camera positions, and the deployment of deep neural networks across various vehicle models by correlating different coordinate systems, including camera, vehicle, and occupant coordinate systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If cameras are placed at arbitrary locations within a vehicle to improve installation flexibility, then ease of manufacture is improved, but measurement precision deteriorates due to uncertainty in relative positions of objects to the camera

Engineering Contradiction:
Improvecamera installation flexibilityVSAvoidgaze direction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces fiducial markers as intermediary objects placed on the vehicle interior at known positions. These markers serve as reference points that enable the system to calculate accurate gaze directions even when cameras are installed at arbitrary locations. The markers mediate between the camera's unknown position and the objects being observed, allowing the neural network to learn accurate spatial relationships without requiring precise camera positioning during installation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive labeled training data with ground truth is collected to improve machine learning model accuracy, then measurement precision is improved, but loss of time increases due to the long and complicated data creation process

Engineering Contradiction:
Improveobject detection accuracyVSAvoidtraining data creation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses fiducial markers as simplified copies or proxies for complex ground truth data. Instead of requiring lengthy manual labeling processes to establish spatial relationships between objects and camera positions, the system uses easily detectable fiducial markers with known positions as substitutes. These markers provide the necessary geometric information in a standardized, machine-readable format that accelerates the training data creation process while maintaining accuracy.

Inventive Principle:
Principle #26Copying

3Reliability

If supervised training with ground truth data is used to improve model reliability, then reliability is improved, but device complexity increases due to the need for coordinate propagation mechanisms and virtual camera space

Engineering Contradiction:
Improvegaze detection reliabilityVSAvoidcoordinate system transformation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service mechanism where the fiducial markers automatically provide the coordinate transformation information needed for reliable gaze detection. The markers contain encoded spatial information that enables the system to self-calibrate and perform coordinate propagation without requiring complex external calibration procedures or manual intervention. This self-service approach maintains reliability while reducing operational complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11803759B2Gaze detection using one or more neural networks
Publication Date: 2023.10.31 NVIDIA CORP
  • US11803759B2 patent drawing
  • US11803759B2 patent drawing
  • US11803759B2 patent drawing

AI summary

Apparatuses, systems, and techniques are described to determine locations of objects using images including digital representations of those objects. In at least one embodiment, a gaze of one or more occupants of a vehicle is determined independently of a location of one or more sensors used to detect those occupants.