Vehicle Gaze Detection Across Arbitrary Camera Positions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches for training machine learning models for object identification require a significant amount of labeled training data, which can be costly and time-consuming to create, especially in environments where camera positions are arbitrary and variable.

Innovation Solution

The use of a virtual camera space and a coordinate propagation mechanism to bridge the physical world with image data captured from arbitrary camera positions, allowing for the generation of ground truth data and training of neural networks across various vehicle configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional supervised training approaches are used for machine learning models, then the model can achieve accurate object identification, but a significant amount of labeled training data is required which is costly and time-consuming to create

Engineering Contradiction:
Improveobject identification accuracyVSAvoidtraining data creation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates virtual 3D representations (digital twins) of physical objects and environments, which can be copied and used as synthetic training data. These virtual models replicate the properties of real objects without requiring physical prototypes or extensive photo collection, thereby reducing the time and cost of creating training datasets while maintaining identification accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables automatic generation of training data through automated 3D scanning and modeling processes. The virtual environment automatically captures and processes physical object data, generating labeled training samples without manual annotation, thus eliminating the time-consuming manual labeling process while preserving measurement precision

Inventive Principle:
Principle #25Self-service

2Measurement precision

If conventional supervised training approaches are used for machine learning models, then the model can achieve accurate object identification, but the process becomes too expensive for various uses

Engineering Contradiction:
Improveobject identification accuracyVSAvoidtraining cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

By creating virtual copies of physical objects and environments, the patent eliminates the need for expensive physical prototypes, multiple camera setups, and manual annotation services. The virtual 3D models can be generated once and reused indefinitely across different training scenarios, significantly reducing costs while maintaining identification accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The virtual 3D environment serves multiple functions: it acts as both the training data source and the simulation environment. A single virtual model can be used for training multiple different object identification tasks and can be adapted to various camera positions and lighting conditions, reducing the need for separate expensive datasets for each scenario

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If training data is created manually for supervised learning, then ground truth data can be obtained, but the process is complicated and results in insufficient amount of training data

Engineering Contradiction:
Improveground truth data qualityVSAvoidtraining data generation rate
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The virtual 3D environment automatically generates ground truth data through programmed object positions, properties, and relationships. The system self-annotates the training data by tracking object states and camera parameters, eliminating manual annotation bottlenecks and enabling unlimited generation of high-quality labeled data at high speed

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

All ground truth information is pre-programmed into the virtual 3D environment before training begins. Object positions, physical properties, and environmental conditions are defined in advance, allowing rapid generation of labeled training samples without time-consuming post-capture annotation processes

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If neural networks are trained for specific vehicle configurations, then accurate gaze detection can be achieved, but the model cannot be deployed in different vehicle models with variable camera positions

Engineering Contradiction:
Improvegaze detection accuracyVSAvoidvehicle model compatibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates virtual 3D models of different vehicle interiors and camera systems. These virtual models can be configured to match any vehicle configuration, allowing the neural network to be trained on diverse virtual datasets that encompass multiple vehicle types and camera positions, thereby achieving both accuracy and adaptability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The virtual 3D environment dynamically adapts to different vehicle configurations through programmable camera positions and vehicle geometry. The training system can simulate various camera placements and vehicle interiors, enabling the model to learn gaze detection across multiple vehicle types without retraining from scratch, thus achieving versatility while maintaining precision

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250156717A1Gaze detection using one or more neural networks
Publication Date: 2025.05.15 NVIDIA CORP
  • US20250156717A1 patent drawing
  • US20250156717A1 patent drawing
  • US20250156717A1 patent drawing

AI summary

Apparatuses, systems, and techniques are described to determine locations of objects using images including digital representations of those objects. In at least one embodiment, a gaze of one or more occupants of a vehicle is determined independently of a location of one or more sensors used to detect those occupants.