Augmenting Driving Data with Virtual Drivers for Interaction Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Perceiving driver interactions through vehicle windows is challenging due to their scarcity and difficulty in obtaining sufficient and quality training data, especially as these interactions often occur at large distances and result in blurry visual data.

Innovation Solution

A method and system that utilize simulation to train machine learning models, such as deep neural networks, by augmenting real-world driving data with virtual drivers performing target interactions, allowing for the generation of a large dataset of photo-realistic driver interactions like gaze and gestures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world driving data is used to train models, then the training data reflects actual driving scenarios, but the data is insufficient and of low quality due to scarcity of driver interactions

Engineering Contradiction:
Improvequality of training dataVSAvoidquantity of training data
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates virtual copies of real drivers through photorealistic rendering, generating synthetic training data that mimics real-world driver interactions. This copying approach allows unlimited generation of high-quality interaction data (gestures, gaze, expressions) without being constrained by the scarcity of actual recorded interactions

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-renders virtual driver interactions under various lighting, weather, and viewing conditions before they are needed for training. This preliminary preparation of diverse interaction scenarios ensures that the model receives comprehensive training data covering edge cases that would be difficult to capture in real-world recording

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If driver interactions are captured at large distances through windows, then the model can detect interactions from ego vehicle perspective, but the visual data becomes blurry and difficult to detect

Engineering Contradiction:
Improvedetection capability through windowVSAvoidclarity of visual data
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The rendering system applies local quality enhancement by selectively optimizing the resolution and detail of driver interaction regions (face, hands, upper body) while maintaining overall scene realism. This ensures that critical interaction features remain detectable even when viewed through window distortions at distance

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces a virtual camera intermediary that simulates the ego vehicle's viewing perspective through the window, applying appropriate distortions, reflections, and atmospheric effects. This intermediary allows the model to be trained on data that accurately represents the challenging viewing conditions while maintaining detectability of interactions

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If photorealistic virtual drivers are generated, then the training data quality improves, but the computational resources and time required for rendering increase

Engineering Contradiction:
Improvequality of training dataVSAvoiddata generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary rendering of virtual driver interactions, pre-computing photorealistic images under various conditions before they are needed for training. This advance preparation creates a reusable dataset that can be applied across multiple training iterations without repeated rendering costs

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The rendering pipeline uses dynamic parameter adjustment to balance photorealism quality with generation speed. Parameters such as resolution, lighting complexity, and environmental detail can be adjusted based on training needs, allowing the system to adapt between high-fidelity rendering for quality-critical samples and faster rendering for quantity-critical samples

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11574468B2Simulation-based learning of driver interactions through a vehicle window
Publication Date: 2023.02.07 TOYOTA JIDOSHA KK
  • US11574468B2 patent drawing
  • US11574468B2 patent drawing

AI summary

A model can be trained to detect interactions of other drivers through a window of their vehicle. A human driver behind a window (e.g., front windshield) of a vehicle can be detected in a real-world driving data. The human driver can be tracked over time through the window. The real-world driving data can be augmented by replacing at least a portion of the human driver with at least a portion of a virtual driver performing a target driver interaction to generate an augmented real-world driving dataset. The target driver interaction can be a gesture or a gaze. Using the augmented real-world driving data set, a machine learning model can be trained to detect the target driver interactions. Thus, simulation can be leveraged to provide a large set of useful training data without having to acquire real-world data of drivers performing target driver interactions as viewed from outside the vehicle.