Augmenting Driving Data with Virtual Drivers for Interaction Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Perceiving driver interactions through vehicle windows is challenging due to their scarcity and difficulty in obtaining sufficient and quality training data, especially as these interactions often occur at large distances and result in blurry visual data.
Innovation Solution
A method and system that utilize simulation to train machine learning models, such as deep neural networks, by augmenting real-world driving data with virtual drivers performing target interactions, allowing for the generation of a large dataset of photo-realistic driver interactions like gaze and gestures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-world driving data is used to train models, then the training data reflects actual driving scenarios, but the data is insufficient and of low quality due to scarcity of driver interactions
Solution Approach 1:
The patent creates virtual copies of real drivers through photorealistic rendering, generating synthetic training data that mimics real-world driver interactions. This copying approach allows unlimited generation of high-quality interaction data (gestures, gaze, expressions) without being constrained by the scarcity of actual recorded interactions
Solution Approach 2:
The system pre-renders virtual driver interactions under various lighting, weather, and viewing conditions before they are needed for training. This preliminary preparation of diverse interaction scenarios ensures that the model receives comprehensive training data covering edge cases that would be difficult to capture in real-world recording
2Adaptability or versatility
If driver interactions are captured at large distances through windows, then the model can detect interactions from ego vehicle perspective, but the visual data becomes blurry and difficult to detect
Solution Approach 1:
The rendering system applies local quality enhancement by selectively optimizing the resolution and detail of driver interaction regions (face, hands, upper body) while maintaining overall scene realism. This ensures that critical interaction features remain detectable even when viewed through window distortions at distance
Solution Approach 2:
The patent introduces a virtual camera intermediary that simulates the ego vehicle's viewing perspective through the window, applying appropriate distortions, reflections, and atmospheric effects. This intermediary allows the model to be trained on data that accurately represents the challenging viewing conditions while maintaining detectability of interactions
3Reliability
If photorealistic virtual drivers are generated, then the training data quality improves, but the computational resources and time required for rendering increase
Solution Approach 1:
The system performs preliminary rendering of virtual driver interactions, pre-computing photorealistic images under various conditions before they are needed for training. This advance preparation creates a reusable dataset that can be applied across multiple training iterations without repeated rendering costs
Solution Approach 2:
The rendering pipeline uses dynamic parameter adjustment to balance photorealism quality with generation speed. Parameters such as resolution, lighting complexity, and environmental detail can be adjusted based on training needs, allowing the system to adapt between high-fidelity rendering for quality-critical samples and faster rendering for quantity-critical samples
Data Source
AI summary
A model can be trained to detect interactions of other drivers through a window of their vehicle. A human driver behind a window (e.g., front windshield) of a vehicle can be detected in a real-world driving data. The human driver can be tracked over time through the window. The real-world driving data can be augmented by replacing at least a portion of the human driver with at least a portion of a virtual driver performing a target driver interaction to generate an augmented real-world driving dataset. The target driver interaction can be a gesture or a gaze. Using the augmented real-world driving data set, a machine learning model can be trained to detect the target driver interactions. Thus, simulation can be leveraged to provide a large set of useful training data without having to acquire real-world data of drivers performing target driver interactions as viewed from outside the vehicle.

