Motion Planner Training Data Balancing for Rare Driving Behaviors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating training data for motion planners of autonomous vehicles is tedious, costly, and often results in imbalanced datasets that can lead to biased simulations, particularly in near-collision scenarios, which increases the risk of accidents and reduces the comprehensiveness of vehicle motion simulation.
Innovation Solution
A method and system that maps behaviors associated with trajectories into a feature space, clusters them, and adjusts the initial distribution to a target distribution by generating additional trajectories using a predictive model, ensuring a more uniform distribution of training data to minimize bias and enhance simulation comprehensiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If training data is generated from real-life settings with diverse behaviors, then simulation comprehensiveness is improved, but data collection cost and time increase significantly
Solution Approach 1:
The patent uses synthetic data generation through simulation environments to create copies of realistic driving scenarios without requiring actual physical data collection. The system generates training trajectories by simulating vehicle motions and environmental interactions, producing labeled data that mirrors real-life conditions while avoiding the time and cost of physical data gathering campaigns
Solution Approach 2:
The patent performs preliminary data generation and balancing before actual model training. By pre-generating diverse trajectories and applying re-sampling techniques to achieve balanced class distributions in advance, the system prepares comprehensive training data without needing to collect real-life data during the training phase, thus saving time while maintaining simulation comprehensiveness
2Adaptability or versatility
If training data is generated from real-life settings with diverse behaviors, then simulation comprehensiveness is improved, but data collection cost increases significantly
Solution Approach 1:
The patent replaces expensive real-life data collection with cost-effective synthetic data generation through simulation. The system creates training datasets by programmatically generating driving scenarios, vehicle trajectories, and environmental conditions, eliminating the need for costly human driver studies, sensor hardware deployment, and data annotation processes while maintaining comprehensive behavioral coverage
Solution Approach 2:
The patent implements self-service data generation where the simulation system automatically produces and balances training data without requiring external human assessors or data collection teams. The automated trajectory generation and re-sampling processes enable the system to service its own training data needs, significantly reducing operational costs while achieving diverse behavior representation
3Manufacturing precision
If training trajectories are labeled by human assessors to achieve behavior diversity, then data quality is improved, but generation cost increases with the number of behaviors
Solution Approach 1:
The patent replaces human assessor labeling with automated synthetic labeling through simulation physics and predefined behavior models. The system generates trajectories with inherent semantic labels based on simulation parameters and environmental conditions, producing high-quality labeled data without human intervention, thus eliminating the linear cost increase associated with adding more behavior categories
Solution Approach 2:
The patent implements self-service automated labeling where the simulation system automatically assigns behavior labels to generated trajectories based on embedded logic and environmental context. This automated labeling process maintains consistent data quality across all behavior types without requiring human assessors, thereby preventing generation cost from increasing with the number of behaviors
4Productivity
If training data is imbalanced towards certain behaviors, then training efficiency is improved, but motion planner simulation becomes biased
Solution Approach 1:
The patent implements feedback-based data balancing through re-sampling techniques that analyze the distribution of behaviors in the training dataset and adjust sample frequencies accordingly. The system calculates class imbalances and applies over-sampling to underrepresented behaviors or under-sampling to overrepresented ones, creating a balanced dataset that maintains training efficiency while eliminating simulation bias through iterative distribution correction
Solution Approach 2:
The patent changes the distribution parameters of training data by applying re-sampling techniques that modify the frequency weights of different behavior classes. By adjusting these distribution parameters to achieve uniform or target distributions, the system maintains efficient training convergence while ensuring reliable, unbiased motion planner performance across all behavior categories
Data Source
AI summary
A method and server for motion planning of an autonomous vehicle are provided. The method comprises: receiving an original dataset including a plurality of trajectories for the autonomous vehicle, the plurality of trajectories being distributed over a predetermined plurality of behaviors according to an initial distribution; acquiring an indication of a target distribution of the plurality of trajectories over the predetermined plurality of behaviors; determining a difference between the initial and target distributions, thereby identifying behaviors, for which additional trajectories are to be generated for adjusting the initial distribution of the plurality of trajectories to the target distribution; generating, using the plurality of trajectories in the original dataset and the additional trajectories, an augmented dataset having the target distribution of trajectories over the predetermined plurality of behaviors; and training a motion planner model to simulate motion of the autonomous vehicle based on the augmented plurality of trajectories.


