Motion Planner Training Data Balancing for Rare Driving Behaviors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating training data for motion planners of autonomous vehicles is tedious, costly, and often results in imbalanced datasets that can lead to biased simulations, particularly in near-collision scenarios, which increases the risk of accidents and reduces the comprehensiveness of vehicle motion simulation.

Innovation Solution

A method and system that maps behaviors associated with trajectories into a feature space, clusters them, and adjusts the initial distribution to a target distribution by generating additional trajectories using a predictive model, ensuring a more uniform distribution of training data to minimize bias and enhance simulation comprehensiveness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If training data is generated from real-life settings with diverse behaviors, then simulation comprehensiveness is improved, but data collection cost and time increase significantly

Engineering Contradiction:
Improvesimulation comprehensivenessVSAvoiddata collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation through simulation environments to create copies of realistic driving scenarios without requiring actual physical data collection. The system generates training trajectories by simulating vehicle motions and environmental interactions, producing labeled data that mirrors real-life conditions while avoiding the time and cost of physical data gathering campaigns

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data generation and balancing before actual model training. By pre-generating diverse trajectories and applying re-sampling techniques to achieve balanced class distributions in advance, the system prepares comprehensive training data without needing to collect real-life data during the training phase, thus saving time while maintaining simulation comprehensiveness

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If training data is generated from real-life settings with diverse behaviors, then simulation comprehensiveness is improved, but data collection cost increases significantly

Engineering Contradiction:
Improvesimulation comprehensivenessVSAvoiddata collection cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent replaces expensive real-life data collection with cost-effective synthetic data generation through simulation. The system creates training datasets by programmatically generating driving scenarios, vehicle trajectories, and environmental conditions, eliminating the need for costly human driver studies, sensor hardware deployment, and data annotation processes while maintaining comprehensive behavioral coverage

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements self-service data generation where the simulation system automatically produces and balances training data without requiring external human assessors or data collection teams. The automated trajectory generation and re-sampling processes enable the system to service its own training data needs, significantly reducing operational costs while achieving diverse behavior representation

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If training trajectories are labeled by human assessors to achieve behavior diversity, then data quality is improved, but generation cost increases with the number of behaviors

Engineering Contradiction:
Improvedata qualityVSAvoidgeneration cost
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The patent replaces human assessor labeling with automated synthetic labeling through simulation physics and predefined behavior models. The system generates trajectories with inherent semantic labels based on simulation parameters and environmental conditions, producing high-quality labeled data without human intervention, thus eliminating the linear cost increase associated with adding more behavior categories

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements self-service automated labeling where the simulation system automatically assigns behavior labels to generated trajectories based on embedded logic and environmental context. This automated labeling process maintains consistent data quality across all behavior types without requiring human assessors, thereby preventing generation cost from increasing with the number of behaviors

Inventive Principle:
Principle #25Self-service

4Productivity

If training data is imbalanced towards certain behaviors, then training efficiency is improved, but motion planner simulation becomes biased

Engineering Contradiction:
Improvetraining efficiencyVSAvoidsimulation bias
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements feedback-based data balancing through re-sampling techniques that analyze the distribution of behaviors in the training dataset and adjust sample frequencies accordingly. The system calculates class imbalances and applies over-sampling to underrepresented behaviors or under-sampling to overrepresented ones, creating a balanced dataset that maintains training efficiency while eliminating simulation bias through iterative distribution correction

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the distribution parameters of training data by applying re-sampling techniques that modify the frequency weights of different behavior classes. By adjusting these distribution parameters to achieve uniform or target distributions, the system maintains efficient training convergence while ensuring reliable, unbiased motion planner performance across all behavior categories

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12565234B2Method and a system for generating training data for training a motion planner
Publication Date: 2026.03.03 HUAWEI TECH CO LTD
  • US12565234B2 patent drawing
  • US12565234B2 patent drawing
  • US12565234B2 patent drawing

AI summary

A method and server for motion planning of an autonomous vehicle are provided. The method comprises: receiving an original dataset including a plurality of trajectories for the autonomous vehicle, the plurality of trajectories being distributed over a predetermined plurality of behaviors according to an initial distribution; acquiring an indication of a target distribution of the plurality of trajectories over the predetermined plurality of behaviors; determining a difference between the initial and target distributions, thereby identifying behaviors, for which additional trajectories are to be generated for adjusting the initial distribution of the plurality of trajectories to the target distribution; generating, using the plurality of trajectories in the original dataset and the additional trajectories, an augmented dataset having the target distribution of trajectories over the predetermined plurality of behaviors; and training a motion planner model to simulate motion of the autonomous vehicle based on the augmented plurality of trajectories.