AV Cost Function Learning with Pareto Dominance Margins
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine-learned cost function learning methods for autonomous vehicles are susceptible to demonstration noise and suboptimality, making it difficult to optimize motion plans and leading to suboptimal performance, such as increased lateral nudging and jerkiness.
Innovation Solution
The use of Pareto dominance-based imitation learning techniques to minimize subdominance by adding a margin to the cost function, ensuring that the learned plans dominate human demonstrations by a significant margin, thereby reducing the influence of outlier demonstrations and improving the optimization of the initial cost function.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional cost function learning methods are used to optimize motion plans, then the system can learn from human demonstrations, but the system becomes highly sensitive to suboptimal outliers and demonstration noise
Solution Approach 1:
The patent converts the harmful effect of suboptimal outliers and demonstration noise into a beneficial filtering mechanism. By intentionally designing the cost function to penalize deviations from demonstrated behavior, the system causes outliers to naturally produce higher cost values, allowing them to be identified and filtered out as non-Pareto optimal solutions. This transforms the problem of noise sensitivity into a solution where noise automatically marks itself for rejection.
Solution Approach 2:
The patent applies preliminary filtering of demonstrations before final cost function optimization. By pre-processing the demonstration data to identify and remove obvious outliers based on initial cost evaluations, the system prepares cleaner training data that reduces the impact of noise on subsequent optimization. This preliminary action prevents noise from corrupting the final learned cost function.
2Productivity
If maximum margin planning or maximum entropy inverse reinforcement learning is used, then the system can learn optimal policies, but the system becomes highly sensitive to suboptimal outliers requiring careful data filtering
Solution Approach 1:
The patent implements self-service data filtering where the learning system automatically identifies and filters out suboptimal outliers through its own cost function evaluations. The system uses its learned cost function to assess each demonstration, automatically marking suboptimal ones as non-Pareto optimal without requiring external data cleaning intervention. This self-service mechanism eliminates the need for manual or complex automated data filtering pipelines.
Solution Approach 2:
The patent incorporates feedback loops where the cost function continuously evaluates demonstrations and adjusts its learned parameters based on Pareto optimality assessments. This feedback mechanism allows the system to automatically adapt to the quality of incoming demonstrations, reinforcing learning from high-quality data while naturally downweighting or rejecting suboptimal outliers through iterative refinement.
3Manufacturing precision
If the system tries to make demonstrated behavior optimal relative to all other possible behaviors, then the system can achieve better motion plans, but the system requires extensive data cleaning to remove suboptimal outliers
Solution Approach 1:
The patent performs preliminary cost function evaluation and Pareto optimality assessment on demonstrations before they are used for final learning. By pre-screening demonstrations to identify obvious suboptimal outliers based on initial cost calculations, the system prepares a filtered dataset that requires minimal further cleaning. This preliminary action significantly reduces the time and computational resources needed for subsequent data cleaning and refinement steps.
Data Source
AI summary
Techniques for improving the performance of an autonomous vehicle (AV) are described herein. A system can determine a plan for the AV in a driving scenario that optimizes an initial cost function of a control algorithm of the AV. The system can obtain data describing an observed human driving path in the driving scenario. Additionally, the system can determine for each cost dimension in the plurality of cost dimensions, a quantity that compares the estimated cost to the observed cost of the observed human driving path. Moreover, the system can determine a function of a sum of the quantities determined for each cost dimension in the plurality of cost dimensions. Subsequently, the system can use an optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities.


