AV Cost Function Learning with Pareto Dominance Margins

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learned cost function learning methods for autonomous vehicles are susceptible to demonstration noise and suboptimality, making it difficult to optimize motion plans and leading to suboptimal performance, such as increased lateral nudging and jerkiness.

Innovation Solution

The use of Pareto dominance-based imitation learning techniques to minimize subdominance by adding a margin to the cost function, ensuring that the learned plans dominate human demonstrations by a significant margin, thereby reducing the influence of outlier demonstrations and improving the optimization of the initial cost function.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional cost function learning methods are used to optimize motion plans, then the system can learn from human demonstrations, but the system becomes highly sensitive to suboptimal outliers and demonstration noise

Engineering Contradiction:
Improvecost function optimization accuracyVSAvoidsensitivity to demonstration noise
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent converts the harmful effect of suboptimal outliers and demonstration noise into a beneficial filtering mechanism. By intentionally designing the cost function to penalize deviations from demonstrated behavior, the system causes outliers to naturally produce higher cost values, allowing them to be identified and filtered out as non-Pareto optimal solutions. This transforms the problem of noise sensitivity into a solution where noise automatically marks itself for rejection.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent applies preliminary filtering of demonstrations before final cost function optimization. By pre-processing the demonstration data to identify and remove obvious outliers based on initial cost evaluations, the system prepares cleaner training data that reduces the impact of noise on subsequent optimization. This preliminary action prevents noise from corrupting the final learned cost function.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If maximum margin planning or maximum entropy inverse reinforcement learning is used, then the system can learn optimal policies, but the system becomes highly sensitive to suboptimal outliers requiring careful data filtering

Engineering Contradiction:
Improvelearning efficiencyVSAvoiddata cleaning complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-service data filtering where the learning system automatically identifies and filters out suboptimal outliers through its own cost function evaluations. The system uses its learned cost function to assess each demonstration, automatically marking suboptimal ones as non-Pareto optimal without requiring external data cleaning intervention. This self-service mechanism eliminates the need for manual or complex automated data filtering pipelines.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback loops where the cost function continuously evaluates demonstrations and adjusts its learned parameters based on Pareto optimality assessments. This feedback mechanism allows the system to automatically adapt to the quality of incoming demonstrations, reinforcing learning from high-quality data while naturally downweighting or rejecting suboptimal outliers through iterative refinement.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If the system tries to make demonstrated behavior optimal relative to all other possible behaviors, then the system can achieve better motion plans, but the system requires extensive data cleaning to remove suboptimal outliers

Engineering Contradiction:
Improvemotion plan qualityVSAvoiddata cleaning time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary cost function evaluation and Pareto optimality assessment on demonstrations before they are used for final learning. By pre-screening demonstrations to identify obvious suboptimal outliers based on initial cost calculations, the system prepares a filtered dataset that requires minimal further cleaning. This preliminary action significantly reduces the time and computational resources needed for subsequent data cleaning and refinement steps.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250368223A1Systems and Methods for Pareto Domination-Based Learning
Publication Date: 2025.12.04 UATC LLC
  • US20250368223A1 patent drawing
  • US20250368223A1 patent drawing
  • US20250368223A1 patent drawing

AI summary

Techniques for improving the performance of an autonomous vehicle (AV) are described herein. A system can determine a plan for the AV in a driving scenario that optimizes an initial cost function of a control algorithm of the AV. The system can obtain data describing an observed human driving path in the driving scenario. Additionally, the system can determine for each cost dimension in the plurality of cost dimensions, a quantity that compares the estimated cost to the observed cost of the observed human driving path. Moreover, the system can determine a function of a sum of the quantities determined for each cost dimension in the plurality of cost dimensions. Subsequently, the system can use an optimization algorithm to adjust one or more weights of the plurality of weights applied to the plurality of cost dimensions to optimize the function of the sum of the quantities.