Reinforcement Learning Scenario Variation for Rare Autonomous Driving Events

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in preparing for anomalous human behavior and rare scenarios due to unpredictable human actions, leading to inefficient training and poorly trained components, with existing techniques failing to effectively handle a wide range of situations and requiring repetitive training.

Innovation Solution

The implementation of reinforcement learning techniques to score autonomous vehicle performance across various behaviors and scenarios, using a reward system to modify machine-learned models, thereby increasing the vehicle's interaction capabilities and training efficiency by reducing repetitive training and discovering latent relations between training scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional simulation training is used to prepare autonomous vehicles for rare scenarios, then the vehicle can handle some anomalous behavior, but the training is inefficient and requires repetitive training on the same scenarios

Engineering Contradiction:
Improveability to handle anomalous behaviorVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic scenario selection by continuously monitoring performance metrics and automatically adjusting which scenarios are selected for training. The system transitions from static, repetitive training to dynamic adaptation where training scenarios change based on identified performance deficiencies, thereby improving training efficiency while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of scenario selection by using performance metrics to determine which scenarios to train on next. Instead of fixed scenario sets, the patent dynamically adjusts scenario parameters based on performance data, allowing the system to focus training on rare and problematic scenarios rather than repetitively training on common scenarios.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If reinforcement learning is used to train autonomous vehicles on all possible scenarios, then the vehicle can handle a broader range of situations, but the training time increases significantly

Engineering Contradiction:
Improverange of situations handledVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies partial action by selectively training on specific scenarios that are identified as rare, problematic, or performance-deficient rather than training on all possible scenarios. The performance metric system identifies which partial set of scenarios needs attention, reducing training time while maintaining adaptability to critical situations.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs self-service by automatically monitoring its own performance metrics and autonomously selecting which scenarios require training. The autonomous vehicle system uses its performance data to identify its own deficiencies and selectively train on those specific scenarios, reducing unnecessary training time while improving adaptability to problematic situations.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If repetitive training on the same scenarios is conducted, then the vehicle becomes proficient in those scenarios, but it fails to discover latent relations between different training scenarios

Engineering Contradiction:
Improveperformance proficiencyVSAvoidlatent relations between scenarios
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent implements universality by using a unified performance metric system that evaluates across multiple scenario types and identifies transferable performance patterns. The system detects latent relations between different scenarios by analyzing performance metrics across diverse situations, allowing proficiency in one scenario to inform performance in related scenarios, thereby preventing information loss.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses feedback from performance metrics to identify latent relations between scenarios. By continuously monitoring performance across different scenario types and providing feedback on performance patterns, the system can detect underlying relationships and transfer learnings between scenarios, maintaining measurement precision while discovering cross-scenario patterns.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12037013B1Automated reinforcement learning scenario variation and impact penalties
Publication Date: 2024.07.16 ZOOX INC
  • US12037013B1 patent drawing
  • US12037013B1 patent drawing
  • US12037013B1 patent drawing

AI summary

Automating reinforcement learning for autonomous vehicles may include assigning a probability with a scenario and varying that probability based at least in part on changes in performance by the autonomous vehicle associated with that scenario. The amount of time and computational bandwidth required to train a machine-learned component of an autonomous vehicle and the accuracy of the machine-learned component may be improved by determining a reward for performance of the autonomous vehicle in a scenario based at least in part on an severity metric. The impact severity metric may be determined based at least in part on a velocity, angle, and/or interaction area associated with the impact.