Reinforcement Learning Training Using Comparable Vehicle Behavior

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning agents for autonomous systems require a large number of training data records, which are time-consuming to collect, especially in automated driving scenarios where a test vehicle experiences limited relevant traffic situations.

Innovation Solution

A device and method that detect and analyze the behavior of objects comparable to the autonomous system in its environment, using sensor data from cameras, radar, and lidar to generate an environment model and train the reinforcement learning agent based on the detected behavior, without requiring extensive data from the autonomous system itself.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If training data is collected using a test vehicle in automated driving scenarios, then the training data reflects real-world traffic situations, but the number of relevant training data records is limited due to the restricted number of traffic situations experienced in a given time

Engineering Contradiction:
Improverealism of training dataVSAvoidnumber of training data records
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies copying by using simulation environments to generate virtual training data that replicates real-world traffic situations. Instead of relying solely on physical test vehicles, the system creates copies of real traffic scenarios in simulated environments, allowing the reinforcement learning agent to train on numerous varied situations without the time constraints of physical testing.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies preliminary action by pre-generating and storing training data in simulation environments before actual deployment. Traffic scenarios, including edge cases and rare situations, are pre-simulated and saved as training data records, allowing the system to have access to a comprehensive dataset without needing to encounter every situation in the real world during limited testing periods.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a test vehicle is used to collect training data records in real-world scenarios, then the data quality is high, but the time required to collect sufficient training data is very long

Engineering Contradiction:
Improvedata qualityVSAvoidtraining data collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses copying to create virtual replicas of real-world traffic scenarios in simulation environments. These simulated copies maintain the essential characteristics and complexity of real traffic situations while allowing for rapid data generation. The simulation can replay and vary scenarios indefinitely, providing high-quality training data without the time constraints of physical testing.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes by systematically varying conditions in the simulation environment (weather, traffic density, vehicle behavior, road conditions) to generate diverse training data. By changing simulation parameters, the system can efficiently explore a wide range of traffic situations and generate comprehensive training datasets much faster than physical testing would allow.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240037447A1Training a Reinforcement Learning Agent to Control an Autonomous System
Publication Date: 2024.02.01 BAYERISCHE MOTOREN WERKE AG
  • US20240037447A1 patent drawing
  • US20240037447A1 patent drawing

AI summary

One aspect of the invention relates to a device for training a reinforcement learning agent to control an autonomous system, wherein the device is designed to detect the environment of the autonomous system, to detect at least one object in the environment of the autonomous system that can be compared to the autonomous system, to detect a behaviour of the at least one object that can be compared to the autonomous system, and to train the reinforcement learning agent in accordance with the detected behaviour of the at least one object that can be compared to the autonomous system.