Autonomous Driving Parameter Tuning With Preference-Guided RL

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning methods for autonomous driving robots require retraining for different scenarios and are inefficient in adapting to various parameters, especially when interacting with humans, as they rely on fixed parameters and extensive preference data collection.

Innovation Solution

The development of a deep reinforcement learning-based autonomous driving system that uses a neural network with a fully-connected layer and gated recurrent unit (GRU) to learn and optimize autonomous driving parameters through simulation, allowing for adaptive policy learning and preference modeling with a small amount of data, utilizing Bayesian neural networks for uncertainty-based query generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning is used for autonomous driving with fixed parameters, then the system can be trained initially, but it requires retraining for different scenarios and cannot adapt to various parameters efficiently

Engineering Contradiction:
Improveadaptability to various parametersVSAvoidretraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies dynamics by making the driving parameters changeable and adaptive rather than fixed. The system dynamically adjusts autonomous driving parameters based on user preferences and environmental conditions, allowing the same trained model to adapt to different scenarios without retraining. This is achieved through a preference model that learns user preferences and generates optimized parameters on-demand.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements parameter changes by introducing a preference model that optimizes autonomous driving parameters based on learned user preferences. Instead of retraining the entire reinforcement learning model for different scenarios, the system changes parameters such as driving style, safety margins, and responsiveness by querying the preference model, which generates optimized parameter sets based on user feedback and environmental context.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extensive preference data collection is used to optimize autonomous driving parameters, then user preference alignment can be improved, but data collection becomes time-consuming and resource-intensive

Engineering Contradiction:
Improvepreference alignment accuracyVSAvoidamount of preference data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training the preference model on a small initial dataset of user preferences. This pre-trained model can then generate optimized driving parameters without requiring extensive additional data collection. The system performs preliminary learning to capture general user preference patterns, enabling subsequent parameter optimization with minimal new data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The preference model serves itself by generating optimized driving parameters based on its learned preferences without requiring continuous external data input. Once trained on a small dataset, the model can independently query and generate parameter recommendations, reducing the need for ongoing data collection and manual optimization efforts.

Inventive Principle:
Principle #25Self-service

3Reliability

If traditional reinforcement learning with fixed parameters is used, then the system structure remains simple, but it cannot handle unpredictable real-world scenarios effectively

Engineering Contradiction:
Improvehandling of unpredictable scenariosVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary preference model that acts as a mediator between the reinforcement learning agent and the environment. This preference model translates user preferences and environmental conditions into optimized driving parameters, allowing the system to handle unpredictable scenarios without fundamentally changing the core reinforcement learning architecture. The intermediary layer adds adaptability while maintaining the original system's simplicity.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If autonomous driving parameters are manually adjusted for different use cases, then customization is possible, but the process becomes complex and time-consuming

Engineering Contradiction:
Improvecustomization capabilityVSAvoidparameter adjustment ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements feedback by using a preference model that learns from user preferences and automatically generates optimized driving parameters. Instead of manual adjustment, the system collects user feedback on driving behavior and uses this feedback to automatically tune parameters for different use cases. This feedback loop enables easy customization without requiring users to manually adjust complex parameters.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20220229435A1Method and system for optimizing reinforcement-learning-based autonomous driving according to user preferences
Publication Date: 2022.07.21 NAVER CORP
  • US20220229435A1 patent drawing
  • US20220229435A1 patent drawing
  • US20220229435A1 patent drawing

AI summary

A method for optimizing autonomous driving includes applying different autonomous driving parameters to a plurality of robot agents in a simulation through an automatic setting by means of the system or a direct setting by means of a manager, so that the robot agents learn robot autonomous driving; and optimizing the autonomous driving parameters by using preference data for the autonomous driving parameters.