Autonomous Driving DRL Training With Adaptive Human Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep reinforcement learning (DRL) methods for autonomous driving face challenges in learning efficiency and require extensive human guidance, which is exhausting and limited by the need for expert-level demonstrations, leading to high manpower costs and data-processing inefficiencies.

Innovation Solution

The proposed Hug-DRL method incorporates real-time human guidance into the training process using an actor-critic architecture with a priority experience replay buffer and adaptively weighted human guidance, allowing for efficient learning and performance improvement without requiring expert-level human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing DRL methods use real-time human guidance to improve learning performance, then the performance of the learned policies is improved, but long-term supervision becomes exhausting for human participants and learning efficiency decreases

Engineering Contradiction:
Improveperformance of learned policiesVSAvoidlearning efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic adjustment of human guidance intensity through an adaptive guidance mechanism that modulates the guidance strength parameter based on the agent's current performance level. As the agent improves, the system automatically reduces human guidance intensity, creating a dynamic training process that maintains effectiveness while reducing long-term human supervision requirements

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system applies partial human guidance only when necessary during the training process, rather than continuous full guidance. The adaptive mechanism determines the optimal level of guidance intervention at each training step, providing just enough human input to maintain learning effectiveness while minimizing human workload and maximizing autonomous learning

Inventive Principle:
Principle #16Partial or excessive action

2Reliability

If existing DRL methods require expert-level demonstrations to ensure data quality, then the performance improvement is ideal, but costly manpower and shortage of professionals limit practical usage

Engineering Contradiction:
Improvequality of collected dataVSAvoidfeasibility in large-scale applications
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent replaces expensive expert-level human demonstrators with readily available non-expert participants. The adaptive guidance mechanism compensates for the lower initial quality of demonstrations from non-experts, enabling effective training using inexpensive, easily recruitable human subjects who do not require specialized training or expertise

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The system enables non-expert human participants to provide effective guidance through the adaptive framework that automatically adjusts guidance parameters. The mechanism allows ordinary users to contribute meaningful training data without requiring expert knowledge, as the system adapts to their input quality and adjusts accordingly

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If the training process is slowed down to adapt to human driver's physical reactions, then real-world adaptability is improved, but the extensive training process decreases learning efficiency and leads to negative subjective feelings

Engineering Contradiction:
Improveadaptability to human driver reactionsVSAvoidlearning efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The training speed is dynamically adjusted based on the phase of training and the agent's performance. During early stages, the system may slow down to accommodate human reaction patterns, but as training progresses and the agent improves, the system automatically increases training speed, creating an optimized balance between realism and efficiency

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240160945A1Autonomous driving methods and systems
Publication Date: 2024.05.16 NANYANG TECH UNIV
  • US20240160945A1 patent drawing
  • US20240160945A1 patent drawing
  • US20240160945A1 patent drawing

AI summary

A method of training a deep reinforcement learning model for autonomous control of a machine, such as autonomous vehicles, the model being configured to output, by a policy network, an agent action in response to input of state information and a value function, the agent action representing a control signal for the machine. The method comprises minimizing a loss function of the policy network; wherein the loss function of the policy network comprises an autonomous guidance component and a human guidance component (human intervention); and wherein the autonomous guidance component is zero when the state information is indicative of a human input signal.