Autonomous Driver Agents Self-Training via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current autonomous vehicle control systems require external supervision and labeled data, making them non-scalable and labor-intensive, especially for achieving automotive reliability in various environments.

Innovation Solution

The system employs Deep Reinforcement Learning (DRL) algorithms to process driving experiences and generate policies that control autonomous vehicles, using driving environment processors and driver agents to learn and optimize neural network parameters without relying on external supervision or labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If deep learning systems are used for autonomous vehicle control, then automation level is improved, but scalability deteriorates due to reliance on labeled data

Engineering Contradiction:
Improveautomation levelVSAvoidscalability
Core Design Contradiction:
Extent of automationVSProductivity

Solution Approach 1:

The system enables autonomous vehicles to self-train by collecting their own driving experiences and using reinforcement learning to improve their policies without external supervision. Each vehicle serves as its own training data source, eliminating the need for manually labeled datasets and enabling scalable deployment across multiple vehicles.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where driving experiences are collected, policies are updated based on this feedback, and improved policies are deployed back to the vehicles. This closed-loop reinforcement learning approach allows the system to automatically improve over time without external intervention.

Inventive Principle:
Principle #23Feedback

2Reliability

If external supervision and labeled data are used for training, then reliability is improved, but loss of time increases due to labor-intensive data creation

Engineering Contradiction:
Improvetraining reliabilityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system eliminates the need for external supervisors and manual data labeling by enabling vehicles to autonomously collect, process, and learn from their own driving experiences in real-time, dramatically reducing both time and labor requirements while maintaining training quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system enables continuous learning during normal vehicle operation rather than requiring separate training phases. Vehicles continuously collect driving experiences and update their policies in real-time, transforming idle training time into productive operational time.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If neural networks are trained to achieve automotive reliability in all environments, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improveautomotive reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the learning process into modular components: experience collection modules, policy learning modules, and deployment modules. This segmentation allows each component to be independently optimized and validated, reducing overall system complexity while maintaining comprehensive reliability across diverse environments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10845815B2Systems, methods and controllers for an autonomous vehicle that implement autonomous driver agents and driving policy learners for generating and improving policies based on collective driving experiences of the autonomous driver agents
Publication Date: 2020.11.24 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US10845815B2 patent drawing
  • US10845815B2 patent drawing
  • US10845815B2 patent drawing

AI summary

Systems and methods are provided autonomous driving policy generation. The system can include a set of autonomous driver agents, and a driving policy generation module that includes a set of driving policy learner modules for generating and improving policies based on the collective experiences collected by the driver agents. The driver agents can collect driving experiences to create a knowledge base. The driving policy learner modules can process the collective driving experiences to extract driving policies. The driver agents can be trained via the driving policy learner modules in a parallel and distributed manner to find novel and efficient driving policies and behaviors faster and more efficiently. Parallel and distributed learning can enable accelerated training of multiple autonomous intelligent driver agents.