Deep Reinforcement Learning for Autonomous Parking Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in navigating diverse parking scenarios, including parallel parking, due to unpredictable environmental factors and the need for precise control, which is costly and requires extensive training.

Innovation Solution

A system utilizing deep reinforcement learning, specifically combining a deep Q-network and an asynchronous advantage actor-critic network, trains an autonomous vehicle to navigate efficiently and accurately by processing sensor data and learning from mistakes in a simulated environment, enabling real-time control of steering, acceleration, and braking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional step-by-step automatic parking systems are used, then collision-free orientation is ensured, but the system cannot handle diverse parking scenarios and unpredictable environmental factors

Engineering Contradiction:
Improvehandling diverse parking scenariosVSAvoidcollision-free operation
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent replaces traditional rule-based control systems with deep reinforcement learning neural networks that learn optimal parking maneuvers through trial-and-error training. The neural network processes sensor data and environmental factors to determine steering angle, accelerator, and brake controls, enabling adaptation to diverse parking scenarios while maintaining safety through learned collision avoidance strategies

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the control parameters from fixed rule-based decisions to dynamic neural network outputs that continuously adjust steering angle, accelerator position, and brake application based on real-time sensor inputs and learned patterns from training data, allowing flexible adaptation to various parking conditions

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extensive trial-and-error training is conducted to improve navigation accuracy, then learning from mistakes is enabled, but training time and computational resources increase

Engineering Contradiction:
Improvenavigation accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses simulated environments to create virtual copies of real-world parking scenarios for training the neural networks. This allows extensive trial-and-error learning to occur in silico without consuming real vehicle time or risking actual collisions, significantly reducing training time while maintaining high navigation accuracy through realistic simulation physics and sensor models

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary training in simulated environments before deploying to real vehicles. By pre-training the neural networks on diverse virtual parking scenarios, the system accumulates learning experience in advance, reducing the need for extensive real-world trial-and-error and accelerating deployment readiness

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11613249B2Automatic navigation using deep reinforcement learning
Publication Date: 2023.03.28 FORD GLOBAL TECH LLC
  • US11613249B2 patent drawing
  • US11613249B2 patent drawing
  • US11613249B2 patent drawing

AI summary

A method for training an autonomous vehicle to reach a target location. The method includes detecting the state of an autonomous vehicle in a simulated environment, and using a neural network to navigate the vehicle from an initial location to a target destination. During the training phase, a second neural network may reward the first neural network for a desired action taken by the autonomous vehicle, and may penalize the first neural network for an undesired action taken by the autonomous vehicle. A corresponding system and computer program product are also disclosed and claimed herein.