Autonomous Driving Policy Control With Adaptive Exploration Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic driving systems using deep reinforcement learning face instability and safety risks due to uncontrolled exploration noise, which is not correlated with environmental states, leading to unpredictable decisions and difficulty in determining issues with neural networks or external disturbances.

Innovation Solution

A method that initializes a deep-reinforcement-learning automatic driving decision system with both noiseless and noisy strategic networks, adjusts noise parameters within a disturbance threshold, and optimizes system parameters to generate an optimized noisy strategy for improved exploration and stability, ensuring that environmental states and driving strategies are fully considered in decision-making.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If exploration noise is added to selected actions in each decision-making process to facilitate full exploration of action space, then exploration capability is improved, but decision stability and safety deteriorate due to randomness and uncontrollability of noise size

Engineering Contradiction:
Improveexploration capabilityVSAvoiddecision stability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the parameter of noise from fixed Gaussian distribution to adaptive noise whose magnitude is dynamically adjusted based on the difference between noisy and noiseless strategy outputs. This allows the noise level to be controlled and adapted to specific situations, resolving the contradiction between exploration capability and decision stability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a feedback mechanism where the noise magnitude is determined by comparing the outputs of noisy and noiseless strategy networks. The difference between these outputs feeds back to control the noise injection level, ensuring that noise remains within acceptable bounds while still enabling exploration.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If random exploration noise is added to actions, then action space exploration is enhanced, but the ability to determine whether problems originate from neural networks or external disturbances deteriorates

Engineering Contradiction:
Improveaction space explorationVSAvoidproblem source identification
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces a noiseless strategy network as an intermediary reference system. By comparing the outputs of the noisy strategy network with the noiseless strategy network, the system can identify whether deviations are caused by neural network issues or external disturbances, thus solving the problem source identification difficulty.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If Gaussian distribution sampling is used for exploration noise, then exploration is simplified, but noise becomes uncontrollable and unpredictable

Engineering Contradiction:
Improvenoise generation simplicityVSAvoidnoise controllability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent transforms the static Gaussian noise into a dynamic noise mechanism where the noise magnitude is continuously adjusted based on real-time strategy differences. This dynamic adaptation maintains ease of implementation while achieving controllable and predictable noise behavior.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11887009B2Autonomous driving control method, apparatus and device, and readable storage medium
Publication Date: 2024.01.30 INSPUR SUZHOU INTELLIGENT TECH CO LTD
  • US11887009B2 patent drawing
  • US11887009B2 patent drawing

AI summary

The present application discloses an automatic driving control method. In the method, parameters are optimally set by using a noisy and noiseless dual-strategy network, identical vehicle traffic environment state information is input into the noisy and noiseless dual-strategy network, a motion space perturbation threshold is set by using a noiseless strategy network as a comparison and a benchmark so as to adaptively adjust noise parameters, and motion noise is indirectly added by adaptively injecting noise into a strategy network parameter space, such that exploration of an environment and a motion space by a deep reinforcement learning algorithm may be effectively improved, automatic driving exploration performance and stability based on deep reinforcement learning is improved, and full consideration of influence of an environment state and driving strategies in vehicle decision-making and motion selection is ensured, thereby improving the stability and safety of an automatic vehicle.