Reinforcement Learning Reward Deriver for Autonomous Vehicle Action Planning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning in automated driving is limited in its application scope, restricting flexible learning for action plans in autonomous vehicles.

Innovation Solution

A learning device and method that includes a planner to generate vehicle actions and a reward deriver to evaluate feedback from simulators or actual environments using diverse reward functions, optimizing rewards for vehicle actions based on speed, acceleration, and risk assessment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning is applied to automated driving using conventional methods, then the system can determine vehicle operations, but the application scope is limited and flexible learning for action plans cannot be performed

Engineering Contradiction:
Improveflexible learning capabilityVSAvoidlearning system structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the reward evaluation into multiple independent reward functions, each responsible for evaluating different aspects of vehicle actions (e.g., safety, comfort, efficiency). This segmentation enables flexible learning by allowing each reward function to independently optimize specific action plan components while maintaining overall system coherence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal learning framework that can handle multiple types of evaluations through a single multi-functional reward deriver. This reward deriver integrates various reward functions that can evaluate different action aspects, making the system adaptable to diverse automated driving scenarios without requiring separate learning systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple reward functions with different evaluation characteristics are applied, then flexible learning is enabled, but the complexity of reward derivation increases

Engineering Contradiction:
Improveevaluation flexibilityVSAvoidreward derivation process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges multiple individual rewards derived from different reward functions into a single comprehensive reward signal. The reward deriver combines these individual rewards through weighted summation or other aggregation methods, enabling flexible multi-aspect evaluation while maintaining a unified and manageable reward derivation process that does not become prohibitively complex.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11498574B2Learning device, learning method, and storage medium
Publication Date: 2022.11.15 HONDA MOTOR CO LTD
  • US11498574B2 patent drawing
  • US11498574B2 patent drawing
  • US11498574B2 patent drawing

AI summary

A learning device includes a planner configured to generate information indicating an action of a vehicle, and a reward deriver configured to derive a plurality of individual rewards obtained by evaluating each of a plurality of pieces of information to be evaluated, which include feedback information obtained from a simulator or an actual environment by inputting information based on the information indicating the action of the vehicle to the simulator or the actual environment, and derive a reward for the action of the vehicle on the basis of the plurality of individual rewards. The planner performs reinforcement learning that optimizes the reward derived by the reward deriver.