Reinforcement Learning Reward Deriver for Autonomous Vehicle Action Planning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reinforcement learning in automated driving is limited in its application scope, restricting flexible learning for action plans in autonomous vehicles.
Innovation Solution
A learning device and method that includes a planner to generate vehicle actions and a reward deriver to evaluate feedback from simulators or actual environments using diverse reward functions, optimizing rewards for vehicle actions based on speed, acceleration, and risk assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If reinforcement learning is applied to automated driving using conventional methods, then the system can determine vehicle operations, but the application scope is limited and flexible learning for action plans cannot be performed
Solution Approach 1:
The patent segments the reward evaluation into multiple independent reward functions, each responsible for evaluating different aspects of vehicle actions (e.g., safety, comfort, efficiency). This segmentation enables flexible learning by allowing each reward function to independently optimize specific action plan components while maintaining overall system coherence.
Solution Approach 2:
The patent creates a universal learning framework that can handle multiple types of evaluations through a single multi-functional reward deriver. This reward deriver integrates various reward functions that can evaluate different action aspects, making the system adaptable to diverse automated driving scenarios without requiring separate learning systems for each function.
2Adaptability or versatility
If multiple reward functions with different evaluation characteristics are applied, then flexible learning is enabled, but the complexity of reward derivation increases
Solution Approach 1:
The patent merges multiple individual rewards derived from different reward functions into a single comprehensive reward signal. The reward deriver combines these individual rewards through weighted summation or other aggregation methods, enabling flexible multi-aspect evaluation while maintaining a unified and manageable reward derivation process that does not become prohibitively complex.
Data Source
AI summary
A learning device includes a planner configured to generate information indicating an action of a vehicle, and a reward deriver configured to derive a plurality of individual rewards obtained by evaluating each of a plurality of pieces of information to be evaluated, which include feedback information obtained from a simulator or an actual environment by inputting information based on the information indicating the action of the vehicle to the simulator or the actual environment, and derive a reward for the action of the vehicle on the basis of the plurality of individual rewards. The planner performs reinforcement learning that optimizes the reward derived by the reward deriver.


