Vehicle Control Data Generation via Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing vehicle control systems require extensive manual effort and skilled labor to adapt operation amounts of electronic devices on vehicles, making it time-consuming and inefficient, especially when distinguishing between self-driving and manual driving modes.
Innovation Solution
A vehicle control data generation method using reinforcement learning to adjust the relationship between vehicle states and action variables, providing different rewards for self-driving and manual driving modes to optimize propelling force generation, thereby reducing manual intervention and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual adaptation of filter configuration is used to set operation amounts of electronic devices, then the control precision can be adjusted to meet vehicle state requirements, but the time and labor required for adaptation increases significantly
Solution Approach 1:
The system performs self-adaptation through reinforcement learning, where the execution device automatically learns and adjusts the relationship defining data between vehicle states and action variables without requiring skilled workers to manually configure filters. The system serves itself by using reward signals from vehicle operation to automatically optimize control parameters.
Solution Approach 2:
The invention changes the approach from manual parameter adjustment to automated parameter learning. The relationship defining data, which defines the relationship between vehicle state and action variable, is automatically updated through reinforcement learning by inputting state values, action variable values, and reward values to an update map, thereby optimizing control precision without manual intervention.
2Device complexity
If a unified reward system is used for both self-driving and manual driving modes, then the system simplicity is maintained, but the ability to learn mode-specific optimal control behavior is reduced
Solution Approach 1:
The reward system is segmented into mode-specific reward functions. The execution device provides different rewards based on the driving mode (self-driving or manual driving) and the meeting of respective standards. This segmentation allows the reinforcement learning system to learn distinct optimal control behaviors for each mode while maintaining a unified overall architecture.
3Measurement precision
If extensive manual adaptation is performed to optimize electronic device operation, then the control accuracy improves, but the productivity of the vehicle setup process decreases
Solution Approach 1:
The system eliminates the need for manual adaptation by implementing self-learning through reinforcement learning. The execution device automatically optimizes the relationship defining data by processing vehicle state data, action variable data, and reward data, thereby achieving high control accuracy without sacrificing setup productivity.
Solution Approach 2:
The invention replaces the mechanical process of manual filter configuration with an automated computational system. Instead of skilled workers manually adjusting parameters, the system uses reinforcement learning algorithms to automatically learn and optimize the relationship between vehicle states and control actions, significantly improving setup efficiency while maintaining or enhancing control accuracy.
Data Source
AI summary
A vehicle control data generation method is provided. A self-driving mode of a vehicle automatically generates a command value of a propelling force produced by a propelling force generator independently of an accelerator operation. An execution device provides a greater reward based on an obtained state of the vehicle when a standard of a characteristic of the vehicle is met than when the standard is not met. Providing the reward changes a reward that is provided when the characteristic of the vehicle is a predetermined characteristic in the self-driving mode such that the changed reward differs from a reward that is provided when the characteristic of the vehicle is the predetermined characteristic in a manual driving mode.


