Vehicle Control Device Reinforcement Learning Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Optimizing the operation amount of a throttle valve in an internal combustion engine requires extensive manual effort and time, as existing methods lack efficient automation for setting appropriate relationships between vehicle states and action variables.
Innovation Solution
A vehicle control device with a processor and storage unit that uses reinforcement learning to update relationship prescription data, determining optimal operation based on sensor detection values and reward calculations, thereby reducing manual effort and computation load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual optimization of filtering parameters is performed, then control precision is improved, but development time and labor cost increase significantly
Solution Approach 1:
The system performs self-optimization by automatically learning optimal filtering parameters through reinforcement learning. The neural network autonomously adjusts control parameters based on vehicle state data and reward signals, eliminating the need for manual expert optimization while achieving high control precision.
Solution Approach 2:
The system implements feedback mechanisms where the neural network receives reward signals based on control performance, uses this feedback to update its internal models, and continuously optimizes filtering parameters. This closed-loop feedback enables automatic refinement of control precision without manual intervention.
2Productivity
If reinforcement learning update process is executed continuously, then control optimization is improved, but computation load increases excessively
Solution Approach 1:
The system executes reinforcement learning update processes periodically rather than continuously. The processor determines appropriate update timing based on vehicle operating conditions, executing updates only when computation load is within acceptable thresholds, thus balancing optimization effectiveness with computational resource consumption.
Solution Approach 2:
The system dynamically adjusts the frequency and intensity of reinforcement learning updates based on real-time computation load monitoring and vehicle state assessment. The update process adapts its behavior to current system conditions, reducing computation during high-load periods while maintaining optimization during low-load periods.
Data Source
AI summary
A vehicle control device includes a storage device and a processor. The storage device is configured to store relationship prescription data that prescribe a relationship between a state of a vehicle and an action variable that is a variable related to an operation of an electronic device in the vehicle. The processor is configured to calculate a reward corresponding to the operation of the electronic device. The processor is configured to update the relationship prescription data using, as inputs to updated mapping determined in advance, the state of the vehicle that is based on a detection value that is acquired, a value of the action variable that is used to operate the electronic device, and the reward corresponding to the operation of the electronic device when a computation load on the processor is equal to or less than a predetermined load.


