Vehicle Control Device Dynamic Learning Update
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Reinforcement learning in vehicle control systems can result in significant deviations from optimal values when drivetrain devices are subject to constraints or abnormalities, such as high or low operating oil temperatures, leading to inappropriate learning results.
Innovation Solution
Implementing a vehicle control device with a storage device and an executing device that restricts the updating of relation-defining data when drivetrain devices are under constraints or abnormalities, ensuring a smaller updating amount or setting it to zero, thereby preventing significant deviations in learning results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reinforcement learning is performed continuously to improve drivetrain operation, then learning accuracy improves, but learning results deviate significantly when constraints or abnormalities occur
Solution Approach 1:
The patent changes the parameter of updating amount from a fixed value to a variable that adapts based on drivetrain state. When the drivetrain is in a normal state, a larger updating amount is applied to improve learning accuracy. When constraints or abnormalities are detected, the updating amount is reduced or set to zero to prevent learning result deviation, thus resolving the contradiction between learning accuracy and result stability.
2Productivity
If the updating amount is increased to accelerate learning convergence, then learning speed improves, but learning results become unstable under constraint conditions
Solution Approach 1:
The patent makes the updating amount dynamic rather than static. The updating amount is adjusted in real-time based on the drivetrain state and whether constraints or abnormalities are present. This dynamic adjustment allows fast learning during normal operations while preventing instability when constraints occur, resolving the contradiction between learning speed and result stability.
3Adaptability or versatility
If reinforcement learning adapts to various driving conditions, then adaptability improves, but computational load increases
Solution Approach 1:
The patent applies partial updating action by selectively updating the relation-defining data only when necessary (when drivetrain state is normal and no constraints are present). Instead of continuously updating under all conditions, the system performs updating only in appropriate states, reducing computational load while maintaining adaptability to various driving conditions through accumulated learning over time.
Data Source
AI summary
A vehicle control device includes: a storage device that stores relation-defining data that is data for defining a relation between a state of a vehicle and an action variable; and an executing device configured to acquire the state, operate a drivetrain device based on a value of the action variable, derive a reward such that the reward is larger when the state of the drivetrain device based on the acquired state satisfies a predetermined criterion, perform an updating of the relation-defining data using an updating map, and restrict the updating of the relation-defining data such that an updating amount of the relation-defining data is smaller when the drivetrain device is subject to a predetermined restriction.


