Agent ECU Reinforced Learning for In-Vehicle Operation Proposals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing in-vehicle information provision systems increase driver burden by requiring mode changes from manual to sound input and offer limited operation burden reduction, with current systems only simplifying speech dialog functions without significant improvement.
Innovation Solution
An information providing device equipped with an agent electronic control unit that uses reinforced learning to construct state and action spaces, calculate probability distributions, and compute dispersion degrees, allowing for definitive or trial-and-error operation proposals based on driver intent, thereby reducing driver input burden and enhancing operation proposal accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the system switches from manual input mode to sound input mode to reduce driver burden, then the ease of operation is improved, but the device complexity increases due to mode switching mechanisms
Solution Approach 1:
The system automatically detects driver state and vehicle conditions to autonomously determine the appropriate input mode, eliminating the need for manual mode switching by the driver. The operation proposal generation unit automatically selects between manual and sound input modes based on real-time conditions, allowing the system to serve itself in mode selection.
Solution Approach 2:
The system dynamically adapts the operation mode based on real-time driver state and vehicle conditions. The operation proposal generation unit flexibly switches between manual input mode and sound input mode according to the calculated appropriateness, making the system adaptable rather than static in its operational characteristics.
2Ease of operation
If the system uses speech dialog functions to simplify operation entrance, then the ease of operation is improved, but the productivity remains limited as it only achieves functions similar to existing systems
Solution Approach 1:
The system incorporates a reinforcement learning unit that accumulates history data on driver responses to operation proposals and uses this feedback to continuously improve the appropriateness calculation. The system learns from past interactions to refine its operation proposals, moving beyond static speech dialog functions to adaptive, improving performance over time.
Solution Approach 2:
The operation proposal generation unit proactively generates operation proposals before the driver needs to perform actions, based on predicted driver intent from accumulated history data. This preliminary generation of context-aware proposals reduces the cognitive load on the driver by presenting relevant options in advance.
3Measurement precision
If the system accumulates and learns history data on driver responses to improve operation proposal accuracy, then the measurement precision is improved, but the loss of time increases due to data processing requirements
Solution Approach 1:
The system processes only the most relevant features from accumulated history data that are necessary for determining driver intent, rather than analyzing all possible data points. The reinforcement learning unit focuses on extracting key patterns from history data to calculate appropriateness, performing partial processing that balances accuracy with efficiency.
Data Source
AI summary
An information providing device includes an agent ECU that sets a reward function through the use of history data on a response, from a driver, to an operation proposal for an in-vehicle component, and calculates a probability distribution of performance of each of actions constructing an action space in each of states constructing a state space, through reinforced learning based on the reward function. The agent ECU calculates a dispersion degree of the probability distribution. The agent ECU makes a trial-and-error operation proposal to select a target action from a plurality of candidates and output the target action when the dispersion degree of the probability distribution is equal to or larger than a threshold, and makes a definitive operation proposal to fix and output a target action when the value of the dispersion degree of the probability distribution is smaller than the threshold.


