Vehicle Controller Reinforcement Learning for Throttle Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle control systems require significant manual effort to adapt the operation of electronic devices, such as throttle valves, to various vehicle states, leading to inefficiencies and potential inappropriate operations.

Innovation Solution

A vehicle controller that employs processing circuitry and storage devices to utilize reinforcement learning, updating relationship specifying data to optimize the operation of electronic devices based on vehicle states, action variables, and rewards, thereby reducing manual adaptation work and preventing inappropriate operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual adaptation of filter parameters is performed to set appropriate throttle valve operation, then control accuracy is improved, but the amount of work required increases significantly

Engineering Contradiction:
Improvecontrol accuracyVSAvoidamount of work
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-adaptation through reinforcement learning, where the filter parameters are automatically adjusted based on observed vehicle states and control outcomes. The controller learns optimal filter settings through trial and error, eliminating the need for manual parameter tuning while maintaining high control accuracy.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements a feedback mechanism where the actual vehicle response to throttle valve operations is observed and used to update the filter parameters. The reinforcement learning algorithm continuously refines the parameters based on the difference between expected and actual vehicle behavior, enabling automatic adaptation without manual intervention.

Inventive Principle:
Principle #23Feedback

2Loss of time

If reinforcement learning is used to automatically update relationship specifying data, then manual adaptation work is reduced, but the complexity of the control system increases

Engineering Contradiction:
Improvemanual adaptation workVSAvoidcontrol system complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The reinforcement learning framework provides a universal solution that can adapt to various vehicle states and electronic device operations through a single unified approach. The same learning mechanism handles different filter parameters, vehicle conditions, and control scenarios, reducing the need for multiple specialized adaptation systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically changes filter parameters based on learned relationships between vehicle states and optimal control actions. Instead of fixing parameters manually, the reinforcement learning algorithm continuously adjusts them according to the current operating conditions, enabling automatic adaptation while maintaining system manageability.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the action variable is restricted to prevent inappropriate operations, then system reliability is improved, but the flexibility of control is reduced

Engineering Contradiction:
Improvesystem reliabilityVSAvoidflexibility of control
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts the scope of action variables based on the current vehicle state and learned knowledge. The reinforcement learning algorithm determines appropriate restrictions in real-time, allowing maximum flexibility when the vehicle is in stable, well-understood conditions while imposing necessary constraints when approaching unsafe or inappropriate operating regions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system pre-defines boundary conditions and constraint rules for action variables based on safety requirements and physical limitations. These preliminary restrictions ensure that even during exploration phases of reinforcement learning, the system operates within safe parameters, maintaining reliability while allowing flexibility within bounded regions.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11603111B2Vehicle controller, vehicle control system, and learning device for vehicle
Publication Date: 2023.03.14 TOYOTA JIDOSHA KK
  • US11603111B2 patent drawing
  • US11603111B2 patent drawing
  • US11603111B2 patent drawing

AI summary

A vehicle controller includes processing circuitry and a storage device. The storage device stores relationship specifying data that specifies a relationship between a vehicle state and an action variable. The processing circuitry is configured to execute an obtaining process obtaining the vehicle state, an operating process operating an electronic device based on a value of the action variable, a reward calculation process assigning a reward based on the vehicle state, an updating process updating the relationship specifying data using the vehicle state, the value of the action variable, and the reward as inputs to an update mapping. When a value of the action variable designated by the relationship specifying data is a first value, a process in which the operating process operates the electronic device in accordance with the first value is executable in a first situation and is not executable in a second situation.