Vehicle Controller Reinforcement Learning Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle control systems require extensive manual effort and man-hours to adapt the operation of electronic devices based on vehicle state, and fail to efficiently update relationships between vehicle states and action variables, especially after functional recovery measures are taken.

Innovation Solution

A vehicle controller with an execution device and memory device that uses reinforcement learning to update relationship defining data, providing rewards for optimal operations and switching back to initial data post-functional recovery to maintain performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual adaptation of operation amounts by skilled workers is used, then the relationship between vehicle state and action variable can be set appropriately, but the process requires a great number of man-hours

Engineering Contradiction:
Improveadaptation precisionVSAvoidman-hours
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The controller automatically adapts operation amounts by itself using reinforcement learning, eliminating the need for skilled workers to manually adjust parameters. The system learns optimal operation amounts through self-experience by accumulating rewards based on vehicle state and operational outcomes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical adjustment processes with an automated electronic learning system. The reinforcement learning mechanism substitutes the mechanical process of skilled workers physically adjusting components with an electronic algorithm that automatically optimizes operation amounts through data processing and reward calculation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If reinforcement learning updates relationship defining data continuously, then the system adapts to vehicle deterioration, but the data becomes inappropriate after functional recovery measures are taken

Engineering Contradiction:
Improveadaptation to deteriorationVSAvoidperformance after recovery
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically switches between different relationship defining data sets based on the operational state. When functional recovery is detected, the system transitions from using learned data (adapted to deterioration) to initial data (optimized for recovered state), allowing the system to adapt its behavior according to the current condition of components

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from functional recovery detection to trigger data switching. When recovery measures are detected, the system receives feedback that component state has changed, and accordingly switches from deterioration-adapted data to initial data, ensuring optimal performance after recovery

Inventive Principle:
Principle #23Feedback

3Productivity

If relationship defining data is updated based on reinforcement learning, then manual adaptation work is reduced, but the complexity of the control system increases

Engineering Contradiction:
Improveadaptation efficiencyVSAvoidcontrol system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The controller integrates multiple functions into a single device: it performs both the original control operations and the reinforcement learning processes. The controller executes obtaining processes, operation processes, reward calculating processes, and update processes all within one integrated system, eliminating the need for separate adaptation equipment

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11453375B2Vehicle controller, vehicle control system, vehicle learning device, vehicle control method, and memory medium
Publication Date: 2022.09.27 TOYOTA JIDOSHA KK
  • US11453375B2 patent drawing
  • US11453375B2 patent drawing
  • US11453375B2 patent drawing

AI summary

A vehicle controller, a vehicle control system, a vehicle learning device, a vehicle control method, and a memory medium are provided. A switching process switches relationship defining data used in an operation process to post-measure data, when a detection process detects that a functional recovery measure has been taken. The switching process includes a process that uses, as the post-measure data, initial data that is the relationship defining data of a state before an update process is executed as the vehicle travels.