Vehicle Control Data Generation via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle control systems require extensive manual effort and skilled labor to adapt operation amounts of electronic devices on vehicles, making it time-consuming and inefficient, especially when distinguishing between self-driving and manual driving modes.

Innovation Solution

A vehicle control data generation method using reinforcement learning to adjust the relationship between vehicle states and action variables, providing different rewards for self-driving and manual driving modes to optimize propelling force generation, thereby reducing manual intervention and improving efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual adaptation of filter configuration is used to set operation amounts of electronic devices, then the control precision can be adjusted to meet vehicle state requirements, but the time and labor required for adaptation increases significantly

Engineering Contradiction:
Improvecontrol precisionVSAvoidadaptation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs self-adaptation through reinforcement learning, where the execution device automatically learns and adjusts the relationship defining data between vehicle states and action variables without requiring skilled workers to manually configure filters. The system serves itself by using reward signals from vehicle operation to automatically optimize control parameters.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention changes the approach from manual parameter adjustment to automated parameter learning. The relationship defining data, which defines the relationship between vehicle state and action variable, is automatically updated through reinforcement learning by inputting state values, action variable values, and reward values to an update map, thereby optimizing control precision without manual intervention.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If a unified reward system is used for both self-driving and manual driving modes, then the system simplicity is maintained, but the ability to learn mode-specific optimal control behavior is reduced

Engineering Contradiction:
Improvereward system complexityVSAvoidmode-specific learning capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The reward system is segmented into mode-specific reward functions. The execution device provides different rewards based on the driving mode (self-driving or manual driving) and the meeting of respective standards. This segmentation allows the reinforcement learning system to learn distinct optimal control behaviors for each mode while maintaining a unified overall architecture.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If extensive manual adaptation is performed to optimize electronic device operation, then the control accuracy improves, but the productivity of the vehicle setup process decreases

Engineering Contradiction:
Improvecontrol accuracyVSAvoidvehicle setup efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system eliminates the need for manual adaptation by implementing self-learning through reinforcement learning. The execution device automatically optimizes the relationship defining data by processing vehicle state data, action variable data, and reward data, thereby achieving high control accuracy without sacrificing setup productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The invention replaces the mechanical process of manual filter configuration with an automated computational system. Instead of skilled workers manually adjusting parameters, the system uses reinforcement learning algorithms to automatically learn and optimize the relationship between vehicle states and control actions, significantly improving setup efficiency while maintaining or enhancing control accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11679784B2Vehicle control data generation method, vehicle controller, vehicle control system, vehicle learning device, vehicle control data generation device, and memory medium
Publication Date: 2023.06.20 TOYOTA JIDOSHA KK
  • US11679784B2 patent drawing
  • US11679784B2 patent drawing
  • US11679784B2 patent drawing

AI summary

A vehicle control data generation method is provided. A self-driving mode of a vehicle automatically generates a command value of a propelling force produced by a propelling force generator independently of an accelerator operation. An execution device provides a greater reward based on an obtained state of the vehicle when a standard of a characteristic of the vehicle is met than when the standard is not met. Providing the reward changes a reward that is provided when the characteristic of the vehicle is a predetermined characteristic in the self-driving mode such that the changed reward differs from a reward that is provided when the characteristic of the vehicle is the predetermined characteristic in a manual driving mode.