Vehicle Control Data Generation via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing methods for adapting the throttle valve opening degree of an internal combustion engine in vehicles require extensive manual effort and man-hours, as they need to be set appropriately based on the accelerator pedal operation amount, leading to inefficiencies in vehicle control.

Innovation Solution

A method using reinforcement learning to generate vehicle control data by defining a relationship between the vehicle's state and action variables, where a processor calculates rewards for different traveling control modes, updating the relationship data to optimize the operation of electronic equipment such as the throttle valve and ignition timing, reducing the need for manual adaptation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If filter processing is applied to accelerator pedal operation amount to control throttle valve opening degree, then the throttle valve can be operated based on processed operation data, but extensive manual adaptation time and expert intervention are required

Engineering Contradiction:
Improvethrottle valve controlVSAvoidadaptation time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The control device automatically performs adaptation by executing the reinforcement learning program to generate relationship definition data between accelerator pedal operation amount and throttle valve opening degree. This self-service mechanism eliminates the need for expert intervention and manual adaptation processes, allowing the system to autonomously optimize control parameters based on filter processing results.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual adaptation process performed by experts is replaced by an automated computational system using reinforcement learning algorithms. The program automatically processes accelerator pedal operation data, applies filter processing, and generates optimized control relationships, substituting the mechanical expert adjustment process with an automated information processing system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If expert adaptation is performed for each vehicle state, then appropriate control parameters can be set, but significant man-hours are consumed in the process

Engineering Contradiction:
Improvecontrol parameter accuracyVSAvoidadaptation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The control device autonomously generates optimized control parameters by executing reinforcement learning programs that process vehicle state data and accelerator pedal operation amounts. This self-service capability maintains high manufacturing precision in control parameter settings while eliminating the need for expert intervention, thereby significantly improving adaptation efficiency and productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts control parameters including throttle valve opening degree, ignition timing, and injection amount based on processed accelerator pedal operation data and vehicle state. These parameter changes are automatically optimized through reinforcement learning, maintaining precision while improving efficiency by eliminating manual adaptation processes.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If relationship definition data is updated using reinforcement learning with different rewards for different traveling control modes, then the expected return on reward is increased, but the system complexity increases

Engineering Contradiction:
Improvecontrol optimizationVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The control device uses a universal reinforcement learning framework that handles multiple traveling control modes (economic mode, sport mode, snow mode) through a single integrated system. The relationship definition data structure and reward calculation mechanism are designed to be mode-agnostic, automatically adapting to different modes without requiring separate complex subsystems, thereby maintaining reliability while managing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The reward calculation mechanism dynamically adjusts rewards based on the selected traveling control mode, with different reward values for economic mode, sport mode, and snow mode. This dynamic adaptation allows the system to optimize control parameters for each mode while using a single flexible framework, improving reliability without significantly increasing structural complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11654915B2Method of generating vehicle control data, vehicle control device, and vehicle control system
Publication Date: 2023.05.23 TOYOTA JIDOSHA KK
  • US11654915B2 patent drawing
  • US11654915B2 patent drawing
  • US11654915B2 patent drawing

AI summary

Provided is a method of generating vehicle control data. The method is applied to a vehicle configured to select one of a plurality of traveling control modes and is executed by a processor in a state in which relationship definition data defining a relationship between a state of the vehicle and an action variable as a variable relating to an operation of electronic equipment in the vehicle is stored in a memory. The method includes operation processing for operating the electronic equipment, acquisition processing for acquiring a detection value of a sensor configured to detect the state of the vehicle, reward calculation processing for providing reward, and update processing for updating the relationship definition data.