Vehicle Control Data Generation via Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for adapting the throttle valve opening degree of an internal combustion engine in vehicles require extensive manual effort and man-hours, as they need to be set appropriately based on the accelerator pedal operation amount, leading to inefficiencies in vehicle control.
Innovation Solution
A method using reinforcement learning to generate vehicle control data by defining a relationship between the vehicle's state and action variables, where a processor calculates rewards for different traveling control modes, updating the relationship data to optimize the operation of electronic equipment such as the throttle valve and ignition timing, reducing the need for manual adaptation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If filter processing is applied to accelerator pedal operation amount to control throttle valve opening degree, then the throttle valve can be operated based on processed operation data, but extensive manual adaptation time and expert intervention are required
Solution Approach 1:
The control device automatically performs adaptation by executing the reinforcement learning program to generate relationship definition data between accelerator pedal operation amount and throttle valve opening degree. This self-service mechanism eliminates the need for expert intervention and manual adaptation processes, allowing the system to autonomously optimize control parameters based on filter processing results.
Solution Approach 2:
The manual adaptation process performed by experts is replaced by an automated computational system using reinforcement learning algorithms. The program automatically processes accelerator pedal operation data, applies filter processing, and generates optimized control relationships, substituting the mechanical expert adjustment process with an automated information processing system.
2Manufacturing precision
If expert adaptation is performed for each vehicle state, then appropriate control parameters can be set, but significant man-hours are consumed in the process
Solution Approach 1:
The control device autonomously generates optimized control parameters by executing reinforcement learning programs that process vehicle state data and accelerator pedal operation amounts. This self-service capability maintains high manufacturing precision in control parameter settings while eliminating the need for expert intervention, thereby significantly improving adaptation efficiency and productivity.
Solution Approach 2:
The system dynamically adjusts control parameters including throttle valve opening degree, ignition timing, and injection amount based on processed accelerator pedal operation data and vehicle state. These parameter changes are automatically optimized through reinforcement learning, maintaining precision while improving efficiency by eliminating manual adaptation processes.
3Reliability
If relationship definition data is updated using reinforcement learning with different rewards for different traveling control modes, then the expected return on reward is increased, but the system complexity increases
Solution Approach 1:
The control device uses a universal reinforcement learning framework that handles multiple traveling control modes (economic mode, sport mode, snow mode) through a single integrated system. The relationship definition data structure and reward calculation mechanism are designed to be mode-agnostic, automatically adapting to different modes without requiring separate complex subsystems, thereby maintaining reliability while managing system complexity.
Solution Approach 2:
The reward calculation mechanism dynamically adjusts rewards based on the selected traveling control mode, with different reward values for economic mode, sport mode, and snow mode. This dynamic adaptation allows the system to optimize control parameters for each mode while using a single flexible framework, improving reliability without significantly increasing structural complexity.
Data Source
AI summary
Provided is a method of generating vehicle control data. The method is applied to a vehicle configured to select one of a plurality of traveling control modes and is executed by a processor in a state in which relationship definition data defining a relationship between a state of the vehicle and an action variable as a variable relating to an operation of electronic equipment in the vehicle is stored in a memory. The method includes operation processing for operating the electronic equipment, acquisition processing for acquiring a detection value of a sensor configured to detect the state of the vehicle, reward calculation processing for providing reward, and update processing for updating the relationship definition data.


