Reinforcement learning method

The reinforcement learning method addresses the issue of discontinuous crank angle inputs by updating the relational specification model to optimize torque command values, ensuring proper learning and efficient crankshaft stopping.

JP2026001945APending Publication Date: 2026-01-08TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024099555
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

The crank angle sensor in vehicles detects crank angles with discontinuous values, which can lead to improper learning of the relationship defining model due to the input of such values.

Method used

A reinforcement learning method is employed to repeatedly apply torque from the motor generator to stop the crankshaft within a specified angle range, using a relational specification model that updates based on the reward for selecting torque command values that minimize actions and align with the specified angle range.

Benefits of technology

This approach prevents improper learning of the relationship defining model due to discontinuous crank angle inputs and enables efficient stopping of the crankshaft within the specified angle range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001945000001_ABST
    Figure 2026001945000001_ABST
Patent Text Reader

Abstract

To suppress that learning of a relationship regulation model is not appropriately executed.SOLUTION: In the reinforcement learning method, a computer repeatedly attempts to stop a crankshaft within a specified angle range by applying torque from a motor generator to the crankshaft. The computer stores a relationship defining model that outputs an index value indicating a value of an action corresponding to a plurality of candidates for a torque command value when a state variable is input. The state variables include the engine speed, a first component obtained by converting the crank angle by a sine function, and a second component obtained by converting the crank angle by a cosine function. The computer selects a torque command value based on the acquired state variable and the relationship defining model. The computer calculates the reward such that the profit increases as the crankshaft is stopped with fewer actions. The computer updates the relationship defining model such that a torque command value that increases the profit is selected.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a reinforcement learning method. [Background technology]

[0002] The vehicle disclosed in Patent Document 1 includes an internal combustion engine, a motor generator, a transmission, a crank angle sensor, and a control device. The crank angle sensor detects the crank angle, which is the angular position of the crankshaft of the internal combustion engine. The control device controls the internal combustion engine, the motor generator, and the transmission by outputting control signals to the internal combustion engine, the motor generator, and the transmission.

[0003] The control device also stores a trained relational specification model that has been trained by machine learning. The control device calculates the engine rotation speed, which is the rotation speed of the crankshaft, based on the crank angle acquired from the crank angle sensor. The control device then inputs the engine rotation speed, crank angle, etc. into the relational specification model, and outputs variables that indicate vehicle noise. The control device then estimates a sensory level related to the vehicle noise based on the variables output from the relational specification model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-076162 Summary of the Invention [Problem to be solved by the invention]

[0005] A crank angle sensor for a vehicle such as that disclosed in Patent Document 1 generally detects the crank angle as a value in a range greater than or equal to 0° and less than 360°. Therefore, when the crankshaft is rotating, the crank angle sensor detects crank angles with discontinuous values, such as . . . , 358°, 359°, 0°, 1°, . . . Therefore, the relationship defining model of Patent Document 1 receives crank angles with discontinuous values. When learning such a relationship defining model, there is a risk that the learning of the relationship defining model will not be performed appropriately due to the input of crank angles with discontinuous values ​​to the model. [Means for solving the problem]

[0006] A reinforcement learning method for solving the above-mentioned problems is a reinforcement learning method by a computer for a vehicle having an internal combustion engine and a motor generator connected to a crankshaft of the internal combustion engine, the method repeatedly performing trials to stop the rotating crankshaft within a predetermined specified angle range by applying torque from the motor generator to the crankshaft, the computer having an execution device and a storage device, the storage device storing a relational specification model that outputs an index value indicating the value of an action corresponding to each of a plurality of candidates for a torque command value of the motor generator in the trial when a state variable indicating a state of the vehicle is input, the state variables including an engine rotation speed which is the rotation speed of the crankshaft, a first component obtained by converting a crank angle which is the angular position of the crankshaft using a sine function, and a second component obtained by converting the crank angle using a cosine function, The device executes the following operations for each trial: acquires the state variables; selects the torque command value corresponding to one of the plurality of index values ​​based on the plurality of index values ​​output by inputting the acquired state variables into the relationship specification model; controls the motor generator according to the selected torque command value; acquires a stop crank angle, which is the crank angle when the crankshaft stops, after controlling the motor generator; calculates the reward for the selected torque command value so that the closer the acquired stop crank angle is to the specified angle range and the fewer actions are required to stop the crankshaft, the greater the profit, which is the total sum of rewards when the trial is completed; and updates the relationship specification model so that the torque command value that increases the profit is selected based on the calculated reward and the selected torque command value. [Effects of the Invention]

[0007] According to the above configuration, it is possible to prevent a situation in which learning of the relationship defining model is not properly executed due to input of a crank angle with a discontinuous value. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a schematic diagram of a vehicle and a learning device. [Figure 2] FIG. 2 is a flowchart showing the learning control. [Figure 3] FIG. 3 is an explanatory diagram of the relationship definition model. DETAILED DESCRIPTION OF THE INVENTION

[0009] <Vehicle Overview> An embodiment of the present invention will now be described with reference to Figures 1 to 3. First, a schematic configuration of a vehicle 100 will be described.

[0010] 1, the vehicle 100 includes an internal combustion engine 10, a power split mechanism 20, an automatic transmission 30, a plurality of drive wheels 69, a first motor generator 61, and a second motor generator 62. The internal combustion engine 10, the first motor generator 61, and the second motor generator 62 are the drive sources of the vehicle 100.

[0011] The internal combustion engine 10 has a crankshaft 11 and four cylinders (not shown). Each cylinder undergoes an intake stroke, a compression stroke, a combustion stroke, and an exhaust stroke as the crankshaft 11 rotates twice. The crankshaft 11 is the output shaft of the internal combustion engine 10. The crankshaft 11 is connected to a power split mechanism 20.

[0012] The power split mechanism 20 is a planetary gear mechanism having a sun gear S, a ring gear RG, and a carrier C. The carrier C of the power split mechanism 20 is connected to the crankshaft 11. The sun gear S is connected to a rotating shaft 61A of the first motor generator 61. In other words, the rotating shaft 61A of the first motor generator 61 is connected to the crankshaft 11 of the internal combustion engine 10 via the power split mechanism 20. The ring gear RG has a ring gear shaft RA. The ring gear shaft RA is an output shaft of the ring gear RG. The ring gear shaft RA is connected to a rotating shaft 62A of the second motor generator 62. The ring gear shaft RA is also connected to the automatic transmission 30. The automatic transmission 30 is connected to left and right drive wheels 69 via a differential gear (not shown). An example of the automatic transmission 30 is a stepped automatic transmission. Therefore, the automatic transmission 30 changes the gear ratio by changing the gear position.

[0013] When the internal combustion engine 10 operates and torque from the internal combustion engine 10 is input to the carrier C of the power split mechanism 20, the torque is split between the sun gear S side and the ring gear RG side. Furthermore, when the first motor generator 61 operates as an electric motor and torque from the first motor generator 61 is input to the sun gear S of the power split mechanism 20, the torque is split between the carrier C side and the ring gear RG side. Therefore, the first motor generator 61 can apply torque to the crankshaft 11 of the internal combustion engine 10 via the power split mechanism 20.

[0014] When torque from second motor generator 62 is input to ring gear shaft RA as a result of second motor generator 62 operating as an electric motor, the torque is transmitted to automatic transmission 30. Furthermore, when torque from the drive wheels 69 side is input to second motor generator 62 via ring gear shaft RA, second motor generator 62 functions as a generator. As a result, second motor generator 62 can generate regenerative braking force for vehicle 100.

[0015] The vehicle 100 is equipped with a first inverter 66, a second inverter 67, and a battery 68 as devices for exchanging electric power. The first inverter 66 adjusts the amount of electric power exchanged between the first motor generator 61 and the battery 68. The second inverter 67 adjusts the amount of electric power exchanged between the second motor generator 62 and the battery 68.

[0016] As shown in Fig. 1, vehicle 100 is equipped with a crank angle sensor 71, an accelerator operation amount sensor 72, and a vehicle speed sensor 73. Crank angle sensor 71 detects crank angle SC, which is the angular position of crankshaft 11. Here, crank angle sensor 71 detects crank angle SC as a value in the range of 0° or more and less than 360°. Accelerator operation amount sensor 72 detects accelerator operation amount ACC, which is the amount of operation of the accelerator pedal operated by the driver. Vehicle speed sensor 73 detects vehicle speed SP, which is the speed of vehicle 100.

[0017] The vehicle 100 is equipped with a control device 90. The control device 90 acquires various information from a crank angle sensor 71, an accelerator operation amount sensor 72, and a vehicle speed sensor 73. The control device 90 calculates an engine rotation speed NE, which is the rotation speed of the crankshaft 11, based on the crank angle SC.

[0018] The control device 90 includes an execution device 91 and a storage device 92. An example of the execution device 91 is a CPU. The storage device 92 includes a read-only ROM, a readable / writable volatile RAM, and a readable / writable non-volatile storage. The storage device 92 stores various programs and data in advance. Specifically, the storage device 92 stores a control program 92A in advance as one of the various programs. The storage device 92 also stores a relationship definition model M in advance as one of the various data. The relationship definition model M is a model used to stop the rotating crankshaft 11 within a predetermined specified angle range RR by applying torque from the first motor-generator 61 to the crankshaft 11. The relationship definition model M describes, in a format recognizable by the execution device 91, the relationship between predetermined input data and a Q value, which is an index value indicating the value of an action corresponding to each of multiple candidates for the torque command value CT of the first motor-generator 61. The relationship defining model M receives a plurality of pieces of input data and outputs a Q value corresponding to each of a plurality of candidates for the torque command value CT of the first motor generator 61. In this embodiment, the relationship defining model M stored in the storage device 92 is a model that has been learned by machine learning. In other words, the storage device 92 stores the relationship defining model M generated by learning control, which will be described later. The relationship defining model M will be described in detail later. The execution device 91 executes a control program 92A stored in the storage device 92 to perform various processes, which will be described later.

[0019] The execution unit 91 of the control device 90 controls the internal combustion engine 10, the first motor generator 61, the second motor generator 62, the automatic transmission 30, etc. Specifically, the execution unit 91 calculates a vehicle required driving force, which is a required value of driving force necessary for the vehicle 100 to travel, based on the accelerator operation amount ACC and the vehicle speed SP. The execution unit 91 determines a torque distribution among the internal combustion engine 10, the first motor generator 61, and the second motor generator 62 based on the vehicle required driving force. The execution unit 91 controls the output of the internal combustion engine 10 and the power running and regeneration of the first motor generator 61 and the second motor generator 62 based on the torque distribution among the internal combustion engine 10, the first motor generator 61, and the second motor generator 62. Specifically, the execution unit 91 controls the internal combustion engine 10 by outputting a control signal to the internal combustion engine 10. The execution unit 91 also controls the first motor generator 61 via the first inverter 66 by outputting a control signal to the first inverter 66. Furthermore, the execution device 91 controls the second motor generator 62 via the second inverter 67 by outputting a control signal to the second inverter 67 .

[0020] Furthermore, the execution unit 91 calculates a target gear position, which is a target gear position for the automatic transmission 30, based on the vehicle speed SP and the vehicle required driving force. The execution unit 91 outputs a control signal to the automatic transmission 30 based on the target gear position. As a result, the gear position of the automatic transmission 30 is controlled.

[0021] The vehicle 100 is equipped with a connector 80. The connector 80 is connected to a control device 90. The connector 80 is equipped with a general-purpose terminal group that enables two-way communication. The connector 80 enables communication between the control device 90 and other devices. In this embodiment, the control device 90 can communicate with a learning device 200, which will be described later, via the connector 80.

[0022] <Overview of the learning device> Next, a description will be given of the learning device 200 that learns the relational specification model M. The learning device 200 is a device that is used in the development stage of the vehicle 100, for example.

[0023] As shown in FIG. 1, the learning device 200 includes a microphone 210 , an acceleration sensor 220 , a connector 240 , an input device 250 , a display 260 , and a computer 290 .

[0024] The microphone 210 detects sound pressure NS as noise generated in the vehicle 100. In this embodiment, the microphone 210 is attached to a predetermined position in the interior of the vehicle 100. An example of the predetermined position is the driver's seat of the vehicle 100. The microphone 210 is connected to the computer 290.

[0025] The acceleration sensor 220 is a so-called three-axis sensor. That is, the acceleration sensor 220 can detect longitudinal acceleration GX, lateral acceleration GY, and vertical acceleration GZ. The longitudinal acceleration GX is acceleration along the longitudinal axis of the vehicle 100. The lateral acceleration GY is acceleration along the lateral axis of the vehicle 100. The vertical acceleration GZ is acceleration along the vertical axis of the vehicle 100. In this embodiment, the acceleration sensor 220 is attached to a predetermined position in the interior of the vehicle 100. An example of the predetermined position is the driver's seat of the vehicle 100. The acceleration sensor 220 is connected to the computer 290.

[0026] Input device 250 is a device for inputting commands from an engineer or the like using learning device 200 into computer 290. Input device 250 is connected to computer 290. Input device 250 is, for example, a keyboard, a pointing device, or the like. Display 260 is a device for transmitting information from computer 290 to an engineer or the like. Display 260 is capable of displaying various types of information. Display 260 is connected to computer 290.

[0027] The computer 290 includes an execution device 291 and a storage device 292. An example of the execution device 291 is a CPU. The storage device 292 includes a read-only ROM, a readable / writable volatile RAM, and a readable / writable non-volatile storage. The storage device 292 stores various programs and various data in advance. Specifically, the storage device 292 stores a control program 292A in advance as one of the various programs. The storage device 292 also stores a relationship definition model M in advance as one of the various data. The relationship definition model M stored in the storage device 292 is a model of a learning process by machine learning. The execution device 291 executes the control program 292A stored in the storage device 292 to perform various processes described below. In other words, the execution device 291 executes the control program 292A stored in the storage device 292 to perform various processes related to a reinforcement learning method. An example of the computer 290 is a personal computer.

[0028] Connector 240 is connected to computer 290. Connector 240 has a group of general-purpose terminals that enable bidirectional communication. Connector 240 enables communication between computer 290 and other devices. In this embodiment, for example, when an engineer connects connector 240 of learning device 200 to connector 80 of vehicle 100 during learning of relationship specification model M, computer 290 of learning device 200 becomes able to communicate with control device 90 of vehicle 100. Then, execution device 291 of computer 290 can access control device 90 to obtain various information from control device 90.

[0029] <Learning control> Next, with reference to FIG. 2, the learning control executed by the computer 290 of the learning device 200 will be described. This learning control is control related to reinforcement learning, in which attempts are repeatedly made to stop the rotating crankshaft 11 within a predetermined specified angle range RR by applying torque from the first motor-generator 61 to the crankshaft 11. Note that the learning control is control executed by the computer 290 of the learning device 200, for example, during the development stage of the vehicle 100. Note that, as described above, each of the four cylinders of the internal combustion engine 10 undergoes an intake stroke, a compression stroke, a combustion stroke, and an exhaust stroke when the crankshaft 11 rotates twice. Therefore, when the crankshaft 11 rotates twice, there are four specified angle ranges RR corresponding to the four cylinders. For example, the four specified angle ranges RR are predetermined angle ranges based on 180°, 360°, 540°, and 720°, respectively. As a specific example, the specified angle range RR with 180° as the reference is an angle range of 150° or more and 210° or less.

[0030] In this learning control, the Q value, which is an index value indicating the value of an action, is updated. Here, the Q value is expressed by the following equation (1).

[0031]

number

[0032] Here, Q(s, a) is the expected future profit when action a is selected in state s. "r" is the reward. The subscript "t" in state s, action a, and reward r in the above formula (1) is a number indicating one step in the trial. When state s changes after the action is decided, the step number t, which is the number indicating the step, is incremented by one. Note that the value of step number t at the start of the trial is "0".

[0033] In the following, subscripts will be written with "_". Therefore, the reward r_t+1 in equation (1) is the reward obtained when state s becomes state s_t+1 by selecting action a_t in state s_t. Also, "α" is the learning rate. "γ" is the discount rate. The learning rate α is a value greater than "0" and less than "1". The discount rate γ is a value greater than "0" and less than "1". Also, maxQ(s_t+1,a) is the Q value maximized by selecting action a. In this section, action a indicates the action that maximizes Q(s_t+1,a_t+1) among the actions a_t+1 that can be taken in state s_t+1.

[0034] For example, the execution device 291 of the computer 290 starts learning control when an engineer performs an operation via the input device 250 requesting the execution of learning control, with the necessary condition being that the connector 240 of the learning device 200 and the connector 80 of the vehicle 100 are connected.

[0035] As shown in FIG. 2, when the execution unit 291 of the computer 290 starts learning control, it executes the processing of step S11. In step S11, the execution unit 291 operates the internal combustion engine 10 by outputting a control signal to the internal combustion engine 10 via the control device 90. Specifically, the execution unit 291 controls the internal combustion engine 10 so that the engine rotation speed NE becomes a predetermined reference rotation speed. An example of the reference rotation speed is about 1000 rpm. The reference rotation speed is determined as the engine rotation speed NE during idling of the internal combustion engine 10 in the vehicle 100. After step S11, the execution unit 291 advances the processing to step S21.

[0036] In step S21, the execution unit 291 acquires the state s of the vehicle 100. Specifically, the execution unit 291 acquires the engine speed NE, a first component SC1, and a second component SC2 as state variables indicating the state s of the vehicle 100. At this time, for example, when the number of trial steps at the time of processing in step S21 is "t," the execution unit 291 acquires the state variables in state s_t. Here, the first component SC1 is a value obtained by converting the crank angle SC using a sine function. Therefore, the first component SC1 is expressed as "sin(SC)." Furthermore, the second component SC2 is a value obtained by converting the crank angle SC using a cosine function. Therefore, the second component SC2 is expressed as "cos(SC)." Note that, for example, if the execution unit 291 acquires the state variables in state s_t in the current step S21, in the next step S21, the execution unit 291 acquires the state variables in state s_t+1. After step S21, the execution device 291 advances the process to step S31.

[0037] In step S31, the execution device 291 selects a torque command value CT corresponding to one of the multiple Q values ​​based on the multiple Q values ​​output by inputting the state variables acquired in step S21 into the relational definition model M.

[0038] In this embodiment, the relationship definition model M is a model generated by DQN (Deep Q-Network), which is a method for approximately calculating a Q value. In DQN, a Q value is estimated using a multi-layer neural network. Therefore, as shown in FIG. 3, the relationship definition model M is a neural network that receives a state variable indicating a state s as input and outputs a value Q(s, a), which is a Q value corresponding to the number of selectable actions a. This neural network receives m values ​​X, which are the state variables, and outputs n Q values ​​corresponding to n actions a. In FIG. 3, the m input values ​​in the number of trial steps t are indicated as X_1 to X_m.

[0039] In this embodiment, the state variables input to the relationship specification model M are the state variables acquired in step S21, i.e., the engine speed NE, the first component SC1, and the second component SC2. Therefore, "m" is "3." The Q value output by the relationship specification model M indicates the Q value when a specific action a is selected in the input state s, i.e., the value of the action a. The number of types of Q values ​​output by the relationship specification model M is the same as the number of candidates for the torque command value CT of the first motor generator 61. Here, examples of the candidates are five: a first command value, a second command value, a third command value, a fourth command value, and a fifth command value. Specifically, the third command value is a torque command value CT that is the same as the torque command value CT at a point in time a predetermined period PA before. In other words, the third command value is a value for maintaining the torque command value CT. The predetermined period PA is the same as the cycle at which engine stop control, which will be described later, is executed. An example of the predetermined period PA is 10 msec. The fourth command value is a value that is larger than the third command value by a predetermined constant value. The fifth command value is a value that is larger than the fourth command value by a predetermined constant value. The second command value is a value that is smaller than the third command value by a predetermined constant value. The first command value is a value that is smaller than the second command value by a predetermined constant value. In other words, "n" is "5".

[0040] The neural network shown in Figure 3 has an input layer, multiple hidden layers, and an output layer. This neural network is a multi-layered neural network in which each node in each layer multiplies the input of the previous layer by a weight w and adds a bias b to obtain an output that has passed through an activation function. Note that in Figure 3, the transmission lines connecting nodes in adjacent layers are not shown.

[0041] The structure of a neural network is specified by information such as the weights w and bias b in each layer, the activation function, and the order of layers. During training, the weights w, which are variable values ​​within the neural network, are updated.

[0042] Specifically, in step S31, the execution device 291 selects the torque command value CT as follows. First, the execution device 291 outputs multiple Q values ​​by inputting the state variables acquired in step S21, i.e., the engine speed NE, the first component SC1, and the second component SC2, into the relational definition model M. Then, the execution device 291 selects an action a using the ε-greedy method based on the multiple Q values. As shown in FIG. 2, after step S31, the execution device 291 advances the process to step S32.

[0043] In step S32, the execution device 291 controls the first motor generator 61 in accordance with the action a selected in step S31. That is, the execution device 291 controls the first motor generator 61 in accordance with the torque command value CT selected in step S31. Specifically, the execution device 291 controls the first motor generator 61 via the first inverter 66 by outputting a control signal in accordance with the torque command value CT selected in step S31 to the first inverter 66. After step S32, the execution device 291 advances the process to step S33.

[0044] In step S33, the execution device 291 acquires a state s after the action a. For example, if the execution device 291 selects an action a_t in step S31 in accordance with the state s_t acquired in step S21, the execution device 291 acquires a state s_t+1 in step S33. Note that the execution device 291 acquires a state s_t+1 in the same manner as in step S21. Here, the state s_t+1 is a state s that is a predetermined period PA after the state s_t. After step S33, the execution device 291 advances the process to step S34.

[0045] In step S34, the execution unit 291 calculates the reward r for the action a selected in step S31. That is, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31. In this embodiment, the execution unit 291 acquires the engine rotation speed NE, the sound pressure NS, the longitudinal acceleration GX, the lateral acceleration GY, the vertical acceleration GZ, and the stop-time crank angle SCA. At this time, for example, if the execution unit 291 selects the action a_t in step S31 in accordance with the state s_t acquired in step S21, in step S34, the execution unit 291 acquires various values ​​at the same timing as the state s_t+1 in step S33. Then, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 based on the various acquired values. Here, the stop-time crank angle SCA is the crank angle SC when the crankshaft 11 stops.

[0046] Specifically, the execution unit 291 calculates the reward r as follows. First, the execution unit 291 calculates a first reward for the torque command value CT selected in step S31 so that the profit rs increases as the engine rotation speed NE decreases. For example, the execution unit 291 calculates a larger first reward as the engine rotation speed NE decreases. The profit rs is the sum of the rewards r when the trial is completed. In other words, the profit rs is the sum of multiple rewards r repeatedly calculated in the trial. Furthermore, the execution unit 291 calculates a second reward for the torque command value CT selected in step S31 so that the profit rs increases when the sound pressure NS is equal to or lower than a predetermined specified sound pressure compared to when the sound pressure NS is higher than the specified sound pressure. For example, when the sound pressure NS is equal to or lower than a predetermined specified sound pressure, the execution unit 291 calculates a larger second reward compared to when the sound pressure NS is higher than the specified sound pressure. For example, the specified sound pressure is the upper limit of the sound pressure NS allowed in the design of the vehicle 100. The execution unit 291 determines whether the absolute values ​​of the longitudinal acceleration GX, lateral acceleration GY, and vertical acceleration GZ are all equal to or less than a predetermined specified acceleration. If the above determination is affirmative, the execution unit 291 calculates a third reward for the torque command value CT selected in step S31 so as to increase the profit rs. For example, if the above determination is affirmative, the execution unit 291 calculates a larger third reward than if the above determination is negative. Note that, for example, the specified acceleration is the upper limit value of the longitudinal acceleration GX, etc., allowed in the design of the vehicle 100. Furthermore, the execution unit 291 calculates a fourth reward for the torque command value CT selected in step S31 so as to increase the profit rs as the stop-time crank angle SCA approaches the specified angle range RR. For example, if the stop-time crank angle SCA is outside the specified angle range RR, the execution unit 291 calculates a smaller fourth reward than if the stop-time crank angle SCA is within the specified angle range RR. At this time, if the stop-time crank angle SCA is outside the specified angle range RR, the execution unit 291 calculates a fourth reward with a smaller value as the stop-time crank angle SCA deviates from the specified angle range RR. Then, the execution unit 291 calculates the sum of the first reward, the second reward, the third reward, and the fourth reward as the reward r for the torque command value CT selected in step S31.In this embodiment, the execution unit 291 calculates the first reward, second reward, third reward, and fourth reward as negative values ​​in principle. On the other hand, when the stop-time crank angle SCA is within the specified angle range RR, that is, when the crankshaft 11 is stopped with the crank angle SC within the specified angle range RR, the execution unit 291 calculates the fourth reward as a constant positive value. Therefore, the reward r for the torque command value CT selected in step S31 is, in principle, a negative value.

[0047] The above processing can be summarized as follows. The execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the lower the engine rotation speed NE, the larger the profit rs. Furthermore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs is larger when the sound pressure NS is equal to or lower than a predetermined specified sound pressure, compared to when the sound pressure NS is higher than the specified sound pressure. Furthermore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs is larger when the longitudinal acceleration GX, etc. is equal to or lower than a predetermined specified acceleration, compared to when the longitudinal acceleration GX, etc. is higher than the specified acceleration. Furthermore, as described above, the reward r for the torque command value CT selected in step S31 is, in principle, a negative value. Therefore, the profit rs tends to increase as the number of times the reward r is calculated in the trial decreases. Therefore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions. After step S34, the execution device 291 advances the process to step S35.

[0048] In step S35, the execution device 291 stores various data. For example, if the execution device 291 selects action a_t in step S31 in accordance with state s_t acquired in step S21, the execution device 291 stores data on state s_t, action a_t, state s_t+1, and reward r_t+1 in the storage device 292. After step S35, the execution device 291 proceeds to step S51.

[0049] In step S51, the execution device 291 determines whether or not a predetermined goal condition is satisfied. The execution device 291 determines that the goal condition is satisfied when, for example, one or more of the following conditions (1) and (2) are satisfied:

[0050] Condition (1): The crankshaft 11 is stopped. Condition (2): The number of steps t in the trial is equal to or greater than a predetermined number. Here, the specified number of times is, for example, the upper limit of the number of times that can be allowed as the number of steps t in one trial.

[0051] If the execution device 291 determines in step S51 that the goal condition is not satisfied (S51: NO), the execution device 291 proceeds to step S21 again. On the other hand, if the execution device 291 determines in step S51 that the goal condition is satisfied (S51: YES), the execution device 291 proceeds to step S52.

[0052] In step S52, the execution device 291 learns the relationship definition model M based on the various data stored in step S35. Specifically, the execution device 291 updates the relationship definition model M based on the reward r calculated in step S34 and the torque command value CT selected in step S31 so that a torque command value CT that increases the profit rs is selected.

[0053] In the learning in step S52, the weight w of the neural network of the relationship definition model M is updated and optimized. The Q value is updated according to the Q value update formula shown in the above formula (1) in order to approach the relationship expressed by the following formula (2).

[0054]

number

[0055] Therefore, the weights w of the neural network need to be updated so that Q(s_t, a_t) output from the output layer approaches the action value shown on the right side of the above formula (2). Therefore, in this embodiment, the execution device 291 defines the error function E as shown in the following formula (3) and updates the weights w of the neural network by backpropagation so that the error function E becomes smaller.

[0056]

number

[0057] Here, the value of maxQ(s_t+1, a) needs to be calculated by inputting the state s_t+1 into the neural network. Therefore, in this embodiment, the execution device 291 uses learning methods such as Experience Replay and Fixed Target Q-Network to train the neural network. After step S52, the execution device 291 proceeds to step S53.

[0058] In step S53, the execution device 291 updates the episode number EN, which indicates the number of attempts. Specifically, the execution device 291 adds "1" to the episode number EN at the time of processing in step S53, and sets the new episode number EN to that value. The initial value of the episode number EN is "0". After step S53, the execution device 291 advances the process to step S54.

[0059] In step S54, the execution device 291 determines whether the number of episodes EN is equal to or greater than a predetermined specified number. Here, the specified number is a threshold for determining whether enough trials have been repeated to determine that learning of the relational specified model M has been completed. An example of the specified number is "1000." In step S54, if the execution device 291 determines that the number of episodes EN is less than the specified number (S54: NO), the execution device 291 advances the process back to step S11. In other words, the execution device 291 executes the next trial.

[0060] On the other hand, if the execution device 291 determines in step S54 that the number of episodes EN is equal to or greater than the specified number (S54: YES), the execution device 291 proceeds to step S55. In other words, the execution device 291 proceeds to step S55 when the number of trials has been repeated enough to determine that learning of the relationship specified model M has been completed.

[0061] In step S55, the execution device 291 stores the relational definition model M as a learned model in the storage device 292. After step S55, the execution device 291 ends the learning control.

[0062] <Engine stop control> Next, the engine stop control executed by the control device 90 will be described. This engine stop control is control for stopping the rotating crankshaft 11 within a predetermined specified angle range RR. Note that the engine stop control is control executed by the control device 90 of the vehicle 100, for example, in a post-development stage of the vehicle 100. In this embodiment, the execution device 91 of the control device 90 starts the engine stop control every predetermined period PA, assuming that there is a request to change the internal combustion engine 10 from an operating state to a stopped state. When starting the engine stop control, the execution device 91 of the control device 90 repeatedly executes the processes of steps S21, S31, and S32 described above. Specifically, similar to step S21, the execution device 91 acquires the state s of the vehicle 100. Next, as in step S31, the execution device 91 inputs the acquired state variables into a relational specification model M to output multiple Q values. Furthermore, the execution device 91 identifies the maximum Q value among the multiple Q values. Then, the execution device 91 selects a torque command value CT corresponding to the maximum Q value. Furthermore, similar to step S32, the execution device 91 controls the first motor generator 61 in accordance with the selected torque command value CT. An example of the predetermined period PA for the engine stop control is 10 msec.

[0063] <Operation of this embodiment> 2, the execution device 291 of the computer 290 of the learning device 200 executes learning control. Through this learning control, attempts are repeatedly made to stop the rotating crankshaft 11 within a predetermined specified angle range RR, and the relationship specification model M is updated so that a torque command value CT that increases the profit rs is selected. In other words, the relationship specification model M is learned so that the crankshaft 11 can be stopped quickly when the crank angle SC is within the specified angle range RR.

[0064] <Effects of this embodiment> (1) In this embodiment, a first component SC1 and a second component SC2 obtained by converting the crank angle SC are used as state variables indicating the state s of the vehicle 100 to be input to the relationship defining model M. The crank angle sensor 71 detects the crank angle SC as a value in a range greater than or equal to 0° and less than 360°. Therefore, when the crankshaft 11 is rotating, the crank angle sensor 71 detects the crank angle SC as a discontinuous value, such as , 358°, 359°, 0°, 1°, and so on. In contrast, the first component SC1 is a value obtained by converting the crank angle SC using a sine function. The second component SC2 is a value obtained by converting the crank angle SC using a cosine function. Therefore, the first component SC1 and the second component SC2 each change continuously as values ​​in a range greater than or equal to "-1" and less than or equal to "+1." As a result, the continuously changing first component SC1 and second component SC2 are input to the relationship defining model M. As a result, it is possible to prevent the occurrence of a situation in which learning of the relationship defining model M is not properly executed due to the input of a crank angle SC with a discontinuous value.

[0065] <Example of change> This embodiment can be modified as follows: This embodiment and the following modifications can be combined and implemented within the scope of technical compatibility.

[0066] In the above embodiment, the learning control may be changed. For example, the state variables in step S21 may be changed. Specifically, the execution unit 291 may acquire other values ​​as state variables indicating the state s of the vehicle 100 in addition to the engine speed NE, the first component SC1, and the second component SC2.

[0067] For example, the number of types of Q values ​​output from the relationship specification model M in step S31 may be changed. As a specific example, the number of candidates for the torque command value CT of the first motor generator 61 may be less than five or more than five. In this case, the number of types of Q values ​​output from the relationship specification model M may be changed to match the number of candidates for the torque command value CT.

[0068] For example, in step S34, the way in which the reward r for the action a selected in step S31 is calculated may be changed. As a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S31 regardless of whether the longitudinal acceleration GX or the like is equal to or less than a predetermined specified acceleration. Also, as a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S31 regardless of whether the sound pressure NS is equal to or less than a predetermined specified sound pressure. Furthermore, as a specific example, the execution device 291 may calculate an additional reward in addition to the first reward, second reward, third reward, and fourth reward. As an example, the execution device 291 calculates a larger additional reward the shorter the period from the start of a trial to the completion of the trial. Then, the execution device 291 calculates the sum of the first reward, second reward, third reward, fourth reward, and additional reward as the reward r for the torque command value CT selected in step S31. Even in this case, the execution device 291 can calculate the reward r for the torque command value CT selected in step S31 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions.

[0069] In the above embodiment, the configuration of the learning device 200 may be changed. For example, the relationship definition model M is not limited to a model generated by DQN. As a specific example, the relationship definition model M may be a so-called Q table or map. Furthermore, for example, the relationship definition model M may be generated by machine learning other than reinforcement learning.

[0070] For example, the configuration of the computer 290 may be changed. Specifically, the computer 290 may be configured as a circuit including one or more processors that execute various processes according to a computer program (software). Note that the computer 290 may also be configured as a circuit including one or more dedicated hardware circuits, such as an application-specific integrated circuit (ASIC), that execute at least some of the various processes, or a combination thereof. The processor includes a CPU and memory such as RAM and ROM. The memory stores program code or instructions configured to cause the CPU to execute processes. The memory, i.e., computer-readable medium, includes any medium that can be accessed by a general-purpose or dedicated computer.

[0071] In the above embodiment, the configuration of the vehicle 100 may be changed. For example, the vehicle 100 may include only one of the first motor generator 61 and the second motor generator 62. In other words, the present technology can be applied to any vehicle 100 in which a motor generator is connected to the crankshaft 11 of the internal combustion engine 10.

[0072] The present technology, which inputs the first and second components to the relationship specification model M, may also be applied to discontinuous values ​​other than the crank angle SC. For example, a value based on the motor angle, which is the angular position of the rotating shaft 61A of the first motor-generator 61, is input to the relationship specification model M. In this case, it is effective for the execution device 291 to input to the relationship specification model M a first component obtained by converting the motor angle using a sine function and a second component obtained by converting the motor angle using a cosine function. The above configuration may also be applied to the rotating shaft 62A of the second motor-generator 62, the drive wheels 69, etc. Furthermore, for example, it is effective for the execution device 291 to input to the relationship specification model M a value based on the latitude of a point on Earth where the vehicle 100 is located. In this case, it is effective for the execution device 291 to input to the relationship specification model M a first component obtained by converting the latitude using a sine function and a second component obtained by converting the latitude using a cosine function. The above configuration may also be applied to the longitude of a point on Earth where the vehicle 100 is located. [Explanation of symbols]

[0073] 10...internal combustion engine 11...crankshaft 20...power split mechanism 30...automatic transmission 61...first motor generator 62...second motor generator 66...first inverter 67...second inverter 68...battery 69...drive wheel 71...crank angle sensor 72...accelerator operation amount sensor 73...vehicle speed sensor 80...connector 90...control device 91...execution device 92...storage device 92A...control program 100...vehicle 200...learning device 210...microphone 220...acceleration sensor 240...connector 250...input device 260...display 290...computer 291...execution device 292...storage device 292A...control program M...relationship specification model

Claims

[Claim 1] A computer-implemented reinforcement learning method for a vehicle including an internal combustion engine and a motor generator coupled to a crankshaft of the internal combustion engine, the method repeatedly performing attempts to stop the rotating crankshaft within a predetermined specified angle range by applying torque from the motor generator to the crankshaft, the computer comprises an execution unit and a storage unit; the storage device stores a relational definition model that outputs an index value indicating a value of an action corresponding to each of a plurality of candidates for a torque command value of the motor generator in the trial in response to an input of a state variable indicating a state of the vehicle; the state variables include an engine rotation speed, which is a rotation speed of the crankshaft, a first component obtained by converting a crank angle, which is an angular position of the crankshaft, using a sine function, and a second component obtained by converting the crank angle using a cosine function; The execution device: For each of the trials, obtaining the state variables; selecting the torque command value corresponding to one of the plurality of index values ​​based on the plurality of index values ​​output by inputting the acquired state variables into the relationship definition model; controlling the motor generator in accordance with the selected torque command value; acquiring a crank angle at a stop, which is the crank angle when the crankshaft stops, after controlling the motor generator; calculating the reward for the selected torque command value such that the profit, which is the sum of rewards when the trial is completed, increases as the acquired stop-time crank angle approaches the specified angle range and the crankshaft is stopped with fewer actions; updating the relationship definition model based on the calculated reward and the selected torque command value so that the torque command value that increases the profit is selected; Run Reinforcement learning methods.

Citation Information

Patent Citations

  • Noise estimation device and vehicular control device

    JP2022076162A