Reinforcement learning method
The reinforcement learning method optimizes crankshaft stopping in vehicles by generating a relational specification model to control the motor generator, addressing the inefficiencies in existing systems and achieving rapid crankshaft shutdown.
Patent Information
- Application Number
- JP2024099554
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-20
- Publication Date
- 2026-01-08
AI Technical Summary
Existing vehicle systems lack an efficient method to quickly stop the rotating crankshaft within a specified angle range during engine shutdown, with existing control maps focusing only on battery temperature without addressing the efficient stopping of the crankshaft.
A reinforcement learning method using a computer to control the motor generator, which includes an execution unit and a storage unit to generate a relational specification model that learns to stop the crankshaft within a specified angle range by applying torque from the motor generator, optimizing the torque command values based on state variables and rewards.
The method enables quick and efficient stopping of the crankshaft within the specified angle range, reducing the human effort required to develop a control model compared to traditional experimental methods.
Smart Images

Figure 2026001944000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a reinforcement learning method. [Background technology]
[0002] The vehicle disclosed in Patent Document 1 includes an engine, a motor generator, an inverter, a battery, and a control device. The motor generator is connected to the crankshaft of the engine. The battery supplies power to the motor generator via the inverter.
[0003] The control device controls the engine by outputting a control signal to the engine. The control device stores a first map and a second map. When the temperature of the battery is below a predetermined switching temperature, the control device calculates a torque command value for the motor generator using the first map. When the temperature of the battery is equal to or higher than the switching temperature, the control device calculates a torque command value for the motor generator using the second map. The control device then controls the inverter according to the calculated torque command value, thereby controlling the motor generator. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2020-011530 Summary of the Invention [Problem to be solved by the invention]
[0005] In a vehicle such as that described in Patent Document 1, when stopping an internal combustion engine from a running state, it may be desirable to stop the rotating crankshaft within a predetermined specified angle range. However, Patent Document 1 only discloses that different control maps for the motor generator are used depending on the battery temperature. On the other hand, Patent Document 1 does not pay any attention to the efficient method of generating a relational specified model such as a control map for quickly stopping the crankshaft within the specified angle range. [Means for solving the problem]
[0006] A reinforcement learning method for solving the above-described problems is a reinforcement learning method by a computer, which targets a vehicle equipped with an internal combustion engine and a motor generator connected to a crankshaft of the internal combustion engine, and repeatedly performs attempts to stop the rotating crankshaft within a predetermined specified angle range by applying torque from the motor generator to the crankshaft, wherein the computer comprises an execution unit and a storage unit, and the storage unit stores a relational specification model that receives as input a state variable indicating a state of the vehicle and outputs an index value indicating the value of an action corresponding to each of a plurality of candidates for a torque command value of the motor generator in the attempt, and the state variables include an engine speed which is the rotational speed of the crankshaft, a crank angle which is the angular position of the crankshaft, the torque command value used in the control of the motor generator one step before the torque command value corresponding to the current index value output by the relational specification model, and The execution device executes the following steps for each trial, including a period that has elapsed since the start of the trial: acquire the state variables; select the torque command value corresponding to one of the plurality of index values based on the plurality of index values output by inputting the acquired state variables into the relationship specification model; control the motor generator according to the selected torque command value; acquire a stop crank angle that is the crank angle when the crankshaft stops after controlling the motor generator; calculate the reward for the selected torque command value so that the closer the acquired stop crank angle is to the specified angle range and the fewer actions are required to stop the crankshaft, the greater the profit that is the total of rewards when the trial is completed; and update the relationship specification model so that the torque command value that increases the profit is selected based on the calculated reward and the selected torque command value. [Effects of the Invention]
[0007] According to the above configuration, it is possible to generate a relationship specification model for quickly stopping the crankshaft at the crank angle at which it should be stopped, while reducing the amount of human effort required compared to when a human generates a relationship specification model based on experiments, for example. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a schematic diagram of a vehicle and a learning device. [Figure 2] FIG. 2 is a flowchart showing the learning control. [Figure 3] FIG. 3 is an explanatory diagram of the first relationship definition model and the second relationship definition model. [Figure 4] Fig. 4(a) is a time chart showing changes in engine rotation speed, Fig. 4(b) is a time chart showing changes in the operating state of the internal combustion engine, and Fig. 4(c) is a time chart showing changes in the torque command value of the first motor generator. [Figure 5] FIG. 5 is an explanatory diagram showing changes in the torque command value of the first motor generator. [Figure 6] FIG. 6 is a flowchart showing the engine stop control. DETAILED DESCRIPTION OF THE INVENTION
[0009] <Vehicle Overview> An embodiment of the present invention will now be described with reference to Figures 1 to 6. First, a schematic configuration of a vehicle 100 will be described.
[0010] 1, the vehicle 100 includes an internal combustion engine 10, a power split mechanism 20, an automatic transmission 30, a plurality of drive wheels 69, a first motor generator 61, and a second motor generator 62. The internal combustion engine 10, the first motor generator 61, and the second motor generator 62 are the drive sources of the vehicle 100.
[0011] The internal combustion engine 10 has a crankshaft 11 and four cylinders (not shown). Each cylinder undergoes an intake stroke, a compression stroke, a combustion stroke, and an exhaust stroke as the crankshaft 11 rotates twice. The crankshaft 11 is the output shaft of the internal combustion engine 10. The crankshaft 11 is connected to a power split mechanism 20.
[0012] The power split mechanism 20 is a planetary gear mechanism having a sun gear S, a ring gear RG, and a carrier C. The carrier C of the power split mechanism 20 is connected to the crankshaft 11. The sun gear S is connected to a rotating shaft 61A of the first motor generator 61. In other words, the rotating shaft 61A of the first motor generator 61 is connected to the crankshaft 11 of the internal combustion engine 10 via the power split mechanism 20. The ring gear RG has a ring gear shaft RA. The ring gear shaft RA is an output shaft of the ring gear RG. The ring gear shaft RA is connected to a rotating shaft 62A of the second motor generator 62. The ring gear shaft RA is also connected to the automatic transmission 30. The automatic transmission 30 is connected to left and right drive wheels 69 via a differential gear (not shown). An example of the automatic transmission 30 is a stepped automatic transmission. Therefore, the automatic transmission 30 changes the gear ratio by changing the gear position.
[0013] When the internal combustion engine 10 operates and torque from the internal combustion engine 10 is input to the carrier C of the power split mechanism 20, the torque is split between the sun gear S side and the ring gear RG side. Furthermore, when the first motor generator 61 operates as an electric motor and torque from the first motor generator 61 is input to the sun gear S of the power split mechanism 20, the torque is split between the carrier C side and the ring gear RG side. Therefore, the first motor generator 61 can apply torque to the crankshaft 11 of the internal combustion engine 10 via the power split mechanism 20.
[0014] When torque from second motor generator 62 is input to ring gear shaft RA as a result of second motor generator 62 operating as an electric motor, the torque is transmitted to automatic transmission 30. Furthermore, when torque from the drive wheels 69 side is input to second motor generator 62 via ring gear shaft RA, second motor generator 62 functions as a generator. As a result, second motor generator 62 can generate regenerative braking force for vehicle 100.
[0015] The vehicle 100 is equipped with a first inverter 66, a second inverter 67, and a battery 68 as devices for exchanging electric power. The first inverter 66 adjusts the amount of electric power exchanged between the first motor generator 61 and the battery 68. The second inverter 67 adjusts the amount of electric power exchanged between the second motor generator 62 and the battery 68.
[0016] 1, vehicle 100 is equipped with a crank angle sensor 71, an accelerator operation amount sensor 72, and a vehicle speed sensor 73. Crank angle sensor 71 detects a crank angle SC, which is the angular position of crankshaft 11. Accelerator operation amount sensor 72 detects an accelerator operation amount ACC, which is the amount of operation of the accelerator pedal operated by the driver. Vehicle speed sensor 73 detects a vehicle speed SP, which is the speed of vehicle 100.
[0017] The vehicle 100 is equipped with a control device 90. The control device 90 acquires various information from a crank angle sensor 71, an accelerator operation amount sensor 72, and a vehicle speed sensor 73. The control device 90 calculates an engine rotation speed NE, which is the rotation speed of the crankshaft 11, based on the crank angle SC.
[0018] The control device 90 includes an execution device 91 and a storage device 92. An example of the execution device 91 is a CPU. The storage device 92 includes a read-only ROM, a readable / writable volatile RAM, and a readable / writable non-volatile storage. The storage device 92 stores various programs and data in advance. Specifically, the storage device 92 stores a control program 92A in advance as one of the various programs. The storage device 92 also stores a relationship definition model M in advance as one of the various data. The relationship definition model M is a model used to stop the rotating crankshaft 11 within a predetermined specified angle range RR by applying torque from the first motor-generator 61 to the crankshaft 11. The relationship definition model M describes, in a format recognizable by the execution device 91, the relationship between predetermined input data and a Q value, which is an index value indicating the value of an action corresponding to each of multiple candidates for the torque command value CT of the first motor-generator 61. The relationship specifying model M receives a plurality of pieces of input data and outputs a Q value corresponding to each of a plurality of candidates for the torque command value CT of the first motor generator 61. The relationship specifying model M includes a first relationship specifying model M1 and a second relationship specifying model M2. In this embodiment, the relationship specifying model M stored in the storage device 92 is a model for which learning by machine learning has been completed. In other words, the storage device 92 stores the relationship specifying model M generated by learning control, which will be described later. The relationship specifying model M will be described in detail later. The execution device 91 executes a control program 92A stored in the storage device 92 to perform various processes, which will be described later.
[0019] The execution unit 91 of the control device 90 controls the internal combustion engine 10, the first motor generator 61, the second motor generator 62, the automatic transmission 30, etc. Specifically, the execution unit 91 calculates a vehicle required driving force, which is a required value of driving force necessary for the vehicle 100 to travel, based on the accelerator operation amount ACC and the vehicle speed SP. The execution unit 91 determines a torque distribution among the internal combustion engine 10, the first motor generator 61, and the second motor generator 62 based on the vehicle required driving force. The execution unit 91 controls the output of the internal combustion engine 10 and the power running and regeneration of the first motor generator 61 and the second motor generator 62 based on the torque distribution among the internal combustion engine 10, the first motor generator 61, and the second motor generator 62. Specifically, the execution unit 91 controls the internal combustion engine 10 by outputting a control signal to the internal combustion engine 10. The execution unit 91 also controls the first motor generator 61 via the first inverter 66 by outputting a control signal to the first inverter 66. Furthermore, the execution device 91 controls the second motor generator 62 via the second inverter 67 by outputting a control signal to the second inverter 67 .
[0020] Furthermore, the execution unit 91 calculates a target gear position, which is a target gear position for the automatic transmission 30, based on the vehicle speed SP and the vehicle required driving force. The execution unit 91 outputs a control signal to the automatic transmission 30 based on the target gear position. As a result, the gear position of the automatic transmission 30 is controlled.
[0021] The vehicle 100 is equipped with a connector 80. The connector 80 is connected to a control device 90. The connector 80 is equipped with a general-purpose terminal group that enables two-way communication. The connector 80 enables communication between the control device 90 and other devices. In this embodiment, the control device 90 can communicate with a learning device 200, which will be described later, via the connector 80.
[0022] <Overview of the learning device> Next, a description will be given of the learning device 200 that learns the relational specification model M. The learning device 200 is a device that is used in the development stage of the vehicle 100, for example.
[0023] As shown in FIG. 1, the learning device 200 includes a microphone 210 , an acceleration sensor 220 , a connector 240 , an input device 250 , a display 260 , and a computer 290 .
[0024] The microphone 210 detects sound pressure NS as noise generated in the vehicle 100. In this embodiment, the microphone 210 is attached to a predetermined position in the interior of the vehicle 100. An example of the predetermined position is the driver's seat of the vehicle 100. The microphone 210 is connected to the computer 290.
[0025] The acceleration sensor 220 is a so-called three-axis sensor. That is, the acceleration sensor 220 can detect longitudinal acceleration GX, lateral acceleration GY, and vertical acceleration GZ. The longitudinal acceleration GX is acceleration along the longitudinal axis of the vehicle 100. The lateral acceleration GY is acceleration along the lateral axis of the vehicle 100. The vertical acceleration GZ is acceleration along the vertical axis of the vehicle 100. In other words, the longitudinal acceleration GX is acceleration in the longitudinal direction relative to the vehicle 100. Furthermore, the lateral acceleration GY is acceleration in the lateral direction relative to the vehicle 100. Furthermore, the vertical acceleration GZ is acceleration in the vertical direction relative to the vehicle 100. In this embodiment, the acceleration sensor 220 is attached to a predetermined position in the interior of the vehicle 100. An example of the predetermined position is the driver's seat of the vehicle 100. The acceleration sensor 220 is connected to the computer 290.
[0026] Input device 250 is a device for inputting commands from an engineer or the like using learning device 200 into computer 290. Input device 250 is connected to computer 290. Input device 250 is, for example, a keyboard, a pointing device, or the like. Display 260 is a device for transmitting information from computer 290 to an engineer or the like. Display 260 is capable of displaying various types of information. Display 260 is connected to computer 290.
[0027] The computer 290 includes an execution device 291 and a storage device 292. An example of the execution device 291 is a CPU. The storage device 292 includes a read-only ROM, a readable / writable volatile RAM, and a readable / writable non-volatile storage. The storage device 292 stores various programs and various data in advance. Specifically, the storage device 292 stores a control program 292A in advance as one of the various programs. The storage device 292 also stores a relationship specification model M in advance as one of the various data. As described above, the relationship specification model M includes a first relationship specification model M1 and a second relationship specification model M2. The relationship specification model M stored in the storage device 292 is a model of a learning process based on machine learning. The execution device 291 executes the control program 292A stored in the storage device 292 to perform various processes described below. In other words, the execution device 291 executes the control program 292A stored in the storage device 292 to perform various processes related to a reinforcement learning method. An example of the computer 290 is a personal computer.
[0028] Connector 240 is connected to computer 290. Connector 240 has a group of general-purpose terminals that enable bidirectional communication. Connector 240 enables communication between computer 290 and other devices. In this embodiment, for example, when an engineer connects connector 240 of learning device 200 to connector 80 of vehicle 100 during learning of relationship specification model M, computer 290 of learning device 200 becomes able to communicate with control device 90 of vehicle 100. Then, execution device 291 of computer 290 can access control device 90 to obtain various information from control device 90.
[0029] <Learning control> Next, with reference to FIG. 2, the learning control executed by the computer 290 of the learning device 200 will be described. This learning control is control related to reinforcement learning, in which attempts are repeatedly made to stop the rotating crankshaft 11 within a predetermined specified angle range RR by applying torque from the first motor-generator 61 to the crankshaft 11. Note that the learning control is control executed by the computer 290 of the learning device 200, for example, during the development stage of the vehicle 100. Note that, as described above, each of the four cylinders of the internal combustion engine 10 undergoes an intake stroke, a compression stroke, a combustion stroke, and an exhaust stroke when the crankshaft 11 rotates twice. Therefore, when the crankshaft 11 rotates twice, there are four specified angle ranges RR corresponding to the four cylinders. For example, the four specified angle ranges RR are predetermined angle ranges based on 180°, 360°, 540°, and 720°, respectively. As a specific example, the specified angle range RR with 180° as the reference is an angle range of 150° or more and 210° or less.
[0030] In this learning control, the Q value, which is an index value indicating the value of an action, is updated. Here, the Q value is expressed by the following equation (1).
[0031]
number
[0032] Here, Q(s, a) is the expected future profit when action a is selected in state s. "r" is the reward. The subscript "t" in state s, action a, and reward r in the above formula (1) is a number indicating one step in the trial. When state s changes after the action is decided, the step number t, which is the number indicating the step, is incremented by one. Note that the value of step number t at the start of the trial is "0".
[0033] In the following, subscripts will be written with "_". Therefore, the reward r_t+1 in equation (1) is the reward obtained when state s becomes state s_t+1 by selecting action a_t in state s_t. Also, "α" is the learning rate. "γ" is the discount rate. The learning rate α is a value greater than "0" and less than "1". The discount rate γ is a value greater than "0" and less than "1". Also, maxQ(s_t+1,a) is the Q value maximized by selecting action a. In this section, action a indicates the action that maximizes Q(s_t+1,a_t+1) among the actions a_t+1 that can be taken in state s_t+1.
[0034] For example, the execution device 291 of the computer 290 starts learning control when an engineer performs an operation via the input device 250 requesting the execution of learning control, with the necessary condition being that the connector 240 of the learning device 200 and the connector 80 of the vehicle 100 are connected.
[0035] As shown in FIG. 2, when the execution unit 291 of the computer 290 starts learning control, it executes the processing of step S11. In step S11, the execution unit 291 operates the internal combustion engine 10 by outputting a control signal to the internal combustion engine 10 via the control device 90. Specifically, the execution unit 291 controls the internal combustion engine 10 so that the engine rotation speed NE becomes a predetermined reference rotation speed. An example of the reference rotation speed is about 1000 rpm. The reference rotation speed is determined as the engine rotation speed NE during idling of the internal combustion engine 10 in the vehicle 100. After step S11, the execution unit 291 advances the processing to step S21.
[0036] In step S21, the execution unit 291 acquires the state s of the vehicle 100. Specifically, the execution unit 291 acquires the engine speed NE, the crank angle SC, the previous command value CTA, and the elapsed period PE as state variables indicating the state s of the vehicle 100. At this time, for example, if the number of trial steps at the time of processing step S21 is "t", the execution unit 291 acquires the state variables in the state s_t. Here, the previous command value CTA is the torque command value CT at a time point a predetermined period PA before the time of processing step S21. In other words, the previous command value CTA is the torque command value CT used in the control of the first motor generator 61 immediately before the torque command value CT corresponding to the current Q value output by the relationship specification model M. The predetermined period PA is the same as the cycle at which the engine stop control described below is executed. An example of the predetermined period PA is 10 msec. The elapsed period PE is the period that has elapsed since the crankshaft 11 started to stop. In other words, the elapsed period PE is the period that has elapsed since the start of the current trial. For example, if the execution device 291 acquires the state variables in state s_t in the current step S21, the execution device 291 acquires the state variables in state s_t+1 in the next step S21. After step S21, the execution device 291 proceeds to step S22.
[0037] In step S22, the execution unit 291 determines whether the engine speed NE in the state variables acquired in step S21 is equal to or greater than a predetermined specified speed NEA. Here, an example of the specified speed NEA is 200 rpm. In step S22, if the execution unit 291 determines that the engine speed NE is equal to or greater than the specified speed NEA (S22: YES), the execution unit 291 proceeds to step S31.
[0038] In step S31, the execution device 291 selects a torque command value CT corresponding to one of the multiple Q values based on the multiple Q values output by inputting the state variables acquired in step S21 into the first relationship definition model M1.
[0039] In this embodiment, the first relationship definition model M1 is a model generated by DQN (Deep Q-Network), which is a method for approximately calculating a Q value. In DQN, a Q value is estimated using a multi-layer neural network. Therefore, as shown in FIG. 3, the first relationship definition model M1 is a neural network that receives a state variable indicating a state s as input and outputs a value Q(s, a), which is a Q value corresponding to the number of selectable actions a. This neural network receives m values X, which are the state variables, and outputs n Q values corresponding to n actions a. In FIG. 3, the m input values in the number of steps t of a trial are indicated as X_1 to X_m.
[0040] In this embodiment, the state variables input to the first relationship defining model M1 are the engine speed NE and the elapsed time PE among the state variables acquired in step S21. That is, "m" is "2." The Q value output by the first relationship defining model M1 indicates the Q value when a specific action a is selected in the input state s, i.e., the value of the action a. The number of types of Q values output by the first relationship defining model M1 is the same as the number of candidates for the torque command value CT of the first motor generator 61. Here, examples of the multiple candidates are five: a first command value, a second command value, a third command value, a fourth command value, and a fifth command value. Specifically, the third command value is the torque command value CT that is the same as the previous command value CTA acquired in step S21. In other words, the third command value is a value for maintaining the torque command value CT. The fourth command value is a value that is greater than the third command value by a predetermined constant value. The fifth command value is a value that is greater than the fourth command value by a predetermined constant value. The second command value is a value that is smaller than the third command value by a predetermined constant value. The first command value is a value that is smaller than the second command value by a predetermined constant value. In other words, "n" is "5."
[0041] The neural network shown in Figure 3 has an input layer, multiple hidden layers, and an output layer. This neural network is a multi-layered neural network in which each node in each layer multiplies the input of the previous layer by a weight w and adds a bias b to obtain an output that has passed through an activation function. Note that in Figure 3, the transmission lines connecting nodes in adjacent layers are not shown.
[0042] The structure of a neural network is specified by information such as the weights w and bias b in each layer, the activation function, and the order of layers. During training, the weights w, which are variable values within the neural network, are updated.
[0043] Specifically, in step S31, the execution unit 291 selects the torque command value CT as follows. First, the execution unit 291 extracts the engine speed NE and the elapsed period PE from the state variables acquired in step S21. Next, the execution unit 291 inputs the extracted engine speed NE and elapsed period PE into the first relationship defining model M1, thereby outputting multiple Q values. Then, the execution unit 291 selects an action a using the ε-greedy method based on the multiple Q values.
[0044] In the ε-greedy method, the execution device 291 generally selects action a with the largest Q-value. On the other hand, the execution device 291 randomly selects action a at a predetermined rate. At this time, the execution device 291 decreases the predetermined rate as learning progresses. Therefore, at the beginning of learning, the execution device 291 randomly selects action a without following the Q-value much. Then, as learning progresses, the execution device 291 begins to select action a based on the Q-value. Eventually, the execution device 291 selects action a based only on the Q-value. As shown in FIG. 2, after step S31, the execution device 291 advances the process to step S32.
[0045] In step S32, the execution device 291 controls the first motor generator 61 in accordance with the action a selected in step S31. That is, the execution device 291 controls the first motor generator 61 in accordance with the torque command value CT selected in step S31. Specifically, the execution device 291 controls the first motor generator 61 via the first inverter 66 by outputting a control signal in accordance with the torque command value CT selected in step S31 to the first inverter 66. After step S32, the execution device 291 advances the process to step S33.
[0046] In step S33, the execution device 291 acquires a state s after the action a. For example, if the execution device 291 selects an action a_t in step S31 in accordance with the state s_t acquired in step S21, the execution device 291 acquires a state s_t+1 in step S33. Note that the execution device 291 acquires the state s_t+1 in the same manner as in step S21. Here, the state s_t+1 is the state s that is a predetermined period PA after the state s_t. In this embodiment, the predetermined period PA has the same value as a predetermined delay period. Here, the delay period is, for example, a period from when the execution device 291 selects an action a_t in step S31 in accordance with the state s_t acquired in step S21 until one or more changes in the detected sound pressure and the detected acceleration, which will be described later, begin to occur in accordance with the action a_t. In other words, the delay period is a period from when the execution device 291 acquires the state variables to when one or more changes in the detected sound pressure and the detected acceleration begin to occur due to control of the first motor generator 61 in accordance with the state variables. The delay period is determined in advance through experiments, simulations, etc. After step S33, the execution device 291 advances the process to step S34.
[0047] In step S34, the execution device 291 calculates a reward r for the action a selected in step S31. That is, the execution device 291 calculates a reward r for the torque command value CT selected in step S31. In this embodiment, the execution device 291 acquires the engine rotation speed NE, the sound pressure NS, the longitudinal acceleration GX, the lateral acceleration GY, and the vertical acceleration GZ. At this time, for example, if the execution device 291 selects the action a_t in step S31 in accordance with the state s_t acquired in step S21, in step S34, the execution device 291 acquires various values at the same timing as the state s_t+1 in step S33. Then, the execution device 291 calculates a reward r for the torque command value CT selected in step S31 based on the various acquired values. Here, the sound pressure NS is a detected sound pressure detected as noise generated in the vehicle 100. Furthermore, the longitudinal acceleration GX, the lateral acceleration GY, and the vertical acceleration GZ are detected accelerations detected as vibrations generated in the vehicle 100. As described above, state s_t+1 is state s that is a predetermined period PA after state s_t. The predetermined period PA has the same value as the delay period. Therefore, after controlling the first motor generator 61, the execution device 291 acquires the detected sound pressure and detected acceleration after the delay period from when the state variables were acquired.
[0048] Specifically, the execution unit 291 calculates the reward r as follows. First, the execution unit 291 calculates a first reward for the torque command value CT selected in step S31 so that the profit rs increases as the engine rotation speed NE decreases. For example, the execution unit 291 calculates a larger value of the first reward as the engine rotation speed NE decreases. The profit rs is the sum of the rewards r when the trial is completed. In other words, the profit rs is the sum of multiple rewards r repeatedly calculated in the trial. Furthermore, the execution unit 291 calculates a second reward for the torque command value CT selected in step S31 so that the profit rs increases when the sound pressure NS is equal to or less than a predetermined specified sound pressure compared to when the sound pressure NS is greater than the specified sound pressure. For example, when the sound pressure NS is equal to or less than a predetermined specified sound pressure, the execution unit 291 calculates a larger value of the second reward compared to when the sound pressure NS is greater than the specified sound pressure. For example, the specified sound pressure is the upper limit of the detected sound pressure allowed in the design of the vehicle 100. The execution unit 291 determines whether the absolute values of the longitudinal acceleration GX, lateral acceleration GY, and vertical acceleration GZ are all equal to or less than a predetermined specified acceleration. If the execution unit 291 determines yes, it calculates a third reward for the torque command value CT selected in step S31 so as to increase the profit rs. For example, if the execution unit 291 determines yes, it calculates a larger third reward than if the execution unit 291 determines no. Note that the specified acceleration is, for example, the upper limit of the detected acceleration allowed by the design of the vehicle 100. The execution unit 291 then calculates the sum of the first reward, the second reward, and the third reward as the reward r for the torque command value CT selected in step S31. In this embodiment, the execution unit 291 generally calculates the first reward, the second reward, and the third reward as negative values. On the other hand, if the engine rotation speed NE is less than the specified rotation speed NEA, the execution unit 291 calculates the first reward as a constant positive value. Therefore, the reward r for the torque command value CT selected in step S31 is, in principle, a negative value.
[0049] The above processing can be summarized as follows. The execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the lower the engine rotation speed NE, the larger the profit rs. Furthermore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs is larger when the detected sound pressure is equal to or lower than a predetermined specified sound pressure, compared to when the detected sound pressure is higher than the specified sound pressure. Furthermore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs is larger when the detected acceleration is equal to or lower than a predetermined specified acceleration, compared to when the detected acceleration is higher than the specified acceleration. Furthermore, as described above, the reward r for the torque command value CT selected in step S31 is, in principle, a negative value. Therefore, the profit rs tends to increase as the number of times the reward r is calculated in the trial decreases. Therefore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions. After step S34, the execution device 291 advances the process to step S35.
[0050] In step S35, the execution device 291 stores various data. For example, if the execution device 291 selects action a_t in step S31 in accordance with state s_t acquired in step S21, the execution device 291 stores data on state s_t, action a_t, state s_t+1, and reward r_t+1 in the storage device 292. After step S35, the execution device 291 proceeds to step S51.
[0051] On the other hand, if the execution unit 291 determines in step S22 that the engine rotation speed NE is lower than the specified rotation speed NEA (S22: NO), the execution unit 291 advances the process to step S41.
[0052] In step S41, the execution device 291 selects a torque command value CT corresponding to one of the multiple Q values based on the multiple Q values output by inputting the state variables acquired in step S21 into the second relationship definition model M2.
[0053] In this embodiment, the second relationship specification model M2 is a model generated by DQN (Deep Q-Network), which is a method for approximately calculating a Q value. In DQN, the Q value is estimated using a multilayer neural network. Therefore, as shown in FIG. 3, the second relationship specification model M2 is a neural network that receives a state variable indicating a state s as input and outputs a value of Q(s, a), which is a Q value corresponding to the number of selectable actions a, similar to the first relationship specification model M1.
[0054] In this embodiment, the state variables input to the second relationship specifying model M2 are the engine speed NE, the crank angle SC, and the previous command value CTA from among the state variables acquired in step S21. That is, "m" is "3." The Q value output by the second relationship specifying model M2 indicates the Q value when a specific action a is selected in the input state s, i.e., the value of the action a. The number of types of Q values output by the second relationship specifying model M2 is the same as the number of candidates for the torque command value CT of the first motor generator 61. Here, examples of the plurality of candidates are five, namely, the first command value, the second command value, the third command value, the fourth command value, and the fifth command value, as in the first relationship specifying model M1. That is, "n" is "5."
[0055] Specifically, in step S41, the execution unit 291 selects the torque command value CT as follows. First, the execution unit 291 extracts the engine speed NE, the crank angle SC, and the previous command value CTA from the state variables acquired in step S21. Next, the execution unit 291 outputs a plurality of Q values by inputting the extracted engine speed NE, the crank angle SC, and the previous command value CTA into the second relationship definition model M2. Then, similar to step S31, the execution unit 291 selects an action a using the ε-greedy algorithm based on the plurality of Q values. As shown in FIG. 2, after step S41, the execution unit 291 advances the process to step S42.
[0056] In step S42, the execution device 291 controls the first motor generator 61 in accordance with the action a selected in step S41. That is, the execution device 291 controls the first motor generator 61 in accordance with the torque command value CT selected in step S41. Specifically, the execution device 291 controls the first motor generator 61 via the first inverter 66 by outputting a control signal in accordance with the torque command value CT selected in step S41 to the first inverter 66. After step S42, the execution device 291 advances the process to step S43.
[0057] In step S43, the execution device 291 acquires a state s after the action a. For example, if the execution device 291 selects an action a_t in step S41 in accordance with the state s_t acquired in step S21, the execution device 291 acquires a state s_t+1 in step S43. Note that, similarly to step S21, the execution device 291 acquires a state s_t+1. Here, the state s_t+1 is a state s that is a predetermined period PA after the state s_t. In this embodiment, the predetermined period PA is the same value as the predetermined delay period, similarly to step S33. After step S43, the execution device 291 proceeds to step S44.
[0058] In step S44, the execution device 291 calculates the reward r for the action a selected in step S41. That is, the execution device 291 calculates the reward r for the torque command value CT selected in step S41. In this embodiment, the execution device 291 acquires the stop crank angle SCA, the number of changes NC, and the change period PC. Then, the execution device 291 calculates the reward r for the torque command value CT selected in step S41 based on the various acquired values. Here, the stop crank angle SCA is the crank angle SC when the crankshaft 11 stops.
[0059] The number of changes NC indicates the number of changes in the torque command value CT applied to the crankshaft 11. Here, the torque command value CT in the direction to rotate the crankshaft 11 forward is defined as positive torque. The torque command value CT in the direction to rotate the crankshaft 11 backward is defined as negative torque. For example, as shown in FIG. 4(c), in the period from time T1 to time T2, which is the period immediately after the start of the trial, the torque command value CT continues to be negative torque. On the other hand, in the period from time T2 to time T3, which is the period until the crankshaft 11 is stopped, the torque command value CT becomes positive torque and then negative torque. As shown in FIG. 5, the number of changes NC is the number of times that the torque changes from one of positive torque and negative torque to the other during the period from the stop crank angle SCA to the crank angle SC that is a predetermined fixed angle before. In other words, the number of changes NC is the number of times that the torque changes from one of positive torque and negative torque to the other during a predetermined specified period. Note that an example of the fixed angle is 720°.
[0060] The change period PC indicates the period of change in the torque command value CT applied to the crankshaft 11. For example, assume that the torque command value CT is changing as shown in FIG. 5 . In this situation, the execution unit 291 obtains the change period PC as follows. First, the execution unit 291 calculates, as a first average value, the average value of the angle range of the crank angle SC in which the torque command value CT continues to be positive torque within a period from the stop-time crank angle SCA to the crank angle SC a predetermined angle before. The execution unit 291 also calculates, as a second average value, the average value of the angle range of the crank angle SC in which the torque command value CT continues to be negative torque within a period from the stop-time crank angle SCA to the crank angle SC a predetermined angle before. Next, the execution unit 291 multiplies the first average value by the second average value to calculate the change period PC. Therefore, the change period PC increases as the first average value and the second average value increase. Furthermore, for example, if the sum of the first average value and the second average value is the same, the smaller the absolute value of the difference between the first average value and the second average value, the larger the change period PC. Note that an example of the above-mentioned fixed angle is 720°.
[0061] Specifically, the execution unit 291 calculates the reward r as follows. First, the execution unit 291 calculates a fourth reward for the torque command value CT selected in step S41 so that the profit rs increases as the stop-time crank angle SCA approaches the specified angle range RR. For example, when the stop-time crank angle SCA is outside the specified angle range RR, the execution unit 291 calculates a smaller value of the fourth reward than when the stop-time crank angle SCA is within the specified angle range RR. At this time, when the stop-time crank angle SCA is outside the specified angle range RR, the execution unit 291 calculates a smaller value of the fourth reward as the stop-time crank angle SCA deviates from the specified angle range RR. Furthermore, the execution unit 291 calculates a fifth reward for the torque command value CT selected in step S41 so that the profit rs increases as the number of changes NC decreases. For example, the execution unit 291 calculates a larger value of the fifth reward as the number of changes NC decreases. Furthermore, the execution unit 291 calculates a sixth reward for the torque command value CT selected in step S41 so that the profit rs increases as the change period PC increases. For example, the execution unit 291 calculates a larger sixth reward as the change period PC increases. Then, the execution unit 291 calculates the sum of the fourth, fifth, and sixth rewards as the reward r for the torque command value CT selected in step S41. In this embodiment, the execution unit 291 calculates each of the fourth, fifth, and sixth rewards as a negative value in principle. On the other hand, when the stop-time crank angle SCA is within the specified angle range RR, that is, when the crankshaft 11 is stopped with the crank angle SC within the specified angle range RR, the execution unit 291 calculates the fourth reward as a constant positive value. Therefore, the reward r for the torque command value CT selected in step S41 is a negative value in principle.
[0062] The above process can be summarized as follows. The execution unit 291 calculates the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the stop-time crank angle SCA approaches the specified angle range RR. The execution unit 291 also calculates the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the number of changes NC decreases. The execution unit 291 also calculates the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the change period PC increases. As described above, the reward r for the torque command value CT selected in step S41 is, in principle, a negative value. Therefore, the profit rs tends to increase as the number of times the reward r is calculated in trials decreases. Therefore, the execution unit 291 calculates the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions. As shown in FIG. 2, after step S44, the execution unit 291 proceeds to step S45.
[0063] In step S45, the execution device 291 stores various data. For example, if the execution device 291 selects action a_t in step S41 in accordance with state s_t acquired in step S21, the execution device 291 stores data on state s_t, action a_t, state s_t+1, and reward r_t+1 in the storage device 292. After step S45, the execution device 291 proceeds to step S51.
[0064] In step S51, the execution device 291 determines whether or not a predetermined goal condition is satisfied. The execution device 291 determines that the goal condition is satisfied when, for example, one or more of the following conditions (1) and (2) are satisfied:
[0065] Condition (1): The crankshaft 11 is stopped. Condition (2): The number of steps t in the trial is equal to or greater than a predetermined number. Here, the specified number of times is, for example, the upper limit of the number of times that can be allowed as the number of steps t in one trial.
[0066] If the execution device 291 determines in step S51 that the goal condition is not satisfied (S51: NO), the execution device 291 proceeds to step S21 again. On the other hand, if the execution device 291 determines in step S51 that the goal condition is satisfied (S51: YES), the execution device 291 proceeds to step S52.
[0067] In step S52, the execution device 291 learns the relationship defining model M based on the various data stored in step S35 and the various data stored in step S45. Specifically, the execution device 291 updates the first relationship defining model M1 based on the various data stored in step S35 so that a torque command value CT that increases the profit rs is selected. In other words, the execution device 291 updates the first relationship defining model M1 based on the reward r calculated in step S34 and the torque command value CT selected in step S31 so that a torque command value CT that increases the profit rs is selected. Furthermore, the execution device 291 updates the second relationship defining model M2 based on the various data stored in step S45 so that a torque command value CT that increases the profit rs is selected. In other words, the execution device 291 updates the second relationship defining model M2 based on the reward r calculated in step S44 and the torque command value CT selected in step S41 so that a torque command value CT that increases the profit rs is selected.
[0068] In the learning of step S52, the weights w of the neural networks of the first relationship definition model M1 and the second relationship definition model M2 are updated and optimized. The reason for updating the Q value according to the Q value update formula shown in formula (1) above is to bring it closer to the relationship expressed by formula (2) below.
[0069]
number
[0070] Therefore, the weights w of the neural network need to be updated so that Q(s_t, a_t) output from the output layer approaches the action value shown on the right side of the above formula (2). Therefore, in this embodiment, the execution device 291 defines the error function E as shown in the following formula (3) and updates the weights w of the neural network by backpropagation so that the error function E becomes smaller.
[0071]
number
[0072] Here, the value of maxQ(s_t+1, a) needs to be calculated by inputting the state s_t+1 into the neural network. Therefore, for example, the execution device 291 performs neural network training as follows. First, the execution device 291 performs repeated trials and accumulates multiple pieces of data in the storage device 292 to create mini-batches. Then, the execution device 291 uses the mini-batches to train the neural network. At this time, the execution device 291 stabilizes the training by using a target network separate from the main network. In this embodiment, the execution device 291 uses training methods such as Experience Replay and Fixed Target Q-Network when training the neural network. After step S52, the execution device 291 proceeds to step S53.
[0073] In step S53, the execution device 291 updates the episode number EN, which indicates the number of attempts. Specifically, the execution device 291 adds "1" to the episode number EN at the time of processing in step S53, and sets the new episode number EN to that value. The initial value of the episode number EN is "0". After step S53, the execution device 291 advances the process to step S54.
[0074] In step S54, the execution device 291 determines whether the number of episodes EN is equal to or greater than a predetermined specified number. Here, the specified number is a threshold for determining whether enough trials have been repeated to determine that learning of the relational specified model M has been completed. An example of the specified number is "1000." In step S54, if the execution device 291 determines that the number of episodes EN is less than the specified number (S54: NO), the execution device 291 advances the process back to step S11. In other words, the execution device 291 executes the next trial.
[0075] On the other hand, if the execution device 291 determines in step S54 that the number of episodes EN is equal to or greater than the specified number (S54: YES), the execution device 291 proceeds to step S55. In other words, the execution device 291 proceeds to step S55 when the number of trials has been repeated enough to determine that learning of the relationship specified model M has been completed.
[0076] In step S55, the execution device 291 stores the relational definition model M as a learned model in the storage device 292. After step S55, the execution device 291 ends the learning control.
[0077] <Engine stop control> Next, engine stop control executed by the control device 90 will be described with reference to Fig. 6. This engine stop control is control for stopping the rotating crankshaft 11 within a predetermined specified angle range RR. The engine stop control is control executed by the control device 90 of the vehicle 100, for example, at a stage after the development of the vehicle 100. In this embodiment, the execution device 91 of the control device 90 starts the engine stop control every predetermined period PA, with the necessary condition being the existence of a request to change the internal combustion engine 10 from an operating state to a stopped state of the internal combustion engine 10. An example of the predetermined period PA for the engine stop control is 10 msec.
[0078] As shown in Fig. 6, when the execution unit 91 of the control device 90 starts engine stop control, it executes the processing of step S71. In step S71, the execution unit 91 acquires the state s of the vehicle 100. Specifically, the execution unit 91 acquires the engine rotation speed NE, the crank angle SC, the previous command value CTA, and the elapsed period PE as state variables indicating the state s of the vehicle 100. The processing of step S71 is the same as the processing of step S21 described above. After step S71, the execution unit 91 advances the processing to step S72.
[0079] In step S72, the execution unit 91 determines whether the engine speed NE in the state variables acquired in step S71 is equal to or greater than a predetermined specified speed NEA. The process of step S72 is the same as the process of step S22 described above. In step S72, if the execution unit 91 determines that the engine speed NE is equal to or greater than the specified speed NEA (S72: YES), the execution unit 91 proceeds to step S81.
[0080] In step S81, the execution unit 91 selects a torque command value CT corresponding to one of the multiple Q values based on multiple Q values output by inputting the state variables acquired in step S71 to the first relationship specification model M1. Specifically, the execution unit 91 selects the torque command value CT as follows. First, the execution unit 91 extracts the engine speed NE and the elapsed period PE from the state variables acquired in step S71. Next, the execution unit 91 inputs the extracted engine speed NE and elapsed period PE to the first relationship specification model M1 to output multiple Q values. Furthermore, the execution unit 91 identifies the maximum Q value from the multiple Q values. Then, the execution unit 91 selects the torque command value CT corresponding to the maximum Q value. After step S81, the execution unit 91 advances the process to step S82.
[0081] In step S82, the execution device 91 controls the first motor generator 61 in accordance with the action a selected in step S81. That is, the execution device 91 controls the first motor generator 61 in accordance with the torque command value CT selected in step S81. The processing of step S82 is the same as the processing of step S32 described above. After step S82, the execution device 91 ends the current engine stop control.
[0082] On the other hand, if the execution unit 91 determines in step S72 that the engine rotation speed NE is lower than the specified rotation speed NEA (S72: NO), the execution unit 91 advances the process to step S91.
[0083] In step S91, the execution unit 91 selects a torque command value CT corresponding to one of the multiple Q values based on multiple Q values output by inputting the state variables acquired in step S71 to the second relationship definition model M2. Specifically, the execution unit 91 selects the torque command value CT as follows. First, the execution unit 91 extracts the engine speed NE, the crank angle SC, and the previous command value CTA from the state variables acquired in step S71. Next, the execution unit 91 inputs the extracted engine speed NE, the crank angle SC, and the previous command value CTA to the second relationship definition model M2 to output multiple Q values. Furthermore, the execution unit 91 identifies the maximum Q value from the multiple Q values. Then, the execution unit 91 selects the torque command value CT corresponding to the maximum Q value. After step S91, the execution unit 91 advances the process to step S92.
[0084] In step S92, the execution device 91 controls the first motor generator 61 in accordance with the action a selected in step S91. That is, the execution device 91 controls the first motor generator 61 in accordance with the torque command value CT selected in step S91. The processing of step S92 is the same as the processing of step S42 described above. After step S92, the execution device 91 ends the current engine stop control.
[0085] <Operation of this embodiment> 2, the execution device 291 of the computer 290 of the learning device 200 executes learning control. Through this learning control, attempts are repeatedly made to stop the rotating crankshaft 11 within a predetermined specified angle range RR, and the relationship specification model M is updated so that a torque command value CT that increases the profit rs is selected. In other words, the relationship specification model M is learned so that the crankshaft 11 can be stopped quickly when the crank angle SC is within the specified angle range RR.
[0086] <Effects of this embodiment> (1) In this embodiment, in addition to the engine speed NE and the crank angle SC, the previous command value CTA and the elapsed period PE are used as state variables indicating the state s of the vehicle 100. The inventors have confirmed that by using the above state variables, the relationship defining model M can be appropriately learned even by machine learning. This makes it possible to generate the relationship defining model M for quickly stopping the crankshaft 11 at the crank angle SC at which it should be stopped, while reducing the amount of human effort required compared to when a human generates the relationship defining model M based on experiments, for example.
[0087] (2) For example, as shown in Fig. 4(b), assume that fuel combustion in the cylinders of the internal combustion engine 10 stops at time T1. Then, as shown in Fig. 4(a), after time T1 when an attempt to stop the crankshaft 11 is started, the engine rotation speed NE changes abruptly. Therefore, after time T1, the detected sound pressure detected as noise generated in the vehicle 100 may become excessively large, or the detected acceleration detected as vibration generated in the vehicle 100 may become excessively large.
[0088] In this regard, in step S34, the executing device 291 calculates the reward r for the torque command value CT selected in step S31 when the detected sound pressure is equal to or less than a predetermined specified sound pressure so that the profit rs is larger than when the detected sound pressure is larger than the specified sound pressure. Also, the executing device 291 calculates the reward r for the torque command value CT selected in step S31 so that the profit rs is larger when the detected acceleration is equal to or less than a predetermined specified acceleration so that the profit rs is larger than when the detected acceleration is larger than the specified acceleration. Then, in step S52, the executing device 291 updates the first relationship specifying model M1 based on the reward r calculated in step S34 and the torque command value CT selected in step S31 so that a torque command value CT that increases the profit rs is selected. This makes it possible to generate a relationship specifying model M for quickly stopping the crankshaft 11 at the crank angle SC at which it should be stopped while suppressing the detected sound pressure and detected acceleration in the vehicle 100.
[0089] (3) For example, when an action a_t is selected according to a state s_t depending on the structure of the vehicle 100, the delay period, which is the period until one or more changes in the detected sound pressure and the detected acceleration begin to occur according to that action a_t, differs.
[0090] In this regard, in step S34, the execution device 291 of the computer 290 acquires the detected sound pressure and the detected acceleration after a delay period from when the state variables are acquired after controlling the first motor generator 61. This allows the reward r according to the detected sound pressure and the detected acceleration to be accurately calculated by taking into account the delay in the detected sound pressure and the detected acceleration as described above.
[0091] (4) For example, as shown in FIG. 4B, assume that fuel combustion in the cylinders of the internal combustion engine 10 stops at time T1. In this case, as shown in FIG. 4A, during the period from time T1 onward until time T2, which is the period immediately after the start of an attempt to stop the crankshaft 11, i.e., when the engine speed NE is relatively high, it is required to quickly reduce the engine speed NE. On the other hand, during the period from time T2 onward until the crankshaft 11 is subsequently stopped until time T3, i.e., when the engine speed NE is relatively low, it is required to accurately stop the crankshaft 11 at the crank angle SC at which the crankshaft 11 should be stopped. In other words, the requirements for control of the first motor-generator 61 differ depending on the engine speed NE.
[0092] In this regard, as shown in FIG. 2, in step S22, the execution unit 291 determines whether the engine speed NE in the state variables acquired in step S21 is equal to or greater than a predetermined specified speed NEA. If the execution unit 291 determines in step S22 that the engine speed NE is equal to or greater than the specified speed NEA, the process proceeds to step S31. In step S31, the execution unit 291 selects a torque command value CT corresponding to one of the multiple Q values based on multiple Q values output by inputting the state variables acquired in step S21 into the first relationship definition model M1. Then, the execution unit 291 updates the first relationship definition model M1 based on the torque command value CT selected in step S31 so that a torque command value CT that increases the profit rs is selected. On the other hand, if the execution unit 291 determines in step S22 that the engine speed NE is less than the specified speed NEA, the process proceeds to step S41. In step S41, the execution unit 291 selects a torque command value CT corresponding to one of the multiple Q values based on multiple Q values output by inputting the state variables acquired in step S21 into the second relationship specification model M2. Then, the execution unit 291 updates the second relationship specification model M2 based on the torque command value CT selected in step S41 so that a torque command value CT that maximizes the profit rs is selected. Therefore, according to this embodiment, the first relationship specification model M1 and the second relationship specification model M2 are generated in accordance with the requirements for control of the first motor-generator 61 according to the engine rotation speed NE. By generating two types of models in this way, it is expected that learning can be completed in a shorter period of time and that the accuracy of learning can be improved compared to, for example, generating one type of model to satisfy two requirements.
[0093] (5) As shown in FIG. 4(c), in the period from time T2 onward to time T3, which is the period immediately before the crankshaft 11 is stopped, the torque command value CT becomes positive torque and negative torque. When the torque command value CT changes from one of positive torque and negative torque to the other in this manner, vibrations may occur in the crankshaft 11 and the like due to changes in the direction of the torque applied to the crankshaft 11 from the first motor-generator 61. Therefore, the more times the torque command value CT changes from one of positive torque and negative torque to the other, the more likely the crankshaft 11 and the like are to vibrate.
[0094] In this regard, in step S44, the execution unit 291 acquires, as the number of changes NC, the number of times the torque changes from one of positive torque and negative torque to the other within a period from the stop-time crank angle SCA to the crank angle SC that is a predetermined angle earlier. The execution unit 291 then calculates a reward r for the torque command value CT selected in step S41 so that the profit rs increases as the number of changes NC decreases. In step S52, the execution unit 291 updates the second relationship defining model M2 based on the reward r calculated in step S44 and the torque command value CT selected in step S41 so that a torque command value CT that increases the profit rs is selected. This makes it possible to prevent the number of changes NC from becoming excessively large for the second relationship defining model M2, since the number of changes NC is the subject of evaluation. In other words, it is possible to prevent excessive vibration of the crankshaft 11 and the like caused by the torque command value CT changing from one of positive torque and negative torque to the other.
[0095] (6) In step S44, the execution unit 291 acquires a change period PC indicating the period of change of the torque command value CT applied to the crankshaft 11. This change period PC is a value obtained by multiplying a first average value, which is the average value of the angle range of the crank angle SC in which the torque command value CT is continuously positive torque, by a second average value, which is the average value of the angle range of the crank angle SC in which the torque command value CT is continuously negative torque. Therefore, the change period PC increases as the first average value and the second average value increase. Furthermore, for example, if the sum of the first average value and the second average value is the same, the change period PC increases as the absolute value of the difference between the first average value and the second average value decreases. Furthermore, the larger the change period PC, the slower the rate of change of the torque applied to the crankshaft 11 from the first motor-generator 61 tends to be. Therefore, the larger the change period PC, the easier it is for the crankshaft 11 to stop smoothly. In other words, the larger the change period PC, the less likely the crankshaft 11 to vibrate. Then, in step S44, the execution device 291 calculates the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the change period PC increases. Furthermore, in step S52, the execution device 291 updates the second relationship defining model M2 based on the reward r calculated in step S44 and the torque command value CT selected in step S41 so that a torque command value CT that increases the profit rs is selected. As a result, the change period PC becomes the target of evaluation for the second relationship defining model M2, and the crankshaft 11 can be stopped smoothly.
[0096] <Example of change> This embodiment can be modified as follows: This embodiment and the following modifications can be combined and implemented within the scope of technical compatibility.
[0097] In the above embodiment, the learning control may be changed. For example, the state variable in step S21 may be changed. Specifically, the execution unit 291 may acquire other values as state variables indicating the state s of the vehicle 100 in addition to the engine speed NE, the crank angle SC, the previous command value CTA, and the elapsed period PE. An example of the other value is a past torque command value CT other than the previous command value CTA. As a specific example, the execution unit 291 may acquire a second previous value, which is the torque command value CT used in the control of the first motor generator 61 immediately before the first previous value, which is the previous command value CTA. As a specific example, the execution unit 291 may acquire a third previous value, which is the torque command value CT used in the control of the first motor generator 61 immediately before the second previous value. In other words, the execution unit 291 may acquire multiple past torque command values CT as state variables indicating the state s of the vehicle 100. The number of past torque command values CT acquired by the execution unit 291 as state variables may be adjusted as appropriate.
[0098] For example, the state variables input to the first relationship definition model M1 in step S31 may be changed. As a specific example, the execution device 291 may input all of the state variables acquired in step S21 to the first relationship definition model M1.
[0099] For example, the number of types of Q values output from the first relationship defining model M1 in step S31 may be changed. As a specific example, the number of candidates for the torque command value CT of the first motor generator 61 may be less than five or more than five. The number of types of Q values output from the first relationship defining model M1 may be changed according to the number of candidates for the torque command value CT.
[0100] For example, in step S34, the timing for acquiring the detected sound pressure and the detected acceleration may be changed. As a specific example, the execution device 291 may acquire the detected sound pressure and the detected acceleration at a time point before the delay period has elapsed since the state variables were acquired after the control of the first motor generator 61. Furthermore, as a specific example, the execution device 291 may acquire the detected sound pressure and the detected acceleration at a time point after the delay period has elapsed since the state variables were acquired after the control of the first motor generator 61. In other words, the predetermined period PA does not have to be the same value as the delay period. Therefore, the predetermined period PA can be changed.
[0101] For example, in step S34, the way in which the reward r for the action a selected in step S31 is calculated may be changed. Specifically, the detected acceleration may be one or two predetermined ones of the longitudinal acceleration GX, the lateral acceleration GY, and the vertical acceleration GZ. Furthermore, as a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S31 regardless of whether the detected acceleration is equal to or less than a predetermined specified acceleration. Furthermore, as a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S31 regardless of whether the detected sound pressure is equal to or less than a predetermined specified sound pressure. Furthermore, as a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S31 such that the profit rs increases as the integrated value of the engine rotation speed NE since the start of the trial decreases. Furthermore, as a specific example, the execution device 291 may calculate an additional reward in addition to the first reward, the second reward, and the third reward. As an example, the shorter the period from the start of a trial to the completion of the trial, the larger the additional reward calculated by the execution device 291. Then, the execution device 291 calculates the sum of the first reward, the second reward, the third reward, and the additional reward as the reward r for the torque command value CT selected in step S31. Even in this case, the execution device 291 can calculate the reward r for the torque command value CT selected in step S31 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions.
[0102] For example, the state variables input to the second relationship definition model M2 in step S41 may be changed. As a specific example, the execution device 291 may input all of the state variables acquired in step S21 to the second relationship definition model M2.
[0103] For example, the number of types of Q values output from the second relationship defining model M2 in step S41 may be changed. As a specific example, the number of candidates for the torque command value CT of the first motor generator 61 may be less than five or more than five. The number of types of Q values output from the second relationship defining model M2 may be changed according to the number of candidates for the torque command value CT.
[0104] For example, in step S44, the way in which the reward r for the action a selected in step S41 is calculated may be changed. As a specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S41 regardless of the change period PC. As another specific example, the execution device 291 may calculate the reward r for the torque command value CT selected in step S41 regardless of the change count NC. As a specific example, similar to step S34, the execution device 291 may calculate the reward r for the torque command value CT selected in step S41 so that when the detected sound pressure is equal to or less than a predetermined specified sound pressure, the profit rs is larger than when the detected sound pressure is higher than the predetermined sound pressure. As a specific example, similar to step S34, the execution device 291 may calculate the reward r for the torque command value CT selected in step S41 so that when the detected acceleration is equal to or less than a predetermined specified acceleration, the profit rs is larger than when the detected acceleration is higher than the predetermined specified acceleration. Furthermore, as a specific example, the execution device 291 may calculate an additional reward in addition to the fourth reward, fifth reward, and sixth reward. As an example, the shorter the period from the start of a trial to the completion of the trial, the larger the additional reward calculated by the execution device 291. Then, the execution device 291 calculates the sum of the fourth reward, fifth reward, sixth reward, and additional reward as the reward r for the torque command value CT selected in step S41. Even in this case, the execution device 291 can calculate the reward r for the torque command value CT selected in step S41 so that the profit rs increases as the crankshaft 11 is stopped with fewer actions.
[0105] In the above embodiment, the configuration of the learning device 200 may be changed. For example, the first relationship specifying model M1 is not limited to a model generated by DQN. As a specific example, the first relationship specifying model M1 may be a so-called Q table or map. Note that, when the first relationship specifying model M1 is a Q table or the like, the inventors have confirmed that an appropriate model is easily generated when the predetermined period PA has the same value as the delay period. Furthermore, when the first relationship specifying model M1 is a Q table or the like, it may be possible to prevent the processing load of the control device 90 from becoming excessively large during engine stop control.
[0106] For example, the second relationship specifying model M2 is not limited to a model generated by DQN. Specifically, the second relationship specifying model M2 may be a so-called Q table or map. The inventors have confirmed that when the second relationship specifying model M2 is a Q table or the like, it is easier to generate an appropriate model by setting the predetermined period PA to the same value as the delay period. Furthermore, when the second relationship specifying model M2 is a Q table or the like, it may be possible to prevent the processing load on the control device 90 from becoming excessively large during engine stop control.
[0107] For example, the relationship defining model M does not have to include the first relationship defining model M1 and the second relationship defining model M2 as multiple models. In other words, in the learning control, the execution device 291 may generate one type of model. In this case, the determination process of step S22 in the learning control can be omitted.
[0108] For example, the relationship definition model M may include three or more types of models. In this case, in the learning control, the execution unit 291 may generate three or more types of models according to the magnitude of the engine rotation speed NE.
[0109] For example, the configuration of the computer 290 may be changed. Specifically, the computer 290 may be configured as a circuit including one or more processors that execute various processes according to a computer program (software). Note that the computer 290 may also be configured as a circuit including one or more dedicated hardware circuits, such as an application-specific integrated circuit (ASIC), that execute at least some of the various processes, or a combination thereof. The processor includes a CPU and memory such as RAM and ROM. The memory stores program code or instructions configured to cause the CPU to execute processes. The memory, i.e., computer-readable medium, includes any medium that can be accessed by a general-purpose or dedicated computer.
[0110] In the above embodiment, the configuration of the vehicle 100 may be changed. For example, the vehicle 100 may include only one of the first motor generator 61 and the second motor generator 62. In other words, the present technology can be applied to any vehicle 100 in which a motor generator is connected to the crankshaft 11 of the internal combustion engine 10. [Explanation of symbols]
[0111] 10...internal combustion engine 11...crankshaft 20...power split mechanism 30...automatic transmission 61...first motor generator 62...second motor generator 66...first inverter 67...second inverter 68...battery 69...drive wheels 71...crank angle sensor 72...accelerator operation amount sensor 73...vehicle speed sensor 80...connector 90...control device 91...execution device 92...storage device 92A...control program 100...vehicle 200...learning device 210...microphone 220...acceleration sensor 240...connector 250...input device 260...display 290...computer 291...execution device 292...storage device 292A...control program M...relationship definition model M1...first relationship definition model M2...second relationship definition model
Claims
1. A computer-implemented reinforcement learning method for a vehicle including an internal combustion engine and a motor generator connected to a crankshaft of the internal combustion engine, the method repeatedly performing attempts to stop the rotating crankshaft within a predetermined specified angle range by applying torque from the motor generator to the crankshaft, the computer comprises an execution unit and a storage unit; the storage device stores a relational definition model that outputs an index value indicating a value of an action corresponding to each of a plurality of candidates for a torque command value of the motor generator in the trial in response to an input of a state variable indicating a state of the vehicle; the state variables include an engine rotation speed which is a rotation speed of the crankshaft, a crank angle which is an angular position of the crankshaft, the torque command value used in the previous control of the motor-generator with respect to the torque command value corresponding to the current index value output by the relationship specification model, and a period of time elapsed since the start of the trial; The execution device: For each of the trials, obtaining the state variables; selecting the torque command value corresponding to one of the plurality of index values based on the plurality of index values output by inputting the acquired state variables into the relationship definition model; controlling the motor generator in accordance with the selected torque command value; acquiring a crank angle at a stop, which is the crank angle when the crankshaft stops, after controlling the motor generator; calculating the reward for the selected torque command value such that the profit, which is the sum of rewards when the trial is completed, increases as the acquired stop-time crank angle approaches the specified angle range and the crankshaft is stopped with fewer actions; updating the relationship definition model based on the calculated reward and the selected torque command value so that the torque command value that increases the profit is selected; Run Reinforcement learning methods.
2. The execution device: After controlling the motor generator, in addition to the crank angle at the stop, a detected sound pressure detected as noise generated in the vehicle and a detected acceleration detected as vibration generated in the vehicle are acquired; Calculating the reward for the torque command value selected so that the profit will be larger when the acquired detected sound pressure is equal to or smaller than a predetermined specified sound pressure, compared to when the acquired detected sound pressure is larger than the specified sound pressure, and calculating the reward for the torque command value selected so that the profit will be larger when the acquired detected acceleration is equal to or smaller than a predetermined specified acceleration, compared to when the acquired detected acceleration is larger than the specified acceleration; Run The reinforcement learning method according to claim 1 .
3. When a delay period is defined as a period from when the state variable is acquired until when one or more of the detected sound pressure and the detected acceleration start to change due to control of the motor generator in response to the state variable, The execution device: acquiring the detected sound pressure and the detected acceleration after a delay period from acquiring the state variables after controlling the motor generator; Run The reinforcement learning method according to claim 2 .
4. the relationship definition model includes a first relationship definition model and a second relationship definition model different from the first relationship definition model; The execution device: determining whether the engine rotation speed included in the acquired state variable is equal to or greater than a predetermined specified rotation speed; when it is determined that the engine rotation speed is equal to or higher than the specified rotation speed, selecting the torque command value corresponding to one of the plurality of index values based on the plurality of index values output by inputting the acquired state variables into the first relationship specification model; when it is determined that the engine rotation speed is less than the specified rotation speed, selecting the torque command value corresponding to one of the plurality of index values based on the plurality of index values output by inputting the acquired state variables into the second relationship specification model; Run The reinforcement learning method according to any one of claims 1 to 3.
5. When the torque command value in the direction in which the crankshaft is rotated forward is a positive torque and the torque command value in the direction in which the crankshaft is rotated backward is a negative torque, The execution device: After the control of the motor generator, in addition to the stop crank angle, the number of times that the positive torque and the negative torque change from one to the other within a predetermined specified period is acquired; calculating the reward for the selected torque command value such that the profit increases as the acquired number of changes decreases; Run The reinforcement learning method according to claim 1 .
Citation Information
Patent Citations
Motor control device
JP2020011530A