VEHICLE CONTROL DEVICE, VEHICLE CONTROL SYSTEM, VEHICLE LEARNING DEVICE AND VEHICLE LEARNING PROCEDURE

The vehicle control system addresses inefficiencies in reinforcement learning by limiting the search range and adjusting rewards based on gear shift type, ensuring efficient and timely achievement of optimal gear shift actions that meet high-priority requirements.

DE102021115776B4Active Publication Date: 2026-04-23TOYOTA JIDOSHA KK
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2021-06-18
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Reinforcement learning for vehicle gear ratio control can be inefficient due to an unbounded search range, leading to prolonged time to reach optimal values and difficulty in meeting high-priority requirements such as heat generation, shift time, and speed differences, especially when rewards are not differentiated by gear shift type.

Method used

A vehicle control system that uses a processor to perform detection, operation, and reward calculation processes, including bounding and update processing to limit the search range and adjust rewards based on gear shift type, thereby focusing on high-priority requirements.

Benefits of technology

This approach reduces the time to achieve optimal gear shift actions by narrowing the search range and adjusting rewards, ensuring that high-priority requirements are met efficiently during the learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Vehicle control device, which includes: a processor; and a memory (46) wherein the memory (46) stores relationship-defining data to define a relationship between a state of a vehicle and an action variable, which is a variable relating to operations of a transmission installed in the vehicle, the processor is designed to execute: an investigative processing method to determine the condition of the vehicle based on a sensor reading, an operational processing to control the transmission based on a value of the action variable determined by the state of the vehicle as determined in the investigation processing and the relationship-defining data, a reward calculation process to give a larger reward when the vehicle's characteristics meet a reference, instead of failing to meet the reference, based on the vehicle's state determined in the discovery processing, an update processing to update the relationship-defining data with the state of the vehicle determined in the discovery processing, the value of the action variables used in the operation of the transmission, and the reward corresponding to the operation, as inputs into a pre-set update map, a counting process for counting an update counter by the update processor, and a bounding process to limit downwards a range used by the operation processing, in which a value other than a value that maximizes an expected utility with respect to the reward, of values ​​of the action variable that specify the relationship-defining data, when the update counter is relatively large, and The processor is designed to output updated relationship-defining data, thus increasing the expected benefit when the gearbox is operated according to the relationship-defining data, based on an update map.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a vehicle control device, a vehicle control system, a vehicle learning device and a vehicle learning method.

[0002] For example, JP 2000-250 602 A describes how to set a suitable gear ratio in accordance with the condition of a vehicle by means of reinforcement learning.

[0003] Furthermore, for a better understanding of the present invention, reference is made to DE 689 26 540 T2, in which a “shift control device and a shift control method for the shift control of an automatic transmission” are described.

[0004] The inventors investigated learning an operation to change a translation ratio using reinforcement learning. However, if the search range is not narrowed despite learning progress, the time to reach the optimal value can become lengthy.

[0005] A vehicle control device according to a first aspect of the invention comprises a processor and a memory. The memory stores relationship-defining data for defining a relationship between a vehicle state and an action variable, which is a variable relating to operations of a transmission installed in the vehicle. The processor is designed to perform: a detection process to determine the vehicle state based on a sensor reading; an operation process to control the transmission based on a value of the action variable determined by the vehicle state as determined in the detection process and the relationship-defining data; and a reward calculation process to provide a larger reward when vehicle characteristics meet a reference, rather than failing to meet the reference.Based on the vehicle state determined in the discovery processing, an update processing to update the relationship-defining data, with the vehicle state determined in the discovery processing, the value of the action variable used in the operation of the transmission, and the reward corresponding to the operation as inputs to a preset update map, a count processing to increment an update counter by the update processing, and a bounding processing to limit, to smaller values, a range used by the operation processing in which a value other than a value that maximizes an expected benefit with respect to the reward, of values ​​of the action variable that specify the relationship-defining data when the update counter is relatively large. The processor is designed to output the updated relationship-defining data.so that the expected benefit is increased when the gearbox is operated according to the relationship-defining data, based on the update map.

[0006] As reinforcement learning continues for a certain period, actions that maximize the expected benefit indicated by the relationship-defining data approach actions that actually increase benefit. Therefore, an arbitrarily or indiscriminately continued search that deviates significantly from actions that maximize the expected benefit indicated by the relationship-defining data, when the update counter of the relationship-defining data reaches a certain size, can lead to the execution of unnecessary processes to bring actions that maximize the expected benefit indicated by the relationship-defining data closer to actions that actually increase benefit. Thus, in the configuration above, a range that uses values ​​other than those that maximize the expected benefit indicated by the relationship-defining data is bounded from below when the update counter is relatively large.In other words, the search scope is limited downwards, i.e., in the direction of reduction. Therefore, actions that maximize the expected benefit indicated by the relationship-defining data can be brought close to actions that actually increase the benefit at an early stage.

[0007] In the aspect described above, limiting processing can include processing to limit the update amount when the update counter is relatively large. If the update amount of relationship-defining data is constant based on a reward obtained through a search, the period during which actions that maximize the expected benefit indicated by the relationship-defining data are performed can be long. Therefore, in the configuration described above, the update amount is limited when the update counter is large, thus preventing the relationship-defining data from being extensively updated as a result of a search for progress in reinforcement learning.Accordingly, actions that maximize the expected benefits indicated by the relationship-defining data can be brought close to actions that actually increase benefits at an early stage.

[0008] In the above aspect, the reward calculation process may include: processing to give a larger reward when the amount of heat generated in a gear ratio shift period is relatively small, and processing to change the size of the given reward in accordance with a type of shift operation, even if the amount of heat generated is the same.

[0009] There are various elements required when changing the gear ratio, and the priority level among these elements can vary depending on the type of gear shift. Therefore, it can be difficult to achieve learning outcomes that meet high-priority requirements, such as the amount of heat generated, if the reward for the same amount of heat generated is set regardless of the gear shift. Furthermore, the difficulty of fulfilling each requirement at a predetermined standard can vary depending on the type of gear shift.Accordingly, fulfilling the requirements regarding the amount of heat generated, which is one of these requirement elements, can become difficult if the reward for the same amount of heat generated is set the same regardless of the type of gear shift. Therefore, the configuration described above changes the rewards given for the same amount of heat generated depending on the type of gear shift, thereby increasing the likelihood of obtaining learning outcomes that fulfill high-priority requirement elements and allowing the learning process to proceed smoothly.

[0010] In the aspect above, the reward calculation process may include: processing to give a larger reward when a gear shift time, which is the time required to change the gear ratio, is relatively small, and processing to change the size of the given reward in accordance with a type of gear shift operation, even if the gear shift time is the same.

[0011] There are various elements required when changing the gear ratio, and the priority level among these elements can vary depending on the type of gear shift. Therefore, it can be difficult to achieve learning outcomes related to shift time, which is one of these elements, by fulfilling high-priority elements if the reward for the same shift time is set the same regardless of the type of gear shift. The difficulty of fulfilling each element within a predetermined standard can also vary depending on the type of gear shift.Therefore, with regard to gear shift time, which is one of these requirement elements, it can become difficult to meet the requirements if the size of the reward for the same gear shift time is set regardless of the type of gear shift. The configuration described above therefore changes the rewards given for the same gear shift time depending on the type of gear shift, thereby increasing the likelihood of achieving learning outcomes that meet high-priority requirement elements and allowing the learning process to progress smoothly.

[0012] In the above aspect, the reward calculation process may include: processing to give a larger reward if the excess amount of an input shaft speed of the transmission during a gear ratio shift period exceeding a reference speed is relatively small, and processing to change the size of the given reward in accordance with a type of shift operation, even if the excess amount is the same.

[0013] There are various elements required when changing the gear ratio, and the priority level among these elements can vary depending on the type of gear shift. Therefore, it can be difficult to achieve learning outcomes that meet high-priority requirements if the reward for the same override amount is set the same regardless of the gear shift type. Similarly, the difficulty of meeting each requirement to a predetermined standard can also vary depending on the type of gear shift.Therefore, if the reward size for the same overage amount is set the same regardless of the type of gear shift, it can be difficult to meet the requirements, which are one of these requirement elements. The configuration described above therefore changes the rewards given for the same overage amount depending on the type of gear shift, thus increasing the likelihood of achieving learning outcomes that fulfill high-priority requirement elements and allowing the learning process to progress smoothly.

[0014] In the above aspect, the reward calculation process may include: processing to give a larger reward when the amount of heat generated in a gear ratio shift period is relatively small, and processing to change the size of the given reward in accordance with the amount of accelerator actuation, even if the amount of heat generated is the same.

[0015] There are various elements required when switching the gear ratio, and the priority level among these elements can vary depending on the accelerator actuation amount. Therefore, with respect to the amount of heat generated, which is one of these elements, it can be difficult to achieve learning outcomes that satisfy high-priority elements if the reward for the same amount of heat generated is set to the same regardless of the accelerator actuation amount. Furthermore, the difficulty of fulfilling each element within a predetermined standard can vary depending on the accelerator actuation amount.Therefore, fulfilling the requirements regarding the amount of heat generated, which is one of these requirement elements, can become difficult if the reward size for the same amount of heat generated is set the same regardless of the accelerator actuation amount. The configuration described above therefore changes the rewards given for the same amount of heat generated depending on the accelerator actuation amount, thereby increasing the likelihood of achieving learning outcomes that fulfill high-priority requirement elements and allowing the learning process to progress smoothly.

[0016] In the above aspect, the reward calculation process may include: processing to give a larger reward when a gear shift time, which is the time required to change the gear ratio, is relatively small, and processing to change the size of the given reward in accordance with the amount of accelerator actuation, even if the gear shift time is the same.

[0017] There are various elements required when changing the gear ratio, and the priority level among these elements can vary depending on the amount of accelerator pedal input. Therefore, with regard to shift time, which is one of the required elements, it can be difficult to achieve learning outcomes that meet high-priority requirements if the reward size for the same shift time is set regardless of the amount of accelerator pedal input. Furthermore, the difficulty of meeting each required element to a predetermined standard can vary depending on the amount of accelerator pedal input.Therefore, fulfilling the requirements regarding gear shift time, which is one of these requirements, can become difficult if the reward size for the same gear shift time is set independently of the accelerator pedal input. The configuration described above modifies rewards earned for the same gear shift time based on the accelerator pedal input, thereby increasing the likelihood of achieving learning outcomes that fulfill high-priority requirements and allowing the learning process to progress smoothly.

[0018] In the above aspect, the reward calculation process may include: processing to give a larger reward if the excess speed of a transmission input shaft during a gear ratio shift period exceeding a reference speed is relatively small, and processing to change the size of the given reward in accordance with the amount of accelerator actuation, even if the excess speed is the same.

[0019] There are various elements required when switching the gear ratio, and the priority level among these elements can vary depending on the accelerator actuation amount. Therefore, with respect to the overage amount, which is one of the required elements, it can be difficult to achieve learning outcomes that meet high-priority requirements if the reward size for the same overage amount is set the same regardless of the accelerator actuation amount. Furthermore, the difficulty of setting each required element to a predetermined standard can vary depending on the accelerator actuation amount.Therefore, fulfilling the requirements can become difficult with respect to the overage amount, which represents one of these requirement elements, if the reward for the same overage amount is set to the same value regardless of the accelerator actuation amount. The configuration above therefore modifies rewards given for the same overage amount depending on the accelerator actuation amount, thereby increasing the likelihood of achieving learning outcomes that fulfill high-priority requirement elements and allowing the learning process to progress smoothly.

[0020] A vehicle control system according to a second aspect of the invention comprises the processor and the memory of the vehicle control device according to the first aspect, wherein the processor comprises a first processor which is installed in the vehicle and a second processor which is separate from an on-board device, and the first processor is designed to perform at least the detection processing and the operation processing, and the second processor is designed to perform at least the update processing.

[0021] According to the configuration above, the second processor performs the update processing, and consequently, the workload of the first processor can be reduced compared to when the first processor performs the update processing. It should be noted that saying the second processor is a device separate from an onboard unit means that the second processor is not an onboard unit.

[0022] A vehicle control device according to a third aspect of the invention comprises the first processor of the vehicle control system of the above aspect.

[0023] A vehicle learning device according to a fourth aspect of the invention comprises the second processor of the vehicle control system of the above aspect.

[0024] In a vehicle learning method according to a fifth aspect of the invention, a computer is caused to perform the detection processing, the operation processing, the reward calculation process, the update processing, the counting processing and the limiting processing of the vehicle control device of the above aspect.

[0025] According to the procedure described above, the same advantages as those of the aspect above can be gained.

[0026] The following describes the features, advantages, and technical and industrial significance of exemplary embodiments of the invention with reference to the accompanying drawings, in which the same reference numerals denote the same elements and wherein: Fig. 1 is a diagram showing a control device and a drive train thereof according to a first embodiment; Fig. 2 is a flowchart showing the processing procedures that the control device performs according to the first embodiment; Fig. 3 is a flowchart showing detailed procedures for part of the processing performed by the control device according to the first embodiment; Fig. 4 is a flowchart showing the processing procedures that the control device performs according to the first embodiment; Fig. 5 is a flowchart showing processing procedures performed by a control device according to a second embodiment; Fig. 6 is a diagram showing a configuration of a vehicle control system according to a third embodiment; and Fig. 7 is a flowchart in which section (a) and section (b) show processing procedures that the vehicle control system performs. First embodiment

[0027] A power distribution device 20 is mechanically connected to a crankshaft 12 of an internal combustion engine 10, as shown in Fig. Figure 1 shows the power distribution device 20, which distributes the power of the internal combustion engine 10, a first motor-generator 22, and a second motor-generator 24. The power distribution device 20 includes a planetary gear mechanism. The crankshaft 12 is mechanically connected to a carrier C of the planetary gear mechanism, a rotating shaft 22a of the first motor-generator 22 is mechanically connected to a sun gear S of the first motor-generator 22, and a rotating shaft 24a of the second motor-generator 24 is mechanically connected to a ring gear R of the second motor-generator 24. An output voltage of a first inverter 23 is applied to a terminal of the first motor-generator 22. Furthermore, an output voltage of a second inverter 25 is applied to a terminal of the second motor-generator 24.

[0028] In addition to the rotating shaft 24a of the second motor-generator 24, drive gears 30 are also mechanically connected to the ring gear R of the power distribution device 20 via a gearbox 26. Furthermore, a driven shaft 32a of an oil pump 32 is mechanically connected to the support C. The oil pump 32 is a pump that draws oil into an oil pan 34 and discharges the oil as operating oil into the gearbox 26. It should be noted that the pressure of the operating oil discharged by the oil pump 32 is adjusted by a hydraulic pressure control circuit 28 in the gearbox 26 and is thus used as operating oil. The hydraulic pressure control circuit 28 comprises several solenoid valves 28a and is a circuit that controls the state of the flowing operating oil and the hydraulic pressure of the operating oil by supplying an electric current to the solenoid valves 28a.

[0029] A control device 40 controls the internal combustion engine 10 and different types of operating sections of the internal combustion engine 10 to regulate torque, exhaust gas component ratio, and other controlled variables. The control device 40 further controls the first motor-generator 22 and actuates the first inverter 23 to regulate torque, speed, and other controlled variables. The control device 40 further controls the second motor-generator 24 and actuates the second inverter 25 to regulate torque, speed, and other controlled variables.

[0030] The control of the aforementioned control variables by the control device 40 is based on an output signal Scr from a crank angle sensor 50, an output signal Sm1 from a first rotation angle sensor 52, which detects the rotation angle of the rotating shaft 22a of the first motor generator 22, and an output signal Sm2 from a second rotation angle sensor 54, which detects the rotation angle of the rotating shaft 24a of the second motor generator 24. Furthermore, the control device 40 is based on an oil temperature Toil, which is the temperature of oil detected by an oil temperature sensor 56, a vehicle speed SPD detected by a vehicle speed sensor 58, and an accelerator pedal actuation amount ACCP, which is the depressurization amount of an accelerator pedal 60 detected by an accelerator sensor 62.

[0031] The control device 40 comprises a central processing unit (CPU) 42, a read-only memory (ROM) 44, a memory 46 which is an electrically rewritable, non-volatile memory, and a peripheral circuit 48 that can communicate via a local area network 49. The peripheral circuit 48 includes a circuit that generates clock signals to define internal operations, a power supply circuit, a reset circuit, and so on. The control device 40 regulates the control variables through the CPU 42, which executes programs stored in the ROM 44.

[0032] Fig. Figure 2 shows processing procedures performed by the control device 40. The in Fig. The processing shown in Figure 2 is implemented by a learning program DPL stored in ROM 44, which is executed repeatedly by CPU 42, for example, in a predetermined cycle. Note that numbers following an "S" indicate step numbers for the individual processing operations.

[0033] In the Fig. In the two processing sequences shown, the CPU 42 first determines whether the current situation is a period in which gear ratios are to be changed, i.e., whether the current situation is a gear-shifting period (S10). If it is determined that the current situation is a gear-shifting period (YES in S10), the CPU 42 determines the accelerator actuation amount ACCP, a gear-shifting variable ΔVsft, the oil temperature Toil, a phase variable Vpase, and a rotational speed Nm2 of the second motor generator 24 as a state s (S12). It should be noted that the gear-shifting variable ΔVsft is a variable used to determine, before and after the gear ratio changes, whether a shift from first gear to second gear, from second gear to first gear, or the like, was intended or has occurred. In other words, it is a variable for identifying the type of gear-shifting operation.The phase variable Vpase is a variable used to identify which of the three phases that determine the shift stages in a gear shift period is currently present.

[0034] In the present embodiment, a gear-shifting period is divided into Phase 1, Phase 2, and Phase 3. Phase 1 is the period from the start of the gear ratio switching control until the expiration of a preset time interval. Phase 2 is the period from the end of Phase 1 until the end of a torque phase. In other words, this is the period until the torque transmission capacity through frictional engagement elements, which are switched from an engaged state to a disengaged state by changing the gear ratio, reaches zero. The CPU 42 determines the end time of Phase 2 based on a deviation of the actual input shaft speed from an input shaft speed determined by the speed of an output shaft of the transmission 26 and the gear ratio before the gear ratio change. The input shaft speed can be in Nm².Furthermore, the CPU 42 calculates the output shaft speed according to the vehicle speed SPD. Phase 3 is the period from the end of Phase 2 until the completion of the gear shift. It should be noted that the above speed Nm2 is calculated by the CPU 42 based on the output signals Sm2.

[0035] The state s is determined by values ​​of variables whose relationship to the action variable is defined by relationship-defining data DR, which is contained in the Fig. The storage units shown in Figure 1, 46, are stored. In the first embodiment, a hydraulic pressure command value of an operating oil that actuates the friction engagement elements involved in switching the transmission ratio is represented by way of example as an action variable. In particular, the hydraulic pressure command value with respect to Phase 1 and Phase 2 is a constant value during these periods and is a hydraulic pressure command value that increases at a constant rate in Phase 3. It should be noted that the action variable for Phase 3, which is actually contained in the relationship-defining data DR, can be a pressure increase rate.

[0036] In particular, the relationship-defining data DR includes an action-value function Q. The action-value function Q is a function in which the state s and an action a are independent variables, and an expected benefit with respect to the state s and the action a is a dependent variable. In the present embodiment, the action-value function Q is a function in tabular format.

[0037] The CPU 42 then calculates the value of the action variable based on a strategy (policy) π defined by the relationship-defining data DR (S14). In the present embodiment, the strategy is, by way of example, an ε-greedy strategy (ε-greedy policy). That is, a strategy is described by way of example that determines a rule whereby, given a state s, the largest action of the action value function Q, where one independent variable is the given state s (hereinafter referred to as the greedy action ag), is prioritized, while other actions are simultaneously selected with a predetermined probability. In particular, the probability of taking an action other than the greedy action is ε / |A| whenever the total number of values ​​that the action can take is |A|.

[0038] Since, in the first embodiment, the action value function Q is presented as tabulated data, the state s, which serves as an independent variable, has a specific width. That is, if the action value function Q is defined in 10% increments with respect to the accelerator actuation amount ACCP, the accelerator actuation amount ACCP does not represent different states s simply because the accelerator actuation amount ACCP is sometimes "3%" and sometimes "6%".

[0039] The CPU 42 then regulates an applied electrical current I such that the applied electrical current I of the solenoid valves 28a assumes a value determined on the basis of a hydraulic pressure command value P* (S16). The CPU 42 then calculates a speed difference ΔNm² and a heat generation quantity CV (S18).

[0040] The speed difference ΔNm2 is a quantification of the overshoot of the speed of the input shaft of the transmission 26 during the gear shift period and is calculated as the overshoot of the speed Nm2 relative to the speed Nm2*, which is a preset reference. The CPU 42 sets the reference speed Nm2* according to the accelerator actuation amount ACCP, the vehicle speed SPD, and the gear shift variable ΔVsft. This processing can be implemented by the CPU 42 mapping the reference speed Nm2* in a state where map data, in which the accelerator actuation amount ACCP, the vehicle speed SPD, and the gear shift variable ΔVsft are input variables and the reference speed Nm2* is an output variable, is prestored in ROM 44.It should be noted that map data consists of sets of discrete values ​​of the input variables and values ​​of the output variables that correspond to the respective values ​​of the input variables. A map calculation can also be performed in which, if a value of an input variable matches one of the values ​​of the input variables in the map data, the corresponding value of the output variable in the map data is used as a calculation result; if there is no match, a value obtained by interpolating several values ​​of output variables contained in the map data is used as a calculation result.

[0041] On the other hand, in the present embodiment, the heat generation quantity CV is calculated as a quantity proportional to the product of the rotational speed difference of a pair of friction engagement elements, which switch from a non-engagement state to an engagement state, and the torque acting on them. Specifically, the CPU 42 calculates the heat generation quantity CV based on the rotational speed Nm², which is the rotational speed of the input shaft of the transmission 26, the rotational speed of the output shaft of the transmission 26, which results from the vehicle speed SPD, and the torque resulting from the accelerator actuation amount ACCP.In particular, the CPU 42 performs a map calculation of the heat generation quantity CV in a state in which map data, in which the speed of the input shaft, the speed of the output shaft and the accelerator actuation amount ACCP are input variables and the heat generation quantity CV is an output variable, have been pre-stored in ROM 44.

[0042] CPU 42 processes S16 and S18 until the current phase is complete (NO in S20). If it is determined that the current phase should be completed (YES in S20), CPU 42 updates the relationship-defining data DR through reinforcement learning (S22).

[0043] It should be noted that when the processing of S22 is complete, or a negative determination is made during the processing of S10, CPU 42 will be used. Fig. The two processing sequences shown have been completed once. Fig. Figure 3 shows the details of the processing of S22.

[0044] In the Fig. In the processing flow shown in Figure 3, CPU 42 first determines whether the phase variable Vpase is "3" (S30). If it is determined that the phase variable Vpase is "3" (YES in S30), the gear shift operation is complete, and accordingly, CPU 42 calculates a gear shift time Tsft, which is the time required for the gear shift operation (S32). CPU 42 then calculates a reward r1 in accordance with the gear shift time Tsft (S34). In particular, CPU 42 calculates a larger value for the reward r1 if the gear shift time Tsft is relatively short.

[0045] The CPU 42 then sets the largest value of the rotational speed difference ΔNm2, which was repeatedly calculated in a predetermined cycle during the processing of S18, to a maximum rotational speed difference value ΔNm2max (S36). The CPU 42 then calculates a reward r2 in accordance with the maximum rotational speed difference value ΔNm2max (S38). In particular, the CPU 42 calculates a larger value for the reward r2 if the maximum rotational speed difference value ΔNm2max is relatively small.

[0046] CPU 42 then calculates a heat generation quantity InCV, which is an integer value of the heat generation quantity CV that was repeatedly calculated in the predetermined cycle by processing in S18 (S40). CPU 42 then calculates a reward r3 in accordance with the heat generation quantity InCV (S42). Specifically, CPU 42 calculates a larger value for the reward r3 if the heat generation quantity InCV is relatively small.

[0047] For the action used in processing S16, CPU 42 then sets the sum of reward r1, reward r2, and reward r3 to reward r(S44). However, if it is determined that the phase variable Vpase is "1" or "2" (NO in S30), CPU 42 sets reward r to "0" (S46).

[0048] When the processing of S44 or S46 is complete, CPU 42 updates the action value function Q(s, a) used in the processing of S14 based on the reward r(S48). It should be noted that the action value function Q(s, a) used in the processing of S14 is the action value function Q(s, a) that contains the state s determined by the processing of S12 and the action a set by the processing of S14 as independent variables.

[0049] In the present embodiment, the action value function Q(s, a) is updated by so-called Q-learning, which is an off-policy temporal difference (TD) learning. In particular, the action value function Q(s, a) is updated by the following expression (c1). Q(s,a)←Q+α⋅{r+γ⋅maxQ(s+1,A)−Q(s,a)}

[0050] Here, the discount factor γ and learning rate α are used for the update amount "α · {r + γ · maxQ (s + 1, A) - Q (s, a)}" of the action value function Q (s, a). It is important to note that the discount factor γ and the learning rate α are both constants greater than "0" and not greater than "1". Furthermore, if the current phase is phase 1 or 2, "maxQ (s + 1, a)" represents a state variable at the time of phase completion, i.e., the largest value of the action value function Q, of which an independent variable is the state s, which will be the next time S12 is processed in the Fig. The value of S12 is to be determined from the two processing sequences shown, and the value "1" is added to it. It should be noted that if the current phase is not phase 3, the state s that will be in the next processing of S12 in the sequence shown is... Fig. In the processing sequences shown in 2, the state s is determined, which is used in the processing of S48 and to which the value "1" is added. Furthermore, if the current phase is phase 3, the state s, which this time is determined by the processing of S12 in the Fig. The processing sequences shown in the 2 are determined and set to a state s + 1.

[0051] The CPU 42 then increments an update counter N of the relationship-defining data DR (S50). It should be noted that when the processing of S50 is complete, the CPU 42... Fig. The three processing sequences shown are completed once. It should also be noted that the relationship-defining data DR at the time of shipment of a vehicle VC are data where a learning process is carried out through processing similar to that in Fig. 2 in a vehicle prototype or the like with the same specifications as the vehicle VC. That is to say, the processing of Fig. 2 is a process to update the hydraulic pressure command value P* set before the shipment of the vehicle VC to a value suitable for actual driving of the vehicle VC on the road, through reinforcement learning.

[0052] Fig. Figure 4 shows procedures for processing with respect to a setting of a search range and the learning rate α. The in Fig. The processing shown in step 4 is implemented by a program stored in ROM 44, which is repeatedly executed by the CPU 42 in a predetermined cycle.

[0053] In the Fig. In the 4 processing sequences shown, CPU 42 first determines whether the update counter N is not greater than a first predetermined value N1 (S60). The first predetermined value N1 is set to a counter at which a learning process is assumed to have advanced by a certain amount after the shipment of vehicle VC, in accordance with the individual variability of the vehicle VC.

[0054] If it is determined that the update counter N is not greater than the first predetermined value N1 (JA in S60), the CPU 42 sets an initial value α0 to the learning rate α (S62). The CPU 42 further sets an action range A used for searching to a widest initial range A0 (S64). The initial range A0 allows a maximum number of actions under the conditions that abnormal actions, which would promote deterioration of the gearbox 26, are eliminated.

[0055] On the one hand, if the CPU 42 determines that the update counter N is greater than the first predetermined value N1 (NO in S60), it then determines whether the update counter N is greater than a second predetermined value N2 (S66). The second predetermined value N2 is set to a value greater than the first predetermined value N1. If the CPU 42 determines that the update counter N is not greater than the second predetermined value N2 (YES in S66), it sets a value obtained by multiplying an initial value α0 by a correction coefficient k1 to the learning rate α (S68). The correction coefficient k1 is a value greater than "0" and less than "1". Furthermore, the CPU 42 limits the search action range A to a range in which the difference with respect to the greedy action ag at the current time is not greater than a first assumed value δ1 (S70).However, it should be noted that this area is an area encompassed by the initial area A0.

[0056] On the other hand, if it is determined that the update counter N is greater than the second predetermined value N2 (NO in S66), the CPU 42 determines whether the update counter N is not greater than a third predetermined value N3 (S72). The third predetermined value N3 is set to a value greater than the second predetermined value N2. If it is determined that the update counter N is not greater than the third predetermined value N3 (YES in S72), the CPU 42 sets a value obtained by multiplying the initial value α0 by a correction coefficient k2 to the learning rate α (S74). The correction coefficient k2 here is a value greater than 0 and less than the correction coefficient k1. Furthermore, the CPU 42 limits the search action range A to a range in which the difference with respect to the greedy action ag at the current time is not greater than a second predetermined value δ2 (S76).However, it should be noted that the second given value δ2 is smaller than the first given value δ1. Furthermore, this range is a region encompassed by the initial range A0.

[0057] On the other hand, if the CPU 42 determines that the update counter N is greater than the third predetermined value N3 (NO in S72), it sets a value obtained by multiplying the initial value α0 by a correction coefficient k3 to the learning rate α (S78). The correction coefficient k3 here is a value greater than 0 and less than the correction coefficient k2. Furthermore, the CPU 42 limits the search action range A to a range in which the difference with respect to the greedy action ag at the current time is not greater than a third predetermined value δ3 (S80). However, it should be noted that the third predetermined value δ3 is less than the second predetermined value δ2. This range is also encompassed by the initial range A0.

[0058] It should be noted that CPU 42, when processing of S64, S70, S76, or S80 is complete, is used in Fig. The four processing sequences shown are completed once. The effects and advantages of the present embodiment are described below.

[0059] During a gear-shift cycle, the CPU 42 selects a greedy action ag and regulates an electrical current for the solenoid valves 28a, while simultaneously searching for a better hydraulic pressure command value P* using actions other than greedy actions, in accordance with a predetermined probability. The CPU 42 then updates the action value function Q, which is used to identify the hydraulic pressure command value P* via Q-learning. Therefore, a suitable hydraulic pressure command value P* can be learned through reinforcement learning while the vehicle VC is actually driving.

[0060] As the update counter N of the relationship-defining data DR increases, CPU 42 also reduces the search range of action a to a range that is not very far from the greedy action ag at that time. It is conceivable that, as the update counter N grows large, the greedy action ag, which specifies the relationship-defining data DR, will approximate actions that actually increase utility. Therefore, by narrowing the search range, searches that cannot be optimal can be prevented. Thus, the greedy action ag, which specifies the relationship-defining data DR, can be brought close to actions that actually increase utility at an early stage.

[0061] According to the present embodiment described above, the effects and advantages described below can also be obtained. (1) The learning rate α is changed to a smaller value as the update counter N increases. Therefore, as the update counter N increases, the amount of the update of the relationship-defining data DR can be limited to the smaller side. Furthermore, it can be prevented that the relationship-defining data DR is strongly updated by the search result after the reinforcement learning has progressed. Therefore, the greedy action ag, which specifies the relationship-defining data DR, can be brought close to actions that actually increase utility at an early stage. Second embodiment

[0062] A second embodiment is described below with reference to the drawings, primarily with regard to the differences compared to the first embodiment.

[0063] Fig. Figure 5 shows detailed procedures for processing in S22 according to the present embodiment. The in Fig. The processing shown in Figure 5 is implemented by the CPU 42, which executes the learning program DPL stored in ROM 44.

[0064] In the Fig. In the 5 processing series shown, the CPU 42 uses the accelerator actuation amount ACCP and the gear shift variable ΔVsft when calculating the reward r1 in accordance with the gear shift time Tsft (S34a), the reward r2 in accordance with the speed difference maximum value ΔNm2max (S38a) and the reward r3 in accordance with the heat generation amount InCV (S42a).

[0065] The reasons for the rewards r1, r2, and r3 in accordance with the accelerator pedal actuation amount ACCP and the type of gear shift operation are as follows. First, this is a setting to induce learning of the greedy action ag with different priority levels with respect to the three demand elements of the accelerator pedal response, which is strongly correlated with the gear shift time Tsft; drivability, which is strongly correlated with the maximum engine speed difference value ΔNm2max; and the amount of heat generated InCV, in accordance with the accelerator pedal actuation amount ACCP and the gear shift variable ΔVsft.

[0066] This means that if the priority level for an accelerator response is set higher for shifting from second to first gear than for shifting from first to second gear, the absolute value of the reward with respect to the same shift time Tsft is set to be greater for shifting from second to first gear than for shifting from first to second gear. In this case, the priority level for the heat generation quantity InCV can be set higher with respect to the shift from first to second gear, making the absolute value of the reward r3 with respect to the heat generation quantity InCV greater than for shifting from second to first gear.

[0067] Secondly, this serves to differentiate the values ​​that the maximum speed difference ΔNm2max, the gear shift time Tsft, and the heat generation quantity InCV can assume, in accordance with the accelerator pedal actuation amount ACCP and the gear shift operation, since the torque and speed transmitted to the transmission 26 differ in accordance with the accelerator pedal actuation amount ACCP and the type of gear shift operation. Therefore, there is a concern that awarding the same reward r1 with respect to the gear shift time Tsft or the like, regardless of the accelerator pedal actuation amount ACCP and the type of gear shift operation, would make learning more difficult.

[0068] Thus, in the first embodiment, varying the rewards r1, r2, and r3 in accordance with the accelerator actuation amount ACCP and the gear shift variable ΔVsft enables learning that reflects the difference in priority with respect to the gear shift time Tsft, the speed difference ΔNm2, and the amount of heat generated InCV, in accordance with the accelerator actuation amount ACCP and the type of gear shift operation. Furthermore, the rewards r1 to r3 can be given considering the differential values ​​that the maximum speed difference value ΔNm2max, the gear shift time Tsft, and the amount of heat generated InCV can assume in accordance with the accelerator actuation amount ACCP, resulting in smooth learning progress. Third embodiment

[0069] A third embodiment is described below with reference to the drawings, primarily with regard to the differences compared to the first embodiment.

[0070] Fig. Figure 6 shows a configuration of a system according to the present embodiment. It should be noted that elements in Fig. 6, which correspond to the elements that are in Fig. Figure 1, for the sake of simplicity, is referred to by the same reference numerals and is not described again. The control device 40 of a vehicle VC(1) comprises a communication device 47 and is suitable for communicating with a data analysis center 90, as shown in Figure 1, via an external network 80 that uses the communication device 47. Fig. 6 is shown.

[0071] The data analysis center 90 analyzes data transmitted by several vehicles VC(1), VC(2), etc. The data analysis center 90 comprises a CPU 92, a ROM 94, a memory 96, and a communication device 97, which can communicate via a local network 99. It should be noted that the memory 96 is an electrically rewritable, non-volatile device and stores the relationship-defining data DR.

[0072] Fig. Figure 7 shows processing procedures for reinforcement learning according to the present embodiment. The procedures described in section (a) in Fig. The processing shown in step 7 is implemented by CPU 42, which executes a learning subroutine DPLa that is located in the Fig. ROM 44, shown in section 6, is stored. Furthermore, a [missing information] is shown in section (b) in Fig. The processing shown in Figure 7 is implemented by the CPU 92, which executes a main learning program, DPLb, stored in ROM 94. It should be noted that, for the sake of simplicity, the processing in Figure 7 is shown in Figure 7. Fig. 7, which is in Fig. The processing shown in step 2 corresponds to the same step number. Fig. The processing shown in Figure 7 is described below based on the temporal sequence of reinforcement learning.

[0073] In section (a) in Fig. In the processing sequence shown in section 7, the CPU 42 of the control device 40 first executes the processing from S10 to S18 and then determines whether the gear shifting operation is complete (S80). If it is determined that the gear shifting operation is complete (YES in S80), the CPU 42 controls the communication device 97 and transmits data, along with an identification code of the vehicle VC(1), to the vehicle VC (S82). This data includes the state s, the action a, the speed difference ΔNm², the heat generation quantity CV, and so on.

[0074] In connection with this, the CPU 92 of the data analysis center 90 receives the data to update the relationship-defining data DR (S90), as described in section (b) in Fig. Figure 7 shows that, based on the received data, CPU 92 then performs the processing of S22. CPU 92 then controls the communication device 97 to transmit data to update the relationship-defining data DR to the transmission source of the data received by processing S90 (S92). It should be noted that after completing the processing of S92, CPU 92 performs the operation described in section (b) in Fig. The 7 processing sequences shown have been completed once.

[0075] In response, CPU 42 receives the data for updating, as described in section (a) in Fig. Figure 7 (S84) shows that CPU 42 then updates the relationship-defining data DR used in the processing of S14 based on the received data (S86). It should be noted that upon completion of the processing of S86, or upon a negative determination in the processing of S10 or S80, CPU 42 updates the data described in section (a) in Figure 7. Fig. The processing sequences shown in section 7 are completed once. It should be noted that if the S80 processing is negative and the subsequent re-execution of the sequences shown in section (a) in Fig. In the processing sequence shown in section 7, CPU 42 does not re-update action a by processing S12 to S16, unless the current time is the time at the start of the phase. That is, in this case, only the processing of S18 is executed again.

[0076] In this way, according to the present embodiment, the relationship-defining data DR is updated outside of vehicle VC1, thereby reducing the computational load of the control device 40. Furthermore, by receiving data from vehicles VC(1), VC(2), etc., in the processing of S90 and processing of S22, the number of data used for learning can be easily increased. Correlative relationship

[0077] The correlative relationship between the elements in the above embodiment and the invention is as follows. A processor in the invention corresponds to CPU 42 and ROM 44, and a memory corresponds to memory 46. The detection processing corresponds to the processing of S12, S32, S36, and S40, and an operation processing corresponds to the processing of S16. The reward calculation processing corresponds to the processing of S34, S38, and S42. Fig. 3 and the processing of S34a, S38a and S42a in Fig. 5. The update processing corresponds to the processing of S48. The counting processing corresponds to the processing of S50, and the limiting processing corresponds to the processing of S64, S70, S76, and S80. The update map corresponds to a map defined by the execution processing instructions of S48 in the DPL tutorial. In other words, the update map corresponds to the map defined by the expression (c1) above. In the invention, the limiting processing corresponds to the processing of S62, S68, S74, and S78. In the invention, a first processor corresponds to CPU 42 and ROM 44, and a second processor corresponds to CPU 92 and ROM 94. In the invention, a computer corresponds to CPU 42 in Fig. 1 and the CPUs 42 and 92 in Fig. 6. Other embodiments

[0078] The present embodiment can be modified as follows. It should be noted that the present embodiment and the following modifications can be combined if no technical contradiction arises.

[0079] The state used to select the value of the action variable based on relationship-defining data

[0080] The states used for selecting action variable values ​​based on the relationship-defining data are not limited to those named or explained in the embodiments described above. For example, a state variable that depends on a value of a previous action variable with respect to phase 2 and phase 3 is not limited to the rotational speed Nm², but can be the rotational speed difference ΔNm². The state variable can also be, for example, the amount of heat generated. Firstly, a state variable that depends on a previous value of an action variable with respect to phase 2 and phase 3 need not be included in the states used to select the value of the action variables when a profit-sharing algorithm or the like is used, as described below in the section "On the Update Map".

[0081] Including the accelerator actuation amount ACCP in the state variable is not mandatory. Including the oil temperature Toil in the state variable is not mandatory. Including the phase variable Vpase in the state variable is not mandatory. For example, the time from the start of the gear shift, the input shaft speed, and the gear shift variable ΔVsft can be included in the state variable. An action value function Q can be constructed that directs actions each time, and reinforcement learning can be performed using this action value function. In this arrangement, the gear shift period is not predetermined to three phases. Regarding the action variable

[0082] Although in the above embodiments the action variable for phase 3 was described as the pressure rise rate, this is not restrictive; for example, phase 3 can be further subdivided, and the pressure command values ​​in each stage can be the action variable.

[0083] Although in the above embodiments the pressure command value or the pressure rise rate is described as the action variable, this is not restrictive, but can, for example, be an instruction value of an electrical current supplied to the solenoid valves 28a or a rate of change of an instruction value. Regarding the relationship-defining data

[0084] Although in the above embodiments the action value function Q is described as a table format function, this is not restrictive. For example, a function approximator can be used.

[0085] Instead of using the action value function Q, for example, the strategy π can be expressed by a function approximator in which the state s and the action a are independent variables, and a probability of performing an action a is a dependent variable, and a parameter that sets the function approximator can be updated in accordance with the reward r. For operational processing

[0086] If the action value function Q is a function approximator, as described in the section “On the relationship-defining data”, an action a that maximizes the action value function Q can be selected by inputting each discrete value with respect to actions that are independent variables of the table-type function in the above embodiments, together with the state s, into the action value function Q.

[0087] If the strategy π is a function approximator in which the state s and the action a are independent variables and a probability of performing an action a is a dependent variable, as described in the section “On the relationship-defining data”, an action a can be selected based on a probability specified by the strategy π. To update maps

[0088] Although the aforementioned Q-Learning, which is a policy-off-TD learning approach, is described using S48 as an example, this is not a limitation. For instance, the learning can be performed using the so-called State-Action-Reward-State-Action algorithm (SARSA algorithm), which is a policy-on-TD learning approach. Furthermore, the learning is not limited to the use of TDs; the Monte Carlo method or eligibility traces can also be used.

[0089] A map that follows, for example, a profit-sharing algorithm can be used as the update map for the reward-based relationship-defining data. In particular, if an example of using a map that follows a profit-sharing algorithm is a modification of the processing, exemplified in Fig. As shown in Figure 2, the following is carried out. That is, a calculation of the reward is performed at the stage of completing the gear shift. The calculated reward is then distributed according to a reinforcement function among rules, each of which determines a state-action pair involved in the gear shift. A known geometric distribution function can be used here as the reinforcement function. In particular, in Stage 3, the gear shift time Tsft correlates strongly with the value of the action variable, so that if the reward is distributed according to the gear shift time Tsft, the use of a geometrically diminishing function for a reinforcement function is effective, although this is not limited to a geometrically diminishing function.For example, if a reward is given based on the amount of heat produced, the distribution of the reward in accordance with the amount of heat produced may be greatest for Phase 1, given the strong correlation between the amount of heat produced and the value of the action variable in Phase 1.

[0090] For example, if the strategy π is expressed with a function approximator, as described in the section “On the relationship-defining data”, and the reward r is updated directly on that basis, an update map can be configured using a strategy gradient procedure.

[0091] The arrangement is not limited to an action value function Q and a strategy π, which are the object for direct update by the reward r. For example, the action value function Q and the strategy π can each be updated as in an actor-critic procedure. Furthermore, the actor-critic procedure is not limited to this, and a value function V can be the object of the update, for example, instead of the action value function Q. For reward calculation processing

[0092] Although the reward r is zero in both phase 1 and phase 2 in the embodiments described above, this is not a limitation. For example, the reward may be larger in phase 1 if the heat generation quantity CV in phase 1 is comparatively small. Furthermore, the reward may be larger in phase 2 if the heat generation quantity CV in phase 2 is comparatively small. Additionally, the reward may be larger in phase 2 if the speed difference ΔNm² in phase 2 is comparatively small.

[0093] The processing function of giving a larger reward when the amount of heat generated is comparatively small is not limited to the processing function of giving a larger reward when the amount of heat generated, InCV, is relatively small. For example, a smaller reward can be given if, during the gear-shift period, the largest amount of heat generated, CV, per unit of time is comparatively small.

[0094] The variable that indicates the amount by which the input shaft speed of the transmission exceeds a reference speed is not limited to the maximum value of the speed difference ΔNm2max, but can, for example, be an average value of the speed difference ΔNm2 during the gear shift period. Furthermore, this can, for example, be a variable that quantifies the amount by which the input shaft speed exceeds a reference speed when a gear shift command is issued.

[0095] Although the above embodiments include processing the awarding of a larger reward when the gear shift time Tsft is comparatively short, processing the awarding of a larger reward when the overage amount is comparatively small, and processing the awarding of a larger reward when the heat generation amount InCV is comparatively small, this is not restrictive. For example, only one of these three processes may be executed, or two may be executed.

[0096] Although the processing above in Fig. As described in section 5, the size of the reward r1 is changed according to the accelerator pedal actuation amount ACCP and the type of gear shift, even though the gear shift time Tsft remains the same. This is not a limiting factor. For example, the reward r1 can remain unchanged according to the accelerator pedal actuation amount ACCP and be changed according to the type of gear shift. Alternatively, for example, the reward r1 can remain unchanged according to the type of gear shift and be changed according to the accelerator pedal actuation amount ACCP.

[0097] Although the processing in Fig. As described in section 5, the size of the reward r2 is changed according to the accelerator pedal actuation amount ACCP and the type of gear shift, although the maximum value of the speed difference ΔNm2max is the same, this is not restrictive. For example, the reward r2 can remain unchanged according to the accelerator pedal actuation amount ACCP and be changed according to the type of gear shift. Alternatively, for example, the reward r2 can remain unchanged according to the type of gear shift and be changed according to the accelerator pedal actuation amount ACCP.

[0098] Although the processing in Fig. As described in section 5, the size of the reward r3 is changed according to the accelerator actuation amount ACCP and the type of gear shift, although the heat generation amount InCV is the same, this is not restrictive. For example, the reward r3 can remain unchanged according to the accelerator actuation amount ACCP and be changed according to the type of gear shift. Alternatively, for example, the reward r3 can remain unchanged according to the type of gear shift and be changed according to the accelerator actuation amount ACCP. To the vehicle control system

[0099] The processing of the decision action based on strategy π (processing of S14) is in the Fig. The example shown in section 7 is described as being vehicle-side, but this is not a limiting factor. For example, an arrangement can be implemented in which data determined by processing S12 is transmitted from the vehicle VC1, the data analysis center 90 decides on an action a using the data transmitted there, and transmits the decided action to the vehicle VC1.

[0100] The vehicle control system is not limited to consisting of the control device 40 and the data analysis center 90. For example, a user's mobile connection can be used instead of the data analysis center 90. Furthermore, a vehicle control system can include the control device 40, the data analysis center 90, and the mobile terminal. This can be achieved, for example, by the mobile terminal processing S14. Regarding the processor

[0101] The processor is not limited to comprising the CPU 42 (92) and the ROM 44 (94) and performing the software processing. For example, a dedicated hardware circuit, such as an application-specific integrated circuit (ASIC) or the like, which performs hardware processing, may be provided to perform at least part of what is performed by the software processing in the embodiments above. That is to say, the processor may have one of the following configurations (a) to (c). (a) A processing device that, following a program, performs all of the above processing, and a program memory, such as a ROM or the like, that stores the program, are provided. (b) A processing device that, following a program, performs part of the above processing, and a program memory and a dedicated hardware circuit that performs the remaining processing are provided.(c) A dedicated hardware circuit that performs all of the above processing is provided. Multiple software processors, each comprising a processing device and program memory, and multiple dedicated hardware circuits may be provided. To the computer

[0102] The computer is not on CPU 42 in Fig. 1 and CPUs 42 and 92 in Fig.6. For example, the computer may be a computer for generating the relationship-defining data DR before the shipment of vehicle VC1, and the CPU 42 may be installed in vehicle VC1. In this case, the search range after shipment is preferably a range in which the value that the action variable can take is smaller compared to the search during reinforcement learning performed by the computer for generating the relationship-defining data DR.It should be noted that an arrangement can be implemented in which no vehicle needs to be present during the processing to generate the relationship-defining data DR prior to the vehicle's shipment. The vehicle's state is artificially generated by running the internal combustion engine 10, etc., on a test bench to simulate driving the vehicle. The artificially generated state of the vehicle is then captured through sensor data acquisition, etc., for use in reinforcement learning. In this case, the artificially generated state of the vehicle is considered the vehicle's state based on sensor data acquisition values. To the storage

[0103] In the above embodiments, the memory that stores the relationship-defining data DR and the memory (ROM 44, 94) that stores the learning program DPL, the learning subprogram DPLa and the learning main program DPLb are described as separate memories, but this is not restrictive. To the vehicle

[0104] The vehicle is not limited to a mixed hybrid vehicle, but can, for example, be a series hybrid or a parallel hybrid vehicle. It should be noted that the vehicle is not limited to one equipped with an internal combustion engine and a motor-generator as onboard rotating machinery. For example, the vehicle can be one that includes an internal combustion engine but no motor-generator, or it can be one that includes a motor-generator but no internal combustion engine.

Claims

[1] Vehicle control device comprising: a processor; and a memory (46) wherein the memory (46) stores relationship-defining data to define a relationship between a state of a vehicle and an action variable, which is a variable relating to operations of a transmission installed in the vehicle, the processor is designed to execute: an investigative processing method to determine the condition of the vehicle based on a sensor reading, an operational processing to control the transmission based on a value of the action variable determined by the state of the vehicle as determined in the investigation processing and the relationship-defining data, a reward calculation process to give a larger reward when the vehicle's characteristics meet a reference, instead of failing to meet the reference, based on the vehicle's state determined in the discovery processing, an update processing to update the relationship-defining data with the state of the vehicle determined in the discovery processing, the value of the action variables used in the operation of the transmission, and the reward corresponding to the operation, as inputs into a pre-set update map, a counting process for counting an update counter by the update processor, and a bounding process to limit downwards a range used by the operation processing, in which a value other than a value that maximizes an expected utility with respect to the reward, of values ​​of the action variable that specify the relationship-defining data, when the update counter is relatively large, and The processor is designed to output updated relationship-defining data, thus increasing the expected benefit when the gearbox is operated according to the relationship-defining data, based on an update map. [2] Vehicle control device according to claim 1, wherein the limiting processing comprises processing to limit downwards an update amount in the update processing when the update counter is relatively large. [3] Vehicle control device according to claim 1 or 2, wherein the reward calculation process comprises: processing to give a larger reward when a heat generation quantity in a gear ratio shift period is relatively small, and processing to change a size of the given reward in accordance with a type of gear shift operation, even if the heat generation quantity is the same. [4] Vehicle control device according to any one of claims 1 to 3, wherein the reward calculation process comprises: processing to give a larger reward when a gear shift time, which is a time required to change the gear ratio, is relatively small, and processing to change the size of the reward given in accordance with a type of gear shift operation, even if the gear shift time is the same. [5] Vehicle control device according to any one of claims 1 to 4, wherein the reward calculation process comprises: processing to give a larger reward when the excess amount of a rotational speed of an input shaft of the transmission in a gear ratio shift period exceeding a reference rotational speed is relatively small, and processing to change the size of the reward given in accordance with a type of shift operation, even if the excess amount is the same. [6] Vehicle control device according to any one of claims 1 to 5, wherein the reward calculation process comprises: processing to give a larger reward when a heat generation quantity in a gear ratio shift period is relatively small, and processing to change a size of the given reward in accordance with a level of accelerator actuation amount, even if the heat generation quantity is the same. [7] Vehicle control device according to any one of claims 1 to 6, wherein the reward calculation process comprises: processing to give a larger reward when a gear shift time, which is a time required to change the gear ratio, is relatively small, and processing to change the size of the reward given in accordance with the amount of accelerator actuation, even if the gear shift time is the same. [8] Vehicle control device according to any one of claims 1 to 7, wherein the reward calculation process comprises: processing to give a larger reward when an excess amount of the rotational speed of an input shaft of the transmission in a gear ratio shift period that exceeds a reference rotational speed is relatively small, and processing to change a size of the reward given in accordance with a level of accelerator actuation, even if the excess amount is the same. [9] Vehicle control system, which includes: the processor and the memory (46) of the vehicle control device according to one of claims 1 to 8, wherein The processor comprises a first processor that is installed in the vehicle and a second processor that is separate from an on-board unit and the first processor is designed to perform at least the discovery processing and the operation processing, and the second processor is designed to perform at least the update processing. [10] Vehicle control device comprising the first processor of the vehicle control system according to claim 9. [11] Vehicle learning device comprising the second processor of the vehicle control system according to claim 9. [12] Vehicle learning method in which a computer is caused to perform the detection processing, the operation processing, the reward calculation process, the update processing, the counting processing and the limiting processing of the vehicle control device according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for controlling the reversal of an automatic gearbox

    DE68926540T2

  • Integrated characteristic optimizing device

    JP2000250602A

  • JP002000250602A