A driving strategy determination method, device, system and storage medium
By optimizing the driving strategy of hybrid vehicles through intelligent agents and taking into account battery losses, a more reasonable power distribution between the engine and electric motor is achieved, improving energy utilization and solving the problem of unreasonable power distribution in existing technologies.
Patent Information
- Application Number
- CN202311436311.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-10-31
AI Technical Summary
Existing hybrid vehicles do not properly allocate power between the engine and electric motor when determining driving strategies, and do not fully consider the overall vehicle losses, especially battery losses.
An intelligent agent (composed of a Q-neural network and a target neural network) is used to calculate the overall reward value corresponding to the historical speed curve. The driving strategy is updated cyclically to minimize the overall loss, including battery loss, and the torque distribution is optimized using reinforcement learning.
This achieves a more rational power distribution between the engine and electric motor, improving vehicle energy utilization and reducing overall consumption.
Smart Images

Figure CN119911256B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of driving strategy determination, and particularly relates to a driving strategy determination method, device, system and storage medium. BACKGROUND
[0002] The product of the rotation speed and the torque of an engine is engine power, and the product of the rotation speed and the torque of an electric motor is electric motor power. In existing hybrid vehicles, when a driving strategy is determined, the power is allocated to the engine and the electric motor based on the principle of minimizing the total loss. When the total loss is calculated, only the engine fuel consumption and the electric motor power consumption are considered, and other factors are not considered. This loss calculation method is not comprehensive, and the engine and electric motor power allocation strategy is not reasonable. Therefore, how to provide a driving strategy determination method to more reasonably allocate power to the engine and the electric motor. SUMMARY
[0003] The present application provides a driving strategy determination method, device, system and storage medium to more reasonably allocate power to the engine and the electric motor.
[0004] The present application provides a driving strategy determination method, comprising:
[0005] obtaining a historical speed curve of a vehicle on a target route;
[0006] introducing the historical speed curve into an agent, and calculating a total reward value corresponding to the historical speed curve by the agent, wherein the agent is used to generate a corresponding driving strategy according to the introduced historical speed curve, the agent is composed of a Q neural network and a target neural network, the Q neural network is used to calculate a current Q value according to the driving strategy in the agent and a vehicle state, and the target neural network is used to calculate a maximum Q value; the total reward value is used to represent a total loss of the vehicle in a driving process, the total loss includes a vehicle battery loss, and the total reward value is negatively correlated with the total loss of the vehicle in the driving process;
[0007] when the total reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, cyclically updating the driving strategy in the agent, and determining a current Q value calculated based on the updated driving strategy and the vehicle state, wherein the current Q value is positively correlated with the total reward value;
[0008] when the difference between the Q value calculated based on the updated driving strategy and the vehicle state and the maximum Q value is less than a preset difference value, determining that the agent training is completed, so as to determine the driving strategy of the vehicle.
[0009] The application has the beneficial effects that the battery loss of the vehicle is included in the calculation of the overall loss of the vehicle, so that the calculation result is more accurate and reasonable when determining the overall loss of the vehicle, and when the difference between the Q value calculated by the updated driving strategy and the maximum Q value is less than a preset difference value, it is determined that the agent training is completed, so as to determine the driving strategy of the vehicle, and the vehicle uses the driving strategy included in the agent to drive, which can make the overall consumption of the vehicle in the driving process close to the minimum consumption, thereby realizing a more reasonable torque distribution strategy, that is, the power can be more reasonably allocated to the engine and the motor, and the utilization rate of the energy of the vehicle is improved.
[0010] In one embodiment, the overall reward value corresponding to the historical speed curve is calculated by the agent, including:
[0011] The preset parameter values of each cycle of the vehicle driving according to the historical speed curve are obtained by the agent, and the preset parameter values include the battery loss of the vehicle;
[0012] The preset parameter values of each cycle of the vehicle are introduced into a pre-designed calculation model to obtain the reward value of each cycle;
[0013] The reward values of each cycle are accumulated to obtain the overall reward value corresponding to the historical speed curve.
[0014] In one embodiment, the battery loss of the vehicle is obtained by the following method:
[0015] The cycle life and actual life of the battery under nominal conditions are obtained;
[0016] The cycle life and actual life of the battery under the nominal conditions are substituted into the following first preset formula for calculating the severity factor to obtain the severity factor corresponding to the battery life:
[0017]
[0018] Wherein, σ(I, T b , SOC) is the severity factor; Γ nom is the total ampere-hour throughput of the nominal battery life; Γ represents the actual ampere-hour throughput of the battery; I nom (t) and I(t) are the current values under the nominal conditions and the actual conditions, respectively;
[0019] The preset battery parameters including the severity factor are substituted into the following second preset formula for calculating the battery life loss to obtain the battery loss corresponding to the historical speed curve:
[0020]
[0021] Wherein, C bis the battery loss; c bat is the loss coefficient of replacing the battery; σ is a severity factor; I bat (t) is the corresponding power of the battery, Γ represents the actual ampere-hour throughput of the battery, is the battery life consumption ratio of this vehicle driving.
[0022] In an embodiment, the preset parameter value of the vehicle is introduced into the pre-design calculation model to obtain a result for representing the overall loss of the vehicle, including:
[0023] The power loss, fuel loss, battery loss, vehicle start-stop penalty value and physical constraint penalty value of the vehicle are introduced into the following pre-design calculation model to calculate the reward value representing the overall loss of the vehicle, and the reward value representing the overall loss of the vehicle is taken as the calculation result of the pre-design calculation model, wherein the vehicle start-stop penalty value is determined according to the number of vehicle start-stop times, the more the start-stop times, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery power is in a preset power interval and whether the current torque of the vehicle is in a preset torque interval:
[0024] r(t) = -(C f +C e +C b +C Penalty,1 +C Penalty,2 );
[0025] Wherein, r(t) is the reward value representing the overall loss of the vehicle; C f is the fuel loss; C e is the power loss; C b is the battery loss; C Penalty,1 is the vehicle start-stop penalty value; C Penalty,2 is the physical constraint penalty value.
[0026] In an embodiment, the determination method of the fuel loss is as follows:
[0027] Obtain the engine fuel consumption rate, unit time fuel loss, multiple sets of discrete values corresponding to the engine speed and torque, and fuel density;
[0028] Substitute the engine fuel consumption rate, unit time fuel loss, engine speed, engine torque and fuel density into the following third preset formula to determine the fuel loss:
[0029]
[0030] Wherein, C f is the fuel loss; c f is the unit time fuel loss; b e is the engine fuel consumption rate; Tor eng(t) is the engine torque, ω eng (t) is the engine speed; p is the fuel density.
[0031] In one embodiment, the physical constraint penalty value is determined as follows:
[0032] determine whether the current battery power is in a preset power interval, and determine whether the current torque of the vehicle is in a preset torque interval;
[0033] When the current battery power is in the preset power interval and the current torque of the vehicle is in the preset torque interval, the physical constraint penalty value is determined to be 0;
[0034] When the current battery power exceeds the preset power interval or the current torque of the vehicle exceeds the preset torque interval, the physical constraint penalty value is determined to be a specific value, wherein the specific value makes the reward value of the total loss of the vehicle less than a preset value.
[0035] In one embodiment, the current driving strategy is updated in a cycle, comprising:
[0036] randomly selecting a plurality of four-tuple data from the experience pool corresponding to the agent as learning samples, wherein the four-tuple data contains the state at a historical time, the action taken by the agent at the historical time, the reward at the historical time, and the state at the next historical time;
[0037] training the agent through the learning samples, and updating the weight parameters through the stochastic gradient descent method;
[0038] assigning the updated weight parameters to the Q network to realize the cycle updating of the driving strategy in the agent.
[0039] The application also provides a driving strategy determination device, comprising:
[0040] an acquisition module for acquiring a historical speed curve of a vehicle on a target route;
[0041] an import module for importing the historical speed curve into an agent, and calculating a total reward value corresponding to the historical speed curve through the agent, wherein the agent is used to generate a corresponding driving strategy according to the imported historical speed curve, the agent is composed of a Q neural network and a target neural network, the Q neural network is used to calculate a current Q value according to the driving strategy in the agent and the state of the vehicle, and the target neural network is used to calculate a maximum Q value; the total reward value is used to represent the total loss of the vehicle in the driving process, the total loss includes the battery loss of the vehicle, and the total reward value is negatively correlated with the total loss of the vehicle in the driving process;
[0042] The first determining module is configured to cyclically update the driving strategy in the agent when the overall reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, and determine a current Q value calculated based on the updated driving strategy and the vehicle state, the current Q value being positively correlated with the overall reward value.
[0043] The second determining module is configured to determine that the agent training is completed when a difference between the Q value calculated based on the updated driving strategy and the vehicle state and the maximum Q value is less than a preset difference value, so as to determine the driving strategy of the vehicle.
[0044] In one embodiment, the importing module comprises:
[0045] The obtaining sub-module is configured to obtain, by the agent, preset parameter values of each period when the vehicle drives according to the historical speed curve, the preset parameter values including vehicle battery loss.
[0046] The importing sub-module is configured to import the preset parameter values of each period of the vehicle into a pre-designed calculation model to obtain reward values of each period.
[0047] The accumulating sub-module is configured to accumulate the reward values of each period to obtain the overall reward value corresponding to the historical speed curve.
[0048] In one embodiment, the vehicle battery loss is obtained by the following method:
[0049] The cycle life and the actual life of the battery under the nominal condition are obtained.
[0050] The cycle life and the actual life of the battery under the nominal condition are substituted into the following first preset formula for calculating the severity factor to obtain the severity factor corresponding to the battery life:
[0051]
[0052] wherein, σ(I, T b , SOC) is the severity factor; Γ nom is the total ampere-hour throughput of the nominal battery life; Γ represents the actual ampere-hour throughput of the battery; I nom (t) and I(t) are the current values under the nominal condition and the actual condition, respectively.
[0053] The preset battery parameters including the severity factor are substituted into the following second preset formula for calculating the battery life loss to obtain the battery loss corresponding to the historical speed curve:
[0054]
[0055] wherein, C b is the battery loss; cbat is the loss coefficient of replacing the battery; σ is a serious factor; I bat (t) is the corresponding electricity of the battery, and Г represents the actual ampere-hour passing quantity of the battery, is the battery life consumption proportion of the current vehicle driving.
[0056] In an embodiment, the import submodule comprises:
[0057] The electricity loss, fuel loss, battery loss, vehicle start-stop penalty value and physical constraint penalty value of the vehicle are imported into the following pre-designed calculation model to calculate the reward value representing the overall loss of the vehicle, and the reward value representing the overall loss of the vehicle is taken as the calculation result of the pre-designed calculation model, wherein the vehicle start-stop penalty value is determined according to the number of vehicle start-stops, the more the start-stop times, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery electricity is in a preset electricity interval and whether the current torque of the vehicle is in a preset torque interval:
[0058] r(t) = -(C f + C e + C b + C Penalty,1 + C Penalty,2 );
[0059] Wherein, r(t) is the reward value representing the overall loss of the vehicle; C f is the fuel loss; C e is the electricity loss; C b is the battery loss; C Penalty,1 is the vehicle start-stop penalty value; C Penalty,2 is the physical constraint penalty value.
[0060] In an embodiment, the fuel loss is determined in the following manner:
[0061] Obtain the engine fuel consumption rate, unit time fuel loss, multiple sets of discrete values corresponding to engine speed and torque, and fuel density;
[0062] Substitute the engine fuel consumption rate, unit time fuel loss, engine speed, engine torque and fuel density into the following third preset formula to determine the fuel loss:
[0063]
[0064] Wherein, C f is the fuel loss; c f is the unit time fuel loss; b e is the engine fuel consumption rate; Tor eng (t) is the engine torque, ω eng(t) is the engine speed; p is the fuel density.
[0065] In one embodiment, the physical constraint penalty value is determined as follows:
[0066] determine whether the current battery power is in a preset power interval, and determine whether the current torque of the vehicle is in a preset torque interval;
[0067] When the current battery power is in the preset power interval and the current torque of the vehicle is in the preset torque interval, the physical constraint penalty value is determined to be 0;
[0068] When the current battery power is out of the preset power interval or the current torque of the vehicle is out of the preset torque interval, the physical constraint penalty value is determined to be a specific value, wherein the specific value makes the reward value of the total loss of the vehicle less than a preset value.
[0069] In one embodiment, the current driving strategy is updated in a cycle, comprising:
[0070] randomly selecting a plurality of four-tuple data from an experience pool corresponding to the agent as learning samples, wherein the four-tuple data contains the state at a historical time, the action taken by the agent at the historical time, the reward at the historical time, and the state at the next historical time;
[0071] training the agent in a cycle through the learning samples, and updating the weight parameters through a stochastic gradient descent method;
[0072] assigning the updated weight parameters to the Q network to realize the cycle updating of the driving strategy in the agent.
[0073] The application also provides a driving strategy determination system, comprising:
[0074] at least one processor; and,
[0075] a memory in communication connection with the at least one processor; wherein,
[0076] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to realize the driving strategy determination method as described in any of the above embodiments.
[0077] The application also provides a computer-readable storage medium, when the instructions in the storage medium are executed by the processor corresponding to the driving strategy determination system, the driving strategy determination system can realize the driving strategy determination method as described in any of the above embodiments.
[0078] Other features and advantages of the present application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The purposes and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.
[0079] The technical solutions of the present application are described in further detail below with the aid of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS
[0080] The accompanying drawings are included to provide a further understanding of the present application and are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and serve to explain the present application, and are not intended to limit the present application. In the drawings:
[0081] Figure 1 A flow chart of a driving strategy determination method according to an embodiment of the present application;
[0082] Figure 2 A structural schematic diagram of a driving strategy determination device according to an embodiment of the present application;
[0083] Figure 3 A hardware structural schematic diagram of a driving strategy determination system according to an embodiment of the present application. DETAILED DESCRIPTION
[0084] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to explain and illustrate the present application, and are not intended to limit the present application.
[0085] Figure 1 A flow chart of a driving strategy determination method according to an embodiment of the present application, as shown in Figure 1 the method can be implemented as the following steps S101-S104:
[0086] In step S101, a historical speed curve of a vehicle on a target route is obtained;
[0087] In step S102, the historical speed curve is imported into an agent, and a total reward value corresponding to the historical speed curve is calculated by the agent, wherein the agent is used to generate a corresponding driving strategy according to the imported historical speed curve, the agent is composed of a Q neural network and a target neural network, the Q neural network is used to calculate a current Q value according to the driving strategy in the agent and a vehicle state, and the target neural network is used to calculate a maximum Q value; the total reward value is used to represent a total loss of the vehicle in a driving process, the total loss includes a battery loss of the vehicle, and the total reward value is negatively correlated with the total loss of the vehicle in the driving process;
[0088] In step S103, when the overall reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, the driving strategy in the agent is updated in a loop, and it is determined that the current Q value calculated based on the updated driving strategy and the calculated current Q value is positively correlated with the overall reward value;
[0089] In step S104, when the difference between the Q value calculated by the updated driving strategy and the maximum Q value is less than a preset difference value, it is determined that the agent training is completed, so as to determine the driving strategy of the vehicle.
[0090] In the present application, the historical speed curve of the vehicle on the target route is obtained; the historical speed curve is introduced into an agent, wherein the agent is composed of a Q neural network and a target neural network, the Q neural network is used to calculate the current Q value according to the driving strategy in the agent, and the target neural network is used to calculate the maximum Q value; the maximum Q value is usually the maximum Q value of the next state, i.e. the expected Q value. The overall reward value corresponding to the historical speed curve is calculated by the agent, wherein the overall reward value is used to represent the overall loss of the vehicle in the driving process, the overall loss includes the battery loss of the vehicle, and the overall reward value is negatively correlated with the overall loss of the vehicle in the driving process.
[0091] When the overall reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, the current driving strategy is updated in a loop, and the current Q value calculated by the updated driving strategy is determined, i.e. the agent can select the action with the highest Q value according to the current state, and update the Q value table through the observation reward value. The current Q value represents the expected return obtained by taking the corresponding torque in a specific state. It can be represented as Q(s, a), wherein s represents the state of the vehicle, and a represents the action. The action value function tells the agent the value size of taking different actions in a certain state, thereby helping the agent to make the optimal decision. Specifically, a plurality of four-tuple data are randomly selected from the experience pool corresponding to the agent as learning samples, wherein the four-tuple data include the state at the historical time, the action taken by the agent at the historical time, the reward at the historical time, and the state at the next historical time; the agent is trained in a loop through the learning samples, and the weight parameters are updated through the stochastic gradient descent method; the updated weight parameters are assigned to the Q network, so as to realize the loop updating of the driving strategy in the agent. For example, the data stored in the experience pool is (S t ,A t ,r(t),S t+1 ) four-tuple, wherein S t represents the state at time t, A t represents the action taken by the agent at time t, r(t) represents the reward at time t, and S t+1 represents the state at time t+1, i.e. the action A tThe state of the next moment after the action. Among them, the vehicle state includes the battery SOC of the vehicle, the vehicle speed, the vehicle acceleration, the battery temperature and the remaining mileage ratio of the vehicle.
[0092] When the difference between the Q value calculated by the updated driving strategy and the maximum Q value is less than the preset difference, it is determined that the agent training is completed, so as to determine the driving strategy of the vehicle.
[0093] It should be noted that in this application, after the agent training is completed, the trained agent can be sent to the vehicle that needs to drive on the target route, so that the vehicle driving on the target route can drive according to the driving strategy in the agent, wherein the driving strategy at least includes the torque distribution strategy of the vehicle in the driving process.
[0094] In reinforcement learning, the goal is to find the optimal value function (optimal strategy) so that the agent can obtain the maximum cumulative reward in the process of interacting with the environment. By continuously updating the action value function, the optimal action value function is finally approached, so as to find the optimal value function. It can be understood that when the optimal function is found, the driving strategy in the agent is also the optimal driving strategy that can minimize the loss.
[0095] When calculating the loss, the battery loss is considered in this application, so when formulating the driving strategy, it is based on the goal of minimizing the total loss of fuel consumption, power consumption and battery loss.
[0096] The beneficial effects of the present application are that the battery loss of the vehicle is included when calculating the overall loss of the vehicle, so that the calculation result is more accurate and reasonable when determining the overall loss of the vehicle. When the difference between the Q value calculated by the updated driving strategy and the maximum Q value is less than the preset difference, it is determined that the agent training is completed, and the vehicle adopts the driving strategy contained in the agent to drive, which can make the overall consumption of the vehicle in the driving process close to the minimum consumption, thereby realizing a more reasonable torque distribution strategy, that is, the power can be more reasonably distributed to the engine and the motor, and the utilization rate of the vehicle energy is improved.
[0097] In one embodiment, the step S103 of calculating the overall reward value corresponding to the historical speed curve by the agent includes the following steps A1-A3:
[0098] In step A1, the agent obtains the preset parameter value of each period when the vehicle drives according to the historical speed curve, and the preset parameter value includes the battery loss of the vehicle.
[0099] In step A2, the preset parameter value of each period of the vehicle is introduced into the pre-designed calculation model to obtain the reward value of each period.
[0100] In step A3, the reward values of each cycle are accumulated to obtain a total reward value corresponding to the historical speed curve.
[0101] The vehicle battery loss is obtained by:
[0102] The cycle life and actual life of the battery under nominal conditions are obtained.
[0103] The cycle life of the battery under nominal conditions and the actual battery life are substituted into the following first preset formula for calculating the severity factor to obtain a severity factor corresponding to the battery life:
[0104]
[0105] Wherein, σ(I, T b , SOC) is the severity factor; Γ nom is the total ampere-hour throughput of the nominal battery life; Γ represents the actual ampere-hour throughput of the battery; I nom (t) and I(t) are the nominal condition and actual current values, respectively.
[0106] The preset battery parameters including the severity factor are substituted into the following second preset formula for calculating the battery life loss to obtain the battery loss corresponding to the historical speed curve:
[0107]
[0108] Wherein, C b is the battery loss; c bat is the loss coefficient of replacing the battery; σ is the severity factor; I bat (t) is the battery corresponding electric quantity, Γ represents the actual ampere-hour throughput of the battery, is the battery life consumption ratio of the current vehicle driving.
[0109] In one embodiment, the above step A2 can be implemented as the following steps:
[0110] The electric quantity loss, fuel loss, battery loss, vehicle start-stop penalty value and physical constraint penalty value of the vehicle are introduced into the following preset calculation model to calculate the reward value representing the overall loss of the vehicle, and the reward value representing the overall loss of the vehicle is taken as the calculation result of the preset calculation model, wherein the vehicle start-stop penalty value is determined according to the number of vehicle start-stops, the more the number of start-stops, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery electric quantity is in a preset electric quantity interval and whether the current torque of the vehicle is in a preset torque interval:
[0111] r(t) = -(C f +C e +C b+C Penalty,1 +C Penalty,2 );
[0112] wherein, r(t) is a reward value representing the overall loss of the vehicle; C f is the fuel loss; C e is the power loss; C b is the battery loss; C Penalty,1 is the vehicle start-stop penalty value; C Penalty,2 is the physical constraint penalty value.
[0113] The determination manner of the fuel loss is as follows:
[0114] obtain the engine fuel consumption rate, the unit time fuel loss, the engine speed and the torque corresponding to a plurality of groups of discrete values, and the fuel density;
[0115] substitute the engine fuel consumption rate, the unit time fuel loss, the engine speed, the engine torque and the fuel density into the following third preset formula to determine the fuel loss:
[0116]
[0117] wherein, C f is the fuel loss; c f is the unit time fuel loss; b e is the engine fuel consumption rate; Tor eng (t) is the engine torque, ω eng (t) is the engine speed; and ρ is the fuel density.
[0118] In the present application, preset parameter values of the vehicle are obtained, wherein the preset parameter values include parameters representing the battery loss. In order to improve the energy utilization rate of the vehicle, when determining the overall loss of the vehicle, not only the consumption of power and fuel is considered, but also the hidden cost of battery aging is considered.
[0119] (1) Determine the battery loss: Since the actual working condition of the battery is complex, in order to quantify the influence of complex working conditions on the battery life, the concept of severity factor is introduced, which is defined as follows:
[0120]
[0121] wherein, σ is the severity factor; Γ nom is the total ampere-hour throughput of the nominal battery life; and Г represents the actual ampere-hour throughput of the battery.
[0122] Then, substitute the severity factor, the battery life consumption ratio of this vehicle driving, and the loss coefficient of replacing the battery into the second preset formula for calculating the battery life loss to obtain the battery loss of this driving. Specifically, the second preset formula can be the following formula:
[0123]
[0124] wherein C b is the battery loss; c bat is the loss coefficient of replacing the battery; σ is the severity factor; I bat (t) is the corresponding electricity of the battery, and represents the actual ampere-hour throughput of the battery, is the battery life consumption ratio of this time vehicle driving.
[0125] In an embodiment of the present application, the total ampere-hour throughput of the battery is taken as the life of the battery. In order to determine the ampere-hour throughput of the battery, a battery aging model is constructed. Since the battery aging is related to temperature, current size, SOC (state of charge, which refers to the charging state of the automobile battery, also called the remaining capacity percentage, indicating the ability of the battery to continue to work), etc., the battery aging model can be constructed as follows:
[0126] Q loss = f(T, I, SOC),
[0127] wherein Q loss is the battery capacity loss value, T is the battery temperature, I is the current size, SOC is the remaining capacity of the battery, and f is the prediction function of the battery aging model, which can have various forms and is not limited in the present application. In the embodiment, the specific form of the battery aging model is as follows:
[0128]
[0129] wherein Q loss is the battery capacity loss value; SOC is the remaining capacity of the battery; E a is the battery activation energy; I c is the discharge rate; R gas is the gas constant; T b is the battery temperature; Ah is the ampere-hour throughput of the battery; α and β are fitting constants, and η is a compensation coefficient, which is a power law factor and can be determined in advance by experiment.
[0130] The battery activation energy refers to the energy barrier that needs to be overcome to start the reaction in a chemical reaction, which can be determined by the influence of temperature on the reaction rate of the battery. The commonly used gas constant in the battery is the ideal gas constant. Since it is generally believed that when the battery capacity Q loss loss reaches 20%, it means that the battery life reaches the end. Further, the nominal battery cycle life can be obtained as follows:
[0131]
[0132] where Γ nom is the total ampere-hour throughput of the nominal battery life, SOC nom is the remaining capacity of the nominal battery; E a is the battery activation energy; I c,nom is the discharge rate; R gas is the gas constant; is the battery temperature; a and b are fitting constants, and h is the compensation coefficient, is the power law factor.
[0133] According to the actual use, the actual battery life can be obtained as follows:
[0134]
[0135] where Γ represents the ampere-hour throughput of the actual life; SOC is the remaining capacity of the battery; E a is the battery activation energy; I c is the discharge rate; R gas is the gas constant; T b is the battery temperature; Ah is the ampere-hour throughput of the battery; a and b are fitting constants, and h is the compensation coefficient, is the power law factor, which can be determined in advance by experiment.
[0136] (2) Determine the fuel consumption: obtain the engine fuel consumption rate, unit time fuel consumption, a plurality of discrete values corresponding to the engine speed and torque, and fuel density; the engine fuel consumption rate, unit time fuel consumption, engine speed, engine torque and fuel density are substituted into the third preset formula to determine the fuel consumption.
[0137]
[0138] where C f is the fuel consumption; c f is the unit time fuel consumption; b e is the engine fuel consumption rate; Tor eng (t) is the engine torque, w eng (t) is the engine speed; p is the fuel density.
[0139] (3) Determine the power loss: obtain the unit time power loss and the battery power, and substitute the following formula to determine the corresponding power loss.
[0140]
[0141] where CC e is the power loss; c e is the unit time power loss; P b (t) is the power of the battery.
[0142] Then, after determining the various consumptions of the vehicle, the preset parameters of the vehicle are introduced into a pre-design calculation model, wherein a calculation result of the pre-design calculation model represents the overall consumption of the vehicle, and the greater the calculation result is, the smaller the overall consumption of the vehicle is. Specifically, the electric quantity consumption, the fuel consumption, the battery consumption, the vehicle start-stop penalty value and the physical constraint penalty value of the vehicle are introduced into the following pre-design calculation model to calculate a reward value representing the overall consumption of the vehicle as the calculation result of the pre-design calculation model:
[0143] r(t) = -(C f +C e +C b +C Penalty,1 +C Penalty,2 );
[0144] wherein r(t) is the reward value representing the overall consumption of the vehicle; C f is the fuel consumption; C e is the electric quantity consumption; C b is the battery consumption; C Penalty,1 is the vehicle start-stop penalty value, which avoids frequent start-stop of the engine and can be calibrated; C Penalty,2 is the physical constraint penalty value, which is determined according to whether the current battery electric quantity is in a preset electric quantity interval and whether the current torque of the vehicle is in a preset torque interval. When the current battery electric quantity is in the preset electric quantity interval and the current torque of the vehicle is in the preset torque interval, the physical constraint penalty value is determined to be 0. When the current battery electric quantity is out of the preset electric quantity interval or the current torque of the vehicle is out of the preset torque interval, the physical constraint penalty value is determined to be a specific value, wherein the specific value makes the reward value of the overall consumption of the vehicle less than a preset value.
[0145] The running parameter values of the engine and the running parameter values of the motor are obtained when the calculation result of the pre-design calculation model is maximum, wherein the running parameter values at least include torque and speed. In this embodiment, in order to reasonably allocate the running power of the engine and the motor, the target parameter values of the vehicle driving device corresponding to the throttle opening degree are obtained, wherein the target parameter values are target speed and target torque. The vehicle driving device includes the engine and the motor, and the correspondence between the throttle opening degree and the running parameters of the engine and the motor can be obtained through experiments in advance. Therefore, during the running of the vehicle, the target speed and the target torque corresponding to the engine and the motor can be obtained according to the current throttle opening degree, respectively.
[0146] Finally, the engine and the motor of the vehicle are controlled according to the running parameter values of the engine and the running parameter values of the motor when the overall consumption of the vehicle is minimum.
[0147] Through the present application, for the daily repeated higher route, for example, the daily commuting, the daily route is fixed, and the daily operation working condition (speed curve) coincidence degree is high. The user confirms the working condition data collection in the commuting mode through the interactive interface, and the device collects the vehicle driving speed, driving distance and the like; when the journey is over, the user clicks to confirm this section of route as the commuting route, and sets the oil price and electricity price, and then uploads this section of data to the cloud. The cloud trains the model according to the data uploaded by the user and the set oil price and electricity price; after the training is completed, the trained intelligent agent is updated to the vehicle controller on the user's vehicle. Through the method, the energy in the commuting route is distributed, the approximate global optimum in the commuting route can be realized, the driving cost is optimized, and the energy utilization rate is improved. The energy-saving advantage of the hybrid electric vehicle in the fixed driving route is fully played; in addition, since the battery aging cost is considered, the power distribution of the battery and the engine is reasonable, and the aging speed of the battery is slowed down.
[0148] Figure 2 For the structure of a driving strategy determination device in an embodiment of the present application, as shown in Figure 2 , the device comprises:
[0149] The acquisition module 201 is configured to acquire a historical speed curve of a vehicle on a target route.
[0150] The import module 202 is configured to import the historical speed curve into an intelligent agent, and calculate a total reward value corresponding to the historical speed curve through the intelligent agent, wherein the intelligent agent is configured to generate a corresponding driving strategy according to the imported historical speed curve, the intelligent agent is composed of a Q neural network and a target neural network, the Q neural network is configured to calculate a current Q value according to the driving strategy in the intelligent agent and a vehicle state, and the target neural network is configured to calculate a maximum Q value; the total reward value is configured to represent a total loss of the vehicle in a driving process, the total loss includes a battery loss of the vehicle, and the total reward value is negatively correlated with the total loss of the vehicle in the driving process.
[0151] The first determination module 203 is configured to cyclically update the driving strategy in the intelligent agent when the total reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, and determine a current Q value calculated based on the updated driving strategy and the vehicle state, wherein the current Q value is positively correlated with the total reward value.
[0152] The second determination module 204 is configured to determine that the training of the intelligent agent is completed when the difference between the Q value calculated based on the updated driving strategy and the vehicle state and the maximum Q value is less than a preset difference value, so as to determine the driving strategy of the vehicle.
[0153] In an embodiment, the import module comprises:
[0154] An acquisition sub-module is configured to acquire preset parameter values of each period when the vehicle is driven according to the historical speed curve by the intelligent agent, wherein the preset parameter values include vehicle battery loss;
[0155] An import sub-module is configured to import the preset parameter values of each period of the vehicle into a pre-designed calculation model to obtain reward values of each period;
[0156] An accumulation sub-module is configured to accumulate the reward values of each period to obtain a total reward value corresponding to the historical speed curve.
[0157] In one embodiment, the vehicle battery loss is acquired by the following method:
[0158] The cycle life and actual life of the battery under nominal conditions are acquired;
[0159] The cycle life of the battery under the nominal conditions and the actual life of the battery are substituted into the following first preset formula for calculating a severity factor to obtain a severity factor corresponding to the battery life:
[0160]
[0161] Wherein, σ(I, T b , SOC) is the severity factor; Γ nom is the total ampere-hour throughput of the nominal battery life; Γ represents the actual ampere-hour throughput of the battery; I nom (t) and I(t) are the current values under the nominal conditions and the actual conditions, respectively;
[0162] The preset battery parameters including the severity factor are substituted into the following second preset formula for calculating the battery life loss to obtain the battery loss corresponding to the historical speed curve:
[0163]
[0164] Wherein, C b is the battery loss; c bat is the loss coefficient of replacing the battery; σ is the severity factor; I bat (t) is the battery corresponding electric quantity, Γ represents the actual ampere-hour throughput of the battery, is the battery life consumption ratio of this vehicle driving.
[0165] In one embodiment, the import sub-module includes:
[0166] The power loss, fuel loss, battery loss, vehicle start-stop penalty value and physical constraint penalty value of the vehicle are introduced into the following pre-designed calculation model to calculate a reward value representing the overall loss of the vehicle, and the reward value representing the overall loss of the vehicle is taken as the calculation result of the pre-designed calculation model, wherein the vehicle start-stop penalty value is determined according to the number of vehicle start-stop times, and the more the start-stop times, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery power is in a preset power interval and whether the current torque of the vehicle is in a preset torque interval:
[0167] r(t) = -(C f +C e +C b +C Penalty,1 +C Penalty,2 );
[0168] Wherein, r(t) is a reward value representing the overall loss of the vehicle; C f is the fuel loss; C e is the power loss; C b is the battery loss; C Penalty,1 is the vehicle start-stop penalty value; and C Penalty,2 is the physical constraint penalty value.
[0169] In one embodiment, the fuel loss is determined as follows:
[0170] Obtain the engine fuel consumption rate, unit time fuel loss, multiple sets of discrete values corresponding to the engine speed and torque, and fuel density;
[0171] Substitute the engine fuel consumption rate, unit time fuel loss, engine speed, engine torque and fuel density into the following third preset formula to determine the fuel loss:
[0172]
[0173] Wherein, C f is the fuel loss; c f is the unit time fuel loss; b e is the engine fuel consumption rate; Tor eng (t) is the engine torque, ω eng (t) is the engine speed; and p is the fuel density.
[0174] In one embodiment, the physical constraint penalty value is determined as follows:
[0175] Determine whether the current battery power is in a preset power interval, and determine whether the current torque of the vehicle is in a preset torque interval;
[0176] When the current battery power is in the preset power interval and the current torque of the vehicle is in the preset torque interval, the physical constraint penalty value is determined as 0;
[0177] When the current battery power is out of the preset power interval or the current torque of the vehicle is out of the preset torque interval, the physical constraint penalty value is determined as a specific value, wherein the specific value makes the reward value of the total loss of the vehicle less than a preset value.
[0178] In an embodiment, the current driving strategy is updated in cycles, including:
[0179] Randomly selecting a plurality of quadruple data from an experience pool corresponding to the agent as learning samples, wherein the quadruple data contains a state at a historical time, an action taken by the agent at the historical time, a reward at the historical time, and a state at a next historical time;
[0180] The agent is trained in cycles through the learning samples, and the weight parameters are updated through a stochastic gradient descent method;
[0181] The updated weight parameters are assigned to the Q network to realize the cyclic update of the driving strategy in the agent.
[0182] Figure 3 A hardware structure diagram of a driving strategy determination system in an embodiment of the present application is shown in FIG. 3, which includes: Figure 3
[0183] at least one processor 320; and
[0184] a memory 304 in communication with the at least one processor 320; wherein
[0185] The memory 304 stores instructions executable by the at least one processor 320, and the instructions are executed by the at least one processor 320 to implement the driving strategy determination method described in any of the above embodiments.
[0186] Referring to FIG. 3, Figure 3 The driving strategy determination system 300 can include one or more of the following components: a processing component 302, a memory 304, a power supply component 306, a multimedia component 308, an audio component 310, an input / output (I / O) interface 312, a sensor component 314, and a communication component 316.
[0187] The processing component 302 generally controls the overall operation of the driving policy determination system 300. The processing component 302 can include one or more processors 320 to execute instructions. The instructions can be fetched from the memory 304 or from another computer readable medium. The processing component 302 can further include one or more modules to facilitate the interaction between the processing component 302 and other components. For example, the processing component 302 can include a multimedia module to facilitate the interaction between the multimedia component 308 and the processing component 302.
[0188] The memory 304 is configured to store various types of data to support the operations of the driving policy determination system 300. Examples of these data include instructions for any application or method operating on the driving policy determination system 300, such as text, pictures, videos, etc. The memory 304 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, magnetic or optical disk.
[0189] The power component 306 provides power to the various components of the driving policy determination system 300. The power component 306 can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the driving policy determination system 300.
[0190] The multimedia component 308 includes a screen providing an output interface between the driving policy determination system 300 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, slide and gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 308 can further include a front camera and / or a rear camera. The front camera and / or the rear camera can receive external multimedia data when the driving policy determination system 300 is in an operation mode, such as a photographing mode or a video image mode. Each of the front and rear cameras can be a fixed optical lens system or have a focal length and optical zoom capability.
[0191] The audio component 310 is configured to output and / or input audio signals. For example, the audio component 310 includes a microphone (MIC) that is configured to receive an external audio signal when the driving strategy determination system 300 is in an operational mode, such as an alert mode, a recording mode, a voice recognition mode, and a voice output mode. The received audio signal can be further stored in the memory 304 or transmitted via the communication component 316. In some embodiments, the audio component 310 also includes a speaker for outputting audio signals.
[0192] The I / O interface 312 provides an interface between the processing component 302 and peripheral interface modules, which can include a keypad, a click wheel, buttons, and so on. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0193] The sensor component 314 includes one or more sensors for providing status assessments of various aspects for the driving strategy determination system 300. For example, the sensor component 314 can include a sound sensor. In addition, the sensor component 314 can detect an on / off state of the driving strategy determination system 300, relative positioning of components, such as a display and a keypad for the driving strategy determination system 300, an operational state of the driving strategy determination system 300 or a component of the driving strategy determination system 300, such as a blower operational state, a structure state, a discharge blade operational state, and so on, an orientation or acceleration / deceleration of the driving strategy determination system 300, and a temperature change of the driving strategy determination system 300. The sensor component 314 can include a proximity sensor configured to detect presence of a nearby object without any physical contact. The sensor component 314 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 314 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, a material pile thickness sensor, or a temperature sensor.
[0194] The communication component 316 is configured to enable the driving strategy determination system 300 to provide wired or wireless communication capabilities between the driving strategy determination system 300 and other devices and cloud platforms. The driving strategy determination system 300 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 316 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 316 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0195] In an example embodiment, the driving strategy determination system 300 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements for performing the driving strategy determination method described in any of the above embodiments.
[0196] The present application also provides a computer readable storage medium, when instructions stored in the computer readable storage medium are executed by a processor corresponding to the driving strategy determination system, the driving strategy determination system is enabled to implement the driving strategy determination method described in any of the above embodiments.
[0197] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) including computer-usable program code.
[0198] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device implemented in accordance with the flowcharts and / or block diagrams. Figure 1 The flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one or more flows or blocks.
[0199] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices that implement the flowcharts and / or block diagrams. Figure 1 The flow or multiple flows and / or blocks Figure 1 The device that implements the functions specified in one or more flows or blocks.
[0200] These computer program instructions can also be loaded into a computer or other programmable data processing devices, so that a series of operational steps are performed on the computer or other programmable data processing devices to generate computer-implemented processes, thus the instructions executed on the computer or other programmable data processing devices provide processes for implementing the functions specified in the flowchart Figure 1 one or more flows and / or blocks Figure 1 one or more blocks or steps of the functions specified in the flowchart
[0201] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A driving strategy determination method characterized by comprising: The method comprises the following steps: obtaining a historical speed curve of a vehicle on a target route; introducing the historical speed curve into an agent, and calculating a total reward value corresponding to the historical speed curve by the agent, wherein the agent is used to generate a corresponding driving strategy according to the introduced historical speed curve, the agent is composed of a Q neural network and a target neural network, the Q neural network is used to calculate a current Q value according to the driving strategy in the agent and a vehicle state, and the target neural network is used to calculate a maximum Q value; the total reward value is used to represent a total loss of the vehicle in a driving process, the total loss includes a battery loss of the vehicle, and the total reward value is negatively correlated with the total loss of the vehicle in the driving process; when the total reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, cyclically updating the driving strategy in the agent, and determining a current Q value calculated based on the updated driving strategy and the vehicle state, wherein the current Q value is positively correlated with the total reward value; when a difference between the Q value calculated based on the updated driving strategy and the vehicle state and the maximum Q value is less than a preset difference value, determining that the agent training is completed, so as to determine the driving strategy of the vehicle; calculating the total reward value corresponding to the historical speed curve by the agent, comprising: obtaining preset parameter values of each cycle when the vehicle drives according to the historical speed curve by the agent, wherein the preset parameter values include the battery loss of the vehicle; introducing the preset parameter values of each cycle of the vehicle into a preset calculation model to obtain a reward value of each cycle; accumulating the reward values of each cycle to obtain the total reward value corresponding to the historical speed curve; introducing the preset parameter values of the vehicle into a preset calculation model to obtain a result for representing the total loss of the vehicle, comprising: introducing the power loss, the fuel loss, the battery loss, a vehicle start-stop penalty value and a physical constraint penalty value into the following preset calculation model to calculate a reward value for representing the total loss of the vehicle, and taking the reward value for representing the total loss of the vehicle as a calculation result of the preset calculation model, wherein the vehicle start-stop penalty value is determined according to a start-stop frequency, the more the start-stop frequency, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery power is in a preset power interval and whether the current torque of the vehicle is in a preset torque interval: ; wherein, is a reward value representing the overall loss of the vehicle; is a fuel loss; is an electric power loss; is a battery loss; is a vehicle start-stop penalty value; is a physical constraint penalty value.
2. The method of claim 1, wherein, the battery loss is obtained by the following method: obtaining a cycle life and an actual life of the battery under a nominal condition; substituting the cycle life and the actual life of the battery under the nominal condition into the following first preset formula for calculating a severity factor to obtain a severity factor corresponding to the battery life: ; wherein, is the severity factor; is the total ampere-hour throughput of the nominal battery life; represents the actual ampere-hour throughput of the battery; and are the nominal and actual current values, respectively; substituting preset battery parameters including the severity factor into the following second preset formula for calculating a battery life loss to obtain the battery loss corresponding to the historical speed curve: ; wherein, is the battery loss; is the loss coefficient of replacing the battery; is the severity factor; is the battery corresponding electric quantity, represents the actual ampere-hour throughput of the battery, is the battery life consumption ratio of this vehicle driving.
3. The method of claim 1, wherein, the fuel loss is determined in the following manner: obtaining an engine fuel consumption rate, a unit time fuel loss, a plurality of discrete values corresponding to engine speed and torque, and fuel density; The engine fuel consumption rate, the fuel consumption per unit time, the engine speed, the engine torque and the fuel density are substituted into the following third preset formula to determine the fuel consumption: ; wherein, is fuel consumption; is fuel consumption per unit time; is engine fuel consumption rate; is engine torque, is engine speed; is fuel density.
4. The method of claim 1, wherein, The physical constraint penalty value is determined in the following manner: It is determined whether the current battery power is in a preset power interval and whether the current torque of the vehicle is in a preset torque interval; When the current battery power is in the preset power interval and the current torque of the vehicle is in the preset torque interval, it is determined that the physical constraint penalty value is 0; When the current battery power exceeds the preset power interval or the current torque of the vehicle exceeds the preset torque interval, it is determined that the physical constraint penalty value is a specific value, wherein the specific value makes the reward value of the overall loss of the vehicle less than a preset value.
5. The method of claim 1, wherein, The current driving strategy is updated in cycles, comprising: Randomly selecting a plurality of four-tuple data from an experience pool corresponding to the agent as learning samples, wherein the four-tuple data includes a state at a historical time, an action taken by the agent at the historical time, a reward at the historical time, and a state at a next historical time; The agent is trained in cycles through the learning samples, and the weight parameters are updated through a stochastic gradient descent method; The updated weight parameters are assigned to the Q network to realize the cyclic update of the driving strategy in the agent.
6. A driving strategy determination device characterized by comprising: Comprise: An acquisition module is configured to acquire a historical speed curve of a vehicle on a target route; An import module is configured to import the historical speed curve into an agent, and calculate an overall reward value corresponding to the historical speed curve through the agent, wherein the agent is configured to generate a corresponding driving strategy according to the imported historical speed curve, the agent is composed of a Q neural network and a target neural network, the Q neural network is configured to calculate a current Q value according to the driving strategy in the agent and a vehicle state, and the target neural network is configured to calculate a maximum Q value; the overall reward value is configured to represent an overall loss of the vehicle in a driving process, the overall loss includes a vehicle battery loss, and the overall reward value is negatively correlated with the overall loss of the vehicle in the driving process; A first determination module is configured to cyclically update the driving strategy in the agent when the overall reward value corresponding to the historical speed curve of the vehicle is less than a preset reward value, and determine a current Q value calculated based on the updated driving strategy and the vehicle state, wherein the current Q value is positively correlated with the overall reward value; A second determination module is configured to determine that the training of the agent is completed when a difference between the Q value calculated based on the updated driving strategy and the vehicle state and the maximum Q value is less than a preset difference, so as to determine the driving strategy of the vehicle. The overall reward value corresponding to the historical speed curve is calculated through the agent, comprising: The agent obtains preset parameter values of each cycle when the vehicle drives according to the historical speed curve, wherein the preset parameter values include a vehicle battery loss; The preset parameter values of each cycle are input into a preset calculation model to obtain a reward value of each cycle; The reward values of each cycle are accumulated to obtain the overall reward value corresponding to the historical speed curve; The preset parameter value of the vehicle is introduced into the pre-design calculation model to obtain a result for representing the overall loss of the vehicle, including: The electric quantity loss, fuel loss, battery loss, vehicle start-stop penalty value and physical constraint penalty value of the vehicle are introduced into the following pre-design calculation model to calculate a reward value for representing the overall loss of the vehicle, and the reward value for representing the overall loss of the vehicle is taken as the calculation result of the pre-design calculation model, wherein the vehicle start-stop penalty value is determined according to the number of vehicle start-stop times, the more the start-stop times, the greater the vehicle start-stop penalty value, and the physical constraint penalty value is determined according to whether the current battery electric quantity is in a preset electric quantity interval and whether the current torque of the vehicle is in a preset torque interval: ; wherein, is a reward value representing the overall loss of the vehicle; is a fuel loss; is an electric charge loss; is a battery loss; is a vehicle start-stop penalty value; is a physical constraint penalty value.
7. A driving strategy determination system characterized by comprising: including: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the driving strategy determination method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor corresponding to the driving strategy determination system, the driving strategy determination system can implement the driving strategy determination method according to any one of claims 1-5.
Citation Information
Patent Citations
New energy vehicle ecological driving method based on heterogeneous multi-agent deep reinforcement learning
CN115495997A