An energy management method and system for a hybrid vehicle
By building a working condition information database and SOC reference value model, combined with a model prediction controller, the problem of large amount of energy management strategies for hybrid vehicles is solved and difficult to apply in real time is achieved, and efficient energy management and fuel economy are achieved.
Patent Information
- Application Number
- CN202210799248.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-07-08
AI Technical Summary
When the existing hybrid vehicle energy management strategy seeks global optimal solutions, the large amount of calculation results in delay, making it difficult to realize real-time application, and only computing global optimization is calculated, which has great limitations.
By constructing a working condition information database and vehicle model based on historical working condition information, a deep deterministic strategy gradient algorithm is used to build a SOC reference value model, a SOC reference value trajectory is generated, and the SOC status is updated in real time through the model prediction controller to achieve energy management.
This method can reduce the impact of driver's personal habits on energy management strategies while greatly protecting battery life, improving computing efficiency, making the strategy applicable in real time, and combining instantaneous prediction of energy management strategies with global prediction to improve the fuel economy of the vehicle.
Smart Images

Figure CN115107733B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of energy management of hybrid electric vehicles, and particularly to an energy management method and system for hybrid electric vehicles. Background Art
[0002] Due to problems such as air pollution and depletion of oil resources caused by traditional gasoline and diesel vehicles, the electrification of vehicles has become a popular direction at present. As a key control technology for hybrid electric vehicles, the energy management strategy directly affects the fuel economy of vehicles and has become the research focus of hybrid systems.
[0003] At present, the energy management strategies of hybrid electric vehicles are mainly divided into rule-based and optimization-based energy management strategies. The rule-based energy management strategy mainly designs the system working mode and the energy distribution method of different power sources according to engineering experience and considering the characteristics of each component of the power system. This control strategy has no specific optimization problem, and the formulated rules mainly come from engineering experience, with strong practicability and good real-time effect, but the energy-saving effect is poor, and it can be used as the basis for formulating the optimization-based energy management strategy.
[0004] The existing energy management strategies of hybrid electric vehicles optimize the whole working condition of the hybrid according to the state equation and objective function of the whole vehicle system, using the optimal control theory to obtain the global optimal solution of the feasible region. In theory, it can achieve the optimal global fuel economy, requires the global driving condition of the vehicle, has a large amount of calculation and is easy to cause delay, and is difficult to realize real-time application; and only calculates global optimization, with great limitations. Summary of the Invention
[0005] The present invention provides an energy management method and system for hybrid electric vehicles to solve the technical problems that in the process of the energy management strategy seeking the global optimal solution, the large amount of calculation is easy to cause delay, it is difficult to realize real-time application, and only global optimization is calculated, with great limitations.
[0006] To solve the above technical problems, in a first aspect, an embodiment of the present invention provides an energy management method for hybrid electric vehicles, including:
[0007] Construct a working condition information library and a whole vehicle model of a hybrid electric vehicle based on historical working condition information;
[0008] Based on the working condition information library and the whole vehicle model, construct an SOC reference value model of the deep deterministic policy gradient algorithm, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control;
[0009] During the SOC value control process, based on the historical speed data of the vehicle, the predicted speed of the vehicle is calculated, and based on the predicted speed, a model predictive controller of the vehicle is constructed, and the SOC state is updated in real time through the model predictive controller to achieve energy management.
[0010] In the present invention, by constructing a working condition information library, the driving habits of the driver are fully saved, and the influence of the driver's personal habits on the energy management strategy is reduced; at the same time, by constructing an SOC reference value model, an SOC reference value trajectory is generated, and by controlling the SOC value to travel along the SOC reference value trajectory, the battery life is protected to the greatest extent; secondly, when constructing the SOC reference value model, the deep deterministic policy gradient algorithm is adopted, so that the stability and convergence are greatly improved during calculation, thereby improving the calculation efficiency and enabling the strategy to be applied in real time; and, since the SOC reference value model is constructed before departure, it can be trained offline, which can greatly reduce the memory occupation of the operation; in addition, by updating the SOC state in real time through the model predictive controller, and at the same time through the SOC reference value model control framework, the combination of instantaneous prediction and global prediction of the energy management strategy is realized, and the fuel economy of the vehicle is improved.
[0011] Further, the construction of the working condition information library and the vehicle model of the hybrid vehicle based on the historical working condition information is specifically as follows:
[0012] According to the historical working condition information, the driving conditions are divided into kinematic segments according to the idle state;
[0013] Calculate the characteristic parameters of each kinematic segment, and perform principal component analysis on the characteristic parameters to obtain the contribution rate and cumulative contribution rate of each principal component;
[0014] Select the principal components with a cumulative contribution rate greater than 85% for K-means clustering, and select the clustering center data to construct a working condition information library;
[0015] According to the working condition information library, model the vehicle components, and at the same time establish a reward function to complete the construction of the vehicle model.
[0016] The present invention extracts the characteristics of the driver's historical working condition information to construct a working condition information library for future driving, fully saves the driver's driving habits, and provides an accurate basis for generating the SOC reference value.
[0017] Further, the construction of the SOC reference value model of the deep deterministic policy gradient algorithm based on the working condition information library and the vehicle model is specifically as follows:
[0018] Obtain the working condition information library, and use the vehicle model as the environment of the SOC reference value model;
[0019] Based on the Deep Deterministic Policy Gradient algorithm, an actor network and a critic network are constructed, and the actor network and the critic network together constitute an agent network;
[0020] The agent network and the vehicle model together constitute an SOC reference value model.
[0021] Based on the Deep Deterministic Policy Gradient algorithm, the present invention constructs an actor network and a critic network. The actor network is used to calculate the target policy, the critic network is used to calculate the action value function, evaluate the advantages and disadvantages of the actor network, and finally the agent network interacts with the vehicle model as the environment to achieve the optimal control action;
[0022] Furthermore, constructing the SOC reference value model of the Deep Deterministic Policy Gradient algorithm further includes:
[0023] According to the engine fuel consumption rate curve, obtain the optimal engine operating curve;
[0024] Take the optimal engine operating curve as the constraint condition of the SOC reference value model.
[0025] The present invention takes the optimal engine operating curve as the constraint condition of the SOC reference value model, guiding the SOC reference value model to optimize along the optimal engine operating curve during the training process instead of performing global optimization, greatly improving the calculation efficiency.
[0026] Furthermore, generating the SOC reference value trajectory and controlling the SOC value to change along the SOC reference value trajectory to achieve energy control, specifically:
[0027] Repeat experience replay and reward maximization through the SOC reference value model to generate a target policy set and an SOC reference value trajectory;
[0028] Control the power of the vehicle engine through the target policy set to control the SOC value to change along the SOC reference value trajectory to control energy distribution.
[0029] The present invention can quickly achieve the optimal control action by repeating experience replay and reward maximization to generate a target policy and a reference value trajectory; in addition, since the SOC reference value model has been constructed before departure, it can be trained offline, which can greatly reduce the memory occupation of the operation.
[0030] Furthermore, repeating experience replay and reward maximization through the SOC reference value model to generate a target policy set and an SOC reference value trajectory, specifically:
[0031] Obtain a number of actions to be executed through the vehicle model;
[0032] Execute each execution action in sequence. In each execution, based on the state variables and reward values of the currently pending execution action, and the state variables and reward values of the next pending execution action, combined with the current actor network, generate the target policy for the currently pending execution action. The target policy includes the SOC value, and update the state variables and action variables of the actor network through the chain rule;
[0033] Until each pending execution action is completed, generate a target policy set and an SOC reference value trajectory based on the target policy of each pending execution action.
[0034] The present invention dynamically updates the actor network through the chain rule, continuously optimizes the SOC reference value model, and improves the stability of the SOC reference value model.
[0035] Further, calculating the predicted speed of the vehicle based on the historical speed data of the vehicle specifically includes:
[0036] Construct a forget gate, an input gate, and an output gate, add the forget gate, the input gate, and the output gate to the recurrent neural network, and construct a speed prediction model based on the long short-term memory network algorithm;
[0037] According to the historical speed data of the vehicle, calculate the predicted speed of the vehicle through the speed prediction model.
[0038] The present invention constructs a long short-term memory network by adding a forget gate, an input gate, and an output gate to the recurrent neural network, effectively solving the problem of gradient explosion in the recurrent neural network. At the same time, the long short-term memory network can effectively retain past sequential patterns, which is more conducive to our instantaneous prediction of the speed at a certain moment.
[0039] Further, constructing a model predictive controller for the vehicle based on the predicted speed specifically includes:
[0040] Obtain the predicted speed of the vehicle;
[0041] Calculate the required torque and rotational speed of the vehicle through the SOC reference value model based on the predicted speed;
[0042] Construct an objective function for minimizing the vehicle fuel consumption based on the required torque, rotational speed, SOC reference value, and SOC penalty factor of the vehicle. The objective function is the model predictive controller of the vehicle.
[0043] The present invention constructs a model predictive controller through the predicted speed. The present invention constructs a model predictive controller through the predicted speed, and at the same time introduces a penalty factor in the model predictive controller to achieve accurate following of the SOC reference value.
[0044] Further, updating the SOC state in real time through the model predictive controller to achieve energy management specifically includes:
[0045] Calculate the minimum fuel consumption through the said model prediction controller;
[0046] Within the prediction domain, through the dynamic programming optimizer, calculate the optimal motor torque sequence when the fuel consumption is minimized;
[0047] Adjust the motor torque through the optimal motor torque value corresponding to the optimal motor torque sequence at the current moment, and update the SOC state to achieve energy management.
[0048] The present invention calculates the minimum fuel consumption through the prediction controller to calculate the optimal motor torque, and updates the SOC state by adjusting the motor torque, realizing the instantaneous prediction of energy management, improving the fuel economy of the vehicle, and realizing the application of the intelligent optimization algorithm to actual real-time optimization.
[0049] In a second aspect, an embodiment of the present invention provides an energy management system for a hybrid vehicle, including: a working condition information processing module, an SOC reference value module, and a predictive control module.
[0050] The working condition information processing module is used to construct a working condition information library and a vehicle model of the hybrid vehicle according to historical working condition information.
[0051] The SOC reference value module is used to construct an SOC reference value model of the deep deterministic policy gradient algorithm according to the working condition information library and the vehicle model, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control.
[0052] The predictive control module is used to calculate the predicted speed of the vehicle based on the historical speed data of the vehicle during the SOC value control process, and construct a model predictive controller of the vehicle based on the predicted speed, and update the SOC state in real time through the model predictive controller to achieve energy management.
[0053] The present invention constructs a working condition information library to fully preserve the driving habits of the driver, reducing the influence of the driver's personal habits on the energy management strategy; at the same time, by constructing an SOC reference value model, generating an SOC reference value trajectory, and controlling the SOC value to travel along the SOC reference value trajectory, the battery life is protected to the greatest extent; secondly, the deep deterministic policy gradient algorithm is adopted in constructing the SOC reference value model, so that the stability and convergence are greatly improved during calculation, thereby improving the calculation efficiency and enabling the strategy to be applied in real time; and, since the SOC reference value model is constructed before departure, it can be trained offline, which can greatly reduce the memory occupation of the operation; in addition, the SOC state is updated in real time through the predictive controller, and at the same time through the SOC reference value model control framework, the combination of instantaneous prediction and global prediction of the energy management strategy is realized, improving the fuel economy of the vehicle. Description of the Drawings
[0054] Figure 1 is a schematic flowchart of an energy management method for a hybrid vehicle provided by an embodiment of the present invention;
[0055] Figure 2 is a structural diagram of an SOC reference value model based on the deep deterministic policy gradient algorithm provided by an embodiment of the present invention;
[0056] Figure 3 is a schematic flowchart of step 102 provided by an embodiment of the present invention;
[0057] Figure 4 is a schematic flowchart of step 103 provided by an embodiment of the present invention;
[0058] Figure 5 is a schematic structural diagram of an energy management system for a hybrid vehicle provided by an embodiment of the present invention. Detailed Embodiments
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment 1
[0061] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an energy management method for a hybrid vehicle provided by an embodiment of the present invention, mainly including steps 101 to 103, specifically as follows:
[0062] Step 101: Based on historical driving condition information, construct a driving condition information library and a vehicle model of the hybrid vehicle;
[0063] In this embodiment, the vehicle processor first divides the driving conditions into kinematic segments according to the idle state based on the historical driving condition information; then calculates the characteristic parameters of each kinematic segment, performs principal component analysis on the characteristic parameters, and then obtains the contribution rate and cumulative contribution rate of each principal component; then selects the principal components with a cumulative contribution rate greater than 85% for K-means clustering, selects the clustering center data to construct the driving condition information library; finally, based on the driving condition information library, models the vehicle components and establishes a reward function to complete the construction of the vehicle model.
[0064] As a specific implementation manner of the embodiment of the present invention, the characteristic parameters include: average speed, average driving speed, maximum speed, speed standard deviation, maximum acceleration, minimum acceleration, average acceleration, maximum deceleration, minimum deceleration, average deceleration, acceleration standard deviation, idling time ratio, acceleration time ratio, deceleration time ratio, total running time, and total running mileage characteristic parameter.
[0065] In this embodiment, by performing longitudinal dynamics analysis on the hybrid truck, the vehicle components are modeled. The demand power model of the vehicle is specifically:
[0066]
[0067] Among them, P dem is the demand power, F f is the rolling resistance, F w is the air resistance, F i is the gradient resistance, F j is the acceleration resistance, η is the mechanical transmission efficiency, and u a is the driving speed.
[0068] In this embodiment, the resistance model of the vehicle is specifically:
[0069]
[0070] Among them, m is the vehicle mass, g is the gravitational acceleration, f is the rolling resistance coefficient, α is the road gradient, C D is the air resistance coefficient, A is the frontal area, ρ is the air density, u a is the driving speed, δ is the rotating mass conversion coefficient, is the driving acceleration.
[0071] In this embodiment, the power of the hybrid truck comes from the engine and the motor, and the demand power model is specifically:
[0072] P dem =(P eng +P mot )η (3)
[0073] Among them, P dem is the demand power, P eng is the engine power, P mot is the motor power, and η is the mechanical transmission efficiency.
[0074] In this embodiment, the engine model is established by numerical modeling, and the fuel consumption rate is obtained by interpolation. Specifically, the engine fuel consumption model is specifically:
[0075]
[0076] Among them, is the engine fuel consumption rate, T eng is the engine torque, n eng is the engine speed.
[0077] In this embodiment, the motor adopts a numerical modeling method, and the efficiency is obtained by looking up a table. Specifically, the motor efficiency model is:
[0078] η mot = f(T mot , n mot ) (5)
[0079] Among them, η mot is the motor efficiency, T mot is the motor torque, n mot is the motor speed.
[0080] In this embodiment, the battery model is simplified to an equivalent internal resistance model. Specifically:
[0081]
[0082] Among them, P bat is the battery power, U oc is the open-circuit voltage, R0 is the battery internal resistance, I(t) is the battery current, Q bat is the battery capacity, and SOC0 is the initial state SOC of the battery.
[0083] In this embodiment, the reward function is specifically:
[0084] reward = -{af rate (n eng , T eng ) + bP bat (t)} (7)
[0085] Among them, f rate is the engine fuel consumption rate, a is the fuel consumption weight, and b is the battery power consumption weight.
[0086] Step 102: Based on the working condition information library and the vehicle model, construct an SOC reference value model of the deep deterministic policy gradient algorithm, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control;
[0087] Please refer to Figure 2 , Figure 2 which is the structural diagram of the SOC reference value model based on the deep deterministic policy gradient algorithm provided by the embodiment of the present invention.
[0088] In this embodiment, based on the working condition information library and the vehicle model, the vehicle constructs an SOC reference value model using the deep deterministic policy gradient algorithm. The SOC reference value model includes the environment and the agent. The vehicle model is used as the environment, and the agent is constructed. Finally, the agent network interacts with the vehicle model used as the environment to generate the SOC reference value trajectory.
[0089] In this embodiment, the core of the agent network consists of an actor network and a critic network. The actor network is a deterministic policy function for outputting policies, and the critic network is an action value function for evaluating the quality of the actor network. The value function of the actor network is represented by the Bellman equation as follows:
[0090] Q π (s t ,a t )=E[r(s t ,a t )+γE[Q π (s t+1 ,a t+1 )]] (8)
[0091] Where: E is the environment model, π is the policy, r(s t ,a t ) is the reward function, and γ is the discount factor.
[0092] In this embodiment, to find the optimal policy, the target policy can be expressed as:
[0093] Q μ (s t ,a t )=E[r(s t ,a t )+γQ μ (s t+1 ,μ(s t+1 ))] (9)
[0094] In this embodiment, the loss function of the Q network is:
[0095]
[0096] In this embodiment, after obtaining the target policy, the state variables and action variables of the actor network are updated through the chain rule:
[0097]
[0098] In this embodiment, by controlling the power of the vehicle engine, the SOC value is controlled to change along the SOC reference value trajectory, achieving the optimal control action to control the energy distribution.
[0099] In this embodiment, according to the engine consumption rate curve, the optimal engine operating curve is obtained, and the optimal engine operating curve is used as the constraint condition of the SOC reference value model.
[0100] Step 103: During the SOC value control process, based on the historical speed data of the vehicle, calculate the predicted speed of the vehicle, and based on the predicted speed, construct a model predictive controller for the vehicle, and update the SOC state in real time through the model predictive controller to achieve energy management.
[0101] In this embodiment, first, a forget gate, an input gate, and an output gate are constructed, and the forget gate, the input gate, and the output gate are added to the recurrent neural network, so as to construct a speed prediction model based on the long short-term memory network algorithm; then, according to the historical speed data of the vehicle, the predicted speed of the vehicle is calculated through the speed prediction model.
[0102] In this embodiment, the output value f of the forget gate t is a judgment of the degree of forgetting at the previous moment, and its value ranges from 0 to 1, where 0 represents complete forgetting and 1 represents complete memory; the forget gate is specifically:
[0103] f t =σ(W fh H t-1 +W fx X t +b f ) (12)
[0104] where σ is the sigmoid function, W fh , W fx are the weight matrices of the forget gate, b f is the bias vector of the forget gate, H t-1 is the hidden layer information at time t-1, and X t is the input information.
[0105] In this embodiment, the input gate i t is a process of screening the input state and normalizing it into a recognizable vector, and at the same time calculating the candidate cell state C′ t , and finally calculating the new cell state C t , the input gate is specifically:
[0106]
[0107] where σ is the sigmoid function, tanh is the hyperbolic tangent function, W ih , W ix , W ch , W cx are the weight matrices, b i , bc is the bias vector, ⊙ is the Hadamard product operation, and C t-1 is the cell state at time t-1.
[0108] In this embodiment, the output gate O t is a process of screening cell state information, used to determine which cell states will continue to be retained and transmitted. Finally, through the cell state C t and the output gate O t jointly construct the hidden layer information H at the current moment t . The output gate is specifically:
[0109]
[0110] where σ is the sigmoid function, W oh , W ox are the weight matrices of the input gate, b o is the bias vector of the input gate, tanh is the hyperbolic tangent function, and ⊙ is the Hadamard product operation.
[0111] In this embodiment, according to the predicted speed of the vehicle through the SOC reference value model, the required torque and rotational speed of the vehicle are calculated; and according to the required torque, rotational speed, rotational speed, SOC reference value and SOC penalty factor of the vehicle, an objective function for minimizing the vehicle fuel consumption is constructed. The objective function is the model predictive controller of the vehicle.
[0112] Specifically, the objective function is specifically:
[0113]
[0114] where α is the fuel consumption factor, β is the power consumption factor, γ is the SOC penalty factor, and m fuel is the engine fuel consumption.
[0115] The present invention constructs a working condition information library to fully store the driving habits of the driver, reducing the influence of the driver's personal habits on the energy management strategy; at the same time, by constructing an SOC reference value model to generate an SOC reference value trajectory, and controlling the SOC value to travel along the SOC reference value trajectory, the battery life is protected to the greatest extent; secondly, when constructing the SOC reference value model, the deep deterministic policy gradient algorithm is adopted, so that the stability and convergence are greatly improved during calculation, thereby improving the calculation efficiency and enabling the strategy to be applied in real time; and, since the SOC reference value model is constructed before departure, it can be trained offline, which can greatly reduce the memory occupation of the operation; in addition, by predicting the controller to update the SOC state in real time, and at the same time through the SOC reference value model control framework, the combination of instantaneous prediction and global prediction of the energy management strategy is realized, improving the fuel economy of the vehicle.
[0116] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of step 102 provided by an embodiment of the present invention, and mainly includes steps 301 to 305, specifically as follows:
[0117] Step 301: The vehicle model obtains the first state variable and the first reward value of the current action, and executes the current action to reach the next action.
[0118] Step 302: The SOC reference value model puts the second state variable and the second reward value of the next action into the experience pool.
[0119] Step 303: The actor network generates the target policy of the current action according to the first state variable, the first reward value, the second state variable and the second reward value.
[0120] In this embodiment, the target policy includes a state variable and an action variable. The action variable is the engine power, and the state variables include the SOC value, speed and acceleration of the vehicle. During the energy control process, the SOC state is controlled by controlling the engine power.
[0121] Step 304: The SOC reference value model updates the action variables of the state variables in the actor network according to the target policy.
[0122] In this embodiment, the actor network is dynamically updated by the chain rule, continuously optimizing the SOC reference value model and improving the stability of the SOC reference value model.
[0123] Step 305: Repeat steps 301 to 304 until the working condition ends.
[0124] In this embodiment, by repeating and completing steps 201 to 204, a set of target policies for all actions to be executed within the working condition is generated, and at the same time, an SOC reference value trajectory is generated, and the change of the SOC value is controlled according to the SOC reference value trajectory, extending the protection life of the battery. In addition, since the SOC reference value model is constructed before departure, offline training can be performed, which can greatly reduce the memory occupation of the operation.
[0125] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of step 103 provided by an embodiment of the present invention, and mainly includes steps 401 to 404, specifically as follows:
[0126] Step 401: Construct a model predictive controller for the vehicle according to the predicted speed.
[0127] Step 402: Calculate the minimum fuel consumption through the model predictive controller.
[0128] Step 403: Calculate the optimal motor torque sequence when the fuel consumption is minimized through the dynamic programming optimizer within the prediction domain.
[0129] Step 404: Adjust the motor torque through the optimal motor torque value corresponding to the optimal motor torque sequence at the current moment, and update the SOC state to achieve energy management.
[0130] The present invention calculates the minimum fuel consumption through the predictive controller to calculate the optimal motor torque, and updates the SOC state by adjusting the motor torque, realizing the instantaneous prediction of energy management, improving the fuel economy of the vehicle, and realizing the application of the intelligent optimization algorithm to the actual real-time optimization.
[0131] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an energy management system for a hybrid vehicle provided by an embodiment of the present invention, and mainly includes a working condition information processing module 501, an SOC reference value module 502, and a predictive control module 503;
[0132] In this embodiment, the working condition information processing module 501 is used to construct a working condition information library and a vehicle model of the hybrid vehicle according to historical working condition information;
[0133] The SOC reference value module 502 is used to construct an SOC reference value model of the deep deterministic policy gradient algorithm according to the working condition information library and the vehicle model, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control;
[0134] The predictive control module 503 is used to calculate the predicted speed of the vehicle based on the historical speed data of the vehicle during the SOC value control process, and construct a model predictive controller of the vehicle based on the predicted speed, and update the SOC state in real time through the model predictive controller to achieve energy management.
[0135] The present invention constructs a working condition information library to fully preserve the driving habits of drivers, reducing the impact of driver personal habits on the energy management strategy. At the same time, by constructing an SOC reference value model to generate an SOC reference value trajectory and controlling the SOC value to travel along the SOC reference value trajectory, the battery life is protected to the greatest extent. Secondly, the deep deterministic policy gradient algorithm is adopted when constructing the SOC reference value model, greatly improving the stability and convergence during calculation, thereby improving the calculation efficiency and enabling the real-time application of this strategy. Moreover, since the SOC reference value model is constructed before departure and can be trained offline, it can greatly reduce the memory occupation of the operation. In addition, the SOC state is updated in real time through a predictive controller, and at the same time, through the SOC reference value model control framework, the combination of instantaneous prediction and global prediction of the energy management strategy is realized, improving the fuel economy of the vehicle.
[0136] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An energy management method for a hybrid vehicle, characterized in that, Including: Based on historical operating condition information, construct an operating condition information library and a whole vehicle model of a hybrid vehicle; Based on the operating condition information library and the whole vehicle model, construct an SOC reference value model of the deep deterministic policy gradient algorithm, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control; During the process of controlling the SOC value, based on the historical speed data of the vehicle, calculate the predicted speed of the vehicle, and based on the predicted speed, construct a model predictive controller of the vehicle, and update the SOC state in real time through the model predictive controller to achieve energy management; The constructing a model predictive controller of the vehicle based on the predicted speed specifically includes: obtaining the predicted speed of the vehicle; calculating the required torque and rotational speed of the vehicle through the SOC reference value model according to the predicted speed; constructing an objective function with the minimum vehicle fuel consumption according to the required torque, rotational speed, SOC reference value and SOC penalty factor of the vehicle, and the objective function is the model predictive controller of the vehicle.
2. The energy management method of a hybrid vehicle according to claim 1, wherein The constructing an operating condition information library and a whole vehicle model of a hybrid vehicle based on the historical operating condition information specifically includes: According to the historical operating condition information, divide the driving conditions into kinematic segments according to the idle state; Calculate the characteristic parameters of each kinematic segment, and perform principal component analysis on the characteristic parameters to obtain the contribution rate and cumulative contribution rate of each principal component; Select the principal components with a cumulative contribution rate greater than 85% for K-means clustering, and select the clustering center data to construct an operating condition information library; According to the operating condition information library, model the whole vehicle components, and at the same time establish a reward function to complete the construction of the whole vehicle model.
3. The energy management method of a hybrid vehicle according to claim 1, wherein, The constructing an SOC reference value model of the deep deterministic policy gradient algorithm based on the operating condition information library and the whole vehicle model specifically includes: Obtain the operating condition information library and use the whole vehicle model as the environment of the SOC reference value model; Based on the deep deterministic policy gradient algorithm, construct an actor network and a critic network, and the actor network and the critic network together constitute an agent network; The agent network and the whole vehicle model together constitute the SOC reference value model.
4. The energy management method for a hybrid vehicle according to claim 3, characterized in that, The constructing an SOC reference value model of the deep deterministic policy gradient algorithm further includes: According to the engine consumption rate curve graph, obtain the optimal engine operating curve; Use the optimal engine operating curve as the constraint condition of the SOC reference value model.
5. The energy management method for a hybrid vehicle according to claim 4, characterized in that, The generating an SOC reference value trajectory and controlling the SOC value to change along the SOC reference value trajectory to achieve energy control specifically includes: Repeatedly complete experience replay and reward maximization through the SOC reference value model to generate a target policy set and an SOC reference value trajectory; Control the power of the vehicle engine through the target policy set to control the SOC value to change along the SOC reference value trajectory to control energy distribution.
6. The energy management method of a hybrid vehicle according to claim 5, wherein, The repeatedly completing experience replay and reward maximization through the SOC reference value model to generate a target policy set and an SOC reference value trajectory specifically includes: Obtain a number of actions to be executed through the whole vehicle model; Execute each execution action in sequence, and in each execution, generate the target policy for the currently pending execution action based on the state variable and reward value of the currently pending execution action, the state variable and reward value of the next pending execution action, and the current actor network. The target policy includes the SOC value, and update the state variable and action variable of the actor network through the chain rule; After each pending execution action is completed, generate a target policy set and an SOC reference value trajectory based on the target policy of each pending execution action.
7. The energy management method for a hybrid vehicle according to claim 1, characterized in that, Based on the historical speed data of the vehicle, calculate the predicted speed of the vehicle. Specifically: Construct a forget gate, an input gate, and an output gate, add the forget gate, the input gate, and the output gate to the recurrent neural network, and construct a speed prediction model based on the long short-term memory network algorithm; Based on the historical speed data of the vehicle, calculate the predicted speed of the vehicle through the speed prediction model.
8. The energy management method of a hybrid vehicle according to claim 1, wherein, The SOC state is updated in real time through the model predictive controller to achieve energy management. Specifically: Calculate the minimum fuel consumption through the model predictive controller; Within the prediction domain, calculate the optimal motor torque sequence when the fuel consumption is minimized through a dynamic programming optimizer; Adjust the motor torque through the optimal motor torque value corresponding to the optimal motor torque sequence at the current moment, and update the SOC state to achieve energy management.
9. An energy management system for a hybrid vehicle, characterized in that, It includes: A working condition information processing module, an SOC reference value module, and a predictive control module; The working condition information processing module is used to construct a working condition information library and a vehicle model of a hybrid vehicle based on historical working condition information; The SOC reference value module is used to construct an SOC reference value model of the deep deterministic policy gradient algorithm based on the working condition information library and the vehicle model, generate an SOC reference value trajectory, and control the SOC value to change along the SOC reference value trajectory to achieve energy control; The predictive control module is used to calculate the predicted speed of the vehicle based on the historical speed data of the vehicle during the SOC value control process, and construct a model predictive controller of the vehicle based on the predicted speed. The SOC state is updated in real time through the model predictive controller to achieve energy management; Based on the predicted speed, construct a model predictive controller of the vehicle. Specifically: Obtain the predicted speed of the vehicle; Calculate the required torque and speed of the vehicle through the SOC reference value model according to the predicted speed; Construct an objective function for minimizing the vehicle fuel consumption based on the required torque, speed, SOC reference value, and SOC penalty factor of the vehicle. The objective function is the model predictive controller of the vehicle.
Citation Information
Patent Citations
SOC (state of charge) reference trajectory based energy management method for plug-in hybrid electric vehicles
CN109895760A
Hybrid electric vehicle hierarchical prediction energy management method fused with deep reinforcement learning
CN113525396A
Variable equivalent factor hybrid electric vehicle energy management method based on working condition identification
CN114179777A