An Energy Management Method for Hybrid Power Systems Based on the PA-DDPG Algorithm

By adopting the PA-DDPG algorithm in the energy management of hybrid vehicles and using parameterized action space to process discrete and continuous actions, the problem of insufficient fuel economy and power in the prior art is solved, better fuel economy and power are achieved, and the robustness and adaptability of energy management strategies are improved.

CN115016285BActive Publication Date: 2025-05-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210754519.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-05-27
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

The existing hybrid vehicle energy management system is difficult to achieve optimal fuel economy and power in the case of discrete and continuous control actions while handling simultaneous discrete and continuous control actions, and there are problems of insufficient applicability and robustness based on rules and optimization methods.

Method used

Using the energy management method based on the PA-DDPG algorithm, discrete and continuous actions can be processed simultaneously by parameterizing the action space, and fuel consumption can be optimized while ensuring power. The method includes establishing a PA-DDPG proxy model, setting state, action, action parameters and returns, obtaining training data sets, and obtaining a model for energy management through training.

Benefits of technology

It realizes better fuel economy and power in energy management of hybrid vehicles, improves the control effect of energy management strategies and the robustness and adaptability of algorithms, and solves the learning stability problem in the DDPG method and the problem of only handling discrete actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115016285B_ABST
    Figure CN115016285B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid system energy management method based on a PA-DDPG algorithm, which belongs to the technical field of hybrid vehicle energy management. By using a parameterized action space, not only can discrete actions and continuous actions be used simultaneously, but also better fuel economy can be achieved under the premise of ensuring the power of the vehicle. The present invention includes the following steps: establishing a PA-DDPG agent model; setting the state, action, action parameters and return of the PA-DDPG agent model to obtain the PA-DDPG agent model after setting; obtaining a relevant training data set, and training the PA-DDPG agent model according to the obtained relevant training data set to obtain a trained PA-DDPG agent model; finally, using the trained PA-DDPG agent model to perform energy management of a parallel hybrid vehicle to obtain a better control effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hybrid vehicle energy management, and in particular relates to a hybrid vehicle system energy management method based on a PA-DDPG algorithm. Background Art

[0002] With the development of science and technology, the use of energy in industry is increasing, among which the automobile industry occupies a certain proportion in the industry. In order to solve the automobile industry's dependence on oil, the proportion of hybrid vehicles in the automobile industry continues to increase. Since hybrid vehicles combine the advantages of internal combustion engines and motors, their power sources include at least internal combustion engines and motors. Therefore, the energy management system of hybrid vehicles is of great significance to fuel economy. An effective energy management system can coordinate the cooperation between various power sources to reduce fuel consumption and greenhouse gas emissions.

[0003] At present, the energy management strategy of hybrid vehicles is mainly designed based on three methods: rule-based method, optimization-based method and learning-based method. Among them, the rule-based energy management strategy uses the set rules to calculate the torque distribution, with fast calculation speed and high real-time performance. It is widely used in the current automotive industry and widely distributed. However, when formulating the rule-based energy management strategy, it is necessary to formulate it for a specific model based on the experience of experts, and it cannot be used for other vehicles. It is more dependent on the level of experience, and its applicability and robustness are not ideal, and the fuel economy cannot achieve the best; the optimization-based strategy is calculated based on the global working condition, so it can obtain the best fuel economy, but the formulation of the strategy requires obtaining the globally known driving conditions in advance and requires a long calculation time, so it cannot be applied in practice. In recent years, learning-based algorithms have begun to be applied to the formulation of energy management strategies.

[0004] Learning-based strategies need to rely on reinforcement learning algorithms. These algorithms rely on Markov processes to train agents. Action variables are required when training agents. The action variables used in current reinforcement learning are single discrete actions or continuous actions. When there are both discrete and continuous control actions, the continuous action can only be discretized and processed using a discrete action algorithm, or the discrete action can be processed using rules and the continuous action can be controlled using a continuous algorithm alone. Summary of the invention

[0005] The present invention provides a hybrid power system energy management method based on the PA-DDPG algorithm, which uses a parameterized action space, can not only use discrete actions and continuous actions at the same time, but also can achieve better fuel economy under the premise of ensuring the power of the vehicle.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A hybrid power system energy management method based on PA-DDPG algorithm includes the following steps:

[0008] Step 1: Establish the PA-DDPG agent model;

[0009] Step 2: Set the state, action, action parameters and reward of the PA-DDPG agent model to obtain the set PA-DDPG agent model;

[0010] Step 3: Obtain relevant training data sets, and train the trained PA-DDPG proxy model obtained in step 2 according to the obtained relevant training data sets to obtain the trained PA-DDPG proxy model;

[0011] Step 4: Use the trained PA-DDPG agent model for energy management of parallel hybrid vehicles.

[0012] In the above steps, the PA-DDPG agent model in step 1 includes: an online network and a target network, both of which include an actor network and a critic network. The online network is responsible for outputting actions and interacting with the environment. The target network is responsible for calculating the loss gradient for updating the online network parameters. After the online network parameters are updated, the target network soft-updates the parameters of the online network.

[0013] The state variables in step 2 are: vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear gear. The state variable vector is s = {v, acc, SOC, gear} T , transmission shift = {downShift, sustain, upShift} is the action variable, and the parameter of the action variable is p = {T e down , T e sus , T e up}, the parameterized action variable vector is a = {(downShift, T e down ),(sustain,T e sus ),(upshift,T e up )} T , the reward function is used to evaluate the state s at time t t Next, perform action a tThe reward function is defined as the negative of the sum of the cost function and the duration of a single shift. The cost function is shown in formula (1):

[0014] cost(t)=fuel(t)+α[SOC ref -SOC(t)] 2 (1)

[0015] Among them, fuel(t) is the fuel consumption of the car at the current moment, SOC ref is the reference value of the expected SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight of battery charging maintenance;

[0016] The duration of a single shift is shown in formula (2):

[0017]

[0018] The reward function is shown in formula (3):

[0019] r=-{cost(t)+β*sustainTime(t)} (3)

[0020] Among them, β is the weight of a single shift duration;

[0021] Step 3 specifically includes the following steps:

[0022] Step A: Initializing the PA-DDPG proxy model after the setting to obtain an initialized PA-DDPG proxy model;

[0023] Step B: The initialized PA-DDPG agent model interacts with the HEV to obtain a training dataset;

[0024] Step C: Train the PA-DDPG proxy model according to the training data set, and finally obtain a trained PA-DDPG proxy model.

[0025] Step A specifically includes: initializing the online network parameters and target network parameters in the PA-DDPG agent model after the setting, and the actor network parameters and critic network parameters of the online network are respectively initialized using θ μ and θ Q Indicates that the actor network parameters of the target network and the critic network parameters are θ μ′ and θ Q′ Indicates that the initialized PA-DDPG agent model is obtained;

[0026] Step B specifically includes the following steps: Set the current state set s = {v, acc, SOC, gear} T Input the online actor network, and according to the current online actor network strategy μ, output all the continuous discrete actions logShift and the corresponding action parameters p to form a set of actions a shift ={logShift, p} T , take the discrete action corresponding to the maximum value of logShift, select the shift corresponding to the maximum value of logShift as the discrete action selected at the current moment, and select the action set a shift ={logShift, p} T In order to better explore, add noise N to the action parameter p to generate a new action parameter p = p + N; the current state set s = {v, a, SOC, gear} T and action set a shift ={logShift, p} T Input the online critic network, the online critic network outputs the execution of action set a in state s shift The Q value of the obtained parameterized action a t Act on the hybrid vehicle and get the current moment reward r t And the state set s at the next moment t+1 ; Finally, according to the above relevant data t ,a t ,r t ,s t+1 , get the training data set (s t ,a t ,r t ,s t+1 ), store the data set into the memory bank R for training the neural network;

[0027] In step C, the PA-DDPG proxy model is trained according to the training data set to finally obtain a trained PA-DDPG proxy model, which specifically includes the following steps:

[0028] Step (I): Randomly extract M data sets (s i ,a i ,r i ,s i+1 ), will s i ,a i Enter the onlinecritic network to get the state s i Next action a i The Q value Q(si ,a i |θ Q );Set the state set s i+1 Input the targetactor network and obtain the target action set a i+1 , set the action set a i+1 and the state set s i+1 Input the target critic network and obtain the target network Q value Q′(s i+1 ,a i+1 |θ Q′ ).

[0029] Step (II): Calculate TD target y i ,y i The calculation formula is:

[0030] y i =r i +γQ′(s i+1 ,a i+1 |θ Q′ )

[0031] Among them, γ is the discount factor of the reward

[0032] Calculate the loss function of the online critic network, the calculation formula is:

[0033] L(θ Q )=1 / M(∑[Q(s i ,a i |θ Q )-y i ] 2 )

[0034] Calculate the loss gradient of the online actor network using the following formula:

[0035]

[0036] Step (III): Set the parameter θ Q The loss function L(θ Q ) for θ Q Derivative, using gradient ascent to update the parameter θ Q , update the parameter θ using gradient ascent μ ; Use soft update to update the target network, the update formula is:

[0037]

[0038]

[0039] θ Q′=τθ Q +(1-τ)θ Q′

[0040] θ μ′ =τθ μ +(1-τ)θ μ′

[0041] Where α is the learning rate of the online network, and τ is the soft update coefficient of the target network.

[0042] Step (IV): The agent's current state s t Transfer to t+1 , repeat steps (I) to (III) until the training requirements are met, and finally obtain the trained PA-DDPG agent model;

[0043] In step 4, the trained PA-DDPG agent model is used to manage the energy of the parallel hybrid vehicle, which specifically includes the following steps:

[0044] Step 1: Obtain the current state quantity set s of the car through relevant sensors t ={v,acc,SOC,gear} T ;

[0045] Step 2: Get the current state set s of the car t ={v,acc,SOC,gear} T Input the trained PA-DDPG agent, and then output the control variable gear change shift and the corresponding engine torque Te;

[0046] Step 3: Apply the obtained control quantity shift and engine torque Te to the car to drive it, and then obtain the car state quantity set s at the next moment t+1 ={v,acc,SOC,gear} T ;

[0047] Step 4: Repeat steps 1 to 3 until the car completes the driving task.

[0048] The hybrid vehicle energy management method based on the parameterized deterministic policy gradient algorithm described above is theoretically data-driven and model-free, is generally insensitive to any specific topology of the hybrid system, and is applied to a parallel hybrid system.

[0049] Beneficial effect: The present invention provides a hybrid system energy management method based on the PA-DDPG algorithm, which is suitable for the method of predicting energy management of hybrid vehicles using an intelligent variable time domain model. First, a PA-DDPG agent model is established; second, the state, action, action parameters and return of the PA-DDPG agent model are set to obtain the PA-DDPG agent model after setting; then, a relevant training data set is obtained, and the PA-DDPG agent model is trained according to the obtained relevant training data set to obtain the trained PA-DDPG agent model; finally, the trained PA-DDPG agent model is used to perform energy management of parallel hybrid vehicles to obtain better control effects. The method of the present invention can effectively handle the control problem of energy management of hybrid vehicles in the hybrid action space, and achieve better fuel economy while ensuring the power of the vehicle. At the same time, it can solve the learning stability problem and the problem of only being able to handle discrete actions in the DDPG method, effectively improve the control effect of the energy management strategy and the rapidity of the algorithm, improve the robustness and adaptability of the energy management algorithm to working conditions, and further improve the fuel economy of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic diagram of a parameterized action space model provided in an embodiment of the present invention;

[0051] Figure 2 is a schematic diagram of a PA-DDPG-based hybrid vehicle energy management structure provided in an embodiment of the present invention;

[0052] Figure 3 Schematic diagram of the PA-DDPG proxy model structure provided in an embodiment of the present invention;

[0053] Figure 4 is a flow chart of a hybrid vehicle energy management design method based on PA-DDPG provided in an embodiment of the present invention;

[0054] Figure 5 It is the SOC trajectory diagram under three strategies. DETAILED DESCRIPTION

[0055] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments:

[0056] like Figure 1 As shown in the figure, the energy management action of the parallel hybrid vehicle comes from the parameterized action space. The parameterized action first selects a discrete action shift = {downShift, sustain, upShift}, and then selects the parameter p associated with the discrete action = {T e down , T esus , T e up} T , forming a parameterized action a = {(shift, p)} T ={(downShift, T e down ),(sustain,T e sus ),(upShift,T e up )} T .

[0057] like Figure 2 As shown in the figure, a hybrid vehicle energy management structure based on PA-DDPG is shown. Its basic principle is to obtain the driving state of the vehicle through relevant sensors, obtain relevant state quantities, namely vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear, and form the current state variable vector s t ={v,acc,SOC,gear} T , and then the state variable vector s t ={v,acc,SOC,gear} T Input into the trained PA-DDPG agent, the PA-DDPG agent outputs the action variable a according to its own strategy t ={(shift, p)} T , acting on the parallel hybrid electric vehicle, the state quantity s at the next moment is obtained t+1 ={v,a,SOC,gear} T , until the entire driving condition is completed.

[0058] Figure 3 The figure is a flow chart of a hybrid vehicle energy management design method based on PA-DDPG provided in an embodiment of the present invention. According to the flow chart, the design of the hybrid vehicle energy management structure system based on PA-DDPG is completed, including the following steps:

[0059] Step 201: Establish a PA-DDPG agent model, which specifically includes an online network and a target network. Both include an actor network and a critic network. The online network is responsible for outputting actions and interacting with the environment. The target network is responsible for calculating the loss gradient for updating the online network parameters. After the online network parameters are updated, the target network soft-updates the parameters of the online network.

[0060] Step 202: setting the state, action, action parameters and feedback of the PA-DDPG agent model to obtain the set PA-DDPG agent model;

[0061] When setting the state, action, action parameters and reward of the PA-DDPG agent model, the PA-DDPG agent model after setting is obtained, specifically including: there are 4 state quantities, namely, vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear gear, and the state variable vector is s = {v, acc, SOC, gear} T Since the controlled object is a hybrid vehicle, the shift operation is selected as the action, and the engine torque T e is the action variable p, forming a parameterized action variable vector a = {(shift, p)} T , the reward function is used to evaluate the current state s t Next, perform action a t There are three goals for performance: first, to avoid overcharging and overdischarging of the battery, it is necessary to ensure that the SOC is maintained within a certain range; second, on the basis of ensuring stable operation of the battery, the fuel consumption is minimized to ensure the fuel economy of the engine; third, avoid frequent gear shifting of the transmission to ensure the feasibility of the gear shifting operation and the service life of the transmission. In addition, since deep reinforcement learning determines the optimal control strategy based on the maximum cumulative reward optimization, the reward function is defined as the negative value of the sum of the cost function and the duration of a single gear shift. The cost function is as follows:

[0062] cost(t)=fuel(t)+α[SOC ref -SOC(t)] 2

[0063] Among them, fuel(t) is the fuel consumption of the car at the current moment, SOC ref is the reference value of the expected SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight of the battery charging maintenance, where α is 350.

[0064] The duration of a single shift is as follows:

[0065]

[0066] The reward function is as follows:

[0067] r=-{cost(t)+β*sustainTime(t)}

[0068] Among them, β is the weight of the duration of a single gear shift, and β is taken as 0.6 here;

[0069] Step 203: Acquire a relevant training data set, and train the PA-DDPG proxy model according to the acquired relevant training data set to obtain a trained PA-DDPG proxy model;

[0070] When obtaining relevant training data sets and training the constructed model, the following steps are specifically included:

[0071] Figure 4 The PA-DDPG agent training diagram after the configuration provided in the embodiment of the present invention is as follows: the state variable vector s t ={v,acc,SOC,gear} T Input to the online actor network, through the online actor network strategy μ, output all the continuous discrete actions logShift and the corresponding action parameters p, forming a set of actions a shift ={logShift, p} T , take the discrete action corresponding to the maximum value of logShift, select the shift corresponding to the maximum value of logShift as the discrete action selected at the current moment, and select the action set a shift ={logShift, p} T In order to better explore, add noise N to the action parameter p to generate a new action parameter p = p + N; the current state set s = {v, acc, SOC, gear} T and action set a shift ={logShift, p} T Input onlinecritic network, online critic network output in state s, perform action set a shift At the same time, the obtained parameterized action a t Act on the hybrid vehicle and get the current moment reward r t And the state set s at the next moment t+1 , according to the above relevant data t , a t , r t and t+1 Get the training data set (s t , a t , r t ,s t+1 ), the training data set (s t , a t , r t ,s t+1 ) into the memory bank, randomly extract M data sets from the memory bank, and store M of them (s i , ai ) input into the online network, and obtain M Q values ​​Q(s i ,a i |θ Q ), and M of them (s i+1 ) Input the target network and get the parameterized action a i+1 , and M Q′ values ​​Q′(s i+1 ,a i+1 |θ Q′ ), r i +γQ′(s i+1 ,a i+1 |θ Q′ ) as the TD target y i , get the parameters θ of the online critic network Q The loss function L(θ Q )=1 / M(∑[Q(s i ,a i |θ Q )-y i ] 2 ), do gradient ascent to update the parameter θ Q , calculate the loss gradient of the online actor network Do gradient ascent to update the parameter θ μ , the parameter θ of the target network Q′ and θ μ′ Obtained by online network soft update, θ Q′ =τθ Q +(1-τ)θ Q′ ,θ μ′ =τθ μ +(1-τ)θ μ′ , repeat the above steps until the training goal is completed;

[0072] Step 204: Using the trained PA-DDPG agent model to perform energy management of the parallel hybrid vehicle, specifically including the following steps:

[0073] Step 1: Obtain the current state quantity set s of the car through relevant sensors t ={v,acc,SOC,gear}T;

[0074] Step 2: Get the current state set s of the car t ={v,acc,SOC,gear} T Input the trained PA-DDPG agent, and then output the control variable shift and the corresponding engine torque T e ;

[0075] Step 3: The obtained control amount shift and engine torque T e Act on the car, drive the car to move, and then get the car state set s at the next moment t+1 ={v,acc,SOC,gear} T ;

[0076] Step 4: Repeat steps 1 to 3 until the car completes the driving task.

[0077] Figure 5 This is the SOC trajectory diagram under three strategies. It can be seen from the figure that under the three control strategies, the SOC curve fluctuates between 0.7 and 0.4, indicating that the three methods can well constrain the SOC and ensure that the battery is in the safe SOC range during use. However, the SOC of the energy management strategy based on PA-DDPG fluctuates less than that of the energy management strategy based on DDPG in the previous period, and can obtain better battery SOC stability. At the same time, it has a higher terminal SOC, indicating that the energy management strategy based on PA-DDPG can achieve better fuel economy than the energy management strategy based on DDPG. At the same time, it can be obtained that there is an essential difference in the control idea between the control strategy based on learning and the control strategy based on DP. The control strategy based on the DP algorithm is more inclined to use the engine to maintain the stable decrease of SOC under low load conditions, and is more inclined to use the motor under high load conditions, resulting in a large degree of fluctuation in the SOC in the high load area; and the control strategy based on the PA-DDPG algorithm is different from the DP algorithm. The control strategy based on the PA-DDPG algorithm uses the engine and the motor in a balanced manner under various loads, which not only ensures the power of the vehicle, but also ensures the stability of SOC.

[0078] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0079] The above are only preferred embodiments of the present invention. It is obvious that those familiar with the technology in the field can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the protection scope of the present invention is not limited to the above embodiments. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and purpose of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A method for energy management of a hybrid power system based on the PA-DDPG algorithm, characterized in that, it includes the following steps: Step 1: Establish a PA-DDPG agent model; Step 2: Set the state, action, action parameters, and reward of the PA-DDPG proxy model; the state is: vehicle speed v, vehicle acceleration acc, power battery SOC, and transmission gear, and the state variable vector is s = {v, acc, SOC, gear} T , the transmission shift shift = {downShift, sustain, upShift} is the action variable, and the parameter of the action variable is p = {T e down , T e sus , T e up},the parameterized action variable vector is a = {(downShift, T e down ),(sustain, T e sus ),(upshift, T e up )} T , the reward function is used to evaluate the performance of executing the action a t at time t in the state s t . The reward function is defined as the negative value of the sum of the cost function and the single shift duration; the cost function is: cost(t)= fuel(t)+α[SOC ref -SOC(t)] 2 Among them, fuel(t) is the fuel consumption of the vehicle at the current moment, SOC ref is the reference value of the desired SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight for maintaining battery charging; The single-shift duration is: , The reward function is: r = -{cost(t) + β * sustainTime(t)} where β is the weight of the single-shift duration; Step 3: Obtain relevant training data sets, and train the trained PA-DDPG agent model obtained in Step 2 according to the obtained relevant training data sets to obtain a trained PA-DDPG agent model; specifically including the following steps: Step A: Initialize the set PA-DDPG agent model to obtain an initialized PA-DDPG agent model; Step B: Interact the initialized PA-DDPG agent model with the hybrid vehicle to obtain a training dataset. Let the state set at the current moment be s = {v, acc, SOC, gear} T Input it into the online actor network. According to the policy μ of the current online actor network, output all continuous discrete actions logShift and the corresponding action parameters p, forming a set of action a shift ={logShift, p} T , select the shift corresponding to the maximum value of logShift as the discrete action selected at the current moment, and at the same time select the action parameters p corresponding to the action set a shift ={logShift, p} T ; Let the state set at the current moment be s = {v, a, SOC, gear} T and the action set a shift ={logShift, p} T Input it into the online critic network. The online critic network outputs the Q value of executing the action set a shift in the state s. Apply the parameterized action a t to the hybrid vehicle to obtain the reward r at the current moment t and the state set s at the next moment t+1 ; Finally, according to the above relevant data s t , a t , r t , s t+1 , obtain the training dataset (s t , a t , r t , s t+1 ), store the dataset in the memory bank R for the training of the neural network; Step C: Train the PA-DDPG agent model according to the training data set, and finally obtain a trained PA-DDPG agent model; Step 4: Use the trained PA-DDPG agent model for energy management of a parallel hybrid vehicle.

2. The method for energy management of a hybrid power system based on the PA-DDPG algorithm according to claim 1, characterized in that, the PA-DDPG agent model in Step 1 includes: an online network and a target network, and both the online network and the target network include an actor network and a critic network.

3. The method for energy management of a hybrid power system based on the PA-DDPG algorithm according to claim 1, characterized in that, in order to better explore, noise N is added to the action parameter p to generate a new action parameter p = p + N.

4. The method for energy management of a hybrid power system based on the PA-DDPG algorithm according to claim 1, characterized in that, Step C specifically includes the following steps: Step (Ⅰ): Randomly extract M data sets (s i ,a i ,r i ,s i+1 ) from the memory bank R. Input s i ,a i into the online critic network to obtain the Q-value Q(s i for taking action a i , i.e., Q(s i ,a i |θ Q ). Input the state set s i+1 into the target actor network to obtain the target action set a i+1 . Input the action set a i+1 and the state set s i+1 into the target critic network to obtain the target network Q-value Q′(s i+1 ,a i+1 |θ Q′ ). Step (Ⅱ): Calculate the TD target y i , y i The calculation formula for is: y i =r i +γQ′(s i+1 ,a i+1 |θ Q′ ) where γ is the discount factor of the reward Calculate the loss function of the online critic network, and the calculation formula is: L(θ Q ) = 1 / M (∑[Q(s i , a i |θ Q ) - y i ) 2 ) Calculate the loss gradient of the online actor network, and the calculation formula is: ▽ θ μ μ = ▽ a Q(s i , a i |θ Q )▽ θ μ μ(s t ) Step (Ⅲ): For the parameter θ Q of the loss function L(θ Q ), take the derivative with respect to θ Q and update the parameter θ Q using gradient ascent; update the parameter θ μ using gradient ascent; update the target network using soft update, and the update formula is: , , , , where α is the learning rate of the online network, and τ is the soft update coefficient of the target network; Step (Ⅳ): Agent the state s at the current moment t Transfer to s t+1 , repeat Step (Ⅰ) to Step (Ⅲ) until the training requirements are met, and finally obtain the trained PA-DDPG agent model.

5. The method for energy management of a hybrid power system based on the PA-DDPG algorithm according to claim 1, characterized in that, Step 4 specifically includes the following steps: Step 1: Obtain the set s of current state variables of the vehicle through relevant sensors t ={v, a, SOC, gear} T ; Step 2: Input the set s of current vehicle state variables obtained t = {v, a, SOC, gear} T to the trained PA-DDPG agent, and then output the control variable gear shift and the corresponding engine torque Te; Step 3: Apply the obtained control variables shift and engine torque Te to the vehicle to drive the vehicle, and then obtain the set of vehicle state variables s at the next moment t+1 ={v, a, SOC, gear} T ; The fourth step: Repeat the first step to the third step like this until the vehicle completes the driving task.

6. Application of the method for energy management of a hybrid power system based on the PA-DDPG algorithm according to any one of claims 1-5 in a parallel hybrid system.

Citation Information

Patent Citations

  • Hybrid power system energy management method based on A3C algorithm

    CN112084700A

  • TD3-based heuristic series-parallel hybrid power energy management method

    CN112249002A