A hybrid power system energy management method based on P-DQN algorithm
By using the P-DQN algorithm in the hybrid vehicle energy management system and using the parameterized action space to process discrete and continuous actions, the problem that existing systems are difficult to take into account both fuel economy and power, and better energy management effects are achieved.
Patent Information
- Application Number
- CN202210754220.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The existing hybrid vehicle energy management system is difficult to handle discrete and continuous operations simultaneously, making it difficult to take into account both fuel economy and power.
Using the energy management method based on the P-DQN algorithm, through parameterizing the action space, it is possible to establish a P-DQN proxy model using discrete actions and continuous actions at the same time, train and update the model parameters to achieve better fuel economy and motivation.
On the premise of ensuring the power of the car, better fuel economy can be achieved, the control effect of energy management strategies and the robustness of algorithms can be improved, and different driving conditions can be adapted to different driving conditions.
Smart Images

Figure CN115099148B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hybrid vehicle energy management, and in particular relates to a hybrid vehicle system energy management method based on a P-DQN algorithm. Background Art
[0002] As the energy crisis becomes increasingly serious, hybrid vehicles have begun to occupy a larger proportion in the modern market. Since the power source of a hybrid vehicle includes at least two parts, an internal combustion engine and an electric motor, the energy management system is of great significance to the fuel economy of a hybrid vehicle. The energy management system of a hybrid vehicle can coordinate the cooperation between various power sources to reduce fuel consumption and greenhouse gas emissions.
[0003] At present, the energy management strategies of hybrid vehicles can be divided into three categories according to the design method: rule-based methods, optimization-based methods and learning-based methods. Rule-based energy management strategies are widely used in the current automotive industry due to their fast online calculation and high real-time performance. However, the formulation of rule-based energy management strategies requires the experience of experts and cannot be used for other vehicles. The applicability and robustness are not ideal, and the fuel economy cannot achieve the optimal value. Although the optimization-based strategy can achieve the optimal fuel economy, the formulation of the strategy requires globally known driving conditions and a long calculation time, so it cannot be actually applied.
[0004] In recent years, learning-based algorithms have begun to be applied to the formulation of energy management strategies. Learning-based strategies rely on reinforcement learning algorithms. These algorithms rely on Markov processes to train agents, and action variables are required when training agents. Current reinforcement learning algorithms can basically only handle discrete actions or continuous actions alone, while the energy management of hybrid vehicles involves both discrete and continuous actions. When formulating strategies, current learning-based rules choose to discretize continuous actions and process them with discrete action algorithms, or formulate separate rules for discrete actions, and continuous actions are processed with continuous action algorithms. Summary of the invention
[0005] The present invention provides a hybrid power system energy management method based on the P-DQN algorithm, which uses a parameterized action space, can not only use discrete actions and continuous actions at the same time, but also achieve better fuel economy under the premise of ensuring the power of the vehicle.
[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A hybrid power system energy management method based on a P-DQN algorithm comprises the following steps:
[0008] Step 1: Establish a P-DQN agent model;
[0009] Step 2: Set the state, action, action parameters and reward of the P-DQN agent model to obtain the set P-DQN agent model;
[0010] Step 3: Obtain relevant training data sets, and train the set P-DQN proxy model obtained in step 2 according to the obtained relevant training data sets to obtain a trained P-DQN proxy model;
[0011] Step 4: Use the trained P-DQN agent model for energy management of parallel hybrid vehicles.
[0012] In the above steps, the P-DQN proxy model in step 1 includes: an online network and a target network, both of which include a determined policy network and a deep value network. The online network is responsible for outputting Q values and interacting with the environment. The target network is responsible for calculating the loss gradient for updating the online network parameters. After the online network parameters are updated, the target network soft-updates the parameters of the online network.
[0013] The state variables in step 2 are: vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear. The state variable vector is s={v,acc,SOC,gear} T , transmission shift = {downShift, sustain, upShift} is the action variable, and the parameter of the action variable is p = {T e down , T e sus , T e up}, the parameterized action variable vector is a={(downShift, T e down ),( sustain, T e sus ),( upshift, T e up )} T , the reward function is used to evaluate the state s at time t t Next, perform action a t The reward function is defined as the negative of the sum of the cost function and the duration of a single shift. The cost function is shown in formula (1):
[0014] cost(t)= fuel(t)+α[SOC ref -SOC(t)] 2 (1)
[0015] Among them, fuel(t) is the fuel consumption of the car at the current moment, SOC ref is the reference value of the expected SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight of battery charging maintenance;
[0016] The duration of a single shift is shown in formula (2):
[0017] (2)
[0018] The reward function is shown in formula (3):
[0019] r=-{cost(t)+β*sustainTime(t)} (3)
[0020] Among them, β is the weight of a single shift duration;
[0021] In step 3, a relevant training data set is obtained, and the P-DQN proxy model is trained according to the obtained relevant training data set to obtain a trained P-DQN proxy model, which specifically includes the following steps:
[0022] Step A: Initializing the set P-DQN proxy model to obtain an initialized P-DQN proxy model;
[0023] Step B: The initialized P-DQN agent model interacts with the hybrid electric vehicle to obtain a training data set;
[0024] Step C: training the P-DQN proxy model according to the training data set, and finally obtaining a trained P-DQN proxy model.
[0025] The above step A specifically includes: respectively initializing the online network parameters and the target network parameters in the set P-DQN proxy model, the determination strategy network parameters and the deep value network parameters of the online network are represented by θ and ω respectively, and the determination strategy network parameters and the deep value network parameters of the target network are represented by θ' and ω' respectively, to obtain the initialized P-DQN proxy model;
[0026] The above step B specifically includes: the current state set s={v, acc, SOC, gear} T Input online to determine the policy network, according to the current online policy network π(p t |s t ; θ) obtain the action parameters p corresponding to each discrete action t In order to better explore, the action parameter p t Add noise N to generate new action parameters pt =p t +N; the current state set s={v, acc, SOC, gear}T and the action parameter p t Input the online deep value network, the online deep value network outputs the Q value of each discrete action under the corresponding action parameter, randomly selects the discrete action and its action parameter corresponding to each Q value according to the probability of ε, selects the discrete action and its action parameter corresponding to the maximum Q value according to the probability of (1-ε), and obtains the parameterized action a t Act on the hybrid vehicle and get the current moment reward r t And the state set s at the next moment t+1 ; Finally, according to the above relevant data t ,a t ,r t ,s t+1 , get the training data set (s t ,a t ,r t ,s t+1 ), store the data set in the memory bank R for training the neural network;
[0027] Step C specifically includes the following steps:
[0028] Step (I): Randomly extract M data sets (s i ,a i ,r i ,s i+1 ), will s i , a i Enter the online deep value network to obtain the state s i Next action a i Q value Q (s i ,a i ;ω) and reward r i ; Set the state set s i+1 Input target to determine the policy network and obtain the target action parameter set p i+1 , the action set s i+1 and the target action parameter set p i+1 Input the target deep value network to obtain the target discrete action and the Q value Q(s) corresponding to the target discrete action i+1 ,a i+1 ;ω');
[0029] Step (II): Calculate TD target y i ,y i The calculation formula is:
[0030] y i =r i +γmaxQ(s i+1 ,a i+1 ;ω')
[0031] Among them, γ is the discount factor of the reward
[0032] Calculate the loss function of the online deep value network, the calculation formula is:
[0033] L(ω)=∑[Q(s i ,a i ;ω)-y i ] 2 / M
[0034] Calculate the loss function of the online determined policy network, the calculation formula is:
[0035] L(θ)=∑Q(s t ,a t ;θ)
[0036] Step (III): Derivative the loss function L(ω) of parameter ω with respect to ω, and update the parameter ω using gradient descent. Derivative the loss function L(θ) of parameter θ with respect to θ, and update the parameter θ using gradient ascent. Use soft update to update the target network. The update formula is:
[0037] ω'=τ a ω+(1-τ a )ω'
[0038] θ'=τ p θ+(1-τ p )θ'
[0039] where τ a and τ p Determine the soft update coefficients of the policy network for the target deep value network and target respectively.
[0040] Step (IV): The agent's current state s t Transfer to t+1 , repeat steps (I) to (III) until the training requirements are met, and finally obtain the trained P-DQN agent model.
[0041] The above step 4 specifically includes the following steps:
[0042] Step 1: Obtain the current state quantity set s of the car through relevant sensors t ={v,acc,SOC,gear} T ;
[0043] Step 2: Get the current state set s of the car t ={v, acc, SOC, gear} T Input the trained P-DQN agent, and then output the control variable gear change shift and the corresponding engine torque T e ;
[0044] Step 3: The obtained control amount shift and engine torque T e Act on the car, drive the car to move, and then get the car state set s at the next moment t+1 ={v, acc, SOC, gear} T ;
[0045] Step 4: Repeat steps 1 to 3 until the car completes the driving task.
[0046] The hybrid vehicle energy management method based on parameterized deep Q-learning described above is theoretically data-driven and model-free, generally insensitive to any specific topology of the hybrid system, and is applied to parallel hybrid systems.
[0047] Beneficial effects: The present invention provides a hybrid power system energy management method based on the P-DQN algorithm, which is applicable to the method of predicting energy management of hybrid vehicles using intelligent variable time domain models. First, a P-DQN agent model is established; second, the state, action, action parameters and returns of the P-DQN agent model are set to obtain the set P-DQN agent model; then, relevant training data sets are obtained, and the P-DQN agent model is trained according to the obtained relevant training data sets to obtain the trained P-DQN agent model; finally, the trained P-DQN agent model is used to perform energy management of parallel hybrid vehicles to obtain better control effects. The method of the present invention can effectively handle the control problem of energy management of hybrid vehicles in the hybrid action space, and achieve better fuel economy while ensuring the power of the vehicle;
[0048] At the same time, it can solve the learning stability problem in the DQN method, effectively improve the control effect of the energy management strategy and the speed of the algorithm, improve the robustness and adaptability of the energy management algorithm to working conditions, and further improve the fuel economy of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic diagram of a parameterized action space model provided in an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of a hybrid vehicle energy management structure based on P-DQN provided in an embodiment of the present invention;
[0051] Figure 3 is a schematic flow chart of a hybrid vehicle energy management design method based on P-DQN provided in an embodiment of the present invention;
[0052] Figure 4 Schematic diagram of the structure of the P-DQN proxy model provided in an embodiment of the present invention;
[0053] Figure 5 It is a SOC trajectory diagram under three strategies provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments:
[0055] like Figure 1 As shown in the figure, the energy management action of the parallel hybrid vehicle comes from the parameterized action space. The parameterized action first selects a discrete action shift={downShift, sustain, upShift}, and then selects the parameter p={T e down , T e sus , T e up} T , forming a parameterized action a={(shift, p)} T ={(downShift, T e down ),( sustain, T e sus ),( upShift, T e up )} T .
[0056] like Figure 2 As shown in the figure, a hybrid vehicle energy management structure based on P-DQN is proposed. Its basic principle is to obtain the driving state of the vehicle through relevant sensors, and obtain relevant state quantities, which are vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear, and form the current state variable vector s t ={v,acc,SOC,gear} T , and then the state variable vector s t ={v,acc,SOC,gear} T Input into the trained P-DQN agent, the P-DQN agent outputs the action variable a according to its own strategy t ={(shift, p)} T, acting on the parallel hybrid electric vehicle, the state quantity s at the next moment is obtained t+1 ={v,acc,SOC,gear} T , until the entire driving condition is completed.
[0057] Figure 3 1 is a flow chart of a hybrid vehicle energy management design method based on P-DQN provided in an embodiment of the present invention. According to the flow chart, the design of the hybrid vehicle energy management structure system based on P-DQN is completed, including the following steps:
[0058] Step 201: Establishing a P-DQN proxy model includes, specifically, an online network and a target network, and both include a determination policy network and a deep value network. The online network is responsible for outputting Q values and interacting with the environment, and the target network is responsible for calculating the loss gradient for updating the online network parameters. After the online network parameters are updated, the target network softly updates the parameters of the online network.
[0059] Step 202: setting the state, action, action parameter and reward of the P-DQN agent model to obtain the set P-DQN agent model;
[0060] When setting the state, action, action parameters and reward of the P-DQN agent model, the P-DQN agent model after setting is obtained, specifically including: there are 4 state quantities, namely, vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear, and the state variable vector is s={v,acc,SOC,gear} T , since the controlled object is a hybrid vehicle, the shift operation shift = {downShift, sustain, upShift} is selected as the action, and the engine torque T e For action variable p={T e down , T e sus , T e up}, forming a parameterized action variable vector a={(downShift, T e down ),( sustain, T e sus ),(upshift, T e up )} T , the reward function is used to evaluate the t Next, perform action a tThere are three goals for performance: first, to avoid overcharging and overdischarging of the battery, it is necessary to ensure that the SOC is maintained within a certain range; second, on the basis of ensuring stable operation of the battery, the fuel consumption is minimized to ensure the fuel economy of the engine; third, avoid frequent gear shifting of the transmission to ensure the feasibility of the gear shifting operation and the service life of the transmission. In addition, since deep reinforcement learning determines the optimal control strategy based on the maximum cumulative reward optimization, the reward function is defined as the negative value of the sum of the cost function and the duration of a single gear shift. The cost function is as follows:
[0061] cost(t)= fuel(t)+α[SOC ref -SOC(t)] 2
[0062] Among them, fuel(t) is the fuel consumption of the car at the current moment, SOC ref is the reference value of the expected SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight of the battery charging maintenance, where α is 350.
[0063] The duration of a single shift is as follows:
[0064]
[0065] The reward function is as follows:
[0066] r=-{cost(t)+β*sustainTime(t)}
[0067] Where β is the weight of a single shift duration, and β is 0.6 here;
[0068] Step 203: Acquire a relevant training data set, and train the P-DQN proxy model according to the acquired relevant training data set to obtain a trained P-DQN proxy model;
[0069] When obtaining relevant training data sets and training the constructed model, the following steps are specifically included:
[0070] Figure 4 : is a schematic diagram of the P-DQN proxy model structure provided in an embodiment of the present invention, and its basic principle is: initialize the P-DQN proxy model after the setting to obtain the initialized P-DQN proxy model; set the current state set s={v,acc, SOC, gear} T Input online to determine the policy network, according to the current online policy network π(p t |s t ; θ) obtain the action parameters p corresponding to each discrete action t In order to better explore, the action parameter pt Add noise N to generate new action parameters p t =p t +N; the current state set s={v, acc, SOC, gear}T and the action parameter p t Input the online deep value network, the online deep value network outputs the Q value of each discrete action under the corresponding action parameter, randomly selects the discrete action and its action parameter corresponding to each Q value according to the probability of ε, selects the discrete action and its action parameter corresponding to the maximum Q value according to the probability of (1-ε), and obtains the parameterized action a t Act on the hybrid vehicle and get the current moment reward r t And the state set s at the next moment t+1 ; Finally, according to the above relevant data t ,a t ,r t ,s t+1 , get the training data set (s t ,a t ,r t ,s t+1 ), store the data set in the memory bank R, and randomly extract M data sets (s i ,a i ,r i ,s i+1 ), will s i ,a i Enter the online deep value network to obtain the state s i Next action a i Q value Q (s i ,a i ;ω) and reward r i ; Set the state set s i+1 Input target to determine the policy network and obtain the target action parameter set p i+1 , the action set s i+1 and the target action parameter set p i+1 Input the target deep value network to obtain the target discrete action and the Q value Q(s) corresponding to the target discrete action i+1 ,a i+1 ;ω'), r i +γmaxQ(s i+1 ,a i+1 ; ω') as the TD target y i , we get the loss function of the online deep value network parameter ω L(ω)=∑[Q(s i ,a i ;ω)-y i] 2 / M and online determine the loss function of the policy network parameters θ L(θ) = ∑Q(s t ,a t ; θ), the loss function L(ω) of parameter ω is derived with respect to ω, and the parameter ω is updated using gradient descent. The loss function L(θ) of parameter θ is derived with respect to θ, and the parameter θ is updated using gradient ascent. The parameters ω' and θ' of the target network are obtained by soft updating of the online network, ω'=τ a ω+(1-τ a )ω',θ'=τ p θ+(1-τ p )θ', the agent's current state s t Transfer to t+1 , repeat the above steps until the training goal is completed;
[0071] Step 204: Using the trained PA-DDPG agent model to perform energy management of the parallel hybrid vehicle, specifically including the following steps:
[0072] Step 1: Obtain the current state quantity set s of the car through relevant sensors t ={v,acc,SOC,gear} T ;
[0073] Step 2: Get the current state set s of the car t ={v, acc, SOC, gear} T Input the trained P-DQN agent, and then output the control variable gear change shift and the corresponding engine torque T e ;
[0074] Step 3: The obtained control amount shift and engine torque T e Act on the car, drive the car to move, and then get the car state set s at the next moment t+1 ={v, acc, SOC, gear} T ;
[0075] Step 4: Repeat steps 1 to 3 until the car completes the driving task.
[0076] Figure 5It is a SOC trajectory diagram under three strategies provided in an embodiment of the present invention. It can be seen from the figure that under the three control strategies, the SOC curves fluctuate between 0.6 and 0.4, among which the SOC fluctuation range of the energy management strategy based on learning is between 0.6 and 0.5, indicating that the three methods can well constrain the SOC to ensure that the battery is in a safe SOC range during use, and the energy management strategy based on learning can better maintain the stability of the SOC. At the same time, compared with the energy management strategy based on DDPG, the SOC of the energy management strategy based on P-DQN fluctuates less in the first period of time and decreases more slowly, avoiding rapid discharge, and better maintaining battery health. Compared with the energy management strategy based on DP, the SOC drops to a larger minimum value in the later period of time, avoiding the rapid charging and discharging behavior of the battery in the high load area and the high fuel consumption of the engine in the high load area, ensuring the stability of the SOC and fuel economy. At the same time, it can be concluded that there are essential differences in control ideas between the control strategy based on learning and the control strategy based on DP. The control strategy based on the DP algorithm is more inclined to use the engine to maintain a stable decrease in SOC under low load conditions, and is more inclined to use the motor under high load conditions, resulting in a large degree of fluctuation in SOC in the high load area. The control strategy based on the DDPG algorithm is different from the DP algorithm. The control strategy based on the DDPG algorithm uses both the engine and the motor to maintain the fluctuation range of SOC under various loads. At the same time, due to the use of complete control actions, the control strategy based on P-DQN can balance the use of the engine and the motor under various global loads to maintain a stable change in SOC, and can also maintain a slow decrease in SOC under low loads, thereby ensuring both the vehicle's power and fuel economy and the stability of SOC.
[0077] The above are only preferred embodiments of the present invention. It is obvious that those familiar with the technology in the field can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative work. Therefore, the protection scope of the present invention is not limited to the above embodiments. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and purpose of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A hybrid power system energy management method based on P-DQN algorithm, characterized in that: The following steps are involved: Step 1: Establish a P-DQN agent model; Step 2: Set the state, action, action parameters and reward of the P-DQN agent model to obtain the set P-DQN agent model; the state variables are: vehicle speed v, vehicle acceleration acc, power battery SOC and transmission gear gear, and the state variable vector is s={v,acc,SOC,gear} T , the transmission shift = {downShift, sustain, upShift} is the action variable, and the parameter of the action variable is p = {T e down , T e sus , T e up }, the parameterized action variable vector is a={(downShift, T e down ),(sustain, T e sus ),( upshift, T e up )} T , the reward function is defined as the negative of the sum of the cost function and the duration of a single shift; The cost function is shown in formula (1): cost(t)= fuel(t)+α[SOC ref -SOC(t)] 2 (1) Among them, fuel(t) is the fuel consumption of the car at the current moment, SOC ref is the reference value of the expected SOC, SOC(t) is the battery SOC value at the current moment, and α is the weight of battery charging maintenance; The duration of a single shift is shown in formula (2): (2) The reward function is shown in formula (3): r=-{cost(t)+β*sustainTime(t)} (3) Among them, β is the weight of a single shift duration; Step 3: Obtain relevant training data sets, and train the set P-DQN proxy model obtained in step 2 according to the obtained relevant training data sets to obtain a trained P-DQN proxy model; Step 4: Use the trained P-DQN agent model to perform energy management of the parallel hybrid vehicle, which specifically includes the following steps: Step 1: Obtain the current state quantity set s of the car through relevant sensors t ={v,acc,SOC,gear} T ; Step 2: Get the current state set s of the car t ={v, acc, SOC, gear} T Input the trained P-DQN agent, and then output the control variable gear change shift and the corresponding engine torque T e ; Step 3: The obtained control amount shift and engine torque T e Act on the car, drive the car to move, and then get the car state set s at the next moment t+1 ={v, acc, SOC, gear} T ; Step 4: Repeat steps 1 to 3 until the car completes the driving task.
2. The hybrid power system energy management method based on the P-DQN algorithm according to claim 1 is characterized in that: The P-DQN proxy model in step 1 includes: an online network and a target network, both of which include a determined policy network and a deep value network. The online network is responsible for outputting Q values and interacting with the environment. The target network is responsible for calculating the loss gradient for updating the online network parameters. After the online network parameters are updated, the target network soft-updates the parameters of the online network.
3. The hybrid power system energy management method based on the P-DQN algorithm according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step A: Initialize the P-DQN proxy model to obtain the initialized P-DQN proxy model; Step B: The initialized P-DQN agent model interacts with the hybrid electric vehicle to obtain a training data set; Step C: training the P-DQN proxy model according to the training data set, and finally obtaining a trained P-DQN proxy model.
4. The hybrid power system energy management method based on the P-DQN algorithm according to claim 3 is characterized in that: Step A specifically includes: respectively initializing the online network parameters and the target network parameters in the set P-DQN proxy model, the determination strategy network parameters and the deep value network parameters of the online network are represented by θ and ω respectively, and the determination strategy network parameters and the deep value network parameters of the target network are represented by θ' and ω' respectively, to obtain the initialized P-DQN proxy model.
5. The hybrid power system energy management method based on the P-DQN algorithm according to claim 3 is characterized in that: Step B specifically includes: setting the current state set s={v, acc, SOC, gear} T Input online to determine the policy network, according to the current online policy network π(p t |s t ; θ) obtain the action parameters p corresponding to each discrete action t , in the action parameter p t Add noise N to generate new action parameters p t =p t +N; set the current state set s={v, acc, SOC, gear} T and action parameter p t Input the online deep value network, the online deep value network outputs the Q value of each discrete action under the corresponding action parameter, randomly selects the discrete action and its action parameter corresponding to each Q value according to the probability of ε, selects the discrete action and its action parameter corresponding to the maximum Q value according to the probability of (1-ε), and obtains the parameterized action a t Act on the hybrid vehicle and get the current moment reward r t And the state set s at the next moment t+1 ; Finally, according to the above relevant data t ,a t ,r t ,s t+1 , get the training data set (s t ,a t ,r t ,s t+1 ), store the data set in the memory bank R for training the neural network.
6. The hybrid power system energy management method based on the P-DQN algorithm according to claim 3 is characterized in that: Step C specifically includes the following steps: Step (I): Randomly extract M data sets (s i ,a i ,r i ,s i+1 ), will s i , a i Enter the online deep value network to obtain the state s i Next action a i Q value Q (s i ,a i ;ω) and reward r i ; Set the state set s i+1 Input target to determine the policy network and obtain the target action parameter set p i+1 , the state set s i+1 and the target action parameter set p i+1 Input the target deep value network to obtain the target discrete action and the Q value Q(s) corresponding to the target discrete action i+1 ,a i+1 ;ω'); Step (II): Calculate TD target y i ,y i The calculation formula is: y i =r i +γmaxQ(s i+1 ,a i+1 ; oh) Among them, γ is the discount factor of the reward Calculate the loss function of the online deep value network, the calculation formula is: L(ω)=∑[Q(s i ,a i ;ω)- y i ] 2 / M Calculate the loss function of the online determined policy network, the calculation formula is: L(θ)=∑Q(s t ,a t (i) Step (III): Derivative the loss function L(ω) of parameter ω with respect to ω, and update the parameter ω using gradient descent. Derivative the loss function L(θ) of parameter θ with respect to θ, and update the parameter θ using gradient ascent. Use soft update to update the target network. The update formula is: ω'=τ a ω+(1-τ a )oh' θ'=τ p θ+(1-τ) p )th' where τ a and τ p Determine the soft update coefficients of the policy network for the target deep value network and target respectively; Step (IV): The agent's current state s t Transfer to t+1 , repeat steps (I) to (III) until the training requirements are met, and finally obtain the trained P-DQN agent model.
7. The hybrid power system energy management method based on the P-DQN algorithm according to claim 1, characterized in that: The method is applied to a parallel hybrid power system.
Citation Information
Patent Citations
System and method for constructing molecule reaction force field based on reinforced learning
CN109994158A
Management and control method of hybrid electric vehicle based on digital twinning technology
CN110488629A