A method for optimizing the entire lifecycle operation of photovoltaic energy storage systems based on deep reinforcement learning
By proposing a deep reinforcement learning-based method for optimizing the entire lifecycle of photovoltaic energy storage systems, this study solves the optimization problem of photovoltaic energy storage systems under uncertain environments, achieves efficient and economical operation of photovoltaic energy storage systems, meets the charging needs of electric vehicles and the consumption of photovoltaic power generation, and improves the convergence performance and sampling efficiency of the model.
Patent Information
- Application Number
- CN202411452747.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-10-17
AI Technical Summary
Existing methods for optimizing the operation of photovoltaic energy storage systems suffer from problems such as large computational loads, inaccurate or conservative results when dealing with uncertainties in photovoltaic power generation and electric vehicle charging loads under uncertain environments. Furthermore, the operating efficiency and lifespan of energy storage systems affect economic benefits, making it difficult to effectively solve the problem of optimizing the operation of photovoltaic energy storage systems.
A deep reinforcement learning-based approach is adopted to establish an optimized operation strategy for the entire life cycle of a photovoltaic energy storage system. By finely modeling the energy storage operation efficiency and capacity decay model, and combining deep reinforcement learning for near-end strategy optimization, the energy storage charging and discharging strategy is optimized in real time. Considering the uncertainties of photovoltaic power generation, electric vehicle charging demand and electricity prices, a stochastic diagonal Gaussian strategy is used to process the continuous action space.
This approach enables efficient and optimized operation of photovoltaic energy storage systems under uncertain environments, improves the economic benefits of energy storage systems, meets the charging needs of electric vehicles and the consumption of photovoltaic power generation, avoids the complexity of penalty constraints in the reward function, and improves the convergence performance and sampling efficiency of the model.
Smart Images

Figure CN119647638B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power and energy, and specifically relates to a method for optimizing the operation of a photovoltaic energy storage system throughout its entire life cycle based on deep reinforcement learning. Background Technology
[0002] The widespread adoption of electric vehicles (EVs) is a crucial pathway to achieving carbon neutrality and an effective solution to the global energy crisis and environmental challenges. Utilizing new energy sources as the primary power source for EV charging stations can achieve true "low carbon" emissions. Photovoltaic (PV) systems are new energy charging station systems that use photovoltaic technology to charge EVs. Because PV power generation is closely related to weather factors, EV charging demand also exhibits significant uncertainty. PV systems are typically equipped with energy storage devices to achieve energy regulation, improve PV consumption efficiency, and generate economic profits through "low-storage, high-generation." Some research has been conducted on scheduling strategies for PV energy storage systems. These strategies address the uncertainty of EV charging loads by optimizing the active and reactive power of energy storage. Stochastic dynamic programming is used to solve the scheduling problem of PV systems. Fuel cells and energy storage aim to reduce the operating costs of PV systems and their impact on the power grid. Stochastic optimization and robust optimization methods are used to handle the impact of uncertainties such as EV charging loads and PV output during PV system operation.
[0003] However, stochastic optimization relies on accurate probability distribution models, resulting in poor out-of-sample performance of the optimization results, and large-scale stochastic optimization problems are also quite large. Robust optimization, due to its consideration of worst-case scheduling schemes, can be conservative. Distributed robust optimization methods are used to handle uncertainties in the location and capacity of photovoltaic energy storage systems. While this method addresses the problems of model inaccuracy in stochastic optimization and conservative results in robust optimization, its computational cost is relatively high, limiting its application in real-time scheduling. On the other hand, the operating efficiency and lifespan of energy storage directly affect the economic benefits of photovoltaic energy storage systems, making the consideration of refined energy storage operation models significant. However, in existing photovoltaic energy storage system optimization problems, the complex characteristics of uncertainties such as photovoltaic power generation and electric vehicle charging loads, as well as the nonlinear operating characteristics of energy storage systems, greatly increase the difficulty of solving optimization problems based on analytical mathematical models. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the existing technology and provide an optimization technology that can effectively solve the problem of optimizing the operation of photovoltaic energy storage systems under uncertain environments.
[0005] To achieve the above objectives, the technical solution provided by the present invention is as follows:
[0006] A method for optimizing the entire lifecycle operation of a photovoltaic energy storage system based on deep reinforcement learning, used for real-time optimization control of charging and discharging in a photovoltaic energy storage system for electric vehicles, includes the following steps:
[0007] S1. Perform detailed modeling of the energy storage operation efficiency model and capacity decay model;
[0008] S2. Establish an optimized operation strategy for photovoltaic energy storage systems based on deep reinforcement learning.
[0009] To maximize the revenue of the photovoltaic energy storage system, an objective function for optimizing the operation of the photovoltaic energy storage system is established. The variables of the objective function include charging revenue, grid transaction revenue, and energy storage capacity decay cost. A deep reinforcement learning model is established to determine the state vector using photovoltaic power generation, total electric vehicle charging demand, electricity price, and energy storage SOC for any given time period. The reward function of the deep reinforcement learning model consists of revenue reward and energy storage capacity decay cost.
[0010] S3. Solve the model using deep reinforcement learning based on proximal policy optimization.
[0011] Reinforcement learning based on the current state s t Select action a from the action space. t And evaluate the quality of the current action based on the advantage evaluation function;
[0012] The action space satisfies the following formula:
[0013] a = [a t ]
[0014]
[0015] In the formula, P es,t Let P be the energy storage capacity in the t-th time period. es,max It is the maximum value of the energy storage output power;
[0016] The advantage evaluation function A π,γ The formula is as follows:
[0017] A π,γ (s t ,a t )=Q π,γ (s t ,a t )-V π,γ (s t ),
[0018]
[0019]
[0020] In the formula, Aπ,γ V π,γ and Q π,γ These are the advantage function, value function, and action value function, respectively; π θ The strategy for choosing an action by the actor is represented by θ; γ represents the discount factor, which determines the impact of future rewards on cumulative rewards.
[0021] The network is updated using the following strategy:
[0022]
[0023]
[0024]
[0025] In the above formula, ∈ is a hyperparameter controlling the allowable policy deviation; ρ t It is the ratio coefficient between the updated strategy and the old strategy;
[0026] And a random diagonal Gaussian strategy is used to realize the continuous action space:
[0027] a = μ θ (s)+σ θ (s)⊙x,
[0028] Where x is a sample vector of a standard multivariate normal distribution; μ and σ are the mean and standard deviation of the action vector; ⊙ is the product of elements;
[0029] According to the random diagonal Gaussian policy, π θ (a t |s t This can be deduced as:
[0030]
[0031] This invention uses deep reinforcement learning (IRL) to optimize electric vehicle charging strategies and performs detailed modeling of energy storage operation efficiency and capacity decay models. It solves the optimization operation problem of photovoltaic energy storage systems under uncertain environments and fully considers the nonlinear operation characteristics of energy storage systems. It adopts a stochastic diagonal Gaussian strategy with continuous action space characteristics to optimize energy storage charging and discharging strategies in real time.
[0032] Furthermore, the detailed modeling of the energy storage operation efficiency model in step S1 includes the following steps:
[0033] The charge-discharge efficiency of a battery is calculated using an equivalent model of its steady-state circuit. The steady-state circuit of a single battery includes the open-circuit voltage V. oc The open-circuit voltage is represented by the series resistor R1 connected in series with the open-circuit voltage, the short-time response resistor R2 related to the transient response, and the long-time response resistor R3.
[0034] Each variable and battery S SOC The charge state exhibits a nonlinear relationship, as shown in the following equation:
[0035]
[0036] Among them, I batt P is the battery current. out R is the output power; oc For terminating resistance; g0-g5, b0-b5, c0-c2, and d0-d2 are all coefficients;
[0037] For a specific SOC and P out The current flowing through a single cell is calculated using the following formula:
[0038]
[0039] The charge / discharge efficiency of a single battery is as follows:
[0040]
[0041]
[0042] An energy storage system consists of multiple individual battery cells connected in series and parallel. The calculation methods for equivalent current and charge / discharge efficiency are as follows:
[0043]
[0044]
[0045] Among them, I eq I is the current of the energy storage system; I is the equivalent current of multiple series- and parallel-connected individual cells; N s N is the number of batteries connected in series; p It refers to the number of batteries connected in parallel;
[0046]
[0047] Among them, e 0- e5 and h0-h5 are coefficients.
[0048] Furthermore, in step S1, the energy storage capacity decay model is refined using a semi-empirical battery capacity decay model. The battery capacity decay model over N cycles is shown in the following equation:
[0049]
[0050] Among them, E loss E′ loss The decay capacity before and after N cycles, E′, are respectively.loss =0 indicates that the battery has not been charged or discharged; α s β s f is the influence coefficient of the solid electrolyte interface formation in the new battery; d The battery capacity decay function for a single cycle can be expressed as a function of temperature T, δ, SOC, and time t1, as shown in the following equation:
[0051] f d =(f δ (δ)+f t (t1))f SOC (S SOC )f T (T),
[0052]
[0053] Among them, f δ (δ), f t (t1), f SOC (S SOC ), f T (T) is the coefficient representing the influence of DOD, time, SOC, and temperature on the lifespan of the energy storage battery. ref This is a reference temperature;
[0054] The remaining usable capacity E of the energy storage battery re for:
[0055] E re =E ini -E 1oss ,
[0056] In the formula, E ini It is the maximum capacity of the energy storage battery before it degrades.
[0057] Furthermore, the objective function for optimizing the operation of the photovoltaic energy storage system in step S2 is as follows:
[0058]
[0059] I ev,t =λ ev P ev,t Δt,
[0060]
[0061] Among them, I ev,t I represents the charging revenue of the photovoltaic system that provides electricity to electric vehicles during the t-th time period. grid,t Let C be the revenue generated from the transaction between the photovoltaic system and the power grid during the t-th time period. b,t P is the energy storage capacity degradation cost of photovoltaic power in the t-th time period; T0 is the number of periods in a cycle; ev,tλ is the total power demand for electric vehicle charging via photovoltaics in the t-th time period; ev This refers to the price of charging electric vehicles; Δt is a unit of time; P grid,t ρ is the electricity traded between the photovoltaic charging station and the power grid during the t-th time period; stg,t and ρ bfg, t represents the price at which the photovoltaic energy storage system sells electricity to the grid and purchases electricity during the t-th time period.
[0062] More specifically, the reward function in step S2 consists of revenue rewards and energy storage capacity decay costs, as shown in the following formula:
[0063] r t =(I ev,t +I grid,t -C b,t ).
[0064] The optimized operation strategy of the photovoltaic energy storage system in step S2 above satisfies the following constraints:
[0065] 1) Power balance constraints
[0066] P s,t +P es,t =P ev,t +P grid,t ,
[0067] In the formula, P s,t P represents the photovoltaic power in the t-th time period. es,t P represents the energy storage power in the t-th time period. ev,t P is the total power demand for electric vehicle charging via photovoltaics in the t-th time period; grid,t It represents the electricity traded between the photovoltaic charging station and the power grid during the t-th time period;
[0068] 2) Energy storage constraints
[0069] -P es,max ≤P es,t ≤P es,max ,
[0070]
[0071]
[0072] To ensure that the energy storage operates normally in the next control cycle, the SOC of the energy storage operation at the end of the cycle should be a fixed value, as shown in the following formula:
[0073]
[0074] In the above formula, P es,max It is the maximum value of the energy storage output power; Let SOC and SOC be the minimum and maximum values of the stored energy in the t-th time period, respectively. P represents the remaining capacity of the energy storage battery in the t-th time period; es,t E represents the discharge power of the energy storage battery during the t-th time period. ini This is the maximum capacity of the energy storage battery before it degrades. and These are the charging efficiency and discharging efficiency of the energy storage battery in the t-th time period, respectively.
[0075] The present invention has the following advantages over the prior art:
[0076] This invention proposes a method for optimizing the entire lifecycle operation of a photovoltaic (PV) energy storage system based on deep reinforcement learning. This method involves refined modeling of the energy storage operation efficiency and capacity decay models, and establishes an optimized operation strategy for the PV energy storage system based on deep reinforcement learning, aiming to maximize the system's revenue. It also considers the uncertainties in electric vehicle charging demand, PV power output, and electricity prices to meet both EV charging needs and PV power consumption. Since the energy storage charging and discharging decision-making behavior is continuous, deep reinforcement learning based on proximal policy optimization (PPO) is used to address this issue. The model is trained using real historical data and optimizes the energy storage charging and discharging strategy in real time based on the current state.
[0077] Compared to existing reward functions that rely on meticulous design, the reward function used in this invention is more direct and easier to implement. Battery operation must meet safety constraints, including maximum / minimum SOC and maximum power capacity. If DRL is directly applied to control the operation of the energy storage battery, adding a penalty for violating the constraints to the objective function is indispensable. Although the penalty function method is widely used, determining an appropriate penalty coefficient is difficult; too small a penalty coefficient may lead to non-compliance with constraints, while too large a penalty coefficient will introduce significant errors, resulting in a degraded agent performance. Therefore, the choice of penalty coefficient is crucial for conventional reinforcement learning algorithms, as it is a key parameter dominating the algorithm's optimality and convergence. In contrast, the reward function defined in this invention is very direct and easy to implement. It avoids introducing penalties for constraint violations into the reward function. This is attributed to the excellent convergence performance of the PPO agent.
[0078] The PPO agent used in this invention is characterized by employing a sheared agent objective function, which improves sampling efficiency and convergence speed. Furthermore, unlike most value function-based DRL agents which are only applicable to discrete action spaces, the PPO agent has the characteristic of a continuous action space, thus enabling better utilization of the profitability of overlay services. Attached Figure Description
[0079] Figure 1This is a flowchart of the photovoltaic energy storage system full life cycle optimization operation method based on deep reinforcement learning according to the present invention;
[0080] Figure 2 This is a schematic diagram of the equivalent model of the steady-state circuit of the energy storage battery in an embodiment of the present invention;
[0081] Figure 3 This is a schematic diagram of the operation mode of the photovoltaic energy storage system in an embodiment of the present invention. Detailed Implementation
[0082] The present invention will now be described in detail with reference to the accompanying drawings.
[0083] This invention proposes a method for optimizing the entire lifecycle operation of photovoltaic energy storage systems based on deep reinforcement learning, applicable to photovoltaic energy storage systems for electric vehicles (such as...). Figure 3 The real-time optimization control of charging and discharging (as shown) involves the following steps: Figure 1 As shown, it includes:
[0084] Step 1) Perform detailed modeling of the energy storage operation efficiency model and capacity decay model;
[0085] Step 2) Establish an optimized operation strategy for the photovoltaic energy storage system based on deep reinforcement learning. With the goal of maximizing the benefits of the photovoltaic energy storage system, the strategy takes into account the uncertainties of electric vehicle charging demand, photovoltaic power generation output and electricity price, so as to meet the charging demand of electric vehicles and the consumption of photovoltaic power generation.
[0086] Step 3) Solve the model using deep reinforcement learning based on proximal policy optimization (PPO);
[0087] Step 4) Train using actual historical data and optimize the energy storage charging and discharging strategy in real time based on the current state.
[0088] Furthermore, the refined modeling representation of the energy storage operation efficiency model is as follows:
[0089] The charge and discharge efficiency of a battery can be calculated using an equivalent model of the battery's steady-state circuit. The steady-state circuit of a single battery is as follows: Figure 2 As shown. Where V oc I is the open-circuit voltage. batt P is the battery current. out R1 is the output power; R2 and R3 are the short-time response resistor and long-time response resistor, respectively, which are related to the transient response.
[0090] The above variables and battery S SOC The charge state has a nonlinear relationship, as shown in formula (1).
[0091]
[0092] Among them, R oc For terminating resistors, g0-g5, b0-b5, c 0- c2 and d0-d2 are both coefficients.
[0093] For a specific SOC (State of Charge, which refers to the percentage of a battery's rated capacity remaining) and P out (Output power of the energy storage battery), the current flowing through a single battery can be obtained by solving equation (2).
[0094]
[0095] The charge / discharge efficiency of a single battery can be expressed by a formula.
[0096]
[0097]
[0098] The energy storage battery used consists of multiple individual cells connected in series and parallel. The calculation methods for equivalent current and charge / discharge efficiency are as follows:
[0099]
[0100]
[0101] Among them, I eq I is the current of the energy storage battery; I is the equivalent current of multiple series- and parallel-connected individual cells; N s N is the number of batteries connected in series; p It refers to the number of batteries connected in parallel.
[0102]
[0103] Among them, e0-e5 and h0-h5 are coefficients.
[0104] Furthermore, the refined model of energy storage capacity decay is expressed as follows:
[0105] Battery capacity degradation is a function of ambient temperature, DOD (Depth of Discharge, a parameter used to measure the percentage difference between the battery's discharged capacity and its rated capacity), SOC (State of Charge), and battery runtime. Battery capacity degradation is a non-linear process.
[0106] This example uses a semi-empirical battery capacity decay model. The battery capacity decay model over N cycles is shown in Equation (8).
[0107]
[0108] Among them, Eloss E′ loss E′ represents the decay capacity before and after N cycles. loss =0 indicates the battery has not been charged or discharged. α s β s f represents the influence coefficient of the solid electrolyte interface formation in the new battery. d The battery capacity decay function for a single cycle can be expressed as a function of temperature T, δ, SOC and time t1, as shown in formulas (9) and (10).
[0109] f d =(f δ (δ)+f t (t1))f SOC (S SOC )f T (T) (9)
[0110]
[0111] Among them, f δ (δ), f t (t1), f SOC (S SOC ), f T (T) is the coefficient representing the influence of DOD, time, SOC, and temperature on the lifespan of the energy storage battery. ref This is a reference temperature. The remaining usable capacity E of the energy storage battery. re for:
[0112] E re =E ini -E 1oss (11)
[0113] In the formula, E ini It is the maximum capacity of the energy storage battery before it degrades.
[0114] Furthermore, an optimized operation strategy for photovoltaic energy storage systems based on deep reinforcement learning is established.
[0115] (1) Establish the objective function for optimizing the operation of the photovoltaic energy storage system
[0116] The optimization objective is to maximize the economic benefits of the photovoltaic energy storage system while absorbing photovoltaic power generation and meeting the charging needs of electric vehicles. The photovoltaic energy storage system can generate economic benefits through charging electric vehicles and trading with the grid, while also considering the system's lifespan. The objective function includes charging revenue, grid trading revenue, and the cost of energy storage capacity degradation.
[0117]
[0118] I ev,t =λev P ev,t Δt (13)
[0119]
[0120] Among them, I ev,t Let I represent the charging revenue of the photovoltaic system that provides electricity to electric vehicles during the t-th time period. grid,t Let C be the revenue generated from the transaction between the photovoltaic system and the power grid during the t-th time period. b,t P represents the energy storage capacity degradation cost of the photovoltaic system in time period t. T0 is the number of periods in a cycle. This example uses one day as a cycle, divided into 96 periods, with each period t being 15 minutes. ev,t λ is the total power demand for electric vehicle charging via photovoltaics in the t-th time period; ev This refers to the price of charging electric vehicles. Δt is a unit of time. P grid,t ρ represents the electricity traded between the photovoltaic charging station and the power grid during the t-th time period. Its value can be positive or negative, indicating whether the photovoltaic energy storage system sells electricity to the grid or purchases electricity from the grid. stg,t and ρ bfg,t These represent the prices at which the photovoltaic energy storage system sells electricity to the grid and purchases electricity during the t-th time period, respectively.
[0121] (2) Establish constraints
[0122] 1) Power balance constraints
[0123] P s,t +P es,t =P ev,t +P grid,t (15)
[0124] In the formula, P s,t P represents the photovoltaic power in the t-th time period; es,t Let be the discharge power in the t-th time period.
[0125] 2) Energy storage constraints
[0126] -P es,max ≤P es,t ≤P es,max (16)
[0127]
[0128]
[0129] To ensure that the energy storage operates normally in the next control cycle, the SOC of the energy storage action at the end of the cycle should be a fixed value, as shown in formula (19).
[0130]
[0131] P es,max It is the maximum value of the energy storage output power; Let E be the minimum and maximum SOC of energy storage in the t-th time period. ini This is the maximum capacity of the energy storage battery before it degrades. and These are the charging efficiency and discharging efficiency of the energy storage battery in the t-th time period, respectively.
[0132] Furthermore, a deep reinforcement learning model is established.
[0133] (1) State vector
[0134] S = [P] s,t ,P ev,t ,ρ g,t ,S SOC (20)
[0135] At any given time interval t, the state space consists of photovoltaic power generation, total electric vehicle charging demand, electricity price, and energy storage SOC, ρ g,t It is the real-time electricity price for the t-th time period.
[0136] (2) Action Vector
[0137] The energy storage output is simulated as an action vector, and the range of the action vector is limited by the constraint equations. The control problem of this invention is characterized by a continuous action space, which is more suitable for controlling energy storage batteries.
[0138] a = [a t ]
[0139]
[0140] (3) Reward function
[0141] The reward-penalty function determines the immediate reward of the environment for energy storage charging and discharging behavior within a certain period, thus influencing the agent's choice of behavior. The reward-penalty function consists of revenue rewards and energy storage capacity decay costs.
[0142] The revenue reward corresponds to the charging revenue and grid trading revenue obtained by photovoltaic power generation in the objective function equations (13) and (14), as shown in equation (22).
[0143] r t =(I ev,t +I grid,t -C b,t ) (twenty two)
[0144] Furthermore, a proximal policy optimization method (PPO) is proposed.
[0145] 1) Define the advantage evaluation function A π,γ—Used to measure how much better an action is on average than other actions.
[0146] A π,γ (s t ,a t )=Q π,γ (s t ,a t )-V π,γ (s t ) (twenty three)
[0147]
[0148]
[0149] Among them, A π,γ V π,γ and Q π,γ These are the advantage function, the value function (also known as the critique function), and the action value function, respectively. π θ The strategy representing the agent's choice of action is θ, and γ represents the discount factor, which determines the impact of future rewards on cumulative rewards.
[0150] The reinforcement learning agent determines the current state s based on the current state. t Select action a from the action space. t And evaluate the merits (value) of the current action based on the advantage evaluation function.
[0151] 2) Proximal Policy Optimization (PPO): By employing importance sampling techniques, PPO utilizes an innovative policy gradient expression. This expression allows for multiple updates to the policy network after collecting the trajectory set, significantly improving PPO's sampling efficiency and training stability.
[0152]
[0153]
[0154]
[0155] ∈ is a hyperparameter that controls the allowable policy deviation; ρ t It is the ratio coefficient between the updated strategy and the old strategy. Specifically, ρ t The further the value deviates from 1, the further the updated policy deviates from the original policy. Equation (26) can prevent drastic changes in the policy (actor) network, thereby degrading the performance of the PPO agent. Equation (28) is the update function for the θ parameters of the policy network.
[0156] Furthermore, a stochastic diagonal Gaussian strategy with a continuous action space is proposed.
[0157] PPO agents employ a randomized diagonal Gaussian strategy to achieve a continuous action space. Specifically: assuming an actor network π... θ The output is a mean vector of actions, which follows a multivariate normal distribution with a diagonal covariance matrix. The diagonal elements of the covariance matrix are the variances of each action. When the PPO agent attempts to determine an action based on observations, it depends on sampling actions from a multivariate normal distribution:
[0158] a = μ θ (s)+σ θ (s)⊙x (29)
[0159] Where x is the sample vector of the standard multivariate normal distribution; μ and σ are the mean and standard deviation of the action vector; and ⊙ is the product of elements. The exploration of the PPO agent is achieved by sampling the Gaussian distribution in (29).
[0160] According to the random diagonal Gaussian policy, π θ (a t |s t This can be deduced as
[0161]
Claims
1. A method for optimizing the entire lifecycle operation of a photovoltaic energy storage system based on deep reinforcement learning, used for real-time optimization control of charging and discharging in a photovoltaic energy storage system for electric vehicles, characterized in that... Includes the following steps: S1. Perform detailed modeling of the energy storage operation efficiency model and capacity decay model; S2. Establish an optimized operation strategy for photovoltaic energy storage systems based on deep reinforcement learning. To maximize the revenue of the photovoltaic energy storage system, an objective function for optimizing the operation of the photovoltaic energy storage system is established. The variables of the objective function include charging revenue, grid transaction revenue, and energy storage capacity decay cost. A deep reinforcement learning model is established to determine the state vector using photovoltaic power generation, total electric vehicle charging demand, electricity price, and energy storage SOC for any given time period. The reward function of the deep reinforcement learning model consists of revenue reward and energy storage capacity decay cost. S3. Solve the model using deep reinforcement learning based on proximal policy optimization. Reinforcement learning based on the current state s t Select action a from the action space. t And evaluate the quality of the current action based on the advantage evaluation function; The action space satisfies the following formula: a=[a t ] In the formula, P es,t Let P be the energy storage capacity in the t-th time period. es,max It is the maximum value of the energy storage output power; The advantage evaluation function A π,γ The formula is as follows: A π,γ (s t ,a t )=Q π,γ (s t ,a t )-V π,γ (s t ), In the formula, A π,γ V π,γ and Q π,γ These are the advantage function, value function, and action value function, respectively; π θ The strategy for choosing an action by the actor is represented by θ; γ represents the discount factor, which determines the impact of future rewards on cumulative rewards. The network is updated using the following strategy: In the above formula, ∈ is a hyperparameter controlling the allowable policy deviation; ρ t It is the ratio coefficient between the updated strategy and the old strategy; And a random diagonal Gaussian strategy is used to realize the continuous action space: a=μ θ (s)+σ θ (s)⊙x, Where x is a sample vector of a standard multivariate normal distribution; μ and σ are the mean and standard deviation of the action vector; ⊙ is the product of elements; According to the random diagonal Gaussian policy, π θ (a t |s t This can be deduced as:
2. The method for optimizing the operation of a photovoltaic energy storage system throughout its entire lifecycle according to claim 1, characterized in that, The detailed modeling of the energy storage operation efficiency model in step S1 includes the following steps: The charge-discharge efficiency of a battery is calculated using an equivalent model of its steady-state circuit. The steady-state circuit of a single battery includes the open-circuit voltage V. oc The open-circuit voltage is represented by the series resistor R1 connected in series with the open-circuit voltage, the short-time response resistor R2 related to the transient response, and the long-time response resistor R3. Each variable and battery S SOC The charge state exhibits a nonlinear relationship, as shown in the following equation: Among them, I batt P is the battery current. out R is the output power; oc For terminating resistance; g0-g5, b0-b5, c0-c2, and d0-d2 are all coefficients; For a specific SOC and P out The current flowing through a single cell is calculated using the following formula: The charge / discharge efficiency of a single battery is as follows: An energy storage system consists of multiple individual battery cells connected in series and parallel. The calculation methods for equivalent current and charge / discharge efficiency are as follows: Among them, I eq I is the current of the energy storage system; I is the equivalent current of multiple series- and parallel-connected individual cells; N s N is the number of batteries connected in series; p It refers to the number of batteries connected in parallel; Among them, e0-e5 and h0-h5 are coefficients.
3. The method for optimizing the operation of a photovoltaic energy storage system throughout its entire lifecycle according to claim 2, characterized in that, In step S1, the energy storage capacity decay model is modeled using a semi-empirical battery capacity decay model. The battery capacity decay model over N cycles is shown in the following equation: Among them, E loss E′ loss The decay capacity before and after N cycles, E′, are respectively. loss =0 indicates that the battery has not been charged or discharged; α s β s f is the influence coefficient of the solid electrolyte interface formation in the new battery; d The battery capacity decay function for a single cycle can be expressed as a function of temperature T, δ, SOC, and time t1, as shown in the following equation: f d =(f δ (δ)+f t (t1))f SOC (S SOC )f T (T), Among them, f δ (δ), f t (t1), f SOC (S SOC ), f T (T) is the coefficient representing the influence of DOD, time, SOC, and temperature on the lifespan of the energy storage battery. ref This is a reference temperature; The remaining usable capacity E of the energy storage battery re for: AND re =And ini -AND 1oss , In the formula, E ini It is the maximum capacity of the energy storage battery before it degrades.
4. The method for optimizing the operation of a photovoltaic energy storage system throughout its entire lifecycle according to claim 3, characterized in that, The objective function for optimizing the operation of the photovoltaic energy storage system in step S2 is as follows: I ev,t =λ ev P ev,t Δt, Among them, I ev,t I represents the charging revenue of the photovoltaic system that provides electricity to electric vehicles during the t-th time period. grid,t Let C be the revenue generated from the transaction between the photovoltaic system and the power grid during the t-th time period. b,t P is the energy storage capacity degradation cost of photovoltaic power in the t-th time period; T0 is the number of periods in a cycle; ev,t λ is the total power demand for electric vehicle charging via photovoltaics in the t-th time period; ev This refers to the price of charging electric vehicles; Δt is a unit of time; P grid,t ρ is the electricity traded between the photovoltaic charging station and the power grid during the t-th time period; stg,t and ρ bfg,t These represent the prices at which the photovoltaic energy storage system sells electricity to the grid and purchases electricity during the t-th time period, respectively.
5. The method for optimizing the operation of a photovoltaic energy storage system throughout its entire lifecycle according to claim 4, characterized in that, In step S2, the reward function consists of revenue reward and energy storage capacity decay cost, as shown in the following formula: r t =(I ev,t +I grid,t -C b,t )。 6. The method for optimizing the operation of a photovoltaic energy storage system throughout its entire lifecycle according to claim 5, characterized in that, The optimized operation strategy of the photovoltaic energy storage system in step S2 satisfies the following constraints: 1) Power balance constraints P s,t +P es,t =P ev,t +P grid,t , In the formula, P s,t P represents the photovoltaic power in the t-th time period; es,t P represents the energy storage capacity during the t-th time period. ev,t P is the total power demand for electric vehicle charging via photovoltaics in the t-th time period; grid,t It represents the electricity traded between the photovoltaic charging station and the power grid during the t-th time period; 2) Energy storage constraints To ensure that the energy storage operates normally in the next control cycle, the SOC of the energy storage operation at the end of the cycle should be a fixed value, as shown in the following formula: In the above formula, P es,max It is the maximum value of the energy storage output power; Let SOC and SOC be the minimum and maximum values of the stored energy in the t-th time period, respectively. P represents the remaining capacity of the energy storage battery in the t-th time period; es,t E represents the discharge power of the energy storage battery during the t-th time period. ini This is the maximum capacity of the energy storage battery before it degrades. and These are the charging efficiency and discharging efficiency of the energy storage battery in the t-th time period, respectively.
Citation Information
Patent Citations
Optimization method of optical storage charging station system based on reinforcement learning and terminal
CN117993647A
Hydrogen energy storage unit power distribution method based on multi-agent deep reinforcement learning
CN118589501A