Operation method of photovoltaic storage and charging integrated station for power grid power balance and new energy consumption

By using the DDPG algorithm to train a reinforcement learning network in an integrated photovoltaic-storage-charging station, the problems of energy storage system output coordination and new energy consumption were solved, achieving optimized operation of grid power balance and new energy consumption, and improving dispatch economy.

CN115940289BActive Publication Date: 2025-11-18BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211625303.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2025-11-18
Estimated Expiration
2042-12-16

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively coordinate the output of energy storage systems in integrated photovoltaic-storage-charging stations, and cannot effectively solve the problem of local consumption of new energy. Furthermore, traditional reinforcement learning methods suffer from the curse of dimensionality when dealing with continuous state variables, making them ineffective in solving problems.

Method used

The reinforcement learning network is trained using the Deep Deterministic Policy Gradient (DDPG) algorithm and combined with the Markov Decision Process method. The joint operation optimization model of the photovoltaic-storage-charging integrated station is transformed into a policy decision problem. The system is trained offline using historical system state data to achieve real-time self-optimizing operation of the energy storage system.

Benefits of technology

It improves the economic efficiency of the scheduling and operation of integrated photovoltaic, energy storage and charging stations, realizes the local consumption of new energy resources, avoids the dimensionality disaster and suboptimal scheduling strategies, and achieves optimized operation of power grid balance and new energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115940289B_ABST
    Figure CN115940289B_ABST
Patent Text Reader

Abstract

The application provides a kind of photovoltaic-storage-charging integrated station operation method for power grid power balance and new energy consumption, comprising: based on the comprehensive service scene of photovoltaic-storage-charging integrated station, with the minimum operation cost of photovoltaic-storage-charging integrated station as the target, with power balance constraint, main grid interactive power constraint and equipment operation constraint as constraint condition, with the real-time output of energy storage unit as decision variable, the joint operation optimization model of photovoltaic-storage-charging integrated station is established;Using Markov decision process method to transform into the reinforcement learning network of strategy decision problem;Using historical system state data, based on DDPG algorithm, the reinforcement learning network is trained offline, the joint operation optimization model of photovoltaic-storage-charging integrated station is scheduled with the lowest economic optimization calculation of scheduling cost, the real-time self-optimizing operation output result of energy storage system in photovoltaic-storage-charging integrated station is obtained, and the dynamic optimization operation of photovoltaic-storage-charging integrated station is realized, and dimension disaster in calculation can also be effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electric vehicle charging station operation and scheduling technology, and in particular to an operation method for an integrated photovoltaic-storage-charging station oriented towards grid power balance and new energy consumption. Background Technology

[0002] In the context of transportation electrification, the planning and design of charging facilities and power distribution networks need to adapt to the large-scale development trend of electric vehicles. Meanwhile, energy storage and photovoltaics have received widespread attention due to improved technical performance and reduced costs. Photovoltaic power generation has the advantages of low cost and cleanliness, but it is greatly affected by the external environment, resulting in some fluctuations in output. Configuring energy storage systems can further enhance the local compensation effect for electric vehicle charging loads. By managing the charging and discharging behavior of energy storage batteries, energy can be shifted in time and space, alleviating the power supply pressure on the grid during peak hours. Therefore, a diversified application model that organically combines photovoltaics, energy storage, and electric vehicles is an important way to connect electric vehicle charging stations with renewable energy sources, effectively reducing the impact of electric vehicle charging behavior on the power grid.

[0003] Current research on the optimal scheduling of integrated photovoltaic (PV) and energy storage power plants is relatively limited, primarily relying on various control algorithms based on traditional optimization modeling to address the coordination and complementarity of resources within the plant. These include, for example, an adaptive robust day-ahead energy-reserve collaborative optimization scheduling method for PV-storage charging towers aiming to minimize total daily operating costs, and a four-stage intelligent optimization control algorithm integrating PV power generation and stationary battery storage bidirectional charging stations for electric vehicles with commercial buildings. These algorithms aim to minimize operating costs related to customer satisfaction while considering potential uncertainties, and balance real-time supply and demand among the "source-storage-load" system by adjusting the schedule. However, their drawbacks include the fact that optimization models are often constructed with the goal of maximizing charging station revenue. For integrated PV-storage charging and discharging stations, the local consumption of renewable energy must be considered in addition to the economic operation within the station. Furthermore, current optimization scheduling methods are mostly focused on day-ahead scheduling, thus limiting them to fixed scheduling plans and failing to dynamically respond to random changes in sources and loads. Additionally, existing optimization operation models are largely based on traditional mathematical optimization modeling, which still relies on accurate predictions of renewable energy and load. With the promotion of electricity consumption information collection systems, the application of data-driven machine learning methods in the optimization of power system operation has attracted widespread attention from scholars at home and abroad. To optimize the output plans of various resources within a power station and reduce operating costs, existing reinforcement learning algorithms describe the scheduling problem of various resources within the station as a constrained Markov decision process, and propose a model-free method based on deep reinforcement learning. This method generates a constrained optimal output plan for resources within the station by directly learning from deep neural networks. However, while this traditional reinforcement learning method performs well in handling small-scale discrete space problems, when dealing with continuous state variable tasks, the number of discretized states increases exponentially with the increase of spatial dimension, resulting in the curse of dimensionality and hindering effective learning. In the joint economic operation problem of integrated photovoltaic-storage-charging power stations, since the load, photovoltaic power generation, and state of charge in its state space are all continuous variables, traditional reinforcement learning methods often cannot solve the problem effectively. Simultaneously, the actions of the energy storage system within the station are also continuous variables; discretizing the action space will obscure much information in the decision-making action domain.

[0004] Therefore, for integrated photovoltaic-storage-charging stations, while considering the economic operation within the station, it is still necessary to further consider the local consumption of new energy. How to coordinate the output of all energy storage systems within the station, effectively avoid the dimensionality disaster and preserve the information of the entire action domain, and thus achieve optimized operation of photovoltaic-storage-charging stations for grid power balance and new energy consumption, based on a comprehensive consideration of the integrated service scenario of electric vehicle charging stations that integrate photovoltaic and energy storage systems, is a problem that existing technologies urgently need to solve. Summary of the Invention

[0005] This invention provides an operation method for integrated photovoltaic-storage-charging stations that addresses grid power balance and renewable energy consumption, thereby solving the problems existing in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution.

[0007] This invention provides an operational method for an integrated photovoltaic-storage-charging station oriented towards grid power balance and renewable energy consumption, comprising:

[0008] S1 is based on the integrated service scenario of photovoltaic-storage-charging integrated station. With the goal of minimizing the operating cost of photovoltaic-storage-charging integrated station, and with power balance constraints, main grid interaction power constraints and equipment operation constraints as constraints, and real-time output of energy storage units as decision variables, a joint operation optimization model for photovoltaic-storage-charging integrated station is established.

[0009] S2 transforms the dynamic scheduling problem in the joint operation optimization model into a reinforcement learning network for a policy decision problem using the Markov decision process method.

[0010] S3 uses historical system state data and trains the reinforcement learning network offline based on the Deep Deterministic Policy Gradient (DDPG) algorithm to obtain the trained reinforcement learning network.

[0011] S4 uses the trained reinforcement learning network and the time-of-use pricing mechanism to perform the minimum scheduling cost economic optimization calculation on the joint operation optimization model of the photovoltaic-storage-charging integrated station, and obtains the real-time self-optimizing operation output of the energy storage system in the photovoltaic-storage-charging integrated station, thereby achieving the optimal operation of the photovoltaic-storage-charging integrated station.

[0012] Preferably, the joint operation optimization model includes:

[0013] The objective function is shown in equation (1) below:

[0014] F = min(C) E +C BES (1)

[0015] Where F represents the operating cost of the integrated photovoltaic, energy storage, and charging station, and C... E The cost of purchasing electricity from the grid for integrated photovoltaic, energy storage, and charging stations. P grid (t) represents the power exchanged between the system and the main grid during time period t. A positive value indicates that the system purchases electricity from the main grid, while a negative value indicates that the system sells surplus electricity back to the grid. e (t) represents the electricity price for time period t; Δt represents the length of the time interval; C BES The depreciation cost of charging and discharging electrical energy storage. P BES (t) represents the charging or discharging power of the energy storage during time period t. A positive value indicates that the energy storage is in a discharging state, and a negative value indicates that it is in a charging state; ρ BES This is the depreciation cost coefficient for energy storage.

[0016] The constraints are as follows:

[0017] 1) Power balance constraints:

[0018] In time period t, the power balance constraint is as shown in equation (2):

[0019] P grid (t)+P pv (t)+P BES (t)=P load (t) (2)

[0020] Among them, P pv (t) represents the photovoltaic power generation; P load (t) represents the user's electricity load demand during time period t;

[0021] 2) The power constraint of the main power grid interaction is shown in equation (3) below:

[0022]

[0023] in, and These are the lower and upper limits of the power exchange between the system and the main power grid, respectively.

[0024] 3) The equipment operation constraints are shown in equations (2)-(3) below:

[0025]

[0026] in, and These are the lower and upper limits of the charging / discharging power of electrical energy storage, respectively.

[0027] For energy storage devices, the constraints are as shown in equations (3)-(4):

[0028]

[0029] in, and These are the lower and upper limits of the state of charge of electrical energy storage, respectively; C SOC (t) represents the state of charge of the energy storage during time period t; Q BES The capacity of electrical energy storage; The initial state of charge of the electrical energy storage; η BES The charge / discharge coefficient for electrical energy storage; η ch and η dis These represent the charging efficiency and discharging efficiency of electrical energy storage, respectively.

[0030] Preferably, step S2 includes:

[0031] The problem of minimizing the operating cost of a photovoltaic-storage-charging integrated power station is transformed into a reward maximization problem for the intelligent agent, as shown in equation (5):

[0032]

[0033] Where, r t (s t a t ) represents the total reward value obtained by the agent during the scheduling period t, s t For the observation status of the photovoltaic-storage-charging integrated power station during the dispatch period t, s t ={P load (t),P pv (t),C soc (t-1),t},P load (t), P pv (t), C soc (t-1), where t represents the user's electricity load demand, photovoltaic power generation, energy storage state of charge, and the current dispatch period, respectively; a t For the dynamic economic dispatching of the energy storage system in a photovoltaic-storage-charging integrated power station, the actions within the integrated power station during time period t can be determined by the power output P of the equipment. BES (t) indicates that a t ={P BES (t)}; It is the scaling factor for the cost value, (C) E (s t a t ) represents the agent in state s t Below, the cost of purchasing electricity from the grid for the photovoltaic-storage-charging integrated station during time period t; C BES (s t a t ) represents the agent in state s i The charging and discharging depreciation cost of energy storage during time period t;

[0034] Use the action-value function Q of equation (6) below. π (s,a) The energy storage system of the photovoltaic-storage-charging integrated power station is in state s l Dynamic economic scheduling action a l The action-value function Q is evaluated. π The larger (s,a) is, the greater a is. l The better:

[0035]

[0036] Among them, E π (.) represents the expectation under the optimal objective policy π; γ k ∈[0,1], where γ is the discount factor, representing the proportion of the reward at a future point in the cumulative reward. k The larger the size, the more emphasis is placed on future rewards; t+k s represents the total reward value obtained by the agent in time period t+k; t+k For the integrated photovoltaic, energy storage, and charging station, the state during time period t+k; a t+k For the actions performed by the integrated photovoltaic, energy storage, and charging station during time period t+k, k∈N * , indicating the generation of the agent's iterative learning;

[0037] The optimal objective policy π is obtained according to the following equation (7) to maximize the action-value function:

[0038]

[0039] Where A is the set of actions of the agent.

[0040] Preferably, step S3 includes:

[0041] The historical system status data refers to the observed status of the photovoltaic-storage-charging integrated power station, including the system power load demand, photovoltaic power generation, energy storage charge status, and scheduling period of the integrated photovoltaic-storage-charging power station.

[0042] The reinforcement learning network includes a value network and a policy network, through which the policy network π(s|θ) π Sum-value network Q(s,a|θ) Q Create two independent target networks π'(s|θ) respectively. π' ) and Q'(s,a|θ Q' As shown in equations (8) and (9):

[0043]

[0044] Set an optimization cycle T, input historical system state data into the policy network and the target network of the value network. After training a batch of data, the DDPG algorithm updates the parameters of the current network of the policy network and the value network through gradient ascent or gradient descent, and then updates the parameters of the target network of the policy network and the value network through a soft update method. After T iterations, the offline learning of the DDPG algorithm is completed, and the trained reinforcement learning network is obtained.

[0045] Preferably, step S4 includes:

[0046] When a scheduling task is received, in each time period, based on the current system state s t Using a trained reinforcement learning network to select and schedule action a t ;

[0047] Perform action a t And enter the next environmental state, while receiving the reward r. t ;

[0048] Then, the status information s within the integrated station during the time period t+1 is collected. t+1 The time-of-use electricity price information is used as a new sample, and dynamic scheduling decisions are made for that period, which is the real-time self-optimizing operation output result of the energy storage system in the integrated photovoltaic-storage-charging station.

[0049] Preferably, when training the reinforcement learning network offline, the action a is based on the following formula (10). t Training the reinforcement learning network:

[0050] a t =π(s) t |θ π )+v t (10)

[0051] Among them, v t It is random noise.

[0052] As can be seen from the technical solution provided by the above-mentioned photovoltaic-storage-charging integrated station operation method for grid power balance and new energy consumption of the present invention, the present invention improves the economy of photovoltaic-storage-charging integrated station scheduling operation by converting the mathematical model into a reinforcement learning network that can be solved by reinforcement learning algorithm, by inputting historical data and training with DDPG algorithm, and by inputting real-time parameters within the optimization period based on the trained reinforcement learning network, and realizes the local consumption of new energy resources. At the same time, the use of DDPG algorithm also effectively avoids the dimensionality curse and suboptimal scheduling strategy selection problem in the discretization process.

[0053] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1This is a schematic diagram of the operation method of an integrated photovoltaic-storage-charging station for grid power balance and new energy consumption provided by an embodiment of the present invention;

[0056] Figure 2 This is a schematic diagram of the operation framework of an integrated photovoltaic-storage-charging station for grid power balance and new energy consumption, provided by an embodiment of the present invention.

[0057] Figure 3 This is a schematic diagram of the reinforcement learning process in this embodiment;

[0058] Figure 4 This is a schematic diagram of the reinforcement learning process in this embodiment;

[0059] Figure 5 This is a schematic diagram illustrating historical data statistics for an example.

[0060] Figure 6 This is a time-of-use electricity pricing information chart;

[0061] Figure 7 A trend chart of DDPG algorithm training results;

[0062] Figure 8 This is a diagram showing the scheduling results of the energy storage system.

[0063] Figure 9 This is a schematic diagram showing the power exchange between the integrated substation and the main power grid. Detailed Implementation

[0064] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0065] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0066] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0067] To facilitate understanding of the embodiments of the present invention, further explanations and descriptions will be provided below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0068] Example

[0069] This invention provides an operational method for integrated photovoltaic-storage-charging stations aimed at grid power balance and renewable energy consumption, such as... Figure 1 and Figure 2 As shown, the specific steps include the following:

[0070] S1 is based on the integrated service scenario of photovoltaic-storage-charging stations. With the goal of minimizing the operating cost of photovoltaic-storage-charging stations, and with power balance constraints, main grid interaction power constraints, and equipment operation constraints as constraints, and with the real-time output of energy storage units as decision variables, a joint operation optimization model for photovoltaic-storage-charging stations is established.

[0071] The joint operation optimization model includes:

[0072] The objective of the economic dispatch problem for integrated photovoltaic-storage-charging stations is to minimize the station's operating costs, which include the cost of purchasing electricity from the grid and the depreciation costs of charging and discharging energy storage. The objective function is shown in equation (1) below:

[0073] F = min(C) E +C BES (1)

[0074] Where F represents the operating cost of the integrated photovoltaic, energy storage, and charging station, and C... E The cost of purchasing electricity from the grid for integrated photovoltaic, energy storage, and charging stations. P grid (t) represents the power exchanged between the system and the main grid during time period t. A positive value indicates that the system purchases electricity from the main grid, while a negative value indicates that the system sells surplus electricity back to the grid. e (t) represents the electricity price for time period t; Δt represents the length of the time interval; C BES The depreciation cost of charging and discharging electrical energy storage. P BES (t) represents the charging or discharging power of the energy storage during time period t. A positive value indicates that the energy storage is in a discharging state, and a negative value indicates that it is in a charging state; ρ BES This is the depreciation cost coefficient for energy storage.

[0075] To ensure that the output power of the integrated photovoltaic-storage-charging system within the station meets the equipment's operating conditions and achieves real-time self-optimization as much as possible, the following constraints need to be considered: power balance constraints, main grid interaction power constraints, and equipment operation constraints. Specific constraints are shown below:

[0076] Power balance constraints:

[0077] In time period t, the power balance constraint is as shown in equation (2):

[0078] P grid (t)+P pv (t)+P BES (t)=P load (t) (2)

[0079] Among them, P pv (t) represents the photovoltaic power generation; P load (t) represents the user's electricity load demand during time period t;

[0080] Considering the operational stability of the power grid, the main grid imposes upper and lower limits on the power interaction of the photovoltaic-storage-charging integrated station. The main grid interaction power constraints are shown in equation (3) below:

[0081]

[0082] in, and These are the lower and upper limits of the power exchange between the system and the main power grid, respectively.

[0083] Each piece of equipment in the integrated photovoltaic, energy storage, and charging station has an upper and lower limit range for operation, and the equipment operation constraints are shown in the following formula (4):

[0084]

[0085] in, and These are the lower and upper limits of the charging / discharging power of electrical energy storage, respectively.

[0086] For energy storage devices, it is also necessary to avoid damage to the energy storage from deep charging and discharging. Therefore, the state of charge (SOC) of the energy storage is limited to a certain range. In addition, in order to ensure the continuous and stable operation of the energy storage, the energy storage capacity at the beginning and end of a scheduling cycle is required to be equal, as shown in the following equations (5)-(6):

[0087]

[0088] in, and These are the lower and upper limits of the state of charge of electrical energy storage, respectively; CSOC (t) represents the state of charge of the energy storage during time period t; Q BES The capacity of electrical energy storage; The initial state of charge of the electrical energy storage; η BES The charge / discharge coefficient for electrical energy storage; η ch and η dis These represent the charging efficiency and discharging efficiency of electrical energy storage, respectively.

[0089] By modeling the optimized operation of the photovoltaic-storage-charging integrated station, the coordination and scheduling relationships among the various devices within the integrated station can be determined, and the system power constraints and equipment operation constraints can be satisfied.

[0090] S2 transforms the dynamic scheduling problem in the joint operation optimization model into a reinforcement learning network for the policy decision problem using the Markov decision process method.

[0091] The joint operation optimization model is shown in equation (7) below:

[0092]

[0093] The monitoring status of an integrated photovoltaic-storage-charging power station includes user electricity load demand, photovoltaic power generation, energy storage state of charge, and the current dispatch period. For this integrated station, its status is represented as: s t ={P load (t),P pv (t),C soc (t-1),t},P load (t), P pv (t), C soc (t-1), where t represents the user's electricity load demand, photovoltaic power generation, energy storage charge status, and the current dispatch period, respectively.

[0094] During time period t, the actions in the integrated photovoltaic, energy storage, and charging station can be represented by the power output of the equipment, and its actions can be represented by P. BES (t) represents:

[0095] a t ={P BES (t)} (8)

[0096] The goal of optimizing the operation of a photovoltaic-storage-charging integrated power station is to minimize the operating cost of the power station. The problem of minimizing the operating cost of a photovoltaic-storage-charging integrated power station can be transformed into a reward maximization problem for the agent, as shown in equation (9):

[0097]

[0098] Where, r t (s ta t ) represents the total reward value obtained by the agent during the scheduling period t, s t The observation status of the integrated photovoltaic, energy storage, and charging station during scheduling period t; a t The dynamic economic dispatching action of the energy storage system in a photovoltaic-storage-charging integrated power station; It is the scaling factor for the cost value, (C) E (s t ,a t ) represents the agent in state s t Below, the cost of purchasing electricity from the grid for the photovoltaic-storage-charging integrated station during time period t; C BES (s t a t ) represents the agent in state S t The depreciation cost of charging and discharging of electrical energy storage during time period t.

[0099] During the operation of the integrated photovoltaic, energy storage, and charging station, a certain state s t When determined, the action-value function Q of equation (10) is used. π (s,a) The energy storage system of the photovoltaic-storage-charging integrated station is in state s l Dynamic economic scheduling action a l The action-value function Q is evaluated. π The larger (s,a) is, the greater a is. l The better:

[0100]

[0101] Among them, E π (.) represents the expectation under the optimal objective policy π; γ k ∈[0,1], where γ is the discount factor, representing the proportion of the reward at a future point in the cumulative reward. k The larger the size, the more emphasis is placed on future rewards; t+k s represents the total reward value obtained by the agent in time period t+k; t+k For the integrated photovoltaic, energy storage, and charging station, the state during time period t+k; a t+k For the actions performed by the integrated photovoltaic, energy storage, and charging station during time period t+k, k∈N * , indicating the generation in which the agent learns cyclically.

[0102] The goal of the joint operation optimization model of the photovoltaic-storage-charging integrated station is to find the optimal strategy π to maximize the action-value function. The optimal target strategy π is obtained according to the following equation (11):

[0103]

[0104] Where A is the set of actions of the agent.

[0105] S3 uses historical system state data and offline trains a reinforcement learning network based on the Deep Deterministic Policy Gradient (DDPG) algorithm to obtain a trained reinforcement learning network. Specifically, as follows... Figure 3 A diagram illustrating the reinforcement learning process and Figure 4 The reinforcement learning process is illustrated in the diagram.

[0106] Historical system status data represents the observed status of the integrated photovoltaic-storage-charging power station, including the system's electrical load demand, photovoltaic power generation, energy storage charge status, and scheduling periods.

[0107] The basic components of training include a set of states S representing the environment, a set of actions A representing the agent's actions, and a reward r for the agent. During time interval t, the environment provides the agent with observed states s. t ∈S, the agent is based on policy π* and the state s of the integrated station. t Generate action state a t .

[0108] Because reinforcement learning data exhibits the Markov property, it does not satisfy the assumption that training neural networks requires samples to be independent and identically distributed. To ensure learning effectiveness, when generating sample data, DDPG stores the data explored in the environment in the replay pool R. Each time it is updated, the value network and policy network randomly select a portion of the samples from it for optimization to reduce instability.

[0109] Reinforcement learning networks consist of value networks and policy networks, with the policy network π(s|θ) π Sum-value network Q(s,a|θ) Q Create two independent target networks π'(s|θ) respectively. π' ) and Q'(s,a|θ Q' As shown in equations (12) and (13):

[0110]

[0111] The DDPG algorithm sets an optimization cycle T, inputs historical system state data into the target network of the policy network and value network, and after training a mini-batch of data, updates the parameters of the online network of the policy network and value network using gradient ascent or gradient descent. Then, it updates the parameters of the target network of the policy network and value network using a soft update method. After T iterations, the offline learning of the DDPG algorithm is completed, and the trained DDPG network is obtained.

[0112] The specific training content is as follows:

[0113] 1) Value network training

[0114] In a value network, by minimizing the loss function L(θ) Q To optimize the parameters, as shown in equation (14):

[0115] L(θ Q )=E(y t -Q(s t ,a t |θ Q ) 2 (14)

[0116] Where, θ Q The parameter y is the current network parameter in the value network; t Let Q be the target Q value; E(.) be the expected function.

[0117] y t =r t +γQ'(s t+1 ,π'(s t+1 |θ π' )|θ Q' (15)

[0118] Where, r t γ is the total reward value obtained by the agent in time period t; γ∈[0,1] is the discount factor; Q' is the target Q value before the update; π' is the target policy; θ π' θ represents the parameters of the target network in the policy network. Q' These are the parameters of the target network in the value network.

[0119] During time period t, the integrated photovoltaic, energy storage, and charging station performs action a. t Then it will enter the next state s t+1 This refers to the updated state of charge of the energy storage, the electrical load observed over a period of time, and the photovoltaic power generation.

[0120] L(θ Q Regarding θ Q The gradient is given by the following equation (16):

[0121]

[0122] Among them, y t -Q(s t ,a t |θ Q This refers to the temporal difference error. Based on the gradient rule, the network is updated using the following update formula:

[0123]

[0124] Where, μ Q The value is the network learning rate.

[0125] 2) Policy Network Training

[0126] In the policy network, it provides gradient information. This serves as a direction for action improvement. To update the policy network, the sampled policy gradient is used as follows:

[0127]

[0128] Update the policy network parameters θ based on the deterministic policy gradient. π :

[0129]

[0130] Where, μ π is the learning rate of the policy network.

[0131] Furthermore,

[0132] θ Q' ←τθ Q +(1-τ)θ Q' (20)

[0133] θ π' ←τθ π +(1-τ)θ π' (twenty one)

[0134] Where τ is the soft update coefficient, and τ << 1.

[0135] S4 uses a trained reinforcement learning network and a time-of-use pricing mechanism to perform a minimum scheduling cost optimization calculation on the joint operation optimization model of the photovoltaic-storage-charging integrated station. This results in the real-time self-optimizing power output of the energy storage system within the integrated station, thereby achieving optimal operation of the integrated photovoltaic-storage-charging station.

[0136] When a scheduling task is received, in each time period, based on the current system state s t Using a trained reinforcement learning network to select and schedule action a t ;

[0137] Perform action a t And enter the next environmental state, while receiving the reward r. t ;

[0138] Then, the status information s within the integrated station during the time period t+1 is collected. t+1 The time-of-use electricity price information is used as a new sample, and dynamic scheduling decisions are made for that period, which is the real-time self-optimizing operation output result of the energy storage system in the integrated photovoltaic-storage-charging station.

[0139] Preferably, when training the reinforcement learning network offline, the action a is based on the following formula (18). t Training the reinforcement learning network:

[0140] a t =π(s) t |θ π )+v t (twenty two)

[0141] Among them, v t It is random noise. Through action a t ={P BES (t)} Add random noise v t This aims to enhance the DDPG algorithm's ability to explore the environment during interactions with integrated photovoltaic, energy storage, and charging stations, thereby learning more optimized dynamic scheduling strategies.

[0142] The following is a specific example of the operation method of the photovoltaic-storage-charging integrated station for grid power balance and new energy consumption in this embodiment:

[0143] Figure 5 The document provides statistical charts of some historical monitoring sample data from February 1, 2021, specifically including user-required electricity load data and photovoltaic power generation data for that day. Electricity pricing adopts a time-of-use pricing mechanism, such as... Figure 6 As shown, the peak period is 12:00-19:00, the normal period is 07:00-12:00 and 19:00-23:00, and the valley period is 23:00-07:00. Based on existing historical data of photovoltaic output and electricity load, combined with time-of-use pricing information, the scheduling cycle length of the photovoltaic-storage-charging integrated station is set to 24 hours, with an interval of 15 minutes between two adjacent time periods. After training the agent with the DDPG algorithm for 5000 episodes and then converging, the optimal operation strategy of the energy storage system in the power station is obtained. Figure 5 The average reward curve during the agent's training process is shown. The algorithm converged after 5000 episodes of training, obtaining the optimal dynamic economic scheduling strategy. It can be observed that, due to the agent's initial unfamiliarity with the environment, the reward value obtained after executing scheduling decisions is small. As the training process continues, the agent continuously interacts with the environment and gains experience, therefore the overall trend of the reward value gradually increases and eventually converges. This indicates that the agent has learned the optimal scheduling strategy that minimizes the system's operating cost. Figure 7 It can be seen that energy storage operates under the guidance of electricity prices, charging and discharging during off-peak hours when electricity prices are low and the load is low to prepare for subsequent peak periods, such as 00:00-00:30 and 03:45-04:00; and discharging during peak hours when electricity prices are high and the load is high to reduce operating costs, such as 12:00-12:15 and 17:30-18:45. Figure 8It can be seen that during off-peak and flat electricity price periods, the photovoltaic-storage-charging integrated station purchases electricity from the main grid to meet its electricity demand. When the electricity price is at its peak, the station generates electricity through its on-site photovoltaic power generation system and energy storage equipment to avoid purchasing electricity from the main grid, thereby reducing the station's operating costs.

[0144] In summary, the embodiments of this application focus on the integrated service scenario of electric vehicle charging stations that integrate photovoltaic and energy storage systems. By combining the time-of-use pricing mechanism, the output of all energy storage systems within the station is coordinated, thereby achieving optimal operational economy for the integrated photovoltaic-energy storage-charging station and realizing real-time optimization of the integrated photovoltaic-energy storage-charging station. By estimating the optimal strategy function based on the DDPG algorithm, not only can the curse of dimensionality be effectively avoided, but the information of the entire action domain can also be preserved, realizing the local consumption of new energy resources.

[0145] Those skilled in the art should understand that the above-described input box application types are merely examples. Other existing or future input box application types that are applicable to the embodiments of the present invention should also be included within the scope of protection of the present invention, and are hereby incorporated by reference.

[0146] Those skilled in the art should understand that Figure 2 The number of various network elements shown may be less than that in an actual network for the sake of simplicity, but such omissions are undoubtedly on the premise that they do not affect the clear and sufficient disclosure of the embodiments of the invention.

[0147] Those skilled in the art will understand that the above-described method of determining the invocation strategy based on user information is merely for better illustrating the technical solutions of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention. Any method for determining the invocation strategy based on user attributes is included within the scope of the embodiments of the present invention.

[0148] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the processes depicted in the drawings are not necessarily essential for implementing the present invention.

[0149] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0150] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for operating an integrated photovoltaic-storage-charging station oriented towards grid power balance and new energy consumption, characterized in that, include: S1, based on the integrated service scenario of photovoltaic-storage-charging stations, aims to minimize the operating cost of these stations. It uses power balance constraints, main grid interaction power constraints, and equipment operation constraints as constraints, and real-time output of energy storage units as decision variables to establish a joint operation optimization model for the integrated photovoltaic-storage-charging stations. This model includes: The objective function is shown in equation (1) below: F=min(C E +C BES )(1) Where F represents the operating cost of the integrated photovoltaic, energy storage, and charging station, and C... E The cost of purchasing electricity from the grid for integrated photovoltaic, energy storage, and charging stations. P grid (t) represents the power exchanged between the system and the main grid during time period t. A positive value indicates that the system purchases electricity from the main grid, while a negative value indicates that the system sells surplus electricity back to the grid. e (t) represents the electricity price for time period t; Δt represents the length of the time interval; C BES The depreciation cost of charging and discharging electrical energy storage. P BES (t) represents the charging or discharging power of the energy storage during time period t. A positive value indicates that the energy storage is in a discharging state, and a negative value indicates that it is in a charging state; ρ BES This is the depreciation cost coefficient for energy storage. The constraints are as follows: 1) Power balance constraints: In time period t, the power balance constraint is as shown in equation (2): P grid (t)+P pv (t)+P BES (t)=P load (t)(2) Among them, P pv (t) represents the photovoltaic power generation; P load (t) represents the user's electricity load demand during time period t; 2) The power constraint of the main power grid interaction is shown in equation (3) below: in, and These are the lower and upper limits of the power exchange between the system and the main power grid, respectively. 3) The equipment operation constraints are shown in equations (2)-(3) below: in, and These are the lower and upper limits of the charging / discharging power of electrical energy storage, respectively. For energy storage devices, the constraints are as shown in equations (3)-(4): in, and These are the lower and upper limits of the state of charge of electrical energy storage, respectively; C SOC (t) represents the state of charge of the energy storage during time period t; Q BES The capacity of electrical energy storage; The initial state of charge of the electrical energy storage; η BES The charge / discharge coefficient for electrical energy storage; η ch and η dis These are the charging efficiency and discharging efficiency of electrical energy storage, respectively. S2 transforms the dynamic scheduling problem in the joint operation optimization model into a reinforcement learning network for a policy decision problem using a Markov decision process method; including: The problem of minimizing the operating cost of a photovoltaic-storage-charging integrated power station is transformed into a reward maximization problem for the intelligent agent, as shown in equation (5): Where, r t (s t a t ) represents the total reward value obtained by the agent during the scheduling period t, s i For the observation status of the photovoltaic-storage-charging integrated power station during the dispatch period t, s t ={P load (t),P pv (t),C soc (t-1),t},P load (t), P pv (t), C soc (t-1), where t represents the user's electricity load demand, photovoltaic power generation, energy storage state of charge, and the current dispatch period, respectively; a t For the dynamic economic dispatching of the energy storage system in a photovoltaic-storage-charging integrated power station, the actions within the integrated power station during time period t can be determined by the power output P of the equipment. BES (t) indicates that a t ={P BES (t)}; It is the scaling factor for the cost value, (C) E (s t a t ) represents the agent in state s t Below, the cost of purchasing electricity from the grid for the photovoltaic-storage-charging integrated station during time period t; C BES (s t a t ) represents the agent in state S t The charging and discharging depreciation cost of energy storage during time period t; Use the action-value function Q of equation (6) below. π (s,a) For the energy storage system of the photovoltaic-storage-charging integrated power station in state S l Dynamic economic scheduling action a l The action-value function Q is evaluated. π The larger (s,a) is, the greater a is. l The better: Among them, E π (.) represents the expectation under the optimal objective policy π; γ k ∈[0,1], where γ is the discount factor, representing the proportion of the reward at a future point in the cumulative reward. k The larger the size, the more emphasis is placed on future rewards; t+k s represents the total reward value obtained by the agent in time period t+k; t+k For the integrated photovoltaic, energy storage, and charging station, the state during time period t+k; a t+k The actions performed by the integrated photovoltaic, energy storage, and charging station during time period t+k. k∈N * , indicating the generation of the agent's iterative learning; The optimal objective policy π is obtained according to the following equation (7) to maximize the action-value function: Where A is the set of actions of the agent; S3 uses historical system state data and trains the reinforcement learning network offline based on the Deep Deterministic Policy Gradient (DDPG) algorithm to obtain a trained reinforcement learning network; including: The historical system status data refers to the observed status of the photovoltaic-storage-charging integrated power station, including the system power load demand, photovoltaic power generation, energy storage charge status, and scheduling period of the integrated photovoltaic-storage-charging power station. The reinforcement learning network includes a value network and a policy network, through which the policy network π(s|θ) π Sum-value network Q(s,a|θ) Q Create two independent target networks π'(s|θ) respectively. π' ) and Q'(s,a|θ Q' As shown in equations (8) and (9): Policy Network Value Network The DDPG algorithm sets an optimization cycle T, inputs historical system state data into the policy network and the target network of the value network. After training a batch of data, the DDPG algorithm updates the parameters of the current network of the policy network and the value network using gradient ascent or gradient descent, and then updates the parameters of the target network of the policy network and the value network using a soft update method. After T iterations, the offline learning of the DDPG algorithm is completed, resulting in a trained reinforcement learning network. Specifically, this includes: By minimizing the loss function L(θ) Q )Mode L(θ Q )=E(y t -Q(s t ,a t |θ Q ) 2 (14) Optimize parameters; where θ Q y represents the parameters of the current network in the value network; E(.) is the expectation function; t For the target Q value, through the formula y t =r t +γQ'(s t+1 ,π'(s t+1 |θ π' )|θ Q' (15) Calculated; in equation (15), r t γ is the total reward value obtained by the agent in time period t; γ∈[0,1] is the discount factor; Q' is the target Q value before the update; π' is the target policy; θ π' θ represents the parameters of the target network in the policy network. Q' These are the parameters of the target network in the value network; Through L(θ) Q Regarding θ Q gradient and Update the network; in equation (17), μ Q The value is the network learning rate; Gradient sampling strategy Update the policy network; Through i Q' ←tth Q +(1-τ)θ Q' (20) i π' ←tth π +(1-τ)θ π' (21) Update strategy network parameters θ π ; where μ π τ is the learning rate of the policy network; τ is the soft update coefficient, τ << 1; S4 uses the trained reinforcement learning network and the time-of-use pricing mechanism to perform minimum scheduling cost optimization calculations on the joint operation optimization model of the photovoltaic-storage-charging integrated station, obtaining the real-time self-optimizing operation output of the energy storage system within the integrated station, thereby achieving optimal operation of the integrated photovoltaic-storage-charging station; including: When a scheduling task is received, in each time period, based on the current system state s t Using a trained reinforcement learning network to select and schedule action a t ; Perform action a t And enter the next environmental state, while receiving the reward r. t ; Then, the status information s within the integrated station during the time period t+1 is collected. t+1 The time-of-use electricity price information is used as a new sample, and dynamic scheduling decisions are made for that period, which is the real-time self-optimizing operation output result of the energy storage system in the integrated photovoltaic-storage-charging station.

2. The method according to claim 1, characterized in that, When training the reinforcement learning network offline, the action a is based on the following equation (10). t Training the reinforcement learning network: a t =π(s t |θ π )+v t (10) Among them, v t It is random noise.