Energy storage system control method, device and medium based on dynamic electricity price

Through the energy storage system control method based on dynamic electricity prices, the charging and discharging strategies of the energy storage system are optimized, and the problems of electricity price changes and power supply coordination are solved, system cost minimization and energy utilization efficiency are achieved, and the system flexibility and responsiveness are enhanced.

CN119340974BActive Publication Date: 2025-08-26JIANGYIN FURUI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411393348.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-08
Publication Date
2025-08-26
Estimated Expiration
2044-10-08

AI Technical Summary

Technical Problem

The existing energy storage system control methods fail to effectively consider the real-time changes in electricity prices and the coordination of different power supplies, resulting in low system efficiency and limited energy utilization.

Method used

The energy storage system control method based on dynamic electricity prices is adopted. By obtaining the time-sharing electricity price and power load in the future period, the charging and discharging power of the energy storage module and the purchasing and selling power of the grid interactive module are optimized, and constraints such as power balance and SOC restrictions are established. The Actor-Critic network is used for dynamic adjustments to adapt to electricity price fluctuations and coordinate wind energy and photovoltaic power generation.

Benefits of technology

It minimizes the total cost of the energy storage system, improves system efficiency and energy utilization, enhances the flexibility and response speed of the system, and meets the maximum utilization of renewable energy and the dynamic changes in power load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119340974B_ABST
    Figure CN119340974B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention disclose a method, device, and medium for controlling an energy storage system based on dynamic electricity prices, and relate to the field of energy storage system control. The method comprises: obtaining a time-of-use electricity price for at least one future time period, and the power load of the energy storage system; optimizing the charging power and discharging power of the energy storage module, and the power selling power and power purchasing power of the grid interaction module in each future time period based on the time-of-use electricity price and power load, so that the total cost of the energy storage system is minimized; wherein, the constraints in the optimization process include: the sum of the grid exchange power, photovoltaic power generation power, and wind power generation power is equal to the sum of the charging and discharging power of the energy storage module and the power load. This embodiment improves system efficiency and energy utilization capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of energy storage system control, and in particular to a method, device, and medium for controlling an energy storage system based on dynamic electricity pricing. Background Art

[0002] An energy storage system refers to a system that can store energy from the power grid. It is generally managed based on a fixed electricity price strategy, that is, the charging and discharging operations of the energy storage system are determined according to a preset fixed electricity price curve.

[0003] Existing energy storage system control methods usually do not consider real-time changes in electricity prices and lack effective coordination of different power sources (such as wind power and photovoltaics), resulting in low overall system efficiency and limited effective energy utilization. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device, and medium for controlling an energy storage system based on dynamic electricity pricing to solve the above technical problems.

[0005] In a first aspect, an embodiment of the present invention provides an energy storage system control method based on dynamic electricity pricing, which is applied to an energy storage system including a photovoltaic power generation module, a wind power generation module, an energy storage module, and a grid exchange module;

[0006] The method comprises:

[0007] Obtaining a time-of-use electricity price for at least one future period and the power load of the energy storage system;

[0008] Based on the time-of-use electricity price and power load, the charging power and discharging power of the energy storage module, as well as the electricity selling power and electricity purchasing power of the grid interaction module, are optimized in each future time period to minimize the total cost of the energy storage system;

[0009] Among them, the constraints in the optimization process include: the sum of the grid exchange power, photovoltaic power generation power and wind power generation power is equal to the sum of the charging and discharging power of the energy storage module and the power load.

[0010] In a second aspect, an embodiment of the present invention provides an electronic device, comprising:

[0011] one or more processors;

[0012] a memory for storing one or more programs,

[0013] When the one or more programs are executed by the one or more processors, the one or more processors implement the energy storage system control method based on dynamic electricity price described in any embodiment.

[0014] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the energy storage system control method based on dynamic electricity prices described in any embodiment.

[0015] In summary, the embodiments of the present invention provide a method for controlling an energy storage system based on dynamic electricity prices. By introducing a dynamic electricity price control mechanism, the charging and discharging strategies of the energy storage system are optimized, the overall operating costs of the system are minimized, and the efficiency and energy utilization capacity of the energy storage system are improved. This method establishes detailed power balance conditions, which can be combined with constraints such as charging and discharging restrictions and SOC restrictions to ensure the stable operation of the system. At the same time, it can dynamically adjust the power purchase and sales strategies to adapt to real-time electricity price fluctuations. The system can coordinate wind power, photovoltaic power generation and energy storage systems to maximize the utilization rate of renewable energy. It also takes into account the dynamic changes in power load, improves the flexibility and response speed of the system, and meets sudden power demands. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 This is a flow chart of a method for controlling an energy storage system based on dynamic electricity pricing provided by an embodiment of the present invention;

[0018] Figure 2 This is a schematic diagram of the structure of an Actor-Critic network provided by an embodiment of the present invention;

[0019] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0021] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0022] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0023] An embodiment of the present invention provides a method for controlling an energy storage system based on dynamic electricity pricing. To illustrate this method, the energy storage system to which this method is applied is first described. Optionally, the energy storage system in this embodiment of the present invention includes an energy storage module, a photovoltaic power generation module, a wind power generation module, and a grid exchange module. The photovoltaic power generation module is used to access photovoltaic power, the wind power generation module is used to access wind power, and the grid exchange module is used to exchange power with a trading grid. The energy storage module may include a battery for storing surplus photovoltaic power, wind power, and power purchased from the trading grid within the system, and may also sell the stored power to the trading grid. Furthermore, the energy storage system may also be connected to electrical loads, such as industrial charging stations. These loads consume power from the energy storage system, and the consumed power may come from at least one of the photovoltaic power generation module, the wind power generation module, and the energy storage module.

[0024] For the above energy storage systems, Figure 1 This is a flow chart of a method for controlling an energy storage system based on dynamic electricity prices provided by an embodiment of the present invention. The method dynamically controls the charge and discharge power of the energy storage module according to the real-time electricity price to minimize the cost of the entire energy storage system and achieve real-time response capabilities. The method is executed by an electronic device, such as Figure 1 As shown, specifically including:

[0025] S110: Obtain a time-of-use electricity price for at least one future period and the power load of the energy storage system.

[0026] Time-of-use electricity prices can be derived from the electricity trading market. For example, the European Electricity Trading Market announces electricity prices for the next 24 hours at 2:00 PM daily. The power load is a forecast of the amount of electricity consumed by the energy storage system's loads. Optionally, the power load for the same period the previous day can be used as the power load for the next period. This embodiment periodically acquires these parameters as data sources for the entire method.

[0027] S120. Based on the time-of-use electricity price and power load, optimize the battery charging power and discharging power, as well as the power sales and purchase power of the grid interaction module, for each future time period to minimize the total cost of the energy storage system. The optimization process includes constraints such as: the sum of grid exchange power, photovoltaic power generation power, and wind power generation power equals the sum of battery power and power load, i.e., a power balance constraint.

[0028] This embodiment comprehensively considers real-time dynamic electricity prices, as well as predicted values ​​for power load, photovoltaic power generation, and wind power generation, to optimize the energy storage system's charge and discharge strategy. This optimized strategy enables the energy storage system to purchase and store electricity during periods of low electricity prices while meeting its own power load, and sell the stored electricity during periods of higher electricity prices, thereby reducing the total cost of the entire system. Optionally, the charge and discharge strategy can be represented by the battery's charging and discharging power, as well as the grid interaction module's power sales and purchase power, for each future time period. Arranging these four power values ​​in chronological order yields a power curve for the battery and the charge and discharge strategy, which is used to guide the battery and charge and discharge strategies' charging and discharging actions over multiple future time periods.

[0029] In a specific implementation, the constraints of the optimization process can be constructed from the following aspects:

[0030] 1) Battery-controlled type:

[0031]

[0032] P t cha ≤P t cs max

[0033] P t dis ≤P t cs max (2)

[0034] Among them, P t cs maxt Indicates the rated power of the battery, P t bat Represents the battery power during period t, P t chaand P t dis They represent the battery charging power and battery discharging power during period t, respectively. t cs max 、P t cha and P t dis are all non-negative numbers, and P in the same period t cha and P t dis At most one of them is non-zero.

[0035] At the same time, the battery capacity meets:

[0036] soc=E t / E max (3)

[0037] E t =E t-1 +P t cha -P t dis (4)

[0038] Among them, SOC represents the charging state of the battery in the t period, which can be read through the battery manager; E max Indicates the rated capacity of the battery, E t and E t-1 They represent the remaining capacity of the battery in period t and period t-1 respectively. In formula (4), the unit time for calculating power is taken as one period.

[0039] Divide both sides of formula (4) by E max , we can get:

[0040]

[0041] and then

[0042] E max SOC = E max ·SOC t-1 +P t cha -P t dis (6)

[0043] The SOC itself needs to meet the following requirements:

[0044] SOC min ≤SOC≤SOC max (7)

[0045] Among them, SOC min and SOCmax Respectively represent the minimum SOC value and maximum SOC value that the battery should meet for normal operation. Based on formula (7), the at least one period is accumulated and summed, and the initial energy storage capacity SOC of the first period is considered. 0 And formula (6), we can get:

[0046] E max (SOC min -SOC 0 ))≤sum(P t cha -P t dis )≤E max (SOC max -SOC 0 ) (8)

[0047] In addition, t in the above variables can also be understood as time t, and formula (8) can be understood as summing each time. The following variables are similar and will not be repeated here.

[0048] 2) Grid exchange constraints:

[0049]

[0050] P t buy ≤P t net max

[0051] P t sell ≤P t net max (10)

[0052] Among them, P t net represents the grid exchange power during period t, P t buy and P t sell They represent the power purchased from the grid and the power sold to the grid during period t, respectively. t net max Indicates the maximum exchange power of the power grid. Similarly, P t net max 、P t buy and P t sell are all non-negative numbers, and P in the same period t buy and P t sell At most one of them is non-zero. There is a certain approximation in formula (9), which assumes that when the battery is in the charging state (P tbat Greater than or equal to 0) means purchasing electricity; battery discharge state (P t bat Less than 0) means electricity is sold. This approximation will only slightly prolong the convergence time of the optimization process and will not affect the optimization results.

[0053] 3) Power balance constraints:

[0054]

[0055] in, and They represent the wind power generation and photovoltaic power generation in period t, The left side of the equation represents the power supply of the energy storage system, and the right side represents the power consumption of the energy storage system. Power supply = power consumption.

[0056] 4) Objective function: Minimum total cost

[0057]

[0058] Among them, F represents the total cost of T periods, s w 、s pv 、s buy and s ess Represent the unit costs of wind power generation, photovoltaic power generation, electricity purchase, electricity sales and battery discharge, s sell Indicates the unit revenue of electricity sales, which is a fixed constant. buy and s sell This is the real-time electricity price for the current period.

[0059] In summary, the overall model is:

[0060] min F

[0061] st(1)-(12)

[0062] To solve the optimization problem of the above model, this embodiment provides the following two optional implementations based on the changes in wind turbine output and photovoltaic output:

[0063] The first optional implementation method is suitable for situations where the wind turbine power generation and photovoltaic power generation are relatively stable. and If the system is relatively stable over a longer period, such as 24 hours or longer, then at the beginning of the stable period, an exhaustive algorithm can be used to traverse each set of solutions in the solution space and select the optimal solution that meets the constraints as the control strategy for the stable period. Alternatively, a recursive algorithm function or an exhaustive algorithm function in the Java language can be called, and the objective function min F and all constraints are input as known conditions into the function, which automatically returns the optimal strategy that meets the conditions. This method has a large amount of computation, but after a single calculation, it can be used continuously for a long period of time. Therefore, it can be applied when the wind turbine power generation and photovoltaic power generation are relatively stable.

[0064] The second optional implementation is suitable for situations where the wind turbine power generation and photovoltaic power generation are relatively frequent, for example, due to frequent changes in light and / or wind power, and / or It changes every hour, or fluctuates at a faster frequency. In this case, if an exhaustive algorithm is executed every time a change occurs, it will cause strategy lag and waste computing resources. Therefore, the Actor-Critic algorithm based on reinforcement learning can be used to dynamically predict the optimal strategy for the next cycle based on the current status and historical strategy of the energy storage system. Optionally, the duration of the at least one period (T periods) can be used as an optimization cycle, which can also be based on and The following description will be made using T=2 as an example.

[0065] In a specific embodiment, first, the action variables and state variables of the Actor-Critic network are constructed. Optionally, the charging power and discharging power of the battery, as well as the distribution of the power sold and purchased by the grid interaction module in a certain optimization cycle can be used as action variables:

[0066]

[0067] Among them, a N Represents the action of optimization period N, t and t+1 are two time periods in optimization period N. In order to reduce the amount of calculation, the action space can be discretized to obtain candidate actions for each discretization. Specifically, the action space is determined by the rated power of the battery and the maximum exchange power of the grid, as shown in formula (2) and formula (9). The four power variables in the action can be and Discretize them at certain intervals, combine the discrete values ​​of each power in each period of an optimization cycle one by one, and then exclude the values ​​that do not meet the "P t buy and P tsell At most one of the P values ​​in the same period is non-zero. t cha and P t dis There is at most one non-zero constraint in the combination (e.g., excluding P t buy ×P t sell Not equal to 0, or P t cha ×P t dis not equal to 0), we can get each discrete candidate action.

[0068] At the same time, the initial SOC value of the battery, time-of-use electricity price, power load, photovoltaic power generation and wind power generation of an optimization cycle, as well as the action, time-of-use electricity price, power load, photovoltaic power generation and wind power generation of the previous optimization cycle are arranged in sequence and used as the state variables of the optimization cycle. Taking the optimization cycle N as an example, the S N Including: initial SOC value Time-of-use electricity price Power load Photovoltaic power generation and wind power generation And the action a of the optimization cycle N-1 N-1 , time-of-use electricity prices Power load Photovoltaic power generation and wind power generation Wherein, t-2 and t-1 are two time periods in the optimization cycle N-1.

[0069] Based on the above action variables and state variables, we can construct Figure 2 The Actor-Critic network shown. When this network is used, the state S of the current optimization cycle N can be N Input the Actor network, which outputs the probability of selecting each candidate action within the cycle, and selects the candidate action with the highest probability as the optimal action. Then, the reward value R is calculated according to formula (14), and the Critic network evaluates the value of the candidate action based on the reward value and updates the parameters of the Actor network and Critic network based on the value:

[0070]

[0071] Among them, F N and F N-1 Respectively represent the total cost of the current optimization cycle N and the previous optimization cycle N-1; -A means less than -F maxA fixed negative constant, F max Indicates the maximum fee threshold, F max >0.

[0072] Specifically, the parameters of the Actor network are updated according to the following formula:

[0073]

[0074] in, and They represent the network parameters of the Actor network in the optimization cycle N+1 and cycle N, α represents the learning rate, π(a N |s N ) means in state s N Select each action a N The probability vector, A(s N ,a N ) represents state s N and action a N Advantage function under. When the advantage value is positive, it means that the current action is feasible, and its probability of being selected is gradually increased through gradient ascent. On the contrary, when the advantage value is negative, the action probability will decrease. A(s N ,a N ) can be obtained by:

[0075] A(s N ,a N )=Q(s N ,a N )-V(s N ) (18)

[0076] Among them, Q(s N ,a N ) represents the state action value, that is, s N with a N The value of the combination; V represents the state value. The Critic network can output V or Q(s) for each action. N ,a N ), which is not specifically limited in this embodiment. Optionally, according to the Bellman equation, the action value Q function can be approximately expressed as:

[0077] Q(s N ,a N )≈r N +γV(s N+1 ) (19)

[0078] Among them, r N Represents the reward value of the current cycle, γ represents the discount rate, which is a fixed constant.

[0079] Loss function of the Critic network It can be expressed as:

[0080]

[0081] The critic updates the network parameters based on this loss function. The parameter update process described above is based on existing techniques and will not be further detailed. In the early stages of network operation, the actor network may not be able to accurately identify the optimal action that minimizes the total system cost. However, as network parameters are updated, the actor network will gradually favor the action that minimizes the total cost, ultimately achieving dynamic, real-time optimal control of the energy storage system.

[0082] Furthermore, it can be seen that the number of candidate actions in the action space is still relatively large, such as Figure 2 As shown, the number of output nodes in the Actor network is the same as the number of candidate actions. Therefore, the computational complexity within the network is still very large, especially when T takes a larger value, as the number of candidate actions increases exponentially. To accelerate the learning rate of the agent, this embodiment can further reduce the size of the action space based on the power variation pattern of the energy storage system and the unique input-output relationship of the Actor-Critic network of this embodiment, thereby reducing the computational complexity of the agent and the complexity of the Actor network.

[0083] In one specific embodiment, each candidate action variable can be clustered into multiple clusters based on action similarity. Alternatively, each candidate action in the original action space can be arranged as a vector, and the distance between candidate actions can be measured by vector similarity. A K-Means clustering algorithm is then used to cluster all candidate actions, with candidate actions with close distances grouped together.

[0084] Then, based on the change rate threshold of each power in the action variable, determine the action range that each candidate action in the cluster can achieve in the next optimization cycle, and use the actions within each action range as the final multiple candidate actions of each cluster. In the power system, due to the limitations of the device itself or the consideration of stable operation of the system, the input and output power change rates of each module have certain threshold limits, and excessive power mutations should not occur (except in the start or stop state). Therefore, the four power change rate thresholds Pm in the action variable can be obtained separately, such as Pm = 10% × Pe / min, where Pe represents the rated power of a certain power, and 10% × Pe / min means that the output power change per minute reaches 10% of the rated output power. According to this threshold, for any power, the value in the current period is P t In the case of t -Pm×⊿t,P t+Pm×⊿t] range, where ⊿t represents the duration of an optimization cycle. The discrete power that meets this range can be used together with the 0 power (corresponding to the next period of stop) and the starting power (corresponding to the next period of start) as P t The power value can be reached in the next period. Accordingly, for any cluster, all candidate actions within the cluster are similar, so the ranges of the four powers are also relatively concentrated. Assuming a power range is [a, b], the set consisting of [a-Pm×⊿t, b+Pm×⊿t], zero power, and the start-up power can be used to determine the power range that all candidate actions in the cluster can reach in the next period. If the power range includes multiple intervals [a, b], the interval [a-Pm×⊿t, b+Pm×⊿t] can be determined for each [a, b] interval. The union of all intervals, together with zero power and the start-up power, is then taken to form the power range that can be reached in the next period. After performing the above operation on each of the four powers, the four power ranges can be used to obtain the action subspace of the cluster in the next period. Of course, other methods can also be used to determine the action subspace of each cluster, as long as the power change rate constraint is satisfied. Each action subspace will be much smaller than the original action space.

[0085] After clustering is complete, an Actor network can be constructed for each cluster to make decisions based on the state and actions of each cluster. This embodiment provides two optional implementations based on the differences in the Critic networks of each cluster:

[0086] Optional implementation method A: construct an Actor-Critic network for each cluster, so that the number of output nodes of the Actor network is equal to the size of the action subspace of the corresponding cluster. The number of output nodes of the Critic network can be 1, or it can be equal to the size X of the action subspace of the corresponding cluster. When the number of output nodes of the Critic network is 1, the input of the Critic network is the new state brought about by the current optimal action, and the output is the value V of the new state, such as Figure 2 As shown, the action state value corresponding to the optimal action is Q = R + γV. When the number of output nodes of the Critic network is X ( Figure 2 (not shown in the figure), the critic network inputs the new state and reward value brought by each candidate action in the subspace, and outputs the value Q of each candidate action combined with the current state.

[0087] Correspondingly, the Actor network and Critic network of each cluster are updated separately. The state and action of a cluster will only update the network parameters of the cluster and will not affect the network parameters of other clusters. Optionally, when a certain optimization cycle arrives, the target cluster to which the state of the current optimization cycle belongs can be first determined, the state is input into the Actor network of the target cluster, and the Actor network of the target cluster outputs the probability of selecting each candidate action in the action subspace of the target cluster during the cycle, and the action with the highest probability is selected as the optimal action. Then, the reward of the optimal action is calculated according to formula (14), or the reward of all candidate actions in the action subspace of this cluster is calculated; the Critic network of the target cluster evaluates the value of the combination of the optimal action and the current state based on the reward value, or evaluates the value of the combination of all actions in the action subspace of this cluster with the current state, and updates the parameters of the Actor network and Critic network of the target cluster based on the value.

[0088] In this approach, the network structure of each cluster is relatively streamlined, the amount of computation is relatively small, and each cluster network performs training, decision-making, and evaluation separately. The time required for each Actor network to be trained and stabilized is longer than that of optional implementation method B.

[0089] Optional implementation B: Construct an actor network for each cluster. Each cluster's actor network is independent, and the number of output nodes in each actor network is equal to the size of the corresponding cluster's action subspace. At the same time, each cluster shares a critic network, and the number of output nodes in this critic network is equal to the size of the union of all clusters' action subspaces. The critic network inputs the new state and reward value brought about by each candidate action in this union, and outputs the value Q of each candidate action combined with the current action.

[0090] Accordingly, the Actor network of each cluster is updated separately, and the Critic network is also updated after each Actor network is updated. Optionally, when a certain optimization cycle arrives, the target cluster to which the state of the current optimization cycle belongs can be first determined; the state is input into the Actor network of the target cluster, and the Actor network of the target cluster outputs the probability of selecting each candidate action in the action subspace of the target cluster within the cycle, and the action with the highest probability is selected as the optimal action. Then, the reward value R of each candidate action in the action subspace of the target cluster is calculated according to formula (14), and the reward value of the remaining actions in the union is set to a negative constant less than or equal to -A; the Critic network evaluates the value Q of each action and state combination in the union according to each reward value; and according to the Q corresponding to each action in the action subspace of the target cluster, the Actor network of the target cluster is updated; according to the Q corresponding to each action in the union, the parameters of the Critic network are updated.

[0091] In this approach, the actor network structure is relatively streamlined, requiring relatively little computation. The critic network is also somewhat streamlined compared to the original action space output critic network, but is more complex than the network structure of Alternative Implementation A. The critic network of this embodiment evaluates the action states of all actions in the union of each cluster's action subspaces. Each actor network decision triggers an update to the critic network, helping the critic network more quickly and objectively evaluate action state values ​​and state values, thereby promoting stable training of the actor network.

[0092] The above two optional implementations both improve the Actor-Critic network based on the power variation characteristics of the power system, reducing the number of output nodes and the amount of computation required in the network. In practical applications, they can be flexibly selected as needed.

[0093] Furthermore, to ensure the stability of power system power variations, whether using an Actor-Critic network in the original action space, an independent Actor-Critic network for each cluster, or a network structure with independent Actor networks in each cluster and a shared Critic network, after the Actor network outputs the probability of selecting each candidate action within a certain period, a mask with the same size as the output action space is constructed, with each mask element corresponding to a candidate action. Specifically, the mask size varies for each of the three Actor-Critic networks described above and can be the same as the actual number of output nodes in the network.

[0094] Then, according to the change rate threshold of each power in the action variable, determine the multiple actions that can be achieved in the current optimization cycle for the action of the previous optimization cycle, and set the mask elements corresponding to the multiple actions to 1, and the mask elements of other actions to 0. Optionally, the power of the last period of the previous optimization cycle is P t In the case of , the power range that can be achieved in the first period of the current optimization cycle is [P t -Pm×⊿t,P t +Pm×⊿t], the maximum power range that can be achieved in the next period is [P t -Pm×2⊿t,P t +Pm×2⊿t] (this can also be a portion of this range, depending on the actual situation), and so on. The power range for each period can be gradually determined. Finally, adding the zero power and the starting power gives the range of actions that the action in the previous cycle can reach in this cycle. This range is smaller than the action subspace described by the current cluster. Mask elements corresponding to actions within this range are set to 1, and mask elements corresponding to actions outside the range are set to 0.

[0095] Finally, multiply the probability of each candidate action by the corresponding mask element to obtain the final probability of each candidate action. The candidate action with the highest probability is selected as the optimal action. Alternatively, you can skip the mask and start directly with the candidate with the highest probability. Then, determine whether it falls within the range of actions that the action in the previous cycle can reach in this cycle. The first candidate action that meets this range is the optimal action.

[0096] In summary, this embodiment provides a method for controlling an energy storage system based on dynamic electricity pricing. By introducing a dynamic electricity price control mechanism, it optimizes the charging and discharging strategies of the energy storage system, minimizes the overall operating costs of the system, and improves the efficiency and energy utilization capabilities of the energy storage system. This method establishes detailed constraints such as power balance, charge and discharge limits, and SOC limits to ensure stable system operation. It can also dynamically adjust power purchase and sales strategies to adapt to real-time electricity price fluctuations. The system can coordinate wind power, photovoltaic power generation, and energy storage systems to maximize the utilization of renewable energy. It also takes into account the dynamic changes in power load, improving the system's flexibility and response speed to meet sudden power demands.

[0097] In particular, when photovoltaic and wind energy fluctuations are relatively stable, multiple constraints and objective functions can be established to perform global optimization over relatively long periods of time, ensuring the efficient and economical operation of the system. When photovoltaic and wind energy fluctuations are relatively frequent, the intelligent agent can adaptively update the control strategy for each shorter period, further improving the system's dispatch intelligence level. During the intelligent agent construction process, this embodiment provides various forms of actor-critic networks and matching training methods based on the power stability requirements of the power system and specific input-output relationships. It also provides an output mask to further constrain the power change rate, jointly achieving optimal and stable control of the energy storage system.

[0098] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more. Figure 3 In the embodiment, a processor 60 is used as an example; the processor 60, the memory 61, the input device 62 and the output device 63 in the device can be connected by a bus or other means. Figure 3 The bus connection is taken as an example.

[0099] Memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for controlling an energy storage system based on dynamic electricity pricing in the embodiments of the present invention. Processor 60 executes the software programs, instructions, and modules stored in memory 61 to execute various functional applications and data processing functions of the device, thereby implementing the aforementioned method for controlling an energy storage system based on dynamic electricity pricing.

[0100] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0101] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.

[0102] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the energy storage system control method based on dynamic electricity pricing according to any embodiment.

[0103] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by an instruction execution system, device or device or used in combination with it.

[0104] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0105] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0106] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for controlling an energy storage system based on dynamic electricity pricing, characterized in that: Applicable to energy storage systems including photovoltaic power generation modules, wind power generation modules, batteries and grid exchange modules; The method comprises: Obtaining a time-of-use electricity price for at least one future period and the power load of the energy storage system; Based on the time-of-use electricity price and power load, the charging power and discharging power of the battery, as well as the electricity selling power and electricity purchasing power of the grid exchange module in each future period are optimized to minimize the total cost of the energy storage system; Specifically, the duration of the at least one time period is used as an optimization cycle; The charging power and discharging power of the battery, and the distribution of the power selling power and power purchasing power of the grid exchange module in an optimization cycle are used as action variables; The initial SOC of the battery, time-of-use electricity price, power load, photovoltaic power generation and wind power generation of an optimization cycle, as well as the action, time-of-use electricity price, power load, photovoltaic power generation and wind power generation of the previous optimization cycle are used as state variables; The action space is determined based on the rated power of the battery and the maximum exchange power of the grid. Each power in the action variable is discretized at a certain interval. The discretized values ​​of each power in each time period of an optimization cycle are combined one by one, and the combinations that do not meet the constraints are eliminated to obtain the discrete candidate actions. Based on the action similarity, each candidate action variable is clustered into multiple clusters; based on the change rate threshold of each power in the action variable, multiple actions that each candidate action in the cluster can achieve in the next optimization cycle are determined as the final multiple candidate actions of each cluster; An Actor-Critic network is constructed for each cluster. The target cluster to which the state of the current optimization cycle belongs is determined. The state is input into the Actor network of the target cluster, which then outputs the probability of selecting each candidate action of the target cluster within the cycle. The final candidate action with the highest probability is selected as the optimal action. The reward value R is calculated according to formula (14), and the Critic network evaluates the value of the state action based on the reward value, and updates the parameters of the Actor network and Critic network of the target cluster according to the value: C1:E max (SOC min -SOC 0 )≤sum(P t cha -P t dis )≤E max ((SOC max -SOC 0 ) Among them, F N and F N-1 Respectively represent the total cost of the current optimization cycle and the previous optimization cycle; T represents the number of time periods included in each optimization cycle, and They represent the wind power generation and photovoltaic power generation in period t, P t buy and P t sell They represent the power purchased from the grid and the power sold to the grid during period t, respectively. t cha and P t dis They represent the battery charging power and battery discharging power during period t, s w 、s pv 、s buy and s ess Represent the unit costs of wind power generation, photovoltaic power generation, electricity purchase, electricity sales and battery discharge, s sell Represents the revenue from selling electricity, P t net represents the grid exchange power during period t, represents the power load during period t, P t bat Indicates the battery power during period t; SOC min , SOC max and SOC 0 Respectively represent the minimum SOC, maximum SOC, and initial SOC of the battery in the current optimization cycle; E max Indicates the rated capacity of the battery; -A means less than -F max A fixed negative constant, F max Indicates the maximum fee threshold, F max >0; α and β represent preset weights between [0,1].

2. The method according to claim 1, characterized in that The constraints of the optimization process include: From max (SOC min -SOC 0 )≤sum(P t cha -P t dis )≤E max (SOC max -SOC 0 )(8) P t buy ≤P t netmax P t sell ≤P t netmax (10) Among them, P t csmax Indicates the rated power of the battery; P t netmax Indicates the maximum exchange power of the power grid.

3. The method according to claim 1, characterized in that The number of output nodes of each Actor network is equal to the final number of candidate actions of the corresponding cluster, and the number of output nodes of each Critic network is equal to 1 or equal to the final number of candidate actions of the corresponding cluster; The Critic network evaluates the value of the state action according to the reward value, and updates the parameters of the Actor network and the Critic network according to the value, including: the Critic network of the target cluster evaluates the value of the final state and action combination corresponding to the target cluster according to the reward value, and updates the parameters of the Actor network and the Critic network of the target cluster according to the value.

4. The method according to claim 1, wherein The Actor network of each cluster is independent, and the number of output nodes of each Actor network is equal to the number of candidate actions of the corresponding cluster; each cluster shares a Critic network, and the number of output nodes of the Critic network is equal to the size of the union of the final candidate actions of all clusters; The formula (14) calculates the reward value R, the Critic network evaluates the value of the state action according to the reward value, and updates the parameters of the Actor network and the Critic network according to the value, including: calculating the reward value R of each final candidate action corresponding to the target cluster according to formula (14), and setting the reward value of the remaining actions in the union to a negative constant less than or equal to -A, and the Critic network evaluating the value of each action and state combination in the union according to each reward value; and updating the Actor network of the target cluster according to the value of each final candidate action and state combination corresponding to the target cluster, and updating the parameters of the Critic network according to the value of each action in the union.

5. The method according to claim 1, wherein The Actor network of the target cluster outputs the probability of selecting the final candidate actions of the target cluster within the period, and takes the final candidate action with the highest probability as the optimal action, including: The Actor network outputs the probability of selecting each candidate action in this cycle; Construct a mask of the same size as the action space, with each mask element corresponding to each candidate action; According to the change rate threshold of each power in the action variable, determine multiple actions that can be achieved in the current optimization cycle in the action of the previous optimization cycle, set the mask elements corresponding to the multiple actions to 1, and set the mask elements of other actions to 0; Multiply the probability of each candidate action by the corresponding mask element to obtain the final probability of each candidate action; The final candidate action with the highest probability is taken as the optimal action.

6. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the energy storage system control method based on dynamic electricity price as described in any one of claims 1-5.

7. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the method for controlling an energy storage system based on dynamic electricity prices as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Dynamic economic dispatching method and system for power distribution network

    CN117674114A