Energy storage system scheduling model training method and device, electronic equipment and storage medium

By introducing mainline reward data and branchline penalty data into the energy storage system scheduling model, and using policy gradient and time-series difference methods to update parameters, the problem of low convergence accuracy of the energy storage system scheduling model is solved, and the economic benefits of the energy storage system and the user's electricity cost are optimized.

CN116384479BActive Publication Date: 2026-01-06SUNGROW (SHANGHAI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310318448.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2026-01-06
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

Existing energy storage system scheduling models are easily affected by factors such as the charging and discharging actions of energy storage systems, electricity prices, and electricity consumption during the training process using near-end strategy optimization algorithms. This results in low convergence accuracy and makes it difficult to achieve the global optimum for economic scheduling.

Method used

By introducing mainline reward data and branchline penalty data into the energy storage system scheduling model, and using policy gradient and time-series difference methods to update model parameters, the consistency between the charging and discharging actions of the energy storage system and the direction of state variable changes is ensured, thereby improving the convergence accuracy and speed of the model.

Benefits of technology

This improved the convergence accuracy and speed of the energy storage system scheduling model, ensuring the optimal economic benefits of the energy storage system and reducing users' electricity expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116384479B_ABST
    Figure CN116384479B_ABST
Patent Text Reader

Abstract

The present specification relates to a method and device for training an energy storage system scheduling model, an electronic device, and a storage medium. The method comprises: determining main line reward data of an energy storage system scheduling model according to electricity cost data of the energy storage system in a time period between a first time and a second time, wherein the first time is earlier than the second time; determining branch line penalty data of the energy storage system scheduling model according to charge-discharge action deviation data of the energy storage system in the time period between the first time and the second time, wherein the charge-discharge action deviation data is used to describe consistency between a charge-discharge action of the energy storage system and a change direction of a state variable of the energy storage system; and updating parameters of the energy storage system scheduling model according to the main line reward data and the branch line penalty data to obtain a target energy storage system scheduling model. Embodiments of the present specification improve the convergence accuracy of the energy storage system scheduling model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of power dispatching technology, and in particular to a method, apparatus, electronic device and storage medium for training a dispatching model of an energy storage system. Background Technology

[0002] Because real-time electricity pricing and load consumption processes are highly random and uncertain, it is crucial to predetermine the scheduling strategy for energy storage systems in order to optimize electricity consumption and reduce electricity costs.

[0003] Given the global exploration capability of near-end policy optimization algorithms, some related technologies utilize these algorithms to construct scheduling models and obtain scheduling strategies. However, due to the complexity of energy storage system operation, the convergence accuracy of scheduling models constructed using near-end policy optimization algorithms still needs improvement. Summary of the Invention

[0004] This specification aims to at least partially address one of the technical problems in the related art. To this end, one objective of this specification is to propose a training method for an energy storage system scheduling model that can improve the convergence accuracy of the energy storage system scheduling model.

[0005] The second objective of this specification is to propose a training device for an energy storage system scheduling model.

[0006] The third objective of this specification is to propose an electronic device.

[0007] The fourth objective of this specification is to provide a computer-readable storage medium.

[0008] To achieve the above objectives, a training method for an energy storage system scheduling model is proposed in the first aspect of this specification. The training method includes: determining the main reward data of the energy storage system scheduling model based on electricity cost data of the energy storage system during a period between a first time and a second time; wherein the first time is earlier than the second time; determining the branch penalty data of the energy storage system scheduling model based on charging and discharging action deviation data of the energy storage system during the period between the first time and the second time; wherein the charging and discharging action deviation data describes the consistency between the charging and discharging actions of the energy storage system and the direction of change of the state variables of the energy storage system; and updating the parameters of the energy storage system scheduling model based on the main reward data and the branch penalty data to obtain a target energy storage system scheduling model.

[0009] In some embodiments of this specification, the environmental state variables of the energy storage system include a power quantity variable; before determining the branch penalty data of the energy storage system scheduling model based on the charge and discharge action deviation data of the energy storage system during the period between the first time and the second time, the method further includes: determining the power quantity action deviation data of the energy storage system during the period between the first time and the second time based on the first power quantity data of the power quantity variable at the first time, the second power quantity data of the power quantity variable at the second time, and the action data of the action variable of the energy storage system at the first time; wherein, the charge and discharge action deviation data includes the power quantity action deviation data.

[0010] In some embodiments of this specification, the environmental state variables of the energy storage system include actual electricity price variables; before determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first time and the second time, the method further includes: determining the electricity price action deviation data of the energy storage system during the period between the first time and the second time based on the first actual electricity price data of the actual electricity price variable at the first time, the second actual electricity price data of the actual electricity price variable at the second time, and the action data of the action variable of the energy storage system at the first time; wherein, the charging and discharging action deviation data includes the electricity price action deviation data.

[0011] In some embodiments of this specification, the environmental state variables of the energy storage system include actual electricity price variables and predicted electricity price variables. Before determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first time and the second time, the method further includes: determining the electricity price prediction action deviation data of the energy storage system during the period between the first time and the second time based on the predicted electricity price data of the predicted electricity price variable at the second time, the second actual electricity price data of the actual electricity price variable at the second time, and the action data of the action variable of the energy storage system at the first time; wherein, the charging and discharging action deviation data includes the electricity price prediction action deviation data.

[0012] In some embodiments of this specification, the charging and discharging action deviation data includes power consumption action deviation data, electricity price action deviation data, and electricity price prediction action deviation data. The step of determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the time period between the first and second times includes: performing a weighted summation calculation on the power consumption action deviation data, the electricity price action deviation data, and the electricity price prediction action deviation data according to a first weighting coefficient corresponding to the power consumption action deviation data, a second weighting coefficient corresponding to the electricity price action deviation data, and a third weighting coefficient corresponding to the electricity price prediction action deviation data, to obtain the branch penalty data of the energy storage system scheduling model.

[0013] In some embodiments of this specification, the environmental state variables of the energy storage system include actual electricity price variables and load power variables; before determining the mainline reward data of the energy storage system scheduling model based on the electricity price data of the energy storage system during the period between the first time and the second time, the method further includes: determining the electricity price data based on the first actual electricity price data of the actual electricity price variable at the first time, the first load power of the load power variable at the first time, the second load power of the load power variable at the second time, and the action data of the action variables of the energy storage system at the first time.

[0014] In some embodiments of this specification, determining the main reward data of the energy storage system scheduling model based on the electricity cost data of the energy storage system during the time period between the first time and the second time includes: multiplying the electricity cost data with the fourth weight coefficient corresponding to the electricity cost data to obtain the main reward data of the energy storage system scheduling model.

[0015] In some embodiments of this specification, the predicted electricity price data corresponding to the second time point is obtained by predicting using an electricity price prediction model. The training process of the electricity price prediction model includes: constructing a training sample set for electricity prices; wherein the training sample set includes multiple training samples; the training samples use historical electricity price-related data within a specified time interval; the historical electricity price-related data has multiple feature variables; the training samples include data on the multiple feature variables; the multiple feature variables include the highest daily electricity price, the lowest daily electricity price, the average daily electricity price, and the daily electricity price volatility; training an initial prediction model using the training sample set to obtain the electricity price prediction model; wherein the initial prediction model uses a long short-term memory network.

[0016] In some embodiments of this specification, updating the parameters of the energy storage system scheduling model based on the main reward data and the branch penalty data includes: using the environmental state variables and action variables of the energy storage system as inputs to the value network and action network of the energy storage system scheduling model, and using the policy gradient and temporal difference method to update the parameters of the energy storage system scheduling model with the objective of maximizing the difference between the main reward data and the branch penalty data.

[0017] In some embodiments of this specification, the method further includes: determining the objective function and constraints of the minimum system operating electricity cost of the energy storage system scheduling model; and constructing a reward function for a near-end strategy optimization algorithm using the objective function and the constraints; wherein the reward function includes a main reward function for outputting the main reward data and a branch penalty function for outputting the branch penalty data.

[0018] In some embodiments of this specification, the constraints include the energy storage system's power constraint and the energy storage system's actual output power constraint; the energy storage system's power constraint is constructed based on the upper and lower limits of the energy storage system's power; the energy storage system's actual output power constraint is constructed based on the energy storage system's actual output power being less than or equal to the energy storage system's rated power.

[0019] To achieve the above objectives, a third aspect of this specification provides a training device for an energy storage system scheduling model. The training device includes: a mainline reward data determination module, used to determine the mainline reward data of the energy storage system scheduling model based on electricity cost data of the energy storage system during a period between a first time and a second time; wherein the first time is earlier than the second time; a branch line penalty data determination module, used to determine the branch line penalty data of the energy storage system scheduling model based on charging and discharging action deviation data of the energy storage system during the period between the first time and the second time; wherein the charging and discharging action deviation data describes the consistency between the charging and discharging actions of the energy storage system and the direction of change of the state variables of the energy storage system; and a training module, used to update the parameters of the energy storage system scheduling model based on the mainline reward data and the branch line penalty data to obtain a target energy storage system scheduling model.

[0020] To achieve the above objectives, a third aspect of this specification provides an electronic device, characterized in that it includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the method as described in any of the embodiments of the first aspect.

[0021] To achieve the above objectives, a fourth aspect of this specification provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device containing the computer-readable storage medium to perform the method described in any of the embodiments of the first aspect.

[0022] Through the above embodiments, when updating the parameters of the initially constructed energy storage system scheduling model, in addition to ensuring that the main reward data guarantees the main objective of the energy storage system, supplementary penalty data is added. This reduces the possibility of poor model training performance due to changes in or inaccuracies of state variables in the environment where the energy storage system is located. Iteratively updating the parameters of the energy storage system scheduling model using an immediate reward composed of main reward data and supplementary penalty data improves the model's convergence accuracy and speed, resulting in a more accurate target energy storage system scheduling model.

[0023] Additional aspects and advantages of this specification will be set forth in part in the description which follows, and in part will be obvious from the description or may be learned by practice of this specification. Attached Figure Description

[0024] Figure 1a This is a schematic diagram of the energy storage system scheduling model training system provided in the embodiments of this specification.

[0025] Figure 1b This is a schematic diagram of a user storage system provided in one embodiment of this specification.

[0026] Figure 2 This is a flowchart illustrating the energy storage system scheduling model training method provided in the embodiments of this specification.

[0027] Figure 3 This is a flowchart illustrating the training method for the electricity price prediction model provided in the embodiments of this specification.

[0028] Figure 4 This is a schematic diagram of the network structure of the Long Short-Term Memory network provided in the embodiments of this specification.

[0029] Figure 5 This is a flowchart illustrating a method for training an energy storage system scheduling model according to another embodiment of this specification.

[0030] Figure 6 This is a schematic diagram of the structure of the energy storage system scheduling model training device provided in the embodiments of this specification.

[0031] Figure 7 This is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification. Detailed Implementation

[0032] The embodiments of this specification are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this specification, and should not be construed as limiting this specification.

[0033] An energy storage system is a defined area of ​​objects or space used to determine the object of study when analyzing energy storage processes. Energy storage systems often involve multiple energies, devices, substances, and processes, and are complex energy systems that change over time.

[0034] The energy storage system proposed in this specification is an energy storage system for electricity. Due to its peak-shaving and valley-filling capabilities, the installed capacity of energy storage systems is increasing to cope with fluctuations in market electricity prices. However, because the design of peak-shaving and valley-filling strategies for energy storage systems is relatively simple, it is difficult to achieve optimal economic benefits in real time by using these strategies to regulate the power energy of the energy storage system. Therefore, several energy storage system scheduling models have been constructed in related technologies to obtain the optimal scheduling strategy. Among the various energy storage system scheduling models constructed in related technologies, the commonly used heuristic optimization algorithms are prone to getting trapped in local optima, which limits the convergence speed and accuracy of the model. Compared with heuristic optimization algorithms, the near-end strategy optimization algorithm is increasingly used in energy storage system scheduling models due to its global exploration capability. The near-end strategy optimization algorithm can overcome the shortcomings of heuristic optimization algorithms in getting trapped in local optima, achieve the global optimum of economic scheduling of the energy storage system, improve the economic efficiency of the energy storage system, and thus improve the return on investment of the energy storage system. However, due to the complexity of the operation of energy storage systems, the scheduling model is easily affected by various factors such as the charging and discharging actions of the energy storage system, electricity price, and electricity quantity during the training process using the near-end strategy optimization algorithm, resulting in deviations and thus lower convergence accuracy of the energy storage system scheduling model.

[0035] Therefore, to improve the convergence accuracy of the energy storage system scheduling model and ensure its convergence speed, this specification proposes a training method, apparatus, electronic device, and storage medium for an energy storage system scheduling model. This specification's embodiments determine the main reward data for the energy storage system scheduling model based on the electricity cost data of the energy storage system during the period between a first time point and a second time point. The first time point is earlier than the second time point. The branch penalty data for the energy storage system scheduling model are determined based on the charging and discharging action deviation data of the energy storage system during the period between the first and second time points. The charging and discharging action deviation data describes the consistency between the charging and discharging actions of the energy storage system and the direction of change of the energy storage system's state variables. The parameters of the energy storage system scheduling model are updated based on the main reward data and the branch penalty data to obtain the target energy storage system scheduling model.

[0036] This specification provides a scenario example of a method for training an energy storage system scheduling model, using a residential energy storage system as an example. This method is applied to... Figure 1a The energy storage system scheduling model training system shown includes a user-level energy storage system 102, a client 104, and a server 106. The server is used to train the energy storage system scheduling model. In this scenario example, the near-end policy optimization algorithm can be used to build the energy storage system scheduling model.

[0037] For details on the specific components of the household storage system 102, please refer to [link / reference]. Figure 1b The household energy storage system 102 must include at least an energy storage system 110 and a household electrical load 120. The electrical load 120 is connected to a load meter 130; the load power of the electrical load can be obtained by multiplying the current and power supply voltage collected by the load meter. The energy storage system is typically equipped with a local controller 150, which can collect energy storage output data and determine the energy storage capacity, among other data. The energy storage output data can be understood as the output power. A grid connection point meter 140 is also generally configured on the grid bus.

[0038] In this scenario example, the home energy storage system can communicate with the server to provide sample data for training the energy storage system scheduling model. The sample data includes environmental state variables and operational variables of the energy storage system. The environmental state variables can include electricity consumption variables, actual electricity price variables, predicted electricity price variables, and load power variables. The operational variables of the energy storage system represent its charging and discharging actions.

[0039] The server can use the environmental state variables and action variables of the sample data as inputs to the value network and action network of the energy storage system scheduling model. Based on the environmental state variables and action variables, the electricity cost data for the energy storage system during the period between the first and second time points is determined. This electricity cost data is then used to determine the main reward data for the energy storage system scheduling model. The first time point is earlier than the second time point. The first and second time points can be the time intervals before and after a scheduling interval within a scheduling cycle. Similarly, based on the environmental state variables and action variables, the charging and discharging action deviation data for the energy storage system during the period between the first and second time points is determined. This charging and discharging action deviation data is then used to determine the branch penalty data for the energy storage system scheduling model. The charging and discharging action deviation data describes the consistency between the charging and discharging actions of the energy storage system and the direction of change of the energy storage system's state variables. It can be understood that the main reward data is to ensure the minimum electricity cost for users in the energy storage system. The branch penalty data ensures the consistency between the charging and discharging actions of the energy storage system and the state variables, reducing the occurrence of unreasonable actions or minor charging and discharging actions that could lead to suboptimal energy storage system revenue and increased user electricity costs.

[0040] Subsequently, with the objective of maximizing the difference between the main reward data and the branch penalty data, the parameters of the energy storage system scheduling model are updated using policy gradient and temporal difference methods. This process is repeated iteratively to update the energy storage system scheduling model until it converges, yielding the target energy storage system scheduling model. In this scenario example, the trained target energy storage system scheduling model can be deployed on the client side.

[0041] In this scenario example, a target energy storage system scheduling model is deployed on the client. The client can communicate with the user's energy storage system to receive the system's current environmental state variables. These environmental state variables serve as input to the target energy storage system scheduling model to derive the scheduling strategy for the next scheduling interval. The scheduling strategy includes the energy storage system's action variables. The client schedules the user's energy storage system according to the scheduling strategy to optimize its economic efficiency, thereby minimizing the user's electricity costs.

[0042] This specification provides an embodiment of a method for training an energy storage system scheduling model. Please refer to [link / reference]. Figure 2 The model training method includes the following steps:

[0043] S210, Based on the electricity cost data of the energy storage system during the period between the first and second time points, determine the main reward data of the energy storage system scheduling model.

[0044] The first moment is earlier than the second moment.

[0045] S220, Based on the charging and discharging deviation data of the energy storage system during the period between the first and second moments, determine the branch penalty data of the energy storage system scheduling model.

[0046] Among them, the charge and discharge action deviation data is used to describe the consistency between the charge and discharge actions of the energy storage system and the direction of change of the energy storage system's state variables.

[0047] S230, update the parameters of the energy storage system scheduling model based on the main line reward data and the branch line penalty data to obtain the target energy storage system scheduling model.

[0048] In the embodiments of this specification, the actual power environment in which the energy storage system can be scheduled is used as the theoretical basis, and an initial energy storage system scheduling model is built based on the near-end strategy optimization algorithm. Simultaneously, the scheduling cycle of the energy storage system is determined according to actual needs. For example, in a residential scheduling system, a scheduling cycle is generally set to one day. Within each scheduling cycle, the energy storage system can be scheduled multiple times. The scheduling time interval can be determined based on the scheduling frequency of the energy storage system within a scheduling cycle. For example, using one day as a scheduling cycle, if the scheduling time interval is set to 15 minutes based on a preset scheduling frequency within a scheduling cycle, then the energy storage system will be scheduled 95 times within a scheduling cycle, that is, power scheduling will be performed on the energy storage system once every 15 minutes.

[0049] Therefore, in this embodiment, when training the energy storage system scheduling model, the relevant data for each scheduling time interval is used as sample data to update the parameters of the energy storage system scheduling model. The first time point and the second time point can be preceding or following a scheduling time interval, with the first time point being earlier than the second time point.

[0050] Specifically, when training the energy storage system scheduling model, the energy storage system can act as an agent in a deep reinforcement learning environment. The agent executes the optimal action and interacts with the agent's environment, obtaining immediate rewards and the next environment state variable. The current environment state variable, the optimal action variable, the immediate reward, and the next environment state variable constitute the transition experience, which is stored in a buffer.

[0051] In residential energy storage systems, the primary goal of scheduling the system is to minimize user electricity costs. The immediate reward is determined by the reward function of the energy storage system scheduling model and represents the electricity cost over a scheduling time interval. Therefore, when updating the parameters of the energy storage system scheduling model using the immediate reward, the parameters can be adjusted based on the reward function settings, aiming to maximize or minimize the immediate reward.

[0052] In the embodiments described in this specification, the reward is determined using both mainline reward data and branchline penalty data. Specifically, the mainline reward data is determined based on the electricity cost data of the energy storage system during the time interval between a first moment and a second moment. Thus, the mainline reward data is determined by the electricity cost data within a scheduling time interval. The mainline reward data can be used to maintain the primary objective of the energy storage system scheduling model: ensuring the lowest possible user electricity cost. For example, to ensure the lowest possible user electricity cost over a scheduling cycle (e.g., one day), the mainline reward data can be determined from the electricity cost data within a scheduling time interval of that cycle. In this way, as long as the electricity cost data represented by the mainline reward data for each scheduling time interval tends to be the lowest, the user electricity cost for a given cycle can be guaranteed to tend to be the lowest.

[0053] However, in some cases, when determining the optimal action of an energy storage system in the next scheduling time interval using an energy storage system scheduling model, it is easily affected by other data, leading to suboptimal or unreasonable scheduling strategies. For example, changes in electricity consumption, electricity prices, or the accuracy of predicted electricity prices can all affect model training. Therefore, if the reward is determined solely by the main reward data, the model's convergence accuracy will be low, and the convergence speed will also be relatively slow. To reduce the influence of other data on the charging and discharging actions of the energy storage system, the embodiments in this specification incorporate branch penalty data.

[0054] Specifically, the branch penalty data is determined based on the charging and discharging deviation data of the energy storage system during the period between the first and second time points. This deviation data can be determined using state variables that influence the relevant environment of the energy storage system, such as the energy storage system's power quantity, electricity price variables before and after the current scheduling interval, and the predicted electricity price variable for the next time point. The charging and discharging deviation data describes the consistency between the energy storage system's charging and discharging actions and the direction of change of its state variables. By adding branch penalty data to the main reward data, the inconsistency between the direction of change of the energy storage system's state variables and its charging and discharging actions can be reduced or avoided, preventing poor model training and insufficient convergence accuracy of the energy storage system scheduling model. This can lead to unreasonable or inaccurate scheduling strategies, which, when used to schedule the energy storage system in the next scheduling interval, may not optimize the system's benefits, resulting in increased electricity costs for users.

[0055] Through the above embodiments, when updating the parameters of the initially constructed energy storage system scheduling model, in addition to ensuring that the main reward data guarantees the main objective of the energy storage system, supplementary penalty data is added. This reduces the possibility of poor model training performance due to changes in or inaccuracies of state variables in the environment where the energy storage system is located. Iteratively updating the parameters of the energy storage system scheduling model using an immediate reward composed of main reward data and supplementary penalty data improves the model's convergence accuracy and speed, resulting in a more accurate target energy storage system scheduling model.

[0056] In some embodiments of this specification, the environmental state variables of the energy storage system include the energy quantity variable. Before determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first and second time points, the model training method further includes: determining the energy quantity action deviation data of the energy storage system during the period between the first and second time points based on the first energy quantity data of the energy quantity variable at the first time point, the second energy quantity data of the energy quantity variable at the second time point, and the action data of the energy storage system's action variable at the first time point.

[0057] In the embodiments of this specification, charge / discharge action deviation data includes energy level action deviation data. Environmental state variables are the relevant state variables of the energy storage system in the deep reinforcement learning environment, including energy level variables. Energy level variables represent the SOC value of the energy storage system at different times, which can be expressed as SOC. t This represents the energy storage system's charge data at time t. The action variables of the energy storage system represent its charging and discharging actions at different times, which can be represented by a. t This represents the action data of the energy storage system at time t. Positive action data indicates that the energy storage system is performing a charging action, and negative action data indicates that the energy storage system is performing a discharging action. Generally, a t The value range is [-1, 1].

[0058] In some cases, if the direction of the energy storage system's power change and the direction of its charging and discharging actions are inconsistent during a scheduling time interval, the energy storage system may experience additional minor charging and discharging actions.

[0059] Therefore, based on the energy quantity variables of the energy storage system before and after a scheduling time interval and the action variables of the energy storage system during that scheduling time interval, the energy quantity action deviation data during that scheduling time interval can be determined. Specifically, within a scheduling time interval, the product of the difference between the first energy quantity data of the energy storage system at the first moment and the second energy quantity data at the second moment, and the action data of the energy storage system at the first moment, is used to determine whether the charging and discharging direction of the energy storage system is consistent with the direction of energy quantity change during that scheduling time interval.

[0060] The energy storage system's energy variable is the first energy data at the first moment, using SOC. t This means that the second battery level data can be represented by the SoC. t+1 Represented by 'a'. Correspondingly, the action data of the energy storage system's action variables at the first moment are represented by 'a'. t This is indicated by the product of the difference between the first and second battery level data and the action data at the first moment (soc). t -soc t+1 )×a t This allows us to determine the deviation data of the energy storage system's charge / discharge behavior during the time interval between the first and second moments. By adding this penalty factor for the charge / discharge deviation data, we can ensure the consistency between the direction of the energy storage battery's charging / discharging behavior and the direction of charge / discharge change. This reduces the likelihood of decreased energy storage system efficiency and increased user electricity costs due to minor charging / discharging actions.

[0061] In some embodiments of this specification, the environmental state variables of the energy storage system include the actual electricity price variable. Before determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first and second time points, the model training method further includes: determining the electricity price action deviation data of the energy storage system during the period between the first and second time points based on the first actual electricity price data of the actual electricity price variable at the first time point, the second actual electricity price data of the actual electricity price variable at the second time point, and the action data of the energy storage system's action variable at the first time point. The charging and discharging action deviation data includes the electricity price action deviation data.

[0062] In the embodiments of this specification, the charging and discharging action deviation data includes electricity price action deviation data. Environmental state variables are the state variables related to the energy storage system in the deep reinforcement learning environment, and also include actual electricity price variables. Actual electricity price variables represent actual electricity price data at different times, which can be represented by the price variable. t This represents the actual electricity price data at time t.

[0063] In some situations, the scheduling of charging and discharging operations of energy storage systems is primarily influenced by changes in actual electricity prices. In the event of a sudden surge in electricity prices, the optimal action of the energy storage system should be discharging to supply power to the user's load. Therefore, a penalty term based on electricity price deviation data can be added to ensure that the optimal action of the energy storage system aligns with the direction of actual electricity price changes.

[0064] Specifically, within a scheduling time interval, the optimal action of the energy storage system and whether it is consistent with the direction of change of the actual electricity price can be determined by multiplying the difference between the first actual electricity price data at the first moment and the second actual electricity price data at the second moment with the action data of the energy storage system at the first moment.

[0065] The first actual electricity price data at the first moment can be represented by the price. t This means that the second actual electricity price data can be represented by price. t+1 This can be expressed as follows: The price can be calculated by multiplying the difference between the second and first actual electricity prices by the energy storage system's action data at the first moment (price). t+1 -price t )×a t This method determines the electricity price deviation data for the energy storage system during the period between the first and second time points. Including this electricity price deviation data ensures that the optimal action of the energy storage system aligns with the direction of electricity price changes.

[0066] In some embodiments of this specification, the environmental state variables of the energy storage system include actual electricity price variables and predicted electricity price variables. Before determining the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first and second time points, the model training method further includes: determining the electricity price prediction action deviation data of the energy storage system during the period between the first and second time points based on the predicted electricity price data of the predicted electricity price variable at the second time point, the second actual electricity price data of the actual electricity price variable at the second time point, and the action data of the energy storage system's action variable at the first time point. The charging and discharging action deviation data includes the electricity price prediction action deviation data.

[0067] In the embodiments of this specification, the charging and discharging action deviation data includes electricity price prediction action deviation data. Environmental state variables are the state variables related to the energy storage system in the deep reinforcement learning environment, and also include predicted electricity price variables. The predicted electricity price variable represents the predicted electricity price data for the next time step, which can be used... This represents the predicted electricity price data for the next time step, as predicted at time t.

[0068] In some cases, predicted electricity prices can be obtained from historical electricity price data using a price prediction model, but their accuracy is highly uncertain. Furthermore, if the predicted electricity price data deviates significantly from the actual electricity price data at any given moment, it can lead to unreasonable charging and discharging actions in the energy storage system. Therefore, adding data on the deviation of electricity price prediction actions can help avoid or reduce the occurrence of such situations.

[0069] Specifically, within a scheduling time interval, a penalty term in the direction of the predicted electricity price can be constructed by multiplying the absolute difference between the predicted electricity price data of the predicted electricity price variable at the second time moment and the second actual electricity price data of the actual electricity price variable at the second time moment with the action data of the energy storage system at the first time moment, so as to reduce unreasonable charging and discharging actions of the energy storage system.

[0070] The actual electricity price variable is represented by the second actual electricity price data at the second time point, using the price variable. t+1 This means that the predicted electricity price data at the second time point can be used. This can be expressed as follows: The absolute difference between the predicted electricity price data and the actual electricity price data at the second time step can be multiplied by the product of the energy storage system's action variables and the action data at the first time step. To determine the deviation data of the electricity price forecast action of the energy storage system during the period between the first and second time points.

[0071] In some embodiments of this specification, the branch penalty data of the energy storage system scheduling model is determined based on the charging and discharging action deviation data of the energy storage system during the time period between the first time and the second time. This includes: calculating the branch penalty data of the energy storage system scheduling model by weighting and summing the energy action deviation data, the electricity price action deviation data, and the electricity price prediction action deviation data according to the first weighting coefficient corresponding to the energy action deviation data, the second weighting coefficient corresponding to the electricity price action deviation data, and the third weighting coefficient corresponding to the electricity price prediction action deviation data.

[0072] In the embodiments of this specification, the charging and discharging operation deviation data includes power operation deviation data, electricity price operation deviation data, and electricity price prediction operation deviation data. Power operation deviation data, electricity price operation deviation data, and electricity price prediction operation deviation data can be considered as three penalty terms in the energy storage system scheduling model. When determining the branch penalty data using these three penalty terms, corresponding weight coefficients can be assigned to the three penalty terms based on the emphasis of the environment in which the energy storage system is located.

[0073] Using r branch This represents the branch penalty data. The branch penalty data can be obtained by weighted summing of the three penalty terms. Where α2 represents the first weighting coefficient; α3 represents the second weighting coefficient; and α4 represents the third weighting coefficient.

[0074] In some embodiments of this specification, the environmental state variables of the energy storage system include actual electricity price variables and load power variables. Before determining the main reward data of the energy storage system scheduling model based on the electricity price data of the energy storage system during the period between the first and second time points, the model training method further includes: determining the electricity price data based on the first actual electricity price data of the actual electricity price variable at the first time point, the first load power of the load power variable at the first time point, the second load power of the load power variable at the second time point, and the action data of the energy storage system's action variables at the first time point.

[0075] In the embodiments of this specification, the environmental state variables of the energy storage system include the actual electricity price variable and the load power variable. The load power variable represents the load power of the electrical load.

[0076] Electricity charges are determined by multiplying electricity consumption by the electricity price. Therefore, in the embodiments of this specification, to determine the electricity charges consumed by a user within a scheduling time interval, it is necessary to first determine the user's electricity consumption within that scheduling time interval. Electricity consumption can be determined based on the load power of the electrical load and the duration of electricity consumption. The duration of electricity consumption is the length of a scheduling time interval. The load power of the electrical load can be determined based on the first load power variable at a first moment, the second load power variable at a second moment, and the action data of the energy storage system at the first moment.

[0077] Specifically, within a scheduling time interval, the first load power at the first moment and the second load power at the second moment can be determined using load meters. However, since the load power at either moment cannot represent the load power over the entire scheduling time interval, the average of the first and second load powers can be taken as the load power of the electrical load without external power influence. In some cases, the charging and discharging operations of the energy storage system directly affect the load power of the electrical load. Therefore, it is also necessary to determine the power impact of the energy storage system on the electrical load based on the system's operation data within this scheduling time interval. The operation data of the energy storage system at the first moment is a. t This indicates the charging and discharging actions of the energy storage system during this scheduling interval. It utilizes the rated power P of the energy storage system. ess Action data of energy storage system a t The product of these factors determines the power impact of the energy storage system on the electrical load. Based on the above theory, the load power during a scheduling time interval can be expressed as... Where, p load,t p represents the first load power; load,t+1 This indicates the power of the second load.

[0078] Using Δt to represent the duration of a scheduling interval, the electricity cost data for a scheduling interval is determined based on the first actual electricity price data at the first moment, the first load power data at the first moment, the second load power data at the second moment, and the action data of the energy storage system at the first moment.

[0079] In some embodiments of this specification, the main reward data of the energy storage system scheduling model is determined based on the electricity cost data of the energy storage system during the time period between the first time and the second time. This includes: multiplying the electricity cost data with the fourth weight coefficient corresponding to the electricity cost data to obtain the main reward data of the energy storage system scheduling model.

[0080] In the embodiments of this specification, the mainline reward data and the branchline penalty data constitute the immediate reward during the training of the energy storage system scheduling model. Each of the three penalty terms in the branchline penalty data has a corresponding weight coefficient. Correspondingly, the electricity cost data constituting the mainline reward data can also be assigned a corresponding weight coefficient based on the importance of the mainline reward data during the training process. Then, the mainline reward data of the energy storage system scheduling model is obtained by multiplying the electricity cost data with the fourth weight coefficient corresponding to the electricity cost data.

[0081] In some cases, the mainline reward data is related to the primary objective of energy storage system scheduling. However, in the embodiments of this specification, the immediate reward is used to update the parameters of the energy storage system scheduling model with the objective of maximizing the reward value. Therefore, the mainline reward data is obtained by multiplying the negative of the electricity cost data with the corresponding fourth weighting coefficient.

[0082] Specifically, main storyline reward data Where α1 represents the fourth weighting coefficient.

[0083] In some embodiments of this specification, the predicted electricity price data corresponding to the second time point is obtained by predicting using an electricity price prediction model. Please refer to... Figure 3 The training process for the electricity price prediction model includes:

[0084] S310, construct a training sample set for electricity prices.

[0085] S320 uses the training sample set to train the initial prediction model, resulting in an electricity price prediction model. The initial prediction model employs a long short-term memory network.

[0086] Since the predicted electricity price variable is one of the environmental state variables in the energy storage system scheduling model, in order to improve the accuracy of the predicted electricity price data and enhance the agent's ability to perceive the environment during training, the embodiments in this specification use a long short-term memory (LSTM) network from deep learning to construct an initial electricity price prediction model.

[0087] Specifically, a training sample set for electricity prices is constructed, comprising multiple training samples. These training samples can utilize historical electricity price data within a specified time interval. The specified time interval is the period closest to the point in time to be predicted. For example, historical electricity price data from the past 30 days can be used to construct the training samples.

[0088] In the embodiments of this specification, historical electricity price data includes at least historical real electricity price data within a specified time interval. To improve prediction accuracy, multiple feature variables can be constructed based on the historical real electricity price data within the specified time interval, such as the highest daily electricity price, the lowest daily electricity price, the average daily electricity price, and the daily electricity price volatility within the specified time interval. The training samples include historical real electricity price data within the specified time interval and data on the constructed feature variables. In addition, the training samples may also include time variables within the specified time interval, such as date variables, date-related variables used to represent working days or non-working days, etc.

[0089] The constructed training sample set is used as input to the Long Short-Term Memory (LSTM) network, and the predicted electricity price data at the time to be predicted is used as output. The LSM network is trained using the Keras open-source neural network library. Training termination conditions, such as the number of iterations, are set, and the electricity price prediction model is obtained after training. It can be understood that when training the initial electricity price prediction model, the time to be predicted is a past time, and the actual electricity price data at the time to be predicted is used as the predicted electricity price data for that time to train the initial electricity price prediction model.

[0090] Using this deep learning-based electricity price prediction model, the predicted electricity price data for the time to be predicted can be obtained based on historical real electricity price data within a specified time interval before the time to be predicted, multiple feature variables, and time variables.

[0091] The following is combined Figure 4 A detailed explanation of the network structure of Long Short-Term Memory (LSTM) networks is provided. Please refer to [link / reference]. Figure 4Long short-term memory (LSTM) networks are an improved structure of recurrent neural networks (RNNs). By adding an extra unit C in the hidden layers to store long-term states, they possess long-term memory capabilities and can effectively avoid the gradient vanishing problem, making them suitable for handling time series prediction problems.

[0092] The input to an LSTM unit includes the current input vector X. t The previous cell state C t-1 and the hidden layer state h at the previous time step t-1 The output is the current cell state C. t and the current hidden layer state h t Its interior consists of three doors, namely the forgetting door and the forgetting door. t Output gate i t and output gate o t To control the discarding and inheritance of information.

[0093] Their calculation formulas are as follows: f t =σ(W f ·[h t-1 ,x t ]+b f );i t =σ(W i ·[h t-1 ,x t ]+b i );o t =σ(W o ·[h t-1 ,x t ]+b0). Among them, W f W i W o These are the weight matrices for the forget gate, input gate, and output gate, respectively. t-1 ,x t The symbol ] indicates that the two vectors are concatenated into a longer vector. f ,b i ,b o These are the bias vectors for the forget gate, input gate, and output gate, respectively. σ is the sigmoid activation function, which transforms the output to the [0,1] interval. It can be understood that the above training samples constitute the output vector X. t .

[0094] Current state of memory cell C t From the previous time unit state C t-1It is calculated together with the intermediate unit state of the current input. tanh is the hyperbolic tangent activation function, which transforms the output to the interval [-1, 1]. * is the Hadamard product. For temporary states, W C and b C The weight matrix and bias vector of the tanh layer are calculated using the following formulas:

[0095]

[0096]

[0097] Finally, the current output h of the LSTM unit t For: h t =o t *tanh(C t ).

[0098] The LSTM model training uses the backpropagation algorithm for error, with hidden layer h. t-1 The error is due to h t Decision. LSTM cell state C t-1 By C t Decision. And C t The error consists of two parts, one part being C t+1 The other part is h t h t When updating, h needs to be considered t+1 This allows us to calculate the gradient at any time step from time t, and then update the weight coefficients using stochastic gradient descent.

[0099] According to the above embodiments, by using a deep learning-embedded electricity price prediction model to predict the predicted electricity price data at the second moment, and using it as one of the environmental state variables for training the energy storage system scheduling model, the agent's ability to perceive the environment can be enhanced, thereby improving the stability and adaptability of the energy storage system scheduling model.

[0100] In some embodiments of this specification, updating the parameters of the energy storage system scheduling model based on the main reward data and the branch penalty data includes: using the environmental state variables and action variables of the energy storage system as inputs to the value network and action network of the energy storage system scheduling model, and using the policy gradient and time-series difference method to update the parameters of the energy storage system scheduling model with the goal of maximizing the difference between the main reward data and the branch penalty data.

[0101] In the embodiments of this specification, before constructing the initial energy storage system scheduling model using the near-end policy optimization algorithm, the state observation space, action space, and reward function required by the near-end policy optimization algorithm are first constructed.

[0102] The state observation space is used to acquire the environmental state variables for training the initial energy storage system scheduling model. In the embodiments of this specification, the environmental state variables include load power variables, actual electricity price variables, predicted electricity price variables, and energy storage system power quantity variables. During a training process, the environmental state variables, as inputs to the energy storage system scheduling model, may include the load power at the first time step, the actual electricity price data at the first time step, the predicted electricity price data at the second time step, and the energy storage system power quantity data at the first time step. s t Represents environmental state variables.

[0103] The action space is used to determine the range of action variables of the energy storage system. In the embodiments of this specification, the action space a t ∈[-1,1],a t A positive value indicates that the energy storage system is performing a charging action; a t A negative value indicates that the energy storage system is performing a discharge action.

[0104] The difference between the main reward data and the branch penalty data is the value of the reward function, i.e., the immediate reward. The immediate reward is determined by the model based on the reward function after the environmental state variables and action variables of the energy storage system are input into the value network and action network of the energy storage system scheduling model.

[0105] In the embodiments described in this specification, the mainline reward data is determined by multiplying the negative of the electricity cost data by the fourth weighting coefficient, and the primary objective of constructing the energy storage system scheduling model is to minimize the user's electricity expenditure. Therefore, during model training, the objective can be to maximize the reward value in real time, that is, to maximize the difference between the mainline reward data and the branch penalty data, and the parameters of the energy storage system scheduling model can be updated using policy gradient and time-series difference methods.

[0106] Energy storage systems can function as intelligent agents in deep reinforcement learning environments. The agent executes the optimal action a. t Interacting with the agent's environment allows it to obtain immediate rewards and the next environmental state variables. Current environment state variable, optimal action variable a t Even the reward and the next environment state variable s t+1 This constitutes a transition experience, which is then stored in a buffer. The soc variable in the next environment state variable... t+1 It is determined during the training process in the following ways:

[0107]

[0108] Where, η charge E represents the charging efficiency of the energy storage system. ess Indicates the rated capacity of the energy storage system; SOCmax Indicates the maximum energy capacity of the energy storage system; SOC min Indicates the minimum energy capacity of the energy storage system; η discharge This indicates the discharge efficiency of the energy storage system.

[0109] In some embodiments of this specification, the model training method further includes: determining the objective function and constraints for the minimum system operating electricity cost of the energy storage system scheduling model; and constructing the reward function for the near-end policy optimization algorithm using the objective function and constraints. The reward function includes a main reward function for outputting main reward data and a branch penalty function for outputting branch penalty data.

[0110] Before constructing the energy storage system scheduling model, the objective function and constraints of the energy storage system scheduling model can be constructed based on the actual power environment in which the energy storage system is located and the main objective of minimizing the system's operating electricity costs.

[0111] In the embodiments of this specification, the primary objective of minimizing system operating electricity costs is to minimize electricity costs over a scheduling cycle. For example, if a scheduling cycle is one day, the objective function of the energy storage system scheduling model is constructed with the goal of minimizing daily electricity costs. The scheduling time interval Δt within a scheduling cycle is determined based on the scheduling frequency. Based on the aforementioned process of determining electricity cost data, the objective function can be constructed as follows: Where T represents the number of scheduling cycles in that scheduling period. For example, if Δt is 15 minutes, then a 24-hour day includes 96 Δt cycles, therefore T = 96.

[0112] The objective function represents minimizing the electricity cost over a scheduling cycle, while the main reward function represents the electricity cost over a scheduling time interval. Therefore, the main reward function can be constructed based on the objective function. Simultaneously construct the branch penalty function determined by the three penalty terms.

[0113] The value of the reward function is determined by the difference between the main reward data and the side penalty data; therefore, the reward function r... t =r master -r branch During the training of the energy storage system scheduling model, the main reward data can be obtained using the main reward function based on the input environmental state variables and action variables. Similarly, the branch penalty data can be obtained using the branch penalty function based on the input environmental state variables and action variables.

[0114] In some embodiments of this specification, the constraints include energy storage system capacity constraints and energy storage system actual output power constraints. The energy storage system capacity constraints are based on the upper and lower limits of the energy storage system's state of charge (SOC). max and soc minThe actual output power constraint of the energy storage system is constructed based on the premise that the actual output power of the energy storage system is less than or equal to the rated power of the energy storage system.

[0115] The actual output power of the energy storage system can be determined by its rated power P. ess With the action variable a of the energy storage system t The product of these factors determines the actual output power. The absolute value of the actual output power should be less than the rated power P of the energy storage system. ess .

[0116] Specifically, the constraints of the energy storage system scheduling model include the following:

[0117]

[0118] In the embodiments of this specification, constraints are used to impose conditions on the calculation of the reward function.

[0119] It should be noted that the energy storage system scheduling model in the embodiments of this specification is applicable to energy storage systems without photovoltaic (PV) capabilities. For energy storage systems with PV functionality, the load power variable in the environmental state variables also needs to include PV power generation data.

[0120] Please see Figure 5 In some embodiments of this specification, the energy storage system scheduling model training method may further include the following steps:

[0121] S510, construct a training sample set for electricity prices.

[0122] S520 uses the training sample set to train the initial prediction model to obtain the electricity price prediction model.

[0123] S530 defines the objective function and constraints for the minimum system operating electricity cost of the energy storage system scheduling model.

[0124] S540 utilizes the objective function and constraints to construct the reward function for the near-end policy optimization algorithm.

[0125] The reward function includes a main reward function for outputting main reward data and a branch penalty function for outputting branch penalty data.

[0126] S550 determines the main reward data for the energy storage system scheduling model based on the electricity cost data of the energy storage system during the period between the first and second time points.

[0127] S560 determines the branch penalty data of the energy storage system scheduling model based on the charging and discharging deviation data of the energy storage system during the period between the first and second time points.

[0128] S570 uses the environmental state variables and action variables of the energy storage system as inputs to the value network and action network of the energy storage system scheduling model. With the goal of maximizing the difference between the main reward data and the branch penalty data, it uses policy gradient and time-series difference methods to update the parameters of the energy storage system scheduling model.

[0129] In one specific embodiment, an energy storage system without photovoltaic (PV) power is used as an example. First, historical real electricity price data and load power data for the past year are acquired. Using historical real electricity price data from the past 90 days, feature variables are constructed, such as the highest daily electricity price, the lowest daily electricity price, the average daily electricity price, and the daily electricity price volatility, to obtain the input and output set for training the electricity price prediction model. The input set includes time variables, historical real electricity price data from the past 90 days, and the constructed feature variables; the output set is the electricity price for the next household-storage dispatch cycle. The model is trained based on a deep learning LSTM (long short term memory) neural network.

[0130] Here we can assume that the update interval of each Long Short-Term Memory network is 30 days, that is, a new electricity price prediction model is obtained every 30 days.

[0131] A deep reinforcement learning environment for the energy storage system was built using Python. The state observation space, action space, and reward function were established, with the latter three defined as described in the previous embodiment. The energy storage system scheduling model was iteratively trained using policy gradient and temporal difference methods. The optimal training model was saved to obtain the target energy storage system scheduling model. During actual scheduling using this model, real-time environmental state variables were substituted into the model, and intelligent control of the energy storage system was performed based on the model's output.

[0132] Corresponding to the above embodiments, this specification also proposes an energy storage system scheduling model training device. Please refer to... Figure 6 The training device includes:

[0133] The mainline reward data determination module 610 is used to determine the mainline reward data of the energy storage system scheduling model based on the electricity cost data of the energy storage system during the period between the first and second time points.

[0134] The first moment is earlier than the second moment.

[0135] The branch penalty data determination module 620 is used to determine the branch penalty data of the energy storage system scheduling model based on the charging and discharging action deviation data of the energy storage system during the period between the first time and the second time.

[0136] Among them, the charge and discharge action deviation data is used to describe the consistency between the charge and discharge actions of the energy storage system and the direction of change of the energy storage system's state variables.

[0137] Training module 630 is used to update the parameters of the energy storage system scheduling model based on the main reward data and the branch penalty data to obtain the target energy storage system scheduling model.

[0138] According to the energy storage system scheduling model training device of the embodiments of this specification, when updating the parameters of the initially constructed energy storage system scheduling model, in addition to ensuring that the main reward data can guarantee the main objective of the energy storage system, branch penalty data is added to reduce the possibility of poor model training effect due to changes in state variables or inaccuracies of state variables in the environment in which the energy storage system is located. Iteratively updating the parameters of the energy storage system scheduling model through an instantaneous reward composed of main reward data and branch penalty data can improve the convergence accuracy and convergence speed of the model, resulting in a target energy storage system scheduling model with higher accuracy.

[0139] Specific limitations regarding the energy storage system scheduling model training device can be found in the limitations of the energy storage system scheduling model training method described above, and will not be repeated here. Each module in the aforementioned energy storage system scheduling model training device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0140] Corresponding to the above embodiments, this specification also provides an electronic device.

[0141] Figure 7 This is a structural block diagram of an electronic device according to one embodiment of this specification. For example... Figure 7 As shown, the electronic device 700 includes a memory 704, a processor 702, and an energy storage system scheduling model training program 706 stored in the memory 704 and capable of running on the processor 702. When the processor 702 executes the energy storage system scheduling model training program 706, it implements the energy storage system scheduling model training method of any of the above embodiments.

[0142] According to the embodiments of this specification, when the processor 702 executes the energy storage system scheduling model training program 706, the parameters of the energy storage system scheduling model are iteratively updated using an instant reward composed of main reward data and branch penalty data. This can improve the convergence accuracy and convergence speed of the model, resulting in a target energy storage system scheduling model with higher accuracy.

[0143] Corresponding to the above embodiments, this specification also proposes a computer-readable storage medium storing an energy storage system scheduling model training program thereon. When the energy storage system scheduling model training program is executed by a processor, it implements the energy storage system scheduling model training method of any of the above embodiments.

[0144] According to the computer-readable storage medium of the embodiments of this specification, when the energy storage system scheduling model training program is executed by the processor, the parameters of the energy storage system scheduling model are iteratively updated by the instantaneous reward composed of main reward data and branch penalty data, which can improve the convergence accuracy and convergence speed of the model, and obtain a target energy storage system scheduling model with higher accuracy.

[0145] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0146] It should be understood that various parts of this specification can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0147] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0148] In the description of this specification, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this specification and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this specification.

[0149] Furthermore, the terms "first," "second," etc., used in the embodiments of this specification are for descriptive purposes only and should not be construed as indicating or implying relative importance, or implicitly specifying the number of technical features indicated in this embodiment. Therefore, features defined with terms such as "first" and "second" in the embodiments of this specification can explicitly or implicitly indicate that the embodiment includes at least one of those features. In the description of this specification, the word "multiple" means at least two or more, such as two, three, four, etc., unless otherwise explicitly specified in the embodiments.

[0150] In this specification, unless otherwise explicitly specified or limited in the embodiments, the terms "installation," "connection," "joining," and "fixing," etc., appearing in the embodiments, should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral part; it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication between two components, or the interaction between two components. Those skilled in the art will be able to understand the specific meaning of the above terms in this specification based on the specific implementation.

[0151] In this specification, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, "above," "over," and "on top" of the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.

[0152] Although embodiments of this specification have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting this specification. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this specification.

Claims

1. A method for training a dispatch model of an energy storage system, characterized in that, The method comprises: determining a minimum system operation electricity cost objective function and constraint conditions of an energy storage system scheduling model; using the objective function and the constraint conditions to construct a reward function of a proximal policy optimization algorithm; wherein the reward function comprises a mainline reward function for outputting mainline reward data and a sideline penalty function for outputting sideline penalty data; determining mainline reward data of the energy storage system scheduling model according to electricity cost data of the energy storage system in a period between a first time and a second time; wherein the first time is earlier than the second time; determining sideline penalty data of the energy storage system scheduling model according to charge-discharge action deviation data of the energy storage system in the period between the first time and the second time; wherein the charge-discharge action deviation data is used to describe consistency of a charge-discharge action of the energy storage system with a change direction of an environmental state variable of the energy storage system; updating parameters of the energy storage system scheduling model according to the mainline reward data and the sideline penalty data to obtain a target energy storage system scheduling model; determining mainline reward data of the energy storage system scheduling model according to electricity cost data of the energy storage system in a period between a first time and a second time and the mainline reward function, comprises: performing product operation on the electricity cost data by using a fourth weight coefficient corresponding to the electricity cost data to obtain the mainline reward data of the energy storage system scheduling model; updating parameters of the energy storage system scheduling model according to the mainline reward data and the sideline penalty data, comprises: taking an environmental state variable of the energy storage system and an action variable of the energy storage system as inputs of a value network and an action network of the energy storage system scheduling model, taking maximum difference of the mainline reward data and the sideline penalty data as a target, and updating the parameters of the energy storage system scheduling model by using a policy gradient and a time difference method.

2. The method of claim 1, wherein, The environmental state variable of the energy storage system comprises an electricity variable; before determining the sideline penalty data of the energy storage system scheduling model according to the charge-discharge action deviation data of the energy storage system in the period between the first time and the second time, the method further comprises: determining electricity action deviation data of the energy storage system in the period between the first time and the second time according to first electricity data of the electricity variable at the first time, second electricity data of the electricity variable at the second time, and action data of the action variable of the energy storage system at the first time; wherein the charge-discharge action deviation data comprises the electricity action deviation data.

3. The method of claim 1, wherein, The environmental state variable of the energy storage system comprises an actual electricity price variable; before determining the sideline penalty data of the energy storage system scheduling model according to the charge-discharge action deviation data of the energy storage system in the period between the first time and the second time, the method further comprises: According to the first actual electricity price data of the actual electricity price variable at the first time, the second actual electricity price data of the actual electricity price variable at the second time, and the action data of the action variable of the energy storage system at the first time, determine electricity price action deviation data of the energy storage system in a time period between the first time and the second time; wherein the charge-discharge action deviation data comprises the electricity price action deviation data.

4. The method of claim 1, wherein, The environmental state variable of the energy storage system comprises an actual electricity price variable and a predicted electricity price variable; before the step of determining the branch penalty data of the energy storage system scheduling model according to the charge-discharge action deviation data of the energy storage system in a time period between the first time and the second time, the method further comprises: According to the predicted electricity price data of the predicted electricity price variable at the second time, the second actual electricity price data of the actual electricity price variable at the second time, and the action data of the action variable of the energy storage system at the first time, determine electricity price prediction action deviation data of the energy storage system in a time period between the first time and the second time; wherein the charge-discharge action deviation data comprises the electricity price prediction action deviation data.

5. The method of claim 1, wherein, The charge-discharge action deviation data comprises electricity quantity action deviation data, electricity price action deviation data, and electricity price prediction action deviation data; the step of determining the branch penalty data of the energy storage system scheduling model according to the charge-discharge action deviation data of the energy storage system in a time period between the first time and the second time comprises: According to the first weight coefficient corresponding to the electricity quantity action deviation data, the second weight coefficient corresponding to the electricity price action deviation data, and the third weight coefficient corresponding to the electricity price prediction action deviation data, perform weighted summation calculation on the electricity quantity action deviation data, the electricity price action deviation data, and the electricity price prediction action deviation data to obtain the branch penalty data of the energy storage system scheduling model.

6. The method of claim 1, wherein, The environmental state variable of the energy storage system comprises an actual electricity price variable and a load power variable; before the step of determining the main line reward data of the energy storage system scheduling model according to the electricity fee data of the energy storage system in a time period between the first time and the second time, the method further comprises: According to the first actual electricity price data of the actual electricity price variable at the first time, the first load power of the load power variable at the first time, the second load power of the load power variable at the second time, and the action data of the action variable of the energy storage system at the first time, determine the electricity fee data.

7. The method according to any one of claims 1 to 6, characterized in that, The predicted electricity price data corresponding to the second time is obtained by predicting through an electricity price prediction model, and a training process of the electricity price prediction model comprises: Construct a training sample set for electricity price; wherein the training sample set comprises a plurality of training samples; the training samples use historical electricity price related data in a specified time interval; the historical electricity price related data has a plurality of feature variables; the training samples comprise data on the plurality of feature variables; the plurality of feature variables comprise daily electricity price maximum value, daily electricity price minimum value, daily electricity price average value, and daily electricity price fluctuation rate; The initial prediction model is trained by using the training sample set, and the electricity price prediction model is obtained; wherein, the initial prediction model adopts a long short-term memory network.

8. The method of claim 1, wherein, The constraint conditions include the energy storage system power constraint condition and the actual output power constraint condition of the energy storage system; the energy storage system power constraint condition is constructed based on upper and lower limit values of the power of the energy storage system; and the actual output power constraint condition of the energy storage system is constructed based on that the actual output power of the energy storage system is less than or equal to the rated power of the energy storage system.

9. An energy storage system dispatch model training apparatus, comprising: The device comprises: The main line reward data determination module is configured to determine main line reward data of the energy storage system scheduling model according to electricity cost data of the energy storage system in a time period between a first time and a second time, including: multiplying the electricity cost data by a fourth weight coefficient corresponding to the electricity cost data to obtain the main line reward data of the energy storage system scheduling model; wherein, the first time is earlier than the second time. The branch line penalty data determination module is configured to determine branch line penalty data of the energy storage system scheduling model according to charge-discharge action deviation data of the energy storage system in the time period between the first time and the second time; wherein, the charge-discharge action deviation data is used to describe the consistency of the charge-discharge action of the energy storage system with the change direction of the environmental state variable of the energy storage system. The training module is configured to update parameters of the energy storage system scheduling model according to the main line reward data and the branch line penalty data to obtain a target energy storage system scheduling model, including: taking the environmental state variable of the energy storage system and the action variable of the energy storage system as inputs of a value network and an action network of the energy storage system scheduling model, maximizing the difference between the main line reward data and the branch line penalty data as a target, and updating the parameters of the energy storage system scheduling model by using a policy gradient and a time difference method. The reward function construction module is configured to determine a lowest system operation electricity cost objective function and constraint conditions of the energy storage system scheduling model, and construct a reward function of a proximal policy optimization algorithm by using the objective function and the constraint conditions; wherein, the reward function includes a main line reward function for outputting the main line reward data and a branch line penalty function for outputting the branch line penalty data.

10. An electronic device, comprising: The computer program is configured to be executed by the processor, and the processor implements the method according to any one of claims 1 to 8 when executing the computer program.

11. A computer readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the method according to any one of claims 1 to 8 when the computer program is running.

Citation Information

Patent Citations

  • Photovoltaic power station bidding optimization method based on consideration of energy storage demand response effect

    CN112348252A