A control method and device for a grid-based user energy storage system
By combining a reinforcement learning model with a policy network and a value network, the regulation and control problem of user energy storage systems in complex environments was solved, achieving efficient and stable battery charging and discharging decisions, and improving the system's operating efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2026-03-27
AI Technical Summary
Existing user energy storage systems are difficult to effectively regulate and control when facing complex and ever-changing operating environments, resulting in low operating efficiency and insufficient stability.
A pre-built energy storage control model trained with reinforcement learning is used, combined with a policy network and a value network, to make battery charging and discharging decisions by comprehensively considering electricity price, load demand, ambient temperature and battery health status.
By leveraging the decision-making capabilities of reinforcement learning models, complex factors such as electricity price fluctuations, load changes, and ambient temperature fluctuations can be effectively addressed, ensuring the efficient operation and long-term stability of users' energy storage systems.
Smart Images

Figure CN120387619B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric power, in particular to a control method and device for user energy storage system based on power grid. BACKGROUND
[0002] With the rapid development of new energy technology, especially the application of renewable energy such as wind and solar energy gradually popular, distributed generation and energy storage system in the modern power network plays an increasingly important role, the development of these technologies provides a new way to solve the environmental pollution problem caused by traditional energy supply, at the same time also promotes the energy structure to a more clean and sustainable direction.
[0003] Among them, the user energy storage system usually refers to the energy storage device installed at the user end of electric power, which can store and release electric energy when needed, for example, the main functions include peak clipping, improving power supply reliability, promoting renewable energy consumption and participating in power grid services, etc.
[0004] At present, the running environment of user energy storage system is very complex and changeable, such as the comprehensive influence of factors such as price fluctuation, load demand change, environmental temperature fluctuation, battery health condition, etc. In order to cope with such environment, how to provide an effective regulation and control method on energy storage to ensure the efficient operation and long-term stability of user energy storage system is a technical problem to be solved. SUMMARY
[0005] The present application provides a control method and device for user energy storage system based on power grid, the main purpose is to use the preset energy storage control model trained by reinforcement learning to make decision, considering the complex and changeable environmental factors such as price fluctuation, load demand change, environmental temperature fluctuation and battery health condition, so as to provide an effective regulation and control method on energy storage to ensure the efficient operation and long-term stability of user energy storage system.
[0006] In order to achieve the above purpose, the present application mainly provides the following technical scheme:
[0007] The first aspect of the present application provides a control method for user energy storage system based on power grid, the user energy storage system is composed of energy storage facilities deployed at the user end, the method comprises:
[0008] determining the current environment information of the energy storage facility, the current environment information at least includes: price information, power load information, environmental temperature information and battery state of charge information;
[0009] The preset energy storage control model is a pre-trained reinforcement learning model, the preset energy storage control model comprises a policy network and a value network, the policy network is used to provide a control action for the energy storage facility according to a perception interaction with the current environment information, and the value network is used to evaluate an expected return of the control action to assist in making a control decision for the energy storage facility.
[0010] According to the decision result, the energy storage facility is controlled to store or release electric energy.
[0011] The second aspect of the present application provides a control device of a user energy storage system based on a power grid, the user energy storage system is composed of an energy storage facility deployed at a user end, and the device comprises:
[0012] A determination unit is configured to determine current environment information of the energy storage facility, the current environment information at least comprises price information, power load information, environment temperature information and battery state of charge information.
[0013] A processing unit is configured to process the current environment information by using a preset energy storage control model, and output a decision result, the decision result is a charging and discharging control strategy for a battery in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model, the preset energy storage control model comprises a policy network and a value network, the policy network is used to provide a control action for the energy storage facility according to a perception interaction with the current environment information, and the value network is used to evaluate an expected return of the control action to assist in making a control decision for the energy storage facility.
[0014] A control unit is configured to control the energy storage facility to store or release electric energy according to the decision result.
[0015] The third aspect of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the control method of the user energy storage system based on the power grid.
[0016] The fourth aspect of the present application provides an electronic device, the device comprises at least one processor, and at least one memory connected with the processor, and a bus;
[0017] The processor, the memory and the bus complete communication with each other through the bus.
[0018] The processor is configured to invoke program instructions in the memory to perform the method for controlling a user energy storage system based on a power grid as described above.
[0019] By means of the technical solutions described above, the technical solutions provided by the application have at least the following advantages:
[0020] The application provides a method and device for controlling a user energy storage system based on a power grid. For a user energy storage system composed of energy storage facilities deployed at a user end, the application pre-trains a preset energy storage control model using a reinforcement learning framework. The model includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facilities based on perception interaction with current environmental information. The value network is used to evaluate the expected returns of the control actions to assist in making control decisions for the energy storage facilities. Then, in the process of using such a model to make energy storage decisions on the user energy storage system, the application first determines the current environmental information in which the energy storage facilities at the user end operate, such as at least including electricity price information, power load information, environmental temperature information, and battery state of charge information. Then, the preset energy storage control model pre-trained as above is used to perceive and interact with the current environmental information to obtain a decision result for energy storage.
[0021] Compared with existing methods for dealing with the complex and variable requirements of the operating environment of a user energy storage system, the application utilizes the reinforcement learning model to have interaction with the environment to learn optimization according to environmental feedback reward signals. The application uses a reinforcement learning model to make decisions on the user energy storage system, and more state variables (electricity price, power load, environmental temperature, and battery state of charge) are used in the model training process. This will enable the model to make decisions that take into account the fluctuations in electricity prices, changes in load demand, fluctuations in environmental temperature, and battery health conditions, which are complex and variable environmental factors. Thus, an effective regulation and control method for energy storage is provided to ensure the efficient operation and long-term stability of the user energy storage system.
[0022] The above description is only a summary of the technical solutions of the application. In order to more clearly understand the technical means of the application, the application can be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0023] Various other advantages and benefits will become apparent to those of ordinary skill in the art, upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of preferred embodiments, and are not meant to limit the present application. Moreover, the same reference numerals in the attached drawings indicate the same or similar components. In the drawings:
[0024] Figure 1A flow chart of a control method of a user energy storage system based on a power grid is provided for an embodiment of the present application.
[0025] Figure 2 A flow chart of another control method of a user energy storage system based on a power grid is provided for an embodiment of the present application.
[0026] Figure 3 A composition block diagram of a control device of a user energy storage system based on a power grid is provided for an embodiment of the present application.
[0027] Figure 4 A composition block diagram of another control device of a user energy storage system based on a power grid is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0028] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the application to those skilled in the art.
[0029] A user energy storage system generally refers to an energy storage device installed at the end of a power user. Such a system can store electrical energy and release it for use when needed. For example, at the user side, especially at the residential, commercial and industrial user side, the installed energy storage device can, but is not limited to, provide the following energy storage needs for the user:
[0030] (1) Residential users; solar + energy storage solution: Many families install a combination of solar photovoltaic panels (PV) and battery energy storage systems. For example, during the day when the sun is shining, the amount of solar power generated may exceed the immediate consumption needs of the family, and the excess energy can be stored in the battery for use at night or on cloudy days. This not only reduces dependence on grid power, but also further reduces electricity bills by participating in net metering programs.
[0031] (2) Commercial users; peak shaving: Commercial sites such as shopping malls, office buildings, etc. usually have obvious peak periods of electricity consumption. By installing an energy storage system, electricity can be purchased at a low price during off-peak hours and stored for use during peak hours to avoid high electricity bills. For example, a large supermarket may use the space under its parking lot to install a large-scale lithium-ion battery energy storage system to achieve this goal.
[0032] (3) Industrial users; self-sufficiency / microgrid: Some geographically remote enterprises or those that wish to be completely independent of the public power grid may choose to build their own microgrids, combining renewable energy power generation devices (such as wind turbines and solar panels) and energy storage facilities to form a closed-loop energy supply system. For example, mines located in desert areas may adopt this method to ensure a continuous and stable power supply.
[0033] In the embodiments of this application, the physical model of the user energy storage system typically includes hardware components such as batteries, inverters, grid interfaces, and management control units.
[0034] (1) Battery; The battery is the core component of an energy storage system, responsible for storing electrical energy and releasing it when needed. Currently, lithium-ion batteries are widely used due to their high energy density, long lifespan, and relatively low cost. In addition, with technological advancements, new battery technologies such as sodium-sulfur batteries and flow batteries are gradually entering the market, providing more options.
[0035] Functions provided: Stores excess electricity generated by solar panels or wind turbines and provides power support during peak electricity demand periods or when the grid is down.
[0036] (2) Inverter; An inverter is used to convert direct current (DC) to alternating current (AC) for compatibility with household or industrial equipment, or vice versa. It is a key component connecting the battery to the power grid.
[0037] Functionality provided: Enables bidirectional energy flow between the battery and the power grid, ensuring the quality and stability of electrical energy.
[0038] (3) Grid interface; The grid interface refers to the connection point between the energy storage system and the public power grid, which allows the energy storage system to obtain power from the grid or feed excess power back to the grid.
[0039] Functionality provided: Provides a secure and reliable interface that enables energy storage systems to seamlessly interact with the grid and participate in services such as demand response and frequency regulation.
[0040] (4) Management and control unit; The management and control unit is the "brain" of the entire energy storage system, responsible for monitoring and regulating the working status of each component and optimizing the system's operating efficiency.
[0041] Features include: real-time monitoring of battery status (such as SOC, State of Charge), grid conditions, load demand, etc.; automatic adjustment of charging and discharging strategies based on set goals (such as minimizing electricity costs and maximizing self-use rate); and support for remote monitoring and maintenance, improving system operability and maintenance convenience.
[0042] According to the user energy storage system composed of the energy storage facilities provided on the user side as above, the embodiments of the present application mainly take the battery as the control object, such as the lithium iron phosphate battery pack in the energy storage facility, and the embodiments of the present application provide a control method of a user energy storage system based on a power grid, as shown in Figure 1 The embodiments of the present application provide the following specific steps for this purpose:
[0043] 101. Determine the current environmental information of the energy storage facility, and the current environmental information at least includes: electricity price information, power load information, environmental temperature information and battery state of charge information.
[0044] The operating environment of the user energy storage system is very complex and variable, such as the comprehensive influence of factors such as the fluctuation of electricity price, the change of load demand, the fluctuation of environmental temperature, the battery health condition, etc.
[0045] (1) Electricity price fluctuation: The price of the power market will fluctuate according to the supply and demand relationship, the time period (such as peak and valley electricity price), the weather condition (especially for the case of relying on renewable energy), etc. Therefore, the energy storage system needs to have an intelligent algorithm to predict the electricity price change, and optimize the charging and discharging strategy accordingly to reduce the electricity cost or participate in the power market transaction to obtain the benefit.
[0046] (2) Change of load demand: The power consumption mode of the user will change over time, which depends on many factors, such as daily activity mode, seasonal change or special event, etc. The energy storage system needs to be able to adapt to such changes, to ensure sufficient power support in high demand period, and to charge in low demand period.
[0047] (3) Fluctuation of environmental temperature: The performance and life of the battery and other electronic components are greatly affected by temperature. Extreme temperature can cause the battery efficiency to decrease, and even can cause permanent damage. Therefore, the design of the energy storage system needs to consider an effective thermal management system to maintain the optimal working temperature range.
[0048] (4) Battery health condition: With the increase of use time, the capacity and efficiency of the battery will gradually decrease. Regular monitoring and evaluation of the state of the battery is crucial to ensure the long-term reliable operation of the system. In addition, reasonable charging and discharging strategy can also help to prolong the service life of the battery.
[0049] In full consideration of the influence of the above factors on the control of the user energy storage system, the embodiments of the present application quantify each influencing factor into different detection indexes, such as the fluctuation of electricity price -> detection of electricity price information, the change of load demand -> detection of power load information, the fluctuation of environmental temperature -> detection of environmental temperature information, and the battery health condition -> detection of battery state of charge information.
[0050] Therefore, at each running moment, the embodiment of the application obtains current environment information of the energy storage facility, at least including: electricity price information, power load information, environment temperature information and battery state of charge information.
[0051] 102. processing the current environment information by using the preset energy storage control model, and outputting a decision result.
[0052] The decision result is a charging and discharging control strategy for the battery in the energy storage facility.
[0053] The embodiment of the application mainly takes the battery as the control object, for example, the battery in the energy storage facility is a lithium iron phosphate battery pack, and the control strategy indicated by the decision result is the control of charging or discharging of the battery and the control of the allowed power in the control process.
[0054] The preset energy storage control model is a pre-trained reinforcement learning model, and the preset energy storage control model includes a policy network and a value network. The policy network is used to provide a control action for the energy storage facility according to the perception interaction with the current environment information, and the value network is used to evaluate the expected return of the control action to assist in making a control decision for the energy storage facility.
[0055] The main task of the policy network is to decide the best action according to the current environment information. For the user energy storage system, this means deciding whether the battery should be charged, discharged or not operated based on the current electricity price information, power load information, environment temperature information and battery state of charge information.
[0056] The role of the value network is to evaluate the expected return that can be obtained by following a certain strategy in a given state. It provides information about the pros and cons of the current strategy for the policy network, helping to improve the decision-making process.
[0057] 103. controlling the energy storage facility to store or release electric energy according to the decision result output by the preset energy storage control model.
[0058] The embodiment of the application uses the reinforcement learning model to make decisions on the user energy storage system, and more state variables (electricity price, power load, environment temperature and battery state of charge) are used in the model training process. The model makes decisions by comprehensively considering the fluctuations of electricity price, changes of load demand, fluctuations of environment temperature and battery health status, which are complex and variable environmental factors, thereby providing an effective regulation and control method for energy storage to ensure efficient operation and long-term stability of the user energy storage system.
[0059] In some alternative embodiments, for the training process of the preset energy storage control model, the embodiments of the present application provide the following specific implementation steps:
[0060] A1, set the state space, action space and reward function under the reinforcement learning framework.
[0061] Set the state space under the reinforcement learning framework, the state space contains a plurality of preset state variables, the state variables are used to represent different influencing factors on the energy storage facility during operation; set the action space under the reinforcement learning framework, the action space contains a plurality of preset actions, the preset actions are used to represent different control actions that need to be taken on the energy storage facility through perception interaction with different state variables, the action space is a discrete action space or a continuous action space; set the reward function under the reinforcement learning framework, the reward function at least includes: positive reward, neutral reward and negative reward performed when controlling the energy storage facility to perform control charging and discharging.
[0062] Next, taking the user energy storage system provided by the embodiments of the present application as an example, the role of the state space, action space and reward function in training a preset energy storage control model (reinforcement learning model) is explained.
[0063] Under the reinforcement learning framework, the state space, action space and reward function are the basic elements that define the interaction between the agent (i.e., can be simply referred to as the entity that executes decision-making and interacts with the environment in the process of controlling decision-making on the user energy storage system) and its environment. Their respective roles and relationships with each other are as follows:
[0064] (1) State space; role: the state space defines the set of all possible states in the environment. Each state reflects the environmental conditions at a certain time for the agent. For an agent, the state provides information about the current situation of the environment, based on which the agent can make decisions.
[0065] For example: in order to design a reinforcement learning model applied to decision-making on the user energy storage system, the embodiments of the present application design the state space to include the following several preset state variables:
[0066] (1.1) State of Charge (SOC): SOC is a key parameter reflecting the remaining capacity of the battery, usually ranging from 0% to 100%. It directly affects the discharging capacity and strategy selection of the energy storage system. .
[0067] (1.2) Current electricity price: the change of electricity price determines the economic benefit of charging and discharging, the energy storage system usually charges at low electricity price and discharges at high electricity price, so the electricity price is a key influencing factor for optimization strategy.
[0068] Let the current electricity price be denoted as .
[0069] (1.3) Load demand (power load status): Load demand directly affects the working status of the energy storage system and load balancing, and needs to provide power support during high load periods and reduce discharging during low load periods. Let the current load demand be denoted as .
[0070] (1.4) Ambient temperature: Temperature affects the life and performance of the battery. In high or low temperature environment, the charging and discharging efficiency and safety of the battery may be reduced.
[0071] (1.5) Time characteristics: Intraday period, day of the week, and other time characteristics have a significant impact on electricity prices and load changes, which can be used to help the model predict demand peaks and price fluctuations.
[0072] (2) Action space; Role: Action space defines the set of all possible actions that an agent can perform in the environment. Actions are the ways an agent chooses to affect the environment according to its current strategy. The choice of action is based on the agent's understanding of the current state and aims to achieve a certain goal or maximize a certain cumulative reward.
[0073] For example: In a user energy storage system, the action space defines all possible operations (i.e. preset actions) that the agent can take, such as mainly including the charging and discharging operations of the battery and the charging and discharging rate control. Based on different action space design requirements, actions can be divided into the following two categories:
[0074] (2.1) Discrete action space: The preset action can be designed as several fixed charging and discharging levels (such as charging, discharging, and idle), which is suitable for simple control tasks.
[0075] (2.2) Continuous action space: In order to control the charging and discharging power more accurately, the preset action can be designed as a continuous value, for example, selecting a charging and discharging rate between 0 and the maximum charging and discharging power. This design can help the model adapt flexibly in complex demand environments. For more complex demand environments, continuous action space can provide more flexible scheduling strategies, enabling the agent to better cope with price fluctuations, load changes, and other uncertainty factors. The design of continuous action space enables the model to adaptively adjust in a wider range of situations, optimizing the economic benefits and battery life of the energy storage system. However, the corresponding computational cost will also increase significantly.
[0076] (3) Reward function; role: the reward function provides an feedback signal for the agent to evaluate the effect of taking a certain action. The reward is usually immediate, which tells the agent whether the action is good or not. By maximizing the long-term cumulative reward, the agent learns the optimal behavior strategy. In the application scenario of making decisions on user energy storage systems, the reward function design needs to consider factors such as economy, stability and safety.
[0077] (3.1) Economic reward: design the reward by calculating the income obtained from the price difference. For example, discharging the energy storage system at high electricity price can obtain a larger positive reward, while discharging at low electricity price will produce a negative reward.
[0078] (3.2) Stability reward: encourage peak clipping and valley filling behavior to reduce the impact of load fluctuations on the power grid. When the energy storage system discharges at load peak and charges at load valley, the system will obtain a positive reward.
[0079] (3.3) Safety penalty: to protect the battery life, the system will be punished when the charging and discharging behavior causes the SOC to exceed the reasonable range or the temperature exceeds the safety interval, so as to encourage the agent to learn reasonable charging and discharging behavior.
[0080] As above (1)-(3) and B1-B3, it can be seen that in the reinforcement learning framework, the state space, action space and reward function are designed in advance, so that the agent (which can be simply referred to as an entity that controls decision-making on user energy storage systems and interacts with the environment) can learn how to take the optimal behavior in the environment by constantly exploring the environment, performing actions and receiving feedback.
[0081] A2 uses historical data obtained on historical time series to build a simulation running environment for the energy storage facility, the historical data at least including historical electricity price information, historical power load information, historical environmental temperature information and historical battery state of charge information, and the historical time series including multiple consecutive time steps.
[0082] The model trained by the embodiments of the application makes decisions, and the historical electricity price information, historical power load information, historical environmental temperature information and historical battery state of charge information on historical time series are used in the training, so as to comprehensively consider the complex and variable environmental factors such as the fluctuation of electricity price, the change of load demand, the fluctuation of environmental temperature and the battery health condition.
[0083] A3 analyzes and processes the historical data by combining the preset state variables in the state space preset in the reinforcement learning framework, to obtain the target state variables corresponding to each historical data information in the historical data, so as to constitute the target state space required for training the preset energy storage control model.
[0084] The multiple preset state variables can be pre-designed in the state space as above, but during the model training, the required state variables can be trained according to the required state variables, such as selecting the required types of state variables from the state variables, so that the trained model can consider the state factors represented by the required state variables when making decisions, to output the decision results.
[0085] In the embodiments of the present application, for the historical data including historical electricity price information, historical power load information, historical environmental temperature information and historical battery state of charge information, the data is parsed, and compared with the preset state variables in the state space pre-designed under the reinforcement learning framework, it can be seen that the corresponding historical data is converted into state variables including at least electricity price state, power load state, environmental temperature state and battery state of charge, which are used to constitute the target state space required for training the preset energy storage control model.
[0086] A4, at each time step, according to the target state information corresponding to each target state variable in the target state space, selects different preset actions from the action space pre-set under the reinforcement learning framework, to form at least one target data group, each target data group including a state set composed of each target state variable and a preset action.
[0087] At each time step, the agent perceives the target state space to select different preset actions from the action space to form a state-action pair, which can be called a target data group, and then at each time step, based on the different selected preset actions, multiple state-action pairs are formed.
[0088] A5, at least one data group is put into a simulation running environment to simulate the charging or discharging control operation of the energy storage facility.
[0089] By putting the state-action pair into the simulation running environment, the energy storage facility will simulate the charging or discharging action.
[0090] A6, a reward function pre-set under the reinforcement learning framework is used to evaluate the simulated operation to obtain a reward value.
[0091] The reward function provides a learning signal for the agent to guide its learning direction. The reward is a quantitative evaluation of the results after the agent takes an action, so that for any state-action pair, a corresponding reward value will be obtained.
[0092] A7, at least one data group corresponding to each time step, the reward value obtained based on the simulation operation of the data group, obtains at least one data set corresponding to each time step, each data set includes a data group and a reward value obtained based on the simulation operation of the data group.
[0093] Whenever an agent performs an action in a state, it receives an immediate reward that reflects how well that action has served the long-term goal. The design of the reward function directly impacts the learning efficiency and final performance of the agent. During training, the agent tries to maximize the cumulative reward, which means it needs to find a policy that allows it to achieve the highest possible sum of rewards in the future.
[0094] A8 takes each data set at each time step as a training sample and puts it into the preset experience pool. The training sample is a data sequence containing three data dimensions: state, action, and reward value.
[0095] A9 trains and updates the policy network and the value network based on the training samples in the experience pool to obtain the preset energy storage control model. Specifically, it includes but is not limited to the following (1)-(4):
[0096] (1) Extract a preset number of samples from the experience pool for a round of training operation. The preset number includes training samples representing data sets at different time steps.
[0097] (2) During the execution of a round of training operation, iteratively use different training samples to train the value network to evaluate the long-term benefits of the training samples through cumulative reward values.
[0098] (3) Use the evaluation results of the value network on the training samples to provide feedback to guide the update of the training policy network.
[0099] (4) Train and update the internal parameters of the policy network and the value network by iteratively performing multiple rounds of training operations until the preset energy storage control model achieves the expected benefits.
[0100] As can be seen from the above (1)-(4), in the reinforcement learning model applied in the user energy storage system, the update process of the policy network and the value network usually involves continuous optimization of the current policy and state value evaluation. This process is based on the experience obtained by the agent interacting with the environment, and adjusts the network parameters through a series of algorithms to achieve better performance. The following is the general mechanism of updating the policy network and the value network:
[0101] Policy Network Update: Experience Collection: The agent (energy storage system) selects actions according to the current policy and performs these actions with the environment, collecting data about states, actions, and their outcomes (rewards); Gradient Calculation: Using policy gradient methods (such as the REINFORCE algorithm or its variants like the Actor-Critic method), calculate the direction in which to adjust the policy network parameters to maximize the expected return. This is usually done by computing the gradient of a loss function with respect to the network weights, which is based on the collected experience data; Parameter Update: Using the calculated gradient information, update the policy network's weights using optimization algorithms (like Stochastic Gradient Descent, SGD, or Adam) to make the policy more inclined to select actions that lead to higher cumulative rewards.
[0102] Value Network Update: Target Value Estimation: For each experienced state-action pair, calculate a target value, which can be the immediate reward plus the estimated value of the next state (TD learning), or the value backpropagated from the actual return after multiple steps (Monte Carlo method); Error Calculation: Compare the difference between the predicted value of the current state's value by the value network and the target value calculated above, which is the error of the value function; Parameter Update: Based on the calculated error, adjust the parameters of the value network using optimization algorithms similar to those in the policy network to reduce the prediction error and improve the ability to accurately predict future cumulative rewards.
[0103] The relationship and collaborative work between the policy network and the value network have the following three characteristics:
[0104] Joint Action: In many advanced algorithms, such as the Actor-Critic framework, the policy network (Actor) is responsible for learning what actions to take, while the value network (Critic) evaluates the goodness of these actions. The two work closely together, with the Critic's feedback guiding the learning direction of the Actor.
[0105] Shared Representation: Sometimes, to improve efficiency, the policy network and the value network share some structure or feature representation layers, especially in deep reinforcement learning. This can speed up the learning process and improve generalization performance.
[0106] Simultaneous Update: In some implementations, the two networks may not be updated independently, but are adjusted simultaneously so that they can better support each other and move towards the optimal solution together.
[0107] In some modified embodiments, to make a more detailed description of the above embodiments, the application embodiment also provides an implementation process of the preset energy storage control model for perceiving and interacting with the current environment information to output a decision result, as shown in Figure 2 As shown, the application embodiment provides the following specific steps:
[0108] 201. By parsing and processing the current environmental information, the state variables and their corresponding state information are obtained. The state variables include at least the electricity price state, the power load state, the ambient temperature state, and the battery charge state. The state variables come from multiple pre-set state variables in the state space pre-designed when training the model using the reinforcement learning framework.
[0109] Multiple pre-defined state variables can be designed in the state space. However, during model training, the model can be trained according to the required state variables, such as selecting the required type of state variables. This allows the trained model to comprehensively consider the state factors represented by the required state variables when making decisions, and output the decision results.
[0110] In this embodiment of the application, if the current environmental information includes electricity price information, power load information, ambient temperature information and battery state of charge information, then by parsing these data and comparing them with the pre-set state variables in the pre-designed state space under the reinforcement learning framework, it can be seen that the current environmental information is correspondingly converted into a state variable that includes at least electricity price state, power load state, ambient temperature state and battery state of charge.
[0111] 202. In the process of perceiving and interacting with the state information of the state variables, the probability of each preset action being selected and executed is calculated using a policy network. The preset actions are used to represent the pre-designed operation of performing charge and discharge control on the energy storage facility. The preset actions are derived from the action space pre-designed when training the model using a reinforcement learning framework.
[0112] 203. Based on the probability of each preset action being selected for execution, select at least one target action from multiple preset actions.
[0113] 204. Combine state variables and different target actions into at least one data group, with each data group including a state variable and a target action.
[0114] In the embodiments of this application, each data group represents a pair of data consisting of a state variable and a target action, i.e., a state-action pair.
[0115] 205. Combining the pre-designed reward function under the reinforcement learning framework, the value network is used to evaluate each data group to obtain the expected return for each data group in the future preset time range. The expected return is used to characterize the benefit brought to the user's energy storage system when the target action is adopted.
[0116] 206. Based on the expected return for each data set, select the decision action from at least one target action.
[0117] In the embodiments of the present application, according to the expected income in the future preset time range corresponding to each data group evaluated by the value network, the target action corresponding to the maximum income is selected from the different expected incomes as the decision-making action output by the preset energy storage control model.
[0118] 207. Based on the charging and discharging control strategy provided by the decision-making action to the energy storage facility, the decision result output by the preset energy storage control model is determined.
[0119] The decision-making action can be but is not limited to providing charging / discharging control and charging / discharging power in the control process, so that the charging / discharging control strategy provided by the decision-making action is taken as the decision result output by the model.
[0120] As can be seen from 201-207 above, in the process of executing the model, the strategy network and the value network participate in the work including the following:
[0121] The work of the strategy network: state evaluation: when receiving the state of the current environment (such as the current electricity price state, the power load state, the environment temperature state and the battery charge state), the strategy network first processes these inputs; action selection: based on the input state information, the strategy network outputs a probability distribution representing the selection possibility of each possible action (such as charging, discharging or keeping the status quo). In actual deployment, the action with the highest probability of bringing high returns will usually be selected according to this probability distribution; real-time decision-making: the selected action is then implemented on the user's energy storage system, such as adjusting the charging and discharging rate of the battery to respond to the current grid conditions or the user's power demand.
[0122] The work of the value network: state evaluation: at the same time, the value network also receives the same environmental state (such as the current electricity price state, the power load state, the environment temperature state and the battery charge state) as input and calculates the expected cumulative reward or long-term income that can be obtained by following the current strategy in this state; auxiliary decision-making: although the final action is determined by the strategy network, the evaluation provided by the value network helps to determine the quality of the current strategy and can serve as an additional information layer to help understand the consequences of certain actions. For example, in some implementations, the value estimate can be used in combination with the action preferences generated by the strategy network to make more refined decisions.
[0123] Collaborative work between the strategy network and the value network: direct and indirect influence: the strategy network directly influences the decision-making process, while the value network indirectly supports this process by providing feedback on the quality of different states. The combined action of the two ensures that the energy storage system can respond quickly and accurately based on the latest environmental information.
[0124] Further, as an implementation of the method shown in the above Figure 1 , Figure 2 , the embodiment of the present application provides a control device of a user energy storage system based on a power grid. The device embodiment corresponds to the foregoing method embodiment, for the convenience of reading, the details of the foregoing method embodiment will not be described one by one, but it should be clear that the device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment. The device is applied to make energy storage decisions for a user energy storage system, specifically as shown in Figure 3 , the device comprises:
[0125] A determination unit 31 is configured to determine current environment information of the energy storage facility, wherein the current environment information at least includes price information, power load information, environmental temperature information and battery state of charge information;
[0126] A processing unit 32 is configured to process the current environment information by using a preset energy storage control model, and output a decision result, wherein the decision result is a charging and discharging control strategy for the battery in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model, the preset energy storage control model includes a policy network and a value network, the policy network is configured to provide a control action for the energy storage facility according to a perception interaction with the current environment information, and the value network is configured to evaluate an expected return of the control action to assist in making a control decision for the energy storage facility;
[0127] A control unit 33 is configured to control the energy storage facility to store or release electric energy according to the decision result.
[0128] Further, as shown in Figure 4 , the processing unit 32 comprises:
[0129] An analysis module 321 is configured to obtain state variables and state information corresponding to the state variables by analyzing the current environment information, wherein the state variables at least include price state, power load state, environmental temperature state and battery state of charge, and the state variables come from a plurality of preset state variables in a state space pre-designed when a reinforcement learning framework is used for model training;
[0130] A calculation module 322 is configured to calculate a probability of each preset action being selected and executed by using the policy network in a process of perceiving interaction with the state information of the state variables, wherein the preset action is used to represent a pre-designed operation of performing charging and discharging control on the energy storage facility, and the preset action comes from an action space pre-designed when a reinforcement learning framework is used for model training;
[0131] The first selection module 323 is configured to select at least one target action from the plurality of preset actions according to a probability of each preset action being selected for execution.
[0132] The composition module 324 is configured to compose the state variable and different target actions into at least one data group, each data group including the state variable and one target action.
[0133] The evaluation module 325 is configured to evaluate each data group by using the value network in combination with a reward function pre-designed under the reinforcement learning framework, to obtain an expected return of each data group over a preset time range in the future, the expected return being used to represent a predicted return brought by the target action to the user energy storage system.
[0134] The second selection module 326 is configured to select a decision-making action from the at least one target action based on the expected return of each data group.
[0135] The determination module 327 is configured to determine a decision result output by the preset energy storage control model based on a charging and discharging control strategy provided by the decision-making action to the energy storage facility.
[0136] Further, as shown in Figure 4 in the process of training the preset energy storage control model, the apparatus comprises:
[0137] The setting unit 34 is configured to set a state space under a reinforcement learning framework, the state space including a plurality of preset state variables, the state variables being used to represent different influencing factors to which the energy storage facility is subjected when in operation.
[0138] The setting unit 34 is further configured to set an action space under the reinforcement learning framework, the action space including a plurality of preset actions, the preset actions being used to represent different control actions that need to be taken on the energy storage facility through perceptual interaction with different state variables, the action space being a discrete action space or a continuous action space.
[0139] The setting unit 34 is further configured to set a reward function under the reinforcement learning framework, the reward function including at least positive reward, neutral reward and negative reward performed when the energy storage facility is controlled to perform control charging and discharging.
[0140] Further, as shown in Figure 4 in the process of training the preset energy storage control model, the apparatus further comprises a training unit 35; and the training unit is specifically configured to:
[0141] The historical data obtained on the historical time sequence is used to build a simulation running environment of the energy storage facility, and the historical data at least includes historical electricity price information, historical power load information, historical environment temperature information and historical battery state of charge information, and the historical time sequence includes a plurality of continuous time steps;
[0142] In combination with the preset state variables in the state space under the reinforcement learning framework, the historical data is analyzed and processed to obtain the target state variables corresponding to each historical data information in the historical data, so as to constitute a target state space required for training the preset energy storage control model;
[0143] At each time step, different preset actions are selected from the action space preset under the reinforcement learning framework according to the target state information corresponding to each target state variable in the target state space, to form at least one target data group, and each target data group includes a state set composed of each target state variable and a preset action;
[0144] At least one data group is put into the simulation running environment to simulate the charging or discharging control operation of the energy storage facility;
[0145] The reward function preset under the reinforcement learning framework is used to evaluate the simulated operation to obtain a reward value;
[0146] At least one data group corresponding to each time step, the reward value obtained based on the simulation operation of the data group, at least one data set corresponding to each time step is obtained, and each data set includes a data group and a reward value obtained based on the simulation operation of the data group;
[0147] Each data set at each time step is put into a preset experience pool as a training sample, and the training sample is a data sequence including three data dimensions of state, action and reward value;
[0148] Based on the training samples in the experience pool, the policy network and the value network are trained and updated to obtain the preset energy storage control model.
[0149] Further, as shown in Figure 4 Based on the training samples in the experience pool, the policy network and the value network are trained and updated to obtain the preset energy storage control model, and the training unit 35 is further specifically used for:
[0150] A preset number of samples are extracted from the experience pool for performing a round of training operation, and the preset number includes training samples representing the data sets at different time steps.
[0151] In the process of performing a round of training operation, different training samples are iteratively adopted to train the value network to evaluate long-term benefits of the training samples by accumulating reward values;
[0152] The evaluation results of the training samples by the value network are fed back to guide the training of the policy network to update;
[0153] By iteratively performing multiple rounds of training operation, the respective internal parameters of the policy network and the value network are trained and updated until the preset energy storage control model achieves the expected benefits.
[0154] As described above, the control device of the user energy storage system based on the power grid comprises a processor and a memory, and the determination unit, the processing unit and the control unit are stored in the memory as program units, and the corresponding functions are realized by the processor executing the program units stored in the memory.
[0155] The processor comprises a core, and the core retrieves the corresponding program units from the memory. The core can be set to one or more, and by adjusting the core parameters, the preset energy storage control model trained by reinforcement learning is used to make decisions by comprehensively considering complex and variable environmental factors such as fluctuations in electricity prices, changes in load demand, fluctuations in environmental temperature, and battery health status, thereby providing an effective adjustment and control method for energy storage to ensure efficient operation and long-term stability of the user energy storage system.
[0156] The embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to realize the control method of the user energy storage system based on the power grid.
[0157] The embodiment of the present application provides an electronic device, which comprises at least one processor, at least one memory connected with the processor, and a bus. The processor, the memory and the bus complete communication with each other through the bus. The processor is used to call program instructions in the memory to execute the control method of the user energy storage system based on the power grid. The present application also provides a computer program product which, when executed on a data processing device, is suitable for executing the program initialized with the steps of the control method of the user energy storage system based on the power grid.
[0158] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each flowchart block and / or block in the Figures can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that, in some alternative implementations, the flowchart blocks and / or blocks in the Figures can represent a Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments. In this regard, each flowchart block and / or block in the Figures can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that, in some alternative implementations, the flowchart blocks and / or blocks in the Figures can represent a
[0159] In one typical configuration, the device includes one or more processors (CPU), memory, and a bus. The device can also include an input / output interface, a network interface, and the like.
[0160] The memory can include non-persistent memory and / or volatile memory, e.g., random access memory (RAM) comprising a number of memory locations that can be read and / or written on the fly. The memory can also include non-volatile memory, e.g., read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. The memory includes at least one memory chip. The memory is an example of computer-readable media.
[0161] Computer-readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer-readable media does not include transitory media, such as modulated data signals and carrier waves.
[0162] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0163] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code.
[0164] The embodiments of the present application are only illustrative and are not intended to limit the present application. Various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A control method for a grid-based user energy storage system, wherein the user energy storage system comprises energy storage facilities deployed at the user end, characterized in that, The method includes: Determine the current environmental information of the energy storage facility, which includes at least: electricity price information, power load information, ambient temperature information, and battery state of charge information; The current environmental information is processed using a pre-set energy storage control model, and a decision result is output. The decision result is a charging and discharging control strategy for the batteries in the energy storage facility. The pre-set energy storage control model is a pre-trained reinforcement learning model, which includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facility based on the perception interaction with the current environmental information, and the value network is used to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facility. The step of processing the current environmental information using a pre-set energy storage control model and outputting decision results includes: By parsing the current environmental information, the state variables and the state information corresponding to the state variables are obtained. The state variables include at least the electricity price state, the power load state, the ambient temperature state, and the battery charge state. The state variables are derived from multiple preset state variables in the state space pre-designed when training the model using a reinforcement learning framework. During the process of perceiving and interacting with the state information of the state variables, the probability of each preset action being selected for execution is calculated using the policy network. The preset action is used to characterize a pre-designed operation for performing charge and discharge control on the energy storage facility. The preset action comes from the action space pre-designed when training the model using a reinforcement learning framework. Based on the probability that each of the preset actions is selected for execution, at least one target action is selected from the plurality of preset actions; The state variables and different target actions are grouped into at least one data group, and each data group includes the state variables and one target action; By combining the pre-designed reward function under the reinforcement learning framework, the value network is used to evaluate each data group to obtain the expected return for each data group in the future preset time range. The expected return is used to characterize the predicted return to the user energy storage system when the target action is adopted. Based on the expected return corresponding to each of the data groups, a decision action is selected from at least one of the target actions; Based on the charging and discharging control strategy provided to the energy storage facility by the decision-making action, the decision result output by the preset energy storage control model is determined; Based on the decision result, the energy storage facility is controlled to store or release electrical energy.
2. The method according to claim 1, characterized in that, The method for training the pre-set energy storage control model includes: A state space is set up under the reinforcement learning framework. The state space contains multiple preset state variables, which are used to characterize the different influencing factors that the energy storage facility is subjected to during operation. An action space is set up under the reinforcement learning framework. The action space contains multiple preset actions. The preset actions are used to represent the different control actions that need to be taken for the energy storage facility through perception interaction with different state variables. The action space is a discrete action space or a continuous action space. A reward function is set within the reinforcement learning framework, and the reward function includes at least: positive rewards, neutral rewards, and negative rewards for controlling the energy storage facility to perform controlled charging and discharging.
3. The method according to claim 2, characterized in that, During the training of the pre-set energy storage control model, the method further includes: Using historical data obtained from historical time series, a simulated operating environment for the energy storage facility is constructed. The historical data includes at least historical electricity price information, historical power load information, historical ambient temperature information, and historical battery state of charge information. The historical time series contains multiple consecutive time steps. By combining the preset state variables in the state space pre-set under the reinforcement learning framework, the historical data is analyzed and processed to obtain the target state variables corresponding to each historical data information in the historical data, so as to form the target state space required for training the preset energy storage control model. At each time step, based on the target state information corresponding to each target state variable in the target state space, different preset actions are selected from the action space pre-set under the reinforcement learning framework to form at least one target data group. Each target data group includes a set of states composed of each target state variable and a preset action. At least one of the data sets is deployed into the simulation environment to simulate the charging or discharging control operation of the energy storage facility. The simulated operation is evaluated using the reward function pre-set under the reinforcement learning framework to obtain a reward value; At least one data set corresponding to each time step is obtained by taking at least one data group corresponding to each time step and the reward value obtained by performing a simulation operation based on the data group. Each data set includes one data group and the reward value obtained by performing a simulation operation based on the data group. Each set of data at each time step is used as a training sample and placed into a preset experience pool. The training sample is a data sequence containing three data dimensions: state, action, and reward value. Based on the training samples in the experience pool, the policy network and value network are trained and updated to obtain the preset energy storage control model.
4. The method according to claim 3, characterized in that, The step of training and updating the policy network and value network based on the training samples in the experience pool to obtain the pre-set energy storage control model includes: A preset number of samples are drawn from the experience pool to perform one round of training operations. The preset number includes training samples that represent the data sets at different time steps. During a round of training, different training samples are used iteratively to train the value network to evaluate the long-term benefits of the training samples by accumulating reward values. The evaluation results of the training samples by the value network are used to guide the training of the policy network for updates; By iteratively executing multiple rounds of training operations to train and update the internal parameters of the policy network and the value network, the preset energy storage control model achieves the expected returns.
5. A control device for a grid-based user energy storage system, characterized in that, The user energy storage system consists of energy storage facilities deployed at the user end, and the device includes: The determining unit is used to determine the current environmental information of the energy storage facility, wherein the current environmental information includes at least: electricity price information, power load information, ambient temperature information, and battery state of charge information; The processing unit is used to process the current environmental information using a pre-set energy storage control model and output a decision result, which is a charging and discharging control strategy for the batteries in the energy storage facility. The pre-set energy storage control model is a pre-trained reinforcement learning model, which includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facility based on the perception interaction with the current environmental information, and the value network is used to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facility. The processing unit includes: The parsing module is used to parse and process the current environmental information to obtain the included state variables and the state information corresponding to the state variables. The state variables include at least the electricity price state, the power load state, the ambient temperature state, and the battery charge state. The state variables are derived from multiple preset state variables in the state space pre-designed when training the model using the reinforcement learning framework. The calculation module is used to calculate the probability of each preset action being selected for execution using the policy network during the process of perceiving and interacting with the state information of the state variables. The preset action is used to characterize a pre-designed operation for performing charge and discharge control on the energy storage facility. The preset action is derived from the action space pre-designed during model training using a reinforcement learning framework. The first selection module is used to select at least one target action from a plurality of preset actions based on the probability that each preset action is selected for execution; A composition module is used to assemble the state variables and different target actions into at least one data group, each data group including the state variables and one target action; An evaluation module is used to evaluate each data group by combining a pre-designed reward function under the reinforcement learning framework and the value network, so as to obtain the expected return of each data group in the future preset time range. The expected return is used to characterize the predicted return to the user energy storage system when the target action is adopted. The second selection module is used to select a decision action from at least one of the target actions based on the expected benefit corresponding to each of the data groups. The determination module is used to determine the decision result output by the preset energy storage control model based on the charging and discharging control strategy provided to the energy storage facility by the decision action; A control unit is configured to control the energy storage facility to store or release electrical energy based on the decision result.
6. The apparatus according to claim 5, characterized in that, During the training process of the pre-set energy storage control model, the device further includes: The setting unit is used to set up a state space under the reinforcement learning framework. The state space contains multiple preset state variables, which are used to characterize the different influencing factors that the energy storage facility is subjected to during operation. The setting unit is also used to set an action space under the reinforcement learning framework. The action space contains multiple preset actions. The preset actions are used to characterize the different control actions that need to be taken for the energy storage facility through perception interaction with different state variables. The action space is a discrete action space or a continuous action space. The setting unit is further configured to set a reward function within a reinforcement learning framework, the reward function including at least: positive rewards, neutral rewards, and negative rewards for controlling the energy storage facility to perform controlled charging and discharging.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the control method for a grid-based user energy storage system as described in any one of claims 1-4.
8. An electronic device, characterized in that, The device includes at least one processor, and at least one memory and bus connected to the processor; The processor and the memory communicate with each other via the bus. The processor is used to call program instructions in the memory to execute the control method of the grid-based user energy storage system as described in any one of claims 1-4.
Citation Information
Patent Citations
Optical storage charging station operation optimization method and system based on near-end strategy optimization algorithm
CN115986834A
D2D user resource allocation method based on deep reinforcement learning algorithm and storage medium
CN116456493A