Control method and device of user energy storage system based on power grid

Through the reinforcement learning model's strategy network and value network, combining electricity price, load, temperature and battery health status, the battery charging and discharging strategy of the user's energy storage system is optimized, and the problem of unstable operation of the user's energy storage system is solved and efficient and stable energy storage control is achieved.

CN120387619AActive Publication Date: 2025-07-29INFORMATION & COMMUNICATION BRANCH STATE GRID JIBEI ELECTRIC POWER CO LTD +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510394750.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The operating environment of the user's energy storage system is complex and changeable, and it is difficult to effectively deal with factors such as electricity price fluctuations, load demand changes, ambient temperature fluctuations and battery health, resulting in unstable system operation and inefficient efficiency.

Method used

A pre-installed energy storage control model with reinforcement learning training, including a policy network and a value network, is adopted to formulate a battery charging and discharging strategy through perceived interaction and expected income assessment, comprehensively considering electricity prices, loads, temperatures and battery health status, and optimizing the operation of energy storage facilities.

Benefits of technology

It improves the efficient operation and long-term stability of the user's energy storage system, and can make optimization decisions in complex and changeable environments to ensure the economic, stability and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387619A_ABST
    Figure CN120387619A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and device of a user energy storage system based on a power grid, and relates to the technical field of electric power. According to the main technical scheme, a reinforcement learning framework is adopted in advance to train a preset energy storage control model, the model comprises a strategy network and a value network, and the strategy network is used for providing a control action on an energy storage facility according to perception interaction with current environment information; the value network is used for evaluating the expected income of the control action to assist in making a control decision on the energy storage facility, and then in the process of using the model to cope with making the energy storage decision on the user energy storage system, the method comprises the following steps: firstly, determining the information of the current environment where the energy storage facility of the user side runs; and if the current environment information at least comprises the electricity price information, the power load information, the environment temperature information and the battery charge state information, performing perception interaction processing on the current environment information by using the pre-trained preset energy storage control model so as to obtain an energy storage decision result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electric power technology, and in particular to a control method and device for a user energy storage system based on a power grid. Background Art

[0002] With the rapid development of new energy technologies, especially the increasing popularity of renewable energy such as wind and solar energy, the role of distributed power generation and energy storage systems in modern power networks has become increasingly important. The development of these technologies has provided new ways to solve the environmental pollution problems caused by traditional energy supply, and has also promoted the transformation of energy structure towards a cleaner and more sustainable direction.

[0003] Among them, user energy storage systems usually refer to energy storage devices installed at the power user end. This system can store electrical energy and release it for use when needed. For example, its main functions include peak shaving and valley filling, improving power supply reliability, promoting the consumption of renewable energy, and participating in grid services, etc.

[0004] Currently, the operating environment of user energy storage systems is very complex and changeable. For example, it involves the combined influence of factors such as fluctuations in electricity prices, changes in load demand, fluctuations in ambient temperature, and battery health. In order to cope with such an environment, how to provide an effective regulation and control method for energy storage to ensure the efficient operation and long-term stability of user energy storage systems is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present application provides a control method and device for a grid-based user energy storage system. The main purpose is to utilize a preset energy storage control model trained through reinforcement learning to make decisions that comprehensively consider complex and variable environmental factors such as fluctuations in electricity prices, changes in load demand, fluctuations in ambient temperature, and battery health, thereby providing an effective regulation and control method for energy storage to ensure the efficient operation and long-term stability of the user energy storage system.

[0006] In order to achieve the above objectives, this application mainly provides the following technical solutions:

[0007] In a first aspect, the present application provides a control method for a user energy storage system based on a power grid, wherein the user energy storage system is composed of energy storage facilities deployed at a user end, and the method comprises:

[0008] Determining current environmental information of the energy storage facility, the current environmental information including at least electricity price information, power load information, ambient temperature information, and battery state of charge information;

[0009] The preset energy storage control model is used to process the current environmental information and output a decision result, where the decision result is a charge and discharge control strategy for the battery in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model, and the preset energy storage control model includes a policy network and a value network. The policy network is used to provide a control action for the energy storage facility according to the perception interaction with the current environmental information, and the value network is used to evaluate the expected benefit of the control action to assist in making a control decision for the energy storage facility;

[0010] According to the decision result, control the energy storage facility to store or release electrical energy.

[0011] A second aspect of the present application provides a control device for a user energy storage system based on a power grid. The user energy storage system is composed of energy storage facilities deployed at the user side. The device includes:

[0012] A determination unit for determining the current environmental information where the energy storage facility is located. The current environmental information at least includes: electricity price information, power load information, environmental temperature information, and battery state of charge information;

[0013] A processing unit for using a preset energy storage control model to process the current environmental information and output a decision result. The decision result is a charge and discharge control strategy for the battery in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model, and the preset energy storage control model includes a policy network and a value network. The policy network is used to provide a control action for the energy storage facility according to the perception interaction with the current environmental information, and the value network is used to evaluate the expected benefit of the control action to assist in making a control decision for the energy storage facility;

[0014] A control unit for controlling the energy storage facility to store or release electrical energy according to the decision result.

[0015] A third aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the control method for the user energy storage system based on the power grid as described above.

[0016] A fourth aspect of the present application provides an electronic device, which includes at least one processor, at least one memory connected to the processor, and a bus;

[0017] Wherein, the processor and the memory communicate with each other through the bus;

[0018] The processor is used to call program instructions in the memory to execute the control method of the user energy storage system based on the power grid as described above.

[0019] By means of the above technical solution, the technical solution provided by this application has at least the following advantages:

[0020] This application provides a control method and device for a user energy storage system based on the power grid. For a user energy storage system composed of energy storage facilities deployed at the user side, this application pre-uses a reinforcement learning framework to train a preset energy storage control model, which includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facilities based on the perception interaction with the current environmental information, and the value network is used to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facilities. Subsequently, in the process of using such a model to make energy storage decisions on the user energy storage system, this application first determines the current environmental information where the energy storage facilities at the user side are operating, such as at least including electricity price information, power load information, environmental temperature information, and state of charge of the battery, and then uses the above-mentioned pre-trained preset energy storage control model to perform perception interaction processing on the current environmental information to obtain the decision result for energy storage.

[0021] Compared with the existing requirements for the complex and variable operating environment of the user energy storage system, this application makes use of the characteristics of the reinforcement learning model to interact with the environment and learn and optimize according to the environmental feedback reward signal. This application uses a reinforcement learning model to make decisions on the user energy storage system, and if more state variables (electricity price, power load, environmental temperature, and state of charge of the battery) are used in the model training process, it will make the model's decision-making comprehensively consider complex and variable environmental factors such as electricity price fluctuations, changes in load demand, environmental temperature fluctuations, and battery health status, thereby providing an effective regulation and control method for energy storage to ensure the efficient operation and long-term stability of the user energy storage system.

[0022] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specific embodiments of this application are specifically given. Description of the Drawings

[0023] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0024] Figure 1Flowchart of a control method for a user energy storage system based on the power grid provided by an embodiment of the present application;

[0025] Figure 2 Another flowchart of a control method for a user energy storage system based on the power grid provided by an embodiment of the present application;

[0026] Figure 3 Block diagram of the composition of a control device for a user energy storage system based on the power grid provided by an embodiment of the present application;

[0027] Figure 4 Another block diagram of the composition of a control device for a user energy storage system based on the power grid provided by an embodiment of the present application. Detailed implementation manners

[0028] Hereinafter, the exemplary embodiments of the present application will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0029] A user energy storage system generally refers to an energy storage device installed at the power user side. Such a system can store electrical energy and release it for use when needed. For example, at the user side, especially at the residential, commercial, and industrial user sides, the installed energy storage devices can, but are not limited to, meet the following energy storage requirements of users:

[0030] (1) Residential users; adopting a solar + energy storage solution: Many households have installed a solution that combines solar photovoltaic panels (PV) with a battery energy storage system. For example, when there is sufficient sunlight during the day, the solar power generation may exceed the immediate consumption demand of the household, and the excess energy can be stored in the battery for use at night or on cloudy days. This not only reduces the dependence on grid power but also further reduces electricity bills by participating in the net metering program.

[0031] (2) Commercial users; peak shaving: Commercial premises such as shopping malls and office buildings usually have obvious peak electricity consumption periods. By installing an energy storage system, electricity can be purchased at a low price during off-peak hours and stored, and the stored electricity can be used during peak hours to avoid high electricity bills. For example, a large supermarket may install a large-scale lithium-ion battery energy storage system in the space under its parking lot to achieve this goal.

[0032] (3) Industrial users; Self-sufficient / microgrid: Some enterprises in remote geographical locations or those wishing to be completely independent of the public power grid will choose to establish their own microgrids, combining renewable energy generation devices (such as wind turbines, solar panels) and energy storage facilities to form a closed-loop energy supply system. For example, mines located in desert areas may adopt this method to ensure continuous and stable power supply.

[0033] In the embodiments of the present application, the physical model of the provided user energy storage system generally includes hardware components such as batteries, inverters, grid interfaces, and management and control units.

[0034] (1) Battery; The battery is the core component of the energy storage system, responsible for storing electrical energy and releasing it when needed. Currently, lithium-ion batteries are widely used due to their high energy density, long lifespan, and relatively low cost. In addition, with the progress of technology, new battery technologies such as sodium-sulfur batteries and flow batteries are gradually entering the market, providing more options.

[0035] Provided function: Store the excess power generated by solar panels or wind turbines, and provide power support during peak electricity consumption periods or when the power grid is out of power.

[0036] (2) Inverter; The inverter is used to convert direct current (DC) to alternating current (AC) to be compatible with household or industrial equipment, or vice versa. It is a key component connecting the battery to the grid.

[0037] Provided function: Achieve bidirectional energy flow between the battery and the grid, ensuring the quality and stability of electrical energy.

[0038] (3) Grid interface; The grid interface refers to the connection point between the energy storage system and the public power grid, which allows the energy storage system to obtain power from the grid or feed back excess power to the grid.

[0039] Provided function: Provide a safe and reliable interface, enabling the energy storage system to interact seamlessly with the grid and participate in services such as demand response and frequency regulation.

[0040] (4) Management and control unit; The management and control unit is the "brain" of the entire energy storage system, responsible for monitoring and regulating the working states of each component, and optimizing the operating efficiency of the system.

[0041] Provided function: Real-time monitor information such as the state of the battery (such as SOC, State of Charge), grid conditions, and load demands; Automatically adjust the charge and discharge strategies according to set goals (such as minimizing electricity bills, maximizing self-use rate); Support remote monitoring and maintenance, improving the operability and maintenance convenience of the system.

[0042] According to the user energy storage system composed of providing energy storage facilities on the user side as above, the embodiments of the present application mainly take the battery as the control object. For example, the battery in the energy storage facility is a lithium iron phosphate battery pack. The embodiments of the present application provide a control method for the user energy storage system based on the power grid, such as Figure 1 As shown, the embodiments of the present invention provide the following specific steps:

[0043] 101. Determine the current environmental information where the energy storage facility is located. The current environmental information at least includes: electricity price information, power load information, environmental temperature information, and state of charge information of the battery.

[0044] The operating environment of the user energy storage system is very complex and changeable. For example, it will involve the comprehensive influence of factors such as fluctuations in electricity prices, changes in load demands, fluctuations in environmental temperatures, and battery health conditions.

[0045] (1) Electricity price fluctuations: The price in the electricity market fluctuates according to factors such as supply and demand relationships, time periods (such as peak-valley electricity prices), and weather conditions (especially for cases relying on renewable energy). Therefore, the energy storage system needs to have intelligent algorithms to predict electricity price changes and optimize the charge-discharge strategy accordingly to reduce electricity costs or participate in electricity market transactions to obtain benefits.

[0046] (2) Changes in load demands: The power consumption pattern of users changes over time, which depends on various factors, such as daily activity patterns, seasonal changes, or special events. The energy storage system needs to be able to adapt to this change, ensure sufficient power support during high-demand periods, and charge during low-demand periods.

[0047] (3) Fluctuations in environmental temperature: The performance and lifespan of batteries and other electronic components are greatly affected by temperature. Extreme temperatures can lead to a decrease in battery efficiency and may even cause permanent damage. Therefore, the design of the energy storage system needs to consider an effective thermal management system to maintain the optimal operating temperature range.

[0048] (4) Battery health condition: As the usage time increases, the capacity and efficiency of the battery gradually decrease. Regularly monitoring and evaluating the battery state is crucial for ensuring the long-term reliable operation of the system. In addition, a reasonable charge-discharge strategy can also help extend the battery lifespan.

[0049] Fully considering the impacts of the above factors on the control of the user energy storage system, the embodiments of the present application quantify each influencing factor into different detection indicators. For example, fluctuations in electricity prices -> detect electricity price information, changes in load demands -> detect power load information, fluctuations in environmental temperature -> detect environmental temperature information, battery health condition -> detect state of charge information of the battery.

[0050] Therefore, at each running moment, the embodiments of the present application obtain the current environmental information of the energy storage facility, including at least: electricity price information, power load information, environmental temperature information, and state of charge information of the battery.

[0051] 102. Process the current environmental information using a preset energy storage control model and output a decision result.

[0052] Among them, the decision result is the charge and discharge control strategy for the battery in the energy storage facility.

[0053] The embodiments of the present application mainly take the battery as the control object. For example, if the battery in the energy storage facility is a lithium iron phosphate battery pack, the control strategy referred to by the decision result is the control of charging or discharging the battery and the control of the allowed power during the control process.

[0054] The preset energy storage control model is a pre-trained reinforcement learning model. The preset energy storage control model includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facility based on the perceptual interaction with the current environmental information, and the value network is used to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facility.

[0055] The main task of the policy network is to determine the best action to take based on the current environmental information. For a user energy storage system, this means making a decision on whether to charge, discharge, or do nothing to the battery based on factors such as the current electricity price information, power load information, environmental temperature information, and state of charge information of the battery.

[0056] The role of the value network is to evaluate the expected return that can be obtained by following a certain policy in a given state. It provides information about the quality of the current policy for the policy network and helps improve the decision-making process.

[0057] 103. Control the energy storage facility to store or release electrical energy according to the decision result output by the preset energy storage control model.

[0058] The embodiments of the present application utilize the characteristics of the reinforcement learning model to interact with the environment and learn and optimize according to the environmental feedback reward signal. The embodiments of the present application adopt a reinforcement learning model to make decisions on the user energy storage system. And if more state variables (electricity price, power load, environmental temperature, and state of charge of the battery) are used during the model training process, it will enable the model to make decisions by comprehensively considering complex and variable environmental factors such as electricity price fluctuations, load demand changes, environmental temperature fluctuations, and battery health conditions, thereby providing an effective regulation and control method for energy storage to ensure the efficient operation and long-term stability of the user energy storage system.

[0059] In some modified embodiments, for the training process of the preset energy storage control model, the embodiments of the present application provide the following specific implementation steps:

[0060] A1. Set the state space, action space, and reward function under the reinforcement learning framework.

[0061] Set the state space under the reinforcement learning framework. The state space contains multiple preset state variables, and the state variables are used to represent different influencing factors that the energy storage facility is subject to during operation; set the action space under the reinforcement learning framework. The action space contains multiple preset actions, and the preset actions are used to represent different control actions that need to be taken for the energy storage facility by perceiving and interacting with different state variables. The action space is a discrete action space or a continuous action space; set the reward function under the reinforcement learning framework. The reward function includes at least: positive rewards, neutral rewards, and negative rewards when controlling the energy storage facility to perform controlled charging and discharging.

[0062] Next, taking the user energy storage system provided by the embodiments of the present application as an example, the roles of the state space, action space, and reward function in training a preset energy storage control model (reinforcement learning model) are explained.

[0063] Under the reinforcement learning framework, the state space, action space, and reward function are the basic elements that define the interaction between an agent (which can be simply referred to as the entity that executes decisions and interacts with the environment during the process of making control decisions on the user energy storage system) and its environment. Their respective roles and the relationships between them are as follows:

[0064] (1) State space; Role: The state space defines all possible state sets in the environment. Each state reflects the environmental situation where the agent is at a certain moment. For an agent, the state provides information about the current situation of the environment, and based on this information, the agent can make decisions.

[0065] For example: To design a reinforcement learning model for making decisions on the user energy storage system, the state space designed in the embodiments of the present application may include the following several preset state variables:

[0066] (1.1) State of Charge (SOC) of the battery: SOC is a key parameter reflecting the remaining battery charge, usually ranging from 0% to 100%. It directly affects the discharge capacity and strategy selection of the energy storage system.

[0067]

[0068] (1.2) Current electricity price: The change in the electricity price determines the economic benefits of charging and discharging. The energy storage system usually charges at low electricity prices and discharges at high electricity prices. Therefore, the electricity price is a key influencing factor for optimizing strategies.

[0069] Record the current electricity price as e(t).

[0070] (1.3) Load demand (power load status): Load demand directly affects the working status and load balance of the energy storage system. It is necessary to give priority to providing power support during high-load periods and reduce discharging during low-load periods. Record the current load demand as P(t).

[0071] (1.4) Ambient temperature: Temperature affects the battery life and performance. In high or low temperature environments, the charge and discharge efficiency and safety of the battery may both decrease.

[0072] (1.5) Time characteristics: Time characteristics such as intra-day periods and days of the week have a significant impact on electricity price and load changes, and can be used to help the model predict demand peaks and electricity price fluctuations.

[0073] (2) Action space; Function: The action space defines the set of all possible actions that an agent can perform in the environment. An action is the way an agent chooses to affect the environment according to its current policy. The choice of action is based on the agent's understanding of the current state and aims to achieve a certain goal or maximize a certain cumulative reward.

[0074] For example: In a user energy storage system, the action space defines all possible operations (i.e., preset actions) that the agent can take, such as mainly including the charge and discharge operations of the battery and the control of charge and discharge rates. Based on different action space design requirements, actions can be divided into the following two types:

[0075] (2.1) Discrete action space: The preset actions can be designed as several fixed charge and discharge levels (such as charge, discharge, idle), which is suitable for simple control tasks.

[0076] (2.2) Continuous action space: In order to more precisely control the charge and discharge power, the preset actions can be designed as continuous values. For example, select a charge and discharge rate between 0 and the maximum charge and discharge power. This design can help the model flexibly adapt to complex demand environments. For more complex demand environments, the continuous action space can provide a more flexible scheduling strategy, enabling the agent to better cope with uncertainties such as electricity price fluctuations and load changes. The design of the continuous action space enables the model to make adaptive adjustments in a wider range of scenarios, optimizing the economic benefits and battery life of the energy storage system. However, correspondingly, the computational cost will also increase significantly.

[0077] (3) Reward function; Function: The reward function provides a feedback signal for the agent to evaluate the effect after taking a certain action. Rewards are usually immediate, telling the agent whether the action is good or bad. By maximizing the long-term cumulative reward, the agent learns the optimal behavior strategy. In the application scenario of making decisions on the user energy storage system, the design of the reward function needs to consider factors such as economy, stability, and safety.

[0078] (3.1) Economic reward: Design the reward by calculating the revenue obtained from the electricity price difference. For example, when the energy storage system discharges at a high electricity price, it can obtain a large positive reward, while discharging at a low electricity price will result in a negative reward.

[0079] (3.2) Stability reward: Encourage the behavior of peak shaving and valley filling to reduce the impact of load fluctuations on the power grid. When the energy storage system discharges during peak load and charges during valley load, the system will receive a positive reward.

[0080] (3.3) Safety penalty: To protect the battery life, when the charge and discharge behavior causes the state of charge (SOC) to exceed the reasonable range or the temperature to exceed the safe range, the system will be penalized, thereby encouraging the agent to learn reasonable charge and discharge behavior.

[0081] As described above in (1)-(3) and B1-B3, it can be seen that by pre-designing the state space, action space, and reward function in the reinforcement learning framework, the agent (which can be simply referred to as the entity that makes decisions and interacts with the environment during the process of controlling decisions on the user energy storage system) can learn how to take optimal actions in the environment by continuously exploring the environment, taking actions, and receiving feedback.

[0082] A2 Utilize the historical data obtained in the historical time series to construct a simulation operation environment where the energy storage facility is located. The historical data includes at least historical electricity price information, historical power load information, historical environmental temperature information, and historical battery state of charge information. The historical time series contains multiple consecutive time steps.

[0083] In the embodiments of this application, when training the model to make decisions, considering the complex and variable environmental factors such as the fluctuation of electricity prices, the change of load demand, the fluctuation of environmental temperature, and the battery health status, historical electricity price information, historical power load information, historical environmental temperature information, and historical battery state of charge information in the historical time series are used during training.

[0084] A3 Combine the preset state variables in the state space preset under the reinforcement learning framework, analyze and process the historical data, and obtain the target state variables corresponding to each historical data information in the historical data to form the target state space required for training the preset energy storage control model.

[0085] As described above, multiple preset state variables can be designed in advance in the state space. However, during model training, training can be performed according to the required state variables. For example, the required types of state variables can be selected from them, so that when the trained model makes a decision, it will comprehensively consider the state factors represented by the required state variables to output a decision result.

[0086] In the embodiment of the present application, for historical data including historical electricity price information, historical power load information, historical environmental temperature information, and historical battery state of charge information, then these data are parsed and compared with the preset state variables in the state space pre-designed under the reinforcement learning framework. It can be seen that the historical data is correspondingly converted into a target state space for training a preset energy storage control model, where the state variables at least include electricity price state, power load state, environmental temperature state, and battery state of charge.

[0087] A state At each time step, according to the target state information corresponding to each target state variable in the target state space, different preset actions are selected from the action space pre-set under the reinforcement learning framework to form at least one target data group. Each target data group includes a state set composed of each target state variable and a preset action.

[0088] At each time step, the agent perceives the target state space to select different preset actions from the action space to form a state-action pair, which can be called a target data group. Then, at each time step, based on different selected preset actions, multiple state-action pairs will be formed.

[0089] A5 Put at least one data group into the simulation running environment to simulate the charging or discharging control operation of the energy storage facility.

[0090] By putting the state-action pair into the simulation running environment, the charging or discharging action of the energy storage facility will be simulated.

[0091] A6 Use the reward function pre-set under the reinforcement learning framework to evaluate the simulated operation and obtain a reward value.

[0092] The reward function provides a learning signal for the agent to guide its learning direction. The reward is a quantitative evaluation of the result generated after the agent takes a certain action. Therefore, for any state-action pair, a corresponding reward value will be obtained.

[0093] A7 Obtain at least one data set corresponding to each time step, where each data set includes a data group and a reward value obtained by performing a simulation operation based on the data group, based on at least one data group corresponding to each time step and the reward value obtained by performing a simulation operation based on the data group.

[0094] Whenever the agent executes an action in a state, it receives an immediate reward, which reflects how good or bad the action is for achieving the long-term goal. The design of the reward function directly affects the learning efficiency and ultimate performance of the agent. During the training process, the agent tries to maximize the cumulative reward, which means it needs to find a strategy that can obtain the highest possible total reward in the future.

[0095] A8 takes each data set at each time step as a training sample and puts it into the pre-set experience pool. The training sample is a data sequence containing three data dimensions: state, action, and reward value.

[0096] A9 trains and updates the policy network and the value network based on the training samples in the experience pool to obtain the pre-set energy storage control model. Specifically, it includes but is not limited to the following (1)-(4);

[0097] (1) Draw a pre-set number of samples from the experience pool for performing a round of training operations. The pre-set number includes training samples representing data sets at different time steps;

[0098] (2) During the process of performing a round of training operations, iteratively use different training samples to train the value network to evaluate the long-term benefits of the training samples by accumulating the reward values;

[0099] (3) Use the evaluation results of the value network for the training samples to feedback and guide the update of the training policy network;

[0100] (4) Through iteratively performing multiple rounds of training operations, train and update the internal parameters of the policy network and the value network respectively until the pre-set energy storage control model reaches the expected benefits.

[0101] Based on the above (1)-(4), it can be seen that in the reinforcement learning model applied to the user energy storage system, the update process of the policy network and the value network usually involves the continuous optimization of the current policy and state value evaluation. This process is based on the experience obtained from the interaction between the agent and the environment, and adjusts the network parameters through a series of algorithms to achieve better performance. The following is the general mechanism for updating the policy network and the value network:

[0102] Policy network update: Experience collection: The agent (energy storage system) selects actions according to the current policy, executes these actions to interact with the environment, and collects data on states, actions, and their results (rewards); Gradient calculation: Using policy gradient methods (such as the REINFORCE algorithm or its variant Actor-Critic method), calculate the direction of how to adjust the policy network parameters to maximize the expected return. This is usually done by calculating the gradient of the loss function with respect to the network weights, and the loss function is based on the collected experience data; Parameter update: Utilize the calculated gradient information and adopt an optimization algorithm (such as Stochastic Gradient Descent SGD or Adam) to update the weights of the policy network, making the policy more inclined to select actions that can bring higher cumulative rewards.

[0103] Value network update: Target value estimation: For each experienced state-action pair, calculate a target value, which can be the immediate reward plus the value estimate of the next state (TD learning), or the value calculated by backtracking from the actual return several steps later (Monte Carlo method); Error calculation: Compare the difference between the predicted value of the value network for the current state value and the target value calculated above, that is, the error of the value function; Parameter update: According to the calculated error, use an optimization algorithm similar to that in the policy network to adjust the parameters of the value network to reduce the prediction error and improve the ability to accurately predict future cumulative rewards.

[0104] There are three characteristics of the relationship and collaborative work between the policy network and the value network as follows:

[0105] Cooperative action: In many advanced algorithms, such as the Actor-Critic framework, the policy network (Actor) is responsible for learning what actions to take, while the value network (Critic) evaluates the quality of these actions. The two cooperate closely, and the feedback provided by the Critic is used to guide the learning direction of the Actor.

[0106] Shared representation: Sometimes, to improve efficiency, the policy network and the value network will share some structural or feature representation layers, especially in deep reinforcement learning, which can accelerate the learning process and improve the generalization performance.

[0107] Synchronous update: In some implementations, these two networks may not be updated independently, but adjusted synchronously so that they can better support each other and jointly move towards the optimal solution.

[0108] In some modified embodiments, to make a more detailed description of the above embodiments, the embodiments of the present application also provide an implementation process in which a pre-set energy storage control model senses and interacts with the current environmental information to output a decision result, such as Figure 2 shown. For this, the embodiments of the present application provide the following specific steps:

[0109] 201. By parsing the current environmental information, the state variables contained therein and the state information corresponding to the state variables are obtained. The state variables include at least the electricity price state, the power load state, the ambient temperature state and the battery charge state. The state variables come from multiple preset state variables in the state space pre-designed when the reinforcement learning framework is used for model training.

[0110] Multiple preset state variables can be pre-designed in the state space, but during model training, training can be performed based on the required state variables, such as selecting the required types of state variables, so that the trained model will comprehensively consider the state factors represented by the required state variables when making decisions to output the decision results.

[0111] In an embodiment of the present application, the current environmental information includes electricity price information, power load information, ambient temperature information and battery state of charge information, so these data are parsed and compared with the preset state variables in the state space pre-designed under the reinforcement learning framework. It can be seen that the current environmental information is correspondingly converted into state variables including at least electricity price status, power load status, ambient temperature status and battery state of charge.

[0112] 202. In the process of perceiving and interacting with the state information of the state variables, the policy network is used to calculate the probability of each preset action being selected for execution. The preset action is used to represent the pre-designed operation for performing charge and discharge control on the energy storage facility. The preset action comes from the action space pre-designed when the reinforcement learning framework is used for model training.

[0113] 203. Select at least one target action from multiple preset actions according to the probability of each preset action being selected for execution.

[0114] 204. The state variables and different target actions are combined into at least one data group, where each data group includes a state variable and one target action.

[0115] In the embodiment of the present application, each data group represents a pair of data consisting of a state variable and a target action, that is, a state-action pair.

[0116] 205. Combined with the pre-designed reward function under the reinforcement learning framework, the value network is used to evaluate each data group to obtain the expected benefit of each data group in the future preset time range. The expected benefit is used to represent the predicted benefit to the user's energy storage system when the target action is adopted.

[0117] 206. Based on the expected benefit corresponding to each data group, select a decision action from at least one target action.

[0118] In the embodiments of the present application, according to the expected benefits corresponding to each data group evaluated by the value network over a preset future time range, the target action corresponding to the maximum benefit is selected from these different expected benefits as the decision-making action output by the preset energy storage control model of the present application.

[0119] 207. Determine the decision result output by the preset energy storage control model based on the charge and discharge control strategy provided by the energy storage facility for the decision-making action.

[0120] This decision-making action can, but is not limited to, providing charge / discharge control, as well as the charge and discharge power during the control process. Thus, based on the charge and discharge control strategy that can be given by such a decision-making action, it is used as the decision result output by the model.

[0121] As can be seen from 201-207 above, during the execution of the model processing by the preset energy storage control model (including the policy network and the value network), the participation of the policy network and the value network in the work includes the following:

[0122] Work of the policy network: State evaluation: When receiving the state of the current environment (such as the current electricity price state, power load state, environmental temperature state, and battery charge state), the policy network first processes these inputs; Action selection: Based on the input state information, the policy network outputs a probability distribution representing the selection possibility for each possible action (such as charging, discharging, or maintaining the current state). In actual deployment, usually, the action most likely to bring high returns is selected according to this probability distribution; Real-time decision-making: The selected action is then implemented on the user energy storage system, such as adjusting the charge and discharge rate of the battery to respond to the current grid conditions or the user's power demand.

[0123] Work of the value network: State evaluation: At the same time, the value network also receives the same environmental state (such as the current electricity price state, power load state, environmental temperature state, and battery charge state) as input and calculates the expected cumulative reward or long-term benefit that can be obtained by following the current policy in this state; Auxiliary decision-making: Although the final action is determined by the policy network, the evaluation provided by the value network helps to determine the quality of the current policy and can be used as an additional information layer to help understand the possible consequences of certain actions. For example, in some implementations, the value estimation can be used in combination with the action preferences generated by the policy network to make more refined decisions.

[0124] Collaborative work between the policy network and the value network: Direct and indirect impacts: The policy network directly affects the decision-making process, while the value network indirectly supports this process by providing feedback on the quality of different states. The two work together to ensure that the energy storage system can respond quickly and accurately based on the latest environmental information.

[0125] Furthermore, as a response to the above Figure 1 , Figure 2 The embodiment of the present application provides a control device for a user energy storage system based on a power grid. This device embodiment corresponds to the aforementioned method embodiment. For ease of reading, this device embodiment will no longer describe the details of the aforementioned method embodiment one by one, but it should be clear that the device in this embodiment can correspond to all the contents of the aforementioned method embodiment. This device is used to make energy storage decisions for user energy storage systems, specifically as follows Figure 3 As shown, the device includes:

[0126] A determining unit 31 is configured to determine current environmental information of the energy storage facility, wherein the current environmental information includes at least electricity price information, power load information, ambient temperature information, and battery state of charge information;

[0127] a processing unit 32 configured to process the current environmental information using a preset energy storage control model and output a decision result, wherein the decision result is a charge and discharge control strategy for the batteries in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model comprising a policy network and a value network, wherein the policy network is configured to provide control actions for the energy storage facility based on perceptual interaction with the current environmental information, and the value network is configured to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facility;

[0128] The control unit 33 is used to control the energy storage facility to store or release electric energy according to the decision result.

[0129] Further, such as Figure 4 As shown, the processing unit 32 includes:

[0130] A parsing module 321 is configured to parse the current environmental information to obtain state variables and state information corresponding to the state variables, wherein the state variables include at least electricity price status, power load status, ambient temperature status, and battery state of charge. The state variables are derived from a plurality of preset state variables in a state space pre-designed when a reinforcement learning framework is used for model training.

[0131] a calculation module 322 for calculating, using the policy network, a probability of each preset action being selected for execution during the process of perceiving and interacting with the state information of the state variable, the preset action being used to represent a pre-designed operation for performing charge and discharge control on the energy storage facility, the preset action being from an action space pre-designed during model training using a reinforcement learning framework;

[0132] The first selection module 323 is configured to select at least one target action from multiple preset actions according to the probability that each preset action is selected for execution.

[0133] The composition module 324 is configured to compose the state variable and different target actions into at least one data group, and each data group includes the state variable and one target action.

[0134] The evaluation module 325 is configured to evaluate each data group by using the value network in combination with a pre-designed reward function under the reinforcement learning framework, and obtain the expected benefit of each data group over a future preset time range, where the expected benefit is used to characterize the benefit brought to the user energy storage system when adopting the target action.

[0135] The second selection module 326 is configured to select a decision-making action from at least one target action based on the expected benefit corresponding to each data group.

[0136] The determination module 327 is configured to determine the decision result output by the preset energy storage control model based on the charge and discharge control strategy provided by the energy storage facility for the decision-making action.

[0137] Further, as Figure 4 shown, during the training process of the preset energy storage control model, the device includes:

[0138] The setting unit 34 is configured to set a state space under the reinforcement learning framework, where the state space includes multiple preset state variables, and the state variables are used to characterize different influencing factors on the energy storage facility during operation.

[0139] The setting unit 34 is further configured to set an action space under the reinforcement learning framework, where the action space includes multiple preset actions, and the preset actions are used to characterize different control actions that need to be taken for the energy storage facility by interacting with different state variables. The action space is a discrete action space or a continuous action space.

[0140] The setting unit 34 is further configured to set a reward function under the reinforcement learning framework, and the reward function includes at least: positive reward, neutral reward, and negative reward when controlling the energy storage facility to perform controlled charging and discharging.

[0141] Further, as Figure 4 shown, during the training process of the preset energy storage control model, the device further includes: a training unit 35; specifically, the training unit is used for:

[0142] Using the historical data obtained from the historical time series, a simulation operation environment where the energy storage facility is located is constructed. The historical data at least includes historical electricity price information, historical power load information, historical environmental temperature information, and historical battery state of charge information. The historical time series contains a plurality of consecutive time steps;

[0143] Combined with the preset state variables in the state space preset under the reinforcement learning framework, the historical data is analyzed and processed to obtain the target state variables corresponding to each historical data information in the historical data, so as to constitute the target state space required for training the preset energy storage control model;

[0144] At each of the time steps, according to the target state information corresponding to each target state variable in the target state space, different preset actions are selected from the action space preset under the reinforcement learning framework to form at least one target data group. Each target data group includes a state set composed of each target state variable and one of the preset actions;

[0145] Put at least one of the data groups into the simulation operation environment to simulate the charging or discharging control operation of the energy storage facility;

[0146] Use the reward function preset under the reinforcement learning framework to evaluate the simulated operation to obtain a reward value;

[0147] For each time step, at least one of the data groups and the reward value obtained by performing the simulation operation based on the data group are used to obtain at least one data set corresponding to each time step. Each data set includes one of the data groups and the reward value obtained by performing the simulation operation based on the data group;

[0148] For each data set at each time step, it is used as a training sample and put into a preset experience pool. The training sample is a data sequence containing three data dimensions: state, action, and reward value;

[0149] Based on the training samples in the experience pool, the policy network and the value network are trained and updated to obtain the preset energy storage control model.

[0150] Further, as Figure 4 shown, for the training of the policy network and the value network based on the training samples in the experience pool to obtain the preset energy storage control model, the training unit 35 is further specifically configured to:

[0151] Extract a preset number of the samples from the experience pool for performing a round of training operation. The preset number includes training samples representing the data sets at different time steps;

[0152] During the execution of a round of training operations, different training samples are iteratively used, and the value network is trained to evaluate the long-term benefits of the training samples by accumulating reward values.

[0153] Using the evaluation results of the value network for the training samples, feedback is provided to guide the update of the training of the policy network.

[0154] By iteratively executing multiple rounds of training operations, the internal parameters of the policy network and the value network are trained and updated until the preset energy storage control model achieves the expected benefits.

[0155] As described above, the control device of the user energy storage system based on the power grid includes a processor and a memory. The above-mentioned determination unit, processing unit, control unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.

[0156] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set. By adjusting the kernel parameters, using the preset energy storage control model obtained through reinforcement learning training, the decision-making takes into account complex and variable environmental factors such as power price fluctuations, load demand changes, environmental temperature fluctuations, and battery health conditions, thereby providing an effective method for regulating and controlling energy storage to ensure the efficient operation and long-term stability of the user energy storage system.

[0157] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the control method of the user energy storage system based on the power grid as described above.

[0158] The embodiment of the present application provides an electronic device, which includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory communicate with each other through the bus; the processor is used to call the program instructions in the memory to execute the control method of the user energy storage system based on the power grid as described above. The present application also provides a computer program product, which is suitable for executing a program initialized with the steps of the control method of the user energy storage system based on the power grid when executed on a data processing device.

[0159] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0160] In a typical configuration, the device includes one or more processors (CPUs), a memory, and a bus. The device may also include an input / output interface, a network interface, etc.

[0161] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip. The memory is an example of computer-readable media.

[0162] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0163] It should also be noted that the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A control method for a user energy storage system based on the power grid, wherein the user energy storage system is composed of energy storage facilities deployed at the user side, characterized in that, The method comprises: Determining current environmental information of the energy storage facility, the current environmental information including at least electricity price information, power load information, ambient temperature information, and battery state of charge information; The current environmental information is processed using a preset energy storage control model to output a decision result, wherein the decision result is a charge and discharge control strategy for the batteries in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model, and the preset energy storage control model includes a policy network and a value network. The policy network is used to provide control actions for the energy storage facility based on perceptual interaction with the current environmental information, and the value network is used to evaluate the expected benefits of the control actions to assist in making control decisions for the energy storage facility. According to the decision result, the energy storage facility is controlled to store electric energy or release electric energy.

2. The method according to claim 1, wherein The process of processing the current environmental information using the preset energy storage control model and outputting a decision result includes: The current environmental information is parsed to obtain state variables and state information corresponding to the state variables, wherein the state variables include at least electricity price status, power load status, ambient temperature status, and battery state of charge. The state variables are derived from a plurality of preset state variables in a state space pre-designed when a reinforcement learning framework is used for model training. In the process of perceiving and interacting with the state information of the state variable, the policy network is used to calculate the probability of each preset action being selected for execution, the preset action being used to represent a pre-designed operation for performing charge and discharge control on the energy storage facility, the preset action being from an action space pre-designed when training a model using a reinforcement learning framework; selecting at least one target action from a plurality of the preset actions according to a probability of each of the preset actions being selected for execution; Grouping the state variables and different target actions into at least one data group, each data group including the state variables and one target action; In combination with the pre-designed reward function under the reinforcement learning framework, the value network is used to evaluate each data group to obtain the expected benefit corresponding to each data group over a preset future time range. The expected benefit is used to represent the predicted benefit to the user's energy storage system when the target action is adopted; Selecting a decision action from at least one target action based on the expected benefit corresponding to each data group; Based on the charge and discharge control strategy provided by the decision-making action to the energy storage facility, a decision result output by the preset energy storage control model is determined.

3. The method according to claim 1 or 2, characterized in that, During the training of the preset energy storage control model, the method includes: Setting a state space under a reinforcement learning framework, wherein the state space includes a plurality of preset state variables, and the state variables are used to represent different influencing factors of the energy storage facility during operation; Set the action space under the reinforcement learning framework. The action space contains multiple preset actions, which are used to represent different control actions that need to be taken for the energy storage facility by perceiving and interacting with different state variables. The action space is a discrete action space or a continuous action space; Set the reward function under the reinforcement learning framework. The reward function at least includes positive rewards, neutral rewards, and negative rewards when controlling the energy storage facility to perform controlled charging and discharging.

4. The method according to claim 3, wherein During the process of training the preset energy storage control model, the method further includes: Use the historical data obtained in the historical time series to construct a simulation operation environment where the energy storage facility is located. The historical data at least includes historical electricity price information, historical power load information, historical environmental temperature information, and historical battery state of charge information. The historical time series contains multiple consecutive time steps; Combine the preset state variables in the state space preset under the reinforcement learning framework to analyze and process the historical data, and obtain the target state variables corresponding to each historical data information in the historical data, so as to constitute the target state space required for training the preset energy storage control model; At each time step, according to the target state information corresponding to each target state variable in the target state space, select different preset actions from the action space preset under the reinforcement learning framework to form at least one target data group. Each target data group includes a state set composed of each target state variable and one preset action; Put at least one data group into the simulation operation environment to simulate the charging or discharging control operation of the energy storage facility; Use the reward function preset under the reinforcement learning framework to evaluate the simulated operation and obtain a reward value; Obtain at least one data set corresponding to each time step by using at least one data group corresponding to each time step and the reward value obtained from the simulated operation based on the data group. Each data set includes a data group and the reward value obtained from the simulated operation based on the data group; Take each data set at each time step as a training sample and put it into a preset experience pool. The training sample is a data sequence containing three data dimensions: state, action, and reward value; Based on the training samples in the experience pool, train and update the policy network and the value network to obtain the preset energy storage control model.

5. The method according to claim 4, characterized in that, The training and updating of the policy network and the value network based on the training samples in the experience pool to obtain the preset energy storage control model includes: Extract a preset number of samples from the experience pool for performing a round of training operations. The preset number includes training samples representing data sets at different time steps; During the process of performing a round of training operations, iteratively use different training samples to train the value network to evaluate the long-term benefits of the training samples by accumulating reward values; Using the evaluation results of the training samples by the value network to provide feedback and guide the training of the strategy network for updating; Multiple rounds of training operations are iteratively performed to train and update the internal parameters of the strategy network and the value network until the preset energy storage control model achieves the expected benefit.

6. A control device for a user energy storage system based on a power grid, characterized in that, The user energy storage system is composed of energy storage facilities deployed at the user end, and the device includes: a determining unit, configured to determine current environmental information of the energy storage facility, wherein the current environmental information includes at least electricity price information, power load information, ambient temperature information, and battery state of charge information; a processing unit configured to process the current environmental information using a preset energy storage control model and output a decision result, wherein the decision result is a charge and discharge control strategy for batteries in the energy storage facility; the preset energy storage control model is a pre-trained reinforcement learning model comprising a policy network and a value network, wherein the policy network is configured to provide a control action for the energy storage facility based on a perceptual interaction with the current environmental information, and the value network is configured to evaluate the expected benefit of the control action to assist in making a control decision for the energy storage facility; A control unit is used to control the energy storage facility to store or release electrical energy according to the decision result.

7. The device according to claim 6, characterized in that, The processing unit comprises: a parsing module, configured to parse the current environmental information to obtain state variables and state information corresponding to the state variables, wherein the state variables include at least electricity price status, power load status, ambient temperature status, and battery state of charge, and the state variables are derived from a plurality of preset state variables in a state space pre-designed when a reinforcement learning framework is used for model training; a calculation module configured to calculate, using the policy network, a probability of each preset action being selected for execution during a process of perceiving and interacting with the state information of the state variable, the preset action being used to represent a pre-designed operation for performing charge and discharge control on the energy storage facility, the preset action being from an action space pre-designed during model training using a reinforcement learning framework; A first selection module is configured to select at least one target action from a plurality of preset actions according to a probability of each preset action being selected for execution; a composition module, configured to combine the state variables and different target actions into at least one data group, each data group including the state variable and one target action; An evaluation module, configured to evaluate each data group using the value network in combination with a pre-designed reward function under the reinforcement learning framework to obtain an expected benefit corresponding to each data group over a preset future time range, wherein the expected benefit is used to represent the predicted benefit to the user's energy storage system when the target action is adopted; a second selection module, configured to select a decision-making action from at least one of the target actions based on the expected benefit corresponding to each of the data groups; A determination module, configured to determine a decision result output by the preset energy storage control model based on the charge and discharge control strategy provided by the energy storage facility for the decision-making action.

8. The device according to claim 6 or 7, characterized in that, During the process of training the preset energy storage control model, the device further includes: A setting unit, configured to set a state space in a reinforcement learning framework, where the state space includes a plurality of preset state variables, and the state variables are used to characterize different influencing factors on the energy storage facility during operation; The setting unit is further configured to set an action space in a reinforcement learning framework, where the action space includes a plurality of preset actions, and the preset actions are used to characterize different control actions that need to be taken for the energy storage facility by sensing and interacting with different state variables, and the action space is a discrete action space or a continuous action space; The setting unit is further configured to set a reward function in a reinforcement learning framework, and the reward function at least includes: positive rewards, neutral rewards, and negative rewards during controlling the energy storage facility to perform controlled charging and discharging.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the control method of the user energy storage system based on the power grid according to any one of claims 1-5.

10. An electronic device, characterized in that, The device includes at least one processor, and at least one memory and a bus connected to the processor; Wherein, the processor and the memory complete communication with each other through the bus; The processor is configured to call program instructions in the memory to execute the control method of the user energy storage system based on the power grid according to any one of claims 1-5.

Citation Information

Patent Citations

  • Optical storage charging station operation optimization method and system based on near-end strategy optimization algorithm

    CN115986834A

  • D2D user resource allocation method based on deep reinforcement learning algorithm and storage medium

    CN116456493A

  • New energy microgrid optimization operation method based on improved model predictive control

    CN116488150A

  • Equalization method of energy storage battery pack management system based on neural network and medium

    CN117613421A

  • Automatic driving method, device and equipment based on reinforcement learning and storage medium

    CN119356310A