Energy storage coordination and regulation method and device
By encapsulating the operation modes of energy storage stations as options and combining them with a semi-Markov model and a hierarchical option evaluation model, the problem of collaborative optimization across application scenarios in energy storage coordination and control is solved, achieving efficient and reliable energy storage control in diverse scenarios.
Patent Information
- Application Number
- CN202511659536.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-11-13
AI Technical Summary
Existing energy storage coordination and control technologies lack a collaborative optimization framework across application scenarios, making it difficult to adapt to the complex coupling challenges in diverse scenarios, especially in terms of power dynamic balance, state of charge equilibrium, and economic optimization.
By encapsulating the operation modes of energy storage stations as options and combining them with a semi-Markov model and a hierarchical option evaluation model, a hierarchical decision-making mechanism is constructed to achieve optimal coordinated control of energy storage under diverse application scenarios.
It improves the economy and stability of energy storage systems, enhances the reliability and optimization effect of regulation in diverse application scenarios, reduces computing costs and decision-making difficulty, and achieves efficient collaborative optimization across scenarios.
Smart Images

Figure CN121124140B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of integrated energy management technology, specifically to a method and device for coordinated regulation of energy storage. Background Technology
[0002] Large-scale grid connection of renewable energy sources, such as wind and solar power, has become an inevitable trend. However, the intermittency and volatility of renewable energy pose serious challenges to the stability of the power system, especially in terms of frequency regulation, peak shaving, and energy balance. Energy storage technology, as a core carrier of flexible resources, has become a key means to smooth out renewable energy fluctuations and improve grid reliability through its "power support and power transfer" characteristics. However, in diverse application scenarios, energy storage systems must simultaneously meet multiple objectives, including dynamic power balance, state of charge (SOC) equilibrium, and economic optimization, leading to complex coupling issues in their coordinated control. Furthermore, factors such as renewable energy output prediction errors, load randomness, and battery aging characteristics further exacerbate the difficulty of control.
[0003] Existing research has made some progress in the coordinated regulation of energy storage. Traditional power control strategies are mostly based on fixed threshold algorithms, which are difficult to adapt to the dynamic demands of high-dimensional nonlinear scenarios. Existing research mostly focuses on single application scenarios and lacks a cross-scenario collaborative optimization framework: user-side energy storage economic models often ignore the impact of battery capacity degradation on the total life cycle cost, while grid-connected control strategies are rarely deeply coupled with demand response and renewable energy forecasting. Summary of the Invention
[0004] In view of this, the present invention provides an energy storage coordination and control method and apparatus to solve the problem of the lack of a cross-application scenario collaborative optimization framework in the existing energy storage coordination and control.
[0005] In a first aspect, the present invention provides an energy storage coordinated control method, the method comprising:
[0006] Determine the operating mode of the energy storage station and encapsulate the operating mode of the energy storage station as an option;
[0007] Construct a semi-Markov model incorporating options based on options and a semi-Markov model;
[0008] A hierarchical option evaluation model is constructed based on a semi-Markov model with integrated options, and the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios is solved based on the hierarchical option evaluation model.
[0009] This invention provides an energy storage coordination and control method that, by clearly defining and encapsulating the operating modes of energy storage stations as options, provides a standardized and scenario-based control strategy foundation for diverse application scenarios, avoiding the problem of insufficient targeting in traditional control strategies. The semi-Markov model incorporating options can accurately characterize the randomness of mode durations in actual operation, overcoming the simplification and distortion of operating conditions by fixed-time-step models, and better reflecting the real uncertainty of energy storage system operation. Based on this, a hierarchical option evaluation model simplifies the solution difficulty of complex control problems and improves decision-making efficiency through a hierarchical decision-making mechanism (divided into optimizing operating mode selection and refining action execution). Ultimately, it can address the challenges of diverse applications. By accurately matching the appropriate operating mode with different scenarios (such as peak shaving and valley filling, frequency regulation, peak-valley arbitrage, etc.), and outputting the best energy storage coordination and control scheme, the system achieves efficient synergy between the economy (reducing overall costs), stability (smoothing power fluctuations), and scenario adaptability of the energy storage system. This significantly improves the reliability and optimization effect of energy storage control in diverse application scenarios. At the same time, by encapsulating the station operation mode in the options, the algorithm solution efficiency is improved. In addition, the hierarchical model reduces the number of parameters and the computational cost, enabling coordinated control of energy storage in diverse application scenarios in a short time. This solves the problem of the lack of a cross-application scenario collaborative optimization framework in the existing technology for coordinated control of energy storage.
[0010] In one alternative implementation, the energy storage station operation modes include a rapid response mode and a long-term balancing mode.
[0011] Determine the operating mode of the energy storage site and encapsulate the operating mode of the energy storage site as options, including:
[0012] Based on the core objectives and constraints of different application scenarios, determine the rapid response mode and long-term balance mode of energy storage stations;
[0013] The fast response mode and the long-term equilibrium mode are encapsulated as corresponding identifiable, selectable, and executable options within the reinforcement learning framework.
[0014] This invention provides an energy storage coordinated control method that determines a rapid response mode and a long-term balance mode by combining the core objectives (such as short-term power fluctuation mitigation and long-term energy supply and demand balance) and constraints (such as response speed requirements and cost control boundaries) of different application scenarios. This ensures that the two operating modes can accurately match the scenario requirements, avoiding the problems of insufficient adaptability of traditional single modes to multiple scenarios and deviation of control objectives. The two modes are encapsulated into corresponding options that can be identified, selected, and executed by the reinforcement learning framework, building a standardized bridge between the actual operating mode and the algorithm decision-making logic. This allows reinforcement learning to carry out hierarchical decision-making based on the options without having to build control strategies from scratch. This reduces the decision-making difficulty of the algorithm in complex scenarios and provides a clear and reusable decision unit for subsequent integration of semi-Markov models and construction of hierarchical option evaluation models, ensuring the efficiency and accuracy of the entire energy storage coordinated control process.
[0015] In one alternative implementation, a semi-Markov model incorporating the options is constructed based on the options and the semi-Markov model, including:
[0016] The components of a semi-Markov model incorporating options are determined, including the state space, action space, option space, reward function, and state transition function.
[0017] Run the semi-Markov model in both fast response and long-term equilibrium modes to generate a semi-Markov model incorporating options.
[0018] In one alternative implementation, the components of the semi-Markov model incorporating the options are represented as follows:
[0019] ;
[0020] in, For state space, For the action space, For the option space, For the reward function, This is the state transition function.
[0021] This invention provides an energy storage coordination and control method that, by clearly defining the complete components of the state space, action space, option space (associating two modes: rapid response and long-term balance), reward function, and state transition function, ensures that the model can comprehensively and accurately map the operating characteristics and control requirements of the energy storage system, avoiding model distortion caused by missing or vague elements. By generating the final model through running a semi-Markov model in both rapid response and long-term balance modes, this method fully leverages the ability of the semi-Markov model to characterize the stochastic duration of mode operation (e.g., the duration of rapid response mode changes with the power fluctuation smoothing speed, and the long-term balance mode dynamically adjusts with the load curve), overcoming the limitations of traditional fixed-time-step models that are difficult to adapt to actual operational uncertainties. Furthermore, it allows the model to deeply integrate the differentiated operating logic of the two modes, providing a reliable model foundation for subsequent hierarchical option evaluation models that fits the actual operating conditions of energy storage and can accurately support mode decision-making and action optimization.
[0022] In one alternative implementation, a hierarchical option evaluation model is constructed based on a semi-Markov model of incorporating options, including:
[0023] Construct an option evaluation layer for a hierarchical option evaluation model. The option evaluation layer is used to randomly select an option with a preset probability or to select the option that maximizes the value function of the option action.
[0024] An action evaluation layer is constructed in the hierarchical option evaluation model. The action evaluation layer is used to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under the constraints of multiple application scenarios, under the options selected in the option evaluation layer.
[0025] The option evaluation layer and the action evaluation layer are trained separately to obtain a hierarchical option evaluation model.
[0026] This invention provides an energy storage coordinated control method. The option evaluation layer employs a dual strategy: randomly selecting options with preset probabilities or selecting the option that maximizes the value function of the selected action. This approach balances the exploration of potential optimal options (avoiding local optima) with prioritizing the optimal option after value evaluation, ensuring flexibility and accuracy in mode selection and guaranteeing the selection of energy storage operation modes suitable for diverse application scenarios. The action evaluation layer focuses on detailed control under the selected option, solving for the optimal solution with constraints from diverse scenarios. This ensures that the control actions align with the core logic of the option and meet the actual operational requirements of the scenario, avoiding ineffective solutions that deviate from constraints. Furthermore, training the option evaluation layer and the action evaluation layer separately allows each layer to focus on the core objectives of mode selection optimization and action detail optimization, reducing inter-layer coupling interference, improving training efficiency and model accuracy. The resulting hierarchical model efficiently connects mode decision-making and action execution, laying a solid foundation for outputting scientific and feasible optimal energy storage coordinated control solutions for diverse scenarios.
[0027] In one optional implementation, the option evaluation layer of the hierarchical option evaluation model is constructed, including:
[0028] The constituent elements of an option are determined, including the option's start state, action strategy, and termination function.
[0029] The value function of the selected action;
[0030] The option evaluation layer of the hierarchical option evaluation model is constructed based on the constituent elements of the options and the value function of the option actions.
[0031] In one alternative implementation, the constituent elements of the option are represented as follows:
[0032] ;
[0033] in, For options, , This is the initial state of the options. For action strategy, This is a terminating function;
[0034] The option action value function is expressed as:
[0035] ;
[0036] in, For the action strategy in the options, In the state Select action The expected cumulative rewards.
[0037] This invention provides an energy storage coordination and control method that, by clearly defining the three main components of an option—its start state, action strategy, and termination function—ensures that the option has clear operational boundaries and execution logic. Combined with the quantitative evaluation basis provided by the option action value function, the constructed hierarchical option evaluation layer can accurately anchor options suitable for the scenario, laying a logically clear and scientifically evaluated foundation for subsequent efficient decision-making.
[0038] In one alternative implementation, an action evaluation layer of the hierarchical option evaluation model is constructed, including:
[0039] Based on the energy storage station operation mode selected by the option evaluation layer, a deep deterministic policy gradient algorithm is used to generate action strategies within the options and evaluate the action strategies to obtain the action evaluation layer. The structure of the deep deterministic policy gradient algorithm includes an action network and an evaluation network. The action network is used to generate action strategies within the options, and the evaluation network is used to evaluate the action strategies.
[0040] This invention provides an energy storage coordination and control method that constructs an action evaluation layer. Based on the energy storage station operation mode selected by the option evaluation layer, it uses a deep deterministic strategy gradient algorithm containing an action network (generating action strategies within the options) and an evaluation network (evaluating the merits of the strategies). This method can ensure that the generated action strategies accurately meet the core requirements of the selected mode, and can also achieve strategy optimization through the collaboration of the two networks, thus ensuring the accuracy and adaptability of energy storage control actions in diverse scenarios.
[0041] In one optional implementation, the option evaluation layer and the action evaluation layer are trained separately to obtain a hierarchical option evaluation model, including:
[0042] Acquire historical data from diverse application scenarios and build an energy storage system simulation environment based on this historical data.
[0043] An experience pool was built based on the energy storage system simulation environment. The option evaluation layer and action evaluation layer were trained under different energy storage station operation modes and different diverse application scenarios to obtain the energy storage coordination and control model.
[0044] This invention provides an energy storage coordination and control method. By acquiring historical data from diverse application scenarios to build an energy storage system simulation environment, it can highly reproduce the real operating conditions of different scenarios, providing a realistic data foundation for model training and avoiding ineffective training detached from actual operating conditions. Based on the simulation environment, an experience pool is constructed to integrate different energy storage station operation modes (rapid response, long-term balance) and training samples from diverse scenarios, ensuring comprehensive sample coverage and providing rich learning materials for the model. Simultaneously, training the option evaluation layer and the action evaluation layer separately can reduce the target coupling interference between the two layers, allowing each layer to focus on its own core task and improving the training accuracy and efficiency of each layer. Finally, through targeted training under different modes and scenarios, the generated hierarchical option evaluation model has strong scenario adaptability and control reliability, and can stably output scientific energy storage coordination and control schemes in diverse application scenarios.
[0045] In one optional implementation, the option evaluation layer and the action evaluation layer are trained separately to obtain a hierarchical option evaluation model, which further includes:
[0046] During the training of the option evaluation layer, the termination function of the option evaluation layer is updated using gradient descent.
[0047] During the training of the action evaluation layer, the evaluation network of the action evaluation layer is updated by minimizing the loss function, and the action network of the action evaluation layer is updated by the policy gradient method.
[0048] This invention provides an energy storage coordination and control method. For the option evaluation layer, the gradient descent method is used to update the termination function, which can accurately optimize the termination timing of the option and make the mode switching more adaptable to the dynamics of the scenario. For the action evaluation layer, the evaluation network is updated by minimizing the loss function to improve the accuracy of action evaluation, and the action network is updated by the policy gradient method to optimize action generation, so as to achieve precise iteration of action policy. The targeted updates of the two layers work together to improve the training accuracy of the hierarchical model and the adaptability of the control strategy, ensuring that the model outputs a better energy storage coordination scheme.
[0049] In a second aspect, the present invention provides an energy storage coordination and control device, the device comprising:
[0050] The energy storage site operation mode encapsulation module is used to determine the energy storage site operation mode and encapsulate the energy storage site operation mode as an option;
[0051] A semi-Markov model building module incorporating options, used to build semi-Markov models incorporating options based on options and semi-Markov models;
[0052] The module for constructing a hierarchical option evaluation model and solving control schemes is used to construct a hierarchical option evaluation model based on a semi-Markov model incorporating options, and to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios based on the hierarchical option evaluation model.
[0053] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the energy storage coordination and control method of the first aspect or any corresponding embodiment described above.
[0054] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the energy storage coordination and control method of the first aspect or any corresponding embodiment described above.
[0055] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the energy storage coordination and control method of the first aspect or any corresponding embodiment described above. Attached Figure Description
[0056] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0057] Figure 1 This is a flowchart illustrating the energy storage coordinated control method according to an embodiment of the present invention;
[0058] Figure 2 This is a flowchart illustrating another energy storage coordinated control method according to an embodiment of the present invention;
[0059] Figure 3 This is a flowchart illustrating another energy storage coordinated control method according to an embodiment of the present invention;
[0060] Figure 4 This is a flowchart illustrating another energy storage coordinated control method according to an embodiment of the present invention;
[0061] Figure 5 This is a structural block diagram of the hierarchical option evaluation model in the energy storage coordinated control method according to an embodiment of the present invention;
[0062] Figure 6 This is a structural block diagram of an energy storage coordination and control device according to an embodiment of the present invention;
[0063] Figure 7 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. According to the embodiments of the present invention, an embodiment of an energy storage coordination and control method is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0065] This embodiment provides an energy storage coordination and control method, which can be used in the aforementioned computer equipment. Figure 1 This is a flowchart of an energy storage coordinated control method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:
[0066] Step S101: Determine the operating mode of the energy storage station and encapsulate the operating mode of the energy storage station as an option.
[0067] Specifically, considering the diverse application scenarios that energy storage systems need to adapt to, such as grid frequency regulation, peak shaving and valley filling, and peak-valley arbitrage on the user side, and targeting the core objectives (such as short-term power fluctuation mitigation, long-term energy supply and demand balance, and energy cost reduction) and constraints (such as response speed requirements, battery charging and discharging power limits, and SOC safety range) of different scenarios, two types of core energy storage station operation modes are defined: one is the rapid response mode, which focuses on instantaneous regulation at the second to minute level (such as mitigating sudden changes in renewable energy output and grid frequency fluctuations), and adopts a high-power, short-duration charging and discharging strategy; the other is the long-term balance mode, which focuses on continuous regulation at the hour to day level (such as day and night peak-valley load regulation and wind and solar power curtailment storage), and adopts a stable charging and discharging strategy with low losses.
[0068] The option is a standardized decision-making and execution unit with clear boundaries and execution logic, built under the reinforcement learning framework to adapt to the diverse control needs of energy storage systems. Its core is to transform a certain type of operation mode of energy storage stations (such as rapid response and long-term balance mode) into a modular strategy package that the algorithm can identify, select and execute, so as to achieve hierarchical decomposition and efficient optimization of complex energy storage control tasks.
[0069] The option is essentially an algorithmic encapsulation of a certain type of operation mode of an energy storage station. That is, the option transforms the abstract control objectives (such as smoothing power fluctuations and reducing energy costs) into standardized modules with start-up conditions, execution rules, and termination boundaries. This allows the model to quickly identify which operation mode to select in different scenarios and provides clear strategy constraints for the subsequent generation of specific control actions (such as charging and discharging power and duration). It is the core bridge connecting the macro-level operation mode selection and the micro-level action execution.
[0070] Step S102: Construct a semi-Markov model incorporating options based on the options and the semi-Markov model.
[0071] Specifically, based on the semi-Markov model, the options encapsulated in step S101 are deeply integrated into the semi-Markov model, and the actions in the model are extended to have random durations.
[0072] Step S103: Construct a hierarchical option evaluation model based on the semi-Markov model with integrated options, and solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios based on the hierarchical option evaluation model.
[0073] Specifically, based on the semi-Markov model with integrated options constructed in step S102, a hierarchical option evaluation model with option evaluation layer and action evaluation layer is built and trained. The trained hierarchical option evaluation model can select the appropriate station operation mode in the option evaluation layer; under the constraints of the energy storage station operation mode and application scenario, the action evaluation layer solves the optimal energy storage coordination and control scheme under multiple application scenarios.
[0074] The energy storage coordination and control method provided in this embodiment clarifies and encapsulates the operating modes of energy storage stations as options, providing a standardized and scenario-based control strategy foundation for diverse application scenarios, thus avoiding the problem of insufficient targeting in traditional control strategies. The semi-Markov model incorporating options can accurately characterize the randomness of mode duration in actual operation, overcoming the simplification and distortion of operating conditions by fixed-time-step models, and better reflecting the real uncertainty of energy storage system operation. The hierarchical option evaluation model constructed based on this simplifies the solution difficulty of complex control problems and improves decision-making efficiency through a hierarchical decision-making mechanism (divided into optimizing operating mode selection and refining action execution). Ultimately, it can address diverse application scenarios... By accurately matching the appropriate operating mode with different scenarios (such as peak shaving and valley filling, frequency regulation, peak-valley arbitrage, etc.), and outputting the best energy storage coordination and control scheme, the system achieves efficient synergy between the economy (reducing overall costs), stability (smoothing power fluctuations), and scenario adaptability of the energy storage system. This significantly improves the reliability and optimization effect of energy storage control in diverse application scenarios. At the same time, by encapsulating the station operation mode in the options, the algorithm solution efficiency is improved. In addition, the hierarchical model reduces the number of parameters and the computational cost, enabling coordinated control of energy storage in diverse application scenarios in a short time. This solves the problem of the lack of a cross-application scenario collaborative optimization framework in the existing technology for coordinated control of energy storage.
[0075] This embodiment provides an energy storage coordination and control method, which can be used in the aforementioned computer equipment. Figure 2 This is a flowchart of an energy storage coordinated control method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:
[0076] Step S201: Determine the operating mode of the energy storage station and encapsulate the operating mode of the energy storage station as an option.
[0077] Specifically, the energy storage station operation modes include a rapid response mode and a long-term balance mode; the above step S201 includes:
[0078] Step S2011: Determine the rapid response mode and long-term balance mode of the energy storage station based on the core objectives and constraints of different application scenarios.
[0079] Different application scenarios include, but are not limited to, grid frequency regulation, peak shaving and valley filling, and user-side peak-valley arbitrage. The core objectives of different application scenarios include, but are not limited to, mitigating short-term power fluctuations, achieving long-term energy supply and demand balance, and reducing energy costs. Constraints for different application scenarios include, but are not limited to, response speed requirements, battery charging and discharging power limits, and SOC safety range.
[0080] Because energy storage stations may face application scenarios such as grid frequency fluctuations, short-term load surges, and the suppression of renewable energy fluctuations, two operating modes are determined for energy storage stations: a rapid response mode and a long-term balancing mode. Specifically:
[0081] The rapid response mode is a dispatching strategy that uses rapid charging and discharging to cope with large fluctuations in energy demand in a short period of time and quickly adjusts the energy supply and demand changes; the long-term balance mode is a dispatching strategy that uses stable charging and discharging to maintain the long-term energy balance of the system and smooths the dispatching of electricity demand.
[0082] Step S2012: The fast response mode and the long-term equilibrium mode are encapsulated into corresponding identifiable, selectable and executable options in the reinforcement learning framework.
[0083] Specifically, the fast response mode is encapsulated as options, including: clearly defining the three core options of the fast response mode reinforcement learning options: for example, starting with the grid power fluctuation exceeding the threshold and the battery SOC being in the safe range of 30% to 80%; using high-power charging and discharging calculated based on a power deficit of 90% and adjustments made in 100ms increments as the action strategy; and using the fluctuation calming down in 30 seconds and the battery reaching the safe boundary as the termination function. These rules are parameterized (mapped to state / action space parameters), and start-up judgment, action generation, and termination judgment code modules are written, along with the fluctuation suppression rate and the reward function of battery loss, making them options that the framework can recognize and execute on demand.
[0084] The long-term balancing model is encapsulated as options, including: defining its option elements, such as: starting with a low electricity price, a renewable energy curtailment rate exceeding 5%, and a low SOC before peak load; using a strategy of smoothly allocating charging and discharging power based on 24-hour forecasts and adjusting at the hourly level; and using the completion of the peak-valley charging and discharging plan and departure from the target scenario as the termination function. Similarly, the rules are parameterized, corresponding code modules are developed, and reward functions for "cost savings and stable SOC" are bound to them, forming a framework with selectable options that can be executed stably over the long term.
[0085] Step S202: Construct a semi-Markov model incorporating options based on the options and the semi-Markov model.
[0086] Specifically, the options are incorporated into a semi-Markov model, and the actions are extended to have random durations. Step S202 above includes:
[0087] Step S2021: Determine the components of the semi-Markov model incorporating options. The components include the state space, action space, option space, reward function, and state transition function.
[0088] In one alternative implementation, the components of the semi-Markov model incorporating the options are represented as follows:
[0089] (1);
[0090] in, For state space, For the action space, For the option space, For the reward function, This is the state transition function. Where:
[0091] The state space is represented as:
[0092] (2);
[0093] in, For electricity purchase price, These represent the power generated by photovoltaic power generation and wind power generation, respectively. To meet user power requirements, This refers to the state of charge of the lithium battery.
[0094] The action space is represented as:
[0095] (3);
[0096] in, The charging and discharging power of the lithium battery;
[0097] The option space is represented as:
[0098] (4);
[0099] in, These are respectively the rapid response mode and the long-term balance mode;
[0100] The reward function is expressed as:
[0101] (5);
[0102] in, Power fluctuation index Normalized value, Comprehensive cost index Normalized value, Exclusive rewards for the option Normalized value, , All are weighting factors.
[0103] The aforementioned power fluctuation index Represented as:
[0104] (6);
[0105] in, For lithium battery and flywheel multi-point energy storage stations and the main grid, time Power of internal switching, This represents the average switching power.
[0106] Comprehensive Cost Index Represented as:
[0107] (7);
[0108] in, For system transaction costs, This refers to the degradation cost of lithium batteries.
[0109] Option-specific rewards Represented as:
[0110] (8);
[0111] in, Exclusive rewards for Quick Response Mode Normalized value, Exclusive rewards for long-term balanced mode Normalized value, For when (Right now and The value is 1 if it is true, and 0 otherwise. For the power that needs to be adjusted, This is the fluctuation deviation threshold.
[0112] Exclusive rewards for Quick Response Mode for:
[0113] (9);
[0114] Long-term Balance Mode Exclusive Rewards for:
[0115] (10);
[0116] in, To achieve target balance for lithium batteries value.
[0117] Step S2022: Run the semi-Markov model in fast response mode and long-term equilibrium mode to generate a semi-Markov model with incorporation options.
[0118] Specifically, first, ensure that the semi-Markov model incorporates the option elements corresponding to both the fast response and long-term equilibrium modes (e.g., the state space covers the operating parameters adapting to both modes, and the state transition function can characterize the random duration of mode operation). Then, run the model in each of the two modes: in the fast response mode, the model learns the state transition rules of high-power charging and discharging actions and the random distribution of mode duration as the fluctuation smoothing speed changes under short-term power fluctuation scenarios; in the long-term equilibrium mode, the model learns the state transition rules of stable charging and discharging actions and the random distribution of mode duration as the load / electricity price changes under long-term energy balance scenarios. Finally, by integrating the operating data and learning results from the two modes, a semi-Markov model with integrated options is generated that can accurately adapt to the two options and characterize the state dynamics and random time characteristics under different modes.
[0119] Step S203: Construct a hierarchical option evaluation model based on the semi-Markov model incorporating options, and solve for the optimal energy storage coordination and control scheme for the corresponding energy storage station operation modes under various application scenarios based on the hierarchical option evaluation model. For details, please refer to [link to relevant documentation]. Figure 1 Step S103 of the illustrated embodiment will not be described again here.
[0120] The energy storage coordination and control method provided in this embodiment ensures that the model can comprehensively and accurately map the operating characteristics and control requirements of the energy storage system by clearly defining the complete components of the state space, action space, option space (associating the two modes of rapid response and long-term balance), reward function, and state transition function. This avoids model distortion caused by missing or vague elements. By generating the final model through running a semi-Markov model in both rapid response and long-term balance modes, the method fully leverages the ability of the semi-Markov model to characterize the stochastic duration of mode operation (such as the duration of rapid response mode changing with the power fluctuation smoothing speed, and long-term balance mode dynamically adjusting with the load curve). This overcomes the limitations of traditional fixed-time-step models in adapting to the uncertainties of actual operation. Furthermore, it allows the model to deeply integrate the differentiated operating logic of the two modes, providing a reliable model foundation for subsequent hierarchical option evaluation models that fits the actual operating conditions of energy storage and can accurately support mode decision-making and action optimization.
[0121] This embodiment provides an energy storage coordination and control method, which can be used in the aforementioned computer equipment. Figure 3 This is a flowchart of an energy storage coordinated control method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps:
[0122] Step S301: Determine the operating mode of the energy storage site and encapsulate the operating mode as an option. For details, please refer to [link to relevant documentation]. Figure 2 Step S201 of the illustrated embodiment will not be described again here.
[0123] Step S302: Construct a semi-Markov model incorporating options based on the options and the semi-Markov model. See details below. Figure 2 Step S202 of the illustrated embodiment will not be described again here.
[0124] Step S303: Construct a hierarchical option evaluation model based on the semi-Markov model with integrated options, and solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios based on the hierarchical option evaluation model.
[0125] Specifically, the hierarchical option evaluation model is constructed based on the hierarchical option-evaluation reinforcement learning algorithm.
[0126] Step S303 above includes:
[0127] Step S3031: Construct the option evaluation layer of the hierarchical option evaluation model. The option evaluation layer is used to randomly select an option with a preset probability or select the option that maximizes the value function of the option action.
[0128] Specifically, the option evaluation layer is used to randomly select an option with a preset probability or to select the option that maximizes the value function of the option action. In other words, the option evaluation layer is used to select a suitable operating mode for the energy storage station.
[0129] In one alternative implementation, an option is randomly selected with a certain probability (i.e., a preset probability, which is set according to the actual situation and is not specifically limited here), or an option is selected that makes the action value function of the option. The options to maximize, such as Figure 5 The option evaluation layer shown above includes the following steps in step S3031:
[0130] Step a1: Determine the constituent elements of the option. The constituent elements of the option include the option's start state, action strategy, and termination function.
[0131] The constituent elements of the options are represented as follows:
[0132] (11);
[0133] in, For options, , This is the initial state of the options. For action strategy, This is the terminating function.
[0134] Step a2: Select the action value function.
[0135] The option action value function is expressed as:
[0136] (12);
[0137] in, For the action strategy in the options, In the state Select action The expected cumulative reward, The expression is:
[0138] (13);
[0139] in, For instant rewards, As a discount factor, For in the options The next state from arrive The value function, For state from arrive The corresponding probability. The expression is:
[0140] (14);
[0141] in, A parameterized neural network for a termination function.
[0142] Step a3: Construct the option evaluation layer of the hierarchical option evaluation model based on the constituent elements of the options and the option action value function.
[0143] The constructed option evaluation layer, such as Figure 5 As shown, it is possible to randomly select an option with a certain probability or select an option whose action value function is affected by the option. The purpose of maximizing the options.
[0144] Step S3032: Construct the action evaluation layer of the hierarchical option evaluation model. The action evaluation layer is used to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under the constraints of multiple application scenarios, under the options selected in the option evaluation layer.
[0145] Specifically, an action evaluation layer is constructed, and a deep deterministic policy gradient algorithm is used to generate in-option policies, implementing detailed scheduling policies under specific options, such as... Figure 5 The action evaluation layer is shown.
[0146] In an optional implementation, step S3032 includes:
[0147] Step b: Based on the energy storage station operation mode selected by the option evaluation layer, the deep deterministic policy gradient algorithm is used to generate action strategies within the options, and the action strategies are evaluated to obtain the action evaluation layer; wherein, the deep deterministic policy gradient algorithm includes an action network and an evaluation network; the action network is used to generate action strategies within the options, and the evaluation network is used to evaluate the action strategies.
[0148] Specifically, such as Figure 5 As shown, the structure of the deep deterministic policy gradient algorithm includes an action network and an evaluation network. The action network generates action policies under the guidance of the evaluation network, and the evaluation network evaluates the action policies generated by the action network.
[0149] Step S3033: Train the option evaluation layer and the action evaluation layer respectively to obtain the hierarchical option evaluation model.
[0150] In an optional implementation, step S3033 includes:
[0151] Step c1: Obtain historical data from various application scenarios and build an energy storage system simulation environment based on the historical data from various application scenarios.
[0152] Specifically, for the diverse application scenarios that energy storage systems need to adapt to (such as grid frequency regulation, peak-valley arbitrage, renewable energy consumption, and user-side load smoothing), corresponding historical operational data should be collected. Data types include: external environmental data (such as electricity price curves for different time periods, photovoltaic / wind power output time-series data, and user load fluctuation data), energy storage system data (such as battery state of charge (changes, charge / discharge power records, battery temperature and health status monitoring data), and scenario characteristic data (such as power fluctuation amplitude, curtailment rate, peak-valley time markers, etc.). Data sources may include the grid dispatch center database, energy storage site monitoring systems, and renewable energy power plant operation and maintenance records, and should cover typical scenarios of different seasons, different workdays / holidays, and different extreme weather conditions (such as cloudy / rainy days and windy days) to ensure the comprehensiveness and representativeness of the data.
[0153] Subsequently, a simulation environment for the energy storage system was built based on this historical data. This environment needs to have three main functions: first, scenario reproduction, accurately simulating system dynamics (such as power gap changes and electricity price fluctuations) under corresponding scenarios by inputting key parameters from historical data (such as daily power output curves and load curves); second, physical constraint simulation, embedding battery charge-discharge characteristic models (such as the state of charge (SOC)-power relationship and charge-discharge efficiency curves) and safety constraint rules (such as upper and lower limits of SOC and power thresholds) to ensure that the simulation actions conform to the actual equipment capabilities; and third, state feedback, outputting the system state at the next moment (such as the updated SOC and power gap changes) and reward values (such as fluctuation suppression effect and cost changes) based on the patterns in historical data when a certain control action (such as charge-discharge power) is input. The final simulation environment needs to be validated using historical data (e.g., the deviation between the simulated control action and the historical actual operating results should be ≤5%) to ensure it can realistically map the operating characteristics of diverse scenarios and provide a reliable virtual training ground for subsequent model training.
[0154] Step c2 involves constructing an experience pool based on the energy storage system simulation environment, and training the option evaluation layer and action evaluation layer under different energy storage station operation modes and different diverse application scenarios to obtain the energy storage coordination and control model.
[0155] Specifically, an experience pool is built based on the established energy storage system simulation environment. For example... Figure 5 As shown, the experience pool is used to store a five-tuple sample of state, option, action, reward, and next state. ,in: This refers to the current system status (such as SOC, power deficit, and electricity price). Selected energy storage operation mode (fast response / long-term balance). For the action to be performed (such as charging and discharging power). Rewards for actions (such as fluctuation suppression rewards, cost-saving rewards). This represents the next state after the action is executed. The specific construction method is as follows: In a simulation environment, by randomly selecting options and actions (in the initial exploration phase), or by calling the model to be trained to generate options and actions (in the later optimization phase), the operation process of different scenarios (such as frequency modulation scenarios and arbitrage scenarios) and different modes (high-power actions in fast response mode and stable actions in long-term equilibrium mode) is simulated. The five-tuple samples generated by each interaction are stored in the experience pool until the sample size covers a sufficient number of state spaces and scenario types (usually the experience pool capacity needs to reach 100,000 or more).
[0156] Next, the option evaluation layer and action evaluation layer were trained under different energy storage station operation modes and diverse application scenarios: Option evaluation layer training: focusing on selecting the optimal option (energy storage station operation mode), inputting the current real-time state of the simulation environment. Options are generated using a combination of random exploration and greedy selection (selecting options randomly with a certain probability, or selecting the option that maximizes the value function of the options). Apply the options to the simulation environment to obtain rewards. and the next state Based on samples in the experience pool, the termination function and value function of the options are updated using gradient descent (optimizing the termination timing and selection strategy of the options), enabling the option evaluation layer to more accurately select suitable options in different scenarios (such as prioritizing the fast response mode when power fluctuations are large). Action evaluation layer training focuses on generating the optimal action under the selected option, based on the options output by the option evaluation layer. and current real-time status Actions are generated using an action network based on a deep deterministic policy gradient algorithm. Apply actions to the simulation environment to obtain rewards. and the next state .
[0157] In one optional implementation, the option evaluation layer and the action evaluation layer are trained separately to obtain a hierarchical option evaluation model, which further includes:
[0158] When training the option evaluation layer, the termination function of the option evaluation layer is updated using gradient descent; when training the action evaluation layer, the evaluation network of the action evaluation layer is updated by minimizing the loss function, and the action network of the action evaluation layer is updated using policy gradient.
[0159] Specifically, samples are randomly drawn from the experience pool. The evaluation network is updated by minimizing the loss function (improving the accuracy of action value assessment), and the action network is updated by the policy gradient method (optimizing the action generation logic). This ensures that the actions generated by the action network both meet option constraints (such as the high-power characteristics of the fast response mode) and maximize the scenario objectives (such as smoothing fluctuations and reducing costs). Through multiple rounds of iterative training (usually requiring thousands to tens of thousands of rounds), when the policies of the two-layer model tend to stabilize (such as option selection accuracy ≥90% and action reward values converge), an energy storage coordination and control model that can adapt to diverse scenarios is obtained.
[0160] During training, the parameters of the termination function of the option evaluation layer are... The gradient is updated using gradient descent, and its gradient is:
[0161] (15);
[0162] in, From arrive The probability, For the dominant function, Let be a parameter vector used to parameterize the policy, value function, or termination rule of the options (such as the weights of a neural network, or the parameters of the option termination probability), and its expression is:
[0163] (16);
[0164] in, In the state Select option The expected cumulative reward, This is an option strategy.
[0165] Evaluation network of deep deterministic policy gradient algorithm for action evaluation layer By updating the loss function by minimizing the loss function, the loss function is defined as:
[0166] (17);
[0167] in, For instant rewards, As a discount factor, In the state The strategy for outputting the next action network. In the state Take action The expected return.
[0168] Action networks with deep deterministic policy gradient algorithm for action evaluation layer The policy gradient is defined as follows: (Updated via policy gradient)
[0169] (18);
[0170] in, This refers to the batch size.
[0171] exist Figure 5 The paper demonstrates a hierarchical reinforcement learning model for coordinated regulation of energy storage. Through a hierarchical architecture consisting of an option evaluation layer (upper layer, selecting the operating mode) and an action evaluation layer (lower layer, generating regulation actions), combined with a deep deterministic policy gradient algorithm, it achieves coordinated optimization of energy storage operating modes and specific actions under different scenarios. The functions and interaction logic of each module are as follows:
[0172] 1. The environment module includes user load, photovoltaic power generation, wind power generation, etc., representing the dynamic external (grid, load) and internal (power source, battery status) environment in which the energy storage system operates. The environment generates system states s (such as real-time power deficit, battery state of charge, electricity price, etc.), providing input for model decision-making.
[0173] 2. The experience tuples generated by the interaction between the experience replay pool storage environment and the model. in: This refers to the current system status (such as SOC, power deficit, and electricity price). Selected energy storage operation mode (fast response / long-term balance). For the action to be performed (such as charging and discharging power). Rewards for actions (such as fluctuation suppression rewards, cost-saving rewards). This represents the next state after the action is executed. By replaying experiences, data correlations are broken down, providing stable and diverse samples for the training of the action evaluation layer network, thereby improving training efficiency and stability.
[0174] 3. The option evaluation layer (upper layer: select operating mode) is responsible for selecting the energy storage operating mode option that best suits the current scenario from the rapid response mode (option 1) and the long-term balance mode (option 3). Core logic: Based on an option function, a random exploration + greedy selection strategy is adopted: random: to ensure that the model has a probability of exploring new options, avoiding getting trapped in local optima. Using the option value function Choose the option that yields the highest expected return under the current state s.
[0175] 4. Action Evaluation Layer (Lower Layer: Generate Control Actions): Based on the deep deterministic policy gradient algorithm, under the constraints of the operating mode (option) selected by the option evaluation layer, specific energy storage control actions (such as charging and discharging power) are generated.
[0176] The Deep Deterministic Policy Gradient Algorithm (DPRK) uses a dual-network (action network + evaluation network) + dual-objective (training network + target network) structure to stably optimize actions: Action network: includes the training network. and target network Training the network: Real-time adjustments based on the current state. Generate Actions Target network: Periodically copies parameters from the training network to provide a stable action target reference for training, avoiding training oscillations. Evaluation network: Includes the training network. and target network Training the network: Evaluating the current state and action pairs The value (expected reward); the target network: provides the value target, updates the parameters of the training network by minimizing the loss function (the loss is calculated by the difference between the value of the target network and the training network); at the same time, it optimizes the action network through policy gradients to make the generated actions more in line with the option policy and the scene target.
[0177] 5. Module interaction logic environment generates state s → Option evaluation layer selects option → Action evaluation layer based on and Generate Actions →Action It acts on the environment, and the environment provides feedback for the next state. The reward R → experience is stored in the replay pool → the replay pool provides samples for the network training of the action evaluation layer, driving the model to iteratively optimize and form a closed loop of perception, decision-making, execution and feedback.
[0178] The energy storage coordination and control method provided in this embodiment employs a dual strategy: the option evaluation layer randomly selects options with preset probabilities or selects the option that maximizes the value function of the selected action. This strategy balances the exploration of potential optimal options (avoiding local optima) with the priority selection of the optimal option after value evaluation, ensuring the flexibility and accuracy of mode selection and guaranteeing the selection of energy storage operation modes suitable for diverse application scenarios. The action evaluation layer focuses on the detailed control under the selected option, solving for the best solution with constraints from diverse scenarios. This ensures that the control actions are both consistent with the core logic of the option and meet the actual operational requirements of the scenario, avoiding the ineffectiveness of the solution deviating from the constraints. Furthermore, training the option evaluation layer and the action evaluation layer separately allows each layer to focus on the core objectives of mode selection optimization and action detail optimization, reducing inter-layer coupling interference, improving training efficiency and model accuracy. The resulting hierarchical model can efficiently connect mode decision-making and action execution, laying a solid foundation for outputting scientific and feasible optimal energy storage coordination and control solutions for diverse scenarios.
[0179] As one or more specific application embodiments of the present invention, combined with Figure 4 The energy storage coordinated control method provided by this invention will be further described in detail, such as... Figure 4 As shown, the process includes designing the operating mode, designing the semi-Markov model, and constructing the hierarchical option-evaluation reinforcement learning algorithm. The specific steps are as follows:
[0180] Step S1: Design the operating modes of the energy storage station and encapsulate them in the options. Since energy storage stations may face application scenarios such as grid frequency fluctuations, short-term load surges, and suppression of renewable energy fluctuations, two operating modes are designed for the energy storage station: a fast response mode and a long-term balancing mode, which are then encapsulated as two options.
[0181] The rapid response mode is a scheduling strategy that uses rapid charging and discharging to cope with large fluctuations in energy demand in a short period of time and quickly adjusts energy supply and demand changes.
[0182] The long-term balance mode is a scheduling strategy that uses stable charging and discharging to maintain the long-term energy balance of the system and smooth the power demand.
[0183] Step S2: Incorporating Options into a Semi-Markov Model Is it the option It incorporates Markov models and extends the actions to give them random durations.
[0184] Step S3: Construct and train a hierarchical option evaluation model based on a hierarchical option-evaluation reinforcement learning algorithm. The hierarchical option evaluation model includes an option evaluation layer and an action evaluation layer. After training, the hierarchical option-evaluation reinforcement learning algorithm can select an appropriate energy storage operation mode at the option evaluation layer; the action evaluation layer, under the constraints of the energy storage operation mode and the scenario, solves for the optimal energy storage coordination and control scheme under multiple application scenarios. Wherein:
[0185] Step S3.1: Construct an option evaluation layer and randomly select an option with a certain probability. Alternatively, you can choose to make the option action value function. Maximize the options This refers to the operation mode of energy storage stations.
[0186] Step S3.2: Construct an action evaluation layer and use the deep deterministic policy gradient algorithm to generate in-option policies to achieve specific options. The following is a solution for a multi-point energy storage aggregation and collaborative operation scheme.
[0187] Step S3.3: Train the algorithm to achieve coordinated regulation of energy storage in various application scenarios.
[0188] Termination function parameters of the option evaluation layer The evaluation network of the action evaluation layer is updated using the gradient descent method and the depth-deterministic policy gradient algorithm. The action network is updated by minimizing the loss function. By updating the policy gradient and iteratively calculating, a hierarchical option evaluation model based on the hierarchical option-evaluation reinforcement learning algorithm is obtained after training.
[0189] The energy storage coordination and control method provided in this embodiment significantly improves the algorithm's solution efficiency by encapsulating the station operation mode into options. Simultaneously, the hierarchical structure not only reduces the number of required parameters but also lowers computational costs, thereby enabling coordinated control of energy storage in diverse application scenarios in a shorter time.
[0190] This embodiment also provides an energy storage coordination and control device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0191] This embodiment provides an energy storage coordination and control device, such as... Figure 6 As shown, it includes:
[0192] The energy storage site operation mode encapsulation module 601 is used to determine the energy storage site operation mode and encapsulate the energy storage site operation mode as an option.
[0193] The semi-Markov model building module 602 incorporates options for building a semi-Markov model incorporating options based on options and a semi-Markov model.
[0194] The 603 module for constructing a hierarchical option evaluation model and solving control schemes is used to construct a hierarchical option evaluation model based on a semi-Markov model incorporating options, and to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios based on the hierarchical option evaluation model.
[0195] In some optional implementations, the energy storage site operation modes include a rapid response mode and a long-term balancing mode; the energy storage site operation mode encapsulation module 601 includes:
[0196] The operation mode determination unit is used to determine the rapid response mode and long-term balance mode of the energy storage station based on the core objectives and constraints of different application scenarios.
[0197] The operating mode is encapsulated as an option unit, which is used to encapsulate the fast response mode and the long-term equilibrium mode into corresponding identifiable, selectable and executable options in the reinforcement learning framework.
[0198] In some alternative implementations, the semi-Markov model building module 602 incorporating options includes:
[0199] The component determination unit is used to determine the components of the semi-Markov model with incorporated options. The components include the state space, action space, option space, reward function, and state transition function.
[0200] The semi-Markov model building blocks with incorporation options are used to run semi-Markov models in fast response mode and long-term equilibrium mode, generating semi-Markov models with incorporation options.
[0201] In some alternative implementations, the components of the semi-Markov model incorporating the options are represented as follows:
[0202] ;
[0203] in, For state space, For the action space, For the option space, For the reward function, This is the state transition function.
[0204] In some optional implementations, the hierarchical option evaluation model construction and control scheme solution module 603 includes:
[0205] The option evaluation layer construction unit is used to construct the option evaluation layer of the hierarchical option evaluation model. The option evaluation layer is used to randomly select an option with a preset probability or select the option that maximizes the value function of the option action.
[0206] The action evaluation layer construction unit is used to construct the action evaluation layer of the hierarchical option evaluation model. The action evaluation layer is used to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under the constraints of multiple application scenarios, under the options selected by the option evaluation layer.
[0207] The model training unit is used to train the option evaluation layer and the action evaluation layer separately to obtain a hierarchical option evaluation model.
[0208] In some optional implementations, the option evaluation layer building unit includes:
[0209] The component determination sub-unit is used to determine the components of an option. The components of an option include the option's start state, action strategy, and termination function.
[0210] The action value function determines the sub-unit, which is used to select the action value function.
[0211] The option evaluation layer construction subunit is used to construct the option evaluation layer of the hierarchical option evaluation model based on the constituent elements of the option and the option action value function.
[0212] In one alternative implementation, the constituent elements of the option are represented as follows:
[0213] ;
[0214] in, For options, , This is the initial state of the options. For action strategy, This is a terminating function;
[0215] The option action value function is expressed as:
[0216] ;
[0217] in, For the action strategy in the options, In the state Select action The expected cumulative rewards.
[0218] In some optional implementations, the action evaluation layer construction unit includes:
[0219] The action evaluation layer is a sub-unit used to generate action strategies within the options based on the energy storage station operation mode selected by the option evaluation layer, and to evaluate the action strategies, thus obtaining the action evaluation layer. The structure of the deep deterministic policy gradient algorithm includes an action network and an evaluation network. The action network is used to generate action strategies within the options, and the evaluation network is used to evaluate the action strategies.
[0220] In some alternative implementations, the model training unit includes:
[0221] The simulation environment construction sub-unit is used to acquire historical data under various application scenarios and build an energy storage system simulation environment based on the historical data under various application scenarios.
[0222] The model training subunit is used to build an experience pool based on the energy storage system simulation environment, and to train the option evaluation layer and action evaluation layer under different energy storage station operation modes and different multi-application scenarios to obtain the energy storage coordination and control model.
[0223] In some optional implementations, the model training unit further includes:
[0224] The termination function update sub-unit is used to update the termination function of the option evaluation layer using gradient descent during training of the option evaluation layer.
[0225] The evaluation network and action network update sub-units are used to update the evaluation network of the action evaluation layer by minimizing the loss function and to update the action network of the action evaluation layer by using the policy gradient method during the training of the action evaluation layer.
[0226] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0227] In this embodiment, the energy storage coordination and control device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0228] This invention also provides a computer device having the above-described features. Figure 6 The energy storage coordination and control device shown.
[0229] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 7 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 7 Take a processor 10 as an example.
[0230] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0231] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0232] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0233] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0234] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0235] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some optional embodiments, the display device may be a touchscreen. Embodiments of the present invention also provide a computer-readable storage medium in which the methods described above according to embodiments of the present invention can be implemented in hardware, firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and will be stored on a local storage medium, so that the methods described herein can be stored on such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; furthermore, the storage medium can also include combinations of the above types of memory. It is understood that a computer, processor, microprocessor controller, or programmable hardware includes storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0236] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0237] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for coordinated regulation of energy storage, characterized in that, The method includes: Determine the operation mode of the energy storage station and encapsulate the operation mode of the energy storage station as an option; the operation mode of the energy storage station includes a fast response mode and a long-term balance mode. Determine the operating mode of the energy storage site and encapsulate the operating mode of the energy storage site as options, including: Based on the core objectives and constraints of different application scenarios, determine the rapid response mode and long-term balance mode of energy storage stations; The fast response mode and the long-term equilibrium mode are respectively encapsulated into corresponding options that are identifiable, selectable, and executable within the reinforcement learning framework; Construct a semi-Markov model incorporating the options based on the aforementioned options and the semi-Markov model; A hierarchical option evaluation model is constructed based on a semi-Markov model with integrated options, and the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios is solved based on the hierarchical option evaluation model.
2. The method according to claim 1, characterized in that, The construction of a semi-Markov model incorporating the options based on the options and the semi-Markov model includes: The components of a semi-Markov model incorporating options are determined, including a state space, an action space, an option space, a reward function, and a state transition function. Run the semi-Markov model in both fast response and long-term equilibrium modes to generate a semi-Markov model incorporating options.
3. The method according to claim 2, characterized in that, The components of the semi-Markov model for the incorporation options are represented as follows: ; in, For state space, For the action space, For the option space, For the reward function, This is the state transition function.
4. The method according to claim 3, characterized in that, The hierarchical option evaluation model constructed based on the semi-Markov model of integration options includes: Construct an option evaluation layer for a hierarchical option evaluation model, wherein the option evaluation layer is used to randomly select an option with a preset probability or select an option that maximizes the value function of the option action; An action evaluation layer of a hierarchical option evaluation model is constructed. The action evaluation layer is used to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under the constraints of multiple application scenarios, under the options selected by the option evaluation layer. The option evaluation layer and the action evaluation layer are trained separately to obtain a hierarchical option evaluation model.
5. The method according to claim 4, characterized in that, The option evaluation layer of the hierarchical option evaluation model includes: The constituent elements of the option are determined, including the option start state, action strategy, and termination function; Determine the value function of the action options; The option evaluation layer of the hierarchical option evaluation model is constructed based on the constituent elements of the options and the option action value function.
6. The method according to claim 5, characterized in that, The constituent elements of the option are represented as follows: ; in, For options, , This is the initial state of the options. For action strategy, This is a terminating function; The value function of the option action is expressed as follows: ; in, For the action strategy in the options, In the state Select action The expected cumulative rewards.
7. The method according to claim 4, characterized in that, The action evaluation layer of the hierarchical option evaluation model includes: Based on the energy storage station operation mode selected by the option evaluation layer, a deep deterministic policy gradient algorithm is used to generate action strategies within the options and evaluate the action strategies to obtain the action evaluation layer. The structure of the deep deterministic policy gradient algorithm includes an action network and an evaluation network. The action network is used to generate action strategies within the options, and the evaluation network is used to evaluate the action strategies.
8. The method according to claim 4, characterized in that, The option evaluation layer and action evaluation layer are trained separately to obtain a hierarchical option evaluation model, including: Acquire historical data from diverse application scenarios and build an energy storage system simulation environment based on this historical data. An experience pool is constructed based on the energy storage system simulation environment. The option evaluation layer and action evaluation layer are trained under different energy storage station operation modes and different multi-application scenarios to obtain the energy storage coordination and control model.
9. The method according to claim 4, characterized in that, The hierarchical option evaluation model is obtained by training the option evaluation layer and the action evaluation layer respectively, and also includes: During the training of the option evaluation layer, the termination function of the option evaluation layer is updated using gradient descent. During the training of the action evaluation layer, the evaluation network of the action evaluation layer is updated by minimizing the loss function, and the action network of the action evaluation layer is updated by the policy gradient method.
10. An energy storage coordination and control device, characterized in that, The device includes: An energy storage site operation mode encapsulation module is used to determine the energy storage site operation mode and encapsulate the energy storage site operation mode as an option; the energy storage site operation modes include a fast response mode and a long-term balance mode; the energy storage site operation mode encapsulation module includes: The operation mode determination unit is used to determine the rapid response mode and long-term balance mode of the energy storage station based on the core objectives and constraints of different application scenarios. The operating mode is encapsulated as an option unit, which is used to encapsulate the fast response mode and the long-term equilibrium mode into corresponding identifiable, selectable and executable options in the reinforcement learning framework; A semi-Markov model building module incorporating options is used to construct a semi-Markov model incorporating options based on the options and the semi-Markov model. The hierarchical option evaluation model construction and control scheme solution module is used to construct a hierarchical option evaluation model based on a semi-Markov model incorporating options, and to solve the optimal energy storage coordination and control scheme for the corresponding energy storage station operation mode under multiple application scenarios based on the hierarchical option evaluation model.
Citation Information
Patent Citations
Dynamic power dispatching method for assisting user travel by charging and discharging strategy
CN112348387A
Photovoltaic off-grid hydrogen production system and control method thereof
CN119482343A