Energy storage power station and power grid coordinated complementary regulation method and system
By decomposing power grid data at multiple scales and using distributed model predictive control, combined with power grid topology for disturbance prediction and optimization control, the problem of coordinated regulation of energy storage systems in complex power grid environments has been solved, improving power grid stability and the economy and service life of energy storage devices.
Patent Information
- Application Number
- CN202511100684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing energy storage control methods cannot effectively identify and specifically adjust grid disturbances in different frequency bands, resulting in low energy storage resource utilization efficiency. Furthermore, they lack distributed coordination mechanisms, making it difficult to maintain efficient collaboration when the grid topology changes in a complex manner, and they fail to balance system stability control with the economics of energy storage devices.
By performing multi-scale decomposition of real-time frequency and voltage data from the grid side, a distributed model predictive control framework is constructed. By combining the grid topology, disturbance propagation prediction is performed, a local optimization controller is configured, and global optimization is achieved through a consensus protocol. An adaptive dynamic programming controller is designed to evaluate the remaining lifetime cost of the energy storage unit and adjust the charging and discharging strategy.
It enables precise sensing and early response to power grid fluctuations, improves power grid stability and anti-disturbance capabilities, rationally allocates energy storage resources, enhances the overall economy and operating efficiency of the system, and extends the service life of energy storage equipment.
Smart Images

Figure CN120638422B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to smart grid technology, and more particularly to a method and system for coordinated and complementary regulation of energy storage power stations and power grids. Background Art
[0002] With the increasing proportion of renewable energy generation, the safe and stable operation of the power grid faces new challenges. The intermittent and volatile nature of renewable energy sources exacerbates frequency and voltage fluctuations in the power grid, affecting system stability. Energy storage systems, as an important power regulation resource, can mitigate these fluctuations through rapid charge and discharge responses, improving the stability and flexibility of the power grid. In power grid frequency and voltage regulation, energy storage power stations have advantages such as fast response speed and high regulation accuracy, and are gradually becoming an important support for the safe and stable operation of the power grid.
[0003] Currently, the coordinated regulation of power grids and energy storage systems faces multiple challenges. Existing energy storage control methods typically employ a single control strategy, failing to effectively identify and address grid disturbances across different frequency bands. This results in low energy storage resource utilization efficiency and hinders the full realization of the technical characteristics of different energy storage types. Existing energy storage control systems mostly adopt a centralized control architecture, lacking distributed coordination mechanisms. When faced with complex changes in grid topology, the adaptability and scalability of the control system are poor, making it difficult to achieve efficient coordination among multiple energy storage units. Most energy storage control methods lack a comprehensive consideration of the lifespan cost of energy storage, failing to achieve a balance between system stability control and the economics of energy storage devices. This leads to increased operating costs for energy storage systems and poor overall economic performance.
[0004] Against the backdrop of rapid development of power systems and energy transition, there is an urgent need for an advanced control method that can achieve coordinated and complementary regulation between energy storage power stations and the power grid, in order to improve the stability, flexibility and economy of the power system. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for coordinated and complementary regulation between energy storage power stations and power grids, which can solve the problems in the prior art.
[0006] A first aspect of the present invention provides a method for coordinated and complementary regulation between an energy storage power station and the power grid, comprising:
[0007] Multi-scale decomposition is performed on real-time frequency and voltage data from the power grid to obtain disturbance components in different frequency bands. Combined with the power grid topology, disturbance propagation prediction data is calculated.
[0008] Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to the response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol.
[0009] An adaptive dynamic programming controller is designed, and the control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group, and the policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results.
[0010] When a change in grid topology is detected, the state assessment is recalculated based on the new network structure, and the optimal charging and discharging power output of the strategy network is updated. Based on the optimized optimal charging and discharging power, each energy storage group is controlled to perform charging and discharging operations.
[0011] In one alternative implementation,
[0012] The steps for performing multi-scale decomposition on real-time frequency and voltage data from the power grid side to obtain disturbance components in different frequency bands, and then calculating disturbance propagation prediction data based on the power grid topology, include:
[0013] The real-time frequency data and voltage data are decomposed into perturbation components in different frequency bands through multi-scale decomposition.
[0014] A dynamic topology identification matrix is established, which includes a set of nodes, a set of edges, and a weight matrix. The value of each element in the weight matrix is calculated based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance. The weight matrix reflects the real-time electrical connection strength between power grid nodes.
[0015] Based on the disturbance components of the different frequency bands and the dynamic topology identification matrix, the prediction parameter vector is calculated by recursive least squares method;
[0016] The prediction error is calculated within the sliding time window. When the prediction error is greater than a first preset threshold or the change in the dynamic topology recognition matrix is greater than a second preset threshold, the update calculation of the prediction parameter vector is triggered.
[0017] Based on the updated prediction parameter vector, the dynamic topology identification matrix, and the real-time disturbance data of the disturbance source node, disturbance propagation prediction data is calculated. The disturbance propagation prediction data includes the predicted disturbance amplitude, prediction time, and propagation path of each target node.
[0018] In one alternative implementation,
[0019] Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charging and discharging efficiency, and response speed of the energy storage. The steps to achieve global optimization through a consensus protocol include:
[0020] Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed, and energy storage resources are divided into inertial response, transient stability and power balance energy storage groups according to the response time.
[0021] The local optimization controller of the inertial response energy storage group constructs a first objective function based on frequency deviation, frequency change rate, and energy storage output power; the local optimization controller of the transient stable energy storage group constructs a second objective function based on voltage deviation, tie-line power deviation, and energy storage output power; and the local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second, and third objective functions are all constrained by the remaining capacity, charge / discharge efficiency, and response speed of the corresponding energy storage group.
[0022] A basic communication topology is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology is optimized based on the weighted adjacency matrix. A distributed consensus protocol considering delay compensation is adopted to achieve global optimization control of the energy storage groups.
[0023] In one alternative implementation,
[0024] The steps of collecting communication state parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consensus protocol that considers delay compensation include:
[0025] The communication quality factor, latency factor, and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method.
[0026] The network connectivity is calculated based on the weighted adjacency matrix. When the network connectivity is lower than a preset connectivity threshold, the communication topology is reconstructed using the minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topology with optimal latency performance.
[0027] Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative computation, and the adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iteration convergence speed.
[0028] When the state deviation of an energy storage node exceeds a preset deviation threshold, the state information is broadcast, and the adjacent nodes that receive the broadcast information update their respective optimization variables.
[0029] The communication connectivity of energy storage nodes is periodically calculated. When an abnormal node communication is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
[0030] In one alternative implementation,
[0031] The steps of designing an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal, constructing a value network and a policy network, wherein the value network evaluates the grid state and the remaining lifetime cost of the energy storage unit, and the policy network outputs the optimal charging and discharging power of the energy storage unit based on the value evaluation results, include:
[0032] The frequency deviation, frequency change rate and energy storage output power of the inertial response energy storage group, the voltage deviation, tie-line power deviation and energy storage output power of the transient stable energy storage group, and the load power deviation and operating cost of the power balance energy storage group in the distributed model predictive control framework are respectively constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function.
[0033] A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs the state value evaluation result based on the operating status of the energy storage group, the remaining lifetime cost, and the grid status; the strategy network updates the network parameters and outputs the charging and discharging power of the energy storage group based on the state value evaluation result using the strategy gradient method.
[0034] The charging and discharging power of the energy storage group, the comprehensive reward signal, and the grid status are used to construct an experience playback pool. A value network and a policy network are trained based on the data samples in the experience playback pool. The value network uses time-series differential error for parameter updates, and the learning rate of the policy network is adaptively adjusted according to historical gradients.
[0035] In one alternative implementation,
[0036] The adaptive adjustment steps for the weights of each evaluation function in the comprehensive reward signal include:
[0037] A control performance index is constructed, which is calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within the evaluation time window; the sensitivity of the weights of the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function to the control performance index is calculated.
[0038] The root mean square of frequency deviation, root mean square of voltage deviation, and root mean square of power deviation are collected to construct a state evaluation vector. The product of the state evaluation vector and the sensitivity is exponentially decayed and used as the weight update amount.
[0039] The energy storage life loss assessment value is calculated based on the energy storage group's discharge depth and charge / discharge power, and the energy storage life loss assessment value is mapped to weight adjustment constraints.
[0040] Construct a multi-objective optimization function that includes the control performance index, the energy storage lifetime loss assessment value, and the weight change amount; calculate the weight optimization direction based on the gradient of the multi-objective optimization function; and determine the weight update value according to the preset weight adjustment step size.
[0041] The weight update values are smoothed using an exponential smoothing method.
[0042] In one alternative implementation,
[0043] The steps for updating the parameters of the value network and the policy network include:
[0044] A dual-network structure of a target network and an evaluation network is constructed for the value network. The temporal difference error is calculated based on the current state value, the next state value, and the immediate reward. The evaluation network constructs a loss function based on the temporal difference error and continuously updates the network parameters. The target network determines the parameter update period based on the prediction error change rate of the evaluation network.
[0045] The action probability distribution output by the policy network is calculated. The importance sampling weights are calculated based on the probability ratio of the new and old policies. The cutoff range is determined according to the variance of the importance sampling weights. The evaluation result of the value network is used as the baseline function. The weighted sum of the temporal difference error is calculated using the generalized advantage estimation method. The policy network parameters are updated based on the weighted sum.
[0046] A hybrid strategy is constructed that includes exploration and exploitation items. The weight of the exploration items gradually decreases as the training process progresses, while the weight of the exploitation items is dynamically adjusted based on the value evaluation results.
[0047] The timeliness assessment value and importance assessment value of the data samples are calculated. The timeliness assessment value is obtained by an exponential function of the sample storage time, and the importance assessment value is obtained by the absolute value of the time series difference error. Based on the timeliness assessment value and the importance assessment value, the sample priority is constructed. The data samples in the experience pool are sorted according to the sample priority. The capacity of the experience pool is dynamically adjusted according to the distribution range of the data sample priority.
[0048] A second aspect of the present invention provides a coordinated and complementary regulation system for energy storage power stations and power grids, comprising:
[0049] The first unit is used to perform multi-scale decomposition on real-time frequency and voltage data from the power grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data in combination with the power grid topology.
[0050] The second unit is used to construct a distributed model prediction control framework based on the disturbance propagation prediction data. According to the response time, the energy storage resources are divided into multiple energy storage groups. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol.
[0051] The third unit is used to design an adaptive dynamic programming controller. The control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group. The policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results.
[0052] The fourth unit is used to recalculate the state assessment based on the new network structure when a change in the power grid topology is detected, update the optimal charging and discharging power output of the strategy network, control each energy storage group to perform charging and discharging operations according to the optimized optimal charging and discharging power, and feed back the operating data to the adaptive dynamic programming controller.
[0053] A third aspect of the present invention provides an electronic device, comprising:
[0054] processor;
[0055] Memory used to store processor-executable instructions;
[0056] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0057] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0058] This invention achieves accurate perception and early response to power grid fluctuations by decomposing power grid disturbance data at multiple scales and combining it with topology for prediction, thus significantly improving power grid stability and anti-disturbance capability.
[0059] This invention employs a distributed model predictive control framework and a consensus protocol to perform group optimization based on the remaining capacity, charging and discharging efficiency, and response speed of energy storage, thereby achieving reasonable scheduling and global optimization of energy storage resources and improving the overall economy and operating efficiency of the system.
[0060] The adaptive dynamic programming controller designed in this invention can evaluate the remaining lifetime cost of the energy storage unit and adjust the charging and discharging strategy in real time. At the same time, it can quickly recalculate the optimal control strategy when the grid topology changes, effectively extending the service life of the energy storage equipment and enhancing the system's adaptability and robustness. Attached Figure Description
[0061] Figure 1 This is a flowchart illustrating the energy storage power station and grid coordinated and complementary regulation method according to an embodiment of the present invention;
[0062] Figure 2 A comparison chart of comprehensive control performance under frequency disturbance conditions;
[0063] Figure 3 This is a flowchart illustrating the adaptive adjustment of the comprehensive reward signal weights according to the present invention.
[0064] Figure 4 A performance comparison chart for updating parameters of the value network and the policy network. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0066] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0067] Figure 1 This is a flowchart illustrating the coordinated and complementary regulation method between energy storage power stations and the power grid according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0068] Multi-scale decomposition is performed on real-time frequency and voltage data from the power grid to obtain disturbance components in different frequency bands. Combined with the power grid topology, disturbance propagation prediction data is calculated.
[0069] Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to the response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol.
[0070] An adaptive dynamic programming controller is designed, and the control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group, and the policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results.
[0071] When a change in grid topology is detected, the state assessment is recalculated based on the new network structure, and the optimal charging and discharging power output by the strategy network is updated. Based on the optimized optimal charging and discharging power, each energy storage group is controlled to perform charging and discharging operations, and the operating data is fed back to the adaptive dynamic programming controller.
[0072] In one optional implementation, the steps of performing multi-scale decomposition on real-time frequency and voltage data from the power grid side to obtain disturbance components in different frequency bands, and calculating disturbance propagation prediction data in conjunction with the power grid topology include:
[0073] The real-time frequency data and voltage data are decomposed into perturbation components in different frequency bands through multi-scale decomposition.
[0074] A dynamic topology identification matrix is established, which includes a set of nodes, a set of edges, and a weight matrix. The value of each element in the weight matrix is calculated based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance. The weight matrix reflects the real-time electrical connection strength between power grid nodes.
[0075] Based on the disturbance components of the different frequency bands and the dynamic topology identification matrix, a prediction parameter vector is calculated by recursive least squares method. The prediction parameter vector is used as the weight coefficients corresponding to the disturbance components of the different frequency bands.
[0076] The prediction error is calculated within the sliding time window. When the prediction error is greater than a first preset threshold or the change in the dynamic topology recognition matrix is greater than a second preset threshold, the update calculation of the prediction parameter vector is triggered.
[0077] Based on the updated prediction parameter vector, the dynamic topology identification matrix, and the real-time disturbance data of the disturbance source node, disturbance propagation prediction data is calculated. The disturbance propagation prediction data includes the predicted disturbance amplitude, prediction time, and propagation path of each target node.
[0078] For example, real-time frequency and voltage data are collected from various monitoring points in the power grid. This data can be acquired using a wide-area measurement system (WAMS), with a sampling frequency typically of 100 Hz and a collection duration of multiple consecutive time windows, each lasting 10 seconds.
[0079] Multi-scale decomposition of the acquired real-time frequency and voltage data can be performed using the Empirical Mode Decomposition (EMD) method. Specifically, EMD decomposition of the frequency signal f(t) at a certain monitoring point yields n intrinsic mode function (IMF) components c1(t), c2(t), ..., cn(t) and a residual term rn(t). These IMF components represent the disturbance characteristics of different frequency bands. For example, for the frequency data of a 110kV substation, EMD decomposition yields five IMF components, corresponding to the disturbance characteristics of the 0.5-2.5Hz, 0.2-0.5Hz, 0.05-0.2Hz, 0.01-0.05Hz, and 0-0.01Hz frequency bands, respectively. Similarly, voltage data is processed in a similar manner to obtain the disturbance components of the corresponding frequency bands.
[0080] The dynamic topology identification matrix includes a node set V, an edge set E, and a weight matrix W. The node set V represents all bus nodes in the power grid. For example, if a regional power grid contains 15 500kV nodes and 30 220kV nodes, then |V| = 45. The edge set E represents the connections between nodes, such as lines and transformers. The weight matrix W reflects the electrical connection strength between nodes. The calculation of each element wij considers the node voltage amplitudes Vi and Vj, voltage phase angles θi and θj, and the inter-node reactance Xij. Specifically, wij can be expressed as the product of Vi and Vj, divided by Xij, and considering the sinθij factor. For example, for two nodes with voltages of 525kV and 515kV respectively, a phase angle difference of 5 degrees, and a connection reactance of 0.1 ohms, their weight value is approximately 2625. A larger weight value indicates a stronger electrical connection between nodes and easier disturbance propagation. For nodes that are not directly connected, the weight value is set to 0.
[0081] The prediction parameter vector α contains multiple elements, each corresponding to the weighting coefficient of the disturbance component in different frequency bands. During calculation, the disturbance data within the historical observation window is constructed into a matrix, and then combined with the dynamic topology identification matrix, the value of α is estimated using the recursive least squares method. For example, for a disturbance in the 0.5-2.5Hz frequency band of a 500kV substation, α1=0.85, α2=0.72, α3=0.63, etc., are obtained, indicating that the disturbance in this frequency band attenuates relatively quickly during propagation. However, for the disturbance in the 0-0.01Hz frequency band, α1=0.98, α2=0.95, α3=0.91, etc., are obtained, indicating that low-frequency disturbances attenuate more slowly during propagation.
[0082] In practical applications, the prediction error needs to be continuously calculated within a sliding time window. The time window length is typically set to 5-10 seconds, and the sliding step size is 1 second. The prediction error ε is defined as the root mean square error between the predicted value and the actual observed value. When ε is greater than a first preset threshold (e.g., 0.05), or when the change in the dynamic topology identification matrix is greater than a second preset threshold (e.g., the change in weight matrix elements exceeds 15%), the prediction parameter vector update calculation is triggered. This mechanism can adapt to changes in the power grid operating state and improve prediction accuracy.
[0083] For a disturbance ds(t) in a certain frequency band from the source node s, predict the disturbance di(t+Δt) propagating to the target node i, considering the prediction parameter vector α and the propagation delay Δt between nodes. The specific calculation method is as follows: determine the real-time disturbance data ds(t) of the source node s, and then calculate the attenuation coefficient of the disturbance during propagation based on the weight coefficient αsi corresponding to the frequency band in the prediction parameter vector α. For directly connected nodes, the predicted disturbance amplitude di(t+Δt) = αsi × ds(t) × wsi, where wsi is the corresponding weight value in the dynamic topology identification matrix; for cases requiring propagation through multiple nodes, the cumulative attenuation effect is calculated using matrix multiplication. The propagation delay Δt is calculated based on the propagation speed of electromagnetic waves in the power system (approximately 60% of the speed of light). For example, for two nodes 200 kilometers apart, the disturbance propagation delay is approximately 6.67 milliseconds (200km ÷ 0.6 × 300,000km / s). By combining the weight information in the dynamic topology identification matrix, the optimal propagation path is determined using Dijkstra's shortest path algorithm, i.e., selecting the path with the smallest ∑(1 / wij) among all paths. For complex network structures, suboptimal paths are also considered as alternative propagation paths. The final output disturbance propagation prediction data includes: the predicted disturbance amplitude of each target node (accurate to 0.01Hz or 0.01kV), the predicted arrival time (accurate to milliseconds), and a detailed propagation path description (including the sequence of all traversed nodes and their corresponding propagation times). This prediction data will serve as an important basis for subsequent energy storage regulation decisions.
[0084] This invention enables accurate prediction of power grid disturbance propagation, which is of great significance for improving the safe and stable operation of the power grid. Through multi-scale decomposition and dynamic topology identification, it can adapt to the propagation characteristics of disturbances in different frequency bands and has good adaptability to changes in power grid topology.
[0085] In one optional implementation, based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charging and discharging efficiency, and response speed of the energy storage, and achieves global optimization through a consensus protocol. The steps include:
[0086] Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed, and energy storage resources are divided into inertial response, transient stability and power balance energy storage groups according to the response time.
[0087] The local optimization controller of the inertial response energy storage group constructs a first objective function based on frequency deviation, frequency change rate, and energy storage output power; the local optimization controller of the transient stable energy storage group constructs a second objective function based on voltage deviation, tie-line power deviation, and energy storage output power; and the local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second, and third objective functions are all constrained by the remaining capacity, charge / discharge efficiency, and response speed of the corresponding energy storage group.
[0088] A basic communication topology is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology is optimized based on the weighted adjacency matrix. A distributed consensus protocol considering delay compensation is adopted to achieve global optimization control of the energy storage groups.
[0089] For example, based on the response time characteristics of energy storage resources, energy storage resources can be classified into three categories: inertial response energy storage groups, transient stability energy storage groups, and power balance energy storage groups. The response time of inertial response energy storage groups is typically in the millisecond range, such as supercapacitors and flywheel energy storage devices; the response time of transient stability energy storage groups is in the second range, such as lithium-ion batteries and lead-acid batteries; and the response time of power balance energy storage groups is in the minute range, such as pumped hydro storage and compressed air energy storage.
[0090] For the inertial response energy storage unit, a local optimization controller constructs a first objective function. This objective function primarily considers frequency deviation, frequency change rate, and energy storage output power. Frequency deviation refers to the difference between the current system frequency and the rated frequency, usually measured in Hertz; the frequency change rate represents the change in frequency per unit time, measured in Hertz per second; and the energy storage output power is the active power supplied by the energy storage device to the grid, measured in megawatts. In specific implementation, when the system frequency is detected to be below the rated value of 49.8 Hz and the frequency change rate is -0.2 Hz / s, the inertial response energy storage unit will rapidly increase its output power to 80% of the rated power, completing the response within 200 milliseconds to provide system inertia support.
[0091] For transiently stable energy storage units, a second objective function is constructed by the local optimization controller. This objective function is primarily based on voltage deviation, tie-line power deviation, and energy storage output power. Voltage deviation refers to the difference between the node voltage and the rated voltage, expressed in per-unit values; tie-line power deviation represents the difference between the actual transmitted power and the planned transmitted power, expressed in megawatts (MW); energy storage output power is also expressed in megawatts. In practical applications, when a system voltage drop to 0.92 per-unit value and a tie-line power deviation reaching 50 MW are detected, the transiently stable energy storage unit will gradually increase its output power to 65% of the rated power within 2 seconds to stabilize the system voltage and power transmission.
[0092] For power balancing energy storage units, a third objective function is constructed by the local optimization controller. This objective function mainly considers load power deviation and energy storage operating costs. Load power deviation refers to the difference between the actual load and the predicted load, expressed in megawatts (MW). Energy storage operating costs include factors such as equipment depreciation and charging / discharging losses, typically calculated in yuan per kilowatt-hour. In practical applications, when fluctuations in renewable energy output cause a load power deviation of 100MW, the power balancing energy storage unit will gradually adjust its output power within 5 minutes to minimize system power imbalance while simultaneously optimizing operating costs.
[0093] All objective functions are constrained by the physical constraints of their respective energy storage units. The remaining capacity constraint ensures that the state of charge (SOC) of the energy storage device remains within a safe range, typically 20% to 80%; the charge / discharge efficiency constraint takes into account losses during energy conversion, such as the charge / discharge efficiency of lithium-ion batteries, which is approximately 95%; and the response speed constraint considers the physical characteristics of the device, such as power ramp-up limits, where supercapacitors can reach 100% of rated power per second, while pumped hydro storage can only reach 10% of rated power per minute.
[0094] To achieve coordinated control among energy storage groups, a basic communication topology needs to be established. Communication status parameters between energy storage nodes are collected, including communication latency, packet loss rate, and bandwidth utilization. A weighted adjacency matrix is constructed based on these parameters. Matrix element aij represents the communication quality weight between node i and node j, with weight values ranging from 0 to 1, where 0 represents no connection and 1 represents an ideal connection. For example, when the communication latency between two nodes is 15ms and the packet loss rate is 2%, the weight can be set to 0.9; while when the communication latency reaches 50ms and the packet loss rate is 10%, the weight can be reduced to 0.6.
[0095] The communication topology is optimized based on a weighted adjacency matrix. The optimization process includes removing low-quality connections (such as those with a weight below 0.4), enhancing communication capabilities on critical paths (such as increasing the bandwidth of the central node), and establishing redundant communication paths to improve system reliability. The optimized communication topology should meet network connectivity requirements, ensuring that there is at least one communication path between any two nodes.
[0096] A distributed consensus protocol considering latency compensation is employed to achieve global optimal control of the energy storage group. This protocol uses an iterative approach to bring the control variables of each energy storage group towards consistency. The latency compensation mechanism reduces the impact of communication delays by predicting future state values.
[0097] This invention divides energy storage resources into different energy storage groups based on response time, enabling precise matching of the different time-scale characteristics of grid disturbances. Dedicated local optimization controllers are configured for different energy storage groups, allowing each group to address specific grid problems such as frequency deviation, voltage deviation, and power balance. A distributed consensus protocol enables coordinated control among energy storage groups, avoiding the communication bottlenecks and single-point-of-failure risks of centralized control, thus improving the reliability and robustness of the system control.
[0098] In one optional implementation, the steps of collecting communication state parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consensus protocol that considers delay compensation include:
[0099] The communication quality factor, latency factor and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method. The weighted adjacency matrix is used to characterize the communication connection strength between energy storage nodes.
[0100] The network connectivity is calculated based on the weighted adjacency matrix. When the network connectivity is lower than a preset connectivity threshold, the communication topology is reconstructed using the minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topology with optimal latency performance.
[0101] Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative computation. The adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iteration convergence speed, and the delay compensation term is used to reduce the impact of communication delay on global optimization performance.
[0102] When the state deviation of an energy storage node exceeds a preset deviation threshold, the state information is broadcast, and the adjacent nodes that receive the broadcast information update their respective optimization variables.
[0103] The communication connectivity of energy storage nodes is periodically calculated. When an abnormal node communication is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
[0104] For example, during the stage of collecting communication status parameters and constructing the weighted adjacency matrix between energy storage nodes, the communication quality factor, delay factor, and bandwidth utilization between each node are monitored and collected in real time. The communication quality factor can be obtained by measuring the signal-to-noise ratio (SNR). For example, when the SNR is higher than 35dB, the communication quality factor can be set to 0.9. The delay factor is obtained by recording the time difference between sending and receiving data packets. For example, when the delay is less than 15ms, the delay factor can be set to 0.8. The bandwidth utilization is obtained by calculating the ratio of the actual data transmission volume to the theoretical bandwidth of the communication link. When the utilization is less than 50%, it can be set to 0.7. Fuzzy comprehensive evaluation is used to process these three parameters. First, a membership function is established to map each parameter to the interval [0,1]. For example, the communication quality factor can be set to three fuzzy subsets: "excellent," "good," and "poor." The corresponding membership functions take their maximum values when the SNR is higher than 30dB, between 20-30dB, and lower than 20dB, respectively. The delay factor and bandwidth utilization are also constructed using a similar method to construct fuzzy subsets. The weight vector is then determined, typically setting the weights for communication quality factor, latency factor, and bandwidth utilization to 0.4, 0.35, and 0.25, respectively. Finally, a comprehensive evaluation value is obtained through fuzzy synthesis calculation, which serves as an element of the weighted adjacency matrix. For example, a weighted value of 0.76 between nodes i and j indicates a relatively strong communication connection.
[0105] During the communication topology optimization phase, network connectivity is calculated based on the constructed weighted adjacency matrix. Network connectivity can be obtained by calculating the eigenvalues of the weighted adjacency matrix, specifically the magnitude of the second smallest eigenvalue. When the calculated network connectivity is lower than a preset threshold (e.g., 0.15), the communication topology reconstruction process is triggered. The reconstruction process employs a minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes. This algorithm first calculates the spatiotemporal distance between each pair of nodes, which is a weighted combination of physical distance and communication latency. For example, for two nodes 50 meters apart with a communication latency of 25ms, setting the physical distance weight to 0.4 and the latency weight to 0.6, the calculated spatiotemporal distance is 0.4 × 50 + 0.6 × 25 = 35. Using the spatiotemporal distance as the edge weight, an improved Kruskal algorithm is applied to construct the minimum spanning tree, gradually adding edges starting from the edge with the minimum spatiotemporal distance until all nodes are connected and no loops are formed. Furthermore, to ensure the communication stability of important nodes, critical backup links need to be added, typically selecting the path with the second highest connectivity as a backup. After reconstruction, a topology with optimal latency performance is obtained.
[0106] In the distributed consensus protocol implementation phase, based on the reconstructed communication topology, a consensus protocol with latency compensation is used for iterative computation. The core of this protocol is to ensure consistency in the key states (such as power allocation ratios) of each energy storage unit through information exchange between nodes. In practical applications, the state update formula for each energy storage node i at time t includes an adaptive step size term and a latency compensation term. The adaptive step size is dynamically adjusted according to the iterative convergence speed, initially set to 0.05. Each iteration calculates the state change rate; when the change rate is less than 80% of the previous iteration, the step size increases by 10%; when the change rate is greater than 120% of the previous iteration, the step size decreases by 10%, ensuring both rapid and stable convergence. The latency compensation term is used to reduce the impact of communication latency on optimization performance, implemented through a predictive compensation method. For example, when a communication latency of 30ms is detected between node j and node i, node i predicts its current actual state based on the historical state change trend of node j, thereby reducing the error caused by latency. Experimental results show that after applying latency compensation, the convergence time can be reduced by approximately 25%, and the global optimization accuracy can be improved by approximately 18%.
[0107] In the energy storage node state deviation handling mechanism, the deviation of each node's state from the global target is continuously monitored. When a node's state deviation (such as power allocation error) exceeds a preset threshold (usually set to 5%), a state information broadcasting mechanism is triggered. For example, when the charging power of an energy storage unit deviates from the target value by 8%, the node will immediately broadcast its latest state information to the network. Neighboring nodes that receive the broadcast information will update their respective optimization variables accordingly, such as charging and discharging power setpoints and voltage regulation parameters. This event-triggered communication mechanism significantly reduces network communication load, and in actual measurements, it can reduce data transmission volume by approximately 40%.
[0108] In the communication anomaly handling mechanism, the communication connectivity of energy storage nodes, i.e., the number of effective connections between a node and other nodes, is calculated periodically (e.g., every 500ms). When a node communication anomaly is detected (e.g., failure to receive expected data packets three consecutive times), the energy storage node with the highest connectivity is automatically selected as the backup communication path. For example, when communication between node A and node B is interrupted, node C, which has good connections with both A and B, is selected as the relay point to establish the communication path ACB, maintaining the continuity of the global optimization process.
[0109] This invention can adaptively evaluate and optimize the communication quality between energy storage nodes, ensuring the efficient operation of the distributed control system in real-world communication environments. By employing a distributed consensus protocol with delay compensation and an adaptive step size adjustment mechanism, the adverse effects of communication delay on control performance are effectively overcome. Through a broadcast mechanism triggered by state deviations and a backup communication path selection strategy, the system's fault tolerance and continuous operation capability under communication anomalies are significantly improved, ensuring the stability and reliability of grid control.
[0110] In one optional implementation, an adaptive dynamic programming controller is designed, using the control effect of the distributed model predictive control framework as a reward signal, and a value network and a policy network are constructed. The value network evaluates the grid state and the remaining lifetime cost of the energy storage unit, and the policy network outputs the optimal charging and discharging power of the energy storage unit based on the value evaluation results. The steps include:
[0111] The frequency deviation, frequency change rate and energy storage output power of the inertial response energy storage group, the voltage deviation, tie-line power deviation and energy storage output power of the transient stable energy storage group, and the load power deviation and operating cost of the power balance energy storage group in the distributed model predictive control framework are respectively constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function.
[0112] A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs the state value evaluation result based on the operating status of the energy storage group, the remaining lifetime cost, and the grid status; the strategy network updates the network parameters and outputs the charging and discharging power of the energy storage group based on the state value evaluation result using the strategy gradient method.
[0113] The charging and discharging power of the energy storage group, the comprehensive reward signal, and the grid status are used to construct an experience playback pool. A value network and a policy network are trained based on the data samples in the experience playback pool. The value network uses time-series differential error for parameter updates, and the learning rate of the policy network is adaptively adjusted according to historical gradients.
[0114] Based on the aforementioned design requirements, this invention provides a specific implementation of an adaptive dynamic programming controller, which uses the control effect of a distributed model predictive control framework as a reward signal to construct a value network and a policy network.
[0115] For example, to construct a comprehensive reward signal, the control effects of different types of energy storage groups need to be evaluated. For inertial response energy storage groups, frequency deviation, frequency change rate, and energy storage output power are used to construct the first control effect evaluation function. When the system frequency deviation is 0.1Hz, the frequency change rate is 0.05Hz / s, and the energy storage output power is 0.8MW, the value of the first control effect evaluation function is 0.85; while when the system frequency deviation increases to 0.2Hz, the frequency change rate is 0.1Hz / s, and the energy storage output power is 1.2MW, the value of the first control effect evaluation function drops to 0.65.
[0116] For transiently stable energy storage units, voltage deviation, tie-line power deviation, and energy storage output power are used to construct a second control effect evaluation function. For example, when the voltage deviation is 2%, the tie-line power deviation is 5%, and the energy storage output power is 1.5MW, the value of the second control effect evaluation function is 0.78; when the voltage deviation increases to 4%, the tie-line power deviation is 8%, and the energy storage output power is 2.0MW, the value of the second control effect evaluation function decreases to 0.62.
[0117] For power balancing energy storage units, load power deviation and operating cost are used to construct a third control effect evaluation function. When the load power deviation is 3MW and the operating cost is 500 yuan / hour, the value of the third control effect evaluation function is 0.82; when the load power deviation increases to 5MW and the operating cost is 700 yuan / hour, the value of the third control effect evaluation function decreases to 0.70.
[0118] Based on the three control effect evaluation functions mentioned above, a comprehensive reward signal is constructed. The comprehensive reward signal considers the weighting factors of various energy storage groups; for example, the weights of the inertial response energy storage group, the transient stable energy storage group, and the power balance energy storage group are set to 0.4, 0.3, and 0.3, respectively. The comprehensive reward signal is obtained by weighted summation. In the example above, the comprehensive reward signal value is 0.82×0.4+0.78×0.3+0.82×0.3=0.808.
[0119] The inputs to the value network include the operating status of the energy storage unit, its remaining lifetime cost, and the grid status. The operating status of the energy storage unit includes its current state of charge (SOC), depth of charge / discharge, and number of charge / discharge cycles. For example, an energy storage unit may have a current SOC of 65%, have completed 350 charge / discharge cycles, and have a depth of charge / discharge of 80%. The remaining lifetime cost is calculated through depreciation and efficiency losses of the energy storage equipment; in this example, the remaining lifetime cost is 0.35 yuan per kilowatt-hour. The grid status includes bus voltage, line load factor, and system frequency. In this example, a test system has 10 buses, an average voltage of 220 kV, an average line load factor of 75%, and a system frequency of 49.9 Hz.
[0120] The value network employs a multilayer perceptron structure, comprising an input layer, two hidden layers, and an output layer. The number of nodes in the input layer is determined by the number of state variables; in this example, it has 12 nodes. The first hidden layer has 24 nodes, the second hidden layer has 16 nodes, and the output layer has one node representing the state value. The network uses the ReLU activation function, and the state value evaluation result is calculated through forward propagation; in this example, the state value evaluation result is 0.76.
[0121] The policy network updates its parameters using the policy gradient method based on the state value assessment results. The policy network also employs a multilayer perceptron structure, with the same number of nodes in the input layer as the value network. The two hidden layers have 32 and 24 nodes respectively, and the number of nodes in the output layer is twice the number of energy storage groups, representing charging and discharging power. In this example, there are 3 energy storage groups, therefore the output layer has 6 nodes. The policy network uses the Softmax activation function and outputs the probability distribution of charging and discharging power for each energy storage group.
[0122] The learning rate of the policy network is adaptively adjusted based on historical gradients, with an initial learning rate of 0.01. When the gradient direction is consistent for five consecutive iterations, the learning rate increases by 20%; when the gradient direction changes continuously, the learning rate decreases by 15%. In the example, the learning rate is adjusted to 0.0135 during the 30th iteration of training.
[0123] The construction process of the experience replay pool is as follows: the charging and discharging power of the energy storage group, the comprehensive reward signal, and the grid status are stored as data samples. The capacity of the experience replay pool is set to 10,000 samples, and a random sampling method is used to draw data in batches of 128 for training each time. When the experience replay pool is full, a first-in-first-out strategy is used to update the data.
[0124] The value network uses temporal difference error for parameter updates. The temporal difference objective consists of the current reward and the next state value estimate, with a discount factor set to 0.95. In the example, the temporal difference objective for a certain state transition is calculated as 0.808 + 0.95 × 0.76 = 1.53, and the temporal difference error is 1.53 - 0.76 = 0.77.
[0125] Figure 2 This is a comparison chart of the comprehensive control performance under frequency disturbance conditions. The thick solid black dot represents the system frequency response of this invention, showing a faster frequency recovery speed, smaller fluctuation amplitude, and almost zero steady-state error. In contrast, the traditional MPC control method (dashed lines, marked with squares) exhibits a larger frequency deviation and a longer recovery time. Simultaneously, the comprehensive reward signal of this invention (dashed black dot) is significantly higher than that of the traditional method, rapidly recovering and stabilizing at a higher level, indicating that this control strategy has significant advantages in frequency control capability, system stability, and economy. Especially in the early stages of disturbance, the control strategy of this invention demonstrates superior anti-disturbance capability and rapid recovery characteristics, fully proving the technological innovation and practical value of the adaptive dynamic programming controller in power system energy storage control applications.
[0126] This invention constructs a value network and a policy network by using the control effect of a distributed model predictive control framework as a reward signal, thereby achieving autonomous learning and optimization of energy storage control strategies. This method combines traditional control theory with deep reinforcement learning techniques, overcoming the limitations of single control methods in handling the complex dynamic characteristics of the power grid. The value network comprehensively evaluates the grid state and the remaining lifetime cost of the energy storage, while the policy network optimizes energy storage charging and discharging decisions based on the value assessment results, achieving a balance between short-term control effectiveness and long-term economic efficiency. Through an experience replay pool and an adaptive learning rate adjustment mechanism, learning efficiency and convergence performance are improved, enabling the control strategy to continuously optimize and adapt to changes in the power grid operating environment, significantly enhancing the economy and control performance of the energy storage system.
[0127] In one optional implementation, the adaptive adjustment step of the weights of each evaluation function in the comprehensive reward signal includes:
[0128] A control performance index is constructed, which is calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within the evaluation time window; the sensitivity of the weights of the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function to the control performance index is calculated.
[0129] The root mean square of frequency deviation, root mean square of voltage deviation, and root mean square of power deviation are collected to construct a state evaluation vector. The product of the state evaluation vector and the sensitivity is exponentially decayed and used as the weight update amount.
[0130] The energy storage life loss assessment value is calculated based on the energy storage group's discharge depth and charge / discharge power, and the energy storage life loss assessment value is mapped to weight adjustment constraints.
[0131] Construct a multi-objective optimization function that includes the control performance index, the energy storage lifetime loss assessment value, and the weight change amount; calculate the weight optimization direction based on the gradient of the multi-objective optimization function; and determine the weight update value according to the preset weight adjustment step size.
[0132] The weight update values are smoothed using an exponential smoothing method to obtain the final values of the first control effect evaluation function weight, the second control effect evaluation function weight, and the third control effect evaluation function weight. The sum of the first control effect evaluation function weight, the second control effect evaluation function weight, and the third control effect evaluation function weight is 1.
[0133] For example, in combination Figure 3The flowchart for adaptive adjustment of the integrated reward signal weights is used to illustrate this. The control performance index is obtained by calculating the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within the evaluation time window. Specifically, a 10-second evaluation time window is selected, with a sampling interval of 0.1 seconds to collect frequency, voltage, and power data in the microgrid. Assuming that within a certain evaluation time window, the cumulative sum of squares of frequency deviation is 0.25, the cumulative sum of squares of voltage deviation is 0.36, and the cumulative sum of squares of power deviation is 0.49, the control performance index can be expressed as a weighted combination of these three factors, with initial weights set to 0.4, 0.3, and 0.3, respectively.
[0134] The sensitivity calculation for the control performance index is based on perturbation analysis. Small perturbations (e.g., 0.01) are applied to the weights of the first, second, and third control performance evaluation functions, respectively, and the changes in the control performance index are observed. For example, when the weight of the first control performance evaluation function increases from 0.4 to 0.41, if the control performance index increases from 0.345 to 0.348, then the sensitivity of the first control performance evaluation function weight to the control performance index is 0.3. Similarly, the sensitivity of the other weights is calculated, resulting in a sensitivity vector of [0.3, 0.2, 0.25].
[0135] The state assessment vector is constructed by collecting the root mean square (RMS) of frequency deviation, voltage deviation, and power deviation. The RMS of frequency deviation is obtained by taking the square root of the average of the squares of the differences between the measured frequency and the rated frequency (e.g., 50Hz) within the assessment window; the RMS of voltage deviation is obtained by taking the square root of the average of the squares of the per-unit values of the differences between the measured voltage and the rated voltage (e.g., 220kV) within the assessment window; and the RMS of power deviation is obtained by taking the square root of the average of the squares of the per-unit values of the differences between the actual power and the planned power within the assessment window. In practical implementation, assuming the RMS of frequency deviation is 0.05, the RMS of voltage deviation is 0.06, and the RMS of power deviation is 0.07 at a certain moment, the state assessment vector is [0.05, 0.06, 0.07]. The product of the state assessment vector and the sensitivity is [0.015, 0.012, 0.0175]. Introducing an exponential decay factor of 0.9, the product result is subjected to exponential decay processing, resulting in weight update amounts of [0.0135, 0.0108, 0.01575].
[0136] Energy storage lifetime loss assessment is a crucial constraint for weight adjustment. The energy storage lifetime loss assessment value is calculated based on the energy storage group's depth of discharge and charge / discharge power. Assuming the current depth of discharge is 20% and the charge / discharge power is 100kW, and the cycle life is determined to be 5000 cycles using a lookup table, the energy storage lifetime loss assessment value is 0.0002 (i.e., 1 / 5000). This assessment value is mapped to a weight adjustment constraint: when the energy storage lifetime loss assessment value is greater than 0.0005, the change in the weight of the first control effect evaluation function is limited to no more than 0.05; when the energy storage lifetime loss assessment value is less than 0.0005, the restriction is relaxed, allowing a maximum change in the weight of the first control effect evaluation function of 0.1.
[0137] The multi-objective optimization function comprises three parts: a control performance index, an energy storage lifetime loss assessment value, and a weight change. The weight of the control performance index is 0.6, the weight of the energy storage lifetime loss assessment value is 0.3, and the weight change value is 0.1. The weight optimization direction is calculated based on the gradient descent method, assuming the obtained gradient direction is [-0.05, 0.03, 0.02]. With a preset weight adjustment step size of 0.02, the updated weight values are [-0.001, 0.0006, 0.0004].
[0138] The updated weight values are exponentially smoothed using a smoothing factor of 0.8. Assuming the weights of the first, second, and third control effect evaluation functions at the previous time step were 0.4, 0.3, and 0.3 respectively, the smoothed updated weight values are [-0.0008, 0.00048, 0.00032]. These updated values are added to the original weights to obtain new weight values [0.3992, 0.30048, 0.30032]. Since the sum of the three weights must be 1, normalization is required, resulting in final weight values [0.399, 0.3005, 0.3005].
[0139] This invention achieves dynamic optimization of the reward signal by constructing control performance indicators and calculating the sensitivity of the weights of each evaluation function. By incorporating the energy storage lifetime loss assessment value into the weight adjustment constraints, it ensures both grid stability and energy storage asset lifetime protection. Through the construction of a multi-objective optimization function and a smoothing mechanism, a balance is achieved between control performance, energy storage lifetime, and weight changes, avoiding control instability caused by drastic weight fluctuations, ensuring a smooth transition and continuity of grid control, and improving the overall operating efficiency of the system.
[0140] In one optional implementation, the parameter update steps for the value network and the policy network include:
[0141] A dual-network structure of a target network and an evaluation network is constructed for the value network. The temporal difference error is calculated based on the current state value, the next state value, and the immediate reward. The evaluation network constructs a loss function based on the temporal difference error and continuously updates the network parameters. The target network determines the parameter update period based on the prediction error change rate of the evaluation network.
[0142] The action probability distribution output by the policy network is calculated. The importance sampling weights are calculated based on the probability ratio of the new and old policies. The cutoff range is determined according to the variance of the importance sampling weights. The evaluation result of the value network is used as the baseline function. The weighted sum of the temporal difference error is calculated using the generalized advantage estimation method. The policy network parameters are updated based on the weighted sum.
[0143] A hybrid strategy is constructed that includes exploration and exploitation items. The weight of the exploration items gradually decreases as the training process progresses, while the weight of the exploitation items is dynamically adjusted based on the value evaluation results.
[0144] The timeliness assessment value and importance assessment value of the data samples are calculated. The timeliness assessment value is obtained by an exponential function of the sample storage time, and the importance assessment value is obtained by the absolute value of the time series difference error. Based on the timeliness assessment value and the importance assessment value, the sample priority is constructed. The data samples in the experience pool are sorted according to the sample priority. The capacity of the experience pool is dynamically adjusted according to the distribution range of the data sample priority.
[0145] For example, a dual-network structure of a target network and an evaluation network is constructed for the value network. The value network adopts a multilayer perceptron structure. The number of nodes in the input layer is the state dimension (including the state of charge of the energy storage group, charging and discharging power, grid frequency, voltage, power deviation, and remaining lifetime characteristics of the energy storage, a total of 12 features). The hidden layers have 24 nodes and 16 nodes respectively, and the output layer has 1 node representing the state value evaluation result. The evaluation network and the target network have the same network structure, but the parameter update mechanism is different. The parameters of the evaluation network are updated in real time, while the parameters of the target network are periodically copied from the evaluation network. Specifically, the calculation process of the time-series difference error is as follows: batch data (batch size of 128) is randomly sampled from the experience playback pool. Each data point includes the current state, action (charging and discharging power of the energy storage group), reward (comprehensive reward signal), next state, and termination flag. For each data point, the target network is used to calculate the value estimate of the next state, and the evaluation network is used to calculate the value estimate of the current state. The time-series difference error is equal to the immediate reward plus a discount factor (set to 0.95) multiplied by the next state value estimate, and then subtracted from the current state value estimate. For example, when the immediate reward is 0.808, the current state value estimate is 0.76, and the next state value estimate is 0.81, the temporal difference error is 0.808 + 0.95 × 0.81 - 0.76 = 0.82. The loss function of the evaluation network is the sum of squares of the temporal difference errors, and the Adam optimizer is used for parameter updates with a learning rate of 0.001. The target network does not update directly through gradient descent but periodically copies parameters from the evaluation network. The update cycle is not fixed but dynamically adjusted based on the rate of change of the prediction error of the evaluation network. When the rate of change of the prediction error (the relative change in prediction error over 10 consecutive iterations) is less than 0.05, the target network parameters are updated. In practical applications, the target network update cycle is approximately once every 500 iterations in the initial stage, and gradually extends to once every 2000 iterations as training progresses.
[0146] The action probability distribution output by the policy network is calculated and the parameters are updated. The policy network also adopts a multilayer perceptron structure, with the same input layer as the value network, 32 and 24 hidden layers respectively, and twice the number of nodes in the output layer (in the example, there are 3 energy storage groups, so the output layer has 6 nodes, representing charging and discharging power respectively). The output is converted into an action probability distribution through a Softmax function. During training, the network parameters before each policy update are saved to calculate the probability ratio between the old and new policies. Specifically, for each action sample in the experience pool, the ratio of its probability under the new policy to its probability under the old policy is calculated to obtain the importance sampling weight. For example, if the probability of a certain action under the old policy is 0.25 and the probability under the new policy is 0.30, then the importance sampling weight is 1.2. When the variance of the importance sampling weight exceeds a preset threshold (such as 0.2), the importance sampling weight is truncated: weights exceeding twice the mean are set to twice the mean, and weights below 0.5 times the mean are set to 0.5 times the mean. The generalized dominance estimation method is used to calculate the weighted sum of temporal difference errors. The specific steps are as follows: Select a continuous state-action sequence (length 10), and calculate the dominance value for each time step backwards from the end of the sequence. The dominance value equals the temporal difference error of the current time step, plus a decay factor (set to 0.97), multiplied by the dominance value of the next time step. For example, assuming the temporal difference errors of the last three time steps of a sequence are 0.82, 0.75, and 0.68 respectively, then the dominance value of the third-to-last time step is calculated as 0.82 + 0.97 × 0.75 + 0.97. 2 ×0.68=2.19. The loss function of the policy network is the negative of the product of the importance sampling weights and the advantage value, plus a regularization term of the policy entropy (coefficient of 0.01). The Adam optimizer is used to update the policy network parameters, with an initial learning rate of 0.01. The learning rate is adaptively adjusted based on historical gradients. When the gradient direction is consistent for 5 consecutive iterations, the learning rate increases by 20%; when the gradient direction changes continuously, the learning rate decreases by 15%.
[0147] A hybrid policy is constructed, comprising exploration and exploitation terms. The output probability distribution of the hybrid policy is a weighted sum of the exploration and exploitation terms. The exploration terms are uniformly distributed to ensure the policy network can explore the unknown action space. The exploitation terms directly use the output probability distribution of the policy network. The initial weight of the exploration terms is set to 0.3, decreasing exponentially with each training epoch, with a decay coefficient of 0.995 and a minimum value not lower than 0.05, ensuring the policy maintains a certain level of exploration capability. For example, at the 100th training epoch, the exploration term weight is 0.3 × 0.995^100 ≈ 0.18. The weight of the exploitation terms is equal to 1 minus the exploration term weight, and is dynamically adjusted based on the value assessment results. When the value assessment result is higher than the historical average, the exploitation term weight increases by 0.01; when the value assessment result is lower than the historical average, the exploitation term weight decreases by 0.01, but the total weight does not exceed 0.95. Suppose that by the 200th training round, the weight of the exploration term has decayed to 0.12. If the value assessment result at this point is higher than the historical average, then the weight of the exploit term increases from 0.88 to 0.89, while the weight of the exploration term decreases accordingly to 0.11. This hybrid strategy mechanism allows for sufficient exploration in the early stages of training, shifts towards more exploitation as the strategy matures, and dynamically adjusts the balance between exploration and exploitation based on actual performance.
[0148] The experience playback pool stores samples of energy storage group charging and discharging power, integrated reward signals, and grid status data. The timeliness assessment value of the data samples is calculated using the exponential decay function exp(-t / T), where t is the number of time steps the sample is stored in the experience pool, and T is the time constant (set to 1000). A new sample has a timeliness assessment value of 1, decaying to 0.61 after 500 time steps and to 0.37 after 1000 time steps. Next, the importance assessment value of the data samples is calculated, normalized using the absolute value of the time-series difference error. For example, if the absolute values of the time-series difference errors for a batch of samples are 0.82, 0.65, and 0.43, the normalized importance assessment values are 0.43, 0.34, and 0.23, respectively. Sample priority is a weighted sum of the timeliness assessment value and the importance assessment value, with a weight ratio of 3:7. If a sample has a timeliness assessment value of 0.8 and an importance assessment value of 0.6, then its priority is 0.8 × 0.3 + 0.6 × 0.7 = 0.66. Data samples in the experience replay pool are sorted according to their priority, with higher-priority samples being used for training first. The size of the experience pool is not fixed but dynamically adjusted based on the distribution range of data sample priorities. Specifically, the difference between the 90th percentile and the 10th percentile of the sample priority is calculated. When the difference is greater than a preset threshold (e.g., 0.6), the experience pool size is increased to 1.2 times the original size, with a maximum of 20,000; when the difference is less than the preset threshold (e.g., 0.2), the experience pool size is decreased to 0.8 times the original size, with a minimum of 5,000. For example, if the initial experience pool size is 10,000, and the 90th percentile of the sample priority is 0.85, the 10th percentile is 0.15, and the difference is 0.7, which is greater than the threshold of 0.6, then the experience pool size is increased to 12,000. By dynamically adjusting the capacity of the experience pool, we can retain a sufficiently diverse sample while saving memory usage.
[0149] Figure 4 The performance comparison chart of parameter updates for the value network and policy network clearly demonstrates the superiority of the dual-network structure of this invention compared to the traditional DQN network and the PPO algorithm. This invention achieves significant advantages in four key metrics: a 46.5% improvement in convergence speed; a 44.0% improvement in policy stability; a 50.0% improvement in sample utilization efficiency; and most notably, a 54.0% improvement in control accuracy, significantly outperforming the 32.0% of traditional DQN and the 36.0% of PPO. This proves the significant effects of the dual-network structure design, dynamic update mechanism, hybrid policy design, and experience pool optimization mechanism based on timeliness and importance adopted in this invention on improving training stability, convergence efficiency, and control accuracy.
[0150] This invention improves the training stability and convergence efficiency of the value network through a dual-network structure design and a dynamic update mechanism for the target network. The policy network update method, combining importance sampling and generalized advantage estimation, effectively reduces the variance of the policy gradient and accelerates the policy optimization process. The dynamically balanced hybrid policy design achieves a good balance between exploration and utilization. The experience pool optimization mechanism based on timeliness and importance significantly improves sample utilization efficiency, making the entire learning process more efficient and stable. Ultimately, it achieves autonomous optimization of the energy storage group control strategy, improving the accuracy and economy of grid regulation.
[0151] A second aspect of the present invention provides a coordinated and complementary regulation system for energy storage power stations and power grids, comprising:
[0152] The first unit is used to perform multi-scale decomposition on real-time frequency and voltage data from the power grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data in combination with the power grid topology.
[0153] The second unit is used to construct a distributed model prediction control framework based on the disturbance propagation prediction data. According to the response time, the energy storage resources are divided into multiple energy storage groups. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol.
[0154] The third unit is used to design an adaptive dynamic programming controller. The control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group. The policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results.
[0155] The fourth unit is used to recalculate the state assessment based on the new network structure when a change in the power grid topology is detected, update the optimal charging and discharging power output of the strategy network, control each energy storage group to perform charging and discharging operations according to the optimized optimal charging and discharging power, and feed back the operating data to the adaptive dynamic programming controller.
[0156] A third aspect of the present invention provides an electronic device, comprising:
[0157] processor;
[0158] Memory used to store processor-executable instructions;
[0159] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0160] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0161] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for coordinated and complementary regulation between energy storage power stations and power grids, characterized in that, include: Multi-scale decomposition is performed on real-time frequency and voltage data from the power grid to obtain disturbance components in different frequency bands. Combined with the power grid topology, disturbance propagation prediction data is calculated. Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to the response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol. An adaptive dynamic programming controller is designed, and the control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group, and the policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results. When a change in grid topology is detected, the state assessment is recalculated based on the new network structure, and the optimal charging and discharging power output of the strategy network is updated. Based on the optimized optimal charging and discharging power, each energy storage group is controlled to perform charging and discharging operations.
2. The method according to claim 1, characterized in that, The steps for performing multi-scale decomposition on real-time frequency and voltage data from the power grid side to obtain disturbance components in different frequency bands, and then calculating disturbance propagation prediction data based on the power grid topology, include: The real-time frequency data and voltage data are decomposed into perturbation components in different frequency bands through multi-scale decomposition. A dynamic topology identification matrix is established, which includes a set of nodes, a set of edges, and a weight matrix. The value of each element in the weight matrix is calculated based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance. The weight matrix reflects the real-time electrical connection strength between power grid nodes. Based on the disturbance components of the different frequency bands and the dynamic topology identification matrix, the prediction parameter vector is calculated by recursive least squares method; The prediction error is calculated within the sliding time window. When the prediction error is greater than a first preset threshold or the change in the dynamic topology recognition matrix is greater than a second preset threshold, the update calculation of the prediction parameter vector is triggered. Based on the updated prediction parameter vector, the dynamic topology identification matrix, and the real-time disturbance data of the disturbance source node, the disturbance propagation prediction data is calculated.
3. The method according to claim 1, characterized in that, Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charging and discharging efficiency, and response speed of the energy storage. The steps to achieve global optimization through a consensus protocol include: Based on the disturbance propagation prediction data, a distributed model prediction control framework is constructed, and energy storage resources are divided into inertial response, transient stability and power balance energy storage groups according to the response time. The local optimization controller of the inertial response energy storage group constructs a first objective function based on frequency deviation, frequency change rate, and energy storage output power; the local optimization controller of the transient stable energy storage group constructs a second objective function based on voltage deviation, tie-line power deviation, and energy storage output power; and the local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second, and third objective functions are all constrained by the remaining capacity, charge / discharge efficiency, and response speed of the corresponding energy storage group. A basic communication topology is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology is optimized based on the weighted adjacency matrix. A distributed consensus protocol considering delay compensation is adopted to achieve global optimization control of the energy storage groups.
4. The method according to claim 3, characterized in that, The steps of collecting communication state parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consensus protocol that considers delay compensation include: The communication quality factor, latency factor, and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method. The network connectivity is calculated based on the weighted adjacency matrix. When the network connectivity is lower than a preset connectivity threshold, the communication topology is reconstructed using the minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topology with optimal latency performance. Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative computation, and the adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iteration convergence speed. When the state deviation of an energy storage node exceeds a preset deviation threshold, the state information is broadcast, and the adjacent nodes that receive the broadcast information update their respective optimization variables. The communication connectivity of energy storage nodes is periodically calculated. When an abnormal node communication is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
5. The method according to claim 3, characterized in that, The steps of designing an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal, constructing a value network and a policy network, wherein the value network evaluates the grid state and the remaining lifetime cost of the energy storage unit, and the policy network outputs the optimal charging and discharging power of the energy storage unit based on the value evaluation results, include: The frequency deviation, frequency change rate and energy storage output power of the inertial response energy storage group, the voltage deviation, tie-line power deviation and energy storage output power of the transient stable energy storage group, and the load power deviation and operating cost of the power balance energy storage group in the distributed model predictive control framework are respectively constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function. A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs the state value evaluation result based on the operating status of the energy storage group, the remaining lifetime cost, and the grid status; the strategy network updates the network parameters and outputs the charging and discharging power of the energy storage group based on the state value evaluation result using the strategy gradient method. The charging and discharging power of the energy storage group, the comprehensive reward signal, and the grid status are used to construct an experience playback pool. A value network and a policy network are trained based on the data samples in the experience playback pool. The value network uses time-series differential error for parameter updates, and the learning rate of the policy network is adaptively adjusted according to historical gradients.
6. The method according to claim 5, characterized in that, The adaptive adjustment steps for the weights of each evaluation function in the comprehensive reward signal include: A control performance index is constructed, which is calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within the evaluation time window; the sensitivity of the weights of the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function to the control performance index is calculated. The root mean square of frequency deviation, root mean square of voltage deviation, and root mean square of power deviation are collected to construct a state evaluation vector. The product of the state evaluation vector and the sensitivity is exponentially decayed and used as the weight update amount. The energy storage life loss assessment value is calculated based on the energy storage group's discharge depth and charge / discharge power, and the energy storage life loss assessment value is mapped to weight adjustment constraints. Construct a multi-objective optimization function that includes the control performance index, the energy storage lifetime loss assessment value, and the weight change amount; calculate the weight optimization direction based on the gradient of the multi-objective optimization function; and determine the weight update value according to the preset weight adjustment step size. The weight update values are smoothed using an exponential smoothing method.
7. The method according to claim 5, characterized in that, The steps for updating the parameters of the value network and the policy network include: A dual-network structure of a target network and an evaluation network is constructed for the value network. The temporal difference error is calculated based on the current state value, the next state value, and the immediate reward. The evaluation network constructs a loss function based on the temporal difference error and continuously updates the network parameters. The target network determines the parameter update period based on the prediction error change rate of the evaluation network. The action probability distribution output by the policy network is calculated. The importance sampling weights are calculated based on the probability ratio of the new and old policies. The cutoff range is determined according to the variance of the importance sampling weights. The evaluation result of the value network is used as the baseline function. The weighted sum of the temporal difference error is calculated using the generalized advantage estimation method. The policy network parameters are updated based on the weighted sum. A hybrid strategy is constructed that includes exploration and exploitation items. The weight of the exploration items gradually decreases as the training process progresses, while the weight of the exploitation items is dynamically adjusted based on the value evaluation results. The timeliness assessment value and importance assessment value of the data samples are calculated. The timeliness assessment value is obtained by an exponential function of the sample storage time, and the importance assessment value is obtained by the absolute value of the time series difference error. Based on the timeliness assessment value and the importance assessment value, the sample priority is constructed. The data samples in the experience pool are sorted according to the sample priority. The capacity of the experience pool is dynamically adjusted according to the distribution range of the data sample priority.
8. A coordinated and complementary regulation system for energy storage power stations and power grids, used to implement the method described in any one of claims 1-7, characterized in that, include: The first unit is used to perform multi-scale decomposition on real-time frequency and voltage data from the power grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data in combination with the power grid topology. The second unit is used to construct a distributed model prediction control framework based on the disturbance propagation prediction data. According to the response time, the energy storage resources are divided into multiple energy storage groups. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity of the energy storage, the charging and discharging efficiency, and the response speed, and achieves global optimization through a consensus protocol. The third unit is used to design an adaptive dynamic programming controller. The control effect of the distributed model predictive control framework is used as a reward signal to construct a value network and a policy network. The value network evaluates the system state and the remaining lifetime cost of the energy storage group. The policy network outputs the optimal charging and discharging power of the energy storage group based on the value evaluation results. The fourth unit is used to recalculate the state assessment based on the new network structure when a change in the power grid topology is detected, update the optimal charging and discharging power output of the strategy network, control each energy storage group to perform charging and discharging operations according to the optimized optimal charging and discharging power, and feed back the operating data to the adaptive dynamic programming controller.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Power grid side energy storage optimal configuration and operation control method and system
CN119171486A
Multi-scale energy storage system control method and system
CN119419889A