Energy storage power station and power grid collaborative complementary regulation method and system
By performing multi-scale decomposition of grid data and distributed model predictive control, combined with an adaptive dynamic programming controller, the inefficiency of energy storage systems in grid disturbance identification and coordinated regulation is solved, the stability and economy of the grid are improved, and the service life of energy storage equipment is extended.
Patent Information
- Application Number
- CN202511100684.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-07
AI Technical Summary
Existing energy storage control methods are unable to effectively identify and specifically adjust grid disturbances in different frequency bands, resulting in inefficient utilization of energy storage resources. They also lack a distributed coordination mechanism, making it difficult to maintain efficient coordination when the grid topology changes in a complex manner, and failing to achieve a balance between system stability control and the economic efficiency of energy storage equipment.
By performing multi-scale decomposition on the real-time frequency and voltage data on the grid side and calculating the disturbance propagation prediction data in combination with the grid topology, a distributed model predictive control framework is constructed. A local optimization controller is configured and global optimization is achieved through a consistency protocol. An adaptive dynamic programming controller is designed to evaluate the remaining life cost of the energy storage group and adjust the charging and discharging strategy.
It achieves accurate perception and early response to grid fluctuations, improves grid stability and anti-disturbance capabilities, optimizes energy storage resource allocation, improves the overall economy and operating efficiency of the system, and extends the service life of energy storage equipment.
Smart Images

Figure CN120638422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to smart grid technology, and in particular to a method and system for coordinated and complementary regulation of energy storage power stations and power grids. Background Art
[0002] As the proportion of renewable energy generation continues to increase, the safe and stable operation of the power grid faces new challenges. The intermittent and fluctuating nature of renewable energy exacerbates fluctuations in grid frequency and voltage, impacting system stability. Energy storage systems, as a key power regulation resource, can smooth these fluctuations through rapid charge and discharge responses, improving grid stability and flexibility. Energy storage power stations, with their fast response speed and high regulation accuracy, are becoming a crucial support for the safe and stable operation of the power grid.
[0003] Currently, the coordinated regulation of power grids and energy storage systems faces multiple challenges. Existing energy storage control methods usually adopt a single control strategy, which is unable to effectively identify and specifically regulate grid disturbances in different frequency bands, resulting in inefficient utilization of energy storage resources and difficulty in fully leveraging the technical characteristics of different types of energy storage. Existing energy storage control systems mostly adopt a centralized control architecture and lack a distributed coordination mechanism. When faced with complex changes in the grid topology, the control system has poor adaptability and scalability, making it difficult to achieve efficient coordination between multiple energy storage units. Most energy storage control methods lack comprehensive consideration of the life cycle cost of energy storage and fail to achieve a balance between system stability control and the economic efficiency of energy storage equipment, resulting in increased operating costs of energy storage systems and poor overall economic efficiency.
[0004] Against the backdrop of rapid power system development and energy transformation, there is an urgent need for an advanced control method that can achieve coordinated and complementary regulation between energy storage power stations and power grids to improve the stability, flexibility, and economy of the power system. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for coordinated and complementary regulation of an energy storage power station and a power grid, which can solve the problems in the prior art.
[0006] A first aspect of an embodiment of the present invention provides a method for coordinated and complementary regulation of an energy storage power station and a power grid, comprising: Perform multi-scale decomposition on the real-time frequency and voltage data on the grid side to obtain disturbance components in different frequency bands. Combined with the grid topology calculation, disturbance propagation prediction data is obtained. Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieves global optimization through a consistency protocol. Design an adaptive dynamic programming controller, use the control effect of the distributed model predictive control framework as a reward signal, and construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; When a change in the grid topology is detected, the state evaluation is recalculated based on the new network structure, and the optimal charge and discharge power output by the strategy network is updated; and each energy storage group is controlled to perform charge and discharge operations according to the optimized optimal charge and discharge power.
[0007] In an optional embodiment, The steps of performing multi-scale decomposition on the real-time frequency and voltage data on the grid side to obtain disturbance components in different frequency bands and then calculating the disturbance propagation prediction data based on the grid topology include: Performing multi-scale decomposition on the real-time frequency data and voltage data to obtain disturbance components in different frequency bands; Establishing a dynamic topology identification matrix, the dynamic topology identification matrix including a node set, an edge set, and a weight matrix, calculating the value of each element in the weight matrix based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance, the weight matrix reflecting the real-time electrical connection strength between the grid nodes; Calculating a prediction parameter vector by a recursive least squares method according to the disturbance components of the different frequency bands and the dynamic topology identification matrix; Calculating a prediction error within a sliding time window, and triggering an update calculation of the prediction parameter vector when the prediction error is greater than a first preset threshold or a change in the dynamic topology identification matrix is greater than a second preset threshold; According to the updated prediction parameter vector, the dynamic topology identification matrix and the real-time disturbance data of the disturbance source node, the disturbance propagation prediction data is calculated, and the disturbance propagation prediction data includes the predicted disturbance amplitude, predicted time and propagation path of each target node.
[0008] In an optional embodiment, Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage. The steps of achieving global optimization through a consistency protocol include: Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed to divide energy storage resources into inertia response, transient stability, and power balance energy storage groups according to response time; The local optimization controller of the inertia response energy storage group constructs a first objective function based on frequency deviation, frequency change rate and energy storage output power. The local optimization controller of the transient stability energy storage group constructs a second objective function based on voltage deviation, tie line power deviation and energy storage output power. The local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second and third objective functions are all constrained by the remaining capacity, charge and discharge efficiency and response speed of the corresponding energy storage group. A basic communication topology structure is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology structure is optimized based on the weighted adjacency matrix. A distributed consistency protocol that considers delay compensation is used to achieve global optimization control of the energy storage group.
[0009] In an optional embodiment, The steps of collecting communication status parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consistency protocol that considers delay compensation include: The communication quality factor, delay factor and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method. Calculating network connectivity based on the weighted adjacency matrix; and when the network connectivity is lower than a preset connectivity threshold, reconstructing the communication topology using a minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topological structure with optimal latency performance; Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative calculation. The adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iterative convergence speed. When the state deviation of the energy storage node exceeds the preset deviation threshold, the broadcast of the state information is triggered, and the adjacent nodes that receive the broadcast information update their respective optimization variables; The communication connectivity of energy storage nodes is calculated periodically. When a node communication anomaly is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
[0010] In an optional embodiment, The steps of designing an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal, and constructing a value network and a strategy network, wherein the value network evaluates the grid status and the remaining life cost of the energy storage group, and the strategy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results, include: The frequency deviation, frequency change rate and energy storage output power of the inertia response energy storage group in the distributed model predictive control framework, the voltage deviation, tie line power deviation and energy storage output power of the transient stability energy storage group, and the load power deviation and operating cost of the power balance energy storage group are constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function respectively; A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs a state value evaluation result based on the operating state, remaining life cost, and grid state of the energy storage group; the strategy network uses a policy gradient method to update network parameters based on the state value evaluation result and output the charge and discharge power of the energy storage group; An experience replay pool is constructed using the charge and discharge power of the energy storage group, the comprehensive reward signal, and the grid status. The value network and the policy network are trained based on the data samples in the experience replay pool. The value network uses temporal difference error to update parameters, and the learning rate of the policy network is adaptively adjusted according to the historical gradient.
[0011] In an optional embodiment, The step of adaptively adjusting the weights of each evaluation function in the comprehensive reward signal includes: Constructing a control performance index, the control performance index being calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within an evaluation time window; calculating the sensitivity of a first control effect evaluation function weight, a second control effect evaluation function weight, and a third control effect evaluation function weight to the control performance index; Collecting the root mean square of frequency deviation, the root mean square of voltage deviation, and the root mean square of power deviation to construct a state evaluation vector, and using the product of the state evaluation vector and the sensitivity as a weight update amount after exponential decay; Calculating an energy storage life loss assessment value based on the discharge depth and charge and discharge power of the energy storage group, and mapping the energy storage life loss assessment value into a weight adjustment constraint condition; Constructing a multi-objective optimization function including the control performance index, the energy storage life loss assessment value, and the weight change amount, calculating the weight optimization direction based on the gradient of the multi-objective optimization function, and determining the weight update value according to the preset weight adjustment step size; An exponential smoothing method is used to smooth the weight update value.
[0012] In an optional embodiment, The step of updating the parameters of the value network and the policy network includes: Constructing a dual-network structure of a target network and an evaluation network for the value network, calculating the temporal difference error based on the current state value, the next state value, and the immediate reward, the evaluation network constructing a loss function based on the temporal difference error and continuously updating the network parameters, and the target network determining the parameter update period based on the prediction error change rate of the evaluation network; Calculate the action probability distribution output by the policy network, calculate the importance sampling weight based on the probability ratio of the new and old policies, determine the truncation range based on the variance of the importance sampling weight, use the evaluation result of the value network as the baseline function, calculate the weighted sum of the temporal difference error using the generalized advantage estimation method, and update the policy network parameters based on the weighted sum; Construct a hybrid strategy that includes an exploration term and an exploitation term, where the weight of the exploration term gradually decays as the training progresses, and the weight of the exploitation term is dynamically adjusted based on the value assessment results; Calculate the timeliness evaluation value and importance evaluation value of the data sample, the timeliness evaluation value is obtained by the exponential function of the sample storage time, and the importance evaluation value is obtained by the absolute value of the time series difference error. Construct the sample priority based on the timeliness evaluation value and the importance evaluation value, sort the data samples in the experience pool according to the sample priority, and dynamically adjust the capacity of the experience pool according to the distribution range of the data sample priority.
[0013] A second aspect of an embodiment of the present invention provides a coordinated and complementary regulation system for an energy storage power station and a power grid, including: The first unit is used to perform multi-scale decomposition on the real-time frequency data and voltage data on the grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data in combination with the grid topology structure; The second unit is used to build a distributed model predictive control framework based on the disturbance propagation prediction data, divide the energy storage resources into multiple energy storage groups according to the response time, configure a local optimization controller for each energy storage group, and construct an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieve global optimization through a consistency protocol; The third unit is used to design an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal to construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; The fourth unit is used to recalculate the state evaluation based on the new network structure when a change in the grid topology is detected, and to update the optimal charge and discharge power output by the strategy network; control each energy storage group to perform charge and discharge operations based on the optimized optimal charge and discharge power, and feed back the operating data to the adaptive dynamic programming controller.
[0014] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0015] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0016] The present invention achieves accurate perception and early response to grid fluctuations by decomposing grid disturbance data at multiple scales and combining it with topological structure for prediction, significantly improving grid stability and anti-disturbance capability.
[0017] The present invention adopts a distributed model predictive control framework and consistency protocol to perform group optimization according to the remaining capacity, charging and discharging efficiency and response speed of energy storage, thereby achieving reasonable scheduling and global optimal configuration of energy storage resources, and improving the overall economy and operating efficiency of the system.
[0018] The adaptive dynamic programming controller designed in this invention can evaluate the remaining life cost of the energy storage group and adjust the charging and discharging strategy in real time. At the same time, it can quickly recalculate the optimal control strategy when the grid topology changes, effectively extending the service life of the energy storage equipment and enhancing the system's adaptability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the process of the coordinated complementary regulation method between the energy storage power station and the power grid according to an embodiment of the present invention; Figure 2 This is a comparison chart of comprehensive control performance under frequency disturbance conditions; Figure 3 This is a flow chart of the adaptive adjustment of the comprehensive reward signal weight of the present invention; Figure 4 The performance comparison chart of parameter update of value network and policy network is shown. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0021] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0022] Figure 1 FIG. 1 is a flow chart of a method for cooperatively and complementaryly regulating an energy storage power station and a power grid according to an embodiment of the present invention. Figure 1 As shown, the method includes: Perform multi-scale decomposition on the real-time frequency and voltage data on the grid side to obtain disturbance components in different frequency bands. Combined with the grid topology calculation, disturbance propagation prediction data is obtained. Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieves global optimization through a consistency protocol. Design an adaptive dynamic programming controller, use the control effect of the distributed model predictive control framework as a reward signal, and construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; When a change in the grid topology is detected, the state evaluation is recalculated based on the new network structure, and the optimal charge and discharge power output by the strategy network is updated; based on the optimized optimal charge and discharge power, each energy storage group is controlled to perform charge and discharge operations, and the operating data is fed back to the adaptive dynamic programming controller.
[0023] In an optional embodiment, the steps of performing multi-scale decomposition on the real-time frequency data and voltage data on the grid side to obtain disturbance components in different frequency bands and calculating disturbance propagation prediction data based on the grid topology include: Performing multi-scale decomposition on the real-time frequency data and voltage data to obtain disturbance components in different frequency bands; Establishing a dynamic topology identification matrix, the dynamic topology identification matrix including a node set, an edge set, and a weight matrix, calculating the value of each element in the weight matrix based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance, the weight matrix reflecting the real-time electrical connection strength between the grid nodes; Calculating a prediction parameter vector using a recursive least squares method according to the disturbance components of the different frequency bands and the dynamic topology identification matrix, wherein the prediction parameter vector is used for weight coefficients corresponding to the disturbance components of the different frequency bands; Calculating a prediction error within a sliding time window, and triggering an update calculation of the prediction parameter vector when the prediction error is greater than a first preset threshold or a change in the dynamic topology identification matrix is greater than a second preset threshold; According to the updated prediction parameter vector, the dynamic topology identification matrix and the real-time disturbance data of the disturbance source node, the disturbance propagation prediction data is calculated, and the disturbance propagation prediction data includes the predicted disturbance amplitude, predicted time and propagation path of each target node.
[0024] For example, real-time frequency and voltage data from each monitoring point in the power grid can be collected using a wide area measurement system (WAMS), typically with a sampling frequency of 100 Hz and a collection time of multiple consecutive time windows, each of which is 10 seconds long.
[0025] Empirical mode decomposition (EMD) can be used to perform multi-scale decomposition on the collected real-time frequency and voltage data. Specifically, EMD decomposition is performed on the frequency signal f(t) at a monitoring point to obtain n intrinsic mode function (IMF) components c1(t), c2(t), ..., cn(t) and a residual term rn(t). These IMF components represent the disturbance characteristics of different frequency bands. For example, for the frequency data of a 110kV substation, EMD decomposition yields five IMF components, corresponding to the disturbance characteristics of the 0.5-2.5Hz, 0.2-0.5Hz, 0.05-0.2Hz, 0.01-0.05Hz, and 0-0.01Hz frequency bands. Similarly, voltage data is processed similarly to obtain the disturbance components of the corresponding frequency bands.
[0026] The dynamic topology identification matrix consists of a node set V, an edge set E, and a weight matrix W. The node set V represents all busbar nodes in the power grid. For example, if a regional power grid contains 15 500 kV nodes and 30 220 kV nodes, |V| = 45. The edge set E represents the connections between nodes, such as lines and transformers. The weight matrix W reflects the strength of the electrical connections between nodes. The calculation of each element wij considers the node voltage amplitudes Vi and Vj, the voltage phase angles θi and θj, and the internode reactance Xij. Specifically, wij can be expressed as the product of Vi and Vj, divided by Xij, and taking into account the sinθij factor. For example, for two nodes with voltages of 525 kV and 515 kV, a phase angle difference of 5 degrees, and a connecting reactance of 0.1 ohm, the weight value is approximately 2625. A larger weight value indicates a stronger electrical connection between nodes and easier disturbance propagation. For nodes that are not directly connected, the weight value is set to 0.
[0027] The prediction parameter vector α contains multiple elements, each of which corresponds to the weight coefficient of the disturbance component in different frequency bands. During the calculation, the disturbance data within the historical observation window is constructed into a matrix form, and then combined with the dynamic topology identification matrix, the value of α is estimated by recursive least squares method. For example, for the disturbance in the 0.5-2.5Hz frequency band of a certain 500kV substation, α1=0.85, α2=0.72, α3=0.63, etc. are obtained, which indicates that the disturbance in this frequency band decays faster during propagation. For the disturbance in the 0-0.01Hz frequency band, α1=0.98, α2=0.95, α3=0.91, etc. are obtained, indicating that the propagation and decay of low-frequency disturbances are slower.
[0028] In practical applications, the prediction error needs to be continuously calculated within a sliding time window. The time window length is typically set to 5-10 seconds, with a sliding step size of 1 second. The prediction error, ε, is defined as the root mean square error between the predicted value and the actual observed value. When ε is greater than a first preset threshold (e.g., 0.05), or the change in the dynamic topology identification matrix is greater than a second preset threshold (e.g., the change in the weight matrix elements exceeds 15%), an update calculation of the prediction parameter vector is triggered. This mechanism can adapt to changes in the grid's operating state and improve prediction accuracy.
[0029] For a disturbance ds(t) in a certain frequency band at a source node s, the disturbance di(t+Δt) propagating to a target node i is predicted, taking into account the prediction parameter vector α and the inter-node propagation delay Δt. The specific calculation method is to determine the real-time disturbance data ds(t) at the source node s. Then, based on the weight coefficient αsi corresponding to the frequency band in the prediction parameter vector α, the attenuation coefficient of the disturbance during propagation is calculated. For directly connected nodes, the predicted disturbance amplitude di(t+Δt) = αsi × ds(t) × wsi, where wsi is the corresponding weight value in the dynamic topology identification matrix. For propagation through multiple nodes, a matrix multiplication method is used to calculate the cumulative attenuation effect. The propagation delay Δt is calculated based on the propagation speed of electromagnetic waves in the power system (approximately 60% of the speed of light). For example, for two nodes 200 kilometers apart, the disturbance propagation delay is approximately 6.67 milliseconds (200 km ÷ 0.6 × 300,000 km / s). Combining the weight information in the dynamic topology identification matrix, the Dijkstra shortest path algorithm determines the optimal propagation path. Specifically, the path with the smallest ∑(1 / wij) is selected from all available paths. For complex network structures, suboptimal paths are also considered as alternative propagation paths. The final output of disturbance propagation prediction data includes the predicted disturbance amplitude at each target node (accurate to 0.01 Hz or 0.01 kV), the predicted arrival time (accurate to the millisecond level), and a detailed propagation path description (including the sequence of all nodes passed and the corresponding propagation time). This prediction data will serve as an important basis for subsequent energy storage control decisions.
[0030] This invention can accurately predict the propagation of power grid disturbances, which is of great significance for improving the safe and stable operation of power grids. Through multi-scale decomposition and dynamic topology recognition, it can adapt to the propagation characteristics of disturbances in different frequency bands and has good adaptability to changes in power grid topology.
[0031] In an optional embodiment, a distributed model predictive control framework is constructed based on the disturbance propagation prediction data. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage. The steps of achieving global optimization through a consistency protocol include: Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed to divide energy storage resources into inertia response, transient stability, and power balance energy storage groups according to response time; The local optimization controller of the inertia response energy storage group constructs a first objective function based on frequency deviation, frequency change rate and energy storage output power. The local optimization controller of the transient stability energy storage group constructs a second objective function based on voltage deviation, tie line power deviation and energy storage output power. The local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second and third objective functions are all constrained by the remaining capacity, charge and discharge efficiency and response speed of the corresponding energy storage group. A basic communication topology structure is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology structure is optimized based on the weighted adjacency matrix. A distributed consistency protocol that considers delay compensation is used to achieve global optimization control of the energy storage group.
[0032] For example, energy storage resources are classified according to their response time characteristics into three types: inertia response energy storage, transient stability energy storage, and power balancing energy storage. Inertia response energy storage typically has a response time in milliseconds, such as supercapacitors and flywheel energy storage devices; transient stability energy storage has a response time in seconds, such as lithium-ion batteries and lead-acid batteries; and power balancing energy storage has a response time in minutes, such as pumped hydro and compressed air energy storage.
[0033] For the inertia response energy storage group, the local optimization controller constructs the first objective function. This objective function takes frequency deviation, frequency change rate and energy storage output power as the main considerations. Frequency deviation refers to the difference between the current system frequency and the rated frequency, usually in Hertz; the frequency change rate represents the change in frequency per unit time, in Hertz / second; the energy storage output power is the active power provided by the energy storage device to the grid, in megawatts. In specific implementation, when it is detected that the system frequency is lower than the rated value of 49.8Hz and the frequency change rate is -0.2Hz / s, the inertia response energy storage group will quickly increase the output power to 80% of the rated power, and the response will be completed within 200 milliseconds to provide system inertia support.
[0034] For the transient stability energy storage group, the local optimization controller constructs a second objective function. This objective function is primarily based on voltage deviation, tie-line power deviation, and energy storage output power. Voltage deviation refers to the difference between the node voltage and the rated voltage, expressed in per-unit values; tie-line power deviation represents the difference between the actual transmission power and the planned transmission power, expressed in megawatts; and energy storage output power is also expressed in megawatts. In practice, when the system voltage is detected to have dropped to 0.92 per-unit values and the tie-line power deviation reaches 50MW, the transient stability energy storage group will gradually increase its output power to 65% of the rated power within 2 seconds to stabilize the system voltage and power transmission.
[0035] For the power-balancing energy storage group, the local optimization controller constructs a third objective function. This objective function primarily considers load power deviation and energy storage operating costs. Load power deviation refers to the difference between the actual load and the predicted load, measured in megawatts; energy storage operating costs include factors such as equipment depreciation and charge and discharge losses, and are typically calculated in units of yuan per kilowatt-hour. In actual application scenarios, when fluctuations in renewable energy output cause the load power deviation to reach 100MW, the power-balancing energy storage group will gradually adjust its output power within 5 minutes to minimize system power imbalance while optimizing operating costs.
[0036] All objective functions are subject to the physical constraints of their respective energy storage groups. Remaining capacity constraints ensure that the state of charge (SOC) of the energy storage device remains within a safe range, typically between 20% and 80%. Charge and discharge efficiency constraints account for losses during energy conversion, such as the approximately 95% charge and discharge efficiency of lithium-ion batteries. Response speed constraints consider the physical limitations of the device, such as the power ramp rate limit. Supercapacitors can reach 100% of their rated power per second, while pumped hydro storage can only reach 10% of its rated power per minute.
[0037] To achieve coordinated control between energy storage groups, a basic communication topology must be established. Communication status parameters between energy storage nodes are collected, including communication delay, data packet loss rate, and bandwidth utilization. Based on these parameters, a weighted adjacency matrix is constructed. The matrix element aij represents the communication quality weight between node i and node j, with weight values ranging from 0 to 1, where 0 indicates no connection and 1 indicates an ideal connection. For example, when the communication delay between two nodes is 15ms and the packet loss rate is 2%, the weight can be set to 0.9; when the communication delay reaches 50ms and the packet loss rate is 10%, the weight can be reduced to 0.6.
[0038] Optimize the communication topology based on the weighted adjacency matrix. This optimization process involves removing low-quality connections (e.g., connections with a weight below 0.4), enhancing critical path communication capabilities (e.g., increasing the bandwidth connecting central nodes), and establishing redundant communication paths to improve system reliability. The optimized communication topology should meet network connectivity requirements, ensuring that at least one communication path exists between any two nodes.
[0039] A distributed consensus protocol with delay compensation is used to achieve global optimal control of energy storage groups. This protocol iteratively converges the control variables of each energy storage group. The delay compensation mechanism reduces the impact of communication delays by predicting future state values.
[0040] By dividing energy storage resources into different groups based on response time, this method accurately matches the varying timescales of grid disturbances. Dedicated local optimization controllers are configured for each group, enabling each to address specific grid issues such as frequency deviation, voltage deviation, and power balance. Coordinated control between energy storage groups is achieved through a distributed consensus protocol, avoiding the communication bottlenecks and single-point failure risks of centralized control and improving the reliability and robustness of system control.
[0041] In an optional embodiment, the steps of collecting communication status parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consensus protocol that considers delay compensation include: The communication quality factor, delay factor, and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method. The weighted adjacency matrix is used to characterize the communication connection strength between energy storage nodes. Calculating network connectivity based on the weighted adjacency matrix; and when the network connectivity is lower than a preset connectivity threshold, reconstructing the communication topology using a minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topological structure with optimal latency performance; Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative calculation. The adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iterative convergence speed. The delay compensation term is used to reduce the impact of communication delay on global optimization performance. When the state deviation of the energy storage node exceeds the preset deviation threshold, the broadcast of the state information is triggered, and the adjacent nodes that receive the broadcast information update their respective optimization variables; The communication connectivity of energy storage nodes is calculated periodically. When a node communication anomaly is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
[0042] For example, during the phase of collecting communication status parameters and constructing a weighted adjacency matrix between energy storage nodes, the communication quality factor, delay factor, and bandwidth utilization between each node are monitored and collected in real time. The communication quality factor can be obtained by measuring the signal-to-noise ratio. For example, when the signal-to-noise ratio is above 35dB, the communication quality factor can be set to 0.9. The delay factor is obtained by recording the time difference between the transmission and reception of a data packet. For example, when the delay is below 15ms, the delay factor can be set to 0.8. The bandwidth utilization is calculated by calculating the ratio of the actual data transmission volume to the theoretical bandwidth of the communication link. When the utilization rate is below 50%, it can be set to 0.7. A fuzzy comprehensive evaluation method is used to process these three parameters. First, a membership function is established, mapping each parameter to the interval [0, 1]. For example, the communication quality factor can be set to three fuzzy subsets: "excellent," "good," and "poor." The corresponding membership functions are maximized when the signal-to-noise ratio is above 30dB, between 20-30dB, and below 20dB, respectively. A similar method is used to construct fuzzy subsets for the delay factor and bandwidth utilization. Next, a weight vector is determined. Typically, the weights for the communication quality factor, latency factor, and bandwidth utilization are set to 0.4, 0.35, and 0.25, respectively. Finally, a fuzzy synthesis calculation is performed to obtain a comprehensive evaluation value, which serves as an element of the weighted adjacency matrix. For example, the weighted value between nodes i and j is 0.76, indicating a strong communication connection.
[0043] During the communication topology optimization phase, network connectivity is calculated based on the constructed weighted adjacency matrix. Network connectivity can be obtained by calculating the eigenvalues of the weighted adjacency matrix, specifically the magnitude of the second smallest eigenvalue. When the calculated network connectivity falls below a preset threshold (e.g., 0.15), the communication topology reconstruction process is triggered. This reconstruction process utilizes a minimum spanning tree algorithm that incorporates the spatiotemporal correlation characteristics of energy storage nodes. This algorithm first calculates the spatiotemporal distance between each pair of nodes. Spatiotemporal distance is a weighted combination of physical distance and communication latency. For example, for two nodes 50 meters apart with a communication latency of 25ms, if the physical distance weight is set to 0.4 and the latency weight is set to 0.6, the calculated spatiotemporal distance is 0.4 × 50 + 0.6 × 25 = 35. Using spatiotemporal distance as edge weights, an improved Kruskal algorithm is applied to construct a minimum spanning tree, starting with the edge with the minimum spatiotemporal distance and gradually adding edges until all nodes are connected and no loops are formed. Furthermore, to ensure communication stability for important nodes, critical backup links are added, typically using the path with the second highest connectivity as the backup. After the reconstruction is completed, a topology with optimal delay performance is obtained.
[0044] During the distributed consensus protocol implementation phase, an iterative calculation is performed using a consensus protocol with delay compensation based on the reconstructed communication topology. The core of this protocol is to ensure consistency in the key states of each energy storage unit (such as the power allocation ratio) through inter-node information exchange. In practice, the state update formula for each energy storage node i at time t includes an adaptive step size term and a delay compensation term. The adaptive step size is dynamically adjusted based on the iterative convergence rate, initially set to 0.05. The state change rate is calculated at each iteration. When the change rate is less than 80% of the previous iteration, the step size is increased by 10%; when the change rate is greater than 120%, the step size is decreased by 10%, ensuring both fast and stable convergence. The delay compensation term is used to mitigate the impact of communication delay on optimization performance and is implemented through a predictive compensation method. For example, if the communication delay from node j to node i is detected to be 30ms, node i predicts its current actual state based on node j's historical state change trends, thereby reducing the error caused by delay. Experimental measurements show that applying delay compensation can reduce convergence time by approximately 25% and improve global optimization accuracy by approximately 18%.
[0045] The energy storage node state deviation processing mechanism continuously monitors the deviation of each node's state from the global target. When a node's state deviation (such as power allocation error) exceeds a preset threshold (typically set at 5%), a state information broadcast mechanism is triggered. For example, when the charging power of a certain energy storage unit deviates by 8% from the target value, the node immediately broadcasts its latest state information to the network. Neighboring nodes that receive this broadcast information then update their respective optimization variables, such as charge and discharge power setpoints and voltage regulation parameters. This event-triggered communication mechanism significantly reduces network communication load, and has been measured to reduce data transmission volume by approximately 40%.
[0046] In the communication anomaly handling mechanism, the energy storage node's connectivity—the number of valid connections it has with other nodes—is periodically calculated (e.g., every 500ms). When a node communication anomaly is detected (e.g., failure to receive an expected data packet three times in a row), the energy storage node with the highest connectivity is automatically selected as the backup communication path. For example, if communication between nodes A and B is interrupted, node C, which has a good connection to both A and B, is selected as a relay point to establish the ACB communication path and maintain the continuity of the global optimization process.
[0047] This invention adaptively evaluates and optimizes communication quality between energy storage nodes, ensuring efficient operation of distributed control systems in practical communication environments. It employs a distributed consensus protocol with delay compensation and an adaptive step-size adjustment mechanism to effectively mitigate the adverse effects of communication delays on control performance. Through a broadcast mechanism triggered by state deviations and a backup communication path selection strategy, the system significantly improves its fault tolerance and continuous operation capabilities in the event of communication anomalies, ensuring the stability and reliability of grid control.
[0048] In an optional embodiment, an adaptive dynamic programming controller is designed, the control effect of the distributed model predictive control framework is used as a reward signal, and a value network and a strategy network are constructed. The value network evaluates the grid status and the remaining life cost of the energy storage group, and the strategy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results. The steps include: The frequency deviation, frequency change rate and energy storage output power of the inertia response energy storage group in the distributed model predictive control framework, the voltage deviation, tie line power deviation and energy storage output power of the transient stability energy storage group, and the load power deviation and operating cost of the power balance energy storage group are constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function respectively; A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs a state value evaluation result based on the operating state, remaining life cost, and grid state of the energy storage group; the strategy network uses a policy gradient method to update network parameters based on the state value evaluation result and output the charge and discharge power of the energy storage group; An experience replay pool is constructed using the charge and discharge power of the energy storage group, the comprehensive reward signal, and the grid status. The value network and the policy network are trained based on the data samples in the experience replay pool. The value network uses temporal difference error to update parameters, and the learning rate of the policy network is adaptively adjusted according to the historical gradient.
[0049] According to the design requirements, the present invention provides a specific implementation of an adaptive dynamic programming controller, which uses the control effect of a distributed model predictive control framework as a reward signal to construct a value network and a policy network.
[0050] For example, to construct a comprehensive reward signal, it is necessary to evaluate the control effects of different types of energy storage groups. For the inertia response energy storage group, the frequency deviation, frequency change rate, and energy storage output power are constructed as the first control effect evaluation function. When the system frequency deviation is 0.1 Hz, the frequency change rate is 0.05 Hz / s, and the energy storage output power is 0.8 MW, the value of the first control effect evaluation function is 0.85. When the system frequency deviation increases to 0.2 Hz, the frequency change rate is 0.1 Hz / s, and the energy storage output power is 1.2 MW, the value of the first control effect evaluation function drops to 0.65.
[0051] For the transient stability energy storage group, the voltage deviation, tie-line power deviation, and energy storage output power are constructed into a second control effect evaluation function. For example, when the voltage deviation is 2%, the tie-line power deviation is 5%, and the energy storage output power is 1.5MW, the second control effect evaluation function value is 0.78. When the voltage deviation increases to 4%, the tie-line power deviation is 8%, and the energy storage output power is 2.0MW, the second control effect evaluation function value decreases to 0.62.
[0052] For the power-balancing energy storage group, the load power deviation and operating cost are used to construct a third control effect evaluation function. When the load power deviation is 3MW and the operating cost is 500 yuan / hour, the third control effect evaluation function value is 0.82. When the load power deviation increases to 5MW and the operating cost is 700 yuan / hour, the third control effect evaluation function value decreases to 0.70.
[0053] Based on the three control effect evaluation functions described above, a comprehensive reward signal is constructed. This signal takes into account the weighting of various energy storage groups. For example, the weights of the inertia response energy storage group, the transient stability energy storage group, and the power balance energy storage group are set to 0.4, 0.3, and 0.3, respectively. The comprehensive reward signal is calculated through weighted summation. In the above example, the comprehensive reward signal value is 0.82 × 0.4 + 0.78 × 0.3 + 0.82 × 0.3 = 0.808.
[0054] The inputs to the value network include the energy storage group's operating status, remaining life cost, and grid status. The energy storage group's operating status includes its current state of charge (SOC), charge / discharge depth, and number of charge / discharge cycles. For example, a certain energy storage group has a current SOC of 65% and has been charged and discharged 350 times, resulting in a depth of 80%. The remaining life cost is calculated using depreciation and efficiency loss of the energy storage equipment. In this example, the remaining life cost is 0.35 yuan per kilowatt-hour. Grid status includes bus voltage, line load factor, and system frequency. In this example, a test system has 10 buses with an average voltage of 220 kV, an average line load factor of 75%, and a system frequency of 49.9 Hz.
[0055] The value network uses a multilayer perceptron architecture, consisting of an input layer, two hidden layers, and an output layer. The number of input layer nodes is determined by the number of state variables; in this example, it has 12 nodes. The first hidden layer has 24 nodes, the second hidden layer has 16 nodes, and the output layer has one node representing the state value. The network uses the ReLU activation function, and the state value is calculated through forward propagation. In this example, the state value is 0.76.
[0056] The policy network uses a policy gradient method to update network parameters based on the state value assessment results. The policy network also uses a multi-layer perceptron architecture. The input layer has the same number of nodes as the value network, and the two hidden layers have 32 and 24 nodes, respectively. The output layer has twice the number of energy storage groups, representing the charge and discharge power, respectively. In this example, there are three energy storage groups, so the output layer has six nodes. The policy network uses a softmax activation function to output the probability distribution of the charge and discharge power of the energy storage groups.
[0057] The policy network's learning rate is adaptively adjusted based on historical gradients, with an initial learning rate of 0.01. If the gradient direction remains consistent for five consecutive iterations, the learning rate is increased by 20%. If the gradient direction continuously changes, the learning rate is decreased by 15%. In this example, the learning rate is adjusted to 0.0135 at the 30th iteration of training.
[0058] The experience replay pool is constructed by storing the energy storage group's charge and discharge power, integrated reward signals, and grid status as data samples. The pool is set to hold 10,000 samples, and training is performed using a random sampling method with batches of 128 samples. When the pool is full, data is updated using a first-in, first-out strategy.
[0059] The value network uses temporal difference error (TDE) for parameter updates. The TD target consists of the current reward and the estimated next-state value, with a discount factor of 0.95. In this example, the TD target for a state transition is calculated as 0.808 + 0.95 × 0.76 = 1.53, and the TD error is 1.53 - 0.76 = 0.77.
[0060] Figure 2 It is a comparison chart of comprehensive control performance under frequency disturbance conditions. The thick solid line with black dots in the figure represents the system frequency response of the present invention. It can be seen that its frequency recovery speed is faster, the fluctuation amplitude is smaller, and the steady-state error is almost zero; while the traditional MPC control method (dashed line, square mark) shows a larger frequency deviation and a longer recovery time. At the same time, the comprehensive reward signal (black dotted line) of the present invention is significantly higher than that of the traditional method, and it recovers quickly and stabilizes at a higher level, indicating that this control strategy has significant advantages in frequency control capability, system stability and economy. Especially in the early stage of disturbance, the control strategy of the present invention shows excellent anti-disturbance ability and fast recovery characteristics, which fully proves the technical innovation and practical value of the adaptive dynamic programming controller in the application of power system energy storage control.
[0061] The present invention uses the control effect of the distributed model predictive control framework as a reward signal to construct a value network and a policy network, thereby realizing autonomous learning and optimization of the energy storage control strategy. This method combines traditional control theory with deep reinforcement learning technology, overcoming the limitation that a single control method is difficult to cope with the complex dynamic characteristics of the power grid. The value network can comprehensively evaluate the state of the power grid and the remaining life cost of energy storage, while the policy network optimizes the energy storage charging and discharging decisions based on the value evaluation results, achieving a balance between short-term control effect and long-term economy. Through the experience replay pool and adaptive learning rate adjustment mechanism, the learning efficiency and convergence performance are improved, so that the control strategy can be continuously optimized and adapted to changes in the power grid operating environment, significantly improving the economy and control performance of the energy storage system.
[0062] In an optional embodiment, the step of adaptively adjusting the weights of each evaluation function in the comprehensive reward signal includes: Constructing a control performance index, the control performance index being calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within an evaluation time window; calculating the sensitivity of a first control effect evaluation function weight, a second control effect evaluation function weight, and a third control effect evaluation function weight to the control performance index; Collecting the root mean square of frequency deviation, the root mean square of voltage deviation, and the root mean square of power deviation to construct a state evaluation vector, and using the product of the state evaluation vector and the sensitivity as a weight update amount after exponential decay; Calculating an energy storage life loss assessment value based on the discharge depth and charge and discharge power of the energy storage group, and mapping the energy storage life loss assessment value into a weight adjustment constraint condition; Constructing a multi-objective optimization function including the control performance index, the energy storage life loss assessment value, and the weight change amount, calculating the weight optimization direction based on the gradient of the multi-objective optimization function, and determining the weight update value according to the preset weight adjustment step size; The weight update value is smoothed by using an exponential smoothing method to obtain the final values of the first control effect evaluation function weight, the second control effect evaluation function weight, and the third control effect evaluation function weight. The sum of the first control effect evaluation function weight, the second control effect evaluation function weight, and the third control effect evaluation function weight is 1.
[0063] For example, combined Figure 3 The adaptive adjustment flow chart for the integrated reward signal weights illustrates this. The control performance index is calculated by calculating the cumulative sum of squares of frequency, voltage, and power deviations within an evaluation time window. Specifically, a 10-second evaluation window is selected, and frequency, voltage, and power data from the microgrid are collected at a sampling interval of 0.1 seconds. Assuming that within a given evaluation time window, the cumulative sum of squares of frequency deviation is 0.25, the cumulative sum of squares of voltage deviation is 0.36, and the cumulative sum of squares of power deviation is 0.49, the control performance index can be expressed as a weighted combination of these three factors, with initial weights set to 0.4, 0.3, and 0.3, respectively.
[0064] The sensitivity calculation of the control performance index is based on perturbation analysis. By applying a small perturbation (e.g., 0.01) to the weights of the first, second, and third control effect evaluation functions, the changes in the control performance index are observed. For example, if the weight of the first control effect evaluation function increases from 0.4 to 0.41, and the control performance index increases from 0.345 to 0.348, the sensitivity of the first control effect evaluation function weight to the control performance index is 0.3. Similarly, the sensitivity of the other weights is calculated, resulting in a sensitivity vector of [0.3, 0.2, 0.25].
[0065] The RMS frequency, voltage, and power deviations are collected to construct the state assessment vector. The RMS frequency deviation is calculated by taking the square root of the sum of the squares of the difference between the measured frequency value and the rated value (e.g., 50 Hz) within the assessment window. The RMS voltage deviation is calculated by taking the square root of the sum of the squares of the per-unit differences between the measured voltage value and the rated value (e.g., 220 kV) within the assessment window. The RMS power deviation is calculated by taking the square root of the sum of the squares of the per-unit differences between the actual power and the planned power within the assessment window. In actual implementation, assuming the RMS frequency deviation, voltage deviation, and power deviation collected at a certain moment are 0.05, 0.06, and 0.07, the state assessment vector is [0.05, 0.06, 0.07]. The product of the state assessment vector and the sensitivity is [0.015, 0.012, 0.0175]. An exponential decay factor of 0.9 is introduced to perform exponential decay processing on the product result, and the weight update amount is obtained as [0.0135, 0.0108, 0.01575].
[0066] Energy storage life loss assessment is a key constraint for weight adjustment. The energy storage life loss assessment value is calculated based on the energy storage group's depth of discharge and charge / discharge power. Assuming the energy storage group's current depth of discharge is 20% and its charge / discharge power is 100kW, and a table lookup method determines that its cycle life at this time is 5000 cycles, the energy storage life loss assessment value is 0.0002 (i.e., 1 / 5000). This assessment value is mapped to a weight adjustment constraint. When the energy storage life loss assessment value is greater than 0.0005, the change in the weight of the first control effect evaluation function is limited to no more than 0.05. When the energy storage life loss assessment value is less than 0.0005, the restriction is relaxed, allowing the maximum change in the weight of the first control effect evaluation function to be 0.1.
[0067] The multi-objective optimization function consists of three components: a control performance indicator, an estimated energy storage lifespan loss, and a weight change. The weight of the control performance indicator is 0.6, the weight of the estimated energy storage lifespan loss is 0.3, and the weight change is 0.1. The weight optimization direction is calculated using the gradient descent method. Assume the resulting gradient direction is [-0.05, 0.03, 0.02]. If the preset weight adjustment step size is set to 0.02, the weight update value is [-0.001, 0.0006, 0.0004].
[0068] Exponential smoothing is performed on the updated weights, using a smoothing factor of 0.8. Assuming the weights of the first, second, and third control effect evaluation functions at the previous moment are 0.4, 0.3, and 0.3, respectively, the smoothed updated weights are [-0.0008, 0.00048, 0.00032]. Adding these updated values to the original weights yields new weights of [0.3992, 0.30048, 0.30032]. Since the sum of the three weights must be 1, normalization is performed, resulting in a final weight value of [0.399, 0.3005, 0.3005].
[0069] This invention achieves dynamic optimization of reward signals by constructing control performance indicators and calculating the sensitivity of the weights of each evaluation function. The energy storage life loss assessment is incorporated into the weight adjustment constraints, ensuring grid stability while also protecting the life of energy storage assets. By constructing a multi-objective optimization function and a smoothing mechanism, a balance is achieved between control performance, energy storage life, and weight changes. This avoids control instability caused by drastic weight fluctuations, ensures a smooth transition and continuity of grid control, and improves the overall operational efficiency of the system.
[0070] In an optional embodiment, the step of updating the parameters of the value network and the policy network includes: Constructing a dual-network structure of a target network and an evaluation network for the value network, calculating the temporal difference error based on the current state value, the next state value, and the immediate reward, the evaluation network constructing a loss function based on the temporal difference error and continuously updating the network parameters, and the target network determining the parameter update period based on the prediction error change rate of the evaluation network; Calculate the action probability distribution output by the policy network, calculate the importance sampling weight based on the probability ratio of the new and old policies, determine the truncation range based on the variance of the importance sampling weight, use the evaluation result of the value network as the baseline function, calculate the weighted sum of the temporal difference error using the generalized advantage estimation method, and update the policy network parameters based on the weighted sum; Construct a hybrid strategy that includes an exploration term and an exploitation term, where the weight of the exploration term gradually decays as the training progresses, and the weight of the exploitation term is dynamically adjusted based on the value assessment results; Calculate the timeliness evaluation value and importance evaluation value of the data sample, the timeliness evaluation value is obtained by the exponential function of the sample storage time, and the importance evaluation value is obtained by the absolute value of the time series difference error. Construct the sample priority based on the timeliness evaluation value and the importance evaluation value, sort the data samples in the experience pool according to the sample priority, and dynamically adjust the capacity of the experience pool according to the distribution range of the data sample priority.
[0071] For example, a dual-network structure, target network and evaluation network, is constructed for the value network. The value network uses a multi-layer perceptron architecture. The input layer contains state dimensions (including the energy storage group's state of charge, charge and discharge power, grid frequency, voltage, power deviation, and energy storage remaining life characteristics, a total of 12 features). The hidden layers have 24 and 16 nodes, respectively. The output layer consists of one node, representing the state value assessment result. The evaluation network and target network have the same network structure, but different parameter update mechanisms. The evaluation network parameters are updated in real time, while the target network parameters are periodically copied from the evaluation network. Specifically, the temporal difference error calculation process is as follows: batches of data (batch size 128) are randomly sampled from the experience replay pool. Each data entry contains the current state, action (energy storage group charge and discharge power), reward (comprehensive reward signal), next state, and a termination flag. For each data entry, the target network calculates the value estimate of the next state, and the evaluation network calculates the value estimate of the current state. The temporal difference error is equal to the immediate reward plus a discount factor (set to 0.95) multiplied by the value estimate of the next state, minus the value estimate of the current state. For example, when the immediate reward is 0.808, the current state value estimate is 0.76, and the next state value estimate is 0.81, the temporal difference error is 0.808 + 0.95 × 0.81 - 0.76 = 0.82. The evaluation network's loss function is the sum of squared temporal difference errors. Parameters are updated using the Adam optimizer with a learning rate of 0.001. The target network is not updated directly via gradient descent, but rather periodically copies parameters from the evaluation network. The update cycle is not fixed but dynamically adjusted based on the rate of change of the evaluation network's prediction error. When the prediction error rate of change (the relative change in prediction error over 10 consecutive iterations) is less than 0.05, the target network parameter update is triggered. In practice, the target network is updated approximately every 500 iterations initially, and this cycle is gradually increased to every 2000 iterations as training progresses.
[0072] The policy network outputs the action probability distribution and updates its parameters. The policy network also uses a multi-layer perceptron architecture. The input layer is identical to the value network, with 32 and 24 hidden layers, respectively. The output layer has twice the number of energy storage groups (in this example, there are three energy storage groups, so the output layer has six nodes, representing the charging and discharging power, respectively). The output is converted to an action probability distribution using a softmax function. During training, the network parameters before each policy update are saved and used to calculate the probability ratio between the new and old policies. Specifically, for each action sample in the experience pool, the ratio of its probability under the new policy to its probability under the old policy is calculated to obtain the importance sampling weight. For example, if the probability of an action under the old policy is 0.25 and the probability under the new policy is 0.30, the importance sampling weight is 1.2. When the variance of the importance sampling weight exceeds a preset threshold (e.g., 0.2), the importance sampling weight is truncated: weights exceeding twice the mean are set to twice the mean, and weights below 0.5 times the mean are set to 0.5 times the mean. The generalized advantage estimation method is used to calculate the weighted sum of the time series difference error. The specific steps are as follows: select a continuous state-action sequence (length 10) and calculate the advantage value of each time step backward from the end of the sequence. The advantage value is equal to the time series difference error of the current time step, plus a decay factor (set to 0.97) multiplied by the advantage value of the next time step. For example, if the time series difference errors of the last three time steps of a sequence are 0.82, 0.75, and 0.68 respectively, then the advantage value of the third-to-last time step is calculated as 0.82 + 0.97 × 0.75 + 0.97 2 ×0.68=2.19. The policy network's loss function is the negative of the product of the importance sampling weight and the advantage value, plus a regularization term for the policy entropy (with a coefficient of 0.01). The Adam optimizer is used to update the policy network parameters, with an initial learning rate of 0.01. The learning rate is adaptively adjusted based on the historical gradient. When the gradient direction is consistent for five consecutive iterations, the learning rate is increased by 20%; when the gradient direction continuously changes, the learning rate is decreased by 15%.
[0073] A hybrid policy is constructed that includes an exploration term and an exploitation term. The output probability distribution of the hybrid policy is the weighted sum of the exploration term and the exploitation term. The exploration term uses a uniform distribution to ensure that the policy network can explore the unknown action space. The exploitation term directly uses the output probability distribution of the policy network. The weight of the exploration term is initially set to 0.3 and decreases exponentially with each training round, with a decay coefficient of 0.995 and a minimum value of no less than 0.05, to ensure that the policy maintains a certain level of exploration capability. For example, at the 100th training round, the weight of the exploration term is 0.3 × 0.995^100 ≈ 0.18. The weight of the exploitation term is equal to 1 minus the weight of the exploration term and is dynamically adjusted based on the value assessment results. When the value assessment results are above the historical average, the weight of the exploitation term is increased by 0.01; when the value assessment results are below the historical average, the weight of the exploitation term is decreased by 0.01, but the total weight does not exceed 0.95. Assume that by the 200th training round, the exploration weight has decayed to 0.12. If the value assessment result at this time is higher than the historical average, the utilization weight will increase from 0.88 to 0.89, and the exploration weight will decrease accordingly to 0.11. This hybrid strategy mechanism allows for full exploration in the early stages of training, shifts to more utilization as the strategy matures, and dynamically adjusts the balance between exploration and utilization based on actual performance.
[0074] The experience replay pool stores data samples of the energy storage group's charge and discharge power, integrated reward signals, and grid status. The timeliness evaluation value of each data sample is calculated using the exponential decay function exp(-t / T), where t is the number of time steps the sample has been stored in the experience pool, and T is the time constant (set to 1000). The timeliness evaluation value of a new sample is 1, which decays to 0.61 after 500 time steps of storage and to 0.37 after 1000 time steps of storage. Next, the importance evaluation value of each data sample is calculated, normalized by the absolute value of the time series difference error. For example, if the absolute values of the time series difference error for a batch of samples are 0.82, 0.65, and 0.43, the normalized importance evaluation values are 0.43, 0.34, and 0.23, respectively. Sample priority is the weighted sum of the timeliness evaluation value and the importance evaluation value, with a weight ratio of 3:7. For example, if a sample has a timeliness assessment value of 0.8 and an importance assessment value of 0.6, its priority is 0.8 × 0.3 + 0.6 × 0.7 = 0.66. Data samples in the experience replay pool are sorted according to their priority, with higher-priority samples prioritized for training. The experience pool capacity is not fixed but dynamically adjusted based on the distribution of data sample priorities. Specifically, the difference between the 90th and 10th percentiles of a sample's priority is calculated. If the difference is greater than a preset threshold (e.g., 0.6), the experience pool capacity is increased by 1.2 times the original capacity, with a maximum of 20,000. If the difference is less than a preset threshold (e.g., 0.2), the experience pool capacity is reduced by 0.8 times the original capacity, with a minimum of 5,000. For example, if the initial experience pool capacity is 10,000, and the 90th and 10th percentiles of the sample's priority are 0.85 and 0.15, respectively, and the difference is 0.7, which is greater than the threshold of 0.6, the experience pool capacity is increased to 12,000. By dynamically adjusting the capacity of the experience pool, we can retain samples with sufficient diversity while saving memory usage.
[0075] Figure 4 This is a performance comparison chart of the parameter update of the value network and the policy network, which clearly shows the superiority of the dual network structure of the present invention compared to the traditional DQN network and PPO algorithm. In four key indicators, the present invention has achieved significant advantages: in terms of convergence speed, the present invention has achieved a 46.5% improvement; in terms of policy stability, the present invention has achieved a 44.0% improvement; in terms of sample utilization efficiency, the present invention has achieved a 50.0% improvement; the most outstanding is the control accuracy indicator, the present invention has achieved a 54.0% improvement, which is significantly ahead of the traditional DQN's 32.0% and PPO's 36.0%. It proves that the dual network structure design, dynamic update mechanism, hybrid strategy design and experience pool optimization mechanism based on timeliness and importance adopted by the present invention have significant effects in improving training stability, convergence efficiency and control accuracy.
[0076] The present invention improves the training stability and convergence efficiency of the value network through dual-network structure design and dynamic update mechanism of the target network; combines the policy network update method with importance sampling and generalized advantage estimation to effectively reduce the variance of the policy gradient and accelerate the policy optimization process; the dynamic balanced hybrid strategy design achieves a good balance between exploration and utilization; the experience pool optimization mechanism based on timeliness and importance significantly improves the sample utilization efficiency, making the entire learning process more efficient and stable, and ultimately realizing the autonomous optimization of the energy storage group control strategy, improving the accuracy and economy of power grid regulation.
[0077] A second aspect of an embodiment of the present invention provides a coordinated and complementary regulation system for an energy storage power station and a power grid, including: The first unit is used to perform multi-scale decomposition on the real-time frequency and voltage data on the grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data based on the grid topology. The second unit is used to build a distributed model predictive control framework based on the disturbance propagation prediction data, divide the energy storage resources into multiple energy storage groups according to the response time, configure a local optimization controller for each energy storage group, and construct an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieve global optimization through a consistency protocol; The third unit is used to design an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal to construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; The fourth unit is used to recalculate the state evaluation based on the new network structure when a change in the grid topology is detected, and to update the optimal charge and discharge power output by the strategy network; control each energy storage group to perform charge and discharge operations based on the optimized optimal charge and discharge power, and feed back the operating data to the adaptive dynamic programming controller.
[0078] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0079] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0080] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for cooperative and complementary regulation of energy storage power stations and power grids, characterized in that: include: Perform multi-scale decomposition on the real-time frequency and voltage data on the grid side to obtain disturbance components in different frequency bands. Combined with the grid topology calculation, disturbance propagation prediction data is obtained. Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieves global optimization through a consistency protocol. Design an adaptive dynamic programming controller, use the control effect of the distributed model predictive control framework as a reward signal, and construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; When a change in the grid topology is detected, the state evaluation is recalculated based on the new network structure, and the optimal charge and discharge power output by the strategy network is updated; and each energy storage group is controlled to perform charge and discharge operations according to the optimized optimal charge and discharge power.
2. The method according to claim 1, characterized in that The steps of performing multi-scale decomposition on the real-time frequency and voltage data on the grid side to obtain disturbance components in different frequency bands and then calculating the disturbance propagation prediction data based on the grid topology include: Performing multi-scale decomposition on the real-time frequency data and voltage data to obtain disturbance components in different frequency bands; Establishing a dynamic topology identification matrix, the dynamic topology identification matrix including a node set, an edge set, and a weight matrix, calculating the value of each element in the weight matrix based on the node voltage amplitude, the node voltage phase angle, and the inter-node reactance, the weight matrix reflecting the real-time electrical connection strength between the grid nodes; Calculating a prediction parameter vector by a recursive least squares method according to the disturbance components of the different frequency bands and the dynamic topology identification matrix; Calculating a prediction error within a sliding time window, and triggering an update calculation of the prediction parameter vector when the prediction error is greater than a first preset threshold or a change in the dynamic topology identification matrix is greater than a second preset threshold; The disturbance propagation prediction data is calculated based on the updated prediction parameter vector, the dynamic topology identification matrix and the real-time disturbance data of the disturbance source node.
3. The method according to claim 1, characterized in that Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed. Energy storage resources are divided into multiple energy storage groups according to response time. Each energy storage group is configured with a local optimization controller. The local optimization controller constructs an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage. The steps of achieving global optimization through a consistency protocol include: Based on the disturbance propagation prediction data, a distributed model predictive control framework is constructed to divide energy storage resources into inertia response, transient stability, and power balance energy storage groups according to response time; The local optimization controller of the inertia response energy storage group constructs a first objective function based on frequency deviation, frequency change rate and energy storage output power. The local optimization controller of the transient stability energy storage group constructs a second objective function based on voltage deviation, tie line power deviation and energy storage output power. The local optimization controller of the power balance energy storage group constructs a third objective function based on load power deviation and energy storage operating cost. The first, second and third objective functions are all constrained by the remaining capacity, charge and discharge efficiency and response speed of the corresponding energy storage group. A basic communication topology structure is established between energy storage groups. Communication status parameters between energy storage nodes are collected to construct a weighted adjacency matrix. The communication topology structure is optimized based on the weighted adjacency matrix. A distributed consistency protocol that considers delay compensation is used to achieve global optimization control of the energy storage group.
4. The method according to claim 3, characterized in that The steps of collecting communication status parameters between energy storage nodes to construct a weighted adjacency matrix, optimizing the communication topology based on the weighted adjacency matrix, and implementing global optimization control of the energy storage group using a distributed consistency protocol that considers delay compensation include: The communication quality factor, delay factor and bandwidth utilization between energy storage nodes are collected, and a weighted adjacency matrix is constructed using a fuzzy comprehensive evaluation method. Calculating network connectivity based on the weighted adjacency matrix; and when the network connectivity is lower than a preset connectivity threshold, reconstructing the communication topology using a minimum spanning tree algorithm that combines the spatiotemporal correlation characteristics of energy storage nodes to obtain a topological structure with optimal latency performance; Based on the reconstructed communication topology, a distributed consensus protocol with delay compensation is used for iterative calculation. The adaptive step size of the distributed consensus protocol is dynamically adjusted according to the iterative convergence speed. When the state deviation of the energy storage node exceeds the preset deviation threshold, the broadcast of the state information is triggered, and the adjacent nodes that receive the broadcast information update their respective optimization variables; The communication connectivity of energy storage nodes is calculated periodically. When a node communication anomaly is detected, the energy storage node with the highest connectivity is selected as the backup communication path to maintain the continuity of the global optimization process.
5. The method according to claim 3, characterized in that The steps of designing an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal, and constructing a value network and a strategy network, wherein the value network evaluates the grid status and the remaining life cost of the energy storage group, and the strategy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results, include: The frequency deviation, frequency change rate and energy storage output power of the inertia response energy storage group in the distributed model predictive control framework, the voltage deviation, tie line power deviation and energy storage output power of the transient stability energy storage group, and the load power deviation and operating cost of the power balance energy storage group are constructed as the first control effect evaluation function, the second control effect evaluation function and the third control effect evaluation function respectively; A comprehensive reward signal is constructed based on the first control effect evaluation function, the second control effect evaluation function, and the third control effect evaluation function; the value network outputs a state value evaluation result based on the operating state, remaining life cost, and grid state of the energy storage group; the strategy network uses a policy gradient method to update network parameters based on the state value evaluation result and output the charge and discharge power of the energy storage group; An experience replay pool is constructed using the charge and discharge power of the energy storage group, the comprehensive reward signal, and the grid status. The value network and the policy network are trained based on the data samples in the experience replay pool. The value network uses temporal difference error to update parameters, and the learning rate of the policy network is adaptively adjusted according to the historical gradient.
6. The method according to claim 5, characterized in that The step of adaptively adjusting the weights of each evaluation function in the comprehensive reward signal includes: Constructing a control performance index, the control performance index being calculated based on the cumulative sum of squares of frequency deviation, voltage deviation, and power deviation within an evaluation time window; calculating the sensitivity of a first control effect evaluation function weight, a second control effect evaluation function weight, and a third control effect evaluation function weight to the control performance index; Collecting the root mean square of frequency deviation, the root mean square of voltage deviation, and the root mean square of power deviation to construct a state evaluation vector, and using the product of the state evaluation vector and the sensitivity as a weight update amount after exponential decay; Calculating an energy storage life loss assessment value based on the discharge depth and charge and discharge power of the energy storage group, and mapping the energy storage life loss assessment value into a weight adjustment constraint condition; Constructing a multi-objective optimization function including the control performance index, the energy storage life loss assessment value, and the weight change amount, calculating the weight optimization direction based on the gradient of the multi-objective optimization function, and determining the weight update value according to the preset weight adjustment step size; An exponential smoothing method is used to smooth the weight update value.
7. The method according to claim 5, characterized in that The step of updating the parameters of the value network and the policy network includes: Constructing a dual-network structure of a target network and an evaluation network for the value network, calculating the temporal difference error based on the current state value, the next state value, and the immediate reward, the evaluation network constructing a loss function based on the temporal difference error and continuously updating the network parameters, and the target network determining the parameter update period based on the prediction error change rate of the evaluation network; Calculate the action probability distribution output by the policy network, calculate the importance sampling weight based on the probability ratio of the new and old policies, determine the truncation range based on the variance of the importance sampling weight, use the evaluation result of the value network as the baseline function, calculate the weighted sum of the temporal difference error using the generalized advantage estimation method, and update the policy network parameters based on the weighted sum; Construct a hybrid strategy that includes an exploration term and an exploitation term, where the weight of the exploration term gradually decays as the training progresses, and the weight of the exploitation term is dynamically adjusted based on the value assessment results; Calculate the timeliness evaluation value and importance evaluation value of the data sample, the timeliness evaluation value is obtained by the exponential function of the sample storage time, and the importance evaluation value is obtained by the absolute value of the time series difference error. Construct the sample priority based on the timeliness evaluation value and the importance evaluation value, sort the data samples in the experience pool according to the sample priority, and dynamically adjust the capacity of the experience pool according to the distribution range of the data sample priority.
8. A coordinated and complementary regulation system of an energy storage power station and a power grid, for implementing the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to perform multi-scale decomposition on the real-time frequency data and voltage data on the grid side, obtain disturbance components in different frequency bands, and calculate disturbance propagation prediction data in combination with the grid topology structure; The second unit is used to build a distributed model predictive control framework based on the disturbance propagation prediction data, divide the energy storage resources into multiple energy storage groups according to the response time, configure a local optimization controller for each energy storage group, and construct an objective function based on the remaining capacity, charge and discharge efficiency, and response speed of the energy storage, and achieve global optimization through a consistency protocol; The third unit is used to design an adaptive dynamic programming controller, using the control effect of the distributed model predictive control framework as a reward signal to construct a value network and a policy network. The value network evaluates the system status and the remaining life cost of the energy storage group, and the policy network outputs the optimal charge and discharge power of the energy storage group based on the value evaluation results; The fourth unit is used to recalculate the state evaluation based on the new network structure when a change in the grid topology is detected, and to update the optimal charge and discharge power output by the strategy network; control each energy storage group to perform charge and discharge operations based on the optimized optimal charge and discharge power, and feed back the operating data to the adaptive dynamic programming controller.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Power grid side energy storage optimal configuration and operation control method and system
CN119171486A
Multi-scale energy storage system control method and system
CN119419889A
Networking type energy storage hierarchical optimization site selection method considering multi-scale stability influence
CN119765403A
Power-factor-corrected resonant converter and parallel power-factor-corrected resonant converter
US20130154372A1
Cited By
Intelligent power regulation and control and power grid interaction system and method for energy storage battery box
CN120914861A
Energy storage rapid frequency modulation control method and system based on power grid frequency response
CN121216515A
Coordination control method for flywheel energy storage array
CN121308049A
Micro-grid energy coordination control method considering dynamic change of communication topology
CN121529792A
A micro-grid energy coordination control method considering dynamic change of communication topology
CN121529792B