A distributed energy storage coordinated frequency modulation parameter fast self-adaptive setting system and method

CN122553213APending Publication Date: 2026-08-11HANGZHOU E ENERGY ELECTRIC POWER TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-27
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

控制器参数整定依赖人工经验:多数DESS调频控制器仍采用固定或离线整定的比例-积分-微分(Proportional-Integral-Derivative,PID)参数,在面对复杂多变的运行工况(如负荷突变、新能源剧烈波动、储能荷电状态(State of Charge,SOC)差异大等)时,难以兼顾响应速度、超调抑制与稳态精度,易引发振荡或调节失效

Benefits of technology

1.大幅提升调频响应速度与精度:通过构建毫秒级模糊PID快速响应层,结合改进型深度强化学习(DRL)驱动的参数自整定算法,可在45ms内完成PID参数动态优化(较传统方法缩短62.5%),有效抑制频率突变,频率偏差恢复时间缩短40%以上。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122553213A_ABST
    Figure CN122553213A_ABST
Patent Text Reader

Abstract

This invention discloses a system and method for rapid adaptive tuning of distributed energy storage cooperative frequency regulation parameters. The system includes: a data acquisition module for real-time acquisition of grid frequency deviation, SOC status of distributed energy storage units, communication data between adjacent nodes, and historical frequency regulation response data; a multi-timescale hierarchical control module for generating preliminary control commands based on the acquired data; a parameter self-tuning algorithm module that dynamically generates proportional-integral-derivative controller parameters using a hybrid algorithm based on an improved deep reinforcement learning (DRL) framework, incorporating Q-learning and a neural network containing a long short-term memory network; a distributed cooperative communication module for real-time data synchronization and cooperative decision-making among nodes in the distributed energy storage system; and a safety verification module for providing over-limit warnings and dynamic constraints on the parameter tuning results generated by cooperative decision-making. This invention improves the frequency regulation response speed and accuracy while ensuring the safe operation of the distributed energy storage system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field This invention belongs to the field of power system automation control technology, specifically a system and method for rapid adaptive tuning of distributed energy storage cooperative frequency regulation parameters based on multiple time scales. Background Technology

[0001] Currently, the penetration rate of renewable energy sources such as wind power and photovoltaics in the power system continues to increase; however, the output of renewable energy is highly volatile and intermittent, significantly weakening the inertia support and primary frequency regulation capabilities provided by traditional synchronous generator units, leading to severe challenges to grid frequency stability. Against this backdrop, Distributed Energy Storage Systems (DESS) are widely regarded as a key supporting resource for improving the frequency stability of new power systems due to their advantages such as fast response speed, flexible deployment, and strong bidirectional regulation capabilities.

[0002] Currently, research on distributed energy storage systems participating in grid frequency regulation mainly focuses on control strategy design and coordination mechanism construction. Common methods include droop control, virtual inertia control, and optimization scheduling strategies based on Model Predictive Control (MPC) or artificial intelligence. However, existing technologies generally suffer from the following shortcomings: Controller parameter tuning relies on human experience: Most DESS frequency controllers still use fixed or offline tuned proportional-integral-derivative (PID) parameters. When faced with complex and ever-changing operating conditions (such as sudden load changes, drastic fluctuations in new energy sources, and large differences in the state of charge (SOC) of energy storage), it is difficult to balance response speed, overshoot suppression, and steady-state accuracy, which can easily lead to oscillations or regulation failures.

[0003] Lack of multi-timescale coordination mechanism: Grid frequency disturbances range from instantaneous impacts at the millisecond level to continuous power imbalances at the minute level, while existing control architectures often focus only on a single time scale (such as only fast response or only economic dispatch), failing to achieve effective decoupling and coordination of control objectives at different time scales, resulting in limited overall system performance.

[0004] Insufficient efficiency and reliability of distributed collaboration: Collaboration between large-scale DESS nodes depends on the communication network, but existing solutions mostly adopt centralized or simple distributed architectures, which have problems such as high communication latency, poor consistency, and weak anti-interference ability. In particular, in the 5G / fiber hybrid networking environment, there is a lack of robust handling mechanisms for uncertain factors such as latency and packet loss.

[0005] Insufficient consideration of safety constraints: If SOC balancing, equipment capacity limits and grid safety boundaries are ignored during the parameter adaptive adjustment process, individual energy storage units may be overcharged / over-discharged or the system may operate beyond its limits, affecting equipment lifespan and grid safety.

[0006] Therefore, there is an urgent need for a fast adaptive tuning method for distributed energy storage frequency regulation parameters that can integrate multi-source real-time information, support multi-timescale collaboration, have online self-learning capabilities, and meet high reliability communication and security constraints, in order to address the new challenges of grid frequency regulation under the high proportion of renewable energy access. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a fast adaptive tuning system and method for distributed energy storage cooperative frequency regulation parameters based on multiple time scales. It dynamically optimizes the frequency regulation parameters in real time to adapt to the multi-time scale changes of the grid frequency, improves the frequency regulation response speed and accuracy, and ensures the safe operation of the distributed energy storage system.

[0008] Therefore, the present invention adopts the following technical solution.

[0009] In a first aspect, the present invention provides a distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system, comprising: Data acquisition module: used to acquire grid frequency deviation, SOC status of distributed energy storage units, communication data between adjacent nodes, and historical frequency regulation response data in real time; Multi-timescale hierarchical control module: Generates preliminary control commands based on the data collected by the data acquisition module. This module includes a millisecond-level fast response layer, a second-level dynamic adjustment layer, and a minute-level global optimization layer. The layers are coupled together through dynamic weight factors. The parameter self-tuning algorithm module is associated with the initial control commands generated by the millisecond-level fast response layer, the second-level dynamic adjustment layer, and the minute-level global optimization layer. Based on the improved deep reinforcement learning (DRL) framework, it dynamically generates proportional-integral-derivative controller parameters through a hybrid algorithm that integrates Q-learning with a neural network containing a long short-term memory network. Distributed collaborative communication module: associated with the parameters of the proportional-integral-derivative controller, it adopts a fiber-optic-5G hybrid communication network based on a consensus protocol and is networked through a clustered topology to realize real-time data synchronization and collaborative decision-making among nodes of the distributed energy storage system; Safety verification module: It is associated with the collaborative decision parameter tuning results output by the distributed collaborative communication module, and is used to provide over-limit warnings and dynamic constraints on the parameter tuning results generated by collaborative decision-making, so as to ensure that all distributed energy storage units operate within the preset safety domain.

[0010] This invention achieves rapid adaptive tuning of frequency regulation parameters by integrating data acquisition, hierarchical control, parameter self-tuning, collaborative communication, and security verification modules. It has the advantages of being able to dynamically optimize frequency regulation parameters in real time, adapt to multi-timescale changes in grid frequency, improve frequency regulation response speed and accuracy, and ensure the safe operation of distributed energy storage units.

[0011] This invention is an adaptive tuning technique specifically for frequency control controller parameters, applicable to frequency stability control of power grids with a high proportion of renewable energy integration.

[0012] Furthermore, the multi-timescale hierarchical control module specifically includes: 1) Millisecond-level fast response layer: adopts fuzzy PID control algorithm, with a response time ≤50ms, and prioritizes the suppression of frequency mutations; 2) Second-level dynamic adjustment layer: The frequency regulation responsibility coefficient of each distributed energy storage unit is calculated through dynamic sensitivity analysis and the droop coefficient is adjusted to compensate for the fluctuation of renewable energy output; 3) Minute-level global optimization layer: Based on model predictive control, optimize the SOC balance of the entire network, with an optimization cycle of 1-5 minutes.

[0013] Furthermore, the millisecond-level fast response layer is deployed on the local controller and uses hardware-in-the-loop technology to achieve fuzzy PID fast frequency modulation with a control cycle of 15-25ms.

[0014] Furthermore, the second-level dynamic adjustment layer calculates the frequency regulation responsibility coefficient of each distributed energy storage unit in real time based on the dynamic sensitivity matrix.

[0015] Furthermore, the minute-level global optimization layer establishes a global optimization model based on model predictive control, and uses a distributed alternating direction multiplier method algorithm to solve the global optimization model in a distributed manner to optimize the SOC balance of the entire network.

[0016] Furthermore, the parameter self-tuning algorithm module includes: 1) State-space definition unit: The input vectors are the grid frequency deviation Δf, df / dt, the SOC state of each distributed energy storage unit and the parameter deviation of adjacent nodes. df / dt represents the rate of change of grid frequency with time. 2) Reward function design unit: R=α·(1 / |Δf|)+β·(SOC balance)+γ·(communication delay penalty term), where R represents the reward value, and α, β and γ are all weight coefficients; 3) Policy learning unit: A neural network containing a long short-term memory network is used as a function approximator, and the network parameters are updated using a dual-delay deep deterministic policy gradient algorithm.

[0017] Furthermore, the distributed collaborative communication module adopts a clustered topology for networking, with each cluster containing one master node and N slave nodes; the master node uses an improved Raft consensus algorithm to verify the consistency of data within the cluster and ensures that the communication latency within the cluster is ≤10ms.

[0018] Secondly, this invention provides a method for rapid adaptive tuning of distributed energy storage collaborative frequency regulation parameters, employing the aforementioned system for rapid adaptive tuning of distributed energy storage collaborative frequency regulation parameters, comprising the following steps: S1. Through the data acquisition module, the grid frequency deviation, the SOC status of the distributed energy storage unit, the communication data of adjacent nodes, and the historical frequency regulation response data are collected in real time; S2. Input the data into the multi-timescale hierarchical control module to generate preliminary hierarchical control instructions for different time scales; S3. Call the parameter self-tuning algorithm module to dynamically optimize the PID controller parameters based on the current system state and the initial hierarchical control instructions; S4. Through the distributed collaborative communication module, the optimized PID controller parameters are distributed and collaboratively decided and synchronized based on the consensus protocol; S5. The security verification module is used to perform security verification on the final parameters after collaborative decision-making, and the security control command that passes the verification is output to each distributed energy storage unit for execution.

[0019] Furthermore, in step S5, after outputting the safety control command to each distributed energy storage unit, the experience data consisting of the current control cycle's state, action, reward, and next state is stored in the improved deep reinforcement learning (DRL) experience replay pool for online learning, and then the next control cycle begins.

[0020] Furthermore, in step S5, if the security check fails, parameter correction or warning will be issued.

[0021] Compared with the prior art, the present invention has the following beneficial effects: 1. Significantly improve frequency modulation response speed and accuracy: By constructing a millisecond-level fuzzy PID fast response layer and combining it with an improved deep reinforcement learning (DRL) driven parameter self-tuning algorithm, dynamic optimization of PID parameters can be completed within 45ms (62.5% shorter than traditional methods), effectively suppressing frequency mutations and reducing frequency deviation recovery time by more than 40%.

[0022] 2. Achieve multi-timescale collaborative optimization: Innovatively designed a three-level hierarchical control architecture of "millisecond-second-minute" to deal with instantaneous disturbances, dynamic fluctuations and long-term energy management needs respectively. Each layer is coupled through dynamic weight factors to ensure speed while taking into account SOC balance and system economy. The maximum SOC difference of the entire network is reduced from 23% to 7.8%.

[0023] 3. Strong adaptability and generalization ability: Based on the state perception and reward mechanism of the hybrid DRL framework of neural network (LSTM-TD3) including long short-term memory network, the system can learn the optimal control strategy in different operating scenarios online with parameter tuning error ≤3% and without relying on an accurate system model. It is suitable for various power grid topologies and new energy penetration scenarios.

[0024] 4. Highly reliable distributed collaborative communication: It adopts a fiber-optic-5G hybrid communication network based on a consensus protocol and a clustered Raft consensus algorithm, with a communication latency of ≤10ms and a packet loss rate reduced from 0.15% to 0.02%. It also supports dual-channel redundant transmission and clock synchronization error compensation (accuracy ≤1μs), which significantly improves the real-time performance and robustness of collaborative decision-making for large-scale nodes (supporting ≥1000 nodes).

[0025] 5. Built-in multiple safety verification mechanisms: The safety verification module dynamically constrains the setting parameters and provides over-limit warnings to ensure that all distributed energy storage units always operate within the safe SOC range and equipment capacity range, effectively preventing overcharging / over-discharging and system over-limit risks, and ensuring long-term stable operation. Attached Figure Description

[0026] Figure 1 This is a diagram showing the composition of the distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system of the present invention; Figure 2 This is a flowchart of the distributed energy storage coordinated frequency regulation parameter fast adaptive tuning method of the present invention. Detailed Implementation

[0027] The technical solutions of this invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of this invention. The components of this invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0028] For ease of understanding, some key terms in this embodiment are explained below.

[0029] Deep reinforcement learning framework: This framework integrates the perception capabilities of deep learning with the decision-making capabilities of reinforcement learning. Through interactive learning between the agent and the environment, it aims to maximize cumulative rewards and autonomously optimize control strategies.

[0030] Q-learning: As a model-free reinforcement learning algorithm, Q-learning guides the agent to choose the optimal action in a given state by learning the action-value function (Q-function), thereby gradually approaching the optimal policy.

[0031] Neural networks: This computational model consists of interconnected nodes (neurons) that can learn complex patterns in input data to perform tasks such as function approximation, classification, or prediction.

[0032] Proportional-Integral-Derivative Controller: This controller is a classic feedback control algorithm that combines proportional, integral, and derivative terms to process system errors, thereby achieving precise and stable control of the controlled object.

[0033] Consistency Protocol: This protocol aims to ensure that multiple nodes in a distributed system can reach a consensus on a certain data or decision when faced with communication delays, data loss, or node failures, thereby maintaining the synchronization and consistency of the system state.

[0034] SOC Status: This status represents the percentage of the current charge of a distributed energy storage unit and is an important indicator for assessing the available capacity and operational health of the energy storage unit.

[0035] Dynamic weighting factor: This factor is used in multi-level or multi-objective control strategies to flexibly adjust the priority and degree of influence between different control levels or different control objectives based on the real-time operating status of the system and preset objectives.

[0036] Example 1 This embodiment describes a distributed energy storage collaborative frequency regulation parameter fast adaptive tuning system, such as... Figure 1 As shown, it includes: Data acquisition module: used to acquire in real time the grid frequency deviation Δf, the SOC status of the distributed energy storage unit, the communication data of adjacent nodes, and historical frequency regulation response data; Multi-timescale hierarchical control module: Generates preliminary control commands based on the data collected by the data acquisition module. This module includes a millisecond-level fast response layer, a second-level dynamic adjustment layer, and a minute-level global optimization layer. The layers are coupled together through dynamic weight factors. The parameter self-tuning algorithm module is associated with the initial control commands generated by the millisecond-level fast response layer, second-level dynamic adjustment layer, and minute-level global optimization layer in the multi-time-scale hierarchical control module. Based on the improved deep reinforcement learning (DRL) framework, it dynamically generates proportional-integral-derivative (PID) controller parameters (Kp, Ki, Kd) through a hybrid algorithm that integrates Q-learning with a neural network containing a long short-term memory network. Distributed collaborative communication module: It is associated with the proportional-integral-derivative controller parameters output by the parameter self-tuning algorithm module, and adopts a fiber-optic-5G hybrid communication network based on a consensus protocol. It is networked through a clustered topology to realize real-time data synchronization and collaborative decision-making among nodes of the distributed energy storage system. Safety verification module: It is associated with the collaborative decision parameter tuning results output by the distributed collaborative communication module, and is used to provide over-limit warnings and dynamic constraints on the parameter tuning results generated by collaborative decision-making, so as to ensure that all distributed energy storage units operate within the preset safety domain.

[0037] This system achieves rapid adaptive tuning of distributed energy storage frequency regulation parameters by combining multi-timescale hierarchical control with deep reinforcement learning, effectively solving the problems of traditional solutions that rely on manual experience and lack multi-timescale coordination in parameter tuning. Simultaneously, based on a fiber-optic-5G hybrid communication network and consensus protocol, the efficiency and reliability of distributed coordination are improved. Furthermore, a security verification module ensures the safe and stable operation of the energy storage unit under complex operating conditions, comprehensively enhancing the stability and resilience of the grid frequency under high-proportion renewable energy access.

[0038] The data acquisition module is configured to acquire key parameters of the power grid operation in real time. For example, grid frequency deviation can be measured using frequency sensors installed at key nodes in the grid, and the analog signals are converted into digital signals for transmission. The State of Charge (SOC) of distributed energy storage units can be estimated using algorithms within the Battery Storage Management System (BMS) and periodically uploaded to the data acquisition module. Neighbor node communication data can be received via standard communication interfaces. Historical frequency regulation response data can be stored in a local database and retrieved as needed. This data can be collected centrally or independently by each distributed energy storage unit and then uploaded to a central aggregation point.

[0039] The multi-timescale hierarchical control module is designed to generate initial control commands based on data provided by the data acquisition module. This module can contain multiple logically independent control levels; for example, one level might focus on rapidly suppressing instantaneous frequency disturbances, another on compensating for power imbalances at medium timescales, and yet another on achieving long-term system optimization goals. These levels can operate independently or be coordinated through preset logical rules or algorithms. The levels can be coupled through dynamic weighting factors; for instance, the weight of the fast response layer can be increased when the system frequency fluctuates drastically, while the weight of the global optimization layer can be increased when the system stabilizes. These weighting factors can be adjusted based on system state or preset strategies.

[0040] Specifically, the multi-timescale hierarchical control module includes: 1) Millisecond-level fast response layer: adopts fuzzy PID control algorithm, with a response time ≤50ms, and prioritizes the suppression of frequency mutations.

[0041] The millisecond-level fast response layer is designed to react rapidly to sudden and drastic changes in grid frequency. This layer employs a fuzzy PID control algorithm, the core of which combines fuzzy logic reasoning with traditional proportional-integral-derivative (PID) control. Specifically, the fuzzy PID control algorithm can adaptively adjust the parameters of the PID controller online based on real-time monitored grid frequency deviations and their rates of change, using a pre-defined fuzzy rule base and membership functions. This method does not require a precise system mathematical model, effectively handling nonlinearities and uncertainties in the system, thereby generating control commands in an extremely short time (response time ≤ 50ms), prioritizing the suppression of sudden frequency fluctuations, and ensuring rapid system stability in the face of instantaneous shocks. In addition to the fuzzy PID control algorithm, this layer can also employ strategies based on high-gain proportional control or feedforward control to quickly detect frequency abrupt changes and immediately issue reverse power commands, achieving rapid suppression of instantaneous frequency deviations.

[0042] 2) Second-level dynamic adjustment layer: The frequency regulation responsibility coefficient of each distributed energy storage unit is calculated through dynamic sensitivity analysis and the droop coefficient is adjusted to compensate for the fluctuation of renewable energy output.

[0043] The second-level dynamic adjustment layer is primarily responsible for handling dynamic frequency changes caused by fluctuations in renewable energy output in the power grid. This layer adjusts the droop coefficient through dynamic sensitivity analysis. Dynamic sensitivity analysis refers to the real-time assessment of the impact of grid operating conditions (such as load levels and renewable energy output forecasts) on the frequency response, and dynamically adjusting the droop control parameters of the energy storage units accordingly. For example, when a large fluctuation trend in renewable energy output is detected, the system will adjust the droop coefficient in a timely manner based on the sensitivity analysis results, enabling the energy storage units to participate in frequency regulation more actively or conservatively, thereby effectively compensating for the impact of renewable energy output fluctuations on the grid frequency. In addition, this layer can also employ an adaptive droop control strategy, dynamically adjusting the droop coefficient by estimating the system inertia or damping characteristics online to adapt to changes in grid operating conditions and ensure dynamic frequency stability.

[0044] 3) Minute-level global optimization layer: Based on model predictive control, optimize the SOC balance of the entire network, with an optimization cycle of 1-5 minutes.

[0045] The minute-level global optimization layer focuses on the long-term optimization of the distributed energy storage system as a whole, particularly the state of charge (SOC) balance of all energy storage units in the network. This layer is based on model predictive control (MPC), which works by using models of the energy storage system and the power grid to predict the system behavior over a future period (e.g., 15-30 minutes) within each optimization cycle (1-5 minutes). Through rolling optimization, MPC can calculate the optimal charging and discharging plan for energy storage units to maximize the balance of SOC across the network, while meeting the grid frequency regulation requirements and the operational constraints of the energy storage units themselves (such as power limits, SOC upper and lower limits, etc.). This optimization method ensures that all distributed energy storage units can work collaboratively, avoiding overcharging or over-discharging of individual units, extending equipment lifespan, and reserving sufficient regulation margin for subsequent frequency regulation tasks. In addition to model predictive control, this layer can also employ distributed optimization algorithms, such as the alternating direction multiplier method (ADMM), enabling each energy storage unit to achieve global SOC balance optimization through iterative calculation based on local communication.

[0046] The millisecond-level fast response layer is deployed on the local controller and uses hardware-in-the-loop technology to achieve fuzzy PID fast frequency modulation with a control cycle of 15-25ms.

[0047] The aforementioned second-level dynamic adjustment layer calculates the frequency regulation responsibility coefficient of each distributed energy storage unit in real time based on the dynamic sensitivity matrix, using the following formula:

[0048] In the formula, Indicates the first The current state of charge of each distributed energy storage unit Indicates the first The current state of charge of each distributed energy storage unit Indicates the first The maximum available frequency regulation power of each distributed energy storage unit Indicates the first The maximum available frequency regulation power of each distributed energy storage unit The minute-level global optimization layer establishes a global optimization model based on model predictive control, and uses a distributed alternating direction multiplier method algorithm to solve the global optimization model in a distributed manner to optimize the SOC balance of the entire network. The objective function is:

[0049] In the formula, The weighting coefficients representing the frequency deviation term. The weighting coefficients of the SOC deviation term are represented by... Represents the grid frequency deviation at time step t. The SOC deviation vector of the distributed energy storage units in the entire network (i.e., the difference between the SOC of each unit and the average SOC of the entire network) is represented by T, which represents the total number of time steps in the optimization time domain.

[0050] Through the above technical solution, this invention achieves effective decoupling and coordinated control of grid frequency disturbances at different time scales. The millisecond-level fast response layer can quickly suppress frequency abrupt changes, ensuring instantaneous system stability; the second-level dynamic adjustment layer dynamically compensates for renewable energy output fluctuations, maintaining system dynamic balance; and the minute-level global optimization layer optimizes the SOC balance of all grid energy storage units, ensuring long-term efficient system operation. This hierarchical control mechanism allows each layer to focus on addressing frequency issues at specific time scales, avoiding the limitations of a single control strategy when facing complex and variable operating conditions. Simultaneously, this multi-time-scale hierarchical control module provides the parameter self-tuning algorithm module with more stable and representative system state information and preliminary control commands, enabling the deep reinforcement learning-based parameter self-tuning algorithm to optimize proportional-integral-derivative controller parameters more efficiently and accurately. This significantly improves the overall response speed, regulation accuracy, and operational reliability of the distributed energy storage coordinated frequency regulation system, effectively addressing the severe challenges of grid frequency regulation under high-proportion renewable energy access.

[0051] The parameter self-tuning algorithm module is configured to connect with the multi-timescale hierarchical control module to dynamically generate proportional-integral-derivative (PDT) controller parameters. This module employs a deep reinforcement learning framework, where Q-learning is used to learn the optimal parameter tuning strategy, while a neural network is used as a function approximator to handle high-dimensional state inputs and output corresponding actions (i.e., parameter adjustments). This hybrid algorithm continuously optimizes its internal strategy through constant interaction with the power grid environment, enabling the controller parameters to adaptively adjust to changes in power grid operating conditions. For example, during the training phase, the module can receive power grid state information as input and output a set of PID parameters. It then obtains a reward signal based on the performance of these parameters in the simulation environment, thereby updating its internal model.

[0052] Specifically, the parameter self-tuning algorithm module includes: 1) State-space definition unit: The input vectors are the grid frequency deviation Δf, df / dt, the SOC state of each distributed energy storage unit and the parameter deviation of adjacent nodes. df / dt represents the rate of change of grid frequency with time, i.e. the frequency change rate.

[0053] The df / dt state-space definition unit is the foundation for reinforcement learning agents to perceive their environment. It abstracts the current state of the environment into a set of numerical or symbolic values ​​that can be processed by the learning algorithm. Its role is to provide the learning algorithm with comprehensive and crucial system information so that the agent can understand the current operating conditions and make decisions. In addition to the input vectors mentioned above, this unit can also use active power, reactive power, voltage amplitude, load forecast data, etc., as part of the input vector to more comprehensively reflect the grid operating status. Alternatively, the state-space definition unit can also adopt a multidimensional tensor form to encode different types of data (such as scalars, time series, and topology information) and extract features through convolutional neural networks or graph neural networks to adapt to more complex grid environments.

[0054] 2) Reward function design unit: R=α·(1 / |Δf|)+β·(SOC balance)+γ·(communication delay penalty term).

[0055] Here, R represents the reward value, which is the feedback signal obtained by the reinforcement learning agent from the environment after performing an action, used to evaluate the merits of the action; α, β, and γ are weighting coefficients used to balance the importance of different optimization objectives. These coefficients are usually positive and can be dynamically adjusted through expert experience, heuristic methods, or meta-learning algorithms. The reward function design unit is a core component of reinforcement learning, defining the agent's objectives and behavioral guidelines. By designing a reasonable reward function, the agent can be guided to learn the desired policy, enabling it to optimize specific objectives (such as frequency stability and SOC balance) while avoiding undesirable behaviors (such as excessive communication latency). For example, the reward function can further introduce penalty terms, such as limiting the charging and discharging power of energy storage units, penalizing equipment lifespan depletion, or penalizing states exceeding the safe operating range, to enhance the robustness and safety of the system. Furthermore, the reward function can also take the form of a piecewise or nonlinear function, for example, imposing a small penalty when the frequency deviation is small, and a large penalty when the frequency deviation exceeds the safe threshold, to more finely guide the learning process.

[0056] 3) Policy learning unit: A neural network containing a long short-term memory network is used as a function approximator, and the network parameters are updated using a dual-delay deep deterministic policy gradient algorithm.

[0057] The policy learning unit (LSM) employs a neural network containing a long short-term memory (LSTM) network as a function approximator and uses a dual-delay deep deterministic policy gradient (TDPG) algorithm to update network parameters. The LSM is the core decision-making mechanism of the reinforcement learning agent, responsible for outputting actions based on the current state, i.e., PID controller parameters. By continuously interacting with the environment and receiving reward signals, the LSM optimizes its decision-making policy, enabling it to generate optimal controller parameters under various operating conditions. The LSTM network, a special type of recurrent neural network, excels at processing and predicting time-series data. By introducing a gating mechanism, it effectively solves the gradient vanishing and gradient exploding problems in traditional recurrent neural networks, enabling it to capture long-term dependencies and thus more accurately predict future states and generate highly adaptive controller parameters. The dual-delay deep deterministic policy gradient (TDPG) algorithm is an off-policy reinforcement learning algorithm based on an Actor-Critic architecture. By introducing a dual Q-network, target policy smoothing, and delayed policy updates, it effectively solves the problems of overestimation of Q-values ​​and policy oscillations in the deep deterministic policy gradient (DDPG) algorithm, thereby improving learning stability and convergence efficiency. In addition to neural networks incorporating Long Short-Term Memory (LSTM) networks, the policy learning unit can also employ Transformer networks or graph neural networks as function approximators to better handle complex temporal dependencies and distributed topological data. Furthermore, besides the dual-delay deep deterministic policy gradient algorithm, the policy learning unit can also use soft Actor-Critic (SAC) algorithms or proximal policy optimization (PPO) algorithms to update network parameters. These algorithms each have their own advantages in terms of the balance between exploration and exploitation, sample efficiency, and convergence stability.

[0058] The specific algorithm implementation of the policy learning unit in the parameter self-tuning algorithm module is as follows: DRL framework design: A. The Actor network outputs the adjustment values ​​of the PID parameters ΔKp, ΔKi, and ΔKd; B. The Critic network evaluates the state-action value function and uses the Prioritized Experience Replay (PER) mechanism to accelerate convergence; C. Exploration strategy: Introduce OU (Ornstein–Uhlenbeck) noise perturbation, with the noise attenuation coefficient decreasing exponentially with the number of training rounds.

[0059] Offline training and online updates: A. Pre-training phase: Generate initial strategies based on a historical fault scenario library; B. Online Phase: The policy network parameters are updated every 5 minutes, using the following formula:

[0060] In the formula, The parameter vector of the Actor policy network, Indicates learning rate, The operator that calculates the gradient with respect to the network parameter θ This represents the objective function for strategy performance.

[0061] Through the above technical solution, the state-space definition unit uses the grid frequency deviation Δf, frequency change rate df / dt, energy storage state of charge (SOC), and adjacent node parameter deviation as input vectors. This allows for comprehensive capture of the dynamic changes in the grid, the operating status of energy storage units, and the collaborative information between distributed systems. This ensures that the reinforcement learning agent, when performing parameter tuning, can rely on sufficient and critical real-time data, avoiding decision-making biases caused by missing information, thereby improving the accuracy and adaptability of parameter tuning. The reward function design unit incorporates multiple key performance indicators such as frequency stability, energy storage SOC balance, and communication efficiency into the optimization objectives through a composite form R=α•(1 / |Δf|)+β•(SOC balance)+γ•(communication delay penalty term). This multi-objective reward mechanism guides the policy learning unit to comprehensively consider the trade-offs between different objectives when optimizing PID controller parameters, avoiding system imbalances caused by single-objective optimization, thus achieving more comprehensive and robust system performance. The policy learning unit employs a neural network incorporating a long short-term memory (LSTM) network as a function approximator, enabling it to effectively process time-dependent grid operation data and capture long-term dynamic characteristics. Simultaneously, a dual-delay deep deterministic policy gradient algorithm is used to update network parameters, significantly improving learning stability and convergence efficiency, and effectively suppressing overestimation of Q-values ​​and policy oscillations. This allows the parameter self-tuning process to quickly adapt to changes in grid operating conditions and generate stable and reliable controller parameters, thus solving the problems of slow response and instability. Overall, through the refined design of the parameter self-tuning algorithm module, the system overcomes the limitations of traditional methods, such as incomplete state perception, singular objectives, and low learning efficiency. This allows the distributed energy storage collaborative frequency regulation system to quickly, accurately, and stably adaptively tune PID controller parameters when facing complex and changing grid operating conditions, effectively improving grid frequency stability and the operating efficiency of energy storage units, while also considering the communication performance of distributed collaboration and the overall robustness of the system.

[0062] The distributed collaborative communication module is configured to connect with the parameter self-tuning algorithm module to achieve real-time data synchronization and collaborative decision-making among ESS nodes. This module can utilize a hybrid network architecture combining fiber optic and 5G communication technologies. Fiber optic communication provides high-bandwidth, low-latency backbone transmission capabilities, while 5G communication offers flexible, wide-coverage access capabilities. In this hybrid network, a consensus protocol can be deployed to ensure that all participating ESS nodes can reach a consensus on shared data or decision results. For example, when an ESS node generates a new set of PID parameters, these parameters can be broadcast to other nodes through this communication module, verified by the consensus protocol, and ultimately adopted by all nodes.

[0063] Specifically, the distributed collaborative communication module adopts a clustered topology for networking, with each cluster containing one master node and N slave nodes; the master node uses an improved Raft consensus algorithm to verify the consistency of data within the cluster and ensures that the communication latency within the cluster is ≤10ms.

[0064] Specifically, the clustered topology is a network organization structure designed to divide a large number of distributed nodes into several logical or physical clusters. Nodes within each cluster cooperate closely, while inter-cluster communication occurs through specific mechanisms. This structure effectively reduces the complexity of inter-node communication in large-scale networks, improving communication efficiency and scalability. For example, adjacent energy storage units can be grouped into a cluster based on their geographical distribution, such as energy storage units within the same substation area forming a cluster; or nodes that are close together or have good communication quality can be dynamically grouped into a cluster based on the communication distance or signal strength between ESS nodes.

[0065] Each cluster consists of one master node and N slave nodes. This master-slave node architecture is a common distributed system design pattern. One node is designated as the master node, responsible for coordination, management, and decision-making, while the other nodes act as slave nodes, executing the tasks or instructions issued by the master node. This architecture aims to simplify cluster management, avoid multi-node decision-making conflicts, and improve decision-making efficiency and system stability. For example, during system deployment, a fixed master node can be pre-assigned to each cluster. This master node typically has stronger computing and communication capabilities; alternatively, cluster nodes can dynamically elect a master node using an election algorithm (such as Paxos, Bully, etc.), and a new election can be held when the master node fails.

[0066] The master node implements cluster data consistency verification using an improved Raft consensus algorithm. The improved Raft consensus algorithm is an optimization of the standard Raft algorithm to adapt to distributed consensus protocols in specific application scenarios (such as 5G / fiber hybrid networks). This algorithm ensures that all nodes in the distributed system maintain consistent data replicas through mechanisms such as leader election, log replication, and security guarantees. For example, real-time evaluation of communication link quality can be added to the Raft algorithm's heartbeat mechanism, dynamically adjusting the heartbeat interval or retransmission strategy based on link conditions to address potential latency and packet loss in 5G / fiber hybrid networks; or the master node can be allowed to submit multiple log entries to slave nodes in batches using an asynchronous replication mechanism, reducing the master node's waiting time and increasing throughput while ensuring eventual consistency, while simultaneously ensuring data integrity through additional verification mechanisms.

[0067] Meanwhile, ensuring intra-cluster communication latency is ≤10ms means designing and optimizing communication protocols, network architecture, and hardware configurations to ensure that the time required for data to travel from the sender to the receiver does not exceed a preset threshold. This aims to meet the high real-time requirements of distributed energy storage coordinated frequency regulation, ensuring the system can quickly respond to changes in grid frequency and avoid frequency regulation failures or performance degradation due to communication lags. For example, in a fiber-optic-5G hybrid communication network, a Quality of Service (QoS) policy can be configured to set high priority for key frequency regulation-related data streams, ensuring they receive priority transmission and lower latency even during network congestion; or some data processing and decision-making functions can be offloaded to intra-cluster edge nodes to reduce the distance and time for data transmission to the central server, thereby reducing end-to-end latency.

[0068] The above technical solution employs a clustered topology for networking, dividing large-scale distributed energy storage units into several management units, effectively reducing the overall network complexity and communication load. Each cluster has a master node and slave nodes; the master node is responsible for coordination and decision-making, while the slave nodes execute instructions. This master-slave architecture avoids decision-making conflicts between multiple nodes, simplifies the data synchronization process, and enhances the system's anti-interference capability. The master node uses an improved Raft consensus algorithm to verify the consistency of data within the cluster. This algorithm optimizes the handling of uncertainties such as latency and packet loss compared to traditional Raft, ensuring the consistency and reliability of data transmission between nodes. It is particularly suitable for 5G / fiber hybrid network environments, thus solving the problems of high communication latency, poor consistency, and weak anti-interference capability. Simultaneously, by ensuring that the intra-cluster communication latency is ≤10ms, the real-time performance of data synchronization is strictly controlled, enabling the distributed energy storage collaborative frequency regulation system to quickly respond to grid frequency changes, avoiding frequency regulation failures caused by communication delays, significantly improving collaborative efficiency and reliability, and thus ensuring the overall frequency regulation performance of the system.

[0069] The safety verification module is configured to connect with the distributed collaborative communication module to perform safety assessments on the parameter tuning results generated by collaborative decision-making. This module can include a series of preset safety rules and constraints, such as the upper and lower limits of the energy storage unit's state of charge (SOC), charging and discharging power limits, and grid frequency stability margins. Upon receiving the parameter tuning results generated by collaborative decision-making, the module performs real-time checks to determine whether these parameters will cause any ESS unit or the entire system to exceed its safe operating range. Once a potential over-limit risk is detected, the module can immediately issue a warning signal and dynamically adjust or limit the parameters according to preset strategies to ensure that all ESS units always operate within the preset safety domain, avoiding equipment damage or system instability.

[0070] Example 2 This embodiment provides a method for rapid adaptive tuning of distributed energy storage collaborative frequency regulation parameters, using the distributed energy storage collaborative frequency regulation parameter rapid adaptive tuning system described in Embodiment 1, which includes the following steps: S1. Through the data acquisition module, the grid frequency deviation, the SOC status of the distributed energy storage unit, the communication data of adjacent nodes, and the historical frequency regulation response data are collected in real time; This step aims to provide real-time and accurate input information for subsequent adaptive tuning of frequency regulation parameters. The data acquisition module can employ a distributed sensor network, deploying sensors at key nodes in the power grid and at each ESS unit to monitor the grid frequency deviation (Δf) and the operating parameters of each ESS unit, such as state of charge (SOC), power output, voltage, and current, in real time. These sensors can transmit data to a local controller or central data processing unit via wired or wireless means. Alternatively, existing power system SCADA (Supervisory Control and Data Acquisition) systems or PMUs (Phasor Measurement Units) can be used to acquire grid frequency deviation data, and the ESS's operating status data can be obtained through its own Battery Management System (BMS) or Energy Management System (EMS). The data acquisition frequency can be set according to system response requirements, for example, once every millisecond or every tens of milliseconds, to ensure real-time data accuracy.

[0071] S2. Input the data into the multi-timescale hierarchical control module to generate preliminary hierarchical control instructions for different time scales; This step involves generating preliminary response commands for frequency disturbances at different time scales based on real-time acquired data and a hierarchical control strategy, thereby achieving coordinated frequency modulation across multiple time scales. The multi-time-scale hierarchical control module receives frequency deviation and ESS status data acquired by S1 and sends them to the millisecond-level fast response layer, the second-level dynamic adjustment layer, and the minute-level global optimization layer, respectively. The millisecond-level layer rapidly generates suppression commands based on frequency mutations, the second-level layer generates compensation commands based on frequency fluctuation trends, and the minute-level layer generates optimization commands based on the overall network SOC balance. These commands are preliminary and will subsequently undergo parameter self-tuning and security verification. Alternatively, the module integrates a data preprocessing unit to filter and normalize the acquired raw data, and then calculates preliminary power regulation commands independently or collaboratively at each time scale layer according to preset control logic or algorithms. For example, the millisecond-level layer might focus on rapid power injection / absorption to suppress instantaneous frequency drops / rises, while the minute-level layer might focus on adjusting the ESS's charging and discharging strategy to maintain system frequency stability and SOC balance.

[0072] S3. Call the parameter self-tuning algorithm module to dynamically optimize the PID controller parameters based on the current system state and the initial hierarchical control instructions; This step is the core of this method, aiming to solve the problem that traditional PID parameters fixed or offline tuning cannot adapt to complex and changing operating conditions. Through online learning and optimization, the PID controller parameters are dynamically adjusted to improve the adaptability and robustness of the frequency regulation response. The parameter self-tuning algorithm module can receive system states collected by S1 (such as grid frequency deviation Δf, df / dt, SOC, and adjacent node parameter deviations) and preliminary control commands generated by S2 as inputs. Based on an improved deep reinforcement learning framework, this module maps the current system state to a reinforcement learning state space, using the preliminary commands as a reference for the action space. Through interaction with the environment (simulated or actual system response), it learns and updates the optimal PID parameters (Kp, Ki, Kd). Alternatively, this module includes a policy learning unit that uses a neural network containing a long short-term memory network as a function approximator and employs a double-delay deep deterministic policy gradient algorithm to update the network parameters. By defining a reward function, the algorithm evaluates the impact of the current PID parameters on the system frequency stability and ESS operating state in each iteration and adjusts the weights of the neural network according to the reward signal, thereby outputting the optimal PID parameters adapted to the current operating conditions.

[0073] S4. Through the distributed collaborative communication module, the optimized PID controller parameters are distributed and collaboratively decided and synchronized based on the consensus protocol; This step aims to ensure that in a distributed energy storage system, each ESS unit can efficiently and reliably share and synchronize optimized PID parameters, achieving global coordinated frequency regulation and avoiding local conflicts or overall performance degradation caused by inconsistencies in information. The distributed coordinated communication module can employ a fiber-5G hybrid communication network based on a consensus protocol to broadcast or transmit the S3-optimized PID parameters between ESS nodes. Through the consensus protocol, each node verifies and confirms the received parameters, ensuring that all ESS units participating in frequency regulation use a unified and up-to-date set of PID parameters. Alternatively, this module can utilize a clustered topology for networking, with each cluster containing one master node and N slave nodes. The master node is responsible for collecting the optimized parameters from the slave nodes within the cluster and achieving parameter consistency within the cluster through a consensus algorithm. Then, the master nodes of each cluster synchronize parameters through a higher-level consensus protocol, ultimately achieving parameter coordinated decision-making and synchronization across all ESS units in the network. Communication delays and data packet loss issues can be handled through the protocol's fault tolerance and retransmission mechanisms.

[0074] S5. The security verification module is used to perform security verification on the final parameters after collaborative decision-making, and the security control command that passes the verification is output to each distributed energy storage unit for execution; after outputting the security control command to each distributed energy storage unit, the experience data consisting of the current control cycle's state, action, reward and the next state is stored in the improved deep reinforcement learning (DRL) experience replay pool for online learning, and then the next control cycle is entered; if the security verification fails, parameter correction or warning is performed.

[0075] This step is crucial for ensuring the safe operation of the system, guaranteeing that the optimized and coordinated PID parameters will not cause ESS unit overload, overcharge / over-discharge, or exceed the grid's safe operating boundaries in practical applications. After receiving the final PID parameters from the S4 coordinated decision, the safety verification module evaluates these parameters based on preset safety domains (such as the ESS's upper and lower SOC limits, maximum charging and discharging power, and grid frequency stability range). If a parameter might cause any ESS unit or the entire system to exceed the safety domain, the module will trigger an over-limit warning and dynamically adjust or limit these parameters, for example, by reducing the PID gain or adjusting the SOC reference value, until the parameters meet safety constraints. Alternatively, the safety verification module can incorporate a rule-based or model-based predictive safety assessment algorithm. This algorithm simulates the ESS unit's operating state and grid frequency response under current parameters to predict potential safety risks. Once a risk is detected, the module immediately initiates a dynamic constraint mechanism, such as truncating or saturating the PID parameters, or providing adjustment suggestions to the multi-timescale hierarchical control module, until a final control command that meets safety requirements is generated. These commands are then sent to each ESS unit to guide its actual frequency regulation operations.

[0076] Through the above technical solution, this invention acquires real-time data on grid frequency deviation and the operating status of each ESS (Electrical Controller System) via a data acquisition module, ensuring the real-time nature and comprehensiveness of the input information. This provides an accurate foundation for subsequent processing and avoids tuning deviations caused by data delays or incompleteness. Subsequently, the acquired data is input to a multi-time-scale hierarchical control module to generate preliminary control commands for different time scales. This step generates commands based on a hierarchical control mechanism (millisecond, second, and minute levels), achieving coordinated response across multiple time scales and solving the problem that single-time-scale control cannot simultaneously address instantaneous impacts and continuous power imbalances. Furthermore, the parameter self-tuning algorithm module is invoked to dynamically optimize the PID controller parameters based on the current system state and preliminary commands. This feature utilizes the current state and preliminary commands as input, combined with an improved deep reinforcement learning framework for online optimization, ensuring the real-time nature and context-dependent nature of parameter adaptive adjustment and avoiding oscillations or failures caused by reliance on human experience. Simultaneously, through a distributed collaborative communication module, the optimized parameters are distributed for collaborative decision-making and synchronization based on a consensus protocol. This step, leveraging the consensus protocol, achieves real-time data synchronization between nodes, improving the efficiency and reliability of collaborative decision-making and reducing the impact of communication delays and packet loss. Finally, the final parameters after collaborative decision-making are verified using a safety verification module, and control commands that pass verification are output to each ESS unit for execution. This feature, based on a safety constraint mechanism, performs dynamic verification to ensure that all operations are within a preset safety domain, preventing risks such as overcharging / over-discharging and enhancing system security. Overall, this method, by defining ordered execution steps and coordinating the collaborative operation of various system modules, achieves an efficient, reliable, and safe adaptive tuning process for frequency modulation parameters, effectively solving problems such as low parameter tuning efficiency, high collaborative decision-making delay, and insufficient safety verification.

[0077] Using the distributed energy storage collaborative frequency regulation parameter fast adaptive tuning system described in Example 1 or the distributed energy storage collaborative frequency regulation parameter fast adaptive tuning method described in Example 2, taking a regional power grid with 30 or more ESS nodes as an example: 1) Frequency change scenario (Δf=0.5Hz): Parameter tuning time is reduced from 120ms in the traditional method to 45ms; 2) Improved SOC balance: The maximum SOC difference decreased from 23% to 7.8%; 3) Communication reliability: Packet loss rate decreased from 0.15% to 0.02%.

[0078] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system, characterized in that, include: Data acquisition module: used to acquire grid frequency deviation, SOC status of distributed energy storage units, communication data between adjacent nodes, and historical frequency regulation response data in real time; Multi-timescale hierarchical control module: Generates preliminary control commands based on the data collected by the data acquisition module. This module includes a millisecond-level fast response layer, a second-level dynamic adjustment layer, and a minute-level global optimization layer. The layers are coupled together through dynamic weight factors. The parameter self-tuning algorithm module is associated with the initial control commands generated by the millisecond-level fast response layer, the second-level dynamic adjustment layer, and the minute-level global optimization layer. Based on the improved deep reinforcement learning (DRL) framework, it dynamically generates proportional-integral-derivative controller parameters through a hybrid algorithm that integrates Q-learning with a neural network containing a long short-term memory network. Distributed collaborative communication module: associated with the parameters of the proportional-integral-derivative controller, it adopts a fiber-optic-5G hybrid communication network based on a consensus protocol and is networked through a clustered topology to realize real-time data synchronization and collaborative decision-making among nodes of the distributed energy storage system; Safety verification module: It is associated with the collaborative decision parameter tuning results output by the distributed collaborative communication module, and is used to provide over-limit warnings and dynamic constraints on the parameter tuning results generated by collaborative decision-making, so as to ensure that all distributed energy storage units operate within the preset safety domain.

2. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 1, characterized in that, The multi-timescale hierarchical control module specifically includes: 1) Millisecond-level fast response layer: adopts fuzzy PID control algorithm, with a response time ≤50ms, and prioritizes the suppression of frequency mutations; 2) Second-level dynamic adjustment layer: The frequency regulation responsibility coefficient of each distributed energy storage unit is calculated through dynamic sensitivity analysis and the droop coefficient is adjusted to compensate for the fluctuation of renewable energy output; 3) Minute-level global optimization layer: Based on model predictive control, optimize the SOC balance of the entire network, with an optimization cycle of 1-5 minutes.

3. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 2, characterized in that, The millisecond-level fast response layer is deployed on the local controller and uses hardware-in-the-loop technology to achieve fuzzy PID fast frequency modulation with a control cycle of 15-25ms.

4. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 2, characterized in that, The aforementioned second-level dynamic adjustment layer calculates the frequency regulation responsibility coefficient of each distributed energy storage unit in real time based on the dynamic sensitivity matrix.

5. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 2, characterized in that, The minute-level global optimization layer establishes a global optimization model based on model predictive control, and uses a distributed alternating direction multiplier method algorithm to solve the global optimization model in a distributed manner to optimize the SOC balance of the entire network.

6. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 1, characterized in that, The parameter self-tuning algorithm module includes: 1) State-space definition unit: The input vectors are the grid frequency deviation Δf, df / dt, the SOC state of each distributed energy storage unit and the parameter deviation of adjacent nodes. df / dt represents the rate of change of grid frequency with time. 2) Reward function design unit: R=α·(1 / |Δf|)+β·(SOC balance)+γ·(communication delay penalty term), where R represents the reward value, and α, β and γ are all weight coefficients; 3) Policy learning unit: A neural network containing a long short-term memory network is used as a function approximator, and the network parameters are updated using a dual-delay deep deterministic policy gradient algorithm.

7. The distributed energy storage coordinated frequency regulation parameter fast adaptive tuning system according to claim 1, characterized in that, The distributed collaborative communication module adopts a clustered topology for networking, with each cluster containing one master node and N slave nodes; the master node uses the improved Raft consensus algorithm to verify the consistency of data within the cluster and ensures that the communication latency within the cluster is ≤10ms.

8. A method for rapid adaptive tuning of distributed energy storage coordinated frequency regulation parameters, employing the rapid adaptive tuning system for distributed energy storage coordinated frequency regulation parameters as described in any one of claims 1-7, characterized in that, Including the following steps: S1. Through the data acquisition module, the grid frequency deviation, the SOC status of the distributed energy storage unit, the communication data of adjacent nodes, and the historical frequency regulation response data are collected in real time; S2. Input the data into the multi-timescale hierarchical control module to generate preliminary hierarchical control instructions for different time scales; S3. Call the parameter self-tuning algorithm module to dynamically optimize the PID controller parameters based on the current system state and the initial hierarchical control instructions; S4. Through the distributed collaborative communication module, the optimized PID controller parameters are distributed and collaboratively decided and synchronized based on the consensus protocol; S5. The security verification module is used to perform security verification on the final parameters after collaborative decision-making, and the security control command that passes the verification is output to each distributed energy storage unit for execution.

9. The method for rapid adaptive tuning of distributed energy storage coordinated frequency regulation parameters according to claim 8, characterized in that, In step S5, after outputting safety control commands to each distributed energy storage unit, the experience data consisting of the current control cycle's state, actions, rewards, and the next state is stored in the improved deep reinforcement learning (DRL) experience replay pool for online learning, and then the next control cycle begins.

10. The method for rapid adaptive tuning of distributed energy storage coordinated frequency regulation parameters according to claim 8, characterized in that, In step S5, if the security check fails, parameter correction or warning will be issued.