Distributed new energy and energy storage primary frequency modulation multi-agent collaborative control method and system

By employing a multi-agent cooperative control method, the problems of frequency regulation response lag and security risks in distributed renewable energy systems have been solved, achieving efficient and safe frequency regulation cooperative control and improving the frequency stability of the power system and the utilization efficiency of renewable energy.

CN121332573BActive Publication Date: 2026-07-03POWER RES INST OF STATE GRID SHAANXI ELECTRIC POWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
POWER RES INST OF STATE GRID SHAANXI ELECTRIC POWER CO LTD
Filing Date
2025-12-09
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

With the large-scale grid connection of distributed new energy sources such as wind power and photovoltaics, the proportion of traditional synchronous generator sets has gradually decreased, resulting in a significant decline in the inertial support and frequency regulation capability of the power system. Existing control methods are difficult to adapt to complex and ever-changing operating conditions and the needs of multi-device coordination, and there are problems such as high communication delay and insufficient utilization of local information, leading to risks such as delayed frequency regulation response, voltage exceeding limits, or overcharging and over-discharging of energy storage.

Method used

A multi-agent cooperative control method is adopted. By establishing a multi-agent power grid simulation environment, an independent agent is assigned to each distributed energy and energy storage unit. The MADDPG algorithm framework is constructed, a comprehensive reward function is designed to guide the agents to optimize action decisions, and efficient frequency regulation coordination between distributed new energy and energy storage is achieved through reinforcement learning.

Benefits of technology

It achieves multi-objective dynamic priority frequency regulation, taking into account both rapid response and global security, improves the environmental perception and collaborative decision-making accuracy of intelligent agents, constructs a closed loop for security and efficiency assurance throughout the entire "training-application" cycle, provides a quantitative and multi-dimensional collaborative performance evaluation system, and significantly improves the frequency stability of the power system and the utilization efficiency of new energy sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121332573B_ABST
    Figure CN121332573B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent collaborative control method and system for primary frequency regulation of distributed renewable energy and energy storage, belonging to the field of power system frequency regulation technology. The method includes: constructing a multi-agent distribution network simulation environment based on grid topology and integrating wind power / photovoltaic / energy storage models and constraints; assigning independent agents to each unit, defining an observation space and output adjustment action space containing information such as frequency deviation; building a MADDPG framework containing a shared Critic and independent Actor networks, and achieving collaborative training through experience replay; designing a comprehensive reward function that integrates frequency, voltage, energy storage SOC, and curtailment rate; and using the trained model to output real-time output commands to achieve primary frequency regulation collaborative control. This invention can significantly improve the response speed and accuracy of distributed energy participating in primary frequency regulation, ensure system safety, reduce renewable energy curtailment rate, and adapt to the grid security requirements of high-penetration renewable energy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system frequency regulation control technology, specifically to a multi-agent collaborative control method and system for primary frequency regulation of distributed new energy and energy storage. Background Technology

[0002] With the large-scale grid connection of distributed renewable energy sources such as wind power and photovoltaics, the proportion of traditional synchronous generator units is gradually decreasing, leading to a significant decline in the inertial support and frequency regulation capabilities of the power system. The intermittency and volatility of distributed renewable energy further exacerbate system power imbalances and trigger frequency fluctuation risks.

[0003] Existing frequency regulation methods for new energy sources mostly employ fixed-parameter control strategies (such as virtual inertial control and droop control), relying on experience to tune parameters, making them difficult to adapt to complex and ever-changing operating conditions and the collaborative needs of multiple devices. Furthermore, traditional centralized control methods struggle to accommodate the dispersed, small, and numerous topological characteristics of distributed power sources, resulting in high communication delays and insufficient utilization of local information, easily leading to risks such as delayed frequency regulation response, voltage exceeding limits, or overcharging and over-discharging of energy storage. Simultaneously, existing single-agent control schemes cannot achieve multi-unit collaborative optimization, making it difficult to balance the contradiction between system frequency stability and the operational constraints of each unit, necessitating a breakthrough in efficient distributed collaborative control technology. Moreover, the widespread distribution and large number of distributed energy and energy storage units present challenges such as high communication pressure and slow response speeds, while traditional decentralized control lacks global collaborative optimization capabilities, easily leading to poor frequency regulation performance or equipment operational conflicts. Therefore, an intelligent control method capable of achieving distributed decision-making and global collaboration is urgently needed to improve the primary frequency regulation performance of high-proportion new energy power systems. Summary of the Invention

[0004] The purpose of this invention is to provide a multi-agent collaborative control method and system for primary frequency regulation of distributed new energy and energy storage. Through multi-agent collaborative decision-making and reinforcement learning adaptive optimization, it achieves efficient frequency regulation collaboration between distributed new energy and energy storage, thereby improving the frequency stability of the power system and the utilization efficiency of new energy.

[0005] To address the aforementioned problems, this invention provides a method and system for multi-agent collaborative control of primary frequency regulation in distributed new energy and energy storage systems. The specific technical solution is as follows:

[0006] A multi-agent power grid simulation environment is established, which includes distributed new energy sources and energy storage systems. The simulation environment is based on the power grid topology and integrates equipment models and operational constraints of wind power, photovoltaics and energy storage.

[0007] Each distributed energy source and energy storage unit is assigned an independent intelligent agent, and the local observation space and action space of each intelligent agent are defined. The observation space includes the system frequency deviation, its own operating status and local power grid information, and the action space corresponds to the power output adjustment amount.

[0008] A multi-agent reinforcement learning framework is constructed, and the MADDPG algorithm is used to realize the collaborative training of agents. The framework includes a shared Critic network and an independent Actor network, and stores and samples interaction data through an experience replay mechanism.

[0009] Design a comprehensive reward function that integrates frequency deviation, voltage limit exceedance, energy storage SOC constraints, and renewable energy curtailment rate to guide the agent in optimizing action decisions;

[0010] The trained multi-agent reinforcement learning model outputs real-time power output adjustment commands for each distributed energy source and energy storage unit, thereby achieving primary frequency regulation and coordinated control.

[0011] The comprehensive reward function introduces an adaptive weight coefficient adjustment mechanism based on dynamic scheduling priority to respond to the differentiated requirements for frequency regulation speed and adjustment accuracy under different operating conditions. Specifically, when the absolute value of the system frequency deviation is greater than the preset emergency threshold, the coefficient of the frequency deviation weight term is automatically increased, while the penalty intensity on the energy storage SOC is dynamically reduced to prioritize rapid frequency recovery.

[0012] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent collaborative control method is characterized in that the distributed new energy equipment model includes output characteristic models of wind power and photovoltaics, and the energy storage system model includes charging and discharging power constraints, SOC (State of Charge) upper and lower limits constraints, and energy conversion efficiency; the power grid simulation environment supports forward-backward substitution power flow calculation and can update node voltage and system power balance state in real time; the output characteristic models of wind power and photovoltaics not only include their maximum available power, but also integrate the predicted power fluctuation range based on ultra-short-term prediction confidence intervals; the SOC constraints of the energy storage system model are dynamically divided into a high-response priority zone, an economic optimization zone, and a prohibited action zone; and the forward-backward substitution power flow calculation and the power-frequency dynamic equations of the simulation environment are synchronously iteratively decoupled and solved on a sub-second time scale.

[0013] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method is characterized in that the local observation space of the agent is defined as:

[0014] ;

[0015] in, For power system frequency deviation, This refers to the operating status of the device corresponding to the intelligent agent. This represents the average operating status of adjacent distributed energy sources. The observation space is a local node voltage sequence; an adaptive sensing mechanism is embedded within the observation space, wherein the range of adjacent distributed energy sources is dynamically determined based on the real-time topological connectivity and electrical distance of the power grid, and the self-operating state... For wind / solar power units, this includes the deviation between their ultra-short-term predicted power and their current actual output; for energy storage units, it includes their... The current dynamic partition status and estimated remaining adjustable time.

[0016] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent collaborative control method is characterized in that the action space of the agent is a continuous space, and the action value is mapped to the equipment output adjustment amount:

[0017] ;

[0018] in, The standardized action output by the agent (range [-1,1]). Rated power of the equipment The output adjustment scaling factor is used as the actual adjustment command after the action is clipped by the equipment constraints; the output adjustment scaling factor Instead of being a fixed constant, these are time-varying parameters dynamically generated based on a deep belief network. The input to this deep belief network is the local observation space vector of the corresponding agent, and the output is a parameter adapted to the current urgency of the system and the state of the device itself. The value is calculated and jointly optimized through policy gradients during training.

[0019] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method is characterized in that the system frequency deviation is dynamically updated through the power balance equation:

[0020] ;

[0021] in, This represents the frequency deviation from the previous moment. For the total output of distributed new energy and energy storage, This represents the total system load (including random disturbances). The system damping coefficient is... Let be the system's inertial time constant. This is for simulating step size;

[0022] In the power balance equation, the system inertial time constant With damping coefficient The parameters are modeled as dynamic parameters that change in real time with the penetration rate of distributed renewable energy, and their dynamic relationships are obtained through fitting offline simulation data. Simultaneously, a disturbance prediction and correction module based on a feedforward neural network is embedded in the frequency deviation update loop. This module takes the short-term variation trends of load and renewable energy output as input and adjusts accordingly. The item provides advance compensation to simulate and achieve predictive primary frequency modulation actions.

[0023] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method is characterized in that the comprehensive reward function is expressed as:

[0024] ;

[0025] in, , , , These are frequency deviation, voltage exceeding limit, The weighting coefficients for constraints and curtailment rates, This represents the cumulative value of node voltage exceeding the limit. For energy storage Penalties for deviations from the reasonable range The curtailment rate of renewable energy; the energy storage Penalty items The calculation is performed using a piecewise nonlinear function: when When in the high-response priority zone, the penalty is zero to encourage action; when When in the economically optimal zone, the penalty term is a mild quadratic function; when As the penalty term approaches the boundary of the prohibited action zone, it becomes a steep exponential function to impose a strong constraint. The frequency deviation weight... Based on the interval to which its absolute value belongs, under the control of the adaptive weighting coefficient adjustment mechanism, The range is continuously varying, where k>1 is the amplification factor for emergency conditions.

[0026] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method is characterized in that the network update process of the MADDPG algorithm includes:

[0027] The Critic network receives the joint state and joint actions of all agents, outputs a global Q-value, and updates the network parameters by minimizing the Q-value prediction error. Each Actor network outputs action values ​​based on its own local observations, and updates its parameters by maximizing the Q-value evaluated by the Critic network using the policy gradient method. A soft update mechanism is used to update the parameters of the target Critic network and the target Actor network to ensure training stability. The training process of the MADDPG algorithm is divided into two stages: the first stage is offline centralized pre-training based on historical typical working condition data to initialize the policy network parameters of all agents; the second stage is online collaborative training based on real-time simulation data, and a policy distillation mechanism is introduced in this stage, that is, training a lightweight global policy evaluator, periodically extracting and fusing effective experiences from each independent Actor network to form a rapidly deployable suboptimal collaborative policy as a backup.

[0028] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent collaborative control method is characterized by further including setting primary frequency regulation performance evaluation indicators to quantify the collaborative control effect. The evaluation indicators include the average frequency deviation, the maximum frequency deviation, the number of voltage over-limit times, and the new energy curtailment rate. The evaluation indicators further include: a frequency regulation contribution fairness index, used to quantify the matching degree between the actual contribution and capacity ratio of distributed energy and energy storage at different regulation rates during the frequency regulation process; and a control command smoothness index, used to evaluate the rate of change of output commands at adjacent times to ensure equipment safety and lifespan.

[0029] Furthermore, the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method is characterized in that the average frequency deviation evaluation index is expressed as:

[0030] ;

[0031] in, The mean of frequency deviation, To assess the time window, for The system frequency deviation at a given time; the calculation method for the frequency modulation contribution fairness index SF is as follows: Where N is the total number of units participating in frequency modulation. For the first The ratio of the actual frequency modulation energy contribution of each unit to its rated adjustable capacity. This is the average of the ratios across all units. The closer this index is to 1, the better the fairness of the coordinated control.

[0032] Furthermore, the distributed multi-agent cooperative control system for primary frequency regulation of new energy and energy storage based on multi-agent reinforcement learning includes the following five modules:

[0033] Environment construction module: Establishes a multi-agent power grid simulation environment including distributed new energy sources and energy storage systems, integrating equipment models, power flow calculations, and operational constraints;

[0034] Intelligent agent configuration module: Assigns intelligent agents to each distributed energy source and energy storage unit, and defines the observation space and action space;

[0035] Reinforcement learning framework module: Constructs the MADDPG algorithm framework, including an experience replay buffer, a shared Critic network, and independent Actor networks;

[0036] Reward function design module: Constructs a comprehensive reward function that integrates multi-dimensional indicators to guide the agent in optimizing decision-making;

[0037] The collaborative control module loads the trained model and outputs real-time power output adjustment commands for each device to achieve primary frequency modulation collaborative control. The performance evaluation module quantifies the collaborative control effect through preset evaluation indicators. The collaborative control system also includes an independent safety verification and command correction module, which is located between the collaborative control module and the physical devices. After receiving the power output command from the collaborative control module, it performs a millisecond-level rapid verification of the command based on a set of simplified and deterministic safety rules. If a risk that may cause the device to exceed its limits or the local voltage to momentarily exceed its limits is detected, the command is slightly conservatively corrected using the safety rules within that time step, and the correction record is fed back to the experience playback buffer for subsequent training.

[0038] This invention achieves primary frequency regulation coordinated control of distributed new energy and energy storage through multi-agent reinforcement learning, and has the following beneficial technical effects compared with the prior art:

[0039] This system achieves multi-objective dynamic priority frequency regulation, balancing rapid response and global safety. By introducing an adaptive weight adjustment mechanism based on dynamic scheduling priorities, the system can automatically focus on rapid recovery in frequency emergency situations, and smoothly transition to economic optimization and equipment protection in normal situations. Combined with the dynamic zoning and piecewise nonlinear penalty function of the energy storage SOC, it fundamentally avoids the inherent contradiction between frequency regulation requirements and energy storage safety in traditional methods. While ensuring primary frequency regulation speed, it significantly extends the lifespan of energy storage equipment and maintains the system's long-term regulation capability.

[0040] The accuracy of environmental perception and collaborative decision-making by intelligent agents has been improved: By introducing embedded adaptive perception mechanisms (such as dynamic neighbor range and device status with prediction bias) into the observation space, and dynamic action scaling coefficients based on deep belief networks, each agent possesses an "intuition" of condition adaptation, enabling it to more accurately assess the value of its own actions. In particular, by modeling the system inertia H and damping D as dynamic parameters that change with the penetration rate of new energy sources, the simulation environment and decision-making model can realistically reflect the vulnerability of power grids with a high proportion of new energy sources, thereby training more grid-adaptive strategies.

[0041] A closed loop ensuring safety and efficiency throughout the entire "training-application" lifecycle has been constructed. On the one hand, through a two-stage training and strategy distillation mechanism, offline training efficiency is improved while generating a lightweight backup strategy that can be quickly deployed, enhancing the system's practicality and reliability. On the other hand, an innovative independent safety verification and instruction correction module has been established. This module acts as a "security firewall" between AI decision-making and physical devices, capable of intercepting and correcting potentially risky instructions in milliseconds and feeding the correction experience back to the training loop. This forms a virtuous cycle of "AI optimizing decision-making, rules ensuring safety, and safety experience feeding back into AI," representing a significant breakthrough in the safe and reliable application of cutting-edge artificial intelligence algorithms in the critical field of power control.

[0042] A quantitative, multi-dimensional collaborative performance evaluation system is provided: by introducing the frequency modulation contribution fairness index and control command smoothness index, the evaluation dimensions are expanded from a single system frequency index to a comprehensive multi-dimensional evaluation including the rationality of equipment collaboration and the smoothness of control behavior. This not only provides more refined feedback for the optimization algorithm, but also provides measurable and comparable objective evidence for the technical effects of this invention (such as a fairness index of 0.93 or higher), strongly demonstrating the substantial effect of collaborative optimization.

[0043] In summary, this invention is not a simple application of multi-agent reinforcement learning, but a systematic engineering innovation for the specific scenario of distributed renewable energy primary frequency regulation. Through a series of interconnected and progressively layered technical features, it effectively solves the intertwined problems of response lag, security risks, and coordination mismatch, achieving a comprehensive leap in frequency regulation performance, equipment safety, and system adaptability. Attached Figure Description

[0044] Figure 1 The flowchart illustrates the multi-agent collaborative control method for primary frequency regulation of distributed new energy and energy storage provided in this embodiment of the invention.

[0045] Figure 2 This is a block diagram of a distributed new energy and energy storage primary frequency regulation multi-agent collaborative control system provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0047] Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0048] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0049] The invention will now be described in more detail with reference to the accompanying drawings. In the various drawings, the same elements are indicated by similar reference numerals. For clarity, the various parts in the drawings are not drawn to scale.

[0050] Example 1:

[0051] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Environment Setup: Based on the power grid topology, a base capacity of 100MVA and a rated voltage of 12.66kV are set. Three photovoltaic units (rated power 1MW / unit), two wind power units (rated power 1.5MW / unit), and three energy storage units (rated power 0.5MW / unit, rated capacity 2MWh / unit) are configured, distributed across different nodes. Typical wind and solar power output curve data are loaded to simulate the intermittency of renewable energy; random load disturbances are set to simulate power fluctuations in actual operation.

[0053] Intelligent agent configuration: An intelligent agent is assigned to each photovoltaic, wind power, and energy storage unit, for a total of 8 intelligent agents. The observation space is defined as 12-dimensional (including 1-dimensional frequency deviation, 2-dimensional self-state, 4-dimensional average state of adjacent devices, and 5-dimensional local node voltage); the action space is a 1-dimensional continuous space, and the output adjustment scaling factor α is set to 0.2.

[0054] Reinforcement learning framework training: The network parameters of the MADDPG algorithm are set as follows: hidden layer dimension 256, Actor learning rate 1e-4, Critic learning rate 1e-3, discount factor γ=0.95, soft update coefficient τ=0.01, experience replay buffer capacity 100000, batch sampling size 256. The total number of training episodes is 1000, and the maximum number of steps per episode is 200.

[0055] Reward function parameter settings: frequency deviation weight w_f=10.0, voltage over-limit weight w_v=5.0, SOC constraint weight w_s=2.0, curtailment rate weight w_c=1.0, action smoothing penalty weight w_a=0.1.

[0056] In a preferred embodiment of the present invention, the technical solution has been deeply optimized to further enhance the intelligence, security, and engineering practicality of the collaborative control. First, a dynamic, high-fidelity simulation environment is constructed: the inertial time constant H and damping coefficient D of the power grid are modeled as dynamic parameters that change in real time with the penetration rate by fitting simulation data under different new energy penetration rates; power flow calculation and frequency dynamic equations are solved using sub-second synchronous iterative decoupling to ensure the realism of the dynamic process simulation. In the agent design, the observation space is endowed with adaptive perception capabilities: the neighbor range is dynamically determined based on real-time electrical distance; the state of wind and solar power units includes their ultra-short-term prediction bias; and the state of energy storage units clearly defines the "high-response priority zone," "economic optimization zone," or "prohibited action zone" to which their SOC belongs. The scaling factor α of the action space is not a fixed value but is dynamically generated by a deep belief network with local observations as input and jointly optimized with the Actor policy network, enabling the agent to possess adaptive action sensitivity under operating conditions.

[0057] The core reward function employs an adaptive design with dynamic scheduling capabilities. The frequency deviation weight ω_f is designed to smoothly vary between a base value ω_f_base and an amplified value k*ω_f_base as the absolute value of the frequency deviation changes. When the frequency deviation exceeds the emergency threshold of 0.25Hz, the weight automatically increases to prioritize frequency recovery. The energy storage SOC penalty term SOC_pen uses a piecewise nonlinear function: zero penalty in the high response priority zone, a quadratic function in the economic optimization zone, and an exponential function for strong constraints near the boundary, thereby precisely guiding the safe and efficient use of energy storage.

[0058] The training process employs a two-stage progressive strategy: the first stage utilizes historical data for offline pre-training of the MADDPG network, laying the foundation for the basic strategy; the second stage involves online fine-tuning training and introduces policy distillation technology, periodically fusing knowledge from each independent Actor network to generate a lightweight global evaluator as a high-performance backup strategy. At the control system level, a safety verification and instruction correction module has been added. This module, located between the agent and the device, performs millisecond-level verification and conservative correction of the output instructions at each moment based on a set of defined fast safety rules (such as instantaneous power limits, SOC hard boundaries, etc.), and feeds the corrected examples back to the training buffer, thereby continuously improving the reliability of the AI ​​strategy while ensuring real-time safety.

[0059] In addition to traditional frequency and voltage indicators, the performance evaluation system adds a "frequency modulation contribution fairness index" and a "control command smoothness index." The fairness index quantifies the rationality of coordination by calculating the complement of the standardized variance of the actual frequency modulation contribution of each unit and its own capacity, thus avoiding over-adjustment of individual devices. Ultimately, in test scenarios containing random disturbances and fluctuations, the optimized scheme not only further reduced the average frequency deviation by approximately 15%, but also stabilized the maximum frequency deviation within ±0.15Hz, achieving a fairness index of over 0.92, and improving the smoothness of control commands by 30%, fully verifying the comprehensive advantages of this invention in dynamic coordination, safety constraints, and multi-objective optimization.

[0060] To quantitatively verify the comprehensive benefits of the collaborative enhancement mechanisms proposed in this invention, such as adaptive weight adjustment, dynamic perception, and security closed loop, multiple sets of comparative simulation experiments were designed and executed.

[0061] 1. Experimental setup

[0062] Scenario: In the standard IEEE 33-node distribution network model, the same random load disturbance and wind and solar fluctuation scenarios are set up, and a fault with an instantaneous power deficit of 5% of the total load is introduced to simulate the stringent frequency regulation requirements.

[0063] Comparison of options:

[0064] Option A (Traditional Method): Employs a droop control strategy with fixed parameters.

[0065] Option B (Baseline Intelligent Agent Method): The basic MADDPG framework is adopted without incorporating the optimization features of this invention (such as dynamic weights, adaptive perception, security verification, etc.).

[0066] Solution C (the present invention): The complete optimized solution proposed in the present invention is adopted, including all the features described in the present invention.

[0067] 2. The comparison between the experimental results and the key performance indicators is shown in the table below:

[0068]

[0069] 3. Experimental Conclusions The experimental data fully demonstrate that, compared to traditional methods and basic intelligent agent methods, the proposed solution exhibits comprehensive and significant superiority in multiple dimensions, including frequency modulation speed, steady-state accuracy, voltage safety, equipment collaborative fairness, and control command smoothness. This is not the result of a single technological improvement, but rather the systematic effect of the "adaptive perception-dynamic decision-making-safe execution" collaborative enhancement closed loop constructed by the present invention, verifying the non-obviousness and synergistic gain of the combination of various technical features in the proposed solution.

[0070] Simulation verification of coordinated control: Under scenarios of random load disturbance and fluctuations in renewable energy output, the frequency regulation effect of the method of this invention is compared with that of the traditional droop control method. The results show that the average frequency deviation of the method of this invention is reduced by more than 35%, the maximum frequency deviation is controlled within ±0.2Hz, the number of voltage over-limit incidents is reduced by 60%, and the renewable energy curtailment rate is reduced by 20%, verifying the effectiveness and superiority of coordinated control.

[0071] In summary, the embodiments of this application provide a collaborative control method and system for new energy generating units and power plants participating in primary frequency regulation. The core is to construct a distribution network simulation environment that integrates wind power, photovoltaic, and energy storage models and constraints, assign independent intelligent agents to each unit, and define an observation space and output adjustment action space containing frequency deviation. Collaborative training is completed based on the MADDPG framework containing shared Critic and independent Actor networks and a comprehensive reward function that integrates multiple objectives. Finally, real-time output commands are output to achieve frequency regulation control. This can improve the frequency regulation response speed and accuracy, ensure system safety, reduce the curtailment rate of new energy, adapt to the needs of high-penetration new energy power grids, and effectively solve problems such as poor coordination and unreasonable power allocation when new energy participates in primary frequency regulation.

[0072] The multi-agent collaborative control method for primary frequency regulation of distributed new energy and energy storage also includes: constructing a primary frequency regulation model of the power system containing distributed new energy and energy storage units, dynamically allocating agent decision weights based on an improved consensus algorithm incorporating Vague set theory, and realizing the initial collaborative framework of multiple agents.

[0073] In the process of constructing the primary frequency modulation model, a multi-scale morphological frequency divider filter is embedded to divide the frequency signal into a high-frequency band of 0.1-0.5Hz and a low-frequency band of 0.01-0.1Hz, which are then matched with energy storage units with different response characteristics.

[0074] The multi-agent system includes a new energy agent, an energy storage agent, and a coordination agent. Each agent has a built-in GRNN neural network prediction module that can predict frequency fluctuation trends in the next 30 seconds in real time.

[0075] During the coordinated control process, the IHFAA operator is used to fuse multi-source monitoring data to eliminate control deviations caused by asynchronous plant and grid signals.

[0076] In the energy storage unit control strategy, a SOC tiered threshold adaptive adjustment mechanism is set up to automatically adjust the frequency modulation output power when the SOC is in the range of 20%-30% or 70%-80%.

[0077] The distributed new energy and energy storage primary frequency regulation multi-agent collaborative control system also includes: a data acquisition module, a high-frequency synchronous communication module (communication delay ≤10ms), a multi-agent decision-making module, an execution control module, and a feedback optimization module.

[0078] The data acquisition module is configured with dual backup acquisition channels of fiber optic sensing and wireless transmission to synchronously acquire parameters such as frequency deviation, active power, and energy storage SOC.

[0079] The multi-agent decision-making module has a built-in primary frequency regulation performance evaluation and diagnostic model, and outputs real-time reverse regulation warnings and contribution power calculation results.

[0080] The execution control module has a virtual power plant cluster collaboration interface and supports parallel control of no less than 100 distributed units.

[0081] The feedback optimization module continuously iterates the control parameters based on the reinforcement learning algorithm, updating the collaborative control strategy every 5 minutes.

[0082] This invention further optimizes the accuracy, stability, and scalability of collaborative control based on existing technical solutions: It achieves dynamic adaptation of agent decision weights through a combination of an improved consensus algorithm and Vague set theory; it utilizes multi-scale morphological frequency division filters to perform hierarchical processing of frequency signals, matching the response characteristics of different energy storage units; a new GRNN neural network prediction module predicts frequency fluctuations in advance, and IHFAA operator data fusion eliminates signal interference, ensuring data reliability; it protects energy storage devices through SOC hierarchical threshold adaptive adjustment, and a high-frequency synchronous communication module ensures control synchronization; it employs dual backup acquisition channels to improve system stability, relies on a performance evaluation and diagnostic model to achieve dynamic management and control, expands the application scale through a virtual power plant cluster collaborative interface, and finally utilizes reinforcement learning algorithms to continuously iterate control parameters. These improvements make the system superior to existing technologies in terms of frequency regulation accuracy, equipment lifespan, and applicability, effectively solving the technical problems of slow response, poor coordination, and excessive losses in traditional systems.

[0083] A real-time communication delay compensation module is added. This module, based on a timestamp synchronization mechanism, performs millisecond-level delay correction on the interaction data between intelligent agents to ensure the real-time nature of collaborative decision-making.

[0084] The SOC dynamic zoning further introduces a temperature compensation factor, which adjusts the zoning boundaries according to the real-time temperature of the energy storage unit to cope with capacity decay under high and low temperature conditions.

[0085] The observation space is expanded to include the inertial center frequency offset of the power grid, and the global consistency of frequency deviation perception is improved through multi-source data fusion calculation.

[0086] A motion smoothing constraint is introduced, and a sliding window algorithm is used to filter continuous motion commands, limiting the rate of change of output between adjacent time points and preventing mechanical shock to the equipment.

[0087] The integrated frequency second derivative feedback term enhances the system's ability to suppress sudden disturbances by calculating the acceleration due to frequency changes.

[0088] A new frequency modulation contribution balancing reward item has been added, which calculates the fairness score based on the ratio of the actual output to the capacity of each unit, thus avoiding overload of individual devices.

[0089] Adding a network sparsity regularization term and introducing L1 norm penalty into the Critic network loss function reduces overfitting and improves generalization ability.

[0090] The evaluation index has been expanded to include the system recovery resilience index, which quantifies the grid's shock resistance performance by calculating the ratio of frequency recovery time to disturbance amplitude.

[0091] A dynamic weighting adaptive mechanism for indicators is introduced to automatically adjust the proportion of each evaluation indicator in the total score based on the importance of the working condition.

[0092] Upgrade the security verification module to support blockchain distributed ledger records to ensure the immutability and traceability of instruction correction history.

[0093] To further enhance the intelligence and security of collaborative control, the specific implementation of this invention is optimized as follows: In the real-time communication layer, a delay compensation module based on timestamp synchronization is added to perform millisecond-level correction on the agent's interactive data, ensuring decision synchronization. For energy storage units, a temperature compensation factor is introduced to dynamically adjust the SOC partition boundary to cope with capacity decay under high and low temperature conditions. Simultaneously, the observation space is expanded to include the grid inertial center frequency offset, improving global perception consistency through multi-source data fusion. In the action control layer, a new action smoothing constraint is added, employing a sliding window algorithm to filter output commands and limit the rate of change to prevent equipment impact. A frequency second derivative feedback term is integrated into the power balance equation to enhance the ability to suppress sudden disturbances. A frequency modulation contribution balancing term is added to the reward function, calculating a fairness score based on the capacity ratio to avoid equipment overload. A network sparsity regularization term is added during MADDPG training, improving generalization through L1 norm penalty. The evaluation system expands to include a system recovery resilience index and introduces a dynamic weight adaptive mechanism to adjust weights according to operating conditions. The security verification module is upgraded to support a blockchain distributed ledger, ensuring the immutability of command correction history. These optimizations, through algorithm refinement and hardware adaptation, significantly improve the reliability, fairness, and system resilience of the frequency modulation response.

[0094] Finally, it should be noted that the above description is merely one embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-agent cooperative control method for primary frequency regulation of distributed new energy and energy storage, characterized in that, include: A multi-agent power grid simulation environment is established, which includes distributed new energy sources and energy storage systems. The simulation environment is based on the power grid topology and integrates equipment models and operational constraints of wind power, photovoltaics and energy storage. Each distributed energy source and energy storage unit is assigned an independent intelligent agent, and the local observation space and action space of each intelligent agent are defined. The observation space includes the system frequency deviation, its own operating status and local power grid information, and the action space corresponds to the power output adjustment amount. A multi-agent reinforcement learning framework is constructed, and the MADDPG algorithm is used to realize the collaborative training of agents. The framework includes a shared Critic network and an independent Actor network, and stores and samples interaction data through an experience replay mechanism. Design a comprehensive reward function that integrates frequency deviation, voltage limit exceedance, energy storage SOC constraints, and renewable energy curtailment rate to guide the agent in optimizing action decisions; The trained multi-agent reinforcement learning model outputs real-time power output adjustment commands for each distributed energy source and energy storage unit, thereby achieving primary frequency regulation and coordinated control. The comprehensive reward function introduces an adaptive weight coefficient adjustment mechanism based on dynamic scheduling priority to respond to the differentiated requirements for frequency modulation speed and adjustment accuracy under different operating conditions. The comprehensive reward function is expressed as follows: ; in, , , , These are frequency deviation, voltage exceeding limit, The weighting coefficients for constraints and curtailment rates, This represents the cumulative value of node voltage exceeding the limit. For energy storage Penalties for deviations from the reasonable range For the curtailment rate of renewable energy; The SOC constraint of the energy storage system model is dynamically divided into a high response priority zone, an economic optimization zone, and a prohibited action zone; The energy storage Penalty items The calculation is performed using a piecewise nonlinear function: when When in the high-response priority zone, the penalty is zero to encourage action; when When in the economically optimal zone, the penalty term is a mild quadratic function; when As the penalty term approaches the boundary of the prohibited action zone, it becomes a steep exponential function to impose a strong constraint; the frequency deviation weight... Based on the interval to which its absolute value belongs, under the control of the adaptive weighting coefficient adjustment mechanism, The frequency deviation changes smoothly within a certain range, where k>1 is the emergency condition amplification factor; when the absolute value of the system frequency deviation is greater than the preset emergency threshold, the coefficient of the frequency deviation weighting term is automatically increased.

2. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, The distributed new energy equipment model includes output characteristic models for wind power and photovoltaics, and the energy storage system model includes charging and discharging power constraints, SOC upper and lower limit constraints, and energy conversion efficiency. The power grid simulation environment supports forward-backward substitution power flow calculation and can update node voltage and system power balance status in real time. The output characteristic models for wind power and photovoltaics not only include their maximum available power but also integrate the predicted power fluctuation range based on ultra-short-term prediction confidence intervals. The forward-backward substitution power flow calculation and the power-frequency dynamic equations of the simulation environment are solved synchronously and iteratively in a decoupled manner on a sub-second time scale.

3. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, The local observation space of the intelligent agent is defined as: ; in, For power system frequency deviation, This refers to the operating status of the device corresponding to the intelligent agent. This represents the average operating status of adjacent distributed energy sources. The observation space is a local node voltage sequence; an adaptive sensing mechanism is embedded within the observation space, wherein the range of adjacent distributed energy sources is dynamically determined based on the real-time topological connectivity and electrical distance of the power grid, and the self-operating state... For wind / solar power units, this includes the deviation between their ultra-short-term predicted power and their current actual output; for energy storage units, it includes their... The current dynamic partition status and estimated remaining adjustable time.

4. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, The action space of the intelligent agent is a continuous space, and the action values ​​are mapped to the device output adjustment amount: ; in, The standardized action output by the agent (range [-1,1]). Rated power of the equipment The output adjustment scaling factor is used as the actual adjustment command after the action is clipped by the equipment constraints; the output adjustment scaling factor Instead of being a fixed constant, these are time-varying parameters dynamically generated based on a deep belief network. The input to this deep belief network is the local observation space vector of the corresponding agent, and the output is a parameter adapted to the current urgency of the system and the state of the device itself. The value is calculated and jointly optimized through policy gradients during training.

5. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, The system frequency deviation is dynamically updated through the power balance equation: ; in, This represents the frequency deviation from the previous moment. For the total output of distributed new energy and energy storage, This represents the total system load (including random disturbances). The system damping coefficient is... Let be the system's inertial time constant. This is for simulating step size; In the power balance equation, the system inertial time constant With damping coefficient The parameters are modeled as dynamic parameters that change in real time with the penetration rate of distributed renewable energy, and their dynamic relationships are obtained through fitting offline simulation data. Simultaneously, a disturbance prediction and correction module based on a feedforward neural network is embedded in the frequency deviation update loop. This module takes the short-term variation trends of load and renewable energy output as input and adjusts accordingly. The item provides advance compensation to simulate and achieve predictive primary frequency modulation actions.

6. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, The network update process of the MADDPG algorithm includes: The Critic network receives the joint state and joint actions of all agents, outputs a global Q-value, and updates the network parameters by minimizing the Q-value prediction error. Each Actor network outputs action values ​​based on its own local observations, and updates its parameters by maximizing the Q-value evaluated by the Critic network using the policy gradient method. A soft update mechanism is used to update the parameters of the target Critic network and the target Actor network to ensure training stability. The training process of the MADDPG algorithm is divided into two stages: the first stage is offline centralized pre-training based on historical typical working condition data to initialize the policy network parameters of all agents; the second stage is online collaborative training based on real-time simulation data, and a policy distillation mechanism is introduced in this stage, that is, training a lightweight global policy evaluator, periodically extracting and fusing effective experiences from each independent Actor network to form a rapidly deployable suboptimal collaborative policy as a backup.

7. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 1, characterized in that, It also includes setting primary frequency regulation performance evaluation indicators to quantify the effect of coordinated control. The evaluation indicators include the average frequency deviation, the maximum frequency deviation, the number of voltage over-limit times, and the renewable energy curtailment rate. The evaluation indicators further include: a frequency regulation contribution fairness index, used to quantify the matching degree between the actual contribution and capacity ratio of distributed energy and energy storage at different regulation rates during the frequency regulation process; and a control command smoothness index, used to evaluate the rate of change of output commands at adjacent times to ensure equipment safety and lifespan.

8. The distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method according to claim 7, characterized in that, The mean frequency deviation evaluation index is expressed as follows: ; in, The mean of frequency deviation, To assess the time window, for The system frequency deviation at a given time; the calculation method for the frequency modulation contribution fairness index SF is as follows: , Where N is the total number of units participating in frequency modulation. For the first The ratio of the actual frequency modulation energy contribution of each unit to its rated adjustable capacity. This is the average of the ratios across all units; the closer the index is to 1, the better the fairness of the collaborative control.

9. A distributed new energy and energy storage primary frequency regulation multi-agent cooperative control system, used to implement the distributed new energy and energy storage primary frequency regulation multi-agent cooperative control method as described in claim 1, characterized in that, include: Environment construction module: Establishes a multi-agent power grid simulation environment including distributed new energy sources and energy storage systems, integrating equipment models, power flow calculations, and operational constraints; Intelligent agent configuration module: Assigns intelligent agents to each distributed energy source and energy storage unit, and defines the observation space and action space; Reinforcement learning framework module: Constructs the MADDPG algorithm framework, including an experience replay buffer, a shared Critic network, and independent Actor networks; Reward function design module: Constructs a comprehensive reward function that integrates multi-dimensional indicators to guide the agent in optimizing decision-making; The collaborative control module loads the trained model and outputs real-time power output adjustment commands for each device to achieve primary frequency modulation collaborative control. The performance evaluation module quantifies the collaborative control effect through preset evaluation indicators. The collaborative control system also includes an independent safety verification and command correction module, which is located between the collaborative control module and the physical devices. After receiving the power output command from the collaborative control module, it performs a millisecond-level rapid verification of the command based on a set of simplified and deterministic safety rules. If a risk of causing the device to exceed its limits or a momentary local voltage exceedance is detected, the command is slightly conservatively corrected using the safety rules, and the correction record is fed back to the experience playback buffer for subsequent training.