Virtual power plant gradeability dynamic approval method based on AGC performance index

By using phase space reconstruction, optimal transmission theory, and multi-agent reinforcement learning, the high-frequency transient ramping capability of the virtual power plant is dynamically reduced, solving the transient response problem of the virtual power plant in a very short time scale and realizing safe and accurate virtual power plant scheduling.

CN122065016AActive Publication Date: 2026-05-19HEFEI POWER SUPPLY COMPANY OF STATE GRID ANHUI ELECTRIC POWER +2
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI POWER SUPPLY COMPANY OF STATE GRID ANHUI ELECTRIC POWER
Filing Date
2026-04-21
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies cannot accurately quantify the transient burst potential of virtual power plants on extremely short time scales. Fixed-ratio allocation is difficult to overcome phase delay distortion when multiple nodes respond on a large scale. Traditional AI scheduling algorithms lack physical underlying constraints, leading to falsely advertised frequency regulation capabilities and scheduling default risks. They also cannot dynamically map capacity reduction caused by equipment lifespan degradation.

Method used

High-frequency transient ramp-up features of the equipment are extracted by phase space reconstruction, and multi-source node power allocation manifolds are generated by combining optimal transmission theory. The optimal response strategy is solved within the safety boundary by using multi-agent reinforcement learning, and fatigue damage is quantified by rainflow counting method to dynamically reduce the remaining available capacity.

Benefits of technology

It accurately captures the high-frequency transient ramp-up potential of heterogeneous equipment, eliminates phase delay and spatial allocation distortion, embeds physical safety constraints, ensures that control commands are safe and do not violate regulations, dynamically reduces equipment fatigue damage, and improves the real reliability and scheduling accuracy of virtual power plant cluster control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065016A_ABST
    Figure CN122065016A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual power plant gradeability dynamic approval method based on AGC performance indexes, and relates to the technical field of virtual power plant scheduling and power system control. The method comprises the following steps: performing phase-space reconstruction on distributed resources to extract chaotic features and generate a basic climbing feature set; generating a multi-source node power distribution manifold by using an optimal transmission theory; the AGC index is converted into a dynamic repulsion potential field based on a control barrier function, and an optimal response strategy is solved in a manifold through multi-agent reinforcement learning; and calculating the power fluctuation fatigue damage degree by using a rain flow counting method, and dynamically reducing the residual available capacity. The method is used for solving the problems that traditional static approval is low in precision, massive resources respond to phase distortion, AI regulation and control lacks physical security constraints, and scheduling default is caused by the fact that service life attenuation of equipment under the high-frequency frequency modulation working condition cannot be sensed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual power plant dispatching and power system control technology, and more specifically, to a method for dynamic approval of the ramp-up capability of a virtual power plant based on AGC performance indicators. Background Technology

[0002] Virtual power plants, by aggregating massive heterogeneous distributed resources to participate in grid frequency regulation, have become a core means of improving the flexibility of new power systems. In Automatic Generation Control (AGC) scenarios, the dispatch master station places extremely high demands on the response rate and control accuracy of frequency regulation resources. To ensure that the aggregation cluster can accurately track the issued high-frequency transient commands, it is necessary to conduct precise and continuous verification and evaluation of the actual ramp-up capability and response boundaries of the virtual power plant under real complex operating conditions.

[0003] Currently, for the approval of virtual power plant ramp-up capabilities, mainstream technologies mostly adopt static proportional allocation methods or linear time-domain smoothing prediction methods based on equipment nameplate capacity. These schemes typically extract historical low-frequency output curves of the equipment, calculate aggregated capacity by combining simple fixed delay parameters, and use conventional linear programming or model-free reinforcement learning algorithms to generate power allocation strategies, which serve as evaluation benchmarks for cluster regulation capabilities and AGC response performance.

[0004] However, existing technologies have significant limitations: First, linear time-domain characteristics cannot accurately quantify the transient burst potential of heterogeneous resources on extremely short time scales, and fixed-ratio allocation is difficult to overcome the phase delay distortion caused by large-scale responses from multiple nodes; second, traditional black-box AI scheduling algorithms lack strong constraints at the physical level, making them prone to deviating from the grid AGC assessment baseline and even exceeding the physical safety boundaries of equipment when exploring control strategies; finally, existing static evaluation mechanisms ignore the microscopic fatigue damage caused to underlying equipment by continuous high-frequency frequency regulation actions, and cannot dynamically map the actual capacity derating caused by lifespan decay, which can easily lead to falsely advertised frequency regulation capabilities and scheduling default risks. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of existing technologies, embodiments of the present invention provide a dynamic verification method for the ramping capability of a virtual power plant based on AGC performance indicators. This method extracts high-frequency transient ramping features of equipment through phase space reconstruction and generates a power allocation manifold by combining optimal transmission theory. Based on the control barrier function, it guides multi-agent reinforcement learning to solve the optimal response strategy within the safety boundary. Finally, it uses the rainflow counting method to quantify power fluctuation fatigue damage to dynamically reduce the remaining available capacity. This addresses the technical problems of traditional virtual power plants, such as falsely inflated static capacity assessment, phase delay distortion in massive multi-node collaborative response, lack of underlying physical security constraints in AI black-box scheduling leading to over-limit violations, and inability to truly perceive the actual lifespan degradation of equipment under continuous high-frequency modulation conditions.

[0006] To achieve the above objectives, the present invention provides the following technical solution: Phase space reconstruction is performed on the historical operating time series data and meteorological series data of distributed resources, and feature parameters are extracted to generate a basic climbing feature set. By combining the basic ramping feature set, communication response delay time and overall ramping requirements, a multi-source node power allocation manifold is generated based on optimal transmission theory. Based on the control barrier function, the target AGC comprehensive performance index requirement is transformed into a dynamic repulsive potential field, and the optimal response strategy is solved within the manifold using a multi-agent reinforcement learning algorithm. The fatigue damage degree of the power fluctuation data for implementing the strategy is calculated using the rainflow counting method, and the remaining available capacity is dynamically reduced.

[0007] In a preferred embodiment, the step of reconstructing the phase space of historical operating time series data and meteorological series data of distributed resources, and extracting feature parameters to generate a basic climbing feature set includes: Determine the delay time and embedding dimension of the historical running time series data; Based on the delay time and embedding dimension, historical runtime time series data are mapped to a high-dimensional topological phase space; The maximum Lyapunov exponent of each evolution trajectory in the high-dimensional topological phase space is calculated and used as the characteristic parameter.

[0008] In a preferred embodiment, after calculating the maximum Lyapunov exponent of each evolution trajectory in the high-dimensional topological phase space, the method further includes: Determine whether the maximum Lyapunov exponent is greater than zero; If the value is greater than zero, the time series of the corresponding device is determined to have chaotic characteristics, and the feature parameter is added to the basic climbing feature set. If the value is less than or equal to zero, the corresponding device is determined not to have high-frequency transient ramping potential and will not be included in the basic ramping feature set.

[0009] In a preferred embodiment, the step of generating a multi-source node power allocation manifold based on optimal transmission theory, by combining the basic ramp feature set, communication response delay time, and overall ramp requirements, includes: The feature parameters are weighted with the communication response delay time to obtain the AGC multi-source contribution weight of each distributed resource; wherein, the feature parameters are positively correlated with the AGC multi-source contribution weight, and the communication response delay time is negatively correlated with the AGC multi-source contribution weight. With the goal of minimizing the power transmission cost of all distributed resources, and with the overall ramping demand and the multi-source contribution weight of AGC as constraints, an optimal transmission objective function based on Wasserstein distance is constructed. Solve for the optimal transmission objective function to obtain the power allocation manifold of the multi-source nodes that satisfies the minimum power transmission cost.

[0010] In a preferred embodiment, solving the optimal transmission objective function to obtain the multi-source node power allocation manifold that satisfies the minimum power transmission cost includes: By introducing an entropy regularization term into the optimal transmission objective function, the optimal transmission objective function based on Wasserstein distance is transformed into a strictly convex optimization problem. The Sinkhorn alternating projection algorithm is used to iteratively solve the strictly convex optimization problem until the preset convergence condition is met. The converged transmission plan matrix is ​​used as the optimal power allocation strategy to generate a multi-source node power allocation manifold.

[0011] In a preferred embodiment, the optimal response policy is solved using a multi-agent reinforcement learning algorithm within the manifold, including: Analyze the target AGC comprehensive performance index requirements and extract parameters such as adjustment rate, adjustment accuracy and response time; The physical safety boundary in the Hamilton-Jacobi reachability control model is determined based on the parameters of regulation rate, regulation accuracy and response time, and a continuous-time control barrier function is constructed. The control obstacle function is used as a hard constraint layer and embedded into the action output of a multi-agent reinforcement learning algorithm for solution.

[0012] In a preferred embodiment, the step of embedding the action output of the multi-agent reinforcement learning algorithm for solving includes: The multi-agent reinforcement learning algorithm is a deep dual-Q network reinforcement learning algorithm. Real-time acquisition of candidate action instructions output by the deep double-Q network reinforcement learning algorithm; Calculate the control obstacle function value corresponding to the candidate action command; If the control barrier function value is less than zero, a dynamic repulsion potential field is triggered to perform spatial projection correction on the candidate action command, pull the corrected action command back into the physical safety boundary, and output it as the optimal response strategy. If the value of the control barrier function is greater than or equal to zero, the candidate action command is directly output as the optimal response strategy.

[0013] In a preferred embodiment, calculating the fatigue damage degree of the power fluctuation data for implementing the strategy using the rainflow counting method includes: Extract the high-frequency power fluctuation curves of each distributed resource in the current scheduling cycle from the optimal response strategy; Extreme value extraction is performed on the high-frequency power fluctuation curve to filter out invalid fluctuations with amplitudes smaller than a preset dead zone threshold, thereby obtaining an effective stress time series; The effective stress time series is extracted using the rainflow counting method, and the effective stress time series is transformed into an equivalent alternating stress cyclic spectrum containing different mean values ​​and amplitudes.

[0014] In a preferred embodiment, the dynamic reduction of remaining available capacity includes: Based on Miner's linear fatigue cumulative damage theory, the single fatigue damage degree corresponding to the equivalent alternating stress cycle spectrum in the current scheduling cycle is calculated. The single fatigue damage degree is accumulated into the historical damage record of the corresponding equipment to obtain the cumulative fatigue damage degree; According to the preset nonlinear attenuation model, the initial ramp capacity of the distributed resource is reduced based on the cumulative fatigue damage degree to obtain the remaining available capacity.

[0015] This invention provides a dynamic verification system for the ramping capability of a virtual power plant based on AGC performance indicators, including: a feature extraction module, used to reconstruct the phase space of historical operating time series data and meteorological series data of distributed resources, and extract feature parameters to generate a basic ramping feature set; The manifold allocation module is used to combine the basic ramp feature set, communication response delay time and overall ramp requirements to generate a multi-source node power allocation manifold based on optimal transmission theory. The strategy generation module is used to transform the target AGC comprehensive performance index requirements into a dynamic repulsive potential field based on the control barrier function, and to solve the optimal response strategy within the manifold using a multi-agent reinforcement learning algorithm. The dynamic reduction module is used to calculate the fatigue damage degree of the power fluctuation data for implementing the strategy using the rainflow counting method, and dynamically reduce the remaining available capacity.

[0016] The technical effects and advantages of the present invention regarding the dynamic verification method for the ramp-up capability of virtual power plants based on AGC performance indicators are as follows: This invention extracts core feature parameters by reconstructing the phase space of historical operating time series and meteorological sequence data of distributed resources, accurately capturing the high-frequency transient ramping potential of heterogeneous equipment. Combining communication latency and overall control requirements, it innovatively uses optimal transmission theory to generate a multi-source node power allocation manifold, effectively eliminating phase delay and spatial allocation distortion problems during massive concurrent node responses. Simultaneously, based on a control barrier function, the strict AGC comprehensive performance evaluation index is transformed into a dynamic repulsive potential field, guiding a multi-agent reinforcement learning algorithm to solve for the optimal response strategy within this allocation manifold. This embeds robust underlying physical security hard constraints into the black-box AI optimization process, ensuring absolute safety and compliance of control commands. Finally, it utilizes rainflow counting to quantify the fatigue damage caused by high-frequency power fluctuations in the execution of response strategies, thereby scientifically and dynamically reducing the remaining available capacity of the equipment. This solution completely breaks through the limitations of traditional static capacity assessment based on nameplate parameters, constructing a dynamic approval closed loop that considers multi-node collaborative resistance reduction, absolute physical boundary safety, and equipment lifecycle health perception, greatly improving the real reliability and scheduling accuracy of virtual power plant clusters participating in high-frequency AGC control of new power systems. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the process for dynamically approving the ramp-up capability of a virtual power plant based on AGC performance indicators, as provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram comparing the power response trajectories of multi-agent reinforcement learning under extreme scheduling instructions, provided in an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of the three-dimensional data histogram of the equivalent alternating stress cyclic spectrum provided by the rainflow counting method in an embodiment of the present invention.

[0020] Figure 4 This is a schematic diagram of a virtual power plant ramp-up capability dynamic approval system based on AGC performance indicators, provided in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0022] Example 1, Figure 1 This invention presents a method for dynamically verifying the ramp-up capability of a virtual power plant based on AGC performance indicators, comprising the following steps: Step S1: Reconstruct the phase space of the historical operating time series data and meteorological series data of the distributed resources, and extract feature parameters to generate a basic climbing feature set.

[0023] It should be noted that, in order to obtain high-quality underlying multi-source heterogeneous device data, the underlying data acquisition terminal acquires real-time apparent power, active power output, and voltage transient data of photovoltaic inverters, energy storage converters, and adjustable loads at a high-frequency sampling rate of hundreds of milliseconds using smart meters based on the IEC 61850 protocol. Simultaneously, meteorological sequence data including ambient temperature, instantaneous irradiance, wind speed, and humidity are collected through a micro-weather station array deployed at the plant side. This multi-source heterogeneous data is aggregated through the network to the edge computing gateway for time-series alignment and cleaning, forming the original time-series dataset.

[0024] This embodiment introduces chaotic dynamics theory for phase space reconstruction to uncover the implicit nonlinear dynamic evolution laws of time series. Specifically, the phase space reconstruction of historical operating time series data and meteorological sequence data of distributed resources, and the extraction of feature parameters to generate a basic climbing feature set, includes: determining the delay time and embedding dimension of the historical operating time series data. When determining the delay time, mutual information is used to measure the nonlinear correlation between the time series and its delayed sequence. By calculating the mutual information between the original time series and its sequences at different delay step sizes, the first local minimum point in the mutual information function graph is found, and the step size corresponding to this minimum point is the optimal delay time. Based on this, the False Nearest Neighbors method is used to determine the optimal embedding dimension. The algorithm calculates the rate of change of the Euclidean distance between adjacent points in the phase space by gradually increasing the dimension of the phase space. When the dimension increases to a certain value, the proportion of spurious neighbor points caused by low-dimensional projection drops sharply and falls below a preset tolerance threshold (e.g., 5%). At this point, the corresponding dimension is strictly selected as the embedding dimension. .

[0025] Next, using the aforementioned delay time and embedding dimension, the historical operating time series data is mapped to a high-dimensional topological phase space. Considering that this embodiment involves multivariate meteorological sequences such as active power, temperature, and instantaneous irradiance, before mapping, a multivariate phase space reconstruction method based on principal component analysis (PCA) or direct state splicing is used to bring the master sequence after dimensionality reduction and fusion of multi-source heterogeneous data into the following reconstruction process. Based on Takens' embedding theorem, the dimensionality-reduced one-dimensional time series is losslessly expanded and reconstructed into an equivalent high-dimensional dynamic system state space to fully reveal the intrinsic evolutionary dynamic structure hidden behind equipment operation. The calculation formula for the reconstructed state vector of the historical operating time series data mapped to the high-dimensional topological phase space is as follows: (1) In the formula, Represents the first in a high-dimensional topological phase space A reconstructed state vector, This indicates that the time series after dimensionality reduction and fusion is at time 10:00. The sampled values, This represents the optimal delay time determined by the mutual information method. This represents the optimal embedding dimension determined by the spurious nearest neighbor method. This represents the vector transpose operation.

[0026] Subsequently, the maximum Lyapunov exponent of each evolutionary trajectory in the high-dimensional topological phase space is calculated, and this maximum Lyapunov exponent is used as the characteristic parameter. To quantify the exponential separation rate of adjacent trajectories in the phase space, the Rosenstein small-data-quantity method is used to find the nearest neighbor of each state vector in the reconstructed high-dimensional phase space, and close points on the same trajectory are excluded by limiting the initial time interval between the two. The evolution distance between these two state vectors over time is continuously tracked, and the ensemble average of the average divergence logarithms of all reference trajectories is calculated. Specifically, the evolution is calculated using the following formula. Average divergence after discrete time steps : (2) In the formula, This represents the total number of reference trajectories (i.e., the number of effective state vectors) participating in the ensemble averaging in phase space. Indicates the first A state vector and its nearest neighbor pass through The Euclidean distance evolved over discrete time steps. Plotting... With time step The evolution curve is analyzed, and the slope of the initial linearly rising region of the curve is extracted as the maximum Lyapunov exponent of the evolution trajectory. That is, the characteristic parameters.

[0027] Further, after calculating the maximum Lyapunov exponent of each evolution trajectory in the high-dimensional topological phase space, the method further includes: determining whether the maximum Lyapunov exponent is greater than zero; if the maximum Lyapunov exponent is greater than zero, then the time series of the corresponding device is determined to have chaotic characteristics, and the feature parameter is added to the basic climbing feature set; if the maximum Lyapunov exponent is less than or equal to zero, then the corresponding device is determined not to have high-frequency transient climbing potential, and it is not included in the basic climbing feature set. Let the evaluated... The characteristic parameters corresponding to each distributed resource are: The basic climbing feature set is specifically represented by the following formula: (3) In the formula, This represents the set of basic climbing features generated after rigorous screening based on chaotic properties. This represents the total number of distributed resources under the virtual power plant that participated in the performance evaluation.

[0028] It should be noted that, at the physics level, a maximum Lyapunov exponent greater than zero indicates that the distributed resource (such as wind power or photovoltaics) is severely affected by the external meteorological environment, and its output sequence exhibits nonlinear and drastic fluctuation characteristics (i.e., chaotic characteristics). Extracting this characteristic and incorporating it into the basic ramp-up feature set allows the system to identify these sources with extremely high power fluctuation potential in advance during subsequent steps. When allocating power to multiple source nodes in a virtual power plant, the system can specifically utilize the inherent high-frequency fluctuation characteristics of these resources to respond to transient ramp-up demands, or plan strategies in advance to smooth out such fluctuations under opposite demands. Conversely, if the exponent is not greater than zero, it indicates that the equipment output tends to be stable or the system exhibits strong damping characteristics, lacking the chaotic physical basis to cause drastic power fluctuations. To more intuitively illustrate the physical effectiveness of the chaotic dynamics screening mechanism in this embodiment, Table 1 presents the measured simulation data of the phase space reconstruction characteristic parameters of some typical heterogeneous devices in a virtual power plant.

[0029] Table 1

[0030] As shown in Table 1, the maximum Lyapunov exponents of both flywheel energy storage and electrochemical energy storage are significantly greater than zero, exhibiting extremely strong transient burst capabilities, and were successfully included in the system's basic ramp-up characteristic set. However, due to the strong damping characteristics of their internal thermodynamic systems, large temperature-controlled loads (such as central air conditioning clusters) have negative maximum Lyapunov exponent values, exhibiting a state of convergence, and were accurately eliminated by the algorithm. This data distribution objectively demonstrates that this method can accurately quantify and differentiate the true high-frequency response potential of different devices from a physical perspective. This step, by introducing chaotic dynamics theory and extracting the maximum Lyapunov exponent as the core criterion, not only breaks the limitation that conventional linear time-domain features cannot accurately characterize the power mutation capability of devices in extremely short time scales, but also achieves accurate capture of the high-frequency transient response potential of devices and early physical hard isolation of hysteresis ineffective devices, thus laying a high-quality data foundation for the subsequent construction of a high-precision multi-source AGC power allocation manifold.

[0031] S2 combines the basic ramp feature set, communication response delay time and overall ramp requirements to generate a multi-source node power allocation manifold based on optimal transmission theory.

[0032] It should be noted that, in response to the phase delay and time domain response dead zone problems caused by the physical space span, network topology differences, and local response lag when a large number of heterogeneous devices receive scheduling instructions simultaneously, this embodiment introduces the optimal transmission theory in fluid mechanics to physically map the power allocation process of the virtual power plant into a probabilistic fluid transport process.

[0033] Specifically, the step of combining the basic ramp feature set, communication response delay time, and overall ramp demand to generate a multi-source node power allocation manifold using optimal transmission theory includes: weighting the feature parameters with the communication response delay time to obtain the AGC multi-source contribution weight for each of the distributed resources; wherein the feature parameters are positively correlated with the AGC multi-source contribution weight, and the communication response delay time is negatively correlated with the AGC multi-source contribution weight. After obtaining the basic ramp feature set, the underlying communication gateway, through network slicing supporting IEEE 1588 PTP, calculates the power allocation manifold from the master station to the next node. The bidirectional communication delay time of each distributed resource. To balance the high-frequency transient response potential of the device with the physical latency of network transmission, this embodiment establishes a multi-dimensional evaluation system through a normalized exponential function with an attenuation factor. The calculation formula for the AGC multi-source contribution weight of each distributed resource is as follows: (4) In the formula, Indicates the first AGC multi-source contribution weights for distributed resources; Indicates the first The characteristic parameters of a distributed resource (i.e., the maximum Lyapunov exponent). The largest Lyapunov exponent in the basic climbing feature set is used to perform maximum normalization; Indicates the dispatching master station to the number Communication response latency of distributed resources; and It is an importance adjustment coefficient and always satisfies ; A positive delay decay hyperparameter is used to control the sensitivity of time delay to exponential decay boundary.

[0034] After obtaining the weight constraints of each node, the optimization objective is further optimized by minimizing the power transmission cost of all distributed resources, with the overall ramping demand and the AGC multi-source contribution weights as constraints. An optimal transmission objective function based on Wasserstein distance is then constructed. Physically, optimal transmission theory treats the overall ramping demand issued by the scheduling master station as the initial source domain probability quality distribution and each distributed device as the target domain probability quality distribution. This objective function represents the minimum transmission work cost required to "transfer" and cut the total ramping demand in the multi-dimensional state space and inject it into each physical node. This process fully considers the heterogeneity of each node in terms of network distance and dynamic performance. The calculation formula for the optimal transmission objective function based on Wasserstein distance is as follows: (5) In the formula, This represents the function that minimizes the total power transmission cost, i.e., the Wasserstein distance; This represents the multi-source node power transmission plan matrix (i.e., the set of decision variables) to be optimized; This represents the number of discretized source nodes representing the overall ramp demand (in this power grid dispatching scenario, the master station demand is usually discretized into a small number of source nodes according to the topological region to ensure vector dimension matching). Indicates the source node Transmission unit power requirement up to the The physical transmission cost matrix element of each distributed resource is defined as the network electrical distance (specifically, in this embodiment, it is taken as the equivalent impedance of the AC tie line from the scheduling master station to the corresponding distributed resource node) multiplied by the inverse of the multi-source contribution weight. The product; Indicates the source node Assigned to the The power transmission plan matrix elements of each distributed resource are the decision variables to be optimized. These decision variables are constrained by the overall system ramp-up requirements and the capacity boundary constraints of the individual devices.

[0035] To address the "dimensionality explosion" problem inevitably encountered by traditional linear programming algorithms such as the simplex method or interior point method when solving problems involving tens of thousands of high-dimensional distributed nodes, this step utilizes geometric smoothing to project the optimization space into a lower dimension. Solving the optimal transmission objective function to obtain the multi-source node power allocation manifold that satisfies the minimum power transmission cost includes: introducing an entropy regularization term into the optimal transmission objective function, transforming the Wasserstein distance-based optimal transmission objective function into a strictly convex optimization problem. The introduction of entropy regularization expands the originally sparse and non-smooth discrete solution into a dense flow field distribution with non-zero probabilities, greatly improving the differentiability of the objective function and the strict convexity of the Hessian matrix. The objective function of the strictly convex optimization problem, transformed from the Wasserstein distance-based optimal transmission objective function, is specifically expressed by the following formula: (6) In the formula, This represents the smooth optimal transmission cost after introducing entropy regularization. Let represent the entropy penalty hyperparameter that controls the degree of regularization shrinkage, and satisfy . ; This refers to the defined independent Shannon information entropy regularization term element.

[0036] Considering that in actual virtual power plant scheduling, the total available capacity of the system is usually greater than the overall ramp-up demand (i.e., an unbalanced optimal transmission problem), a virtual absorbing node with a capacity equal to the difference between the total capacity and the total demand needs to be introduced before constructing the constraint boundary vector and performing iterative solutions. Specifically, this virtual node is added to the following... At the end of the vector, expand its dimensions to... And set the corresponding physical transmission cost matrix elements to zero or a uniform constant penalty value to ensure This satisfies the quality conservation premise of the standard optimal transmission theory.

[0037] This embodiment utilizes the Sinkhorn alternating projection algorithm to iteratively solve the strictly convex optimization problem until a preset convergence condition is met. The Sinkhorn algorithm cleverly transforms the original problem into an alternating scaling projection process of two diagonal matrices onto a Gibbs kernel matrix. Let the Gibbs kernel matrix be... The elements are The iterative process introduces a left-scaled column vector. With right-scaled column vectors The specific matrix multiplication form for iterative updates is expressed as: (7) (8) In the formula, and They represent the first The left-scaled column vector and the right-scaled column vector are updated in the next iteration step; in this embodiment... This is the normalized column vector of the overall ramp demand quota distribution. This represents a column vector containing the capacity boundary distribution of each distributed resource according to its weight. The Gibbs kernel matrix is ​​represented by its transpose; the division fraction bar indicates digit-by-digit division of the corresponding elements in the vector (Hadamard division). In this iterative projection calculation, the algorithm monitors the iteration error of the edge distribution in real time until the L1 norm of the error is less than a preset convergence threshold (e.g., set to...). If convergence is reached, iteration stops. Then, the converged transmission plan matrix is ​​used as the optimal power allocation strategy to generate the multi-source node power allocation manifold. The final high-dimensional manifold is... High-dimensional surfaces under mapping; where, This represents the optimal multi-source node power allocation strategy matrix after convergence. This represents the diagonalization operator for column vector elements as the main diagonal elements. and These represent the left-scaled column vector and the right-scaled column vector, respectively, of the final output after satisfying the preset convergence conditions. This manifold constitutes a smooth geometric energy channel for distributing global control commands to lower-level physical nodes.

[0038] This step innovatively maps the evolution mechanism of probabilistic flow field in hydrodynamics to power grid allocation, which not only completely eliminates the geometric rigid aggregation blind zone caused by traditional fixed capacity ratio control, but also optimally solves the distortion and overlap problem caused by communication delay during the response of massive heterogeneous devices with extremely low time complexity and computing power overhead.

[0039] Step S3: Based on the control barrier function, the target AGC comprehensive performance index requirement is transformed into a dynamic repulsive potential field, and the optimal response strategy is solved within the manifold using a multi-agent reinforcement learning algorithm.

[0040] After obtaining the physical mapping manifold, the dynamic scheduling solution phase begins. Specifically, within the manifold, a multi-agent reinforcement learning algorithm is used to solve for the optimal response strategy, including: analyzing the target AGC comprehensive performance index requirements and extracting the regulation rate, regulation accuracy, and response time parameters; the dispatch control center, by analyzing the standard AGC command messages issued by the power grid, extracts the assessment benchmark for the overall response performance of the virtual power plant, namely, the regulation rate required to be achieved by the unit, the steady-state allowable regulation accuracy (power deviation percentage), and the response time limit from receiving the command to crossing the response dead zone.

[0041] Traditional model-free reinforcement learning is prone to causing physical equipment damage or failing AGC (Automatic Guided Vehicle) assessments due to outputting actions that exceed limits during the exploration process. Therefore, this paper determines the physical safety boundary in the Hamilton-Jacobi reachability control model based on the aforementioned adjustment rate, adjustment accuracy, and response time parameters, and constructs a continuous-time control barrier function. This process rigorously maps the extracted abstract performance index parameters to multidimensional geometric boundaries in the state space. The calculation formula for the control barrier function is as follows: (9) In the formula, This is the mapped continuous-time control barrier function; To start from the current control model state vector The actual combined active power output of the virtual power plant extracted from it; The active power reference value required by the target AGC command; The steady-state allowable power error limit is obtained by analyzing the target AGC adjustment accuracy. The adjustment rate parameter required by the target AGC instruction; The response time parameter required by the target AGC command; From the state vector The combined state of charge (or equivalent adjustable margin) of the available devices at the current moment is extracted. This represents the system's rated maximum state of charge. This is a dimensionless dynamic margin adjustment coefficient.

[0042] Considering the actual evolution of the underlying physical equipment in the virtual power plant, this embodiment pre-constructs an affine nonlinear dynamic model of the control system based on the energy storage capacity, charge / discharge rate, and response damping of each heterogeneous device. ,in For the system's internal drift vector field, To control the input vector field, a control model state vector is defined, containing physical quantities such as the real-time power and state of charge of each device. and build an absolutely secure set In a physical sense, This indicates that the current combined output status of the virtual power plant is entirely within the safe operating range that meets the AGC assessment indicators and the equipment does not experience overload; when Approaching zero signifies that the current operating state is dangerously close to the physical boundary. To ensure that the operating trajectory of the virtual power plant cluster always evolves within the safe set without exceeding the limits, the continuous-time control barrier function must satisfy the Lie derivative constraint condition, specifically expressed by the following formula: (10) In the formula, Represents a control barrier function for continuous time; Represents the state vector of the control model; This represents the control action vector (i.e., power regulation command) applied to the controlled object. This indicates the physical limits of the operating space allowed for the underlying heterogeneous devices; The Lie derivative of the control barrier function along the drift vector field of the nonlinear dynamic model is given. This represents the Lie derivative of the control barrier function along the input control vector field; This represents a strictly increasing extended K-class function used to adjust the dynamic buffer repulsion strength when the running state approaches the safety boundary.

[0043] Next, the control obstacle function is used as a hard constraint layer and embedded into the action output of the multi-agent reinforcement learning algorithm for solution. Specifically, the embedding into the action output of the multi-agent reinforcement learning algorithm for solution includes: the multi-agent reinforcement learning algorithm is a deep double Q-network reinforcement learning algorithm; candidate action instructions output by the deep double Q-network reinforcement learning algorithm are acquired in real time; the scheduling platform deploys a deep double Q-network (Double DQN) based on a centralized training and distributed execution architecture, using the decoupled dual structure of the main network and the target network to eliminate the inherent action overestimation bias in traditional Q-learning. The state space of the agent specifically includes the current AGC overall ramp demand instruction, the real-time active power output of each distributed resource, and the state of charge (SOC) of the energy storage device. In terms of network architecture, both the main network and the target network adopt multilayer fully connected neural networks (MLP), with the joint state vector as input, and the output layer nodes corresponding to the Q-value evaluation of each discrete action index. Meanwhile, to guide the agent to learn autonomously and approach the globally optimal strategy, a comprehensive reward function for the Markov decision process was designed. This reward function mainly consists of a linearly weighted performance penalty term negatively correlated with AGC command tracking error and an economic penalty term negatively correlated with the marginal cost of equipment power adjustment. The specific expression of the reward function is as follows: (11) In the formula, Indicates time The comprehensive reward function value; Indicates time The overall ramp-up requirement instruction for AGC; Indicates the first Distributed resources at time The actual power regulation amount; Indicates the first The marginal cost coefficient for power regulation of a distributed resource; and These represent the weights of the performance penalty and the economic penalty, respectively.

[0044] Within each extremely short control cycle, candidate action commands output by the deep double-Q network reinforcement learning algorithm are acquired in real time. Since the output of the deep double-Q network is a discrete action index, the scheduling platform pre-constructs a mapping dictionary from the discrete action space to continuous physical commands. The discrete action indexes output by the algorithm are then transformed into continuous reference power adjustment vectors using the mapping dictionary. This is so that it can be input into the subsequent continuous-time security verification layer.

[0045] Subsequently, the control obstacle function value corresponding to the candidate action command is calculated; that is, the mapped continuous action command. Substitute the values ​​into the Lie derivative inequality for calculation. For ease of engineering description, this embodiment includes the CBF inequality constraint terms containing the action. The calculation results are uniformly and equivalently defined as "the control obstacle function value corresponding to the candidate action instruction" (i.e., the action safety assessment margin), and are subject to mathematical verification and underlying safety verification.

[0046] Under this verification mechanism, if the control barrier function value is less than zero, the dynamic repulsion potential field is triggered to perform spatial projection correction on the candidate action command, pulling the corrected action command back to the physical safety boundary and outputting it as the optimal response strategy. When a candidate action causes the safety constraint to be violated (i.e., the evaluation result is less than zero, indicating that an AGC default or equipment overload danger is about to occur), the system essentially transforms the target AGC comprehensive performance index requirement into a dynamic repulsion potential field in the mathematical action space based on the control barrier function. Specifically, the hard constraint layer transforms it into an online solveable quadratic programming (QP) optimization problem. This underlying mechanism aims to minimize the deviation from the original candidate action. Under the linear inequality hard constraint of the control barrier function, it searches for the closest and absolutely safe alternative action in the multidimensional action space, thereby forming an invisible repulsion potential field that forcibly bounces the overstepping action back to the safety domain. The calculation formula of the quadratic programming objective function for spatial projection correction is as follows: (12) (13) In the formula, This represents the optimal response strategy output after spatial projection correction (i.e., the final command issued after safety verification). This represents the continuous reference power adjustment vector after mapping the original output of the deep double-Q network reinforcement learning algorithm; The squared Euclidean norm representing the deviation between action vectors; inequality constraints. This refers to the underlying constraints of the control barrier function that ensure the corrective action is absolutely pulled back within the physical safety boundary.

[0047] Conversely, if the control barrier function value is greater than or equal to zero, the candidate action command is directly output as the optimal response strategy. In this case, it is proven that the initial exploration action generated by the neural network perfectly meets all AGC performance index requirements and equipment operation physical constraints, without the need to initiate a repulsive potential field intervention, thus issuing the command unchanged (i.e., ...). As the ultimate optimal response strategy, it preserves reinforcement learning's black-box optimization capabilities for economic costs and global control objectives to the greatest extent possible.

[0048] Figure 2 Simulation comparisons of the power response trajectories of the multi-agent reinforcement learning algorithm with CBF hard constraint layer described in this invention and the traditional unconstrained Double DQN algorithm are presented when facing consecutive out-of-limit extreme scheduling instructions. Figure 2 As shown in the data curves, the X-axis represents the AGC scheduling time (seconds), and the Y-axis represents the actual output of the device after standardization (PU). During the extreme ramp-up demand surge interval from the 15th to the 25th second, the instructions output by the traditional unconstrained reinforcement learning algorithm (dashed curve in the figure) caused the actual output of the distributed energy storage unit to reach 1.12 PU, exceeding the physical safety limit of the nameplate capacity (1.0 PU), resulting in a physical over-limit danger. However, the solution of this embodiment, which introduces a dynamic repulsion potential field by a control barrier function (solid curve in the figure), successfully triggered the quadratic planning spatial projection correction at the microsecond instant of touching the physical safety boundary (1.0 PU), strictly and smoothly clamping the output instructions within the safety boundary, achieving global zero over-limit operation.

[0049] This step, by deeply integrating Hamilton-Jacobi reachability control theory with deep double-Q networks, endows artificial intelligence with an insurmountable security mindset at its core. Without requiring a major reconstruction of the neural network architecture, it achieves transient fine-tuning of dangerous scheduling instructions with minimal microsecond-level computational overhead, resulting in an unexpected technical effect that balances global economic deep optimization with absolute non-violation of physical control.

[0050] Step S4: Calculate the fatigue damage degree of the power fluctuation data for implementing the strategy using the rainflow counting method, dynamically reduce the remaining available capacity, and generate a digital capability certificate.

[0051] To ensure the authenticity of the entire lifecycle of capability approval and the physical security of dispatching, this embodiment innovatively introduces materials mechanics and fatigue damage theory, mapping the degradation of energy storage batteries during charge-discharge cycles or the mechanical wear of flywheel equipment to the fatigue cost under high-frequency modulation operations. Specifically: The calculation of fatigue damage degree of power fluctuation data for implementing the strategy using the rainflow counting method includes: extracting high-frequency power fluctuation curves of each distributed resource in the current scheduling cycle from the optimal response strategy; subsequently, the edge computing gateway performs extreme value extraction on the high-frequency power fluctuation curves, filtering out invalid fluctuations with amplitudes less than a preset dead zone threshold, and obtaining an effective stress time series. In actual operation, small power fluctuations (e.g., white noise level disturbances below 2% of rated power) have negligible impact on equipment lifespan. By setting a dead zone threshold (e.g., setting the dead zone range to 0.02 per unit) for preliminary filtering, subsequent computing power overhead can be significantly reduced and the signal-to-noise ratio of lifespan assessment can be improved.

[0052] After obtaining the extreme value sequence, the effective stress time series is extracted using the rainflow counting method, transforming it into an equivalent alternating stress cyclic spectrum containing different means and amplitudes. The physical process of the rainflow counting method is analogous to the flow of raindrops on the eaves of a multi-story tower: using the peaks and valleys of the time series as eaves, "raindrops" fall from each extreme point, following specific stopping conditions (such as encountering a peak larger than the initial extreme value or encountering the trajectory of rainwater flowing down from above). This precisely decomposes the complex and irregular random power stress history (in this power grid scenario, the stress is specifically physically represented as the amplitude of active power fluctuations in equipment or the depth of fluctuations in the state of charge of energy storage) into a series of complete hysteresis closed loops (full cycles) and divergent semi-cycles. These extracted cyclic events collectively constitute an equivalent alternating stress cyclic spectrum with clear stress amplitude and mean characteristics. After obtaining the cyclic spectrum, the dynamic reduction of remaining available capacity includes: calculating the single fatigue damage degree corresponding to the equivalent alternating stress cyclic spectrum within the current scheduling cycle based on Miner's linear fatigue cumulative damage theory; here, Miner's rule is used to linearly superimpose fatigue damage at different amplitudes, and the formula for calculating the single fatigue damage degree is as follows: (14) In the formula, This represents the single fatigue damage degree corresponding to the equivalent alternating stress cycle spectrum within the current scheduling cycle. This represents the total number of discrete stress amplitude levels contained in the equivalent alternating stress cycle spectrum extracted by the rainflow counting method; Indicates the first The actual number of stress cycles occurring at each stress amplitude level; Indicates the corresponding device in the 1st The theoretical limit number of cycles at which physical failure occurs under a stress amplitude level; for mechanical equipment, this can be obtained through standard SN fatigue curve calibration; for electrochemical energy storage equipment, it is equivalently mapped to the Wöhler conversion curve (Wöhler curve) of "discharge depth / power amplitude - cycle life" in its factory parameters.

[0053] Subsequently, the single fatigue damage degree is accumulated into the historical damage record of the corresponding equipment to obtain the cumulative fatigue damage degree. Let the historical damage record of the corresponding equipment at the end of the previous scheduling cycle be... The cumulative fatigue damage is then... Next, according to the preset nonlinear decay model, the initial ramp-up capacity of the distributed resource is reduced based on the cumulative fatigue damage to obtain the remaining available capacity. Considering that the performance of electrochemical energy storage or mechanical equipment often exhibits accelerated nonlinear degradation at the end of its lifespan, the remaining available capacity is specifically expressed by the following formula: (15) In the formula, This represents the remaining available capacity after dynamic reduction; This indicates the pre-determined initial ramp-up capacity of the distributed resource (i.e., the nameplate capacity or admission capacity before any aging occurs). Indicates the cumulative degree of fatigue damage; This represents the fatigue derating rate coefficient determined by the physical properties of the equipment materials. A nonlinear shape factor (typically) representing the steepness of the capacity decay trajectory. (To characterize the accelerated decay effect at the end of its lifespan). For example, when an electrochemical energy storage station frequently participates in high-frequency AGC response, it leads to cumulative fatigue damage. When the ramp rate reaches 20%, due to the nonlinear increase in internal ohmic resistance, the remaining available ramp capacity calculated according to this model will decrease. This could potentially be reduced by more than 35%, thus strictly preventing the issuance of instructions exceeding its actual physical carrying capacity during subsequent scheduling. To further verify the evolutionary pattern of the dynamic reduction mechanism, Figure 3 This paper presents a three-dimensional histogram of the equivalent alternating stress cyclic spectrum using the rainflow counting method for a 5MW electrochemical energy storage power station after 720 consecutive high-frequency regulation cycles. Figure 3 As shown, the X-axis represents the mean stress, the Y-axis represents the stress amplitude, and the Z-axis represents the number of closed-loop cycles. This graph clearly shows that the equipment was subjected to a large number of low-to-medium amplitude, extremely high frequency fatigue impacts during actual operation.

[0054] Finally, the results are solidified on the blockchain. The generation of the digital capability certificate includes: sending the remaining available capacity and the corresponding device's adjustment rate and adjustment precision to the distributed authentication node; and using the blockchain smart contract of the distributed authentication node to digitally sign and encrypt the remaining available capacity, the adjustment rate, and the adjustment precision to generate the digital capability certificate. The specific on-chain data packets are encapsulated in a JSON structured format, which, in addition to containing the aforementioned core performance parameters, also integrates the device's unique MAC physical address identifier and timestamp. This underlying node data is automatically verified by smart contracts in blockchain frameworks such as the Ethereum Virtual Machine (EVM) or Hyperledger Fabric, signed with a private key using an asymmetric encryption algorithm, and packaged into blocks of the distributed ledger using the SHA-256 hash algorithm to generate a unique data fingerprint.

[0055] This step effectively solves the industry pain points of traditional virtual power plants, such as falsely labeled static capacity and the exhaustion of frequency regulation capability after long-cycle high-frequency response of physical equipment, by dynamically converting the mechanical fatigue life loss of materials into usable capacity reduction and combining the decentralized and tamper-proof characteristics of blockchain for on-chain solidification. It achieves unexpected technical effects from static physical capacity assessment to high-frequency dynamic life cycle tracking closed loop.

[0056] Example 2, Figure 4 A dynamic approval system for the ramp-up capability of virtual power plants based on AGC performance indicators is presented, including: The feature extraction module is used to reconstruct the phase space of historical operating time series data and meteorological series data of distributed resources, and extract feature parameters to generate a basic climbing feature set. The manifold allocation module is used to combine the basic ramp feature set, communication response delay time and overall ramp requirements to generate a multi-source node power allocation manifold based on optimal transmission theory. The strategy generation module is used to transform the target AGC comprehensive performance index requirements into a dynamic repulsive potential field based on the control barrier function, and to solve the optimal response strategy within the manifold using a multi-agent reinforcement learning algorithm. The dynamic reduction module is used to calculate the fatigue damage degree of the power fluctuation data for implementing the strategy using the rainflow counting method, and dynamically reduce the remaining available capacity.

[0057] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0058] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0059] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0060] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0061] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0062] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for dynamically verifying the ramp-up capability of a virtual power plant based on AGC performance indicators, characterized in that, Includes the following steps: Phase space reconstruction is performed on the historical operating time series data and meteorological series data of distributed resources, and feature parameters are extracted to generate a basic climbing feature set. By combining the basic ramping feature set, communication response delay time and overall ramping requirements, a multi-source node power allocation manifold is generated based on optimal transmission theory. Based on the control barrier function, the target AGC comprehensive performance index requirement is transformed into a dynamic repulsive potential field, and the optimal response strategy is solved within the manifold using a multi-agent reinforcement learning algorithm. The fatigue damage degree of the power fluctuation data for implementing the strategy is calculated using the rainflow counting method, and the remaining available capacity is dynamically reduced.

2. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 1, characterized in that, The process of reconstructing the phase space of historical operating time series data and meteorological series data of distributed resources, and extracting feature parameters to generate a basic climbing feature set includes: Determine the delay time and embedding dimension of the historical running time series data; Based on the delay time and embedding dimension, historical runtime time series data are mapped to a high-dimensional topological phase space; The maximum Lyapunov exponent of each evolution trajectory in the high-dimensional topological phase space is calculated and used as the characteristic parameter.

3. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 2, characterized in that, After calculating the maximum Lyapunov exponent of each evolution trajectory in the high-dimensional topological phase space, the method further includes: Determine whether the maximum Lyapunov exponent is greater than zero; If the value is greater than zero, the time series of the corresponding device is determined to have chaotic characteristics, and the feature parameter is added to the basic climbing feature set. If the value is less than or equal to zero, the corresponding device is determined not to have high-frequency transient ramping potential and will not be included in the basic ramping feature set.

4. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 1, characterized in that, The process of combining the basic ramp feature set, communication response delay time, and overall ramp requirements to generate a multi-source node power allocation manifold based on optimal transmission theory includes: The feature parameters are weighted with the communication response delay time to obtain the AGC multi-source contribution weight of each distributed resource; wherein, the feature parameters are positively correlated with the AGC multi-source contribution weight, and the communication response delay time is negatively correlated with the AGC multi-source contribution weight. With the goal of minimizing the power transmission cost of all distributed resources, and with the overall ramping demand and the multi-source contribution weight of AGC as constraints, an optimal transmission objective function based on Wasserstein distance is constructed. Solve for the optimal transmission objective function to obtain the power allocation manifold of the multi-source nodes that satisfies the minimum power transmission cost.

5. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 4, characterized in that, Solving the optimal transmission objective function to obtain the multi-source node power allocation manifold that satisfies the minimum power transmission cost includes: By introducing an entropy regularization term into the optimal transmission objective function, the optimal transmission objective function based on Wasserstein distance is transformed into a strictly convex optimization problem. The Sinkhorn alternating projection algorithm is used to iteratively solve the strictly convex optimization problem until the preset convergence condition is met. The converged transmission plan matrix is ​​used as the optimal power allocation strategy to generate a multi-source node power allocation manifold.

6. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 1, characterized in that, Solving for the optimal response policy within the manifold using a multi-agent reinforcement learning algorithm includes: Analyze the target AGC comprehensive performance index requirements and extract parameters such as adjustment rate, adjustment accuracy and response time; The physical safety boundary in the Hamilton-Jacobi reachability control model is determined based on the parameters of regulation rate, regulation accuracy and response time, and a continuous-time control barrier function is constructed. The control obstacle function is used as a hard constraint layer and embedded into the action output of a multi-agent reinforcement learning algorithm for solution.

7. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 6, characterized in that, The algorithm is embedded in the action output of a multi-agent reinforcement learning algorithm for solving, including: The multi-agent reinforcement learning algorithm is a deep dual-Q network reinforcement learning algorithm. Real-time acquisition of candidate action instructions output by the deep double-Q network reinforcement learning algorithm; Calculate the control obstacle function value corresponding to the candidate action command; If the control barrier function value is less than zero, a dynamic repulsion potential field is triggered to perform spatial projection correction on the candidate action command, pull the corrected action command back into the physical safety boundary, and output it as the optimal response strategy. If the value of the control barrier function is greater than or equal to zero, the candidate action command is directly output as the optimal response strategy.

8. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 1, characterized in that, The calculation of fatigue damage degree of power fluctuation data for implementing the strategy using the rainflow counting method includes: Extract the high-frequency power fluctuation curves of each distributed resource in the current scheduling cycle from the optimal response strategy; Extreme value extraction is performed on the high-frequency power fluctuation curve to filter out invalid fluctuations with amplitudes smaller than a preset dead zone threshold, thereby obtaining an effective stress time series; The effective stress time series is extracted using the rainflow counting method, and the effective stress time series is transformed into an equivalent alternating stress cyclic spectrum containing different mean values ​​and amplitudes.

9. The method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators according to claim 8, characterized in that, The dynamic reduction of remaining available capacity includes: Based on Miner's linear fatigue cumulative damage theory, the single fatigue damage degree corresponding to the equivalent alternating stress cycle spectrum in the current scheduling cycle is calculated. The single fatigue damage degree is accumulated into the historical damage record of the corresponding equipment to obtain the cumulative fatigue damage degree; According to the preset nonlinear attenuation model, the initial ramp capacity of the distributed resource is reduced based on the cumulative fatigue damage degree to obtain the remaining available capacity.

10. A system using the method for dynamic verification of virtual power plant ramp-up capability based on AGC performance indicators as described in any one of claims 1-9, characterized in that, include: The feature extraction module is used to reconstruct the phase space of historical operating time series data and meteorological series data of distributed resources, and extract feature parameters to generate a basic climbing feature set. The manifold allocation module is used to combine the basic ramp feature set, communication response delay time and overall ramp requirements to generate a multi-source node power allocation manifold based on optimal transmission theory. The strategy generation module is used to transform the target AGC comprehensive performance index requirements into a dynamic repulsive potential field based on the control barrier function, and to solve the optimal response strategy within the manifold using a multi-agent reinforcement learning algorithm. The dynamic reduction module is used to calculate the fatigue damage degree of the power fluctuation data for implementing the strategy using the rainflow counting method, and dynamically reduce the remaining available capacity.