A virtual power plant distributed collaborative control method and system based on multi-agent game and impedance perception

By using real-time monitoring and online impedance identification, a two-level game model is constructed to optimize the distributed collaborative control of the virtual power plant. This solves the problems of scheduling capability assessment bias and insufficient adaptive adjustment in traditional methods, and improves the system's economy and flexibility.

CN121367330BActive Publication Date: 2026-05-15ZHEJIANG JIFENG ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG JIFENG ENERGY TECH CO LTD
Filing Date
2025-12-23
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional centralized control methods are difficult to adapt to the dynamic aggregation and coordination needs of massive heterogeneous distributed units. Existing distributed collaborative control methods do not fully consider the dynamic impedance characteristics and equipment aging of energy storage units, resulting in deviations in scheduling capability assessment and a lack of adaptive adjustment strategies for real-time response to game states.

Method used

Real-time monitoring of the state of charge of distributed energy storage units and electrical parameters at grid connection points within a virtual power plant; generation of dispatchable potential profiles through online impedance identification algorithms; construction of a two-layer non-cooperative game model; optimization of decision-making strategies using a multi-agent reinforcement learning framework; and generation of adaptive cooperative control commands.

Benefits of technology

It achieves accurate profiling of the dispatchable potential of distributed energy storage units within a virtual power plant and efficient solution of game equilibrium, thereby improving the system's economy, stability, and response flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121367330B_ABST
    Figure CN121367330B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-agent game and impedance perception's virtual power plant distributed collaborative control method and system, method includes: real-time monitoring the state of charge of each distributed energy storage unit in virtual power plant and the electrical parameter of point of common coupling, synchronous acquisition scheduling instruction of upper layer power grid;Based on electrical parameter, state of charge and scheduling instruction, generate the unit schedulable potential image characterized with impedance spectrum;According to unit schedulable potential image, construct the double-layer non-cooperative game model of internal aggregator of virtual power plant and external power grid, output the collaborative scheduling scheme under dynamic Nash equilibrium;Based on collaborative scheduling scheme, the decision strategy of each aggregator is optimized online using multi-agent reinforcement learning framework, and the collaborative control instruction with adaptive capacity is generated. Using the embodiment of the application, the schedulable potential of distributed energy storage unit in virtual power plant can be accurately imaged, the game equilibrium can be efficiently solved, and the online adaptive optimization of control strategy can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual power plant technology, and in particular to a distributed collaborative control method and system for virtual power plants based on multi-agent game theory and impedance sensing. Background Technology

[0002] Virtual power plants, as a key technology for integrating distributed energy storage resources, play a crucial role in improving grid flexibility and promoting renewable energy consumption. Traditional centralized control methods struggle to meet the dynamic aggregation and coordination needs of massive heterogeneous distributed units, particularly when dealing with grid disturbances and participating in multi-level markets, exhibiting issues such as response lag and insufficient scalability. Existing distributed collaborative control methods primarily focus on optimizing power allocation, while neglecting factors such as the dynamic impedance characteristics of each energy storage unit, equipment aging, and changes in operating conditions, leading to biased assessments of the unit's actual dispatchable capacity. Furthermore, in scenarios where aggregators and the grid engage in a power struggle, there is a lack of collaborative mechanisms capable of responding in real-time to the power struggle state and adaptively adjusting strategies. Summary of the Invention

[0003] The purpose of this invention is to provide a distributed collaborative control method and system for virtual power plants based on multi-agent game theory and impedance sensing, in order to overcome the shortcomings of the prior art. This method can accurately profile the dispatchable potential of distributed energy storage units in the virtual power plant, efficiently solve the game equilibrium, and achieve online adaptive optimization of control strategies, thereby improving the overall economic efficiency, stability, and response flexibility of the system.

[0004] One embodiment of this application provides a distributed cooperative control method for a virtual power plant based on multi-agent game theory and impedance sensing, the method comprising:

[0005] Real-time monitoring of the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point, and synchronous collection of dispatch instructions from the upper-level power grid;

[0006] Based on the electrical parameters, the state of charge, and the scheduling command, the power regulation capability and dynamic response characteristics of each energy storage unit are dynamically evaluated through an online impedance identification algorithm, generating a unit dispatchable potential profile characterized by impedance spectrum.

[0007] Based on the schedulable potential profile of the unit, a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid is constructed. The optimal bid and output strategy of each game participant are solved by a distributed iterative algorithm, and a collaborative scheduling scheme under dynamic Nash equilibrium is output.

[0008] Based on the aforementioned collaborative scheduling scheme, the decision-making strategies of each aggregator are optimized online using a multi-agent reinforcement learning framework, and the game behavior is dynamically adjusted by predicting the opponent's strategy, thereby generating adaptive collaborative control instructions.

[0009] Optionally, the real-time monitoring of the state of charge of each distributed energy storage unit within the virtual power plant and the electrical parameters of the grid connection point, and the synchronous collection of dispatch instructions from the upper-level power grid, includes:

[0010] Within the virtual power plant, a status monitoring module is deployed in each distributed energy storage unit. The system collects state-of-charge data in real time through smart meters and collects active and reactive power data through power sensors to generate raw status data streams.

[0011] Deploy electrical parameter monitoring devices at the grid connection point of the virtual power plant to measure voltage amplitude, phase angle, frequency deviation and harmonic content in real time, and generate time-series data of electrical parameters at the grid connection point.

[0012] The system establishes a communication connection with the upper-level power grid energy management system through the power dispatch data network interface, receives active power regulation instructions and voltage regulation instructions issued by the power grid dispatch center in real time, and generates a power grid dispatch instruction dataset.

[0013] Establish a unified time synchronization system to align the timestamps and unify the sampling rates of the original state data stream, the time-series data of electrical parameters at the grid connection point, and the power grid dispatch instruction dataset, and generate synchronous monitoring data packets.

[0014] Optionally, based on the electrical parameters, the state of charge, and the dispatch command, the power regulation capability and dynamic response characteristics of each energy storage unit are dynamically evaluated using an online impedance identification algorithm to generate a unit dispatchable potential profile characterized by impedance spectrum, including:

[0015] Extract the voltage and current timing data of the grid connection point from the synchronous monitoring data packet, and use the frequency domain analysis method to calculate the amplitude and phase of the fundamental and harmonic components to generate a frequency domain electrical parameter feature set;

[0016] Based on the frequency domain electrical parameter feature set, the equivalent impedance amplitude and phase angle at each harmonic frequency are calculated in real time using the recursive least squares method to generate dynamic impedance spectrum data.

[0017] By combining state-of-charge data and grid dispatch command datasets, the power regulation range and response time characteristics of energy storage units under different states of charge are analyzed, and a power regulation capability evaluation matrix is ​​generated.

[0018] Based on dynamic impedance spectrum data and power regulation capability assessment matrix, a multi-dimensional feature vector is constructed to represent the dynamic response characteristics of each energy storage unit, and finally a unit dispatchable potential profile characterized by impedance spectrum is generated.

[0019] Optionally, based on the schedulable potential profile of the unit, a two-layer non-cooperative game model is constructed between the internal aggregator of the virtual power plant and the external power grid. A distributed iterative algorithm is used to solve for the optimal bid and output strategies of each game participant, outputting a collaborative scheduling scheme under dynamic Nash equilibrium, including:

[0020] Based on the unit schedulable potential profile, the set of game participants is defined to include each aggregator and the power grid dispatch center, and the strategy space and payoff function of each participant are determined to generate the set of game participants and the strategy space.

[0021] Based on the set of game participants and the strategy space, a two-layer game framework structure is constructed, with the upper layer consisting of the power grid dispatch center and the virtual power plant, and the lower layer consisting of the aggregators.

[0022] Based on a two-level game framework, the alternating direction multiplier method is used for distributed iterative solution to update the bids and output strategies of each participant and generate a strategy iteration sequence.

[0023] The convergence of the monitoring strategy iteration sequence is determined. When the policy changes of all participants are less than a preset threshold, a dynamic Nash equilibrium is reached, and a collaborative scheduling scheme is output.

[0024] Optionally, based on the cooperative scheduling scheme, the decision-making strategies of each aggregator are optimized online using a multi-agent reinforcement learning framework, and the game behavior is dynamically adjusted by predicting the opponent's strategy to generate adaptive cooperative control instructions, including:

[0025] Based on the collaborative scheduling scheme, historical decision data and environmental state data of each aggregator are collected to construct a multi-agent state-action pair dataset containing state, action, reward and next state;

[0026] A network for predicting opponent policies is trained using a multi-agent state-action dataset to learn the decision-making patterns of other agents and generate an opponent behavior prediction model.

[0027] The opponent behavior prediction model is integrated into a multi-agent reinforcement learning framework. The decision policy network of each aggregator is optimized by the policy gradient method to generate an optimized decision policy network.

[0028] The optimized decision-making strategy network is deployed to the actual operating environment, and output instructions for each aggregator are dynamically generated based on real-time monitoring data, ultimately generating adaptive collaborative control instructions.

[0029] Another embodiment of this application provides a distributed collaborative control system for a virtual power plant based on multi-agent game theory and impedance sensing, the system comprising:

[0030] The monitoring module is used to monitor the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point in real time, and to collect the dispatch instructions of the upper power grid simultaneously.

[0031] The evaluation module is used to dynamically evaluate the power regulation capability and dynamic response characteristics of each energy storage unit based on the electrical parameters, the state of charge and the scheduling command, and generate a unit dispatchable potential profile characterized by impedance spectrum.

[0032] The module is used to construct a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid based on the schedulable potential profile of the unit. It solves the optimal bid and output strategy of each game participant through a distributed iterative algorithm and outputs a collaborative scheduling scheme under dynamic Nash equilibrium.

[0033] The optimization module is used to optimize the decision-making strategies of each aggregator online based on the cooperative scheduling scheme using a multi-agent reinforcement learning framework, and dynamically adjust the game behavior of the network through the opponent's strategy prediction to generate cooperative control instructions with adaptive capabilities.

[0034] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0035] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0036] Compared with existing technologies, this invention provides a distributed collaborative control method for virtual power plants based on multi-agent game theory and impedance sensing. This method monitors the state of charge (SOC) of each distributed energy storage unit within the virtual power plant and the electrical parameters of the grid connection point in real time, while simultaneously collecting dispatch commands from the upper-level power grid. Based on the electrical parameters, SOC, and dispatch commands, it generates a unit dispatchable potential profile characterized by impedance spectrum. According to the unit dispatchable potential profile, it constructs a two-layer non-cooperative game model between the aggregators within the virtual power plant and the external power grid, outputting a collaborative dispatch scheme under dynamic Nash equilibrium. Based on the collaborative dispatch scheme, it utilizes a multi-agent reinforcement learning framework to optimize the decision-making strategies of each aggregator online, and dynamically adjusts the game behavior through opponent strategy prediction, generating adaptive collaborative control commands. This enables accurate profiling of the dispatchable potential of distributed energy storage units within the virtual power plant, efficient solution of game equilibrium, and online adaptive optimization of control strategies, improving the overall system's economy, stability, and response flexibility. Attached Figure Description

[0037] Figure 1The hardware structure block diagram of the computer terminal for a distributed collaborative control method for a virtual power plant based on multi-agent game theory and impedance sensing provided in an embodiment of the present invention.

[0038] Figure 2 A flowchart illustrating a distributed collaborative control method for a virtual power plant based on multi-agent game theory and impedance sensing, provided in an embodiment of the present invention.

[0039] Figure 3 This is a schematic diagram of the structure of a distributed collaborative control system for a virtual power plant based on multi-agent game theory and impedance sensing, provided in an embodiment of the present invention. Detailed Implementation

[0040] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0041] This invention first provides a distributed collaborative control method for a virtual power plant based on multi-agent game theory and impedance sensing. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0042] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a distributed cooperative control method for a virtual power plant based on multi-agent game theory and impedance sensing, provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0043] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, enable the processor to perform any multi-agent game theory and impedance-sensing virtual power plant distributed cooperative control method.

[0044] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0045] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any multi-agent game and impedance-sensing virtual power plant distributed cooperative control method.

[0046] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0047] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0048] See Figure 2 The present invention provides a distributed cooperative control method for a virtual power plant based on multi-agent game theory and impedance sensing, which may include the following steps:

[0049] S201 monitors the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point in real time, and simultaneously collects the dispatch instructions of the upper-level power grid.

[0050] Specifically, a status monitoring module can be deployed in each distributed energy storage unit within the virtual power plant to collect state of charge data in real time through smart meters and active and reactive power data through power sensors to generate raw status data streams.

[0051] This step is the foundational data acquisition stage for the collaborative control of the virtual power plant. Its core is to accurately capture the key operating status of each distributed energy storage unit through a dedicated monitoring module, providing raw data support for subsequent impedance identification and game theory decision-making. The specific implementation method is as follows:

[0052] The status monitoring module is an integrated unit that combines data acquisition, preprocessing, and preliminary transmission functions. Each distributed energy storage unit (such as a lithium battery energy storage cabinet or a vanadium redox flow storage system) is equipped with an independent module. The module communicates with the energy storage unit's battery management system (BMS) via an industrial bus and is also directly connected to the signal output terminals of the smart meter and power sensor. The smart meter is a three-phase smart meter with real-time metering capabilities. Its sampling frequency is set to 1Hz (once per second), and its measurement accuracy is 0.2% (error ≤ ±0.2%). It supports the acquisition of core parameters of the energy storage unit, such as charging and discharging current, voltage, and remaining capacity. The state of charge (SOC) data is calculated by the ratio of the remaining capacity collected by the meter to the rated capacity, using the formula: SOC = (remaining capacity / rated capacity) × 100%. The rated capacity is set according to the energy storage unit's nameplate parameters (e.g., 100kWh), and the remaining capacity is calculated in real-time by the meter using the current integration method (integration time interval 0.1 seconds). For example, if a certain energy storage unit has a rated capacity of 100kWh and the current meter reads a remaining power of 68.5kWh, then SOC = (68.5 / 100) × 100% = 68.5%. This data includes the data collection timestamp (accurate to milliseconds) and the unit number (e.g., ESS-001).

[0053] The power sensor is a Hall effect type, installed at the grid connection point of the energy storage unit. The sampling frequency is set to 10Hz (sampling once every 0.1 seconds), with a measurement range of -500kW to 500kW (negative values ​​indicate charging, positive values ​​indicate discharging). The measurement accuracy is 0.1%. It simultaneously collects active power (P) and reactive power (Q) data. Active power reflects the actual electrical energy output or absorbed by the energy storage unit, measured in kW; reactive power reflects the energy exchange between the energy storage unit and the grid, measured in kVar. For example, if the sensor collects P=120.3kW and Q=15.7kVar at a certain moment, it means that the energy storage unit is currently outputting 120.3kW of active power to the grid, while simultaneously exchanging 15.7kVar of reactive power with the grid.

[0054] The raw state data stream is generated using a structured format. Each data record contains six core fields: "timestamp-unit number-SOC-active power-reactive power-data quality tag". The timestamp format is YYYY-MM-DDHH:MM:SS.sss, and the data quality tag is used to mark data validity (0 indicates normal, 1 indicates abnormal). For example, a complete data record is "2025-09-0108:30:00.123-ESS-001-68.5%-120.3kW-15.7kVar-0", where the data quality tag is 0, indicating that the data acquisition was normal and there was no abnormal interference. All data is transmitted in real time to the virtual power plant's local data acquisition server via industrial Ethernet, and a circular buffer is used for storage, retaining the raw data of the most recent 24 hours to ensure data real-time performance and traceability.

[0055] Deploy electrical parameter monitoring devices at the grid connection point of the virtual power plant to measure voltage amplitude, phase angle, frequency deviation and harmonic content in real time, and generate time-series data of electrical parameters at the grid connection point.

[0056] This step focuses on monitoring the electrical condition of the virtual power plant's connection nodes with the upper-level power grid. High-precision monitoring devices capture the dynamic changes of key electrical parameters, providing a basis for assessing the grid's operational status and the energy storage unit's adjustment needs. The specific implementation method is as follows:

[0057] The grid connection point electrical parameter monitoring device is an integrated unit combining a synchronous phasor measurement unit (PMU) and a harmonic analyzer. It is deployed within the switchgear of the high-voltage side of the virtual power plant's grid connection point, at a distance of ≤3 meters from the grid connection point circuit breaker, ensuring the authenticity and stability of the measurement signals. The device's sampling frequency is set to 50Hz (synchronized with the grid frequency), and the data output frequency is 10Hz (outputting one set of effective data every 0.1 seconds). The measurement accuracy meets the requirements of the IEC61850 standard, specifically ±0.1% for voltage amplitude measurement, ±0.1° for phase angle measurement, ±0.001Hz for frequency deviation measurement, and ±0.5% for harmonic content measurement.

[0058] Voltage amplitude refers to the effective value of the three-phase voltage at the grid connection point, in kV. The measurement range is from 0kV to 35kV (adapted to the grid connection point voltage level of medium-voltage virtual power plants). For example, if the collected voltage amplitudes are 10.5kV for phase A, 10.48kV for phase B, and 10.52kV for phase C, the three-phase voltage imbalance is (10.52-10.48) / 10.5×100%≈0.38%, which meets the grid operation standards. Phase angle refers to the phase difference between the three-phase voltages. With phase A voltage as the reference, the phase angle of phase B is usually -120°, and phase C is 120°. The device measures the deviation between the actual phase angle and the reference value in real time. For example, if the measured phase angle of phase B is -119.8°, the deviation is 0.2°, reflecting the phase synchronization of the three-phase voltages.

[0059] Frequency deviation refers to the difference between the actual frequency at the grid connection point and the rated frequency (50Hz), measured in Hz. For example, if the actual measured frequency is 50.02Hz, the frequency deviation is +0.02Hz. If the deviation exceeds ±0.2Hz, it may trigger a grid dispatching command for adjustment. Harmonic content refers to the amplitude percentage of each harmonic in the voltage signal other than the fundamental frequency (50Hz). The monitoring range is from the 2nd to the 50th harmonic. For example, if the measured 3rd harmonic content is 1.2% and the 5th harmonic content is 0.8%, the total harmonic distortion (THD) is √(1.2%² + 0.8%²) ≈ 1.44%, which meets the requirements for grid harmonic control (total harmonic distortion ≤ 5%).

[0060] The time-series data of electrical parameters at the grid connection point are generated in the structure of "timestamp-voltage amplitude (A / B / C phase)-phase angle (A / B / C phase)-frequency deviation-harmonic content-total harmonic distortion rate". The timestamp is consistent with the original state data stream (accurate to milliseconds) to ensure data time synchronization. For example, a time-series data record is "2025-09-01 08:30:00.123-10.5kV / 10.48kV / 10.52kV-0° / -119.8° / 120.1°-+0.02Hz-3rd harmonic 1.2% / 5th harmonic 0.8% / 7th harmonic 0.3%-1.44%". The data is stored continuously in chronological order to form a time-series data sequence, providing complete electrical parameter input for subsequent online impedance identification.

[0061] The system establishes a communication connection with the upper-level power grid energy management system through the power dispatch data network interface, receives active power regulation instructions and voltage regulation instructions issued by the power grid dispatch center in real time, and generates a power grid dispatch instruction dataset.

[0062] This step is a crucial communication link for enabling collaboration between the virtual power plant and the upper-level power grid. It receives grid dispatch instructions through standardized interfaces and protocols, clarifies the grid's regulation requirements for the virtual power plant, and provides a basis for subsequent collaborative dispatch scheme development. The specific implementation method is as follows:

[0063] The power dispatch data network interface adopts a fiber optic Ethernet interface based on the IEC61850 standard, deployed in the dispatch control center of the virtual power plant. The interface transmission rate is 1000Mbps, and the communication latency is ≤50ms, meeting the real-time requirements of dispatch commands. The interface supports bidirectional communication, capable of receiving dispatch commands issued by the power grid energy management system (EMS) and also feeding back the real-time operating status of the virtual power plant (such as total output, average SOC, etc.) to the power grid. The communication process uses an encrypted transmission protocol (TLS 1.3) and verifies the legitimacy of the command source through a digital signature mechanism to prevent commands from being tampered with or forged.

[0064] Power grid dispatch instructions mainly include two categories: active power regulation instructions and voltage regulation instructions, both encapsulated using the standardized IEC 61850-9-2 protocol format. Active power regulation instructions specify the total active power output target or regulation amount of the virtual power plant, in kW. Instruction parameters include the target value, regulation time limit, and regulation accuracy requirements. For example, "Active power regulation instruction: target output 800kW, regulation time limit 30 seconds, regulation accuracy ±5kW" means that the virtual power plant is required to adjust its total active power to 800kW within 30 seconds, and the actual output deviation from the target value must not exceed ±5kW. If the instruction is in the form of a regulation amount, such as "active power increase by 50kW," then the virtual power plant needs to increase its current total output by 50kW.

[0065] The voltage regulation command specifies the voltage control target at the grid connection point, in kV. Command parameters include the target voltage value, allowable voltage deviation range, and regulation response time. For example, "Voltage Regulation Command: Target voltage 10.5kV, allowable deviation range ±0.1kV, response time 20 seconds" indicates that the virtual power plant is required to stabilize the grid connection point voltage between 10.4kV and 10.6kV by adjusting the reactive power output of each energy storage unit, and must reach the target range within 20 seconds. In addition, the dispatch command also includes auxiliary information such as command number, issuance time, and validity period. For example, command number "DISP-20250901083000", issuance time "2025-09-01 08:30:00", and validity period "1 hour". The command automatically expires after the validity period.

[0066] The power grid dispatch instruction dataset is generated using a structured storage method. Each instruction record contains eight core fields: "Instruction Number - Issuance Time - Instruction Type - Target Parameter - Adjustment Time Limit - Accuracy Requirement - Validity Period - Instruction Status". The instruction status includes four categories: "Not Executed", "Executing", "Completed", and "Expired". For example, an instruction record might be "DISP-20250901083000-2025-09-0108:30:00-Active Power Adjustment-Target Output 800kW-30 seconds-±5kW-1 hour-Not Executed". The dataset is updated in real time. When a virtual power plant completes the instruction adjustment, the instruction status is automatically updated to "Completed", and the actual execution result (e.g., actual output 798.5kW, deviation 1.5kW) is recorded, providing a basis for subsequent dispatch effect evaluation.

[0067] Establish a unified time synchronization system to align the timestamps and unify the sampling rates of the original state data stream, the time-series data of electrical parameters at the grid connection point, and the power grid dispatch instruction dataset, and generate synchronous monitoring data packets.

[0068] This step is the core of achieving multi-source data fusion. By unifying the time base, it eliminates time deviations from different data sources, and at the same time, it standardizes the sampling rate to ensure the spatiotemporal consistency of the data. This provides regular input data for subsequent steps such as impedance identification and game theory modeling. The specific implementation method is as follows:

[0069] The unified time synchronization system is built on the Precision Time Protocol (PTP, IEEE 1588-2008), using the BeiDou satellite time server of the power grid dispatch center as the time reference source, achieving a time accuracy of nanoseconds (±100ns). All data acquisition devices within the virtual power plant (status monitoring modules, grid connection point monitoring devices, and dispatch interface devices) support the PTP protocol and are connected to the time synchronization network via industrial Ethernet. They synchronize with the reference source every 10 seconds to ensure that the deviation between the local clock of each device and the reference time does not exceed 1ms. During synchronization, the devices automatically correct for local clock drift. For example, if the local clock of a status monitoring module deviates from the reference time by 0.5ms, it will automatically adjust to be completely consistent after synchronization.

[0070] The specific operation of timestamp alignment is as follows: Using the timestamp of the base time as an index, a unified time axis (interval of 0.1ms) is constructed. The original state data stream (sampling frequency 1Hz / 10Hz), the time-series data of electrical parameters at the grid connection point (sampling frequency 10Hz), and the power grid dispatch instruction dataset (triggered issuance) are associated and matched according to their timestamps. For signals with a sampling frequency higher than 0.1ms (e.g., 10Hz corresponding to a 100ms interval), resampling is performed along the time axis, taking one data point every 0.1ms, and supplementing missing data through linear interpolation. For signals with a sampling frequency lower than 0.1ms (e.g., 1Hz corresponding to a 1000ms interval), the value of the data point is filled into 10 consecutive 0.1ms intervals on the corresponding time axis. For triggered data such as dispatch instructions, the instruction issuance timestamp is mapped to the unified time axis, and the instruction parameters remain valid within their validity period until the next instruction overrides them or the validity period ends.

[0071] The sampling rate is uniformly aligned with the highest sampling rate. All data is resampled using the highest sampling rate (10Hz, corresponding to a 0.1-second interval) from each data source. The 1Hz sampling data from the smart meters in the original state data stream is upsampled to 10Hz using linear interpolation. For example, if an energy storage unit has a SOC of 68.5% at 08:30:00 and 68.3% at 08:30:01, then the interpolated SOC data from 08:30:00.1 to 08:30:00.9 will be 68.48%, 68.46%, ..., 68.32%, respectively, ensuring a uniform sampling rate. During resampling, the accuracy of the original data is strictly preserved, with interpolation errors controlled within ±0.1% to avoid data distortion.

[0072] The generation of synchronous monitoring data packets adopts an integrated structured format. Each data packet record contains four core parts: "unified timestamp - energy storage unit status data (multi-unit aggregation) - grid connection point electrical parameters - currently valid dispatch instructions". The unified timestamp is accurate to 0.1ms (format YYYY-MM-DDHH:MM:SS.sss). For example, a synchronous monitoring data packet might be: "2025-09-01 08:30:00.120-ESS-001: SOC 68.5%, P 120.3kW, Q 15.7kVar; ESS-002: SOC 72.1%, P 95.8kW, Q 12.3kVar-Voltage 10.5kV / 10.48kV / 10.52kV, Phase Angle 0° / -119.8° / 120.1°, Frequency Deviation +0.02Hz, THD 1.44%-Active Power Regulation: Target 800kW, Time Limit 30 seconds, Accuracy ±5kW". These data packets are generated sequentially and transmitted to the core control server of the virtual power plant. They serve as the unified data input for subsequent online impedance identification, two-layer game modeling, and multi-agent optimization, ensuring the consistency and coordination of data across all stages.

[0073] S202, based on the electrical parameters, the state of charge and the scheduling command, the power regulation capability and dynamic response characteristics of each energy storage unit are dynamically evaluated through the impedance online identification algorithm, and a unit scheduling potential profile characterized by impedance spectrum is generated.

[0074] Specifically, the voltage and current timing data at the grid connection point can be extracted from the synchronous monitoring data packet, and the amplitude and phase of the fundamental and harmonic components can be calculated using frequency domain analysis methods to generate a frequency domain electrical parameter feature set.

[0075] This step is the preliminary data processing stage for online impedance identification. Its core is to extract the voltage and current signals reflecting the electrical characteristics of the energy storage unit from the synchronous monitoring data, and to decompose the fundamental and harmonic components through frequency domain analysis, providing accurate frequency domain characteristic input for subsequent impedance calculations. The specific implementation method is as follows:

[0076] The synchronous monitoring data packet is a structured collection containing multi-source data. Among them, the time-series data of the three-phase voltage (u_A, u_B, u_C) at the grid connection point and the branch current (i_ESS1, i_ESS2, ..., i_ESSn) of each energy storage unit are directly related to impedance identification. The data format is "unified timestamp-instantaneous voltage value-instantaneous current value", with a sampling frequency of 10Hz and a timestamp accuracy of 0.1ms. During the extraction process, the system filters data according to the logic of "unit number-phase sequence-time sequence". For example, when extracting the A-phase current time sequence data of energy storage unit ESS-001, the filtering condition is "unit number = ESS-001 and phase sequence = A", which yields the instantaneous current values ​​(in A) at 100 consecutive time points: [120.1, 121.3, 119.8, ..., 120.5]. The corresponding voltage time sequence data is the instantaneous voltage value of phase A (in kV): [10.52, 10.51, 10.53, ..., 10.52]. This ensures that the timestamps of the voltage and current data are completely aligned (deviation ≤ 0.1ms).

[0077] The frequency domain analysis method employs Fast Fourier Transform (FFT), which essentially converts the instantaneous voltage and current signals in the time domain into amplitude-frequency and phase-frequency characteristics in the frequency domain. This allows for the separation of the fundamental frequency (50Hz) from various harmonics (2nd order 100Hz, 3rd order 150Hz, ..., 5th order 250Hz; the harmonic analysis range is 2nd-5th, covering common harmonic frequency bands in a virtual power plant). Before FFT calculation, the time-series data needs preprocessing: first, a Hanning window is used to suppress spectral leakage (window function length N=1024, satisfying an integer power of 2 to improve computational efficiency); then, zero-padding is performed on the data (filling 100 data points to 1024, ensuring a frequency resolution Δf = sampling frequency / N = 10Hz / 1024 ≈ 0.0098Hz, improving frequency identification accuracy).

[0078] Taking the A-phase voltage signal of ESS-001 as an example, the FFT calculation process is as follows: Let the time-domain voltage signal be u(t), and after FFT transformation, the frequency domain complex sequence U(k) = U_m(k) × e^(jφ(k)) is obtained, where U_m(k) is the amplitude at the k-th frequency point, and φ(k) is the phase angle. By traversing the frequency points, the fundamental frequency (50Hz) with the largest amplitude is found, with an amplitude of U_m1 = 10.5kV (effective value, calculated by dividing the time domain peak value by √2) and a phase of φ1 = 0° (based on the phase of phase A voltage); the amplitude of the second harmonic (100Hz) is U_m2 = 0.12kV, and the phase of φ2 = -30°; the amplitude of the third harmonic (150Hz) is U_m3 = 0.08kV, and the phase of φ3 = 60°; the amplitude of the fourth harmonic (200Hz) is U_m4 = 0.05kV, and the phase of φ4 = -45°; and the amplitude of the fifth harmonic (250Hz) is U_m5 = 0.03kV, and the phase of φ5 = 30°. The FFT calculation process for current signals is the same as that for voltage. For example, the fundamental amplitude of phase A current is I_m1=120A, the phase is φ_i1=-30°, the amplitude of the second harmonic is I_m2=1.5A, and the phase is φ_i2=-60°.

[0079] The frequency domain electrical parameter feature set is generated in a structured format of "energy storage unit number - phase sequence - harmonic order - voltage amplitude - voltage phase - current amplitude - current phase", covering all monitored energy storage units and three-phase data. For example, the A-phase feature set segment of ESS-001 is "ESS-001-A-1 (fundamental) -10.5kV -0° -120A -30°; ESS-001-A-2 (2nd harmonic) -0.12kV -30° -1.5A -60°; ESS-001-A-3 (3rd harmonic) -0.08kV -60° -0.9A -30°". The feature set also labels the data confidence level (based on the signal-to-noise ratio calculated by FFT, SNR≥30dB is 100% confidence level, SNR<20dB is invalid data) to ensure the reliability of subsequent impedance calculations.

[0080] Based on the frequency domain electrical parameter feature set, the equivalent impedance amplitude and phase angle at each harmonic frequency are calculated in real time using the recursive least squares method to generate dynamic impedance spectrum data.

[0081] This step is the core calculation process for online impedance identification. The Recursive Least Squares (RLS) method, with its advantages of "real-time updates, forgetting old data, and tracking dynamic changes," adapts to the characteristics of energy storage unit impedance changing with operating conditions, accurately calculating the equivalent impedance at each harmonic frequency. The specific implementation method is as follows:

[0082] The physical meaning of equivalent impedance is the opposition of an energy storage unit to current at a specific frequency. Its core calculation formula is Z(k) = U(k) / I(k), where U(k) and I(k) are the complex components of voltage and current of the k-th harmonic, respectively. The impedance Z(k) is also expressed in complex form as Z(k) = Z_m(k) × e^(jθ(k)), where Z_m(k) is the impedance amplitude (unit Ω), and θ(k) is the impedance phase angle (the phase difference between voltage and current, θ(k) = φ_u(k) - φ_i(k), a positive phase angle indicates inductive, and a negative phase angle indicates capacitive). The role of the RLS algorithm is to optimize the impedance calculation results through recursive iteration, reducing the impact of measurement noise (such as sensor error and power grid interference) on the calculation accuracy.

[0083] The core parameters of the RLS algorithm include the forgetting factor λ, the initial covariance matrix P(0), and the initial impedance estimate Z(0). The forgetting factor λ is set to 0.98, which gives new data a higher weight and the weight of old data decreases with iteration (the closer λ is to 1, the slower the forgetting of old data, which is suitable for slowly changing systems; the smaller λ is, the stronger the ability to track dynamic changes); the initial covariance matrix P(0) is set to 10^4×I (I is the identity matrix), which indicates that the uncertainty of the initial impedance estimate is relatively large; the initial impedance estimate Z(0) is set based on the nameplate parameters of the energy storage unit (e.g., the fundamental initial impedance of the lithium battery energy storage unit is set to 87.5Ω, which is calculated from the rated voltage of 10.5kV / rated current of 120A).

[0084] The recursive process of the RLS algorithm is taken as an example of the third harmonic: Let the input of the t-th iteration be the complex component of current I(t) and the output be the complex component of voltage U(t). Then the impedance estimate Z(t) = Z(t-1) + K(t) × [U(t) - I(t) × Z(t-1)], where K(t) is the gain matrix, K(t) = P(t-1) × I*(t) / [λ + I(t) × P(t-1) × I*(t)], I*(t) is the conjugate complex number of I(t), and the covariance matrix update formula is P(t) = [P(t-1) - K(t) × I(t) × P(t-1)] / λ. For example, at t=1, I(1)=0.9A (amplitude of the 3rd harmonic current) and phase 30°, i.e., I(1)=0.9×(cos30°+jsin30°)=0.779+j0.45; U(1)=0.08kV=80V and phase 60°, i.e., U(1)=80×(cos60°+jsin60°)=40+j69.28; substituting these values, we get K(1)=10^4×(0.779-j0.45) / [0.98+(0. 779+j0.45)×10^4×(0.779-j0.45)]≈98.5; Z(1)=87.5+98.5×[80+j69.28-(0.779+j0.45)×87.5]≈88.9+j51.2Ω, impedance amplitude Z_m=√(88.9²+51.2²)≈102.5Ω, phase angle θ=arctan(51.2 / 88.9)≈30°, consistent with the theoretical phase difference (60°-30°).

[0085] Through the above iterative process, the equivalent impedance amplitude and phase angle of each energy storage unit at the fundamental (50Hz), second (100Hz), third (150Hz), fourth (200Hz), and fifth (250Hz) harmonic frequencies are calculated respectively, and dynamic impedance spectrum data are generated. The dynamic impedance spectrum uses "frequency-impedance amplitude-impedance phase angle-time stamp" as its core dimensions, continuously recording impedance changes in a time series. For example, the impedance spectrum data of ESS-001 at t=08:30:00 is "50Hz-87.5Ω-30°; 100Hz-92.3Ω-28°; 150Hz-102.5Ω-30°; 200Hz-110.1Ω-25°; 250Hz-118.7Ω-22°", and the impedance spectrum data at t=08:30:01 is "50Hz-87.8Ω-30°; 100Hz-92.1Ω-27°; 150Hz-102.3Ω-29°; 200Hz-110.3Ω-26°; 250Hz-118.5Ω-23°", reflecting the dynamic change characteristics of impedance.

[0086] The dynamic impedance spectrum data contains health status markers at the same time. By comparing the percentage deviation of the current impedance from the initial impedance (ΔZ% = |Z_current - Z_initial| / Z_initial × 100%), the state of the energy storage unit is judged: ΔZ% ≤ 5% is "healthy", 5% < ΔZ% ≤ 15% is "sub - healthy", and ΔZ% > 15% is "abnormal". For example, the initial value of the fundamental wave impedance of ESS - 001 is 87.5 Ω, and the current value is 90.0 Ω, ΔZ% = 2.86%, which is marked as "healthy"; if the current value of the fundamental wave impedance of a certain unit is 101.6 Ω and ΔZ% = 16%, it is marked as "abnormal", indicating that there may be problems such as battery aging or poor contact.

[0087] Combined with the state - of - charge data and the power grid dispatching instruction dataset, analyze the power regulation range and response time characteristics of the energy storage unit under different states of charge, and generate a power regulation ability evaluation matrix;

[0088] This step is a key link in combining impedance characteristics with operating status and dispatching requirements. By quantifying the constraints of the state of charge (SOC) on power regulation and the requirements of dispatching instructions for regulation targets, a quantitative evaluation of the regulation ability of the energy storage unit is formed. The specific implementation method is as follows:

[0089] The state - of - charge data is extracted from the synchronous monitoring data packet. The sampling frequency of the SOC data of each energy storage unit is 1 Hz, which reflects the relative level of the remaining battery power (0% - 100%). Combined with the rated capacity of the energy storage unit (such as the rated capacity of ESS - 001 is 100 kWh), the actual remaining power can be calculated (remaining power = SOC × rated capacity). The constraints of SOC on power regulation ability are determined based on the charge - discharge characteristics of the battery and are divided into five intervals: SOC ≤ 20% (low - power interval), 20% < SOC ≤ 40% (medium - low power), 40% < SOC ≤ 80% (optimal power), 80% < SOC ≤ 90% (medium - high power), SOC > 90% (high - power interval). The charge - discharge power limits in different intervals are different - the regulation range in the optimal power interval is the largest, and the regulation ability is limited in the low / high - power intervals to protect the battery.

[0090] Taking ESS-001 (rated charge-discharge power ±200kW, rated capacity 100kWh) as an example, the power adjustment range for each SOC interval is set as follows: when SOC≤20%, the discharge power limit is 0-50kW (to avoid over-discharge), and the charge power limit is 100-150kW (allowing balanced charging); when 20%<SOC≤40%, the discharge power is 0-100kW, and the charge power is 80-180kW; when 40%<SOC≤80%, the discharge power is 0-200kW, and the charge power is 0-200kW (full adjustment range); when 80%<SOC≤90%, the discharge power is 50-200kW, and the charge power is 0-80kW (to avoid over-charging); when SOC>90%, the discharge power is 100-200kW, and the charge power is 0-50kW. If the current SOC of ESS-001 is 65% (optimal interval), then its active power adjustment range is -200kW (charging) to 200kW (discharging), and the reactive power adjustment range is determined based on the impedance characteristic to be -50kVar to 50kVar (inductive reactive power is positive, capacitive is negative).

[0091] The dynamic response characteristic is measured through a step response experiment, that is, sending an active power step command (such as suddenly increasing from 100kW to 150kW) to the energy storage unit, recording the time (response time t_r) from the command issuance to the power stabilizing within the range of ±5% of the target value, and the maximum overshoot of the power fluctuation (σ%=(maximum peak - target value) / target value×100%). The response time reflects the adjustment speed, and the overshoot reflects the adjustment stability. For example, when the SOC of ESS-001 is 65%, for the step command 100kW→150kW, the response time t_r = 0.8 seconds, and the overshoot σ% = 3.2%; when the SOC = 25% (medium to low battery level), the response time t_r = 1.2 seconds, and the overshoot σ% = 5.1%, which reflects the influence of SOC on the response characteristic.

[0092] The power grid dispatching instruction dataset defines the target boundary for the regulation capacity. For example, if the dispatching instruction requires "the total active power of the virtual power plant to be adjusted to 800 kW, the adjustment time limit is 30 seconds, and the accuracy is ±5 kW", then the regulation capacity of each energy storage unit needs to meet "the sum of the adjustment amounts of all units = 800 kW - the current total output", and the adjustment amount of a single unit shall not exceed the limit of its own SOC range. Combining the above information, the power regulation capacity evaluation matrix is constructed with dimensions of "energy storage unit number - SOC range - active power adjustment range (kW) - reactive power adjustment range (kVar) - response time (s) - overshoot (%) - dispatching instruction adaptability". The matrix elements are specific quantified values. For example, the evaluation matrix fragment of ESS-001 is "ESS-001 - 20% < SOC ≤ 40% - 0~100 / -80~180 - (-50)~50 - 1.2 - 5.1 - adaptable (adjustment amount ≤ 100 kW)", where the "dispatching instruction adaptability" is determined based on whether the unit adjustment range covers the adjustment amount required by the instruction.

[0093] Based on the dynamic impedance spectrum data and the power regulation capacity evaluation matrix, a multi-dimensional feature vector is constructed to represent the dynamic response characteristics of each energy storage unit, and finally a unit dispatchable potential portrait characterized by the impedance spectrum is generated.

[0094] This step is a comprehensive refinement of the dispatchable potential of the energy storage unit. By constructing a multi-dimensional feature vector, the impedance characteristics and regulation capacity are integrated to form a standardized potential portrait, providing a clear unit capacity input for subsequent game modeling. The specific implementation method is as follows:

[0095] The dimension design of the multi-dimensional feature vector is based on the three core dimensions of "impedance characteristics - regulation capacity - operating state", and it includes a total of 12 feature components, namely: 1. Fundamental wave impedance amplitude (Z_1); 2. Fundamental wave impedance phase angle (θ_1); 3. Second harmonic impedance amplitude (Z_2); 4. Third harmonic impedance amplitude (Z_3); 5. Maximum active power adjustment amount (P_max, taking the maximum value of the absolute value of charge and discharge); 6. Maximum reactive power adjustment amount (Q_max); 7. Dynamic response time (t_r); 8. Power overshoot (σ); 9. Current SOC value; 10. Impedance deviation percentage (ΔZ%); 11. Dispatching instruction adaptability (quantified as 0 - 1, adaptable is 1, partially adaptable is 0.5, not adaptable is 0); 12. Remaining adjustment duration (calculated based on the current SOC and adjustment power, for example, when discharging at 100 kW, the remaining duration = (current SOC - 20%) × rated capacity / 100 kW).

[0096] The construction of eigenvectors requires standardization of each eigencomponent to eliminate dimensional differences. The standardization formula is x_norm=(x-x_min) / (x_max-x_min), where x is the original eigenvalue, and x_min and x_max are the global minimum / maximum values ​​of the eigenvalue (based on historical data statistics from all energy storage units). For example, if the fundamental impedance Z_1 has x_min=80Ω and x_max=120Ω, and Z_1=87.5Ω for ESS-001, then Z_1_norm=(87.5-80) / (120-80)=0.1875; if the response time t_r has x_min=0.5 seconds and x_max=2.0 seconds, and t_r=0.8 seconds for ESS-001, then t_r_norm=(0.8-0.5) / (2.0-0.5)=0.2. For features with a clear physical range, such as phase angle and SOC, standardize directly according to the range (e.g., θ_1∈[-90°,90°], standardize to (θ_1+90°) / 180°).

[0097] The unit's dispatchable potential profile is based on multi-dimensional feature vectors, integrating basic information, dynamic impedance spectrum data, and power regulation capability assessment results into a structured document containing four parts: 1. Basic information (unit number, type, rated parameters, installation time); 2. Core dynamic impedance spectrum data (frequency-impedance amplitude / phase angle table, indicating health status); 3. Power regulation capability assessment (SOC range-regulation range correspondence table, response characteristic parameters); 4. Multi-dimensional feature vectors and comprehensive rating (divided into four levels: A, B, C, and D based on the mean of the feature vectors, with a mean ≥0.8 being A, 0.6-0.8 being B, 0.4-0.6 being C, and <0.4 being D).

[0098] For example, the dispatchable potential profile of ESS-001 is as follows: "Unit Number: ESS-001; Type: Lithium-ion battery energy storage; Rated Capacity: 100kWh; Rated Power: ±200kW; Dynamic Impedance Spectrum (Healthy): 50Hz-87.5Ω / 30°, 100Hz-92.3Ω / 28°, 150Hz-102.5Ω / 30°; Power Regulation Capability: at SOC 40%-80%, P±200kW, Q±50kVar, response time 0.8 seconds; Multidimensional Feature Vector:" [0.1875,0.6667,0.3075,0.5625,1.0,1.0,0.2,0.32,0.65,0.0715,1.0,0.9]; Overall rating: A (mean 0.68); Summary of dispatchable potential: Wide adjustment range, fast response speed, stable impedance characteristics, adaptable to the active / reactive power regulation needs of the power grid, making it a core dispatchable unit. This profile clearly quantifies the dispatchable potential of the energy storage unit, providing a direct basis for determining the strategy space of participants in the subsequent two-layer game model.

[0099] S203. Based on the schedulable potential profile of the unit, a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid is constructed. The optimal bid and output strategy of each game participant are solved by a distributed iterative algorithm, and a collaborative scheduling scheme under dynamic Nash equilibrium is output.

[0100] Specifically, based on the unit's schedulable potential profile, the set of game participants can be defined, including each aggregator and the power grid dispatch center, and the strategy space and payoff function of each participant can be determined to generate the set of game participants and the strategy space.

[0101] This step is fundamental to building the game theory model. Its core is to clarify "who participates in the game," "what actions they can take," and "what goals they pursue." By quantifying the participants' capability boundaries through unit schedulable potential profiles, it provides clear definitions of subjects and rules for subsequent game theory modeling. The specific implementation method is as follows:

[0102] The unit dispatchable potential profile defines the core constraints on the participants' capabilities. The "comprehensive rating," "active / reactive power adjustment range," and "impedance deviation percentage" in the profile directly determine the feasibility of the participants' strategies. First, the set of game participants is defined, clarifying the types and numbers of participants. In the virtual power plant scenario, participants are divided into two categories: one is the "external participant," the power grid dispatch center (denoted as participant G), which is responsible for setting electricity purchase prices and issuing dispatch requirements, and is the dominant party in the game; the other is the "internal participant," each distributed energy storage aggregator (denoted as participants A_1, A_2, ..., A_n, where n is the number of aggregators, and in this example, n=3). Each aggregator is responsible for managing a group of energy storage units (e.g., A_1 manages ESS-001 to ESS-005, and A_2 manages ESS-006 to ESS-010), and participates in the game by adjusting the output and bidding of its units, and is the responder in the game. The participant set is defined in the form of "participant ID-type-management resources-capability constraints", such as "A_1-aggregator-ESS-001~005-active power regulation range 0~800kW, comprehensive rating A" and "G-power grid dispatch center-regional power supply network-electricity purchase budget ≤50,000 yuan / hour".

[0103] A participant's strategy space is the set of all possible actions they can take, which must be combined with capability constraints and game objectives. An aggregator's strategy space comprises two core dimensions: bidding strategy p_i (unit: yuan / kWh) and output strategy P_i (unit: kW). The bidding strategy must satisfy the regional electricity market price range (example set as 0.3~0.8 yuan / kWh), and the output strategy must satisfy the total regulation capacity of its energy storage units (determined by the dispatchable potential profile, such as P_i∈[0,800]kW for A_1, [0,600]kW for A_2, and [0,700]kW for A_3). For example, A_1's strategy space is "p_1∈[0.3,0.8] yuan / kWh, P_1∈[0,800]kW", and a specific strategy is "p_1=0.52 yuan / kWh, P_1=650kW". The strategy space of the power grid dispatch center includes the electricity purchase price strategy p_G (unit: yuan / kWh) and the total dispatch demand strategy P_G (unit: kW). p_G needs to match the aggregator's bid (usually p_G ≥ the aggregator's lowest bid to attract power output). P_G is determined by the power grid load gap (in the example, P_G ∈ [1000, 2000] kW). For example, the specific strategy of the power grid is "p_G = 0.5 yuan / kWh, P_G = 1800 kW".

[0104] The revenue function is a quantitative expression of the participants' objectives, with the core being "revenue = income - cost". The revenue composition differs for different participants. The aggregator's revenue function is U_i = p_i × P_i - C_i(P_i) - L_i(P_i, Z_i), where p_i × P_i is the electricity sales revenue (bid price × output); C_i(P_i) is the operating cost, which has a quadratic function relationship with output: C_i(P_i) = a_i × P_i² + b_i × P_i + c_i (a_i is the cost coefficient, in example A_1 a_1 = 0.0001 yuan / kW²). b_1=0.1 yuan / kW, c_1=50 yuan, reflecting that the greater the output, the higher the marginal cost); Li(P_i,Z_i) is the loss cost, which is related to the output and impedance. Li(P_i,Z_i)=d_i×P_i×Z_i (d_i is the loss coefficient, d_1=0.0005 yuan / (kW・Ω) for A_1, and Z_i is the average fundamental impedance of the unit under the aggregator, Z_1=88Ω for A_1). Using strategy A_1, "p_1=0.52 yuan / kWh, P_1=650kW", the calculation is as follows: U_1=0.52×650-(0.0001×650²+0.1×650+50)-(0.0005×650×88)=338-(42.25+65+50)-(28.6)=338-157.25-28.6=152.15 yuan.

[0105] The revenue function of the power grid dispatch center is U_G=λ×P_G-p_G×P_total-γ×|P_G-P_total|, where λ is the power grid's electricity sales price (example: λ=0.8 yuan / kWh, reflecting the power grid's revenue benchmark); P_total is the sum of the output of each aggregator (P_total=P_1+P_2+P_3); p_G×P_total is the electricity purchase cost; and γ is the supply-demand deviation penalty coefficient (γ=2 yuan / kWh, reflecting the power grid's requirements for power supply reliability; the larger the deviation, the higher the penalty). If the grid strategy is "p_G=0.5 yuan / kWh, P_G=1800kW", and the aggregator's total output P_total=650+550+600=1800kW, then U_G=0.8×1800-0.5×1800-2×|1800-1800|=1440-900-0=540 yuan; if P_total=1750kW, with a deviation of 50kW, then U_G=1440-0.5×1750-2×50=1440-875-100=465 yuan, the penalty significantly reduces the revenue.

[0106] The generated set of game participants and strategy space are presented in a structured document, clearly labeling each participant's ID, type, strategy dimension, value range, and payoff function expression. For example, "A_1: Type = Aggregator, Strategy Space = (p_1∈[0.3,0.8] yuan / kWh, P_1∈[0,800]kW), Payoff Function U_1=0.52P_1-(0.0001P_1²+0.1P_1+50)-0.0005P_1×88", providing a clear basis for the subsequent construction of the game framework.

[0107] Based on the set of game participants and the strategy space, a two-layer game framework structure is constructed, with the upper layer consisting of the power grid dispatch center and the virtual power plant, and the lower layer consisting of the aggregators.

[0108] This step is the core architecture design of the game theory model. Through a layered logic of "upper-lower level," it simulates the two-layer interest game relationship between the power grid and the virtual power plant, as well as the aggregators within the virtual power plant, to achieve a balance between global and local objectives. The specific implementation method is as follows:

[0109] The core logic of the two-layer game framework is "the upper layer sets the rules, and the lower layer determines the allocation." The main players in the upper-layer game are the power grid dispatch center (G) and the virtual power plant representative (denoted as A_0, elected by each aggregator, responsible for summarizing the lower-layer output information and negotiating with the power grid). The main players in the lower-layer game are the aggregators (A_1, A_2, A_3). The two layers of the game are not independent, but form a closed-loop interaction through "price-output" signals: the upper-layer power grid transmits the electricity purchase price p_G and the total dispatch demand P_G to the lower-layer aggregators through A_0; the lower-layer aggregators formulate their own p_i and Pi based on p_G, summarize them into P_total through A_0, and feed it back to the power grid; the power grid adjusts p_G based on the deviation between P_total and P_G, and then distributes it back to the lower layer until a supply-demand balance is achieved.

[0110] The goal of the upper-level game is to achieve an equilibrium between "minimizing the grid's electricity purchase cost" and "maximizing the revenue of the virtual power plant." Constraints include grid purchase budget constraints (p_G×P_total≤B, B=50,000 yuan / hour), power supply reliability constraints (|P_G-P_total|≤ΔP_max, ΔP_max=50kW, meaning the supply-demand deviation does not exceed 50kW), and voltage quality constraints (based on a grid connection voltage amplitude of 10.5±0.1kV, achieved through reactive power output adjustment by the virtual power plant). Taking the initial stage of the upper-level game as an example, the grid determines P_G=1800kW based on the load gap and, combined with market conditions, gives an initial electricity purchase price of p_G=0.5 yuan / kWh. A_0 aggregates the initial output of the lower-level aggregators, P_total=1700kW (not meeting P_G), and A_0 requests a price increase from the grid (the reason being that insufficient output requires incentives). The grid, considering the deviation penalty cost, adjusts p_G to 0.53 yuan / kWh, incentivizing the lower-level aggregators to increase output.

[0111] The goal of the lower-level game is for each aggregator to maximize its own revenue. Constraints include individual output constraints (P_i ≤ P_i_max, e.g., A_1 ≤ 800kW, A_2 ≤ 600kW, A_3 ≤ 700kW, derived from the dispatchable potential profile), total output constraints (P_1 + P_2 + P_3 = P_total, which must match the aggregation requirements of the upper-level A_0), and price consistency constraints (aggregator price p_i ≤ p_G + 0.05 yuan / kWh, to avoid excessively high prices being rejected by the grid). The core of the lower-level game is the competition for output allocation—under the grid's given p_G, aggregators with higher output and lower costs have higher revenue, but must avoid penalties resulting from total output exceeding P_G. For example, when p_G = 0.53 yuan / kWh, A_1 has the highest unit revenue (p_G - marginal cost) (marginal cost C_i' = 2a_iP_i + b_i, C_1' of A_1 = 2 × 0.0001 × P_1 + 0.1), so A_1 tends to increase its output to 700kW; A_2 has the second highest marginal cost and an output of 550kW; A_3 has a relatively high marginal cost and an output of 550kW, with a total output P_total = 1800kW, which meets the needs of the upper layer.

[0112] The interaction mechanism of the two-layer game is realized through "information transmission - strategy adjustment - feedback iteration". The specific process is as follows: 1. The upper-level power grid publishes P_G and the initial p_G; 2. A_0 transmits the signal to the lower-level aggregator; 3. The aggregator calculates the optimal P_i based on p_G and reports it to A_0; 4. A_0 summarizes P_total. If |P_G-P_total|≤ΔP_max is satisfied, P_total and the p_i of each A_i are reported to the power grid; if not satisfied, the power grid needs to adjust p_G; 5. The power grid adjusts p_G according to P_total and its own revenue U_G and repeats steps 2-4. For example, when the initial p_G = 0.5 yuan / kWh, P_total = 1700kW (deviation 100kW). The grid adjusts p_G to 0.53 yuan / kWh. After the lower layer recalculates, P_total = 1800kW (deviation 0). A_0 reports the information of A_1 (p_1 = 0.55 yuan / kWh, P_1 = 700kW), A_2 (p_2 = 0.54 yuan / kWh, P_2 = 550kW), and A_3 (p_3 = 0.56 yuan / kWh, P_3 = 550kW) to the grid. The upper-level game then enters the stage of determining the benefit equilibrium.

[0113] The framework also needs to clearly define conflict coordination rules—when there is a conflict in output allocation between aggregators (such as both A_1 and A_2 wanting to increase output, causing the total output to exceed P_G), A_0 coordinates according to "unit revenue priority": calculating the unit revenue of each aggregator (p_G-C_i(P_i) / P_i), the aggregator with higher unit revenue gets priority in obtaining output quota. For example, A_1's unit revenue is 0.53-(0.0001×700²+0.1×700+50) / 700≈0.53-(49+70+50) / 700≈0.53-0.241≈0.289 yuan / kWh, and A_2's unit revenue is ≈0.53-0.25≈0.28 yuan / kWh. Therefore, A_1 gets priority in obtaining a higher output quota, and A_2 appropriately reduces its output to ensure that the total output complies with regulations.

[0114] Based on a two-level game framework, the alternating direction multiplier method is used for distributed iterative solution to update the bids and output strategies of each participant and generate a strategy iteration sequence.

[0115] This step is the core of solving the game theory model. The Alternating Direction Multiplier Method (ADMM), with its advantages of "distributed computing, fast convergence speed, and applicability to non-convex problems," is well-suited to the hierarchical structure of two-level games. By iteratively updating the strategies of each participant, it gradually approaches the optimal solution. The specific implementation is as follows:

[0116] The core principle of the alternating direction multiplier method is to decompose the global optimization problem of a two-level game into multiple local subproblems (policy optimization for each participant). By introducing a penalty factor and a dual variable, the consistency between local optima and global optima is coordinated. The basic algorithm flow is "initialize parameters → alternately update the policies of each participant → update the dual variable → check for convergence," where "alternating update" is the key—in each iteration, only the policy of one participant is updated, while the policies of other participants remain fixed, ensuring that the solution to each subproblem is simple and feasible.

[0117] First, parameter initialization is performed to determine the core parameters of ADMM: 1. Penalty factor ρ, used to balance local optimization and global constraints. A value that is too large will lead to convergence oscillations, while a value that is too small will lead to slow convergence. Based on the virtual power plant scenario, ρ is set to 10 yuan / (kW²); 2. Dual variables λ_i (corresponding to the output constraints of each aggregator) and λ_G (corresponding to the supply and demand constraints of the power grid), both are initially set to 0; 3. Initial strategies of each participant: the initial bid of the aggregator is p_i^0 = 0.4 yuan / kWh, the initial output is P_i^0 = P_i_max × 50% (A_1 = 400kW, A_2 = 300kW, A_3 = 350kW), the initial power purchase price of the power grid is p_G^0 = 0.45 yuan / kWh, and the initial dispatch demand is P_G^0 = 1800kW; 4. Iteration termination threshold ε = 0.01 (the iteration stops when the change in strategy is less than ε), and the maximum number of iterations is T = 50.

[0118] The iterative process is divided into three stages: "aggregator strategy update → grid strategy update → dual variable update", which are performed alternately. The first stage is aggregator strategy update: the grid strategy (p_G^k, P_G^k) and other aggregator strategies (P_j^k, j≠i) are fixed. Each aggregator A_i solves the subproblem of maximizing its own revenue - maxU_i=p_i×P_i-C_i(P_i)-L_i(P_i,Z_i), with the constraints P_i≤P_i_max, p_i≤p_G^k+0.05, and P_1+P_2+P_3=P_G^k+(λ_G^k) / ρ (constraint relaxation introduced by the dual variable). Taking the k=0 iteration as an example, the subproblem of A_1 is maxU_1=0.4P_1-(0.0001P_1²+0.1P_1+50)-0.0005P_1×88, with constraints P_1≤800, p_1≤0.5, P_1+300+350=1800+0 / 10→P_1=1150 (exceeding Pi_i_max=800, so we take P_1=800). Substituting these values, we get U_1=0.4×800-(64+80+50)-35.2=320-194-35.2=90.8 yuan. We update P_1^1=800kW, p_1^1=0.5 yuan / kWh (close to p_G^0+0.05).

[0119] The second phase of grid strategy update: fixed aggregator strategy (P_i^k, p_i^k), the grid G ​​solves the subproblem of maximizing its own revenue - maxU_G=λ×P_G-p_G×P_total-γ×|P_G-P_total|, with the constraints p_G≥max(p_i^k)-0.02 (to ensure that the electricity purchase price is attractive) and p_G×P_total≤B. Taking the k=1 iteration as an example, the total output of the aggregator is P_total^1=800+350+400=1550kW, max(p_i^1)=0.5 yuan / kWh, the grid constraint is p_G≥0.48, 0.48×1550≈744 yuan≤50000 yuan. Solving the subproblem, we get p_G^1=0.52 yuan / kWh (balancing the cost of electricity purchase and the deviation penalty), P_G^1=1550kW (matching the current P_total), U_G=0.8×1550-0.52×1550-2×0=1240-806=434 yuan.

[0120] The third stage involves updating the dual variables: The dual variables λ_i and λ_G are updated based on the deviation between the current policy and the constraints. The update formula is λ_i^(k+1)=λ_i^k+ρ×(P_i^(k+1)-P_i_max) (if P_i^(k+1)≤P_i_max, then λ_i remains unchanged), and λ_G^(k+1)=λ_G^k+ρ×(P_G^(k+1)-P_total^(k+1)). Taking k=1 as an example, for A_1, P_1^1=800=P_i_max, λ_1^1=0+10×(800-800)=0; P_G^1=1550, P_total^1=1550, λ_G^1=0+10×(1550-1550)=0. If P_2^2=650>P_i_max=600 for A_2 when k=2, then λ_2^3=λ_2^2+10×(650-600)=0+500=500. In subsequent iterations, A_2 will reduce its output due to the penalty effect of λ_2.

[0121] Through the above iterations, a strategy iteration sequence is generated, which is recorded in the form of "iteration number k - aggregator strategy (A_1:P_1,p_1;A_2:P_2,p_2;A_3:P_3,p_3) - grid strategy (p_G,P_G) - dual variables (λ_1,λ_2,λ_3,λ_G)". For example, the sequence fragments of the first three iterations are: k=0: A_1(400,0.4);A_2(300,0.4);A_3(350,0.4);G(0.45,1800);λ=(0,0,0,0); k=1: A_1(800,0.5);A_2(350,0.48);A_3(400,0.49);G(0.52,1550);λ=(0,0,0,0); k=2: A_1(750,0.53);A_2(600,0.51);A_3(550,0.52);G(0.54,1900);λ=(0,0,0,1000). It can be seen that as the iterations proceed, the aggregator's output gradually aligns with the grid's demand, and the quoted price fluctuates around the grid's electricity purchase price.

[0122] During the iteration process, the distributed nature must be guaranteed—each aggregator only needs to obtain its own dispatchable potential profile and the grid's public strategy (p_G, P_G), without needing to know the private information of other aggregators such as costs and losses, thus avoiding information leakage and meeting the distributed management requirements of virtual power plants. For example, when A_1 updates its strategy, it only needs to know p_G = 0.52 yuan / kWh and P_G = 1550kW, without needing to know the cost coefficient a_2 of A_2. Global constraint information is passed through the dual variables of ADMM, achieving "information localization and decision globalization".

[0123] The convergence of the monitoring strategy iteration sequence is determined. When the policy changes of all participants are less than a preset threshold, a dynamic Nash equilibrium is reached, and a collaborative scheduling scheme is output.

[0124] This step is the final stage of game theory problem-solving. It involves monitoring convergence to determine the stability of the strategy, confirming the dynamic Nash equilibrium state, and finally transforming the equilibrium strategy into an executable collaborative scheduling scheme. The specific implementation is as follows:

[0125] The core indicator for convergence monitoring is the "strategy change," which is the absolute difference between the strategies of each participant in two adjacent iterations. This includes the output change of the aggregator ΔP_i = |P_i^(k+1) - P_i^k|, the price change Δp_i = |p_i^(k+1) - p_i^k|, the power grid purchase price change Δp_G = |p_G^(k+1) - p_G^k|, and the dispatch demand change ΔP_G = |P_G^(k+1) - P_G^k|. The preset convergence thresholds are set according to the virtual power plant's operational accuracy requirements: ΔP_i ≤ 0.5kW (output change less than 0.5kW is considered stable), Δp_i ≤ 0.01 yuan / kWh (price change less than 0.01 yuan is considered stable), Δp_G ≤ 0.01 yuan / kWh, and ΔP_G ≤ 1kW. When all indicators are simultaneously satisfied, the iteration is considered converged.

[0126] Taking the iterative process as an example, the strategies for k=6 are: A_1 (P=680kW, p=0.52 yuan / kWh), A_2 (P=560kW, p=0.51 yuan / kWh), A_3 (P=560kW, p=0.53 yuan / kWh), G (p=0.52 yuan / kWh, P=1800kW); the strategies for k=7 are: A_1 (P=680.3kW, p=0.52 yuan / kWh), A_2 (P=559.8kW, p=0.51 yuan / kWh), A_3 (P=559.9k... W (p=0.53 yuan / kWh), G (p=0.52 yuan / kWh, P=1800kW); Calculate the changes: ΔP_1=0.3kW≤0.5kW, Δp_1=0≤0.01 yuan / kWh, ΔP_2=0.2kW≤0.5kW, Δp_2=0≤0.01 yuan / kWh, ΔP_3=0.1kW≤0.5kW, Δp_3=0≤0.01 yuan / kWh, Δp_G=0≤0.01 yuan / kWh, ΔP_G=0≤1kW. All indicators meet the threshold requirements, and the iteration converges.

[0127] After convergence, the strategy state needs to be verified to see if a dynamic Nash equilibrium has been reached. The core definition of Nash equilibrium is that "all participants, with the strategies of other participants unchanged, can no longer obtain higher returns by adjusting their own strategies." The verification method is as follows: fix the strategies of other participants, adjust the strategy of one participant alone, and calculate the change in its returns. If the returns decrease, then equilibrium is satisfied. Taking A_1 as an example, fix the strategies of A_2 (A_3) as (559.8kW, 0.51 yuan / kWh) and (559.9kW, 0.53 yuan / kWh), and the grid strategy as (0.52 yuan / kWh, 1800kW). Adjust the P_1 of A_1 from 680.3kW to 690kW, and calculate the returns U_1 = 0.52×690 - (0.0001×690² + 0.1×690 + 50) - 0.0005×690×88 = 358.8 - (4 7.61+69+50)-30.36=358.8-166.61-30.36=161.83 yuan, which is 0.27 yuan less than the original profit of 162.1 yuan; the p_1 is adjusted from 0.52 yuan / kWh to 0.53 yuan / kWh. The grid exceeds p_G+0.05 (0.52+0.05=0.57, which is not exceeded) because p_1 exceeds p_G+0.05. However, the output demand of A_1 causes the total output to exceed 1800kW. After the grid penalty, the profit of A_1 decreases. Therefore, the strategy of A_1 has reached the optimal level.

[0128] The "dynamic" nature of dynamic Nash equilibrium is reflected in the fact that the equilibrium state adjusts with changes in the external environment. When the grid dispatch demand P_G changes from 1800kW to 1900kW, the original equilibrium is broken, and a new equilibrium needs to be solved iteratively. For example, when P_G=1900kW, the iteration converges at k=12, and the new strategy is A_1 (730kW, 0.54 yuan / kWh), A_2 (580kW, 0.53 yuan / kWh), A_3 (590kW, 0.55 yuan / kWh), G (p=0.54 yuan / kWh, P=1900kW), ensuring that the equilibrium state adapts to the dynamic grid demand.

[0129] The coordinated dispatch scheme is an engineering transformation of the dynamic Nash equilibrium strategy, comprising three parts: "grid dispatch instructions," "aggregator execution strategies," and "safeguard measures." The grid dispatch instructions specify the total demand and electricity purchase price: "The total active power target of the virtual power plant is 1800kW, the electricity purchase price is 0.52 yuan / kWh, the adjustment time limit is 10 seconds, and the output deviation is allowed ±0.5kW." The aggregator execution strategies specify the specific tasks of each aggregator: A_1 "manages ESS-001~005, with a total output of 680.3kW, and the output allocation of individual units is ESS-001 (150kW), ESS-002 (140kW), and ESS-003 (130kW)." ESS-004 (130kW) and ESS-005 (130.3kW) are priced at RMB 0.52 / kWh; A_2 has a total output of 559.8kW and is priced at RMB 0.51 / kWh; A_3 has a total output of 559.9kW and is priced at RMB 0.53 / kWh. The safeguards include reactive power regulation assistance (each aggregator adjusts reactive power output according to impedance spectrum to maintain grid connection voltage of 10.5kV) and fault backup (50kW of A_1 output is reserved as a backup to cope with sudden loads).

[0130] The plan also specifies the revenue distribution and responsibility allocation: each aggregator's revenue (A_1=162.1 yuan, A_2=138.5 yuan, A_3=140.2 yuan), and the grid's revenue of 540 yuan. If the aggregator's output deviation exceeds the threshold, revenue will be deducted based on the deviation amount (2 yuan will be deducted for every 1kW of deviation). If the grid's electricity purchase price is not implemented as agreed, the aggregator's losses will be compensated (0.01 yuan / kWh will be compensated for every 0.01 yuan / kWh decrease), ensuring the plan's feasibility and binding force.

[0131] S204. Based on the aforementioned collaborative scheduling scheme, the decision-making strategies of each aggregator are optimized online using a multi-agent reinforcement learning framework, and the game behavior is dynamically adjusted through the opponent's strategy prediction network to generate adaptive collaborative control instructions.

[0132] Specifically, based on the collaborative scheduling scheme, historical decision data and environmental state data of each aggregator can be collected to construct a multi-agent state-action pair dataset containing state, action, reward and next state;

[0133] This step is the foundational data preparation stage for multi-agent reinforcement learning. Its core is to extract the correlation data between "agent behavior and environment feedback" from the cooperative scheduling scheme and operational history, constructing a standardized state-action pair dataset to provide high-quality training samples for subsequent policy optimization and opponent prediction. The specific implementation method is as follows:

[0134] The collaborative scheduling scheme clarifies the basic output targets and pricing ranges for each aggregator, serving as the core reference for data collection. Its "aggregator-output quota-pricing ceiling" correspondence (e.g., A_1 output 680kW, pricing ≤ 0.52 yuan / kWh; A_2 output 560kW, pricing ≤ 0.51 yuan / kWh) provides a benchmark for data filtering. Historical decision data collection covers the past 90 days of operational records, extracted in a structured manner as "aggregator ID-decision timestamp-pricing-actual output-output deviation." Output deviation refers to the difference between actual output and the scheduling scheme quota (e.g., A_1's actual output at a certain moment is 682kW, quota is 680kW, deviation +2kW). This indicator directly reflects the effectiveness of decision execution.

[0135] Environmental status data is an external factor influencing aggregators' decisions and needs to be precisely aligned with the decision data timestamps (deviation ≤ 0.1 seconds). The data collection dimensions include four core parameters: first, grid dispatch information (total demand P_G, electricity purchase price p_G, adjustment time limit), such as "P_G = 1800kW, p_G = 0.52 yuan / kWh, time limit 10 seconds"; second, energy storage unit status (average SOC of energy storage units under each aggregator, average fundamental impedance Z_avg, maximum regulation power P_ma). x), such as A_1's average SOC of 65%, Z_avg=88Ω, P_max=800kW; third, grid connection point electrical parameters (voltage amplitude U, frequency f, power factor cosφ), such as "U=10.5kV, f=50.02Hz, cosφ=0.98"; fourth, market environment parameters (regional electricity market real-time price p_market, load peak and valley marking), such as "p_market=0.55 yuan / kWh, marking=peak segment". All environmental data are extracted in real time from synchronous monitoring data packets and the electricity market data platform to ensure timeliness and authenticity.

[0136] The core structure of the multi-agent state-action pair dataset is "state s-action a-reward r-next state s'", and these four must form a complete causal chain. State s is a fusion vector of the environment and the agent's own state, containing 15 feature components: total grid demand P_G, electricity purchase price p_G, market price p_market, grid connection point voltage U, frequency f, power factor cosφ, SOC and Z_avg of A_1, SOC and Z_avg of A_2, SOC and Z_avg of A_3, and current output of each aggregator P_1 / P_2 / P_3. All components are standardized (the formula is x_norm=(x-x_min) / (x_max-x_min), for example, x_min=1000kW and x_max=2000kW of P_G, and 1800kW is standardized to 0.8).

[0137] Action 'a' is the aggregator's decision variable, a two-dimensional vector for each aggregator (bid price p_i, output P_i), such as action a_1 of A_1 = (0.52 yuan / kWh, 680kW). The action space must meet the constraints of the dispatchable potential profile (P_i≤P_max, p_i∈[0.3,0.8] yuan / kWh). Reward 'r' is the environmental feedback to the action, calculated using a weighted summation formula: r_i = w1×r_income + w2×r_deviation + w3×r_stability, where w1 = 0.6 (income weight), w2 = -0.3 (deviation penalty weight), and w3 = 0.1 (stability weight). r_income is the unit revenue (p_i - marginal cost). For example, if the marginal cost of A_1 is 0.24 yuan / kWh, then r_income = 0.52 - 0.24 = 0.28 yuan / kWh. r_deviation is the deviation penalty (1 when |P_i - quota| ≤ 1kW, otherwise 1 - |P_i - quota| / 5, for example, r_deviation = 1 - 2 / 5 = 0.6 when the deviation is 2kW). r_stability is the stability of the operation (1 when the output change between adjacent times is ≤ 5kW, otherwise 0.8). For example, r = 0.6 × 0.28 + (-0.3) × 0.6 + 0.1 × 1 = 0.168 - 0.18 + 0.1 = 0.088 for A_1.

[0138] The next state s' is the new state of the environment after the execution of action a. It has the same dimension as s. The difference is that the output P_1 / P_2 / P_3 is updated to the actual output, the SOC changes due to charging and discharging (e.g., A_1 discharges 680kW, and after 1 hour the SOC drops from 65% to 58.2%), and the grid connection point parameters are slightly adjusted due to the change in total output (e.g., when the total output is 1800kW, U=10.5kV, and when it rises to 1810kW, U=10.51kV). The final dataset contains 100,000 state-action pairs. For example, a complete sample might be: s=[0.8,0.52,0.55,1.0,0.51,0.98,0.65,0.18,0.72,0.22,0.68,0.68,0.56,0.56]; a_1=(0.52,680), a_2=(0.51,560), a_3 =(0.53,560); r_1=0.088, r_2=0.075, r_3=0.082; s'=[0.8,0.52,0.55,1.01,0.52,0.98,0.64,0.18,0.71,0.22,0.67,0.682,0.559,0.561], providing rich associated data for subsequent model training.

[0139] A network for predicting opponent policies is trained using a multi-agent state-action dataset to learn the decision-making patterns of other agents and generate an opponent behavior prediction model.

[0140] This step is crucial for enhancing the agent's game-playing capabilities. By capturing the decision-making patterns of other aggregators through the opponent's strategy prediction network, the agent can anticipate the opponent's behavior during the game and adjust its own strategy in advance to obtain better returns. The specific implementation method is as follows:

[0141] The core function of the opponent strategy prediction network is to "take the current state s and the historical action sequence as input, and output the prediction value of the opponent's next action". Its essence is a time series data prediction model. It is constructed using a Long Short-Term Memory (LSTM) network because LSTM can effectively capture long-term dependencies in time series data and is suitable for the characteristic that aggregator decisions are "affected by historical behavior". The network structure consists of an input layer, two LSTM hidden layers, and a fully connected output layer. The input layer dimension is "historical window length × feature dimension". The historical window length is set to 10 (i.e., using the state and action of the previous 10 time steps to predict the action of the 11th time step). The feature dimension is 23 (15-dimensional state s + 3 aggregate quotients of 2-dimensional action a, totaling 15 + 3 × 2 = 23). The first LSTM hidden layer has 128 units, and the second layer has 64 units. The activation function for both is tanh (to solve the gradient vanishing problem). The output layer dimension is 4 (predicting the actions of two main opponents, each opponent is 2-dimensional, such as predicting p_2 and P_2 of A_2, and p_3 and P_3 of A_3). The activation function is sigmoid (mapping the output to the action space range, such as mapping the price to [0.3, 0.8] yuan / kWh, and the output to [0, P_max]).

[0142] Before training, the dataset needs to be preprocessed temporally: after sorting by timestamp, training samples are constructed in a sliding window manner, each sample containing "the (s_t, a_t) sequence of the first 10 time steps - the opponent's action a_opponent at the 11th time step"; the input features are normalized (state features are normalized according to the standardization method, the price in the action features is normalized to the range [0.3, 0.8], and the output is normalized according to each aggregator P_max); the dataset is divided into a training set (70,000 records), a validation set (20,000 records), and a test set (10,000 records) in a 7:2:1 ratio. The training set is used for parameter updates, the validation set is used for adjusting hyperparameters, and the test set is used to evaluate the final prediction accuracy.

[0143] The training process employs the Adaptive Moment Estimation (Adam) optimizer, with the core hyperparameters set as follows: learning rate η = 0.001 (controlling the parameter update speed to avoid oscillations due to excessive η and slow convergence due to excessive η), batch size 64 (64 samples are input per iteration to balance training efficiency and stability), number of iterations epochs = 50 (ensuring the model fully learns the patterns), and weight decay coefficient λ = 0.0001 (preventing overfitting). The loss function is the mean squared error (MSE), calculated as Loss = 1 / N × Σ(a_pred - a_true)², where a_pred is the predicted action, a_true is the actual action, and N is the number of samples. For example, in a certain training batch, the predicted action of A_2 is a_pred=(0.51 yuan / kWh, 560kW), and the actual action is a_true=(0.51 yuan / kWh, 559kW). Then the loss is 1 / 64×[(0.51-0.51)²+(560-559)²+...]≈0.015, which reflects that the prediction deviation is small.

[0144] During training, model performance is monitored using a validation set. When the validation set loss stops decreasing after five consecutive epochs, an early stopping strategy is employed to prevent overfitting and to save the optimal model parameters. After training, prediction accuracy is evaluated on the test set using the Mean Absolute Percentage Error (MAPE) metric, calculated as MAPE = 1 / N × Σ|(a_pred - a_true) / a_true| × 100%. For example, on the test set, A_2's bid MAPE is 1.2%, and its output MAPE is 0.8%; A_3's bid MAPE is 1.5%, and its output MAPE is 1.0%, both below the preset accuracy threshold of 5%, indicating that the model can accurately capture the opponent's decision-making patterns.

[0145] The generated opponent behavior prediction model can directly output quantified opponent action prediction results. For example, if the input is the state and action sequence of the previous 10 time steps (including A_1's own actions and the historical actions of A_2 and A_3), the model outputs "A_2's next bid is 0.51±0.005 yuan / kWh, and its output is 559±1kW; A_3's next bid is 0.53±0.006 yuan / kWh, and its output is 560±1.2kW", providing A_1 with a clear expectation of opponent behavior for its decision.

[0146] The opponent behavior prediction model is integrated into a multi-agent reinforcement learning framework. The decision policy network of each aggregator is optimized by the policy gradient method to generate an optimized decision policy network.

[0147] This step is the core of strategy optimization. By integrating the opponent prediction model into the reinforcement learning framework, each aggregator can make decisions based on the expected behavior of the opponent. Then, the decision network parameters are updated through the policy gradient method, realizing a closed loop of "prediction-optimization-profit improvement". The specific implementation method is as follows:

[0148] The multi-agent reinforcement learning framework adopts the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework, which is specifically designed for multi-agent cooperation and game theory in partially observable environments. Its core advantage is that "each agent has an independent policy network, while a centralized critic network evaluates the global action value," adapting to the characteristics of virtual power plant aggregators who "make independent decisions but whose returns are affected by the global environment." The framework structure consists of three core modules: the independent policy networks (Actor networks) of each aggregator, the centralized critic network (Critic network), and an integrated opponent policy prediction model. The interaction logic among the three is as follows: the policy network outputs its own action → the opponent prediction model outputs the opponent's action → the critic network calculates the action value by combining the global action and state → the policy network updates its parameters based on the value signal.

[0149] The policy network (Actor) employs a fully connected neural network structure. Its inputs are the current state *s* and the predicted opponent's action *a_opponent_pred* (a total of 15+4=19 dimensions of features). The output is its own action *a_i* (quote *p_i* and force *P_i*, 2 dimensions). The network structure is: "Input layer (19 dimensions) → Hidden layer 1 (64 dimensions, ReLU activation) → Hidden layer 2 (32 dimensions, ReLU activation) → Output layer (2 dimensions, sigmoid activation)". The output is denormalized and mapped to the actual action space (e.g., the denormalization formula for the quotation is *p_i* = 0.3 + *a_pred* × 0.5, where *a_pred* is the output value). For example, in the input s, P_G=1800kW (0.8) and p_G=0.52 yuan / kWh (0.55), the opponent predicts a_opponent_pred=(A_2:0.51,560; A_3:0.53,560), and the network outputs a_pred=(0.44,0.85). After denormalization, p_i=0.3+0.44×0.5=0.52 yuan / kWh, and P_i=0+0.85×800=680kW, which is consistent with the initial scheduling scheme.

[0150] The Critic network is a centralized network. Its inputs are the global state *s* and the actions of all agents (including the agent's own action *a_i* and the opponent's actual action *a_opponent_true*, totaling 15 + 3 × 2 = 21 dimensions). The output is the action value Q(s, a_1, a_2, a_3) (reflecting the long-term reward of the current global action combination). The network structure is: "Input layer (21 dimensions) → Hidden layer 1 (128 dimensions, ReLU activation) → Hidden layer 2 (64 dimensions, ReLU activation) → Output layer (1 dimension, linear activation)". Value calculation combines the immediate reward *r* with a discounted sum of future rewards, expressed as Q(s, a) = *r* + γ × maxQ(s', a'), where γ = 0.95 is the discount factor (emphasizing the importance of the current reward while considering future rewards).

[0151] The policy gradient method is the core algorithm for updating the policy network parameters. It employs the REINFORCE algorithm with an advantage function, where the advantage function A(s,a) = Q(s,a) - V(s) measures how much better the current action is than the average action, and V(s) is the state value (estimated through a separate value network). The goal of policy update is to maximize the cumulative advantage. The parameter update formula is θ_i←θ_i+α×∇θ_ilogπ_i(a_i|s,a_opponent_pred)×A(s,a), where α=0.0005 is the policy learning rate, and π_i is the action probability distribution of the policy network. For example, for an action a_i in A_1, A(s,a) = 0.12 (better than the average action), and logπ_i(a_i|s) = -1.5 (logarithm of action probability), then the parameter update is 0.0005×(-1.5)×0.12≈-0.00009, fine-tuning the network parameters to increase the probability of that action.

[0152] The optimization process is divided into two phases: exploration and utilization. The first 200,000 steps are the exploration phase, where Gaussian noise (mean 0, standard deviation 0.02) is added to the action output to encourage the agent to try new actions (such as A_1's price fluctuating between 0.51-0.53 yuan / kWh) and avoid getting trapped in local optima. After 200,000 steps, the utilization phase begins, gradually reducing the noise to 0, and the agent outputs actions based on the optimized strategy. During the optimization process, the average revenue of each aggregator is monitored in real time. Initially, A_1's average revenue was 162 yuan / hour, which increased to 175 yuan / hour after 100,000 steps of optimization, and stabilized at 178 yuan / hour after 200,000 steps, representing a 9.9% improvement over the initial amount, indicating that the strategy optimization is effective.

[0153] The generated optimized decision-making strategy network has the ability to "predict the opponent and dynamically adjust". For example, when the opponent's prediction model outputs "A_2 will raise the bid to 0.53 yuan / kWh", the strategy network of A_1 will, while ensuring the profit, fine-tune the bid to 0.525 yuan / kWh (lower than A_2, to increase its own probability of winning the bid), and at the same time increase the output to 685kW (to seize more output quotas by taking advantage of its own lower marginal cost), so as to further improve its own profit.

[0154] The optimized decision-making strategy network is deployed to the actual operating environment, and output instructions for each aggregator are dynamically generated based on real-time monitoring data, ultimately generating adaptive collaborative control instructions.

[0155] This step is the engineering aspect of strategy implementation. By deploying the optimized network to the real-time control system and combining it with dynamic monitoring data to generate executable output commands, while also possessing adaptive adjustment capabilities in response to environmental changes, the collaborative control effect of the virtual power plant is ensured. The specific implementation method is as follows:

[0156] The deployment architecture adopts an "edge computing + cloud collaboration" model. The optimized decision-making strategy network (including the Actor network and the competitor prediction model) is deployed on the edge control nodes of the virtual power plant (response latency ≤50ms, meeting real-time control requirements). The model update service and data storage center are deployed in the cloud. The edge nodes communicate with the local controllers of each aggregator via industrial Ethernet (transmission protocol is Modbus-TCP), and communicate with the synchronous monitoring system via the IEC61850 protocol to ensure high-speed transmission of real-time data and reliable command issuance. Before deployment, the network needs to be engineered and adapted, the model needs to be converted into a lightweight format supported by the edge nodes (such as ONNX format), and the input and output interfaces need to be standardized (input is real-time monitoring data in JSON format, output is control commands in XML format).

[0157] The real-time monitoring data processing flow is "data acquisition - preprocessing - feature extraction - network input". The data acquisition frequency is 10Hz (consistent with the synchronous monitoring data packets). The acquired content includes real-time grid dispatch instructions (P_G, p_G), energy storage unit status (SOC, Z_avg), grid connection point electrical parameters (U, f, cosφ), and real-time actions of competitors (current bids and outputs of A_2 and A_3). The preprocessing stage removes abnormal data (such as voltage exceeding 10.5±0.1kV, which is filled with data from the previous moment). The feature extraction stage generates a state vector s according to its dimensions and standardizes it. Finally, s and the competitor's real-time action sequence are input into the policy network, and the network outputs the action prediction value within 5ms.

[0158] The generation of output commands follows a process of "network output - constraint verification - smoothing." The predicted action value of the network output first undergoes constraint verification to ensure it conforms to the dispatchable potential profile and grid rules: output P_i ≤ P_max (e.g., A_1 ≤ 800kW), quoted price p_i ≤ p_G + 0.05 yuan / kWh (e.g., p_i ≤ 0.57 when p_G = 0.52), and total output P_1 + P_2 + P_3 ≤ P_G + ΔP_max (ΔP_max = 50kW, to avoid exceeding the total output limit). If the network output exceeds the constraints, it is corrected using the "nearest adjustment" principle. For example, if the network output P_i of A_1 is 810kW (exceeding 800kW), it is corrected to 800kW; if the quoted price p_i is 0.58 yuan / kWh (exceeding 0.57), it is corrected to 0.57 yuan / kWh.

[0159] The smoothing process is used to avoid equipment shocks caused by sudden changes in operation. It employs a first-order low-pass filter algorithm, with the formula P_i(t) = β × P_i_pred + (1-β) × P_i(t-1), where β = 0.2 is the smoothing coefficient (the smaller β is, the smoother the change in operation), P_i(t) is the current command output, P_i_pred is the network output, and P_i(t-1) is the command output from the previous moment. For example, if A_1's command output was 680kW and the network output was 685kW at the previous moment, after smoothing, P_i(t) = 0.2 × 685 + 0.8 × 680 = 681kW, thus avoiding a 5kW sudden change that could impact the energy storage unit. The final output command includes "aggregator ID - command timestamp - target output - price ceiling - execution time limit", for example, "A_1-2025-09-0110:00:00.123-target output 681kW-price ceiling 0.525 yuan / kWh-execution time limit 2 seconds".

[0160] The adaptive capability is based on a mechanism of "environmental change detection - rapid policy adjustment - online model update". Environmental change detection is achieved by comparing the changes in state characteristics over 10 consecutive time periods. When the change exceeds a threshold (e.g., P_G change ≥ 100kW, p_G change ≥ 0.05 yuan / kWh), it is determined to be an environmental abrupt change. The edge node immediately triggers the rapid adjustment mode of the policy network, shortening the historical window length from 10 to 5, thus improving the response speed. For example, if the grid dispatch demand suddenly increases from 1800kW to 1900kW (a change of 100kW), the A_1 policy network adjusts the target output from 681kW to 730kW within 10ms, while simultaneously increasing the price to 0.54 yuan / kWh to adapt to the new dispatch demand.

[0161] The online model update mechanism ensures the long-term adaptability of the strategy. Edge nodes upload their daily status and action data to the cloud at 2:00 AM (during off-peak load). The cloud then uses the new data to incrementally train the strategy network (1000 iterations, learning rate 0.0001), and distributes the updated model back to the edge nodes, achieving a continuous cycle of "run-learn-optimize". The final generated collaborative control instruction set includes output and pricing instructions from all aggregators, as well as global grid coordination instructions. For example, "Collaborative Control Instruction-202509011000: A_1 output 681kW, price 0.525 yuan / kWh; A_2 output 580kW, price 0.53 yuan / kWh; A_3 output 639kW, price 0.54 yuan / kWh; grid purchase price 0.52 yuan / kWh, total output 1900kW, voltage maintained at 10.5kV". This instruction set has dynamic adjustment capabilities, adapting to changes in grid demand and competitor behavior.

[0162] As can be seen, real-time monitoring of the state of charge (SOC) and electrical parameters of the grid connection points of each distributed energy storage unit within the virtual power plant, along with synchronous acquisition of dispatch commands from the upper-level power grid, allows for the generation of a unit dispatchable potential profiles characterized by impedance spectra. Based on these profiles, a two-layer non-cooperative game model is constructed between the aggregators within the virtual power plant and the external power grid, outputting a collaborative dispatch scheme under dynamic Nash equilibrium. Based on this collaborative dispatch scheme, a multi-agent reinforcement learning framework is used to optimize the decision-making strategies of each aggregator online. Furthermore, the network dynamically adjusts its game behavior by predicting the opponent's strategy, generating adaptive collaborative control commands. This enables accurate profiling of the dispatchable potential of distributed energy storage units within the virtual power plant, efficient solution of game equilibrium, and online adaptive optimization of control strategies, thereby improving the overall system's economy, stability, and response flexibility.

[0163] Another embodiment of the present invention provides a distributed collaborative control system for a virtual power plant based on multi-agent game theory and impedance sensing, see [link to relevant documentation]. Figure 3 The system may include:

[0164] The monitoring module 301 is used to monitor the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point in real time, and to collect the dispatch instructions of the upper power grid simultaneously.

[0165] The evaluation module 302 is used to dynamically evaluate the power regulation capability and dynamic response characteristics of each energy storage unit based on the electrical parameters, the state of charge and the scheduling command, and generate a unit dispatchable potential profile characterized by impedance spectrum.

[0166] The construction module 303 is used to construct a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid based on the schedulable potential profile of the unit, solve the optimal bid and output strategy of each game participant through a distributed iterative algorithm, and output a collaborative scheduling scheme under dynamic Nash equilibrium.

[0167] The optimization module 304 is used to optimize the decision-making strategies of each aggregator online based on the cooperative scheduling scheme using a multi-agent reinforcement learning framework, and dynamically adjust the game behavior of the network through the opponent's strategy prediction to generate cooperative control instructions with adaptive capabilities.

[0168] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0169] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0170] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0171] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A distributed cooperative control method for a virtual power plant based on multi-agent game theory and impedance sensing, characterized in that, The method includes: Real-time monitoring of the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point, and synchronous collection of dispatch instructions from the upper-level power grid; Voltage and current time-series data at the grid connection point are extracted from synchronous monitoring data packets. The amplitude and phase of the fundamental and harmonic components are calculated using frequency domain analysis methods to generate a frequency domain electrical parameter feature set. Based on the frequency domain electrical parameter feature set, the equivalent impedance amplitude and phase angle at each harmonic frequency are calculated in real time using the recursive least squares method to generate dynamic impedance spectrum data. Combining state-of-charge data and grid dispatch command datasets, the power regulation range and response time characteristics of energy storage units under different states of charge are analyzed to generate a power regulation capability evaluation matrix. Based on the dynamic impedance spectrum data and the power regulation capability evaluation matrix, a multi-dimensional feature vector is constructed to represent the dynamic response characteristics of each energy storage unit, and finally, a unit dispatchable potential profile characterized by impedance spectrum is generated. Based on the schedulable potential profile of the unit, a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid is constructed. The optimal bid and output strategy of each game participant are solved by a distributed iterative algorithm, and a collaborative scheduling scheme under dynamic Nash equilibrium is output. Based on the aforementioned collaborative scheduling scheme, the decision-making strategies of each aggregator are optimized online using a multi-agent reinforcement learning framework, and the game behavior is dynamically adjusted by predicting the opponent's strategy, thereby generating adaptive collaborative control instructions.

2. The method according to claim 1, characterized in that, The system monitors the state of charge (SOC) of each distributed energy storage unit within the virtual power plant and the electrical parameters of the grid connection point in real time, and simultaneously collects dispatch instructions from the upper-level power grid, including: Within the virtual power plant, a status monitoring module is deployed in each distributed energy storage unit. The system collects state-of-charge data in real time through smart meters and collects active and reactive power data through power sensors to generate raw status data streams. Deploy electrical parameter monitoring devices at the grid connection point of the virtual power plant to measure voltage amplitude, phase angle, frequency deviation and harmonic content in real time, and generate time-series data of electrical parameters at the grid connection point. The system establishes a communication connection with the upper-level power grid energy management system through the power dispatch data network interface, receives active power regulation instructions and voltage regulation instructions issued by the power grid dispatch center in real time, and generates a power grid dispatch instruction dataset. Establish a unified time synchronization system to align the timestamps and unify the sampling rates of the original state data stream, the time-series data of electrical parameters at the grid connection point, and the power grid dispatch instruction dataset, and generate synchronous monitoring data packets.

3. The method according to claim 2, characterized in that, Based on the schedulable potential profile of the unit, a two-layer non-cooperative game model is constructed between the internal aggregator of the virtual power plant and the external power grid. A distributed iterative algorithm is used to solve for the optimal bid and output strategies of each game participant, outputting a collaborative scheduling scheme under dynamic Nash equilibrium, including: Based on the unit schedulable potential profile, the set of game participants is defined to include each aggregator and the power grid dispatch center, and the strategy space and payoff function of each participant are determined to generate the set of game participants and the strategy space. Based on the set of game participants and the strategy space, a two-layer game framework structure is constructed, with the upper layer consisting of the power grid dispatch center and the virtual power plant, and the lower layer consisting of the aggregators. Based on a two-level game framework, the alternating direction multiplier method is used for distributed iterative solution to update the bids and output strategies of each participant and generate a strategy iteration sequence. The convergence of the monitoring strategy iteration sequence is determined. When the policy changes of all participants are less than a preset threshold, a dynamic Nash equilibrium is reached, and a collaborative scheduling scheme is output.

4. The method according to claim 3, characterized in that, The method based on the cooperative scheduling scheme utilizes a multi-agent reinforcement learning framework to optimize the decision-making strategies of each aggregator online, and dynamically adjusts the game behavior of the network through opponent policy prediction to generate adaptive cooperative control instructions, including: Based on the collaborative scheduling scheme, historical decision data and environmental state data of each aggregator are collected to construct a multi-agent state-action pair dataset containing state, action, reward and next state; A network for predicting opponent policies is trained using a multi-agent state-action dataset to learn the decision-making patterns of other agents and generate an opponent behavior prediction model. The opponent behavior prediction model is integrated into a multi-agent reinforcement learning framework. The decision policy network of each aggregator is optimized by the policy gradient method to generate an optimized decision policy network. The optimized decision-making strategy network is deployed to the actual operating environment, and output instructions for each aggregator are dynamically generated based on real-time monitoring data, ultimately generating adaptive collaborative control instructions.

5. A distributed collaborative control system for a virtual power plant based on multi-agent game theory and impedance sensing, characterized in that, The system includes: The monitoring module is used to monitor the state of charge of each distributed energy storage unit in the virtual power plant and the electrical parameters of the grid connection point in real time, and to collect the dispatch instructions of the upper power grid simultaneously. The evaluation module extracts grid-connected voltage and current time-series data from synchronous monitoring data packets, calculates the amplitude and phase of the fundamental and harmonic components using frequency domain analysis, and generates a frequency domain electrical parameter feature set. Based on the frequency domain electrical parameter feature set, it uses the recursive least squares method to calculate the equivalent impedance amplitude and phase angle at each harmonic frequency in real time, generating dynamic impedance spectrum data. Combining state-of-charge data and grid dispatch command datasets, it analyzes the power regulation range and response time characteristics of energy storage units under different states of charge, generating a power regulation capability evaluation matrix. Based on the dynamic impedance spectrum data and the power regulation capability evaluation matrix, it constructs a multi-dimensional feature vector to represent the dynamic response characteristics of each energy storage unit, and finally generates a unit dispatchable potential profile characterized by impedance spectrum. The module is used to construct a two-layer non-cooperative game model between the internal aggregator of the virtual power plant and the external power grid based on the schedulable potential profile of the unit. It solves the optimal bid and output strategy of each game participant through a distributed iterative algorithm and outputs a collaborative scheduling scheme under dynamic Nash equilibrium. The optimization module is used to optimize the decision-making strategies of each aggregator online based on the cooperative scheduling scheme using a multi-agent reinforcement learning framework, and dynamically adjust the game behavior of the network through the opponent's strategy prediction to generate cooperative control instructions with adaptive capabilities.

6. The system according to claim 5, characterized in that, The monitoring module is specifically used for: Within the virtual power plant, a status monitoring module is deployed in each distributed energy storage unit. The system collects state-of-charge data in real time through smart meters and collects active and reactive power data through power sensors to generate raw status data streams. Deploy electrical parameter monitoring devices at the grid connection point of the virtual power plant to measure voltage amplitude, phase angle, frequency deviation and harmonic content in real time, and generate time-series data of electrical parameters at the grid connection point. The system establishes a communication connection with the upper-level power grid energy management system through the power dispatch data network interface, receives active power regulation instructions and voltage regulation instructions issued by the power grid dispatch center in real time, and generates a power grid dispatch instruction dataset. Establish a unified time synchronization system to align the timestamps and unify the sampling rates of the original state data stream, the time-series data of electrical parameters at the grid connection point, and the power grid dispatch instruction dataset, and generate synchronous monitoring data packets.

7. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-4 when it is run.

8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-4.