Multi-agent collaborative power system multi-time scale balancing method and device
By employing a multi-agent collaborative power system multi-timescale balancing method, and utilizing agent models for pre-disaster early warning, in-disaster scheduling, and post-disaster recovery, combined with reinforcement learning and Markov game frameworks, the problem of insufficient power system scheduling under extreme weather conditions is solved, enabling rapid response and recovery optimization of the power grid.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST BRANCH OF STATE GRID POWER GRID CO
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-01
AI Technical Summary
The existing power system lacks data fusion and a full-stage disaster coordination mechanism under extreme weather conditions, resulting in insufficient flexible dispatch capabilities and difficulty in effectively coping with complex and ever-changing power system operation scenarios.
By employing a multi-agent collaborative approach, intelligent agent models for pre-disaster early warning, in-disaster scheduling, and post-disaster recovery are pre-constructed. Combined with reinforcement learning and Markov game frameworks, early warning signals and scheduling sequences are generated to achieve power system balance across multiple time scales.
It improves the safety, reliability, and dispatch efficiency of the power system under extreme weather conditions, enhances the overall operational resilience, enables rapid response and recovery of the power grid, and supports optimized management of load and generating units at different time scales and regional ranges.
Smart Images

Figure CN121965622A_ABST
Abstract
Description
Multi-agent cooperative power system multi-timescale balancing method and device Technical Field
[0001] This application relates to the field of power system control technology, and in particular to a multi-agent cooperative power system multi-timescale balancing method and apparatus. Background Technology
[0002] In related technologies, power system balancing methods utilize reinforcement learning algorithms to learn the mapping relationship between the grid operating state and control actions during the interaction process. This enables dynamic adjustment of scheduling strategies to adapt to random fluctuations in renewable energy output and, to a certain extent, enhance the resilience of the power system in the face of frequent extreme weather conditions.
[0003] However, the data integration of related technologies is low, and there is a lack of unified modeling and deep correlation between different types of operational data, new energy forecast data and meteorological disaster-related data. At the same time, the lack of a collaborative mechanism that runs through the entire disaster process leads to fragmented control processes and insufficient information transmission, resulting in insufficient flexible scheduling capabilities of the system. This makes it difficult to fully integrate and leverage the advantages of each control link, and it is difficult to effectively cope with the complex and ever-changing power system operation scenarios under extreme weather conditions, which urgently needs to be addressed. Summary of the Invention
[0004] This application provides a multi-agent collaborative power system multi-timescale balancing method and apparatus to solve the problems in related technologies, such as the low degree of data fusion in power system balancing methods and the lack of a collaborative mechanism for all stages of disasters, which leads to insufficient flexible scheduling capability of power systems, making it difficult to achieve efficient linkage and optimized utilization of various resources, and failing to effectively cope with complex and ever-changing power system operation scenarios under extreme weather conditions.
[0005] The first aspect of this application provides a multi-agent collaborative power system multi-timescale balancing method, comprising the following steps: generating a pre-disaster early warning signal for the power system based on a pre-built pre-disaster early warning agent model, combined with historical and predicted data on wind and solar power generation and load; determining the scheduling sequence of thermal power units during a disaster based on the pre-disaster early warning signal and a pre-built mid-disaster scheduling agent model, and obtaining the post-disaster thermal power unit scheduling sequence using a pre-built post-disaster recovery agent model; determining the state value, action value, and reward value of reinforcement learning based on a pre-built multi-agent collaborative Markov game framework, according to the pre-disaster early warning signal, the mid-disaster thermal power unit scheduling sequence, and the post-disaster thermal power unit scheduling sequence; and performing phased collaborative reinforcement learning using the state value, action value, and reward value of reinforcement learning to determine the experience-sharing parameters for multi-agent collaboration, thereby enabling the power system to achieve multi-timescale balancing.
[0006] Through the above technical means, the embodiments of this application can generate the required pre-disaster early warning signals, mid-disaster thermal power unit scheduling sequences, and post-disaster thermal power unit correction sequences based on the pre-disaster early warning intelligent agent model, the mid-disaster scheduling intelligent agent model, and the post-disaster recovery intelligent agent model, respectively. Furthermore, through the Markov game framework of multi-agent cooperation, the state value, action value, and reward value of each agent are determined, and based on this, experience-sharing parameters for multi-agent cooperation are obtained. This enables dynamic control and phased collaborative optimization of the power system at each stage before, during, and after a disaster, allowing multiple agents to make flexible decisions for different extreme weather and disaster types, thereby improving the safety, reliability, and scheduling efficiency of the power system and enhancing its overall operational resilience.
[0007] Optionally, in one embodiment of this application, the method includes: sorting the disaster-affected thermal power units according to their power upper limit to obtain a power sequence of the disaster-affected thermal power units; calculating the full-load cost, start-up cost, and shutdown cost of the disaster-affected thermal power units according to the power sequence, and calculating the unit full-load cost of the disaster-affected thermal power units based on the full-load cost, start-up cost, and shutdown cost; sorting the unit full-load cost to obtain an initial sequence of disaster-affected thermal power units; obtaining the maximum and minimum values of at least one time period of fluctuating load data, and dividing the initial sequence of disaster-affected thermal power units according to the maximum and minimum values to obtain a first part sequence of disaster-affected thermal power units; obtaining a new initial sequence of disaster-affected thermal power units based on the sequence number of the first part sequence; obtaining a second part sequence of disaster-affected thermal power units according to the new initial sequence of disaster-affected thermal power units, the pre-disaster energy storage accumulation value, the start-up cost, and the shutdown cost; and obtaining a scheduling sequence of disaster-affected thermal power units according to the first part sequence and the second part sequence.
[0008] Through the above technical means, the embodiments of this application can determine the scheduling sequence of thermal power units during a disaster based on the early warning signal output by the pre-disaster early warning intelligent agent model, the initial full-load cost sequence of thermal power units, and the load fluctuation constraints during the disaster. This enables optimized control of the start-up and shutdown sequence, output adjustment range, and reserve capacity allocation of thermal power units, thereby maintaining stable power supply to critical loads during a disaster, improving the operational reliability and scheduling flexibility of the power system, and providing a reliable scheduling basis and reference data for the post-disaster recovery phase.
[0009] Optionally, in one embodiment of this application, obtaining the post-disaster thermal power unit scheduling sequence using a pre-built post-disaster recovery intelligent agent model includes: inputting the in-disaster thermal power unit scheduling sequence into the pre-built post-disaster recovery intelligent agent model to calculate the change values of post-disaster fluctuating load data and in-disaster fluctuating load data; correcting the initial sequence of post-disaster thermal power units according to the change values to obtain a third part sequence of post-disaster thermal power units; obtaining a new initial sequence of post-disaster thermal power units based on the sequence number of the third part sequence; obtaining a fourth part sequence of post-disaster thermal power units according to the new initial sequence of post-disaster thermal power units, the start-up cost of post-disaster thermal power units, and the shutdown cost of post-disaster thermal power units; and obtaining the post-disaster thermal power unit scheduling sequence according to the third part sequence and the fourth part sequence.
[0010] Through the above technical means, the embodiments of this application can determine the post-disaster thermal power unit scheduling sequence based on the early warning information output by the pre-disaster early warning intelligent agent model and the scheduling sequence of thermal power units during the disaster. This allows for the optimization of the start-up and shutdown sequence, output adjustment, and correction strategies of thermal power units during the post-disaster recovery phase, enabling the system to gradually recover to normal operation based on load recovery, line availability, and equipment health status. This method can improve the recovery speed and operational reliability of the power system after a disaster, while optimizing resource scheduling, reducing recovery costs, and providing a continuous and stable power supply guarantee for the power system under extreme weather or disaster conditions.
[0011] Optionally, in one embodiment of this application, the formula for calculating the state value is: ,in, Indicates the area Internal disaster pre-warning intelligent agent in time period The state vector, Indicates the area Disaster-related scheduling agents during time periods The state vector, Indicates the area Internal disaster recovery agents during the time period The state vector, Indicates the area node During the period Temperature information; Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time; Indicates the area exist Real-time status alerts; Representing regions Transmission power, disaster mitigation power, historical power generation, and thermal power unit output power at time t; Representing regions The charging power, discharging power, disaster resilience power, and historical power of the internal battery at time t; Indicates the total duration; The region is represented; the formula for calculating the action value is: ,in, These represent wind power within the region. field Photovoltaic power stations thermal power plant During the period The power action value, Indicates the area Inner Time Period At the node The load shedding power action value at the location, Indicates the area Internal energy storage devices During the period The percentage of its real-time dispatch capacity to its total power capacity; The state space is represented; the formula for calculating the reward value is: ,in, , , These represent the reward values of the agent before, during, and after the disaster, respectively. , , , , , , , These represent the wind power output penalty coefficient, the photovoltaic power output penalty coefficient, and the regional penalty coefficient, respectively. The start-up status and region of the internal thermal power unit at time t The power of the internal battery e at time t, the start-up cost of the thermal power unit g, the shutdown cost of the thermal power unit, and the secondary and primary output coefficients of the thermal power unit.
[0012] Through the above technical means, the embodiments of this application can transform the decision-making process of multiple agents in the pre-disaster, during-disaster and post-disaster stages of the power system into computable reinforcement learning input and output indicators by quantifying state values, action values and reward values. Under the Markov game framework of multi-agent cooperation, experience sharing and strategy iterative updates can be carried out, thereby optimizing the decision-making strategies of agents at each stage, realizing cross-stage collaborative scheduling, and improving the operational resilience, scheduling efficiency and reliability of the power system under extreme weather or disaster conditions.
[0013] Optionally, in one embodiment of this application, the step of using the state values, action values, and reward values of the reinforcement learning to perform phased collaborative reinforcement learning to determine the experience-sharing parameters for multi-agent cooperation includes: calculating the prediction accuracy of the pre-disaster early warning agent model, the total cost ratio of the in-disaster scheduling agent model, and the degree of power shortage improvement of the post-disaster recovery agent model based on the state values, action values, and reward values, respectively; and calculating the experience-sharing parameters based on the prediction accuracy, the total cost ratio, and the degree of power shortage improvement.
[0014] Through the above technical means, the embodiments of this application can calculate the experience sharing parameters in the multi-agent collaboration process based on the state value, action value, and reward value, and finally obtain the experience sharing parameters. This enables each agent to share the experience information accumulated at different stages in the Markov game framework, thereby referring to the experience of other agents in the policy update process and improving the accuracy and coordination of decision-making.
[0015] A second aspect of this application provides a multi-agent collaborative power system multi-timescale balancing device, comprising: a generation module, used to generate a pre-disaster early warning signal for the power system based on a pre-built pre-disaster early warning agent model and combined with historical and predicted data on wind and solar power generation and load; an acquisition module, used to determine the scheduling sequence of thermal power units during a disaster based on the pre-disaster early warning signal and a pre-built mid-disaster scheduling agent model, and to acquire the scheduling sequence of thermal power units after a disaster using a pre-built post-disaster recovery agent model; and a balancing module, used to determine the state value, action value, and reward value of reinforcement learning based on a pre-built multi-agent collaborative Markov game framework, according to the pre-disaster early warning signal, the mid-disaster thermal power unit scheduling sequence, and the post-disaster thermal power unit scheduling sequence, so as to determine the experience-sharing parameters for multi-agent collaboration, thereby enabling the power system to achieve multi-timescale balancing.
[0016] Optionally, in one embodiment of this application, the mutual acquisition module includes: a first sorting unit, configured to sort the disaster-affected thermal power units according to their power upper limit to obtain a power sequence of the disaster-affected thermal power units; a calculation unit, configured to calculate the full-load cost, start-up cost, and shutdown cost of the disaster-affected thermal power units according to the power sequence, and to calculate the unit full-load cost of the disaster-affected thermal power units based on the full-load cost, start-up cost, and shutdown cost; a second sorting unit, configured to sort the unit full-load cost to obtain an initial sequence of disaster-affected thermal power units; and a partitioning unit, configured to acquire at least one time-period fluctuation negative The system obtains the maximum and minimum values of the load data and divides the initial sequence of the disaster-affected thermal power units according to the maximum and minimum values to obtain a first part sequence of the disaster-affected thermal power units; a first acquisition unit is used to obtain a new initial sequence of the disaster-affected thermal power units based on the sequence number of the first part sequence; a second acquisition unit is used to obtain a second part sequence of the disaster-affected thermal power units according to the new initial sequence of the disaster-affected thermal power units, the pre-disaster energy storage accumulation value, the start-up cost, and the shutdown cost; a third acquisition unit is used to obtain a scheduling sequence of the disaster-affected thermal power units according to the first part sequence and the second part sequence.
[0017] Optionally, in one embodiment of this application, the acquisition module includes: a second calculation unit, configured to input the disaster-affected thermal power unit scheduling sequence into the pre-built post-disaster recovery intelligent agent model to calculate the change values of post-disaster fluctuating load data and disaster-affected fluctuating load data; a correction unit, configured to correct the post-disaster thermal power unit initial sequence according to the change values to obtain a third part sequence of the post-disaster thermal power units; a fourth acquisition unit, configured to obtain a new initial sequence of the post-disaster thermal power units based on the sequence number of the third part sequence; a fifth acquisition unit, configured to obtain a fourth part sequence of the post-disaster thermal power units according to the new initial sequence of the post-disaster thermal power units, the start-up cost of the post-disaster thermal power units, and the shutdown cost of the post-disaster thermal power units; and a sixth acquisition unit, configured to obtain the post-disaster thermal power unit scheduling sequence according to the third part sequence and the fourth part sequence.
[0018] Optionally, in one embodiment of this application, the formula for calculating the state value is: ,in, Indicates the area Internal disaster pre-warning intelligent agent in time period The state vector, Indicates the area Disaster-related scheduling agents during time periods The state vector, Indicates the area Internal disaster recovery agents during the time period The state vector, Indicates the area node During the period Temperature information; Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time; Indicates the area exist Real-time status alerts; Representing regions Transmission power, disaster mitigation power, historical power generation, and thermal power unit output power at time t; Representing regions The charging power, discharging power, disaster resilience power, and historical power of the internal battery at time t; Indicates the total duration; The region is represented; the formula for calculating the action value is: ,in, These represent wind power within the region. field Photovoltaic power stations thermal power plant During the period The power action value, Indicates the area Inner Time Period At the node The load shedding power action value at the location, Indicates the area Internal energy storage devices During the period The percentage of its real-time dispatch capacity to its total power capacity; The state space is represented; the formula for calculating the reward value is: ,in, , , These represent the reward values of the agent before, during, and after the disaster, respectively. , , , , , , , These represent the wind power output penalty coefficient, the photovoltaic power output penalty coefficient, and the regional penalty coefficient, respectively. The start-up status and region of the internal thermal power unit at time t The power of the internal battery e at time t, the start-up cost of the thermal power unit g, the shutdown cost of the thermal power unit, and the secondary and primary output coefficients of the thermal power unit.
[0019] Optionally, in one embodiment of this application, the balancing module includes: a third calculation unit, configured to calculate the prediction accuracy of the pre-disaster early warning agent model, the total cost ratio of the in-disaster scheduling agent model, and the degree of power shortage improvement of the post-disaster recovery agent model based on the state value, the action value, and the reward value, respectively; and a fourth calculation unit, configured to calculate the experience sharing parameters based on the prediction accuracy, the total cost ratio, and the degree of power shortage improvement.
[0020] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a multi-agent cooperative power system multi-timescale balancing method as described in the above embodiments.
[0021] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-agent cooperative power system multi-timescale balancing method.
[0022] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, is used to implement the above-described multi-agent cooperative power system multi-timescale balancing method.
[0023] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0024] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 is a flowchart of a multi-agent cooperative power system multi-timescale balancing method provided according to an embodiment of this application; Figure 2 is a flowchart of a multi-agent cooperative power system multi-timescale balancing method according to an embodiment of this application; Figure 3 is a block diagram of a multi-agent cooperative power system multi-timescale balancing device provided according to an embodiment of this application; and Figure 4 is a structural diagram of an electronic device provided according to an embodiment of this application.
[0025] Figure reference numerals: 10-Multi-agent collaborative power system multi-timescale balancing device; 100-Generation module, 200-Acquisition module, 300-Balancing module; 401-Memory, 402-Processor, 403-Communication interface. Detailed Implementation
[0026] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0027] The following describes, with reference to the accompanying drawings, a multi-agent collaborative power system multi-timescale balancing method and apparatus according to embodiments of this application. Addressing the technical problems mentioned in the background art, such as low data fusion levels and a lack of a full-stage disaster coordination mechanism, which leads to insufficient flexible dispatching capabilities of the power system and difficulty in achieving efficient linkage and optimized utilization of various resources, thus hindering effective responses to complex and variable power system operation scenarios under extreme weather conditions, this application provides a multi-agent collaborative power system multi-timescale balancing method. In this method, based on pre-constructed pre-disaster early warning agent model, in-disaster dispatch agent model, and post-disaster recovery agent model, the required early warning signals, in-disaster thermal power unit dispatch sequences, and post-disaster thermal power unit dispatch sequences are generated respectively. Then, using a Markov game framework for multi-agent collaboration, the state values, action values, and reward values of reinforcement learning are determined to obtain experience-sharing parameters for multi-agent collaboration. This approach enables power systems to achieve multi-timescale balance, fully considering the temporal and spatial characteristics of extreme weather, and achieving dynamic control of power system operation, thus improving the reliability of the power system under extreme weather conditions. Simultaneously, it effectively enhances the scheduling efficiency of multi-agent models, allowing agents at each stage to collaboratively optimize decisions before, during, and after disasters, dynamically adjusting unit output, load allocation, and reserve resource scheduling to achieve rapid grid response and recovery. Furthermore, it supports optimized management of load and units at different time scales and regional scopes, achieving stable operation of the entire power grid. Through rapid response and rolling optimization mechanisms, it shortens disaster response and recovery time, significantly enhancing the overall resilience, reliability, and economic operation level of the power system under extreme weather conditions, providing strong support for grid safety and sustainable energy dispatch. This addresses the problem of insufficient flexible dispatch capabilities of the power system due to low data fusion levels and a lack of collaborative mechanisms across all stages of disasters, making it difficult to achieve efficient linkage and optimized utilization of various resources and effectively cope with complex and ever-changing power system operation scenarios under extreme weather conditions.
[0028] Specifically, Figure 1 is a flowchart illustrating a multi-agent cooperative power system multi-timescale balancing method provided in an embodiment of this application.
[0029] As shown in Figure 1, the multi-agent collaborative power system multi-timescale balancing method includes the following steps: In step S101, based on the pre-constructed pre-disaster early warning agent model, a pre-disaster early warning signal for the power system is generated by combining historical and predicted data of wind and solar power generation and load.
[0030] It can be explained that the pre-disaster early warning intelligent agent model can be a deep learning model based on multimodal data fusion, such as meteorological image feature extraction models based on convolutional neural networks and time-series meteorological prediction models based on LSTM (Long Short-Term Memory). It can jointly model meteorological data, satellite remote sensing data, power grid operation data and geographical environment data to identify key precursor features that may lead to disasters such as storms, rainstorms, high temperatures, and freezing, and predict the time, intensity and scope of disasters, thereby realizing proactive early warning in the pre-disaster stage and improving the power system's ability to respond to potential extreme weather.
[0031] Warning signals may include, but are not limited to: meteorological warning signals that characterize extreme weather risks, such as strong wind warnings, rainstorm warnings, high temperature warnings, freezing warnings, or lightning warnings; operation warning signals that characterize abnormal trends in power grid operation, such as load change warnings, line overload warnings, abnormal temperature rise warnings of key equipment, or severe fluctuations in new energy output; and risk warning signals that characterize environmental changes and external disturbances, such as geological disaster warnings, forest fire warnings, or flood risk warnings.
[0032] As one possible implementation, embodiments of this application can generate early warning signals based on a pre-disaster early warning intelligent agent model, using a long short-term memory network, and combining historical and predicted data on wind and solar power generation and load; it can be expressed as: , , ,in, Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time Indicates the area internal nodes In Temperature value at any given time; For the region The hidden layer in time The state vector, , , They are respectively regions The input gate, forget gate, and output gate are at time... The output vector, For the sigmoid function, Represents element-wise multiplication. For the region Candidate memory units in time The state vector, For the region The memory unit at time The state vector, , and They are respectively regions The weight matrices corresponding to the input gate, forget gate, and output gate. This indicates a region. The weight matrix used to update the state of memory cells. area Mid-memory network at time The final state value, Indicates the area China's early warning information is timely State value ( ).
[0033] In step S102, based on the pre-disaster early warning signal and the pre-built mid-disaster scheduling intelligent agent model, the mid-disaster thermal power unit scheduling sequence is determined, and the post-disaster thermal power unit scheduling sequence is obtained using the pre-built post-disaster recovery intelligent agent model.
[0034] The disaster-stricken dispatching intelligent agent model can be used for real-time control of the power system during a disaster. It can be built based on frameworks such as deep reinforcement learning and graph neural networks, and utilizes grid topology, equipment status data, load change information, and real-time output data of new energy sources for dynamic decision-making. It can work in conjunction with the risk level output by the pre-disaster early warning intelligent agent model to ensure the timeliness and effectiveness of dispatching actions. In the embodiments of this application, the disaster-stricken thermal power unit dispatching sequence may include, but is not limited to, unit start-up and shutdown sequence, output adjustment range, response time window, and reserve capacity allocation strategy, which can be used to ensure stable power supply to critical loads during a disaster. This sequence can be set by those skilled in the art according to actual conditions, and no specific limitations are imposed here.
[0035] Post-disaster recovery intelligent agent models can be used to conduct power system recovery and reconstruction control after a disaster. They can take multi-source data, such as a list of damaged equipment, fault location distribution, line accessibility, maintenance resource status, and load restoration needs, as input to generate optimized recovery plans for the post-disaster phase. Post-disaster thermal power unit scheduling sequences refer to the set of time sequences and control strategies for the start-up, shutdown, and output adjustment of thermal power units after a disaster to achieve gradual power system recovery and stable power supply.
[0036] Optionally, in one embodiment of this application, determining the scheduling sequence of thermal power units during a disaster includes: sorting the thermal power units during a disaster according to their power upper limit to obtain a power sequence of the thermal power units during a disaster; calculating the full-load cost, start-up cost, and shutdown cost of the thermal power units during a disaster based on the power sequence, and calculating the unit full-load cost of the thermal power units during a disaster based on the full-load cost, start-up cost, and shutdown cost; sorting the unit full-load costs to obtain an initial sequence of the thermal power units during a disaster; obtaining the maximum and minimum values of at least one period of fluctuating load data, and dividing the initial sequence of the thermal power units during a disaster based on the maximum and minimum values to obtain a first part sequence of the thermal power units during a disaster; obtaining a new initial sequence of the thermal power units during a disaster based on the sequence number of the first part sequence; obtaining a second part sequence of the thermal power units during a disaster based on the new initial sequence of the thermal power units during a disaster, the pre-disaster accumulated energy storage value, the start-up cost, and the shutdown cost; and obtaining a scheduling sequence of the thermal power units during a disaster based on the first part sequence and the second part sequence.
[0037] As one possible implementation method, the embodiments of this application may include the following steps: (1) First, the embodiments of this application may be applied to thermal power units. Arranged in descending order of maximum power output, calculate the full-load cost of each thermal power unit based on its sequence. Start-up costs and stopping costs Thus, the unit full load cost is obtained. It can be represented as: ,in, Indicates thermal power unit The corresponding full-load power.
[0038] Furthermore, the embodiments of this application can address the unit full-load cost of each thermal power unit. Sort in ascending order to obtain the initial sequence of thermal power units. .
[0039] (2) Secondly, the embodiments of this application can obtain the maximum and minimum values of the fluctuating load data for each time period, and use the maximum and minimum values to constrain the cumulative power value of the thermal power unit, which can be expressed as: ,in, express Time Node Minimum fluctuating load at the location; express Time Node The maximum value of the fluctuating load at that location; thus, the embodiments of this application can obtain value( This allows us to divide the initial sequence, resulting in the first part of the thermal power unit sequence. .
[0040] (3) Next, the embodiments of this application can sort the thermal power unit serial numbers After elimination, a new initial sequence for the thermal power units is obtained. Based on the accumulated energy storage value before the disaster The start-up cost and shutdown cost yield the latter half of the thermal power unit sequence, which can be expressed as: , , , ,in, Indicates energy storage devices At any moment The discharge power, Indicates thermal power unit Number of startups Indicates thermal power unit Number of stops; unit start-stop cost Sort in ascending order to obtain the second part of the sequence of thermal power units. .
[0041] Combining the above steps, the embodiments of this application can obtain a thermal power unit sequence, which can be represented as: Optionally, in one embodiment of this application, obtaining the post-disaster thermal power unit scheduling sequence using a pre-built post-disaster recovery intelligent agent model includes: inputting the in-disaster thermal power unit scheduling sequence into the pre-built post-disaster recovery intelligent agent model to calculate the change values of post-disaster fluctuating load data and in-disaster fluctuating load data; correcting the initial sequence of post-disaster thermal power units based on the change values to obtain the third part sequence of post-disaster thermal power units; obtaining a new initial sequence of post-disaster thermal power units based on the sequence number of the third part sequence; obtaining a fourth part sequence of post-disaster thermal power units based on the new initial sequence of post-disaster thermal power units, the start-up cost of post-disaster thermal power units, and the shutdown cost of post-disaster thermal power units; and obtaining the post-disaster thermal power unit scheduling sequence based on the third part sequence and the fourth part sequence.
[0042] Specifically, firstly, the embodiments of this application can obtain post-disaster phase fluctuation load data. Compared to the fluctuating load data during the disaster phase Change value It can be represented as: Furthermore, embodiments of this application can establish new dynamic optimization scheduling strategies for thermal power units and energy storage devices, that is, adding energy storage devices and power corrections for thermal power units that consider load changes on the basis of the original thermal power unit sequence. The changes in the load will cause the maximum and minimum values of the original fluctuating load to exhibit new characteristics. The constraints on the correction strategies for energy storage devices and thermal power units can be expressed as follows: Based on the above constraints, the embodiments of this application can be obtained. This value is used to correct the initial sequence of thermal power units after the disaster, i.e. Next, in this embodiment of the application, the sorted thermal power unit serial numbers can be excluded to obtain a new initial sequence of thermal power units after the disaster. Based on the start-up and shutdown costs, the fourth part of the sequence of post-disaster thermal power units is obtained, which can be expressed as: , ,in, This indicates the number of times the thermal power unit has been started after the correction. Indicates the revised thermal power unit The number of stops, for the corrected unit start The remaining correction sequence for thermal power units is obtained by sorting the stopping costs in ascending order. Finally, the modified post-disaster thermal power unit scheduling sequence expression obtained from the embodiments of this application is as follows: In step S103, based on the pre-constructed Markov game framework of multi-agent cooperation, the state value, action value and reward value of reinforcement learning are determined according to the pre-disaster early warning signal, the scheduling sequence of thermal power units during the disaster and the scheduling sequence of thermal power units after the disaster, so as to determine the experience sharing parameters of multi-agent cooperation, so that the power system can achieve multi-time scale balance.
[0043] Understandably, a multi-agent collaborative Markov game framework can introduce multiple agents—such as a pre-disaster early warning agent, a disaster scheduling agent, and a post-disaster recovery agent—at different stages—before, during, and after a disaster. Through sharing environmental states, mutual interaction, and collaborative decision-making, dynamic game theory and strategy optimization are conducted within a unified Markov game framework to achieve resilient power system scheduling throughout its entire lifecycle. Each agent can perform risk identification, resource allocation, scheduling optimization, and recovery strategy formulation based on stage-specific state spaces, action spaces, and reward functions, achieving phased collaborative responses to disaster scenarios and thus improving the overall disturbance rejection capability and operational stability of the power system. State values, action values, and reward values can be used to describe the agent's state representation, behavioral selection, and strategy optimization results in the environment, respectively. Through reinforcement learning based on state values, action values, and reward values, and through strategy evaluation and improvement, the agent can continuously adjust its decision-making strategies to maximize long-term cumulative gains.
[0044] Optionally, in one embodiment of this application, the formula for calculating the state value can be expressed as: ,in, Indicates the area Internal disaster pre-warning intelligent agent in time period The state vector, Indicates the area Disaster-related scheduling agents during time periods The state vector, Indicates the area Internal disaster recovery agents during the time period The state vector, Indicates the area node During the period Temperature information; Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time; Indicates the area exist Real-time status alerts; Representing regions Transmission power, disaster mitigation power, historical power generation, and thermal power unit output power at time t; Representing regions The charging power, discharging power, disaster resilience power, and historical power of the internal battery at time t; Indicates the total duration; The region is represented; the formula for calculating the action value is: ,in, These represent wind power within the region. field Photovoltaic power stations thermal power plant During the period The power action value, Indicates the area Inner Time Period At the node The load shedding power action value at the location, Indicates the area Internal energy storage devices During the period The percentage of its real-time dispatch capacity to its total power capacity; Representing the state space; the formula for calculating the reward value is: ,in, , , These represent the reward values of the agent before, during, and after the disaster, respectively. , , , , , , , These represent the wind power output penalty coefficient, the photovoltaic power output penalty coefficient, and the regional penalty coefficient, respectively. The start-up status and region of the internal thermal power unit at time t The power of the internal battery e at time t, the start-up cost of the thermal power unit g, the shutdown cost of the thermal power unit, and the secondary and primary output coefficients of the thermal power unit.
[0045] The total reward value in this embodiment can be expressed as: .
[0046] In actual implementation, the embodiments of this application can utilize the state value function, action value function and reward feedback formed during the interaction process, and uniformly incorporate them into the multi-agent cooperation framework. By sharing experience pools, joint policy gradients or value iterations, experience sharing parameters for multi-agent cooperation can be calculated, so that agents at different stages can complete policy collaborative optimization under common goals, thereby improving the overall response capability and recovery efficiency of the power system before, during and after disasters.
[0047] Optionally, in one embodiment of this application, reinforcement learning is performed in stages using state values, action values, and reward values to determine experience-sharing parameters for multi-agent collaboration. This includes: calculating the prediction accuracy of the pre-disaster early warning agent model, the total cost ratio of the in-disaster scheduling agent model, and the degree of power shortage improvement of the post-disaster recovery agent model based on the state values, action values, and reward values, respectively; and calculating experience-sharing parameters based on the prediction accuracy, total cost ratio, and degree of power shortage improvement.
[0048] In the embodiments of this application, the experience-sharing parameters for agent collaboration can be expressed as: ,in, , , These represent the experience-sharing parameters of the intelligent agent before, during, and after the disaster, respectively, satisfying... ; This indicates the accuracy of the predictions made by the pre-disaster early warning intelligent agent, i.e. , This represents the difference between the predicted time of a disaster and the actual time of its occurrence. Indicates the allowable deviation range; This represents the ratio of the change in the total cost per unit time of the disaster-stricken scheduling agent to the total cost per unit time at that time. ; This indicates the degree of improvement in power shortage for the post-disaster recovery agent, and , Indicates the area Post-disaster intelligent agents and during-disaster intelligent agents in time period The difference in power shortage; This represents the weighting coefficient of the experience-sharing parameter.
[0049] As shown in Figure 2, the following specific example further illustrates the multi-agent cooperative power system multi-timescale balancing method of this application embodiment; the embodiment of this application may include the following steps: in step S201, the early warning state value is output based on the pre-disaster early warning agent.
[0050] This application embodiment can extract features from wind and solar power output data and load time series data based on a pre-disaster early warning intelligent agent and combined with a Long Short-Term Memory (LSTM) network. By capturing the time series correlation and short-term mutation characteristics, it outputs an early warning status value that represents the disaster risk level or sudden fluctuation trend, providing input basis for the subsequent scheduling stage.
[0051] In step S202, a corresponding thermal power unit scheduling sequence is generated based on the disaster-stricken scheduling agent.
[0052] This application embodiment can be based on a disaster-stricken dispatching agent, taking early warning status values, initial full-load cost sequences of thermal power units, disaster-stricken load fluctuation constraints, and unit start-up and shutdown characteristics as inputs, and combining the risk level output during the early warning stage, to divide the dispatching process into a first half and a second half, so as to achieve the dispatching optimization focus in different time periods; through the strategy iteration process in reinforcement learning, this application embodiment can learn the impact of each dispatching action on system stability in multiple rounds of interaction, and finally output a thermal power unit dispatching sequence that meets the requirements of economy, stability and emergency response capability.
[0053] In step S203, a corresponding thermal power unit correction sequence is generated based on the post-disaster recovery agent.
[0054] As one possible approach, embodiments of this application can be based on a post-disaster recovery agent, taking early warning information, the thermal power unit scheduling sequence generated during the disaster scheduling phase, and post-disaster load changes as inputs. By modeling the characteristics of the system recovery phase, the recovery process is divided into a first half and a second half, enabling the correction actions to adapt to the different strategy requirements of the rapid recovery period and the stable adjustment period, respectively, so as to finally output the thermal power unit correction sequence, thereby achieving rapid stabilization and economic recovery of the power system in the post-disaster phase.
[0055] In step S204, a Markov game process involving multi-agent cooperation is performed.
[0056] In the embodiments of this application, a pre-disaster early warning agent, a disaster scheduling agent, and a post-disaster recovery agent constitute a multi-agent collaborative system. Each agent can complete state perception, action output, and reward feedback within a Markov game framework. The embodiments of this application can use the state information of the three types of agents as a joint state input, generate control actions based on their respective selectable action spaces, and receive multi-dimensional reward values from the environment, including indicators such as system stability, scheduling cost, and recovery speed.
[0057] In step S205, phased collaborative reinforcement learning is performed.
[0058] Specifically, in this application embodiment, the state values, action values, and reward values of the three stages of pre-disaster early warning, in-disaster scheduling, and post-disaster recovery can be used as inputs. Multi-agent experience sharing can be achieved through mechanisms such as joint experience pool, cooperative strategy gradient, or centralized training-distributed execution. Furthermore, in this application embodiment, experience sharing parameters for multi-agent cooperation can be calculated based on the above-mentioned experience data, so that the agent strategy can continuously iterate to a better solution in cross-stage collaborative training.
[0059] According to the multi-agent collaborative power system multi-timescale balancing method proposed in this application, the required early warning signals, mid-disaster scheduling sequences, and post-disaster recovery sequences are generated based on pre-constructed pre-disaster early warning agent models, mid-disaster scheduling agent models, and post-disaster recovery agent models, respectively. Then, using a Markov game framework for multi-agent collaboration, the state values, action values, and reward values of reinforcement learning are determined to obtain experience-sharing parameters for multi-agent collaboration. This enables the power system to achieve multi-timescale balancing, fully considering the temporal and spatial characteristics of extreme weather, realizing dynamic control of the power system's operating state, and improving the efficiency of extreme weather response. The system enhances the reliability of the power system under extreme weather conditions. Simultaneously, it effectively improves the scheduling efficiency of multi-agent models, enabling agents at each stage to collaboratively optimize decisions before, during, and after disasters, dynamically adjusting unit output, load allocation, and reserve resource scheduling to achieve rapid response and recovery of the power grid. Furthermore, it supports optimized management of load and units at different time scales and regional scopes, achieving stable operation of the entire power grid. Through rapid response and rolling optimization mechanisms, it shortens disaster response and recovery time, thereby significantly enhancing the overall resilience, reliability, and economic operation level of the power system under extreme weather conditions, providing strong support for safe grid operation and sustainable energy dispatch.
[0060] Next, referring to the accompanying drawings, a multi-agent cooperative power system multi-timescale balancing device according to an embodiment of this application is described.
[0061] Figure 3 is a block diagram of a multi-agent cooperative power system multi-timescale balancing device according to an embodiment of this application.
[0062] As shown in Figure 3, the multi-agent collaborative power system multi-timescale balancing device 10 includes: a pre-disaster module 100, an acquisition module 200, and an evaluation module 300.
[0063] Among them, the generation module 100 is used to generate a pre-disaster early warning signal for the power system based on a pre-built pre-disaster early warning intelligent agent model and combined with historical and predicted data on wind and solar power generation and load.
[0064] The acquisition module 200 is used to determine the scheduling sequence of thermal power units during a disaster based on the pre-disaster early warning signal and the pre-built mid-disaster scheduling intelligent agent model, and to obtain the post-disaster thermal power unit scheduling sequence using the pre-built post-disaster recovery intelligent agent model.
[0065] Evaluation module 300 is used to determine the state value, action value and reward value of reinforcement learning based on the pre-built multi-agent cooperative Markov game framework, according to the pre-disaster early warning signal, the scheduling sequence of thermal power units during the disaster and the scheduling sequence of thermal power units after the disaster, so as to determine the experience sharing parameters of multi-agent cooperation, so as to enable the power system to achieve multi-time scale balance.
[0066] Optionally, in one embodiment of this application, the acquisition module 200 includes: a first sorting unit, a calculation unit, a second sorting unit, a partitioning unit, a first acquisition unit, a second acquisition unit, and a third acquisition unit.
[0067] The first sorting unit is used to sort the thermal power units affected by the disaster according to their power upper limit in order to obtain the power sequence of the thermal power units affected by the disaster.
[0068] The calculation unit is used to calculate the full-load cost, start-up cost, and shutdown cost of the thermal power unit in the disaster according to the power sequence, so as to calculate the unit full-load cost of the thermal power unit in the disaster based on the full-load cost, start-up cost, and shutdown cost.
[0069] The second sorting unit is used to sort the unit full-load cost to obtain the initial sequence of thermal power units in the disaster area.
[0070] A partitioning unit is used to obtain the maximum and minimum values of at least one time period of fluctuating load data, and to partition the initial sequence of thermal power units in the disaster area according to the maximum and minimum values to obtain the first part of the sequence of thermal power units in the disaster area.
[0071] The first acquisition unit is used to obtain a new initial sequence of thermal power units in the disaster area based on the sequence number of the first part of the sequence.
[0072] The second acquisition unit is used to obtain the second part of the sequence of thermal power units in the disaster based on the new initial sequence of thermal power units in the disaster, the accumulated energy storage value before the disaster, the start-up cost, and the shutdown cost.
[0073] The third acquisition unit is used to obtain the disaster-affected thermal power unit scheduling sequence based on the first part sequence and the second part sequence.
[0074] Optionally, in one embodiment of this application, the acquisition module 200 includes: a second calculation unit, a correction unit, a fourth acquisition unit, a fifth acquisition unit, and a sixth acquisition unit.
[0075] The second calculation unit is used to input the disaster-affected thermal power unit scheduling sequence into a pre-built post-disaster recovery intelligent agent model to calculate the changes in post-disaster fluctuating load data and disaster-affected fluctuating load data.
[0076] The correction unit is used to correct the initial sequence of the post-disaster thermal power units based on the change value, so as to obtain the third part of the post-disaster thermal power unit sequence.
[0077] The fourth acquisition unit is used to obtain the new initial sequence of the post-disaster thermal power units based on the sequence number of the third part.
[0078] The fifth acquisition unit is used to obtain the fourth part sequence of the post-disaster thermal power units based on the new initial sequence of the post-disaster thermal power units, the start-up cost of the post-disaster thermal power units, and the shutdown cost of the post-disaster thermal power units.
[0079] The sixth acquisition unit is used to obtain the post-disaster thermal power unit scheduling sequence based on the third and fourth part sequences.
[0080] Optionally, in one embodiment of this application, the formula for calculating the state value is: ,in, Indicates the area Internal disaster pre-warning intelligent agent in time period The state vector, Indicates the area Disaster-related scheduling agents during time periods The state vector, Indicates the area Internal disaster recovery agents during the time period The state vector, Indicates the area node During the period Temperature information; Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time; Indicates the area exist Real-time status alerts; Representing regions Transmission power, disaster mitigation power, historical power generation, and thermal power unit output power at time t; Representing regions The charging power, discharging power, disaster resilience power, and historical power of the internal battery at time t; Indicates the total duration; The region is represented; the formula for calculating the action value is: ,in, These represent wind power within the region. field Photovoltaic power stations thermal power plant During the period The power action value, Indicates the area Inner Time Period At the node The load shedding power action value at the location, Indicates the area Internal energy storage devices During the period The percentage of its real-time dispatch capacity to its total power capacity; Representing the state space; the formula for calculating the reward value is: ,in, , , These represent the reward values of the agent before, during, and after the disaster, respectively. , , , , , , , These represent the wind power output penalty coefficient, the photovoltaic power output penalty coefficient, and the regional penalty coefficient, respectively. The start-up status and region of the internal thermal power unit at time t The power of the internal battery e at time t, the start-up cost of the thermal power unit g, the shutdown cost of the thermal power unit, and the secondary and primary output coefficients of the thermal power unit.
[0081] Optionally, in one embodiment of this application, the balancing module 300 includes: a third calculation unit and a fourth calculation unit.
[0082] The third calculation unit is used to calculate the prediction accuracy of the pre-disaster early warning intelligent agent model, the total cost ratio of the mid-disaster scheduling intelligent agent model, and the degree of power shortage improvement of the post-disaster recovery intelligent agent model based on the state value, action value, and reward value, respectively.
[0083] The fourth calculation unit is used to calculate experience-sharing parameters based on prediction accuracy, total cost ratio, and the degree of improvement in power shortages.
[0084] It should be noted that the foregoing explanation of the multi-agent cooperative power system multi-timescale balancing method embodiment also applies to the multi-agent cooperative power system multi-timescale balancing device of this embodiment, and will not be repeated here.
[0085] According to the multi-agent collaborative power system multi-timescale balancing device proposed in this application, based on pre-constructed pre-disaster early warning agent model, in-disaster scheduling agent model, and post-disaster recovery agent model, the required early warning signals, in-disaster thermal power unit scheduling sequences, and post-disaster thermal power unit scheduling sequences are generated respectively. Then, using a Markov game framework of multi-agent collaboration, the state values, action values, and reward values of reinforcement learning are determined to obtain experience-sharing parameters for multi-agent collaboration. This enables the power system to achieve multi-timescale balancing, fully considering the temporal and spatial characteristics of extreme weather, realizing dynamic control of the power system's operating state, and improving the efficiency of extreme weather response. The system enhances the reliability of the power system under extreme weather conditions. Simultaneously, it effectively improves the scheduling efficiency of multi-agent models, enabling agents at each stage to collaboratively optimize decisions before, during, and after disasters, dynamically adjusting unit output, load allocation, and reserve resource scheduling to achieve rapid response and recovery of the power grid. Furthermore, it supports optimized management of load and units at different time scales and regional scopes, achieving stable operation of the entire power grid. Through rapid response and rolling optimization mechanisms, it shortens disaster response and recovery time, thereby significantly enhancing the overall resilience, reliability, and economic operation level of the power system under extreme weather conditions, providing strong support for safe grid operation and sustainable energy dispatch.
[0086] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: a memory 401, a processor 402, and a computer program stored in the memory 401 and executable on the processor 402.
[0087] When the processor 402 executes the program, it implements the multi-agent cooperative power system multi-timescale balancing method provided in the above embodiments.
[0088] Furthermore, the electronic device also includes a communication interface 403 for communication between the memory 401 and the processor 402.
[0089] The memory 401 is used to store computer programs that can run on the processor 402.
[0090] Memory 401 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0091] If the memory 401, processor 402, and communication interface 403 are implemented independently, they can be interconnected via a bus to communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one thick line is used in Figure 4, but this does not indicate that there is only one bus or one type of bus.
[0092] Optionally, in a specific implementation, if the memory 401, processor 402, and communication interface 403 are integrated on a single chip, then the memory 401, processor 402, and communication interface 403 can communicate with each other through an internal interface.
[0093] Processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0094] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-agent cooperative power system multi-timescale balancing method.
[0095] This application also provides a computer program product, including a computer program that can run computer instructions. When the computer instructions are executed by a processor, they implement the multi-agent cooperative power system multi-timescale balancing method provided in this application.
[0096] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0097] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0098] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0099] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0100] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0101] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0102] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0103] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A multi-agent cooperative method for balancing power systems across multiple time scales, characterized in that, Includes the following steps: Based on a pre-built pre-disaster early warning intelligent agent model, a pre-disaster early warning signal for the power system is generated by combining historical and predicted data on wind and solar power generation and load. Based on the pre-built pre-disaster scheduling intelligent agent model, the scheduling sequence of thermal power units during the disaster is determined, and the scheduling sequence of thermal power units after the disaster is obtained using a pre-built post-disaster recovery intelligent agent model. Based on a pre-built multi-agent collaborative Markov game framework, the state value, action value, and reward value of reinforcement learning are determined according to the pre-disaster early warning signal, the scheduling sequence of thermal power units during the disaster, and the scheduling sequence of thermal power units after the disaster, so as to determine the experience sharing parameters for multi-agent collaboration, enabling the power system to achieve multi-timescale balance.
2. The method according to claim 1, characterized in that, The process of determining the scheduling sequence of thermal power units during a disaster includes: sorting the thermal power units during a disaster according to their maximum power capacity to obtain a power sequence; calculating the full-load cost, start-up cost, and shutdown cost of each thermal power unit during a disaster based on the power sequence, and calculating the unit full-load cost of each thermal power unit during a disaster based on the full-load cost, start-up cost, and shutdown cost; sorting the unit full-load costs to obtain an initial sequence of thermal power units during a disaster; obtaining the maximum and minimum values of at least one period of fluctuating load data, and dividing the initial sequence of thermal power units during a disaster based on the maximum and minimum values to obtain a first part sequence of thermal power units during a disaster; obtaining a new initial sequence of thermal power units during a disaster based on the sequence number of the first part sequence; obtaining a second part sequence of thermal power units during a disaster based on the new initial sequence of thermal power units during a disaster, the pre-disaster accumulated energy storage value, the start-up cost, and the shutdown cost; and obtaining the scheduling sequence of thermal power units during a disaster based on the first part sequence and the second part sequence.
3. The method according to claim 1, characterized in that, The step of obtaining the post-disaster thermal power unit scheduling sequence using a pre-built post-disaster recovery intelligent agent model includes: inputting the in-disaster thermal power unit scheduling sequence into the pre-built post-disaster recovery intelligent agent model to calculate the changes in post-disaster fluctuating load data and in-disaster fluctuating load data; correcting the initial sequence of post-disaster thermal power units based on the changes to obtain a third part sequence of post-disaster thermal power units; obtaining a new initial sequence of post-disaster thermal power units based on the sequence number of the third part sequence; obtaining a fourth part sequence of post-disaster thermal power units based on the new initial sequence, the start-up cost, and the shutdown cost; and obtaining the post-disaster thermal power unit scheduling sequence based on the third part sequence and the fourth part sequence.
4. The method according to claim 1, characterized in that, The formula for calculating the state value is: ,in, Indicates the area Internal disaster pre-warning intelligent agent in time period The state vector, Indicates the area Disaster-related scheduling agents during time periods The state vector, Indicates the area Internal disaster recovery agents during the time period The state vector, Indicates the area node During the period Temperature information; Indicates the area Inland wind farm exist Predicted power at time, Indicates the area Inland wind farm exist Actual scheduling power at any given time; Indicates the area Internal photovoltaic power station exist Predicted power at time, Indicates the area Internal photovoltaic power station exist Actual scheduling power at any given time; Indicates the area internal nodes In The actual load at any given time; Indicates the area exist Real-time status alerts; Representing regions Transmission power, disaster mitigation power, historical power generation, and thermal power unit output power at time t; Representing regions The charging power, discharging power, disaster resilience power, and historical power of the internal battery at time t; Indicates the total duration; The region is represented; the formula for calculating the action value is: ,in, These represent wind power within the region. field Photovoltaic power stations thermal power plant During the period The power action value, Indicates the area Inner Time Period At the node The load shedding power action value at the location, Indicates the area Internal energy storage devices During the period The percentage of its real-time dispatch capacity to its total power capacity; The state space is represented; the formula for calculating the reward value is: ,in, 、 、 These represent the reward values of the agent before, during, and after the disaster, respectively. 、 、 、 、 、 、 、 These represent the wind power output penalty coefficient, the photovoltaic power output penalty coefficient, and the regional penalty coefficient, respectively. The start-up status and region of the internal thermal power unit at time t The power of the internal battery e at time t, the start-up cost of the thermal power unit g, the shutdown cost of the thermal power unit, and the secondary and primary output coefficients of the thermal power unit.
5. The method according to claim 1, characterized in that, The step of using the state values, action values, and reward values of the reinforcement learning to perform phased collaborative reinforcement learning to determine the experience-sharing parameters for multi-agent cooperation includes: calculating the prediction accuracy of the pre-disaster early warning agent model, the total cost ratio of the in-disaster scheduling agent model, and the degree of power shortage improvement of the post-disaster recovery agent model based on the state values, action values, and reward values, respectively; and calculating the experience-sharing parameters based on the prediction accuracy, the total cost ratio, and the degree of power shortage improvement.
6. A multi-agent cooperative power system multi-timescale balancing device, characterized in that, include: The generation module is used to generate pre-disaster warning signals for the power system based on a pre-built pre-disaster warning intelligent agent model, combined with historical and predicted data on wind and solar power generation and load. The acquisition module is used to determine the scheduling sequence of thermal power units during a disaster based on the pre-disaster early warning signal and the pre-built mid-disaster scheduling intelligent agent model, and to obtain the post-disaster thermal power unit scheduling sequence using the pre-built post-disaster recovery intelligent agent model. The balancing module is used to determine the state value, action value, and reward value of reinforcement learning based on the pre-built multi-agent cooperative Markov game framework, the pre-disaster early warning signal, the mid-disaster thermal power unit scheduling sequence, and the post-disaster thermal power unit scheduling sequence, so as to determine the experience sharing parameters of multi-agent cooperation, so that the power system can achieve multi-time scale balance.
7. The apparatus according to claim 6, characterized in that, The acquisition module includes: a first sorting unit, used to sort the disaster-affected thermal power units according to their power upper limit to obtain a power sequence of the disaster-affected thermal power units; a calculation unit, used to calculate the full-load cost, start-up cost, and shutdown cost of the disaster-affected thermal power units according to the power sequence, and to calculate the unit full-load cost of the disaster-affected thermal power units based on the full-load cost, start-up cost, and shutdown cost; a second sorting unit, used to sort the unit full-load cost to obtain an initial sequence of disaster-affected thermal power units; a partitioning unit, used to acquire the maximum and minimum values of at least one time period fluctuation load data, and to partition the initial sequence of disaster-affected thermal power units according to the maximum and minimum values to obtain a first part sequence of the disaster-affected thermal power units; a first acquisition unit, used to obtain a new initial sequence of disaster-affected thermal power units based on the sequence number of the first part sequence; a second acquisition unit, used to obtain a second part sequence of the disaster-affected thermal power units according to the new initial sequence of the disaster-affected thermal power units, the pre-disaster energy storage accumulation value, the start-up cost, and the shutdown cost; and a third acquisition unit, used to obtain a scheduling sequence of the disaster-affected thermal power units according to the first part sequence and the second part sequence.
8. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the multi-agent cooperative power system multi-timescale balancing method as described in any one of claims 1-5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the multi-agent cooperative power system multi-timescale balancing method as described in any one of claims 1-5.
10. A computer program product, comprising a computer program, characterized in that, The computer program is executed to implement a multi-agent cooperative power system multi-timescale balancing method as described in any one of claims 1-5.