Integrated energy multi-agent collaborative optimization scheduling method and system under extreme event
By constructing a spatiotemporal feature extraction method based on mutual information and fast Fourier transform, and a multi-agent deep reinforcement learning framework, the shortcomings of the scheduling model of integrated energy system under extreme events are solved, and more efficient and safer multi-energy flow collaborative optimization is achieved.
Patent Information
- Application Number
- CN202511510542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing integrated energy system scheduling models under extreme events suffer from problems such as a single decision center being susceptible to disasters, a lack of dynamic response in multi-energy flow coupling mechanisms, and insufficient safety and reliability.
A spatiotemporal feature extraction method based on mutual information and fast Fourier transform is adopted, combined with a multi-agent deep reinforcement learning framework, to construct a four-dimensional dynamic energy flow coupling model of electricity-gas-heat-storage, and optimize the scheduling strategy through distributed decision-making and risk perception mechanisms.
It enhances the collaborative optimization capability and security of the integrated energy system under extreme events, reduces load loss, and improves the economy and adaptability of dispatching.
Smart Images

Figure CN120996517B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of integrated energy system scheduling technology, and more specifically, relates to a multi-agent collaborative optimization scheduling model for integrated energy systems under the influence of extreme events. Background Technology
[0002] Extreme events, including natural disasters such as typhoons and thunderstorms, pose a dual and severe challenge to integrated energy systems: First, the risk of energy supply disruptions, as these events often damage energy infrastructure, leading to regional power outages, affecting residents' lives and impacting industrial production; second, system instability, as extreme events disrupt the system's operational balance, causing imbalances between power generation and consumption, significant fluctuations in frequency and voltage, and in severe cases, system collapse; heating and gas systems also suffer from abnormal heating and gas supply due to unstable pipeline pressure and heat / gas sources, even leading to safety accidents that endanger lives and property.
[0003] As a complex energy supply system with multiple energy flows coupled together, the scheduling and optimization technology of integrated energy systems has become a key support for ensuring energy security. Current mainstream optimization scheduling models mostly adopt a centralized decision-making architecture, which can achieve energy supply and demand balance under normal operating conditions, but has significant limitations under extreme events: First, a single decision-making center is susceptible to chain-like failures caused by disasters, has insufficient system topology reconfiguration capabilities, and struggles to cope with emergencies such as equipment damage and communication interruptions; second, existing models focus on steady-state analysis in characterizing the coupling mechanism of electric, gas, and heat flows, lacking refined modeling of dynamic responses under extreme disturbances, resulting in weak multi-timescale coordination capabilities; third, traditional optimization objectives often focus on economic indicators, insufficiently considering safety factors such as energy supply reliability and fault recovery speed under extreme scenarios.
[0004] In conclusion, given the crucial role of integrated energy systems in modern energy supply and the severe challenges posed by extreme events, and considering the numerous shortcomings of existing technologies in responding to extreme events, there is an urgent need for an innovative multi-agent collaborative optimization scheduling method for integrated energy systems. This method aims to improve the collaborative optimization capabilities of integrated energy systems under extreme events and ensure the safety, reliability, efficiency, and sustainability of energy supply.
[0005] Existing technical document 1 (CN119539424A) discloses a resource regulation method, device, equipment, and medium for adjustable loads at the park level. Its shortcomings include an inability to cope with extreme event disturbances, reliance on a centralized communication architecture, and a lack of a dynamic security risk quantification mechanism. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a spatiotemporal feature extraction method based on mutual information and fast Fourier transform. It uses mutual information weights to select key variables and fuses frequency domain features to achieve adaptive dimensionality reduction of input data. For four types of extreme events—typhoons, extreme cold, extreme heat, and thunderstorms—a multi-factor coupled power outage probability model is established to quantify the nonlinear correlation between meteorological parameters and energy system vulnerability. A four-dimensional dynamic energy flow coupled differential equation system (electricity-gas-heat-storage) is proposed to accurately characterize the cross-medium propagation mechanism and spatiotemporal evolution of multi-energy flows under extreme disturbances. A risk perception action selection mechanism based on MADDPG (Multi Agent Deep Deterministic Policy Gradient) is constructed, embedding a dynamically adjusted demand matrix as a constraint into the agent's decision-making process. Distributed collaborative optimization of safety constraints and economic goals is achieved through Nash equilibrium solution.
[0007] The present invention adopts the following technical solution.
[0008] The first aspect of this invention provides a method for multi-agent cooperative optimization scheduling of integrated energy under extreme events, comprising the following steps:
[0009] The spatiotemporal characteristics of multiple agents in the integrated energy system are extracted by mutual information and fast Fourier transform. After knowledge distillation and lightweighting, a lightweight general model of agents is constructed and solved to obtain the global system state and the optimal action strategy of each agent under the global system state.
[0010] Construct power outage probability models for various extreme events, including typhoon events, extreme cold events, high temperature events, and thunderstorm events;
[0011] Based on the base power loss of the power system, gas system, heat system, and energy storage system caused by extreme events, a multi-energy disturbance propagation model driven by extreme events is constructed, and the model is solved to generate a regulation demand matrix.
[0012] By combining the global system state and the optimal action strategies of each agent under the global system state with the power outage probability model and adjustment demand matrix of each extreme event, a multi-agent cooperative strategy model is constructed using multi-agent deep reinforcement learning.
[0013] By inputting real-time monitoring data of integrated energy multi-agents into the multi-agent collaborative strategy model, a collaborative optimization scheduling strategy for integrated energy multi-agents under extreme events is obtained.
[0014] Preferably, the construction and solution of the lightweight general model of the intelligent agent includes:
[0015] The energy physical equipment of the integrated energy system is divided into multiple types of intelligent agents according to its functional attributes, resulting in a multi-type intelligent agent cluster of the integrated energy system.
[0016] Data from a cluster of multiple intelligent agents in an integrated energy system is acquired, and spatiotemporal features are extracted using mutual information and fast Fourier transform to obtain spatiotemporal features.
[0017] The spatiotemporal features are transferred from the teacher model to the student model through knowledge distillation to obtain lightweight student model data.
[0018] The lightweight student model data is set as the global system state. A lightweight general model of agents is constructed by combining multi-agent Nash equilibrium. The optimal action strategy of each agent in the global system state is obtained by solving the lightweight general model of agents.
[0019] Preferably, the method of solving the lightweight general model of the agents to obtain the optimal action strategies of each agent in the global system state includes:
[0020] Construct a reward function for each agent based on its own actions and the actions of other agents;
[0021] Set the lightweight student model data as the global system state, input the security constraint function, and if the security constraint function is violated, set the security constraint relaxation term to the output of the security constraint function; otherwise, set the security constraint relaxation term to 0.
[0022] The optimal action strategy of the current agent is obtained by subtracting the product of the balance coefficient and the safety constraint relaxation term from the reward function of each agent.
[0023] The optimal action strategies of each agent are summed to obtain the optimal action strategies of each agent in the global system state.
[0024] Preferably, the construction of the power outage probability model for each extreme event includes:
[0025] The probability of power outages caused by precipitation, the cumulative risk probability corresponding to the degree of wind speed deviation, and the joint risk degree of precipitation and wind speed are solved by wind speed and precipitation amount. Based on the solution results, a power outage probability model for typhoon events is constructed.
[0026] The probability of power outage caused by extreme cold temperature and the probability of power outage caused by the duration of extreme cold event are solved by using temperature and duration, and a power outage probability model for extreme cold event is constructed based on the solution results.
[0027] The probability of power outages caused by high temperatures is calculated by temperature, and a power outage probability model for high-temperature events is constructed by combining it with air conditioning load rate.
[0028] The probability of power outages caused by lightning density and the nonlinear saturation risk of rainfall are solved by using lightning density and rainfall amount, and a power outage probability model for thunderstorm events is constructed based on the solution results.
[0029] Preferably, the generation of the adjustment demand matrix includes:
[0030] Based on the power system base power loss caused by extreme events, construct the power system energy flow model; based on the gas system base power loss caused by extreme events, construct the gas system energy flow model; based on the thermal system base power loss caused by extreme events, construct the thermal system energy flow model; based on the energy storage system base power loss caused by extreme events, construct the energy storage system energy flow model.
[0031] A multi-energy disturbance propagation model for electricity-gas-heat-storage is constructed based on the energy flow models of power systems, gas systems, heat systems, and energy storage systems.
[0032] Solve the multi-energy disturbance propagation model of electricity-gas-heat-storage to generate the power replenishment required by the power system, the power replenishment required by the gas system, the heat replenishment required by the thermal system, and the charging and discharging demand of the thermal system and the energy storage system. Construct a regulation demand matrix based on the generated results.
[0033] Preferably, the construction of the multi-agent cooperative strategy model includes:
[0034] The basic state space and basic action space of the multi-agent are obtained based on the global system state and the optimal action strategy of each agent under the global system state. The extreme event state space and extreme event action space are solved by combining the power outage probability model of extreme events.
[0035] The single reward value is calculated based on the cost-effectiveness, global voltage security status and regulation efficiency calculation results. Then, a risk-sensitive reward function is constructed by combining the power outage probability model and regulation demand matrix of each extreme event.
[0036] Based on the state space and action space of extreme events, as well as the risk-sensitive reward function, a multi-agent collaborative optimization model is constructed using a multi-agent deep reinforcement learning algorithm.
[0037] Preferably, the extreme event action space includes:
[0038] The action is differentiated by the power outage probability model, and the rate of change of the power outage probability after the action is executed is solved. The rate of change of the power outage probability after the action and the amount of load cut off after the action are multiplied by the corresponding weight coefficients and summed to obtain the load loss risk quantification score of the action.
[0039] If the load loss risk quantification score of the action corresponding to the action located in the basic action space is less than a set threshold, the extreme event action space is obtained.
[0040] Preferably, the construction of the risk-sensitive reward function includes:
[0041] Multiply the risk sensitivity coefficient by the power outage probability model to obtain the risk penalty term. Take the negative value of the risk penalty term and use the exponential function as the base to obtain the exponential decay function. Then, subtract 1 to obtain the linear penalty value for the power outage.
[0042] The absolute error between the actual power adjustment amount generated by the agent's actions and the adjustment demand matrix is calculated. This error is divided by the sum of the adjustment demand matrix and the parameter to prevent division by zero, to obtain the relative error rate. Subtracting the relative error rate from 1 yields the generation adjustment accuracy score.
[0043] The risk-sensitive reward function is obtained by multiplying the linear penalty value for power outages and the adjustment accuracy score by their respective weights, summing the results, and then adding them to a single reward value.
[0044] Preferably, the construction of the multi-agent cooperative optimization model includes:
[0045] Construct the Actor network for each agent based on the state space of extreme events, and construct the Critic network for each agent based on the action space and state space of extreme events.
[0046] Summing the actions output by each agent from the Actor network with the exploration noise, we obtain the actions of each agent that satisfy the extreme event action space after the summation result.
[0047] Construct joint action of agents based on the actions of each agent, execute joint action of agents, obtain risk-sensitive reward of multi-agent and new multi-agent state, and construct transition tuple with extreme event state space and joint action of agents, and store it in experience pool.
[0048] The transition tuples are obtained from the experience pool and a multi-agent cooperative policy model is trained. The preliminary multi-agent cooperative policy model is then obtained by solving the problem.
[0049] The action policies of the multi-agents output by the preliminary multi-agent cooperative strategy model are subjected to Nash equilibrium to solve the multi-agent cooperative strategy and construct the multi-agent cooperative strategy model.
[0050] The second aspect of the present invention provides a comprehensive energy multi-agent cooperative optimization scheduling system under extreme events, which runs the comprehensive energy multi-agent cooperative optimization scheduling method under extreme events described in the first aspect, including:
[0051] The multi-agent state extraction module is used to extract the spatiotemporal characteristics of multiple agents in the integrated energy system through mutual information and fast Fourier transform. After knowledge distillation and lightweighting, a lightweight general model of agents is constructed and solved to obtain the global system state and the optimal action strategy of each agent under the global system state.
[0052] The power outage probability solution module is used to build power outage probability models for various extreme events, including typhoon events, extreme cold events, high temperature events, and thunderstorm events.
[0053] The demand solving module is used to construct an extreme event-driven multi-energy disturbance propagation model of electricity-gas-heat-storage based on the base power loss of the power system, gas system, thermal system, and energy storage system caused by extreme events, solve the model, and generate the regulation demand matrix.
[0054] The collaborative strategy model construction module is used to combine the global system state and the optimal action strategies of each agent under the global system state with the power outage probability model and adjustment demand matrix of each extreme event, and to construct a multi-agent collaborative strategy model using multi-agent deep reinforcement learning.
[0055] The output module is used to input real-time monitoring data of the integrated energy multi-agent into the multi-agent collaborative strategy model to obtain the integrated energy multi-agent collaborative optimization scheduling strategy under extreme events.
[0056] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0057] This invention uses knowledge distillation technology, supports edge deployment, and ensures that the integrated energy system can still operate independently when communication is interrupted due to extreme events, thereby improving the security of integrated energy system scheduling. It adopts the Nash equilibrium method of distributed decision architecture to reward and punish each agent, forcing agents to consider the actions of other agents when pursuing their own benefits, reducing multi-agent policy conflicts and improving the efficiency of multi-agent collaborative optimization of the integrated energy system.
[0058] This invention enhances the basic state space of multiple agents by constructing a power outage probability model for extreme events, thereby improving the agents' ability to perceive risks under extreme events. By constraining the basic action space of multiple agents through the power outage probability model for extreme events, the integrated energy system restricts the agents' action choices to extreme events, thereby constraining high-risk actions under extreme events and improving the integrated energy system's adaptive ability and security under extreme events.
[0059] This invention generates scheduling requirements by constructing a multi-energy disturbance propagation model of electricity-gas-heat-storage, and combines the scheduling requirements with a power outage probability model of extreme events to generate a risk-sensitive reward function for extreme events, thereby reducing the load loss of the integrated energy system during extreme events and improving the economic efficiency of integrated energy system regulation.
[0060] This invention reduces the conflict rate of a comprehensive energy system under multi-agent strategy regulation by balancing the strategies of multiple agents through Nash equilibrium, thereby improving the security of the comprehensive energy system under multi-agent strategy regulation. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of the integrated energy multi-agent collaborative optimization scheduling under extreme events provided in accordance with the embodiments of the present invention. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0063] like Figure 1 As shown, Embodiment 1 of the present invention provides a comprehensive energy multi-agent cooperative optimization scheduling method under extreme events, including the following steps:
[0064] Step 1: Divide the multi-agent system according to the functional attributes of the equipment. Extract the spatiotemporal features of the multi-agent system through mutual information and fast Fourier transform. After knowledge distillation and lightweighting, construct a lightweight general model of the agents and solve it to obtain the global system state and the optimal action strategy of each agent under the global system state.
[0065] In a preferred but non-limiting embodiment of the present invention, step 1 includes:
[0066] Step 1.1: Based on the functional attributes of the energy physical equipment of the integrated energy system, such as generator sets, hydrogen production devices, gas storage tanks, and industrial flexible loads, the energy physical equipment is divided into four types of intelligent agents: energy production, energy conversion, energy storage, and adjustable loads, thus obtaining a multi-type intelligent agent cluster of the integrated energy system.
[0067] More preferably, step 1.1 includes:
[0068] Producer Agents (PAs) encompass all energy production units, enabling the conversion of primary energy into secondary energy sources such as electricity, heat, and cooling. These primarily include photovoltaic arrays, gas turbine generator sets, industrial waste heat boilers, solar thermal towers, and combined cooling, heating, and power (CCHP) systems.
[0069] Energy conversion agents (CA) include all energy form conversion units, realizing multi-energy coupling and energy flow direction control. They mainly involve electric boilers for electric-to-thermal conversion, absorption chillers for electric-to-cooling conversion, gas boilers for gas-to-heat conversion, and heat pump systems for multi-directional coupling conversion.
[0070] Storage Agents (SAs) encompass all energy storage units, enabling energy shifting and power buffering across time scales, including lithium-ion battery packs, thermal storage tanks, and gas storage tanks.
[0071] Adjustable Load Agents (ALA) cover all energy end-use units and have demand response capabilities and energy consumption pattern optimization functions, including industrial load, commercial load, residential load, traffic load and special load.
[0072] Step 1.2: Obtain real-time monitoring data, prediction and scheduling data, system coupling status, and environmental and event signals from the multi-type intelligent agent cluster of the integrated energy system obtained in Step 1.1. Spatiotemporal features are extracted using mutual information and Fast Fourier Transform to obtain the spatiotemporal features, expressed by the following formula:
[0073]
[0074] In the formula, H t express t The spatiotemporal characteristics of time, Concat represents concatenation, k represents the sliding index within the time window, its range is from the start time k=tT of the time window to the end time k=t of the time window, t represents the end time, x k The input data at time k includes real-time monitoring data such as, but not limited to, power and temperature; forecast data such as, but not limited to, weather forecasts; scheduling data such as, but not limited to, scheduling instructions; and system coupling status such as, but not limited to, SOC and health status. T represents the adaptive sliding window length, which dynamically adjusts the time range and is set to 5-30 minutes. This represents the input data within the time window, starting from time... The feature sequence to t; FFT(.) represents Fourier transform, MutualInfo(x, y) represents the mutual information between feature x and target variable y, used to calculate feature importance weights. Feature x represents the input item in Table 1, and target variable y represents the output item in Table 1, including optimized control instructions, state information of participating agents, performance indicators, etc. Among them, performance indicators include economic performance, collaborative performance and safety performance. Economic performance includes, but is not limited to, local scheduling plan and market bidding capacity. Collaborative performance includes, but is not limited to, collaborative boundary constraints and cross-agent coordination parameters. Safety performance includes, but is not limited to, emergency adjustment strategy and safety limit alarm.
[0075] Table 1 Input and output data of the lightweight general model for intelligent agents
[0076]
[0077] Based on the selection of key variables based on mutual information, features strongly correlated with target control are retained, such as, but not limited to, the irradiance-output nonlinear relationship features of photovoltaic smart agents. Fast Fourier transform is performed on x within the time window to extract frequency domain features, such as, but not limited to, periodicity and fluctuation intensity.
[0078] It is worth noting that this invention filters time-domain correlation features that are strongly correlated with the control objectives of multi-agent systems through mutual information, and extracts periodic frequency-domain features from multi-agent data using fast Fourier transform. This can better capture the temporal patterns of energy load. By combining time-domain correlation feature filtering and frequency-domain pattern feature extraction, it is more comprehensive and efficient than simply using time-domain or frequency-domain methods.
[0079] Step 1.3: The spatiotemporal features obtained in Step 1.2 are transferred from the teacher model to the student model through knowledge distillation for lightweight processing, resulting in lightweight student model data.
[0080] More preferably, step 1.3 includes:
[0081] A teacher-student distillation-based LSTM (Long Short-Term Memory) compression model is adopted. Through knowledge distillation, knowledge from the complex teacher model built with BiGRU (Bidirectional Gated Recurrent Unit), such as but not limited to output distribution and hidden states, is transferred to the lightweight student model built with PrunedLSTM (Pruned Long Short-Term Memory). This allows the student model to maintain high performance while reducing the number of parameters. The full-size teacher model uses BiGRU to extract temporal features, resulting in a large parameter scale and strong performance. The lightweight student model reduces LSTM parameters through pruning, lowering computational and storage requirements, thus achieving lightweighting, as expressed by the following formula:
[0082]
[0083] In the formula, For teacher models in t The hidden state at all times For teacher models in t The hidden state at time -1 H t for t The spatiotemporal characteristics of a moment i tea For teacher model weights, For student models in t The hidden state at all times For student models in t The hidden state at time -1 i stu For student model weights; L distill This is the distillation loss function, used to ensure that the lightweight model can still approximate the performance of the original model; l This is used as a distillation weight to balance the contribution of knowledge distillation and task loss. KL (·∥·) represents the KL divergence, which measures the difference between the output distributions of teachers and students; s (·) is the Softmax function, which transforms the hidden state into a probability distribution; t It is a temperature coefficient used to soften probability distributions and enhance knowledge transfer effects; L task This represents the loss of the student model on the target task.
[0084] Step 1.4: Set the lightweight student model data as the global system state, construct a lightweight general model of agents by combining multi-agent Nash equilibrium, and solve the lightweight general model of agents to obtain the optimal action strategy of each agent in the global system state.
[0085] More preferably, step 1.4 includes:
[0086] Nash equilibrium is suitable for non-cooperative game scenarios, aiming to maximize the interests of each agent. It more closely reflects the decentralized decision-making needs of real-world multi-agent systems. Nash equilibrium enables distributed collaborative optimization of the economics and security of multiple agents, solving for the optimal action strategies of each agent in the global system state, including:
[0087] Construct a reward function for each agent based on its own actions and the actions of other agents;
[0088] Set the lightweight student model data as the global system state, input the security constraint function, and if the security constraint function is violated, set the security constraint relaxation term to the output of the security constraint function; otherwise, set the security constraint relaxation term to 0.
[0089] The optimal action strategy of the current agent is obtained by subtracting the product of the balance coefficient and the safety constraint relaxation term from the reward function of each agent.
[0090] Summing the optimal action policies of each agent yields the optimal action policies of each agent in the global system state. The lightweight general model of agents is expressed by the following formula:
[0091]
[0092] In the formula, N The number of intelligent agents; in this invention, N is 4. For the first i The reward function of each agent depends on its own actions. Actions with other intelligent agents Margin(g) j ( S )) represents a safety constraint relaxation term, such as, but not limited to, voltage over-limit penalties. When the safety constraint function g j ( S ) is violated, that is g j ( S A penalty is incurred when )>0. j ( S ))=g j ( S Otherwise, Margin(g) j (S ))=0;g j ( S ) indicates the first j The specific form of the safety constraint function is determined by the physical safety boundary of the integrated energy system, such as, but not limited to, voltage over-limit constraints in power systems and temperature constraints in thermal systems. S The global system state, containing the basic state space of all agents, is the lightweight student model data output from step 1.3. , set as the input state for Nash equilibrium solution, contains compressed spatiotemporal features, used to describe the global system state, such as, but not limited to, power, temperature, etc.; c This is a balancing coefficient used to adjust the weighting of safety constraints and economic rewards, thereby increasing... c ↑ indicates stricter security constraints, prioritizing system stability and reducing [risks / damage]. c The ↓ indicates a greater focus on economic efficiency, while allowing for a certain degree of security risk.
[0093] Step 2: Construct a power outage probability model for each extreme event, which includes typhoon events, extreme cold events, high temperature events, and thunderstorm events.
[0094] In a preferred but non-limiting embodiment of the present invention, step 2 includes:
[0095] Step 2.1: Solve the probability that precipitation directly causes power outage, the cumulative risk probability corresponding to the degree of wind speed deviation, and the joint risk degree of precipitation and wind speed by wind speed and precipitation. Based on the solution results, construct a power outage probability model for typhoon events to quantify the power outage risk caused by typhoon events.
[0096] More preferably, step 2.1 includes:
[0097] The risk-weighted value of precipitation is obtained by multiplying the sensitivity of the precipitation impact by the precipitation amount. Using an exponential function as the base and the negative of the risk-weighted value as the exponent, a saturation effect of precipitation impact is simulated to obtain the attenuation trend of precipitation-induced power outages. Subtracting this attenuation trend from 1 yields the probability that precipitation directly causes a power outage. When the precipitation amount R is small, such as light rain, the risk-weighted value of precipitation is small, and the impact of precipitation on the probability of power outages increases approximately linearly. During light rain, the risk of power outages increases slowly with increasing rainfall. When the precipitation amount is very large, such as during torrential rain, the risk-weighted value of precipitation is very large. When the value approaches 1, the contribution of precipitation to the probability of power outage reaches saturation, describing a scenario where the risk of power outage due to excessive rainfall is close to its maximum value, and further increasing rainfall has limited effect on increasing the risk.
[0098] Subtract the wind speed threshold from the actual wind speed and then divide by the dispersion of the wind speed risk distribution to obtain the standardized deviation of the actual wind speed from the wind speed threshold. Input the cumulative distribution function of the standard normal distribution to obtain the cumulative risk probability corresponding to the wind speed deviation.
[0099] Multiplying the precipitation amount by the actual wind speed yields the intensity of the combined effect of precipitation and wind speed. Normalizing this effect using a normalization coefficient gives the combined risk level of precipitation and wind speed.
[0100] The probability of a power outage due to typhoon events is calculated by multiplying the probability of a direct power outage caused by precipitation by a precipitation weight, the cumulative risk probability corresponding to the degree of wind speed deviation by a wind speed weight, and the combined risk of precipitation and wind speed by a combined weight of precipitation and wind speed. The sum of these multiplications yields the power outage probability of the typhoon event. The power outage probability model for typhoon events is expressed by the following formula:
[0101]
[0102] In the formula, P outage,1 Indicates the probability of power outages during a typhoon event; R Rainfall is measured in mm / h and is used to quantify the direct impact of rainfall on power facilities, such as, but not limited to, short circuits and flooding. V Φ(·) represents the actual wind speed in m / s, reflecting the physical damage caused by wind to power facilities, such as, but not limited to, pole collapse and wire breakage; Φ(·) is the cumulative distribution function of the standard normal distribution, used to describe wind speeds exceeding a threshold. V The risk spike at 0; α 1, β 1, c 1 represents the weight of precipitation, the weight of wind speed, and the weight of the combined effect of precipitation and wind speed, respectively. These are calibrated based on historical data and are used to reflect the relative importance of precipitation, wind speed, and their combined effect. k 1, V 0, s , k 2 represents the sensitivity (h / mm), wind speed threshold (m / s), and dispersion (m / s) of wind speed risk distribution, respectively, and their normalization coefficients. .
[0103] Step 2.2: Solve the probability of power outage caused by extreme cold temperature and the probability of power outage caused by the duration of extreme cold event by using temperature and duration, and construct a power outage probability model for extreme cold event based on the solution results, so as to quantitatively assess the impact of extreme cold weather on the power system.
[0104] More preferably, step 2.2 includes:
[0105] The temperature deviation is obtained by subtracting the extreme cold temperature threshold from the actual extreme cold temperature and dividing by the dispersion parameter of the temperature distribution. The cumulative distribution function of the standard normal distribution is then input to obtain the probability of a power outage caused by extreme cold temperature.
[0106] Multiplying the decay rate of the duration of the extreme cold event by the duration of the extreme cold event yields the risk weighted value of the duration of the extreme cold event. Using an exponential function as the base and the negative of the risk weighted value of the duration of the extreme cold event as the exponent, we obtain the decay trend of the impact of the duration of the extreme cold event on the power outage. Subtracting the decay trend of the duration of the extreme cold event on the power outage from 1 yields the probability that the duration of the extreme cold event will cause a power outage.
[0107] The probability of a power outage caused by extreme cold temperature is obtained by multiplying the probability of a power outage caused by temperature by a weighting coefficient, and the probability of a power outage caused by the duration of an extreme cold event is obtained by multiplying the probability of a power outage caused by the duration of the event by a weighting coefficient. The sum of these multiplications yields the probability of a power outage during an extreme cold event. The power outage probability model for an extreme cold event is expressed by the following formula:
[0108]
[0109] In the formula, P outage,2 α represents the probability of a power outage during an extreme cold event; α² and β² are the weighting coefficients for the effects of temperature and duration, respectively; Φ is the cumulative distribution function of the standard normal distribution, used to quantify the risk when the temperature is below a threshold; T is the actual extreme cold temperature, in °C; T c The extreme cold temperature threshold is expressed in °C. An actual temperature below this threshold is considered the actual extreme cold temperature. In this invention, T... c Take -5℃; σ T is the parameter representing the dispersion of the temperature distribution, in °C; t is the duration of the extreme cold event, in hours. k t The decay rate over time, in hours (h). -1 ; It is an exponentially decaying function that describes the dynamic impact of duration on risk.
[0110] Step 2.3: Solve for the probability of power outage caused by high temperature by temperature calculation, and construct a power outage probability model for high temperature events by combining the air conditioning load rate.
[0111] More preferably, step 2.3 includes:
[0112] The actual high temperature is subtracted from the high temperature threshold, and multiplied by the temperature sensitivity coefficient to obtain the temperature deviation. This deviation is then input into an S-shaped function to calculate the probability of a power outage caused by high temperature. The S-shaped function is used to describe the smooth transition of temperature exceeding the limit and avoid false alarms caused by sudden temperature threshold changes.
[0113] The probability of a power outage due to a high-temperature event is obtained by multiplying the square of the air conditioning load rate by its weighted effect, and by multiplying the probability of a power outage caused by high temperature by its weighted effect coefficient. The sum of these multiplications yields the probability of a power outage during a high-temperature event. The square of the air conditioning load rate is used to describe the overload caused by a surge in air conditioning usage, thus identifying the actual overload risk. The power outage probability model for high-temperature events is expressed by the following formula:
[0114]
[0115] In the formula, P outage,3 k represents the probability of a power outage during a high-temperature event. T Sensitivity coefficient for temperature risk, in °C -1 α3 and β3 are the weighting coefficients for the effect of high temperature events and the effect of air conditioning load rate, respectively. T’ This represents the actual high temperature, expressed in °C, which directly affects transformer heat dissipation and photovoltaic efficiency. T 0 ’ This is the high temperature threshold, expressed in °C. An actual temperature greater than the high temperature threshold is considered the actual high temperature. L AC The air conditioning load factor reflects the load pressure on the power grid exerted by air conditioning use. (sigmoid( . ) is a sigmoid function. x )=1 / (1+ e -x This is used to smooth out the sharp increase in risk caused by temperature.
[0116] Step 2.4: Solve for the probability of power outage caused by lightning density and the nonlinear saturation risk of rainfall by using lightning density and rainfall, and construct a power outage probability model for thunderstorm events based on the solution results.
[0117] More preferably, step 2.4 includes:
[0118] The lightning density is normalized by multiplying it by the normalization coefficient of the lightning density to obtain the risk-weighted value of the lightning density. Using an exponential function as the base and the negative of the risk-weighted value of the lightning density as the exponent, the attenuation trend of the impact of lightning density on power outages is obtained. Subtracting the attenuation trend of the impact of lightning density on power outages from 1 gives the probability of lightning density causing power outages, which is used to describe the cumulative effect of lightning.
[0119] The hyperbolic tangent function is set as the rainstorm saturation function. The rainfall is normalized by the normalization coefficient of the rainfall. The nonlinear saturation risk of the rainfall is obtained by inputting the hyperbolic tangent function.
[0120] The nonlinear saturation risk of rainfall is multiplied by the weight of rainfall on the thunderstorm event, and the probability of power outage caused by lightning density is multiplied by the weight of lightning density. The sum of these multiplications yields the power outage probability of the thunderstorm event. The power outage probability model for thunderstorm events is expressed by the following formula:
[0121]
[0122] In the formula, P outage,4 α4 represents the probability of power outage during a thunderstorm event; β4 and α4 are the weighting coefficients of lightning density and rainfall, respectively, in relation to the thunderstorm event. l Lightning density, expressed in times per km²·h. k λ This is the normalization factor for lightning density, in km². 2 h / time; R Rainfall amount, in mm / h. k R This is the normalization coefficient for rainfall, expressed in h / mm; Let be an exponentially decaying function, describing the cumulative effect of lightning-induced line tripping. l As we approach infinity, 1- →1, meaning that the probability of power outage approaches 100% when lightning strikes are frequent. l When =0, the risk of lightning is 0, tanh( k R R ) is a hyperbolic tangent function, used to simulate the nonlinear saturation effect of rainfall on power outage risk.
[0123] Step 3: Based on the base power loss of the power system, the base power loss of the gas system, the base power loss of the heating system, and the base power loss of the energy storage system caused by extreme events.
[0124] In a preferred but non-limiting embodiment of the present invention, step 3 includes:
[0125] Step 3.1: Construct a multi-energy disturbance propagation model encompassing the power system energy flow model, gas system energy flow model, thermal system energy flow model, and energy storage system energy flow model.
[0126] More preferably, step 3.1 includes:
[0127] Step 3.1.1: Construct a power system energy flow model based on the power system base power loss caused by extreme events, including:
[0128] Through voltage diffusion coefficient G e With voltage gradient Multiply to obtain the power flux density Solve for the divergence of the power flux density and take its negative value to obtain the power change driven by voltage diffusion. ;
[0129] The power contribution of gas input converted into electrical energy is obtained by multiplying the gas-to-electricity efficiency by the gas input power. ;
[0130] The heat-to-electricity conversion efficiency is multiplied by the heat input power to obtain the contribution of the heat input to the conversion of heat into electrical energy. ;
[0131] The discharge efficiency of energy storage converted into electrical energy is multiplied by the discharge power to obtain the output power contribution during energy storage discharge. ;
[0132] The charging efficiency of the electro-to-energy storage process is multiplied by the charging power to obtain the input power contribution during energy storage charging. ;
[0133] The power system vulnerability coefficient to extreme events is multiplied by the power system base power loss caused by the extreme event to obtain the power loss caused by the extreme event.
[0134] The required power replenishment for the power system is obtained by summing the power change driven by voltage diffusion with the power contributions from gas input to electrical energy, thermal input to electrical energy, and the output power contribution during energy storage discharge. This summation is then subtracted from the input power contribution during energy storage charging and the power losses caused by extreme events in the power system. The power system energy flow model is expressed by the following formula:
[0135]
[0136] In the formula, dQ e / dt The power required to supplement the power system is used to describe the energy stored in the power system. Q e Rate of change over time; G e is the voltage diffusion coefficient, used to describe the power diffusion driven by the voltage gradient; V Voltage, unit is V; For gradient operators; It is a divergence operator; P g This refers to the gas input power, measured in watts (W). or eg For gas-to-electricity efficiency, oreh For heat-to-electricity conversion efficiency; Q h The thermal input power is expressed in W. or sd The discharge efficiency of energy storage to electrical energy reflects the conversion loss of energy storage to electrical energy. or se The charging efficiency for converting electrical energy into energy storage reflects the conversion loss during the conversion process. This refers to the charging power, measured in watts (W), which is the energy input from the power grid for energy storage. This refers to the discharge power, measured in W, which is output from the energy storage system to the power grid. x e This represents the vulnerability coefficient of the power system to extreme events. This represents the base power loss of the power system caused by extreme events, measured in W.
[0137] Step 3.1.2: Construct a gas system energy flow model based on the base power loss of the gas system caused by extreme events, including:
[0138] The heat transfer power is obtained by multiplying the gas density, specific heat capacity at constant pressure, flow characteristic cross-sectional area, and characteristic temperature difference, and then dividing by the characteristic length of the gas flow.
[0139] The enthalpy flow density vector is obtained by multiplying the gas density, the enthalpy per unit mass of the gas, and the gas flow velocity. The spatial divergence of the enthalpy flow is calculated by using the gradient operator on the enthalpy flow density vector. The convective transport power is obtained by integrating the volume of the gas system control volume and multiplying it with the spatial divergence of the enthalpy flow.
[0140] The power-to-gas conversion efficiency is obtained by multiplying the power input by the power output.
[0141] The gas system extreme event vulnerability coefficient is multiplied by the gas system base power loss caused by the gas system extreme event to obtain the gas system extreme event loss;
[0142] The required additional power for the gas system is obtained by summing the heat transfer power and the power converted from electricity to gas, and then subtracting the convective transport power and losses due to extreme events in the gas system. The energy flow model of the gas system is expressed by the following formula:
[0143]
[0144] In the formula, dQ g / dt The gas system needs additional power. Q g ρ represents the total energy of the gas system. g This refers to the density of the gas, expressed in kg / m³. 3 C pg This refers to the specific heat capacity of the gas at constant pressure, in units of... ; A The characteristic cross-sectional area of the gas flow is expressed in meters (m²). 2 ; This refers to the gas temperature difference, expressed in Kelvin (K). L The characteristic length of the gas flow is given in meters (m); h is the enthalpy of the gas per unit mass (J / kg); v g V represents the gas flow velocity, measured in m / s. g The volume of the gas system control unit is in meters (m). 3 ; or ge Indicates the efficiency of electricity-to-gas conversion; P e For electrical input power; x g This represents the vulnerability coefficient of the gas system to extreme events. This represents the basic power loss of the gas system caused by extreme events, measured in watts (W).
[0145] Step 3.1.3: Construct a thermal system energy flow model based on the fundamental power loss of the thermal system caused by extreme events, including:
[0146] Integrating the volume of the gas system control volume, and multiplying the integral result by the Laplace operator of thermal conductivity and temperature field, we obtain the heat diffusion power caused by temperature unevenness inside the gas system.
[0147] Multiplying the electro-thermal conversion efficiency by the electrical input power yields the power that directly converts electrical energy into heat energy.
[0148] Multiplying the energy storage heat conversion efficiency by the energy storage input power yields the power released by the energy storage system to the heating network.
[0149] The difference between the system average temperature and the ambient temperature is multiplied by the thermal conductivity between the heating network and the environment to obtain the irreversible heat loss from the heating network to the environment through the equipment.
[0150] The power loss caused by extreme events in a thermal system is obtained by multiplying the vulnerability coefficient of the thermal system to extreme events by the base power loss of the thermal system.
[0151] The required heat power to be replenished by the gas system is obtained by summing the power of heat diffusion due to uneven temperature within the gas system, the power of direct conversion of electrical energy into heat energy, and the power released from the energy storage system to the heating network, and then subtracting the irreversible heat loss from the heating network to the environment through equipment and the power loss caused by extreme events in the thermal system. The energy flow model of the thermal system is expressed by the following formula:
[0152]
[0153] In the formula, dQ h / dt The heat power that needs to be supplemented to the thermal system, Q h The total energy of the thermal system; α h Thermal conductivity, in units of ;T h The system average temperature; The Laplace operator for the temperature field, with units of K / m. 2 V h The volume of the gas system control unit is in meters (m). 3 ; or he P represents the electro-thermal conversion efficiency. s This refers to the energy storage input power, measured in watts (W). or hs Indicates energy storage to heat conversion efficiency; β h T0 represents the thermal conductivity between the heating network and the environment, measured in W / K; T0 represents the ambient temperature, measured in K. x h This represents the vulnerability coefficient of a thermal system to extreme events. This represents the basic power loss of a thermal system caused by extreme events, measured in W.
[0154] Step 3.1.4: Construct an energy flow model for the energy storage system based on the base power loss caused by extreme events, including:
[0155] The charging efficiency of the electro-to-energy storage is multiplied by the charging power to obtain the input power contribution during energy storage charging.
[0156] The discharge efficiency of energy storage to electrical energy is multiplied by the discharge power to obtain the output power contribution during energy storage discharge.
[0157] The energy storage self-loss coefficient is multiplied by the total energy of the energy storage system to obtain the self-loss of the energy storage body;
[0158] The extreme event vulnerability coefficient of an energy storage system is multiplied by the base power loss of the energy storage system caused by an extreme event to obtain the power loss of the energy storage system caused by an extreme event.
[0159] The charging or discharging demand of a thermal energy storage system is obtained by subtracting the output power contribution during energy storage discharge, the self-loss of the energy storage unit, and the power loss caused by extreme events in the energy storage system from the input power contribution during energy storage charging. The energy flow model of the energy storage system is expressed by the following formula:
[0160]
[0161] In the formula, d Q s / dtQ is needed to meet the charging and discharging requirements of thermal energy storage systems. s γ represents the total energy of the energy storage system, expressed in J. s This is the energy storage self-loss coefficient, in seconds. -1 ξ s This represents the vulnerability coefficient of the energy storage system to extreme events. This represents the base power loss of the energy storage system caused by extreme events, measured in W.
[0162] Step 3.2: Solve the multi-energy disturbance propagation model of electricity-gas-heat-storage to obtain the power replenishment required by the power system, the power replenishment required by the gas system, the heat replenishment required by the thermal system, and the charging and discharging demand of the thermal system and energy storage system. Generate the regulation demand matrix based on the solution results, expressed by the following formula:
[0163]
[0164] In the formula, This represents the demand adjustment matrix.
[0165] Step 4: Combine the global system state and the optimal action strategies of each agent under the global system state obtained in Step 1 with the power outage probability model obtained in Step 2 and the adjustment demand matrix obtained in Step 3, and construct a multi-agent cooperative strategy model using MADDPG.
[0166] In a preferred but non-limiting embodiment of the present invention, step 4 includes:
[0167] Step 4.1: Based on the global system state and the optimal action strategy of each agent under the global system state obtained in Step 1, obtain the basic state space and basic action space of the multi-agent, as shown in Table 2. Combine the power outage probability model of extreme events to solve the state space and action space of extreme events.
[0168] More preferably, step 4.1 includes:
[0169] Step 4.1.1: Based on the global system state and the optimal action policy of each agent under the global system state obtained in Step 1, set the basic state space and basic action space of the multi-agent system. The global system state includes the basic state space of each multi-agent system, and the optimal action policy of each agent under the global system state includes the basic action space of each agent, as shown in Table 2.
[0170] Table 2 Basic State-Movement Space
[0171]
[0172] Step 4.1.2: Connect the basic state space of the multi-agent system with the blackout probability model, the rate of change of the blackout probability over time, and the spatial rate of change of the blackout probability obtained in Step 2, and solve for the extreme event state space, expressed by the following formula:
[0173]
[0174] In the formula, This represents the state space for extreme events, used to describe the state space after adding the characteristic of extreme event regulation requirements. This represents the basic state space of multiple agents. It contains the basic state space of all intelligent agents. The probability of a power outage at the current moment is derived from a power outage probability model based on extreme events. Real-time calculation, This represents the rate of change of the probability of a power outage over time, and is the time derivative of the probability of a power outage. This represents the spatial rate of change of the probability of power outages, used to describe high-risk areas.
[0175] Step 4.1.3, use the power outage probability model to determine the action. a i Perform differentiation to solve for the action to be performed. a i The rate of change of the probability of a subsequent power outage will trigger an action. a i Rate of change of probability of subsequent power outage and execution action a i The load amount removed afterward is multiplied by the corresponding weighting coefficient and summed to obtain the load loss risk quantification score of the action. C(P outage , a i ) It can be expressed by the following formula:
[0176]
[0177] In the formula, α represents the weighting coefficient, which is between 0 and 1 and is used to adjust the weight ratio between load reduction and risk-sensitive items. LoadShed( a i ) to perform the action a i The load after resection; The marginal impact of the action on the probability of power outage is calculated by differentiating the power outage probability model for extreme events.
[0178] Step 4.1.4: Set the load loss risk quantification score of the action corresponding to the action located in the basic action space to be less than a set threshold to obtain the extreme event action space, and realize risk adaptive action selection, expressed by the following formula:
[0179]
[0180] In the formula, Represents the action space for extreme events. Represents the basic action space, γ th The threshold represents the upper limit of the constraint function Γ, which determines whether an action is allowed.
[0181] Step 4.2: Solve for the single reward value based on the cost-effectiveness, global voltage security status and regulation efficiency calculation results, and then construct the risk-sensitive reward function by combining the power outage probability model of each extreme event obtained in Step 2 and the regulation demand matrix obtained in Step 3.
[0182] More preferably, step 4.2 includes:
[0183] Step 4.2.1: Divide the change in current operating cost relative to the benchmark value by the benchmark operating cost to obtain the cost-benefit calculation result;
[0184] By subtracting the rated voltage from the actual voltage value at each monitoring point and taking the absolute value of the difference, the absolute deviation of the actual voltage from the standard value is obtained. Dividing this by the voltage safety threshold yields the voltage risk level. Subtracting the voltage risk level from 1 yields the positive evaluation index of the voltage safety status. The positive evaluation indexes of the voltage safety status of all monitoring points are summed to obtain the comprehensive calculation result of the global voltage safety status.
[0185] The regulation efficiency is calculated by comparing the actual response of the adjustable load with the maximum adjustable capacity.
[0186] The calculation results of cost-effectiveness, global voltage security status, and regulation efficiency are multiplied by their respective weighting coefficients to obtain a single reward value, thus unifying the three objectives of economy, safety, and regulation efficiency into a single reward value. R t It can be expressed by the following formula:
[0187]
[0188] In the formula, R t Represents a single reward value, Δ C op Δ represents the change in current operating costs relative to the baseline value. C op When the value is less than 0, costs decrease and rewards increase; conversely, they increase and penalties are imposed. Cbase Baseline operating costs; V i For the first i The actual voltage value at each monitoring point
[0189] V nom Rated voltage, V th For voltage safety threshold, Q DR This represents the actual response of the adjustable load. Q max This represents the maximum adjustable capability.
[0190] Step 4.2.2: Multiply the risk sensitivity coefficient by the power outage probability model to obtain the risk penalty term. Take the negative value of the risk penalty term and use the exponential function as the base to obtain the exponential decay function, which is used to convert the power outage probability model into a decay factor in the interval [0, 1]. Then, perform a subtraction operation to shift the exponential decay function to the interval [-1, 0] to obtain the linear penalty value for the power outage.
[0191] The absolute error between the actual power adjustment amount generated by the agent's actions and the adjustment demand matrix is calculated by norm. The result is divided by the sum of the adjustment demand matrix and the parameter to prevent division by zero, and the relative error rate is obtained. 1 minus the relative error rate equals the generated adjustment accuracy score.
[0192] The linear penalty value for power outages and the adjustment accuracy score are multiplied by their respective weights, and the sum of these multiplications is then added to the single reward value to obtain the risk-sensitive reward. The risk-sensitive reward function is expressed by the following formula:
[0193]
[0194] In the formula, Indicates risk-sensitive rewards; Δ P x,act The actual power adjustment amount generated for the actions of multiple agents. β The risk sensitivity coefficient increases with the penalty index when the probability of a power outage is high; w 1~ w 5 is the weighting coefficient, which must satisfy... w 1+ w 2+ w 3+ w 4+ w 5=1, reflecting the priority of different objectives; e To prevent division by zero, .
[0195] Step 4.3: Based on the state space and action space of the extreme event obtained in Step 4.1 and the risk-sensitive reward function obtained in Step 4.2, a multi-agent cooperative optimization model is constructed using a multi-agent deep reinforcement learning algorithm to obtain a comprehensive energy multi-agent cooperative strategy under extreme events, thereby realizing the comprehensive energy multi-agent cooperative optimization scheduling under extreme events.
[0196] More preferably, step 4.3 includes:
[0197] Step 4.3.1: Construct the Actor network for each agent (PA / CA / SA / ALA) based on the extreme event state space. Construct the Critic network for each agent based on the optimal action strategy of each agent under the extreme event action space, global power grid state, and global system state. Initialize the multi-agent Actor-Critic network of the multi-agent collaborative optimization model.
[0198] More preferably, step 4.3.1 includes:
[0199] Construct an Actor network for agent i, and set the action policy function as follows: m i( ; Input the extreme event state space of the i-th agent. , This represents the weight parameters of the i-th agent in the Actor network, and the output action. a i .
[0200] Construct a Critic network for agent i, and set the action value evaluation function as follows: Q i ( , a 1,... a N ; The input is a global extreme event state space containing all agents. and joint actions a 1,... a N Among them, joint actions a 1,... a N This represents the set of actions of all agents, conforming to the extreme event action space. , For the i-th agent, the weight parameters of the Critic network are used to output the action value. Q .
[0201] Step 4.3.2: Based on the sum of the actions and exploration noise output by each agent in Step 4.3.1, select the actions of each agent from the extreme event action space. The actions of each agent that satisfy the summation result of the extreme event action space are expressed by the following formula:
[0202]
[0203] in, N represents the action of the agent. t To explore noise, Clip() stands for Clip function, which is used to ensure that actions satisfy feasible region constraints.
[0204] Step 4.3.3: Construct agent joint actions based on the agent actions obtained in Step 4.3.2, and execute the agent joint actions. a ={ a 1,multi ,... a N,multi Obtain risk-sensitive rewards for multi-agent interactions. and the new global multi-agent state Transfer tuples ( , a , , Store it in the experience pool.
[0205] Step 4.3.4: Obtain the transition tuples from the experience pool in Step 4.3.3 to perform centralized network training of the multi-agent cooperative strategy model, and solve to obtain the preliminary multi-agent cooperative strategy model.
[0206] More preferably, step 4.3.4 includes:
[0207] Sample batch data from the experience pool and update it using the following steps:
[0208] For each intelligent agent i The Critic network is updated by minimizing the temporal difference error, as expressed by the following formula:
[0209]
[0210] In the formula, This represents the loss function of the Critic network. Expressing expectations, Q i This represents the current Q-value of the Critic network's evaluation of the joint action. y i The target Q-value for time-difference is expressed by the following formula:
[0211]
[0212] In the formula, r i For immediate rewards, risk-sensitive rewards will be implemented. Set as instant reward Real-time feedback on action effects Q i ’ This represents the Q-value of the target network at the next time step. This indicates the output state of agent N at the next time step. The predicted action vector is below. , This is the discount factor.
[0213] The Actor network is updated using the policy gradient ascent method to maximize the expected return, expressed by the following formula:
[0214]
[0215] In the formula, Indicates adjustment The direction of change of expected return J This indicates that the Q-value of the Critic network output corresponds to the action. The gradient; Indicates the output of the Actor network. For parameters The gradient.
[0216] The target network parameters are updated using the following formula:
[0217]
[0218] In the formula, Indicates the first i The weight parameters of the Critic target network for each agent at the next time step. Indicates the first i The weight parameters of the target network for each agent (Actor) at the next time step. This is the soft update coefficient.
[0219] The updated target network i The weight parameters of the Critic target network for each agent at the next time step are minimized using gradient descent to reduce the temporal difference error. Generate new actions a i '= m i( ; ), through policy gradient Maximize cumulative rewards;
[0220] Store after each update and If in N consecutive iterations, satisfy , The L2 norm satisfies If the network converges, proceed to step 4.3.5; otherwise, return to step 4.3.2 and regenerate the action based on the optimized network.
[0221] Step 4.3.5: Based on the action policies of the multi-agents output by the preliminary multi-agent cooperative policy model obtained in Step 4.3.4, perform Nash equilibrium to solve for the multi-agent cooperative policy, and obtain the multi-agent cooperative policy model, expressed by the following formula:
[0222]
[0223] in, Indicates the first i Update strategy for each agent Represents intelligent agents i Expected cumulative reward Indicates the weight of the strength of strategy collaboration. Indicates the first i The policy function of each agent m i( ; ), used to describe the optimal action policy output after training in step 4.3.4, can be regarded as the optimal action of the i-th agent after training, and KL() is the KL divergence, used to measure the agent's performance. i strategy π i Average strategy π avg The differences between them π avg =∑ π j / N represents the average strategy, which promotes strategy consensus.
[0224] Step 5: Input the real-time monitoring data of the integrated energy multi-agent into the multi-agent collaborative strategy model constructed in Step 4 to obtain the integrated energy multi-agent collaborative optimization scheduling strategy under extreme events, and realize the integrated energy multi-agent collaborative optimization scheduling under extreme events.
[0225] Embodiment 2 of the present invention provides a comprehensive energy multi-agent cooperative optimization scheduling system under extreme events, which runs the comprehensive energy multi-agent cooperative optimization scheduling method under extreme events described in Embodiment 1, including:
[0226] The multi-agent state extraction module is used to extract the spatiotemporal characteristics of multiple agents in the integrated energy system through mutual information and fast Fourier transform. After knowledge distillation and lightweighting, a lightweight general model of agents is constructed and solved to obtain the global system state and the optimal action strategy of each agent under the global system state.
[0227] The power outage probability solution module is used to build power outage probability models for various extreme events, including typhoon events, extreme cold events, high temperature events, and thunderstorm events.
[0228] The demand solving module is used to construct an extreme event-driven multi-energy disturbance propagation model of electricity-gas-heat-storage based on the base power loss of the power system, gas system, thermal system, and energy storage system caused by extreme events, solve the model, and generate the regulation demand matrix.
[0229] The collaborative strategy model construction module is used to combine the global system state and the optimal action strategies of each agent under the global system state with the power outage probability model and adjustment demand matrix of each extreme event, and to construct a multi-agent collaborative strategy model using multi-agent deep reinforcement learning.
[0230] The output module is used to input real-time monitoring data of the integrated energy multi-agent into the multi-agent collaborative strategy model to obtain the integrated energy multi-agent collaborative optimization scheduling strategy under extreme events.
[0231] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0232] This invention uses knowledge distillation technology, supports edge deployment, and ensures that the integrated energy system can still operate independently when communication is interrupted due to extreme events, thereby improving the security of integrated energy system scheduling. It adopts the Nash equilibrium method of distributed decision architecture to reward and punish each agent, forcing agents to consider the actions of other agents when pursuing their own benefits, reducing multi-agent policy conflicts, and improving the efficiency and accuracy of multi-agent collaborative optimization of the integrated energy system.
[0233] This invention enhances the basic state space of multiple agents by constructing a power outage probability model for extreme events, thereby improving the agents' ability to perceive risks under extreme events. By constraining the basic action space of multiple agents through the power outage probability model for extreme events, the integrated energy system restricts the agents' action choices to extreme events, thereby constraining high-risk actions under extreme events and improving the integrated energy system's adaptive ability and security under extreme events.
[0234] This invention generates scheduling requirements by constructing a multi-energy disturbance propagation model of electricity-gas-heat-storage, and combines the scheduling requirements with a power outage probability model of extreme events to generate a risk-sensitive reward function for extreme events, thereby reducing the load loss of the integrated energy system during extreme events and improving the economic efficiency of integrated energy system regulation.
[0235] This invention reduces the conflict rate of a comprehensive energy system under multi-agent strategy regulation by balancing the strategies of multiple agents through Nash equilibrium, thereby improving the security of the comprehensive energy system under multi-agent strategy regulation.
[0236] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A comprehensive energy multi-agent cooperative optimization scheduling method under extreme events, characterized in that, Includes the following steps: The spatiotemporal characteristics of multiple agents in the integrated energy system are extracted by mutual information and fast Fourier transform. After knowledge distillation and lightweighting, a lightweight general model of agents is constructed and solved to obtain the global system state and the optimal action strategy of each agent under the global system state. Construct power outage probability models for various extreme events, including typhoon events, extreme cold events, high temperature events, and thunderstorm events; Based on the base power loss of the power system, gas system, heat system, and energy storage system caused by extreme events, a multi-energy disturbance propagation model driven by extreme events is constructed, and the model is solved to generate a regulation demand matrix. By combining the global system state and the optimal action strategies of each agent under the global system state with the power outage probability model and adjustment demand matrix of each extreme event, a multi-agent cooperative strategy model is constructed using multi-agent deep reinforcement learning. By inputting real-time monitoring data of integrated energy multi-agents into the multi-agent collaborative strategy model, a collaborative optimization scheduling strategy for integrated energy multi-agents under extreme events is obtained.
2. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 1, characterized in that: The construction and solution of the lightweight general model of the intelligent agent includes: The energy physical equipment of the integrated energy system is divided into multiple types of intelligent agents according to its functional attributes, resulting in a multi-type intelligent agent cluster of the integrated energy system. Data from a cluster of multiple intelligent agents in an integrated energy system is acquired, and spatiotemporal features are extracted using mutual information and fast Fourier transform to obtain spatiotemporal features. The spatiotemporal features are transferred from the teacher model to the student model through knowledge distillation to obtain lightweight student model data. The lightweight student model data is set as the global system state. A lightweight general model of agents is constructed by combining multi-agent Nash equilibrium. The optimal action strategy of each agent in the global system state is obtained by solving the lightweight general model of agents.
3. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 2, characterized in that: The solution to the lightweight general model for agents yields the optimal action strategies for each agent under the global system state, including: Construct a reward function for each agent based on its own actions and the actions of other agents; Set the lightweight student model data as the global system state, input the security constraint function, and if the security constraint function is violated, set the security constraint relaxation term to the output of the security constraint function; otherwise, set the security constraint relaxation term to 0. The optimal action strategy of the current agent is obtained by subtracting the product of the balance coefficient and the safety constraint relaxation term from the reward function of each agent. The optimal action strategies of each agent are summed to obtain the optimal action strategies of each agent in the global system state.
4. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 1, characterized in that: The proposed power outage probability models for each extreme event include: The probability of power outages caused by precipitation, the cumulative risk probability corresponding to the degree of wind speed deviation, and the joint risk degree of precipitation and wind speed are solved by wind speed and precipitation amount. Based on the solution results, a power outage probability model for typhoon events is constructed. The probability of power outage caused by extreme cold temperature and the probability of power outage caused by the duration of extreme cold event are solved by using temperature and duration, and a power outage probability model for extreme cold event is constructed based on the solution results. The probability of power outages caused by high temperatures is calculated by temperature, and a power outage probability model for high-temperature events is constructed by combining it with air conditioning load rate. The probability of power outages caused by lightning density and the nonlinear saturation risk of rainfall are solved by using lightning density and rainfall amount, and a power outage probability model for thunderstorm events is constructed based on the solution results.
5. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 1, characterized in that: The generated adjustment demand matrix includes: include: Based on the power system base power loss caused by extreme events, construct the power system energy flow model; based on the gas system base power loss caused by extreme events, construct the gas system energy flow model; based on the thermal system base power loss caused by extreme events, construct the thermal system energy flow model; based on the energy storage system base power loss caused by extreme events, construct the energy storage system energy flow model. A multi-energy disturbance propagation model for electricity-gas-heat-storage is constructed based on the energy flow models of power systems, gas systems, heat systems, and energy storage systems. Solve the multi-energy disturbance propagation model of electricity-gas-heat-storage to generate the power replenishment required by the power system, the power replenishment required by the gas system, the heat replenishment required by the thermal system, and the charging and discharging demand of the thermal system and the energy storage system. Construct a regulation demand matrix based on the generated results.
6. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 1, characterized in that: The construction of the multi-agent cooperative strategy model includes: The basic state space and basic action space of the multi-agent are obtained based on the global system state and the optimal action strategy of each agent under the global system state. The extreme event state space and extreme event action space are solved by combining the power outage probability model of extreme events. The single reward value is calculated based on the cost-effectiveness, global voltage security status and regulation efficiency calculation results. Then, a risk-sensitive reward function is constructed by combining the power outage probability model and regulation demand matrix of each extreme event. Based on the state space and action space of extreme events, as well as the risk-sensitive reward function, a multi-agent collaborative optimization model is constructed using a multi-agent deep reinforcement learning algorithm.
7. The integrated energy multi-agent cooperative optimization scheduling method under extreme events as described in claim 6, characterized in that: The extreme event action space includes: The action is differentiated by the power outage probability model, and the rate of change of the power outage probability after the action is executed is solved. The rate of change of the power outage probability after the action and the amount of load cut off after the action are multiplied by the corresponding weight coefficients and summed to obtain the load loss risk quantification score of the action. If the load loss risk quantification score of the action corresponding to the action located in the basic action space is less than a set threshold, the extreme event action space is obtained.
8. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 6, characterized in that: The construction of the risk-sensitive reward function includes: Multiply the risk sensitivity coefficient by the power outage probability model to obtain the risk penalty term. Take the negative value of the risk penalty term and use the exponential function as the base to obtain the exponential decay function. Then, subtract 1 to obtain the linear penalty value for the power outage. The absolute error between the actual power adjustment amount generated by the agent's actions and the adjustment demand matrix is calculated. This error is divided by the sum of the adjustment demand matrix and the parameter to prevent division by zero, to obtain the relative error rate. Subtracting the relative error rate from 1 yields the generation adjustment accuracy score. The risk-sensitive reward function is obtained by multiplying the linear penalty value for power outages and the adjustment accuracy score by their respective weights, summing the results, and then adding them to a single reward value.
9. The integrated energy multi-agent cooperative optimization scheduling method under extreme events according to claim 6, characterized in that: The construction of the multi-agent cooperative optimization model includes: Construct the Actor network for each agent based on the state space of extreme events, and construct the Critic network for each agent based on the action space and state space of extreme events. Summing the actions output by each agent from the Actor network with the exploration noise, we obtain the actions of each agent that satisfy the extreme event action space after the summation result. Construct joint action of agents based on the actions of each agent, execute joint action of agents, obtain risk-sensitive reward of multi-agent and new multi-agent state, and construct transition tuple with extreme event state space and joint action of agents, and store it in experience pool. The transition tuples are obtained from the experience pool and a multi-agent cooperative policy model is trained. The preliminary multi-agent cooperative policy model is then obtained by solving the problem. The action policies of the multi-agents output by the preliminary multi-agent cooperative strategy model are subjected to Nash equilibrium to solve the multi-agent cooperative strategy and construct the multi-agent cooperative strategy model.
10. A multi-agent cooperative optimization scheduling system for integrated energy under extreme events, comprising the multi-agent cooperative optimization scheduling method for integrated energy under extreme events as described in any one of claims 1-9, characterized in that: The multi-agent state extraction module is used to extract the spatiotemporal characteristics of multiple agents in the integrated energy system through mutual information and fast Fourier transform. After knowledge distillation and lightweighting, a lightweight general model of agents is constructed and solved to obtain the global system state and the optimal action strategy of each agent under the global system state. The power outage probability solution module is used to build power outage probability models for various extreme events, including typhoon events, extreme cold events, high temperature events, and thunderstorm events. The demand solving module is used to construct an extreme event-driven multi-energy disturbance propagation model of electricity-gas-heat-storage based on the base power loss of the power system, gas system, thermal system, and energy storage system caused by extreme events, solve the model, and generate the regulation demand matrix. The collaborative strategy model construction module is used to combine the global system state and the optimal action strategies of each agent under the global system state with the power outage probability model and adjustment demand matrix of each extreme event, and to construct a multi-agent collaborative strategy model using multi-agent deep reinforcement learning. The output module is used to input real-time monitoring data of the integrated energy multi-agent into the multi-agent collaborative strategy model to obtain the integrated energy multi-agent collaborative optimization scheduling strategy under extreme events.
Citation Information
Patent Citations
Park-level adjustable load-oriented resource regulation and control method, device and equipment and medium
CN119539424A
Emergency response method for extreme event of complex energy interconnection system
CN111276976A
Power failure processing method, device and equipment for electric power guarantee area and medium
CN118504991A