A power distribution network power router optimization scheduling method for extreme weather scenarios

CN122553150APending Publication Date: 2026-08-11FOSHAN GUYUXUAN BRAND MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

求解技术方面,传统确定性优化方法依赖精确的负荷预测模型,而寒潮属于低概率高影响事件,其到达时间、持续时长和强度均存在较大不确定性,传统方法在此类极端场景下的调度效果往往不够理想

Benefits of technology

本发明通过构建包含多个台区的互联拓扑架构并为每个台区配置电能路由器和分布式储能系统,利用天气状况指标、负荷需求、分布式光伏出力、储能荷电状态及时间特征构建马尔可夫决策过程的状态空间,并以储能系统充放电功率和互联端口功率交换量为动作空间,结合包含网损惩罚、电压偏移惩罚及功率越限惩罚的奖励函数,在内部功率平衡方程和充放电功率条件的约束下对智能体进行训练,使得训练后的各台区智能体能够仅基于局部观测状态独立输出优化的调度决策;该方法通过电能路由器和分布式储能系统实现了不同负荷特性台区间的功率互济,引入天气状况指标使智能体能够直接感知极端气象强度与演变趋势,通过奖励函数引导智能体在平抑负荷尖峰的同时优先满足电压安全约束,并在满足物理约束的前提下有效缓解变压器过载及电压偏移问题,实现了极端气象期间配电网的负荷平抑和电压稳定,提升了极端气象场景下配电网的运行韧性与调度决策质量,保障了极端气象下配电系统的安全可靠运行。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122553150A_ABST
    Figure CN122553150A_ABST
Patent Text Reader

Abstract

This invention relates to an optimized scheduling method for power routers in distribution networks under extreme weather conditions. The method includes: constructing an interconnected topology architecture comprising multiple transformer substations with heterogeneous load characteristics, each substation equipped with a power router and a distributed energy storage system, and each substation corresponding to an intelligent agent; defining the load demand and distributed photovoltaic output of each substation, and defining weather condition indicators characterizing extreme weather; establishing internal power balance equations and interconnected power balance equations for each substation; establishing charging and discharging power conditions and state of charge update equations for the distributed energy storage system; constructing the scheduling behavior of the power router and the distributed energy storage system as a Markov decision process; training the intelligent agents to obtain trained agents for each substation; and using the trained intelligent agents for each substation to output scheduling actions based on their respective local observation states. This invention achieves load smoothing and voltage stability in the distribution network during extreme weather events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power engineering, specifically relating to an optimized scheduling method for power distribution network routers in extreme weather scenarios. Background Technology

[0002] In the process of comprehensive building electrification, electric heating equipment such as heat pumps have gradually replaced traditional coal-fired boilers, and the load structure and peak characteristics of building power distribution systems have also changed accordingly. Under extreme cold waves, the heating and power load rises sharply, impacting the carrying capacity of regional building power grids. Problems such as transformer overload and voltage deviation occur frequently, and the load peak impact caused by cold waves poses a direct threat to the resilience of building power distribution systems.

[0003] Existing power grid dispatching methods have several shortcomings in dealing with cold wave scenarios. Regarding load modeling, most studies focus on macro-level overall load analysis or the load characteristics of single-type users, paying less attention to the heterogeneity of building load characteristics in different areas during cold waves. Residential, industrial, and commercial building areas exhibit different load variation patterns under the impact of cold waves. This spatiotemporal heterogeneity contains the potential for power exchange, but existing methods mostly treat it as equal. In terms of equipment utilization, power routers, as flexible interconnection components in smart grids, possess flexible power flow management and voltage regulation capabilities. Existing work has verified their role in alleviating energy congestion and optimizing energy distribution, but research on using power routers to connect building areas with different characteristics and explore load heterogeneity to achieve cross-regional power exchange remains limited. Regarding solution techniques, traditional deterministic optimization methods rely on accurate load forecasting models, while cold waves are low-probability, high-impact events with significant uncertainties in their arrival time, duration, and intensity. Traditional methods often fail to achieve ideal dispatching results in such extreme scenarios. Deep reinforcement learning has certain advantages in dealing with uncertainty, but existing mainstream solutions are constrained by the difficulty of credit allocation when dealing with long-term window scheduling decisions. The long-term impact of early scheduling actions on the later system state is difficult to be accurately attributed, and this problem is particularly evident in long-term chain decision-making in power systems.

[0004] Therefore, how to achieve load stabilization and voltage stability of the power distribution network during extreme weather events, and ensure the safe and reliable operation of the power distribution system under extreme weather conditions, has become an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned problems in existing technologies, this invention provides a power router optimization scheduling method for distribution networks in extreme weather scenarios. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a method for optimizing the scheduling of power routers in distribution networks for extreme weather scenarios, comprising the following steps: Construct an interconnected topology architecture that includes multiple transformer substations with heterogeneous load characteristics. Each substation is equipped with an energy router and a distributed energy storage system, and each substation corresponds to an intelligent agent. The load demand and distributed photovoltaic output of each transformer area are defined based on historical operating data, and weather condition indicators characterizing extreme weather are defined based on meteorological data of the corresponding time period. Establish the internal power balance equations and interconnection power balance equations for each transformer area; Establish the charging and discharging power conditions and state of charge update equations for the distributed energy storage system; The scheduling behavior of the power router and the distributed energy storage system is constructed as a Markov decision process. The state space of the Markov decision process includes weather condition indicators, load demand, distributed photovoltaic output, energy storage state of charge and time characteristics, the action space includes the charging and discharging power of the energy storage system and the power exchange of the interconnection ports of the power router, and the reward function includes network loss penalty, voltage deviation penalty and power limit violation penalty established based on the interconnection power balance equation. The agent is trained based on the Markov decision process, constrained by the internal power balance equation and the charging and discharging power conditions, to obtain the trained agent for each power station area; the trained agent for each power station area is used to output scheduling actions based on its local observation state.

[0006] In one embodiment of the present invention, the interconnection topology includes multiple transformer substations with heterogeneous load characteristics, a central multi-port converter, and a modular multilevel converter. Each transformer substation is equipped with a power router, a distributed energy storage system, and a distributed photovoltaic system. The DC terminal of each of the power routers is connected to the DC terminal of the modular multilevel converter via a central multiport converter; the AC terminal of the modular multilevel converter is connected to the external AC power grid via a transformer and the system central bus. In each distribution area, the terminal power load is connected to the AC terminal of the power router, and the distributed energy storage system and the distributed photovoltaic system are connected to the DC terminal of the power router. Each distribution area exchanges energy with the external power distribution system through a transformer.

[0007] In one embodiment of the present invention, the load demand and distributed photovoltaic output of each transformer area are defined based on historical operating data, and weather condition indicators characterizing extreme weather are defined based on meteorological data for the corresponding time period, including: Define the first based on historical operational data. Taiwan area at all times The load demand is , define the first Taiwan area at all times The photovoltaic output is ; Defining time based on meteorological data for the corresponding time period The weather condition index is ; in, , , The number of interconnected stations, The number of time periods divided into the scheduling cycle.

[0008] In one embodiment of the present invention, establishing the internal power balance equations and interconnection power balance equations for each transformer area includes: Based on the principle of energy conservation, the DC bus power balance equations for each transformer substation are established. By coupling the AC-side exchange power, DC-side power, and internal conversion power, the internal power balance equations for each transformer substation are obtained:

[0009] in, For the first Taiwan area at all times Photovoltaic power output; For the first The distribution area energy storage system is always The discharge power, For the first The energy storage system in the distribution area is always The charging power; For the first The DC-DC port of the transformer power router is at the time The power flowing from the DC bus to the central multiport converter DC-DC module, For the first The DC-DC port of the transformer power router is at the time Power flowing into the DC bus from the central multi-port converter DC-DC module; This refers to the power consumption of the power router interacting with the AC side. For power electronic conversion efficiency, For the discharge efficiency of the energy storage system The charging efficiency of the energy storage system; Based on the energy interaction logic between the DC interconnection system and the external AC grid in the aforementioned interconnection topology, a central DC-DC power balance equation is established, resulting in the interconnection power balance equation:

[0010] in, For the first The DC-DC port of the transformer power router is at the time Power exchange quantity For modular multilevel converters at time Transmission power, To improve the transmission efficiency of modular multilevel converters, This refers to the number of interconnected stations.

[0011] In one embodiment of the present invention, the charging and discharging power conditions of the distributed energy storage system are as follows:

[0012] in, For the first The energy storage system in the distribution area is always The charging and discharging power, For the first The maximum charging and discharging power of the energy storage system in the distribution area; The state-of-charge update equation for the distributed energy storage system is:

[0013] in, For the first The energy storage system in the distribution area is always The state of charge; For the first The energy storage system in the distribution area is always The state of charge of , with values ​​ranging from [0,1]; For the first The rated capacity of the energy storage system in the distribution area The step size for each scheduling period.

[0014] In one embodiment of the present invention, the Markov decision process includes a state space, an action space, state transition probabilities, a reward function, and a discount factor; The state space is formed by the local observation states of each station area:

[0015] in, For the global state space, For the first The intelligent agent in the district is at all times The local observation status, , , The number of interconnected stations, As a weather condition indicator, For all within the distribution network Each interconnected platform is at any time The set of load demands For all within the distribution network Each interconnected platform is at any time Photovoltaic power generation collection, For the first The energy storage system in the distribution area is always The state of charge, This represents the normalized time characteristics of the current scheduling period; The action space is formed by the joint action space of the intelligent agents in each station area:

[0016] in, For the joint action space of intelligent agents in each station area, For the first The intelligent agent in the district is at all times The action vector, , , , For the first The energy storage system in the distribution area is always The charging and discharging power, For the first The maximum charging and discharging power of the energy storage system in the distribution area, For the first The DC-DC port of the transformer power router is at the time Power exchange quantity This represents the maximum transmission power of the DC-DC port. The reward function is:

[0017] in, For a moment Instant reward value, Basic rewards, For network damage penalties, As a penalty for voltage offset, As a penalty for system crash, Penalty for exceeding power limits; , For the system at time Total network loss per unit value, This is the network loss penalty coefficient; , For the system time The minimum node voltage per unit value, For the system time Maximum node voltage per unit value, This is the per-unit value of the lower voltage threshold. This is the per-unit value of the upper voltage threshold. This is the voltage offset penalty coefficient; , For the first The transformer in the distribution area is at the time The per-unit value of the transmission power, For the first The rated power per unit value of the transformer in the distribution area. For modular multilevel converters at time The per-unit value of transmission power, This refers to the per-unit rated power of the modular multilevel converter. This is the power over-limit penalty factor; when Less than 0.90 pu or System crash penalty when PU exceeds 1.10. Set as a constant ; State transition probability Discount factor .

[0018] In one embodiment of the present invention, the intelligent agent is implemented based on a multi-agent soft actor-critic algorithm; The multi-agent soft actor-critic algorithm includes an actor network and a critic network, both of which are constructed based on a Transformer encoder.

[0019] In one embodiment of the present invention, the actor network includes: The historical data collection input module is used to construct a historical state sequence from the local observation states of each transformer area agent; wherein, the local observation states include weather condition indicators, load demand, distributed photovoltaic power output, energy storage state of charge, and normalized time characteristics; The auxiliary feature input module is used to concatenate the current state of charge with the action at the previous time step, and project the concatenated auxiliary input vector onto a unified feature space through a linear mapping to obtain the auxiliary feature projection. An observation state encoding module is used to encode the historical state sequence to obtain an encoded sequence. A linear projection layer is used to perform a linear projection on the encoded sequence to obtain a projection result; The overlay module is used to overlay the projection result of the last time step onto the auxiliary feature projection to form the input sequence; The first position encoding module is used to obtain the final input sequence of the first encoder after adding position encoding to the input sequence; The first Transformer encoder is used to extract long-range dependency features of the final input sequence of the first encoder using a multi-head self-attention mechanism to obtain the encoded first hidden state sequence. The first linear feature extraction module is used to extract the feature vector corresponding to the last position of the first hidden state sequence and obtain the shared features through linear mapping. The mean output header and standard deviation output header are used to generate the mean parameter and standard deviation parameter of the action distribution based on the shared features, respectively. The action distribution module is used to construct an action distribution based on the mean parameter and the standard deviation parameter; The action sampling module is used to sample continuous actions from the action distribution using reparameterization techniques. The sampling results are then subjected to hyperbolic tangent compression and scaled according to physical ratings based on the charging and discharging power conditions and the maximum transmission power of the interconnect port to obtain the action vector of the agent.

[0020] In one embodiment of the present invention, the critic network includes a first online critic network, a second online critic network, a first target critic network, and a second target critic network, wherein the first online critic network, the second online critic network, the first target critic network, and the second target critic network each include: The state and action merging module is used to concatenate the local observation state at each time step with the sampled action vector to form a state-action coupled vector. The second linear feature extraction module is used to map the state-action coupling vector to a unified dimensional space through linear projection to obtain a state-action feature sequence. The second position encoding module is used to add position encoding to the state action feature sequence to obtain the final input sequence of the second encoder; The second Transformer encoder is used to extract the long-range spatiotemporal features of the final input sequence of the second encoder using a multi-head self-attention mechanism to obtain the second hidden state sequence. The sequence post-processing module is used to extract the feature representation corresponding to the current time from the second hidden state sequence to obtain refined features; Q output layer, used to output state action values ​​based on the refined features; Specifically, the smaller value between the state action value output by the first online commentator network and the state action value output by the second online commentator network is taken as the final state action value of the online commentator network. The smaller value between the state-action value output by the first target critic network and the state-action value output by the second target critic network is taken as the target state-action value of the target critic network.

[0021] In one embodiment of the present invention, the process of training the intelligent agent includes: Reset the power distribution area environment; Obtain environmental information, including weather conditions, load demand, distributed photovoltaic output, energy storage status of charge, and normalized time characteristics; Extract the local observation state vector based on the environmental information; The actor network of each station area agent performs forward propagation based on the local observation state vector, outputs action distribution parameters and samples to obtain continuous actions; Based on the charging and discharging power conditions of the distributed energy storage system and the maximum transmission power of the power router interconnection port, the continuous action is subjected to hyperbolic tangent compression and scaled proportionally according to the physical rating of the equipment to obtain the action vector of the intelligent agent. The interconnection port between the distributed energy storage system and the power router executes the action vector; The actual power values ​​of each port are obtained through power flow calculation and substituted into the internal power balance equation for physical consistency verification. If the action vector causes power imbalance or violates the quasi-steady-state operation law of the power system, the agent is negatively corrected through the power limit penalty term in the reward function. Based on the reward function, determine whether the node voltage exceeds the safety threshold or whether the transformer is severely overloaded; if the operating status is safe, proceed to the cumulative operating reward branch; if an over-limit is triggered, proceed to the calculation failure penalty branch and terminate this round of exploration. Store the experience tuples formed at the current moment into the experience replay pool; The process involves sequentially updating the critic network, updating the actor network, automatically adjusting the temperature parameters, and performing a soft update to the target critic network.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention constructs an interconnected topology architecture comprising multiple distribution substations and configures a power router and a distributed energy storage system for each substation. It utilizes weather condition indicators, load demand, distributed photovoltaic output, energy storage state of charge, and temporal characteristics to construct the state space of a Markov decision process. The action space is defined by the charging and discharging power of the energy storage system and the power exchange at the interconnection ports. A reward function, incorporating network loss penalties, voltage deviation penalties, and power limit violation penalties, is used to train the agent under the constraints of internal power balance equations and charging / discharging power conditions. This allows the trained agent in each substation to independently output optimized scheduling decisions based solely on local observations. This method achieves power balance between substations with different load characteristics through power routers and distributed energy storage systems. The introduction of weather condition indicators enables the agent to directly perceive the intensity and evolution of extreme weather conditions. The reward function guides the agent to prioritize voltage safety constraints while smoothing load peaks, effectively mitigating transformer overload and voltage deviation problems while meeting physical constraints. This achieves load smoothing and voltage stability of the distribution network during extreme weather events, improving the operational resilience and scheduling decision quality of the distribution network under extreme weather conditions, and ensuring the safe and reliable operation of the distribution system under extreme weather conditions. Attached Figure Description

[0023] Figure 1 A flowchart illustrating an optimized scheduling method for power distribution network routers in extreme weather scenarios, provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of a multi-building distribution transformer area interconnection topology based on a power router, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the power conversion process of a power router provided in an embodiment of the present invention; Figure 4 This is an overall framework diagram of the intelligent agent implementation algorithm constructed according to an embodiment of the present invention; Figure 5 This is a flowchart illustrating the training process of an intelligent agent according to an embodiment of the present invention; Figure 6 A schematic diagram of load optimization allocation during normal periods; Figure 7 A schematic diagram illustrating the suppression of peak loads during cold waves; Figure 8 This is a schematic diagram illustrating the principle of how this invention performs well in long-range scheduling. Detailed Implementation

[0024] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0025] Example 1 Please see Figure 1 , Figure 1This is a flowchart illustrating a power router optimization scheduling method for distribution networks in extreme weather scenarios, provided by an embodiment of the present invention. The method includes the following steps: S1. Construct an interconnected topology architecture containing multiple transformer substations with heterogeneous load characteristics, wherein each transformer substation is equipped with an energy router and a distributed energy storage system, and each transformer substation corresponds to an intelligent agent; S2. Define the load demand and distributed photovoltaic output of each transformer area based on historical operating data, and define weather condition indicators characterizing extreme weather based on meteorological data of the corresponding time period; in this embodiment, extreme weather includes extreme cold waves; S3. Establish the internal power balance equations and interconnection power balance equations for each transformer area; S4. Establish the charging and discharging power conditions and state of charge update equations for the distributed energy storage system. S5. The scheduling behavior of the power router and the distributed energy storage system is constructed as a Markov decision process. The state space of the Markov decision process includes weather condition indicators, load demand, distributed photovoltaic output, energy storage state of charge and time characteristics. The action space includes the charging and discharging power of the energy storage system and the power exchange of the interconnection port of the power router. The reward function includes network loss penalty, voltage deviation penalty and power limit violation penalty established based on the interconnection power balance equation. S6. Using the internal power balance equation and charging / discharging power conditions as constraints, the agent is trained based on the Markov decision process to obtain the trained agent for each power station area; the trained agent for each power station area is used to output scheduling actions based on its own local observation state.

[0026] In one specific embodiment, step S1 includes: Please see Figure 2 , Figure 2This is a schematic diagram of a multi-building distribution transformer area interconnection topology based on an electric energy router, provided in an embodiment of the present invention. The interconnection topology includes multiple transformer areas with heterogeneous load characteristics, a central multi-port converter (MPC), and a modular multilevel converter (MMC). Each transformer area is equipped with an electric energy router (EER), an energy storage system (ESS), and a distributed photovoltaic system. The DC terminal of each electric energy router is connected to the DC terminal of the modular multilevel converter via the central MPC. The AC terminal of the modular multilevel converter is connected to the external AC power grid via a transformer and the system's central bus. In each transformer area, the terminal power loads are connected to the AC terminal of the electric energy router, while the distributed energy storage system and the distributed photovoltaic system are connected to the DC terminal of the electric energy router. Each transformer area exchanges energy with the external power distribution system.

[0027] Specifically, this interconnected topology uses the central multi-port converter (MPC) at the bottom as the core DC interconnection hub. The external AC grid 1 is connected to the system's central bus via transformer 2, while the overall energy balance and AC / DC conversion of the system are achieved by modular multilevel converters (MMCs), which are coupled to the topology center through branch node 15. This interconnected topology includes four typical building distribution areas with heterogeneous load characteristics, corresponding to loads 3, 4, 5, and 6 in the diagram. These areas cover different types of terminal power loads, including residential, commercial, and industrial loads. Each distribution area exchanges energy with the external power distribution system through transformers 7, 8, 9, and 10. The refined energy distribution within each distribution area is handled by four power routers (EER1, EER2, EER3, and EER4) deployed at the hub location. These power routers, as the core control units connecting the AC side of the distribution area, the internal DC side, and the inter-area interconnection channel, provide the physical basis for power balance among heterogeneous loads in each area.

[0028] Furthermore, this interconnection topology includes a total of This embodiment connects specific areas such as residential, industrial, and commercial building areas with heterogeneous load characteristics, each equipped with a power router and a distributed energy storage system. The distributed energy storage system is physically connected to the power router within the area via DC coupling. Each area corresponds to an independent intelligent agent. It should be noted that the power router in this embodiment can be a smart soft switch or other flexible interconnection device capable of flexible power distribution between multiple areas. Accordingly, the power conversion process and efficiency parameters need to be adjusted according to the specific device characteristics.

[0029] In this embodiment, the scheduling period is divided into: There are 10 time periods, and the step size for each time period is 1. .

[0030] In one specific embodiment, step S2 includes: Historical operating data for each transformer substation is collected, including load data and photovoltaic output data. Based on this historical operating data, the first... Taiwan area at all times The load demand is , define the first Taiwan area at all times The photovoltaic output is ,in , .

[0031] Simultaneously, meteorological data for the corresponding time period is collected, and time is defined based on the meteorological data for that time period. The weather condition index is , It is an environmental perception vector or normalized index that takes real-time temperature as the core quantitative feature and combines it with meteorological factors such as temperature change rate and wind speed to comprehensively generate an environmental perception vector or normalized index, which is used to characterize the intensity and evolution trend of extreme weather in multiple dimensions.

[0032] In one specific embodiment, step S3 includes: Please see Figure 3 , Figure 3 This is a schematic diagram of the power conversion process of a power router provided in an embodiment of the present invention. In this embodiment, the internal power conversion process of the power router is equivalent to a simplified multi-port model with specific efficiency and transmission capacity constraints, in order to focus on the top-level steady-state scheduling strategy.

[0033] To eliminate the ambiguity of power attributes, Figure 3 The text clearly defines and distinguishes three types of core power through different energy paths. The first is AC-power (alternating current exchange power), indicated by the symbol... This indicates that it represents the energy flow between the AC port of the power router and the transformer and AC load in the distribution area; secondly, it represents the DC power (DC-Power), which includes the DC power connected to the photovoltaic port, i.e., the photovoltaic output. DC power for charging and discharging connected to the energy storage port And DC switching power that can be exchanged across regions with an external multi-port controller via a DC-DC port through a DC bus. Finally, there is the internal power conversion, which refers to the energy flow logic between the central DC bus and the ports of each converter module inside the power router. Figure 3 middle , , , It refers to the first The relevant power of the transformer area.

[0034] Based on the aforementioned power categories, this embodiment further defines relevant conversion efficiency parameters to quantify conversion losses. Among them, the power electronic conversion efficiency is defined as... It is used to correct energy loss during AC / DC conversion; the charging efficiency of an energy storage system is defined as... Discharge efficiency is defined as These are used to constrain the actual output of the energy storage system under different operating conditions. Based on the principle of energy conservation, this embodiment establishes the DC bus power balance equation for each transformer area. This equation couples the AC-side switching power, DC-side power, and internal conversion power together to ensure the... Taiwan area at all times The input and output power reach a dynamic balance, thus providing accurate physical constraints for subsequent reinforcement learning algorithms.

[0035] Specifically, for the first The internal power balance equation of the DC bus in the transformer substation is as follows:

[0036] in, For the first Taiwan area at all times Photovoltaic power output; For the first The energy storage system in the distribution area is always The discharge power, i.e. The absolute value of the negative part; For the first The energy storage system in the distribution area is always The charging power, i.e. The positive part; For the first The DC-DC port of the transformer power router is at the time The power flowing from the DC bus to the central DC-DC module, i.e. The positive part; For the first The DC-DC port of the transformer power router is at the time The power flowing into the DC bus from the central DC-DC module, i.e. The absolute value of the negative part; This represents the power exchanged between the power router and the AC side. A positive value indicates power flowing into the AC side, while a negative value indicates power flowing into the power router. For power electronic conversion efficiency, For the discharge efficiency of the energy storage system The charging efficiency of the energy storage system.

[0037] A central DC-DC power balance equation is established based on the energy interaction logic between the DC interconnection system and the external AC grid in the interconnected topology architecture. Specifically, the transmission efficiency of the modular multilevel converter is defined as follows: Then, the power transmission equation through the modular multilevel converter is the interconnection power balance equation, and its expression is:

[0038] in, For the first The DC-DC port of the transformer power router is at the time The power exchange rate is positive, indicating that energy flows from the DC bus of the distribution area to the central multi-port converter DC-DC module, while a negative value indicates that energy is injected from the central multi-port converter DC-DC module into the DC bus of the distribution area. For modular multilevel converters at time The transmission power; This refers to the number of interconnected stations.

[0039] In one specific embodiment, step S4 includes: Definition of the first The energy storage equipment in the distribution area is always The charging and discharging power is Positive values ​​indicate charging, and negative values ​​indicate discharging; the first The maximum charging and discharging power of the transformer area energy storage system is Rated capacity is .

[0040] To ensure the safe operation of distributed energy storage systems during extreme weather dispatching, their charging and discharging power must be strictly limited within the allowable range of the equipment's physical characteristics, and must meet the following requirements:

[0041] This constraint constitutes the key boundary for the subsequent action output of the deep reinforcement learning agent.

[0042] The state-of-charge (POC) update equation for a distributed energy storage system is:

[0043] in, For the first The energy storage system in the distribution area is always The state of charge; For the first The energy storage system in the distribution area is always The state of charge of , with values ​​ranging from [0,1]; For the first The rated capacity of the energy storage system in the distribution area The step size for each scheduling period.

[0044] In one specific embodiment, step S5 includes: The scheduling behavior of power routers and distributed energy storage is constructed as a Markov decision process, consisting of a quintuple. Description, in which For state space, For the action space, Let be the state transition probability. For the reward function, This is the discount factor.

[0045] S51. Define the state space of a Markov decision process.

[0046] No. The intelligent agent in the district is at all times Local observation status Defined as:

[0047] in, It serves as a weather condition indicator, used to provide environmental intensity perception; For all within the distribution network Each interconnected platform is at any time The set of load demands For all within the distribution network Each interconnected platform is at any time Photovoltaic power generation collection, and This indicates that the observation vector contains load and energy information for the entire distribution area; For the first The energy storage system in the distribution area is always The state of charge is used as a local feedback constraint for decision-making; This is a normalized time feature for the current scheduling period, used to endow the agent with time perception capabilities.

[0048] This embodiment uses composite state coding that includes environmental indicators, global energy status, local device status, and time information to enable the agent to capture the spatiotemporal heterogeneity of load under extreme weather conditions.

[0049] The global state space is formed by the local observation states of each station area:

[0050] in, This is the global state space.

[0051] S52. Define the action space of a Markov decision process.

[0052] The action space is formed by the joint action space of the agents in each station area:

[0053] in, For the joint action space of intelligent agents in each station area, For the first The intelligent agent in the district is at all times The action vector, , , , For the first The energy storage system in the distribution area is always The charging and discharging power, For the first The maximum charging and discharging power of the energy storage system in the distribution area, For the first The DC-DC port of the transformer power router is at the time Power exchange quantity This represents the maximum transmission power of the DC-DC port.

[0054] S53. Define the state transition probabilities of a Markov decision process.

[0055] In extreme weather scenarios, the arrival time, intensity, and load response to extreme weather are all uncertain, and the state transition probability function characterizes these uncertainties. At any given time period... In a given state, the distribution network Execute action Then transition to the next state The state transition probability is:

[0056] in, For any time Distribution network in a given state Execute action Then transition to the next state The transition probability; It represents probability operations, used to measure the likelihood of an event occurring under given conditions; Representative moment State random variable, and Representing time respectively Random variables relating to state and action; For a moment The specific state observations of the system. For a moment The specific scheduling action value executed by the intelligent agent.

[0057] S54. Design a multi-objective reward function for extreme weather scenarios.

[0058] The reward function, taking into account the system's economy, reliability, and security, is defined as follows:

[0059] in, For a moment Instant reward value; The basic reward represents the fundamental compensation for successfully maintaining the system's operation and is set to a constant value. Penalties for network damage; Penalty for voltage offset; As a system crash penalty, when the system voltage exceeds the limit severely, a large penalty is imposed and the current training round is terminated. Penalty for exceeding power limits.

[0060] Network loss penalty is defined as:

[0061] in, For the system at time Total network loss per unit value; This represents the network loss penalty coefficient.

[0062] Voltage offset penalty is defined as:

[0063] in, For the system time The minimum node voltage per unit value, For the system time Maximum node voltage per unit value, This is the per-unit value of the lower voltage threshold. This is the per-unit value of the upper voltage threshold. This is the voltage offset penalty coefficient.

[0064]

[0065] in, For the first The transformer in the distribution area is at the time The per-unit value of the transmission power, For the first The rated power per unit value of the transformer in the distribution area. For modular multilevel converters at time The per-unit value of transmission power, This refers to the per-unit rated power of the modular multilevel converter. This is the power over-limit penalty coefficient.

[0066] System crash penalty items This is to address the risk of voltage instability under extreme cold waves. When the minimum node voltage per unit value in the system... Less than 0.90 pu, or the maximum node voltage per unit value When the CPU usage exceeds 1.10 PU, the system is considered to have entered a severe over-limit state. Once this state is triggered, the system will apply a quantified "large penalty," for example, the penalty value is set to a constant. (This value is usually much larger than the immediate reward value during normal periods), and the current training round is terminated immediately. This tiered reward design forces the agent to develop a sense of awe for voltage safety boundaries in the early stages of training, ensuring that it prioritizes maintaining the voltage within a safe range through power distribution via the power router during subsequent fluctuations in cold wave intensity.

[0067] It is understandable that in the calculation of the above reward function, both power and voltage are calculated using per-unit values.

[0068] In one specific embodiment, step S6 includes: The agent is implemented based on any one of the following algorithms: Multi-Agent Soft Actor-Critic (MASAC), Dual-Delay Deep Deterministic Policy Gradient Algorithm, and Proximal Policy Optimization Algorithm. Among them, the soft actor-critic algorithm includes an actor network and a critic network, both of which are constructed based on any one of the following network architectures: Transformer encoder, Mamba state space model, and Long Short-Term Memory network.

[0069] This embodiment further illustrates the specific implementation of the agent using a multi-agent soft actor-critic algorithm, where both the actor network and the critic network are constructed based on a Transformer encoder. Please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a diagram illustrating the overall framework of the intelligent agent implementation algorithm constructed for an embodiment of the present invention.

[0070] S61. Construct an actor network.

[0071] Actor Network Corresponds Figure 4The intelligent agent online actor network includes: a history collection input module, an auxiliary feature input module, an observation state encoding module, a linear projection layer, an overlay module, a first position encoding module, a first Transformer encoder, a first linear feature extraction module, a mean output head, a standard deviation output head, an action distribution module, and an action sampling module.

[0072] The historical data collection and input module is used to construct a historical state sequence from the local observation states of each transformer area agent. These local observation states include weather condition indicators, load demand, and distributed photovoltaic output (i.e.,...). Figure 4 The photovoltaic (PV) and energy storage state of charge and normalized time characteristics.

[0073] Specifically, for the first Each intelligent agent in the platform area, at any time The local observation state adopts the local observation state in step S51. :

[0074] According to time The local observation state construction length is Historical state sequence:

[0075] in, This represents the final state of the sequence, corresponding to the input at the last time step.

[0076] The auxiliary feature input module is used to concatenate the current state of charge with the action at the previous time step, and project the concatenated auxiliary input vector onto a unified feature space through a linear mapping to obtain the auxiliary feature projection.

[0077] Specifically, the definition of the first Each intelligent agent in the platform area is at all times Execution action for:

[0078] The current state of charge is concatenated with the action from the previous time step to form an auxiliary input vector. :

[0079] Then, it is projected onto a unified feature space through a linear mapping to obtain the auxiliary feature projection. :

[0080] in, and These are the auxiliary feature projection matrix and the bias term, respectively.

[0081] The observation state encoding module is used to encode the historical state sequence to obtain the encoded sequence; the linear projection layer is used to perform linear projection on the encoded sequence to obtain the projection result; the overlay module is used to overlay the projection result of the last time step with auxiliary feature projection to form the input sequence.

[0082] Specifically, for any moment in the historical state sequence Observation status First, state encoding is performed to obtain the encoded sequence. Then perform linear projection to obtain the projection result. Specifically, it is expressed as:

[0083]

[0084] in, , These are the encoding matrix and bias term for the observation state encoding, respectively; , These are the linear projection matrix and bias term for the linear projection layer, respectively.

[0085] Then, input the last time step. Projection results Superimposed auxiliary feature projection Due to superposition features Forming the input sequence :

[0086]

[0087] The first position encoding module is used to obtain the final input sequence of the first encoder after adding position encoding to the input sequence.

[0088] Specifically, to enable the model to recognize temporal order, positional encoding is added to the input sequence. Position is defined. and dimensional index The corresponding sine position code is:

[0089]

[0090] in, This represents the feature dimension of the model.

[0091] Encode the position With input sequence By adding element by element, the final input sequence of the first encoder is obtained. :

[0092] The first Transformer encoder is used to extract long-range dependency features of the final input sequence of the first encoder using a multi-head self-attention mechanism, so as to obtain the encoded first hidden state sequence.

[0093] Specifically, for the first Each attention head first maps the final input sequence of the first Transformer encoder to a query matrix. Key matrix Sum matrix :

[0094] In the formula, , , The first The query, key, and value projection matrices for each attention head. The corresponding single-head self-attention output. for:

[0095] In the formula, For the key vector dimension. All The attention heads are concatenated and linearly transformed to obtain the multi-head self-attention output. :

[0096] In the formula, For multi-head output projection matrix, For splicing.

[0097] Based on this, each Transformer encoder layer includes residual connections, layer normalization, and a feedforward network, the calculation process of which is expressed as follows:

[0098]

[0099]

[0100] In the formula, For activation function, , , , These are the weights and biases of the feedforward network. For layer normalization, This is the output of the feedforward network.

[0101] If the actor network includes The Transformer encoding layer, then the first... Layer output The recursive representation is as follows:

[0102]

[0103] Finally, the encoded first hidden state sequence is obtained. :

[0104] The first linear feature extraction module is used to extract the feature vector corresponding to the last position of the first hidden state sequence and obtain the shared features through linear mapping.

[0105] Specifically, the feature vector corresponding to the last position of the first hidden state sequence is extracted, which is the sequence representation at the current decision time. :

[0106] In the formula, The length of the input sequence is given. Shared features are then obtained through linear mapping. express:

[0107] In the formula, , This refers to the linear extraction matrix and bias term of the linear feature extraction module.

[0108] The mean output header and standard deviation output header are used to generate the mean and standard deviation parameters of the action distribution based on shared features, respectively.

[0109] Specifically, the mean output header is used to generate the mean parameter of the action distribution. :

[0110] in, , These are the weights and biases of the mean output head.

[0111] The standard deviation output header first outputs the logarithmic standard deviation, and then the standard deviation parameter is obtained through exponential mapping. :

[0112]

[0113] in, , The weights and biases of the standard deviation output head.

[0114] Mean parameter and standard deviation parameter Both are two-dimensional vectors, corresponding to the charging and discharging power actions of the energy storage system and the power exchange actions of the interconnection ports of the power router (i.e., the power exchange actions of the DC-DC ports).

[0115] The Action Distribution module is used to construct action distributions based on mean and standard deviation parameters.

[0116] Specifically, the first An intelligent agent at time Action distribution For Gaussian policy distribution:

[0117] In the formula, This represents all parameters of the actor network.

[0118] The action sampling module is used to sample continuous actions from the action distribution using reparameterization techniques. The sampling results are hyperbolic tangent compression and scaled according to physical ratings based on the charging and discharging power conditions and the maximum transmission power of the interconnect port to obtain the action vector of the agent.

[0119] Specifically, reparameterization techniques are used to sample continuous motions from the motion distribution. :

[0120]

[0121] To meet the motion space constraints, the sampling results are hyperbolic tangent compression and scaled to physical rated values ​​based on the charging / discharging power conditions and interconnect port power exchange conditions to obtain the final control action:

[0122]

[0123] in, The charging and discharging power of the energy storage system is activated. This refers to the power exchange quantity of the interconnected ports of the power router.

[0124] Thus, the first Each intelligent agent in the platform area is at all times Action vector:

[0125] In the formula, For ESS charging and discharging power output, i.e. the first The energy storage system in the distribution area is always The charging and discharging power; For EER DC mutual power output, i.e. the first The DC-DC port of the transformer power router is at the time The amount of power exchange.

[0126] This embodiment achieves flexible control of energy flow across multiple nodes in a building power distribution network by transforming distributed sampling into physical execution, thereby ensuring the resilience of the system operation under extreme weather conditions through joint scheduling in both spatiotemporal dimensions.

[0127] S62. Construct a network of critics.

[0128] The critic network is used to evaluate the state action values ​​of the actor network output actions, and the value overestimation is suppressed by the dual critic minimum mechanism.

[0129] The critic network comprises a first online critic network, a second online critic network, a first target critic network, and a second target critic network. Each of these networks includes: a state and action merging module, a second linear feature extraction module, a second position encoding module, a second Transformer encoder, a sequence post-processing module, a first Q-output layer, and a second Q-output layer.

[0130] The state and action merging module is used to concatenate the local observation state at each time step with the sampled action vector to form a state-action coupled vector.

[0131] Specifically, for the first Each intelligent agent in the platform area, at any time Forming length is Historical state action sequence :

[0132] in, For a moment The local observation status, For a moment The action vector, for to At any time.

[0133] Then, for any given moment By concatenating the state and action, we obtain the state-action coupling vector:

[0134] The input sequence of the critic network is composed of the coupling vectors at all times. :

[0135] The second linear feature extraction module is used to map the state-action coupling vector to a unified dimensional space through linear projection to obtain the state-action feature sequence.

[0136] Specifically, firstly, the state-action coupling vector is mapped to a unified high-dimensional feature space to obtain the mapped features. :

[0137] in, , This represents the linear mapping matrix and bias term of the second linear feature extraction module.

[0138] Based on mapping features Constituting a sequence of state-action features :

[0139] The second position encoding module is used to add position encoding to the state action feature sequence to obtain the final input sequence of the second encoder.

[0140] Specifically, in order to preserve sequence position information, position encoding is added to the above embedded sequence. The position encoding format is the same as that of the actor network position encoding, that is:

[0141]

[0142] The summation yields the final input sequence for the second encoder. :

[0143] The second Transformer encoder is used to extract long-range spatiotemporal features of the final input sequence of the second encoder using a multi-head self-attention mechanism, thereby obtaining the second hidden state sequence. .

[0144] The structure of the second Transformer encoder is the same as that of the first Transformer encoder, and will not be described again here.

[0145] The sequence post-processing module is used to extract the feature representation corresponding to the current time step from the second hidden state sequence to obtain refined features.

[0146] Specifically, the sequence post-processing module starts from the second hidden state sequence. The refined features are obtained by extracting the feature representation corresponding to the current time step. The feature of the last time step is preferred and is represented as follows:

[0147] in, The sequence length is given.

[0148] In other implementations, the second hidden state sequence can be aggregated to obtain refined features, but terminal feature extraction is preferred.

[0149] The Q-output layer is used to output state action values ​​based on refined features.

[0150] Specifically, Q-output layers are constructed for both the first and second online critic networks, meaning two structurally identical but parameter-independent online critic output branches are built. The first online critic network... The output layer's state action values ​​are:

[0151] The Second Critics Network The output layer's state action values ​​are:

[0152] In the formula, and These represent the network parameters of the first and second online commentators, respectively. , These represent the weights and biases of the first online commentator network, respectively. , These represent the weights and biases of the second online commentator network, respectively.

[0153] To suppress overestimation of the Q value, according to Figure 4 The min module takes the smaller value between the state / action values ​​output by the first online commentator network and the second online commentator network as the final state / action value of the online commentator network.

[0154] Furthermore, for the first target critic network and the second target critic network, their parameters are respectively... and Then, the target commentator outputs for the next state-action pair are as follows:

[0155]

[0156] The smaller of the two values ​​is taken to obtain the target state action value:

[0157] Furthermore, by combining immediate rewards and target commentator output, a TD target value is defined. for:

[0158] In the formula, For a moment Instant rewards As a discount factor, For temperature parameters, The strategy distribution for the actor network output.

[0159] Correspondingly, the TD errors of the two online commentator networks are defined as follows:

[0160]

[0161] Their loss functions are as follows:

[0162]

[0163] in, This represents the expectation operation. This is an experience replay pool.

[0164] Total loss function of the critic network for:

[0165] S63. Implement a centralized training process.

[0166] Specifically, the training method for the agents includes: based on Markov decision processes, multiple agents are centrally optimized using global state and joint action information. The training process includes experience collection and storage, commentator network updates, actor network updates, automatic temperature parameter adjustment, and soft updates of the target commentator network, resulting in a trained agent for each distribution area. Furthermore, during agent training, the charging and discharging power conditions are hyperbolic tangent compression and linear scaling according to physical rated values ​​applied to the sampled results of the actor network output, ensuring that the generated scheduling commands are strictly within the limits of the energy storage hardware from the source. The internal power balance equation is verified through power flow calculation and a reward / penalty mechanism. After each training step, the actual power values ​​of each port are obtained through power flow calculation and substituted into the internal power balance equation for physical consistency verification. If the action vector leads to power imbalance or violates the quasi-steady-state operation law of the power system, the agent is negatively corrected through the power limit violation penalty term in the reward function, thereby guiding the strategy to achieve optimized scheduling of the distribution network while satisfying energy conservation constraints.

[0167] Please see Figure 5 , Figure 5 This is a flowchart illustrating the training process for an intelligent agent according to an embodiment of the present invention. The training process includes: 1) Initialize network parameters and experience replay pool.

[0168] 2) Perform the main training loop, which mainly includes two stages: experience collection and storage, and updating the MASAC core network.

[0169] During the experience acquisition and storage phase, each agent continuously interacts with the power distribution network environment, encapsulating the generated scheduling experience into experience tuples and storing them in the experience replay pool. The specific steps include: a. Reset the distribution area environment. Before each training round, initialize the simulation environment to simulate various extreme weather conditions. The reset includes the initial values ​​of the load of each transformer area, the initial value of the energy storage state of charge (SOC), the initial value of the distributed photovoltaic output, as well as the grid topology and simulation clock, etc. By setting heterogeneous initial scenarios, the generalization ability and robustness of the agent under changing conditions are improved.

[0170] b. Acquire environmental information to perceive the multi-dimensional operational characteristics of the current power system in real time. Environmental information includes weather condition indicators characterizing the intensity of extreme weather events, and the current status of each distribution area. t The load demand, real-time distributed photovoltaic output, energy storage status of charge, and normalized time characteristics provide data support for scheduling decisions that combine global and local data.

[0171] c. Extracting State Vectors. The raw environmental information is preprocessed and feature-encoded, converting it into local observation state vectors that the agent can directly read. This step aims to map high-dimensional physical environment parameters to the algorithm's feature space, enabling the agent to capture long-range spatiotemporal dependency features of historical sequences through its internal Transformer architecture.

[0172] d. Agent generates action vectors. The agent network of each substation performs forward propagation based on the local observed state vector, outputs action distribution parameters, and samples them to obtain preliminary continuous actions. This action represents the agent's optimal energy scheduling intention based on the current weather and load conditions.

[0173] e. Action pruning. To ensure that scheduling instructions comply with hardware physical constraints, the sampled action results are subjected to hyperbolic tangent compression based on the charging and discharging power conditions of the distributed energy storage system and the maximum transmission power of the power router interconnection port, and are scaled proportionally according to the physical rating of the equipment, thereby limiting the numerical output at the algorithm level within the hardware safe operating boundary.

[0174] f. ESS charging / discharging and power exchange operation execution. The trimmed physical action vectors are sent to the distribution network simulation environment. By controlling the charging / discharging behavior of the distributed energy storage system and the power exchange amount of the DC-DC port of the power router, power mutual assistance and spatiotemporal regulation between substations with different load characteristics are realized at the physical level.

[0175] g. Power Flow Calculation. The standard solution module for distribution network dispatch simulation is invoked to calculate the voltage magnitude, phase angle, branch power, and network loss of each node after the action is executed. The calculated actual power values ​​of each port are substituted into the internal power balance equation for physical consistency verification, ensuring that the agent's dispatch scheme strictly adheres to basic power grid physical constraints such as Kirchhoff's current law. If the action vector leads to power imbalance or violates the quasi-steady-state operation law of the power system, negative feedback correction is applied to the agent through the power limit violation penalty term in the reward function.

[0176] h. Voltage over-limit or severe overload judgment, real-time monitoring of system operation safety indicators. Combining a multi-objective reward function, determine whether the node voltage exceeds the safety threshold or whether the transformer is severely overloaded; if not, it indicates that the operating state is safe, and enters the cumulative operating reward branch to guide the agent to maintain stable scheduling; if yes, it indicates that an over-limit has been triggered, and enters the calculation failure penalty branch and terminates the current round of exploration, forcing the agent to perceive and avoid the safety boundary through a tiered negative reward.

[0177] i. Store experiences in the experience replay pool The current state-action-next state-instant reward sequence is encapsulated into an explicit four-tuple experience tuple and stored in the experience replay pool.D These scheduling experience data provide sample support for subsequent offline gradient updates of the actor-critic network, ensuring that the agent can continuously optimize its scheduling strategy based on historical experience.

[0178] Specifically, an empirical tuple is defined as... ,in Represents the state of the intelligent agent. Next action Then, the system provides an immediate reward value at the next moment. This explicit quadruplet storage method can completely record the dynamic evolution of the distribution network state, providing high-quality sample support for subsequent offline strategy optimization.

[0179] Updating the MASAC core network includes updating the critic network, updating the actor network, automatically adjusting the temperature parameter, and soft updating the target critic network.

[0180] During the commentator network update phase, the system moves from the experience replay pool. Sampling of random batches of data and calculation of the objective value of the Q function. Then, based on the target value, the TD error is calculated, and the total loss function of the commentator network is calculated to update the commentator network.

[0181] During the actor network update phase, the optimization objective of the actor network is to maximize the weighted sum of cumulative reward and policy entropy:

[0182] in, For policy entropy, For strategy State-action distribution.

[0183] Furthermore, the policy gradient of the actor network is calculated using a reparameterization method to update the actor network:

[0184] in, This is the reparameterized action sampling function. Indicates the network parameters of the actor gradient calculation, Optimize the objective function for the actor network strategy. For parameters The strategy function, .

[0185] During the automatic temperature parameter adjustment phase, the temperature parameter Through the following objective function Adaptive update:

[0186] in, Let the target entropy value be the negative of the action space dimension. When the policy entropy is higher than the target entropy... Reduce to encourage exploitation when policy entropy is lower than target entropy. Increase to encourage exploration.

[0187] During the soft update phase of the target commentator network, the target commentator network is updated using an exponential moving average method:

[0188] in, For the first Parameters of a target critic network; This is the soft update coefficient, and a small positive value is chosen to ensure training stability.

[0189] 3) Determine if convergence has been achieved. If not, return to step 2) and repeat the main training loop. If yes, verify the model performance on the validation set and end the training to obtain the trained agent for each substation.

[0190] Furthermore, after training, the actor network of intelligent agents in each station area... Each agent independently outputs scheduling actions based solely on its local observation state. During each scheduling period, the agent generates an action vector. It has a clear hardware mapping logic: the first component is the charging and discharging power of the energy storage system. The first component controls the physical actuators of the distributed energy storage system; the second component corresponds to the power exchange quantity at the DC-DC port of the power router. This is used to regulate the energy flow between the power router in the distribution area and the central multi-port converter. The effectiveness of this decision-making process in different scenarios is as follows: Figure 6 and Figure 7 As shown, Figure 6 This is a schematic diagram of load optimization during normal periods. Figure 7 This diagram illustrates peak load smoothing during cold waves. During normal operation, the dispatching strategy focuses on smoothing the load within each transformer substation, improving operational efficiency through local regulation of the energy storage system. However, under extreme cold wave scenarios, this strategy exhibits significant spatiotemporal coordination characteristics: spatially, through the DC-DC ports of the power router, the remaining line capacity in lightly loaded areas is precisely allocated to heavily loaded residential or industrial areas, achieving cross-regional power sharing; temporally, by dynamically coordinating the charging and discharging sequence of the energy storage system, energy is stored in advance before the load truly reaches its peak or fully supported during the peak period, thereby jointly achieving peak load smoothing and system voltage stability.

[0191] Please see Figure 8 , Figure 8This is a schematic diagram illustrating the superior performance of this invention in long-range scheduling. Compared to traditional architectures where gradient signals decay with step size and struggle to handle long-distance causal relationships, the Transformer architecture employed in this embodiment achieves all-to-all correlation modeling through its unique self-attention mechanism. This mechanism endows each distribution area agent with the ability to perceive long-range credit allocation, enabling it to identify the direct causal impact of early scheduling actions on later system states. By assigning more accurate credit signals to historical decisions with high weights, the Transformer architecture effectively improves scheduling quality over long time windows. Ultimately, this scheduling scheme, combining precise control of physical variables, flexible cross-regional power allocation, and long-range causal perception capabilities, provides a solid technical guarantee for building power distribution networks to cope with the impact of extreme cold waves.

[0192] This invention connects multiple building distribution substations with different load characteristics through an energy router. It utilizes the differences in load variation patterns of residential, industrial, and commercial areas under the impact of cold waves to achieve cross-regional load transfer through power complementarity between building areas with different load characteristics. The remaining line capacity of the less loaded areas is allocated to the more loaded areas. In terms of time, the load peaks are transferred through the charging and discharging strategy of the distributed energy storage system. In both time and space, peak shaving and valley filling are carried out in a coordinated manner to achieve a smooth load curve and eliminate transformer overload problems during cold waves.

[0193] The reward function of the Markov decision process of this invention includes a voltage offset penalty term and a system crash penalty term. When the node voltage approaches the over-limit threshold, a penalty proportional to the offset is applied. When the voltage is severely over-limit, the training is terminated and a large penalty is applied. The trained strategy can control voltage fluctuations while smoothing the load, guiding the agent to prioritize meeting voltage safety constraints.

[0194] This invention employs a model-free deep reinforcement learning scheme, where the agent learns scheduling strategies through direct interaction with the environment, without relying on precise load prediction models. The state space includes weather condition indicators, allowing the agent to perceive changes in cold wave intensity and adjust scheduling actions accordingly. Simultaneously, the actor and critic networks are built based on Transformer encoders, utilizing self-attention mechanisms to directly model the dependencies between different positions in the historical scheduling sequence. Compared to traditional recurrent neural networks, the Transformer architecture enables the agent to perceive the impact of early scheduling actions on later system states, improving credit allocation in long-window scheduling.

[0195] In this invention, each distribution substation corresponds to an independent intelligent agent. During the training phase, global information is used to optimize the coordination strategy. During the execution phase, each intelligent agent makes independent decisions based only on local observations. This division corresponds to the physical topology of the distribution network, reduces the state-action decision space dimension of a single intelligent agent, and retains the coordination and optimization capabilities between substations.

[0196] In summary, the method of this invention achieves power mutual assistance between distribution networks with different load characteristics through power routers and distributed energy storage systems. The introduction of weather condition indicators enables the agent to directly perceive the intensity and evolution trend of extreme weather. The reward function guides the agent to prioritize voltage safety constraints while smoothing load peaks. Under the premise of satisfying physical constraints, it effectively alleviates transformer overload and voltage deviation problems, realizes load smoothing and voltage stability of the distribution network during extreme weather, improves the operational resilience and scheduling decision quality of the distribution network under extreme weather scenarios, and ensures the safe and reliable operation of the distribution system under extreme weather conditions.

[0197] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A power distribution network power router optimal scheduling method for extreme weather scenarios, characterized in that, Including the following steps: Construct an interconnected topology architecture that includes multiple transformer substations with heterogeneous load characteristics. Each substation is equipped with an energy router and a distributed energy storage system, and each substation corresponds to an intelligent agent. The load demand and distributed photovoltaic output of each transformer area are defined based on historical operating data, and weather condition indicators characterizing extreme weather are defined based on meteorological data of the corresponding time period. Establish the internal power balance equations and interconnection power balance equations for each transformer area; Establish the charging and discharging power conditions and state of charge update equations for the distributed energy storage system; The scheduling behavior of the power router and the distributed energy storage system is constructed as a Markov decision process. The state space of the Markov decision process includes weather condition indicators, load demand, distributed photovoltaic output, energy storage state of charge and time characteristics, the action space includes the charging and discharging power of the energy storage system and the power exchange of the interconnection ports of the power router, and the reward function includes network loss penalty, voltage deviation penalty and power limit violation penalty established based on the interconnection power balance equation. The agent is trained based on the Markov decision process, constrained by the internal power balance equation and the charging and discharging power conditions, to obtain the trained agent for each power station area; the trained agent for each power station area is used to output scheduling actions based on its local observation state.

2. The optimal scheduling method of power distribution network energy router for extreme weather scenarios according to claim 1, characterized in that, The interconnected topology includes multiple transformer substations with heterogeneous load characteristics, a central multi-port converter, and a modular multilevel converter. Each transformer substation is equipped with a power router, a distributed energy storage system, and a distributed photovoltaic system. The DC terminal of each of the power routers is connected to the DC terminal of the modular multilevel converter via a central multiport converter; the AC terminal of the modular multilevel converter is connected to the external AC power grid via a transformer and the system central bus. In each distribution area, the terminal power load is connected to the AC terminal of the power router, and the distributed energy storage system and the distributed photovoltaic system are connected to the DC terminal of the power router. Each distribution area exchanges energy with the external power distribution system through a transformer.

3. The method of claim 1, wherein, Based on historical operational data, the load demand and distributed photovoltaic output of each transformer area are defined, and weather condition indicators characterizing extreme weather conditions are defined based on meteorological data for the corresponding time periods, including: Define the first based on historical operational data. Taiwan area at all times The load demand is , define the first Taiwan area at all times The photovoltaic output is ; defining a weather condition indicator for a time instant based on meteorological data for a corresponding time period ;​ in, , , The number of interconnected stations, The number of time periods divided into the scheduling cycle.

4. The method of claim 2, wherein, Establish the internal power balance equations and interconnection power balance equations for each transformer area, including: Based on the principle of energy conservation, the DC bus power balance equations for each transformer substation are established. By coupling the AC-side exchange power, DC-side power, and internal conversion power, the internal power balance equations for each transformer substation are obtained: ; in, For the first Taiwan area at all times Photovoltaic power output; For the first The distribution area energy storage system is always The discharge power, For the first The energy storage system in the distribution area is always The charging power; For the first The DC-DC port of the transformer power router is at the time The power flowing from the DC bus to the central multiport converter DC-DC module, For the first The DC-DC port of the transformer power router is at the time Power flowing into the DC bus from the central multi-port converter DC-DC module; This refers to the power consumption of the power router interacting with the AC side. For power electronic conversion efficiency, For the discharge efficiency of the energy storage system, The charging efficiency of the energy storage system; Based on the energy interaction logic between the DC interconnection system and the external AC grid in the aforementioned interconnection topology, a central DC-DC power balance equation is established, resulting in the interconnection power balance equation: ; in, For the first The DC-DC port of the transformer power router is at the time Power exchange quantity For modular multilevel converters at time Transmission power, To improve the transmission efficiency of modular multilevel converters, This refers to the number of interconnected stations.

5. The method of claim 2, wherein, The charging and discharging power conditions of the distributed energy storage system are as follows: ; in, For the first The energy storage system in the distribution area is always The charging and discharging power, For the first The maximum charging and discharging power of the energy storage system in the distribution area; The state-of-charge update equation for the distributed energy storage system is: ; in, For the first The distribution area energy storage system is always The state of charge; For the first The distribution area energy storage system is always The state of charge of , with values ​​ranging from [0,1]; For the first The rated capacity of the energy storage system in the distribution area The step size for each scheduling period.

6. The method of claim 2, wherein, The Markov decision process includes a state space, an action space, state transition probabilities, a reward function, and a discount factor. The state space is formed by the local observation states of each station area: ; in, For the global state space, For the first The intelligent agent in the district is at all times The local observation status, , , The number of interconnected stations, As a weather condition indicator, For all within the distribution network Each interconnected platform is at any time The set of load demands For all within the distribution network Each interconnected platform is at any time Photovoltaic power generation collection, For the first The energy storage system in the distribution area is always The state of charge, This represents the normalized time characteristics of the current scheduling period; The action space is formed by the joint action space of the intelligent agents in each station area: ; in, For the joint action space of intelligent agents in each station area, For the first The intelligent agent in the district is at all times The action vector, , , , For the first The energy storage system in the distribution area is always The charging and discharging power, For the first The maximum charging and discharging power of the energy storage system in the distribution area, For the first The DC-DC port of the transformer power router is at the time Power exchange quantity, This represents the maximum transmission power of the DC-DC port. The reward function is: ; in, For a moment Instant reward value, Basic rewards, For network damage penalties, As a penalty for voltage offset, As a penalty for system crash, Penalty for exceeding power limits; , For the system at time Total network loss per unit value, This is the network loss penalty coefficient; , For the system time The minimum node voltage per unit value, For the system time Maximum node voltage per unit value, This is the per-unit value of the lower voltage threshold. This is the per-unit value of the upper voltage threshold. This is the voltage offset penalty coefficient; , For the first The transformer in the distribution area is at the time The per-unit value of the transmission power, For the first The rated power per unit value of the transformer in the distribution area. For modular multilevel converters at time The per-unit value of transmission power, This refers to the per-unit rated power of the modular multilevel converter. This is the power over-limit penalty factor; when Less than 0.90 pu or System crash penalty when PU exceeds 1.

10. Set as a constant ; state transition probabilities ; discount factor .

7. The method of claim 1, wherein, The intelligent agent is implemented based on a multi-agent soft actor-critic algorithm; The multi-agent soft actor-critic algorithm includes an actor network and a critic network, both of which are constructed based on a Transformer encoder.

8. The method of claim 7, wherein, The actor network includes: The historical data collection input module is used to construct a historical state sequence from the local observation states of each transformer area agent; wherein, the local observation states include weather condition indicators, load demand, distributed photovoltaic power output, energy storage state of charge, and normalized time characteristics; The auxiliary feature input module is used to concatenate the current state of charge with the action at the previous time step, and project the concatenated auxiliary input vector onto a unified feature space through a linear mapping to obtain the auxiliary feature projection. An observation state encoding module is used to encode the historical state sequence to obtain an encoded sequence. A linear projection layer is used to perform a linear projection on the encoded sequence to obtain a projection result; The overlay module is used to overlay the projection result of the last time step onto the auxiliary feature projection to form the input sequence; The first position encoding module is used to obtain the final input sequence of the first encoder after adding position encoding to the input sequence; The first Transformer encoder is used to extract long-range dependency features of the final input sequence of the first encoder using a multi-head self-attention mechanism to obtain the encoded first hidden state sequence. The first linear feature extraction module is used to extract the feature vector corresponding to the last position of the first hidden state sequence and obtain the shared features through linear mapping. The mean output header and standard deviation output header are used to generate the mean parameter and standard deviation parameter of the action distribution based on the shared features, respectively. The action distribution module is used to construct an action distribution based on the mean parameter and the standard deviation parameter; The action sampling module is used to sample continuous actions from the action distribution using reparameterization techniques. The sampling results are then subjected to hyperbolic tangent compression and scaled according to physical ratings based on the charging and discharging power conditions and the maximum transmission power of the interconnect port to obtain the action vector of the agent.

9. The method of claim 7, wherein, The commentator network includes a first online commentator network, a second online commentator network, a first target commentator network, and a second target commentator network. Each of the first online commentator network, the second online commentator network, the first target commentator network, and the second target commentator network includes: The state and action merging module is used to concatenate the local observation state at each time step with the sampled action vector to form a state-action coupled vector. The second linear feature extraction module is used to map the state-action coupling vector to a unified dimensional space through linear projection to obtain a state-action feature sequence. The second position encoding module is used to add position encoding to the state action feature sequence to obtain the final input sequence of the second encoder; The second Transformer encoder is used to extract the long-range spatiotemporal features of the final input sequence of the second encoder using a multi-head self-attention mechanism to obtain the second hidden state sequence. The sequence post-processing module is used to extract the feature representation corresponding to the current time from the second hidden state sequence to obtain refined features; Q output layer, used to output state action values ​​based on the refined features; Specifically, the smaller value between the state action value output by the first online commentator network and the state action value output by the second online commentator network is taken as the final state action value of the online commentator network. The smaller value between the state-action value output by the first target critic network and the state-action value output by the second target critic network is taken as the target state-action value of the target critic network.

10. The power router optimization scheduling method for distribution networks in extreme weather scenarios according to claim 7, characterized in that, The process of training the agent includes: Reset the power distribution area environment; Obtain environmental information, including weather conditions, load demand, distributed photovoltaic output, energy storage status of charge, and normalized time characteristics; Extract the local observation state vector based on the environmental information; The actor network of each station area agent performs forward propagation based on the local observation state vector, outputs action distribution parameters and samples to obtain continuous actions; Based on the charging and discharging power conditions of the distributed energy storage system and the maximum transmission power of the power router interconnection port, the continuous action is subjected to hyperbolic tangent compression and scaled proportionally according to the physical rating of the equipment to obtain the action vector of the intelligent agent. The interconnection port between the distributed energy storage system and the power router executes the action vector; The actual power values ​​of each port are obtained through power flow calculation and substituted into the internal power balance equation for physical consistency verification. If the action vector causes power imbalance or violates the quasi-steady-state operation law of the power system, the agent is negatively corrected through the power limit penalty term in the reward function. Based on the reward function, determine whether the node voltage exceeds the safety threshold or whether the transformer is severely overloaded; if the operating status is safe, proceed to the cumulative operating reward branch; if an over-limit is triggered, proceed to the calculation failure penalty branch and terminate this round of exploration. Store the experience tuples formed at the current moment into the experience replay pool; The process involves sequentially updating the critic network, updating the actor network, automatically adjusting the temperature parameters, and performing a soft update to the target critic network.