A method for intelligent control of power distribution network safety that takes into account the health and operating status of equipment.
Patent Information
- Application Number
- CN202610039395.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-26
Smart Images

Figure CN122092205A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system operation control, and in particular to a method for intelligent control of distribution network safety that takes into account the healthy operating status of equipment. Background Technology
[0002] As the link between the transmission system and users, the power distribution network is a crucial component for ensuring the smooth operation of the power system. With the acceleration of urbanization, urban high-voltage distribution networks face the dual pressures of high load density and high power supply reliability requirements. In distribution network operation, N-1 faults, i.e., the failure of a single major component such as the main transformer or a critical line, are the most common risk scenario. When an N-1 fault occurs, especially during periods of high temperature and high load, the remaining operating equipment often faces severe overload pressure. If not controlled in time, it can easily trigger cascading trips, leading to large-scale power outages.
[0003] However, existing technologies have significant limitations. First, the overload capacity of physical equipment is not fully utilized. Oil-immersed equipment such as transformers exhibit significant thermal inertia, and their insulation life depends on the winding hotspot temperature rather than the instantaneous current. In the short period following a fault, the equipment can withstand a certain degree of overload without damaging the insulation; this "short-term emergency load tolerance time" is an extremely valuable buffer period for dispatching, but traditional methods typically ignore this dynamic thermal characteristic. Second, equipment risk is probabilistic rather than binary. As temperature and load rate increase, the instantaneous failure rate of equipment rises exponentially, and simple current over-limit criteria cannot quantify this dynamic risk.
[0004] In recent years, Deep Reinforcement Learning (DRL) has been introduced into power system control. However, applying DRL to safety-critical distribution network fault recovery faces significant challenges. The core issue lies in the phenomenon of "cost underestimation." In Constrained Markov Decision Processes (CMDPs), agents, in pursuit of high rewards, often underestimate the potential costs of high-risk states, causing the policy to frequently cross safety red lines in the early stages of training and application. Existing constrained reinforcement learning algorithms, such as CPO and Lagrangian PPO, often exhibit slow convergence and severe oscillations when dealing with sparse and high-cost grid fault constraints, making it difficult to meet the practical engineering requirement of "absolute safety."
[0005] In summary, current emergency control methods for N-1 faults in distribution networks still face significant technical bottlenecks in terms of equipment safety boundary characterization, dynamic risk quantification, and security assurance through reinforcement learning. Practical engineering urgently requires a method for safe, rapid, and economical emergency control of distribution networks that can accurately characterize the physical and thermal characteristics of equipment and fault risks, effectively overcome the underestimation of reinforcement learning costs, and achieve this goal. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a power distribution network safety intelligent control method that takes into account the healthy operating status of equipment.
[0007] The objective of this invention is achieved through the following technical solution: a power distribution network safety intelligent control method considering the healthy operating status of equipment, the method comprising the following steps,
[0008] S1. Construct a digital simulation environment for the distribution network based on the real power grid topology;
[0009] S2. Establish an equipment health status assessment model, using the transformer's short-term emergency load capacity and the equipment's instantaneous failure rate as dual physical indicators to quantify the equipment's health status and form safety constraints.
[0010] S3. Model the emergency control process of N-1 fault in the distribution network as a constrained Markov decision process (CMDP), and define the state space, mixed action space, hierarchical reward function and external cost function.
[0011] S4. Construct a secure reinforcement learning agent that calculates the intrinsic cost based on historical high-cost states and corrects the total cost by combining the extrinsic cost.
[0012] S5. Train a safety reinforcement learning agent under the N-1 fault scenario, so that it learns to prioritize emergency control strategies such as topology transfer and minimum load shedding under the premise of meeting safety constraints.
[0013] S6. Deploy the trained security reinforcement learning agent in the power distribution network control system for online emergency control.
[0014] Specifically, the calculation of the transformer's short-term emergency load capacity in S2 is as follows:
[0015] Calculate the winding hot spot temperature:
[0016] ;
[0017] In the formula, Current load rate; Ambient temperature; This refers to the loss ratio; For winding index; The temperature rise of the top oil temperature relative to the ambient temperature under rated loss; The temperature rise of the top oil layer due to the hot spot temperature under rated current;
[0018] Based on the current winding hot spot temperature, calculate the remaining time the transformer can maintain operation until it reaches the upper limit of insulation withstand capability:
[0019] ;
[0020] In the formula, The equivalent thermal time constant of the winding; To ensure the transformer can continue operating until the insulation withstands its maximum temperature;
[0021] Minimum emergency response capability of the computing system:
[0022] ;
[0023] when Less than the current scheduling period At that time, it was determined to be a violation of the safety constraints on short-term emergency medical load capacity.
[0024] Specifically, the quantification of the instantaneous failure rate of the equipment in S2 includes:
[0025] Transformer failure rate quantification employs a Weibull + Arrhenius joint model to characterize the effect of temperature on the failure rate.
[0026] Convert the winding hot spot temperature to Kelvin temperature:
[0027]
[0028] Calculate temperature-related lifetime parameters:
[0029] ;
[0030] In the formula, A coefficient related to the nominal lifespan; These are parameters related to activation energy;
[0031] Assuming the fault time follows a Weibull distribution, the instantaneous fault rate of the transformer is calculated as follows:
[0032] ;
[0033] In the formula, For shape parameters; Equivalent runtime;
[0034] To quantify the line failure rate, the current carrying capacity correction factor is first calculated based on the ambient temperature and the reference temperature.
[0035] ;
[0036] In the formula, Ambient temperature; Reference temperature; This is the current carrying capacity correction factor; This refers to the maximum allowable temperature of the circuit.
[0037] The line operating temperature is approximated as:
[0038] ;
[0039] In the formula, For the first The rated current carrying capacity of the line; Real-time current obtained from power flow calculations;
[0040] The instantaneous fault rate of the line is calculated as follows:
[0041] ;
[0042] In the formula, Basic failure rate; .
[0043] Specifically, the state space is as follows:
[0044] ;
[0045] ;
[0046] ;
[0047] ;
[0048] ;
[0049] ;
[0050] ;
[0051] In the formula, Total number of main transformers; The main variable's operating state sub-vector; The state subvector representing the equipment failure rate; For electrical quantity state vectors; For the operation state sub-vector; This is the load state sub-vector; For environmental measurement subvectors; For the first Real-time load rate of the main transformer; This refers to the hot spot temperature of the winding. For the remaining short-term emergency tolerance time; The instantaneous failure rate of the main transformer at its hot spot temperature; This represents the instantaneous failure rate of the line under current or temperature rise conditions. For each bus voltage; For each line current; This indicates the real-time open / closed status of the interconnecting switch; This indicates the open / closed state of the feeder sectionalizing switch; Real-time active power of each node; Real-time reactive power of each node; This represents the maximum allowable load shedding ratio for each node; The current ambient temperature; This refers to the time points corresponding to the daily load curve or temperature curve.
[0052] Specifically, the hybrid action space is defined as:
[0053] The topology transfer action is used to control the opening and closing of critical tie switches. The output of each tie switch action is a continuous value:
[0054] ;
[0055] The final action is obtained through threshold mapping:
[0056] ;
[0057] Load shedding action is used to control the reduction ratio of the shearable load. Each controllable load node outputs a reduction rate:
[0058] ;
[0059] In the formula, To reduce the upper limit of the load; The reduction ratio instruction corresponding to the j-th controllable load node in the continuous action vector output by the agent, with a value range of [0,1].
[0060] The power after reduction is:
[0061] .
[0062] Specifically, the external cost function is calculated as follows:
[0063] The cost of power flow failure is determined by imposing the maximum penalty if the power flow calculation diverges or any node voltage exceeds the limit.
[0064] ;
[0065] Costs due to hotspot temperature exceeding limits, if any main transformer Added fixed penalties:
[0066] ;
[0067] Emergency response time is insufficient to cover costs, if Superimposed linear penalty:
[0068] ;
[0069] In the formula, One scheduling cycle;
[0070] Failure rate risk cost, if the average failure rate Exceeding the warning value Adding normalized risk costs:
[0071] ;
[0072] In the formula, This represents the average failure rate. This is a warning value; For extreme anchor points;
[0073] The warning value Based on the basic failure rate of the line It is deduced that 1.5 times the basic failure rate corresponds to the warning threshold for line risk just deviating from the normal level; the extreme anchor point The corresponding line operating temperature has reached the maximum allowable temperature. The failure rate at that time; the average failure rate The arithmetic mean of the instantaneous failure rates of the entire network is calculated using the following formula:
[0074]
[0075] in This represents the total number of distribution network lines. For the first The line at the time Instantaneous failure rate.
[0076] The final external cost is calculated as follows:
[0077] .
[0078] Specifically, the hierarchical reward function is:
[0079] ;
[0080] ;
[0081] ;
[0082] ;
[0083] ;
[0084] In the formula, The reward for improving emergency response time is used to encourage agents to improve the system's minimum remaining short-term emergency response tolerance time. As a reward weight; This represents the minimum remaining short-term emergency response tolerance time of the system after the current update. This represents the minimum remaining short-term emergency tolerance time of the system in the previous step. The switching operation is used as a penalty and reward mechanism to constrain the agent to reduce unnecessary switching changes. This is the penalty coefficient; The load reduction penalty reward is used to minimize the amount of load reduction in order to ensure power supply reliability. This is the penalty coefficient; This represents the total load shedding. As a reward for a complete successful rescue, when all main transformer load rates are... A one-time reward is given when the value is ≤1.0 (no overload) to encourage the agent to achieve global optimal control; Let be the real-time load rate of the i-th main transformer; num_ops is the number of switch changes in this step; This represents the total load shedding.
[0085] Specifically, the calculation of intrinsic costs based on historical high-cost conditions and the correction of total costs by combining extrinsic costs include:
[0086] Calculate the intrinsic cost:
[0087] ;
[0088] In the formula, For kernel functions; It is the nearest neighbor set; For embedding vectors; For state embedding vectors;
[0089] Total cost correction:
[0090] ;
[0091] In the formula, This is the adaptive balance coefficient.
[0092] The present invention has the following advantages:
[0093] 1. This invention constructs a quantitative model of the healthy operating state of equipment by calculating two physical indicators: the short-term emergency load capacity of the transformer and the instantaneous failure rate of the equipment. Based on this model, safety constraints are formed to achieve a unified constraint on thermal safety and risk safety, thereby avoiding premature or excessive load shedding in traditional control and improving the safety and sustainability of the system.
[0094] 2. This invention generates topology operation modes by thresholding continuous actions, thereby achieving coordinated optimization of topology transfer and continuous load shedding, enabling the agent to perform hybrid decision-making in a single strategy.
[0095] 3. The high-risk state memory and similarity intrinsic cost mechanism of the present invention can identify potentially dangerous states in advance, avoid the common problem of "cost underestimation" in reinforcement learning, and improve the stability and engineering usability of the strategy.
[0096] 4. Under the premise of meeting safety constraints, this invention prioritizes the use of inter-regional power transfer to reduce load loss and improve the control effect; the load reduction is significantly reduced, ensuring the economic efficiency of power supply. Attached Figure Description
[0097] Figure 1 This is a schematic diagram of the control method of the present invention;
[0098] Figure 2 This is a schematic diagram of the hybrid action space construction and thresholding mapping of the present invention;
[0099] Figure 3 This is a schematic diagram of the security enhancement decision-making mechanism of the present invention. Detailed Implementation
[0100] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described embodiments are merely some embodiments of the invention, and not all embodiments. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0101] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0102] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0103] The present invention will be further described below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0104] like Figures 1 to 3 As shown, a method for intelligent control of power distribution network safety that takes into account the healthy operating status of equipment includes the following steps:
[0105] S1. Based on the real power grid topology, construct a digital simulation environment for the distribution network. Specifically, this includes reading the topology, line parameters, transformer parameters, and initial load data of the actual urban distribution network to construct a digital model of the distribution network, including: high-voltage busbars, main transformers, distribution lines, sectionalizing switches and tie switches, important loads, and interruptible loads. Initialize physical parameters, including transformer thermal model parameters (such as thermal time constant). (Rated temperature rise, etc.) Line thermal stability limits and dispatching cycles Based on real power grid data, the agent recreates an actual N-1 fault scenario and learns how to adjust the line topology and load under these fault conditions to bring the system back to a safe operating range.
[0106] S2. Establish an equipment health status assessment model, using the transformer's short-term emergency load capacity and the equipment's instantaneous failure rate as dual physical indicators to quantify the equipment's health status and form safety constraints. This invention introduces the equipment health status as a criterion for safety boundaries, selecting the transformer's short-term emergency load capacity and the equipment's instantaneous failure rate as two core indicators characterizing the equipment's health status, and constructing constraints accordingly.
[0107] Health status characterization and constraints based on short-term emergency medical load capacity:
[0108] The short-term emergency load capacity of a transformer reflects the health potential of the equipment in the thermal overload dimension. The main transformer, its winding hot spot temperature Discretization update using the following differential form:
[0109] Calculate the winding hot spot temperature:
[0110] ;
[0111] In the formula, Current load rate; Ambient temperature; This refers to the loss ratio; For winding index; The temperature rise of the top oil temperature relative to the ambient temperature under rated loss; The temperature rise of the top oil layer due to the hot spot temperature under rated current;
[0112] Based on the current winding hot spot temperature The calculation shows that the transformer can continue to operate until the insulation withstand limit is reached. The remaining time, i.e., short-term emergency tolerance time. :
[0113] ;
[0114] In the formula, The equivalent thermal time constant of the winding; To ensure the transformer can continue operating until the insulation withstands its maximum temperature;
[0115] Minimum emergency response capability of the computing system:
[0116] ;
[0117] Minimum remaining withstand time for all main transformers It must be greater than or equal to one scheduling cycle. To ensure that the equipment does not burn out before the next dispatch instruction is issued; when Less than the current scheduling period At that time, it was determined to be a violation of the safety constraints on short-term emergency medical load capacity.
[0118] The instantaneous failure rate of equipment reflects the real-time failure risk of the equipment under the current operating environment and is a probabilistic representation of the equipment's healthy operating state. This invention constructs failure rate models for transformers and lines respectively;
[0119] Transformer failure rate quantification employs a Weibull + Arrhenius joint model to characterize the effect of temperature on the failure rate.
[0120] Winding hot spot temperature Convert to Kelvin temperature:
[0121]
[0122] Introducing equivalent running time Parameters related to activation energy Calculate temperature-related lifetime parameters:
[0123] ;
[0124] In the formula, A coefficient related to the nominal lifespan; These are parameters related to activation energy;
[0125] Assume the failure time follows a Weibull distribution, with shape parameter... Scale parameters , The instantaneous failure rate of the transformer is calculated as follows:
[0126] ;
[0127] In the formula, For shape parameters; This is the equivalent runtime.
[0128] Line failure rate quantification is first based on ambient temperature. Compared with reference temperature Calculate the current carrying capacity correction factor :
[0129] ;
[0130] In the formula, Ambient temperature; Reference temperature; This is the current carrying capacity correction factor; This refers to the maximum allowable temperature of the circuit.
[0131] The line operating temperature is approximated as:
[0132] ;
[0133] In the formula, For the first The rated current carrying capacity of the line; Real-time current obtained from power flow calculations;
[0134] The instantaneous fault rate of the line is modeled using a segmented approach:
[0135] ;
[0136] In the formula, Basic failure rate; .
[0137] Construct piecewise nonlinear functions based on line temperature or current, at line temperature The risk increases significantly when approaching the limit.
[0138] S3. Model the emergency control process of N-1 fault in the distribution network as a constrained Markov decision process (CMDP), and define the state space, mixed action space, hierarchical reward function and external cost function.
[0139] This invention describes the operating state of the distribution network under N-1 fault conditions using a parameterized vector form, constructing the system state vector at the current moment. The state space is as follows:
[0140] ;
[0141] In the formula, This represents the current power grid state; the sub-vectors are structured as follows:
[0142] Main transformer operating state sub-vector :
[0143] ;
[0144] In the formula, Total number of main transformers; For the first The real-time load rate of the main transformer is used to reflect the transformer's load-bearing capacity. This refers to the winding hot spot temperature, reflecting the current thermal stress of the main transformer; The remaining short-term emergency withstand time is used to characterize the remaining safety margin of the main transformer from the insulation temperature limit under the current operating conditions; this sub-vector is used to comprehensively characterize the thermal dynamics of the main transformer in a short-term overload scenario after an N-1 fault.
[0145] Equipment failure rate state subvector :
[0146] ;
[0147] In the formula, The instantaneous failure rate of the main transformer at its hot spot temperature is used to quantify the risk of insulation failure. This represents the instantaneous failure rate of the line under current current or temperature rise conditions, used to characterize the failure probability of conductors and connectors. This sub-vector can reflect the operational risk level of equipment in real time, enabling probability-based safety assessment.
[0148] Electrical quantity state subvector :
[0149] ;
[0150] In the formula, These are the voltages of each bus, used to reflect whether the node voltages are within the allowable range; This represents the current of each line, used to determine the power flow distribution and line load conditions. This sub-vector is used to ensure power flow stability and provides necessary reference for topology adjustment.
[0151] Operation state subvector :
[0152] ;
[0153] In the formula, This refers to the real-time open / closed status of the interconnection switch, used to describe the cross-regional power transfer capability. This represents the open / closed state of the feeder sectionalizing switch, used to record the current feeder topology. This sub-vector reflects the controllable topology configuration of the power grid.
[0154] Load state subvector :
[0155] ;
[0156] In the formula, Real-time active power of each node; Real-time reactive power of each node; This subvector represents the maximum allowable load shedding ratio for each node, reflecting the load adjustability. It provides a constraint range for calculating load shedding actions.
[0157] Environmental Measurement Subvector :
[0158] ;
[0159] In the formula, The current ambient temperature affects the calculation of equipment hotspot temperature and failure rate. These are time points corresponding to the daily load curve or temperature curve, used to characterize periodic fluctuations.
[0160] By constructing a state vector from the above six types of parameters, this invention can comprehensively reflect the operating status, equipment risks, topology, load demand, and environmental impact of the distribution network under N-1 fault conditions, providing complete and accurate input information for reinforcement learning decision-making of subsequent emergency control strategies.
[0161] Hybrid Action Space: Agents based on the current policy network Output the action vector, which is divided into:
[0162] The topology transfer action is used to control the opening and closing of critical tie switches. Each tie switch action output is a continuous value. By using continuous value output combined with a thresholding mapping mechanism, the continuous action is converted into discrete switch opening / closing commands:
[0163] ;
[0164] The final action is obtained through threshold mapping:
[0165] ;
[0166] It indicates "closing" and "opening".
[0167] Load shedding action is used to control the reduction ratio of shearable loads. Each controllable load node outputs a reduction rate, which is a continuous value used to directly control the adjustment amount of load power.
[0168] ;
[0169] In the formula, To reduce the upper limit of the load; The reduction ratio instruction corresponding to the j-th controllable load node in the continuous action vector output by the agent, with a value range of [0,1].
[0170] The power after reduction is:
[0171] .
[0172] The hierarchical reward function is used to guide policy preferences. It adopts a multi-level priority design to guide the agent to make optimal decisions. Specifically:
[0173] ;
[0174] In the formula, The reward for improving emergency response time is used to encourage agents to improve the system's minimum remaining short-term emergency response tolerance time. The switching operation is used as a penalty and reward mechanism to constrain the agent and reduce unnecessary switching changes. The load reduction penalty reward is used to minimize the amount of load reduction in order to ensure power supply reliability. A reward for a complete successful rescue; num_ops is the number of times the switch was changed in this step.
[0175] First level (highest priority): Enhance safety margin. For main transformers within the safety zone, reward them with increased emergency response time, encouraging further improvements in system robustness within the safety margin, which includes:
[0176] ;
[0177] In the formula, As a reward weight; This represents the minimum remaining short-term emergency response tolerance time of the system after the current update. This represents the minimum remaining short-term emergency tolerance time of the system in the previous step.
[0178] The second layer: Reduce the number of operations, encouraging strategies to minimize switching actions while achieving the goal, aligning with the "minimize operations" principle in engineering. Let num_ops be the number of switch changes in this step, then:
[0179] ;
[0180] In the formula, This is the penalty coefficient;
[0181] Third layer: Reduce load shedding, for the total load shedding amount Impose severe penalties to ensure power supply reliability; apply total load shedding. Apply linear penalty:
[0182] ;
[0183] In the formula, This is the penalty coefficient; This represents the total load shedding.
[0184] Fourth layer: Success reward, when all main transformer load rates... Upon successful (restoring all safe zones), an additional one-time success bonus will be awarded:
[0185] ;
[0186] In the formula, num_ops represents the number of times the switch changes position in this step; This represents the total load shedding; when the load rate of all main transformers... A one-time reward is given when the value is ≤1.0 (no overload) to encourage the agent to achieve global optimal control.
[0187] Multi-layered structures naturally encode the following priority order:
[0188] Prioritize ensuring and improving emergency tolerance time to guarantee safety margins;
[0189] Provided there is sufficient safety margin, try to solve the problem with a small number of topology operations;
[0190] Only when the first two conditions cannot be met is load reduction forced, and significant penalties imposed.
[0191] Combined with hard safety constraints in cost, this design allows intelligent agents to find strategies with "minimum load reduction and minimum actions" without "crossing the safety red line".
[0192] In the CMDP framework, the reward function Responsible for guiding policy preferences, cost function This part is responsible for defining the safety boundaries. It determines the agent's regulatory preferences and safety boundaries.
[0193] The external cost function is used to define the inviolable safety red line. It is the sole carrier of safety and risk, and the calculation logic is as follows:
[0194] The cost of power flow failure is determined by imposing the maximum penalty if the power flow calculation diverges or any node voltage exceeds the limit.
[0195] ;
[0196] Costs due to hotspot temperature exceeding limits, if any main transformer Added fixed penalties:
[0197] ;
[0198] Emergency response time is insufficient to cover costs, if Superimposed linear penalty:
[0199] ;
[0200] In the formula, One scheduling cycle;
[0201] Failure rate risk cost, if the average failure rate Exceeding the warning value Adding normalized risk costs:
[0202] ;
[0203] In the formula, This represents the average failure rate. This is a warning value; For extreme anchor points;
[0204] The warning value Based on the basic failure rate of the line It is deduced that 1.5 times the basic failure rate corresponds to the warning threshold for line risk just deviating from the normal level; the extreme anchor point The corresponding line operating temperature has reached the maximum allowable temperature. The failure rate at that time; the average failure rate The arithmetic mean of the instantaneous failure rates of the entire network is calculated using the following formula:
[0205]
[0206] in This represents the total number of distribution network lines. For the first The line at the time Instantaneous failure rate.
[0207] The final external cost is calculated as follows:
[0208] .
[0209] All costs are normalized to the interval [0,5] for training stability; the external cost design is intended to mark unacceptable dangerous states (power flow collapse, hotspot overrun, insufficient emergency response time), and at the same time, a three-stage risk step is constructed through failure rate anchors: safe zone - warning zone - danger zone, so that RL can perceive that the cost increases rapidly when approaching danger, and thus learn to stay away from these areas.
[0210] S4. Construct a secure reinforcement learning agent that calculates intrinsic cost based on historical high-cost states and corrects the total cost by combining extrinsic costs. To address the issue that reinforcement learning may underestimate the costs caused by the aforementioned complex physical constraints during the exploration process, this invention designs a memory-driven intrinsic cost correction module.
[0211] High-cost sample memory; the agent maintains a memory. High external costs are incurred in storing historical interactions. ) state embedding vector To reduce dimensionality while preserving features, a random projection layer is used to transform the original high-dimensional state. The mapping is a low-dimensional vector to record the states that trigger high costs (such as exceeding temperature limits or voltage limits) during training;
[0212] Internal cost calculation, for the current state Calculate its embedding vector With memory bank middle The sum of the kernel similarities of the nearest neighbor samples is used as a "pseudo-count" of the state's access to dangerous regions. Intrinsic cost. Defined as:
[0213] ;
[0214] In the formula, For kernel functions; It is the nearest neighbor set; For embedding vectors; This is the state embedding vector; the cost intuitively reflects the similarity between the current state and historical dangerous states. If the current state is highly similar to the fault state in memory, the intrinsic cost will increase significantly.
[0215] Total cost adjustment adds internal costs to external costs to form the adjusted total cost:
[0216] ;
[0217] In the formula, This is the adaptive balancing coefficient. The total cost is used to correct estimation biases in the value function, preventing the agent from taking risky actions due to underestimating the potential costs of high-risk areas.
[0218] During the strategy update phase, the optimization objective is to maximize cumulative rewards. While satisfying the total cost constraint In this way, even before the agent actually triggers external penalties, the intrinsic cost will provide an early warning signal based on the similarity to historical dangerous states, correcting the underestimation bias of the value function, thereby proactively avoiding potential N-1 failure deterioration paths in action selection.
[0219] S5. Train a safety reinforcement learning agent under the N-1 fault scenario, so that it learns to prioritize emergency control strategies such as topology transfer and minimizing load shedding under the premise of meeting safety constraints; use deep reinforcement learning algorithm for multi-round training so that the strategy gradually learns to prioritize topology transfer and reduce load shedding under the premise of meeting safety constraints.
[0220] S6. Deploy the trained security reinforcement learning agent in the power distribution network control system to read the power grid status in real time and generate switching operation and load shedding commands.
[0221] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solution of the present invention, or modify it into equivalent embodiments, without departing from the scope of the present invention. Therefore, any modifications, equivalent changes, and alterations made to the above embodiments based on the technology of the present invention without departing from the scope of the present invention are within the protection scope of the present invention.
Claims
1. A method for intelligent control of power distribution network safety that takes into account the healthy operating status of equipment, characterized in that: The method includes the following steps: S1. Construct a digital simulation environment for the distribution network based on the real power grid topology; S2. Establish an equipment health status assessment model, using the transformer's short-term emergency load capacity and the equipment's instantaneous failure rate as dual physical indicators to quantify the equipment's health status and form safety constraints. S3. Model the emergency control process of N-1 fault in the distribution network as a constrained Markov decision process (CMDP), and define the state space, mixed action space, hierarchical reward function and external cost function. S4. Construct a secure reinforcement learning agent that calculates the intrinsic cost based on historical high-cost states and corrects the total cost by combining the extrinsic cost. S5. Train a safety reinforcement learning agent under the N-1 fault scenario, so that it learns to prioritize emergency control strategies such as topology transfer and minimum load shedding under the premise of meeting safety constraints. S6. Deploy the trained security reinforcement learning agent in the power distribution network control system for online emergency control. 2.The power grid security intelligent control method considering the health state of equipment according to claim 1, wherein: The calculation of the transformer's short-term emergency load capacity in S2 is as follows: Calculate the winding hot spot temperature: ; wherein, is the current load ratio; is the ambient temperature; is the loss ratio; is the winding factor; is the temperature rise of the top oil temperature over the ambient temperature at the rated loss; is the temperature rise of the hot spot temperature over the top oil temperature at the rated current; Based on the current winding hot spot temperature, calculate the remaining time the transformer can maintain operation until it reaches the upper limit of insulation withstand capability: ; wherein is the winding equivalent thermal time constant; is the transformer is able to maintain operation up to the upper limit temperature of the insulation resistance; Minimum emergency response capability of the computing system: ; When Less than the current dispatch period is determined to violate the short-term emergency load capability safety constraint. 3.The power grid security intelligent control method considering the health state of equipment according to claim 1, wherein: The quantification of the instantaneous failure rate of the equipment in S2 includes: Transformer failure rate quantification employs a Weibull + Arrhenius joint model to characterize the effect of temperature on the failure rate. Convert the winding hot spot temperature to Kelvin temperature: ; Calculate temperature-related lifetime parameters: ; wherein is a coefficient related to the nominal lifetime; is an activation energy related parameter; Assuming the fault time follows a Weibull distribution, the instantaneous fault rate of the transformer is calculated as follows: ; wherein is a shape parameter; is an equivalent run time; To quantify the line failure rate, the current carrying capacity correction factor is first calculated based on the ambient temperature and the reference temperature. ; wherein is the ambient temperature; is the base temperature; is the flow rate correction factor; is the maximum allowable line temperature; The line operating temperature is approximated as: ; In the formula, is the rated current of the line; is the rated current of the line; is the real-time current obtained by power flow calculation; The instantaneous fault rate of the line is calculated as follows: ; In the formula, is the base failure rate; .
4. The power distribution network safety intelligent control method considering device health operation state according to claim 1, characterized in that: The state space is as follows: ; ; ; ; ; ; ; In the formula, is the total number of main transformers; is the main transformer operation state sub-vector; is the equipment failure rate state sub-vector; is the electrical quantity state sub-vector; is the operation state sub-vector; is the load state sub-vector; is the environmental quantity measurement sub-vector; is the first real-time load rate of the main transformer; is the winding hot spot temperature; is the remaining short-term emergency tolerance time; is the instantaneous failure rate of the main transformer at its hot spot temperature; is the instantaneous failure rate of the line under the current or temperature rise condition; is the voltage of each bus; is the current of each line; is the real-time on-off state of the tie switch; is the on-off state of the feeder section switch; is the real-time active power of each node; is the real-time reactive power of each node; is the maximum allowable cuttable load ratio of each node; is the current environmental temperature; is the time point corresponding to the load daily curve or temperature curve.
5. The power distribution network safety intelligent control method considering the health state of equipment according to claim 1, characterized in that: The hybrid action space is defined as: The topology transfer action is used to control the opening and closing of critical tie switches. The output of each tie switch action is a continuous value: ; The final action is obtained through threshold mapping: ; Load shedding action is used to control the reduction ratio of the shearable load. Each controllable load node outputs a reduction rate: ; In the formula, is the reduction ratio instruction of the corresponding jth controllable load node in the continuous action vector output by the intelligent agent; is the upper limit of load shedding; The power after reduction is: 。 6. The power distribution network safety intelligent control method considering device health operation state according to claim 2, characterized in that: The external cost function is calculated as follows: The cost of power flow failure is determined by imposing the maximum penalty if the power flow calculation diverges or any node voltage exceeds the limit. ; Hotspot temperature overrun cost, if any main transformer , superimposed fixed penalty: ; Insufficient first aid time cost, if , superimposed linear penalty: ; In the formula, is a primary scheduling period; Failure rate risk cost, if average failure rate Exceeds warning value , superimposed normalized risk cost: ; In the formula, is the average failure rate; is the early warning value; is the limit anchor point; The final external cost is calculated as follows: 。 7. The power distribution network safety intelligent control method considering device health operation state according to claim 1, characterized in that: The hierarchical reward function is as follows: ; ; ; ; ; wherein, is the first aid time improvement reward; is the reward weight; is the updated system minimum remaining short-term first aid tolerance time for the current step; is the system minimum remaining short-term first aid tolerance time for the previous step; is the switch operation penalty reward; is the penalty coefficient; is the load shedding penalty reward; is the penalty coefficient, is the complete rescue success reward; is the real-time load rate of the i-th main transformer; num_ops is the number of switch position changes in the current step; is the total amount of load shedding.
8. The intelligent control method for power distribution network safety considering the healthy operating status of equipment as described in claim 6, characterized in that: The process of calculating intrinsic costs based on historical high-cost conditions and adjusting the total cost by incorporating extrinsic costs includes: Calculate the intrinsic cost: ; In the formula, For kernel functions; It is the nearest neighbor set; For embedding vectors; For state embedding vectors; Total cost correction: ; In the formula, This is the adaptive balance coefficient.