Power distribution network active operation optimization method based on multi-level security domain
Patent Information
- Application Number
- CN202610614179.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-07
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]传统的配电网优化方法多依赖于预设的固定运行预案,当实际运行状态与预案设定条件出现偏差时,预案往往无法有效适配,导致控制效果下降甚至引发安全风险,整体适应性较差
[0048] The aforementioned active operation optimization method for distribution networks based on multi-level safety domains first constructs a multi-level safety domain model of the distribution network based on historical operating data and fault simulation data. This model characterizes the allowable operating range of the distribution network at the feeder, transformer, and substation levels. Based on this model, day-ahead operation optimization is performed to generate a day-ahead operation plan. During intraday operation, the day-ahead plan is corrected for safety based on the multi-level safety domain model. In the event of a real-time fault in the distribution network, the multi-level safety domain model serves as a safety constraint to generate a fault repair strategy. Thus, this scheme effectively overcomes the shortcomings of traditional fixed contingency plans that fail or trigger risks under deviation conditions, significantly improving the adaptive capability and safety of the distribution network in the face of high uncertainty and complex operating environments.
Smart Images

Figure CN122659923A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power grids, and in particular to a method for active operation optimization of distribution networks based on multi-level security domains. Background Technology
[0002] With the large-scale grid connection of distributed photovoltaic, wind power and other new energy sources and the popularization of electric vehicles, the operating environment of modern power distribution networks is becoming increasingly complex. The random fluctuations in new energy output and the spatiotemporal uncertainty of electric vehicle charging loads bring unprecedented challenges to the safe and economical operation of power distribution networks.
[0003] Traditional power distribution network optimization methods often rely on pre-set fixed operation plans. When the actual operating conditions deviate from the conditions set in the plan, the plan often cannot be effectively adapted, resulting in a decline in control effectiveness or even safety risks, and the overall adaptability is poor. Summary of the Invention
[0004] Therefore, it is necessary to provide a multi-level security domain-based active operation optimization method for distribution networks that can improve the adaptability of distribution network operation control and address the aforementioned technical problems.
[0005] Firstly, this application provides a method for active operation optimization of distribution networks based on multi-level security domains, including:
[0006] Based on historical operation data and fault simulation data of the distribution network, a multi-level security domain model of the distribution network is constructed. The multi-level security domain model is used to characterize the allowable operating range of the distribution network at the feeder level, transformer level and substation level, respectively.
[0007] Based on a multi-level security domain model, the day-ahead operation of the distribution network is optimized to generate the day-ahead operation plan of the distribution network.
[0008] During the intraday operation phase, safety corrections are performed on the daily operation plan based on a multi-level security domain model;
[0009] In the event of a real-time fault in the distribution network, a fault repair strategy for the distribution network is generated using a multi-level security domain model as a security constraint.
[0010] In one embodiment, based on a multi-level security domain model, day-ahead operation optimization of the distribution network is performed to generate a day-ahead operation plan for the distribution network, including:
[0011] At the preset day-ahead scheduling time, based on historical operating data, the first target operating data of the distribution network at the first future time is predicted;
[0012] The first target's operational data is input into the decision model pre-constructed based on the spatiotemporal graph attention network, so that the decision model can analyze the first target's operational data with a multi-level security domain model as a security constraint, and output the opening and closing commands of the switching equipment contained in the feeder level, transformer level, and substation level.
[0013] Under the network topology determined by each opening and closing command, a deep reinforcement learning agent is used to generate the day-ahead operation plan of the power distribution equipment in the power distribution network with a multi-level security domain model as the security constraint.
[0014] The power distribution equipment is operated in accordance with the operation control instructions that match the daily operation plan.
[0015] In one embodiment, based on historical operating data, the first target operating data of the distribution network at a first future moment is predicted, including:
[0016] Outlier identification is performed on historical operational data to obtain intermediate operational data after outlier removal;
[0017] The intermediate running data is denoised to obtain the denoised running data.
[0018] The noise-reduced operating data is then subjected to dimensionality reduction processing to obtain the dimensionality-reduced operating data.
[0019] The dimensionality-reduced operating data is input into a pre-trained prediction model so that the prediction model can predict the first target operating data of the distribution network at the first future moment by analyzing the dimensionality-reduced operating data; wherein, the prediction model is constructed based on a long short-term memory network.
[0020] In one embodiment, outlier identification is performed on historical operational data to obtain intermediate operational data after removing outliers, including:
[0021] Clustering historical operational data yields multiple data clusters;
[0022] The random forest algorithm is used to identify outliers in each data cluster, and intermediate running data after removing outliers is obtained.
[0023] In one embodiment, during the intraday operation phase, a safety correction is performed on the daily operation plan based on a multi-level security domain model, including:
[0024] Under the preset intraday scheduling time, based on historical operating data, the second target operating data of the distribution network at the second future time is predicted; wherein, the second future time is earlier than the first future time;
[0025] Based on the second target operation data and the multi-level security domain model, a risk assessment of the distribution network under preset faults is conducted to obtain the risk assessment results.
[0026] If the risk assessment results indicate that a preset fault would cause the distribution network's operating state to exceed the allowable operating range, a pre-trained dual-agent reinforcement learning model is invoked.
[0027] A dual-agent reinforcement learning model is used to perform safety corrections on the day-to-day operation plan.
[0028] In one embodiment, the dual-agent reinforcement learning model includes a first agent reinforcement learning model and a second agent reinforcement learning model; the dual-agent reinforcement learning model is used to perform safety corrections on the day-to-day operation plan, including:
[0029] The first agent reinforcement learning model determines the direction of network structure adjustment to bring the operating state back to the allowable operating range based on the deviation information between the operating state and the allowable operating range.
[0030] Using the second agent reinforcement learning model, the set of key switching devices associated with preset faults is selected from the various switching devices of the distribution network according to the network structure adjustment direction, and the opening and closing adjustment instructions for each candidate switching device in the set of key switching devices are generated.
[0031] Based on the switching adjustment commands, the daytime operation plan is corrected for safety, and the candidate switching equipment is controlled according to the corrected operation plan.
[0032] In one embodiment, in the event of a real-time fault in the distribution network, a fault repair strategy for the distribution network is generated using a multi-level security domain model as a security constraint, including:
[0033] When a real-time fault occurs in the distribution network, and the current operating state of the distribution network under the actual fault exceeds the allowable operating range, a multi-level intelligent agent is invoked.
[0034] By using multi-level intelligent agents and a multi-level security domain model as security constraints, fault repair strategies for actual faults are generated.
[0035] The power distribution network is operated and controlled according to the repair strategy.
[0036] In one embodiment, the multi-level intelligent agent includes a feeder-level intelligent agent corresponding to the feeder level, a transformer-level intelligent agent corresponding to the transformer level, and a substation-level intelligent agent corresponding to the substation level.
[0037] In one embodiment, the distribution network active operation optimization method based on multi-level security domains further includes:
[0038] The power distribution network is divided into multiple power grid areas;
[0039] Historical operational data includes regional operational data corresponding to each power grid area; based on the historical operational data of the distribution network and fault simulation data, a multi-level security domain model of the distribution network is constructed, including:
[0040] For each power grid region, based on the corresponding regional operation data and fault simulation data, network topology reconstruction simulation under fault conditions is performed on the power grid region to obtain regional reconstruction simulation results.
[0041] The simulation results of each region are summarized to obtain the topology reconstruction simulation results;
[0042] Based on the simulation results of topology reconfiguration, a multi-level security domain model of the distribution network is constructed.
[0043] In one embodiment, the distribution network is divided into multiple power grid areas, including:
[0044] Obtain the power flow calculation results of the distribution network;
[0045] Based on the power flow calculation results, a sensitivity matrix is constructed to reflect the electrical coupling strength between nodes in the distribution network.
[0046] The electrical distance between each node is determined based on the sensitivity matrix;
[0047] Based on electrical distance, the nodes are clustered to obtain multiple power grid regions.
[0048] The aforementioned active operation optimization method for distribution networks based on multi-level safety domains first constructs a multi-level safety domain model of the distribution network based on historical operating data and fault simulation data. This model characterizes the allowable operating range of the distribution network at the feeder, transformer, and substation levels. Based on this model, day-ahead operation optimization is performed to generate a day-ahead operation plan. During intraday operation, the day-ahead plan is corrected for safety based on the multi-level safety domain model. In the event of a real-time fault in the distribution network, the multi-level safety domain model serves as a safety constraint to generate a fault repair strategy. Thus, this scheme effectively overcomes the shortcomings of traditional fixed contingency plans that fail or trigger risks under deviation conditions, significantly improving the adaptive capability and safety of the distribution network in the face of high uncertainty and complex operating environments. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0050] Figure 1 This is an application environment diagram of a distribution network active operation optimization method based on multi-level security domains in one embodiment;
[0051] Figure 2 This is a flowchart illustrating a distribution network active operation optimization method based on multi-level security domains in one embodiment.
[0052] Figure 3 This is a flowchart illustrating a distribution network active operation optimization method based on multi-level security domains in a specific embodiment.
[0053] Figure 4 This is a schematic diagram of multi-level reconstruction based on a spatiotemporal graph attention network in a specific embodiment;
[0054] Figure 5 This is a schematic diagram of the dual-agent training process for intraday preventive reconstruction in a specific embodiment.
[0055] Figure 6 This is a schematic diagram of a collaborative hierarchical multi-agent reinforcement learning architecture based on real-time emergency control in a specific embodiment. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0057] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0058] This application provides a method for active operation optimization of distribution networks based on multi-level security domains, which can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, and IoT devices. Server 104 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0059] The active operation optimization method for distribution networks based on multi-level security domains provided in this application embodiment can be derived from... Figure 1 The terminal 102 or server 104 can execute the operation independently, or they can interact. The following explanation uses server 104 executing independently as an example. Specifically, server 104 constructs a multi-level security domain model of the distribution network based on historical operating data and fault simulation data. This multi-level security domain model characterizes the allowable operating range of the distribution network at the feeder, transformer, and substation levels. Based on the multi-level security domain model, server 104 performs day-ahead operation optimization and generates a day-ahead operation plan for the distribution network. During the intraday operation phase, server 104 performs safety corrections on the day-ahead operation plan based on the multi-level security domain model. In the event of a real-time fault in the distribution network, server 104 uses the multi-level security domain model as a safety constraint to generate a fault repair strategy for the distribution network.
[0060] In one exemplary embodiment, such as Figure 2 As shown, a method for active operation optimization of distribution networks based on multi-level security domains is provided, which is then applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0061] Step S202: Based on the historical operation data and fault simulation data of the distribution network, construct a multi-level security domain model of the distribution network; wherein, the multi-level security domain model is used to characterize the allowable operating range of the distribution network at the feeder level, transformer level and substation level respectively.
[0062] Historical operational data of the distribution network refers to the operational status data of the distribution network recorded over a period of time (e.g., one year or longer), including but not limited to at least one of the following: voltage amplitude and phase angle of each node (bus), active and reactive power of each branch (line, transformer), actual output of each distributed power source (photovoltaic, wind power), power consumption of each load point, charging power of electric vehicle charging facilities, state of charge of energy storage systems, and switching status of each switchgear. This data is typically collected at regular time intervals (e.g., every 15 minutes or hour) to form time-series data.
[0063] Fault simulation data refers to a collection of fault scenarios generated through computer simulation, simulating faults in different components (such as lines, transformers, and buses) in a power distribution network. Each fault scenario includes at least one of the following: fault type, fault location, operating point characteristics before the fault, and other parameters that may be required. The fault type can include N-1 faults and / or other faults. The fault type can be an N-1 fault, meaning any single device fails to operate. Other types of faults can also be simulated, and this embodiment does not impose any limitations on this. The fault location can be, for example, at least one of a line, a node, or a transformer. The operating point characteristics before the fault can be, for example, at least one of the following: source-load power distribution before the fault, or the status of switching equipment. Other parameters can be, for example, the duration of the fault.
[0064] For example, firstly, the server acquires historical operating data from a power distribution automation system or a data acquisition and monitoring system. Then, a Conditional Generative Adversarial Network (CGAN) is used to learn from the historical operating data, generating a large number of representative typical scenarios of anticipated operating points. A CGAN is a deep learning model consisting of a generator and a discriminator. The generator is responsible for producing realistic operating point data based on input conditions (e.g., season, weather type, load level range), while the discriminator is responsible for distinguishing the generated data from real historical data. Through alternating training, the generator learns to generate marginal scenarios that are consistent with the distribution of historical data but exceed the range of historical samples. Each generated scenario is accompanied by a probability weight, representing the frequency at which the scenario may occur in real operation; the weight is output by the CGAN or obtained through subsequent statistics. Next, for each generated typical scenario, corresponding fault simulation data is constructed.
[0065] In some embodiments, the server performs network topology reconstruction simulation of the distribution network under fault conditions based on historical operating data and fault simulation data, and obtains the topology reconstruction simulation results.
[0066] Network topology reconfiguration refers to altering the power flow path by changing the closed or open state of at least one controllable switch in the distribution network, such as sectionalizing switches or tie switches, thereby isolating faulty sections, transferring loads from non-faulty sections, reducing line losses, or improving voltage distribution. Under fault conditions, the main purpose of network topology reconfiguration is to quickly restore power supply to non-faulty sections. Network topology reconfiguration simulation involves simulating the gradual operation of switches according to certain reconfiguration strategies (such as minimum load shedding, maximum load restoration, minimum number of switch operations, etc.) for each preset fault scenario, calculating the power flow distribution after each operation, and ultimately obtaining a feasible reconfiguration scheme that restores the distribution network to stable operation, along with the corresponding electrical quantity results—the topology reconfiguration simulation results.
[0067] For example, for each fault scenario, the server performs pre-fault power flow calculations and simulates the fault occurrence. Specifically, using power system simulation software, based on the fault simulation data, it simulates disconnecting a device from the network and then recalculates the power flow to obtain the post-fault network state. At this point, the faulty section loses power, and other areas may experience overload. To facilitate protection actions, the server further performs topology reconfiguration simulations, i.e., it initiates a reconfiguration optimization algorithm. The algorithm's input includes at least one of the following: the post-fault network topology, the load priority of each node, and the availability status of each switch. The output is a series of time-sequential switching operation instructions. During algorithm execution, each attempt to close or open a switch triggers a power flow calculation to check for voltage overload, line overload, or loop formation. If a new overload or loop occurs, the operation is rolled back, and other switches are tried. The algorithm iterates continuously until any termination condition is met, such as all non-faulty power-loss areas have been restored to power, all electrical quantities meet safety constraints, all possible switching actions have been adjusted and no further improvements are possible, or the preset maximum number of switching operations has been reached. After successful reconstruction, the server records at least one piece of information, including the final network topology, voltage of each node, current of each branch, restored load power, and operation sequence. If reconstruction fails, it may record the load shedding amount and the reason for failure to restore. Furthermore, multi-level safety labels need to be generated for this simulation result, that is, based on whether the final voltage and current are within the allowable range, the feeder level, transformer level, and substation level are marked as safe. All information (fault scenario identification, reconstructed topology, multi-level safety labels, electrical quantity results, etc.) constitutes the construction stage of the multi-level safety domain model for the topology reconstruction simulation result.
[0068] A multi-level safety domain model is a data-driven classification model, such as a support vector machine, random forest, shallow neural network, or gradient boosting tree. Its input is a feature vector of a distribution network's operating state (mainly including active / reactive power at each node, the current on / off state of each switch, etc.), and its output is a safety or insecurity judgment result for the feeder level, transformer level, and substation level, respectively, or a more refined safety margin value. The safety domain refers to the set of all operating states that satisfy safety constraints in a multi-dimensional space. Due to the large scale and nonlinearity of distribution networks, a data-driven method is used to learn an approximate discrimination boundary from simulation results.
[0069] The permissible operating range at the feeder level is for each individual distribution feeder and is defined by at least one of the following constraints: the voltage amplitude at the feeder terminal node is not lower than the lower limit and not higher than the upper limit; the effective value of the current in any branch of the feeder does not exceed the long-term thermal stability current carrying capacity of the conductor; and the total injected power of the distributed power sources connected to the feeder does not exceed the transmission capacity of the feeder transformer or tie line. When all constraints are met simultaneously, the operating point of the feeder is considered to be within the safe region.
[0070] The permissible operating range for each transformer level is for each distribution transformer and includes at least one of the following: the transformer load rate does not exceed a preset safety threshold; the low-voltage bus voltage of the transformer is within the acceptable range; if multiple transformers are operating in parallel and connected through a bus tie switch, the difference in load rate between the transformers should not exceed a set value. Exceeding these limits is considered unsafe.
[0071] The permissible operating range at the substation level applies to the entire substation and may include multiple main transformers, multiple busbar sections, and inter-station tie lines. The permissible operating range includes at least one of the following: the voltage of critical busbars within the substation remains stable within the permissible range; the active power transmitted by inter-station tie lines does not exceed their thermal stability limits; and when one main transformer fails, the remaining main transformers are able to handle all important loads.
[0072] For example, the core idea of constructing a multi-level safety domain model is to treat the large amount of previously generated topology reconfiguration simulation results as training data and use a supervised learning algorithm to train a classifier that can directly predict safety from operating point features. Specifically, first, input feature vectors and corresponding output labels are extracted from the topology reconfiguration simulation results. Then, a multi-output classification model is selected, i.e., a model that simultaneously predicts labels at three levels, and the model is trained. The goal of the model training is to minimize the classification error. The final output is a multi-level safety domain model of the distribution network.
[0073] Step S204: Based on the multi-level security domain model, perform day-ahead operation optimization on the distribution network to generate the day-ahead operation plan for the distribution network.
[0074] For example, on a current planning timescale, the server generates a baseline operation control scheme for a future full operating day based on historical and forecast data. In this process, a multi-level security domain model is embedded as a security constraint in the generation logic of the scheme.
[0075] Step S206: During the intraday operation phase, the daily operation plan is adjusted for safety based on the multi-level safety domain model.
[0076] For example, on a rolling correction timescale, the server performs a rolling security check on the day-ahead plan based on the updated forecast data using a multi-level security domain model, and makes local adjustments for the predicted risks.
[0077] Step S208: In the event of a real-time fault in the distribution network, a fault repair strategy for the distribution network is generated using a multi-level security domain model as a security constraint.
[0078] For example, on a real-time response timescale, when an actual fault is detected in the distribution network, the server immediately obtains the post-fault operational data and calls a multi-level security domain model to determine whether the current operational state has exceeded the allowable range. If it is determined to be out of bounds, a multi-level coordination mechanism is triggered, and fault isolation and load restoration are achieved through the collaborative operation between the corresponding agents at each level.
[0079] By coordinating the three time scales mentioned above and always using a multi-level security domain model as a unified security constraint, it can be ensured that the real-time operating data of the distribution network is actively constrained within the allowable operating range under any operating scenario, thereby significantly improving the safety and adaptability of the distribution network operation.
[0080] In this embodiment, a multi-level security domain model of the distribution network is first constructed based on historical operating data and fault simulation data. This model is not a static, fixed boundary, but rather a dynamic security envelope that integrates topology reconfiguration under various fault scenarios, thus more realistically reflecting the actual security constraints of the distribution network under different operating modes. Furthermore, this multi-level security domain model characterizes the allowable operating range of the distribution network at the feeder, transformer, and substation levels. Based on the multi-level security domain model, day-ahead operation optimization is performed on the distribution network to generate a day-ahead operation plan. During intraday operation, the day-ahead operation plan is corrected for safety based on the multi-level security domain model. In the event of a real-time fault in the distribution network, a fault repair strategy is generated using the multi-level security domain model as a safety constraint. Thus, this embodiment effectively overcomes the shortcomings of traditional fixed contingency plans that fail or cause risks under deviation conditions, significantly improving the adaptability and safety of the distribution network in the face of high uncertainty and complex operating environments.
[0081] In an exemplary embodiment, based on a multi-level security domain model, day-ahead operation optimization is performed on the distribution network to generate a day-ahead operation plan for the distribution network. This includes: predicting the first target operation data of the distribution network at a first future time based on historical operation data at a preset day-ahead scheduling time; inputting the first target operation data into a decision model pre-constructed based on a spatiotemporal graph attention network, so that the decision model analyzes the first target operation data with the multi-level security domain model as a security constraint, and outputs the opening and closing instructions of the switching equipment contained in the feeder level, transformer level, and substation level; under the network topology determined by each opening and closing instruction, generating a day-ahead operation plan for the distribution equipment in the distribution network through a deep reinforcement learning agent with the multi-level security domain model as a security constraint; and performing operation control on the distribution equipment according to the operation control instructions matching the day-ahead operation plan.
[0082] The preset day-ahead scheduling time refers to a fixed point in time each day when the server uses historical operational data to begin predicting and planning the distribution network's operational status for the next complete operating day. This time is typically set after the electricity market closes and before the day-ahead plan is released. The first future time refers to any time segment within the next 24 hours, usually a discrete time point, such as a segment every 15 minutes or hour. The set of first future times covers the entire operating day. The first target operating data refers to at least one of the following at the predicted first future time: distributed generation output data, conventional load data, electric vehicle charging load data, etc., for each node or region of the distribution network.
[0083] The Spatiotemporal Graph Attention Network (ST-GAT-GRU) is a deep learning architecture that integrates graph neural networks and recurrent neural networks. The Graph Attention Network (GAT) extracts spatial dependencies in the distribution network topology, where each node's features aggregate information from neighboring nodes through an attention mechanism. The Gated Recurrent Unit (GRU) extracts evolutionary patterns over time. ST-GAT-GRU takes the physical graph structure of the distribution network, node dynamic features, and node static features as input, and outputs the closing / opening probabilities of each switch at multiple future time points. The decision model is the ST-GAT-GRU model trained offline. This model, through supervised learning, masters the mapping relationship from the input of the distribution network's physical graph structure, node dynamic features, and node static features to the optimal switch state sequence. Switchgear specifically includes four levels: substation tie switches (switches connecting different substations or different busbars within a substation), transformer tie switches (switches connecting different main transformers within the same substation), feeder tie switches (switches connecting different feeders), and branch sectionalizing switches (switches on the same feeder used for sectionalized operation or fault isolation).
[0084] In this embodiment, the deep reinforcement learning agent primarily employs a Deep Q-Network (DQN), which learns the optimal strategy through interaction with the environment. At each time step, the agent observes the state of the power distribution network, selects an action, and the environment returns a reward and updates the state. Power distribution equipment includes distributed power sources (photovoltaics, wind power), energy storage systems, and electric vehicle charging facilities, etc. The active power output of these devices can be continuously adjusted within a certain range. The day-ahead operating plan refers to the combination of power commands output by the deep reinforcement learning agent, such as the active power output of each photovoltaic inverter, the charging and discharging power of energy storage, and the charging power of electric vehicles. These commands are discretized values. Operating control commands are the conversion of the power commands output by the agent into actually executable analog or digital quantities. For example, if the agent selects "50%", the corresponding inverter's active power output setting is 50% of its rated power.
[0085] For example, at a preset day-ahead scheduling time, the server predicts the first target operating data of the distribution network at a first future time based on historical operating data. The prediction model can employ a Long Short-Term Memory (LSTM) network. LSTM is a special type of recurrent neural network whose core unit includes a forget gate, input gate, output gate, and cell state, effectively capturing long-term dependencies. For photovoltaic and wind power output prediction, this embodiment uses a standard LSTM structure, incorporating a Dropout mechanism during training to prevent overfitting. The Dropout mechanism randomly selects a group of neurons to ignore during the training phase, and a dynamic learning rate adjustment strategy is adopted. For example, when the validation loss does not decrease for five consecutive cycles, the learning rate is halved. For electric vehicle charging load prediction, since charging behavior is affected by factors such as historical charging patterns, time, and holidays, this embodiment introduces a multi-head self-attention mechanism on top of LSTM, enabling the model to automatically focus on historical time points that are particularly important for the current prediction, such as the same time period of the previous day or the same workday of the previous week. Finally, the model outputs the photovoltaic power output curve, wind power output curve, conventional load curve, and electric vehicle charging and discharging load curve for each node or region in the next 24 hours.
[0086] Next, the server calls the ST-GAT-GRU model that has been trained offline. The model input consists of three parts, namely... .in, This is a topology diagram of the power distribution network. V is the set of nodes, and E is the set of edges; dynamic features N is the number of nodes, and T is the length of the time series, such as 24 hours. Dynamic features for each node at each time step, including photovoltaic / wind power output forecasting, load forecasting, and electric vehicle charging power forecasting; static features. , For each node, there is a static feature dimension, such as at least one of node type, rated voltage, line impedance, etc. For a given input, the model outputs a structured tensor. Its mathematical form is:
[0087]
[0088] Where ∏ represents the Cartesian product (i.e., the combination of all elements). m represents the level index, ranging from 1 to 4. m=1 represents the reconfiguration of the substation tie switch, m=2 represents the reconfiguration of the transformer tie switch, m=3 represents the reconfiguration of the feeder tie switch, and m=4 represents the reconfiguration of the branch section switch. k represents the switch index under each level m, ranging from 1 to s, where s is the total number of switches at level m. d represents the time section index, ranging from 1 to T, where T is the length of the time series. It is a switch status variable, taking the value 0 or 1. A value of 1 indicates that the k-th switch at level m should close at time d. A value of 0 indicates that the k-th switch at level m should be disconnected at time section d.
[0089] In the model output, to distinguish between switches at different physical levels, the following piecewise function can be used to classify all switch state variables:
[0090]
[0091] in, Indicates the substation interconnection switch. Indicates the transformer interconnection switch. Indicates the feeder connection switch. This indicates a branch circuit sectionalizing switch.
[0092] In some embodiments, the decision model can be formalized as a multi-label binary classification model:
[0093]
[0094]
[0095] Where X represents the dynamic feature input and Z represents the static feature input. The output is a sequence of 0s and 1s, outputting the state of four switch types for each time segment d. This represents the predicted or maximum available photovoltaic active power output of the k-th node at the d-th time segment. This represents the predicted electric vehicle charging load at the k-th node at the d-th time segment. This represents the predicted conventional active power load of the k-th node at the d-th time segment. This represents the predicted value of the conventional reactive load at the k-th node at the d-th time segment. This represents the active power injected into the substation or root node at the k-th node at the d-th time segment.
[0096] In some embodiments, during training, the ST-GAT-GRU model defines the total loss function as follows to simultaneously ensure classification accuracy and topological physical feasibility:
[0097]
[0098]
[0099]
[0100]
[0101] in, Focal Loss, or cross-entropy loss function, is used to address the problem of imbalanced switch state categories. This is a loop penalty term that forces the model to output a loop-free radial topology. This is a penalty term for connectivity issues to prevent islanding or abnormal parallel connection of multiple power sources. and This is a hyperparameter used to balance various losses. S is the total number of controllable switches in the distribution network. The label represents the actual state of switch s at time t. A value of 1 indicates that the switch is closed, and a value of 0 indicates that the switch is open. It is the probability predicted by the model that switch s will close at time t, and its value is [0, 1]. The hyperparameters, which are used to balance positive and negative samples, are set between 0 and 1. It is a hyperparameter greater than or equal to 0, when When, Focal Loss degenerates into standard cross-entropy; when When the predicted probability is close to the true label, the factor or This will significantly reduce its contribution to the loss, allowing the model to focus more on samples that are difficult to classify. This represents the power grid topology constructed based on the switching states predicted by the model. Used for calculation graph The number of basic rings, i.e., the number of independent rings. In normal operation of a distribution network, a radial topology (loop-free) is required; therefore, the number of rings should be zero. Any ring can cause problems such as circulating current and protection malfunctions, and must be penalized. Used for calculation graph The number of connected components in a graph. A connected component is a group of nodes that are connected to each other in the graph. This indicates the number of power sources in the distribution network, such as the number of substations. Each power source should independently serve a specific area; normally, each connected component should contain exactly one power source, and there should be no islanded loads. If This means that each power source supplies power to an independent connected area, with no isolated islands or abnormal parallel connections; this is the ideal safe state. If This indicates the presence of an islanded load without a power source, which is not allowed, and the penalty is positive. If... This indicates that multiple power supplies are connected in a loop, which usually causes circulating current problems, but this situation has been penalized by the loop. Because of the coverage, this item only penalizes isolated cases.
[0102] In addition, during the training of the ST-GAT-GRU model, the following three metrics are needed to quantify the model's predictive performance. These metrics are all calculated based on the classification confusion matrix, with values ranging from [0, 1]. A larger value indicates better model performance.
[0103]
[0104]
[0105]
[0106] in, This represents the accuracy value, indicating the proportion of samples correctly predicted by the model out of the total number of samples. This represents the number of true cases, i.e., the number of samples that are actually closed and correctly predicted by the model to be closed. This represents the number of true negative examples, i.e., the number of samples that are actually disconnected but the model correctly predicts as disconnected. This indicates the number of false negatives, i.e., the number of samples that are actually closed but the model incorrectly predicts as open. This indicates the number of false positives, that is, the number of samples that are actually open but the model incorrectly predicts as closed. Precision, also known as accuracy, represents the proportion of samples that are predicted to be closed by the model but are actually closed. Recall, also known as the completeness of the sample, represents the proportion of samples that are correctly predicted by the model out of all samples that are actually closed.
[0107] The distribution network topology is constructed based on the switching commands of each switchgear output by the ST-GAT-GRU model. For a given network topology, the continuous power command optimization problem of distributed source-load is modeled as a Markov decision process and solved using a deep Q-network. The state space S is represented as:
[0108]
[0109] in, This represents the network topology state at the current time t, such as the 0 / 1 state vector of each switch, which originates from the opening and closing commands output in the first stage. This represents the predicted maximum available output of photovoltaic power generation at each node at time t, derived from LSTM prediction. This represents the predicted demand for electric vehicle charging load at each node at time t. This represents the predicted conventional load for each node at time t. This represents the state of charge of each energy storage system at time t, with values ranging from 0 to 1, where 1 indicates full charge. This represents the voltage magnitude of each node at time t, obtained through state estimation or power flow calculation. Let i and j represent the current amplitude of each branch at time t, where i and j are the nodes constituting the branches. These state variables constitute the input feature vector of the deep Q-network.
[0110] Action space A is represented as:
[0111]
[0112] in, This represents the power output of each distributed power source (photovoltaic, wind power) at time t. This represents the charging and discharging power of each energy storage system at time t. This represents the charging power of each electric vehicle charging facility at time t. d is the discretized action number. Since deep Q-networks typically output the Q-values of discrete actions and cannot directly handle continuous actions, each continuous quantity is discretized into a finite number of levels. The combinations of all device levels constitute the discrete action space.
[0113] Action space Represented as:
[0114]
[0115] in, It indicates the current state. This represents the action chosen by the agent, i.e., a discrete combination of power commands. f(·) is the state transition function, which is determined by the physical model (such as power flow calculation) in the distribution network.
[0116] The penalty function C is expressed as:
[0117]
[0118]
[0119]
[0120]
[0121] in, This represents the actual photovoltaic power absorbed. This represents the set of all nodes in the distribution network that have installed photovoltaic power generation units. This represents the actual grid-connected power of the photovoltaic unit at node j at time t after the agent makes its decision. This indicates the actual charging power of the electric vehicle that is met. This represents the set of all nodes in the power distribution network that are connected to electric vehicle charging stations. This represents the actual charging power consumed by the charging pile cluster at node j at time t after the agent makes its decision. This represents the cost of network losses. This indicates the cost per unit of loss. This represents the resistance of branch (i, j). This represents the effective value of the current in branch (i, j).
[0122] This is a penalty for exceeding safety domain limits, ensuring that the operating point remains within the safe operating range defined by the multi-level safety domain model. The calculation method involves setting the voltage at the current operating point... Current ,power , Given a multi-level security domain model as input, the model outputs whether the feeder level, transformer level, and substation level are secure. If all levels are secure, then... Otherwise, a penalty will be imposed according to the degree of violation. , , A positive weighting coefficient is used to balance different objectives.
[0123] After offline training, the trained deep Q-network model is deployed to the server. At each time point, the server inputs the predicted state into the model, and the model outputs the Q-values of each discrete action. The power command corresponding to the largest Q-value is selected as the operating control strategy. Then, these power commands are converted into specific equipment control commands, which are executed to complete the adjustment for that time point. This process can be executed sequentially to form a complete day-ahead power scheduling plan.
[0124] In this embodiment, the entire day-ahead scheduling process is constrained by a multi-level security domain model. First, the optimal switching command is quickly generated using a spatiotemporal graph attention network. Then, deep reinforcement learning is used to finely adjust the equipment power under a fixed topology, ensuring that the operating point is always within the safety boundary, which significantly improves the adaptability of the distribution network operation control.
[0125] In an exemplary embodiment, predicting the first target operating data of the distribution network at a first future moment based on historical operating data includes: identifying outliers in the historical operating data to obtain intermediate operating data after removing outliers; performing noise reduction on the intermediate operating data to obtain denoised operating data; performing dimensionality reduction on the denoised operating data to obtain dimensionality-reduced operating data; and inputting the dimensionality-reduced operating data into a pre-trained prediction model so that the prediction model can predict the first target operating data of the distribution network at a first future moment by analyzing the dimensionality-reduced operating data; wherein the prediction model is constructed based on a long short-term memory network.
[0126] Outliers refer to data points whose measured values deviate significantly from the normal range.
[0127] In some embodiments, outlier identification is performed on historical running data to obtain intermediate running data after removing outliers, including: clustering the historical running data to obtain multiple data clusters; and using a random forest algorithm to identify outliers in each data cluster to obtain intermediate running data after removing outliers.
[0128] For example, in the outlier identification stage, this embodiment employs a two-stage method combining K-Means clustering and random forest. First, the historical data for each time segment is considered a high-dimensional dataset. The K-Means algorithm divides the dataset into K clusters, where the value of K can be determined based on the specific circumstances, for example, K=10. Normal data will form several dense clusters, while outliers often form isolated clusters or are far from all cluster centers. The objective function of K-Means clustering is as follows:
[0129]
[0130] in, This represents the i-th data point. Let the centroid of the j-th cluster be denoted as . This is a membership indicator variable, taking the value 0 or 1. If the i-th data point is assigned to the j-th cluster, then... ;otherwise .
[0131] The allocation rules are as follows:
[0132]
[0133] in, This represents the cluster number to which the i-th data point is assigned, i.e., the cluster index. This represents the square of the Euclidean distance, measuring the degree of difference between the data point and the centroid. For each data point... Calculate it to all cluster centroids The point is assigned to the cluster with the smallest squared Euclidean distance.
[0134] The centroid update rule is as follows:
[0135]
[0136] in, This represents the vector sum of all data points within cluster j. This represents the number of data points in cluster j, i.e., the size of the cluster. After all data points have been allocated, for each cluster j, the mean vector of all data points in it is calculated, and the centroid is updated to this mean.
[0137] Next, a random forest classifier is used to perform secondary discrimination on each cluster. First, a feature is randomly selected from the original dataset, such as the voltage of a node, and a split point is randomly selected within the range of this feature value. The data is divided into two subsets, left and right. The above process is recursively repeated for each subset until the stopping condition is met, that is, all data points are isolated, i.e., each leaf node contains only one data point, or the preset maximum depth of the tree is reached, thus constructing an isolated tree. Then, the path length of each data point in the isolated tree is calculated. The calculation formula is:
[0138]
[0139]
[0140]
[0141] Where e is the number of edges that data point x traverses from the root node to the leaf node in the isolated tree. This is a correction term used to estimate the average additional path length required to continue isolating samples when there are still multiple samples in a leaf node. n is the number of leaf nodes. Let i be the i-th harmonic number, where i is a positive integer, i = n - 1. It is the (n-1)th harmonic number.
[0142] Calculate the anomaly score based on the path length:
[0143] S(x,n)= 2 - E[h(x)] C( n )
[0144] in, E[h(x)] Let S be the expected path length of data point x in the isolated tree. The final result is determined based on the anomaly score: when the anomaly score S approaches 0, it indicates that data point x is normal data; when the anomaly score S approaches 1, it indicates that data point x is anomaly data.
[0145] Furthermore, for the intermediate running data after outlier removal, wavelet transform denoising is performed to remove high-frequency random noise. Specifically, let the noisy intermediate running data sequence be x(t), which is composed of the superposition of real data s(t) and noisy data n(t). Multi-scale wavelet decomposition of x(t) yields a series of wavelet coefficients, which can be expressed as:
[0146]
[0147] in, For wavelet transform operators; This represents the approximation coefficient of the Nth layer, corresponding to the low-frequency part of the data, and representing the long-term trend or main component of the data; This represents the detail coefficient of the Nth layer, corresponding to the high-frequency part of the data. The value of j ranges from 1 to N. The larger the j is, the lower the frequency. Usually, noise is mainly concentrated in the first few detail coefficients.
[0148] To filter out noise, the high-frequency detail coefficients at each level are... Thresholding is performed using the following expression:
[0149]
[0150] in, This represents the k-th detail coefficient value of the j-th layer, where j∈[1,N]. This is the new detail coefficient after thresholding. The threshold set for the j-th layer. This is a sign function that returns a positive or negative sign. Coefficients with absolute values less than a threshold are treated as noise and set to zero, while coefficients with absolute values greater than a threshold are shrunk towards zero, thus suppressing noise.
[0151] The original Compared with the processed Align them along the time dimension and stitch them together to form a multidimensional feature data matrix. That is, the operating data after noise reduction:
[0152]
[0153] Furthermore, principal component analysis is used to analyze the multidimensional feature data matrix. Dimensionality reduction is performed to obtain the reduced operational data.
[0154] Finally, the dimensionality-reduced operating data is input into the LSTM model for prediction, and the first target operating data of the distribution network at the first future time point is obtained from the model output.
[0155] In this embodiment, by performing outlier identification, wavelet denoising, and principal component dimensionality reduction on historical operating data, erroneous data can be effectively eliminated, random fluctuations can be suppressed, and key features can be retained, significantly improving the prediction accuracy of the long short-term memory network and providing more reliable target operating data for the operation and control of the power distribution network.
[0156] In an exemplary embodiment, during the intraday operation phase, a safety correction is performed on the daily operation plan based on a multi-level security domain model. This includes: predicting the second target operation data of the distribution network at a second future time based on historical operation data at a preset intraday scheduling time; wherein the second future time is earlier than the first future time; performing a risk assessment on the distribution network under a preset fault based on the second target operation data and the multi-level security domain model, and obtaining a risk assessment result; if the risk assessment result indicates that the preset fault will cause the operating state of the distribution network to exceed the allowable operating range, invoking a pre-trained dual-agent reinforcement learning model; and performing a safety correction on the daily operation plan through the dual-agent reinforcement learning model.
[0157] In some embodiments, the dual-agent reinforcement learning model includes a first agent reinforcement learning model and a second agent reinforcement learning model. The dual-agent reinforcement learning model is used to perform safety correction on the day-ahead operating plan, including: using the first agent reinforcement learning model, determining a network structure adjustment direction to adjust the operating state back to the allowable operating range based on deviation information between the operating state and the allowable operating range; using the second agent reinforcement learning model, selecting a set of key switching devices associated with preset faults from each switching device in the distribution network based on the network structure adjustment direction, and generating opening and closing adjustment instructions for each candidate switching device in the set of key switching devices; performing safety correction on the day-ahead operating plan based on each opening and closing adjustment instruction, and controlling each candidate switching device according to the corrected operating plan.
[0158] The preset intraday scheduling time refers to the time when optimization calculations are triggered at fixed intervals (e.g., every 15 or 30 minutes) on the operating day. This time is usually synchronized with the data acquisition system and used to revise the day-ahead plan using the latest measurements and short-term forecast information. The second future time refers to a future period that is closer to the current time, such as the next 15 minutes, with a time scale shorter than the first future time of the day-ahead scheduling. The second target operating data is the photovoltaic power output, wind power output, conventional load, and electric vehicle charging power data of each node or region within the future period obtained through short-term forecasts.
[0159] Preset faults refer to the set of N-1 possible faults in a distribution network, such as the disconnection of any line or the outage of any distribution transformer. Risk assessment refers to using a multi-level security domain model to simulate whether the operating point after a fault exceeds the allowable operating range for each preset fault.
[0160] The dual-agent reinforcement learning model consists of two agents trained using a competitive dual deep Q-network (D3QN) architecture, one responsible for path planning and the other for switch execution. During the offline phase, the two agents learn through interaction with numerous fault scenarios, enabling them to quickly generate switch adjustment schemes during online inference.
[0161] For example, at each preset intraday scheduling time, the server obtains the latest real-time measurement values (such as voltage, current, power, switch status, etc.) and numerical weather forecasts (high-resolution irradiance, wind speed, etc. for the next 2-4 hours) from the data acquisition and monitoring system. Using an LSTM prediction model, the server takes the real-time data from the past hour as input and outputs the second target operational data for the next 2-4 hours, one section every 15 minutes. The operational data in this process also undergoes outlier identification, wavelet denoising, and PCA dimensionality reduction preprocessing to ensure prediction accuracy. For each preset fault (e.g., line L1 disconnection), the actual topology at the current moment is combined with the second target operational data, and at least one electrical quantity, such as node voltage, branch current, and transformer load rate after the fault, is obtained through power flow calculation. These electrical quantities are input into a multi-level safety domain model, which outputs three safety flags: feeder level, transformer level, and substation level. If any level flag is 0, the fault is recorded as a "risk fault," and the level of violation and the expected time section are noted. After all preset faults have been traversed, a risk assessment result list is obtained.
[0162] When at least one risk fault is identified in the risk assessment, the server inputs the current power grid state and the fault identifier into a pre-trained dual-agent model. The model includes: a first agent reinforcement learning model (path planning agent), which receives the current state and deviation information output by the safety domain model (e.g., which feeder voltage exceeds the limit and by how much), and outputs a network structure adjustment direction. This direction is a sparse vector indicating the feeder area or candidate switch range that needs adjustment; and a second agent reinforcement learning model (switch execution agent), which receives the direction information output by the first agent, combines it with the same power grid state, selects the set of key switching devices most relevant to the current risk from all controllable switches in the distribution network, and generates specific opening and closing adjustment instructions for each switch in this set.
[0163] The state vector observed by the first agent for:
[0164]
[0165] The state vector observed by the second agent for:
[0166]
[0167] in, This represents the N-1 risk indicator, used to indicate whether the preset fault in the current assessment will lead to exceeding the limit. It is usually pre-calculated by the rolling risk assessment module, with a value of 1 indicating the presence of risk and 0 indicating no risk. This represents the distribution network topology at time t, i.e., the current open / closed state of all controllable switches. This represents the voltage magnitude vector of each node at time t. This represents the current amplitude vector of each branch at time t. The result also includes the output of the first agent, that is, the state observed by the second agent includes all the information output by the first agent plus the adjustment direction information output by the first agent.
[0168] During load restoration following an N-1 fault, a series of switch switching operations are required. To reduce learning difficulty, only one switch is operated at a time step. The operation space includes all lines, transformers, bus tie switches, and sectionalizing switches of each feeder that are not faulty.
[0169] The action space of the first intelligent agent for:
[0170] A 1 =[ a 0 , a 1 ,..., a m ]
[0171] The action space of the second intelligent agent for:
[0172] A 2 =[ a 0 , a 1 ,..., a n ]
[0173] in, It is a variable of 0 or 1, indicating whether to select the i-th operation, where i∈[0,m] or i∈[0,n]. Vector Only one element in the array is 1, and the rest are 0. This unique 1 corresponds to the network structure adjustment direction chosen by the first agent, such as "transferring the load to feeder B" or "adjusting the bus tie switch state". m is the total number of possible directions. Similarly, Only one element in the array is 1, and the rest are 0. This unique 1 corresponds to a specific switch selected by the second agent, such as closing the contact switch K or opening the segment switch S. n is the total number of controllable switches.
[0174] The first agent outputs a macroscopic adjustment direction without directly operating the switches. The second agent, based on the direction output by the first agent, selects a specific switch from all available switches to operate. The two agents work alternately, executing single-step actions multiple times to ultimately complete the entire load recovery process.
[0175] After the two agents complete a collaborative operation, the server calculates a comprehensive reward based on the new power distribution network status. The reward consists of positive incentives and multiple penalties, targeting safety domains, radial structures, voltage overruns, branch overloads, and switching operation costs.
[0176] Security Domain Reassessment Rewards :
[0177]
[0178] in, This indicates that after the two agents complete their collaborative operation, a risk reassessment is performed on the new distribution network state using a security domain model. If the assessment result is safe, meaning there are no level violations, then... Choose a large positive number, otherwise choose 0. This is the scaling factor for the safety reward, typically set to 1. This reward is the agent's primary positive incentive, guiding it to prioritize restoring the distribution network to the safety domain.
[0179] Radial structure reward :
[0180]
[0181] in, This indicates the state of the distribution network topology after the current action. It represents the set of all possible combinations of radial structures. It is a positive number, and a reward is given when the topology has a radial structure; otherwise, there is no reward.
[0182] Voltage over-limit penalty :
[0183]
[0184] in, This represents the voltage amplitude at node i. , These represent the minimum and maximum allowed voltage values for node i, respectively. r is the penalty scaling factor, for example, set to 1, where the penalty is proportional to the amount exceeding the limit.
[0185] Branch circuit overload penalty :
[0186]
[0187] in, This represents the active power transmitted on branch ij. This represents the maximum permissible transmission power of branch ij. r is the penalty scaling factor; when overloaded, the penalty value increases linearly with the overload rate.
[0188] Switching action cost penalty :
[0189]
[0190] in, This indicates the number of switches that have been activated in the current collaborative operation. r is the penalty metric factor. hour, .along with The increase is exponential, when When the penalty value reaches r, it increases rapidly thereafter, thereby encouraging the agent to complete the transfer within 8 actions.
[0191] The total rewards for each of the two agents:
[0192]
[0193]
[0194] in, The reward is given to the first agent, who does not directly receive the radial reward. This is because the radial pattern is mainly achieved by the switching agent. As a reward for the second agent, it receives an additional radial reward. The switch sequences that are encouraged to operate directly eventually form a radial topology.
[0195] In some embodiments, to address the problem of low sample efficiency in traditional experience pools, this embodiment sets up three independent experience pools: a boundary exploration experience pool. Risk avoidance experience pool and efficient operation experience pool Among them, the boundary exploration experience pool This is used to store samples where agents attempt to navigate near the safety domain boundary without seriously violating it, thus encouraging exploration. Risk avoidance experience pool. This is used to store samples that successfully pulled runpoints back from out-of-limit states to the safe domain, reinforcing safe recovery experience. High-efficiency runpoint experience pool. It is used to store successful samples with few switching actions, fast recovery and low economic loss, in order to optimize economic efficiency.
[0196] The sampling probabilities of the three experience pools are as follows:
[0197]
[0198]
[0199]
[0200] Where k is the current training iteration round. , The training process is divided into three stages based on a preset round threshold. Initial stage: Mid-term: Later stages: . This represents the probability of sampling from the experience pool during training round k. This represents the probability of sampling from the risk avoidance experience pool at training round k. This represents the probability of sampling from the efficient running experience pool at training epoch k. The three conditions must be met. That is, each sample must be drawn from one of the three pools. This represents the boundary exploration sampling probability in the early stages of training. At this time, it encourages more exploration of the safe region boundary, so this value is relatively high, for example, 0.7. To train the boundary exploration sampling probability in the later stages, exploration should be reduced in the later stages to focus on efficient operation. This value should be low, for example, 0.1. The sampling probability of the risk avoidance pool during the initial training phase is, for example, 0.2. This represents the peak value at a certain point in the middle of the process, such as 0.5. To train the sampling probability of the later risk avoidance pool, for example, 0.3.
[0201] At each daily scheduling interval, the server performs a rolling evaluation of all pre-set N-1 faults based on the latest short-term forecast data (e.g., ultra-short-term forecasts for the next 15 minutes to 2 hours). That is, for each device, the short-term forecast data is combined with the current real-time status to calculate the node voltage, branch current, and transformer load rate after the fault. These electrical quantities are input into a multi-level safety domain model to determine if they exceed the permissible operating range. If a fault causes any level of limit exceedance, it is marked as a risk fault, and the exceedance type and expected timeframe are recorded. Once a specific N-1 fault is predicted to cause risk, the server immediately activates the dedicated set of critical switches corresponding to that fault. This set of critical switches is learned offline for this fault type by the dual-agent model and contains a set of the most relevant and effective switches. The server then executes the opening and closing operations of these switches to perform safety corrections and proactively eliminate the risk.
[0202] In an exemplary embodiment, the operation control of the distribution network based on the multi-level security domain model further includes: when a real-time fault occurs in the distribution network and the current operating state of the distribution network under the actual fault exceeds the allowable operating range, invoking a multi-level intelligent agent; generating a fault repair strategy for the actual fault through the multi-level intelligent agent, using the multi-level security domain model as a security constraint; and performing operation control of the distribution network according to the repair strategy.
[0203] Real-time faults refer to equipment faults that occur in the distribution network in real time, such as at least one of the following: line short-circuit tripping, transformer internal faults, or busbar undervoltage. These faults are detected through protection action signals, circuit breaker status changes, or voltage and current surges. The current operating status refers to the real-time electrical quantities after fault isolation or at the moment the fault occurs, including the voltage of each node, the current of each branch, the load rate of each transformer, and the status of switches.
[0204] In some embodiments, the multi-level intelligent agent consists of three intelligent agents pre-trained through hierarchical reinforcement learning: a feeder-level intelligent agent corresponding to the feeder level, a transformer-level intelligent agent corresponding to the transformer level, and a substation-level intelligent agent corresponding to the substation level. The feeder-level intelligent agent can also be understood as an equipment control intelligent agent, the transformer-level intelligent agent can also be understood as a regional coordination intelligent agent, and the substation-level intelligent agent can also be understood as a central coordination intelligent agent.
[0205] For example, when the server detects a fault, it first automatically isolates the faulty section based on the fault indicator and topology information, such as disconnecting the switches on both sides of the faulty line. After isolation, the server immediately collects real-time operational data and models its multi-level security domains. If the model outputs a security flag of 0 at any level, indicating that the operation is outside the permissible range, it determines that emergency control is needed and immediately invokes the multi-level intelligent agent. If all levels are safe, there is no need to trigger emergency control.
[0206] Specifically, each feeder-level agent is responsible for load restoration within its own feeder. Its state space... Represented as:
[0207]
[0208] in, This refers to the topology connection status of the feeder, i.e., the on / off state of the switch. This represents the node voltage amplitude. This represents the branch current. , It refers to the active and reactive power output of distributed power sources. It refers to the state of charge of the energy storage system connected to the feeder. This is the total load of the feeder. It is a fault indicator, indicating the location of the faulty section. It is a feeder-level security domain over-limit indicator.
[0209] Its action space Represented as:
[0210]
[0211] in, To control the switching action of the sectionalizing switch and the tie switch on this feeder line. To adjust the output of distributed power sources and energy storage connected to this feeder.
[0212] reward function Represented as:
[0213]
[0214] in, This refers to the load power restored after the agent takes action. This is the reward for the satisfaction of the security domain at the feeder level; it is 1 if the security domain is within the target range, and -1 otherwise. This represents the number of times the switch has been activated. This is the load shedding amount. , , , These are the weighting coefficients for different objectives.
[0215] The feeder-level agent first attempts to restore the load through switching operations within the feeder. If, after several steps, such as five steps, the safety domain model still determines that the load is unsafe, the feeder-level agent sends a distress signal to its corresponding transformer-level agent, including the current over-limit type and severity.
[0216] A transformer-level intelligent agent is responsible for one main transformer and its multiple feeders. Its state space... Represented as:
[0217]
[0218] in, It provides the latest status information for all feeder-level intelligent agents under its jurisdiction (such as the load rate of each feeder and voltage over-limit status). This refers to the topology of the area under its jurisdiction (such as the status of the bus tie switch on the low-voltage side of the main transformer). , These represent the active power and reactive power on the low-voltage side of the main transformer, respectively. This is the voltage at the point of common coupling. This refers to the current of all outgoing feeders under its jurisdiction. This is the primary variable level security domain over-limit flag. This is a distress signal sent by a subordinate feeder-level intelligent agent.
[0219] Its action space Represented as:
[0220]
[0221] in, To issue coordination instructions to one or more designated feeder-level agents, such as "feeder A to transfer load to feeder B". For direct operation of regional equipment, such as adjusting the tap changer of an on-load tap-changing transformer, operating a bus tie switch, or at least one other type.
[0222] reward function Represented as:
[0223]
[0224] in, This represents the total load power restored within the jurisdiction. Rewards are given based on the satisfaction level of the primary security domain. This represents the average reward of all feeder-level agents under its jurisdiction. The total cost of the transformer-level intelligent agent itself and the actions of issuing commands. , , , These are the weighting coefficients for different objectives.
[0225] Upon receiving a distress call, the transformer-level agent, based on global information, decides whether to transfer some load from the faulty feeder to an adjacent healthy feeder. For example, the transformer-level agent can instruct the closure of the tie switch between the two feeders and adjust the output of the distributed power sources on both sides. If the transformer-level agent can eliminate all over-limits through regional coordination, the entire recovery process ends. Otherwise, the transformer-level agent continues to send distress signals to the superior substation-level agent.
[0226] The substation-level intelligent agent is responsible for the entire substation and inter-station communication. Its state space... Represented as:
[0227]
[0228] in, It represents the status information of all transformer-level intelligent entities under its jurisdiction (such as the load rate of each main transformer and the voltage level of the area). This indicates the status of critical interconnection lines within and between stations. , These represent the active and reactive power transmitted on the connection lines between this substation and other substations, respectively. This refers to the critical bus voltage. This is a safety domain over-limit indicator for substation levels. This is a distress signal sent by a subordinate transformer-level intelligent agent.
[0229] Its action space Represented as:
[0230]
[0231] in, To coordinate actions, higher-level coordination instructions are issued to one or more designated transformer-level agents, such as "start the interconnection switch between substations S1 and S2, and transfer the 50kW load from main transformer T1 to main transformer T2".
[0232] reward function Represented as:
[0233]
[0234] in, This represents the total load power restored across the entire network. For substation-level security domain satisfaction. This represents the average reward for all transformer-level agents under its jurisdiction. Costs associated with large-scale, global operations (such as inter-station communication switches).
[0235] After receiving a request for assistance from the transformer-level agent, the substation-level agent assesses the remaining capacity of the entire network. It may activate inter-station interconnection switches to transfer some load to adjacent substations, or adjust the tap changers of the main transformer to stabilize voltage. The substation-level agent's decision-making is at the highest level and usually involves large-scale operations; therefore, its reward function includes a high operational cost.
[0236] Each level of the intelligent agent undergoes independent pre-training in the offline phase. The feeder-level agent learns basic load recovery in a single feeder environment; the transformer-level agent learns regional coordination in a regional environment containing multiple feeders; and the substation-level agent learns global resource allocation in a network-wide environment. Then, the pre-trained models are placed in a co-simulation environment for online fine-tuning. Through an event-triggered communication mechanism (i.e., the lower layers only send requests for help to the upper layers when they cannot solve the problem, and the upper layers issue specific coordination instructions based on the status of the lower layers), the agents learn inter-level collaborative strategies. The final trained multi-level agent can quickly generate optimal recovery strategies from local to global perspectives after a real fault occurs.
[0237] In this embodiment, by constructing a three-layer intelligent agent collaborative mechanism of feeder, transformer and substation, and judging the over-limit state in real time based on the multi-level security domain model, an adaptive repair strategy from local recovery to global coordination can be quickly generated after the fault occurs, which can significantly shorten the fault handling time, reduce the load shedding, and significantly improve the adaptability and power supply reliability of the distribution network under complex fault scenarios.
[0238] In an exemplary embodiment, the active operation optimization method for distribution networks based on multi-level security domains further includes: dividing the distribution network into multiple power grid regions.
[0239] In some embodiments, dividing the distribution network into multiple power grid regions includes: obtaining power flow calculation results of the distribution network; constructing a sensitivity matrix based on the power flow calculation results to reflect the electrical coupling strength between nodes in the distribution network; determining the electrical distance between nodes based on the sensitivity matrix; and clustering the nodes based on the electrical distance to obtain multiple power grid regions.
[0240] The power flow calculation results are obtained by numerical methods such as the Newton-Raphson method to solve for the steady-state distribution of voltage amplitude, phase angle, and power of each branch in the distribution network. In this embodiment, the power flow calculation is used to obtain sensitivity information in the Jacobian matrix. The sensitivity matrix is used to describe the degree of influence of node power changes on node voltage. Electrical distance is an indicator used to quantify the degree of electrical coupling between two nodes. The power grid region is a set of nodes obtained after clustering, also known as an equivalent load block. The electrical characteristics within each region are strongly correlated, which can be used as the basic analysis unit in subsequent security domain modeling and fault simulation, thereby reducing the complexity of the overall network calculation.
[0241] For example, firstly, the server performs power flow calculations using the Newton-Raphson method for typical operating modes of the distribution network (e.g., taking a historical operating section during a peak load period on a certain day), and obtains the following results:
[0242]
[0243] in Active power-voltage sensitivity reflects the relationship between voltage amplitude U and active power P. In a power grid, when the active power (such as photovoltaic output or load active power) of a node changes slightly, the voltage amplitude of that node will change accordingly. The ratio of this change is called active power-voltage sensitivity. The reactive-voltage sensitivity is the ratio of the change in voltage amplitude when the reactive power at a node (such as capacitor switching and inverter reactive power support) changes, reflecting the coupling relationship between voltage amplitude U and reactive power Q.
[0244] Furthermore, construct the sensitivity matrix:
[0245]
[0246] in, λ ∈ [0,1] , where is the weighting coefficient. Sensitivity matrix elements in This demonstrates the impact of node j's power on node i's voltage. The larger the value, the closer the electrical distance between the two nodes, and the greater the probability that node j and node i will be assigned to the same region.
[0247] Furthermore, in the distribution network structure, other nodes connected to nodes j and i will also affect the voltage of nodes j and i. Assuming there are N nodes in the distribution network, the electrical distance between nodes j and i is defined as... :
[0248]
[0249] in, is an element in the sensitivity matrix, representing the degree of influence of the unit power change at node s on the voltage amplitude at node i. The larger the value, the more significant the impact of power fluctuations at node s on the voltage at node i. Similarly, This represents the effect of power changes at node s on the voltage at node j.
[0250] Based on electrical distance, the weight of the edge between nodes j and i is defined as follows: :
[0251]
[0252] in, This represents the maximum electrical distance between all node pairs in the entire distribution network, i.e., the global maximum value.
[0253] Furthermore, a modularity algorithm based on the aforementioned weights is used for clustering. Modularity Defined as:
[0254]
[0255]
[0256] in, This is the sum of the weights of the edges connected to node i. Let be the sum of the weights of the edges connected to node j. Let m be the sum of the weights of the network edges. A cluster consisting of individual nodes.
[0257] Maximize through iterative optimization (such as moving nodes to neighboring regions). The final result is several node clusters, each representing a power grid region. These regions are electrically tightly coupled within each other, but weakly coupled between them. Each power grid region can be considered as an equivalent load block, simplifying subsequent analysis.
[0258] In some embodiments, historical operating data includes regional operating data corresponding to each power grid region; based on the historical operating data of the distribution network and fault simulation data, a multi-level security domain model of the distribution network is constructed, including: for each power grid region, based on the regional operating data and fault simulation data corresponding to the power grid region, performing network topology reconstruction simulation under fault conditions to obtain regional reconstruction simulation results; summarizing the regional reconstruction simulation results to obtain topology reconstruction simulation results; and constructing a multi-level security domain model of the distribution network based on the topology reconstruction simulation results.
[0259] Among them, regional operation data is extracted from the original historical operation data of each power grid region, including the time-series data of all nodes and branches within that region, as well as the power data exchanged with other regions at the regional boundary (i.e., equivalent injected power). Regional fault simulation simulates an N-1 fault within or at the boundary of a single power grid region based on its regional operation data, and performs controllable switch reconfiguration operations within the region.
[0260] For example, after the regions are divided, the server reorganizes the original historical operating data by region. The operating data for each region includes the voltage, power, and switch status of all nodes within that region, as well as power exchange data on the interconnecting lines between regions. Simultaneously, fault simulation data is also allocated by region, setting N-1 fault scenarios for the lines and transformers within each region, and for the interconnecting lines between regions. Subsequently, network topology reconstruction simulation is performed independently for each power grid region. Finally, the simulation results from each region are aggregated and integrated into a network-wide topology reconstruction simulation result, thereby significantly reducing the computational scale of the global simulation.
[0261] In this embodiment, electrical distances are constructed by power flow sensitivity and clustered to obtain electrically coupled power grid regions. The large-scale distribution network is decomposed into multiple independent sub-regions. Each region only needs to perform fault simulation based on its own operating data, which greatly reduces the dimensionality and computational load of topology reconfiguration simulation.
[0262] In one specific embodiment, such as Figure 3 As shown, this embodiment constructs a multi-timescale closed-loop optimization system covering day-ahead, intraday, and real-time scenarios, including: dividing the distribution network into regions based on electrical distance and modularity functions, and using a Long Short-Term Memory (LSTM) network to accurately predict the source-load status; generating massive operational scenarios through a Conditional Generative Adversarial Network (CGAN), and obtaining safety labels by combining N-1 fault simulation, thereby training and constructing a data-driven multi-level safety domain model; performing day-ahead economic scheduling based on the constructed safety domain model as a safety constraint to maximize the economic access of distributed source-loads; using the safety domain model to predict risks based on ultra-short-term predictions, and performing preventative topology reconfiguration to proactively avoid potential N-1 fault risks; and in the event of an emergency fault, launching a collaborative hierarchical multi-agent reinforcement learning emergency control architecture based on the safety domain model to achieve rapid recovery.
[0263] Figure 4 This diagram illustrates the multi-level reconstruction based on the Spatiotemporal Graph Attention Network (ST-GAT-GRU) in this embodiment. The process begins by fusing two types of key data—static and dynamic features—as input. The model's internal multi-spatiotemporal feature fusion layer and fully connected layer deeply analyze the spatial topology and electrical coupling relationships of the power grid through the Graph Attention Network (GAT), and utilize gated cyclic units (GRUs) to capture the dynamic patterns of source-load evolution over time. Its output employs a multi-label classification architecture to generate a multi-level reconstruction strategy, namely a structured sequence tensor. This tensor can precisely define the switch closing / opening state sequence at four different levels within the next 24 hours. These four levels correspond to the substation interconnection switch, transformer interconnection switch, feeder interconnection switch, and branch section switch, respectively.
[0264] Figure 5This is a schematic diagram of the training process for the dual-agent D3QN for intraday preventative reconfiguration in this embodiment. Its core is a collaborative decision-making architecture consisting of a path planning agent and a switch execution agent. Both jointly observe the distribution network state and, through a series of coordinated single-step switch switching actions, quickly identify and filter a set of key switch candidates highly relevant to mitigating the current risk from the entire network. To guide the agents in learning the optimal strategy, the training process is driven by a comprehensive reward function. This function not only provides a large positive reward when the system successfully restores to a safe state, but also includes rewards for ensuring the network maintains a radial structure, as well as corresponding penalties for voltage exceedances, line overloads, and frequent switch actions. To improve training efficiency, this process innovatively sets up three independent experience pools: boundary exploration, risk avoidance, and efficient operation. By dynamically adjusting the sampling probability from each pool in the initial, middle, and final stages of training, the agent is guided from exploring the safety boundary to focusing on risk avoidance, and finally to optimizing operational efficiency, thereby achieving a closed-loop learning process for rapid, intelligent, and proactive avoidance of potential intraday risks.
[0265] Figure 6 This diagram illustrates the collaborative hierarchical multi-agent reinforcement learning architecture based on real-time emergency control in this embodiment. Its core is a collaborative hierarchical multi-agent reinforcement learning emergency response system. This architecture decomposes the overall network recovery target into different levels and regions, constructing a command system composed of three levels of agents: substation, transformer, and feeder, to achieve second-level rapid recovery. Specifically, the bottom-level feeder-level agents are responsible for the operation of equipment on a single feeder; the middle-level transformer-level agents coordinate multiple feeders under their respective main transformers; and the top-level substation-level agents perform global coordination. The architecture's collaborative mechanism employs an efficient event-triggered mode. Only when a bottom-level agent cannot resolve the problem independently will it send a distress signal to its superior agent, which then issues specific coordination instructions based on its assessment. Through this bottom-up information feedback and top-down hierarchical decision-making, the architecture can achieve multi-level collaborative emergency response, forming a globally optimal recovery strategy within seconds.
[0266] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0267] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device stores a computer program that, when executed by a processor, implements a method for proactively optimizing the operation of a distribution network based on multi-level security domains. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.
[0268] It should be noted that the data involved in this application (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
Claims
1. A method for active operation optimization of distribution networks based on multi-level security domains, characterized in that, The method includes: Based on historical operation data and fault simulation data of the distribution network, a multi-level security domain model of the distribution network is constructed; wherein, the multi-level security domain model is used to characterize the allowable operating range of the distribution network at the feeder level, transformer level and substation level respectively; Based on the multi-level security domain model, the day-ahead operation optimization of the distribution network is performed to generate the day-ahead operation plan of the distribution network; During the intraday operation phase, the daily operation plan is adjusted for safety based on the multi-level security domain model. In the event of a real-time fault in the distribution network, a fault repair strategy for the distribution network is generated using the multi-level security domain model as a security constraint.
2. The method according to claim 1, characterized in that, The step of optimizing the day-ahead operation of the distribution network based on the multi-level security domain model to generate the day-ahead operation plan of the distribution network includes: At a preset daytime scheduling time, based on the historical operating data, the first target operating data of the distribution network at a first future time is predicted; The first target operation data is input into a decision model pre-constructed based on a spatiotemporal graph attention network, so that the decision model can analyze the first target operation data with the multi-level security domain model as a security constraint, and output the opening and closing instructions of the switching equipment contained in the feeder level, the transformer level and the substation level respectively. Under the network topology determined by the opening and closing instructions, a deep reinforcement learning agent generates a daily operation plan for the power distribution equipment in the power distribution network, using the multi-level security domain model as a security constraint. The power distribution equipment is operated in accordance with the operation control instructions that match the day-ahead operation plan.
3. The method according to claim 2, characterized in that, The prediction of the first target operating data of the distribution network at a first future moment based on the historical operating data includes: The historical operating data is subjected to outlier identification to obtain intermediate operating data after outlier removal; The intermediate running data is subjected to noise reduction processing to obtain noise-reduced running data; The noise-reduced operating data is then subjected to dimensionality reduction processing to obtain dimensionality-reduced operating data; The dimensionality-reduced operating data is input into a pre-trained prediction model, so that the prediction model can predict the first target operating data of the power distribution network at a first future time by analyzing the dimensionality-reduced operating data; wherein, the prediction model is constructed based on a long short-term memory network.
4. The method according to claim 3, characterized in that, The step of identifying outliers in the historical operating data to obtain intermediate operating data after removing outliers includes: The historical operational data is clustered to obtain multiple data clusters; The random forest algorithm is used to identify outliers in each data cluster, and intermediate running data after removing outliers is obtained.
5. The method according to claim 1, characterized in that, During the intraday operation phase, based on the multi-level security domain model, a security correction is performed on the daily operation plan, including: Under a preset intraday scheduling time, based on the historical operating data, the second target operating data of the distribution network at a second future time is predicted; wherein, the second future time is earlier than the first future time; Based on the second target operating data and the multi-level security domain model, a risk assessment is performed on the power distribution network under a preset fault, and the risk assessment result is obtained. If the risk assessment results indicate that the preset fault will cause the operating state of the power distribution network to exceed the allowable operating range, the pre-trained dual-agent reinforcement learning model is invoked. The daily operation plan is corrected for safety using the dual-agent reinforcement learning model.
6. The method according to claim 5, characterized in that, The dual-agent reinforcement learning model includes a first agent reinforcement learning model and a second agent reinforcement learning model; the step of performing safety correction on the current day's operational plan using the dual-agent reinforcement learning model includes: Based on the deviation information between the operating state and the allowed operating range, the first agent reinforcement learning model determines the network structure adjustment direction to adjust the operating state back to the allowed operating range. Using the second agent reinforcement learning model, the set of key switching devices associated with the preset fault is selected from the switching devices of the power distribution network according to the network structure adjustment direction, and the opening and closing adjustment instructions for each candidate switching device in the set of key switching devices are generated. Based on the aforementioned switching adjustment commands, the daytime operation plan is safety-corrected, and the candidate switching devices are controlled according to the corrected operation plan.
7. The method according to claim 1, characterized in that, In the event of a real-time fault in the distribution network, a fault repair strategy for the distribution network is generated using the multi-level security domain model as a security constraint, including: When a real-time fault occurs in the power distribution network, and the current operating state of the power distribution network under the actual fault exceeds the allowable operating range, a multi-level intelligent agent is invoked. Through the multi-level intelligent agent, and using the multi-level security domain model as security constraints, a fault repair strategy is generated for the actual fault.
8. The method according to claim 7, characterized in that, The multi-level intelligent agent includes a feeder-level intelligent agent corresponding to the feeder level, a transformer-level intelligent agent corresponding to the transformer level, and a substation-level intelligent agent corresponding to the substation level.
9. The method according to claim 1, characterized in that, The method further includes: The power distribution network is divided into multiple power grid areas; The historical operational data includes regional operational data corresponding to each of the power grid regions; the construction of a multi-level security domain model for the distribution network based on the historical operational data and fault simulation data includes: For each power grid region, based on the region operation data and fault simulation data corresponding to the power grid region, network topology reconstruction simulation under fault conditions is performed on the power grid region to obtain the region reconstruction simulation results. The simulation results of each region reconstruction are summarized to obtain the topology reconstruction simulation results; Based on the simulation results of the topology reconfiguration, a multi-level security domain model of the distribution network is constructed.
10. The method according to claim 9, characterized in that, The division of the distribution network into multiple power grid areas includes: Obtain the power flow calculation results of the power distribution network; Based on the power flow calculation results, a sensitivity matrix is constructed to reflect the electrical coupling strength between nodes in the distribution network. The electrical distance between each node is determined based on the sensitivity matrix. Based on the electrical distance, the nodes are clustered to obtain multiple power grid regions.