Large-scale power distribution network topology reconstruction method based on DQN algorithm

By adopting a topology reconfiguration method based on the DQN algorithm, the problems of slow response speed and high computational complexity in distribution network topology optimization are solved, realizing automated and dynamic topology optimization of the distribution network and improving economy and security.

CN121923090APending Publication Date: 2026-04-24STATE GRID JIANGSU ELECTRIC POWER CO LTD TAIZHOU POWER SUPPLY BRANCH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD TAIZHOU POWER SUPPLY BRANCH
Filing Date
2025-12-25
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies for distribution network topology optimization suffer from slow response speed, susceptibility to local optima, high computational complexity, and difficulty in balancing multi-objective optimization and real-time constraints. In particular, they struggle to generate the optimal topology structure when load fluctuates and distributed power sources are integrated, resulting in insufficient operational economy and reliability.

Method used

A topology reconstruction method based on the DQN algorithm is adopted. The DQN model is constructed through graph theory modeling and power flow calculation modules. The state space, action space and multi-objective reward function are designed. Combined with the reinforcement learning training mechanism, efficient topology optimization is achieved, and a compliance verification link is set up.

Benefits of technology

It achieves automation and dynamism in distribution network topology reconfiguration, taking into account reduced line losses, improved voltage quality, and topology compliance, thus significantly improving the operational economy and safety of the distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121923090A_ABST
    Figure CN121923090A_ABST
Patent Text Reader

Abstract

The invention discloses a large-scale power distribution network topology reconstruction method based on a DQN algorithm, and belongs to the technical field of power transmission and distribution network optimization. The method comprises the following steps: S1, constructing a topological model based on a graph theory, determining voltage, branch capacity and topological radial constraints, and obtaining operation parameters through load flow calculation; s2, designing a state space, a node voltage, a switching state and the like, an action space, an effective switching operation and a multi-target reward function containing violation punishment of the DQN, and training a model by adopting experience playback, target network updating and a greedy strategy; s3, selecting an optimal switching operation adjustment topology by using the trained model; and S4, outputting an optimization result. According to the method, the high-dimensional state is processed through the DQN, loss reduction and compliance are considered, training is stable, convergence is fast, load fluctuation can be dealt with, dynamic reconstruction is achieved, line loss is reduced after optimization, and power supply economy and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system operation and maintenance technology, and in particular to a method for topology reconfiguration of large distribution networks based on the DQN algorithm. Background Technology

[0002] This invention relates to the field of distribution network operation optimization technology, specifically to the background technology of distribution network topology reconfiguration. Currently, distribution network topology optimization (also known as topology reconfiguration) mainly relies on manual experience to adjust switch states, heuristic optimization methods based on genetic algorithms or particle swarm optimization, and mathematical modeling methods based on mixed integer programming. Among these, manual experience-based adjustment relies on the subjective judgment of professionals regarding the distribution network's operating state, resulting in slow response speed, poor consistency of optimization results, and difficulty in handling real-time dynamic changes in large-scale distribution networks (such as load fluctuations and distributed power source integration). Heuristic algorithms, while achieving some degree of automated optimization, generally suffer from slow convergence speed, susceptibility to local optima, and a lack of dynamic adaptability in handling real-time operational constraints (such as node voltage exceeding limits, topology loops, or island formation). Mathematical programming methods, due to the large number of nodes and complex topology of distribution networks, exhibit exponentially increasing computational complexity, making it difficult to meet the timeliness requirements of real-time distribution network optimization and to fully consider the multi-dimensional constraints in distribution network operation (such as minimizing line losses and improving voltage quality).

[0003] In practical applications, existing technologies often fail to balance the efficiency and accuracy of topology optimization. In particular, when dealing with scenarios such as high load rate branch adjustments in distribution networks and topology reconfiguration after the integration of distributed power sources, it is difficult to quickly generate the optimal topology structure that meets real-time operating constraints, resulting in the inability to effectively improve the economic efficiency and reliability of distribution network operation.

[0004] Patent CN120613716A discloses a partitioned mesh power transmission and distribution network optimization system and method, comprising seven units including topology perception and partitioning, power flow data acquisition and feature extraction, etc. The topology perception and partitioning unit dynamically divides sub-regions based on the connection relationships of power grid nodes and load characteristics; the power flow data acquisition and feature extraction unit deeply extracts data features such as voltage and current; the improved deep Q-network strategy generation unit combines power grid parameters to construct a state and action space to generate a preliminary strategy; the optimized attention mechanism weight allocation unit allocates weights based on node electrical distance and line transmission capacity for strategy fusion; the multi-region collaborative decision fusion unit comprehensively considers sub-region power interaction and stability indicators for global decision-making; the control command generation and issuance unit issues commands based on equipment constraints; and the operation status feedback and update unit achieves real-time status updates, improving the efficiency and safety of power grid operation optimization. However, the above method does not explicitly include a topology compliance verification step. After improving the DQN to generate the preliminary strategy, it only optimizes through attention weight fusion and multi-region collaborative decision-making, but does not perform real-time verification of whether the topology has loops, islands, or voltage exceeding limits. Therefore, those skilled in the art urgently need to solve the above technical problems. Summary of the Invention

[0005] The proposed method for large-scale distribution network topology reconfiguration based on the DQN algorithm effectively solves the problems of slow response of traditional manual experience-based adjustments, easy getting trapped in local optima by heuristic algorithms, and high computational complexity of mathematical programming methods. By using the DQN algorithm to handle high-dimensional states, it achieves efficient training and stable convergence, taking into account multiple objectives such as reducing line losses, topology compliance (no loops, islanding, voltage over-limits), and voltage quality optimization. It can dynamically respond to load fluctuations to achieve real-time reconfiguration, and a compliance verification link is set up to ensure operational safety. After optimization, it significantly improves the economic efficiency and reliability of power supply.

[0006] A method for topology reconfiguration of a large-scale distribution network based on the DQN algorithm includes the following steps:

[0007] Step S1. Distribution network modeling: Construct a distribution network topology model based on graph theory, clarifying voltage constraints, branch capacity constraints, and topology radial constraints; construct a power flow calculation module, inputting switch states and node loads, and outputting power grid operating state parameters;

[0008] Step S2. DQN Model Construction and Training: Design the state space, action space, and multi-objective reward function of the DQN model; train the DQN model using a reinforcement learning training mechanism;

[0009] Step S3. Topology optimization: Use the trained DQN model to select the optimal switching operation and adjust the opening and closing state of the distribution network switches;

[0010] Step S4. Output Results: Output the optimized distribution network topology and operating parameters.

[0011] By adopting the above technical solution and integrating four core steps—distribution network modeling, DQN model construction and training, topology optimization, and result output—this approach effectively solves the problems of strong subjectivity in traditional manual adjustments based on experience, susceptibility of heuristic algorithms to local optima, and high computational complexity of mathematical programming methods. Leveraging the high-dimensional state processing capabilities of the DQN algorithm, it achieves automated and dynamic topology reconfiguration, balancing multiple objectives such as line loss reduction, voltage quality improvement, and topology compliance, significantly enhancing the economy and safety of distribution network operation.

[0012] Furthermore, in step S1, the distribution network modeling specifically includes the following steps:

[0013] S11. Topology Abstraction: Based on graph theory, the distribution network is transformed into a node-edge undirected graph model, where the node set contains 1 power supply node and n-1 load nodes, where n is the total number of nodes in the distribution network, and the edge set corresponds to the lines and switching equipment, with sectionalizing switches being mandatory edges and tie switches being optional edges;

[0014] S12. Constraint Quantization: Establish a mapping relationship between the topology and switch states. When the switch is closed, the corresponding edge exists in the undirected graph; when the switch is open, the corresponding edge is removed from the undirected graph. Voltage constraints are... ,in, Let be the voltage magnitude at node i. The rated voltage and branch capacity constraints are: ,in, branch road Apparent power branch road Rated capacity;

[0015] S13. Power Flow Calculation Model: The topology radial constraint is achieved through graph connectivity determination, requiring that there is a unique path between any two nodes and no loops, and that all nodes are connected to power nodes and there are no islands. The forward and backward substitution method is used to construct the power flow calculation module, which takes switch status and node load data as input and outputs node voltage, branch power and bus loss.

[0016] By adopting the above technical solution, the complex power grid is abstracted into a node-edge undirected graph using graph theory, clearly distinguishing the roles of sectionalizing switches and tie switches. Combined with the quantification of voltage, branch capacity, and radial constraints, a structured and standardized input foundation is provided for subsequent model training. The introduction of the power flow calculation module can accurately output power grid operating state parameters, ensuring the accuracy of DQN model training and topology optimization, and avoiding decision-making biases caused by fuzzy input data.

[0017] Further, in step S13, the power flow calculation step includes:

[0018] S131. Initialization parameters: Set the initial value of the node voltage; the power node voltage is the rated voltage. The initial value of the load node voltage is set to The initial value of the branch power is 0, and the convergence accuracy threshold is... ;

[0019] S132. Forward Calculation: Starting from the power node, calculate the power loss and terminal node voltage of each branch sequentially according to the branch connection order of the radial network topology. For branch j, calculate the power loss and terminal node voltage of each branch according to the terminal node voltage. Branch impedance and load power Calculate the active power loss of the branch. Reactive power loss Thus, the power of the end node is obtained. , And derive the end node voltage. ;

[0020] S133. Backward Substitution Correction: Starting from the farthest load node and working backward to the power supply node, based on the terminal node voltage obtained from the forward calculation, correct the voltage amplitude and phase angle of each node and update the branch power distribution;

[0021] S134. Convergence Judgment: Compare the node voltage difference and branch power difference between two adjacent iterations. If the maximum value of all node voltage differences is less than the convergence accuracy threshold ε, and the branch power difference meets the requirements, then stop the iteration and output the final node voltage, branch power, and bus loss; otherwise, return to the previous calculation step and repeat the iteration.

[0022] By adopting the above technical solution, the forward-backward substitution method is adapted to the characteristics of the radial power grid. The initialization parameters and iterative convergence mechanism ensure the efficiency and accuracy of the calculation. By forward-calculating branch power loss and terminal voltage and backward-substituting to correct node states, key parameters such as node voltage and branch power can be quickly output, providing real-time and reliable state information for topology adjustment and supporting the timeliness requirements of dynamic reconfiguration.

[0023] Furthermore, the DQN model includes a state-space input module, a fully connected hidden layer, an output layer, and a parameter optimization unit, wherein:

[0024] Input distribution network operating state vector Output the Q-value vector Q for each topology adjustment action;

[0025] Model parameters: , where W is the weight matrix and b is the bias vector;

[0026] The fully connected hidden layer comprises three fully connected neural networks: an input layer, a 64-neuron hidden layer 1, a 32-neuron hidden layer 2, and an output layer.

[0027] The input layer directly receives the normalized operating state vector of the distribution network, without parameterized operations, and its mathematical expression is: ;in: The state vector at time t, where d = number of load nodes + number of tie switches + number of critical branches;

[0028] : Input layer output, i.e., input of hidden layer one;

[0029] The first hidden layer consists of 64 neurons, denoted as L1=64, and uses the ReLU activation function. To achieve nonlinear feature extraction of the input state, mathematically expressed as:

[0030] ;

[0031] in: : The weight matrix from the input layer to the first hidden layer;

[0032] : The bias vector of the first hidden layer;

[0033] Output of the first hidden layer;

[0034] The ReLU activation function is used to introduce non-linearity and prevent the model from fitting only a linear relationship.

[0035] Hidden layer 2 consists of 32 neurons, denoted as L2=32. It also uses the ReLU activation function to further abstract and compress the features of the first hidden layer, mathematically expressed as:

[0036] ,in: : Weight matrix from the first hidden layer to the second hidden layer (dimension 64×32); : The bias vector of the second hidden layer;

[0037] : Output of the second hidden layer;

[0038] The output layer is linearly activated, with no activation function, or can be considered as an activation function. Output the Q-value corresponding to each topology adjustment action, mathematically expressed as: ;

[0039] in: : The weight matrix from the second hidden layer to the output layer;

[0040] : The bias vector of the output layer;

[0041] State s at time t t Next, each action The Q-value vector is used for subsequent greedy strategies to select the optimal action;

[0042] The mean squared error loss between the current Q value and the target Q value, combined with the Bellman equation, is mathematically expressed as:

[0043] ;

[0044] Where: D: Experience replay pool, storing historical state-action-reward-next state samples;

[0045] : Target Q value, θ′ is the target network parameter, θ is synchronized every 100 steps;

[0046] The immediate reward at time t;

[0047] Discount factor;

[0048] Iterative updates are achieved using gradient descent, such as the Adam optimizer. , minimize .

[0049] Furthermore, for the current state, the action is determined by measuring the reward based on the Q-value. Combining this with the optimal Bellman equation, the iterative formula for Q-learning can be derived as follows: (1);

[0050] in, This is the state transition function. Is the action value function about Expectations at all times It is the intelligent agent that takes action. The reward provided by the post-environment feedback, where γ is the discount factor. For time t The estimate is given by α, where α is the learning rate.

[0051] By employing the above technical solutions, the long-term benefits of state-action pairs are quantified using Q-values, guiding the model to learn policies that conform to multi-objective optimization and avoiding short-sighted decision-making. The iterative formula of the Bellman equation provides a theoretical basis for model training, ensuring the rationality of Q-value updates and enabling the model to gradually converge to the optimal policy, achieving a balance between objectives such as reduced line loss and topology compliance.

[0052] Furthermore, the construction and training of the DQN model specifically includes the following steps:

[0053] S21. State-space design:

[0054] Constructing a high-dimensional state vector to comprehensively characterize the operating state of the distribution network, specifically including:

[0055] Node voltage characteristics: The voltage amplitude of all load nodes is normalized to the [0,1] interval to eliminate the dimensional differences of different physical quantities;

[0056] Normalization formula: ;

[0057] in: : The normalized voltage value of the i-th load node, with an output range of [0,1]; : The original voltage amplitude of the i-th load node; =0.95 , =1.05 , : System rated voltage.

[0058] Switch status characteristics: The open / closed status of all interconnecting switches is encoded in binary, where 1 indicates closed and 0 indicates open;

[0059] Branch load characteristics: The active power of 3-5 branches with high load rate is selected and normalized to the [0,1] interval to characterize the overload risk status of the branches. The dimensions of the state vector are the number of load nodes, the number of tie switches and the number of critical branches, of which the number of critical branches is 3-5.

[0060] S22. Operational Space Design:

[0061] Filter the set of valid switch operations and exclude invalid actions that could lead to islanding. Specific operation types include:

[0062] Handling switch state toggling: Switches the state of the handling switch from closed to open, or from open to closed;

[0063] Redundant sectionalizing switch disconnection: Only for branches with large power current in the loop, the redundant sectionalizing switch is disconnected;

[0064] No operation: No switching action is performed to avoid equipment damage caused by frequent switching actions;

[0065] The size of the action space is controlled by "number of communication switches + 2", where "2" corresponds to the redundant segment switch disconnection operation and no operation;

[0066] S23. Reward Function Design:

[0067] The mathematical expression for constructing a multi-objective quantitative reward mechanism is as follows:

[0068] ;

[0069] in, The weighting coefficients are determined through engineering examples and prioritize the loss reduction target.

[0070] The average deviation of the node voltage;

[0071] To constrain the penalty for violations: the value is 1 when there are loops, islanding, or voltage exceeding limits in the distribution network; otherwise, the value is 0.

[0072] The average deviation of node voltage is the average of the absolute values ​​of the deviations of all load node voltages from the rated voltage. The smaller the deviation, the smaller the penalty.

[0073] By adopting the above technical solutions, the design of the state space, action space, and reward function of the DQN model is improved. The state space comprehensively covers key information such as load node voltage and tie switch status, ensuring that the model can perceive the overall operation of the power grid. The action space filters effective operations to avoid invalid actions and reduces computational redundancy. The multi-objective reward function balances loss reduction, voltage quality, and compliance, guiding the model to train towards the comprehensive optimal direction and improving the practicality of topology reconfiguration.

[0074] Further, in step S2, the training method for the DQN model is as follows: the DQN model adopts a 3-layer fully connected neural network structure, consisting of an input layer, a first hidden layer with 64 neurons, a second hidden layer with 32 neurons, and an output layer; the first and second hidden layers use the ReLU activation function, the output layer uses a linear activation function, and outputs the long-term cumulative reward expectation Q value of each switching action; an experience replay mechanism is used to break data correlation; a target Q network is introduced, and the current network parameters are synchronized every 100 steps to stabilize training; a greedy strategy is used to balance exploration and utilization.

[0075] The activation function can be used to introduce non-linear feature extraction capabilities or preserve the value quantification range, as detailed below:

[0076] Hidden layer: Nonlinearity is introduced through the ReLU activation function to extract complex state features. Specifically, the operating state of the distribution network—node voltage, switch state, and branch power—is a high-dimensional, nonlinear relationship. The mathematical definition of ReLU is: Where: x: input to the activation function; σReLU(x): output of the activation function;

[0077] Output layer: Uses a linear activation function, the specific formula of which is: Where: x: weighted sum of neurons in the output layer; : The output of the activation function.

[0078] Output logic for the long-term cumulative reward expected Q value:

[0079] Input: The high-dimensional state vector st consists of the load node normalized voltage, tie switch binary state, and high load branch normalized power, containing all the key information of the distribution network operation;

[0080] Feature extraction: The role of ReLU hidden layers: The first hidden layer filters out "invalid associations" in the state, the second hidden layer focuses on "key associations", and finally outputs the core features;

[0081] Value estimation: The role of the linear output layer The output layer maps features to the Q-values ​​of A actions through a weight matrix. Linear activation ensures that the Q-values ​​accurately reflect the true range of "long-term cumulative reward".

[0082] The specific process of the experience replay mechanism is as follows: The experience replay pool stores the "state-action-reward-next state" quadruple generated by the agent when performing actions in the power distribution network environment. During training, a batch of samples is randomly drawn from the pool;

[0083] Sample storage: Each time the agent adjusts the distribution network topology, it will generate... Store in the playback pool;

[0084] Random sampling: When training the model, a batch of samples are randomly drawn from the replay pool. Random sampling can filter independent samples at different times and cut off temporal correlation.

[0085] Model update: Calculate the loss and update the model parameters using randomly sampled unbiased samples to avoid parameter oscillations caused by continuous samples and achieve stable training.

[0086] By adopting the above technical solutions, the DQN model training method is optimized. The experience replay mechanism breaks down data correlations and avoids parameter oscillations; the target network synchronizes parameters every 100 steps to stabilize the training process; and the greedy strategy balances exploration and utilization, fully searching the state space in the early stages and focusing on the optimal action in the later stages. These mechanisms collectively improve the efficiency and stability of model training, enabling the model to quickly converge to a reliable optimization strategy.

[0087] Furthermore, in step S3, when the distribution network triggers topology adjustments due to load fluctuations, excessive line losses, or periodic optimization needs, the specific steps include:

[0088] S31. First, collect the current operating data of the distribution network, including the voltage amplitude of all load nodes, the opening and closing status of tie switches, and the active power of 3-5 high load rate branches. Normalize the voltage amplitude and branch active power to the [0,1] interval, and encode the switch status in binary, where 1 represents closed and 0 represents open. Construct a state feature vector that meets the input requirements of the DQN model.

[0089] S32. Input the state feature vector into the trained DQN model. The DQN model selects the optimal switch operation action from the preset action space, including the state flipping of the handshake switch, the opening of the redundant segment switch and no operation, with a scale of the number of handshake switches + 2, through a greedy strategy.

[0090] S33. Perform compliance verification on the selected action to check for violations such as loops, islanding, or voltage overruns. If the verification passes, generate the corresponding topology adjustment command and execute the switch state adjustment. If the verification fails, the model reselects the action until the optimal compliant operation is obtained, and finally outputs the adjusted distribution network topology to achieve the goals of reducing line losses and optimizing node voltage deviation.

[0091] By adopting the above technical solution, the topology optimization steps are refined, triggering conditions are flexibly adapted to scenarios such as load fluctuations and excessive line losses, state construction conforms to the input requirements of the DQN model, and compliance verification ensures that the adjusted topology is free of loops, islands, or voltage exceedances, thus achieving dynamic and safe topology optimization. The optimal operation is selected by the trained model, reducing manual intervention and improving the efficiency and accuracy of topology adjustments.

[0092] Further, in step S32, the greedy search strategy includes,

[0093] Set the exploration probability ε, with an initial value of 0.9, which gradually decreases with training iterations. At each time an action is selected, generate a random number ξ in the interval [0,1].

[0094] When ξ < ε, the agent randomly selects a valid switch operation from the action space, including the handshake switch state flipping, the redundant segment switch being opened, or no operation;

[0095] When ξ≥ε, the agent selects the switching operation with the largest current Q value, that is, based on the trained Q value estimation network, selects the optimal action that maximizes the long-term cumulative reward.

[0096] Using the exponential decay formula ,in, =0.9 is the initial exploration probability, λ=0.995 is the decay coefficient, and t is the number of iterations. The exploration probability is gradually reduced as training progresses.

[0097] By adopting the above technical solutions, the initial high exploration rate ensures that the model fully searches the state space and avoids getting trapped in local optima; the exponential decay mechanism gradually reduces the exploration rate, focuses on the optimal action, and improves the convergence speed. This strategy of balancing exploration and utilization enables the model to learn the globally optimal strategy and quickly adapt to different operating scenarios, thus enhancing the model's generalization ability.

[0098] Furthermore, in step S4, the output of the optimized distribution network topology and operating parameters specifically includes:

[0099] S41. When the distribution network topology optimization operation is completed and the system enters a stable operating state, this output step is triggered. First, the distribution network real-time monitoring module collects basic operating data such as the adjusted tie switch and sectional switch opening and closing status, voltage amplitude of each node, and active power of the branch. At the same time, the power flow calculation unit is called to recalculate the branch reactive power, line loss value on the system bus, and node voltage phase angle and other derived parameters based on the new topology.

[0100] S42. The topology encoding unit converts the switch state data into a binary switch state matrix, where 1 represents closed and 0 represents open, and generates a list of node-branch connection relationships as a topology descriptor to clarify the physical connection form of the distribution network.

[0101] S43. The operating parameter integration unit performs inverse normalization on the collected and calculated parameters such as voltage, power, and line loss, and organizes them into a standardized set of operating parameters according to the node, branch, and system levels.

[0102] S44. Subsequently, the validity verification unit performs compliance verification on the topology descriptor and optimization target matching degree verification on the running parameter set;

[0103] S45. The formatted output unit integrates the verified topology descriptor and the set of operating parameters into a structured result. The topology is presented as a list of switch statuses and structured network topology data, and the operating parameters are presented as numerical tables and statistical reports. The output carriers include a visual scheduling interface, a printed optimization report, or an API interface called by a third-party system.

[0104] By adopting the above technical solution, the topology and operating parameters are integrated, and the topology is clearly presented through a binary switch matrix and a list of connection relationships; the inverse normalization process restores the actual physical meaning of the parameters, making it easier for operation and maintenance personnel to understand; validity verification ensures the reliability of the results; multiple output formats meet different needs such as visual scheduling and report generation, improving the practicality and operability of the results.

[0105] The beneficial effects of this invention are:

[0106] 1. This invention effectively solves the problems of strong subjectivity in traditional manual adjustment, easy getting trapped in local optima by heuristic algorithms, and high computational complexity of mathematical programming methods by integrating four core steps: distribution network modeling, DQN model construction and training, topology optimization and result output. It leverages the high-dimensional state processing capability of the DQN algorithm to achieve automation and dynamism of topology reconfiguration, taking into account multiple objectives such as reducing line losses, improving voltage quality and topology compliance, and significantly improving the economy and safety of distribution network operation.

[0107] 2. This invention provides a structured and standardized input basis for DQN model training and topology optimization through graph theory-based distribution network modeling (topology abstraction and constraint quantization) and a power flow calculation module based on forward and backward substitution, accurately outputting power grid operating state parameters and avoiding decision-making biases caused by fuzzy input data;

[0108] 3. This invention designs a DQN model that includes a state space, an action space, and a multi-objective reward function, and adopts training methods such as experience replay mechanism, target network synchronization, and greedy strategy to break data correlation, stabilize the training process, balance exploration and utilization, improve the efficiency and stability of model training, ensure accurate Q-value estimation, and provide a reliable decision basis for the optimal switching operation selection.

[0109] 4. This invention achieves dynamic and secure topology optimization through flexible topology optimization triggering conditions, compliance verification mechanisms, and structured result output (binary switch matrix, inverse normalized parameters, and multi-carrier output), reducing manual intervention, restoring the actual physical meaning of parameters, meeting the needs of different scenarios, and improving the practicality and operability of the results. Attached Figure Description

[0110] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0111] Figure 1 This is a diagram of the QN line optimization model proposed in this invention;

[0112] Figure 2 This is a schematic diagram of the model training loss in the circuit optimization results of this invention;

[0113] Figure 3 This is a schematic diagram of the optimized node voltage distribution in the line optimization results of this invention. Detailed Implementation

[0114] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0115] refer to Figure 1 , Figure 2 and Figure 3 A method for topology reconfiguration of a large-scale distribution network based on the DQN algorithm includes the following steps:

[0116] Step S1. Distribution network modeling: Construct a distribution network topology model based on graph theory, clarifying voltage constraints, branch capacity constraints, and topology radial constraints; construct a power flow calculation module, inputting switch states and node loads, and outputting power grid operating state parameters;

[0117] Step S2. DQN Model Construction and Training: Design the state space, action space, and multi-objective reward function of the DQN model; train the DQN model using a reinforcement learning training mechanism;

[0118] Step S3. Topology optimization: Use the trained DQN model to select the optimal switching operation and adjust the opening and closing state of the distribution network switches;

[0119] Step S4. Output Results: Output the optimized distribution network topology and operating parameters.

[0120] In step S1, the distribution network modeling specifically includes the following steps:

[0121] S11. Topology Abstraction: Based on graph theory, the distribution network is transformed into a node-edge undirected graph model, where the node set contains 1 power supply node and n-1 load nodes, where n is the total number of nodes in the distribution network, and the edge set corresponds to the lines and switching equipment, with sectionalizing switches being mandatory edges and tie switches being optional edges;

[0122] S12. Constraint Quantization: Establish a mapping relationship between the topology and switch states. When the switch is closed, the corresponding edge exists in the undirected graph; when the switch is open, the corresponding edge is removed from the undirected graph. Voltage constraints are... ,in, Let be the voltage amplitude at node i. The rated voltage and branch capacity constraints are: ,in, branch road Apparent power branch road Rated capacity;

[0123] S13. Power Flow Calculation Model: The topology radial constraint is achieved through graph connectivity determination, requiring that there is a unique path between any two nodes and no loops, and that all nodes are connected to power nodes and there are no islands. The forward and backward substitution method is used to construct the power flow calculation module, which takes switch status and node load data as input and outputs node voltage, branch power and bus loss.

[0124] In step S13, the power flow calculation model includes:

[0125] Initialization parameters: Set the initial value of the node voltage (the power node voltage is the rated voltage). The initial value of the load node voltage is set to The initial value of the branch power is 0, and the convergence accuracy threshold is... ;

[0126] Forward calculation: Starting from the power node, calculate the power loss and terminal node voltage of each branch sequentially according to the branch connection order of the radial network topology. For branch j, calculate the power loss and terminal node voltage of each branch according to the terminal node voltage. Branch impedance and load power Calculate the active power loss of the branch. Reactive power loss Thus, the power of the end node is obtained. , And derive the end node voltage. ;

[0127] Backward correction: Starting from the farthest load node and working backward to the power node, the voltage amplitude and phase angle of each node are corrected based on the voltage of the terminal node obtained from the forward calculation, and the branch power distribution is updated.

[0128] Convergence judgment: Compare the node voltage difference and branch power difference between two adjacent iterations. If the maximum value of all node voltage differences is less than the convergence accuracy threshold ε, and the branch power difference meets the requirements, then stop the iteration and output the final node voltage, branch power, and bus loss; otherwise, return to the previous calculation steps and repeat the iteration.

[0129] In step S2, the DQN model consists of a state-space input module, a fully connected hidden layer, an output layer (Q-value estimation), and a parameter optimization unit, wherein:

[0130] Input: Distribution network operating state vector (High-dimensional, denoted by d, consisting of load node voltage, tie switch status, and high-load branch power);

[0131] Output: Q-value vector Q( for each topology adjustment action) ;θ) (Dimension denoted as A, i.e., the size of the action space);

[0132] Model parameters: (W is the weight matrix, and b is the bias vector).

[0133] 2. Mathematical expressions for each layer of the structure

[0134] Based on the configuration of a "3-layer fully connected neural network (input layer → 64-neuron hidden layer 1 → 32-neuron hidden layer 2 → output layer)," the mathematical expression of the input-output relationship of each layer is as follows:

[0135] (1) Input layer: State vector mapping

[0136] The input layer directly receives the normalized operating state vector of the distribution network, without parameterized operations, and its mathematical expression is as follows:

[0137] ;in: :t: The state vector at time t, d = number of load nodes + number of tie switches + number of critical branches (typical value in the document: such as 86 + 12 + 5 = 103);

[0138] : Input layer output, i.e., input of the first hidden layer.

[0139] (2) First hidden layer (including ReLU activation):

[0140] The first hidden layer contains 64 neurons (denoted as L1=64), and uses the ReLU activation function: To achieve nonlinear feature extraction of the input state, mathematically expressed as:

[0141] ;

[0142] in: : Weight matrix from input layer to first hidden layer (dimension such as 103×64);

[0143] : The bias vector of the first hidden layer (dimension 64×1);

[0144] : Output of the first hidden layer (dimension 64×1);

[0145] σReLU: The ReLU activation function, used to introduce nonlinearity and prevent the model from fitting only a linear relationship.

[0146] (3) Second hidden layer (including ReLU activation)

[0147] The second hidden layer contains 32 neurons (denoted as L2=32), and also uses the ReLU activation function to further abstract and compress the features of the first hidden layer, mathematically expressed as:

[0148] ,in: : Weight matrix from the first hidden layer to the second hidden layer (dimension 64×32); : The bias vector of the second hidden layer (32×1 dimension).

[0149] (4) Output layer (linear activation)

[0150] The output layer is a linear activation layer (no activation function, or can be considered as having the activation function σlinear(x)=x), outputting the Q-value corresponding to each topology adjustment action, mathematically expressed as: ;

[0151] in: : Weight matrix from the second hidden layer to the output layer (dimension 32×A, A = number of handshake switches + 2);

[0152] : The bias vector of the output layer (dimension A×1);

[0153] State s at time t t Next, each action The Q-value vector is used for subsequent greedy strategies to select the optimal action.

[0154] 3. Mathematical basis for model parameter optimization

[0155] The core of model training is to minimize the mean squared error (MSE) loss between the "current Q value" and the "target Q value". Combining the Bellman equation in claim 5, the loss function is mathematically expressed as:

[0156] D: Experience replay pool (stores historical state-action-reward-next state samples);

[0157] Target Q value (θ′ is the target network parameter, θ is synchronized every 100 steps);

[0158] The immediate reward at time t (calculated by the multi-objective reward function);

[0159] Discount factor (typical value 0.9, balancing current and future rewards);

[0160] The model iteratively updates θ using gradient descent (such as the Adam optimizer) to minimize L(θ) and achieve an accurate estimate of the Q value.

[0161] For the current state, the Q-value is used to measure the benefit and determine the action. The essence of the Q-value is: the DQN model uses a neural network to analyze the current operating state of the distribution network. and topology adjustment action The core function of the "value quantification result" is to estimate the expected long-term cumulative reward of the "state-action pair," providing a decision-making basis for the greedy strategy. Simultaneously, through continuous optimization via training, it ultimately achieves multi-objective optimization (loss reduction, compliance, and voltage qualification) in distribution network topology reconfiguration. Combining the optimal Bellman equation, the iterative formula for Q-learning can be derived as follows: (1);

[0162] in, This is the state transition function. Is the action value function about Expectations at all times It is the reward fed back by the environment after the agent takes action At(at), where γ is the discount factor. For time t The estimate is given by α, where α is the learning rate.

[0163] The construction and training of the DQN model specifically includes the following steps:

[0164] S21. State-space design:

[0165] Constructing a high-dimensional state vector to comprehensively characterize the operating state of the distribution network, specifically including:

[0166] Node voltage characteristics: The voltage amplitude of all load nodes is normalized to the [0,1] interval to eliminate the dimensional differences of different physical quantities;

[0167] Normalization formula: ;

[0168] in: : The normalized voltage value of the i-th load node, with an output range of [0,1]; : The original voltage amplitude of the i-th load node; =0.95 , =1.05 , : System rated voltage.

[0169] Switch status characteristics: The open / closed status of all interconnecting switches is encoded in binary, where 1 indicates closed and 0 indicates open;

[0170] Branch load characteristics: The active power of 3-5 branches with high load rates is selected and normalized to the [0,1] interval to characterize the overload risk status of the branches. The dimensions of the state vector are the number of load nodes, the number of tie switches and the number of critical branches, of which the number of critical branches is 3-5.

[0171] S22. Operational Space Design:

[0172] Filter the set of valid switch operations and exclude invalid actions that could lead to islanding. Specific operation types include:

[0173] Handling switch state toggling: Switches the state of the handling switch from closed to open, or from open to closed;

[0174] Redundant sectionalizing switch disconnection: Only for branches with large power current in the loop, the redundant sectionalizing switch is disconnected;

[0175] No operation: No switching action is performed to avoid equipment damage caused by frequent switching actions;

[0176] The size of the action space is controlled by "number of communication switches + 2" (where "2" corresponds to the redundant segmented switch disconnection operation and no operation) in order to balance exploration efficiency and computational complexity.

[0177] S23. Reward Function Design:

[0178] The mathematical expression for constructing a multi-objective quantitative reward mechanism is as follows:

[0179] ;

[0180] in, The weighting coefficients are determined through engineering examples and prioritize the loss reduction target.

[0181] The average deviation of the node voltage;

[0182] To constrain the penalty for violations: the value is 1 when there are loops, islanding, or voltage exceeding limits in the distribution network; otherwise, the value is 0.

[0183] The average deviation of node voltage is the average of the absolute values ​​of the deviations of all load node voltages from the rated voltage. The smaller the deviation, the smaller the penalty.

[0184] The training method for the DQN model is as follows: The DQN model adopts a 3-layer fully connected neural network structure, consisting of an input layer, a first hidden layer with 64 neurons, a second hidden layer with 32 neurons, and an output layer; the first and second hidden layers use the ReLU activation function, and the output layer uses a linear activation function to output the expected long-term cumulative reward Q value of each switching action; an experience replay mechanism is used to break data correlation; a target Q network is introduced, and the current network parameters are synchronized every 100 steps to stabilize training; a greedy strategy is used to balance exploration and utilization.

[0185] The core function of activation functions is to introduce non-linear feature extraction capabilities into hidden layers or to preserve the value quantization range of output layers, as detailed below:

[0186] Two hidden layers: ReLU activation function (introducing nonlinearity to extract complex state features). The operating states of a distribution network—node voltages, switch states, and branch power—are highly dimensional and nonlinearly related. For example, a change in the power of a certain branch can simultaneously affect the voltages of multiple nodes. The ReLU activation function is needed to overcome the limitation that "linear models cannot fit complex relationships" while avoiding the gradient vanishing problem. The mathematical definition of ReLU is: Where: x: input to the activation function; σReLU(x): output of the activation function.

[0187] Output Layer: The core objective of the linear activation function output layer is to quantify the long-term cumulative reward expectation (Q-value) of the "state-action pair". The Q-value needs to reflect the true numerical range of "immediate reward + future discounted reward", which may be positive or negative. For example, the Q-value of a violation action is negative, while the Q-value of the optimal action is positive. Therefore, a linear activation function is required to avoid Q-value distortion caused by non-linear compression. Mathematically defined as: Where: x: weighted sum of output layer neurons; : The output of the activation function.

[0188] The output logic of the expected Q value of long-term cumulative reward: The complete mapping from "state features" to "value quantification" The output of the Q value is not an isolated numerical calculation, but a closed loop of "distribution network state → feature extraction → value estimation". The specific logic is as follows: Input: The high-dimensional state vector st is composed of "load node normalized voltage ([0,1]) + tie switch binary state (0 / 1) + high load branch normalized power ([0,1])", which contains all the key information of distribution network operation; Feature extraction: The role of the ReLU hidden layer The first hidden layer filters out "invalid associations" in the state, and the second hidden layer focuses on "key associations", and finally outputs the core features; Value estimation: The role of the linear output layer The output layer maps the features to the value (Q value) of A actions through the weight matrix. Linear activation ensures that the Q value can accurately reflect the true range of "long-term cumulative reward".

[0189] The specific process of the experience replay mechanism is as follows: 1. The experience replay pool (ExperienceReplayBuffer, denoted as D) is a "sample warehouse" used to store the "state-action-reward-next state" quadruple generated by the agent when performing actions in the power distribution network environment. During training, batches of samples are randomly drawn from the pool instead of using the latest continuous samples. The core logic for breaking data correlation lies in sample storage: each time the agent adjusts the distribution network topology, it generates... Store in the replay pool; Random sampling: When training the model, randomly draw a batch of samples from the replay pool. Random sampling can filter independent samples at different times and cut off temporal correlation; Model update: Use the randomly sampled unbiased samples to calculate the loss and update the model parameters to avoid parameter oscillations caused by continuous samples and achieve stable training.

[0190] In step S3, when the distribution network triggers topology adjustments due to load fluctuations, excessive line losses, or periodic optimization needs, the specific steps include:

[0191] S31. First, collect the current operating data of the distribution network, including the voltage amplitude of all load nodes, the opening and closing status of tie switches, and the active power of 3-5 high load rate branches. Normalize the voltage amplitude and branch active power to the [0,1] interval, and encode the switch status in binary, where 1 represents closed and 0 represents open, to construct a state feature vector that meets the input requirements of the DQN model; S32. Input the state feature vector into the trained DQN model. The DQN model selects the optimal switch operation action from a preset action space, including tie switch state flipping, redundant segment switch opening, and no operation, with a scale of the number of tie switches + 2, through a greedy strategy.

[0192] S33. Perform compliance verification on the selected action to check for violations such as loops, islanding, or voltage overruns. If the verification passes, generate the corresponding topology adjustment command and execute the switch state adjustment. If the verification fails, the model reselects the action until the optimal compliant operation is obtained, and finally outputs the adjusted distribution network topology to achieve the goals of reducing line losses and optimizing node voltage deviation.

[0193] In step S32, the greedy search strategy includes,

[0194] Strategy definition: Set the exploration probability ε to an initial value of 0.9, which gradually decays with training iterations. At each time an action is selected, generate a random number ξ in the interval [0,1].

[0195] Exploratory behavior: When ξ < ε, the agent randomly selects an effective switch operation from the action space, including flipping the state of the handshake switch, opening the redundant segment switch or no operation, in order to explore new topology adjustment possibilities and avoid getting trapped in local optima;

[0196] Utilizing behavior: When ξ≥ε, the agent selects the switching operation with the largest current Q value, that is, based on the trained Q value estimation network, selects the optimal action that maximizes long-term cumulative reward;

[0197] ε decay mechanism: using the exponential decay formula ,in =0.9 is the initial exploration probability, λ=0.995 is the decay coefficient, and t is the number of iterations. As training progresses, the exploration probability is gradually reduced to ensure that the state space is fully explored in the early stage and the optimal action is utilized in the later stage, thus balancing exploration efficiency and convergence stability.

[0198] In step S4, the optimized distribution network topology and operating parameters are output, specifically including:

[0199] S41. When the distribution network topology optimization operation is completed and the system enters a stable operating state, this output step is triggered. First, the distribution network real-time monitoring module collects basic operating data such as the adjusted tie switch and sectional switch opening and closing status, voltage amplitude of each node, and active power of the branch. At the same time, the power flow calculation unit is called to recalculate the branch reactive power, line loss value on the system bus, and node voltage phase angle and other derived parameters based on the new topology.

[0200] S42. The topology encoding unit converts the switch state data into a binary switch state matrix, where 1 represents closed and 0 represents open, and generates a list of node-branch connection relationships as a topology descriptor to clarify the physical connection form of the distribution network; S43. The operation parameter integration unit performs inverse normalization on the collected and calculated parameters such as voltage, power, and line loss, and organizes them into a standardized set of operation parameters according to the node, branch, and system levels;

[0201] S44. Subsequently, the validity verification unit performs compliance verification on the topology descriptor and optimization target matching degree verification on the running parameter set;

[0202] S45. The formatted output unit integrates the verified topology descriptor and the set of operating parameters into a structured result. The topology is presented as a list of switch statuses and structured network topology data, and the operating parameters are presented as numerical tables and statistical reports. The output carriers include a visual scheduling interface, a printable optimization report, or an API interface for third-party system calls.

[0203] In one embodiment, refer to Figure 2 and Figure 3As can be seen, the training method of this model is based on the DQN framework of deep Q-networks, combined with distribution network graph theory modeling and power flow calculation, to construct an efficient and stable reinforcement learning training mechanism. First, the model adopts a three-layer fully connected neural network architecture. The input layer receives the normalized distribution network state vector, including load node voltage, tie switch binary state, and power of 3-5 high-load branches. Nonlinear features are extracted through ReLU activation of hidden layers with 64 and 32 neurons. The output layer outputs the Q-values ​​of each switch operation through linear activation. Before training, core components need to be designed: a state space that comprehensively covers key power grid operating information; an action space that filters effective operations; tie switch flipping, redundant segment disconnection, and no operation to exclude invalid actions; a reward function for loss reduction, voltage deviation optimization, and violation penalties; and penalties triggered when loops, islands, or voltage exceedances exist. During training, an experience replay mechanism is used to store state-action-reward-next state samples, and random sampling breaks temporal correlations to avoid parameter oscillations. The target network's parameters are synchronized every 100 steps. The loss curve for stabilizing the training process shows that the loss is high in the early stages of training, gradually decreasing and stabilizing with iterations, verifying the effectiveness of this mechanism. A greedy strategy controls exploration and utilization: the initial exploration rate is set to 0.9, decaying exponentially with training. In the early stages, the state space is fully searched, and in the later stages, the focus is on the optimal action with the largest Q value. Furthermore, compliance checks are performed on selected actions during training to ensure no safety risks, ultimately leading the model to converge to a reliable strategy. The voltage compliance curve shows that in the output results of the trained model, the node voltage consistently remains within 0.95-1.05 times the rated voltage threshold, reflecting the role of voltage constraints in the reward function, indicating that the training method effectively balances topology optimization and operational safety. The overall training method, through structured design and dynamic adjustment, enables the model to converge quickly and has the ability to dynamically respond to scenarios such as load fluctuations and excessive line loss. The output results meet the requirements of practicality and operability.

[0204] Working principle: First, in the distribution network modeling stage, the complex distribution network is abstracted into a node-edge undirected graph based on graph theory. The node set contains 1 power node and n-1 load nodes. The edge set corresponds to the lines and switching equipment. Section switches are mandatory edges, and tie switches are optional edges. Quantification constraints include: node voltage must be in the range of 0.95-1.05 times the rated voltage; branch capacity constraints include: branch apparent power not exceeding the rated value; and radial constraints include: no loops, no islands, and all nodes connected to the power source. A power flow calculation module is constructed using the forward and backward substitution method. The power node voltage and load node voltage are initialized to the rated value, and the branch power is 0. A convergence accuracy threshold is set. The forward calculation starts from the power node and calculates the power loss of each branch and the voltage of the terminal node according to the branch connection sequence. The backward substitution correction adjusts the node voltage and phase angle from the farthest load node. The iteration continues until the difference between the node voltage and the branch power meets the convergence condition. The node voltage, branch power, and bus loss are output.

[0205] Secondly, the DQN model's training phase design features a state space that maps normalized load node voltages to a 0-1 range, binary-coded tie switch states (1 for closed, 0 for open), and normalized active power for 3-5 high-load branches. The action space filters effective operations, including tie switch state flipping, redundant segmented switch disconnection, and no operation. The scale is controlled to "number of tie switches + 2" to balance efficiency and complexity. The reward function employs a multi-objective quantization mechanism, including a loss reduction weight coefficient, a penalty for average node voltage deviation, and a penalty of 1 for violations such as loops / islands / voltage exceeding limits.

[0206] The model employs a three-layer fully connected neural network. The input layer receives the state vector, and ReLU hidden layers with 64 and 32 neurons extract nonlinear features. The linear output layer outputs the Q-value of each action, quantifying the long-term cumulative reward expectation of the state-action pair. During training, an experience replay pool is used to store "state-action-reward-next state" samples, and random sampling breaks temporal correlations to avoid parameter oscillations. The target network is introduced to synchronize the current network parameters every 100 steps to stabilize the training process.

[0207] The greedy strategy initially sets the exploration rate to 0.9, which decays exponentially with training to balance exploring new actions with utilizing the optimal action. In the third step, topology optimization, when the distribution network is adjusted due to load fluctuations, excessive line losses, or periodic optimization triggers, the current load node voltage, tie switch status, and power of 3-5 high-load branches are collected, normalized, and encoded into a state vector. This vector is then input into the trained DQN model, and the optimal action is selected using the greedy strategy. The selected action undergoes compliance verification to check for loops, islanding, and voltage exceeding limits. If successful, switch adjustments are executed; otherwise, actions are reselected until compliance is achieved. In the fourth step, results output, adjusted switch status, node voltage, and other data are collected. The power flow calculation module is called to obtain derived parameters such as branch reactive power and bus losses. The topology is encoded as a binary switch matrix and a node-branch connection list. Parameters are inversely normalized and organized hierarchically by node, branch, and system. Topology compliance and parameter optimization matching are verified. The output is formatted as a visual scheduling interface, a printable report, or an API interface for third-party system calls, meeting different operation and maintenance needs.

[0208] To address the issues of slow response speed, poor consistency of optimization results, and difficulty in coping with dynamic changes such as load fluctuations caused by manual methods that rely on subjective judgment, the DQN model is used to automate topology reconstruction without manual intervention. It can respond to triggering conditions such as load fluctuations and excessive line loss in real time, and output consistent optimal topology results, which greatly improves response speed and result stability.

[0209] To address the slow convergence speed, susceptibility to local optima, and poor adaptability to real-time constraints of heuristic methods such as genetic algorithms and particle swarm optimization, this patent employs the experience replay mechanism of the DQN algorithm to randomly sample historical samples, breaking temporal correlations and avoiding local optima. The target network synchronizes parameters every 100 steps to stabilize the training process, accelerating convergence. A greedy strategy with a high initial exploration rate ensures sufficient search of the state space, while later decay focuses on optimal actions, improving optimization efficiency. The violation penalty term in the reward function dynamically adapts to real-time constraints, solving the problem of poor constraint adaptability in heuristic algorithms.

[0210] To address the exponential increase in computational complexity caused by the large number of nodes in mathematical methods such as mixed-integer programming, which makes it difficult to meet real-time optimization requirements, the patent addresses this by controlling the state space dimension to include only load nodes, connection switches, and key branches, rather than all nodes and branches, and by compressing the action space size by adding 2 to the number of connection switches, thus significantly reducing the computational load. The forward-backward substitution method efficiently computes the power flow, meeting real-time optimization requirements. The nonlinear feature extraction capability of the DQN model eliminates the need for complex mathematical modeling, solving the complexity problem of mathematical programming.

[0211] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for topology reconfiguration of a large-scale distribution network based on the DQN algorithm, characterized in that, Includes the following steps: Step S1. Distribution network modeling: Construct a distribution network topology model based on graph theory, and clarify voltage constraints, branch capacity constraints, and topological radial constraints; Construct a power flow calculation module, input switch states and node loads, and output power grid operating status parameters; Step S2. DQN Model Construction and Training: Design the state space, action space, and multi-objective reward function of the DQN model; The DQN model is trained using a reinforcement learning training mechanism; Step S3. Topology optimization: Use the trained DQN model to select the optimal switching operation and adjust the opening and closing state of the distribution network switches; Step S4. Output Results: Output the optimized distribution network topology and operating parameters.

2. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 1, characterized in that, In step S1, the distribution network modeling specifically includes the following steps: S11. Topology Abstraction: Based on graph theory, the distribution network is transformed into a node-edge undirected graph model, where the node set contains 1 power supply node and n-1 load nodes, where n is the total number of nodes in the distribution network, and the edge set corresponds to the lines and switching equipment, with sectionalizing switches being mandatory edges and tie switches being optional edges; S12. Constraint Quantization: Establish a mapping relationship between the topology and switch states. When the switch is closed, the corresponding edge exists in the undirected graph; when the switch is open, the corresponding edge is removed from the undirected graph. Voltage constraints are... ,in, Let be the voltage magnitude at node i. The rated voltage and branch capacity constraints are: ,in, branch road Apparent power branch road Rated capacity; S13. Power Flow Calculation Model: The topology radial constraint is achieved through graph connectivity determination, requiring that there is a unique path between any two nodes and no loops, and that all nodes are connected to power nodes and there are no islands. The forward and backward substitution method is used to construct the power flow calculation module, which takes switch status and node load data as input and outputs node voltage, branch power and bus loss.

3. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 2, characterized in that, In step S13, the power flow calculation steps include: S131. Initialization parameters: Set the initial value of the node voltage; the power node voltage is the rated voltage. The initial value of the load node voltage is set to The initial value of the branch power is 0, and the convergence accuracy threshold is... ; S132. Forward Calculation: Starting from the power node, calculate the power loss and terminal node voltage of each branch sequentially according to the branch connection order of the radial network topology. For branch j, calculate the power loss and terminal node voltage of each branch according to the terminal node voltage. Branch impedance and load power Calculate the active power loss of the branch. Reactive power loss Thus, the power of the end node is obtained. , And derive the end node voltage. ; S133. Backward Substitution Correction: Starting from the farthest load node and working backward to the power supply node, based on the terminal node voltage obtained from the forward calculation, correct the voltage amplitude and phase angle of each node and update the branch power distribution; S134. Convergence Judgment: Compare the node voltage difference and branch power difference between two adjacent iterations. If the maximum value of all node voltage differences is less than the convergence accuracy threshold ε, and the branch power difference meets the requirements, then stop the iteration and output the final node voltage, branch power, and bus loss; otherwise, return to the previous calculation step and repeat the iteration.

4. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 1, characterized in that, In step S2, the DQN model includes a state-space input module, a fully connected hidden layer, an output layer, and a parameter optimization unit, wherein: Input distribution network operating state vector s t Output the Q-value vector Q for each topology adjustment action; Model parameters: , where W is the weight matrix and b is the bias vector; The fully connected hidden layer comprises three fully connected neural networks: an input layer, a 64-neuron hidden layer 1, a 32-neuron hidden layer 2, and an output layer. The input layer directly receives the normalized operating state vector of the distribution network, without parameterized operations, and its mathematical expression is: ;in: The state vector at time t, where d = number of load nodes + number of tie switches + number of critical branches; : Input layer output, i.e., input of hidden layer one; The first hidden layer consists of 64 neurons, denoted as L1=64, and uses the ReLU activation function. To achieve nonlinear feature extraction of the input state, mathematically expressed as: ; in: : The weight matrix from the input layer to the first hidden layer; : The bias vector of the first hidden layer; Output of the first hidden layer; The ReLU activation function is used to introduce non-linearity and prevent the model from fitting only a linear relationship. Hidden layer 2 consists of 32 neurons, denoted as L2=32. It also uses the ReLU activation function to further abstract and compress the features of the first hidden layer, mathematically expressed as: ,in: : Weight matrix from the first hidden layer to the second hidden layer (dimension 64×32); : The bias vector of the second hidden layer; : Output of the second hidden layer; The output layer is linearly activated, with no activation function, or can be considered as an activation function. Output the Q-value corresponding to each topology adjustment action, mathematically expressed as: ; in: : The weight matrix from the second hidden layer to the output layer; : The bias vector of the output layer; State s at time t t Next, each action The Q-value vector is used for subsequent greedy strategies to select the optimal action; The mean squared error loss between the current Q value and the target Q value, combined with the Bellman equation, is mathematically expressed as: ; Where: D: Experience replay pool, storing historical state-action-reward-next state samples; : Target Q value, θ′ is the target network parameter, θ is synchronized every 100 steps; The immediate reward at time t; Discount factor; Iterative updates are achieved using gradient descent, such as the Adam optimizer. , minimize .

5. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 4, characterized in that, For the current state, the action is determined by measuring the reward based on the Q-value. Combining this with the optimal Bellman equation, the iterative formula for Q-learning can be derived as follows: (1); in, This is the state transition function. Is the action value function about Expectations at all times It is the intelligent agent that takes action. The reward provided by the post-environment feedback, where γ is the discount factor. For time t The estimate is given by α, where α is the learning rate.

6. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 4, characterized in that, The construction and training of the DQN model specifically includes the following steps: S21. State-space design: Constructing a high-dimensional state vector to comprehensively characterize the operating state of the distribution network, specifically including: Node voltage characteristics: The voltage amplitude of all load nodes is normalized to the [0,1] interval to eliminate the dimensional differences of different physical quantities; Normalization formula: ; in: : The normalized voltage value of the i-th load node, with an output range of [0,1]; : The original voltage amplitude of the i-th load node; =0.95 , =1.05 , : System rated voltage. Switch status characteristics: The open / closed status of all interconnecting switches is encoded in binary, where 1 indicates closed and 0 indicates open; Branch load characteristics: The active power of 3-5 branches with high load rates is selected and normalized to the [0,1] interval to characterize the overload risk status of the branches. The dimensions of the state vector are the number of load nodes, the number of tie switches and the number of critical branches, of which the number of critical branches is 3-5. S22. Operational Space Design: Filter the set of valid switch operations and exclude invalid actions that could lead to islanding. Specific operation types include: Handling switch state toggling: Switches the state of the handling switch from closed to open, or from open to closed; Redundant sectionalizing switch disconnection: Only for branches with large power current in the loop, the redundant sectionalizing switch is disconnected; No operation: No switching action is performed to avoid equipment damage caused by frequent switching actions; The size of the action space is controlled as "number of communication switches + 2", where "2" corresponds to the redundant segment switch disconnection operation and no operation; S23. Reward Function Design: The mathematical expression for constructing a multi-objective quantitative reward mechanism is as follows: ; in, The weighting coefficients are determined through engineering examples and prioritize the loss reduction target. The average deviation of the node voltage; To constrain the penalty for violations: the value is 1 when there are loops, islanding, or voltage exceeding limits in the distribution network; otherwise, the value is 0. The average deviation of node voltage is the average of the absolute values ​​of the deviations of all load node voltages from the rated voltage. The smaller the deviation, the smaller the penalty.

7. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 1, characterized in that, In step S2, the training method for the DQN model is as follows: the DQN model adopts a 3-layer fully connected neural network structure, consisting of an input layer, a first hidden layer with 64 neurons, a second hidden layer with 32 neurons, and an output layer; the first and second hidden layers use the ReLU activation function, the output layer uses a linear activation function, and outputs the long-term cumulative reward expectation Q value of each switching action, and uses an experience replay mechanism to break data correlation; A target Q-network is introduced, and the current network parameters are synchronized every 100 steps to stabilize training; a greedy strategy is adopted to balance exploration and utilization. The activation function can be used to introduce non-linear feature extraction capabilities or preserve the value quantification range, as detailed below: Hidden layer: Nonlinearity is introduced through the ReLU activation function to extract complex state features. Specifically, the operating state of the distribution network—node voltage, switch state, and branch power—is a high-dimensional, nonlinearly related relationship. The mathematical definition of ReLU is: ;in: σReLU(x): The input to the activation function; σReLU(x): The output of the activation function; Output layer: Uses a linear activation function, the specific formula of which is: Where: x: weighted sum of neurons in the output layer; : The output of the activation function. Output logic for the long-term cumulative reward expected Q value: Input: High-dimensional state vector It consists of the load node normalized voltage, tie switch binary state, and high load branch normalized power, and contains all the key information of the distribution network operation; Feature extraction: The role of ReLU hidden layers: The first hidden layer filters out "invalid associations" in the state, the second hidden layer focuses on "key associations", and finally outputs the core features; Value estimation: The role of the linear output layer The output layer maps features to the Q-values ​​of A actions through a weight matrix. Linear activation ensures that the Q-values ​​accurately reflect the true range of "long-term cumulative reward". The specific process of the experience replay mechanism is as follows: The experience replay pool stores the "state-action-reward-next state" quadruple generated by the agent when performing actions in the power distribution network environment. During training, a batch of samples is randomly drawn from the pool; Sample storage: Each time the agent adjusts the distribution network topology, it will generate... Store in the playback pool; Random sampling: When training the model, a batch of samples are randomly drawn from the replay pool. Random sampling can filter independent samples at different times and cut off temporal correlation. Model update: Calculate the loss and update the model parameters using randomly sampled unbiased samples to avoid parameter oscillations caused by continuous samples and achieve stable training.

8. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 1, characterized in that, In step S3, when the distribution network triggers topology adjustments due to load fluctuations, excessive line losses, or periodic optimization needs, the specific steps include: S31. First, collect the current operating data of the distribution network, including the voltage amplitude of all load nodes, the opening and closing status of tie switches, and the active power of 3-5 high load rate branches. Normalize the voltage amplitude and branch active power to the [0,1] interval, and encode the switch status in binary, where 1 represents closed and 0 represents open. Construct a state feature vector that meets the input requirements of the DQN model. S32. Input the state feature vector into the trained DQN model. The DQN model selects the optimal switch operation action from the preset action space, including the state flipping of the handshake switch, the opening of the redundant segment switch and no operation, with a scale of the number of handshake switches + 2, through a greedy strategy. S33. Perform compliance verification on the selected action to check for violations such as loops, islanding, or voltage overruns. If the verification passes, generate the corresponding topology adjustment command and execute the switch state adjustment. If the verification fails, the model reselects the action until the optimal compliant operation is obtained, and finally outputs the adjusted distribution network topology to achieve the goals of reducing line losses and optimizing node voltage deviation.

9. The large-scale distribution network topology reconfiguration method based on the DQN algorithm according to claim 8, characterized in that, In step S32, the greedy search strategy includes, Set the exploration probability ε, with an initial value of 0.9, which gradually decreases with training iterations. At each time an action is selected, generate a random number ξ in the interval [0,1]. When ξ < ε, the agent randomly selects a valid switch operation from the action space, including the handshake switch state flipping, the redundant segment switch being opened, or no operation; When ξ≥ε, the agent selects the switching operation with the largest current Q value, that is, based on the trained Q value estimation network, selects the optimal action that maximizes the long-term cumulative reward. Using the exponential decay formula ,in, =0.9 is the initial exploration probability, λ=0.995 is the decay coefficient, and t is the number of iterations. The exploration probability is gradually reduced as training progresses.

10. The method for large-scale distribution network topology reconfiguration based on the DQN algorithm according to claim 1, characterized in that, In step S4, the optimized distribution network topology and operating parameters are output, specifically including: S41. When the distribution network topology optimization operation is completed and the system enters a stable operating state, this output step is triggered. First, the distribution network real-time monitoring module collects basic operating data such as the adjusted tie switch and sectional switch opening and closing status, voltage amplitude of each node, and active power of the branch. At the same time, the power flow calculation unit is called to recalculate the branch reactive power, line loss value on the system bus, and node voltage phase angle and other derived parameters based on the new topology. S42. The topology encoding unit converts the switch state data into a binary switch state matrix, where 1 represents closed and 0 represents open, and generates a list of node-branch connection relationships as a topology descriptor to clarify the physical connection form of the distribution network. S43. The operating parameter integration unit performs inverse normalization on the collected and calculated parameters such as voltage, power, and line loss, and organizes them into a standardized set of operating parameters according to the node, branch, and system levels. S44. Subsequently, the validity verification unit performs compliance verification on the topology descriptor and optimization target matching degree verification on the running parameter set; S45. The formatted output unit integrates the verified topology descriptor and the set of operating parameters into a structured result. The topology is presented as a list of switch statuses and structured network topology data, and the operating parameters are presented as numerical tables and statistical reports. The output carriers include a visual scheduling interface, a printed optimization report, or an API interface called by a third-party system.

Citation Information

Patent Citations

  • Partition-based mesh power transmission and distribution network optimization system and method

    CN120613716A