A method and system for coordinated optimization and control of multiple distribution areas in a distribution network, taking into account topology changes.

By using an attention encoder-based Critic network and an intelligent local moving SLM algorithm, clusters are dynamically divided and a multi-cluster collaborative optimization model is constructed. This solves the problem of scaling up and down the observation space caused by changes in the distribution network topology, and achieves stable and reliable optimized control and improved renewable energy consumption capacity.

CN122495591APending Publication Date: 2026-07-31STATE GRID JIANGSU ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD
Filing Date
2026-07-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies are ill-suited to adapting to the scaling of the observation space caused by dynamic changes in the distribution network topology. Furthermore, the efficiency of multi-agent collaborative regulation is low, local voltage fluctuations are difficult to handle efficiently, and the integration of new energy sources is challenging.

Method used

A Critic network based on an attention encoder is used to map the dynamically changing cluster observations in the observation space to a fixed-dimensional space. The clusters are dynamically divided by the intelligent local moving SLM algorithm, and a multi-cluster collaborative optimization model is constructed. The Actor-Critic framework is used to train the agent to learn the optimal control strategy to adapt to different topologies.

Benefits of technology

It achieves stable and reliable optimized control in scenarios with missing distribution network measurement data and frequent topology changes, reduces node voltage deviation and active power loss, improves the capacity for renewable energy absorption, has strong adaptability, and reduces the complexity of network-wide optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122495591A_ABST
    Figure CN122495591A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for collaborative optimization and control of multiple distribution network clusters considering topology changes, belonging to the field of smart grid operation technology. The method includes: dynamically dividing the distribution network into clusters based on topology changes caused by switching changes; minimizing node voltage deviation and active power loss as optimization objectives; and considering power balance constraints, operational safety constraints, distributed generation (DG) regulation performance constraints, and escalator (ESS) charging and discharging power constraints; establishing a multi-cluster collaborative optimization model and transforming it into a time-series decision problem; establishing a locally observable Markov decision model within the cluster; determining the agent's observation space, action space, and reward function; solving the Markov decision model using an Actor-Critic framework; and performing collaborative optimization and control of the multiple distribution network clusters based on the trained agents. This invention can reduce the overall network optimization complexity, reduce node voltage deviation, reduce active power loss, and improve the renewable energy absorption capacity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for coordinated optimization and control of multiple distribution areas in a distribution network that takes into account topology changes, and belongs to the field of smart grid operation technology. Background Technology

[0002] With the large-scale integration of distributed power sources, the distribution network has transformed from a traditional unidirectional radial network to a bidirectional active network, leading to increasingly prominent problems such as bidirectional power flow, voltage exceeding limits, increased network losses, and difficulties in integrating new energy sources. Traditional optimization and control methods are mainly model-driven, heavily reliant on complete measurement data and fixed network topology. However, medium- and low-voltage distribution networks suffer from missing measurements, frequent line switch operations, and dynamic topology changes, significantly reducing their applicability. While data-driven methods based on deep reinforcement learning can achieve rapid decision-making, existing technologies are mostly designed for fixed topologies and static cluster partitioning.

[0003] Existing technologies struggle to adapt to the scaling of observation spaces caused by dynamic changes in distribution network topology. When switching on the distribution network causes topology changes, the cluster boundary alters, and the scale of the agent's observation space dynamically scales. Current methods cannot map this variable-dimensional observation space to a fixed-dimensional space for processing. Furthermore, multi-agent collaborative control is susceptible to increased equipment numbers, leading to reduced collaborative efficiency and difficulty in efficiently handling local voltage fluctuations. Therefore, there is an urgent need for a collaborative optimization control strategy that can adapt to dynamic topology changes and support dynamic cluster partitioning of multiple distribution areas to improve the safety, economy, and renewable energy absorption capacity of the distribution network. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for collaborative optimization and control of multiple distribution network areas that takes into account topology changes. It introduces a Critic network based on attention encoders to map dynamically changing cluster observations in the observation space to a fixed-dimensional space. This solves the problem that existing technologies cannot adapt to the scaling of the observation space caused by dynamic changes in cluster topology. In the hierarchical autonomous collaborative control of multiple distribution areas, it reduces the overall network optimization complexity, reduces node voltage deviation, reduces active power losses, and improves the capacity for renewable energy absorption.

[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution:

[0006] In a first aspect, the present invention provides a method for coordinated optimization and control of multiple distribution network areas taking into account topology changes, comprising:

[0007] Each transformer substation is defined as an intelligent agent, and the distribution network is defined as a reinforcement learning environment. Based on the electrical distance modularity and power balance index, the intelligent local movement (SLM) algorithm is adopted to dynamically divide the cluster according to the topology changes caused by the switching of the distribution network.

[0008] Based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account distribution network power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints, a multi-cluster collaborative optimization model is established.

[0009] The multi-cluster collaborative optimization model is transformed into a time-series decision problem. A locally observable Markov decision model within the cluster is established. By determining the agent's observation space, action space, and reward function, the correlation between the time-series operation state of the distribution network and decision commands is characterized.

[0010] The Actor-Critic framework is used to solve the Markov decision model, which includes constructing a Critic network based on an attention encoder to map the dynamically changing cluster observations in the observation space to a fixed-dimensional space, and training the agent to learn the optimal control strategy to adapt to different topologies.

[0011] Based on the trained intelligent agent, the multi-cluster control of the power distribution network is carried out in a coordinated and optimized manner.

[0012] Furthermore, based on the electrical distance modularity and power balance indices, the Intelligent Local Mobility (SLM) algorithm is adopted to dynamically divide the clusters according to the topology changes caused by the switching of distribution networks, including:

[0013] Based on the electrical distance modularity and power balance indicators, a comprehensive cluster partitioning index is constructed through weighted fusion;

[0014] Initialize each transformer substation in the distribution network as an independent cluster;

[0015] Perform the following movement operation: move each station area to the adjacent cluster in sequence, calculate the change in the comprehensive division index of the cluster before and after the movement, and perform the movement that maximizes the change;

[0016] Repeat the movement operation until the movement of all stations can no longer increase the change in the cluster comprehensive division index, then perform the compression step;

[0017] The compression step includes compressing each current cluster into a supernode and constructing a weighted network, with the weighted network serving as the new topology.

[0018] The movement and compression steps are repeated with the new topology until the change in the cluster comprehensive partitioning index no longer increases, thus obtaining the optimal cluster partitioning.

[0019] Furthermore, the comprehensive cluster partitioning index is expressed as:

[0020] ;

[0021] In the formula, This represents the comprehensive cluster partitioning index. The value range is [0,1]. When the value is 1, it indicates the transformer area. Hetai District They were divided into a cluster. and These represent the electrical distance modularity. Weighting coefficients and power balance index The weighting coefficients, where, , This represents the sum of the edge weights of all branches in the weighted network. Indicates a branch The adjacency matrix, For all regions with Taiwan The sum of the weights of connected edges. For all regions with Taiwan The sum of the weights of connected edges. Indicates the area Cluster Indicates the area The cluster it belongs to.

[0022] Furthermore, the multi-cluster collaborative optimization model is expressed as:

[0023] ;

[0024] In the formula, This represents a multi-cluster collaborative optimization model. This represents the function that takes the minimum value. For cluster number, For all cluster sets, Represents a cluster The gathering of the platform area, Indicates the total time period of the scheduling cycle. express Timetable area voltage amplitude, Indicates the voltage reference value. Represents a cluster exist The active network loss at any given time, among which, , Represents a cluster The set of branch roads within, Indicates a branch The resistance, and They are respectively the inflow branches Active power and reactive power;

[0025] The power balance constraint of the distribution network is expressed as follows:

[0026] ;

[0027] In the formula, express Time-based intelligent agent Injecting active power into the nodes, express Time-based intelligent agent Inject reactive power into the nodes. Indicates the total number of agents. express Time-based intelligent agent voltage amplitude, and Branch roads The real and imaginary parts of admittance, For intelligent agents and intelligent agents The voltage phase angle difference between them Represents the cosine function. Represents the sine function;

[0028] The operational safety constraints are expressed as follows:

[0029] ;

[0030] In the formula, and These represent the upper and lower safe limits of the transformer area voltage, respectively. express Time Branch The current flowing through, express Time Branch The maximum current that is allowed to flow;

[0031] The DG regulation performance constraint is expressed as follows:

[0032] ;

[0033] In the formula, Indicates the maximum capacity of photovoltaic power. This indicates that photovoltaic power generation contributes energy. This indicates the reactive power output of photovoltaic systems.

[0034] The ESS charging and discharging power constraint is expressed as follows:

[0035] ;

[0036] In the formula, express Time-based intelligent agent The amount of electricity stored in the internal ESS, express Time-based intelligent agent The amount of electricity stored in the internal ESS, , This indicates the maximum storage capacity of the ESS. and These represent the ESS charging efficiency and discharging efficiency, respectively. express Time-based intelligent agent The charging power, express Time-based intelligent agent The discharge power, Indicates the time interval, where, , This indicates the ESS discharge status indicator, where, This indicates that the ESS is in a discharging state. This indicates that the ESS is in a charging state. and These represent the maximum charging power and the maximum discharging power of the ESS, respectively.

[0037] Furthermore, the observation space is defined as the power flow information of all nodes containing all transformer areas within the same cluster, expressed as:

[0038] ;

[0039] In the formula, express Time-based intelligent agent The observation space and Representing intelligent agents respectively The connected distribution network nodes are Active power and reactive power at any given time. For cluster The set of all network nodes connected to the intelligent agents in the system;

[0040] The action space is defined as the decision quantity of the agent based on the observation space, and is expressed as:

[0041] ;

[0042] In the formula, Represents the action space. and They are respectively Time-based intelligent agent The active and reactive power interacting with the feeder layer are achieved by the photovoltaic inverter and energy storage inverter on the distribution transformer side, respectively. express Time-based intelligent agent Power output express Time-based intelligent agent Absorbed active power express Time-based intelligent agent Output reactive power express Time-based intelligent agent Absorbing reactive power;

[0043] Among them, after performing an action in the action space, the connected intelligent agent is then... The power of the distribution network nodes is expressed as:

[0044] ;

[0045] In the formula, express Time-based intelligent agent The equivalent active power of the node after the action is performed. express Time-based intelligent agent The equivalent reactive power of the node after the action is performed;

[0046] The reward function measures the value of an agent's actions given an observation space and action space, and includes two parts: objective function reward and constraint reward. The observation space dimension remains constant for each agent, and the state transition process proceeds continuously over time. When the power flow information of all nodes within the same cluster is infeasible, the current scheduling cycle is terminated, and each agent is assigned a negative reward value. The reward function is expressed as follows:

[0047] ;

[0048] In the formula, Represents the reward function, Represents the objective function reward. Indicates constraint and reward. These represent the weighting coefficients between the objective function reward and the constraint reward. Indicates the number of energy storage devices. Indicates in Time-based intelligent agent Penalties for violating energy storage capacity constraints during execution, including: , and This represents the gain and bias terms for the coefficients of the ESS penalty function used for exceeding the lower limit. and These represent the gain and bias terms, respectively, for the coefficients of the ESS penalty function used to exceed the upper limit. For exponential functions that characterize the degree to which the lower limit is exceeded, An exponential function that characterizes the extent to which something exceeds its upper limit. and This indicates the upper limit and lower limit of the ESS's storage capacity.

[0049] Furthermore, if the cluster consists of only 3 agents, the observation space of the agents before the topology change and the observation space of the agents after the topology change are respectively represented as:

[0050] ;

[0051] ;

[0052] In the formula, These represent the observation vectors of the first, second, and third agents before the topology change, respectively. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the observation vectors of the first, second, and third agents after the topology has changed.

[0053] Furthermore, dynamically changing clustered observations in the observation space are mapped to a fixed-dimensional space, including:

[0054] Acquire dynamically changing cluster observations in the observation space. The cluster observations include the agent's own observation-action combinations and the observation-action combinations of other agents in the same cluster. For each agent, construct a Critic network based on an attention encoder to encode and map the dynamically changing cluster observations to a fixed-dimensional space.

[0055] The input to the attention encoder-based Critic network consists of two parts: the first part is the observation-action combination encoding of the agent itself; the second part is the attention encoding of the observation-action combination of other agents in the same cluster.

[0056] The attention encoding process calculates the correlation coefficient between the agent and other agents within the same cluster, and then normalizes it using a softmax function to obtain attention weight coefficients. These attention weight coefficients characterize the importance of the actions of other agents within the same cluster to the current agent's task. The first part and the second part are concatenated and input into a fully connected neural network to output action values. When the power distribution network topology changes, the attention encoder encodes observation-action samples under different topologies into fixed-dimensional features to keep the input dimension of the attention encoder-based Critic network fixed.

[0057] Furthermore, the observation-action combination encoding of the intelligent agent itself is represented as:

[0058] ;

[0059] In the formula, express Time-based intelligent agent Self-observation-action combination encoding, Represented as an intelligent agent The encoding function, Represented as Time-based intelligent agent The observation space express Time-based intelligent agent Action space;

[0060] The attention encoding of the observation-action combination of the other agents in the same cluster is represented as:

[0061] ;

[0062] In the formula, express Agents within the same cluster at any time Attention encoding of observation-action combinations, Indicates will The attention weight coefficients are obtained by normalization using the softmax function. , Represents an exponential function. express Time-based intelligent agent agents within the same cluster The correlation coefficient between them , Represented as an intelligent agent The encoding function, express Time-based intelligent agent The observation space express Time-based intelligent agent The space of motion Indicates transpose. and For mapping Parameters and mappings to key values The parameters to be queried;

[0063] The value of the action is expressed as:

[0064] ;

[0065] In the formula, express Time-based intelligent agent The value of the action, express The state space at any given moment, express The space of action at any moment Represents intelligent agents Fully connected neural networks, To act on The linear transformation matrix, This represents the activation function.

[0066] Furthermore, the trained agent learns optimal control strategies to adapt to different topologies, including:

[0067] Initialize the intelligent agent Actor Network The parameters and the attention encoder-based Critic network The parameters are set, and an experience replay pool is constructed. ;

[0068] In each training step, the agent outputs and executes actions based on the cluster observations in the current observation space and the current distribution network topology;

[0069] After an action is performed, the reinforcement learning environment returns a reward. as well as Moment State Forming transfer samples The transferred samples are then stored in the experience replay pool. In; wherein, the Moment State Include Information on changes in the distribution network topology at any given time, among which, express The state at any given moment, express Actions performed at all times express Rewards earned at any time express The new state that is constantly being transitioned to;

[0070] Once the number of samples in the experience replay pool reaches a preset threshold, a preset number of samples are randomly sampled from the experience replay pool, as follows:

[0071] ;

[0072] In the formula, Represents intelligent agents state, Represents intelligent agents The action performed Represents intelligent agents The rewards received Represents intelligent agents The transition to the new state;

[0073] Using a preset number of samples obtained from sampling, the parameters of the Critic network are updated by minimizing the joint regression function, and the parameters of the Actor network are updated by the policy gradient method, so that the agent can learn the optimal control strategy to adapt to different topologies.

[0074] The minimized joint regression function is expressed as:

[0075] ;

[0076] In the formula, This represents minimizing the joint regression function. express Rewards earned at any time express The state space at any given moment, Represents the mathematical expectation. This represents the mean squared error loss function. , This represents the future reward discount factor. This indicates the predictive value of the Critic network. express The state space at any given moment, express The action space at any given moment;

[0077] The updating of the parameters of the Actor network using the policy gradient method is expressed as follows:

[0078] ;

[0079] In the formula, This represents the gradient of the objective function with respect to the parameters of the Actor network. The gradient operator represents the parameters of the Actor network. Describe the objective function. This represents the gradient operator for the action. Indicates the value of predicted actions. Represents the observation space. Represents the action space. Represents a deterministic strategy. Indicates the first Local observations of individual agents;

[0080] The updated parameters of the Actor network and the updated parameters of the attention encoder-based Critic network are passed to the preset Actor network and attention encoder-based Critic network respectively through a soft update method.

[0081] The parameters of the updated Actor network and the updated Critic network based on the attention encoder are expressed as follows:

[0082] ;

[0083] ;

[0084] In the formula, Represents intelligent agents The parameters of the corresponding target Actor network, This is the soft update coefficient. Represents intelligent agents The corresponding parameters of the current Actor network, Represents intelligent agents The parameters of the corresponding target Critic network, Represents intelligent agents The corresponding parameters of the current Critic network.

[0085] In a second aspect, the present invention provides a multi-distribution area cluster collaborative optimization control system for a distribution network that takes into account topology changes, used to implement the multi-distribution area cluster collaborative optimization control method for a distribution network that takes into account topology changes described in the first aspect, comprising:

[0086] The dynamic cluster partitioning module is used to define each distribution area as an intelligent agent. Based on the electrical distance modularity and power balance index, it adopts the intelligent local movement (SLM) algorithm to dynamically partition the cluster according to the topology changes caused by the change of distribution network switches.

[0087] The multi-cluster collaborative optimization modeling module is used to establish a multi-cluster collaborative optimization model based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account the power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints of the distribution network.

[0088] The Markov decision process modeling module is used to transform the multi-cluster collaborative optimization model into a time-series decision problem, establish a locally observable Markov decision model within the cluster, and characterize the correlation between the time-series operating state of the distribution network and decision commands by determining the agent's observation space, action space and reward function.

[0089] The agent training module is used to solve Markov decision models using the Actor-Critic framework. This includes building an attention encoder-based Critic network to map dynamically changing cluster observations in the observation space to a fixed-dimensional space, and training the agent to learn the optimal control strategy to adapt to different topologies.

[0090] The collaborative optimization and control module is used to perform collaborative optimization and control of multiple distribution network clusters based on a trained intelligent agent.

[0091] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0092] 1. This invention defines each distribution area as an intelligent agent and uses the Intelligent Local Mobility Model (SLM) algorithm to dynamically divide the clusters according to topology changes. It achieves stable and reliable optimization and control in scenarios where distribution network measurement data is missing, topology changes frequently, and cluster boundaries dynamically change. Based on data-driven principles, it does not rely on precise global topology and full measurement information, making it highly adaptable and closely aligned with actual low-voltage distribution network operating conditions. This invention employs a Critic network based on an attention encoder to map dynamically changing cluster observations to a fixed-dimensional space, solving the problem of scaling the intelligent agent's observation space due to topology changes. Simultaneously, through multi-cluster collaborative optimization models and Markov decision process modeling, it achieves hierarchical autonomous collaborative control of multiple distribution areas, reducing the overall network optimization complexity and computational burden. It has outstanding effects in reducing node voltage deviation, reducing active power losses, and improving the absorption capacity of new energy sources.

[0093] 2. This invention constructs a comprehensive cluster partitioning index based on electrical distance modularity and power balance index, and uses the intelligent local moving SLM algorithm to perform iterative moving operations and compression steps until the change in the comprehensive cluster partitioning index no longer increases. This enables dynamic cluster partitioning under the condition of topology changes caused by changes in distribution network switches. It does not rely on the preset number of clusters and can automatically determine the optimal cluster structure, adapting to scenarios with frequent topology changes.

[0094] 3. This invention constructs a Critic network based on an attention encoder, which encodes the agent's own observation-action combination and the observation-action combinations of other agents in the same cluster, and then inputs them in series into a fully connected neural network to output action value. When the distribution network topology changes, the attention encoder encodes the observation-action samples under different topologies into fixed-dimensional features, so that the input dimension of the Critic network remains fixed, thus solving the problem of the scaling of the observation space caused by topology changes.

[0095] 4. This invention defines the observation space as the power flow information of all nodes in the same cluster, defines the action space as the decision quantity of the agent based on the observation space, and uses the objective function reward and constraint reward in the reward function to balance the objective and the operational constraints. In the scenario where there are only 3 agents in the cluster, the dimensions of the observation space before and after the topology change are respectively and . The attention encoder maps the variable dimension input to the fixed dimension to achieve adaptive control of different topologies. Attached Figure Description

[0096] Figure 1 This is a flowchart illustrating a method for coordinated optimization and control of multiple distribution network zones that takes into account topology changes, provided in an embodiment of the present invention.

[0097] Figure 2 This is a schematic diagram of the simulation test system structure provided in an embodiment of the present invention;

[0098] Figure 3 This is a schematic diagram of the time-series reward curve provided in an embodiment of the present invention;

[0099] Figure 4 This is a schematic diagram showing the comparison results of node voltage within one day after regulation by the multi-area cluster collaborative optimization regulation method for distribution networks that takes into account topology changes, as provided in the embodiments of the present invention.

[0100] Figure 5 This is a schematic diagram showing the comparison results of node voltages on the day before regulation by the multi-area cluster collaborative optimization and control method for distribution networks that takes into account topology changes, provided in an embodiment of the present invention. Detailed Implementation

[0101] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0102] Example 1

[0103] like Figure 1 As shown in the figure, this embodiment introduces a method for coordinated optimization and control of multiple distribution network areas that takes into account topology changes, including:

[0104] Step 1: Define each transformer substation as an intelligent agent and the distribution network as a reinforcement learning environment. Based on the electrical distance modularity and power balance index, adopt the intelligent local mobility (SLM) algorithm to dynamically divide the cluster according to the topology changes caused by the switching of the distribution network.

[0105] This embodiment defines each transformer substation as an intelligent agent and the distribution network as a reinforcement learning environment. Based on electrical distance modularity and power balance indices, it employs the Intelligent Local Mobility Learning (SLM) algorithm to dynamically divide the distribution network into clusters according to topology changes caused by switch movements. This allows the cluster division results to automatically adjust with topology changes, eliminating the need for manual pre-setting of cluster numbers and adapting to operating scenarios with frequent switch operations on distribution network lines.

[0106] Step 2: Based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account the power balance constraints, operation safety constraints, DG regulation performance constraints, and ESS charging and discharging power constraints, a multi-cluster collaborative optimization model is established.

[0107] This embodiment establishes a multi-cluster collaborative optimization model with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and takes into account distribution network power balance constraints, operation safety constraints, DG regulation performance constraints, and ESS charging and discharging power constraints. It integrates multi-objective optimization and multiple types of constraints into the model framework, providing a complete optimization boundary for subsequent time-series decision-making.

[0108] Step 3: Transform the multi-cluster collaborative optimization model into a time-series decision problem, establish a locally observable Markov decision model within the cluster, and characterize the correlation between the time-series operating state of the distribution network and decision commands by determining the agent's observation space, action space, and reward function.

[0109] This embodiment transforms the multi-cluster collaborative optimization model into a time-series decision problem, establishes a locally observable Markov decision model within the cluster, and determines the agent's observation space, action space, and reward function. This allows for a quantitative representation of the correlation between the time-series operating state of the distribution network and decision commands, laying the foundation for deep reinforcement learning solutions.

[0110] Step 4: Solve the Markov decision model using the Actor-Critic framework.

[0111] In this embodiment, a Critic network based on an attention encoder is constructed to map dynamically changing cluster observations in the observation space to a fixed-dimensional space, and the agent is trained to learn the optimal control strategy to adapt to different topologies.

[0112] This embodiment uses the Actor-Critic framework to solve the Markov decision model and constructs a Critic network based on an attention encoder to map the dynamically changing cluster observations in the observation space to a fixed-dimensional space. This allows the agent to directly learn the optimal control strategy to adapt to different topologies without redesigning the network structure when the cluster boundary changes or the scale of the observation space expands or contracts.

[0113] Step 5: Based on the trained intelligent agent, perform collaborative optimization and control of multiple distribution network clusters.

[0114] This embodiment achieves stable and reliable real-time control by using a trained intelligent agent to coordinate and optimize the control of multiple distribution network clusters, even in scenarios with missing distribution network measurement data and frequent topology changes. This reduces node voltage deviation and active power loss, and improves the capacity for renewable energy absorption.

[0115] Example 2

[0116] Based on the same inventive concept as Embodiment 1, this embodiment introduces the implementation steps of a multi-area cluster collaborative optimization control method for distribution networks that takes into account topology changes, including:

[0117] Figure 2This is a schematic diagram of the simulation test system structure provided in this embodiment of the invention, where TS1, TS2, TS3, and TS4 represent four tie switches. It demonstrates an application scenario of the multi-area cluster collaborative optimization and control method for distribution networks that takes into account topology changes, as described in this invention. This embodiment is based on a simulation verification of an improved IEEE 33-node system with multiple area connections. The distribution system in the figure connects to five droop-controlled distributed photovoltaic systems, installed at nodes 14, 15, 20, 24, and 32, respectively. The system base capacity is 10 MVA, the voltage base value is 12.66 kV, and the voltage safety upper and lower limits are set from 0.95 pu to 1.05 pu. All nodes in the system are connected to loads. When power disturbances such as load switching occur in the system, the distribution system is prone to voltage fluctuations and limit exceedances, endangering the safe and stable operation of the system. The above problems are solved by the method described in this invention.

[0118] Step 1: Define each transformer area as an intelligent agent and the distribution network as a reinforcement learning environment. Based on the electrical distance modularity and power balance index, adopt the intelligent local mobility (SLM) algorithm to dynamically divide the cluster according to the topology changes caused by the switching of the distribution network.

[0119] Step 1.1: Define each distribution area as an intelligent agent and the distribution network as a reinforcement learning environment.

[0120] Step 1.2: Based on the electrical distance modularity and power balance index, the intelligent local movement (SLM) algorithm is adopted to dynamically divide the clusters according to the topology changes caused by the change of distribution network switches.

[0121] Step 1.2.1: Based on the electrical distance modularity and power balance index, construct a comprehensive cluster partitioning index through weighted fusion.

[0122] In this embodiment, the cluster comprehensive partitioning index is expressed as:

[0123] ;

[0124] In the formula, This represents the comprehensive cluster partitioning index. The value range is [0,1]. When the value is 1, it indicates the transformer area. Hetai District They were divided into a cluster. and These represent the electrical distance modularity. Weighting coefficients and power balance index The weighting coefficients, where, , This represents the sum of the edge weights of all branches in the weighted network. Indicates a branch The adjacency matrix, For all regions with Taiwan The sum of the weights of connected edges. For all regions with Taiwan The sum of the weights of connected edges. Indicates the area Cluster Indicates the area The cluster it belongs to.

[0125] Step 1.2.2: Initialize each transformer substation in the distribution network as an independent cluster.

[0126] Step 1.2.3: Perform the following movement operation: move each area to the adjacent cluster in sequence, calculate the change in the cluster comprehensive division index before and after the movement, and perform the movement that maximizes the change.

[0127] Step 1.2.4: Repeat the moving operation until the moving of all stations can no longer increase the change in the cluster comprehensive division index, and then perform the compression step.

[0128] In this embodiment, the compression step includes compressing each current cluster into a supernode and constructing a weighted network, using the weighted network as a new topology.

[0129] Step 1.2.5: Repeat the moving and compression steps with the new topology until the change in the cluster comprehensive partition index no longer increases, thus obtaining the optimal cluster partition.

[0130] Step 2: Based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account the power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints, a multi-cluster collaborative optimization model is established.

[0131] In this embodiment, the multi-cluster collaborative optimization model is represented as:

[0132] ;

[0133] In the formula, This represents a multi-cluster collaborative optimization model. This represents the function that takes the minimum value. For cluster number, For all cluster sets, Represents a cluster The gathering of the platform area, Indicates the total time period of the scheduling cycle. express Timetable area voltage amplitude, Indicates the voltage reference value. Represents a cluster exist The active network loss at any given time, among which, , Represents a cluster The set of branch roads within, Indicates a branch The resistance, and They are respectively the inflow branches The active power and reactive power.

[0134] In this embodiment, the power balance constraint of the distribution network is expressed as:

[0135] ;

[0136] In the formula, express Time-based intelligent agent Injecting active power into the nodes, express Time-based intelligent agent Inject reactive power into the nodes. Indicates the total number of agents. express Time-based intelligent agent voltage amplitude, and Branch roads The real and imaginary parts of admittance, For intelligent agents and intelligent agents The voltage phase angle difference between them Represents the cosine function. This represents the sine function.

[0137] In this embodiment, the operational safety constraint is represented as:

[0138] ;

[0139] In the formula, and These represent the upper and lower safe limits of the transformer area voltage, respectively. express Time Branch The current flowing through, express Time Branch The maximum current that is allowed to flow.

[0140] In this embodiment, the DG regulation performance constraint is expressed as:

[0141] ;

[0142] In the formula, Indicates the maximum capacity of photovoltaic power. This indicates that photovoltaic power generation contributes energy. This indicates the reactive power output of photovoltaic systems.

[0143] In this embodiment, the ESS charging and discharging power constraint is expressed as:

[0144] ;

[0145] In the formula, express Time-based intelligent agent The amount of electricity stored in the internal ESS, express Time-based intelligent agent The amount of electricity stored in the internal ESS, , This indicates the maximum storage capacity of the ESS. and These represent the ESS charging efficiency and discharging efficiency, respectively. express Time-based intelligent agent The charging power, express Time-based intelligent agent The discharge power, Indicates the time interval, where, , This indicates the ESS discharge status indicator, where, This indicates that the ESS is in a discharging state. This indicates that the ESS is in a charging state. and These represent the maximum charging power and the maximum discharging power of the ESS, respectively.

[0146] Step 3: Transform the multi-cluster collaborative optimization model into a time-series decision problem, establish a locally observable Markov decision model within the cluster, and characterize the correlation between the time-series operating state of the distribution network and decision commands by determining the agent's observation space, action space, and reward function.

[0147] In this embodiment, the observation space is defined as the power flow information of all nodes containing all transformer areas within the same cluster, expressed as:

[0148] ;

[0149] In the formula, express Time-based intelligent agent The observation space and Representing intelligent agents respectively The connected distribution network nodes are Active power and reactive power at any given time. For cluster The set of all network nodes connected to the intelligent agents in the system.

[0150] In this embodiment, if the cluster includes only 3 agents, the observation space of the agents before the topology change and the observation space of the agents after the topology change are respectively represented as follows:

[0151] ;

[0152] ;

[0153] In the formula, These represent the observation vectors of the first, second, and third agents before the topology change, respectively. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the observation vectors of the first, second, and third agents after the topology has changed.

[0154] In this embodiment, the action space is defined as the decision quantity of the agent based on the observation space, expressed as:

[0155] ;

[0156] In the formula, Represents the action space. and They are respectively Time-based intelligent agent The active and reactive power interacting with the feeder layer are achieved by the photovoltaic inverter and energy storage inverter on the distribution transformer side, respectively. express Time-based intelligent agent Power output express Time-based intelligent agent Absorbed active power express Time-based intelligent agent Output reactive power express Time-based intelligent agent Absorb reactive power.

[0157] In this embodiment, after the action is performed in the action space, the intelligent agent is connected. The power of the distribution network nodes is expressed as:

[0158] ;

[0159] In the formula, express Time-based intelligent agent The injected active power, express Time-based intelligent agent The injected reactive power.

[0160] In this embodiment, the reward function is used to measure the value of an agent performing actions under a given observation space and action space, including two parts: objective function reward and constraint reward; the observation space dimension of each agent remains unchanged, and the state transition process proceeds continuously over time; when the power flow information of all nodes in the same cluster is infeasible, the current scheduling cycle is terminated, and each agent is assigned a negative reward value. The reward function is expressed as:

[0161] ;

[0162] In the formula, Represents the reward function, Represents the objective function reward. Indicates constraint and reward. These represent the weighting coefficients between the objective function reward and the constraint reward. Indicates the number of energy storage devices. Indicates in Time-based intelligent agent Penalties for violating energy storage capacity constraints during execution, including: , and This represents the gain and bias terms for the coefficients of the ESS penalty function used for exceeding the lower limit. and These represent the gain and bias terms, respectively, for the coefficients of the ESS penalty function used to exceed the upper limit. For exponential functions that characterize the degree to which the lower limit is exceeded, An exponential function that characterizes the extent to which something exceeds its upper limit. and This indicates the upper limit and lower limit of the ESS's storage capacity.

[0163] Step 4: Solve the Markov decision model using the Actor-Critic framework.

[0164] Step 4.1: Construct an attention encoder-based Critic network.

[0165] Step 4.2: Map the dynamically changing cluster observations in the observation space to a fixed-dimensional space.

[0166] Step 4.2.1: Obtain dynamically changing cluster observations in the observation space. The cluster observations include the observation-action combinations of the agent itself and the observation-action combinations of other agents in the same cluster.

[0167] Step 4.2.2: For each agent, construct an attention encoder-based Critic network to encode and map the dynamically changing cluster observations to a fixed-dimensional space.

[0168] In this embodiment, the input of the attention encoder-based Critic network includes two parts: the first part is the observation-action combination encoding of the agent itself; the second part is the attention encoding of the observation-action combination of other agents in the same cluster.

[0169] In this embodiment, the observation-action combination encoding of the intelligent agent itself is represented as:

[0170] ;

[0171] In the formula, express Time-based intelligent agent Self-observation-action combination encoding, Represented as an intelligent agent The encoding function, Represented as Time-based intelligent agent The observation space express Time-based intelligent agent The action space.

[0172] In this embodiment, the attention encoding of the observation-action combination of the other agents within the same cluster is represented as:

[0173] ;

[0174] In the formula, express Agents within the same cluster at any time Attention encoding of observation-action combinations, Indicates will The attention weight coefficients are obtained by normalization using the softmax function. , Represents an exponential function. express Time-based intelligent agent agents within the same cluster The correlation coefficient between them , Represented as an intelligent agent The encoding function, express Time-based intelligent agent The observation space express Time-based intelligent agent The space of motion Indicates transpose. and For mapping Parameters and mappings to key values The parameters to be queried;

[0175] The attention encoding process calculates the correlation coefficient between the agent and other agents in the same cluster, and then normalizes it using the softmax function to obtain the attention weight coefficient. The attention weight coefficient is used to characterize the importance of the actions of other agents in the same cluster to the current agent's task.

[0176] Step 4.2.3: Connect the first part and the second part in series and input them into a fully connected neural network to output the action value.

[0177] In this embodiment, the value of the action is expressed as:

[0178] ;

[0179] In the formula, express Time-based intelligent agent The value of the action, express The state space at any given moment, express The space of action at any moment Represents intelligent agents Fully connected neural networks, To act on The linear transformation matrix, This represents the activation function.

[0180] Step 4.2.4: When the distribution network topology changes, the attention encoder encodes the observation-action samples under different topologies into fixed-dimensional features in a unified manner, so that the input dimension of the attention encoder-based Critic network remains fixed.

[0181] Step 4.3: Train the agent to learn the optimal control strategy to adapt to different topologies.

[0182] Step 4.3.1: Initialize the agent Actor Network The parameters and the attention encoder-based Critic network The parameters are set, and an experience replay pool is constructed. .

[0183] Step 4.3.2: In each training step, the agent outputs and executes actions based on the cluster observations in the current observation space and the current distribution network topology;

[0184] Step 4.3.3: After the action is performed, the reinforcement learning environment returns a reward. as well as Moment State Forming transfer samples The transferred samples are then stored in the experience replay pool. In; wherein, the Moment State Include Information on changes in the distribution network topology at any given time, among which, express The state at any given moment, express Actions performed at all times express Rewards earned at any time express The moment shifts to the new state.

[0185] Step 4.3.4: When the number of samples in the experience replay pool reaches a preset threshold, a preset number of samples are randomly sampled from the experience replay pool, represented as:

[0186] ;

[0187] In the formula, Represents intelligent agents state, Represents intelligent agents The action performed Represents intelligent agents The rewards received Represents intelligent agents The transition to the new state.

[0188] Step 4.3.5: Using the preset number of samples obtained from sampling, update the parameters of the Critic network by minimizing the joint regression function, and update the parameters of the Actor network by using the policy gradient method, so that the agent learns the optimal control strategy to adapt to different topologies.

[0189] In this embodiment, the minimization of the joint regression function is expressed as:

[0190] ;

[0191] In the formula, This represents minimizing the joint regression function. express Rewards earned at any time express The state space at any given moment, Represents the mathematical expectation. This represents the mean squared error loss function. , This represents the future reward discount factor. This indicates the predictive value of the Critic network. express The state space at any given moment, express The space of action at any given moment.

[0192] In this embodiment, updating the parameters of the Actor network using the policy gradient method is expressed as follows:

[0193] ;

[0194] In the formula, This represents the gradient of the objective function with respect to the parameters of the Actor network. The gradient operator represents the parameters of the Actor network. Describe the objective function. This represents the gradient operator for the action. Indicates the value of predicted actions. Represents the observation space. Represents the action space. Represents a deterministic strategy. Indicates the first Local observation of an agent.

[0195] Step 4.3.6: Pass the updated parameters of the Actor network and the updated parameters of the attention encoder-based Critic network to the preset Actor network and attention encoder-based Critic network respectively through a soft update method.

[0196] The parameters of the updated Actor network and the updated Critic network based on the attention encoder are expressed as follows:

[0197] ;

[0198] ;

[0199] In the formula, Represents intelligent agents The parameters of the corresponding target Actor network, This is the soft update coefficient. Represents intelligent agents The corresponding parameters of the current Actor network, Represents intelligent agents The parameters of the corresponding target Critic network, Represents intelligent agents The corresponding parameters of the current Critic network.

[0200] Step 5: Based on the trained agent, perform collaborative optimization and control of multiple distribution network clusters.

[0201] To verify the effectiveness of the method proposed in this embodiment, a simulation test was conducted. Figure 3 This is a schematic diagram showing the comparison results of time-series reward curves provided in an embodiment of the present invention. Figure 4 This is a schematic diagram showing the comparison of node voltages within one day after regulation by the multi-area cluster collaborative optimization regulation method for distribution networks that takes into account topology changes, as provided in this embodiment of the invention. Figure 5 This is a schematic diagram showing the comparison results of node voltages on the day before regulation by the multi-area cluster collaborative optimization and control method for distribution networks that takes into account topology changes, provided in an embodiment of the present invention.

[0202] The simulation test scenario in this embodiment is a low-voltage distribution station area. For example... Figure 3 As shown, the temporal reward curve during the training process of the reinforcement learning algorithm is illustrated. Each training run consists of 500 rounds, with validation and average reward calculated every 10 rounds on the test set. Figure 4 , Figure 5As shown, after using the method proposed in this invention for regulation, the voltage level during the noon period is significantly reduced, the voltage level during the heavy load period is increased, and the overall voltage fluctuation is smaller.

[0203] Example 3

[0204] Based on the same inventive concept as other embodiments, this embodiment introduces a multi-area cluster collaborative optimization control system for distribution networks that takes into account topology changes, used to implement the multi-area cluster collaborative optimization control method for distribution networks that takes into account topology changes described in Embodiment 1 or 2, including:

[0205] The dynamic cluster partitioning module is used to define each distribution area as an intelligent agent. Based on the electrical distance modularity and power balance index, it adopts the intelligent local movement (SLM) algorithm to dynamically partition the cluster according to the topology changes caused by the change of distribution network switches.

[0206] The multi-cluster collaborative optimization modeling module is used to establish a multi-cluster collaborative optimization model based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account the power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints of the distribution network.

[0207] The Markov decision process modeling module is used to transform the multi-cluster collaborative optimization model into a time-series decision problem, establish a locally observable Markov decision model within the cluster, and characterize the correlation between the time-series operating state of the distribution network and decision commands by determining the agent's observation space, action space and reward function.

[0208] The agent training module is used to solve Markov decision models using the Actor-Critic framework. This includes building an attention encoder-based Critic network to map dynamically changing cluster observations in the observation space to a fixed-dimensional space, and training the agent to learn the optimal control strategy to adapt to different topologies.

[0209] The collaborative optimization and control module is used to perform collaborative optimization and control of multiple distribution network clusters based on a trained intelligent agent.

[0210] The specific functions of each module described above are explained in the relevant content of the method in Embodiment 1, and will not be repeated here.

[0211] In summary, this invention defines each distribution area as an intelligent agent and uses the Intelligent Local Mobility Model (SLM) algorithm to dynamically divide the clusters according to topology changes. This achieves stable and reliable optimization and control in scenarios with missing distribution network measurement data, frequent topology changes, and dynamic changes in cluster boundaries. Based on data-driven principles, it does not rely on precise global topology and full measurement information, exhibiting strong adaptability and closely aligning with actual low- and medium-voltage distribution network operating conditions. Furthermore, this invention employs a Critic network based on an attention encoder to map dynamically changing cluster observations to a fixed-dimensional space, solving the problem of scaling the intelligent agent's observation space due to topology changes. Simultaneously, through a multi-cluster collaborative optimization model and Markov decision process modeling, it achieves hierarchical autonomous collaborative control of multiple distribution areas, reducing the overall network optimization complexity and computational burden. It demonstrates outstanding effects in reducing node voltage deviation, minimizing active power losses, and improving the absorption capacity of new energy sources.

[0212] This invention constructs a comprehensive cluster partitioning index based on electrical distance modularity and power balance, and uses an intelligent local moving SLM algorithm to perform iterative moving operations and compression steps until the change in the comprehensive cluster partitioning index no longer increases. This enables dynamic cluster partitioning under the condition that the topology changes due to the change of distribution network switches. It does not depend on the preset number of clusters and can automatically determine the optimal cluster structure, adapting to scenarios with frequent topology changes.

[0213] This invention constructs a Critic network based on an attention encoder, which encodes the agent's own observation-action combination and the observation-action combinations of other agents in the same cluster, and then concatenates them into a fully connected neural network to output action values. When the distribution network topology changes, the attention encoder encodes the observation-action samples under different topologies into fixed-dimensional features, keeping the input dimension of the Critic network fixed, thus solving the problem of scaling of the observation space caused by topology changes.

[0214] This invention defines the observation space as the power flow information of all nodes in the same cluster, defines the action space as the decision quantity of the agent based on the observation space, and uses the objective function reward and constraint reward in the reward function to balance the objective and the operational constraints. In the scenario where there are only 3 agents in the cluster, the dimensions of the observation space before and after the topology change are respectively and . The attention encoder maps the variable dimension input to a fixed dimension to achieve adaptive control of different topologies.

[0215] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0216] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0217] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for coordinated optimization and control of multiple distribution areas in a distribution network considering topology changes, characterized in that, include: Each transformer substation is defined as an intelligent agent, and the distribution network is defined as a reinforcement learning environment. Based on the electrical distance modularity and power balance index, the intelligent local movement (SLM) algorithm is adopted to dynamically divide the cluster according to the topology changes caused by the switching of the distribution network. Based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account distribution network power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints, a multi-cluster collaborative optimization model is established. The multi-cluster collaborative optimization model is transformed into a time-series decision problem. A locally observable Markov decision model within the cluster is established. By determining the agent's observation space, action space, and reward function, the correlation between the time-series operation state of the distribution network and decision commands is characterized. The Actor-Critic framework is used to solve the Markov decision model, which includes constructing a Critic network based on an attention encoder to map the dynamically changing cluster observations in the observation space to a fixed-dimensional space, and training the agent to learn the optimal control strategy to adapt to different topologies. Based on the trained intelligent agent, the multi-cluster control of the power distribution network is carried out in a coordinated and optimized manner.

2. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 1, characterized in that, Based on electrical distance modularity and power balance indices, the Intelligent Local Mobility (SLM) algorithm is adopted to dynamically divide clusters according to topology changes caused by distribution network switch movements, including: Based on the electrical distance modularity and power balance indicators, a comprehensive cluster partitioning index is constructed through weighted fusion; Initialize each transformer substation in the distribution network as an independent cluster; Perform the following movement operation: move each station area to the adjacent cluster in sequence, calculate the change in the comprehensive division index of the cluster before and after the movement, and perform the movement that maximizes the change; Repeat the movement operation until the movement of all stations can no longer increase the change in the cluster comprehensive division index, then perform the compression step; The compression step includes compressing each current cluster into a supernode and constructing a weighted network, with the weighted network serving as the new topology. The movement and compression steps are repeated with the new topology until the change in the cluster comprehensive partitioning index no longer increases, thus obtaining the optimal cluster partitioning.

3. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 2, characterized in that, The comprehensive cluster partitioning index is expressed as follows: ; In the formula, This represents the comprehensive cluster partitioning index. The value range is [0,1]. When the value is 1, it indicates the transformer area. Hetai District They were divided into a cluster. and These represent the electrical distance modularity. Weighting coefficients and power balance index The weighting coefficients, where, , This represents the sum of the edge weights of all branches in the weighted network. Indicates a branch The adjacency matrix, For all regions with Taiwan The sum of the weights of connected edges. For all regions with Taiwan The sum of the weights of connected edges. Indicates the area Cluster Indicates the area The cluster it belongs to.

4. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 3, characterized in that, The multi-cluster collaborative optimization model is expressed as: ; In the formula, This represents a multi-cluster collaborative optimization model. This represents the function that takes the minimum value. For cluster number, For all cluster sets, Represents a cluster The gathering of the platform area, Indicates the total time period of the scheduling cycle. express Timetable area voltage amplitude, Indicates the voltage reference value. Represents a cluster exist The active network loss at any given time, among which, , Represents a cluster The set of branch roads within, Indicates a branch The resistance, and They are respectively the inflow branches Active power and reactive power; The power balance constraint of the distribution network is expressed as follows: ; In the formula, express Time-based intelligent agent Injecting active power into the nodes, express Time-based intelligent agent Inject reactive power into the nodes. Indicates the total number of agents. express Time-based intelligent agent voltage amplitude, and Branch roads The real and imaginary parts of admittance, For intelligent agents and intelligent agents The voltage phase angle difference between them Represents the cosine function. Represents the sine function; The operational safety constraints are expressed as follows: ; In the formula, and These represent the upper and lower safe limits of the transformer area voltage, respectively. express Time Branch The current flowing through, express Time Branch The maximum current that is allowed to flow; The DG regulation performance constraint is expressed as follows: ; In the formula, Indicates the maximum capacity of photovoltaic power. This indicates that photovoltaic power generation contributes energy. This indicates the reactive power output of photovoltaic systems. The ESS charging and discharging power constraint is expressed as follows: ; In the formula, express Time-based intelligent agent The amount of electricity stored in the internal ESS, express Time-based intelligent agent The amount of electricity stored in the internal ESS, , This indicates the maximum storage capacity of the ESS. and These represent the ESS charging efficiency and discharging efficiency, respectively. express Time-based intelligent agent The charging power, express Time-based intelligent agent The discharge power, Indicates the time interval, where, , This indicates the ESS discharge status indicator, where, This indicates that the ESS is in a discharging state. This indicates that the ESS is in a charging state. and These represent the maximum charging power and the maximum discharging power of the ESS, respectively.

5. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 4, characterized in that, The observation space is defined as the power flow information of all nodes containing stations within the same cluster, expressed as: ; In the formula, express Time-based intelligent agent The observation space and Representing intelligent agents respectively The connected distribution network nodes are Active power and reactive power at any given time. For cluster The set of all network nodes connected to the intelligent agents in the system; The action space is defined as the decision quantity of the agent based on the observation space, and is expressed as: ; In the formula, Represents the action space. and They are respectively Time-based intelligent agent The active and reactive power interacting with the feeder layer are achieved by the photovoltaic inverter and energy storage inverter on the distribution transformer side, respectively. express Time-based intelligent agent Power output express Time-based intelligent agent Absorbed active power express Time-based intelligent agent Output reactive power express Time-based intelligent agent Absorbing reactive power; Among them, after performing an action in the action space, the connected intelligent agent is then... The power of the distribution network nodes is expressed as: ; In the formula, express Time-based intelligent agent The equivalent active power of the node after the action is performed. express Time-based intelligent agent The equivalent reactive power of the node after the action is performed; The reward function measures the value of an agent's actions given an observation space and action space, and includes two parts: objective function reward and constraint reward. The observation space dimension remains constant for each agent, and the state transition process proceeds continuously over time. When the power flow information of all nodes within the same cluster is infeasible, the current scheduling cycle is terminated, and each agent is assigned a negative reward value. The reward function is expressed as follows: ; In the formula, Represents the reward function, Represents the objective function reward. Indicates constraint and reward. These represent the weighting coefficients between the objective function reward and the constraint reward. Indicates the number of energy storage devices. Indicates in Time-based intelligent agent Penalties for violating energy storage capacity constraints during execution, including: , and This represents the gain and bias terms for the coefficients of the ESS penalty function used for exceeding the lower limit. and These represent the gain and bias terms, respectively, for the coefficients of the ESS penalty function used to exceed the upper limit. For exponential functions that characterize the degree to which the lower limit is exceeded, An exponential function that characterizes the extent to which something exceeds its upper limit. and This indicates the upper limit and lower limit of the ESS's storage capacity.

6. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 5, characterized in that, If the cluster contains only 3 agents, then the observation space of the agents before the topology change and the observation space of the agents after the topology change are respectively represented as: ; ; In the formula, These represent the observation vectors of the first, second, and third agents before the topology change, respectively. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the active power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the reactive power measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents before the topology change, respectively. These represent the voltage measurement vectors of the nodes subordinate to the first, second, and third agents after the topology change. These represent the observation vectors of the first, second, and third agents after the topology has changed.

7. The method for coordinated optimization and control of multiple distribution network zones considering topology changes in a distribution network according to claim 6, characterized in that, Mapping dynamically changing clustered observations in the observation space to a fixed-dimensional space includes: Acquire dynamically changing cluster observations in the observation space, the cluster observations including the observation-action combination of the agent itself, and the observation-action combination of other agents in the same cluster; For each agent, an attention encoder-based Critic network is constructed to encode and map the dynamically changing cluster observations to a fixed-dimensional space. The input to the attention encoder-based Critic network consists of two parts: the first part is the observation-action combination encoding of the agent itself; the second part is the attention encoding of the observation-action combination of other agents in the same cluster. The attention encoding process calculates the correlation coefficient between the agent and other agents within the same cluster, and then normalizes it using a softmax function to obtain attention weight coefficients. These attention weight coefficients characterize the importance of the actions of other agents within the same cluster to the current agent's task. The first part and the second part are concatenated and input into a fully connected neural network to output action values. When the power distribution network topology changes, the attention encoder encodes observation-action samples under different topologies into fixed-dimensional features to keep the input dimension of the attention encoder-based Critic network fixed.

8. The method for coordinated optimization and control of multiple distribution network zones considering topology changes according to claim 7, characterized in that, The agent's own observation-action combination encoding is represented as: ; In the formula, express Time-based intelligent agent Self-observation-action combination encoding, Represented as an intelligent agent The encoding function, Represented as Time-based intelligent agent The observation space express Time-based intelligent agent Action space; The attention encoding of the observation-action combination of the other agents in the same cluster is represented as: ; In the formula, express Agents within the same cluster at any time Attention encoding of observation-action combinations, Indicates will The attention weight coefficients are obtained by normalization using the softmax function. , Represents an exponential function. express Time-based intelligent agent agents within the same cluster The correlation coefficient between them , Represented as an intelligent agent The encoding function, express Time-based intelligent agent The observation space express Time-based intelligent agent The space of motion Indicates transpose. and For mapping Parameters and mappings to key values The parameters to be queried; The value of the action is expressed as: ; In the formula, express Time-based intelligent agent The value of the action, express The state space at any given moment, express The space of action at any moment Represents intelligent agents Fully connected neural networks, To act on The linear transformation matrix, This represents the activation function.

9. The method for coordinated optimization and control of multiple distribution network zones considering topology changes in accordance with claim 8, characterized in that, The trained agent learns optimal control strategies to adapt to different topologies, including: Initialize the intelligent agent Actor Network The parameters and the attention encoder-based Critic network The parameters are set, and an experience replay pool is constructed. ; In each training step, the agent outputs and executes actions based on the cluster observations in the current observation space and the current distribution network topology; After an action is performed, the reinforcement learning environment returns a reward. as well as Moment State Forming transfer samples The transferred samples are then stored in the experience replay pool. In; wherein, the Moment State Include Information on changes in the distribution network topology at any given time, among which, express The state at any given moment, express Actions performed at all times express Rewards earned at any time express The new state that is constantly being transitioned to; Once the number of samples in the experience replay pool reaches a preset threshold, a preset number of samples are randomly sampled from the experience replay pool, as follows: ; In the formula, Represents intelligent agents state, Represents intelligent agents The action performed Represents intelligent agents The rewards received Represents intelligent agents The transition to the new state; Using a preset number of samples obtained from sampling, the parameters of the Critic network are updated by minimizing the joint regression function, and the parameters of the Actor network are updated by the policy gradient method, so that the agent can learn the optimal control strategy to adapt to different topologies. The minimized joint regression function is expressed as: ; In the formula, This represents minimizing the joint regression function. express Rewards earned at any time express The state space at any given moment, Represents the mathematical expectation. This represents the mean squared error loss function. , This represents the future reward discount factor. This indicates the predictive value of the Critic network. express The state space at any given moment, express The space of action at any given moment; The updating of the parameters of the Actor network using the policy gradient method is expressed as follows: ; In the formula, This represents the gradient of the objective function with respect to the parameters of the Actor network. The gradient operator represents the parameters of the Actor network. Describe the objective function. This represents the gradient operator for the action. Indicates the value of predicted actions. Represents the observation space. Represents the action space. Represents a deterministic strategy. Indicates the first Local observations of individual agents; The updated parameters of the Actor network and the updated parameters of the attention encoder-based Critic network are passed to the preset Actor network and attention encoder-based Critic network respectively through a soft update method. The parameters of the updated Actor network and the updated Critic network based on the attention encoder are expressed as follows: ; ; In the formula, Represents intelligent agents The parameters of the corresponding target Actor network, This is the soft update coefficient. Represents intelligent agents The corresponding parameters of the current Actor network, Represents intelligent agents The parameters of the corresponding target Critic network, Represents intelligent agents The corresponding parameters of the current Critic network.

10. A multi-area cluster collaborative optimization and control system for a distribution network that takes into account topology changes, characterized in that, The method for collaborative optimization and control of multiple distribution network zones taking into account topology changes, as described in any one of claims 1 to 9, includes: The dynamic cluster partitioning module is used to define each distribution area as an intelligent agent. Based on the electrical distance modularity and power balance index, it adopts the intelligent local movement (SLM) algorithm to dynamically partition the cluster according to the topology changes caused by the change of distribution network switches. The multi-cluster collaborative optimization modeling module is used to establish a multi-cluster collaborative optimization model based on the cluster, with the optimization objectives of minimizing node voltage deviation and minimizing active power loss, and taking into account distribution network power balance constraints, operation safety constraints, DG regulation performance constraints and ESS charging and discharging power constraints. The Markov decision process modeling module is used to transform the multi-cluster collaborative optimization model into a time-series decision problem, establish a locally observable Markov decision model within the cluster, and characterize the correlation between the time-series operating state of the distribution network and decision commands by determining the agent's observation space, action space and reward function. The agent training module is used to solve Markov decision models using the Actor-Critic framework. This includes building an attention encoder-based Critic network to map dynamically changing cluster observations in the observation space to a fixed-dimensional space, and training the agent to learn the optimal control strategy to adapt to different topologies. The collaborative optimization and control module is used to perform collaborative optimization and control of multiple distribution network clusters based on a trained intelligent agent.